Audio processing for temporally mismatched signals
By determining the time mismatch value in stereo coding and using downmixing algorithms and smoothing techniques, the time offset problem of microphone-received audio signals is solved, improving decoding efficiency and coding quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-03-17
- Publication Date
- 2026-03-17
AI Technical Summary
In stereo coding, the time shift caused by the time delay of audio signals received by different microphones increases the magnitude of the side channel signals, affecting decoding efficiency. Furthermore, the shift estimation changes of different frame types lead to sample duplication and artifact skipping at frame boundaries.
By determining the time mismatch value between the first audio signal and the second audio signal, a downmixing algorithm is used to generate the middle channel signal and the side channel signal. Bit allocation and encoding are performed based on the effective mismatch value. The time shift value is dynamically adjusted to compensate for changes in the sound source position, and smoothing techniques are used to reduce fluctuations in the shift estimation.
It improves decoding efficiency, reduces side channel energy, reduces sample repetition and artifact skipping at frame boundaries, and improves the coding quality of audio signals.
Smart Images

Figure CN116721667B_ABST
Abstract
Description
[0001] This application is a divisional application of the application filed on March 17, 2017, with application number 201780017113.4 and entitled "Audio Processing for Temporally Mismatched Signals".
[0002] Priority Claim
[0003] This application claims priority to the following jointly owned applications: U.S. Provisional Patent Application No. 62 / 310,611, filed March 18, 2016, entitled "Audio Processing for Temporally Offset Signals," and U.S. Non-Provisional Patent Application No. 15 / 461,356, filed March 16, 2017, entitled "Audio Processing for Temporally Mismatched Signals," the entire contents of which are expressly incorporated herein by reference. Technical Field
[0004] This invention generally relates to audio processing. Background Technology
[0005] Technological advancements have led to smaller and more powerful computing devices. For example, various portable personal computing devices exist today, including cordless phones (such as mobile phones and smartphones), tablets, and laptops. These portable personal computing devices are small, lightweight, and easily carried by the user. These devices can transmit voice and data packets via wireless networks. Furthermore, many of these devices have additional functionalities, such as digital cameras, digital camcorders, digital recorders, and audio file players. Moreover, these devices can process executable instructions, including software applications that can access the Internet, such as web browser applications. Therefore, these devices can contain significant computing power.
[0006] A computing device may include multiple microphones to receive audio signals. Generally, the sound source is closer to the first microphone than to the second microphone among multiple microphones. Therefore, the second audio signal received from the second microphone may be delayed relative to the first audio signal received from the first microphone. In stereo coding, the audio signal from the microphones may be encoded to produce a center channel signal and one or more side channel signals. The center channel signal may correspond to the sum of the first and second audio signals. The side channel signals may correspond to the difference between the first and second audio signals. Due to the delay in receiving the second audio signal relative to the first audio signal, the first audio signal may not be time-aligned with the second audio signal. This misalignment (or "timing offset") of the first audio signal relative to the second audio signal may increase the magnitude of the side channel signals. Because of the increased magnitude of the side channels, a greater number of bits may be needed to encode the side channel signals.
[0007] Furthermore, different frame types can cause the computing device to generate different temporal offsets or shift estimates. For example, the computing device may determine that the sound frames of the first audio signal are offset by a specific amount relative to the corresponding sound frames in the second audio signal. However, due to the relatively high noise level, the computing device may determine that the transition frames (or silent frames) of the first audio signal are offset by different amounts relative to the corresponding transition frames (or silent frames) of the second audio signal. Variations in shift estimates can lead to sample duplication and artifact skipping at frame boundaries. In addition, variations in shift estimates can result in higher side channel energy, which can reduce decoding efficiency. Summary of the Invention
[0008] According to one embodiment of the technology disclosed herein, a communication apparatus includes a processor and a transmitter. The processor is configured to determine a first mismatch value indicating a first amount of time mismatch between a first audio signal and a second audio signal. The first mismatch value is associated with a first frame to be encoded. The processor is also configured to determine a second mismatch value indicating a second amount of time mismatch between the first audio signal and the second audio signal. The second mismatch value is associated with a second frame to be encoded. The second frame to be encoded follows the first frame to be encoded. The processor is further configured to determine an effective mismatch value based on the first mismatch value and the second mismatch value. The second frame to be encoded includes a first sample of the first audio signal and a second sample of the second audio signal. The second sample is selected at least partially based on the effective mismatch value. The processor is also configured to generate at least one encoded signal having a bit allocation based at least partially on the second frame to be encoded. The bit allocation is at least partially based on the effective mismatch value. The transmitter is configured to transmit the at least one encoded signal to a second device.
[0009] According to another embodiment of the technology disclosed herein, a communication method includes determining, at a device, a first mismatch value indicating a first amount of time mismatch between a first audio signal and a second audio signal. The first mismatch value is associated with a first frame to be encoded. The method further includes determining, at the device, a second mismatch value. The second mismatch value indicates a second amount of time mismatch between the first audio signal and the second audio signal. The second mismatch value is associated with a second frame to be encoded. The second frame to be encoded follows the first frame to be encoded. The method further includes determining, at the device, an effective mismatch value based on the first mismatch value and the second mismatch value. The second frame to be encoded includes a first sample of the first audio signal and a second sample of the second audio signal. The second sample is selected at least partially based on the effective mismatch value. The method further includes generating at least one encoded signal having a bit allocation based at least partially on the second frame to be encoded. The bit allocation is at least partially based on the effective mismatch value. The method further includes transmitting the at least one encoded signal to a second device.
[0010] According to another embodiment of the technology disclosed herein, a computer-readable storage device stores instructions that, when executed by a processor, cause the processor to perform operations, the operations comprising: determining a first mismatch value indicating a first amount of time mismatch between a first audio signal and a second audio signal. The first mismatch value is associated with a first frame to be encoded. The operations further comprise determining a second mismatch value indicating a second amount of time mismatch between the first audio signal and the second audio signal. The second mismatch value is associated with a second frame to be encoded. The second frame to be encoded follows the first frame to be encoded. The operations further comprise determining an effective mismatch value based on the first mismatch value and the second mismatch value. The second frame to be encoded includes a first sample of the first audio signal and a second sample of the second audio signal. The second sample is selected at least partially based on the effective mismatch value. The operations further comprise generating at least one encoded signal having a bit allocation based at least partially on the second frame to be encoded. The bit allocation is at least partially based on the effective mismatch value.
[0011] According to another embodiment of the technology disclosed herein, a communication apparatus includes a processor configured to determine a shift value and a second shift value. The shift value indicates a shift of a first audio signal relative to a second audio signal. The second shift value is based on the shift value. The processor is also configured to determine a bit allocation based on the second shift value and the shift value. The processor is further configured to generate at least one encoded signal based on the bit allocation. The at least one encoded signal is based on a first sample of the first audio signal and a second sample of the second audio signal. The second sample is time-shifted relative to the first sample by an amount based on the second shift value. The apparatus also includes a transmitter configured to transmit the at least one encoded signal to a second apparatus.
[0012] According to another embodiment of the technology disclosed herein, a communication method includes determining a shift value and a second shift value at a device. The shift value indicates a shift of a first audio signal relative to a second audio signal. The second shift value is based on the shift value. The method further includes determining a decoding mode at the device based on the second shift value and the shift value. The method further includes generating at the device at the device at least one encoded signal based on the decoding mode. The at least one encoded signal is based on a first sample of the first audio signal and a second sample of the second audio signal. The second sample is time-shifted relative to the first sample by an amount based on the second shift value. The method further includes transmitting the at least one encoded signal to a second device.
[0013] According to another embodiment of the technology described herein, a computer-readable storage device stores instructions that, when executed by a processor, cause the processor to perform an operation, the operation including determining a shift value and a second shift value. The shift value indicates a shift of a first audio signal relative to a second audio signal. The second shift value is based on the shift value. The operation further includes determining a bit allocation based on the second shift value and the shift value. The operation further includes generating at least one encoded signal based on the bit allocation. The at least one encoded signal is based on a first sample of the first audio signal and a second sample of the second audio signal. The second sample is time-shifted relative to the first sample by an amount based on the second shift value.
[0014] According to another embodiment of the technology described herein, an apparatus includes means for determining a bit allocation based on a shift value and a second shift value. The shift value indicates a shift of a first audio signal relative to a second audio signal. The second shift value is based on the shift value. The apparatus further includes means for transmitting at least one encoded signal generated based on the bit allocation. The at least one encoded signal is based on a first sample of the first audio signal and a second sample of the second audio signal. The second sample is time-shifted relative to the first sample by an amount based on the second shift value. Attached Figure Description
[0015] Figure 1 It is a block diagram of a specific illustrative example of a system containing means operable to encode multiple audio signals;
[0016] Figure 2 This indicates that it includes Figure 1 A diagram of another example of a system of devices;
[0017] Figure 3 This is an explanation that can be made by Figure 1 A diagram of a specific instance of a sample encoded by a device;
[0018] Figure 4 This is an explanation that can be made by Figure 1 A diagram of a specific instance of a sample encoded by a device;
[0019] Figure 5 This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0020] Figure 6 This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0021] Figure 7 This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0022] Figure 8 This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0023] Figure 9A This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0024] Figure 9B This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0025] Figure 9C This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0026] Figure 10A This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0027] Figure 10B This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0028] Figure 11 This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0029] Figure 12 This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0030] Figure 13 This is a flowchart illustrating a specific method for encoding multiple audio signals;
[0031] Figure 14 This is a diagram illustrating another example of a system operable to encode multiple audio signals;
[0032] Figure 15 Create a chart illustrating the comparison values of frames with sound, transition frames, and silent frames;
[0033] Figure 16 This is a flowchart illustrating a method for estimating the temporal offset between audio captured at multiple microphones;
[0034] Figure 17 It is a diagram used to selectively expand the search range of comparison values for shift estimation;
[0035] Figure 18 It is a chart depicting the selective expansion of the search range for comparison values used in shift estimation;
[0036] Figure 19 It is a block diagram of a specific illustrative example of a system containing means operable to encode multiple audio signals;
[0037] Figure 20 This is a flowchart of a method for allocating bits between intermediate signals and side signals;
[0038] Figure 21 This is a flowchart of a method for selecting different decoding modes based on the final shift value and the correction shift value;
[0039] Figure 22 Explain the different decoding modes based on the techniques described in this article;
[0040] Figure 23 Explain the encoder;
[0041] Figure 24This describes the different encoded signals according to the techniques described in this article;
[0042] Figure 25 It is a system for encoding signals according to the techniques described herein;
[0043] Figure 26 This is a flowchart of a method used for communication;
[0044] Figure 27 This is a flowchart of a method used for communication;
[0045] Figure 28 It is a flowchart of a method for communication; and
[0046] Figure 29 It is a block diagram of a specific illustrative example of a device operable to encode multiple audio signals. Detailed Implementation
[0047] Systems and apparatuses operable to encode multiple audio signals are disclosed. The apparatus may include an encoder configured to encode multiple audio signals. The multiple audio signals can be captured simultaneously in time using multiple recording devices (e.g., multiple microphones). In some instances, multiple audio signals (or multichannel audio) can be synthesized (e.g., artificially) by multiplexing several audio channels recorded simultaneously or not simultaneously. As illustrative examples, simultaneous recording or multiplexing of audio channels can result in 2-channel configurations (i.e., stereo: left and right), 5.1-channel configurations (left, right, center, left surround, right surround, and low-frequency accent (LFE) channels), 7.1-channel configurations, 7.1+4-channel configurations, 22.2-channel configurations, or N-channel configurations.
[0048] Audio capture devices in a teleconference room (or telepresence room) may include multiple microphones for acquiring spatial audio. Spatial audio may include speech and encoded, transmitted background audio. Depending on the microphone arrangement and the location of the source (e.g., a speaker) relative to the microphones and the room size, speech / audio from a given source (e.g., a speaker) may arrive at the multiple microphones at different times. For example, a sound source (e.g., a speaker) may be closer to a first microphone associated with the device than a second microphone associated with the device. Therefore, sound from the sound source may arrive at the first microphone earlier than at the second microphone. The device may receive a first audio signal via the first microphone and a second audio signal via the second microphone.
[0049] Mid-Side (MS) decoding and Parametric Stereo (PS) decoding are stereo decoding techniques that offer improved efficiency compared to dual-mono decoding. In dual-mono decoding, the left (L) channel (or signal) and right (R) channel (or signal) are decoded independently without utilizing inter-channel correlation. By transforming the left and right channels into a sum channel and a difference channel (e.g., side channels) before decoding, MS decoding reduces redundancy between correlated L / R channel pairs. The sum signal and difference signal are waveforms decoded using MS decoding. The sum signal requires relatively more bits than the side signals. By transforming the L / R signals into a sum signal and a set of side parameters, PS decoding reduces redundancy in each subband. The side parameters may indicate inter-channel intensity difference (IID), inter-channel phase difference (IPD), inter-channel time difference (ITD), etc. The sum signal is the decoded waveform and is transmitted along with the side parameters. In a hybrid system, the side channels can be waveform-decoded in the lower frequency band (e.g., less than 2 kHz) and PS-decoded in the higher frequency band (e.g., greater than or equal to 2 kHz), where inter-channel phase retention is not perceptibly important.
[0050] MS decoding and PS decoding can be performed in the frequency domain or in the sub-band domain. In some instances, the left and right channels may be uncorrelated. For example, the left and right channels may contain uncorrelated synthesized signals. When the left and right channels are uncorrelated, the decoding efficiency of MS decoding, PS decoding, or both can approach that of dual-mono decoding.
[0051] Depending on the recording configuration, there may be a time shift (or time mismatch) between the left and right channels, as well as other spatial effects such as echo and room reverberation. If the time shift and phase mismatch between channels are not compensated, the sum and difference channels may contain comparable energy that reduces the decoding gain associated with MS or PS techniques. The reduction in decoding gain may be based on the amount of time (or phase) shift. The comparable energy of the sum and difference signals can limit the use of MS decoding in certain frames where the channels are time-shifted but highly correlated. In stereo decoding, the center channel (e.g., the sum channel) and side channels (e.g., the difference channel) can be generated based on the following formula:
[0052] M = (L + R) / 2, S = (LR) / 2, Formula 1
[0053] Where M corresponds to the center channel, S corresponds to the side channel, L corresponds to the left channel and R corresponds to the right channel.
[0054] In some cases, the center channel and side channels can be generated based on the following formula:
[0055] M = c(L + R), S = c(LR), Formula 2
[0056] Where c corresponds to the frequency-dependent composite value. Generating the center and side channels based on Equation 1 or Equation 2 can be called performing a "downmixing" algorithm. The process of reversing the generation of the left and right channels from the center and side channels based on Equation 1 or Equation 2 can be called performing a "upmixing" algorithm.
[0057] A specific approach for selecting between MS decoding and dual-mono decoding for a particular frame may include: generating an intermediate signal and a side signal, calculating the energy of the intermediate signal and the side signal, and determining whether to perform MS decoding based on the energy. For example, MS decoding may be performed in response to determining that the ratio of the energy of the side signal to the energy of the intermediate signal is less than a threshold. For illustration, if the right channel is shifted at least by a first time interval (e.g., approximately 0.001 seconds or 48 samples at 48 kHz), then the first energy of the intermediate signal (corresponding to the sum of the left and right signals) may be equivalent to the second energy of the side signal of the spoken speech frame (corresponding to the difference between the left and right signals). When the first energy is equivalent to the second energy, a higher number of bits can be used to encode the side channel, thereby reducing the decoding efficiency of MS decoding relative to dual-mono decoding. When the first energy is equivalent to the second energy (e.g., when the ratio of the first energy to the second energy is greater than or equal to a threshold), dual-mono decoding can therefore be used. In alternative approaches, the decision between MS decoding and dual-mono decoding for a specific frame can be made based on a threshold and a comparison of the normalized cross-correlation values of the left and right channels.
[0058] In some instances, the encoder may determine a time shift value indicating a shift of the first audio signal relative to the second audio signal. The shift value may correspond to the amount of time delay between the reception of the first audio signal at a first microphone and the reception of the second audio signal at a second microphone. Alternatively, the encoder may determine the shift value on a frame-by-frame basis (e.g., based on each 20-millisecond (ms) speech / audio frame). For example, the shift value may correspond to the amount of time delay of a second frame of the second audio signal relative to a first frame of the first audio signal. Alternatively, the shift value may correspond to the amount of time delay of a first frame of the first audio signal relative to a second frame of the second audio signal.
[0059] When the sound source is closer to the first microphone than the second microphone, the frames of the second audio signal may be delayed relative to the frames of the first audio signal. In this case, the first audio signal may be referred to as the "reference audio signal" or "reference channel," and the delayed second audio signal may be referred to as the "target audio signal" or "target channel." Alternatively, when the sound source is closer to the second microphone than the first microphone, the frames of the first audio signal may be delayed relative to the frames of the second audio signal. In this case, the second audio signal may be referred to as the reference audio signal or reference channel, and the delayed first audio signal may be referred to as the target audio signal or target channel.
[0060] Depending on the location of the sound source (e.g., a speaker) in the conference room or telepresence room and how the sound source's (e.g., the speaker's) position changes relative to the microphone, the reference and target channels can change from one frame to another; similarly, the time delay value can also change from one frame to another. However, in some implementations, the shift value can always be positive to indicate the amount of delay of the "target" channel relative to the "reference" channel. Furthermore, the shift value can correspond to a "non-causal shift" value that promptly "pulls back" the delayed target channel, thereby aligning the target channel with the "reference" channel (e.g., maximizing alignment). The downmixing algorithm used to determine the center and side channels can be performed on the reference channel and the non-causal shifted target channel.
[0061] The encoder can determine the shift value based on a reference audio channel and multiple shift values applied to the target audio channel. For example, the first frame X of the reference audio channel can be received at a first time (m1). The first specific frame Y of the target audio channel can be received at a second time (n1) corresponding to the first shift value (e.g., shift1 = n1 - m1). Furthermore, the second frame of the reference audio channel can be received at a third time (m2). The second specific frame of the target audio channel can be received at a fourth time (n2) corresponding to the second shift value (e.g., shift2 = n2 - m2).
[0062] The device can perform a framing or buffering algorithm at a first sampling rate (e.g., a 32kHz sampling rate (i.e., 640 samples per frame)) to generate frames (e.g., 20ms samples). In response to determining that the first frame of the first audio signal and the second frame of the second audio signal arrive at the device simultaneously, the encoder can estimate a shift value (e.g., shift1) equal to zero samples. The left channel (e.g., corresponding to the first audio signal) and the right channel (e.g., corresponding to the second audio signal) can be aligned in time. In some cases, even when aligned, the left and right channels may differ in energy due to various reasons (e.g., microphone calibration).
[0063] In some instances, the left and right channels may be misaligned in time due to various reasons (e.g., the sound source (e.g., a speaker) may be closer to one microphone than the other, and the two microphones may be separated by a distance greater than a threshold (e.g., 1 to 20 cm)). The position of the sound source relative to the microphones can introduce different delays in the left and right channels. Additionally, there may be gain differences, energy differences, or level differences between the left and right channels.
[0064] In some instances, when multiple speakers speak alternately (e.g., in a non-overlapping manner), the time it takes for the audio signal to reach the microphone from multiple sound sources (e.g., speakers) can vary. In this case, the encoder can dynamically adjust the time shift value based on the speaker to identify a reference channel. In other instances, multiple speakers may speak simultaneously, depending on which speaker is loudest, closest to the microphone, etc., which can lead to varying time shift values.
[0065] In some instances, the first and second audio signals may be synthesized or artificially generated when the two signals may exhibit little (e.g., no) correlation. It should be understood that the examples described herein are illustrative and may be instructive in determining the relationship between the first and second audio signals in similar or different contexts.
[0066] The encoder can generate comparison values (e.g., difference, change, or cross-correlation values) based on comparisons between a first frame of the first audio signal and multiple frames of the second audio signal. Each of the multiple frames can correspond to a specific shift value. The encoder can generate a first estimated shift value based on the comparison values. For example, the first estimated shift value can correspond to a comparison value indicating a higher temporal similarity (or lower difference) between the first frame of the first audio signal and a corresponding first frame of the second audio signal.
[0067] The encoder can determine the final shift value by optimizing a series of estimated shift values in multiple stages. For example, based on comparison values generated from stereo preprocessed and resampled versions of the first and second audio signals, the encoder can first estimate a "trial" shift value. The encoder can generate interpolated comparison values associated with shift values close to the estimated "trial" shift value. The encoder can determine a second estimated "interpolated" shift value based on the interpolated comparison values. For example, the second estimated "interpolated" shift value can correspond to a specific interpolated comparison value indicating a higher temporal similarity (or smaller difference) compared to the remaining interpolated comparison values and the first estimated "trial" shift value. If the second estimated "interpolated" shift value of the current frame (e.g., the first frame of the first audio signal) differs from the final shift value of the previous frame (e.g., a frame of the first audio signal preceding the first frame), then the "interpolated" shift value of the current frame is further "corrected" to improve the temporal similarity between the first audio signal and the shifted second audio signal. Specifically, by searching around a second estimated “interpolated” shift value for the current frame and the final estimated shift value for the previous frame, a third estimated “corrected” shift value corresponds to a more accurate measurement of temporal similarity. The third estimated “corrected” shift value is further tuned to estimate the final shift value by limiting any spurious changes in the shift values between frames, and further controlled to prevent switching from negative to positive shift values (or vice versa) in two successive (or consecutive) frames as described herein.
[0068] In some instances, the encoder may avoid switching between positive and negative shift values in consecutive or adjacent frames, or vice versa. For example, based on an estimated "interpolated" or "corrected" shift value for the first frame and a corresponding estimated "interpolated" or "corrected" or final shift value in a specific frame preceding the first frame, the encoder may set the final shift value to a specific value (e.g., 0) indicating no time shift. For illustration, in response to determining that one of the estimated "experimental" or "interpolated" or "corrected" shift values for the current frame is positive and the other of the estimated "experimental" or "interpolated" or "corrected" or "final" shift values for the previous frame (e.g., a frame preceding the first frame) is negative, the encoder may set the final shift value for the current frame (e.g., the first frame) to indicate no time shift, i.e., shift1 = 0. Alternatively, in response to determining that one of the estimated “experimental” or “interpolation” or “correction” shift values for the current frame is negative and the other of the estimated “experimental” or “interpolation” or “correction” or “final” shift values for the previous frame (e.g., the frame preceding the first frame) is positive, the encoder may also set the final shift value for the current frame (e.g., the first frame) to indicate no time shift, i.e., shift1 = 0.
[0069] The encoder can select a frame of either a first audio signal or a second audio signal as a "reference" or a "target" based on a shift value. For example, in response to determining that the final shift value is positive, the encoder can generate a reference channel or signal indicator having a first value (e.g., 0) indicating that the first audio signal is a "reference" signal and the second audio signal is a "target" signal. Alternatively, in response to determining that the final shift value is negative, the encoder can generate a reference channel or signal indicator having a second value (e.g., 1) indicating that the second audio signal is a "reference" signal and the first audio signal is a "target" signal.
[0070] The encoder can estimate a relative gain (e.g., a relative gain parameter) associated with a reference signal and a non-causally shifted target signal. For example, in response to determining that the final shift value is positive, the encoder can estimate a gain value to normalize or equalize the energy or power level of the first audio signal relative to a second audio signal offset by a non-causally shifted value (e.g., the absolute value of the final shift value). Alternatively, in response to determining that the final shift value is negative, the encoder can estimate a gain value to normalize or equalize the power level of the non-causally shifted first audio signal relative to the second audio signal. In some instances, the encoder can estimate a gain value to normalize or equalize the energy or power level of the "reference" signal relative to the non-causally shifted "target" signal. In other instances, the encoder can estimate a gain value (e.g., a relative gain value) based on a reference signal relative to the target signal (e.g., the unshifted target signal).
[0071] The encoder can generate at least one encoded signal (e.g., an intermediate signal, a side signal, or both) based on a reference signal, a target signal, a non-causal shift value, and a relative gain parameter. The side signal may correspond to the difference between a first sample of a first frame of a first audio signal and a selected sample of a selected frame of a second audio signal. The encoder can select the selected frame based on the final shift value. Compared to other samples of the second audio signal corresponding to a frame of the second audio signal (received by the device simultaneously with the first frame), fewer bits are available to encode the side channel signal because the difference between the first sample and the selected sample is reduced. The transmitter of the device can transmit at least one encoded signal, a non-causal shift value, a relative gain parameter, a reference channel or signal indicator, or a combination thereof.
[0072] Based on a reference signal, a target signal, a noncausal shift value, a relative gain parameter, low-frequency band parameters of a specific frame of the first audio signal, high-frequency band parameters of a specific frame, or a combination thereof, the encoder can generate at least one encoded signal (e.g., an intermediate signal, a side signal, or both). The specific frame may precede the first frame. Certain low-frequency band parameters, high-frequency band parameters, or combinations thereof from one or more previous frames can be used to encode the intermediate signal, side signal, or both of the first frame. Encoding the intermediate signal, side signal, or both based on low-frequency band parameters, high-frequency band parameters, or combinations thereof can improve the estimation of the noncausal shift value and the inter-channel relative gain parameter. Low-frequency band parameters, high-frequency band parameters, or combinations thereof may include spacing parameters, speech parameters, decoder type parameters, low-frequency band energy parameters, high-frequency band energy parameters, tilt parameters, spacing gain parameters, FCB gain parameters, decoding mode parameters, speech activity parameters, noise estimation parameters, signal-to-noise ratio parameters, formant parameters, speech / music decision parameters, noncausal shift, inter-channel gain parameters, or combinations thereof. The transmitter of the device can transmit at least one encoded signal, a non-causal shift value, a relative gain parameter, a reference channel (or signal) indicator, or a combination thereof.
[0073] See Figure 1 This document discloses a specific illustrative example of a system, and the system as a whole is designated as 100. System 100 includes a first device 104 communicatively coupled to a second device 106 via a network 120. The network 120 may include one or more wireless networks, one or more wired networks, or a combination thereof.
[0074] The first device 104 may include an encoder 114, a transmitter 110, one or more input interfaces 112, or a combination thereof. A first input interface of the input interface 112 may be coupled to a first microphone 146. A second input interface of the input interface 112 may be coupled to a second microphone 148. The encoder 114 may include a time equalizer 108 and may be configured to downmix and encode multiple audio signals as described herein. The first device 104 may also include a memory 153 configured to store analysis data 190. The second device 106 may include a decoder 118. The decoder 118 may include a time equalizer 124 configured to upmix and display multiple channels. The second device 106 may be coupled to a first speaker 142, a second speaker 144, or both.
[0075] During operation, the first device 104 may receive a first audio signal 130 from a first microphone 146 via a first input interface, and a second audio signal 132 from a second microphone 148 via a second input interface. The first audio signal 130 may correspond to either a right channel signal or a left channel signal. The second audio signal 132 may correspond to the other of the right channel signal or the left channel signal. A sound source 152 (e.g., a user, speaker, ambient noise, musical instrument, etc.) may be closer to the first microphone 146 than to the second microphone 148. Therefore, an audio signal from the sound source 152 may be received at the input interface 112 via the first microphone 146 at a slightly earlier time than via the second microphone 148. This inherent delay in multi-channel signal acquisition via multiple microphones can introduce a time shift between the first audio signal 130 and the second audio signal 132.
[0076] A time equalizer 108 can be configured to estimate the temporal offset between audio signals captured at microphones 146 and 148. The temporal offset can be estimated based on the delay between a first frame of the first audio signal 130 and a second frame of the second audio signal 132, wherein the second frame contains content substantially similar to the first frame. For example, the time equalizer 108 can determine the cross-correlation between the first and second frames. Cross-correlation measures the similarity of two frames based on the lag of one frame relative to the other. Based on the cross-correlation, the time equalizer 108 can determine the delay (e.g., lag) between the first and second frames. The time equalizer 108 can estimate the temporal offset between the first audio signal 130 and the second audio signal 132 based on the delay and historical delay data.
[0077] Historical data may include delays between frames retrieved from the first microphone 146 and corresponding frames retrieved from the second microphone 148. For example, the time equalizer 108 may determine cross-correlation (e.g., hysteresis) between a previous frame associated with the first audio signal 130 and a corresponding frame associated with the second audio signal 132. Each hysteresis may be represented by a “comparison value.” That is, the comparison value may indicate a time shift (k) between a frame of the first audio signal 130 and a corresponding frame of the second audio signal 132. According to one embodiment, the comparison value of the previous frame may be stored in memory 153. The smoother 192 of the time equalizer 108 may “smooth” (or average) the comparison values within a long-term frame set and use the long-term smoothed comparison values to estimate the temporal offset (e.g., “shift”) between the first audio signal 130 and the second audio signal 132.
[0078] For illustrative purposes, if CompVal N (k) represents the comparison value of frame N at shift k, so frame N can have comparison values from k = T_MIN (minimum shift) to k = T_MAX (maximum shift). Smoothing can be performed to ensure that long-term comparison values... Depend on Let f be the function f in the above equation, which can be a function of all past comparison values (or subsets) under shift (k). Long-term comparison values. The alternative representation can be: Functions f or g can be simple finite impulse response (FIR) filters or infinite impulse response (IIR) filters, respectively. For example, function g can be a single-tap IIR filter such that the long-term comparison value... Depend on Let be the value of , where α∈(0,1.0). Therefore, the long-term comparison value... CompVal can be based on the instantaneous comparison value at frame N. N (k) Long-term comparison value with one or more previous frames The weighted mixing. As the value of α increases, the amount of smoothing in the long-term comparisons increases. In a particular aspect, the function f can be an L-tap FIR filter such that the long-term comparisons... Depend on Let α1, α2, ..., αL correspond to weights. In a particular aspect, each of α1, α2, ..., αL ∈ (0, 1.0) and each specific weight of α1, α2, ..., αL may be the same as or different from another weight of α1, α2, ..., αL. Therefore, long-term comparison values... CompVal can be based on the instantaneous comparison value at frame N. N (k) is compared with the values in the previous (L-1) frames. The weighted mixture.
[0079] The smoothing techniques described above can largely normalize the shift estimates between audio frames, silent frames, and transition frames. Normalized shift estimates reduce sample duplication and artifact skipping at frame boundaries. Furthermore, normalized shift estimates result in reduced side channel energy, which improves decoding efficiency.
[0080] The time equalizer 108 can determine a final shift value 116 (e.g., a non-causal shift value) that indicates the shift (e.g., non-causal shift) of the first audio signal 130 (e.g., "target") relative to the second audio signal 132 (e.g., "reference"). The final shift value 116 may be based on an instantaneous comparison value CompVal. N (k) and long-term comparison For example, the smoothing operation described above can be performed on experimental shift values, interpolated shift values, corrected shift values, or combinations thereof, as per [the relevant context]. Figure 5 As described. The final shift value 116 can be based on experimental shift values, interpolated shift values, and corrected shift values, as per [the description of the shift value]. Figure 5 As described. A first value (e.g., a positive value) of the final shift value 116 may indicate that the second audio signal 132 is delayed relative to the first audio signal 130. A second value (e.g., a negative value) of the final shift value 116 may indicate that the first audio signal 130 is delayed relative to the second audio signal 132. A third value (e.g., 0) of the final shift value 116 may indicate that there is no delay between the first audio signal 130 and the second audio signal 132.
[0081] In some implementations, a third value (e.g., 0) of the final shift value 116 may indicate that the delay between the first audio signal 130 and the second audio signal 132 has been switched sign. For example, a first specific frame of the first audio signal 130 may precede a first frame. The first and second specific frames of the second audio signal 132 may correspond to the same sound emitted by the sound source 152. The delay between the first audio signal 130 and the second audio signal 132 may be switched from delaying the first specific frame relative to the second specific frame to delaying the second frame relative to the first frame. Alternatively, the delay between the first audio signal 130 and the second audio signal 132 may be switched from delaying the second specific frame relative to the first specific frame to delaying the first frame relative to the second specific frame. In response to determining that the delay between the first audio signal 130 and the second audio signal 132 has been switched sign, the time equalizer 108 may set the final shift value 116 to indicate a third value (e.g., 0).
[0082] The time equalizer 108 may generate a reference signal indicator 164 based on the final shift value 116. For example, in response to determining that the final shift value 116 indicates a first value (e.g., a positive value), the time equalizer 108 may generate a reference signal indicator 164 having a first value (e.g., 0) indicating that the first audio signal 130 is a "reference" signal. In response to determining that the final shift value 116 indicates a first value (e.g., a positive value), the time equalizer 108 may determine that the second audio signal 132 corresponds to a "target" signal. Alternatively, in response to determining that the final shift value 116 indicates a second value (e.g., a negative value), the time equalizer 108 may generate a reference signal indicator 164 having a second value (e.g., 1) indicating that the second audio signal 132 is a "reference" signal. In response to determining that the final shift value 116 indicates a second value (e.g., a negative value), the time equalizer 108 may determine that the first audio signal 130 corresponds to a "target" signal. In response to determining that the final shift value 116 indicates a third value (e.g., 0), the time equalizer 108 may generate a reference signal indicator 164 having a first value (e.g., 0) indicating that the first audio signal 130 is a "reference" signal. In response to determining that the final shift value 116 indicates a third value (e.g., 0), the time equalizer 108 may determine that the second audio signal 132 corresponds to a "target" signal. Alternatively, in response to determining that the final shift value 116 indicates a third value (e.g., 0), the time equalizer 108 may generate a reference signal indicator 164 having a second value (e.g., 1) indicating that the second audio signal 132 is a "reference" signal. In response to determining that the final shift value 116 indicates a third value (e.g., 0), the time equalizer 108 may determine that the first audio signal 130 corresponds to a "target" signal. In some embodiments, in response to determining that the final shift value 116 indicates a third value (e.g., 0), the time equalizer 108 may leave the reference signal indicator 164 unchanged. For example, reference signal indicator 164 may be the same as the reference signal indicator corresponding to a first specific frame of the first audio signal 130. Time equalizer 108 may generate a noncausal shift value 162 indicating the absolute value of the final shift value 116.
[0083] The time equalizer 108 can generate a gain parameter 160 (e.g., a codec gain parameter) based on samples of the "target" signal and samples of the "reference" signal. For example, the time equalizer 108 can select samples of the second audio signal 132 based on a non-causal shift value 162. Alternatively, the time equalizer 108 can select samples of the second audio signal 132 independently of the non-causal shift value 162. In response to determining that the first audio signal 130 is the reference signal, the time equalizer 108 can determine the gain parameter 160 of the selected sample based on a first sample of a first frame of the first audio signal 130. Alternatively, in response to determining that the second audio signal 132 is the reference signal, the time equalizer 108 can determine the gain parameter 160 of the first sample based on the selected sample. As an example, the gain parameter 160 can be based on one of the following equations:
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090] Where g D Corresponding to the relative gain parameter 160 used for demixing, Ref(n) corresponds to a sample of the "reference" signal, N1 corresponds to the non-causal shift value 162 of the first frame, and Targ(n+N1) corresponds to a sample of the "target" signal. Gain parameter 160(g D This can be modified, for example, based on one of equations 1a to 1f to incorporate long-term smoothing / hysteresis logic to avoid large jumps in gain between frames. When the target signal contains the first audio signal 130, the first sample may contain a sample of the target signal and the selected sample may contain a sample of the reference signal. When the target signal contains the second audio signal 132, the first sample may contain a sample of the reference signal and the selected sample may contain a sample of the target signal.
[0091] In some implementations, based on treating the first audio signal 130 as a reference signal and the second audio signal 132 as a target signal, the time equalizer 108 can generate a gain parameter 160 independent of the reference signal indicator 164. For example, the time equalizer 108 can generate the gain parameter 160 based on one of equations 1a to 1f where Ref(n) corresponds to a sample (e.g., a first sample) of the first audio signal 130 and Targ(n+N1) corresponds to a sample (e.g., a selected sample) of the second audio signal 132. In an alternative implementation, the time equalizer 108 can generate the gain parameter 160 independent of the reference signal indicator 164, based on treating the second audio signal 132 as a reference signal and the first audio signal 130 as a target signal. For example, based on one of equations 1a to 1f, where Ref(n) corresponds to a sample of the second audio signal 132 (e.g., a selected sample) and Targ(n+N1) corresponds to a sample of the first audio signal 130 (e.g., a first sample), the time equalizer 108 can generate a gain parameter 160.
[0092] Based on the first sample, the selected sample, and the relative gain parameter 160 used for downmixing, the time equalizer 108 can generate one or more encoded signals 102 (e.g., a center channel signal, a side channel signal, or both). For example, the time equalizer 108 can generate an intermediate signal based on one of the following equations:
[0093] M = Ref(n) + g D Targ(n+N1), Equation 2a
[0094] M = Ref(n) + Targ(n + N1), Equation 2b
[0095] M=DMXFAC*Ref(n)+(1-DMXFAC)*g D Targ(n+N1), Equation 2c
[0096] M = DMXFAC*Ref(n) + (1 - DMXFAC)*Targ(n + N1), Equation 2d
[0097] Where M corresponds to the middle channel signal, g D Corresponding to the relative gain parameter 160 used for demixing, Ref(n) corresponds to a sample of the "reference" signal, N1 corresponds to the non-causal shift value 162 of the first frame, and Targ(n+N1) corresponds to a sample of the "target" signal. DMXFAC can correspond to the demixing factor, as shown in [reference]. Figure 19 Further description.
[0098] For example, the time equalizer 108 can generate a side channel signal based on one of the following equations:
[0099] S = Ref(n) - g D Targ(n+N1), Equation 3a
[0100] S = g D Ref(n) - Targ(n+N1), Equation 3b
[0101] S=(1-DMXFAC)*Ref(n)-(DMXFAC)*g D Targ(n+N1), Equation 3c
[0102] S = (1 - DMXFAC) * Ref(n) - (DMXFAC) * Targ(n + N1), Equation 3d
[0103] Where S corresponds to the side channel signal, g D Corresponding to the relative gain parameter 160 used for demixing, Ref(n) corresponds to the sample of the “reference” signal, N1 corresponds to the non-causal shift value 162 of the first frame, and Targ(n+N1) corresponds to the sample of the “target” signal.
[0104] Transmitter 110 may transmit encoded signal 102 (e.g., center channel signal, side channel signal, or both), reference signal indicator 164, non-causal shift value 162, gain parameter 160, or a combination thereof, to second device 106 via network 120. In some embodiments, transmitter 110 may store encoded signal 102 (e.g., center channel signal, side channel signal, or both), reference signal indicator 164, non-causal shift value 162, gain parameter 160, or a combination thereof, at a device of network 120 or a local device for later further processing or decoding.
[0105] Decoder 118 can decode the encoded signal 102. Time equalizer 124 can perform upmixing to produce a first output signal 126 (e.g., corresponding to the first audio signal 130), a second output signal 128 (e.g., corresponding to the second audio signal 132), or both. Second device 106 can output the first output signal 126 via first speaker 142. Second device 106 can output the second output signal 128 via second speaker 144.
[0106] System 100 thus enables time equalizer 108 to encode the side channel signal using fewer bits than the intermediate signal. A first sample of the first frame of the first audio signal 130 and selected samples of the second audio signal 132 can correspond to the same sound emitted by sound source 152, and therefore, the difference between the first sample and the selected sample can be smaller than the difference between the first sample and other samples of the second audio signal 132. The side channel signal can correspond to the difference between the first sample and the selected sample.
[0107] See Figure 2 A specific illustrative example of the system is disclosed, and the system as a whole is designated as 200. System 200 includes a first device 204 coupled to a second device 106 via a network 120. The first device 204 may correspond to Figure 1 The first device 104. System 200 and Figure 1 The system 100 differs from this one because the first device 204 is coupled to more than two microphones. For example, the first device 204 may be coupled to a first microphone 146, an Nth microphone 248, and one or more additional microphones (e.g., Figure 1 The second device 106 may be coupled to the first speaker 142, the second speaker 244, one or more additional speakers (e.g., the second speaker 144), or a combination thereof. The first device 204 may include an encoder 214. The encoder 214 may correspond to... Figure 1 The encoder 114. The encoder 214 may include one or more time equalizers 208. For example, the time equalizer 208 may include Figure 1 Time equalizer 108.
[0108] During operation, the first device 204 may receive more than two audio signals. For example, the first device 204 may receive a first audio signal 130 via a first microphone 146, an Nth audio signal 232 via an Nth microphone 248, and one or more additional audio signals (e.g., the second audio signal 132) via an additional microphone (e.g., the second microphone 148).
[0109] The time equalizer 208 may generate one or more reference signal indicators 264, a final shift value 216, a non-causal shift value 262, a gain parameter 260, an encoded signal 202, or a combination thereof. For example, the time equalizer 208 may determine that a first audio signal 130 is a reference signal and each of the Nth audio signal 232 and additional audio signals is a target signal. The time equalizer 208 may generate a reference signal indicator 164, a final shift value 216, a non-causal shift value 262, a gain parameter 260, and an encoded signal 202 corresponding to each of the first audio signal 130, the Nth audio signal 232, and the additional audio signals.
[0110] Reference signal indicator 264 may include reference signal indicator 164. Final shift value 216 may include a final shift value 116 indicating the shift of the second audio signal 132 relative to the first audio signal 130, a second final shift value indicating the shift of the Nth audio signal 232 relative to the first audio signal 130, or both. Non-causal shift value 262 may include a non-causal shift value 162 corresponding to the absolute value of the final shift value 116, a second non-causal shift value corresponding to the absolute value of the second final shift value, or both. Gain parameter 260 may include a gain parameter 160 for a selected sample of the second audio signal 132, a second gain parameter for a selected sample of the Nth audio signal 232, or both. Encoded signal 202 may include at least one of the encoded signals 102. For example, encoded signal 202 may include side channel signals corresponding to a first sample of the first audio signal 130 and a selected sample of the second audio signal 132, a second side channel signal corresponding to a selected sample of the first sample and the Nth audio signal 232, or both. The encoded signal 202 may include the middle channel signal corresponding to the first sample, the selected sample of the second audio signal 132, and the selected sample of the Nth audio signal 232.
[0111] In some implementations, the time equalizer 208 can determine multiple reference signals and corresponding target signals, as shown in [reference]. Figure 15 As described. For example, reference signal indicator 264 may include a reference signal indicator corresponding to each pair of reference signals and target signals. For illustration, reference signal indicator 264 may include reference signal indicator 164 corresponding to the first audio signal 130 and the second audio signal 132. Final shift value 216 may include a final shift value corresponding to each pair of reference signals and target signals. For example, final shift value 216 may include a final shift value 116 corresponding to the first audio signal 130 and the second audio signal 132. Non-causal shift value 262 may include a non-causal shift value corresponding to each pair of reference signals and target signals. For example, non-causal shift value 262 may include a non-causal shift value 162 corresponding to the first audio signal 130 and the second audio signal 132. Gain parameter 260 may include a gain parameter corresponding to each pair of reference signals and target signals. For example, gain parameter 260 may include a gain parameter 160 corresponding to the first audio signal 130 and the second audio signal 132. The encoded signal 202 may include a center channel signal and a side channel signal corresponding to each pair of reference signals and target signals. For example, the encoded signal 202 may include an encoded signal 102 corresponding to the first audio signal 130 and the second audio signal 132.
[0112] Transmitter 110 may transmit a reference signal indicator 264, a non-causal shift value 262, a gain parameter 260, an encoded signal 202, or a combination thereof to second device 106 via network 120. Based on the reference signal indicator 264, the non-causal shift value 262, the gain parameter 260, the encoded signal 202, or a combination thereof, decoder 118 may generate one or more output signals. For example, decoder 118 may output a first output signal 226 via a first speaker 142, a Y-output signal 228 via a Y-speaker 244, and one or more additional output signals (e.g., a second output signal 128) via one or more additional speakers (e.g., the second speaker 144), or a combination thereof.
[0113] System 200 thus enables time equalizer 208 to encode more than two audio signals. For example, by generating side channel signals based on noncausal shift value 262, encoded signal 202 can contain multiple side channel signals encoded using fewer bits than the corresponding center channel.
[0114] See Figure 3 This demonstrates an illustrative example of a sample, and the sample as a whole is designated as 300. As described herein, at least a subset of sample 300 may be encoded by the first device 104.
[0115] Sample 300 may include a first sample 320 corresponding to the first audio signal 130, a second sample 350 corresponding to the second audio signal 132, or both. First sample 320 may include samples 322, 324, 326, 328, 330, 332, 334, 336, one or more additional samples, or a combination thereof. Second sample 350 may include samples 352, 354, 356, 358, 360, 362, 364, 366, one or more additional samples, or a combination thereof.
[0116] The first audio signal 130 may correspond to multiple frames (e.g., frame 302, frame 304, frame 306, or a combination thereof). Each of the multiple frames may correspond to a subset of samples of the first sample 320 (e.g., corresponding to 20 ms, such as 640 samples at 32 kHz or 960 samples at 48 kHz). For example, frame 302 may correspond to sample 322, sample 324, one or more additional samples, or a combination thereof. Frame 304 may correspond to sample 326, sample 328, sample 330, sample 332, one or more additional samples, or a combination thereof. Frame 306 may correspond to sample 334, sample 336, one or more additional samples, or a combination thereof.
[0117] Sample 322 is available Figure 1The input interface 112 receives the sample at approximately the same time as sample 352. Sample 324 can be received at... Figure 1 The input interface 112 receives data at approximately the same time as sample 354. Sample 326 can be received at... Figure 1 The input interface 112 receives data at approximately the same time as sample 356. Sample 328 can be received at... Figure 1 The input interface 112 receives data at approximately the same time as sample 358. Sample 330 can be received at... Figure 1 The input interface 112 receives data at approximately the same time as sample 360. Sample 332 can... Figure 1 The input interface 112 receives data at approximately the same time as sample 362. Sample 334 can be received at... Figure 1 The input interface 112 receives data at approximately the same time as sample 364. Sample 336 can be received at... Figure 1 The input interface 112 and sample 366 were received at approximately the same time.
[0118] A first value (e.g., a positive value) of the final shift value 116 may indicate a delay of the second audio signal 132 relative to the first audio signal 130. For example, a first value of the final shift value 116 (e.g., +X ms or +Y samples, where X and Y contain positive real numbers) may indicate that frames 304 (e.g., samples 326 to 332) correspond to samples 358 to 364. Samples 326 to 332 and samples 358 to 364 may correspond to the same sound emitted from the sound source 152. Samples 358 to 364 may correspond to frame 344 of the second audio signal 132. Figures 1 to 15 The description of one or more samples with reticular lines indicates that the samples correspond to the same sound. For example, samples 326 to 332 and samples 358 to 364 are... Figure 3 The description is as having a mesh to indicate that samples 326 to 332 (e.g., frame 304) and samples 358 to 364 (e.g., frame 344) correspond to the same sound emitted from sound source 152.
[0119] It should be understood that, such as Figure 3As shown, the time offset of Y samples is illustrative. For example, the time offset may correspond to the number of samples Y, which is greater than or equal to 0. In the first case where the time offset Y = 0 samples, samples 326 to 332 (e.g., corresponding to frame 304) and samples 356 to 362 (e.g., corresponding to frame 344) may exhibit high similarity without any frame offset. In the second case where the time offset Y = 2 samples, frames 304 and 344 may be offset by 2 samples. In this case, the first audio signal 130 may be received at the input interface 112 before the second audio signal 132 by Y = 2 samples or X = (2 / Fs) ms, where Fs corresponds to the sampling rate in kHz. In some cases, the time offset Y may contain non-integer values, for example, Y = 1.6 samples, which corresponds to X = 0.05 ms at 32 kHz.
[0120] Figure 1 The time equalizer 108 can generate the encoded signal 102 by encoding samples 326 to 332 and samples 358 to 364, as shown in [reference]. Figure 1 As described, the time equalizer 108 determines that the first audio signal 130 corresponds to the reference signal and the second audio signal 132 corresponds to the target signal.
[0121] See Figure 4 This demonstrates an illustrative example of a sample, and the sample as a whole is designated as 400. Sample 400 differs from sample 300 in that the first audio signal 130 is delayed relative to the second audio signal 132.
[0122] A second value (e.g., a negative value) of the final shift value 116 may indicate a delay of the first audio signal 130 relative to the second audio signal 132. For example, a second value of the final shift value 116 (e.g., -X ms or -Y samples, where X and Y contain positive real numbers) may indicate that frame 304 (e.g., samples 326 to 332) corresponds to samples 354 to 360. Samples 354 to 360 may correspond to frame 344 of the second audio signal 132. Samples 354 to 360 (e.g., frame 344) and samples 326 to 332 (e.g., frame 304) may correspond to the same sound emitted by the sound source 152.
[0123] It should be understood that, such as Figure 4As shown, the time offset of -Y samples is illustrative. For example, the time offset may correspond to the number of samples -Y, which is less than or equal to 0. In the first case where the time offset Y = 0 samples, samples 326 to 332 (e.g., corresponding to frame 304) and samples 356 to 362 (e.g., corresponding to frame 344) may exhibit high similarity without any frame offset. In the second case where the time offset Y = -6 samples, frames 304 and 344 may be offset by 6 samples. In this case, the first audio signal 130 may be received at the input interface 112 with the second audio signal 132 after Y = -6 samples or X = (-6 / Fs) ms, where Fs corresponds to the sampling rate in kHz. In some cases, the time offset Y may contain non-integer values, for example, Y = -3.2 samples, which corresponds to X = -0.1 ms at 32 kHz.
[0124] Figure 1 The time equalizer 108 can generate the encoded signal 102 by encoding samples 354 to 360 and samples 326 to 332, as shown in [reference]. Figure 1 As described. The time equalizer 108 can determine that the second audio signal 132 corresponds to the reference signal and the first audio signal 130 corresponds to the target signal. Specifically, the time equalizer 108 can estimate the non-causal shift value 162 based on the final shift value 116, as shown in [reference]. Figure 5 As described. Based on the sign of the final shift value 116, the time equalizer 108 can identify (e.g., designate) one of the first audio signal 130 or the second audio signal 132 as a reference signal, and identify the other of the first audio signal 130 or the second audio signal 132 as a target signal.
[0125] See Figure 5 This illustrates an example of the system, and the entire system is designated as 500. System 500 can correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 500. Time equalizer 108 may include resampler 504, signal comparator 506, interpolator 510, shift optimizer 511, shift change analyzer 512, absolute shift generator 513, reference signal designator 508, gain parameter generator 514, signal generator 516, or combinations thereof.
[0126] During operation, resampler 504 may generate one or more resampled signals, as shown in [reference]. Figure 6As further described above. For example, by resampling the first audio signal 130 based on a resampling factor (D) (e.g., ≥1), resampler 504 can generate a first resampled signal 530. By resampling the second audio signal 132 based on a resampling factor (D), resampler 504 can generate a second resampled signal 532. Resampler 504 can provide the first resampled signal 530, the second resampled signal 532, or both to signal comparator 506.
[0127] Signal comparator 506 may generate a comparison value 534 (e.g., difference, change, similarity, coherence, or cross-correlation value), a trial shift value 536, or both, as shown in [reference]. Figure 7 As further described. For example, signal comparator 506 may generate a comparison value 534 based on a first resampled signal 530 and multiple shift values applied to a second resampled signal 532, as shown in [reference]. Figure 7 As further described. Signal comparator 506 can determine the experimental shift value 536 based on the comparison value 534, as shown in [reference]. Figure 7 As further described. According to one embodiment, signal comparator 506 can retrieve comparison values of previous frames of resampled signals 530, 532, and can use the comparison values of previous frames to modify comparison value 534 based on long-term smoothing operations. For example, comparison value 534 may include long-term comparison values of the current frame (N). And can be derived from Let be the value of , where α∈(0,1.0). Therefore, the long-term comparison value... CompVal can be based on the instantaneous comparison value at frame N. N (k) Long-term comparison value with one or more previous frames The weighted mixture. As the value of α increases, the amount of smoothing in the long-term comparisons increases.
[0128] The first resampled signal 530 may contain fewer or more samples than the first audio signal 130. The second resampled signal 532 may contain fewer or more samples than the second audio signal 132. Determining the comparison value 534 based on fewer samples of the resampled signals (e.g., the first resampled signal 530 and the second resampled signal 532) can use fewer resources (e.g., time, number of operations, or both) compared to using samples based on the original signals (e.g., the first audio signal 130 and the second audio signal 132). Determining the comparison value 534 based on more samples of the resampled signals (e.g., the first resampled signal 530 and the second resampled signal 532) compared to using samples based on the original signals (e.g., the first audio signal 130 and the second audio signal 132) can increase accuracy. The signal comparator 506 may provide the comparison value 534, the experimental shift value 536, or both to the interpolator 510.
[0129] Interpolator 510 can expand the experimental shift value 536. For example, interpolator 510 can generate an interpolation shift value 538, as shown in [reference]. Figure 8 As further described. For example, interpolator 510 can generate an interpolated comparison value corresponding to a shift value close to the experimental shift value 536 by interpolating the comparison value 534. Interpolator 510 can determine the interpolated shift value 538 based on the interpolated comparison value and the comparison value 534. The comparison value 534 can be based on a coarser granularity of the shift values. For example, the comparison value 534 can be based on a first subset of the set of shift values such that the difference between a first shift value in the first subset and each second shift value in the first subset is greater than or equal to a threshold (e.g., ≥1). The threshold can be based on a resampling factor (D).
[0130] The interpolated comparison value can be based on a finer granularity of shift values close to the resampled experimental shift value 536. For example, the interpolated comparison value can be based on a second subset of the set of shift values such that the difference between the highest shift value of the second subset and the resampled experimental shift value 536 is less than the threshold (e.g., ≥1), and the difference between the lowest shift value of the second subset and the resampled experimental shift value 536 is less than the threshold. Determining the comparison value 534 based on a coarser granularity of the set of shift values (e.g., a first subset) uses fewer resources (e.g., time, operations, or both) than determining the comparison value 534 based on a finer granularity of the set of shift values (e.g., all). Determining the interpolated comparison value corresponding to the second subset of shift values can extend the experimental shift value 536 based on a finer granularity of a smaller set of shift values close to the experimental shift value 536, without needing to determine a comparison value corresponding to every shift value in the set of shift values. Therefore, determining the experimental shift value 536 based on a first subset of the shift values and determining the interpolated shift value 538 based on the interpolated comparison value can balance the resource utilization and optimization of the estimated shift values. The interpolator 510 can provide the interpolated shift value 538 to the shift optimizer 511.
[0131] According to one implementation, interpolator 510 can retrieve interpolation shift values from previous frames and can use these interpolation shift values to modify interpolation shift value 538 based on a long-term smoothing operation. For example, interpolation shift value 538 may include the long-term interpolation shift value of the current frame (N). And can be derived from Let be the value of the long-term interpolation shift, where α∈(0,1.0). Therefore, the long-term interpolation shift value... InterVal can be interpolated based on the instantaneous interpolation shift value at frame N. N (k) and long-term interpolation shift values of one or more previous frames The weighted mixture. As the value of α increases, the amount of smoothing in the long-term comparisons increases.
[0132] Shift optimizer 511 can generate corrected shift value 540 by optimizing interpolated shift value 538, as shown in [reference]. Figures 9A to 9C As further described. For example, the shift optimizer 511 can determine whether the interpolated shift value 538 indicates that the shift change between the first audio signal 130 and the second audio signal 132 is greater than a shift change threshold, as shown in [reference]. Figure 9A Further described. The shift change can be obtained by interpolating the shift value 538 and... Figure 3The difference (e.g., change) between the first shift values associated with frame 302 is used as an indication. In response to determining that the difference is less than or equal to a threshold, shift optimizer 511 may set a corrected shift value 540 as an interpolated shift value 538. Alternatively, in response to determining that the difference is greater than a threshold, shift optimizer 511 may determine multiple shift values corresponding to differences less than or equal to the shift change threshold, as shown in [reference]. Figure 9A As further described. The shift optimizer 511 can determine a comparison value based on the first audio signal 130 and multiple shift values applied to the second audio signal 132. The shift optimizer 511 can determine a corrected shift value 540 based on the comparison value, as shown in [reference]. Figure 9A As further described. For example, shift optimizer 511 may select one of the plurality of shift values based on a comparison value and an interpolated shift value 538, as shown in [reference]. Figure 9A As further described. The shift optimizer 511 can set a corrected shift value 540 to indicate the selected shift value. A non-zero difference between the first shift value corresponding to frame 302 and the interpolated shift value 538 can indicate that some samples of the second audio signal 132 correspond to two frames (e.g., frames 302 and 304). For example, some samples of the second audio signal 132 may be duplicated during encoding. Alternatively, a non-zero difference can indicate that some samples of the second audio signal 132 do not correspond to either frame 302 or frame 304. For example, some samples of the second audio signal 132 may be lost during encoding. Setting the corrected shift value 540 to one of a plurality of shift values can prevent large shift changes between consecutive (or adjacent) frames, thereby reducing the amount of sample loss or sample duplication during encoding. The shift optimizer 511 can provide the corrected shift value 540 to the shift change analyzer 512.
[0133] According to one implementation, the shift optimizer can retrieve the corrected shift value of the previous frame and modify the corrected shift value 540 using the corrected shift value of the previous frame based on a long-term smoothing operation. For example, the corrected shift value 540 may include the long-term corrected shift value of the current frame (N). And can be derived from Let be the value of , where α∈(0,1.0). Therefore, the long-term correction shift value... The instantaneous correction shift value AmendVal at frame N can be used. N (k) and one or more long-term correction shift values from previous frames The weighted mixture. As the value of α increases, the amount of smoothing in the long-term comparisons increases.
[0134] In some implementations, the shift optimizer 511 can adjust the interpolation shift value 538, as shown in [reference]. Figure 9BAs described. The shift optimizer 511 can determine a corrected shift value 540 based on an adjusted interpolated shift value 538. In some embodiments, the shift optimizer 511 can determine the corrected shift value 540, as shown in [reference]. Figure 9C As described.
[0135] Shift change analyzer 512 can determine whether the corrected shift value 540 indicates a timing switch or reversal between the first audio signal 130 and the second audio signal 132, as shown in [reference]. Figure 1 As described. Specifically, a timing reversal or switching may indicate that, for frame 302, the first audio signal 130 is received at input interface 112 before the second audio signal 132, and for subsequent frames (e.g., frame 304 or frame 306), the second audio signal 132 is received at input interface before the first audio signal 130. Alternatively, a timing reversal or switching may indicate that, for frame 302, the second audio signal 132 is received at input interface 112 before the first audio signal 130, and for subsequent frames (e.g., frame 304 or frame 306), the first audio signal 130 is received at input interface before the second audio signal 132. In other words, a timing switching or reversal may indicate that the final shift value corresponding to frame 302 has a first sign (e.g., a positive-to-negative transition, or vice versa) that is different from the second sign of the corrected shift value 540 corresponding to frame 304. Based on the corrected shift value 540 and the first shift value associated with frame 302, the shift change analyzer 512 can determine whether the delay between the first audio signal 130 and the second audio signal 132 has switched signs, as shown in [reference]. Figure 10A Further described. In response to determining that the delay between the first audio signal 130 and the second audio signal 132 has switched signs, the shift change analyzer 512 may set the final shift value 116 to a value indicating no time shift (e.g., 0). Alternatively, in response to determining that the delay between the first audio signal 130 and the second audio signal 132 has not yet switched signs, the shift change analyzer 512 may set the final shift value 116 to a corrected shift value 540, as described in [reference]. Figure 10A As further described. The shift change analyzer 512 can generate an estimated shift value by optimizing the corrected shift value 540, as shown in [reference]. Figure 10A , 11As further described. The shift change analyzer 512 can set the final shift value 116 as the estimated shift value. Setting the final shift value 116 to indicate no time shift can reduce distortion at the decoder by avoiding time shifts in opposite directions between the first audio signal 130 and the second audio signal 132 for consecutive (or adjacent) frames of the first audio signal 130. The shift change analyzer 512 can provide the final shift value 116 to the reference signal designator 508, to the absolute shift generator 513, or both. In some embodiments, the shift change analyzer 512 can determine the final shift value 116, as shown in [reference]. Figure 10B As described.
[0136] By applying an absolute function to the final shift value 116, the absolute shift generator 513 can generate a non-causal shift value 162. The absolute shift generator 513 can then provide the non-causal shift value 162 to the gain parameter generator 514.
[0137] Reference signal designator 508 can generate reference signal indicator 164, as shown in [reference]. Figures 12 to 13 As further described. For example, the reference signal indicator 164 may have a first value indicating that the first audio signal 130 is the reference signal or a second value indicating that the second audio signal 132 is the reference signal. The reference signal designator 508 may provide the reference signal indicator 164 to the gain parameter generator 514.
[0138] Gain parameter generator 514 can select samples of a target signal (e.g., a second audio signal 132) based on the non-causal shift value 162. For example, in response to determining that the non-causal shift value 162 has a first value (e.g., +X ms or +Y samples, where X and Y contain positive real numbers), gain parameter generator 514 can select samples 358 to 364. In response to determining that the non-causal shift value 162 has a second value (e.g., -X ms or -Y samples), gain parameter generator 514 can select samples 354 to 360. In response to determining that the non-causal shift value 162 has a value indicating no time shift (e.g., 0), gain parameter generator 514 can select samples 356 to 362.
[0139] Gain parameter generator 514 can determine whether the first audio signal 130 or the second audio signal 132 is the reference signal based on reference signal indicator 164. Based on samples 326 to 332 of frame 304 and selected samples of the second audio signal 132 (e.g., samples 354 to 360, samples 356 to 362, or samples 358 to 364), gain parameter generator 514 can generate gain parameter 160, as shown in [reference]. Figure 1 As described. For example, the gain parameter generator 514 can generate a gain parameter 160 based on one or more of equations 1a to 1f, where gD Corresponding to gain parameter 160, Ref(n) corresponds to a sample of the reference signal, and Targ(n+N1) corresponds to a sample of the target signal. For illustration, when the non-causal shift value 162 has a first value (e.g., +X ms or +Y samples, where X and Y contain positive real numbers), Ref(n) may correspond to samples 326 to 332 of frame 304, and Targ(n+N1) may correspond to samples 358 to 364 of frame 344. In some embodiments, Ref(n) may correspond to a sample of the first audio signal 130, and Targ(n+N1) may correspond to a sample of the second audio signal 132, as shown in [reference]. Figure 1 As described. In an alternative implementation, Ref(n) may correspond to a sample of the second audio signal 132, and Targ(n+N1) may correspond to a sample of the first audio signal 130, as shown in [reference]. Figure 1 As described.
[0140] Gain parameter generator 514 can provide gain parameter 160, reference signal indicator 164, non-causal shift value 162, or a combination thereof to signal generator 516. Signal generator 516 can generate encoded signal 102, as shown in [reference]. Figure 1 As described. For example, the encoded signal 102 may include a first encoded signal frame 564 (e.g., a center channel frame), a second encoded signal frame 566 (e.g., a side channel frame), or both. The signal generator 516 may generate the first encoded signal frame 564 based on equation 2a or equation 2b, where M corresponds to the first encoded signal frame 564, g D Corresponding to gain parameter 160, Ref(n) corresponds to a sample of the reference signal, and Targ(n+N1) corresponds to a sample of the target signal. Signal generator 516 can generate a second encoded signal frame 566 based on equation 3a or equation 3b, where S corresponds to the second encoded signal frame 566, g... D Corresponding to the gain parameter 160, Ref(n) corresponds to the sample of the reference signal, and Targ(n+N1) corresponds to the sample of the target signal.
[0141] The time equalizer 108 may store the following in memory 153: a first resampled signal 530, a second resampled signal 532, a comparison value 534, an experimental shift value 536, an interpolated shift value 538, a corrected shift value 540, a non-causal shift value 162, a reference signal indicator 164, a final shift value 116, a gain parameter 160, a first encoded signal frame 564, a second encoded signal frame 566, or a combination thereof. For example, the analysis data 190 may include the first resampled signal 530, the second resampled signal 532, the comparison value 534, the experimental shift value 536, the interpolated shift value 538, the corrected shift value 540, the non-causal shift value 162, the reference signal indicator 164, the final shift value 116, the gain parameter 160, the first encoded signal frame 564, the second encoded signal frame 566, or a combination thereof.
[0142] The smoothing techniques described above can largely normalize the shift estimates between audio frames, silent frames, and transition frames. Normalized shift estimates reduce sample duplication and artifact skipping at frame boundaries. Furthermore, normalized shift estimates result in reduced side channel energy, which improves decoding efficiency.
[0143] See Figure 6 This illustrates an illustrative example of the system, and the system as a whole is designated as 600. System 600 can correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 600.
[0144] Resampler 504 can be used to... Figure 1 The first audio signal 130 is resampled (e.g., reduced or increased sampling) to generate a first sample 620 of the first resampled signal 530. The resampler 504 can generate a first sample 620 of the first resampled signal 530 by resampling the first audio signal 130. Figure 1 The second audio signal 132 is resampled (e.g., reduced or increased sampling) to produce a second sample 650 of the second resampled signal 532.
[0145] The first audio signal 130 can be sampled at a first sampling rate (Fs) to generate Figure 3 The first sample 320. The first sampling rate (Fs) may correspond to a first rate associated with a wideband (WB) bandwidth (e.g., 16 kHz), a second rate associated with an ultra-wideband (SWB) bandwidth (e.g., 32 kHz), a third rate associated with a full-band (FB) bandwidth (e.g., 48 kHz), or another rate. The second audio signal 132 may be sampled at the first sampling rate (Fs) to generate Figure 3 The second sample was 350.
[0146] In some implementations, resampler 504 may preprocess the first audio signal 130 (or the second audio signal 132) before resampling it. Resampler 504 may preprocess the first audio signal 130 (or the second audio signal 132) by filtering it based on an infinite impulse response (IIR) filter (e.g., a first-order IIR filter). The IIR filter may be based on the following equation:
[0147] H pre (z)=1 / (1-αz -1 Equation 4
[0148] Where α is positive, for example, 0.68 or 0.72. Performing de-emphasis before resampling can reduce effects such as frequency overlap, signal conditioning, or both. The first audio signal 130 (e.g., the pre-processed first audio signal 130) and the second audio signal 132 (e.g., the pre-processed second audio signal 132) can be resampled based on a resampling factor (D). The resampling factor (D) can be based on a first sampling rate (Fs) (e.g., D = Fs / 8, D = 2Fs, etc.).
[0149] In alternative implementations, the first audio signal 130 and the second audio signal 132 may be low-pass filtered or decimated using an anti-aliasing filter before resampling. The decimation filter may be based on a resampling factor (D). In a particular instance, in response to determining that a first sampling rate (Fs) corresponds to a specific rate (e.g., 32 kHz), the resampler 504 may select a decimation filter with a first cutoff frequency (e.g., π / D or π / 4). Reducing aliasing by de-emphasing multiple signals (e.g., the first audio signal 130 and the second audio signal 132) is computationally less expensive than applying a decimation filter to multiple signals.
[0150] The first sample 620 may include samples 622, 624, 626, 628, 630, 632, 634, 636, one or more additional samples, or a combination thereof. The first sample 620 may include... Figure 3 A subset of the first sample 320 (e.g., 1 / 8). Samples 622, 624, one or more additional samples, or a combination thereof, may correspond to frame 302. Samples 626, 628, 630, 632, one or more additional samples, or a combination thereof, may correspond to frame 304. Samples 634, 636, one or more additional samples, or a combination thereof, may correspond to frame 306.
[0151] The second sample 650 may include samples 652, 654, 656, 658, 660, 662, 664, 667, one or more additional samples, or a combination thereof. The second sample 650 may include... Figure 3 The second sample 350 is a subset (e.g., 1 / 8). Samples 654 to 660 may correspond to samples 354 to 360. For example, samples 654 to 660 may contain a subset (e.g., 1 / 8) of samples 354 to 360. Samples 656 to 662 may correspond to samples 356 to 362. For example, samples 656 to 662 may contain a subset (e.g., 1 / 8) of samples 356 to 362. Samples 658 to 664 may correspond to samples 358 to 364. For example, samples 658 to 664 may contain a subset (e.g., 1 / 8) of samples 358 to 364. In some embodiments, the resampling factor may correspond to a first value (e.g., 1), where Figure 6 Samples 622 to 636 and samples 652 to 667 can be similar to Figure 3 Samples 322 to 336 and samples 352 to 366.
[0152] The resampler 504 can store the first sample 620, the second sample 650, or both in the memory 153. For example, the analysis data 190 may contain the first sample 620, the second sample 650, or both.
[0153] See Figure 7 This illustrates an illustrative example of the system, and the system as a whole is designated as 700. System 700 may correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 700.
[0154] Memory 153 may store multiple shift values 760. Shift values 760 may include a first shift value 764 (e.g., -X ms or -Y samples, where X and Y contain positive real numbers), a second shift value 766 (e.g., +X ms or +Y samples, where X and Y contain positive real numbers), or both. Shift values 760 may range from a smaller shift value (e.g., a minimum shift value T_MIN) to a larger shift value (e.g., a maximum shift value T_MAX). Shift values 760 may indicate an expected time shift (e.g., a maximum expected time shift) between the first audio signal 130 and the second audio signal 132.
[0155] During operation, signal comparator 506 can determine comparison value 534 based on first sample 620 and shift value 760 applied to second sample 650. For example, samples 626 to 632 may correspond to a first time (t). For illustration, Figure 1The input interface 112 can receive samples 626 to 632 corresponding to frame 304 at approximately the first time (t). The first shift value 764 (e.g., -X ms or -Y samples, where X and Y contain positive real numbers) can correspond to the second time (t-1).
[0156] Samples 654 to 660 may correspond to a second time (t-1). For example, input interface 112 may receive samples 654 to 660 at approximately the second time (t-1). Signal comparator 506 may determine a first comparison value 714 (e.g., difference, change, or cross-correlation value) corresponding to a first shift value 764 based on samples 626 to 632 and samples 654 to 660. For example, the first comparison value 714 may correspond to the absolute value of the cross-correlation between samples 626 to 632 and samples 654 to 660. As another example, the first comparison value 714 may indicate the difference between samples 626 to 632 and samples 654 to 660.
[0157] The second shift value 766 (e.g., +X ms or +Y samples, where X and Y contain positive real numbers) may correspond to a third time (t+1). Samples 658 to 664 may correspond to the third time (t+1). For example, input interface 112 may receive samples 658 to 664 at approximately the third time (t+1). Signal comparator 506 may determine a second comparison value 716 (e.g., difference, change, or cross-correlation value) corresponding to the second shift value 766 based on samples 626 to 632 and samples 658 to 664. For example, the second comparison value 716 may correspond to the absolute value of the cross-correlation between samples 626 to 632 and samples 658 to 664. As another example, the second comparison value 716 may indicate the difference between samples 626 to 632 and samples 658 to 664. Signal comparator 506 may store comparison value 534 in memory 153. For example, analysis data 190 may contain comparison value 534.
[0158] Signal comparator 506 can identify a selected comparison value 736 of comparison value 534 that has a value larger (or smaller) than other values of comparison value 534. For example, in response to determining that a second comparison value 716 is greater than or equal to a first comparison value 714, signal comparator 506 can select the second comparison value 716 as the selected comparison value 736. In some embodiments, comparison value 534 may correspond to a cross-correlation value. In response to determining that the second comparison value 716 is greater than the first comparison value 714, signal comparator 506 can determine that samples 626 to 632 have a higher correlation with samples 658 to 664 than with samples 654 to 660. Signal comparator 506 can select the second comparison value 716, which indicates a higher correlation, as the selected comparison value 736. In other embodiments, comparison value 534 may correspond to a difference (e.g., a change value). In response to determining that the second comparison value 716 is less than the first comparison value 714, the signal comparator 506 may determine that the similarity between samples 626 to 632 and samples 658 to 664 is greater than the similarity with samples 654 to 660 (e.g., the difference with samples 658 to 664 is less than the difference with samples 654 to 660). The signal comparator 506 may select the second comparison value 716, which indicates the smaller difference, as the selected comparison value 736.
[0159] The selected comparison value 736 may indicate a higher correlation (or a smaller difference) than other values of comparison value 534. Signal comparator 506 may identify a trial shift value 536 corresponding to the selected comparison value 736 for shift value 760. For example, in response to determining that a second shift value 766 corresponds to the selected comparison value 736 (e.g., a second comparison value 716), signal comparator 506 may identify the second shift value 766 as the trial shift value 536.
[0160] Signal comparator 506 can determine the selected comparison value 736 based on the following equation:
[0161]
[0162] Where maxXCorr corresponds to the selected comparison value 736 and k corresponds to the shift value. w(n)*l′ corresponds to the first audio signal 130 after de-emphasis, resampling, and windowing, and w(n)*r′ corresponds to the second audio signal 132 after de-emphasis, resampling, and windowing. For example, w(n)*l′ may correspond to samples 626 to 632, w(n-1)*r′ may correspond to samples 654 to 660, w(n)*r′ may correspond to samples 656 to 662, and w(n+1)*r′ may correspond to samples 658 to 664. -K may correspond to a smaller shift value (e.g., the minimum shift value) of shift value 760, and K may correspond to a larger shift value (e.g., the maximum shift value) of shift value 760. In Equation 5, w(n)*l′ corresponds to the first audio signal 130, regardless of whether the first audio signal 130 corresponds to the right (r) channel signal or the left (l) channel signal. In Equation 5, w(n)*r′ corresponds to the second audio signal 132, regardless of whether the second audio signal 132 corresponds to the right (r) channel signal or the left (l) channel signal.
[0163] Signal comparator 506 can determine the experimental shift value 536 based on the following equation:
[0164]
[0165] Where T corresponds to the experimental shift value 536.
[0166] The signal comparator 506 can be based on Figure 6 The experimental shift value 536 is mapped from the resampled sample to the original sample by a resampling factor (D). For example, the signal comparator 506 may update the experimental shift value 536 based on the resampling factor (D). For illustration, the signal comparator 506 may set the experimental shift value 536 as the product (e.g., 12) of the experimental shift value 536 (e.g., 3) and the resampling factor (D) (e.g., 4).
[0167] See Figure 8 This illustrates an illustrative example of the system, and the system as a whole is designated as 800. System 800 can correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 800. Memory 153 may be configured to store shift value 860. Shift value 860 may include first shift value 864, second shift value 866, or both.
[0168] During operation, interpolator 510 may generate a shift value 860 that is close to the experimental shift value 536 (e.g., 12), as described herein. The mapped shift value may correspond to a shift value 760 that maps from a resampled sample to the original sample based on a resampling factor (D). For example, a first mapped shift value corresponds to the product of a first shift value 764 and the resampling factor (D). The difference between the first mapped shift value and each second mapped shift value may be greater than or equal to a threshold (e.g., the resampling factor (D), such as 4). Shift value 860 may have a finer granularity than shift value 760. For example, the difference between a smaller value (e.g., the minimum value) in shift value 860 and the experimental shift value 536 may be less than a threshold (e.g., 4). The threshold may correspond to... Figure 6 The resampling factor (D). The shift value 860 can be in the range of a first value (e.g., experimental shift value 536 - (threshold - 1)) to a second value (e.g., experimental shift value 536 + (threshold - 1)).
[0169] Interpolator 510 can generate an interpolated comparison value 816 corresponding to the shift value 860 by performing interpolation on the comparison value 534, as described herein. Due to the low granularity of the comparison value 534, comparison values corresponding to one or more of the shift values 860 may not be included in the comparison value 534. Using the interpolated comparison value 816, it may be possible to search for interpolated comparison values corresponding to one or more of the shift values 860 to determine whether the interpolated comparison value corresponding to a specific shift value close to the test shift value 536 indicates a difference compared to the previous value. Figure 7 The second comparison value of 716 indicates a higher correlation (or a smaller difference).
[0170] Figure 8 Figure 820 includes examples illustrating interpolation comparison values 816 and 534 (e.g., cross-correlation values). Interpolator 510 can perform interpolation based on Hanning windowed sinusoidal interpolation, IIR filter-based interpolation, spline interpolation, another form of signal interpolation, or combinations thereof. For example, interpolator 510 can perform Hanning windowed sinusoidal interpolation based on the following equation:
[0171]
[0172] in b corresponds to the windowed sine function. This corresponds to the experimental shift value of 536. This can correspond to a specific comparison value of 534. For example, when i corresponds to 4, This can indicate the first comparison value corresponding to the first shift value (e.g., 8) for comparison value 534. When i corresponds to 0, A second comparison value 716 can be indicated corresponding to the experimental shift value 536 (e.g., 12). When i corresponds to -4, The comparison value 534 can be indicated as a third comparison value corresponding to the third shift value (e.g., 16).
[0173] R(k) 32kHz A specific interpolation value may correspond to the interpolation comparison value 816. Each interpolation value of the interpolation comparison value 816 may correspond to the sum of the products of the windowed sine function (b) with each of the first comparison value, the second comparison value 716, and the third comparison value. For example, the interpolator 510 may determine a first product of the windowed sine function (b) with the first comparison value, a second product of the windowed sine function (b) with the second comparison value 716, and a third product of the windowed sine function (b) with the third comparison value. The interpolator 510 may determine a specific interpolation value based on the sum of the first, second, and third products. The first interpolation value of the interpolation comparison value 816 may correspond to a first shift value (e.g., 9). The windowed sine function (b) may have a first value corresponding to the first shift value. The second interpolation value of the interpolation comparison value 816 may correspond to a second shift value (e.g., 10). The windowed sine function (b) may have a second value corresponding to the second shift value. The first value of the windowed sine function (b) may differ from the second value. Therefore, the first interpolated value may differ from the second interpolated value.
[0174] In Equation 7, 8kHz may correspond to the first rate of comparison value 534. For example, the first rate may indicate the frame contained in comparison value 534 (e.g., Figure 3 The number of comparison values (e.g., 8) of the frame 304). 32kHz may correspond to a second rate in the interpolation comparison value 816. For example, the second rate may indicate the number of comparison values included in the interpolation comparison value 816 corresponding to the frame (e.g., 304). Figure 3 The number of interpolation comparison values (e.g., 32) of frame 304.
[0175] Interpolator 510 can select an interpolation comparison value 838 (e.g., a maximum or minimum value) for interpolation comparison value 816. Interpolator 510 can select a shift value (e.g., 14) for shift value 860 corresponding to interpolation comparison value 838. Interpolator 510 can generate an interpolation shift value 538 indicating the selected shift value (e.g., a second shift value 866).
[0176] Using a coarse method to determine the trial shift value 536 and searching around the trial shift value 536 to determine the interpolation shift value 538 can reduce search complexity without compromising search efficiency or accuracy.
[0177] See Figure 9A This illustrates an illustrative example of the system, and the entire system is designated as 900. System 900 can correspond to... Figure 1System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 900. System 900 may include memory 153, shift optimizer 911, or both. Memory 153 may be configured to store a first shift value 962 corresponding to frame 302. For example, analysis data 190 may include the first shift value 962. The first shift value 962 may correspond to an experimental shift value, interpolated shift value, corrected shift value, final shift value, or non-causal shift value associated with frame 302. Frame 302 may precede frame 304 in the first audio signal 130. Shift optimizer 911 may correspond to... Figure 1 The shift optimizer 511.
[0178] Figure 9A It also includes a flowchart illustrating the operating method, which is designated as 920. Method 920 can be performed by: Figure 1 Time equalizer 108, encoder 114, first device 104; Figure 2 Time equalizer 208, encoder 214, first device 204; Figure 5 Shift optimizer 511; shift optimizer 911; or a combination thereof.
[0179] Method 920 includes determining at 901 whether the absolute value of the difference between the first shift value 962 and the interpolated shift value 538 is greater than a first threshold. For example, shift optimizer 911 may determine whether the absolute value of the difference between the first shift value 962 and the interpolated shift value 538 is greater than a first threshold (e.g., a shift change threshold).
[0180] Method 920 further includes, in response to determining at 901 that the absolute value is less than or equal to a first threshold, setting a correction shift value 540 at 902 to indicate an interpolated shift value 538. For example, in response to determining that the absolute value is less than or equal to a shift change threshold, shift optimizer 911 may set the correction shift value 540 to indicate an interpolated shift value 538. In some embodiments, the shift change threshold may have a first value (e.g., 0) indicating that the correction shift value 540 will be set to the interpolated shift value 538 when the first shift value 962 equals the interpolated shift value 538. In alternative embodiments, the shift change threshold may have a second value (e.g., ≥1) indicating that the correction shift value 540 will be set to the interpolated shift value 538 at 902, offering greater flexibility. For example, the correction shift value 540 may be set to the interpolated shift value 538 for a series of differences between the first shift value 962 and the interpolated shift value 538. For example, when the absolute value of the difference (e.g., -2, -1, 0, 1, 2) between the first shift value 962 and the interpolated shift value 538 is less than or equal to the shift change threshold (e.g., 2), the corrected shift value 540 can be set to the interpolated shift value 538.
[0181] Method 920 further includes determining at 904 whether a first shift value 962 is greater than an interpolated shift value 538 in response to determining at 901 that the absolute value is greater than a first threshold. For example, in response to determining that the absolute value is less than a shift change threshold, shift optimizer 911 may determine whether the first shift value 962 is greater than the interpolated shift value 538.
[0182] Method 920 further includes, in response to determining at 904 that a first shift value 962 is greater than an interpolated shift value 538, setting a smaller shift value 930 at 906 as the difference between the first shift value 962 and a second threshold, and setting a larger shift value 932 as the first shift value 962. For example, in response to determining that a first shift value 962 (e.g., 20) is greater than an interpolated shift value 538 (e.g., 14), shift optimizer 911 may set a smaller shift value 930 (e.g., 17) as the difference between the first shift value 962 (e.g., 20) and a second threshold (e.g., 3). Alternatively, or in an alternative example, shift optimizer 911 may set a larger shift value 932 (e.g., 20) as the first shift value 962 in response to determining that a first shift value 962 is greater than an interpolated shift value 538. The second threshold may be based on the difference between the first shift value 962 and the interpolated shift value 538. In some implementations, the smaller shift value 930 may be set as the difference between the interpolation shift value 538 offset and a threshold (e.g., a second threshold), and the larger shift value 932 may be set as the difference between the first shift value 962 and a threshold (e.g., a second threshold).
[0183] Method 920 further includes, in response to determining at 904 that the first shift value 962 is less than or equal to the interpolated shift value 538, setting the smaller shift value 930 as the first shift value 962 at 910, and setting the larger shift value 932 as the sum of the first shift value 962 and a third threshold. For example, in response to determining that the first shift value 962 (e.g., 10) is less than or equal to the interpolated shift value 538 (e.g., 14), the shift optimizer 911 may set the smaller shift value 930 as the first shift value 962 (e.g., 10). Alternatively, or in an alternative example, the shift optimizer 911 may, in response to determining that the first shift value 962 is less than or equal to the interpolated shift value 538, set the larger shift value 932 (e.g., 13) as the sum of the first shift value 962 (e.g., 10) and a third threshold (e.g., 3). The third threshold may be based on the difference between the first shift value 962 and the interpolated shift value 538. In some embodiments, the smaller shift value 930 may be set as the difference between the first shift value 962 and the threshold (e.g., the third threshold), and the larger shift value 932 may be set as the difference between the interpolated shift value 538 and the threshold (e.g., the third threshold).
[0184] Method 920 further includes, at 908, determining a comparison value 916 based on the first audio signal 130 and a shift value 960 applied to the second audio signal 132. For example, a shift optimizer 911 (or signal comparator 506) may generate the comparison value 916 based on the first audio signal 130 and the shift value 960 applied to the second audio signal 132, as shown in [reference]. Figure 7 As described. For illustration, shift value 960 can range from a smaller shift value 930 (e.g., 17) to a larger shift value 932 (e.g., 20). Shift optimizer 911 (or signal comparator 506) can generate a specific comparison value 916 based on a specific subset of samples 326 to 332 and the second sample 350. The specific subset of the second sample 350 can correspond to a specific shift value (e.g., 17) of shift value 960. The specific comparison value can indicate the difference (or correlation) between samples 326 to 332 and the specific subset of the second sample 350.
[0185] Method 920 further includes, at 912, determining a corrected shift value 540 based on a comparison value 916 (which is generated based on the first audio signal 130 and the second audio signal 132). For example, shift optimizer 911 may determine the corrected shift value 540 based on the comparison value 916. For example, in a first case, when the comparison value 916 corresponds to a cross-correlation value, shift optimizer 911 may determine: corresponding to the interpolated shift value 538... Figure 8 The interpolated comparison value 838 is greater than or equal to the largest comparison value of comparison value 916. Alternatively, when comparison value 916 corresponds to a difference (e.g., a change value), shift optimizer 911 may determine that interpolated comparison value 838 is less than or equal to the smallest comparison value of comparison value 916. In this case, shift optimizer 911 may set the corrected shift value 540 to a smaller shift value 930 (e.g., 17) in response to determining that the first shift value 962 (e.g., 20) is greater than the interpolated shift value 538 (e.g., 14). Alternatively, shift optimizer 911 may set the corrected shift value 540 to a larger shift value 932 (e.g., 13) in response to determining that the first shift value 962 (e.g., 10) is less than or equal to the interpolated shift value 538 (e.g., 14).
[0186] In the second case, when the comparison value 916 corresponds to a cross-correlation value, the shift optimizer 911 can determine that the interpolated comparison value 838 is less than the maximum comparison value of the comparison value 916, and can set the correction shift value 540 to a specific shift value (e.g., 18) of the shift value 960 corresponding to the maximum comparison value. Alternatively, when the comparison value 916 corresponds to a difference (e.g., a change value), the shift optimizer 911 can determine that the interpolated comparison value 838 is greater than the minimum comparison value of the comparison value 916, and can set the correction shift value 540 to a specific shift value (e.g., 18) of the shift value 960 corresponding to the minimum comparison value.
[0187] The comparison value 916 can be generated based on the first audio signal 130, the second audio signal 132, and the shift value 960. The corrected shift value 540 can be generated based on the comparison value 916 using a similar process as performed by the signal comparator 506, as shown in [reference needed]. Figure 7 As described.
[0188] Method 920 thus enables the shift optimizer 911 to limit shift value changes associated with consecutive (or adjacent) frames. Reduced shift value changes can reduce sample loss or sample duplication during encoding.
[0189] See Figure 9B This illustrates an illustrative example of the system, and the entire system is designated as 950. System 950 can correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 950. System 950 may include memory 153, shift optimizer 511, or both. Shift optimizer 511 may include interpolation shift adjuster 958. Interpolation shift adjuster 958 may be configured to selectively adjust interpolation shift value 538 based on first shift value 962, as described herein. Shift optimizer 511 may determine a corrected shift value 540 based on interpolation shift value 538 (e.g., adjusted interpolation shift value 538), as see [reference]. Figure 9A , 9C As described.
[0190] Figure 9B It also includes a flowchart illustrating the operating method, which is designated as 951. Method 951 can be performed by the following: Figure 1 Time equalizer 108, encoder 114, first device 104; Figure 2 Time equalizer 208, encoder 214, first device 204; Figure 5 The shift optimizer 511; Figure 9A The shift optimizer 911; the interpolation shift adjuster 958; or a combination thereof.
[0191] Method 951 includes generating an offset 957 at 952 based on the difference between a first shift value 962 and an unrestricted interpolation shift value 956. For example, an interpolation shift adjuster 958 may generate the offset 957 based on the difference between the first shift value 962 and the unrestricted interpolation shift value 956. The unrestricted interpolation shift value 956 may correspond to an interpolation shift value 538 (e.g., before adjustment by the interpolation shift adjuster 958). The interpolation shift adjuster 958 may store the unrestricted interpolation shift value 956 in memory 153. For example, analysis data 190 may include the unrestricted interpolation shift value 956.
[0192] Method 951 further includes determining at 953 whether the absolute value of offset 957 is greater than a threshold. For example, interpolation shift adjuster 958 may determine whether the absolute value of offset 957 meets the threshold. The threshold may correspond to an interpolation shift limit MAX_SHIFT_CHANGE (e.g., 4).
[0193] Method 951 includes setting an interpolation shift value 538 at 954 based on a first shift value 962, the sign of the offset 957, and the threshold, in response to determining at 953 that the absolute value of the offset 957 is greater than a threshold. For example, the interpolation shift adjuster 958 may limit the interpolation shift value 538 in response to determining that the absolute value of the offset 957 does not meet (e.g., is greater than) the threshold. For example, the interpolation shift adjuster 958 may adjust the interpolation shift value 538 based on the first shift value 962, the sign of the offset 957 (e.g., +1 or -1), and the threshold (e.g., interpolation shift value 538 = first shift value 962 + sign (offset 957) * threshold).
[0194] Method 951 includes setting the interpolation shift value 538 to an unrestricted interpolation shift value 956 at 955 in response to determining at 953 that the absolute value of the offset 957 is less than or equal to a threshold. For example, the interpolation shift adjuster 958 may avoid changing the interpolation shift value 538 in response to determining that the absolute value of the offset 957 satisfies (e.g., less than or equal to) a threshold.
[0195] Method 951 may therefore constrain the interpolation shift value 538 so that the change of the interpolation shift value 538 relative to the first shift value 962 satisfies the interpolation shift constraint.
[0196] See Figure 9C This illustrates an illustrative example of the system, and the system as a whole is designated as 970. System 970 can correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 970. System 970 may include memory 153, shift optimizer 921, or both. Shift optimizer 921 may correspond to... Figure 5 The shift optimizer 511.
[0197] Figure 9C It also includes a flowchart illustrating the operating method, which is designated as 971. Method 971 can be performed by the following: Figure 1 The time equalizer 108, encoder 114, and first device 104 are executed; Figure 2 Time equalizer 208, encoder 214, first device 204; Figure 5 The shift optimizer 511; Figure 9AShift optimizer 911; shift optimizer 921; or a combination thereof.
[0198] Method 971 includes determining at 972 whether the difference between the first shift value 962 and the interpolated shift value 538 is non-zero. For example, shift optimizer 921 may determine whether the difference between the first shift value 962 and the interpolated shift value 538 is non-zero.
[0199] Method 971 includes setting a corrected shift value 540 to the interpolated shift value 538 at 973 in response to determining at 972 that the difference between the first shift value 962 and the interpolated shift value 538 is zero. For example, in response to determining that the difference between the first shift value 962 and the interpolated shift value 538 is zero, the shift optimizer 921 may determine the corrected shift value 540 based on the interpolated shift value 538 (e.g., corrected shift value 540 = interpolated shift value 538).
[0200] Method 971 includes determining at 975 whether the absolute value of offset 957 is greater than a threshold, in response to determining at 972 that the difference between the first shift value 962 and the interpolated shift value 538 is nonzero. For example, in response to determining that the difference between the first shift value 962 and the interpolated shift value 538 is nonzero, shift optimizer 921 may determine whether the absolute value of offset 957 is greater than a threshold. Offset 957 may correspond to the difference between the first shift value 962 and the unrestricted interpolated shift value 956, as shown in [reference]. Figure 9B As described. The threshold may correspond to an interpolation shift limit MAX_SHIFT_CHANGE (e.g., 4).
[0201] Method 971 includes, in response to determining at 972 that the difference between the first shift value 962 and the interpolated shift value 538 is non-zero, or at 975 that the absolute value of the offset 957 is less than or equal to a threshold, setting the smaller shift value 930 at 976 as the difference between the first threshold and the minimum of the first shift value 962 and the interpolated shift value 538, and setting the larger shift value 932 as the sum of the second threshold and the maximum of the first shift value 962 and the interpolated shift value 538. For example, in response to determining that the absolute value of the offset 957 is less than or equal to a threshold, shift optimizer 921 may determine the smaller shift value 930 based on the difference between the first threshold and the minimum of the first shift value 962 and the interpolated shift value 538. Shift optimizer 921 may also determine the larger shift value 932 based on the second threshold and the sum of the maximum of the first shift value 962 and the interpolated shift value 538.
[0202] Method 971 further includes generating a comparison value 916 at 977 based on the first audio signal 130 and a shift value 960 applied to the second audio signal 132. For example, a shift optimizer 921 (or signal comparator 506) may generate the comparison value 916 based on the first audio signal 130 and the shift value 960 applied to the second audio signal 132, as shown in [reference]. Figure 7 As described. Shift value 960 can be in the range of a smaller shift value 930 to a larger shift value 932. Method 971 can proceed to 979.
[0203] Method 971 includes generating a comparison value 915 at 978 based on the first audio signal 130 and an unrestricted interpolation shift value 956 applied to the second audio signal 132, in response to determining at 975 that the absolute value of the offset 957 is greater than a threshold. For example, a shift optimizer 921 (or signal comparator 506) may generate the comparison value 915 based on the first audio signal 130 and the unrestricted interpolation shift value 956 applied to the second audio signal 132, as shown in [reference]. Figure 7 As described.
[0204] Method 971 further includes determining a corrected shift value 540 at 979 based on comparison value 916, comparison value 915, or a combination thereof. For example, shift optimizer 921 may determine the corrected shift value 540 based on comparison value 916, comparison value 915, or a combination thereof, as shown in [reference]. Figure 9A As described. In some implementations, the shift optimizer 921 may determine the corrected shift value 540 based on a comparison of comparison value 915 and comparison value 916 to avoid local maxima caused by shift changes.
[0205] In some cases, the inherent spacing of the first audio signal 130, the first resampled signal 530, the second audio signal 132, the second resampled signal 532, or combinations thereof can interfere with the shift estimation process. In these cases, spacing deemphasis or spacing filtering can be performed to reduce interference caused by spacing and improve the reliability of shift estimation between multiple channels. In some cases, background noise may be present in the first audio signal 130, the first resampled signal 530, the second audio signal 132, the second resampled signal 532, or combinations thereof, and this background noise can interfere with the shift estimation process. In these cases, noise suppression or noise cancellation can be used to improve the reliability of shift estimation between multiple channels.
[0206] See Figure 10A This illustrates an illustrative example of the system, with the entire system designated as 1000. System 1000 can correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 1000.
[0207] Figure 10A It also includes a flowchart illustrating the method of operation as designated as 1020. Method 1020 may be performed by a shift change analyzer 512, a time equalizer 108, an encoder 114, a first device 104, or a combination thereof.
[0208] Method 1020 includes determining at 1001 whether a first shift value 962 is equal to 0. For example, shift change analyzer 512 may determine whether the first shift value 962 corresponding to frame 302 has a first value indicating no time shift (e.g., 0). Method 1020 includes proceeding to 1010 in response to determining that the first shift value 962 is equal to 0 at 1001.
[0209] Method 1020 includes determining at 1002 whether the first shift value 962 is greater than 0 in response to determining that the first shift value 962 is non-zero at 1001. For example, shift change analyzer 512 may determine whether the first shift value 962 corresponding to frame 302 has a first value (e.g., a positive value) indicating that the second audio signal 132 is delayed in time relative to the first audio signal 130.
[0210] Method 1020 includes determining, in response to determining at 1002 that a first shift value 962 is greater than 0, whether a corrected shift value 540 is less than 0. For example, in response to determining that the first shift value 962 has a first value (e.g., a positive value), the shift change analyzer 512 may determine whether the corrected shift value 540 has a second value (e.g., a negative value) indicating a time delay of the first audio signal 130 relative to the second audio signal 132. Method 1020 includes proceeding to 1008 in response to determining at 1004 that the corrected shift value 540 is less than 0. Method 1020 includes proceeding to 1010 in response to determining at 1004 that the corrected shift value 540 is greater than or equal to 0.
[0211] Method 1020 includes determining, in response to determining at 1002 that a first shift value 962 is less than 0, whether a corrected shift value 540 is greater than 0 at 1006. For example, in response to determining that the first shift value 962 has a second value (e.g., a negative value), the shift change analyzer 512 may determine whether the corrected shift value 540 has a first value (e.g., a positive value) indicating a time delay of the second audio signal 132 relative to the first audio signal 130. Method 1020 includes proceeding to 1008 in response to determining at 1006 that the corrected shift value 540 is greater than 0. Method 1020 includes proceeding to 1010 in response to determining at 1006 that the corrected shift value 540 is less than or equal to 0.
[0212] Method 1020 includes setting the final shift value 116 to 0 at 1008. For example, the shift change analyzer 512 may set the final shift value 116 to a specific value (e.g., 0) indicating no time shift.
[0213] Method 1020 includes determining at 1010 whether a first shift value 962 is equal to a corrected shift value 540. For example, shift change analyzer 512 may determine whether the first shift value 962 and the corrected shift value 540 indicate the same time delay between the first audio signal 130 and the second audio signal 132.
[0214] Method 1020 includes setting the final shift value 116 to the corrected shift value 540 at 1012 in response to determining at 1010 that the first shift value 962 is equal to the corrected shift value 540. For example, the shift change analyzer 512 may set the final shift value 116 to the corrected shift value 540.
[0215] Method 1020 includes generating an estimated shift value 1072 at 1014 in response to determining at 1010 that the first shift value 962 is not equal to the corrected shift value 540. For example, the shift change analyzer 512 can determine the estimated shift value 1072 by optimizing the corrected shift value 540, as shown in [reference]. Figure 11 Further description.
[0216] Method 1020 includes setting the final shift value 116 to the estimated shift value 1072 at 1016. For example, the shift change analyzer 512 may set the final shift value 116 to the estimated shift value 1072.
[0217] In some implementations, in response to determining that the delay between the first audio signal 130 and the second audio signal 132 has not switched, the shift change analyzer 512 may set a non-causal shift value 162 to indicate a second estimated shift value. For example, in response to determining that the first shift value 962 is equal to 0 at 1001, that the corrected shift value 540 is greater than or equal to 0 at 1004, or that the corrected shift value 540 is less than or equal to 0 at 1006, the shift change analyzer 512 may set a non-causal shift value 162 to indicate the corrected shift value 540.
[0218] In response to determining the delay between the first audio signal 130 and the second audio signal 132, Figure 3 Switching between frames 304 and 302 allows the shift change analyzer 512 to set a non-causal shift value 162 to indicate no time shift. Preventing the non-causal shift value 162 from switching directions (e.g., from positive to negative or from negative to positive) between consecutive frames reduces distortion in downmixing signal generation at encoder 114, avoids additional delay at decoder for upmixing, or both.
[0219] See Figure 10B This illustrates an illustrative example of the system, and the entire system is designated as 1030. System 1030 may correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104 or both may include one or more components of system 1030.
[0220] Figure 10B It also includes a flowchart illustrating the method of operation as designated as 1031. Method 1031 may be performed by a shift change analyzer 512, a time equalizer 108, an encoder 114, a first device 104, or a combination thereof.
[0221] Method 1031 includes determining at 1032 whether the first shift value 962 is greater than zero and the corrected shift value 540 is less than zero. For example, shift change analyzer 512 can determine whether the first shift value 962 is greater than zero and whether the corrected shift value 540 is less than zero.
[0222] Method 1031 includes setting the final shift value 116 to zero at 1033 in response to determining at 1032 that the first shift value 962 is greater than zero and the corrected shift value 540 is less than zero. For example, in response to determining that the first shift value 962 is greater than zero and the corrected shift value 540 is less than zero, the shift change analyzer 512 may set the final shift value 116 to a first value (e.g., 0) indicating no time shift.
[0223] Method 1031 includes determining at 1034 whether the first shift value 962 is less than or equal to zero and the corrected shift value 540 is greater than or equal to zero, in response to determining at 1032 that the first shift value 962 is less than zero and the corrected shift value 540 is greater than zero. For example, in response to determining that the first shift value 962 is less than or equal to zero or the corrected shift value 540 is greater than or equal to zero, the shift change analyzer 512 may determine whether the first shift value 962 is less than zero and the corrected shift value 540 is greater than zero.
[0224] Method 1031 includes proceeding to 1033 in response to determining that the first shift value 962 is less than zero and the corrected shift value 540 is greater than zero. Method 1031 also includes setting the final shift value 116 to the corrected shift value 540 at 1035 in response to determining that the first shift value 962 is greater than or equal to zero or the corrected shift value 540 is less than or equal to zero. For example, in response to determining that the first shift value 962 is greater than or equal to zero or the corrected shift value 540 is less than or equal to zero, the shift change analyzer 512 may set the final shift value 116 to the corrected shift value 540.
[0225] See Figure 11This illustrates an illustrative example of the system, and the system as a whole is designated as 1100. System 1100 may correspond to Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 1100. Figure 11 It also includes a flowchart illustrating the operation method specified as 1120. Method 1120 can be performed by a shift change analyzer 512, a time equalizer 108, an encoder 114, a first device 104, or a combination thereof. Method 1120 can correspond to Figure 10A Step 1014.
[0226] Method 1120 includes determining at 1104 whether a first shift value 962 is greater than a corrected shift value 540. For example, a shift change analyzer 512 may determine whether the first shift value 962 is greater than the corrected shift value 540.
[0227] Method 1120 further includes, in response to determining at 1104 that a first shift value 962 is greater than a correction shift value 540, setting a first shift value 1130 at 1106 as the difference between the correction shift value 540 and a first offset, and setting a second shift value 1132 as the sum of the first shift value 962 and the first offset. For example, in response to determining that a first shift value 962 (e.g., 20) is greater than a correction shift value 540 (e.g., 18), the shift change analyzer 512 may determine a first shift value 1130 (e.g., 17) (e.g., correction shift value 540 - first offset) based on the correction shift value 540. Alternatively or additionally, the shift change analyzer 512 may determine a second shift value 1132 (e.g., 21) (e.g., first shift value 962 + first offset) based on the first shift value 962. Method 1120 may proceed to 1108.
[0228] Method 1120 further includes, in response to determining at 1104 that a first shift value 962 is less than or equal to a corrected shift value 540, setting a first shift value 1130 as the difference between the first shift value 962 and a second offset, and setting a second shift value 1132 as the sum of the corrected shift value 540 and the second offset. For example, in response to determining that a first shift value 962 (e.g., 10) is less than or equal to a corrected shift value 540 (e.g., 12), the shift change analyzer 512 may determine a first shift value 1130 (e.g., 9) (e.g., first shift value 962 - second offset) based on the first shift value 962. Alternatively or additionally, the shift change analyzer 512 may determine a second shift value 1132 (e.g., 13) (e.g., corrected shift value 540 + first offset) based on the corrected shift value 540. The first offset (e.g., 2) may be different from the second offset (e.g., 3). In some embodiments, the first offset may be the same as the second offset. The first offset, the second offset, or a larger value of both can improve the search range.
[0229] Method 1120 further includes generating a comparison value 1140 at 1108 based on the first audio signal 130 and a shift value 1160 applied to the second audio signal 132. For example, see [link to documentation]. Figure 7 As described, the shift change analyzer 512 can generate a comparison value 1140 based on a first audio signal 130 and a shift value 1160 applied to a second audio signal 132. For example, the shift value 1160 can be within the range of a first shift value 1130 (e.g., 17) to a second shift value 1132 (e.g., 21). The shift change analyzer 512 can generate a specific comparison value of the comparison value 1140 based on samples 326 to 332 and a specific subset of the second sample 350. The specific subset of the second sample 350 can correspond to a specific shift value (e.g., 17) of the shift value 1160. The specific comparison value can indicate the difference (or correlation) between samples 326 to 332 and the specific subset of the second sample 350.
[0230] Method 1120 further includes determining an estimated shift value 1072 at 1112 based on the comparison value 1140. For example, when the comparison value 1140 corresponds to a cross-correlation value, the shift change analyzer 512 may select the largest comparison value of the comparison value 1140 as the estimated shift value 1072. Alternatively, when the comparison value 1140 corresponds to a difference (e.g., a change value), the shift change analyzer 512 may select the smallest comparison value of the comparison value 1140 as the estimated shift value 1072.
[0231] Method 1120 thus enables the shift change analyzer 512 to generate an estimated shift value 1072 by optimizing the corrected shift value 540. For example, the shift change analyzer 512 may determine a comparison value 1140 based on the original sample and may select an estimated shift value 1072 corresponding to the comparison value 1140 that indicates the highest correlation (or minimum difference).
[0232] See Figure 12 This illustrates an illustrative example of the system, and the system as a whole is designated as 1200. System 1200 may correspond to... Figure 1 System 100. For example, Figure 1 System 100, first device 104, or both may include one or more components of system 1200. Figure 12 It also includes a flowchart illustrating the operation method of overall indication 1220. Method 1220 can be performed by reference signal designator 508, time equalizer 108, encoder 114, first device 104, or a combination thereof.
[0233] Method 1220 includes determining at 1202 whether the final shift value 116 is equal to 0. For example, reference signal designator 508 may determine whether the final shift value 116 has a specific value (e.g., 0) indicating no time shift.
[0234] Method 1220 includes keeping the reference signal indicator 164 unchanged at 1204 in response to determining that the final shift value 116 is equal to 0 at 1202. For example, in response to determining that the final shift value 116 has a specific value indicating no time shift (e.g., 0), the reference signal designator 508 may keep the reference signal indicator 164 unchanged. For example, the reference signal indicator 164 may indicate that the same audio signal (e.g., the first audio signal 130 or the second audio signal 132) is the reference signal associated with frame 304, as is frame 302.
[0235] Method 1220 includes determining at 1206 whether the final shift value 116 is greater than 0 in response to determining that the final shift value 116 is non-zero at 1202. For example, in response to determining that the final shift value 116 has a specific value indicating a time shift (e.g., a non-zero value), the reference signal designator 508 may determine whether the final shift value 116 has a first value (e.g., a positive value) indicating a delay of the second audio signal 132 relative to the first audio signal 130, or a second value (e.g., a negative value) indicating a delay of the first audio signal 130 relative to the second audio signal 132.
[0236] Method 1220 includes setting a reference signal indicator 164 at 1208 to have a first value (e.g., 0) indicating that the first audio signal 130 is a reference signal, in response to determining that the final shift value 116 has a first value (e.g., positive). For example, in response to determining that the final shift value 116 has a first value (e.g., positive), the reference signal designator 508 may set the reference signal indicator 164 to have a first value (e.g., 0) indicating that the first audio signal 130 is a reference signal. In response to determining that the final shift value 116 has a first value (e.g., positive), the reference signal designator 508 may determine that the second audio signal 132 corresponds to a target signal.
[0237] Method 1220 includes setting a reference signal indicator 164 at 1210 to have a second value (e.g., 1) indicating that the second audio signal 132 is a reference signal, in response to determining that the final shift value 116 has a second value (e.g., a negative value) indicating that the first audio signal 130 is delayed relative to the second audio signal 132. For example, in response to determining that the final shift value 116 has a second value (e.g., a negative value) indicating that the first audio signal 130 is delayed relative to the second audio signal 132, the reference signal designator 508 may set the reference signal indicator 164 to the second value (e.g., 1) indicating that the second audio signal 132 is a reference signal. In response to determining that the final shift value 116 has a second value (e.g., a negative value), the reference signal designator 508 may determine that the first audio signal 130 corresponds to a target signal.
[0238] Reference signal designator 508 can provide reference signal indicator 164 to gain parameter generator 514. Gain parameter generator 514 can determine the gain parameter of the target signal (e.g., gain parameter 160) based on the reference signal, as shown in [reference]. Figure 5 As described.
[0239] The target signal may be time-delayed relative to the reference signal. Reference signal indicator 164 may indicate whether the first audio signal 130 or the second audio signal 132 corresponds to the reference signal. Reference signal indicator 164 may also indicate whether the gain parameter 160 corresponds to the first audio signal 130 or the second audio signal 132.
[0240] See Figure 13 The flowchart illustrating a specific method of operation is shown and is designated as 1300. Method 1300 may be performed by a reference signal designator 508, a time equalizer 108, an encoder 114, a first device 104, or a combination thereof.
[0241] Method 1300 includes determining at 1302 whether the final shift value 116 is greater than or equal to zero. For example, reference signal designator 508 can determine whether the final shift value 116 is greater than or equal to zero. Method 1300 also includes proceeding to 1208 in response to determining at 1302 that the final shift value 116 is greater than or equal to zero. Method 1300 further includes proceeding to 1210 in response to determining at 1302 that the final shift value 116 is less than zero. Method 1300 differs from... Figure 12 Method 1220 is used because, in response to determining that the final shift value 116 has a specific value indicating no time shift (e.g., 0), the reference signal indicator 164 is set to indicate that the first audio signal 130 corresponds to a first value (e.g., 0) of the reference signal. In some embodiments, the reference signal designator 508 may perform method 1220. In other embodiments, the reference signal designator 508 may perform method 1300.
[0242] When the final shift value 116 indicates no time shift, method 1300 can therefore set the reference signal indicator 164 to indicate that the first audio signal 130 corresponds to a specific value of the reference signal (e.g., 0), regardless of whether the first audio signal 130 corresponds to the reference signal for frame 302.
[0243] See Figure 14 This illustrates an illustrative example of the system, and the entire system is designated as 1400. System 1400 includes... Figure 5 Signal comparator 506 Figure 5 Intercalator 510, Figure 5 The shift optimizer 511 and Figure 5 The shift change analyzer 512.
[0244] Signal comparator 506 may generate a comparison value 534 (e.g., a difference, bias, similarity, coherence, or cross-correlation value), a trial shift value 536, or both. For example, signal comparator 506 may generate comparison value 534 based on a first resampled signal 530 and a plurality of shift values 1450 applied to a second resampled signal 532. Signal comparator 506 may determine the trial shift value 536 based on comparison value 534. Signal comparator 506 includes a smoother 1410 configured to retrieve comparison values of previous frames of resampled signals 530, 532, and may use comparison values of previous frames to modify comparison value 534 based on long-term smoothing operations. For example, comparison value 534 may include long-term comparison values of the current frame (N). And can be derived from Let be the value of , where α∈(0,1.0). Therefore, the long-term comparison value... CompVal can be based on the instantaneous comparison value at frame N. N(k) Long-term comparison value with one or more previous frames The weighted mixture. As the value of α increases, the amount of smoothing in the long-term comparison value increases. The signal comparator 506 can provide the comparison value 534, the experimental shift value 536, or both to the interpolator 510.
[0245] Interpolator 510 can extend the experimental shift value 536 to generate an interpolated shift value 538. For example, interpolator 510 can generate an interpolated comparison value corresponding to a shift value close to the experimental shift value 536 by interpolating the comparison value 534. Interpolator 510 can determine the interpolated shift value 538 based on the interpolated comparison value and the comparison value 534. The comparison value 534 can be based on a coarser granularity of the shift values. The interpolated comparison value can be based on a finer granularity of the shift values close to the resampled experimental shift value 536. Determining the comparison value 534 based on a coarser granularity of the set of shift values (e.g., a first subset) can use fewer resources (e.g., time, operations, or both) than determining the comparison value 534 based on a finer granularity (e.g., all) of the set of shift values. Determining interpolation comparison values corresponding to a second subset of shift values can augment the experimental shift value 536 with a finer granularity based on a smaller set of shift values close to the experimental shift value 536, without needing to determine a comparison value for every shift value in the set of shift values. Therefore, determining the experimental shift value 536 based on a first subset of shift values and determining the interpolation shift value 538 based on the interpolation comparison values balances the resource utilization and optimization of the estimated shift values. The interpolator 510 can then provide the interpolated shift value 538 to the shift optimizer 511.
[0246] Interpolator 510 includes a smoother 1420 configured to retrieve interpolated shift values from previous frames, and can modify interpolated shift values 538 based on long-term smoothing operations using the interpolated shift values from previous frames. For example, interpolated shift values 538 may include long-term interpolated shift values from the current frame (N). And can be derived from Let be the value of the long-term interpolation shift, where α∈(0,1.0). Therefore, the long-term interpolation shift value... InterVal can be interpolated based on the instantaneous interpolation shift value at frame N. N (k) and long-term interpolation shift values of one or more previous frames The weighted mixture. As the value of α increases, the amount of smoothing in the long-term comparisons increases.
[0247] The shift optimizer 511 can generate a corrected shift value 540 by improving the interpolated shift value 538. For example, the shift optimizer 511 can determine whether the interpolated shift value 538 indicates that the shift change between the first audio signal 130 and the second audio signal 132 is greater than a shift change threshold. The shift change can be determined by the interpolated shift value 538 and associated with... Figure 3 The difference between the first shift values of frame 302 is used as an indication. In response to determining that the difference is less than or equal to a threshold, shift optimizer 511 may set a corrected shift value 540 as an interpolated shift value 538. Alternatively, in response to determining that the difference is greater than a threshold, shift optimizer 511 may determine multiple shift values corresponding to differences less than or equal to a shift change threshold. Shift optimizer 511 may determine a comparison value based on the first audio signal 130 and multiple shift values applied to the second audio signal 132. Shift optimizer 511 may determine a corrected shift value 540 based on the comparison value. For example, shift optimizer 511 may select a shift value from the multiple shift values based on the comparison value and the interpolated shift value 538. Shift optimizer 511 may set a corrected shift value 540 to indicate the selected shift value. A non-zero difference between the first shift value corresponding to frame 302 and the interpolated shift value 538 indicates that some samples of the second audio signal 132 correspond to two frames (e.g., frames 302 and 304). For example, some samples of the second audio signal 132 may be duplicated during encoding. Alternatively, a non-zero difference indicates that some samples of the second audio signal 132 do not correspond to either frame 302 or frame 304. For example, some samples of the second audio signal 132 may be lost during encoding. Setting the corrected shift value 540 to one of a plurality of shift values prevents large shift changes between consecutive (or adjacent) frames, thereby reducing the amount of sample loss or sample duplication during encoding. The shift optimizer 511 may provide the corrected shift value 540 to the shift change analyzer 512.
[0248] The shift optimizer 511 includes a smoother 1430 configured to retrieve the corrected shift value of a previous frame, and can modify the corrected shift value 540 based on a long-term smoothing operation using the corrected shift value of the previous frame. For example, the corrected shift value 540 may include the long-term corrected shift value of the current frame (N). And can be derived from Let be the value of , where α∈(0,1.0). Therefore, the long-term correction shift value... The instantaneous correction shift value AmendVal at frame N can be used. N (k) and one or more long-term correction shift values from previous frames The weighted mixture. As the value of α increases, the amount of smoothing in the long-term comparisons increases.
[0249] The shift change analyzer 512 can determine whether the corrected shift value 540 indicates a timing switch or reversal between the first audio signal 130 and the second audio signal 132. The shift change analyzer 512 can determine whether the delay between the first audio signal 130 and the second audio signal 132 has switched signs based on the corrected shift value 540 and a first shift value associated with frame 302. In response to determining that the delay between the first audio signal 130 and the second audio signal 132 has switched signs, the shift change analyzer 512 can set the final shift value 116 to a value indicating no time shift (e.g., 0). Alternatively, in response to determining that the delay between the first audio signal 130 and the second audio signal 132 has not switched signs, the shift change analyzer 512 can set the final shift value 116 to the corrected shift value 540.
[0250] The shift change analyzer 512 can generate an estimated shift value by optimizing the corrected shift value 540. The shift change analyzer 512 can set a final shift value 116 as the estimated shift value. Setting the final shift value 116 to indicate no time shift can reduce distortion at the decoder by avoiding time shifts in opposite directions between the first audio signal 130 and the second audio signal 132 for consecutive (or adjacent) frames of the first audio signal 130. The shift change analyzer 512 can provide the final shift value 116 to the absolute shift generator 513. By applying an absolute function to the final shift value 116, the absolute shift generator 513 can generate a non-causal shift value 162.
[0251] The smoothing techniques described above can largely normalize the shift estimates between audio frames, silent frames, and transition frames. Normalized shift estimates reduce sample duplication and artifact skipping at frame boundaries. Furthermore, normalized shift estimates result in reduced side channel energy, which improves decoding efficiency.
[0252] Such as about Figure 14As described, smoothing can be performed at signal comparator 506, interpolator 510, shift optimizer 511, or a combination thereof. If the interpolated shift consistently differs from the experimental shift at the input sampling rate (FSin), then smoothing of the interpolated shift value 538 can be performed in addition to or as an alternative to smoothing of the comparison value 534. During the estimation of the interpolated shift value 538, the interpolation process can be performed on: a smoothed long-term comparison value generated at signal comparator 506, an unsmoothed comparison value generated at signal comparator 506, or a weighted mixture of the interpolated smoothed comparison value and the interpolated unsmoothed comparison value. If smoothing is performed at interpolator 510, then the interpolation can be extended to be performed near multiple samples other than the provisional shift estimated in the current frame. For example, interpolation can be performed close to shifts in previous frames (e.g., one or more of previous experimental shifts, previous interpolated shifts, previous corrected shifts, or previous final shifts) and close to the experimental shift in the current frame. As a result, smoothing can be performed on additional samples of interpolated shift values, which improves the interpolation shift estimation.
[0253] See Figure 15 The chart displays a comparison of values for frames with sound, transition frames, and silent frames. According to... Figure 15 Figure 1502 illustrates the comparison values (e.g., cross-correlation values) of sound frames processed without using the described long-term smoothing technique, Figure 1504 illustrates the comparison values of transition frames processed without using the described long-term smoothing technique, and Figure 1506 illustrates the comparison values of silent frames processed without using the described long-term smoothing technique.
[0254] The cross-correlation represented in each of the charts 1502, 1504, and 1506 may be largely different. For example, chart 1502 illustrates the cross-correlation caused by… Figure 1 The first microphone 146 retrieved the audio frames and by Figure 1 The peak cross-correlation between corresponding spoken frames retrieved by the second microphone 148 occurs at approximately a 17-sample shift. However, Figure 1504 illustrates that the peak cross-correlation between transition frames retrieved by the first microphone 146 and their corresponding transition frames retrieved by the second microphone 148 occurs at approximately a 4-sample shift. Furthermore, Figure 1506 illustrates that the peak cross-correlation between silent frames retrieved by the first microphone 146 and their corresponding silent frames retrieved by the second microphone 148 occurs at approximately a 3-sample shift. Therefore, the shift estimation for transition frames and silent frames may be inaccurate due to the relatively high noise levels.
[0255] according to Figure 15Figure 1512 illustrates comparison values (e.g., cross-correlation values) of sound frames processed using the described long-term smoothing technique, Figure 1514 illustrates comparison values of transition frames processed using the described long-term smoothing technique, and Figure 1516 illustrates comparison values of silent frames processed using the described long-term smoothing technique. The cross-correlation values in each of Figures 1512, 1514, and 1516 may be substantially similar. For example, each of Figures 1512, 1514, and 1516 illustrates comparison values of silent frames processed using the described long-term smoothing technique. Figure 1 The frame retrieved by the first microphone 146 and by Figure 1 The peak cross-correlation between corresponding frames retrieved by the second microphone 148 appears at approximately a 17-sample shift. Therefore, regardless of noise, the shift estimates for transition frames (illustrated by Figure 1514) and silent frames (illustrated by Figure 1516) are relatively accurate (or similar) for the shift estimates of sound frames.
[0256] When estimating comparison values over the same shift range in each frame, reference can be applied. Figure 15 The described comparison value smoothing process. Smoothing logic (e.g., smoothers 1410, 1420, 1430) can be performed based on the generated comparison values before estimating the shift between channels. For example, smoothing can be performed before estimating trial shifts, interpolation shifts, or correction shifts. To reduce the adjustment of comparison values during silent periods (or background noise that could cause shift estimation drift), the comparison values can be smoothed based on a large time constant (e.g., α = 0.995); alternatively, smoothing can be based on α = 0.9. The determination of whether to adjust the comparison values can be based on whether the background energy or long-term energy is below a threshold.
[0257] See Figure 16 This displays a flowchart illustrating a specific operational method, and its overall designation is 1600. Method 1600 can be derived from... Figure 1 The time equalizer 108, encoder 114, first device 104, or a combination thereof are executed.
[0258] Method 1600 includes retrieving a first audio signal at a first microphone at 1602. The first audio signal may contain a first frame. For example, see Figure 1 The first microphone 146 can retrieve the first audio signal 130. The first audio signal 130 may contain a first frame.
[0259] At position 1604, a second audio signal can be retrieved at the second microphone. The second audio signal may contain a second frame, and the second frame may have content substantially similar to the first frame. For example, see... Figure 1The second microphone 148 can retrieve the second audio signal 132. The second audio signal 132 may include a second frame, and the second frame may have content substantially similar to the first frame. The first frame and the second frame may be one of a frame with sound, a transition frame, or a silent frame.
[0260] At position 1606, the delay between the first and second frames can be estimated. For example, see... Figure 1 The time equalizer 108 can determine the cross-correlation between the first and second frames. At 1608, the temporal offset between the first and second audio signals can be estimated based on the delay and historical delay data. For example, see... Figure 1 The time equalizer 108 can estimate the temporal offset between audio samples retrieved at microphones 146 and 148. The temporal offset can be estimated based on the delay between a first frame of the first audio signal 130 and a second frame of the second audio signal 132, where the second frame contains content substantially similar to the first frame. For example, the time equalizer 108 can use a cross-correlation function to estimate the delay between the first and second frames. The cross-correlation function can be used to measure the similarity between two frames based on the lag of one frame relative to the other. Based on the cross-correlation function, the time equalizer 108 can determine the delay (e.g., lag) between the first and second frames. The time equalizer 108 can estimate the temporal offset between the first audio signal 130 and the second audio signal 132 based on the delay and historical delay data.
[0261] Historical data may include delays between frames retrieved from the first microphone 146 and corresponding frames retrieved from the second microphone 148. For example, the time equalizer 108 may determine cross-correlation (e.g., hysteresis) between a previous frame associated with the first audio signal 130 and a corresponding frame associated with the second audio signal 132. Each hysteresis may be represented by a “comparison value.” That is, the comparison value may indicate a time shift (k) between a frame of the first audio signal 130 and a corresponding frame of the second audio signal 132. According to one embodiment, the comparison value of the previous frame may be stored in memory 153. The smoother 192 of the time equalizer 108 may “smooth” (or average) the comparison values within a long-term frame set and use the long-term smoothed comparison values to estimate the temporal offset (e.g., “shift”) between the first audio signal 130 and the second audio signal 132.
[0262] Therefore, historical delay data can be generated based on smoothed comparison values associated with the first audio signal 130 and the second audio signal 132. For example, method 1600 may include smoothing the comparison values associated with the first audio signal 130 and the second audio signal 132 to generate historical delay data. The smoothed comparison values may be based on frames of the first audio signal 130 that occurred earlier than the first frame and frames of the second audio signal 132 that occurred earlier than the second frame. According to one embodiment, method 1600 may include shifting the second frame by a temporal offset.
[0263] For illustrative purposes, if CompVal N (k) represents the comparison value of frame N at offset k, so frame N can have comparison values from k = T_MIN (minimum shift) to k = T_MAX (maximum shift). Smoothing can be performed to ensure that long-term comparison values... Depend on Let f be an expression. The function f in the above equation can be a function of all (or a subset of) the past comparison values under shift (k). The alternative expression for the above equation can be: The functions f or g can be simple finite impulse response (FIR) filters or infinite impulse response (IIR) filters, respectively. For example, function g can be a single-tap IIR filter such that the long-term comparison value... Depend on Let be the value of , where α∈(0,1.0). Therefore, the long-term comparison value... CompVal can be based on the instantaneous comparison value at frame N. N (k) Long-term comparison value with one or more previous frames The weighted mixture. As the value of α increases, the amount of smoothing in the long-term comparisons increases.
[0264] According to one implementation, method 1600 may include adjusting the range of comparison values used to estimate the delay between the first frame and the second frame, as shown in [reference]. Figures 17 to 18 A more detailed description is provided. The delay can be associated with the comparison value that has the highest cross-correlation within the range of comparison values. The adjustment range can include determining whether the comparison values at the range boundaries are monotonically increasing, and expanding the boundaries in response to the determination that the comparison values at the boundaries are monotonically increasing. The boundaries can include left or right boundaries.
[0265] Figure 16 Method 1600 can generally normalize the shift estimation between audio frames, silent frames, and transition frames. The normalized shift estimation reduces sample duplication and artifact skipping at frame boundaries. Furthermore, the normalized shift estimation results in reduced side channel energy, which improves decoding efficiency.
[0266] See Figure 17The flowchart 1700 illustrates a process for selectively expanding the search range of comparison values used for shift estimation. For example, flowchart 1700 can be used to expand the search range of comparison values based on comparison values generated for the current frame, comparison values generated for past frames, or a combination thereof.
[0267] According to flowchart 1700, the detector can be configured to determine whether a comparison value near the right or left boundary increases or decreases. The search range boundary used for generating future comparison values can be extrapolated based on this determination to accommodate more shift values. For example, when comparison values are regenerated, the search range boundary can be extrapolated for comparison values in subsequent frames or comparison values in the same frame. The detector can initiate search boundary expansion based on comparison values generated for the current frame or based on comparison values generated for one or more previous frames.
[0268] At 1702, the detector can determine whether the comparison value at the right boundary is monotonically increasing. As a non-limiting example, the search range can be expanded from -20 to 20 (e.g., from a 20-sample shift in the negative direction to a 20-sample shift in the positive direction). As used herein, the shift in the negative direction corresponds to the first signal (e.g., Figure 1 The first audio signal 130) is a reference signal and the second signal (e.g., Figure 1 The second audio signal (132) is the target signal. The shift in the positive direction corresponds to the first signal being the target signal and the second signal being the reference signal.
[0269] If the comparison value at the right boundary monotonically increases at 1702, then at 1704, the detector can adjust the right boundary outward to increase the search range. For illustration, if the comparison value at sample shift 19 has a specific value and the comparison value at sample shift 20 has a large value, then the detector can expand the search range in the positive direction. As a non-limiting example, the detector can expand the search range from -20 to 25. The detector can expand the search range in increments of one sample, two samples, three samples, etc. According to one implementation, the determination at 1702 can be performed by detecting comparison values at multiple samples towards the right boundary based on spurious jumps at the right boundary to reduce the likelihood of expanding the search range.
[0270] If the comparison value at the right boundary is not monotonically increasing at 1702, then at 1706, the detector can determine whether the comparison value at the left boundary is monotonically increasing. If the comparison value at the left boundary is monotonically increasing at 1706, then at 1708, the detector can adjust the left boundary outward to increase the search range. For illustration, if the comparison value at sample shift -19 has a specific value and the comparison value at sample shift -20 has a large value, then the detector can expand the search range in the negative direction. As a non-limiting example, the detector can expand the search range from -25 to 20. The detector can expand the search range in increments of one sample, two samples, three samples, etc. According to one embodiment, the determination at 1702 can be performed by detecting the comparison values at multiple samples towards the left boundary based on spurious jumps at the left boundary to reduce the possibility of expanding the search range. If the comparison value at the left boundary is not monotonically increasing at 1706, then at 1710, the detector can keep the search range unchanged.
[0271] therefore, Figure 17 Flowchart 1700 can initiate modifications to the search range for future frames. For example, if the comparison value is detected as monotonically increasing within the last ten shift values before a threshold in the past three consecutive frames (e.g., from sample shift 10 to sample shift 20, or from sample shift -10 to sample shift -20), then the search range can be expanded outward by a specific number of samples. This outward expansion of the search range can be continuously implemented for future frames until the comparison value at the boundary no longer monotonically increases. Increasing the search range based on the comparison values of previous frames reduces the possibility that a "true shift" might be very close to the boundary of the search range but only outside of it. Reducing this possibility leads to improved side channel energy minimization and channel decoding.
[0272] See Figure 18 The diagram illustrates the selective expansion of the search range for comparison values used in shift estimation. This diagram can be manipulated in conjunction with the data in Table 1.
[0273]
[0274]
[0275] Table 1: Data on Expanding the Selective Search Scope
[0276] According to Table 1, if a specific boundary increases with three or more consecutive frames, the detector can expand its search range. First Graph 1802 illustrates the comparison value for frame i-2. According to First Graph 1802, for a consecutive frame, the left boundary does not monotonically increase while the right boundary monotonically increases. Therefore, the search range remains unchanged for the next frame (e.g., frame i-1) and the boundary can be in the range of -20 to 20. Second Graph 1804 illustrates the comparison value for frame i-1. According to Second Graph 1804, for two consecutive frames, the left boundary does not monotonically increase while the right boundary monotonically increases. As a result, the search range remains unchanged for the next frame (e.g., frame i) and the boundary can be in the range of -20 to 20.
[0277] The third chart 1806 illustrates the comparison value for frame i. According to the third chart 1806, for three consecutive frames, the left boundary does not monotonically increase while the right boundary monotonically increases. Because the right boundary monotonically increases for three or more consecutive frames, the search range for the next frame (e.g., frame i+1) can be expanded, and the boundary of the next frame can be in the range of -23 to 23. The fourth chart 1808 illustrates the comparison value for frame i+1. According to the fourth chart 1808, for four consecutive frames, the left boundary does not monotonically increase while the right boundary monotonically increases. Because the right boundary monotonically increases for three or more consecutive frames, the search range for the next frame (e.g., frame i+2) can be expanded, and the boundary of the next frame can be in the range of -26 to 26. The fifth chart 1810 illustrates the comparison value for frame i+2. According to the fifth chart 1810, for five consecutive frames, the left boundary does not monotonically increase while the right boundary monotonically increases. Because the right boundary monotonically increases for three or more consecutive frames, the search range for the next frame (e.g., frame i+3) can be expanded and the boundary of the next frame can be in the range of -29 to 29.
[0278] Chart 1812 (Sixth Figure) illustrates the comparison values for frame i+3. According to Chart 1812, both the left and right boundaries do not monotonically increase. As a result, the search range remains constant for the next frame (e.g., frame i+4) and the boundaries are within the range of -29 to 29. Chart 1814 (Seventh Figure) illustrates the comparison values for frame i+4. According to Chart 1814, for a consecutive frame, both the left and right boundaries do not monotonically increase. As a result, the search range remains constant for the next frame and the boundaries are within the range of -29 to 29.
[0279] according to Figure 18 The left boundary expands along with the right boundary. In an alternative implementation, the left boundary may be pushed inward to compensate for the outward push of the right boundary, maintaining a constant number of shift values estimated for each frame for the comparison value. In another implementation, the left boundary may remain constant when the detector indicates that the right boundary will expand outward.
[0280] According to one implementation, when the detector indicates that a specific boundary will expand outwards, the sample size for that outward expansion can be determined based on comparison values. For example, when the detector determines that the right boundary will expand outwards based on comparison values, a new set of comparison values can be generated over a wider shift search range, and the detector can use the newly generated comparison values and existing comparison values to determine the final search range. For example, for frame i+1, a set of comparison values over a wider shift range of -30 to 30 can be generated. The final search range can be limited based on the comparison values generated within the wider search range.
[0281] although Figure 18 The example in the diagram indicates that the right boundary can be expanded outwards, but if the detector determines that the left boundary will expand, then a similarity function can be executed to expand the left boundary outwards. According to some implementations, absolute limits on the search range can be used to prevent the search range from increasing or decreasing indefinitely. As a non-limiting example, the absolute value of the search range may not be allowed to increase beyond 8.75 milliseconds (e.g., in codec predictions).
[0282] See Figure 19 A specific illustrative example of the system is disclosed, and the system as a whole is designated as 1900. System 1900 includes a first device 104 communicatively coupled to a second device 106 via a network 120.
[0283] The first device 104 includes similar components and can be connected with... Figure 1 The operation is generally similar as described. For example, the first device 104 includes an encoder 114, a memory 153, an input interface 112, a transmitter 110, a first microphone 146, and a second microphone 148. In addition to the final shift value 116, the memory 153 may also contain additional information. For example, the memory 153 may contain... Figure 5 The corrected shift value is 540, the first threshold is 1902, the second threshold is 1904, the first HB decoding mode is 1912, the first LB decoding mode is 1913, the second HB decoding mode is 1914, the second LB decoding mode is 1915, the first number of bits is 1916, and the second number of bits is 1918. Except... Figure 1 In addition to the time equalizer 108 depicted, the encoder 114 may also include a bit allocator 1908 and a decoding mode selector 1910.
[0284] Encoder 114 (or another processor at the first device 104) can be based on regarding Figure 5The described technique determines a final shift value 116 and a corrected shift value 540. As described below, the corrected shift value 540 may also be referred to as a "shift value," and the final shift value 116 may also be referred to as a "second shift value." The corrected shift value can indicate the shift (e.g., time shift) of the first audio signal 130 retrieved by the first microphone 146 relative to the second audio signal 132 retrieved by the second microphone 148. (See also: ...) Figure 5 As described, the final shift value 116 can be based on the corrected shift value 540.
[0285] Bit allocator 1908 can be configured to determine bit allocation based on a final shift value 116 and a corrected shift value 540. For example, bit allocator 1908 can determine the change between the final shift value 116 and the corrected shift value 540. After determining the change, bit allocator 1908 can compare the change with a first threshold 1902. As described below, if the change satisfies the first threshold 1902, then the number of bits allocated to the intermediate signal and the number of bits allocated to the side signal can be adjusted during the encoding operation.
[0286] For illustration, encoder 114 may be configured to generate at least one encoded signal (e.g., encoded signal 102) based on bit allocation. Encoded signal 102 may include a first encoded signal and a second encoded signal. According to one embodiment, the first encoded signal may correspond to an intermediate signal and the second encoded signal may correspond to a side signal. Encoder 114 may generate the intermediate signal (e.g., the first encoded signal) based on the sum of the first audio signal 130 and the second audio signal 132. Encoder 114 may generate the side signal based on the difference between the first audio signal 130 and the second audio signal 132. According to one embodiment, the first encoded signal and the second encoded signal may include low-frequency band signals. For example, the first encoded signal may include a low-frequency band intermediate signal, and the second encoded signal may include a low-frequency band side signal. The first encoded signal and the second encoded signal may include high-frequency band signals. For example, the first encoded signal may include a high-frequency band intermediate signal, and the second encoded signal may include a high-frequency band side signal.
[0287] If the final shift value 116 (e.g., the shift amount used to encode the encoded signal 102) differs from the corrected shift value 540 (e.g., the shift amount calculated to reduce side signal energy), then additional bits can be allocated to side signal decoding compared to a scenario similar to the final shift value 116 and the corrected shift value 540. After allocating the additional bits to side signal decoding, the remaining available bits can be allocated to intermediate signal decoding and to side parameters. Having similar final shift values 116 and 540 can substantially reduce the likelihood of sign reversal in successive frames, substantially reduce the occurrence of large shift jumps between audio signals 130 and 132, and / or allow the target signal to be shifted slowly frame by frame in time. For example, the shift can evolve slowly (e.g., change) because the side channels are not completely decorrelated and because large step shifts can produce artifacts. Additionally, if the shift change exceeds a specific amount per frame and the final shift change is limited, then increased side frame energy may occur. Therefore, extra bits can be allocated to side signal decoding to account for the increased side frame energy.
[0288] For illustration, bit allocator 1908 may allocate a first number of bits 1916 to a first encoded signal (e.g., an intermediate signal) and a second number of bits 1918 to a second encoded signal (e.g., a side signal). For example, bit allocator 1908 may determine the change (or difference) between a final shift value 116 and a corrected shift value 540. After determining the change, bit allocator 1908 may compare the change with a first threshold 1902. In response to the change between the corrected shift value 540 and the final shift value 116 satisfying the first threshold 1902, bit allocator 1908 may decrease the first number of bits 1916 and increase the second number of bits 1918. For example, bit allocator 1908 may decrease the number of bits allocated to the intermediate signal and increase the number of bits allocated to the side signal. According to one implementation, the first threshold 1902 may be equal to a relatively small value (e.g., zero or one) such that extra bits are allocated to the side signal when the final shift value 116 and the corrected shift value 540 are not (substantially) similar.
[0289] As described above, encoder 114 can generate encoded signal 102 based on bit allocation. Furthermore, encoded signal 102 can be based on a decoding mode, and the decoding mode can be based on a correction shift value 540 (e.g., a shift value) and a final shift value 116 (e.g., a second shift value). For example, encoder 114 can be configured to determine the decoding mode based on correction shift value 540 and final shift value 116. As described above, encoder 114 can determine the difference between correction shift value 540 and final shift value 116.
[0290] In response to a difference satisfying a threshold, encoder 114 may generate a first encoded signal (e.g., an intermediate signal) based on a first decoding mode and a second encoded signal (e.g., a side signal) based on a second decoding mode. Examples of decoding modes will be seen in [reference needed]. Figures 21 to 22 Further description. For illustration, according to one embodiment, the first encoded signal includes a low-frequency band intermediate signal and the second encoded signal includes a low-frequency band side signal, and the first decoding mode and the second decoding mode include an algebraic code-excited linear prediction (ACELP) decoding mode. According to another embodiment, the first encoded signal includes a high-frequency band intermediate signal and the second encoded signal includes a high-frequency band side signal, and the first decoding mode and the second decoding mode include a bandwidth-extended (BWE) decoding mode.
[0291] According to one embodiment, in response to the difference between the corrected shift value 540 and the final shift value 116 failing to meet a threshold, encoder 114 may generate an encoded low-frequency band intermediate signal (e.g., a first encoded signal) based on an ACELP decoding mode and may generate an encoded low-frequency band side signal (e.g., a second encoded signal) based on a predictive ACELP decoding mode. In this context, encoded signal 102 may include the encoded low-frequency band intermediate signal and one or more parameters corresponding to the encoded low-frequency band side signal.
[0292] According to a specific implementation, encoder 114 may set a shift change tracking flag based on the determination that at least a second shift value (e.g., the corrected shift value 540 or the final shift value 116 of frame 304) exceeds a certain threshold relative to a first shift value 962 (e.g., the final shift of frame 302). Based on the shift change tracking flag, gain parameter 160 (e.g., estimated target gain), or both, encoder 114 may estimate an energy ratio value or a de-mixing factor (e.g., DMXFAC (as in equations 2c to 2d)). Based on the de-mixing factor (DMXFAC) controlled by the shift change, encoder 114 may determine the bit allocation for frame 304, as shown in the following pseudocode.
[0293] Pseudocode: Generate shift change tracking flags
[0294]
[0295] Pseudocode: Adjust the de-mixing factor based on shift changes and target gain.
[0296]
[0297]
[0298] Pseudocode: Adjust bit allocation based on the mixing factor.
[0299] sideChannel_bits=functionof(downmixFactor,coding mode);
[0300] HighBand_bits=functionof(coder_type,core samplerate,total_bitrate)
[0301] midChannel_bits=total_bits-sideChannel_bits-HB_bits;
[0302] "sideChannel_bits" may correspond to the second number of bits, 1918. "midChannel_bits" may correspond to the first number of bits, 1916. Depending on the specific implementation, sideChannel_bits may be estimated based on a downmixing factor (e.g., DMXFAC), decoding mode (e.g., ACELP, TCX, INACTIVE, etc.), or both. High-band bit allocation (HighBand_bits) may be based on the decoder type (ACELP, with sound, without sound), core sampling rate (12.8kHz or 16kHz core), a fixed total bit rate available for side channel decoding, middle channel decoding, and high-band decoding, or a combination thereof. The remaining number of bits after allocation to side channel decoding and high-band decoding may be allocated for middle channel decoding.
[0303] In certain implementations, the final shift value 116 selected for target channel adjustment may differ from the suggested or actual corrected shift value (e.g., corrected shift value 540). In response to determining that the corrected shift value 540 is greater than a threshold and will result in a large shift or adjustment in the target channel, the state machine (e.g., encoder 114) sets the final shift value 116 to an intermediate value. For example, encoder 114 may set the final shift value 116 to an intermediate value between a first shift value 962 (e.g., the final shift value of a previous frame) and the corrected shift value 540 (e.g., the suggested or corrected shift value of the current frame). When the final shift value 116 differs from the corrected shift value 540, the side channels may not be decorrelated to the maximum extent. Setting the final shift value 116 to an intermediate value (i.e., a non-true or actual shift value, such as that represented by the corrected shift value 540) may result in more bits being allocated to side channel decoding. Side channel allocation can be based directly on shift changes, or indirectly on shift changes to track flags, target gain, downmixing factor DMXFAC, or a combination thereof.
[0304] According to another embodiment, in response to the difference between the corrected shift value 540 and the final shift value 116 failing to meet a threshold, the encoder 114 may generate an encoded high-frequency band intermediate signal (e.g., a first encoded signal) based on a BWE decoding mode and may generate an encoded high-frequency band side signal (e.g., a second encoded signal) based on a blind BWE decoding mode. In this scenario, the encoded signal 102 may include the encoded high-frequency band intermediate signal and one or more parameters corresponding to the encoded high-frequency band side signal.
[0305] The encoded signal 102 may be based on a first sample of the first audio signal 130 and a second sample of the second audio signal 132. The second sample may be time-shifted relative to the first sample by an amount based on a final shift value 116 (e.g., a second shift value). The transmitter 110 may be configured to transmit the encoded signal 102 to the second device 106 via network 120. Upon receiving the encoded signal 102, the second device 106 may, as per [the relevant information], [further details]. Figure 1 The operation is generally similar to that described, so as to output a first output signal 126 at the first speaker 142 and a second output signal 128 at the second speaker 144.
[0306] When the final shift value 116 differs from the corrected shift value 540, Figure 19 System 1900 allows encoder 114 to adjust (e.g., increase) the number of bits allocated to side channel decoding. For example, the final shift value 116 can be (via...) Figure 5 The shift change analyzer 512 is limited to a value different from the corrected shift value 540 to avoid sign reversal in successive frames, avoid large shift jumps, and / or to slowly shift the target signal frame by frame in time to align with the reference signal. In these situations, the encoder 114 may increase the number of bits allocated to the side channel decoding to reduce artifacts. It should be understood that the final shift value 116 may differ from the corrected shift value 540 based on other parameters (e.g., inter-channel preprocessing / analysis parameters (e.g., phonation, spacing, frame energy, speech activity, transient detection, speech / music classification, decoder type, noise level estimation, signal-to-noise ratio (SNR) estimation, signal entropy, etc.), based on inter-channel cross-correlation, and / or based on inter-channel spectral similarity.
[0307] See Figure 20 The diagram illustrates a flowchart of a method 2000 for allocating bits between an intermediate signal and a side signal. Method 2000 can be performed by a bit allocator 1908.
[0308] At 2052, method 2000 includes determining the difference 2057 between the final shift value 116 and the corrected shift value 540. For example, bit allocator 1908 can determine the difference 2057 by subtracting the corrected shift value 540 from the final shift value 116.
[0309] At 2053, method 2000 includes comparing difference 2057 (e.g., the absolute value of difference 2057) with a first threshold 1902. For example, bit allocator 1908 may determine whether the absolute value of the difference is greater than the first threshold 1902. If the absolute value of difference 2057 is greater than the first threshold 1902, then at 2054, bit allocator 1908 may decrease a first number of bits 1916 and increase a second number of bits 1918. For example, bit allocator 1908 may decrease the number of bits allocated to the intermediate signal and increase the number of bits allocated to the side signal.
[0310] If the absolute value of the difference 2057 is not greater than the first threshold 1902, then at 2055, the bit allocator 1908 can determine whether the absolute value of the difference 2057 is less than the second threshold 1904. If the absolute value of the difference 2057 is less than the second threshold 1904, then at 2056, the bit allocator 1908 can increase the first number of bits 1916 and decrease the second number of bits 1918. For example, the bit allocator 1908 can increase the number of bits allocated to the center signal and decrease the number of bits allocated to the side channels. If the absolute value of the difference 2057 is not less than the second threshold 1904, then at 2057, the first number of bits 1916 and the second number of bits 1918 can remain unchanged.
[0311] When the final shift value 116 differs from the corrected shift value 540, Figure 20 Method 2000 allows bit allocator 1908 to adjust (e.g., increase) the number of bits allocated to the side channel decoder. For example, the final shift value 116 can be (via...) Figure 5 The shift change analyzer 512 is limited to a value different from the corrected shift value 540 to avoid sign reversal in successive frames, avoid large shift jumps, and / or shift the target signal slowly frame by frame in time to align with the reference signal. In these situations, the encoder 114 may increase the number of bits allocated to the side channel decoding to reduce artifacts.
[0312] See Figure 21 The diagram illustrates a flowchart of a method 2100 for selecting different decoding modes based on a final shift value 116 and a modified shift value 540. Method 2100 can be executed by a decoding mode selector 1910.
[0313] At 2152, method 2100 includes determining the difference 2057 between the final shift value 116 and the corrected shift value 540. For example, bit allocator 1908 can determine the difference 2057 by subtracting the corrected shift value 540 from the final shift value 2052.
[0314] At 2153, method 2100 includes comparing difference 2057 (e.g., the absolute value of difference 2057) with a first threshold 1902. For example, bit allocator 1908 can determine whether the absolute value of the difference is greater than the first threshold 1902. If the absolute value of difference 2057 is greater than the first threshold 1902, at 2154, decoding mode selector 1910 can select BWE decoding mode as the first HB decoding mode 1912, select ACELP decoding mode as the first LB decoding mode 1913, select BWE decoding mode as the second HB decoding mode 1914, and select ACELP decoding mode as the second LB decoding mode 1915. An illustrative decoding implementation scheme according to this scenario is described as follows: Figure 22 The decoding scheme 2202 is described in the text. According to decoding scheme 2202, the high-frequency band can be encoded using either time-division (TD) or frequency-division (FD) BWE decoding modes.
[0315] Return to view Figure 21 If the absolute value of the difference 2057 is not greater than the first threshold 1902, at 2155, the decoding mode selector 1910 can determine whether the absolute value of the difference 2057 is less than the second threshold 1904. If the absolute value of the difference 2057 is less than the second threshold 1904, at 2156, the decoding mode selector 1910 can select the BWE decoding mode as the first HB decoding mode 1912, the ACELP decoding mode as the first LB decoding mode 1913, the blind BWE decoding mode as the second HB decoding mode 1914, and the predictive ACELP decoding mode as the second LB decoding mode 1915. An illustrative decoding implementation scheme according to this scenario is described as follows: Figure 22 The decoding scheme 2206 is described above. According to the decoding scheme 2206, the high-frequency band can be encoded using TD or FD BWE decoding modes for center channel decoding, and the high-frequency band can be encoded using TD or FD blind BWE decoding modes for side channel decoding.
[0316] Return to view Figure 21 If the absolute value of the difference 2057 is not less than the second threshold 1904, at 2157, the decoding mode selector 1910 can select the BWE decoding mode as the first HB decoding mode 1912, the ACELP decoding mode as the first LB decoding mode 1913, the blind BWE decoding mode as the second HB decoding mode 1914, and the ACELP decoding mode as the second LB decoding mode 1915. An illustrative decoding implementation scheme according to this scenario is described as follows: Figure 22The decoding scheme 2204 is described above. According to the decoding scheme 2204, the high-frequency band can be encoded using TD or FD BWE decoding modes for center channel decoding, and the high-frequency band can be encoded using TD or FD blind BWE decoding modes for side channel decoding.
[0317] Therefore, according to method 2100, decoding scheme 2202 can allocate a large number of bits for side channel decoding, decoding scheme 2204 can allocate a smaller number of bits for side channel decoding, and decoding scheme 2206 can allocate even fewer bits for side channel decoding. When signals 130 and 132 are noise-like signals, decoding mode selector 1910 can encode signals 130 and 132 according to decoding scheme 2208. For example, the side channels can be encoded using residual or predictive decoding. The high-frequency and low-frequency side channels can be encoded using transform domain (e.g., Discrete Fourier Transform (DFT) or Modified Discrete Cosine Transform (MDCT) decoding). When signals 130 and 132 have reduced noise (e.g., music-like signals), decoding mode selector 1910 can encode signals 130 and 132 according to decoding scheme 2210. Decoding scheme 2210 may be similar to decoding scheme 2208; however, the middle channel decoding according to decoding scheme 2210 includes transform-decode-excited (TCX) decoding.
[0318] Figure 21 Method 2100 enables the decoding mode selector 1910 to change the decoding mode for the center and side channels based on the difference between the final shift value 116 and the modified shift value 540.
[0319] See Figure 23This illustration shows an example of an encoder 114 of a first device 104. The encoder 114 includes a signal preprocessor 2302 coupled to an inter-frame shift change analyzer 2306, a reference signal designator 2309, or both, via a shift estimator 2304. The signal preprocessor 2302 can be configured to receive audio signals 2328 (e.g., a first audio signal 130 and a second audio signal 132) and process the audio signals 2328 to generate a first resampled signal 2330 and a second resampled signal 2332. For example, the signal preprocessor 2302 can be configured to downsample or resample the audio signals 2328 to generate resampled signals 2330 and 2332. The shift estimator 2304 can be configured to determine a shift value based on a comparison of the resampled signals 2330 and 2332. The inter-frame shift change analyzer 2306 can be configured to identify audio channels as reference and target signals. The inter-frame shift change analyzer 2306 can also be configured to determine the difference between two shift values. The reference signal designator 2309 can be configured to select one audio signal as a reference signal (e.g., an untime-shifted signal) and select another audio signal as a target signal (e.g., a signal time-shifted relative to the reference signal to align the signal with the reference signal in time).
[0320] Inter-frame shift change analyzer 2306 may be coupled to gain parameter generator 2315 via target signal adjuster 2308. Target signal adjuster 2308 may be configured to adjust the target signal based on the difference between shift values. For example, target signal adjuster 2308 may be configured to perform interpolation on a subset of samples to generate an estimated sample of adjusted samples for generating the target signal. Gain parameter generator 2315 may be configured to determine a gain parameter of a reference signal that is "normalized" (e.g., equalized) the power level of the reference signal relative to the power level of the target signal. Alternatively, gain parameter generator 2315 may be configured to determine a gain parameter of the target signal that is "normalized" (e.g., equalized) the power level of the target signal relative to the power level of the reference signal.
[0321] Reference signal designator 2309 may be coupled to inter-frame shift change analyzer 2306, to gain parameter generator 2315, or both. Target signal adjuster 2308 may be coupled to center-side generator 2310, to gain parameter generator 2315, or both. Gain parameter generator 2315 may be coupled to center-side generator 2310. Center-side generator 2310 may be configured to encode the reference signal and the adjusted target signal to generate at least one encoded signal. For example, center-side generator 2310 may be configured to perform stereo coding to generate a center channel signal 2370 and side channel signals 2372.
[0322] The center-side generator 2310 may be coupled to a bandwidth-extended (BWE) spatial equalizer 2312, a center BWE decoder 2314, a low-band (LB) signal regenerator 2316, or a combination thereof. The LB signal regenerator 2316 may be coupled to an LB-side core decoder 2318, an LB center-side core decoder 2320, or both. The center BWE decoder 2314 may be coupled to a BWE spatial equalizer 2312, an LB center-side core decoder 2320, or both. The BWE spatial equalizer 2312, the center BWE decoder 2314, the LB signal regenerator 2316, the LB-side core decoder 2318, and the LB center-side core decoder 2320 may be configured to perform bandwidth extension and additional decoding, such as LB decoding and mid-band decoding, on the center channel signal 2370, the side channel signal 2372, or both. Performing bandwidth extension and additional decoding may include performing additional signal encoding, generating parameters, or both.
[0323] During operation, the signal preprocessor 2302 may receive an audio signal 2328. The audio signal 2328 may include a first audio signal 130, a second audio signal 132, or both. In a particular embodiment, the audio signal 2328 may include a left channel signal and a right channel signal. In other embodiments, the audio signal 2328 may include other signals. The signal preprocessor 2302 may downsample (or resample) the first audio signal 130 and the second audio signal 132 to generate resampled signals 2330, 2332 (e.g., downsampled first audio signal 130 and downsampled second audio signal 132).
[0324] Shift estimator 2304 can generate shift values based on resampled signals 2330, 2332. In a particular embodiment, shift estimator 2304 can generate a non-causal shift value (NC_SHIFT_INDX) 2361 after performing an absolute value operation. In a particular embodiment, shift estimator 2304 can prevent the next shift value from having a different sign (e.g., positive or negative) than the current shift value. For example, when the shift value of the first frame is negative and the shift value of the second frame is determined to be positive, shift estimator 2304 can set the shift value of the second frame to zero. As another example, when the shift value of the first frame is positive and the shift value of the second frame is determined to be negative, shift estimator 2304 can set the shift value of the second frame to zero. Therefore, in this embodiment, the shift value of the current frame has the same sign (e.g., positive or negative) as the shift value of the previous frame, or the shift value of the current frame is zero.
[0325] The reference signal designator 2309 can select one of the first audio signal 130 and the second audio signal 132 as a reference signal for time periods corresponding to the third and fourth frames. The reference signal designator 2309 can determine the reference signal based on the final shift value 116 from the shift estimator 2304. For example, when the final shift value 116 is negative, the reference signal designator 2309 can identify the second audio signal 132 as the reference signal and the first audio signal 130 as the target signal. When the final shift value 116 is positive or zero, the reference signal designator 2309 can identify the second audio signal 132 as the target signal and the first audio signal 130 as the reference signal. The reference signal designator 2309 can generate a reference signal indicator 2365 with a value indicating the reference signal. For example, when the first audio signal 130 is identified as a reference signal, the reference signal indicator 2365 may have a first value (e.g., a logic zero value), and when the second audio signal 132 is identified as a reference signal, the reference signal indicator 2365 may have a second value (e.g., a logic one value). The reference signal designator 2309 may provide the reference signal indicator 2365 to the inter-frame shift change analyzer 2306 and the gain parameter generator 2315.
[0326] Based on the final shift value 116, the first shift value 2363, the target signal 2342, the reference signal 2340, and the reference signal indicator 2365, the inter-frame shift change analyzer 2306 can generate a target signal indicator 2364. The target signal indicator 2364 indicates the adjusted target channel. For example, a first value (e.g., logic zero) of the target signal indicator 2364 can indicate that the first audio signal 130 is an adjusted target channel, and a second value (e.g., logic one) of the target signal indicator 2364 can indicate that the second audio signal 132 is an adjusted target channel. The inter-frame shift change analyzer 2306 can provide the target signal indicator 2364 to the target signal adjuster 2308.
[0327] Target signal adjuster 2308 may correspond to a sample of the adjusted target signal to generate an adjusted sample (adjusted target signal 2352). Target signal adjuster 2308 may provide the adjusted target signal 2352 to gain parameter generator 2315 and mid-side generator 2310. Gain parameter generator 2315 may generate gain parameter 261 based on reference signal indicator 2365 and adjusted target signal 2352. Gain parameter 261 may normalize (e.g., equalize) the power level of the target signal relative to the power level of the reference signal. Alternatively, gain parameter generator 2315 may receive a reference signal (or a sample thereof) and determine gain parameter 261, which normalizes the power level of the reference signal relative to the power level of the target signal. Gain parameter generator 2315 may provide gain parameter 261 to mid-side generator 2310.
[0328] The center-side generator 2310 can generate a center channel signal 2370, a side channel signal 2372, or both, based on an adjusted target signal 2352, a reference signal 2340, and a gain parameter 261. The center-side generator 2310 can provide the side channel signal 2372 to the BWE spatial equalizer 2312, the LB signal regenerator 2316, or both. The center-side generator 2310 can also provide the center channel signal 2370 to the intermediate BWE decoder 2314, the LB signal regenerator 2316, or both. The LB signal regenerator 2316 can generate an LB intermediate signal 2360 based on the center channel signal 2370. For example, the LB signal regenerator 2316 can generate the LB intermediate signal 2360 by filtering the center channel signal 2370. The LB signal regenerator 2316 can provide the LB intermediate core decoder 2320 with the LB intermediate signal 2360. The LB intermediate core decoder 2320 can generate parameters (e.g., core parameter 2371, parameter 2375, or both) based on the LB intermediate signal 2360. Core parameters 2371, 2375, or both can include excitation parameters, sound parameters, etc. The LB intermediate core decoder 2320 can provide core parameter 2371 to the intermediate BWE decoder 2314, and provide parameter 2375 to the LB-side core decoder 2318, or both. Core parameter 2371 can be the same as or different from parameter 2375. For example, core parameter 2371 can include one or more of parameters 2375, may not include one or more of parameters 2375, may include one or more additional parameters, or a combination thereof. Based on the intermediate channel signal 2370, core parameter 2371, or a combination thereof, the intermediate BWE decoder 2314 can generate a decoded intermediate BWE signal 2373. Based on the middle channel signal 2370, core parameters 2371, or a combination thereof, the intermediate BWE decoder 2314 can also generate a set of first gain parameters 2394 and LPC parameters 2392. The intermediate BWE decoder 2314 can provide the decoded intermediate BWE signal 2373 to the BWE spatial equalizer 2312. Based on the decoded intermediate BWE signal 2373, the left HB signal 2396 (e.g., the high-frequency band portion of the left channel signal), the right HB signal 2398 (e.g., the high-frequency band portion of the right channel signal), or a combination thereof, the BWE spatial equalizer 2312 can generate parameters (e.g., one or more gain parameters, spectrum adjustment parameters, other parameters, or combinations thereof).
[0329] The LB signal regenerator 2316 can generate an LB side signal 2362 based on the side channel signal 2342. For example, the LB signal regenerator 2316 can generate the LB side signal 2362 by filtering the side channel signal 2342. The LB signal regenerator 2316 can provide the LB side signal 2362 to the LB side core decoder 2318.
[0330] therefore, Figure 23 System 2300 generates an encoded signal based on the adjusted target channel (e.g., an output signal generated at the LB-side core decoder 2318, the LB intermediate core decoder 2320, the intermediate BWE decoder 2314, the BWE spatial equalizer 2312, or a combination thereof). Adjusting the target channel based on the difference between shift values can compensate for (or hide) inter-frame discontinuities, which can reduce clicks or other audio noises during playback of the encoded signal.
[0331] See Figure 24 Figure 2400 illustrates different encoded signals according to the techniques described herein. For example, encoded HB intermediate signal 2102, encoded LB intermediate signal 2104, encoded HB side signal 2108, and encoded LB side signal 2110 are shown.
[0332] The encoded intermediate signal 2102 includes LPC parameter 2392 and a set of first gain parameters 2394. LPC parameter 2392 may indicate a high-frequency band line spectral frequency (LSF) index. The set of first gain parameters 2394 may indicate a gain frame index, a gain shape index, or both. The encoded HB-side signal 2108 includes LPC parameter 2492 and a set of gain parameters 2494. LPC parameter 2492 may indicate a high-frequency band LSF index. The set of gain parameters 2494 may indicate a gain frame index, a gain shape index, or both. The encoded LB intermediate signal 2104 may include core parameter 2371, and the encoded LB-side signal 2110 may include core parameter 2471.
[0333] See Figure 25 This document demonstrates a system 2500 for encoding signals according to the techniques described herein. The system 2500 includes a downmixer 2502, a preprocessor 2504, an intermediate decoder 2506, a first HB intermediate decoder 2508, a second HB intermediate decoder 2509, a side decoder 2510, and an HB side decoder 2512.
[0334] Audio signal 2528 may be provided to demixer 2502. According to one embodiment, audio signal 2528 may include a first audio signal 130 and a second audio signal 132. Demixer 2502 may perform a demixing operation to generate a center channel signal 2370 and a side channel signal 2372. The center channel signal 2370 may be provided to preprocessor 2504, and the side channel signal 2372 may be provided to side decoder 2510.
[0335] The preprocessor 2504 can generate preprocessing parameters 2570 based on the center channel signal 2370. The preprocessing parameters 2570 may include a first number of bits 1916, a second number of bits 1918, a first HB decoding mode 1912, a first LB decoding mode 1913, a second HB decoding mode 1914, and a second LB decoding mode 1915. The center channel signal 2370 and the preprocessing parameters 2570 can be provided to the intermediate decoder 2506. Based on the decoding mode, the intermediate decoder 2506 can be selectively coupled to the first HB intermediate decoder 2508 or to the second HB intermediate decoder 2509. The side decoder 2510 can be coupled to the HB side decoder 2512.
[0336] See Figure 26 The flowchart illustrates method 2600 for communication. Method 2600 can be derived from... Figure 1 and 19 The first device 104 is executed.
[0337] Method 2600 includes determining a shift value and a second shift value at 2602 at the device. The shift value may indicate a shift of the first audio signal relative to the second audio signal, and the second shift value may be based on the shift value. For example, see Figure 19 The encoder 114 (or another processor at the first device 104) can be based on the information regarding Figure 5 The described technique determines a final shift value 116 and a corrected shift value 540. Regarding method 2600, the corrected shift value 540 may also be referred to as a "shift value," and the final shift value 116 may also be referred to as a "second shift value." The corrected shift value may indicate a shift (e.g., time shift) of the first audio signal 130 retrieved by the first microphone 146 relative to the second audio signal 132 retrieved by the second microphone 148. (See also...) Figure 5 As described, the final shift value 116 can be based on the corrected shift value 540.
[0338] Method 2600 further includes, at 2604, determining a bit allocation at the device based on a second shift value and a shift value. For example, see... Figure 19Bit allocator 1908 can determine bit allocation based on the final shift value 116 and the modified shift value 540. For example, bit allocator 1908 can determine the difference between the final shift value 116 and the modified shift value 540. In the case where the final shift value 116 is different from the modified shift value 540, additional bits can be allocated to side signal decoding compared to a scenario similar to the case where the final shift value 116 and the modified shift value 540 are different. After allocating the additional bits to side signal decoding, the remaining available bits can be allocated to intermediate signal decoding and to side parameters. Having similar final shift values 116 and modified shift values 540 can substantially reduce the likelihood of sign reversal in successive frames, substantially reduce the occurrence of large shift jumps between audio signals 130 and 132, and / or allow the target signal to be shifted slowly frame by frame in time.
[0339] Method 2600 further includes generating at least one encoded signal at the device based on bit allocation at 2606. The at least one encoded signal may be based on a first sample of a first audio signal and a second sample of a second audio signal. The second sample may be time-shifted relative to the first sample by an amount based on a second shift value. For example, see... Figure 19 Encoder 114 can generate at least one encoded signal (e.g., encoded signal 102) based on bit allocation. Encoded signal 102 may include a first encoded signal and a second encoded signal. According to one embodiment, the first encoded signal may correspond to an intermediate signal and the second encoded signal may correspond to a side signal. Encoded signal 102 may be based on a first sample of the first audio signal 130 and a second sample of the second audio signal 132. The second sample may be time-shifted relative to the first sample by an amount based on a final shift value 116 (e.g., a second shift value).
[0340] Method 2600 further includes transmitting at least one encoded signal to the second device at 2608. For example, see... Figure 19 The transmitter 110 can transmit the encoded signal 102 to the second device 106 via the network 120. Upon receiving the encoded signal 102, the second device 106 can, as per... Figure 1 The operation is generally similar to that described, so as to output a first output signal 126 at the first speaker 142 and a second output signal 128 at the second speaker 144.
[0341] According to one embodiment, method 2600 includes determining a bit allocation having a first value in response to the difference between a shift value and a second shift value satisfying a threshold. At least one encoded signal may include a first encoded signal and a second encoded signal. The first encoded signal may correspond to an intermediate signal and the second encoded signal may correspond to a side signal. The bit allocation may indicate that a first number of bits are allocated to the first encoded signal and a second number of bits are allocated to the second encoded signal. Method 2600 may further include decreasing a first number of bits and increasing a second number of bits in response to the difference between the shift value and the second shift value satisfying the first threshold.
[0342] According to one embodiment, method 2600 may include generating an intermediate signal based on the sum of a first audio signal and a second audio signal. Method 2600 may also include generating a side signal based on the difference between the first audio signal and the second audio signal. According to one embodiment of method 2600, the first encoded signal includes a low-frequency band intermediate signal and the second encoded signal includes a low-frequency band side signal. According to another embodiment of method 2600, the first encoded signal includes a high-frequency band intermediate signal and the second encoded signal includes a high-frequency band side signal.
[0343] According to one embodiment, method 2600 includes determining a decoding mode based on a shift value and a second shift value. At least one encoded signal may be based on the decoding mode. Method 2600 may further include generating a first encoded signal based on a first decoding mode and generating a second encoded signal based on a second mode in response to the difference between the shift value and the second shift value satisfying a threshold. At least one encoded signal may include the first encoded signal and the second encoded signal. According to one embodiment, the first encoded signal may include a low-frequency band intermediate signal, and the second encoded signal may include a low-frequency band side signal. The first decoding mode and the second decoding mode may include an ACELP decoding mode. According to another embodiment, the first encoded signal may include a high-frequency band intermediate signal, and the second encoded signal may include a high-frequency band side signal. The first decoding mode and the second decoding mode may include a BWE decoding mode.
[0344] According to one embodiment, method 2600 includes generating an coded low-frequency band intermediate signal based on an ACELP decoding pattern and generating an coded low-frequency band side signal based on a predictive ACELP decoding pattern. At least one coded signal may include the coded low-frequency band intermediate signal and one or more parameters corresponding to the coded low-frequency band side signal.
[0345] According to one embodiment, method 2600 includes generating an encoded high-frequency band intermediate signal based on a BWE decoding mode in response to the difference between the shift value and the second shift value failing to meet a threshold. Method 2600 may further include generating an encoded high-frequency band side signal based on a blind BWE decoding mode in response to the difference failing to meet the threshold. At least one encoded signal may include the encoded high-frequency band intermediate signal and one or more parameters corresponding to the encoded high-frequency band side signal.
[0346] When the final shift value 116 differs from the corrected shift value 540, Figure 6 Method 2600 allows encoder 114 to adjust (e.g., increase) the number of bits allocated to side channel decoding. For example, the final shift value 116 can be (via...) Figure 5 The shift change analyzer 512 is limited to a value different from the corrected shift value 540 to avoid sign reversal in successive frames, avoid large shift jumps, and / or shift the target signal slowly frame by frame in time to align with the reference signal. In these situations, the encoder 114 may increase the number of bits allocated to the side channel decoding to reduce artifacts.
[0347] See Figure 27 The flowchart illustrates method 2700 for communication. Method 2700 can be derived from... Figure 1 and 19 The first device 104 is executed.
[0348] Method 2700 may include determining a shift value and a second shift value at 2702 at the device. The shift value may indicate a shift of the first audio signal relative to the second audio signal, and the second shift value may be based on the shift value. For example, see Figure 19 The encoder 114 (or another processor at the first device 104) can be based on the information regarding Figure 5 The described technique determines a final shift value 116 and a corrected shift value 540. Regarding method 2700, the corrected shift value 540 may also be referred to as a "shift value," and the final shift value 116 may also be referred to as a "second shift value." The corrected shift value may indicate a shift (e.g., time shift) of the first audio signal 130 retrieved by the first microphone 146 relative to the second audio signal 132 retrieved by the second microphone 148. (See also...) Figure 5 As described, the final shift value 116 can be based on the corrected shift value 540.
[0349] Method 2700 may further include determining a decoding mode at the device based on a second shift value and the second shift value at 2704. Method 2700 may further include generating at least one encoded signal at the device based on the decoding mode at 2706. The at least one encoded signal may be based on a first sample of a first audio signal and a second sample of a second audio signal. The second sample may be time-shifted relative to the first sample by an amount based on the second shift value. For example, see [link to documentation]. Figure 19 Encoder 114 may generate at least one encoded signal (e.g., encoded signal 102) based on a decoding mode. Encoded signal 102 may include a first encoded signal and a second encoded signal. According to one embodiment, the first encoded signal may correspond to an intermediate signal and the second encoded signal may correspond to a side signal. Encoded signal 102 may be based on a first sample of the first audio signal 130 and a second sample of the second audio signal 132. The second sample may be time-shifted relative to the first sample by an amount based on a final shift value 116 (e.g., a second shift value).
[0350] Method 2700 may further include transmitting at least one encoded signal to the second device at 2708. For example, see... Figure 19 The transmitter 110 can transmit the encoded signal 102 to the second device 106 via the network 120. Upon receiving the encoded signal 102, the second device 106 can, as per... Figure 1 The operation is generally similar to that described, so as to output a first output signal 126 at the first speaker 142 and a second output signal 128 at the second speaker 144.
[0351] Method 2700 may further include generating a first encoded signal based on a first decoding mode and a second encoded signal based on a second mode in response to the difference between the shift value and the second shift value satisfying a threshold. At least one encoded signal may include the first encoded signal and the second encoded signal. According to one embodiment, the first encoded signal may include a low-frequency band intermediate signal, and the second encoded signal may include a low-frequency band side signal. The first decoding mode and the second decoding mode may include an ACELP decoding mode. According to another embodiment, the first encoded signal may include a high-frequency band intermediate signal, and the second encoded signal may include a high-frequency band side signal. The first decoding mode and the second decoding mode may include a BWE decoding mode.
[0352] According to one embodiment, method 2700 may further include generating an coded low-frequency band intermediate signal based on an ACELP decoding mode and an coded low-frequency band side signal based on a predictive ACELP decoding mode in response to the difference between the shift value and the second shift value failing to meet a threshold. At least one coded signal may include the coded low-frequency band intermediate signal and one or more parameters corresponding to the coded low-frequency band side signal.
[0353] According to another embodiment, method 2700 may further include generating an encoded high-frequency band intermediate signal based on a BWE decoding mode and an encoded high-frequency band side signal based on a blind BWE decoding mode in response to the difference between the shift value and the second shift value failing to meet a threshold. At least one encoded signal may include the encoded high-frequency band intermediate signal and one or more parameters corresponding to the encoded high-frequency band side signal.
[0354] According to one embodiment, in response to the difference between the shift value and the second shift value satisfying a first threshold but failing to satisfy a second threshold, method 2700 may include generating an coded low-frequency band intermediate signal and an coded low-frequency band side signal based on an ACELP decoding mode. Method 2700 may also include generating an coded high-frequency band signal based on a BWE decoding mode and generating an coded high-frequency band side signal based on a blind BWE decoding mode. At least one coded signal may include an coded high-frequency band intermediate signal, an coded low-frequency band intermediate signal, an coded low-frequency band side signal, and one or more parameters corresponding to the coded high-frequency band side signal.
[0355] According to one embodiment, method 2700 may include determining a bit allocation based on a second shift value and a shift value. At least one encoded signal may be generated based on the bit allocation. The at least one encoded signal may include a first encoded signal and a second encoded signal. The bit allocation may indicate that a first number of bits are allocated to the first encoded signal and a second number of bits are allocated to the second encoded signal. Method 2700 may further include decreasing a first number of bits and increasing a second number of bits in response to the difference between the shift value and the second shift value satisfying a first threshold.
[0356] See Figure 28 The flowchart illustrates method 2800 for communication. Method 2800 can be derived from... Figure 1 and 19 The first device 104 is executed.
[0357] Method 2800 includes determining, at 2802, a first mismatch value at the device that indicates a first amount of time mismatch between a first audio signal and a second audio signal. For example, referring to FIG9, encoder 114 (or another processor at the first device 104) may determine a first shift value 962, as described with reference to FIG9. Regarding method 2800, the first shift value 962 may also be referred to as a “first mismatch value.” The first shift value 962 may indicate a first amount of time mismatch between the first audio signal 130 and the second audio signal 132, as described with reference to FIG9. The first shift value 962 may be associated with a first frame to be encoded. For example, the first frame to be encoded may contain… Figure 3Samples 322 to 324 of frame 302 and specific samples of the second audio signal 132. Specific samples can be selected based on a first shift value 962, as shown in [reference]. Figure 1 As described.
[0358] Method 2800 further includes determining a second mismatch value at 2804 at the device, the second mismatch value indicating a second amount of time mismatch between the first audio signal and the second audio signal. For example, encoder 114 (or another processor at the first device 104) may determine an experimental shift value 536, an interpolated shift value 538, a corrected shift value 540, or a combination thereof, as shown in [reference]. Figure 5 As described. Regarding method 2800, the experimental shift value 536, interpolation shift value 538, or correction shift value 540 may also be referred to as a "second mismatch value". One or more of the experimental shift value 536, interpolation shift value 538, or correction shift value 540 may indicate a second amount of time mismatch between the first audio signal 130 and the second audio signal 132. The second mismatch value may be associated with a second frame to be encoded. For example, the second frame to be encoded may contain samples 326 to 332 of the first audio signal 130 and samples 354 to 360 of the second audio signal 132, as shown in [reference]. Figure 4 As described. As another example, the second frame to be encoded may contain samples 326 to 332 of the first audio signal 130 and samples 358 to 364 of the second audio signal 132, as shown in [reference]. Figure 3 As described.
[0359] The second frame to be encoded may follow the first frame to be encoded. For example, at least some samples associated with the second frame to be encoded may follow at least some samples of the first frame to be encoded in the first sample 320 of the first audio signal 130 or in the second sample 350 of the second audio signal 132. In a particular aspect, samples 326 to 332 of the second frame to be encoded may follow samples 322 to 324 of the first frame to be encoded in the first sample 320 of the first audio signal 130. For illustration, each of samples 326 to 332 may be associated with a timestamp indicating a time later than the time indicated by the timestamp associated with any of samples 322 to 324. In some aspects, samples 354 to 360 (or samples 358 to 364) of the second frame to be encoded may follow specific samples of the first frame to be encoded in the second sample 350 of the second audio signal 132.
[0360] Method 2800 further includes, at 2806, determining an effective mismatch value at the device based on a first mismatch value and a second mismatch value. For example, encoder 114 (or another processor at the first device 104) can determine the effective mismatch value based on... Figure 5The described technique determines a corrected shift value 540, a final shift value 116, or both. Regarding method 2800, the corrected shift value 540 or the final shift value 116 may also be referred to as a "valid mismatch value." Encoder 114 may identify one of a first shift value 962 or a second mismatch value as a first value. For example, in response to determining that the first shift value 962 is less than or equal to a second mismatch value, encoder 114 identifies the first shift value 962 as the first value. Encoder 114 may identify the other of the first shift value 962 or the second mismatch value as the second value.
[0361] Encoder 114 (or another processor at the first device 104) can generate a valid mismatch value that is greater than or equal to a first value and less than or equal to a second value. For example, in response to determining that the first shift value 962 is greater than 0 and the corrected shift value 540 is less than 0, or the first shift value 962 is less than 0 and the corrected shift value 540 is greater than 0, encoder 114 can generate a final shift value 116 equal to a specific value indicating no time shift (e.g., 0), as see Figure 10A and 10B As described. In this example, the final shift value 116 may be referred to as the "effective mismatch value" and the corrected shift value 540 may be referred to as the "second mismatch value".
[0362] As another example, encoder 114 can produce a final shift value 116 equal to the estimated shift value 1072, as shown in [reference]. Figure 10A and 11 As described. The estimated shift value 1072 may be greater than or equal to the difference between the corrected shift value 540 and the first offset, and less than or equal to the sum of the first shift value 962 and the first offset. Alternatively, the estimated shift value 1072 may be greater than or equal to the difference between the first shift value 962 and the second offset, and less than or equal to the sum of the corrected shift value 540 and the second offset, as described in [reference]. Figure 11 As described. In this example, the final shift value 116 may be referred to as the "effective mismatch value" and the corrected shift value 540 may be referred to as the "second mismatch value".
[0363] In a particular aspect, encoder 114 may generate a corrected shift value 540 that is greater than or equal to the smaller shift value 930 and less than or equal to the larger shift value 932, as described with reference to FIG. 9. The smaller shift value 930 may be based on the smaller of a first shift value 962 or an interpolated shift value 538. The larger shift value 932 may be based on the other of the first shift value 962 or the interpolated shift value 538. In this aspect, the interpolated shift value 538 may be referred to as the “second mismatch value” and the corrected shift value 540 or the final shift value 116 may be referred to as the “effective mismatch value”. Samples 358 to 364 (or samples 354 to 360) of the second sample 350 may be selected at least in part based on the effective mismatch value, as described with reference to FIG. 9. Figure 1 and3 The above five descriptions are provided.
[0364] Method 2800 further includes generating at least one encoded signal with bit allocation based at least in part on the second frame to be encoded. For example, encoder 114 (or another processor at the first device 104) may generate encoded signal 102 based on the second frame to be encoded, as shown in [reference]. Figure 1 As described. For illustration, encoder 114 can generate encoded signal 102 by encoding samples 326 to 332 and samples 354 to 360, as shown in [reference]. Figure 1 and 4 As described. Alternatively, encoder 114 can generate encoded signal 102 by encoding samples 326 to 332 and samples 358 to 364, as shown in [reference]. Figure 1 and 3 As described.
[0365] The encoded signal 102 may have bit allocations, as described with reference to FIG. 9. For example, the bit allocation may indicate that a first number of bits 1916 is allocated to a first encoded signal (e.g., an intermediate signal), a second number of bits 1918 is allocated to a second encoded signal (e.g., a side signal), or both. The encoder 114 (or another processor at the first device 104) may generate a first encoded signal (e.g., an intermediate signal) having a first bit allocation corresponding to the first number of bits 1916, a second encoded signal (e.g., a side signal) having a second bit allocation corresponding to the second number of bits 1918, or both, as described with reference to FIG. 9.
[0366] Method 2800 further includes transmitting at least one encoded signal to the second device at 2810. For example, see... Figure 19 The transmitter 110 can transmit the encoded signal 102 to the second device 106 via the network 120. Upon receiving the encoded signal 102, the second device 106 can, as per... Figure 1 The operation is generally similar to that described, so as to output a first output signal 126 at the first speaker 142 and a second output signal 128 at the second speaker 144.
[0367] Method 2800 may also include generating a first bit allocation associated with the first frame to be encoded, as shown in [reference]. Figure 19 As described. The first bit allocation can indicate that a second number of bits are allocated to the first encoded signal. The bit allocation associated with the second frame to be encoded can indicate that a specific number is allocated to the encoded signal 102. The specific number can be greater than, less than, or equal to the second number. For example, encoder 114 can generate one or more first encoded signals with the first bit allocation based on a first number of bits 1916, a second number of bits 1918, or both, as shown in [reference]. Figure 1 As described. Encoder 114 can generate a first encoded signal by selecting samples from encoded samples 322 to 324 and second sample 350, as shown in [reference]. Figure 3 As described. Encoder 114 can update a first number of bits 1916, a second number of bits 1918, or both, as shown in [reference]. Figure 20 As described. For example, encoder 114 may generate encoded signal 102 having a bit allocation corresponding to a first number 1916 of updated bits, a second number 1918 of updated bits, or both, as shown in [reference]. Figure 20 As described.
[0368] Method 2800 may further include determining Figure 5 The comparison values are 534, 915, and 916 (Figure 9). Figure 11 The comparison value 1140, the comparison value corresponding to Chart 1502, and the comparison value corresponding to Chart 1504. Figure 15 The comparison value is 1506 or a combination thereof. For example, encoder 114 may determine the comparison value based on a comparison of multiple sets of samples 326 to 332 of the first audio signal 130 with samples of the second audio signal 132, as shown in [reference]. Figures 3 to 4 As described. Each set of multiple sets of samples may correspond to a specific mismatch value from a specific search range. For example, a specific search range may be greater than or equal to a smaller shift value 930 and less than or equal to a larger shift value 932, as described in Figure 9. As another example, a specific search range may be greater than or equal to a first shift value 1130 and less than or equal to a second shift value 1132, as described in Figure 9. The interpolation comparison value 838, the correction shift value 540, the final shift value 116, or a combination thereof may be based on the comparison value, as described in Figure 9. Figure 8 , 9A As described in 9B, 10A and 11.
[0369] Method 2800 may also include determining boundary comparison values, as shown in [reference]. Figure 17 As described. For example, encoder 114 can determine a comparison value at the right boundary (e.g., 20 sample shift / mismatch), a comparison value at the left boundary (-20 sample shift / mismatch), or both, as seen in [reference]. Figure 18 As described. The boundary comparison value may correspond to a mismatch value within a threshold (e.g., 10 samples) of the boundary mismatch value (e.g., -20 or 20) for a specific search range. In response to determining whether the boundary comparison value monotonically increases or decreases, encoder 114 may identify a second frame to be encoded indicating a monotonic trend, as shown in [reference]. Figure 17 As described.
[0370] Encoder 114 can determine that a specific number of frames (e.g., three frames) preceding the second frame to be encoded are identified as indicating a monotonic trend, as shown in [reference]. Figures 17 to 18 As described. In response to determining that a specific number is greater than a threshold, encoder 114 may determine a specific search range (e.g., -23 to 23) corresponding to the second frame to be encoded, as shown in [reference]. Figures 17 to 18 As described. A specific search range containing a second boundary mismatch (e.g., -23) exceeds a first boundary mismatch value (e.g., -20) corresponding to a first search range (e.g., -20 to 20) of the first frame to be encoded. Encoder 114 may generate comparison values based on the specific search range, as shown in [reference]. Figure 18 As described. The second mismatch value can be based on the comparison value.
[0371] Method 2800 may further include determining the decoding mode based at least in part on valid mismatch values. For example, encoder 114 may determine a first LB decoding mode 1913, a second LB decoding mode 1915, a first HB decoding mode 1912, a second HB decoding mode 1914, or a combination thereof, as shown in [reference]. Figure 19 As described. The encoded signal 102 may be based on the first LB decoding mode 1913, the second LB decoding mode 1915, the first HB decoding mode 1912, the second HB decoding mode 1914, or a combination thereof, as shown in [reference]. Figure 19 As described. According to a particular embodiment, encoder 114 may generate an encoded HB intermediate signal based on a first HB decoding mode 1912, an encoded HB side signal based on a second HB decoding mode 1914, an encoded LB intermediate signal based on a first LB decoding mode 1913, an encoded LB side signal based on a second LB decoding mode 1915, or a combination thereof, as shown in [reference]. Figure 19 As described.
[0372] According to some implementation schemes, the first HB decoding mode 1912 may include a BWE decoding mode, and the second HB decoding mode 1914 may include a blind BWE decoding mode, as shown in [reference]. Figure 21 As described. The encoded signal 102 may include an encoded HB intermediate signal and one or more parameters corresponding to the encoded HB side signal.
[0373] According to some implementation schemes, the first HB decoding mode 1912 may include the BWE decoding mode, and the second HB decoding mode 1914 may include the BWE decoding mode, as shown in [reference]. Figure 21 As described. The encoded signal 102 may include an encoded HB intermediate signal and one or more parameters corresponding to the encoded HB side signal.
[0374] According to some implementation schemes, the first LB decoding mode 1913 may include the ACELP decoding mode, the second LB decoding mode 1915 may include the ACELP decoding mode, the first HB decoding mode 1912 may include the BWE decoding mode, the second HB decoding mode 1914 may include the blind BWE decoding mode, or combinations thereof, as shown in [reference]. Figure 21 As described. The encoded signal 102 may include an encoded HB intermediate signal, an encoded LB intermediate signal, an encoded LB side signal, and one or more parameters corresponding to the encoded HB side signal.
[0375] According to some implementation schemes, the first LB decoding mode 1913 may include the ACELP decoding mode, and the second LB decoding mode 1915 may include the predictive ACELP decoding mode, or both, as shown in [reference]. Figure 21 As described. The encoded signal 102 may include an encoded LB intermediate signal and one or more parameters corresponding to the encoded LB side signal.
[0376] refer to Figure 29 A block diagram depicting a specific illustrative example of a device (e.g., a wireless communication device) is shown, and the device as a whole is designated as 2900. In various embodiments, with Figure 29 Compared to the components described herein, device 2900 may have fewer or more components. In the illustrative embodiment, device 2900 may correspond to... Figure 1 The first device 104 or the second device 106. In the illustrative embodiment, device 2900 can perform [see reference]. Figures 1 to 28 The system and methods described one or more operations.
[0377] In a particular embodiment, device 2900 includes a processor 2906 (e.g., a central processing unit (CPU)). Device 2900 may include one or more additional processors 2910 (e.g., one or more digital signal processors (DSPs)). Processor 2910 may include a media (e.g., speech and music) decoder / decoder (CODEC) 2908 and an echo canceller 2912. Media CODEC 2908 may include... Figure 1 The decoder 118, encoder 114, or both. Encoder 114 may include time equalizer 108, bit allocator 1908, and decoding mode selector 1910.
[0378] Device 2900 may include memory 153 and CODEC 2934. Although media CODEC 2908 is described as a component of processor 2910 (e.g., dedicated circuitry and / or executable code), in other embodiments, one or more components of media CODEC 2908 (e.g., decoder 118, encoder 114, or both) may be included in processor 2906, CODEC 2934, another processing component, or a combination thereof.
[0379] Device 2900 may include a transmitter 110 coupled to antenna 2942. Device 2900 may include a display 2928 coupled to display controller 2926. One or more speakers 2948 may be coupled to CODEC 2934. One or more microphones 2946 may be coupled to CODEC 2934 via input interface 112. In a particular embodiment, speaker 2948 may include Figure 1 First speaker 142, second speaker 144 Figure 2 The Y-speaker 244 or a combination thereof. In a particular embodiment, the microphone 2946 may include... Figure 1 First microphone 146, second microphone 148 Figure 2 The Nth microphone 248 Figure 11 The third microphone 1146, the fourth microphone 1148, or a combination thereof. The CODEC 2934 may include a digital-to-analog converter (DAC) 2902 and an analog-to-digital converter (ADC) 2904.
[0380] Memory 153 may contain instructions 2960 executable by processor 2906, processor 2910, CODEC 2934, another processing unit of device 2900, or a combination thereof, to perform [see reference]. Figures 1 to 28 One or more operations are described. Memory 153 can store analysis data 190.
[0381] One or more components of device 2900 may be implemented via dedicated hardware (e.g., circuitry), by executing instructions or a combination thereof via a processor for performing one or more tasks. As an example, one or more components of memory 153 or processors 2906, 2910, and / or CODEC 2934 may be memory devices, such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, or optical disc read-only memory (CD-ROM). The memory device may contain instructions (e.g., instruction 2960) that, when executed by a computer (e.g., the processors 2906 and / or 2910 in CODEC 2934), cause the computer to perform [see reference] Figures 1 to 28 The described one or more operations. As an example, memory 153 or one or more components of processor 2906, processor 2910, and / or CODEC 2934 may be non-transitory computer-readable media containing instructions (e.g., instruction 2960) that, when executed by a computer (e.g., the processor, processor 2906, and / or processor 2910 in CODEC 2934), cause the computer to perform [see reference]. Figures 1 to 28 The one or more operations described.
[0382] In a particular embodiment, device 2900 may be included in a system-in-package or system-on-a-chip device (e.g., a mobile station modem (MSM)) 2922. In a particular embodiment, processor 2906, processor 2910, display controller 2926, memory 153, CODEC 2934, and transmitter 110 are included in the system-in-package or system-on-a-chip device 2922. In a particular embodiment, input devices such as touchscreen and / or keypad 2930 and power supply 2944 are coupled to the system-on-a-chip device 2922. Furthermore, in a particular embodiment, such as... Figure 29 As described, the display 2928, input device 2930, speaker 2948, microphone 2946, antenna 2942, and power supply 2944 are external to the system-on-a-chip device 2922. However, each of the display 2928, input device 2930, speaker 2948, microphone 2946, antenna 2942, and power supply 2944 can be coupled to components (e.g., interfaces or controllers) of the system-on-a-chip device 2922.
[0383] Device 2900 may include a wireless telephone, mobile communication device, mobile phone, smartphone, cellular phone, laptop computer, desktop computer, computer, tablet computer, set-top box, personal digital assistant (PDA), display device, television, game console, music player, radio, video player, entertainment unit, communication device, fixed location data unit, personal media player, digital video player, digital video disc (DVD) player, tuner, camera, navigation device, decoder system, encoder system, base station, vehicle, or any combination thereof.
[0384] In certain embodiments, one or more components and devices 2900 of the system described herein may be integrated into a decoding system or device (e.g., an electronic device, a CODEC, or a processor therein), into an encoding system or device, or both. In other embodiments, one or more components and devices 2900 of the system described herein may be integrated into: a wireless communication device (e.g., a cordless phone), a tablet computer, a desktop computer, a laptop computer, a set-top box, a music player, a video player, an entertainment unit, a television, a game console, a navigation device, a communication device, a personal digital assistant (PDA), a fixed-location data unit, a personal media player, a base station, a vehicle, or another type of device.
[0385] It should be noted that the various functions performed by one or more components and devices 2900 of the system described herein are described as being performed by certain components or modules. This division of components and modules is for illustrative purposes only. In alternative embodiments, functions performed by a particular component or module may be divided among multiple components or modules. Furthermore, in alternative embodiments, two or more components or modules of the system described herein may be integrated into a single component or module. Each component or module described in the system described herein may be implemented using hardware (e.g., field-programmable gate array (FPGA) devices, application-specific integrated circuits (ASICs), DSPs, controllers, etc.), software (e.g., instructions executable by a processor), or any combination thereof.
[0386] In conjunction with the described embodiment, the device includes means for determining bit allocation based on a shift value and a second shift value. The shift value may indicate a shift of a first audio signal relative to a second audio signal, and the second shift value may be based on the shift value. For example, the means for determining bit allocation may include... Figure 19 The bit allocator 1908, one or more means / circuits (e.g., processor execution instructions stored in a computer-readable storage device) or a combination thereof configured to determine bit allocation.
[0387] The device may also include means for transmitting at least one encoded signal generated based on a bit allocation. The at least one encoded signal may be based on a first sample of a first audio signal and a second sample of a second audio signal, and the second sample may be time-shifted relative to the first sample by an amount based on a second shift value. For example, the means for transmission may include... Figure 1 and 19 The transmitter 110.
[0388] Also in conjunction with the described embodiments, the apparatus includes means for determining a first mismatch value indicating a first amount of time mismatch between a first audio signal and a second audio signal. The first mismatch value is associated with a first frame to be encoded. For example, the means for determining the first mismatch value may include... Figure 1 Encoder 114, Time equalizer 108, Figure 2 Time equalizer 208 Figure 5 The signal comparator 506, interpolator 510, shift optimizer 511, shift change analyzer 512, absolute shift generator 513, processor 2910, CODEC 2934, processor 2906, one or more devices / circuits configured to determine a first mismatch value (e.g., processor execution instructions stored in a computer-readable storage device), or a combination thereof.
[0389] The device also includes means for determining a second mismatch value, indicating a second amount of time mismatch between a first audio signal and a second audio signal. The second mismatch value is associated with a second frame to be encoded. The second frame to be encoded follows the first frame to be encoded. For example, the means for determining the second mismatch value may include... Figure 1 Encoder 114, Time equalizer 108, Figure 2 Time equalizer 208 Figure 5 The signal comparator 506, interpolator 510, shift optimizer 511, shift change analyzer 512, absolute shift generator 513, processor 2910, CODEC 2934, processor 2906, one or more means / circuits configured to determine a second mismatch value (e.g., processor execution instructions stored in a computer-readable storage device), or a combination thereof.
[0390] The device further includes means for determining an effective mismatch value based on a first mismatch value and a second mismatch value. The second frame to be encoded includes a first sample of a first audio signal and a second sample of a second audio signal. The second sample is selected at least in part based on the effective mismatch value. For example, the means for determining the effective mismatch value may include... Figure 1 Encoder 114, Time equalizer 108, Figure 2The time equalizer 208, signal comparator 506, interpolator 510, shift optimizer 511, shift change analyzer 512, processor 2910, CODEC 2934, processor 2906, one or more means / circuits configured to determine a valid mismatch value (e.g., processor execution instructions stored in a computer-readable storage device), or a combination thereof.
[0391] The device also includes means for transmitting at least one encoded signal having a bit allocation at least partially based on a valid mismatch value. The at least one encoded signal is generated based on at least partially a second frame to be encoded. For example, the means for transmission may include... Figure 1 and 19 The transmitter 110.
[0392] Those skilled in the art will further appreciate that the various illustrative logic blocks, configurations, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processing device such as a hardware processor, or a combination of both. The foregoing description generally focuses on the functionality of the various illustrative components, blocks, configurations, modules, circuits, and steps. Whether this functionality is implemented as hardware or as executable software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in different ways for each specific application, but these implementation decisions should not be construed as causing a departure from the scope of the invention.
[0393] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly embodied in hardware, in software modules executed by a processor, or a combination of both. The software modules can reside in memory devices such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, or optical disc read-only memory (CD-ROM). Exemplary memory devices are coupled to a processor so that the processor can read information from and write information to the memory device. In alternative examples, the memory device can be integrated with the processor. The processor and storage medium can reside in an application-specific integrated circuit (ASIC). The ASIC can reside in a computing device or user terminal. In alternative examples, the processor and storage medium can reside as discrete components in a computing device or user terminal.
[0394] A prior description of the disclosed embodiments is provided to enable those skilled in the art to make or use the disclosed embodiments. Various modifications to these embodiments will readily become apparent to those skilled in the art, and the principles defined herein can be applied to other embodiments without departing from the scope of the invention. Therefore, the invention is not intended to be limited to the embodiments shown herein, but should be accorded the widest scope that may be consistent with the principles and novel features as defined in the following claims.
Claims
1. A device for communication, comprising: a processor configured to: determine a first mismatch value indicative of a first amount of a time mismatch between a first audio signal and a second audio signal, the first mismatch value associated with a first frame to be encoded; determine a second mismatch value indicative of a second amount of a time mismatch between the first audio signal and the second audio signal, the second mismatch value associated with a second frame to be encoded, wherein the second frame to be encoded is subsequent to the first frame to be encoded; determine an effective mismatch value based on the first mismatch value and the second mismatch value, wherein the second frame to be encoded comprises first samples of the first audio signal and second samples of the second audio signal, and wherein the second samples are selected based at least in part on the effective mismatch value; select a first coding mode and a second coding mode based at least in part on the effective mismatch value; and generate at least one encoded signal having a bit allocation based at least in part on the second frame to be encoded, the bit allocation based at least in part on the effective mismatch value, wherein the at least one encoded signal is based on a first encoded signal and a second encoded signal, wherein the first encoded signal is based on the first coding mode, and wherein the second encoded signal is based on the second coding mode; and a transmitter configured to transmit the at least one encoded signal to a second device.
2. The device of claim 1, wherein the effective mismatch value is greater than or equal to a first value and less than or equal to a second value, wherein the first value is equal to one of the first mismatch value or the second mismatch value, wherein the second value is equal to the other of the first mismatch value or the second mismatch value.
3. The device of claim 1, wherein the processor is further configured to determine the effective mismatch value based on a change between the first mismatch value and the second mismatch value.
4. The device of claim 1, wherein the at least one encoded signal comprises the first encoded signal and the second encoded signal, wherein the first encoded signal comprises an encoded mid signal, wherein the second encoded signal comprises an encoded side signal, and wherein the bit allocation indicates a first number of bits allocated to the encoded mid signal and a second number of bits allocated to the encoded side signal.
5. The device of claim 1, wherein the processor is further configured to generate at least a first particular encoded signal having a first bit allocation based on the first frame to be encoded, and wherein the transmitter is further configured to transmit at least the first particular encoded signal.
6. The device of claim 1, wherein: the bit allocation is different than a first bit allocation associated with the first frame to be encoded based on a change between the first mismatch value and the second mismatch value. 7. The device of claim 1, wherein a particular number of bits are available for signal encoding, wherein a first bit allocation associated with the first frame to be encoded indicates a first ratio, and wherein the bit allocation indicates a second ratio.
8. The device of claim 1, wherein the at least one encoded signal comprises the first encoded signal, wherein the processor is further configured to generate the bit allocation to indicate that a particular number of bits are allocated to the first encoded signal, wherein the first encoded signal comprises an encoded mid signal, wherein a first bit allocation associated with the first frame to be encoded indicates that a first number of bits are allocated to a first encoded mid signal, and wherein the particular number is less than the first number.
9. The device of claim 1, wherein the at least one encoded signal comprises the second encoded signal, wherein the processor is further configured to generate the bit allocation to indicate that a particular number of bits are allocated to the second encoded signal, wherein the second encoded signal comprises an encoded side signal, wherein a first bit allocation associated with the first frame to be encoded indicates that a second number of bits are allocated to a first encoded side signal, and wherein the particular number is greater than the second number.
10. The device of claim 1, wherein the processor is further configured to: determine a change value based on the second mismatch value and the effective mismatch value; and in response to determining that the change value is greater than a first threshold, generate the bit allocation to indicate a first number of bits and a second number of bits, wherein the bit allocation indicates that the first number of bits are allocated to an encoded mid signal and the second number of bits are allocated to an encoded side signal, wherein the first encoded signal comprises the encoded mid signal and the second encoded signal comprises the encoded side signal, and wherein the at least one encoded signal comprises the first encoded signal and the second encoded signal.
11. The device of claim 10, wherein the processor is further configured to, in response to determining that the change value is less than or equal to the first threshold value and less than a second threshold value, generate the bit allocation to indicate a third number of bits and a fourth number of bits, wherein the bit allocation indicates that the third number of bits are allocated to the encoded mid signal and the fourth number of bits are allocated to the encoded side signal, wherein the third number of bits is greater than the first number of bits, wherein the fourth number of bits is less than the second number of bits, wherein the first encoded signal comprises the encoded mid signal, and wherein, the second encoded signal comprises the encoded side signal.
12. The device of claim 1, wherein the processor is further configured to determine a comparison value based on a comparison of a first sample of the first audio signal to a plurality of sets of samples of the second audio signal, wherein each set of the plurality of sets of samples corresponds to a particular mismatch value from a particular search range, and wherein the second mismatch value is based on the comparison value.
13. The device of claim 12, wherein the processor is further configured to: determine a boundary comparison value of the comparison values, the boundary comparison value corresponding to a mismatch value that is within a threshold of a boundary mismatch value of the particular search range; and in response to determining that the boundary comparison value monotonically increases, identify the second frame to be encoded as indicating a monotonic trend.
14. The device of claim 12, wherein the processor is further configured to: determine a boundary comparison value of the comparison values, the boundary comparison value corresponding to a mismatch value that is within a threshold of a boundary mismatch value of the particular search range; and identify the second frame to be encoded as indicating a monotonic trend in response to determining that the boundary comparison value monotonically decreases.
15. The device of claim 1, wherein the processor is further configured to: determine that a particular number of frames to be encoded preceding the second frame to be encoded are identified as indicating a monotonic trend; in response to determining that the particular number is greater than a threshold, determine a particular search range corresponding to the second frame to be encoded, the particular search range including a second boundary mismatch value that exceeds a first boundary mismatch value corresponding to the first frame to be encoded; and generate a comparison value based on the particular search range, wherein the second mismatch value is based on the comparison value.
16. The device of claim 1, wherein the processor is further configured to: generate an intermediate signal based on a sum of the first sample of the first audio signal and the second sample of the second audio signal; and generating an encoded mid signal by encoding the mid signal based on the bit allocation, wherein, the first encoded signal includes the encoded intermediate signal, and wherein the at least one encoded signal includes the first encoded signal.
17. The device of claim 1, wherein the processor is further configured to: generate a side signal based on a difference between the first sample of the first audio signal and the second sample of the second audio signal; and generate an encoded side signal by encoding the side signal based on the bit allocation, wherein the second encoded signal comprises the encoded side signal, and wherein, the at least one encoded signal includes the second encoded signal.
18. The apparatus of claim 1, wherein, the at least one encoded signal includes the first encoded signal and the second encoded signal, and wherein the processor is further configured to generate the at least one encoded signal by: based on the first coding mode, generating the first encoded signal based on a first sample of the first audio signal and a second sample of the second audio signal, wherein the second sample is selected based on the effective mismatch value; and based on the second coding mode, generating the second encoded signal based on the first sample and the second sample.
19. The device of claim 1, wherein the first encoded signal includes a low-band intermediate signal, wherein the second encoded signal includes a low-band side signal, and wherein the first coding mode and the second coding mode include algebraic code excited linear prediction (ACELP) coding modes.
20. The device of claim 1, wherein the first encoded signal includes a high-band intermediate signal, wherein the second encoded signal includes a high-band side signal, and wherein the first coding mode and the second coding mode include bandwidth extension (BWE) coding modes.
21. The device of claim 1, wherein the processor is further configured to: generate an encoded low-band intermediate signal based on an algebraic code excited linear prediction (ACELP) coding mode based at least in part on the effective mismatch value, wherein the first encoded signal includes the encoded low-band intermediate signal; and the first encoded signal includes a low-band intermediate signal, wherein the second encoded signal includes a low-band side signal, and wherein the first coding mode and the second coding mode include algebraic code excited linear prediction (ACELP) coding modes. generating an encoded low-band side signal based on a predictive ACELP coding mode based at least in part on the effective mismatch value, wherein the second encoded signal comprises the encoded low-band side signal, wherein the at least one encoded signal comprises the first encoded signal and one or more parameters corresponding to the second encoded signal.
22. The device of claim 1, wherein the processor is further configured to: generate an encoded high-band mid signal based on a bandwidth extension (BWE) coding mode based at least in part on the effective mismatch value, wherein the first encoded signal comprises the encoded high-band mid signal; and generate an encoded high-band side signal based on a blind BWE coding mode based at least in part on the effective mismatch value, wherein the second encoded signal comprises the encoded high-band side signal, wherein the at least one encoded signal comprises the first encoded signal and one or more parameters corresponding to the second encoded signal.
23. The device of claim 1, further comprising an antenna coupled to the transmitter, wherein the transmitter is configured to transmit the at least one encoded signal via the antenna.
24. The device of claim 1, wherein the processor and the transmitter are integrated into a mobile communication device.
25. The device of claim 1, wherein the processor and the transmitter are integrated into a base station.
26. A communication method, comprising: determining, at a device, a first mismatch value indicative of a first amount of a time mismatch between a first audio signal and a second audio signal, the first mismatch value associated with a first frame to be encoded; determining, at the device, a second mismatch value indicative of a second amount of a time mismatch between the first audio signal and the second audio signal, the second mismatch value associated with a second frame to be encoded, wherein the second frame to be encoded is subsequent to the first frame to be encoded; determining, at the device, an effective mismatch value based on the first mismatch value and the second mismatch value, wherein the second frame to be encoded comprises first samples of the first audio signal and second samples of the second audio signal, and wherein the second samples are selected based at least in part on the effective mismatch value; selecting a first coding mode and a second coding mode based at least in part on the effective mismatch value; generating at least one encoded signal having a bit allocation based at least in part on the effective mismatch value based at least in part on the second frame to be encoded, wherein the at least one encoded signal is based on a first encoded signal and a second encoded signal, wherein the first encoded signal is based on the first coding mode, and wherein the second encoded signal is based on the second coding mode; and sending the at least one encoded signal to a second device.
27. The method of claim 26, wherein the at least one encoded signal comprises the first encoded signal and the second encoded signal, and wherein, generating the at least one encoded signal comprises: generating the first encoded signal based on a first sample of the first audio signal and a second sample of the second audio signal based on the first coding mode, wherein the second sample is selected based on the effective mismatch value; and generating the second encoded signal based on the first sample and the second sample based on the second coding mode.
28. The method of claim 26, wherein the at least one encoded signal comprises the first encoded signal and the second encoded signal, wherein the first encoded signal comprises a low-band mid signal, wherein the second encoded signal comprises a low-band side signal, and wherein the first coding mode and the second coding mode comprise algebraic code excited linear prediction (ACELP) coding modes.
29. The method of claim 26, wherein the at least one encoded signal comprises the first encoded signal and the second encoded signal, wherein the first encoded signal comprises a high-band mid signal, wherein the second encoded signal comprises a high-band side signal, and wherein the first coding mode and the second coding mode comprise bandwidth extension (BWE) coding modes.
30. The method of claim 26, wherein the device comprises a mobile communication device.
31. The method of claim 26, wherein the device comprises a base station.
32. The method of claim 26, further comprising: generating an encoded high-band mid signal based on a bandwidth extension (BWE) coding mode based at least in part on the effective mismatch value, wherein the first encoded signal comprises the encoded high-band mid signal; and generating an encoded high-band side signal based on a blind BWE coding mode based at least in part on the effective mismatch value, wherein the second encoded signal comprises the encoded high-band side signal, wherein the at least one encoded signal comprises the first encoded signal and one or more parameters corresponding to the second encoded signal.
33. The method of claim 26, further comprising: generating an encoded low-band mid signal and an encoded low-band side signal based on algebraic code excited linear prediction (ACELP) coding modes based at least in part on the effective mismatch value, wherein the first encoded signal comprises the encoded low-band mid signal; generating an encoded high-band mid signal based on a bandwidth extension (BWE) coding mode based at least in part on the effective mismatch value, wherein the second encoded signal comprises the encoded high-band mid signal; and generating an encoded high-band side signal based on a blind BWE coding mode based at least in part on the effective mismatch value, wherein the at least one encoded signal comprises the encoded high-band mid signal, the encoded low-band mid signal, the encoded low-band side signal, and one or more parameters corresponding to the encoded high-band side signal.
34. The method of claim 26, wherein the bit allocation indicates a first number of bits allocated to the first encoded signal and a second number of bits allocated to the second encoded signal.
35. The method of claim 34, wherein the first number of bits is less than a first particular number of bits indicated by a first bit allocation indication associated with the first frame to be encoded, and wherein the second number of bits is greater than a second particular number of bits indicated by the first bit allocation indication.
36. A computer-readable storage device storing instructions that when executed by a processor cause the processor to perform operations comprising: determining a first mismatch value indicating a first amount of a time mismatch between a first audio signal and a second audio signal, the first mismatch value associated with a first frame to be encoded; determining a second mismatch value indicating a second amount of a time mismatch between the first audio signal and the second audio signal, the second mismatch value associated with a second frame to be encoded, wherein the second frame to be encoded is subsequent to the first frame to be encoded; determining an effective mismatch value based on the first mismatch value and the second mismatch value, wherein the second frame to be encoded includes a first sample of the first audio signal and a second sample of the second audio signal, and wherein the second sample is selected based at least in part on the effective mismatch value; selecting a first coding mode and a second coding mode based at least in part on the effective mismatch value; and generating at least one encoded signal having a bit allocation based at least in part on the second frame to be encoded, the bit allocation based at least in part on the effective mismatch value, wherein the at least one encoded signal is based on a first encoded signal and a second encoded signal, wherein the first encoded signal is based on the first coding mode, and wherein the second encoded signal is based on the second coding mode.
37. The computer-readable storage device of claim 36, wherein the at least one encoded signal includes the first encoded signal and the second encoded signal, wherein the bit allocation indicates a first number of bits allocated to the first encoded signal and a second number of bits allocated to the second encoded signal.
38. The computer-readable storage device of claim 36, wherein the first encoded signal corresponds to a mid signal and the second encoded signal corresponds to a side signal.
39. The computer-readable storage device of claim 38, wherein the operations further comprise: generating the mid signal based on a sum of the first audio signal and the second audio signal; and generating the side signal based on a difference between the first audio signal and the second audio signal.
40. An apparatus comprising: means for determining a first mismatch value indicating a first amount of a time mismatch between a first audio signal and a second audio signal, the first mismatch value associated with a first frame to be encoded; means for determining a second mismatch value indicating a second amount of a time mismatch between the first audio signal and the second audio signal, the second mismatch value associated with a second frame to be encoded, wherein the second frame to be encoded is subsequent to the first frame to be encoded; means for determining an effective mismatch value based on the first mismatch value and the second mismatch value, wherein the second frame to be encoded comprises first samples of the first audio signal and second samples of the second audio signal, and wherein the second samples are selected based at least in part on the effective mismatch value; means for selecting a first coding mode and a second coding mode based at least in part on the effective mismatch value; and means for transmitting at least one encoded signal having a bit allocation based at least in part on the effective mismatch value, the at least one encoded signal generated based at least in part on the second frame to be encoded, wherein the at least one encoded signal is based on a first encoded signal and a second encoded signal, wherein the first encoded signal is based on the first coding mode, and wherein the second encoded signal is based on the second coding mode.
41. The apparatus of claim 40, wherein the means for determining, the means for selecting, and the means for transmitting are integrated into at least one of a mobile telephone, a communication device, a computer, a music player, a video player, an entertainment unit, a navigation device, a personal digital assistant (PDA), a decoder, or a set-top box.
42. The apparatus of claim 40, wherein the means for determining, the means for selecting, and the means for transmitting are integrated into a mobile communication device.
43. The apparatus of claim 40, wherein the means for determining, the means for selecting, and the means for transmitting are integrated into a base station.
Citation Information
Patent Citations
Audio signal encoder, audio signal decoder, method for providing an encoded representation of an audio content, method for providing a decoded representation of an audio content and computer program for use in low delay applications
CN102859588A
Progressive refinement with temporal scalability support in video coding
CN104969555A