Adaptive gain control
Patent Information
- Application Number
- TW111108914
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-02-11
- Filing Date
- 2022-03-11
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2042-03-10
AI Technical Summary
Existing gain control techniques in audio encoding introduce delays and are prone to errors in error-prone environments, leading to signal distortions and quality loss, especially in cellular transmissions.
Adaptive gain control methods that determine gain parameters with zero delay and apply non-recursive techniques, using lookahead samples and selective encoding of gain parameters only where needed, to maintain signal quality and reduce bit usage.
The solution provides efficient and error-resistant gain control, minimizing signal distortions and improving audio quality by reducing bit usage in critical environments.
Smart Images

Figure TWG2TB001908186_001 
Figure TWG2TB001908186_002 
Figure TWG2TB001908186_003
Abstract
Description
[Technical Field]
[0001] This invention relates to systems, methods and media for adaptive gain control. [Previous Technology]
[0002] Gain control can be used, for example, to attenuate a signal to a range expected by a core codec. Many gain control techniques used to determine a gain to be applied require a delay and / or depend on the gain parameters applied to the previous frame. These gain control techniques can cause problems when used in error-prone situations (such as cellular transmission) and / or when real-time processing is required (such as dialogue). [Summary of the Invention]
[0003] At least some aspects of the present invention can be implemented by methods. Some methods may involve determining a downmixed signal associated with one or more downmixed channels of a current frame of an audio signal to be encoded. Some methods may involve determining whether an encoder of the downmixed signals to be used for encoding at least one of the one or more downmixed channels has an overload condition. Some methods may involve determining a gain parameter of the at least one of the one or more downmixed channels of the current frame of the audio signal in response to determining the existence of the overload condition. Some methods may involve determining at least one gain transition function based on the gain parameter and a gain parameter associated with a previous frame of the audio signal. Some methods may involve applying the at least one gain transition function to one or more of the downmixed signals. Some methods may involve encoding the downmixed signals in conjunction with information indicating gain control applied to the current frame.
[0004] In some instances, a portion of the frame buffer is used to determine the at least one gain transition function. In some instances, using this portion of the frame buffer to determine the at least one gain transition function introduces substantially zero additional delay.
[0005] In some instances, the at least one gain transition function includes a transition portion and a steady-state portion, wherein the transition portion corresponds to a transition from the gain parameter associated with the previous frame of the audio signal to the gain parameter associated with the current frame of the audio signal. In some instances, the transition portion has a fading transition type, wherein the gain increases on a portion of the sample of the current frame in response to an attenuation associated with the gain parameter of the previous frame being greater than an attenuation associated with the gain parameter of the current frame. In some instances, the transition portion has a reverse fading transition type, wherein the gain decreases on a portion of the sample of the current frame in response to an attenuation associated with the gain parameter of the previous frame being less than an attenuation associated with the gain parameter of the current frame. In some instances, the transition portion is determined using a prototype function and a scaling factor, wherein the scaling factor is determined based on the gain parameter associated with the current frame and the gain parameter associated with the previous frame. In some instances, the information indicating the gain control applied to the current frame includes information indicating the transition portion of the at least one gain transition function.
[0006] In some instances, the at least one gain transition function includes a single gain transition function applied to one of all one or more downmixing channels where the overload condition exists. In some instances, the at least one gain transition function includes a single gain transition function applied to one of all one or more downmixing channels, wherein the overload condition exists in a subset of one or more downmixing channels. In some instances, the at least one gain transition function includes a gain transition function for each of the one or more downmixing channels where the overload condition exists. In some instances, the number of bits used to encode the information indicating the gain control applied to the current frame scales substantially linearly with the number of downmixing channels where the overload condition exists.
[0007] In some instances, some methods may further involve: determining a second downmixed signal associated with one or more downmixed channels of a second frame of the audio signal to be encoded; determining, for at least one of the one or more downmixed channels of the second frame, whether the encoder has an overload condition; and, in response to determining that the second frame does not have the overload condition, encoding the second downmixed signals without applying a non-unity gain. In some instances, some methods may further involve setting a flag indicating that gain control is not applied to the second frame, wherein the flag includes one bit.
[0008] In some instances, some methods may further involve: determining a number of bits for encoding information indicating the gain control applied to the current frame; and allocating the number of bits from: 1) bits for encoding subsequent data associated with the current frame; and / or 2) bits for encoding the downmixing signals to encode the information indicating the gain control applied to the current frame. In some instances, the number of bits is allocated from the bits used to encode the downmixing signals, and wherein the bits used to encode the downmixing signals are reduced in order based on one of the spatial directions associated with the one or more downmixing channels.
[0009] Some methods may involve receiving, at a decoder, one encoded frame of an audio signal for one current frame of the audio signal. Some methods may involve decoding the encoded frame of the audio signal to obtain a downmixed signal associated with the current frame of the audio signal and information indicating gain control applied to the current frame of the audio signal by an encoder. Some methods may involve determining, at least in part, an inverse gain function to be applied to one or more downmixed signals associated with the current frame of the audio signal based on the information indicating the gain control applied to the current frame of the audio signal. Some methods may involve applying the inverse gain function to the one or more downmixed signals. Some methods may involve upmixing the downmixed signals to produce upmixed signals, including the one or more downmixed signals to which the inverse gain function is applied, wherein the upmixed signals are suitable for rendering.
[0010] In some instances, the information indicating the gain control applied to the current frame includes a gain parameter associated with the current frame of the audio signal. In some instances, the inverse gain function is determined at least in part based on the gain parameter of the current frame of the audio signal and a gain parameter associated with a previous frame of the audio signal.
[0011] In some instances, the inverse gain function includes a transitional part and a steady-state part.
[0012] In some instances, some methods may further involve: determining at the decoder that a second encoded frame has not yet been received; reconstructing an alternative frame by the decoder to replace the second encoded frame; and applying an inverse gain parameter of a previously encoded frame preceding the second encoded frame to the alternative frame. In some instances, some methods may further involve: receiving at the decoder a third encoded frame following the second encoded frame; decoding the third encoded frame to obtain a downmixed signal associated with the third encoded frame and information indicating gain control applied by the encoder to the third encoded frame; and determining the inverse gain parameters to be applied to the downmixed signal associated with the third encoded frame by smoothing the inverse gain parameters applied to the alternative frame using the inverse gain parameters associated with the gain control applied by the encoder to the third encoded frame. In some instances, methods may further involve: receiving a third encoded frame following the second encoded frame at the decoder; decoding the third encoded frame to obtain a downmixed signal associated with the third encoded frame and information indicating gain control applied by the encoder to the third encoded frame; and determining inverse gain parameters to be applied to the downmixed signals associated with the third encoded frame, such that the inverse gain parameters implement a smooth transition of gain parameters from one of the third encoded frames. In some instances, at least one intermediate frame exists between the unreceived second encoded frame and the received third encoded frame, wherein the at least one intermediate frame is not received at the decoder. In some instances, methods may further involve: receiving a third encoded frame following the second encoded frame at the decoder; decoding the third encoded frame to obtain a downmixed signal associated with the third encoded frame and information indicating gain control applied by the encoder to the third encoded frame; and determining, at least in part, an inverse gain parameter to be applied to the downmixed signal associated with the third encoded frame based on an inverse gain parameter applied to a frame received at the decoder prior to the second encoded frame not received at the decoder. In some instances, methods may further involve: receiving a third encoded frame following the second encoded frame at the decoder; decoding the third encoded frame to obtain a downmixed signal associated with the third encoded frame and information indicating gain control applied by the encoder to the third encoded frame; and rescaling one of the decoder's internal states based on the information indicating the gain control applied to the third encoded frame.
[0013] In some instances, some methods may further involve reproducing the amplified signals to generate reproduced audio data. In some instances, some methods may further involve replaying the reproduced audio data using one or more of a loudspeaker or headphones.
[0014] Some or all of the operations, functions, and / or methods described herein can be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices (such as those described herein), including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Therefore, some novel aspects of the subject matter described herein can be implemented via one or more non-transitory media having software stored thereon.
[0015] At least some aspects of the present invention can be implemented via an apparatus. For example, one or more apparatuses may be capable of performing at least partially the methods disclosed herein. In some embodiments, an apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof.
[0016] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages will become apparent from the description, drawings, and the claims of the invention. Note that the relative dimensions in the following figures may not be drawn to scale.
Implementation Method
[0029] Marking and Naming
[0030] Throughout this invention, and as included in the claims, the terms "loudspeaker," "amplifier," and "audio reproduction transducer" are used synonymously to refer to any sound-emitting transducer or transducer group. A typical set of headphones includes two loudspeakers. A loudspeaker may be implemented to include multiple transducers, such as a woofer and a tweeter, which may be driven by a single common loudspeaker or multiple loudspeaker feeds. In some instances, the loudspeaker feeds may undergo different processing in different circuit branches coupled to different transducers.
[0031] Throughout this invention, and within the scope of the invention claims, the expression “to perform an operation on” a signal or data (such as filtering, scaling, transforming, or applying gain to the signal or data) is used in a broad sense to indicate performing the operation directly on the signal or data or on a processed version of the signal or data. For example, the operation may be performed on a version of a signal that has undergone preliminary filtering or preprocessing before the operation is performed on it.
[0032] Throughout this invention, and within the scope of the invention claims, the term "system" is used in a broad sense to refer to an apparatus, system, or subsystem. For example, a subsystem implementing a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, wherein the subsystem generates M inputs and receives other XM inputs from an external source) may also be referred to as a decoder system.
[0033] Throughout this invention, and as included in the scope of the invention claims, the term "processor" is used broadly to refer to a system or apparatus that is programmable or otherwise configurable (e.g., using software or firmware) to perform operations on data (which may include audio or video or other image data). Examples of processors include a programmable gate array (or other configurable integrated circuit or chipset), a digital signal processor that is programmable and / or otherwise configured to perform pipeline processing on audio or other sound data, a programmable general-purpose processor or computer, and a programmable microprocessor chip or chipset.
[0034] Some coding techniques for scene-based audio, stereo audio, multi-channel audio, and / or object audio depend on coding multiple component signals after a downmixing operation. Downmixing allows coding of a reduced number of audio components using a waveform encoding method that preserves the waveform, and the remaining components can be parameterized. On the receiver side, the remaining components can be reconstructed using parameter post-encoding data that indicates the parameter encoding. Since only a subset of the components is waveform encoded, and the parameter post-encoding data associated with the parameter-encoded components can be encoded efficiently with respect to bit rate, this coding technique can be relatively bit rate efficient while still allowing for high-quality audio.
[0035] One potential problem is that the downmixed channel determined by a spatial encoder may contain signals with levels unsuitable for subsequent processing by a core codec that constructs an audio signal bitstream. For example, in some cases, a downmixed signal may have such a high level that it overloads the core codec, even though the original input signal is not overloaded in any of its component signals. This can lead to severe distortion, such as clipping in the reconstructed signal after decoding and rendering. This can result in significant quality loss in the final rendered signal. One potential solution is to attenuate the input signal to avoid overloading the core codec. However, this solution may have the disadvantage of increasing granular noise because the quantizer used to encode the signal may not operate within an optimal range.
[0036] Figure 1 shows a schematic block diagram of a conventional system for performing gain control on an encoded High-Order High-Fidelity Stereo (HOA) signal. The schematic diagram shown in Figure 1 can be used for encoding and decoding MPEG-H signals. MPEG-H is a set of international standards being developed by the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) Animation Experts Group (MPEG). MPEG-H has various parts, including part 3, MPEG-H 3D Audio. It should be noted that since MPEG-H audio is not a codec designed for error-prone communication applications (such as cellular communication), the MPEG-H audio codec does not need to meet strict write delay requirements and / or strict transmission error recovery requirements. Therefore, gain control for such applications can utilize recursive operations and introduce a delay, as will be discussed in more detail below.
[0037] At encoder 102, an input HOA signal is processed at 104. This processing may include decomposition, for example, where a downmixing channel is generated. The downmixing channel may contain a set of signals bounded by [-max,max] for a given frame. Since a core encoder 108 can encode signals in the range [-1,1), samples of signals associated with downmixing channels that exceed the range of core encoder 108 can cause overload. To avoid overload, a gain control 106 adjusts the gain of the frame so that the associated signals are within the range of core encoder 108 (e.g., within [-1,1)). Core encoder 108 can be considered as a codec that generates an encoded bitstream. Ancillary information generated by decomposition / processing block 104 (which may contain post-processing data associated with the parameter encoding channel or the like) can be combined with the signal generated as one output of core encoder 108 and encoded into a bitstream.
[0038] The encoded bitstream is received by a decoder 112. Decoder 112 can extract accompanying information, and a core decoder 116 can extract the downmixed signal. Next, an inverted gain control block 120 can invert the gain applied by the encoder. For example, the inverted gain control block 120 can amplify the signal attenuated by the gain control 106 of the encoder 102. Then, the HOA signal can be reconstructed by a HOA reconstruction block 122. Depending on the situation, the HOA signal can be rendered and / or replayed by a rendering / replay block 124. The rendering / replay block 124 may include, for example, various algorithms for rendering the reconstructed HOA output into rendered audio data. For example, rendering the reconstructed HOA output may involve distributing one or more HOA output signals across multiple speakers to achieve a specific perceptual impression. Depending on the situation, the rendering / replay block 124 may include one or more loudspeakers, headphones, etc., for presenting the rendered audio data.
[0039] Gain control 106 can be implemented using the following techniques. Gain control 106 can first determine an upper bound of the signal value in a frame. For example, for an MPEG-H audio signal, this bound can be expressed as a product, as specified in the MPEG-H standard. Given the upper bound, the minimum required attenuation ensures that the scaled signal sample is bounded by the interval [-1, 1). In other words, the scaled sample is within the range of the core encoder 108. This can be determined by applying a gain factor, where, by definition, emin can be a negative number. In some embodiments, the amplification can be limited to a maximum amplification factor, where emax is a non-negative integer. Therefore, to perform both attenuation and amplification, a gain factor 2e can be defined, where the gain parameter e is a value within the range [emin, emax]. Therefore, the minimum number of bits required to represent the gain parameter e is determined.
[0040] As described above, the gain factor gn(j) of a specific channel n and frame j can be determined by applying a single frame delay corresponding to a HOA block and using the following recursive operation:
[0041] In the above, gn(j-2) represents a gain factor applied to frame (j-2), and represents the gain factor adjustment required to calculate the gain factor gn(j-1) of frame j-1. To determine the gain factor adjustment, information from the current frame j is used, which introduces a delay of one frame. In other words, using this technique to determine the gain factor introduces a single frame delay and requires a recursive operation.
[0042] The knowledge requirement of gain gn(j-2) can be problematic in cases of potential transmission errors, where a discrepancy may exist between the encoder and decoder states, and therefore, the gain may not be accurately reconstructed by the decoder. Furthermore, in cases where encoded content is accessed at a random location, such as at the beginning of the file, previous frame information may not be accessible. Therefore, the drawbacks of conventional gain control utilizing recursion and a delay are unsuitable for implementation in codecs requiring low latency and in error-prone environments (such as those used for cellular transmission).
[0043] This document discloses techniques for providing adaptive gain control. Specifically, as described herein, a gain parameter with zero latency can be determined because it can be determined based on preview samples generated for use by a codec. It should be noted that the codec can be a codec used by a perceptual encoder. Furthermore, the determined gain parameter can be determined non-recursively, thereby allowing the use of adaptive gain control techniques in error-prone environments where frames may be dropped. The determination of the gain parameter and the application of the associated gain transition function are illustrated and described below with reference to Figures 2 through 6.
[0044] Additionally, in some implementations, adaptive gain control may be applied only in instances where one or more downmixing channels are associated with signals that would cause an overload condition in the codec by exceeding its expected range. As described herein, in instances where gain control is not applied, such as instances where no overload condition exists, gain parameters may not be encoded for the frame. By selectively encoding gain parameters in instances where gain control is applied, rather than for the entire frame, the gain control technique described herein produces a more bit-rate efficient encoding. This more efficient encoding of the gain parameters allows more bits to be used for encoding the downmixing channels, ultimately resulting in better audio quality. The technique for allocating bits among the bits used for encoding gain information, the bits used for encoding post-processing data, and the bits used for encoding the downmixing channels is shown and described below in conjunction with Figures 7 and 8.
[0045] Figure 2 shows a schematic block diagram of an example system 200 for performing low-latency adaptive gain control according to some embodiments. As illustrated, system 200 includes an encoder 202 and a decoder 212. At encoder 202, an input HOA signal (or first-order high-fidelity stereo reproduction (FOA)) is processed by a spatial coding block 204. For an N-channel input, spatial coding block 204 can generate a set of M downmixing channels. The number of downmixing channels in this set can be in the range of 1 to N. For example, for a FOA input, the downmixing channels may include: a main downmixing channel W', which can be generated by mixing the omnidirectional input signal W with directional input signals X, Y, and Z using various mixing gains; and up to 3 residual channels X', Y', and Z', each corresponding to signal components in the X, Y, and Z signals that cannot be predicted from the main downmixing signal. In one instance, spatial coding block 204 utilizes the Spatial Reconstruction (SPAR) technique. SPAR is further described in D. McGrath, S. Bruhn, H. Purnhagen, M. Eckert, J. Torres, S. Brown, and D. Darcy, Immersive Audio Coding for Virtual Reality Using a Metadata-assisted Extension of the 3GPP EVS Codec, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 730–734, the entire contents of which are incorporated herein by reference. In other instances, spatial coding block 204 may utilize any other energy-compressed transform suitable for linear predictive codecs, such as the Karhunen-Loeve transform (KLT) or similar. In some implementations, preview samples to be utilized by a core encoder 208 are used to generate downmixed channels. In some implementations, spatial coding block 204 may additionally generate ancillary information 210 that can be utilized by core encoder 208. Ancillary information 210 may include post-processing data for upmixing the downmixed channel by decoder 212. For example, ancillary information 210 may be used to reconstruct a representation of the original audio input 1 downmixed by spatial coding unit 204.
[0046] Next, an adaptive gain control 206 can be used to analyze the signals associated with the M downmixing channels. The adaptive gain control 206 can determine whether the signal associated with any of the M downmixing channels exceeds the range expected by the core encoder 208, and thus would overload the core encoder 208. In some embodiments, in an example where the adaptive gain control 206 determines that gain is not applied, such as in response to determining that the signals of the M downmixing channels do not exceed an expected range of the core encoder 208, the adaptive gain control 206 can set a flag indicating that gain control is not applied. This flag can be set by setting a single bit value. It should be noted that in some embodiments, in an example where the adaptive gain control 206 determines that gain is not applied, the adaptive gain control 206 may not set the flag, thereby reserving one bit (e.g., the bit associated with the flag). For example, in some implementations, if a spatial post-data bitstream and / or a core encoder bitstream (which may be a sense encoder bitstream) is self-terminating, the presence of a gain control flag can be determined by determining whether any unread bits exist in the bitstream. Unread bits may be the remaining bits in the bitstream. Then, M downmixing channels can be passed to the core encoder 208 to be encoded in a single bitstream in conjunction with the accompanying information 210.
[0047] Conversely, in the example where adaptive gain control 206 determines the gain to be applied, adaptive gain control 206 can determine gain parameters and apply (a number of) gains to M downmixing channels based on the determined gain parameters. Then, the M downmixing channels to which the applied gains are applied can be passed to core encoder 208 to be encoded in the bitstream in conjunction with accompanying information 210. Gain parameters can be included in the accompanying information 210, for example, as a set of bits indicating the gain parameters, as described in more detail below.
[0048] In some embodiments, adaptive gain control 206 may determine a gain to be applied by determining a gain parameter e(j) for a specific channel of a current frame j and M downmixed channels that exceeds the expected range of the core encoder 208 (e.g., would cause an overload condition). In some embodiments, the gain parameter e(j) is the smallest positive integer (inclusive of 0) that causes the signal associated with the channel to be within the expected range when scaled by a gain factor determined based on the gain parameter. As described above, the expected range may be [0, 1]. For example, the gain factor may be [0, 1]. It should be noted that in some embodiments, instead of identifying the gain parameter that causes the scaled channel to avoid an overload condition, the gain parameter may be selected such that when scaled by the gain factor, the signal is within a range smaller than the range associated with the overload condition. In other words, the gain parameter may be selected such that the scaled signal only avoids the overload condition, or is within a predetermined range smaller than the range associated with the overload condition, for example, to allow for a margin.
[0049] In some embodiments, adaptive gain control 206 may determine a gain transition function that transitions between a gain parameter e(j-1) associated with a previous frame (e.g., the (j-1)th frame) and a gain parameter e(j) of the current frame. In some embodiments, the gain transition function may allow the gain parameter to smoothly transition across samples of the j-th frame from the value of the gain parameter (e.g., e(j-1)) at the (j-1)th frame to the value of the gain parameter (e.g., e(j)) of the current frame. Therefore, the gain transition function may include two parts: 1) a transition part, wherein the gain parameter transitions across samples of the transition part from the gain parameter of the previous frame to the gain parameter of the current frame; and 2) a steady-state part, wherein the gain parameter has the value of the gain parameter of the current frame for samples of the steady-state part.
[0050] In some embodiments, in an example where the gain applied to the current frame is less than the gain applied to the previous frame, the transition portion may be referred to as having a transition type of "fading," because the attenuation increases across the samples of the current frame. The case where the gain applied to the current frame is less than the gain applied to the previous frame can be expressed as e(j) > e(j-1). In some embodiments, in an example where the gain applied to the current frame is greater than the gain applied to the previous frame, the transition portion may be referred to as having a transition type of "reverse fading" or "non-fading," because the attenuation decreases across the samples of the current frame. The case where the gain applied to the current frame is greater than the gain applied to the previous frame can be expressed as e(j) < e(j-1). In some embodiments, in an example where the gain applied to the current frame is the same as the gain applied to the current frame, the transition portion may be referred to as having a transition type of "holding," wherein the transition portion is not transitional but has the same value as the steady-state portion. The case where the gain applied to the current frame is the same as the gain applied to the current frame can be expressed as e(j) = e(j-1).
[0051] In some embodiments, a prototype shape of a transition portion of a gain transition function can be used to determine a transition portion of a gain transition function, wherein the prototype shape is scaled based on the difference between the gain parameter of the current frame and the gain parameter of the previous frame. For example, the prototype shape can be scaled based on e(j) - e(j-1). For example, a prototype function p may have the following properties: 1) p(0) = 1 (e.g., 0 dB); and 2) p(lend) = 0.5 (e.g., -6 dB), where lend represents the rightmost index for which p is defined. Continuing this example, a gain transition function using this prototype function p can be expressed as:
[0052] Figure 3A shows examples of gain transition functions, each having a transition portion of a transition type characterized by "fading". In the examples shown in Figure 3A, each gain transition function has a transition portion starting at sample 0, which corresponds to the beginning of the current frame and has a gain of 0 dB, where 0 dB is the gain parameter of the previous frame (e.g., the (j-1)th frame). In the examples shown in Figure 3A, the transition portion of each gain transition function changes to the steady-state portion of the gain transition function over approximately 384 samples. For each of the three gain transition functions shown in Figure 3A, the steady-state portion corresponds to a different gain parameter of the jth frame, where the gain increases by 6 dB, 12 dB, and 18 dB relative to the gain of the previous frame, respectively. In other words, as shown in Figure 3A, for the three gain transition functions, exp = -[e(j) - e(j-1)] = -1, -2, and -3, respectively. It should be noted that for each of the gain transition functions shown in Figure 3A, the transition portion has the same length (e.g., approximately 384 samples). It should also be noted that the length of the steady-state portion may correspond to an offset related to the delay introduced by the codec, for example, 12 milliseconds in the example shown in Figure 3A. Accordingly, the length of the transition portion may be related to the reciprocal of the offset. In the example shown in Figure 3A, the length of the transition portion is the frame length (e.g., 20 milliseconds) minus the codec delay (e.g., 12 milliseconds). It should be noted that the codec delay may be the total writer algorithm delay excluding the frame size delay.
[0053] Additionally, it should be noted that a gain transition function having one of the transition types of "reverse fading" or "non-fading" can be represented as a mirror image of the gain transition function shown in Figure 3A, flipped across a horizontal line. By way of example, the horizontal line can be the x-axis.
[0054] Referring back to Figure 2, decoder 212 can receive an encoded bitstream as an input and can reconstruct the HOA signal, for example, for rendering. In some embodiments, a core decoder 216 receives M downmixed channels to which encoder 202 applies gain and provides the M downmixed channels to an inverse gain control 220. Inverse gain control 220 obtains the gain parameters applied by encoder 202 from accompanying information 210. For example, in some embodiments, inverse gain control 220 may retrieve the gain parameter e(j) applied by encoder 202 from accompanying information 210. Additionally, inverse gain control block 220 may, for example, retrieve the gain parameter applied by encoder to a previous frame, for example, e(j-1), from memory. Then, inverse gain control block 220 can use the obtained gain parameter to inverse the gain applied by encoder 202. For example, in some embodiments, inverse gain control 220 may construct an inverse gain transition function that transitions from the gain parameter of the previous frame to the gain parameter of the current frame. In some implementations, the inverse gain transition function may be a gain transition function mirrored and vertically adjusted across a central vertical line applied by encoder 202. By way of example, the vertical line may be the y-axis.
[0055] Turning to Figure 3B, an example of an inverse gain transition function applied by a decoder in response to a gain transition function shown in Figure 3A by an encoder is shown according to some embodiments. As illustrated, the inverse gain transition function has a steady-state portion and a transition portion. The duration of the steady-state portion and the transition portion of the inverse gain transition function may correspond to (e.g., be the same as) the duration of the corresponding steady-state portion and the transition portion of the gain transition function, as illustrated in Figures 3A and 3B. As illustrated, each inverse gain transition function shown in Figure 3B starts at 0 dB and transitions to the inverse gain to be applied to the current j-th frame. That is, each inverse gain transition function starts at 0 dB, which corresponds to the inverse gain applied to the previous frame j-1. It should be noted that when the gain applied by the encoder corresponds to an attenuation indicated by a gain less than 0 dB (as shown in the gain transition function of Figure 3A), the inverse gain applied by the decoder corresponds to an amplification with a gain greater than 0 dB (as shown in the gain transition function of Figure 3B). Conversely, in an example where the gain applied by the encoder corresponds to an amplification, for example, with a gain greater than 0 dB, the inverse gain applied by the decoder corresponds to an attenuation, for example, with a gain less than 0 dB.
[0056] Referring back to Figure 2, after applying inverse gain, the M downmixing channels with applied inverse gain are provided to a spatial decoding block 222. The spatial decoding block 222 can reconstruct the HOA signal using the accompanying information 210. For example, in the example where the spatial encoding block 204 uses SPAR technology for spatial encoding, the spatial decoding block 222 can use SPAR technology to reconstruct one or more channels using the data encoded subsequently included in the accompanying information 210. Then, the reconstructed HOA output can be rendered by a rendering / replay block 224. The rendering / replay block 224 may include, for example, various algorithms for rendering the reconstructed HOA output into rendered audio data. For example, rendering the reconstructed HOA output may involve distributing one or more signals of the HOA output across multiple speakers to achieve a specific perceptual impression. Depending on the situation, the rendering / replay block 224 may include one or more amplifiers, headphones, etc., for presenting the rendered audio data.
[0057] In some embodiments, a decoder can utilize various techniques to recover dropped or lost frames, which may occur, for example, during cellular transmission or in connection with other error-prone environments. In an example where a frame has not been dropped and the decoder can access gain parameters used in conjunction with previous frames, the decoder can determine the inverse gain transition function based on the gain parameters associated with the previous frame. However, in the case where a frame has been dropped, when processing the first recovered frame after the dropped frame (generally referred to herein as a "recovery frame"), the decoder cannot access the gain parameters of the frame preceding the recovery frame because the previous frame and associated gain parameters have been lost. Therefore, in some embodiments, the decoder can reconstruct a replacement frame for the dropped frame using any suitable frame loss concealment technique. The decoder can then use the gain parameters of the previously received frame for the replacement frame.
[0058] Figure 4 illustrates an example of encoder gain and corresponding decoder gain for a series of frames according to some embodiments. As illustrated, a dropped frame 402 (depicted as an "X" in Figure 4) is preceded by a received frame 401 and followed by a recovered frame 403. The encoder applies an encoder gain GE, as shown in curve 404. Specifically, GE is 0 dB for the received frame 401 and -18 dB for both the dropped frame 402 and the recovered frame 403. As illustrated by core decoder output level curve 406, the dropped frame 402 is reconstructed using a frame loss concealment technique to produce a replacement frame. The replacement frame may have a programmer decoder output level corresponding to the decoder gain of the previous frame (e.g., the gain of the received frame 401 or 0 dB), as shown at 408. Accordingly, as illustrated by decoder gain curve 410, the alternative frame has a decoder gain G* that is equivalent to the previous frame (e.g., received frame 401), as shown at 412.
[0059] A similar procedure can occur with a dropped frame 414. In this case, the encoder gain GE of the dropped frame 414 is 0 dB, while the encoder gain of the previously received frame 413 is -18 dB. In other words, the dropped frame 414 occurs during a gain transition from -18 dB to 0 dB. Therefore, when using the frame loss concealment technique, the core decoder output level reconstructs a gain of -18 dB for a replacement frame. The reconstructed gain of the replacement frame corresponds to the -18 dB encoder gain of the previously received frame 413, as shown at 416. Accordingly, the decoder gain of the replacement frame can be set to the decoder gain of the previously received frame 413 or 18 dB, as shown at 418. It should be noted that setting the decoder gain for a replacement frame corresponding to the discarded frame 420 will not cause a discontinuity in decoder gain, as there is no gain change between the previous frame 419 and the discarded frame 420, where the encoder gain is the same for both the discarded frame 420 and the previous frame 419.
[0060] Additionally, it should be noted that, as shown in the relative output gain curve 422, using a technique to set the decoder gain of one of the alternative frames to be equal to the decoder gain of the previously received frame results in a total relative output gain of 0 dB, thereby indicating that there is no fluctuation between frames. This is desirable in reducing perceived discontinuities attributable to changes in output gain across frames.
[0061] In some embodiments, a decoder may perform a smoothing technique to transition from the gain parameter of a previously received frame to the gain parameter of a recovered frame, for example, smoothing across an alternative frame with no received gain parameter.
[0062] In some embodiments, the smoothing technique may involve a decoder fusing the alternative frame and the recovery frame in such a way that it increases the weight given to the alternative frame during an initial portion of one of the fused samples and increases the weight given to the recovery frame during a subsequent portion of one of the fused samples.
[0063] As another example, in some implementations, the smoothing technique may involve adjusting the decoder state memory to compensate for the gain of the lost frame before decoding the recovery frame. As a more specific example, in an instance where the gain of the recovery frame is determined to be too high, the decoder state memory may be adjusted downwards so that the recovery frame is decoded using a suitably reduced decoder state memory. In other words, the decoder state memory may be scaled downwards in response to a determination that the reconstructed decoder gain G* of the previous frame is less than the decoder gain G of the recovery frame. Conversely, in an instance where the gain of the recovery frame is determined to be too low, the decoder state memory may be adjusted upwards so that the recovery frame is decoded using a suitably increased decoder state memory. In other words, the decoder state memory may be scaled upwards in response to a determination that the reconstructed decoder gain G* of the previous frame is greater than the decoder gain G of the recovery frame. Therefore, the decoder gain G of the recovery frame may be adjusted based on the reconstructed decoder gain G*. It should be noted that since the reconstructed decoder gain G* can be determined based on the gain of the frame before the frame was discarded (e.g., frame 401 in Figure 4), the decoder gain G of the recovered frame can be adjusted at least in part based on the decoder gain of the frame before the frame was discarded.
[0064] As yet another example, in some embodiments, the smoothing technique may involve applying a smoothing function between a previously received frame and a recovered frame. This smoothing function may correspond to a smoothing function implemented and utilized by the decoder, thereby allowing smoothing to be performed without additional additions. Alternatively, in some embodiments, the smoothing function may be a dedicated smoothing function used in the case of dropped frames. In these embodiments, the smoothing function may depend on the duration of packet loss, which can be indicated by seconds, blocks, or the number of frames, which may be advantageous in the case of dropping multiple sequential frames.
[0065] Figure 5 illustrates an example of a procedure 500 for determining a gain parameter and applying a gain to a demixed signal according to the determined gain parameter, according to some embodiments. In some embodiments, blocks of procedure 500 may be executed by an encoder device. In some embodiments, blocks of procedure 500 may be executed in a sequence other than that shown in Figure 5. In some embodiments, two or more blocks of procedure 500 may be executed substantially in parallel. In some embodiments, one or more blocks of procedure 500 may be omitted.
[0066] At 502, program 500 may determine a downmixed signal associated with a frame of an audio signal to be encoded. For example, in some embodiments, program 500 may use any suitable spatial coding technique to determine a set of downmixed channels. Examples of spatial coding techniques include SPAR, a linear prediction technique, or the like. The set of downmixed channels may include any one to N channels, where N is the number of input channels, for example, in the case of a FOA signal, N is 4. The downmixed signal may include audio signals of downmixed channels corresponding to a specific frame of an audio signal. It should be noted that in some embodiments, program 500 may determine a "transmit signal" rather than a downmixed signal. Such transmit signals may refer to the signal to be encoded, which is not necessarily downmixed.
[0067] At 504, the procedure 500 may determine whether an overload condition exists in a codec (such as an Enhanced Voice Service (EVS) codec and / or any other suitable codec). For example, the procedure 500 may determine the existence of an overload condition in response to determining that the signal of at least one downmixing channel exceeds a predetermined range (e.g., [-1, 1) and / or any other suitable range).
[0068] If it is determined at 504 that no overload condition exists ("No" at 504), then procedure 500 may continue to 512 and may encode the downmixed signal. For example, in some embodiments, procedure 500 may generate a bitstream that combines with accompanying information (such as post-processing data) that can be used by a decoder to upmix the downmixed signal (e.g., reconstruct a FOA or HOA output) to encode the downmixed signal.
[0069] Conversely, if an overload condition is determined to exist at 504 ("Yes" at 504), then procedure 500 may continue to 506, and a gain parameter of the frame that causes the overload condition to be avoided may be determined. For example, in some embodiments, procedure 500 may determine a gain parameter by determining a minimum positive integer such that when the downmixed signal of the downmixing channel is scaled by a gain factor determined based on the gain parameter, the downmixed signal is within a predetermined range, for example, within [-1, 1). For example, as described above in conjunction with FIG2, the gain parameter may be represented as a positive integer (inclusive of 0) e(j) of the current frame (j), wherein applying a gain factor 2-e(j) to the downmixed signal causes the downmixed signal to be within the predetermined range.
[0070] At 508, program 500 may determine a gain transition function based on the gain parameter of the current frame (e.g., frame j) determined at block 506 and a gain parameter of a previous frame (e.g., frame j-1). For example, as described above in conjunction with Figure 2, the gain transition function may have a transition portion and a steady-state portion, wherein the steady-state portion corresponds to the gain factor of the current frame, and the transition portion corresponds to a sequence of intermediate gain factors of a subset of samples of the current frame, which transition from the gain factor at the end of the previous frame to the gain factor of the steady-state portion of the current frame.
[0071] In an example where the gain parameter of the previous frame corresponds to an attenuation less than the gain parameter of the current frame, the transition portion may be referred to as having a transition type of "fading". Conversely, in an example where the gain parameter of the previous frame corresponds to an attenuation greater than the gain parameter of the current frame, the transition portion may be referred to as having a transition type of "reverse fading" or "non-fading". In an example where the gain parameter of the previous frame is the same as the gain parameter of the current frame, the transition portion may be referred to as having a transition type of "holding". In an example where the transition portion has a transition type of "holding", the value of the gain transition function during the transition portion may be the same as the value of the gain transition function during the steady-state portion. In some embodiments, a transition portion of the gain transition function may be determined by scaling a prototype function based on the gain parameters of the previous and / or current frames. As described above in conjunction with Figure 2, the duration of the transition portion of the gain transition function may correspond to a delay duration utilized by the codec.
[0072] At 510, procedure 500 may apply a gain transition function to the downmixed signal associated with the frame. For example, in some embodiments, procedure 500 may scale samples of the downmixed signal by a gain factor indicated by the gain transition function. As a more specific example, in some embodiments, a first sample of the current frame may be scaled by a gain factor corresponding to a gain parameter of a previous frame, a last sample of the current frame may be scaled by a gain factor corresponding to a gain parameter of the current frame, and intermediate samples may be scaled by a gain factor corresponding to a gain parameter of the transition or steady-state portion of the gain transition function. It should be noted that in examples where procedure 500 is applied to a transmitted signal, for example, as described above in conjunction with block 502, procedure 500 may apply a gain transition function to the transmitted signal.
[0073] It should be noted that in some implementations, the gain transition function may be applied only to the downmixed signal of the downmixed channel where an overload condition is detected at block 504. For example, in one example where an overload condition is detected for both the Y' and X' channels, a separate gain transition function may be determined for each of the Y' and X' channels, and applied to the signals of the Y' and X' channels. Continuing this example, the gain transition function may not be applied to the W' and Z' channels. In these examples, for example, at block 512, an indication of the channel to which the gain transition function is applied and the corresponding gain parameter for each channel may be encoded. Alternatively, in some implementations where an overload condition exists for only one downmixed channel, the corresponding gain transition function may be applied to all downmixed channels. In these examples, since the gain transition function is applied to all channels, there is no need to transmit an indication of the channel to which the gain has been applied, which results in increased bit rate efficiency.
[0074] At 512, program 500 may encode a downmixed signal, and if gain is applied, encode information indicating (some) gain parameters of the frame(s). In an example where gain is applied, the encoded downmixed signal may be the downmixed signal after the application of a gain transition function at block 510. The downmixed signal and any information indicating gain parameters may be encoded by a codec (such as an EVS codec or similar) in conjunction with any accompanying information (such as post-processing data) that may be used by a decoder to reconstruct or upmix the downmixed signal. It should be noted that in an example where program 500 utilizes a transmitted signal, for example, as described above in conjunction with block 502, program 500 may encode a transmitted signal.
[0075] It should be noted that in some embodiments, program 500 may encode the gain parameter in a set of bits. In some embodiments, an additional bit may be used as an exception flag, for example, to indicate a transition function. In some embodiments, the gain transition function may indicate a prototype function associated with the transition portion of the gain transition function. In some embodiments, the gain transition function may indicate a hard transition, for example, a step function, which occurs in instances where a sudden and relatively large level change occurs between frames and therefore a smooth transition cannot be implemented by gain control. By setting this exception using an exception flag, a decoder can implement a hard transition. A gain parameter may be encoded using x bits, where x depends on one of the number of quantized values of the gain parameter of a current frame, for example, one of the number of quantized values of e(j). For example, x may be determined by ceil(log2(number of quantized values of the gain parameter). In one instance, where e(j) can take values of 0, 1, 2, and 3, x is 2 bits.
[0076] In the example where adaptive gain control is enabled for each channel so that a unique gain transition function is applied to each downmixed channel associated with a signal triggering an overload condition, x bits can be used for each channel with gain control enabled, where an additional bit indicator for each channel indicates that the gain parameter has been encoded. In this example, the total number of bits used to transmit gain control information is Ndmx + (x + 1) * N, where Ndmx represents the number of downmixed channels (and where for each of the Ndmx channels, a single bit is used to indicate whether gain control is enabled), and where N represents the number of channels with gain control enabled. Note that in the example where gain control is not enabled for a particular frame, Ndmx bits can be used to indicate that gain control is not enabled; for example, each of the Ndmx channels uses 1 bit. Note that in the example where the number of downmixed channels is 1, for example, only W channels are waveform encoded, the total number of bits used to transmit gain control information is (x + 1) * N. For example, given a downmixing channel, if gain control is not enabled for that downmixing channel (e.g., N=0), the number of bits used is 0. Continuing this example, if gain control is enabled (e.g., N=1), the number of bits used is x+1. Note that in the term "x+1", 1 represents a 1-bit exception flag (e.g., it can be used to indicate that a hard transition (such as a step function) will be implemented to transition between consecutive frames, as described in more detail below).
[0077] In the example where a single gain transition function associated with one downmixing channel that triggers an overload condition is applied to all downmixing channels, fewer bits can be used to transmit gain control information. For example, x bits are used in conjunction with an exception flag indicating, for example, the transition function to transmit a single gain parameter of the current frame. As a more specific example, in these embodiments, the total number of bits used to transmit gain control information for a frame is represented by x+1.
[0078] In some embodiments, program 500 may allocate bits for transmitting gain control information of the frame from bits normally allocated for transmitting incidental information (such as data for reconstructing the HOA signal) and / or from bits normally allocated for encoding the downmixing channel. Example techniques for allocating gain control bits are shown and described below in conjunction with Figures 7 and 8.
[0079] Figure 6 illustrates an example of a procedure 600 according to some embodiments for obtaining gain parameters utilized by an encoder and applying an inverse gain transition function based on the obtained gain parameters. In some embodiments, blocks of procedure 600 may be executed by a decoder device. In some embodiments, blocks of procedure 600 may be executed in an order other than that shown in Figure 6. In some embodiments, two or more blocks of procedure 600 may be executed substantially in parallel. In some embodiments, one or more blocks of procedure 600 may be omitted.
[0080] Procedure 600 may begin at 602 by receiving one encoded frame of an audio signal. The received frame (e.g., the current frame) is generally referred to herein as the j-th frame. The received frame may immediately follow a previously received frame, or it may be a frame that does not immediately follow a previously received frame.
[0081] At 604, program 600 can decode the encoded frames of the audio signal to obtain a downmixed signal and (if gain control is applied by the encoder) information indicating at least one gain parameter associated with the frame. In some embodiments, program 600 can determine whether gain control is applied by the encoder based on an anomaly flag (e.g., a one-bit anomaly flag) indicating whether a hard transition (e.g., a step function transition) is to be performed. In other words, in examples where no anomaly flag is set, the decoder can determine that a smooth transition will be performed between consecutive frames. In examples where the encoder applies gain control on a per-channel basis, program 600 can further identify which downmixed channels the gain control is applied to.
[0082] At 606, program 600 may determine an inverse gain transition function based on the gain parameter of the current frame (generally referred to herein as e(j)) and a gain parameter of a previous frame (e.g., generally referred to herein as e(j-1)). In some embodiments, program 600 may retrieve the gain parameter of the previous frame from memory (e.g., from decoder state memory). In examples where gain control is not applied to the previous frame, program 600 may set e(j-1) to 0.
[0083] In some embodiments, procedure 600 may determine the inverse gain transition function as the inverse of the gain transition function applied at the encoder. For example, the inverse gain transition function may correspond to a gain transition function that is mirrored and adjusted across a horizontal line. The mirroring and adjustment may be along the x-axis. An example of such an inverse gain transition function is shown and described above in conjunction with Figure 3B. In some embodiments, the inverse gain transition function may have a steady-state portion corresponding to the gain applied to a previous frame (where the gain is determined based on the gain parameter of the previous frame, or where gain control is set to 0 in an example where it is not applied to the previous frame). The inverse gain transition function may then have a transition portion that is the inverse of the transition portion of the gain transition function applied at the encoder. For example, where the gain applied to the current frame corresponds to a greater attenuation relative to the previous frame, the inverse gain transition function may have a transition portion from a smaller amplification to a larger amplification. Conversely, where the gain applied to the current frame corresponds to a lesser attenuation relative to the previous frame, the inverse gain transition function may have a transition portion from a larger amplification to a smaller amplification. The duration of one of the transition portions can be related to the delay introduced by the codec, where the duration of the transition portion is the frame length (e.g., 20 ms) minus the codec delay (e.g., 12 ms). It should be noted that in examples where the delay introduced by the codec is longer than the frame length, an inverse gain transition can be applied to one frame's delay. In some examples, the delay can be obtained from the gain control bits via procedure 600 (e.g., via the decoder). It should be noted that the inverse gain transition function can also be used to attenuate signals amplified by the encoder's gain control.
[0084] At 608, program 600 may apply an inverse gain transition function to the downmixed signal to invert the gain applied by the encoder. For example, the application of the inverse gain transition function causes the downmixed signal attenuated by the encoder to be amplified to invert the attenuation. As another example, the application of the inverse gain transition function causes the downmixed signal amplified by the encoder to be attenuated to invert the amplification.
[0085] At 610, program 600 can upmix the downmixed signal. Upmixing can be performed by a spatial encoder. In some instances, the spatial encoder may utilize SPAR technology. The upmixed signal may correspond to a reconstructed FOA or HOA audio signal. In some implementations, program 600 may upmix the signal using ancillary information (e.g., post-data) encoded in the bitstream, wherein the ancillary information can be used to reconstruct the parametrically encoded signal.
[0086] In some embodiments, at 612, program 600 may render an upmixed signal to generate rendered audio data. In some embodiments, program 600 may utilize any suitable rendering algorithm to render a FOA or HOA audio signal, for example, rendering scene-based audio data. In some embodiments, the rendered audio data may be stored in any suitable format, for example, for future rendering or replay. Note that in some embodiments, block 612 may be omitted.
[0087] In some embodiments, at 614, procedure 600 causes the played audio data to be replayed. For example, in some embodiments, the played audio data may be presented via one or more loudspeakers and / or headphones. In some embodiments, multiple loudspeakers may be used, and the multiple loudspeakers may be positioned relative to each other in any suitable location or orientation in three dimensions. It should be noted that in some embodiments, procedure 614 may be omitted.
[0088] As described above in conjunction with Figure 5, a set of gain control bits can be used to encode gain control information (e.g., information indicating gain parameters). In some implementations, different gain parameters and gain transition functions can be determined for each downmixing channel that detects an overload condition. In these implementations, gain control bits are needed to indicate whether gain control is applied to each downmixing channel, and gain parameters are encoded for each downmixing channel to which gain control is applied, as described above in conjunction with Figure 5. Alternatively, in some implementations, a single gain transition function is determined based on the presence of an overload condition in one downmixing channel and can be applied to all downmixing channels. In these implementations, fewer gain control bits are needed because a separate bit flag is not required to indicate whether gain control has been applied to each downmixing channel, resulting in more efficient bit-rate encoding.
[0089] By applying the same gain transition function to all downmixing channels (including those without overload conditions), higher bit-rate efficient coding can lead to a degradation in perceived quality by, for example, attenuating signals without encoder-decoder overload. In contrast, using more targeted gain control (where gain control is applied in a targeted manner to each downmixing channel) requires more bits to transmit gain control information. However, using additional bits to transmit targeted (e.g., channel-specific) gain control information may require reallocating bits typically used for waveform encoding of downmixing channels, which can degrade perceived quality in some cases. Therefore, there may be a conditional trade-off between applying the same gain transition function to all downmixing channels and applying channel-specific gain control. Regardless of whether gain control is applied to all downmixed channels or based on a per-channel approach, the bits associated with gain control information can be allocated from the bits typically used for waveform encoding of the downmixed channels and / or from the bits typically used for encoding ancillary information (such as post-processing data) for reconstructing an FOA or HOA signal from the downmixed channels, thereby reducing the number of available bits for encoding the downmixed channels or ancillary information.
[0090] The following describes a more detailed technique for the bit distribution used to encode gain control information. For background, Figure 7A illustrates a FOA codec for encoding and decoding audio signals using the SPAR technique, which employs the adaptive gain control techniques described above in conjunction with Figures 2 to 6. It should be noted that although Figure 7A describes spatial coding using the SPAR technique, the techniques described in conjunction with Figures 7A and 8 can be used in conjunction with any suitable spatial coding technique. Figure 8 shows a flowchart of an example procedure 800 for allocating bits for encoding gain control information according to some embodiments.
[0091] Figure 7A is a block diagram of an FOA codec 700 for encoding and decoding FOA in SPAR format according to some embodiments. The FOA codec 700 includes an SPAR encoder 701, a core encoder 705, an adaptive gain control (AGC) encoder 713, an SPAR decoder 706, a core decoder 707, and an AGC decoder 714. In some embodiments, the SPAR encoder 701 converts an FOA input signal into a set of downmixed channels and parameters for reproducing the input signal at the SPAR decoder 706. The downmixed signal can vary between 1 and 4 channels, and the parameters can include prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation coefficients (P). A more detailed description of techniques for reconstructing an audio signal from a downmixed version of an audio signal using SPAR with PR, C, and P parameters is provided below.
[0092] It should be noted that the example implementation shown in Figure 7A illustrates a nominal 2-channel downmixing, where a W (passive prediction) or W' (active prediction) channel is sent to the SPAR decoder 706 along with a single prediction channel Y'. In some implementations, W' may be an active channel. An active W' downmixing channel can be constructed by mixing the X, Y, and Z channels into the W channel based on the mixing gain. In one example, one of the active predictions of the W channel can be determined using the following:
[0093] In the above, f represents a function of the normalized input covariance that allows some mixing of the X, Y, and Z channels into the W channel, and , , and represent prediction coefficients. In some implementations, f can also be a constant, for example, 0.50. In passive W, f=0, and therefore there is no mixing of the X, Y, and Z channels into the W channel.
[0094] In the case where at least one channel is transmitted as a residual channel and at least one channel is transmitted parametrically (i.e., for 2- and 3-channel downmixing), the cross-prediction coefficients (C) allow for the reconstruction of some portions of the parametric channels from the residual channels. For dual-channel downmixing (described in further detail below), the C coefficients allow for the reconstruction of some X and Z channels from Y', and the reconstruction of remaining signal components that cannot be reconstructed from the PR and C parameters by means of the decorrelation version of the W channel, as described in further detail below. In the case of 3-channel downmixing, Y' and X' are used alone to reconstruct Z.
[0095] In some embodiments, the SPAR encoder 701 includes a passive / active predictor unit 702, a remixing unit 703, and an extract / downmixing selection unit 704. In some embodiments, the passive / active predictor can receive 4-channel B-format (W, Y, Z, X) FOA channels and can calculate downmixing channels (representations of W (or W'), Y', Z', X').
[0096] In some embodiments, the extract / downmixing selection unit 704 extracts SPAR FOA post-data from one of the post-data payload segments of a bitstream (e.g., an Immersive Voice and Services (IVAS) bitstream), as described in more detail below. The passive / active predictor unit 702 and the remixing unit 703 use the SPAR FOA post-data to generate remixed FOA channels (W or W' and A'), which are input to the core encoder 705 to be encoded into a core-coded bitstream (e.g., an EVS bitstream), which is encapsulated in the IVAS bitstream sent to the SPAR decoder 706. It should be noted that in this example, the high-fidelity stereo reproduction B-format channels are configured using the AmbiX convention. However, other conventions, such as the Furse-Malham (FuMa) convention (W, X, Y, Z), may also be used.
[0097] Referring to SPAR decoder 706, the core encoded bitstream (e.g., an EVS bitstream) is decoded by core decoder 707, resulting in Ndmx (e.g., Ndmx = 2) downmixed channels. In some implementations, SPAR decoder 706 performs the opposite operation to that performed by SPAR encoder 701. For example, in the example of FIG7A, remixed FOA channels (represented as W', A', B', C') are recovered from the two downmixed channels using SPAR FOA space post-processing data. The remixed SPAR FOA channels are input to inverse mixer 711 to recover the SPAR FOA downmixed channels (represented as W', Y', Z', X'). Then, the predicted SPAR FOA channels are input to inverse predictor 712 to recover the original unmixed SPAR FOA channels (W, Y, Z, X).
[0098] It should be noted that in this dual-channel example, the decorcorrelator blocks 709A (dec1) and 709B (dec2) are used to generate a decorcorrelated version of the W' channel using a time-domain or frequency-domain decorcorrelator. The demixing channel and the decorcorrelated channel are combined with SPAR FOA post-data to parametrically reconstruct the X and Z channels. Block C 708 indicates that the residual channel is multiplied by a 2x1 C coefficient matrix to generate two cross-prediction signals, which are summed into the parametric reconstruction channel, as shown in Figure 7A. Blocks P1 710A and P2 710B indicate that the decorcorrelator output is row-multiplied by a 2x2 P coefficient matrix to generate four outputs, which are summed into the parametric reconstruction channel, as shown in Figure 7A.
[0099] In some implementations, depending on the number of downmixing channels, one of the FOA inputs is sent fully to the SPAR decoder 706 (W channel), and one to three of the other channels (Y, Z, and / or X) are sent to the SPAR decoder 706 as residual channels or fully parameterized. PR coefficients (which remain constant regardless of the number of downmixing channels Ndmx) are used to minimize the predictable energy in the residual downmixing channels. C coefficients are used to further aid in regenerating fully parameterized channels from the residual channels. Therefore, C coefficients are not required in single-channel and four-channel downmixing cases, where there are no residual or parameterized channels available for prediction. P coefficients are used to fill in any remaining energy not compensated by the PR and C coefficients. The number of P coefficients depends on the number of downmixing channels N in a frequency band. In some implementations, the SPAR PR coefficients (passive W only) are determined using the following four steps.
[0100] Step 1: Side signals (e.g., Y, Z, X) can be predicted from the main W signal, which can represent an omnidirectional signal. In some implementations, the side signals are predicted based on prediction parameters associated with the corresponding prediction channels. In one example, the side signals Y, Z, and X can be determined as follows:
[0101] As mentioned above, the prediction parameters for each channel can be determined based on the covariance matrix. In one example:
[0102] In the above text, RAB represents the elements of the input covariance matrix of signals A and B. In some implementations, the covariance matrix can be determined based on the frequency band. Note that the prediction parameters prz and prx can be determined for the residual channels Z' and X' respectively in a similar manner. Note that, as used herein, the vector PR represents the vector of prediction coefficients. For example, the vector PR can be determined as [pry,prz,prx]T.
[0103] Step 2: The W channel and the predicted Y', Z', and X' signals can be remixed. As used herein, remixing can refer to reordering or recombining signals based on a criterion. For example, in some implementations, the W channel and the predicted Y', Z', and X' signals can be remixed from most acoustically correlated to least acoustically correlated. As a more specific example, in some implementations, the signals can be remixed by reordering the input signals to W, Y', X', and Z' because audio cues from the left-right direction (e.g., the Y' signal) are more acoustically correlated than audio cues from the front-back direction (e.g., the X' signal), and audio cues from the front-back direction are in turn more acoustically correlated than audio cues from the top-bottom direction (e.g., the Z' signal). Generally speaking, the following can be used to determine the remixed signal:
[0104] In the above text, [remix] refers to a matrix that indicates one of the criteria used to reorder signals.
[0105] Step 3: The covariance of the 4-channel post-prediction and remixing can be determined. For example, the covariance matrix Rpr of the 4-channel post-prediction and remixing can be determined as follows:
[0106] Using the above conditions, the covariance matrix Rpr can have the following format:
[0107] In the above, d represents the residual channel (e.g., if the number of downmixed channels is represented by Ndmx, the residual channel is the second to the Ndmxth channel), and u represents the parameter channel to be fully reconstructed by the decoder (e.g., the Ndmx+1th to the fourth channel). Following the naming convention of one of the W, A, B, and C channels, where A, B, and C correspond to the remixed X, Y, and / or Z channels, the following table illustrates the d and u channels for different values of Ndmx. N dmx d u 1 ---- A', B', C' 2 A' B'、C' 3 A'、B' C' 4 A', B', C' ----
[0108] In some implementations, by utilizing the Rdd, Rud, and Ruu elements of the Rpr covariance matrix (described above), the FOA codec can determine whether a portion of the full parameter channels can be cross-predicted from the remaining channels transmitted to the decoder. For example, in some implementations, the cross-prediction coefficient C can be determined based on the Rdd, Rud, and Ruu elements of the covariance matrix. In one example, the cross-prediction coefficient C can be determined by:
[0109] It should be noted that C can have a shape (1x2) for a 3-channel downmixer and a shape (2x1) for a 2-channel downmixer.
[0110] Step 4: The remaining energy in the parameterized channel reconstructed by the decorcorrelators 709A and 709B can be determined. In some embodiments, the remaining energy can be represented by a matrix P. Since P can be a covariance matrix and is therefore Hermetian symmetric, in some embodiments, only elements from the upper or lower triangle of matrix P are sent to the decoder. The diagonal elements of matrix P can be real numbers, and the off-diagonal elements can be complex numbers. In some embodiments, the remaining energy represented by matrix P can be determined based on the residual energy Resuu in the upmixing channel. In one example, P can be determined by:
[0111] In another instance, only diagonal elements are used to compute the P-parameters, where the number of P-parameters transmitted to the decoder per frequency band is equal to the number of channels to be parametrically reconstructed at the decoder. Here, P can be determined by: , where
[0112] In the above, scale represents a normalized scaling factor. In some implementations, scale can be a wideband value. In one instance, scale = 0.01. Alternatively, in some implementations, scale can be frequency-dependent. In some such implementations, scale can take different values in different frequency bands. In one instance, the spectrum can be divided into 12 frequency bands, and scale can be determined by, for example, a linearly divided vector (0.5, 0.01, 12).
[0113] In some embodiments, the residual energy Resuu in the upmix channel can be determined based on the actual energy post-prediction (e.g., Ruu) and the regenerated cross-prediction energy Reguu. In one instance, the residual energy in the upmix channel can be the difference between the actual energy post-prediction and the regenerated cross-prediction energy Reguu. In one instance, Resuu = Ruu – Reguu. In some embodiments, the regenerated cross-prediction energy Reguu can be determined based on the cross-prediction coefficients and the prediction covariance matrix. For example, in some embodiments, Reguu can be determined by:
[0114] Referring back to Figure 7A, in some embodiments, signals associated with the downmixing channels (e.g., W', Y', X', and / or Z') are provided to the AGC encoder 713. The AGC encoder 713 can then determine the gain parameters in response to the determination that at least one of the downmixing channels has an overload condition, for example using the techniques described above in conjunction with Figures 2 and 5. The gain parameters and information associated with the PR, C, and / or P matrices can be encoded as ancillary information, such as post-processing data.
[0115] Figure 7B is a block diagram of an IVAS codec 750 for encoding and decoding IVAS bitstreams according to an embodiment. The IVAS codec 750 includes an encoder and a remote decoder. The IVAS encoder includes a spatial analysis and demixing unit 752, a quantization and entropy coding unit 753, an AGC gain control unit 762, a core coding unit 756, and a mode / bitrate control unit 757. The IVAS decoder includes a quantization and entropy decoding unit 754, a core decoding unit 758, an inverse gain control unit 763, a spatial synthesis / reproduction unit 759, and a decorrelator unit 761.
[0116] The spatial analysis and downmixing unit 752 receives an N-channel input audio signal 751 representing an audio scene. The input audio signal 751 includes, but is not limited to: mono signal, stereo signal, binaural signal, spatial audio signal (e.g., multi-channel spatial audio object), FOA, high-order high-fidelity stereo reproduction (HOA), and any other audio data. The spatial analysis and downmixing unit 752 downmixes the N-channel input audio signal 751 to a specified number of downmixing channels (Ndmx). In this example, Ndmx <= N. The spatial analysis and downmixing unit 752 also generates accompanying information (e.g., spatial post-processing data) that can be used by a remote IVAS decoder to synthesize the N-channel input audio signal 751 from the Ndmx downmixing channels, spatial post-processing data, and the decoder-generated decorrelation signal. In some embodiments, the spatial analysis and downmixing unit 752 implements a Complex Advanced Coupling (CACPL) for analyzing / downmixing stereo / FOA audio signals and / or a Spatial Reconstructor (SPAR) for analyzing / downmixing FOA audio signals. In other embodiments, the spatial analysis and downmixing unit 752 implements other formats.
[0117] Ndmx downmixing channels may contain a set of signals bounded by [-max,max] for a given frame. Since a core encoder 756 can encode signals in the range of [-1,1), samples of signals associated with downmixing channels that exceed the range of the core encoder 756 will cause overload. To keep the downmixing channels within the desired range, Ndmx channels are fed to a gain control unit 762, which dynamically adjusts the gain of the frame so that the downmixing channels are within the range of the core encoder. Gain adjustment information (AGC post-processing data) is sent to a quantization and coding unit 753 that writes the AGC post-processing data.
[0118] The Ndmx channels with gain adjustment are written by one or more instances of the core codec included in the core coding unit 756. The accompanying information (e.g., spatial post-data (MD)) and AGC post-data are quantized and written by the quantization and entropy writing unit 753. The written bits are then encapsulated into one or more IVAS bitstreams and sent to the IVAS decoder. In one embodiment, the underlying core codec can be any suitable mono, stereo, or multichannel codec that can be used to generate the encoded bitstream.
[0119] In some embodiments, the core codec is an EVS codec. The EVS coding unit 756 complies with 3GPP TS 26.445 and provides a wide range of functionality, such as enhanced quality and coding efficiency for narrowband (EVS-NB) and wideband (EVS-WB) voice services, enhanced quality of use of ultrawideband (EVS-SWB) voice, enhanced quality of mixed content and music in conversational applications, robustness to packet loss and latency jitter, and backtracking compatibility with AMR-WB codecs.
[0120] At the decoder, the Ndmx channels are decoded by one or more instances of the core codec contained in the core decoding unit 758, and the accompanying information, including AGC post-processing data, is decoded by the quantization and entropy decoding unit 754. A main downmixing channel (such as the W channel in FOA signal format) is fed to the decorrelation unit 761, which generates N to Ndmx decorrelation channels. The Ndmx downmixing channels and the AGC post-processing data are fed to the inverse gain control block 763, which cancels the gain adjustment performed by the gain control unit 762. The Ndmx downmixing channels with inverse gain adjustment, the N to Ndmx decorrelation channels, and the accompanying information are fed to the spatial synthesis / reproduction unit 759, which uses these inputs to synthesize or reproduce the original N-channel input audio signal that can be presented by the audio device 760. In one embodiment, the Ndmx channels are decoded by a mono codec other than EVS. In other embodiments, the Ndmx channels are decoded by a combination of one or more multi-channel core codec units and one or more single-channel core codec units.
[0121] In some implementations, the FOA codec may allocate or distribute bits for gain control between bits used for encoding spatial post-processor data (e.g., bits used for reconstructing parametric encoded channels, such as PR, C, and P parameters in SPAR) and bits used for encoding de-mixing channels. Generally, the number of bits used for encoding post-processor data is typically referred to herein as MDbits, and the number of bits used for encoding de-mixing channels is typically referred to herein as EVSbits, where EVS is a perceptual codec used for encoding de-mixing channels. It should be noted that although the examples given below refer to the use of an EVS codec as a codec, the techniques described below can be applied to any other suitable codec. In some implementations, the FOA codec may allocate bits for gain control by: 1) determining the number of bits to encode gain information; 2) determining the number of bits to encode post-data (e.g., determining MDbits); 3) determining the number of bits to encode down-mixing channels (e.g., determining EVSbits); and 4) allocating gain control bits from the post-data bits and / or EVSbits such that fewer bits are used to encode the post-data and / or down-mixing channels compared to examples where gain control is not applied (and therefore, gain control information is not encoded).
[0122] Figure 8 is a flowchart of an example program 800 for allocating gain control bits according to some embodiments. In some embodiments, program 800 may be executed by an encoder device. In some embodiments, the blocks of program 800 may be executed in a sequence other than that shown in Figure 8. In some embodiments, two or more blocks of program 800 may be executed substantially in parallel. In some embodiments, one or more blocks of program 800 may be omitted.
[0123] At 802, program 800 may determine the number of bits to be used for encoding gain control information. The number of bits used to encode a gain parameter is typically represented herein as x. As described above in conjunction with Figure 5, in some embodiments where a common gain transition function is applied to all downmixing channels, the number of bits used to encode gain control information may be represented as x+1, where x bits are used to encode gain parameter information, and a single bit is used to indicate the transition function. Alternatively, as described above in conjunction with Figure 5, in embodiments where a gain transition function is applied individually to each downmixing channel where an overload condition exists, the number of bits used to encode gain control information may depend on the number of downmixing channels (e.g., Ndmx) and the number N of downmixing channels where an overload condition exists (and therefore, gain control is applied). In these examples, the number of bits used to encode gain control information can be represented by Ndmx + (x + 1) * N, where a single bit is used for each downmixing channel to indicate whether gain control has been applied, and an exception flag is used for each downmixing channel where gain control has been applied to indicate the transition function. It should be noted that in examples where the number of downmixing channels is 1 (e.g., using a single W channel), the number of bits used to encode gain control information can be represented as 1 + (x + 1) * N.
[0124] At 804, program 800 may determine the number of bits to be used for encoding post-processing data, for example, the number of bits to be used by a decoder to reconstruct the post-processing data of the parametric encoded channels, commonly referred to herein as MDbits. In some implementations, MDbits may be determined such that MDbits is a value between a target number of bits to be used for encoding post-processing data (commonly referred to herein as MDtar) and a maximum number of bits available for encoding post-processing data (commonly referred to herein as MDmax). In some implementations, MDtar may be determined based on a target number of bits to be used for encoding a downmixing channel (commonly referred to herein as EVStar), and MDmax may be determined based on a minimum number of bits to be used for encoding a downmixing channel (commonly referred to herein as EVSmin). In one example:
[0125] In the above, IVASbits represents the number of bits available for encoding information associated with the IVAS codec, and headerbits represents the number of bits available for encoding a one-bit stream header. In some implementations, MDbits may be less than or equal to MDmax. In other words, the number of bits used to encode post-processing data may be the number of bits that allow a sufficient number of bits to be used to encode downmixing channels to maintain audio quality.
[0126] In some implementations, an iterative procedure can be used to determine MDbits. One example of such an iterative procedure is as follows:
[0127] Step 1: Based on each frame of the input audio signal, the post-data parameters can be quantized, for example, in a non-time differential manner and encoded, for example, using an arithmetic programmer. If the number of bits MDbits is less than the target number of post-data bits (e.g., MDtar), the iteration can exit, and the post-data bits can be encoded into the bitstream. The downmixed channel can be encoded by the core encoder (e.g., EVS codec) using any additional bits (e.g., MDtar - MDbits), thereby increasing the bit rate of the encoded downmixed audio channel. If MDbits is greater than the target number of bits, the iteration can continue to Step 2.
[0128] Step 2: A subset of the post-set data parameters associated with the frame can be quantized and subtracted from the quantized post-set data parameter values of the previous frame, and differential quantization parameter values can be encoded (e.g., using time differential write coding). If the updated value of MDbits is less than MDtar, the iteration can exit, and the post-set data bits can be encoded into the bitstream. Any additional bits (e.g., MDtar-MDbits) can be utilized by the core encoder (e.g., EVS codec). If MDbits is greater than the target number of bits, the iteration can continue to step 3.
[0129] Step 3: Determine MDbits when quantizing the meta data parameters in an entropy-free manner. Compare the values of MDbits from steps 1, 2, and 3 with the maximum number of bits available for encoding the meta data (e.g., MDmax). If the minimum value of MDbits from steps 1, 2, and 3 is less than MDmax, the iterative process exits, and the meta data can be encoded into the bitstream using the minimum value of MDbits. A target number of bits (e.g., MDbits - MDtar) can be allocated from the bits to be used for encoding the downmixing channel to encode the meta data. However, if in step 3, the minimum value of MDbits from steps 1, 2, and 3 exceeds MDmax, the iterative process continues to step 4:
[0130] Step 4: The meta data parameters can be coarser quantized, and the number of bits associated with the coarser quantized parameters can be analyzed according to steps 1 to 3 above. If the meta data parameters after coarser quantization still do not meet the criterion that the number of meta data bits (MDbits) is less than the maximum number of allocated bits used to encode the meta data, then a quantization scheme that guarantees quantization of the meta data parameters within the maximum number of allocated bits is used.
[0131] Referring back to Figure 8, at block 806, program 800 determines the number of bits used to encode the downmixing channel, commonly referred to herein as EVSbits. As described above in conjunction with block 804, in some implementations, the number of bits to be used to encode the downmixing channel may depend on the number of bits used to encode the meta data. For example, in an example where fewer bits are used to encode the meta data parameters, more bits can be used to encode the downmixing channel. Conversely, in an example where more bits are used to encode the meta data parameters, fewer bits can be used to encode the downmixing channel. In one instance, EVSbits can be determined as follows:
[0132] In some embodiments, if the number of bits available for encoding a downmixing channel (e.g., EVSbits) is less than the target number of bits to be used for encoding a downmixing channel (generally referred to herein as EVStar), bits can be reallocated across different downmixing channels. In some embodiments, bits can be reallocated from channels based on acoustic saliency or acoustic importance. For example, in some embodiments, bits can be obtained from channels in the order of Z', X', Y', and W' because audio signals corresponding to the up-down direction (e.g., channel Z') may be less acoustically relevant than those in other directions (e.g., front-back or channel X', or left-right or channel Y').
[0133] Conversely, in some implementations, if the number of bits available for encoding the downmixing channels (e.g., EVSbits) is greater than the target number of bits (EVStar), additional bits can be distributed across the downmixing channels. In some implementations, the distribution of additional bits can be based on the acoustic importance of various downmixing channels. In one example, additional bits can be distributed in the order of W', Y', X', and Z', such that additional bits are preferentially allocated to omnidirectional channels.
[0134] At 808, program 800 may determine one of the bit allocations among gain control bits, post-data bits, and / or down-mixing channel bits. In other words, program 800 may determine the number of bits used to reduce the number of post-data bits (e.g., MDbits) and / or down-mixing channel bits (e.g., EVSbits) so that the gain control information can be encoded using the number of gain control bits determined in block 802.
[0135] In some embodiments, program 800 may allocate bits for encoding downmixing channels to encode gain control information. For example, in some embodiments, program 800 may reduce the number of EVSbits to be used for encoding gain control information. In some of these embodiments, bits for encoding downmixing channels may be allocated to encode gain control information in an order based on the acoustic importance or relevance of the downmixing channels. In one instance, bits may be obtained from the downmixing channels in the order of Z', X', Y', and W'. In some embodiments, the maximum number of bits available from a single downmixing channel may correspond to the difference between the target number of bits to be used to encode the downmixing channel and the minimum number of bits to be used to encode that channel. In some embodiments, if no bits are available for encoding gain control information from the bits allocated to encode the downmixing channels, program 800 may adjust the bit rate of one or more downmixing channels (e.g., reduce the bit rate by one bit) to free up bits for encoding gain control information. In one instance, if EVSbits is set to the minimum number of bits to be used to encode the downmixed channel for all downmixed channels, then procedure 800 may reduce the bit rate. Alternatively, in some implementations, procedure 800 may allocate bits for encoding gain control information from the bits to be used to encode post-data parameters.
[0136] It should be noted that in some embodiments, program 800 may use bits allocated to encode the downmixing channel and bits allocated to encode the post-data parameters to allocate bits to be used for encoding gain control information. For example, in some embodiments, given the AGC bits required to encode the gain control information, program 800 may allocate m bits from the bits initially allocated to encode the post-data parameters, as determined in block 804, and allocate AGC bits-m bits from the bits initially allocated to encode the downmixing channel, as determined in block 806.
[0137] Next, program 800 may continue to the next frame of the input audio signal.
[0138] Figure 9 illustrates an example use case of an IVAS system 900 according to one embodiment. In some embodiments, various devices communicate via a call server 902, which is configured to receive audio signals from, for example, a PSTN or PLMN illustrated by a Public Switched Telephone Network (PSTN) / other Public Land Mobile Network (PLMN) device 904. The use case supports conventional devices 906 that reproduce and capture audio in mono only, including but not limited to: devices supporting Enhanced Voice Service (EVS), Multi-Rate Wideband (AMR-WB), and Adaptive Multi-Rate Narrowband (AMR-NB). The use case also supports user equipment (UE) 908 and / or 914 that capture and reproduce stereo audio signals, or UE 910 that captures mono signals and reproduces binaural signals as multi-channel signals. The use case also supports immersive and stereo signals captured and reproduced by video conferencing systems 916 and / or 918, respectively. The use case also supports stereo capture and immersive rendering of stereo audio signals for home theater systems 920, and mono capture and immersive rendering of audio signals for virtual reality (VR) devices 922 and immersive content capture 924 by computer 912.
[0139] Figure 10 is a block diagram illustrating one example of components of a device capable of implementing the present invention in various forms. As with other figures provided herein, the types and numbers of elements shown in Figure 10 are provided by way of example only. Other embodiments may include more, fewer, and / or different types and numbers of elements. According to some examples, device 1000 may be configured to perform at least some of the methods disclosed herein. In some embodiments, device 1000 may be or may include one or more components of a television set, an audio system, a mobile device (such as a cellular phone), a laptop computer, a tablet device, a smart speaker, or another type of device.
[0140] According to some alternative embodiments, device 1000 may be or may include a server. In some such instances, device 1000 may be or may include an encoder. Thus, in some examples, device 1000 may be a device configured for use in an audio environment (such as a home audio environment), while in other examples, device 1000 may be a device configured for use in the "cloud," for example, a server.
[0141] In this example, device 1000 includes an interface system 1005 and a control system 1010. In some embodiments, the interface system 1005 may be configured to communicate with one or more other devices in an audio environment. In some embodiments, the audio environment may be a home audio environment. In other embodiments, the audio environment may be another type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, etc. In some embodiments, the interface system 1005 may be configured to exchange control information and associated data with the audio devices in the audio environment. In some embodiments, the control information and associated data may relate to one or more software applications being executed by device 1000.
[0142] In some embodiments, the interface system 1005 may be configured to receive or provide a content stream. The content stream may contain audio data. The audio data may include, but is not limited to, audio signals. In some examples, the audio data may include spatial data, such as channel data and / or spatial post-processing data. In some instances, the content stream may contain video data and audio data corresponding to the video data.
[0143] Interface system 1005 may include one or more network interfaces and / or one or more external device interfaces, such as one or more Universal Serial Bus (USB) interfaces. According to some embodiments, interface system 1005 may include one or more wireless interfaces. Interface system 1005 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some instances, interface system 1005 may include one or more interfaces between control system 1010 and a memory system, such as the selected memory system 1015 shown in FIG. 10. However, in some examples, control system 1010 may include a memory system. In some embodiments, interface system 1005 may be configured to receive input from one or more microphones in an environment.
[0144] The control system 1010 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic and / or discrete hardware components.
[0145] In some embodiments, the control system 1010 may reside in more than one device. For example, in some embodiments, a portion of the control system 1010 may reside in one device within one of the environments described herein, and another portion of the control system 1010 may reside in a device outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet). In other instances, a portion of the control system 1010 may reside in one device within an environment, and another portion of the control system 1010 may reside in one or more other devices within that environment. For example, a portion of the control system 1010 may reside in one device implementing a cloud-based service, such as a server, and another portion of the control system 1010 may reside in another device implementing a cloud-based service, such as another server, a memory device, etc. In some instances, the interface system 1005 may also reside in more than one device.
[0146] In some embodiments, the control system 1010 may be configured to perform at least partially the methods disclosed herein. According to some examples, the control system 1010 may be configured to implement methods such as determining gain parameters, applying gain transient functions, determining inverse gain transient functions, applying inverse gain transient functions, distributing bits for gain control relative to a one-bit stream, or the like.
[0147] Some or all of the methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices, such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. One or more non-transitory media may reside, for example, in the selected memory system 1015 and / or control system 1010 shown in FIG. 10. Therefore, various novel forms of the subject matter described herein can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, include instructions for determining gain parameters, applying gain transition functions, determining inverse gain transition functions, applying inverse gain transition functions, distributing bits for gain control relative to a one-bit stream. The software may be executed, for example, by one or more components of a control system (such as control system 1010 of FIG. 10).
[0148] In some instances, device 1000 may include the optional microphone system 1020 shown in FIG. 10. The optional microphone system 1020 may include one or more microphones. In some embodiments, one or more microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc. In some instances, device 1000 may not include a microphone system 1020. However, in some of these embodiments, device 1000 may still be configured to receive microphone data from one or more microphones in an audio environment via interface system 1010. In some of these embodiments, a cloud-based implementation of device 1000 may be configured to receive microphone data or at least a noise measure corresponding to the microphone data from one or more microphones in an audio environment via interface system 1010.
[0149] According to some embodiments, device 1000 may include the optional loudspeaker system 1025 shown in FIG. 10. The optional loudspeaker system 1025 may include one or more loudspeakers, which, or the like, may also be referred to herein as "loudspeakers" or more generally as "audio reproduction transducers". In some instances, such as cloud-based embodiments, device 1000 may not include a loudspeaker system 1025. In some embodiments, device 1000 may include headphones. The headphones may be connected or coupled to device 1000 via a headphone jack or via a wireless connection (e.g., Bluetooth).
[0150] Some embodiments of the present invention include a system or apparatus configured (e.g., programmed) to perform one or more instances of the disclosed methods and a tangible computer-readable medium (e.g., a magnetic disk) storing program code for implementing one or more instances of the disclosed methods or steps thereof. For example, some disclosed systems may be or include a programmable general-purpose processor, digital signal processor, or microprocessor, programmed to have software or firmware and / or otherwise configured to perform various operations on data, including an embodiment of the disclosed methods or steps thereof. This general-purpose processor may be or include a computer system containing an input device, a memory, and a processing subsystem, the processing subsystem being programmed (and / or otherwise configured) to perform one or more instances of the disclosed methods (or steps thereof) in response to verified data.
[0151] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmable or otherwise configured) to perform desired processing on (a number of) audio signals, including performing one or more instances of the disclosed methods. Alternatively, embodiments of the disclosed system (or its elements) may be implemented as a general-purpose processor, such as a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory programmed with software or firmware and / or otherwise configured to perform any of the various operations including one or more instances of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general-purpose processor or DSP configured (e.g., programmable) to perform one or more instances of the disclosed methods, and the system also includes other elements. Other elements may include one or more loudspeakers and / or one or more microphones. A general-purpose processor configured to perform one or more instances of the disclosed methods may be coupled to an input device. Examples of input devices include, for example, a mouse and / or a keyboard. A general-purpose processor may be coupled to a memory, a display device, etc.
[0152] Another aspect of the present invention is a computer-readable medium, such as a magnetic disk or other tangible storage medium, which stores program code for performing (e.g., executable by a programmer) one or more instances of the disclosed method or its steps.
[0153] Although specific embodiments and applications of the invention have been described herein, those skilled in the art will understand that many variations of the embodiments and applications described herein are possible without departing from the scope of the invention as described and claimed herein. It should be understood that although certain forms of the invention have been shown and described, the invention is not limited to the specific embodiments or methods described and shown. [Simplified Explanation of the Diagram]
[0017] Figure 1 is a schematic block diagram of a system for gain control of providing audio signals according to some embodiments.
[0018] Figure 2 is a schematic block diagram of a system for implementing adaptive gain control according to some embodiments.
[0019] Figures 3A and 3B respectively illustrate examples of a gain function that can be implemented by an encoder and an inverse gain function that can be implemented by a decoder according to some embodiments.
[0020] Figure 4 shows an example diagram of an inverse gain that can be applied by a decoder in response to a dropped frame, according to some embodiments.
[0021] Figure 5 is a flowchart of an example program that can be executed by an encoder to implement adaptive gain control according to some embodiments.
[0022] Figure 6 is a flowchart of an example program that can be executed by a decoder to implement adaptive gain control according to some embodiments.
[0023] Figure 7A is a schematic diagram of an example of an encoder and decoder using a spatial reconstruction coding technique according to some embodiments.
[0024] Figure 7B is a block diagram of an example multichannel codec utilizing adaptive gain control according to some embodiments.
[0025] Figure 8 is a flowchart of an example procedure for bit distribution when implementing adaptive gain control according to some embodiments.
[0026] Figure 9 illustrates an example use case of an Immersive Voice and Services (IVAS) system according to one of some embodiments.
[0027] Figure 10 shows a block diagram illustrating an example of components of a device capable of implementing the present invention in various forms.
[0028] In each drawing, the same element symbol and name indicate the same element.
Claims
1. A method for performing gain control on an audio signal, the method comprising: Determine the downmixed signal associated with one or more downmixed channels associated with a current frame of one of the audio signals to be encoded; The method involves determining whether an overload condition exists in one of the encoders of the downmixed signals to be used for encoding at least one of the one or more downmixed channels; in response to the determination that the overload condition exists, determining a gain parameter of the at least one of the one or more downmixed channels of the current frame of the audio signal; determining at least one gain transition function based on the gain parameter and a gain parameter associated with a previous frame of the audio signal; applying the at least one gain transition function to one or more of the downmixed signals; and encoding the downmixed signals using the applied at least one gain transition function in conjunction with information indicating gain control applied to the current frame; wherein the at least one gain transition function includes a transition portion and a steady-state portion, and wherein the transition portion corresponds to a transition from the gain parameter associated with the previous frame of the audio signal to the gain parameter associated with the current frame of the audio signal; wherein the duration of one of the transition portions is related to a delay introduced by a codec including the encoder, and wherein the duration of the transition portion is the frame length minus the codec delay.
2. The method of claim 1, wherein a portion of the frame buffer is used to determine the at least one gain transition function.
3. The method of claim 2, wherein the use of the partial frame buffer determines that the at least one gain transition function introduces substantially zero additional delay.
4. The method of claim 1, wherein the transition portion has a transition type of fading, wherein the gain increases on a portion of the sample of the current frame in response to a decrease in the attenuation associated with the gain parameter of the previous frame being greater than a decrease associated with the gain parameter of the current frame.
5. The method of claim 1, wherein the transition portion has a transition type of reverse fading, wherein the gain decreases on a portion of the sample of the current frame in response to a decrease in attenuation associated with the gain parameter of the previous frame being less than a decrease associated with the gain parameter of the current frame.
6. The method of claim 1, wherein a prototype function and a scaling factor are used to determine the transition portion, and wherein the scaling factor is determined based on the gain parameter associated with the current frame and the gain parameter associated with the previous frame.
7. The method of claim 1, wherein the information indicating the gain control applied to the current frame includes information indicating the transition portion of the at least one gain transition function.
8. The method of any one of claims 1 to 3, wherein the at least one gain transition function comprises a single gain transition function applied to one of all one or more downmixing channels in the presence of the overload condition.
9. The method of any one of claims 1 to 3, wherein the at least one gain transition function includes a single gain transition function applied to all of the one or more downmixing channels, and wherein the overload condition exists in a subset of the one or more downmixing channels.
10. The method of any one of claims 1 to 3, wherein the at least one gain transition function includes a gain transition function of each of the one or more downmixing channels where the overload condition exists.
11. The method of claim 10, wherein the number of bits of the information used to encode the gain control applied to the current frame is substantially linearly scaled to the number of downmixing channels in which the overload condition exists.
12. The method of any one of claims 1 to 3, further comprising: Determine the second downmixed signal associated with the one or more downmixed channels associated with the second frame of the audio signal to be encoded; For at least one of the one or more downmixed channels of the second frame, determine whether the encoder has an overload condition; and in response to determining that the second frame does not have the overload condition, encode the second downmixed signals without applying a non-unity gain.
13. The method of claim 12, further comprising setting a flag indicating that gain control is not applied to the second frame, wherein the flag includes one bit.
14. The method of any one of claims 1 to 3, further comprising: Determine the number of bits of the information used to encode the gain control applied to the current frame; And allocate the number of bits from: 1) bits for encoding subsequent data associated with the current frame; and / or 2) bits for encoding the downmixing signals to encode the information indicating the gain control applied to the current frame.
15. The method of claim 14, wherein the number of bits is allocated from the bits used to encode the downmixed signals, and wherein the bits used to encode the downmixed signals are reduced in order based on one of the spatial directions associated with the one or more downmixed channels.
16. An apparatus configured to implement the method of any one of claims 1 to 15.
17. One or more non-transitory media having software stored thereon, the software containing instructions for controlling one or more devices to perform the methods as described in any of claims 1 to 15.
Citation Information
Patent Citations
Audio decoder and decoding method using efficient downmixing
TW201142826A
Audio decoder and decoding method using efficient downmixing
TW201443876A
Apparatus For Encoding and Decoding Audio Signal and Method Thereof
US20080212803A1
Reality alternate
US20120069131A1
Methods for Parametric Multi-Channel Encoding
US20160005407A1