Method, apparatus, and system for perceptually motivated gain control
Patent Information
- Application Number
- JP2025519776
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-22
- Filing Date
- 2023-09-01
- Publication Date
- 2026-09-08
AI Technical Summary
Existing gain control methods in audio encoding and decoding systems lead to audible artifacts due to abrupt gain changes between frames, and hard transitions exacerbate these issues, while smoothing functions may not adequately address large gain differences, leading to signal distortion and increased bit requirements.
Implementing a gain transition function with a limited gain transition step size to smoothly transition between frames, allowing some samples to exceed the required signal range, reducing noticeable artifacts and optimizing encoding efficiency.
The method achieves improved perceptual quality in decoded audio by minimizing noticeable artifacts and reducing bit requirements, while maintaining signal integrity and encoding efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 378,678, filed October 6, 2023, and U.S. Provisional Patent Application No. 63 / 503,533, filed May 22, 2022, each of which is incorporated by reference in its entirety.
[0002] [Technical field] SUMMARY The present disclosure relates to systems, methods, and media for adaptive gain control in an audio environment. [Background technology]
[0003] Gain control may be used, for example, to attenuate a signal so that it is within the range expected by an audio codec. To improve the perceptual quality of an audio signal to which gain control is applied in an encoder and inverse gain control is applied in a decoder, gain transition functions have been proposed to smoothly transition between different gains applied to successive frames. This method may lead to audible artifacts if there are abrupt gain changes between successive frames. Furthermore, in some cases, the gain changes between the determined gains of successive frames are too large and / or too abrupt for a smooth transition function to be applied. In this case, a hard transition may be used to ensure that the signal is within the expected range. For example, a single bit may be used to convey information that a hard transition is used between the gains of successive frames. However, this hard transition may also lead to audible artifacts in the decoded and rendered audio signal that are worse than those caused by the original overload condition. Therefore, it is necessary to use a gain transition function to improve the perceptual quality of an encoding / decoding system and reduce the bits required for encoding.
[0004] [Notation and Nomenclature] Throughout this disclosure, including the claims, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used interchangeably to refer to any sound-emitting transducer or set of transducers. A typical headphone set includes two speakers. A speaker may be implemented to include multiple transducers, such as a woofer and a tweeter, and may be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feeds may undergo different processing in different circuit branches coupled to different transducers.
[0005] Throughout this disclosure, including the claims, the phrase "performing an operation on" a signal or data, such as filtering, scaling, transforming, or applying a gain to the signal or data, is used broadly to refer to performing the operation on the signal or data directly or on a processed version of the signal or data. For example, the operation may be performed on a version of the signal that has undergone pre-filtering or pre-processing before the operation is performed.
[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.
[0007] Throughout this disclosure, including the claims, the term "processor" is used broadly to denote a system or device that is programmable, or possibly configurable, such as with software or firmware, to perform operations on data, which may include audio, video, or other image data. Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipeline processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets. Summary of the Invention [Means for solving the problem]
[0008] In view of the above, the present disclosure provides a method, an apparatus, a program, and a computer-readable storage medium for improving automatic gain control, having the features of the respective independent claims.
[0009] According to one aspect of the present disclosure, a method for performing gain control on an audio signal is provided. The audio signal may be a Higher-Order Ambisonics (HOA) audio signal. In the method, a downmixed audio signal of an audio signal to be encoded may be obtained. Obtaining the audio signal may include receiving the downmixed audio signal. Alternatively, it may include determining the downmixed audio signal from the audio signal to be encoded. Furthermore, it may be determined that an overload condition has occurred for a frame of the downmixed audio signal. The overload condition may be a condition in which the frame of the downmixed audio signal exceeds a predefined signal range. The predefined signal range may be a signal range expected by the encoder. The encoder may be a core encoder. In response to determining that the overload condition has occurred, a gain transition function for the frame may be determined. The gain transition function may be based on at least a gain transition step size. The gain transition function may be applied to the frame to generate a gain-adjusted frame of the downmixed audio signal. The gain-adjusted frame may be an attenuated frame or an amplified frame. The gain-adjusted frame and information indicative of the gain transition function may be provided for encoding by the encoder.
[0010] By limiting the gain transition function to a gain transition step size, a smoother, less abrupt transition from continuous gains can be achieved. The gain transition step size may be insufficient to attenuate all samples of a frame into the signal range required by the core encoder. However, artifacts due to small overshoots are less noticeable than very abrupt increases or decreases in the gain parameter. Therefore, by allowing some values to be outside the required signal range, an improved audio experience can be achieved when the signal is decoded, rendered, and played back.
[0011] In some embodiments, the gain adjusted frames may be encoded with information indicative of the gain transition function.
[0012] In some embodiments, the downmixed audio signal may be a spatially coded downmixed signal.
[0013] In some embodiments, the frame of the downmixed audio signal is a current frame, and the gain transition function may be further based on a previous gain transition function applied to a frame preceding the current frame.
[0014] In some embodiments, the gain transition function may further depend on a smoothing function that is based on the gain transition step size.
[0015] In some embodiments, the gain transition function may include a transient portion and a steady-state portion. The transient portion may correspond to a transition from a gain associated with a previous frame to a gain associated with the previous frame adjusted by the gain transition step size.
[0016] In some embodiments, the gain associated with the previous frame adjusted by the gain transition step size may be an attenuation by the gain transition step size of the gain corresponding to the previous frame, or an amplification by the gain transition step size, depending on the gain adjustment target for the current frame.
[0017] In some embodiments, the length of the transient portion may be limited by the delay introduced by the codec utilized by the encoder and decoder.
[0018] This causes the gain control to provide essentially zero additional delay.
[0019] In some embodiments, the length of the transient portion may be less than or equal to the number of samples used for the encoding operation by the encoder.
[0020] In some embodiments, the gain transition function is:
number
[0021] In some embodiments, the gain transition step size may be a predefined value or may be determined from a set of predefined values of increasing size. The predefined value or set of predefined values may be determined based on a perceptual quality listening test or an objective quality measurement test. The perceptual quality listening test may be a Multi-Stimulus Test with Hidden Reference and Anchor (MUSHRA). The perceptual quality listening test may be part of the tuning process of automatic gain control in the encoder and decoder.
[0022] In some embodiments, the method may further include determining an amount of overload caused by the frame of the downmixed audio signal. Further, the gain transition step size may be determined from a set of predefined values of increasing size depending on the amount of overload.
[0023] This allows the gain transition step size to be adapted to the required rate of change between successive frames.
[0024] In some embodiments, applying the gain transition function to the frames to generate gain-adjusted frames of the downmixed signal may include applying the gain transition function to samples of the downmixed audio signal, where the total number of samples may correspond to the frames of the downmixed audio signal.
[0025] In some embodiments, encoding the gain-adjusted frame with information indicative of the gain transition function may include determining a coding scheme based on the gain transition function. In some cases, the coding scheme may be determined based on a gain transition step size. In some cases, the coding scheme may be determined based on whether an overload condition has been removed. The coding scheme may be one of a Modified Discrete Cosine Transformation (MDCT) or an Algebraic Code Excited Linear Prediction (ACELP).
[0026] This allows the encoding scheme to be optimized for the particular audio signal and required gain transition step size.
[0027] According to a further aspect, a method for performing gain control on an audio signal is provided. In the method, encoded frames of the audio signal may be received by a decoder. The encoded frames of the audio signal may be decoded to obtain frames of a downmixed audio signal and information indicative of gain control applied by an encoder. An inverse gain transition function to be applied to the frames of the downmixed audio signal may be determined based at least in part on the information indicative of the gain control applied by the encoder. The information indicative of the gain control applied by the encoder may include a gain transition step size. The inverse gain transition function may be applied to the frames of the downmixed audio signal.
[0028] In some embodiments, the method may further include upmixing the downmixed audio signal to generate an upmixed audio signal, the upmixed audio signal being suitable for rendering.
[0029] In some embodiments, the method may further include rendering the upmixed signal to generate rendered audio data.
[0030] In some embodiments, the method may further include playing the rendered audio data using one or more of loudspeakers or headphones.
[0031] In some embodiments, the inverse gain transition function may be determined by inverting the gain transition function applied by the encoder.
[0032] In some embodiments, the inverse gain transition function may include a transient portion and a steady-state portion.
[0033] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory mediums. Such non-transitory mediums may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some inventive aspects of the subject matter described in this disclosure may be implemented via one or more non-transitory mediums having software stored thereon.
[0034] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of at least partially performing the methods disclosed herein. In some implementations, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single- or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof.
[0035] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. It should be noted that the relative dimensions of the following figures may not be drawn to scale. [Brief explanation of the drawings]
[0036] [Figure 1] 1 is an exemplary schematic block diagram of a system for providing gain control of an audio signal in the prior art;
[0037] [Figure 2A] FIG. 1 is an exemplary schematic block diagram of a system for implementing adaptive gain control according to some embodiments. [Figure 2B] FIG. 1 is an exemplary schematic block diagram of a system for implementing adaptive gain control according to some embodiments.
[0038] [Figure 3A] 4 illustrates an example of a gain transition function that may be implemented by an encoder according to some embodiments. [Figure 3B] 4 illustrates an example of an inverse gain transition function that may be implemented by a decoder according to some embodiments.
[0039] [Figure 4] 1 is a flowchart of an example process that may be performed by an encoder to implement adaptive gain control according to some embodiments.
[0040] [Figure 5] 1 is a flowchart of an example process that may be performed by a decoder to implement adaptive gain control according to some embodiments.
[0041] [Figure 6] FIG. 1 illustrates an example use case for an immersive audio and services (IVAS) system according to some embodiments.
[0042] [Figure 7] FIG. 1 shows a block diagram illustrating example components of an apparatus capable of implementing various aspects of the present disclosure.
[0043] [Figure 8A] FIG. 1 illustrates an exemplary embodiment (part 1) of an audio codec utilizing perceptually motivated gain control of a downmixed signal with uniform gain transition step sizes. [Figure 8B] FIG. 2 illustrates a second exemplary embodiment of an audio codec utilizing perceptually motivated gain control of a downmixed signal with uniform gain transition step sizes.
[0044] [Figure 9A] FIG. 1 illustrates an exemplary embodiment (part 1) of an audio codec utilizing perceptually motivated gain control of a downmixed signal with non-uniform gain transition step sizes. [Figure 9B] FIG. 2 illustrates a second exemplary embodiment of an audio codec utilizing perceptually motivated gain control of a downmixed signal with non-uniform gain transition step sizes. DETAILED DESCRIPTION OF THE INVENTION
[0045] Like reference numbers and designations in the various drawings indicate like elements.
[0046] Some coding techniques for scene-based audio, stereo audio, multi-channel audio, and / or object audio rely on coding multiple component signals after a downmix operation. Downmixing may allow a reduction in the number of audio components to be coded in a waveform-preserving waveform-coded manner, while the remaining components may be parametrically coded. At the receiver side, the remaining components may be reconstructed using parametric metadata indicating the parametric coding. Because only a subset of the components are waveform-coded and the parametric metadata associated with the parametrically coded components can be coded efficiently in terms of bit rate, such coding techniques may be relatively bitrate-efficient while still enabling high-quality audio.
[0047] One issue that may arise is that the downmix channels determined by the spatial encoder may contain signals with levels that are not suitable for subsequent processing by the core codec that constructs the audio signal bitstream. For example, in some cases, the downmix signal may have levels high enough to overload the core codec, even though the original input signal is not overloaded in any of its component signals. This can cause severe distortion, such as clipping, in the reconstructed signal after decoding and rendering. This can cause significant quality degradation in the final rendered signal. One potential solution may be to attenuate the input signal to avoid overloading the core codec. However, this solution may have the disadvantage of increasing graininess because the quantizer used to encode the signal may not be operating in an optimal range.
[0048] FIG. 1 shows a schematic block diagram of a conventional system 100 for performing gain control on an encoded Higher Order Ambisonics (HOA) signal. The schematic diagram shown in FIG. 1 can be used to encode and decode an MPEG-H signal. MPEG-H is a group of international standards being developed by the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). MPEG-H has various parts, including Part 3, MPEG-H 3D Audio.
[0049] In the encoder 102, the input HOA signal is processed at 104. The processing may include, for example, decomposition, where a downmix channel is generated. The downmix channel may include a set of signals bounded by [-max, max] for a given frame. Because the core encoder 108 can encode signals within the range of [-1, 1), samples of the signal associated with the downmix channel that exceed the range of the core encoder 108 may cause overload. To avoid overload, the gain control 106 adjusts the gain of the frame so that the associated signal is within the range of the core encoder 108 (e.g., within [-1, 1]). The core encoder 108 may be considered a codec that generates an encoded bitstream. Side information generated by the decomposition / processing block 104, which may include metadata associated with parametrically encoded channels, etc., may be encoded in the bitstream in association with the signal generated as the output of the core encoder 108.
[0050] The encoded bitstream is received by the decoder 112. The decoder 112 may extract the side information, and the core decoder 116 may extract the downmix signal. The inverse gain control block 120 may then invert the gain applied by the encoder. For example, the inverse gain control block 120 may amplify a signal attenuated by the gain control 106 of the encoder 102. The HOA signal may then be reconstructed by the HOA reconstruction block 122. Optionally, the HOA signal may be rendered and / or reproduced by the rendering / reproduction block 124. The rendering / reproduction block 124 may include various algorithms for, for example, rendering the reconstructed HOA output as, for example, rendered audio data. For example, rendering the reconstructed HOA output may include distributing one or more signals of the HOA output across multiple speakers to achieve a particular perceptual impression. Optionally, the rendering / reproduction block 124 may include one or more loudspeakers, headphones, etc. for presenting the rendered audio data.
[0051] The gain control 106 may implement the gain control using the following technique: The gain control 106 may first determine an upper bound on the signal value within a frame. For example, for an MPEG-H audio signal, the bound may be
number
number
number
number
number
[0052] The gain factor gn(j) for a particular channel n and frame j is calculated by applying a one-frame delay corresponding to one HOA block and performing the following recursive operation:
number
[0053] where gn(j-2) represents the gain factor applied to frame (j-2);
number
[0054] Disclosed herein are techniques for providing adaptive gain control. In particular, as described herein, gain parameters may be determined that do not introduce additional delay because the gain parameters may be determined based on lookahead samples generated for use by a codec that may be used by a perceptual encoder. The determination of the gain transition function is illustrated in and described below in connection with Figures 2-5.
[0055] 2A and 2B show schematic block diagrams of an encoder 202 and a decoder 212, respectively, for performing low-latency adaptive gain control according to an example embodiment. In the encoder 202, an input HOA signal (or first-order Ambisonics (FOA)) signal undergoes processing by a spatial analysis block 204. For an N-channel HOA input, the spatial analysis block 204 may generate and output a set of M downmix channels 204A. The number of downmix channels in the set of M downmix channels 204A may be in the range of 1≦M≦N. In addition, the spatial analysis block 204 may generate and output spatial side information 204B for inverting the downmix operation.
[0056] For example, for an FOA input, the downmix channels may include a primary downmix channel W′, which may be generated by mixing an omnidirectional input signal W with directional input signals X, Y, and Z using various mixing gains, and up to three residual channels X′, Y′, and Z′, respectively, corresponding to signal components in the X, Y, and Z signals that cannot be predicted from the primary downmix signal. In one example, the spatial analysis block 204 utilizes spatial reconstruction (SPAR) techniques. SPAR is further described in “Spatial Reconstruction of Audio Signals,” in “Spatial Reconstruction of Audio Signals,” in Proceedings of the IEEE International Conference on Audio and Video Engineering (ARFE), pp. 1111-1114, 2013, which is incorporated herein by reference in its entirety. In other examples, the spatial analysis block 204 may utilize any other suitable linear predictive codec with an energy-compacting transform, such as the Karhunen-Loeve Transform (KLT). The core encoder 208 may be considered a codec that generates the encoded audio bitstream 208A. In some implementations, the core encoder 208 and the core decoder 216 may introduce some look-ahead samples to be utilized by the adaptive gain control 206 to determine gain parameters to avoid adding extra delay to the overall coding process (zero additional delay). [Non-Patent Document 1] D. McGrath, S. Bruhn, H. Purnhagen, M. Eckert, J. Torres, S. Brown, and D. Darcy, "Immersive Audio Coding for Virtual Reality Using a Metadata-assisted Extension of the 3GPP EVS Codec," in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 730-734.
[0057] The signals associated with the M downmix channels 204A may then be analyzed by the adaptive gain control 206. The adaptive gain control 206 may determine whether a signal associated with any of the M downmix channels 204A exceeds the audio amplitude range expected by the core encoder 208 and therefore will overload the core encoder 208. In some embodiments, if the adaptive gain control 206 determines that no gain should be applied, e.g., in response to determining that none of the signals of the M downmix channels 204A exceed the expected range of the core encoder 208, the adaptive gain control 206 may set a flag indicating that gain control is not applied. Indicating the flag may be done by setting a value for the flag, e.g., setting a single-bit value. If the adaptive gain control 206 determines that no gain should be applied, the adaptive gain control 206 may not set the flag, thereby saving one bit (e.g., the bit associated with the flag). For example, in some implementations, if the spatial metadata bitstream and / or the core encoder bitstream (which may be the perceptual encoder bitstream) are self-terminating, the presence of the gain control flag may be determined by determining whether there are any unread bits in the bitstream. The unread bits may be leftover bits in the bitstream. In cases where an overload condition does not exist, the adaptive gain control 206 may output M downmix channels 206A. The M downmix channels 206A may then be passed to the core encoder 208 for encoding in a bitstream 208A.
[0058] In contrast, if the adaptive gain control 206 determines that gain should be applied, the adaptive gain control 206 may determine gain parameters and apply gains to the M downmix channels according to the determined gain parameters. The M downmix channels 206A with gains applied may then be passed to the core encoder 208 for encoding in the bitstream. Furthermore, the adaptive gain control 206 may output gain control side information 206B. Information about the flags may be included in the gain control side information 206B. The side information encoder 210 may encode the spatial side information 210B along with the gain parameters 206B as metadata 204A for transmission in the bitstream. The decoder 212 may then extract this metadata and use it to upmix the downmixed channels and reverse the gain adjustment. For example, the metadata 210A may later be utilized to reconstruct a representation of the original audio input downmixed by the spatial analysis unit 204. The side information encoder 210 may additionally provide side information 208B to the core encoder 208. The core encoder 208 may then use the side information 208B to select one of the coding techniques. Both the encoded bitstream 208A and the encoded bitstream with metadata 210A may be multiplexed to form the final bitstream output by the encoder 202.
[0059] In some implementations, the adaptive gain control 206 may determine a gain transition function that transitions between a gain parameter e(j−1) associated with a previous frame (e.g., the j−1th frame) and a gain parameter e(j) of the current frame. The gain transition function may be applied frame-by-frame by the adaptive gain control 206, where each frame may be one frame of the M downmix channels 204A. In some implementations, the gain transition function may smoothly transition the gain parameter from the value of the gain parameter in the j−1th frame (e.g., e(j−1)) to the gain parameter of the current frame (e.g., e(j)) over samples of the jth frame. Thus, the gain transition function may include two portions: 1) a transient portion in which the gain parameter is transitioning from the gain parameter of the previous frame to the gain parameter of the current frame over samples of the transition portion; and 2) a steady-state portion in which the gain parameter has the value of the gain parameter of the current frame for samples of the steady-state portion.
[0060] In some embodiments, when the gain applied to the current frame is less than the gain applied to the previous frame, the transition portion may be referred to as having a "fade" transition type because the amount of attenuation increases over the samples of the current frame. The case where the gain applied to the current frame is less than the gain applied to the previous frame can be expressed as e(j)>e(j - 1). In some embodiments, when the gain applied to the current frame is greater than the gain applied to the previous frame, the amount of attenuation decreases over the samples of the current frame, so the transition portion may be referred to as having a "reverse fade" or "un - fade" transition type. The case where the gain applied to the current frame is greater than the gain applied to the previous frame can be expressed as e(j)<e(j - 1). In some embodiments, when the gain applied to the current frame is the same as the gain applied to the current frame, the transition portion may be referred to as having a "hold" transition type where the transition portion is not transient but rather has the same value as the steady - state portion. The case where the gain applied to the current frame is the same as the gain applied to the current frame can be expressed as e(j)=e(j - 1).
[0061] In some embodiments, the gain transition function depends on a gain transition step size. The gain transition step size may limit the amount of transition possible from the previous frame to the current frame. This is motivated by the fact that smaller, smoother gain / attenuation changes, which potentially allow overloads to occur during the transition, are perceptually better than larger changes, especially when subjected to further processing by a lossy core encoder that requires the above-mentioned predefined value range as input. By predefining the parameters of the gain transition function in this way, the effect of the parameters on objective or perceptual quality can be evaluated. Perceptual quality may be measured based on known perceptual quality listening tests, such as the Multi-Stimulus Test with Hidden Reference and Anchor (MUSHRA). The perceptual quality listening test may be part of the tuning process of automatic gain control in encoders and decoders. In particular, parameters such as the gain transition step size may be tuned for a particular audio scenario and codec until an optimal perceived audio quality is reached. The tuned parameters are then used by the encoding / decoding system.
[0062] In an exemplary implementation, the processed output 206A of the automatic gain control 206 is further coded by a lossy core codec based on Algebraic Code Excited Linear Prediction (ACELP) coding, which does not aim to perform waveform reconstruction. It has been observed that applying larger gain steps to the ACELP input and output leads to audible glitches in the reconstructed signal and degrades the overall performance of the codec.
[0063] In some implementations, when an overload for the current frame is detected, the automatic gain control 206 may also determine the amount of attenuation required for the frame to be within the expected range of the core encoder 208. If there is a large difference between the required attenuation between consecutive frames, applying a transition function to achieve the range [-1, 1) required by the core encoder 208 may lead to audible artifacts when the audio signal is rendered at the decoder. Instead of applying a transition function to keep each frame inside or on the boundary of the required range, the transition function may be limited to a specific gain transition step size. This allows the transition function to only attenuate a single frame by an amount equal to the gain transition step size, i.e., ±DBSTEP dB, regardless of the amount of attenuation required to achieve the expected range of the core encoder 208. Thus, as an example, if the attenuation of the previous frame is -10 dB, the attenuation applied to the first sample of the current frame will be -10 dB, and the attenuation applied to the last sample of the current frame will be -10 dB ±DBSTEP. To be precise, if the amount of overload does not change from the previous frame to the current frame, the gain transition will be a constant value, e.g., -10 dB. If the attenuation needs to be changed, the gain transition function transitions by ±DBSTEP from the attenuation of the previous frame to the last sample of the current frame.
[0064] In some implementations, DBSTEP may be selected so that the attenuation applied by the automatic gain control 206 is not sufficient to keep the frame within the expected signal range of the core encoder 208. For example, DBSTEP may be a fixed value. When abrupt changes in attenuation are required, allowing frames to be outside the [-1,1) range can avoid strong attenuation differences between consecutive frames. Thus, instead of forcing frames into the [-1,1) range either by a transition function or a static gain change, frames are attenuated by a fixed amount relative to the attenuation of the previous frame. Using a transition function with a specific gain transition step size can improve perceptual audio quality because distortion due to frames being outside the [-1,1) range is less noticeable compared to distortion caused by sharp attenuation differences between consecutive frames. Furthermore, an exception flag for switching between smooth transitions and static gain changes can be avoided. This can save one bit in the core encoder 208.
[0065] In some implementations, DBSTEP may be a single value, e.g., -1 dB. Alternatively, DBSTEP may be selected from a set of increasing fixed values, e.g., -1 dB, -3 dB, -6 dB. In this case, the value of DBSTEP may be selected depending on the amount of overload caused by frames with no attenuation.
[0066] In some implementations, the automatic gain control is configured with the ability to specify a set of target gain values GT, which can be expressed as a table of numbers, e.g., integers, indicating the multiplier of DBS attenuation provided at each step. This is motivated by the fact that smaller changes provide perceptual benefit, but for some signals, higher levels of potential attenuation may be required. Specifying these unequal absolute steps allows a wider attenuation range to be covered while still providing the benefit of smaller steps for many likely cases. For example, a set GT = {0, 1, 3, 6} with a DBS of -2 dB would take successive steps DBSTEP of {-2 dB, -4 dB, -6 dB}, resulting in an absolute target gain of {DBS * GT} = {0 dB, -2 dB, -6 dB, -12 dB}.
[0067] One or more of such integer tables may be specified, and information regarding the selection of the particular table being used at the encoder side may be signaled / transmitted to the decoder side. In contrast to applying uniform steps, which results in a single uniform gain transition shape, applying non-uniform steps results in non-uniform gain transition shapes (level-dependent transition functions).
[0068] In some implementations, when DBSTEP is insufficient to attenuate the current frame into the [-1,1) range, DBSTEP may be applied to frames following the current frame until the [-1,1) range is achieved.
[0069] In some implementations, output level and attenuation information from the automatic gain control 206 system can be used in decision-making processes in other systems, such as the core encoder 208. Relaxed requirements can provide perceptual benefits but can affect the core encoder 208 by introducing gain changes or by not meeting stricter requirements and allowing overload conditions to persist. Information such as whether the gain control met the requirements (if any) or how much gain was applied can be output and passed to the core encoder. This allows better decisions to be made, such as selecting a coding method that can better handle gain changes or out-of-range samples. As an example, when large gain / attenuation steps are applied, the core encoder 208 can use waveform coding techniques, such as MDCT-based coding, instead of predictive ACELP coding techniques.
[0070] In some embodiments, the transition portion of the gain transition function may be determined using a prototype shape of the transition portion of the gain transition function, where the prototype shape is scaled based on the difference between the gain parameters of the current frame and the gain parameters of the previous frame. For example, the prototype shape may be scaled based on e(j)-e(j-1). A gain transition function utilizing such a prototype function p may be expressed as follows:
number
number
number
[0071] FIG. 3A shows examples of gain transition functions each having a transient portion with a "fade" transient type. In the example shown in FIG. 3A, each gain transition function has a transition portion beginning at sample 0, which may correspond to the beginning of the current frame with a gain of 0 dB, where 0 dB is the gain parameter of the previous frame (e.g., the j-1th frame). In the example shown in FIG. 3A, the transient portion of each gain transition function changes to the steady-state portion of the gain transition function over approximately 384 samples. For each of the three gain transition functions shown in FIG. 3A, the steady-state portion corresponds to a different gain transition step size for the jth frame, with a (negative) gain increase of 6 dB, 12 dB, and 18 dB, respectively, relative to the gain of the previous frame. In other words, as shown in FIG. 3A, for the three gain transition functions, exp = -[e(j) - e(j-1)] = -1, -2, and -3, respectively. For each of the gain transition functions shown in FIG. 3A, the transition portions are the same length (e.g., approximately 384 samples). Note that the length of the steady-state portion may correspond to an offset with respect to the delay introduced by the codec, e.g., 12 ms in the example shown in Figure 3A. Correspondingly, the length of the transient portion may be related to the inverse of the offset. In the example shown in Figure 3A, the length of the transient portion is the frame length (e.g., 20 ms) minus the codec delay (e.g., 12 ms). Note that the codec delay may be the overall coder algorithm delay excluding the frame size delay.
[0072] Furthermore, a gain transition function having a "reverse fade" or "unfade" transient type transient portion may be represented as a mirror image of the gain transition function shown in Figure 3A flipped across a horizontal line. By way of example, the horizontal line may be the x-axis.
[0073] Referring again to FIG. 2B , the decoder 212 may receive as input the coded audio bitstream 208A and the metadata bitstream 210A, and may, for example, reconstruct the HOA signal for rendering or directly render it to a desired output format. In some embodiments, the core decoder 216 receives the coded audio bitstream 208A. Additionally, the core decoder 216 may receive information 214A extracted from the metadata bitstream 210A by the side information decoder 214. The core decoder 216 may decode the coded audio bitstream 208A based on the information 214A or without knowing the side information, and output M gain-adjusted downmixed channels 216A to the inverse gain control 220. The side information decoder 214 further extracts gain parameters and spatial side information and sends this information 214B to the inverse gain control 220 and the spatial synthesis / rendering / reconstruction block 222. The inverse gain control 220 may then obtain the gain parameters applied by the encoder 202 from the information 214B. For example, in some implementations, the inverse gain control 220 may obtain from the information 214B an indication of the gain transition step size DBSTEP and / or an arithmetic coefficient associated with DBSTEP applied by the encoder 202. Furthermore, the inverse gain control block 220 may retrieve, for example, from a memory, the shape of the transition function, i.e., the shape of the prototype function p, also referred to as a smoothing function. The inverse gain control block 220 may then use the obtained gain parameters to invert the gains applied by the encoder 202 and output the M downmixed channels 220A. For example, in some implementations, the inverse gain control 220 may construct an inverse gain transition function that transitions from the gain parameters of the previous frame to the gain parameters of the current frame. In some implementations, the inverse gain transition function may be the gain transition function applied by the encoder 202 mirrored and vertically adjusted across a center vertical line. As an example, the vertical line may be the y-axis.
[0074] Referring to FIG. 3B, examples of inverse gain transition functions applied by a decoder corresponding to the gain transition function shown in FIG. 3A applied by an encoder are shown, according to some implementations. As shown, the inverse gain transition function has a steady-state portion and a transition portion. The durations of the steady-state and transition portions of the inverse gain transition function may correspond to, and may be the same as, the durations of the corresponding steady-state and transition portions of the gain transition function, as shown in FIGS. 3A and 3B. As shown, each inverse gain transition function shown in FIG. 3B starts at 0 dB for the current frame and transitions to −DBSTEP. That is, each inverse gain transition function starts at 0 dB, which corresponds to the inverse gain applied to the previous frame j−1. When the gain applied by the encoder corresponds to an attenuation indicated by a gain less than 0 dB as shown in the gain transition function of FIG. 3A, the inverse gain applied by the decoder corresponds to an amplification with a gain greater than 0 dB as shown in the gain transition function of FIG. 3B. In contrast, if the gain applied by the encoder corresponds to an amplification, eg, with a gain greater than 0 dB, then the inverse gain applied by the decoder corresponds to an attenuation, eg, with a gain less than 0 dB.
[0075] Referring again to FIG. 2B , after the inverse gains are applied, the M downmix channels 220A with the inverse gains applied are provided to a spatial synthesis / rendering / reconstruction block 222. The spatial synthesis / rendering / reconstruction block 222 may reconstruct the HOA signal using the information 214B. For example, if the spatial analysis block 204 utilizes a SPAR technique for spatial encoding, the spatial synthesis / rendering / reconstruction block 222 may utilize a SPAR technique to reconstruct one or more channels encoded using the metadata 210A. The reconstructed HOA output may then be rendered directly or provided to another entity for rendering. The spatial synthesis / rendering / reconstruction block 222 may include various algorithms for rendering the reconstructed HOA output, for example, as rendered audio data. For example, rendering the reconstructed HOA output may include distributing one or more signals of the HOA output across multiple speakers to achieve a particular perceptual impression. Optionally, the spatial synthesis / rendering / playback block 222 may include an audio playback device, eg, one or more loudspeakers, headphones, etc., for presenting the rendered audio data.
[0076] 4 shows an example of a process 400 for determining gain parameters and applying gain to a downmixed signal according to the determined gain parameters, according to some implementations. In some implementations, the blocks of process 400 may be performed by an encoder device. In some implementations, the blocks of process 400 may be performed in an order other than the order shown in FIG. 4. In some implementations, two or more blocks of process 400 may be performed substantially in parallel. In some implementations, one or more blocks of process 400 may be omitted.
[0077] At 402, process 400 may obtain a downmixed audio signal associated with a frame of the audio signal to be encoded. The downmixed audio signal may be associated with the frame of the audio signal to be encoded. For example, in some implementations, process 400 may use any suitable spatial coding technique to determine a set of downmixed channels. Examples of spatial coding techniques include SPAR, linear prediction techniques, etc. The set of downmixed channels may include any of 1 to N channels, where N is the number of input channels; for example, in the case of an FOA signal, N is 4. The downmixed signal may include an audio signal corresponding to the downmixed channels for a particular frame of the audio signal.
[0078] At 404, process 400 may determine whether an overload condition exists for a codec, such as an Enhanced Voice Services (EVS) codec, and / or any other suitable codec. For example, process 400 may determine that an overload condition exists in response to determining that a signal corresponding to a frame of the downmix audio signal exceeds a predetermined range, e.g., [−1, 1), and / or any other suitable range.
[0079] If at 404 it is determined that an overload condition does not exist (“NO” at 404), process 400 may proceed to 412, where the downmixed signal may be encoded. For example, in some implementations, process 400 may generate a bitstream that encodes the downmixed signal in association with side information, such as metadata, that may be utilized by a decoder to upmix the downmixed signal, e.g., to reconstruct an FOA or HOA output.
[0080] In contrast, if it is determined at 404 that an overload condition exists (“Yes” at 404), process 400 can proceed to 406, where a gain transition function for the frame can be determined that avoids the overload condition, or if the change in the overload condition from one frame to the next is greater than the gain transition step size, the overload is at least reduced. Further, at 406, the gain transition function can be based on the gain transition step size. The gain transition function can also be based on the shape of the smoothing function. Furthermore, as described above in connection with FIG. 2, the gain transition function can have a transient portion and a steady-state portion, where the steady-state portion corresponds to the gain factor of the current frame and the transient portion corresponds to a sequence of intermediate gain factors of a subset of samples of the current frame that transitions from the gain factor at the end of the previous frame to the gain factor of the previous frame ±DBSTEP.
[0081] If the gain parameters of the previous frame correspond to a smaller attenuation than the gain parameters of the current frame, the transient portion may be referred to as having a transient type of "fade." In contrast, if the gain parameters of the previous frame correspond to a larger attenuation than the gain parameters of the current frame, the transient portion may be referred to as having a transient type of "reverse fade" or "unfade." If the gain parameters of the previous frame are the same as the gain parameters of the current frame, the transient portion may be referred to as having a transient type of "hold." If the transient portion has a transient type of "hold," the value of the gain transition function during the transient portion may be the same as the value of the gain transition function during the steady-state portion. As described above in connection with FIG. 2, the duration of the transition portion of the gain transition function may correspond to the delay duration utilized by the codec.
[0082] At 408, process 400 may apply a gain transition function to a downmixed signal associated with the frame. For example, in some implementations, process 400 may scale samples of the downmixed signal by a gain factor indicated by the gain transition function. As a more specific example, in some implementations, the first sample of the current frame may be scaled by a gain factor corresponding to the gain parameter of the preceding frame, the last sample of the current frame may be scaled by a gain factor ±DBSTEP corresponding to the gain parameter of the previous frame, and intervening samples may be scaled by gain factors corresponding to gain parameters of a transient or steady-state portion of the gain transition function.
[0083] In some implementations, the gain transition function may be applied only to the downmixed signal of the downmix channel for which an overload condition was detected in block 404. For example, if an overload condition is detected for the Y' channel and the X' channel, separate gain transition functions may be determined for each of the Y' channel and the X' channel and applied to the Y' channel and the X' channel signals. Continuing with this example, the gain transition function may not be applied to the W' and Z' channels. In such a case, an indication of the channel to which the gain transition function is applied, as well as corresponding gain parameters for each channel, may be coded, for example, in block 412. Alternatively, in some implementations, if an overload condition exists for only one downmix channel, the corresponding gain transition function may be applied to all downmix channels. In such a case, since the gain transition function is applied to all channels, an indication of the channel to which the gain is applied does not need to be transmitted, which may lead to increased bitrate efficiency.
[0084] At 410, the process 400 may provide the attenuated signal and information indicative of a gain transition function to an encoder for encoding. The information indicative of the gain transition function may be a gain transition step size and / or a mathematical coefficient related to the gain transition step size. Additionally, the shape of the smoothing function may be provided to the encoder for encoding.
[0085] At 412, process 400 may encode information indicative of the downmixed signal and, if gain was applied, gain parameters for the frame. If gain was applied, the encoded downmix signal may be the downmix signal after application of the gain transition function in block 408. The downmixed signal and any information indicative of the gain parameters, in association with any side information, such as metadata, that may be used by a decoder to reconstruct or upmix the downmixed signal, may be encoded by a codec, such as an EVS codec, to generate an encoded bitstream. The encoded bitstream, along with the metadata, may then be stored and / or transmitted to a receiving device capable of reversing the processing steps of the encoder.
[0086] Note that in some implementations, process 400 may encode the gain parameter into a set of bits. In some implementations, the gain transition function may represent a prototype / smoothing function associated with the transient portion of the gain transition function.
[0087] If adaptive gain control is enabled for each downmix channel associated with the signal that triggers the overload condition so that a unique gain transition function is applied, x bits may be utilized for each channel for which gain control is enabled, and an additional 1-bit indicator per channel indicates that a gain parameter is encoded. In such a case, the total number of bits used to transmit gain control information is Ndmx+x*Nagc, where Ndmx represents the number of downmix channels (and one bit is utilized to indicate whether gain control is enabled for each of the Ndmx channels), and Nagc represents the number of channels for which gain control is enabled. Note that if gain control is not enabled for a particular frame, Ndmx bits may be used to indicate that gain control is not enabled, e.g., 1 bit for each of the Ndmx channels. Note that if the number of downmix channels is 1, e.g., if only the W channel is waveform encoded, the total number of bits used to transmit gain control information is represented by x*Nagc. For example, given one downmix channel, if gain control is not enabled for one downmix channel (e.g., Nagc=0), the number of bits used is 0. Continuing with this example, if gain control is enabled (e.g., Nagc=1), the number of bits used is x.
[0088] If a single gain transition function associated with the downmix channel that triggers the overload condition is applied to all downmix channels, fewer bits may be used to transmit gain control information, e.g., a single gain parameter for the current frame is transmitted using x bits.
[0089] 5 shows an example of a process 500 for obtaining gain parameters utilized by an encoder and applying an inverse gain transition function based on the obtained gain parameters, according to some implementations. In some implementations, the blocks of process 500 may be performed by a decoder device. In some implementations, the blocks of process 500 may be performed in an order other than the order shown in FIG. 5. In some implementations, two or more blocks of process 500 may be performed substantially in parallel. In some implementations, one or more blocks of process 500 may be omitted.
[0090] Process 500 may begin at 502 by receiving an encoded frame of an audio signal. The received frame (e.g., the current frame) is generally referred to herein as the jth frame. The received frame may be immediately after a previously received frame or may be a frame that is not immediately after a previously received frame.
[0091] At 504, process 500 may decode the encoded frame of the audio signal to obtain a downmixed signal and information indicative of the gain control applied to the current frame, if gain control was applied by the encoder. The information indicative of the gain control applied to the current frame may be a gain transition step size applied by the encoder. Also, the information indicative of the gain control applied to the current frame may be a shape of a smoothing function of the gain transition function applied by the encoder. If the encoder applies gain control per channel, process 500 may further identify to which downmix channel the gain control was applied.
[0092] At 506, process 500 may determine an inverse gain transition function based on the gain transition step size. In some implementations, process 500 may further determine the inverse gain transition function based on the shape of the smoothing function. The inverse gain transition function may be calculated based on the gain transition function or may be selected from a number of predefined inverse gain transition functions.
[0093] In some implementations, process 500 may determine that the inverse gain transition function is the inverse of the gain transition function applied in the encoder. For example, the inverse gain transition function may correspond to a gain transition function mirrored and adjusted across a horizontal line. The mirroring and adjustment may be along the x-axis. An example of such an inverse gain transition function is shown in FIG. 3B and described above in connection with FIG. 3B. In some implementations, the inverse gain transition function may have a steady-state portion corresponding to the gain applied to the previous frame. The inverse gain transition function may have a transition portion that is the inverse of the transition portion of the gain transition function applied in the encoder. For example, if the gain applied to the current frame corresponds to a larger attenuation relative to the previous frame, the inverse gain transition function may have a transition portion that transitions from a smaller amplification to a larger amplification. In contrast, if the gain applied to the current frame corresponds to a smaller attenuation relative to the previous frame, the inverse gain transition function may have a transition portion that transitions from a larger amplification to a smaller amplification. The duration of the transient portion may be related to the delay introduced by the codec, where the duration of the transient portion is the frame length (e.g., 20 ms) minus the codec delay (e.g., 12 ms). Note that if the delay introduced by the codec is longer than the frame length, the inverse gain transition may be applied with a delay of one frame. In some cases, the delay may be obtained by process 500 (e.g., by a decoder) from gain control bits. The inverse gain transition function may also serve to attenuate signals amplified by the encoder's gain control.
[0094] At 508, process 500 may apply an inverse gain transition function to the downmixed signal to reverse the gain applied by the encoder. For example, applying the inverse gain transition function may amplify the downmix signal that was attenuated by the encoder, reversing the attenuation. As another example, applying the inverse gain transition function may attenuate the downmix signal that was amplified by the encoder, reversing the amplification. The output of step 508 may then be M downmix channels with the same gains as the M downmix channels after step 402 of process 400.
[0095] At 510, process 500 may upmix the downmixed signal. The upmixing may be performed by a spatial encoder. In some cases, the spatial encoder may utilize SPAR techniques. The upmixed signal may correspond to a reconstructed FOA or HOA audio signal. In some implementations, process 500 may upmix the signal using side information, e.g., metadata, encoded in the bitstream, and the side information may be utilized to reconstruct a parametrically encoded signal. In some implementations, block 510 may be optional, for example, when the downmixed signal can be rendered directly.
[0096] In some implementations, at 512, process 500 may render the upmixed signal to generate rendered audio data. In some implementations, process 500 may utilize any suitable rendering algorithm to render the FOA or HOA audio signal, e.g., to render scene-based audio data. In some implementations, the rendered audio data may be stored in any suitable format, e.g., for future presentation or playback. In some implementations, block 512 is optional and thus may be omitted.
[0097] In some implementations, at 514, process 500 may cause the rendered audio data to be played. For example, in some implementations, the rendered audio data may be presented through one or more of loudspeakers and / or headphones. In some implementations, multiple loudspeakers may be utilized, and the multiple loudspeakers may be positioned in any suitable position or orientation relative to one another in three dimensions. In some implementations, process 514 is optional and may therefore be omitted.
[0098] As described above in connection with FIG. 4, gain control information, e.g., information indicating gain parameters, may be coded using a set of gain control bits. In some implementations, a different gain transition function may be determined for each downmix channel in which an overload condition is detected. In such implementations, gain control bits are needed to indicate whether gain control is applied to each of the downmix channels, and as described above in connection with FIG. 4, gain transition function parameters are coded for each of the downmix channels to which gain control is applied. Alternatively, in some implementations, a single gain transition function determined based on one downmix channel in which an overload condition exists may be applied to all of the downmix channels. In such implementations, because a separate bit flag is not needed to indicate whether gain control is applied for each downmix channel, fewer gain control bits are needed, resulting in more bitrate-efficient coding.
[0099] More bitrate-efficient encoding by applying the same gain transition function to all downmix channels, including downmix channels in which no overload conditions exist, may result in degradation of perceptual quality, e.g., by attenuating signals in which no codec overload exists. In contrast, utilizing more targeted gain control, in which gain control is applied in a targeted manner to each downmix channel, may require more bits to transmit gain control information. However, utilizing additional bits to transmit targeted, e.g., channel-specific, gain control information may require reallocation of bits typically used to waveform-code the downmix channels, which may, in some cases, reduce perceptual quality. Therefore, there may be a situation-dependent trade-off between applying the same gain transition function to all downmix channels and applying channel-specific gain control. Regardless of whether gain control is applied across all downmix channels or per targeted channel, the bits associated with the gain control information may be allocated from bits typically used for waveform coding of the downmix channels and / or from bits typically used to code side information such as metadata used to reconstruct FOA or HOA signals from the downmix channels, thereby reducing the number of bits available for coding either the downmix channels or the side information.
[0100] FIG. 6 illustrates an exemplary use case of an IVAS system 600, according to one embodiment. In some embodiments, various devices communicate via a call server 602 configured to receive audio signals from, for example, a public switched telephone network (PSTN) or mobile network devices (PLMN) 604, indicated by PSTN / OTHER PLMN. The use case supports legacy devices 606 that render and capture audio only in mono, including, but not limited to, devices that support Enhanced Voice Services (EVS), Multi-Rate Wideband (AMR-WB), and Adaptive Multi-Rate Narrowband (AMR-NB). The use case also supports user equipment (UE) 608 and / or 614 that capture and render stereo audio signals, or a UE 610 that captures mono signals and binaurally renders them into a multi-channel signal. The use case also supports immersive and stereo signals captured and rendered by videoconferencing room systems 616 and / or 618, respectively. The use case also supports the computer 612 for stereo capture and immersive rendering of stereo audio signals for a home theater system 620, as well as mono capture and immersive rendering of audio signals for virtual reality (VR) gear 622 and immersive content capture 624.
[0101] 7 is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. As with other figures provided herein, the types and number of elements shown in FIG. 7 are given by way of example only. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 700 may be configured to perform at least some of the methods disclosed herein. In some implementations, device 700 may be or include a television, one or more components of an audio system, a mobile device (such as a cellular phone), a laptop computer, a tablet device, a smart speaker, or other type of device.
[0102] According to some alternative implementations, apparatus 700 may be or include a server. In some such examples, apparatus 700 may be or include an encoder. Thus, in some cases, apparatus 700 may be a device configured for use in an audio environment, such as a home audio environment, while in other cases, apparatus 700 may be a device configured for use in the “cloud,” e.g., a server.
[0103] In this example, device 700 includes an interface system 705 and a control system 710. The interface system 705, in some implementations, may be configured to communicate with one or more other devices in an audio environment. The audio environment, in some examples, may be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 705, in some implementations, may be configured to exchange control information and associated data with audio devices in the audio environment. The control information and associated data, in some examples, may relate to one or more software applications that device 700 is executing.
[0104] The interface system 705, in some implementations, may be configured to receive or provide a content stream. The content stream may include audio data. The audio data may include, but is not limited to, an audio signal. In some cases, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.
[0105] The interface system 705 may include one or more external device interfaces, such as one or more network interfaces and / or one or more universal serial bus (USB) interfaces. According to some implementations, the interface system 705 may include one or more wireless interfaces. The interface system 705 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 705 may include one or more interfaces between the control system 710 and a memory system, such as optional memory system 715 shown in FIG. 7 . However, the control system 710 may include a memory system in some cases. In some implementations, the interface system 705 may be configured to receive input from one or more microphones in the environment.
[0106] The control system 710 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0107] In some implementations, control system 710 may reside on more than one device. For example, in some implementations, a portion of control system 710 may reside on a device within one of the environments described herein, while another portion of control system 710 may reside on a device outside the environment, such as a server, a mobile device (e.g., a smartphone or tablet computer), or the like. In other embodiments, a portion of control system 710 may reside on a device within one environment, while another portion of control system 710 may reside on one or more other devices in the environment. For example, a portion of control system 710 may reside on a device implementing a cloud-based service, such as a server, while another portion of control system 710 may reside on other devices implementing the cloud-based service, such as other servers, memory devices, or the like. Interface system 705 may also reside on more than one device, in some examples.
[0108] In some implementations, the control system 710 may be configured to perform, at least in part, the methods disclosed herein. According to some examples, the control system 710 may be configured to implement methods such as determining gain parameters, applying a gain transition function, determining an inverse gain transition function, applying the inverse gain transition function, allocating bits for gain control with respect to a bitstream, etc.
[0109] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, in any memory system 715 and / or control system 710 shown in FIG. 7. Accordingly, various inventive aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media having software stored thereon. The software may include, for example, instructions for determining gain parameters, instructions for applying a gain transition function, instructions for determining an inverse gain transition function, instructions for applying the inverse gain transition function, instructions for allocating bits for gain control on a bitstream, etc. The software may be executable by one or more components of a control system, such as, for example, control system 710 of FIG. 7.
[0110] In some examples, the device 700 may include an optional microphone system 720 shown in FIG. 7 . The optional microphone system 720 may include one or more microphones. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker of a speaker system, a smart audio device, or the like. In some examples, the device 700 may not include the microphone system 720. However, in some such implementations, the device 700 may still be configured to receive microphone data for one or more microphones in the audio environment via the interface system 710. In some such implementations, a cloud-based implementation of the device 700 may be configured to receive microphone data, or noise metrics corresponding at least in part to the microphone data, from one or more microphones in the audio environment via the interface system 710.
[0111] According to some implementations, device 700 may include an optional loudspeaker system 725 shown in FIG. 7. Optional loudspeaker system 725 may include one or more loudspeakers, which may also be referred to herein as a "speaker" or more generally as an "audio reproduction transducer." In some examples, e.g., cloud-based implementations, device 700 may not include loudspeaker system 725. In some implementations, device 700 may include headphones. Headphones may be connected or coupled to device 700 via a headphone jack or via a wireless connection, e.g., BLUETOOTH®.
[0112] 8A and 8B show an exemplary implementation of perceptually motivated gain control where the uniform gain control of samples at the encoder side is DBSTEP=-1 dB. In this particular example, one frame consists of 1024 samples. The sample amplitudes are represented by dotted lines, and the gain applied per sample is represented by solid lines. As can be seen from FIG. 8A, as soon as a frame creates an overload in the encoder (amplitude greater than 0 dB), the gain function transitions from no attenuation (0 dB) to an attenuation of DBSTEP-0 dB=-1 dB. If the input audio signal exceeds 1 dB, further attenuation by DBSTEP is introduced.
[0113] The resulting attenuated downmixed audio signal is shown in Figure 8B. In this example, the value of DBSTEP is large enough so that each sample is attenuated below the required threshold (0 dB).
[0114] Figures 9A and 9B show an example of "non-uniform" gain control where DBS = -1 dB and GT = {0, 1, 3, 6}, resulting in a DBSTEP set of attenuation values {DBS * GT} = {0, -1, -3, -6} dB or {-1, -2, -3} at the encoder side. As in Figures 8A and 8B, the sample amplitudes are shown as dotted lines, and the gain function is shown as a solid line. With the DBSTEP set, the automatic gain control can react to overloads caused in the encoder by attenuating the signal by increasing amounts each frame. As shown in Figure 9B, the gain transition step size is not large enough to ensure that all samples fall below the required threshold (<0 dB). This may result in distortion when the audio signal is rendered at the decoder, but the distortion caused by overloads in the encoder is less noticeable than the distortion caused by very sudden gain changes.
[0115] Some aspects of the present disclosure include systems or devices, e.g., programmed systems or devices, configured to perform one or more examples of the disclosed methods, and tangible computer-readable media, e.g., disks, storing code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system that includes input devices, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to asserted data.
[0116] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on audio signal(s), including performing one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or elements thereof) may be implemented as a general-purpose processor, e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory, and which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements. The other elements may include one or more loudspeakers and / or one or more microphones. A general-purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device. Examples of input devices include, for example, a mouse and / or a keyboard. The general-purpose processor may be coupled to memory, a display device, etc.
[0117] Another aspect of the present disclosure is a computer-readable medium, such as a disk or other tangible storage medium, that stores code for performing, e.g., by an executable coder, one or more examples of the disclosed methods or steps thereof.
[0118] While particular embodiments of and applications of the present disclosure have been described herein, it will be apparent to those skilled in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the present disclosure as described and claimed herein. While particular forms of the present disclosure have been illustrated and described, it should be understood that the disclosure should not be limited to the specific embodiments described and illustrated, or to the particular manner described.
[0119] Various aspects and implementations of the present disclosure can also be understood from the following enumerated exemplary embodiments (EEE), which are not claims. [EEE1] 1. A method for providing gain control to an audio signal, the method comprising: obtaining a downmixed audio signal of the audio signal to be encoded; determining that an overload condition occurs for a frame of the downmixed audio signal; responsive to determining that the overload condition has occurred, determining a gain transition function for the frame, the gain transition function based at least on a gain transition step size; and applying the gain transition function to the frame to generate a gain-adjusted frame of the downmixed audio signal; providing the gain adjusted frame and information indicative of the gain transition function for encoding by an encoder; A method comprising: [EEE2] encoding the gain adjusted frame together with the information indicative of the gain transition function; The method of EEE1 further comprising: [EEE3] Obtaining a downmixed audio signal of an audio signal to be encoded comprises: receiving the downmixed audio signal; or determining the downmixed audio signal from the audio signal to be encoded. The method described in EEE1 or 2. [EEE4] The method according to any one of EEE1 to 3, wherein the audio signal is a Higher Order Ambisonics (HOA) audio signal. [EEE5] 5. The method according to any of claims 1 to 4, wherein the downmixed audio signal is a spatially coded downmixed signal. [EEE6] The method according to any of EEE1 to 5, wherein the overload condition is a condition in which the frames of the downmixed audio signal exceed a predefined signal range. [EEE7] The method according to EEE6, wherein the predefined signal range is a signal range expected by the encoder. [EEE8] The method of any one of EEE1 to 7, wherein the frame of the downmixed audio signal is a current frame, and wherein the gain transition function is further based on a previous gain transition function applied to a frame preceding the current frame. [EEE9] 9. The method of any of EEE1 to 8, wherein the gain transition function is further dependent on a smoothing function that is based on the gain transition step size. [EEE10] The method of claim 8, wherein the gain transition function includes a transient portion and a steady-state portion, the transient portion corresponding to a transition from a gain associated with the previous frame to the gain associated with the previous frame adjusted by the gain transition step size. [EEE11] The method of claim 8, wherein the gain associated with the previous frame adjusted by the gain transition step size is either an attenuation of the gain corresponding to the previous frame by the gain transition step size or an amplification by the gain transition step size, depending on a gain adjustment target for the current frame. [EEE12] 12. The method of claim 10, wherein the length of the transient portion is limited by the delay introduced by the codec used by the encoder. [EEE13] The method according to EEE12, wherein the length of the transient portion is less than or equal to the number of samples used for a coding operation by the encoder. [EEE14] The method of any one of EEE10 to 13, wherein the length of the transient portion is greater than one sample. [EEE15] The gain transition function is
number
Claims
1. A method for performing gain control on an audio signal, Obtaining a downmixed audio signal of the audio signal to be encoded, It is determined that an overload condition has occurred with respect to the frame of the downmixed audio signal, In response to the determination that the aforementioned overload condition has occurred, a gain transition function for the frame is determined, wherein the gain transition function is based at least on the gain transition step size. Applying the aforementioned gain transition function to the frame to generate a gain-adjusted frame of the downmixed audio signal, For encoding by an encoder, information is provided indicating the gain-adjusted frame and the gain transition function. Methods that include...
2. An apparatus configured to implement the method described in claim 1.
3. A program that, when executed by a processing device, includes instructions causing the processing device to perform the method described in claim 1.
4. A computer-readable storage medium storing the program described in claim 3.