Method, apparatus and system for performing gain control of perceptual excitation

By determining the gain conversion function in the audio signal gain control and limiting the gain conversion step, the problem of severe gain variation between consecutive frames is solved, and smooth conversion and perceived quality is improved.

CN119998872APending Publication Date: 2025-05-13DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071339.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-22
Filing Date
2023-09-01
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the gain control of audio signals, it is difficult to effectively handle dramatic gain changes between consecutive frames, resulting in artifacts and perceived quality degradation.

Method used

By determining the gain conversion function, limiting the gain conversion step, smooth conversion from continuous gain is achieved and certain values ​​are allowed to exceed the predefined signal range to reduce artifacts.

Benefits of technology

Smooth conversion from continuous gain is achieved, reducing artifacts, improving the perceived quality of the audio signal, and reducing the number of bits required during the encoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998872A_ABST
    Figure CN119998872A_ABST
Patent Text Reader

Abstract

Systems, methods, and computer program products are provided for performing gain control on an audio signal. An automatic gain control system obtains a downmix audio signal of an audio signal to be encoded. The system determines that an overload condition has occurred in a frame of the downmix audio signal. In response to an overload condition, the system determines a gain conversion function for the frame, where the gain conversion function is based at least on a gain conversion step size. The system applies a gain conversion function to the frame to generate a gain-adjusted frame of the downmix audio signal. The system provides a gain-adjusted frame and information indicative of a gain conversion function for encoding by an encoder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 378,678, filed on October 6, 2022, and U.S. Provisional Application No. 63 / 503,533, filed on May 22, 2023, the entire contents of each of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to systems, methods, and media for adaptive gain control in an audio environment. Background Art

[0004] For example, gain control can be used to attenuate the signal to within the range expected by the audio codec. In order to improve the perceived quality of the audio signal to which gain control is applied at the encoder and inverse gain control is applied at the decoder, a gain conversion function that smoothly switches between different gains applied to consecutive frames has been proposed. If there is a dramatic gain change between consecutive frames, the method may cause audible artifacts. In addition, in some cases, the gain change between the determined gains of consecutive frames is too large and / or too abrupt to apply a smooth conversion function. In this case, a hard conversion can be used to ensure that the signal is within the expected range. For example, a single bit can be used to convey the information that a hard conversion is used between the gains of consecutive frames. However, this hard conversion may also result in audible artifacts that are worse than the artifacts introduced by the original overload condition in the decoded and reproduced audio signal. Therefore, there is a need to use a gain conversion function to improve the perceived quality of the encoding / decoding system and reduce the bits required for encoding.

[0005] Notation and nomenclature

[0006] Throughout this disclosure, including in the claims, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used synonymously to refer to any transducer or group of transducers that emit sound. A typical set of headphones includes two speakers. Speakers may be implemented to include multiple transducers, such as a woofer and a tweeter, which may be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuit branches coupled to different transducers.

[0007] Throughout this disclosure, including in the claims, statements that an operation is performed "on" a signal or data, such as filtering, scaling, transforming, or applying a gain to the signal or data, are used in a broad sense to mean performing the operation directly on the signal or data or on a processed version of the signal or data. For example, the operation may be performed on a version of the signal that has been initially filtered or preprocessed before the operation is performed on the signal.

[0008] Throughout this disclosure, including in the claims, the expression "system" is used in a broad sense to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M inputs and receives other XM inputs from external sources) may also be referred to as a decoder system.

[0009] Throughout this disclosure, including in the claims, the term "processor" is used in a broad sense to refer to a system or device that is programmable (e.g., using software or firmware) or otherwise configurable to perform operations on data, which may include audio or video or other image data. Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors that are programmed and / or otherwise configured to perform pipeline processing of audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets. Summary of the invention

[0010] In view of this, the present disclosure provides a method, an apparatus, a program, and a computer-readable storage medium for improving automatic gain control, which have the features of the corresponding independent claims.

[0011] According to one aspect of the present disclosure, a method for performing gain control on an audio signal is provided. The audio signal may be a higher order ambisonics (HOA) audio signal. In the method, a downmix audio signal of an audio signal to be encoded may be obtained. Obtaining the audio signal may include receiving the downmix audio signal. Alternatively, it may include determining the downmix audio signal from the audio signal to be encoded. In addition, it may be determined that an overload condition has occurred for a frame of the downmix audio signal. The overload condition may be a condition in which the frame of the downmix audio signal exceeds a predefined signal range. The predefined signal range may be a signal range expected by an encoder. The encoder may be a core encoder. In response to determining that an overload condition has occurred, a gain conversion function for the frame may be determined. The gain conversion function may be based at least on a gain conversion step size. The gain conversion function may be applied to the frame to generate a gain-adjusted frame of the downmix audio signal. The gain-adjusted frame may be an attenuated frame or an amplified frame. The gain-adjusted frame and information indicating the gain conversion function may be provided for encoding by the encoder.

[0012] By limiting the gain transition function to a gain transition step size, a smooth and less abrupt transition from a continuous gain can be achieved. The gain transition step size may not be sufficient to attenuate all samples of a frame to the signal range required by the core encoder. However, artifacts due to small overshoots are less noticeable than very abrupt increases or decreases in the gain parameter. Therefore, by allowing certain values ​​to be outside the required signal range, an improved audio experience can be achieved when decoding, rendering, and playing back the signal.

[0013] In some embodiments, the gain adjusted frame may be encoded with information indicating a gain transfer function.

[0014] In some embodiments, the downmix audio signal may be a spatially coded downmix signal.

[0015] In some embodiments, the frame of the downmix audio signal may be a current frame and the gain transfer function is further based on a previous gain transfer function applied to a frame preceding the current frame.

[0016] In some embodiments, the gain transition function may further rely on a smoothing function based on the gain transition step size.

[0017] In some embodiments, the gain transfer function may include a transient portion and a steady-state portion. The transient portion may correspond to a transition from a gain associated with a previous frame to a gain associated with the previous frame adjusted by a gain transfer step size.

[0018] In some embodiments, according to the gain adjustment target of the current frame, the gain associated with the previous frame adjusted by the gain conversion step may be an attenuation of the gain associated with the previous frame by the gain conversion step or an amplification of the gain associated with the previous frame by the gain conversion step.

[0019] In some embodiments, the length of the transient portion may be limited by the delay introduced by the codec used by the encoder and decoder.

[0020] Therefore, the gain control does introduce essentially zero additional delay.

[0021] In some embodiments, the length of the transient portion may be equal to or less than the number of samples used by the encoder for the encoding operation.

[0022] In some embodiments, the gain transfer function may be defined as

[0023]

[0024] Where DBSTEP is the gain conversion step size, l is the sample index, j is the frame index, p() is the smoothing function, l end denotes the rightmost index for which p() is defined, and L is the number of samples in a frame.

[0025] In some embodiments, the gain transition step size may be a predefined value or may be determined from a set of predefined values ​​increasing in size. The predefined value or the set of predefined values ​​may be determined based on a perceptual quality listening test or an objective quality measurement test. The perceptual quality listening test may be a multiple stimulus test with hidden reference and anchor (MUSHRA). The perceptual quality listening test may be part of the tuning process of the automatic gain control at the encoder and decoder.

[0026] In some embodiments, the method may further comprise determining an amount of overload caused by the frame of the downmix audio signal.Furthermore, the gain transition step size may be determined from a set of predefined values ​​increasing in size depending on the amount of overload.

[0027] Thus, the gain switching step size can be adapted to the desired rate of change between consecutive frames.

[0028] In some embodiments, applying the gain transfer function to the frame to generate the gain adjusted frame of the downmix signal may include applying the gain transfer function to samples of the downmix audio signal.The total number of samples may correspond to the frame of the downmix audio signal.

[0029] In some embodiments, encoding the gain-adjusted frame together with information indicating the gain transfer function may include determining a coding scheme based on the gain transfer function. In some cases, the coding scheme may be determined based on the gain transfer step size. In some cases, the coding scheme may be determined based on whether the overload condition has been eliminated. The coding scheme may be one of modified discrete cosine transform (MDCT) or algebraic code excited linear prediction (ACELP).

[0030] Thereby, the coding scheme can be optimized for a specific audio signal and the required gain transition step size.

[0031] According to another aspect, a method for performing gain control on an audio signal is provided. In the method, a coded frame of an audio signal may be received by a decoder. The coded frame of the audio signal may be decoded to obtain a frame of a downmixed audio signal and information indicating gain control applied by an encoder. An inverse gain conversion function to be applied to the frame of the downmixed audio signal may be determined based at least in part on the information indicating gain control applied by the encoder. The information indicating gain control applied by the encoder may include a gain conversion step size. The inverse gain conversion function may be applied to the frame of the downmixed audio signal.

[0032] In some embodiments, the method may further comprise upmixing the downmixed audio signal to generate an upmixed audio signal.The upmixed audio signal may be suitable for rendering.

[0033] In some embodiments, the method may further include rendering the upmix signal to produce rendered audio data.

[0034] In some embodiments, the method may further include playing back the rendered audio data using one or more of a loudspeaker or headphones.

[0035] In some embodiments, the inverse gain transfer function may be determined by inverting the gain transfer function applied by the encoder.

[0036] In some embodiments, the inverse gain transfer function may include a transient portion and a steady-state portion.

[0037] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transient media. Such non-transient media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Therefore, some innovative aspects of the subject matter described in this disclosure may be implemented by one or more non-transient media having software stored thereon.

[0038] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of at least partially performing the methods disclosed herein. In some implementations, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or a combination thereof.

[0039] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages will become apparent from the specification, drawings, and claims. Please note that the relative sizes of the following drawings may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a schematic block diagram of a system for providing gain control of an audio signal in the prior art.

[0041] Figure 2A and 2B is a schematic block diagram of a system for implementing adaptive gain control according to some embodiments.

[0042] Figure 3A and 3B Examples of a gain conversion function implementable by an encoder and an inverse gain conversion function implementable by a decoder, respectively, are shown according to some embodiments.

[0043] Figure 4 is a flow chart of an example process that may be performed by an encoder to implement adaptive gain control in accordance with some embodiments.

[0044] Figure 5 is a flow chart of an example process that may be performed by a decoder to implement adaptive gain control according to some embodiments.

[0045] Figure 6 An example use case of an immersive voice and services (IVAS) system in accordance with some embodiments is shown.

[0046] Figure 7 A block diagram illustrating an example of components of an apparatus capable of implementing various aspects of the present disclosure is shown.

[0047] Fig. 8A and 8B An example embodiment of an audio codec utilizing perceptually motivated gain control of a downmix signal is shown, wherein the gain transition step size is uniform.

[0048] Fig. 9A and 9B An example embodiment of an audio codec utilizing perceptually motivated gain control of a downmix signal is shown, wherein the gain transition step size is non-uniform.

[0049] Like reference numbers and designations in the various drawings represent like elements. DETAILED DESCRIPTION

[0050] Some coding techniques for scene-based audio, stereo audio, multi-channel audio and / or object audio rely on encoding the multi-component signal after a downmix operation. Downmixing can allow a reduced number of audio components to be encoded in a waveform encoding manner that retains the waveform, and the remaining components can be parameter encoded. On the receiver side, parameterized metadata indicating parameter encoding can be used to reconstruct the remaining components. Because only a subset of the components are waveform encoded, and the parameter metadata associated with the parameter-encoded components can be efficiently encoded in terms of bit rate, such coding techniques can be relatively bit rate efficient while still allowing high-quality audio.

[0051] One problem that may arise is that the downmix channels determined by the spatial encoder may include signals with levels that are not suitable for subsequent processing by the core codec that constructs the audio signal bitstream. For example, in some cases, the downmix signal may have a level that is too high, so that the core codec is overloaded although the original input signal is not overloaded in any of its component signals. This may lead to severe distortions, such as clipping in the reconstructed signal after decoding and rendering. This may cause considerable quality loss in the final rendered signal. One possible solution may be to attenuate the input signal to avoid overloading the core codec. However, this solution may have the disadvantage of increasing granular noise, because the quantizer used to encode the signal may not operate in the optimal range.

[0052] Figure 1 A schematic block diagram of a conventional system 100 for performing gain control on an encoded higher-order ambisonic (HOA) signal is shown. Figure 1 The schematic block diagram shown can be used to encode and decode MPEG-H signals. MPEG-H is a set of international standards being developed by the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) Moving Picture Experts Group (MPEG). MPEG-H consists of various parts, including Part 3, MPEG-H 3D Audio.

[0053] At the encoder 102, the input HOA signal is processed at 104. The processing may include, for example, decomposition, where downmix channels are generated. The downmix channels may include a set of signals constrained by [-max, max] for a given frame. Because the core encoder 108 can encode signals within the range of [-1, 1), samples of signals associated with downmix channels that are beyond the range of the core encoder 108 may cause overload. To avoid overload, the gain control 106 adjusts the gain of the frame so that the associated signal is within the range of the core encoder 108 (e.g., within [-1, 1)). The core encoder 108 can be considered as a codec that generates an encoded bitstream. The side information generated by the decomposition / processing block 104 (which may include metadata associated with the parametrically encoded channels, etc.) can be encoded in the bitstream in conjunction with the signal generated as an output of the core encoder 108.

[0054] The decoder 112 receives the encoded bitstream. The decoder 112 can extract the side information, and the core decoder 116 can extract the downmix signal. Then, the inverse gain control block 120 can reverse the gain applied by the encoder. For example, the inverse gain control block 120 can amplify the signal attenuated by the gain control 106 of the encoder 102. Then, the HOA signal can be reconstructed by the HOA reconstruction block 122. Optionally, the HOA signal can be rendered and / or replayed by the rendering / playback block 124. The rendering / playback block 124 may include, for example, various algorithms for rendering the reconstructed HOA output as, for example, rendered audio data. For example, rendering the reconstructed HOA output may involve distributing one or more signals of the HOA output to multiple speakers to achieve a specific perceptual impression. Optionally, the rendering / playback block 124 may include one or more loudspeakers, headphones, etc. for presenting the rendered audio data.

[0055] The gain control 106 may use the following techniques to implement gain control. The gain control 106 may first determine an upper limit on the signal value in the frame. For example, for an MPEG-H audio signal, the upper limit may be expressed as the product Where the product is specified in the MPEG-H standard. Given the upper limit, the minimum attenuation required can ensure that the scaled signal samples are confined to the interval [-1, 1). In other words, the scaled samples can be within the range of the core encoder 108. This can be achieved by applying The gain factor is determined by By definition, e min Can be a negative number. In some embodiments, the amplification can be subject to a maximum amplification factor The limit of e max is a non-negative integer. Therefore, in order to perform both attenuation and amplification, we can define 2 e The gain factor, where the gain parameter e is [e min ,e max ] range. Therefore, the minimum number of bits required to represent the gain parameter e is determined to be β e =ceil(log 2 (|e min |+e max +1)).

[0056] The gain factor g for a particular channel n and frame j can be determined by applying a frame delay corresponding to one HOA block and using the following recursive operation n (j):

[0057]

[0058] In the above formula, g n(j-2) represents the gain factor applied to frame (j-2), and Indicates the gain factor g for calculating frame j-1 n (j-1) The required gain factor adjustment.

[0059] Techniques for providing adaptive gain control are disclosed herein. Specifically, as described herein, a gain parameter can be determined that does not generate additional delay because the gain parameter can be determined based on a look-ahead sample generated for use by a codec. The codec can be used by a perceptual encoder. The determination of the gain transfer function is shown and described below in conjunction with Figures 2-5.

[0060] Figure 2A and 2B Schematic block diagrams of an encoder 202 and a decoder 212 for performing low-delay adaptive gain control according to exemplary embodiments are shown respectively. At the encoder 202, an input HOA signal (or first-order ambisonics (FOA)) signal is processed by a spatial analysis block 204. For an N-channel HOA input, the spatial analysis block 204 may generate and output a set of M down-mix channels 204A. The number of down-mix channels in the set of M down-mix channels 204A may be in the range of 1≤M≤N. In addition, the spatial analysis block 204 may generate and output spatial side information 204B for reversing the down-mix operation.

[0061] For example, for FOA input, the downmix channels may include a main downmix channel W', which may be generated by mixing the omnidirectional input signal W with the directional input signals X, Y, and Z using various mixing gains and up to 3 residual channels X', Y', and Z', each residual channel corresponding to a signal component in the X, Y, and Z signals that cannot be predicted from the main downmix signal. In one example, the spatial analysis block 204 utilizes a spatial reconstruction (SPAR) technique. SPAR is further described below, which is incorporated herein by reference in its entirety: D. McGrath, S. Bruhn, H. Purnhagen, M. Eckert, J. Torres, S. Brown, and D. Darcy, "Immersive Audio Coding for Virtual Reality Using a Metadata-assisted Extension of the 3GPP EVS Codec," IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2019, pp. 730-734. In other examples, the spatial analysis block 204 may utilize any other suitable linear prediction codec of an energy compression transform, such as the Karl Huning-Löf transform (KLT), etc. The core encoder 208 may be considered as a codec that generates the encoded audio bitstream 208A. In some implementations, the core encoder 208 and the core decoder 216 may introduce some look-ahead samples that will be used by the adaptive gain control 206 to determine the gain parameters to avoid adding additional delay to the overall encoding process (zero additional delay).

[0062] Then, the signals associated with the M downmix channels 204A may be analyzed by the adaptive gain control 206. The adaptive gain control 206 may determine whether the signals associated with any of the M downmix channels 204A exceed the audio amplitude range expected by the core encoder 208 and will therefore overload the core encoder 208. In some embodiments, in the case where the adaptive gain control 206 determines that the gain will not be applied, for example, in response to determining that none of the signals of the M downmix channels 204A exceeds the expected range of the core encoder 208, the adaptive gain control 206 may set a flag indicating that the gain control is not applied. The flag indication may be performed by setting a value of the flag, for example, by setting a value of a single bit. In the case where the adaptive gain control 206 determines that the gain will not be applied, the adaptive gain control 206 may not set the flag, thereby reserving one bit (e.g., a bit associated with the flag). For example, in some implementations, if the spatial metadata bitstream and / or the core encoder bitstream (which may be a perceptual encoder bitstream) is self-terminating, the presence of the gain control flag may be determined by determining whether there are any unread bits in the bitstream. The unread bits may be the remaining bits in the bitstream. In the absence of an overload condition, the adaptive gain control 206 may output M downmix channels 206A. The M downmix channels 206A may then be passed to the core encoder 208 for encoding in the bitstream 208A.

[0063] On the contrary, in the case where the adaptive gain control 206 determines that a gain is to be applied, the adaptive gain control 206 can determine a gain parameter and apply (multiple) gains to the M down-mixed channels according to the determined gain parameter. Then, the M down-mixed channels 206A to which the gain is applied can be passed to the core encoder 208 to be encoded in the bitstream. In addition, the adaptive gain control 206 can output side information 206B about the gain control. Information about the flag can be included in the side information 206B about the gain control. The side information encoder 210 can encode the spatial side information 204B together with the gain parameter 206B as metadata 210A for transmission in the bitstream. Then, the decoder 212 can extract and use the metadata to mix up the down-mixed channels and reverse the gain adjustment. For example, the metadata 210A can be used later to reconstruct a representation of the original audio input down-mixed by the spatial analysis unit 204. The side information encoder 210 can additionally provide the core encoder 208 with side information 208B. The core encoder 208 may then use the side information 208B to select between encoding techniques. Both the encoded bitstream 208A and the encoded bitstream with metadata 210A may be multiplexed to form a final bitstream output by the encoder 202 .

[0064] In some implementations, the adaptive gain control 206 may determine a gain transition function for transitioning between a gain parameter e(j - 1) associated with a previous frame (e.g., the j - 1th frame) and a gain parameter e(j) of the current frame. The gain transition function may be applied by the adaptive gain control 206 on a frame - by - frame basis, where each frame may be a frame of one of the M down - mixed channels 204A. In some implementations, the gain transition function may smoothly transition the gain parameter from the gain parameter value of the j - 1th frame (e.g., e(j - 1)) to the gain parameter of the current frame (e.g., e(j)) across the samples of the jth frame. Thus, the gain transition function may include two parts: 1) an instantaneous part, where the gain parameter transitions from the gain parameter of the previous frame to the gain parameter of the current frame across the samples of the instantaneous part; and 2) a steady - state part, where the gain parameter has the gain parameter value of the current frame for the samples of the steady - state part.

[0065] In some embodiments, in the case where the gain applied to the current frame is less than the gain applied to the previous frame, since the attenuation amount increases across the samples of the current frame, the instantaneous part may be referred to as an instantaneous type with "fading". The case where the gain applied to the current frame is less than the gain applied to the previous frame may be expressed as e(j)>e(j - 1). In some embodiments, in the case where the gain applied to the current frame is greater than the gain applied to the previous frame, since the attenuation amount decreases across the samples of the current frame, the instantaneous part may be referred to as an instantaneous type with "inverse fading" or "non - fading". The case where the gain applied to the current frame is greater than the gain applied to the previous frame may be expressed as e(j)<e(j - 1). In some embodiments, in the case where the gain applied to the current frame is the same as the gain applied to the current frame, the instantaneous part may be referred to as an instantaneous type with "hold", where the instantaneous part is not instantaneous but has the same value as the steady - state part. The case where the gain applied to the current frame is the same as the gain applied to the current frame may be expressed as e(j)=e(j - 1).

[0066] In some embodiments, the gain conversion function depends on the gain conversion step size. The gain conversion step size can limit the amount of possible conversion from the previous frame to the current frame. This is motivated by the fact that smaller and smoother gain / attenuation changes that may allow overload during conversion are better than larger changes in perception, especially when this is further processed by a lossy core encoder that requires the predefined value range mentioned as input. By predefining the parameters of the gain conversion function in this way, the impact of the parameters on objective quality or perceived quality can be evaluated. Perceptual quality can be measured based on known perceptual quality listening tests, such as the multi-stimulus test (MUSHRA) with hidden references and anchors. The perceptual quality listening test can be part of the tuning process of the automatic gain control at the encoder and decoder. Specifically, parameters such as the gain conversion step size can be tuned for specific audio scenes and codecs until the best perceived audio quality is reached. Then, the encoding / decoding system uses the tuned parameters.

[0067] In an example implementation, the processed output 206A of the automatic gain control 206 is further encoded by a lossy core codec based on algebraic code excited linear prediction (ACELP) coding that is not targeted for waveform reconstruction. It has been observed that applying larger gain steps to the ACELP input and output results in audible glitches in the reconstructed signal and degrades the overall performance of the codec.

[0068] In some implementations, when an overload of the current frame is detected, the automatic gain control 206 may also determine the amount of attenuation required for the frame to be within the expected range of the core encoder 208. If there is a large difference between the attenuation required between consecutive frames, the application of the conversion function by the core encoder 208 to achieve the desired range [-1, 1) may result in audible artifacts when the audio signal is rendered at the decoder. Instead of applying the conversion function to keep each frame within or at the limits of the desired range, the conversion function may be limited to a specific gain conversion step size. Therefore, regardless of the amount of attenuation required to achieve the expected range of the core encoder 208, the conversion function can only attenuate a single frame by an amount equal to the gain conversion step size, i.e., ± DBSTEPdB. Therefore, as an example, if the attenuation of the previous frame is -10dB, the attenuation applied to the first sample of the current frame will be -10dB, and the attenuation applied to the last sample of the current frame will be -10dB±DBSTEP. More precisely, if the amount of overload does not change from the previous frame to the current frame, the gain conversion will be a constant value, e.g., -10dB. If the attenuation needs to be changed, the gain transfer function will be converted by ±DBSTEPdB from the attenuation of the previous frame to the last sample of the current frame.

[0069] In some implementations, DBSTEP can be selected so that the attenuation applied by the automatic gain control 206 is not enough to keep the frame within the expected signal range of the core encoder 208. For example, DBSTEP can be a fixed value. When a sharp change in attenuation is required, by allowing the frame to be outside the range of [-1,1), the strong attenuation difference between consecutive frames can be avoided. Therefore, the attenuation of the frame relative to the previous frame is attenuated by a fixed amount, rather than forcing the frame to be within the range of [-1,1) by a transfer function or static gain change. By using a transfer function with a specific gain conversion step, the perceived audio quality can be improved because the distortion caused by the frame being outside the range of [-1,1) is less obvious than the distortion caused by the sharp attenuation difference between consecutive frames. In addition, the exception flag for switching between smooth conversion and static gain change can be avoided. Therefore, 1 bit can be saved at the core encoder 208.

[0070] In some implementations, DBSTEP may be a single value, e.g., -1 dB. Alternatively, DBSTEP may be selected from a set of increasing fixed values, e.g., -1 dB, -3 dB, -6 dB. In this case, the value of DBSTEP may be selected based on the amount of overload caused by frames without attenuation.

[0071] In some implementations, the automatic gain control is configured to have a specified set of target gain values ​​G T The target gain value G T The absolute step sizes may be represented as a table of numbers (e.g., integers) indicating the multiple of the DBS attenuation provided at each step. This is motivated by the fact that smaller changes provide a perceptual benefit, however, some signals may require a higher level of possible attenuation. Specifying these non-uniform absolute step sizes allows a wider attenuation range to be covered while providing the benefits of smaller step sizes for many possible situations. For example, a set of G with -2 dB of DBS T = {0,1,3,6} will have {DBS*G T}={0dB,-2dB,-6dB,-12dB} absolute target gain, take {-2dB,-4dB,-6dB} continuous step size DBSTEP.

[0072] One or more of such integer tables may be specified, and information about the selection of a particular table used at the encoding side may be signaled / sent to the decoder side. In contrast to the application of uniform step sizes that produces a single uniform gain transition shape, the application of non-uniform step sizes results in non-uniform gain transition shapes (level-dependent transfer functions).

[0073] In some implementations, when DBSTEP is insufficient to attenuate the current frame to the range of [-1, 1), DBSTEP may be applied to frames following the current frame until the range of [-1, 1) is reached.

[0074] In some implementations, the output level and attenuation information from the automatic gain control 206 system can be used in the decision making process in other systems such as the core encoder 208. Although the relaxed requirements can provide perceptual benefits, it can affect the core encoder 208 by introducing changes in gain or by not meeting the strict requirements and allowing the overload condition to remain. Information such as whether the gain control meets the requirements, whether or how much gain is applied, etc. can be output and passed to the core encoder. This allows better decisions to be made, such as selecting an encoding method that can better handle gain changes or out-of-range samples. For example, when a large gain / attenuation step size is applied, the core encoder 208 can use a waveform coding technique similar to MDCT-based coding instead of a predictive ACELP coding technique.

[0075] In some embodiments, the instantaneous portion of the gain transfer function can be determined using a prototype shape of the instantaneous portion of the gain transfer function, wherein the prototype shape is scaled based on the difference between the gain parameters of the current frame and the gain parameters of the previous frame. For example, the prototype shape can be scaled based on e(j)-e(j-1). The gain transfer function using this prototype function p can be expressed as:

[0076]

[0077] Among them, l end represents the rightmost index for which p is defined, and L represents the number of samples in a frame. For example, the prototype shape of the instantaneous partial gain can be defined as:

[0078]

[0079] in and L here is the number of samples in the frame for which p is defined. For example, L can be l end +1.

[0080] Figure 3A Examples of gain transfer functions are shown in , each having a transient part with a transient type of "fading". Figure 3A In the example shown, each gain transfer function has a transient portion starting from sample 0, which may correspond to the beginning of the current frame, with a gain of 0 dB, where 0 dB is the gain parameter of the previous frame (e.g., the j-1th frame). Figure 3AIn the example shown, the transient portion of each gain transfer function changes to the steady-state portion of the gain transfer function over the course of approximately 384 samples. Figure 3A For each of the three gain conversion functions shown in , the steady-state portion corresponds to a different gain conversion step size for the jth frame, with the (negative) gain increasing by 6dB, 12dB, and 18dB, respectively, relative to the gain of the previous frame. In other words, Figure 3A As shown, for the three gain conversion functions, exp = -[e(j) - e(j-1)] = -1, -2, and -3 respectively. Figure 3A For each gain transfer function shown in , the transient portion has the same length (e.g., about 384 samples). Note that the length of the steady-state portion may correspond to an offset associated with the delay introduced by the codec, e.g., Figure 3A In the example shown, this is 12 milliseconds. Accordingly, the length of the transient portion can be related to the inverse of the offset. Figure 3A In the example shown, the length of the transient portion is the frame length (eg, 20 milliseconds) minus the codec delay (eg, 12 milliseconds). Note that the codec delay may be the overall encoder algorithm delay excluding the frame size delay.

[0081] In addition, the gain transfer function of the transient part with the transient type of "inverse fading" or "non-fading" can be expressed as Figure 3A The gain transfer function shown in FIG. 1 is a mirror image flipped across a horizontal line. For example, the horizontal line can be the x-axis.

[0082] Return to reference Figure 2B, the decoder 212 may receive the encoded audio bitstream 208A and the metadata bitstream 210A as inputs, and may reconstruct the HOA signal for, for example, rendering, or directly rendering to a desired output format. In some embodiments, the core decoder 216 receives the encoded audio bitstream 208A. In addition, the core decoder 216 may receive information 214A extracted from the metadata bitstream 210A by the side information decoder 214. The core decoder 216 may decode the encoded audio bitstream 208A based on the information 214A or without any knowledge of the side information, and output M gain-adjusted downmix channels 216A to the inverse gain control 220. The side information decoder 214 further extracts the gain parameters and the spatial side information, and sends the information 214B to the inverse gain control 220 and the spatial synthesis / rendering / playback block 222. Then, the inverse gain control 220 may obtain the gain parameters applied by the encoder 202 from the information 214B. For example, in some implementations, inverse gain control 220 can retrieve the gain conversion step DBSTEP applied by encoder 202 and / or the indication of the arithmetic factor associated with DBSTEP from information 214B. In addition, inverse gain control block 220 can, for example, retrieve the shape of the conversion function from memory, i.e., the shape of the prototype function p also referred to as a smooth function. Then, inverse gain control block 220 can use the obtained gain parameter to reverse the gain applied by encoder 202, and output M down-mixed channels 220A. For example, in some implementations, inverse gain control 220 can construct the inverse gain conversion function that converts the gain parameter of the previous frame to the gain parameter of the current frame. In some implementations, the inverse gain conversion function can be the gain conversion function applied by encoder 202 that is mirrored and adjusted vertically across the center vertical line. For example, the vertical line can be the y-axis.

[0083] Go to Figure 3B , according to some implementations, the decoder will respond to Figure 3A An example of an inverse gain transfer function applied by an encoder is shown in FIG. As shown, the inverse gain transfer function has a steady-state portion and a transient portion. Figure 3A and 3B As shown, the duration of the steady-state portion and the transient portion of the inverse gain transfer function may correspond to, for example be the same as, the duration of the corresponding steady-state portion and the transient portion of the gain transfer function. Figure 3B Each inverse gain conversion function shown starts from 0 dB and converts to -DBSTEP of the current frame. That is, each inverse gain conversion function starts from 0 dB corresponding to the inverse gain applied one frame j-1 before. When the gain applied by the encoder corresponds to Figure 3AThe gain transfer function shows that when the attenuation is represented by a gain less than 0 dB, the inverse gain applied by the decoder corresponds to an amplification with a gain greater than 0 dB, as shown in Figure 3B Conversely, when the gain applied by the encoder corresponds to amplification (eg, having a gain greater than 0 dB), the inverse gain applied by the decoder corresponds to attenuation (eg, having a gain less than 0 dB).

[0084] Return to reference Figure 2B , after applying the inverse gain, the M downmix channels 220A to which the inverse gain is applied are provided to a spatial synthesis / rendering / playback block 222. The spatial synthesis / rendering / playback block 222 may reconstruct the HOA signal using the information 214B. For example, in the case where the spatial analysis block 204 uses SPAR technology for spatial encoding, the spatial synthesis / rendering / playback block 222 may utilize SPAR technology to reconstruct one or more channels encoded using metadata 210A. The reconstructed HOA output may then be rendered directly or provided to another entity for rendering. The spatial synthesis / rendering / playback block 222 may include, for example, various algorithms for rendering the reconstructed HOA output as, for example, rendered audio data. For example, rendering the reconstructed HOA output may involve distributing one or more signals of the HOA output to multiple speakers to achieve a particular perceptual impression. Optionally, the spatial synthesis / rendering / playback block 222 may include an audio playback device for presenting the rendered audio data, such as one or more loudspeakers, headphones, and the like.

[0085] Figure 4 An example of a process 400 for determining a gain parameter and applying a gain to a downmix signal according to the determined gain parameter according to some implementations is shown. In some implementations, the blocks of process 400 may be performed by an encoder device. In some implementations, the process may be performed according to Figure 4 The blocks of process 400 may be performed in an order other than shown. In some implementations, two or more blocks of process 400 may be performed substantially in parallel. In some implementations, one or more blocks of process 400 may be omitted.

[0086] At 402, process 400 may obtain (multiple) downmix audio signals associated with a frame of an audio signal to be encoded. (Multiple) downmix audio signals may be associated with a frame of an audio signal to be encoded. For example, in some implementations, process 400 may determine a set of downmix channels using any suitable spatial coding technique. Examples of spatial coding techniques include SPAR, linear prediction techniques, etc. The set of downmix channels may include any number from 1 to N channels, where N is the number of input channels, such as 4 in the case of a FOA signal. The downmix signal may include an audio signal of a downmix channel corresponding to a particular frame of the audio signal.

[0087] At 404, the process 400 can determine whether an overload condition exists for a codec such as an enhanced voice service (EVS) codec and / or for any other suitable codec. For example, in response to determining that a signal corresponding to a frame of the downmix audio signal(s) exceeds a predetermined range (e.g., [-1, 1)) and / or any other suitable range, the process 400 can determine that an overload condition exists.

[0088] If it is determined at 404 that there is no overload condition ("No" at 404), process 400 may proceed to 412 and may encode the downmix signal. For example, in some implementations, process 400 may generate a bitstream that encodes the downmix signal in conjunction with side information such as metadata, which a decoder may utilize to upmix the downmix signal, e.g., to reconstruct a FOA or HOA output.

[0089] On the contrary, if it is determined at 404 that there is an overload condition ("yes" at 404), the process 400 can proceed to 406, and a gain transfer function for a frame that results in avoiding the overload condition can be determined, or at least reducing the overload if the change in the overload condition from one frame to the next is greater than the gain transfer step size. In addition, in 406, the gain transfer function can be based on the gain transfer step size. In addition, the gain transfer function can be based on the shape of a smooth function. In addition, as described above in conjunction with FIG. 2, the gain transfer function can have a transient portion and a steady-state portion, wherein the steady-state portion corresponds to the gain factor of the current frame, and the transient portion corresponds to an intermediate gain factor sequence of a sample subset of the current frame, which intermediate gain factor sequence is converted from the gain factor at the end of the previous frame to the gain factor ± DBSTEP of the previous frame.

[0090] In the case where the gain parameter of the previous frame corresponds to less attenuation than the gain parameter of the current frame, the instantaneous portion may be referred to as having an instantaneous type of "fading". Conversely, in the case where the gain parameter of the previous frame corresponds to more attenuation than the gain parameter of the current frame, the instantaneous portion may be referred to as having an instantaneous type of "reverse fading" or "non-fading". In the case where the gain parameter of the previous frame is the same as the gain parameter of the current frame, the instantaneous portion may be referred to as having an instantaneous type of "holding". In the case where the instantaneous portion has an instantaneous type of "holding", the value of the gain transfer function during the instantaneous portion may be the same as the value of the gain transfer function during the steady-state portion. As described above in conjunction with FIG. 2 , the duration of the instantaneous portion of the gain transfer function may correspond to the delay duration utilized by the codec.

[0091] At 408, the process 400 may apply the gain transfer function to the downmix signal associated with the frame. For example, in some implementations, the process 400 may scale the samples of the downmix signal by a gain factor indicated by the gain transfer function. As a more specific example, in some implementations, the first sample of the current frame may be scaled by a gain factor corresponding to a gain parameter of the previous frame, the last sample of the current frame may be scaled by a gain factor corresponding to a gain parameter ± DBSTEP of the previous frame, and intermediate samples may be scaled by a gain factor corresponding to a gain parameter of a transient or steady-state portion of the gain transfer function.

[0092] In some implementations, the gain conversion function may be applied only to the downmixed signal of the downmixed channel for which an overload condition is detected at block 404. For example, in the case where an overload condition of the Y' channel and the X' channel is detected, a separate gain conversion function may be determined for each of the Y' channel and the X' channel and applied to the signal of the Y' channel and the X' channel. Continuing with this example, the gain conversion function may not be applied to the W' and Z' channels. In this case, for example, in block 412, an indication of the channel to which the gain conversion function is applied and the corresponding gain parameters of each channel may be encoded. Alternatively, in some implementations, in the case where an overload condition exists only for one downmixed channel, the corresponding gain conversion function may be applied to all downmixed channels. In this case, because the gain conversion function is applied to all channels, it is not necessary to send an indication of the channel to which the gain has been applied, which may result in an improvement in bit rate efficiency.

[0093] At 410, process 400 can provide the attenuated signal and the information indicating the gain conversion function to the encoder for encoding. The information indicating the gain conversion function can be the gain conversion step length and / or the arithmetic factor related to the gain conversion step length. In addition, the shape of the smoothing function can be provided to the encoder for encoding.

[0094] At 412, process 400 may encode the downmix signal and, if gain is applied, information indicating gain parameters for the frame is encoded. In the case where gain is applied, at block 408, the encoded downmix signal may be the downmix signal after applying the gain transfer function. The downmix signal and any information indicating the gain parameters may be encoded by a codec (e.g., an EVS codec, etc.) in association with any side information (e.g., metadata) that a decoder may use to reconstruct or upmix the downmix signal to generate an encoded bitstream. The encoded bitstream may then be stored and / or sent together with the metadata to a receiving device having the ability to reverse the processing steps of the encoder.

[0095] It should be noted that in some implementations, process 400 may encode the gain parameter as a set of bits.In some implementations, the gain transfer function may indicate a prototype / smooth function associated with the transient portion of the gain transfer function.

[0096] In the case where adaptive gain control is enabled for each channel so that a unique gain transfer function is applied to each downmix channel associated with the signal that triggers the overload condition, x bits may be used for each channel for which gain control is enabled, with an additional one bit indicator per channel indicating that a gain parameter has been encoded. In this case, the total number of bits used to send the gain control information is N dmx +x*N agc , where N dmx denotes the number of downmix channels (and where for N dmx For each channel, a single bit is used to indicate whether gain control is enabled), and where N agc Indicates the number of channels for which gain control has been enabled. It should be noted that in the case where gain control is not enabled for a particular frame, N dmx bits to indicate that gain control is not enabled, e.g., for N dmx Note that when the number of down-mix channels is 1, for example, only W channels are waveform coded, and the total number of bits used to send gain control information is x*N. agc For example, given a downmix channel, if gain control is not enabled for this downmix channel (e.g., N agc =0), the number of bits used is 0. Continuing with the example, if gain control is enabled (e.g., N agc =1), then the number of bits used is x.

[0097] In case a single gain transfer function associated with the downmix channel triggering the overload condition is applied to all downmix channels, fewer bits may be used to transmit the gain control information. For example, a single gain parameter for the current frame is transmitted using x bits.

[0098] Figure 5 An example of a process 500 for obtaining a gain parameter used by an encoder and applying an inverse gain conversion function based on the obtained gain parameter according to some implementations is shown. In some implementations, the blocks of process 500 may be performed by a decoder device. In some implementations, the process 500 may be performed according to Figure 5 The blocks of process 500 may be performed in an order other than shown. In some implementations, two or more blocks of process 500 may be performed substantially in parallel. In some implementations, one or more blocks of process 500 may be omitted.

[0099] Process 500 may begin at 502 by receiving an encoded frame of an audio signal. A received frame (eg, a current frame) is generally referred to herein as a jth frame. The received frame may be immediately following a previously received frame, or may be a frame that does not immediately follow a previously received frame.

[0100] At 504, process 500 can decode the encoded frames of the audio signal to obtain a downmix signal, and if the encoder applies gain control, obtain information indicating the gain control applied to the current frame. The information indicating the gain control applied to the current frame can be a gain conversion step size applied by the encoder. In addition, the information indicating the gain control applied to the current frame can be the shape of a smooth function of the gain conversion function applied by the encoder. In the case where the encoder applies gain control on a per-channel basis, process 500 can also identify which downmix channels the gain control is applied to.

[0101] At 506, process 500 may determine an inverse gain transfer function based on the gain transfer step size. In some implementations, process 500 may also determine an inverse gain transfer function based on the shape of a smoothing function. The inverse gain transfer function may be calculated based on the gain transfer function, or may be selected from a plurality of predefined inverse gain transfer functions.

[0102] In some implementations, process 500 may determine the inverse gain transfer function as the inverse of the gain transfer function applied at the encoder. For example, the inverse gain transfer function may correspond to a gain transfer function that is mirrored and adjusted across a horizontal line. The mirroring and adjustment may be along the x-axis. Examples of such inverse gain transfer functions are described above in conjunction with Figure 3B Shown and described. In some implementations, the inverse gain conversion function may have a steady-state portion corresponding to the gain applied to the previous frame. Then, the inverse gain conversion function may have an instantaneous portion, which is the inverse of the instantaneous portion of the gain conversion function applied at the encoder. For example, in the case where the gain applied to the current frame corresponds to more attenuation relative to the previous frame, the inverse gain conversion function may have an instantaneous portion that is converted from a smaller amplification to a larger amplification. On the contrary, in the case where the gain applied to the current frame corresponds to a smaller attenuation relative to the previous frame, the inverse gain conversion function may have an instantaneous portion that is converted from a larger amplification to a smaller amplification. The duration of the instantaneous portion may be related to the delay introduced by the codec, wherein the duration of the instantaneous portion is the frame length (e.g., 20 milliseconds) minus the codec delay (e.g., 12 milliseconds). Note that in the case where the delay introduced by the codec is longer than the frame length, the delay of a frame may be used to apply the inverse gain conversion. In some cases, the delay may be obtained from the gain control bit by process 500 (e.g., by a decoder). The inverse gain conversion function may also be used to attenuate the signal amplified by the gain control of the encoder.

[0103] At 508, process 500 can apply an inverse gain conversion function to a downmix signal to reverse the gain applied by the encoder. For example, the application of the inverse gain conversion function can cause amplification of a downmix signal attenuated by the encoder to reverse the attenuation. As another example, the application of the inverse gain conversion function can cause attenuation of a downmix signal amplified by the encoder to reverse the amplification. Then, the output of step 508 can be M downmix channels with the same gain as the M downmix channels after step 402 of process 400.

[0104] At 510, process 500 may upmix the downmix signal. The upmix may be performed by a spatial encoder. In some examples, the spatial encoder may utilize SPAR technology. The upmix signal may correspond to a reconstructed FOA or HOA audio signal. In some implementations, process 500 may upmix the signal using side information (e.g., metadata) encoded in the bitstream, where the side information may be used to reconstruct the parameter-encoded signal. In some implementations, block 510 may be optional, for example, when the downmix signal may be rendered directly.

[0105] In some implementations, at 512, process 500 can render the upmixed signal to generate rendered audio data. In some implementations, process 500 can render the FOA or HOA audio signal using any suitable rendering algorithm, such as rendering scene-based audio data. In some implementations, the rendered audio data can be stored in any suitable format, such as for future presentation or playback. In some implementations, block 512 is optional and can be omitted.

[0106] In some implementations, process 500 may cause the rendered audio data to be played back at 514. For example, in some implementations, the rendered audio data may be presented via one or more loudspeakers and / or headphones. In some implementations, multiple loudspeakers may be utilized, and the multiple loudspeakers may be positioned in any suitable position or orientation relative to each other in three dimensions. In some implementations, process 514 is optional and may be omitted.

[0107] As above combined Figure 4 As described above, a set of gain control bits may be used to encode gain control information, such as information indicating gain parameters. In some implementations, a different gain conversion function may be determined for each downmix channel where an overload condition is detected. In such an implementation, a gain control bit is required to indicate whether gain control is applied to each downmix channel, and the gain conversion function parameters are encoded for each downmix channel to which gain control is applied, as described above in conjunction with Figure 4Alternatively, in some implementations, a single gain conversion function determined based on one downmix channel having an overload condition may be applied to all downmix channels. In such an implementation, fewer gain control bits are required because a separate bit flag is not required to indicate whether gain control has been applied to each downmix channel, resulting in more bit rate efficient encoding.

[0108] By applying the same gain conversion function to more efficient bit rate encoding of all downmix channels, including downmix channels that do not have an overload condition, a degradation in perceived quality may result, for example, by attenuating signals that do not have a codec overload. In contrast, using more targeted gain control (wherein gain control is applied to each downmix channel in a targeted manner) may require more bits to send gain control information. However, using additional bits to send targeted (e.g., channel-specific) gain control information may require reallocation of bits that are typically used to waveform encode the downmix channels, which may reduce perceived quality in some cases. Therefore, there may be a trade-off between applying the same gain conversion function to all downmix channels and applying channel-specific gain control. Whether gain control is applied across all downmix channels or on a per-channel basis, bits associated with gain control information may be allocated from bits that would normally be used to encode the waveforms of the downmix channels and / or from bits that would normally be used to encode side information (e.g., metadata) used to reconstruct a FOA or HOA signal from the downmix channels, thereby reducing the number of available bits for encoding the downmix channels or the side information.

[0109] Figure 6An example use case of an IVAS system 600 according to one embodiment is shown. In some embodiments, various devices communicate through a call server 602, which is configured to receive audio signals from a public switched telephone network (PSTN) or a public land mobile network device (PLMN), such as shown by PSTN / other PLMN 604. The use case supports traditional devices 606 that render and capture audio only in mono, including but not limited to devices that support enhanced voice services (EVS), multi-rate wideband (AMR-WB), and adaptive multi-rate narrowband (AMR-NB). The use case also supports user equipment (UE) 608 and / or 614 that captures and renders stereo audio signals, or UE 610 that captures mono signals and renders them as multi-channel signals. The use case also supports immersive and stereo signals captured and rendered by video conference room systems 616 and / or 618, respectively. The use case also supports stereo capture and immersive rendering of stereo audio signals for a home theater system 620 , and mono capture and immersive rendering of audio signals for a virtual reality (VR) device 622 and immersive content ingestion 624 by a computer 612 .

[0110] Figure 7 is a block diagram showing an example of components of an apparatus capable of implementing various aspects of the present disclosure. Figure 1 Sample, Figure 7 The types and quantities of the elements shown in are provided as examples only. Other implementations may include more, less and / or different types and quantities of elements. According to some examples, the device 700 may be configured to perform at least some of the methods disclosed herein. In some implementations, the device 700 may be or may include one or more components of a television, an audio system, a mobile device (such as a cellular phone), a laptop computer, a tablet device, a smart speaker, or another type of device.

[0111] According to some alternative implementations, the apparatus 700 may be or may include a server. In some such examples, the apparatus 700 may be or may include an encoder. Thus, in some instances, the apparatus 700 may be a device configured for use in an audio environment such as a home audio environment, while in other instances, the apparatus 700 may be a device configured for use in the "cloud", such as a server.

[0112] In this example, the device 700 includes an interface system 705 and a control system 710. In some implementations, the interface system 705 can be configured to communicate with one or more other devices of the audio environment. In some examples, the audio environment can be a home audio environment. In other examples, the audio environment can be another type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, etc. In some implementations, the interface system 705 can be configured to exchange control information and associated data with the audio devices of the audio environment. In some examples, the control information and associated data can be related to one or more software applications being executed by the device 700.

[0113] In some implementations, the interface system 705 may be configured to receive or provide a content stream. The content stream may include audio data. The audio data may include, but is not limited to, an audio signal. In some cases, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.

[0114] The interface system 705 may include one or more network interfaces and / or one or more external device interfaces, such as one or more universal serial bus (USB) interfaces. According to some implementations, the interface system 705 may include one or more wireless interfaces. The interface system 705 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 705 may include a control system 710 and a user interface such as a touch screen. Figure 7 One or more interfaces between memory systems of the optional memory system 715 shown. However, in some cases, the control system 710 may include a memory system. In some implementations, the interface system 705 may be configured to receive input from one or more microphones in the environment.

[0115] For example, control system 710 may include a general purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0116] In some implementations, the control system 710 may reside in more than one device. For example, in some implementations, a portion of the control system 710 may reside in a device within one of the environments depicted herein, and another portion of the control system 710 may reside in a device outside the environment, such as a server, a mobile device (e.g., a smartphone or tablet computer), etc. In other examples, a portion of the control system 710 may reside in a device within an environment, and another portion of the control system 710 may reside in one or more other devices of the environment. For example, a portion of the control system 710 may reside in a device that implements a cloud-based service, such as a server, and another portion of the control system 710 may reside in another device that implements a cloud-based service, such as another server, a memory device, etc. In some examples, the interface system 705 may also reside in more than one device.

[0117] In some implementations, the control system 710 can be configured to at least partially perform the methods disclosed herein. According to some examples, the control system 710 can be configured to implement methods of determining a gain parameter, applying a gain transfer function, determining an inverse gain transfer function, applying an inverse gain transfer function, allocating bits for gain control to a bitstream, and the like.

[0118] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. One or more non-transitory media may reside, for example, on Figure 7 715 and / or control system 710 as shown. Thus, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software can include, for example, instructions for determining a gain parameter, applying a gain transfer function, determining an inverse gain transfer function, applying an inverse gain transfer function, allocating bits for gain control to a bitstream, and the like. For example, the software can be implemented by a program such as Figure 7 The control system 710 may be executed by one or more components of the control system.

[0119] In some examples, apparatus 700 may include Figure 7Optional microphone system 720 is shown. Optional microphone system 720 may include one or more microphones. In some implementations, one or more microphones may be part of or associated with another device such as a speaker of a speaker system, a smart audio device, etc. In some examples, device 700 may not include microphone system 720. However, in some such implementations, device 700 may still be configured to receive microphone data from one or more microphones in an audio environment via interface system 710. In some such implementations, a cloud-based implementation of device 700 may be configured to receive microphone data or a noise metric that at least partially corresponds to the microphone data from one or more microphones in an audio environment via interface system 710.

[0120] According to some implementations, the apparatus 700 may include Figure 7 Optional speaker system 725 is shown in . Optional speaker system 725 may include one or more loudspeakers, which may also be referred to herein as "speakers" or, more generally, "audio reproduction transducers". In some examples (e.g., cloud-based implementations), device 700 may not include speaker system 725. In some implementations, device 700 may include headphones. The headphones may be connected or coupled to device 700 via a headphone jack or via a wireless connection (e.g., Bluetooth).

[0121] Fig. 8A and Figure 8B An example implementation of perceptually excited gain control is shown with a sample-uniform gain control with DBSTEP=-1 dB on the encoder side. In this particular example, a frame consists of 1024 samples. The sample amplitudes are represented by dashed lines, while the gain applied to each sample is represented by solid lines. Fig. 8A As shown, once a frame generates an overload (amplitude greater than 0 dB) at the encoder, the gain function switches from no attenuation (0 dB) to an attenuation of DBSTEP-0 dB = -1 dB. When the input audio signal exceeds 1 dB, further attenuation is introduced according to DBSTEP.

[0122] Figure 8B The resulting attenuated downmix audio signal is shown in . In this particular example, the value of DBSTEP is large enough so that every sample is attenuated below the required threshold (0 dB).

[0123] Fig. 9A and Fig. 9B Shown at DBS = -1dB and G T ={0,1,3,6}, a set of attenuation values ​​{DBS*G T An example of "non-uniform" gain control with} = {0, -1, -3, -6} dB or DBSTEP set {-1, -2, -3}. Fig. 8A and Figure 8B As shown, the sample amplitude is represented by the dotted line and the gain function is represented by the solid line. Using the DBSTEP set, the automatic gain control can react to the overload caused by the encoder by attenuating the signal with an increasing value in each frame. Fig. 9B As shown, the gain transition step size is not large enough so that all samples are below the required threshold (0dB). This may cause distortion when the audio signal is rendered at the decoder, but the distortion caused by overload at the encoder is not as obvious as the distortion caused by very sudden gain changes.

[0124] Some aspects of the present disclosure include systems or devices configured (e.g., programmed) to perform one or more examples of the disclosed methods, and tangible computer-readable media, such as disks, storing code for implementing one or more examples of the disclosed methods or steps. For example, some disclosed systems may be or include a programmable general-purpose processor, digital signal processor, or microprocessor that is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system that includes an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to declared data.

[0125] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor that is configured (e.g., programmed and otherwise configured) to perform the desired processing on (multiple) audio signals, including performing one or more examples of the disclosed method. Alternatively, embodiments of the disclosed system (or its elements) may be implemented as a general-purpose processor, such as a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory that is programmed with software or firmware and / or otherwise configured to perform any of the various operations including one or more examples of the disclosed method. Alternatively, the elements of some embodiments of the system of the present invention are implemented as a general-purpose processor or DSP that is configured (e.g., programmed) to perform one or more examples of the disclosed method, and the system also includes other elements. Other elements may include one or more speakers and / or one or more microphones. A general-purpose processor configured to perform one or more examples of the disclosed method may be coupled to an input device. Examples of input devices include, for example, a mouse and / or a keyboard. The general-purpose processor may be coupled to a memory, a display device, etc.

[0126] Another aspect of the present disclosure is a computer-readable medium, such as a disk or other tangible storage medium, storing code for performing, for example, by an encoder executable to perform one or more examples of the methods or steps disclosed herein.

[0127] Although specific embodiments of the present disclosure and applications of the present disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations of the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that although certain forms of the present disclosure have been shown and described, the present disclosure is not limited to the specific embodiments described and shown or the specific methods described.

[0128] Various aspects and implementations of the present disclosure may also be understood from the following enumerated example embodiments (EEE), which are not the claims.

[0129] EEE1. A method for performing gain control on an audio signal, the method comprising:

[0130] Obtaining a downmixed audio signal of an audio signal to be encoded;

[0131] determining that an overload condition has occurred for a frame of the downmix audio signal;

[0132] responsive to determining that the overload condition has occurred, determining a gain transfer function for the frame, wherein the gain transfer function is based on at least a gain transfer step size;

[0133] applying the gain transfer function to the frame to generate a gain adjusted frame of the downmix audio signal; and

[0134] The gain adjusted frame and information indicative of the gain transfer function are provided for encoding by an encoder.

[0135] EEE2. A method according to statement EEE1, wherein the method further comprises:

[0136] The gain adjusted frame is encoded with information indicative of the gain transfer function.

[0137] EEE3. A method according to the preceding statement, wherein obtaining a downmix audio signal of an audio signal to be encoded comprises:

[0138] receiving the downmixed audio signal; or

[0139] The downmix audio signal is determined from the audio signal to be encoded.

[0140] EEE4. A method according to any preceding statement, wherein the audio signal is a higher-order ambisonics (HOA) audio signal.

[0141] EEE5. A method according to any preceding statement, wherein the downmix audio signal is a spatially coded downmix signal.

[0142] EEE6. A method according to any preceding statement, wherein the overload condition is a condition where the frame of the downmix audio signal exceeds a predefined signal range.

[0143] EEE7. A method according to EEE6, wherein the predefined signal range is a signal range expected by the encoder.

[0144] EEE8. A method according to any preceding statement, wherein the frame of the downmix audio signal is a current frame, and the gain transfer function is further based on a previous gain transfer function applied to a frame previous to the current frame. EEE9.

[0145] EEE9. A method according to any preceding statement, wherein the gain transition function further relies on a smoothing function based on the gain transition step size.

[0146] EEE10. A method according to EEE8, wherein the gain conversion function includes an instantaneous part and a steady-state part, and wherein the instantaneous part corresponds to a conversion from a gain associated with the previous frame to a gain associated with the previous frame adjusted by the gain conversion step size.

[0147] EEE11. A method according to EEE10, wherein, depending on a gain adjustment target of the current frame, the gain associated with the previous frame adjusted by the gain conversion step size is to attenuate the gain associated with the previous frame by the gain conversion step size or to amplify the gain conversion step size.

[0148] EEE12. A method according to EEE10 or 11, wherein the length of the instantaneous part is limited by the delay introduced by the codec used by the encoder. EEE13.

[0149] EEE13. A method according to EEE12, wherein the length of the instantaneous portion is equal to or less than the number of samples used by the encoder for encoding operations. EEE14.

[0150] EEE14. A method according to any one of EEE10 to 13, wherein the length of the instantaneous portion is greater than 1 sample.

[0151] EEE15. A method according to any preceding statement, wherein the gain transfer function is defined as

[0152]

[0153] Where DBSTEP is the gain conversion step size, l is the sample index, j is the frame index, p() is the smoothing function, l end denotes the rightmost index for which p() is defined, and L is the number of samples in a frame.

[0154] EEE16. A method according to any preceding statement, wherein the gain conversion step size is a predefined value.

[0155] EEE17. A method according to any preceding statement, wherein the gain transition step size is determined from a set of predefined values ​​that increase in size.

[0156] EEE18. The method according to EEE17, wherein the method further comprises:

[0157] determining an amount of overload caused by the frame of the downmix audio signal;

[0158] The gain transition step size is determined based on a set of predefined values ​​from which the overload amount increases.

[0159] EEE19. A method according to any preceding statement, wherein the gain transition step size is determined based on a perceptual quality listening test or an objective quality measurement.

[0160] EEE20. A method according to EEE19, wherein the perceptual quality listening test is a multi-stimulus test MUSHRA with hidden reference and anchor. EEE21.

[0161] EEE21. A method according to any preceding statement, wherein applying the gain transfer function to the frame to generate a gain adjusted frame of the downmix signal comprises:

[0162] The gain transfer function is applied to samples of the downmix audio signal, wherein a total number of the samples corresponds to the frame of the downmix audio signal.

[0163] EEE22. A method according to any one of EEE2 or EEE3 to 21, wherein encoding the gain-adjusted frame together with information indicating the gain conversion function comprises:

[0164] A coding scheme is determined based on the gain conversion function.

[0165] EEE23. The method according to EEE22, wherein determining a coding scheme based on the gain conversion function comprises:

[0166] The encoding scheme is determined based on the gain conversion step size.

[0167] EEE24. The method according to EEE22, wherein determining a coding scheme based on the gain conversion function comprises:

[0168] The encoding scheme is determined based on whether the gain conversion function can eliminate the overload condition.

[0169] EEE25. A method according to any one of EEE22 to 24, wherein the coding scheme is one of modified discrete cosine transform (MDCT) or algebraic code excited linear prediction (ACELP).

[0170] EEE26. A method according to any preceding statement, wherein the gain adjusted frame is an attenuated frame or an amplified frame.

[0171] EEE27. A method for performing gain control on an audio signal, the method comprising:

[0172] Receiving, at a decoder, encoded frames of an audio signal;

[0173] decoding said encoded frames of the audio signal to obtain frames of the downmix audio signal and information indicating a gain control applied by the encoder;

[0174] determining an inverse gain transfer function to be applied to the frame of the downmix audio signal based at least in part on information indicative of a gain control applied by the encoder, wherein the information indicative of the gain control applied by the encoder comprises a gain transfer step size; and

[0175] The inverse gain transfer function is applied to the frame of the downmix audio signal.

[0176] EEE28. The method according to EEE27, wherein the method further comprises:

[0177] The downmix audio signal is upmixed to generate an upmix audio signal, wherein the upmix audio signal is suitable for rendering.

[0178] EEE29. The method according to EEE28 also includes rendering the upmix signal to generate rendered audio data.

[0179] EEE30. The method according to EEE29 also includes playing back the rendered audio data using one or more of a loudspeaker or headphones.

[0180] EEE31. A method according to any one of EEE27 to 30, wherein the information indicating the gain control applied by the encoder also includes information indicating a smoothing function.

[0181] EEE32. A method according to any one of EEE27 to 31, wherein the inverse gain transfer function is determined by inverting a gain transfer function applied by the encoder. EEE33.

[0182] EEE33. A method according to any one of EEE27 to 32, wherein the inverse gain transfer function comprises an instantaneous part and a steady-state part.

[0183] EEE34. A method according to EEE33, wherein the length of the instantaneous portion is limited by the delay introduced by the codec used by the decoder. EEE34.

[0184] EEE35. An apparatus configured to implement the method of any one of EEE1-34.

[0185] EEE36. A program comprising instructions which, when executed by a processing device, cause the processing device to perform a method according to any one of EEE1-34. EEE36.

[0186] EEE37. A storage medium storing the program according to EEE36

[0187] EEA38. A method for performing gain control on an audio signal, the method comprising:

[0188] receiving, by an automatic gain control system, a spatially encoded downmix audio signal;

[0189] determining that an overload condition occurs for one or more frames of the received signal;

[0190] In response to the overload condition, generating an attenuation signal by applying a gain function to the received signal to attenuate the overload, the gain function being dependent on (1) an attenuation level parameter, (2) a gain function shape that specifies a corresponding attenuation level for each of the one or more frames, or (3) a combination of the attenuation level parameter and the gain function shape; and

[0191] The attenuation signal and a representation of the attenuation level parameter are provided to a core encoder for encoding.

[0192] EEE39. A method according to EEE38, wherein the attenuation level parameter comprises a table of numbers, each number corresponding to a respective attenuation level to be applied successively to the one or more frames. EEE39.

[0193] EEE40. A method according to EEE39, wherein each number has the same value, indicating that each step of attenuation attenuates the signal by the same amount.

[0194] EEE41. A method according to EEE39, wherein the value of the number increases, indicating that each step of attenuation attenuates the signal by an amount greater than the previous step.

[0195] EEA42. A method according to any one of EEE38-41, comprising directing the core encoder to encode the audio signal using different encoding schemes based on the attenuation level parameter. EEA43.

[0196] EEA43. A method according to any one of EEE38-42, comprising changing the shape of the gain function based on different values ​​of the attenuation level parameter. EEA44.

[0197] EEA44. An apparatus configured to implement the method described in any one of EEE38-43.

[0198] EEA45. One or more non-transitory media having software stored thereon, the software comprising instructions for controlling one or more devices to perform the method described in any of EEE38-43.

Claims

1. A method for performing gain control on an audio signal, the method comprising: Obtaining a downmixed audio signal of an audio signal to be encoded; determining that an overload condition has occurred for a frame of the downmix audio signal; responsive to determining that the overload condition has occurred, determining a gain transfer function for the frame, wherein the gain transfer function is based on at least a gain transfer step size; applying the gain transfer function to the frame to generate a gain adjusted frame of the downmix audio signal; as well as The gain adjusted frame and information indicative of the gain transfer function are provided for encoding by an encoder.

2. The method according to claim 1, wherein the method further comprises: The gain adjusted frame is encoded with information indicative of the gain transfer function.

3. The method according to the preceding claim, wherein obtaining a downmix audio signal of the audio signal to be encoded comprises: receiving the downmixed audio signal; or The downmix audio signal is determined from the audio signal to be encoded.

4. A method according to any preceding claim, wherein the audio signal is a higher order ambisonic (HOA) audio signal.

5. The method according to any preceding claim, wherein the downmix audio signal is a spatially coded downmix signal.

6. A method according to any preceding claim, wherein the overload condition is a condition where the frame of the downmix audio signal exceeds a predefined signal range. The method of claim 6 , wherein the predefined signal range is a signal range expected by the encoder.

8. The method of any preceding claim, wherein the frame of the downmix audio signal is a current frame and the gain transfer function is further based on a previous gain transfer function applied to a frame preceding the current frame.

9. A method according to any preceding claim, wherein the gain switching function further relies on a smoothing function based on the gain switching step size.

10. The method of claim 8, wherein the gain transfer function comprises a transient portion and a steady-state portion, and wherein the transient portion corresponds to a transition from a gain associated with the previous frame to the gain associated with the previous frame adjusted by the gain transfer step size.

11. The method according to claim 10, wherein the gain associated with the previous frame adjusted by the gain conversion step size is an attenuation of the gain associated with the previous frame by the gain conversion step size or an amplification by the gain conversion step size depending on the gain adjustment target of the current frame.

12. A method according to claim 10 or 11, wherein the length of the transient portion is limited by a delay introduced by a codec used by the encoder.

13. The method of claim 12, wherein the length of the transient portion is equal to or less than the number of samples used by the encoder for an encoding operation.

14. The method according to any one of claims 10 to 13, wherein the length of the transient portion is greater than 1 sample.

15. A method according to any preceding claim, wherein the gain transfer function is defined as in, DBSTEP is the gain conversion step size, l is the sample index, j is the frame index, p() is the smoothing function, l end denotes the rightmost index for which p() is defined, and L is the number of samples in one frame.

16. A method according to any preceding claim, wherein the gain switching step size is a predefined value.

17. A method according to any preceding claim, wherein the gain transition step size is determined from a set of predefined values ​​of increasing size.

18. The method according to claim 17, wherein the method further comprises: determining an amount of overload caused by the frame of the downmix audio signal; The gain transition step size is determined based on a set of predefined values ​​by which the overload amount is increased from the magnitude.

19. A method according to any preceding claim, wherein the gain switching step size is determined based on a perceptual quality listening test or an objective quality measurement.

20. The method of claim 19, wherein the perceptual quality listening test is a Multiple Stimulus Test (MUSHRA) with hidden reference and anchor.

21. The method of any preceding claim, wherein applying the gain transfer function to the frame to generate a gain adjusted frame of the downmix signal comprises: The gain transfer function is applied to samples of the downmix audio signal, wherein a total number of the samples corresponds to the frame of the downmix audio signal.

22. A method according to claim 2 or any one of claims 3 to 21 when dependent on claim 2, wherein encoding the gain adjusted frame together with information indicative of the gain conversion function comprises: A coding scheme is determined based on the gain conversion function.

23. The method of claim 22, wherein determining a coding scheme based on the gain conversion function comprises: The encoding scheme is determined based on the gain conversion step size.

24. The method of claim 22, wherein determining a coding scheme based on the gain conversion function comprises: The encoding scheme is determined based on whether the gain conversion function can eliminate the overload condition.

25. The method according to any one of claims 22 to 24, wherein the coding scheme is one of Modified Discrete Cosine Transform (MDCT) or Algebraic Code Excited Linear Prediction (ACELP).

26. A method according to any preceding claim, wherein the gain adjusted frame is an attenuated frame or an amplified frame.

27. A method for performing gain control on an audio signal, the method comprising: Receiving, at a decoder, encoded frames of an audio signal; decoding said encoded frames of the audio signal to obtain frames of the downmix audio signal and information indicating a gain control applied by the encoder; determining an inverse gain transfer function to be applied to the frame of the downmix audio signal based at least in part on information indicative of a gain control applied by the encoder, wherein the information indicative of a gain control applied by the encoder comprises a gain transfer step size; and The inverse gain transfer function is applied to the frame of the downmix audio signal.

28. The method according to claim 27, wherein the method further comprises: The downmix audio signal is upmixed to generate an upmix audio signal, wherein the upmix audio signal is suitable for rendering.

29. The method of claim 28, further comprising rendering the upmix signal to produce rendered audio data.

30. The method of claim 29, further comprising playing back the rendered audio data using one or more of a loudspeaker or headphones.

31. A method according to any one of claims 27 to 30, wherein the information indicative of a gain control applied by the encoder further comprises information indicative of a smoothing function.

32. A method according to any one of claims 27 to 31, wherein the inverse gain transfer function is determined by inverting a gain transfer function applied by the encoder.

33. A method according to any one of claims 27 to 32, wherein the inverse gain transfer function comprises a transient portion and a steady-state portion.

34. The method of claim 33, wherein the length of the transient portion is limited by a delay introduced by a codec used by the decoder.

35. An apparatus configured to implement the method of any one of claims 1-34.

36. A program comprising instructions which, when executed by a processing device, cause the processing device to perform a method according to any one of claims 1 to 34.

37. A storage medium storing the program according to claim 36.