Method and apparatus for applying dynamic range compression to high-order Ambisonics signals
By deriving and applying gain factors directly in the HOA domain or transforming HOA signals for complex DRC, the method efficiently compresses dynamic ranges in Higher Order Ambisonics signals, addressing the challenge of spatially varying gains.
Patent Information
- Application Number
- JP2025039903
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2014-04-15
- Filing Date
- 2025-03-13
- Publication Date
- 2026-02-19
- Estimated Expiration
- 2035-03-24
AI Technical Summary
Dynamic Range Compression (DRC) methods are not straightforward for Higher Order Ambisonics (HOA) signals due to their non-sound field description, making it difficult to apply spatially varying gains effectively.
The HOA signal is analyzed to derive one or more gain factors, which are transmitted with the signal, allowing for direct application in the HOA domain in a simplified mode, or transformed into the spatial domain for more complex compression, reducing computational complexity.
This approach enables efficient DRC for HOA signals with reduced processing requirements, maintaining spatial integrity and dynamic range adjustment without audible artifacts.
Smart Images

Figure 0007818124000121 
Figure 0007818124000122 
Figure 0007818124000123
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and apparatus for performing Dynamic Range Compression (DRC) on Ambisonics signals, in particular Higher Order Ambisonics (HOA) signals. [Background technology]
[0002] The goal of Dynamic Range Compression (DRC) is to reduce the dynamic range of an audio signal. A time-varying gain factor is applied to the audio signal. Typically, this gain factor depends on the amplitude envelope of the signal used to control the gain. The mapping is generally non-linear. Large amplitudes are mapped to smaller amplitudes, while soft sounds are often amplified. Scenarios include noisy environments, late-night listening, and listening with small speakers or mobile headphones.
[0003] A common concept for streaming or broadcasting audio is to generate DRC gains before transmission and apply these gains after reception and decoding. The principle of using DRC, i.e. how DRC is typically applied to an audio signal, is shown in Figure 1a. The signal level, typically the signal envelope, is detected and an associated time-varying gain g DRC is calculated. The gain is used to change the amplitude of the audio signal. Figure 1b) shows the principle of using DRC for encoding / decoding, where the gain factor is transmitted together with the coded audio signal. At the decoder side, the gain is applied to the decoded audio signal to reduce its dynamic range.
[0004] For 3D audio, different gains can be applied to loudspeaker channels representing different spatial positions. These positions then need to be known at the sender to be able to generate a matching set of gains. This is usually only possible for idealized conditions. In the real world, the number of speakers and their placement can vary in many ways. This is influenced more by practical considerations than specifications. Higher-order Ambisonics is an audio format that allows flexible rendering. HOA signals consist of coefficient channels that do not directly represent sound levels. Therefore, DRC cannot simply be applied to HOA-based signals. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2015 / 007889A (PD130040) [Non-patent literature]
[0006] [Non-Patent Document 1] J¨org Fliege, “Integration nodes for the sphere”, 2010, online, accessed 2010-10-05, http: / / www.mathematik.uni-dortmund.de / lsx / research / projects / fliege / nodes / nodes.html [Non-patent document 2] J¨org Fliege and Ulrike Maier, “A two-stage approach for computing cubature formula for the sphere”, Technical report, Fachbereich Mathematik, Universitat Dortmund, 1999 Summary of the Invention [Problem to be solved by the invention]
[0007] The present invention at least solves the problem of how DRC can be applied to HOA signals. [Means for solving the problem]
[0008] The HOA signal is analyzed to obtain one or more gain factors. In one embodiment, at least two gain factors are obtained, and the analysis of the HOA signal includes a transformation to the spatial domain (iDSHT). The one or more gain factors are transmitted together with the original HOA signal. A special indicator can be transmitted to indicate whether all gain factors are equal. This is the case in the so-called simplified mode. In the non-simplified mode, on the other hand, at least two different gain factors are used. At the decoder, the one or more gain factors can be applied to the HOA signal (although this is not required). The user has the option of applying the one or more gain factors. The advantage of the simplified mode is that only one gain factor is used, and the gain factor can be applied directly to the coefficient channel of the HOA signal in the HOA domain, skipping the transformation to the spatial domain and then back to the HOA domain, thereby significantly reducing the required computations. In the simplified mode, the gain factor is obtained by analyzing only the zeroth-order coefficients of the HOA signal.
[0009] According to one embodiment of the present invention, a method for performing DRC on an HOA signal includes transforming the HOA signal into the spatial domain (by an inverse DSHT), analyzing the transformed HOA signal, and deriving from the analysis a gain factor usable for dynamic range compression. In a further step, the obtained gain factor is multiplied with the transformed HOA signal (in the spatial domain) to obtain a gain-compressed transformed HOA signal. Finally, the gain-compressed transformed HOA signal is transformed back into the HOA domain, i.e., the coefficient domain, (by a DSHT) to obtain a gain-compressed HOA signal.
[0010] Furthermore, according to an embodiment of the present invention, a method for performing DRC in simplified mode on an HOA signal includes analyzing the HOA signal and, from the analysis, deriving a gain factor that can be used for dynamic range compression. In a further step, upon evaluation of the index, the derived gain factor is multiplied with a coefficient channel of the HOA signal (in the HOA domain) to obtain a gain-compressed transformed HOA signal. Upon evaluation of the index, it may also be determined that transformation of the HOA signal can be skipped. The index indicating the simplified mode, i.e., that only one gain factor is used, can be set implicitly, e.g., if only the simplified mode is available due to hardware or other constraints, or explicitly, e.g., upon user selection of either the simplified mode or the non-simplified mode.
[0011] Furthermore, a method for applying a DRC gain factor to an HOA signal includes receiving an HOA signal, an index, and a gain factor, determining that the index indicates a non-simplified mode, transforming (using an inverse DSHT) the HOA signal to the spatial domain to obtain a transformed HOA signal, multiplying the transformed HOA signal by the gain factor to obtain a dynamically range-compressed transformed HOA signal, and transforming (using the DSHT) the dynamically range-compressed transformed HOA signal back to the HOA domain (i.e., the coefficient domain) to obtain a dynamically range-compressed HOA signal. The gain factor can be received together with or separately from the HOA signal. Furthermore, according to an embodiment of the present invention, a method for applying a DRC gain factor to an HOA signal includes receiving an HOA signal, an index, and a gain factor, determining that the index indicates a simplified mode, and, upon said determination, multiplying the HOA signal by the gain factor to obtain a dynamically range-compressed HOA signal. The gain factor can be received together with or separately from the HOA signal.
[0012] An apparatus for applying a DRC gain factor to an HOA signal is disclosed in claim 11 .
[0013] In one embodiment, the present invention provides a computer-readable medium having executable instructions for causing a computer to perform a method of applying a DRC gain factor to an HOA signal, the method including the steps as described above.
[0014] In one embodiment, the present invention provides a computer-readable medium having executable instructions for causing a computer to perform a method of performing DRC on an HOA signal, the method including the steps as described above.
[0015] Advantageous embodiments of the invention are disclosed in the dependent claims, the following description and the drawings. [Brief explanation of the drawings]
[0016] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. [Figure 1] FIG. 1 illustrates the general principle of DRC applied to audio. [Figure 2] FIG. 1 illustrates a general approach for applying DRC to HOA-based signals according to the present invention. [Figure 3] 1 shows a spherical speaker grid for N=1 to N=6. [Figure 4] FIG. 1 illustrates the generation of DRC gains for an HOA. [Figure 5] FIG. 1 illustrates the application of DRC to an HOA signal. [Figure 6] FIG. 10 is a diagram illustrating a dynamic range compression process on the decoder side. [Figure 7] FIG. 10 illustrates the DRC for HOA in the QMF domain combined with the rendering stage. [Figure 8] FIG. 10 illustrates the DRC for HOA in the QMF domain combined with the rendering stage for the simple case of a single DRC gain group. DETAILED DESCRIPTION OF THE INVENTION
[0017] This invention describes how DRC can be applied to HOA. This is usually not straightforward since HOA is not a sound field description. Figure 2 illustrates the principle of this approach. At the encoding or transmitting side, as shown in Figure 2a), the HOA signal is analyzed, a DRC gain g is calculated from the analysis of the HOA signal, and the DRC gain is coded and transmitted together with the coded representation of the HOA content. This can be a multiplexed bitstream or two or more separate bitstreams.
[0018] At the decoding or receiving side, a gain g is extracted from such bitstream(s), as shown in Fig. 2b). After decoding the bitstream in the decoder, the gain g is applied to the HOA signal as described below. This results in a gain being applied to the HOA signal, i.e., generally, an HOA signal with a reduced dynamic range. Finally, the dynamic range-adjusted HOA signal is rendered in an HOA renderer.
[0019] The assumptions and definitions used are explained below.
[0020] The assumption is that the HOA renderer is energy-preserving: N3D normalized spherical harmonics are used and the energy of the unidirectional signal, encoded inside the HOA representation, is preserved after rendering. For example, U.S. Patent No. 6,249,999 describes how to achieve this energy-preserving HOA rendering.
[0021] The definitions of the terms used are as follows:
[0022]
number
number
[0023]
number
[0024]
number
[0025] For all HOA truncation orders N, the ideal L L =(N+1) 2 virtual speaker lattices and associated rendering matrices D L is defined. The virtual speaker positions sample the spatial area surrounding the virtual listener. The grid for N=1 to 6 is shown in Figure 3, where the area associated with a speaker is the shaded cell. One sampling position is always associated with the center speaker position (azimuth=0, tilt=π / 2; note that the azimuth is measured from the front direction relative to the listening position). The sampling position, D L , D L -1 is known at the encoder side when the DRC gain is generated. At the decoder side, D is used to apply the gain value. L and D L -1 needs to be known.
[0026] The generation of the DRC gain for an HOA works as follows.
[0027] HOA signal is W L =D L B. By analyzing these signals, L L =(N+1) 2 DRC gain gl is generated. If the content is a combination of HOA and Audio Objects (AO), the AO signal, e.g., a dialogue track, can be used for side-chaining. This is shown in Fig. 4b). When generating different DRC gain values related to different spatial areas, care must be taken to ensure that these gains do not affect the spatial image stability at the decoder side. To avoid this, in the simplest case (the so-called simplified mode), a single gain may be assigned to all L channels. This can be done by analyzing all spatial signals W or by using the zeroth-order HOA coefficient sample block.
number
[0028] In Figure 4, the generation of DRC gains for an HOA is shown. Figure 4a) shows how a single gain g1 (for a single gain group) is applied to the zeroth-order HOA component
number
[0029] The gain value is transmitted to the receiver or decoder side.
[0030] Variable numbers 1 through L related to blocks of τ samples L =(N+1) 2 gain values are transmitted. Gain values can be assigned to groups of channels for transmission. In one embodiment, all equal gains are combined into one group of channels to minimize transmitted data. If a single gain is transmitted, it is L L It concerns all channels. The channel group gain value g lg The use of the channel group is signaled so that the receiver or decoder can apply the gain values correctly.
[0031] The gain values are applied as follows:
[0032] The receiver / decoder determines the number of encoded gain values transmitted, decodes the associated information (51), and converts those gains to L L=(N+1) 2 If only one gain value (one channel group) is transmitted, it can be directly applied to the HOA signal (52) (B DRC =g1B). This is advantageous because decoding is much simpler and requires significantly less processing, since no matrix operations are required and instead the gain values can be applied 52 directly, e.g., multiplied by the HOA coefficients.
[0033] When more than one gain is transmitted, the channel group gains are respectively L channel gains g=[g1,...,g L ] is assigned to.
[0034] For a hypothetical regular loudspeaker grid, the loudspeaker signal with DRC gain applied is
number
number
number
number
[0035] This is (N+1) 2<τ, this solution has advantages over conventional solutions because it is more efficient in terms of the computational operations required, i.e., it is much simpler to decode and requires significantly less processing. The reason is that no matrix operations are required, and instead the gain values can be applied directly, i.e., multiplied to the HOA coefficients, in the gain assignment block 54.
[0036] In one embodiment, a more efficient way to apply the gain matrix is to modify the renderer matrix in the renderer matrix modification block 57 as
number
number
[0037] In summary, Figure 5 shows various embodiments of applying DRC to HOA signals. In Figure 5a), a single channel group gain is transmitted, decoded (51), and applied directly to the HOA coefficients (52). The HOA coefficients are then rendered using a regular rendering matrix (56).
[0038] In Figure 5b), two or more channel group gains are transmitted and decoded (51). The result of the decoding is (N+1) 2 A gain vector g of gain values is obtained. A gain matrix G is generated and applied to the block of HOA samples (54). These are then rendered using the normal rendering matrix.
[0039] In Figure 5c), instead of applying the decoded gain matrix / gain values directly to the HOA signals, we apply them directly to the renderer matrices. This is performed in the renderer matrix modification block 57. This is computationally beneficial if the DRC block size τ is larger than the number of output channels L. In this case, the HOA samples are rendered using the modified rendering matrix (57).
[0040] Below, we describe the calculation of an ideal DSHT (Discrete Spherical Harmonic Transform) matrix for DRC. Such a DSHT matrix is specifically optimized for use in DRC and differs from DSHT matrices used for other purposes, such as data rate compression.
[0041] The ideal rendering and encoding matrix D related to the ideal spherical layout L and D L -1 The requirements for are derived below. Ultimately, these requirements are: (1) Rendering matrix D L must be invertible, i.e., D L -1 must be present; (2) The sum of amplitudes in the spatial domain should be reflected as the zeroth-order HOA coefficient after the transformation from the spatial domain to the HOA domain and should be preserved after the subsequent transformation to the spatial domain (amplitude requirement); (3) The energy of the spatial signal should be preserved when transforming into the HOA domain and back into the spatial domain (energy conservation requirement).
[0042] Even for an ideal rendering layout, requirements 2 and 3 seem to contradict each other. When using a simple approach for deriving DSHT transformation matrices as known from the prior art, only one or the other of requirements (2) and (3) can be satisfied without error. Satisfying one of requirements (2) and (3) without error leads to an error of more than 3 dB for the other, which usually leads to audible artifacts. A way to overcome this problem is described below.
[0043] First, L=(N+1) 2 The ideal spherical layout is selected. The L directions of the (virtual) speaker positions are l and the associated modal matrix is
number
number
[0044] First prototype rendering matrix (D with tilde) L ] is derived by the following formula:
[0045]
number
[0046] Second, a compact singular value decomposition is performed:
number
[0047]
number
number
[0048] Fourth, in the final step, the amplitude error is substituted to satisfy requirement 2: If the row vector e is
number
[0049]
number
number
number
number
[0050] The following describes the detailed requirements for DRC.
[0051] First, L with value g1 in the spatial domain L Applying the same gains is equivalent to applying the gain g1 to the HOA coefficients:
number
[0052] Second, analyzing the sum signal in the spatial domain is equivalent to analyzing the zero-order HOA component. The DRC analyzer uses the energy of the signal as well as its amplitude. Thus, the sum signal is related to amplitude and energy.
[0053] HOA signal model
number
[0054]
number
number
[0055] The zeroth order component HOA signal must be summed with the directional signal to reflect the correct amplitude of the sum signal.
[0056]
number
[0057] In this mix,
number
number
[0058] The sum of amplitudes in the spatial domain is the HOA pan matrix M L =D L Ψ e Using
number
[0059] this is,
number
number
number
[0060] This also ensures that the energy requirement for the sum signal can be met. The energy sum in the spatial domain is
number
number
[0061] this is,
number
[0062] Third, conservation of energy is essential.
number
number
number
number
number
[0063] RequirementsVV T =1 is L≧(N+1) 2 This can be achieved for L<(N+1) 2 can only be approximated.
[0064] As examples, cases with ideal spherical positions (HOA orders N=1 to N=3) are described below (Tables 1-3). Ideal spherical positions for additional HOA orders (N=4 to N=6) are described last (Tables 4-6). All positions below are derived from the modified positions published in [1]. These positions and the associated quadrature / higher order quadrature gains were published in [2]. In these tables, azimuth angles are measured counterclockwise from the front direction relative to the listening position, and tilt angles are measured from the z-axis with tilt 0 above the listening position.
[0065] [Table 1] Table 1: a) Spherical positions of virtual loudspeakers for HOA order N=1, b) Resulting rendering matrix for spatial transformation (DSHT).
[0066] [Table 2] Table 2: a) Spherical positions of virtual loudspeakers for HOA order N=2, b) Resulting rendering matrix for spatial transformation (DSHT).
[0067] [Table 3] Table 3: a) Spherical positions of virtual loudspeakers for HOA order N=2, b) Resulting rendering matrix for spatial transformation (DSHT).
[0068] The term numerical quadrature, often abbreviated to quadrature, is entirely synonymous with numerical integration, especially when applied to one-dimensional integrals. Numerical integration in two or more dimensions is referred to herein as higher-order quadrature.
[0069] A typical application scenario for applying DRC gain to an HOA signal is shown above in Figure 5. For mixed content applications, such as HOA and audio objects, DRC gain application can be realized in at least two ways for flexible rendering.
[0070] Figure 6 exemplarily illustrates dynamic range compression (DRC) processing at the decoder side: in Figure 6a) DRC is applied before rendering and mixing; in Figure 6b) DRC is applied to the loudspeaker signals, i.e. after rendering and mixing.
[0071] In Figure 6a), DRC gains are applied separately to the audio objects and the HOA. DRC gains are applied to the audio objects in the audio object DRC block 610, and DRC gains are applied to the HOA in the HOA DRC block 615. Here, the implementation of the block HOA DRC block 615 corresponds to the one in Figure 5. In Figure 6b), a single gain is applied to all channels of the mixed signal of the rendered HOA and the rendered audio object signals. Here, spatial emphasis and attenuation are not possible. The associated DRC gains cannot be generated by analyzing the sum signal of the rendered mix, since the speaker layout at the consumer site is not known at the time of production at the broadcast or content production site. The DRC gains are
number
[0072]
number
[0073] DRC on HOA Content The DRC may be applied to the HOA signal before rendering or may be combined with rendering. The DRC for the HOA can be applied in the time domain or in the QMF filter bank domain.
[0074] For time-domain DRC, the DRC decoder uses (N+1) HOA coefficient channels depending on the number of HOA signal c. 2 gain values
number
[0075] The DRC gain is
number
number
number
[0076] In one embodiment, (N+1) per sample 4 To reduce the computational load by one operation, the loudspeaker signal is processed in a rendering stage.
number
[0077] As in the simplified mode, all gains g1,…,g (N+1)2 The same value g drc , a single gain group was used to transmit the encoder DRC gains. This case can be flagged by the DRC decoder. In this case, no computations are required in the spatial filter, so the computation is
number
[0078] The above describes how to obtain and apply the DRC gain values. Below we describe the calculation of the DSHT matrix for the DRC.
[0079] In the following, D L is D DSHT The name will be changed to Spatial Filter D DSHT and its inverse D DSHT -1 The matrix for determining is calculated as follows:
[0080] A set of spherical positions indexed by the HOA order N from Tables 1-4
number
number
number
number
number
number
number
number
[0081]
number
number
number
[0082] For DRC in the QMF filter bank domain, the following holds true:
[0083] The DRC decoder is (N+1) 2 Gain values g for all time-frequency tiles n, m for spatial channels ch (n,m). The gain for time slot n and frequency band m is
number
[0084] In the QMF filter bank domain, multi-band DRC is applied. The processing steps are shown in Figure 7. The reconstructed HOA signal is then processed using the (inverse DSHT):W DSHT =D DSHT It is transformed into the spatial domain by C, where
number
number
[0085]
number
number
number
[0086] Figure 7 shows the DRC for HOA in the QMF domain combined with the rendering stage. If only a single set of gains for the DRC is used, this should be flagged by the DRC decoder, as this again allows for computational simplification. In this case, the gains in the vector g(n,m) all have the same value g DRC (n,m). A QMF filter bank can be applied directly to the HOA signal, with gain g DRC(n,m) can be multiplied in the filter bank domain.
[0087] Figure 8 shows the DRC for an HOA in the QMF domain (Quadrature Mirror Filter filter domain) combined with the rendering stage, with a computational simplification for the simple case of a single DRC gain group.
[0088] As will become apparent in view of the above, in one embodiment, the present invention relates to a method for applying a dynamic range compression gain factor to an HOA signal, the method comprising the steps of receiving an HOA signal and one or more gain factors, transforming 40 the HOA signal into the spatial domain using an iDSHT with a transformation matrix derived from the spherical positions of virtual loudspeakers and a quadrature gain q to obtain a transformed HOA signal, multiplying the gain factor by the transformed HOA signal to obtain a dynamically range compressed transformed HOA signal, and transforming the dynamically range compressed transformed HOA signal back into the original HOA domain, which is the coefficient domain, using a Discrete Spherical Harmonic Transform (DSHT) to obtain a dynamically range compressed HOA signal.
[0089] Furthermore, the transformation matrix is
number
number
number
number
number
[0090] Furthermore, in one embodiment, the present invention relates to an apparatus for applying DRC gain factors to an HOA signal, the apparatus comprising a processor or one or more processing elements adapted to receive an HOA signal and one or more gain factors, transforming 40 the HOA signal into the spatial domain using an iDSHT with a transformation matrix derived from the spherical position of a virtual loudspeaker and a quadrature gain q to obtain a transformed HOA signal, multiplying the transformed HOA signal by the gain factors to obtain a dynamically range-compressed transformed HOA signal, and transforming the dynamically range-compressed transformed HOA signal back into the original HOA domain, which is the coefficient domain, using a Discrete Spherical Harmonic Transform (DSHT) to obtain a dynamically range-compressed HOA signal. Furthermore, the transformation matrix is
number
number
number
number
number
[0091] Furthermore, in one embodiment, the present invention relates to a computer-readable storage medium having computer-executable instructions, which, when executed on a computer, cause the computer to perform a method for applying dynamic range compression gain factors to a Higher Order Ambisonics (HOA) signal, the method including the steps of receiving an HOA signal and one or more gain factors, transforming 40 the HOA signal into the spatial domain using an iDSHT with a transformation matrix derived from the spherical positions of virtual loudspeakers and a quadrature gain q to obtain a transformed HOA signal, multiplying the gain factors by the transformed HOA signal to obtain a dynamically range compressed transformed HOA signal, and transforming the dynamically range compressed transformed HOA signal back to the original HOA domain, which is the coefficient domain, using a Discrete Spherical Harmonic Transform (DSHT) to obtain a dynamically range compressed HOA signal. Furthermore, the transformation matrix is
number
number
number
number
number
[0092] Furthermore, in one embodiment, the present invention relates to a method for performing DRC on an HOA signal, the method including the steps of: setting or determining a mode, either a simplified mode or a non-simplified mode; transforming the HOA signal into the spatial domain using an inverse DSHT in the non-simplified mode; analyzing the transformed HOA signal in the non-simplified mode; and deriving from the analysis one or more gain factors usable for dynamic range compression, where only one gain factor is obtained in the simplified mode and two or more different gain factors are obtained in the non-simplified mode; multiplying the obtained gain factor by the HOA signal to obtain a gain-compressed HOA signal in the simplified mode; and multiplying the obtained gain factor by the transformed HOA signal to obtain a gain-compressed transformed HOA signal in the non-simplified mode; and transforming the gain-compressed transformed HOA signal back into the HOA domain to obtain the gain-compressed HOA signal.
[0093] In one embodiment, the method further includes receiving an indicator indicating either a simplified mode or a non-simplified mode, and selecting the non-simplified mode if the indicator indicates the non-simplified mode, and selecting the simplified mode if the indicator indicates the simplified mode, wherein the steps of converting the HOA signal to the spatial domain and converting the dynamic range compressed converted HOA signal back to the HOA domain are performed only in the non-simplified mode, and in the simplified mode, the HOA signal is multiplied by only one gain factor.
[0094] In one embodiment, the method further includes analyzing the HOA signal in a simplified mode and analyzing a converted HOA signal in a non-simplified mode, and then obtaining one or more gain factors that can be used for dynamic range compression from the results of the analysis, wherein two or more different gain factors are obtained in the non-simplified mode and only one gain factor is obtained in the simplified mode, and in the simplified mode, a gain-compressed HOA signal is obtained by multiplying the obtained gain factor by the HOA signal, and in the non-simplified mode, the gain-compressed converted HOA signal is obtained by multiplying the obtained two or more gain factors by the converted HOA signal, and in the non-simplified mode, the converting of the HOA signal to the spatial domain uses an inverse DSHT.
[0095] In some embodiments, the HOA signal is divided into frequency subbands, and the gain factor(s) are derived separately for each frequency subband and applied with an individual gain per subband. In some embodiments, the steps of analyzing the HOA signal (or the transformed HOA signal), deriving one or more gain factors, multiplying the HOA signal (or the transformed HOA signal) by the derived gain factors, and converting the gain-compressed transformed HOA signal back to the HOA domain are applied separately to each frequency subband with an individual gain per subband. Note that the order of dividing the HOA signal into frequency subbands and converting the HOA signal to the spatial domain can be interchanged, and / or the order of combining the subbands and converting the gain-compressed transformed HOA signal back to the HOA domain can be interchanged. These interchanges can be independent of each other.
[0096] In one embodiment, the method further comprises transmitting the converted HOA signal together with the obtained gain factors and the number of these gain factors prior to the step of multiplying by the gain factors.
[0097] In one embodiment, the transformation matrix is a mode matrix ΨDSHT and the corresponding quadrature gains, and the modal matrix Ψ DSHT teeth
number
number
[0098] In one embodiment, the HOA signal B is transformed into the spatial domain to produce a transformed HOA signal W DSHT is obtained, and the converted HOA signal W DSHT is the gain value diag(g)
number
number
number
[0099] In one embodiment, at least (N+1) 2 If <τ, the method further comprises:
number
Number
[0100] In an embodiment, assuming that L is the number of output channels and τ is the DRC block size, when at least L < τ, the method further includes applying the gain matrix G
Number
[0101] In an embodiment, the present invention relates to a method of applying a DRC gain factor to an HOA signal. The method includes receiving an HOA signal together with an index and one or more gain factors, where the index indicates either a simplified mode or a non-simplified mode, and only one gain factor is received when the index indicates the simplified mode, a step of selecting either the simplified mode or the non-simplified mode according to the index, in the simplified mode, multiplying the gain factor by the HOA signal to obtain a dynamically range-compressed HOA signal, in the non-simplified mode, converting the HOA signal to the spatial domain to obtain a converted HOA signal, multiplying the gain factor by the converted HOA signal to obtain a dynamically range-compressed converted HOA signal, and converting the dynamically range-compressed converted HOA signal back to the original HOA domain to obtain a dynamically range-compressed HOA signal, including the step and the step.
[0102] Furthermore, in one embodiment, the present invention relates to an apparatus for performing DRC on an HOA signal, the apparatus having a processor or one or more processing elements adapted to perform the following steps: setting or determining a mode, either a simplified mode or a non-simplified mode; transforming the HOA signal into the spatial domain using an inverse DSHT in the non-simplified mode; analyzing the transformed HOA signal in the non-simplified mode and analyzing the HOA signal in the simplified mode; and deriving from the analysis one or more gain factors usable for dynamic range compression, where a single gain factor is obtained in the simplified mode and two or more different gain factors are obtained in the non-simplified mode; multiplying the HOA signal by the obtained gain factor in the simplified mode to obtain a gain-compressed HOA signal; and multiplying the transformed HOA signal by the obtained gain factor in the non-simplified mode to obtain a gain-compressed transformed HOA signal; and transforming the gain-compressed transformed HOA signal back into the HOA domain to obtain the gain-compressed HOA signal.
[0103] In one embodiment for the non-simplified mode only, the apparatus for performing DRC on the HOA signal comprises a processor or one or more processing elements adapted to perform the steps of: analyzing the transformed HOA signal; deriving from the analysis one or more gain factors usable for dynamic range compression; multiplying the transformed HOA signal by the obtained gain factors to obtain a gain-compressed transformed HOA signal; and converting the gain-compressed transformed HOA signal back to the HOA domain to obtain a gain-compressed HOA signal. In one embodiment, the apparatus further comprises a transmitting unit for transmitting the HOA signal together with the obtained gain factor(s) before multiplying with the obtained gain factor(s).
[0104] Furthermore, in one embodiment, the present invention relates to an apparatus for applying DRC gain factors to an HOA signal, the apparatus comprising a processor or one or more processing elements adapted to receive the HOA signal together with an index and one or more gain factors, the index indicating either a simplified mode or a non-simplified mode, where only one gain factor is received when the index indicates the simplified mode, select either the simplified mode or the non-simplified mode according to the index, and multiply the HOA signal by the gain factor in the simplified mode to obtain a dynamic range-compressed HOA signal, or in the non-simplified mode, transform the HOA signal into the spatial domain to obtain a transformed HOA signal, multiply the gain factor by the transformed HOA signal to obtain a dynamic range-compressed transformed HOA signal, and transform the dynamic range-compressed transformed HOA signal back to the HOA domain to obtain a dynamic range-compressed HOA signal.
[0105] Furthermore, in one embodiment, the apparatus further comprises a transmitting unit for transmitting the HOA signal together with the obtained gain factor before multiplying by the obtained gain factor. In one embodiment, the HOA signal is divided into frequency subbands, and analyzing the transformed HOA signal, obtaining a gain factor, multiplying the transformed HOA signal by the obtained gain factor, and converting the gain-compressed transformed HOA signal back to the HOA domain are applied to each frequency subband separately using individual gains for each subband.
[0106] In one embodiment of an apparatus for applying DRC gain factors to an HOA signal, the HOA signal is divided into multiple frequency subbands, and obtaining one or more gain factors, multiplying the obtained gain factors with the HOA signal or the transformed HOA signal, and in a non-simplified mode, converting the gain-compressed transformed HOA signal back to the original HOA signal is applied to each frequency subband separately, using individual gains for each subband.
[0107] Furthermore, in an embodiment in which only the non-simplified mode is used, the present invention relates to an apparatus for applying a DRC gain factor to an HOA signal, the apparatus comprising a processor or one or more processing elements adapted to receive an HOA signal together with a gain factor, transforming the HOA signal to the spatial domain (using an iDSHT) to obtain a transformed HOA signal, multiplying the gain factor by the transformed HOA signal to obtain a dynamically range-compressed transformed HOA signal, and transforming the dynamically range-compressed transformed HOA signal back to the HOA domain (i.e., the coefficient domain) (using an iDSHT) to obtain a dynamically range-compressed HOA signal.
[0108] Tables 4-6 below list the spherical positions of the virtual loudspeakers for an HOA of order N, where N=4, 5 or 6.
[0109] While the basic novel features of the present invention have been shown, described, and pointed out as applied to its preferred embodiments, it will be understood that various omissions, substitutions, and changes in the described apparatus and methods may be made by those skilled in the art in the form and details of the disclosed devices and their operation without departing from the spirit of the invention. Any combination of elements that perform substantially the same function in substantially the same way to achieve the same results is expressly intended to be within the scope of the invention. The substitution of elements from one described embodiment for another described embodiment is also fully intended and contemplated.
[0110] It will be understood that the present invention has been described purely by way of example and modifications of detail can be made without departing from the scope of the invention. Each feature disclosed in the description and (where appropriate) the claims and drawings may be provided independently or in any appropriate combination. Features may, where appropriate, be implemented in hardware, software or a combination of both.
[0111] [Table 4] Table 4: Spherical positions of virtual loudspeakers for HOA order N=4.
[0112] [Table 5] Table 5: Spherical positions of virtual loudspeakers for HOA order N=5.
[0113] [Table 6] Table 6: Spherical positions of virtual loudspeakers for HOA order N=6.
Claims
1. 1. A method for dynamic range compression (DRC): receiving a reconstructed Higher Order Ambisonics (HOA) audio signal representation; The reconstructed HOA audio signal [Equation 1] Transforming into the spatial domain based on D DSHT is the inverse discrete spherical harmonic transform (DSHT) matrix, C is a block of τ HOA samples, and WDSHT is a block of spatial samples that matches the input time granularity of the quadrature mirror filter (QMF) bank; The DRC gain value g(n,m) corresponding to the time-frequency tile (n,m) is [Equation 2] applying the method in accordance with [Equation 3] is the spatial channel vector for time-frequency tile (n,m), and To the loudspeaker channel [Equation 4] Rendering based on D DSHT -1 The matrix is D DSHT where D is the inverse of the matrix D and D is the HOA rendering matrix. D DSHT -1 matrix and D DSHT The queue is [Equation 5] and [Equation 6] and [1,0,0,..,0] is the element (N+1) where the first element has value 1 and all other elements have value 0. 2 row vector, N is the HOA order, [Equation 7] and the compact singular value decomposition is [Equation 8] and the new prototype matrix is [Equation 9] is calculated by [Equation 10] and the set of spherical positions [0011] and related quadrature (higher order quadrature) gains [0012] is selected and the mode matrix Ψ DSHT relates to the spherical position, method.
2. 1. An apparatus for dynamic range compression (DRC), comprising: a receiver for receiving a reconstructed Higher Order Ambisonics (HOA) audio signal representation; and The reconstructed HOA audio signal [0013] Transforming into the spatial domain based on D DSHT is the inverse discrete spherical harmonic transform (DSHT) matrix, C is a block of τ HOA samples, and WDSHT is a block of spatial samples that matches the input time granularity of the quadrature mirror filter (QMF) bank; The DRC gain value g(n,m) corresponding to the time-frequency tile (n,m) is [0014] applying the method in accordance with [Equation 15] is the spatial channel vector for time-frequency tile (n,m), and To the loudspeaker channel [0016] Rendering based on D DSHT -1 The matrix is D DSHT is the inverse matrix of matrix D, and D is the HOA rendering matrix. an audio decoder configured to perform D DSHT -1 matrix and D DSHT The queue is [Equation 17] and [Equation 18] and [1,0,0,..,0] is the element (N+1) where the first element has value 1 and all other elements have value 0. 2 row vector, N is the HOA order, [Equation 19] and the compact singular value decomposition is [Equation 20] and the new prototype matrix is [Equation 21] is calculated by [Equation 22] and the set of spherical positions [Equation 23] and related quadrature (higher order quadrature) gains [0000] is selected and the mode matrix Ψ DSHT relates to the spherical position, Device.
3. A non-transitory computer-readable storage medium having computer-executable instructions that, when executed on a computer, cause the computer to perform the method of claim 1.
Citation Information
Patent Citations
Method and apparatus for encoding multi-channel HOA audio signals for noise reduction, and method and apparatus for decoding multi-channel HOA audio signals for noise reduction
EP2688066A1
PD130040
Method and device for rendering an audio soundfield representation for audio playback
WO2014012945A1
Method for rendering multi-channel audio signals for l1 channels to a different number l2 of loudspeaker channels and apparatus for rendering multi-channel audio signals for l1 channels to a different number l2 of loudspeaker channels
WO2015007889A2