Method and device for applying dynamic range compression to high order ambisonics signal
By analyzing HOA signals to derive gain factors and applying them during decoding, the method addresses the challenge of dynamic range compression for HOA signals, achieving efficient compression while reducing computational complexity.
Patent Information
- Application Number
- JP2025039903
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-04-15
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2035-03-24
AI Technical Summary
Dynamic range compression (DRC) techniques are challenging to apply directly to higher-order ambisonics (HOA) signals due to their unique composition and lack of direct sound level representation.
The method involves analyzing the HOA signal to obtain one or more gain factors, which are then transmitted along with the signal. On the decoding side, these gain factors are applied to the HOA signal, either in the spatial domain or directly in the HOA domain, depending on the simplification mode, to achieve dynamic range compression.
This approach allows for effective dynamic range compression of HOA signals, reducing computational complexity in the simplification mode by enabling direct application of gain factors in the HOA domain, thus improving processing efficiency.
Smart Images

Figure 2025083483000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for performing dynamic range compression (DRC) on an ambisonics signal, particularly on a higher order ambisonics (HOA) signal.
Background Art
[0002] The purpose of dynamic range compression (DRC) is to reduce the dynamic range of an audio signal. A time-varying gain factor is applied to the audio signal. Typically, this gain factor depends on the amplitude envelope of the signal used to control the gain. The mapping is generally non-linear. Large amplitudes are mapped to smaller amplitudes while small sounds are often amplified. Scenarios are noisy environments, late-night listening, small speaker or mobile headphone listening.
[0003] A common concept for streaming or broadcasting audio is to generate DRC gains before transmission and apply these gains after reception and decoding. The principle of DRC use, i.e., how DRC is typically applied to an audio signal, is shown in Fig. 1 a). The signal level, typically the signal envelope, is detected and the associated time-varying gain g DRC is calculated. The gain is used to vary the amplitude of the audio signal. Fig. 1 b) shows the principle of using DRC for encoding / decoding, where the gain factor is transmitted together with the encoded audio signal. On the decoder side, a gain is applied to the decoded audio signal to reduce its dynamic range.
[0004] For 3D audio, different gains can be applied to loudspeaker channels representing different spatial positions. In this case, in order to generate a matching set of gains, these positions need to be known on the sending side. This is usually only possible for idealized conditions. In a realistic scenario, the number and arrangement of speakers can vary in many ways. This is influenced by practical circumstances rather than specifications. Higher-order ambisonics is an audio format that allows for flexible rendering. The HOA signal is composed of coefficient channels that do not directly represent sound levels. Therefore, DRC cannot simply be applied to HOA-based signals.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Non-Patent Documents
[0006]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0007] The present invention solves at least the problem of how DRC can be applied to the HOA signal.
Means for Solving the Problem
[0008] The HOA signal is analyzed to obtain one or more gain factors. In one embodiment, at least two gain factors are obtained and the analysis of the HOA signal includes a transformation to the spatial domain (iDSHT). The one or more gain factors are transmitted together with the original HOA signal. A special indicator can be transmitted to indicate whether all gain factors are equal. This is the case of the so-called simplification mode. On the other hand, in the non-simplification mode, at least two different gain factors are used. In the decoder, the one or more gains can be applied to the HOA signal (this is not essential). The user has the option to apply the one or more gains or not. The advantage of the simplification mode is that only one gain factor is used and the gain factor can be directly applied to the coefficient channels of the HOA signal in the HOA domain, so the transformation to the spatial domain and the subsequent transformation back to the HOA domain can be skipped, resulting in significantly less required calculation. In the simplification mode, the gain factor is obtained by analyzing only the zero-order coefficient of the HOA signal.
[0009] According to an embodiment of the present invention, a method for performing DRC on an HOA signal includes transforming the HOA signal to the spatial domain (by inverse DSHT), analyzing the transformed HOA signal, and obtaining from the result of the analysis a gain factor that can be used for dynamic range compression. In a further step, the obtained gain factor is multiplied with the transformed HOA signal in the (spatial domain) to obtain a gain-compressed transformed HOA signal. Finally, the gain-compressed transformed HOA signal is transformed back to the HOA domain, i.e., the coefficient domain (by DSHT) to obtain a gain-compressed HOA signal.
[0010] Furthermore, according to an embodiment of the present invention, a method of performing DRC in a simplification mode on an HOA signal includes analyzing the HOA signal and obtaining, from the result of the analysis, a gain factor that can be used for dynamic range compression. In a further step, when evaluating an indicator, the obtained gain factor is multiplied with the coefficient channels of the HOA signal (in the HOA domain) to obtain a gain-compressed and transformed HOA signal. It can also be determined that the transformation of the HOA signal can be skipped when evaluating the indicator. An indicator indicating the simplification mode, i.e., indicating that only one gain factor is used, can be set implicitly, for example, when only the simplification mode can be used due to hardware or other constraints, or explicitly, for example, when making a user selection between the simplification mode or the non-simplification mode.
[0011] Furthermore, a method of applying a DRC gain factor to an HOA signal includes receiving the HOA signal, an indicator, and a gain factor, determining that the indicator indicates a non-simplification mode, converting the HOA signal to the spatial domain (using the inverse DSHT) to obtain a converted HOA signal, multiplying the gain factor with the converted HOA signal to obtain a dynamically range-compressed and converted HOA signal, and converting the dynamically range-compressed and converted HOA signal back to the HOA domain (i.e., the coefficient domain) (using the DSHT) to obtain a dynamically range-compressed HOA signal. The gain factor can be received together with or separately from the HOA signal. Furthermore, according to an embodiment of the present invention, a method of applying a DRC gain factor to an HOA signal includes receiving the HOA signal, an indicator, and a gain factor, determining that the indicator indicates a simplification mode, and multiplying the gain factor with the HOA signal when making the determination to obtain a dynamically range-compressed HOA signal. The gain factor can be received together with or separately from the HOA signal.
[0012] An apparatus for applying a DRC gain factor to an HOA signal is disclosed in claim 11.
[0013] In one embodiment, the present invention provides a computer-readable medium having executable instructions for causing a computer to execute a method for applying a DRC gain factor to an HOA signal, the method including the steps as described above.
[0014] In one embodiment, the present invention provides a computer-readable medium having executable instructions for causing a computer to execute a method for performing DRC on an HOA signal, the method including the steps as described above.
[0015] Advantageous embodiments of the present invention are disclosed in the dependent claims, the following description, and the drawings.
Brief Description of the Drawings
[0016] Exemplary embodiments of the present invention are described with reference to the accompanying drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0017] The present invention describes how DRC can be applied to HOA. This is not straightforward as the HOA is usually not a sound field description. Figure 2 depicts the principle of this approach. On the encoding or transmitting side, as shown in Figure 2 a), the HOA signal is analyzed, the DRC gain g is calculated from the analysis of the HOA signal, the DRC gain is encoded, and is transmitted together with the encoded representation of the HOA content. This can be a multiplexed bitstream or two or more separate bitstreams.
[0018] On the decoding or receiving side, as shown in Figure 2 b), the gain g is extracted from such a (single or multiple) bitstream. After decoding of the bitstream in the decoder, the gain g is applied to the HOA signal as described later. Thereby, the gain is applied to the HOA signal, that is, generally, an HOA signal with reduced dynamic range is obtained. Finally, the HOA signal with adjusted dynamic range is rendered in the HOA renderer.
[0019] The assumptions and definitions used are described below.
[0020] The assumption is that the HOA renderer conserves energy. That is, the N3D normalized spherical harmonic functions are used and the energy of the unidirectional signal encoded within the HOA representation is maintained after rendering. For example, Patent Document 1 describes how to achieve such energy-conserving HOA rendering.
[0021] The definitions of the terms used are as follows.
[0022]
Number
Number
[0023]
Number
[0024]
Number
[0025] For all HOA truncation orders N, ideal L L = (N + 1) 2 virtual speaker grids and associated rendering matrices D L are defined. The virtual speaker positions sample the spatial area surrounding the virtual listener. Grids for N = 1 to 6 are shown in Figure 3. Here, the area associated with a certain speaker is the shaded cell. One sampling position is always related to the central speaker position (azimuth = 0, elevation = π / 2; note that the azimuth is measured from the front direction related to the listening position). The sampling positions, D L D L -1 are known on the encoder side when the DRC gain is generated. On the decoder side, D L and D L -1 need to be known in order to apply the gain values.
[0026] The generation of the DRC gain for HOA functions as follows.
[0027] The HOA signal is WL = D L is converted into a spatial region by B. By analyzing these signals, L L = (N + 1) 2 up to DRC gains g l are generated. If the content is a combination of HOA and audio objects (AO), for example, an AO signal such as a dialog track can be used for side chaining. This is shown in Fig. 4 b). When generating different DRC gain values related to different spatial areas, care must be taken so that these gains do not affect the spatial image stability on the decoder side. To avoid this, in the simplest case (so-called simplification mode), a single gain may be assigned to all L channels. This can be done by analyzing all spatial signals W, or by analyzing the zero-order HOA coefficient sample block [Number] without the need for conversion into the spatial region (Fig. 4 a). This is the same as analyzing the downmix signal of W. Further details will be described later.
[0028] In Fig. 4, the generation of DRC gains for HOA is shown. Fig. 4 a) depicts how a single gain g 1 for (a single gain group) can be derived from the zero-order HOA component [Number] The zero-order HOA component is analyzed in the DRC analysis block 41s, and a single gain g 1 is derived. The single gain g 1is separately encoded in the DRC gain encoder 42s. The encoded gain is then encoded in the encoder 43 together with the HOA signal B, and the encoder 43 outputs the encoded bitstream. Optionally, an additional signal 44 can be included in the encoding. Figure 4 b) depicts how two or more DRC gains are generated by converting the HOA representation into the spatial domain. The converted HOA signal W L is then analyzed in the DRC analysis block 41, the gain values g are extracted, and are encoded in the DRC gain encoder 42. Also here, the encoded gain is encoded in the encoder 43 together with the HOA signal B, and optionally, an additional signal 44 is included in the encoding. As an example, rear sounds (e.g., background noise) may undergo greater attenuation than sounds emitted from the front and lateral directions. This leads to (N + 1) 2 gain values within g that can be transmitted within two gain groups. Optionally, here, side chaining based on the audio object waveforms and their direction information is also possible. Side chaining means that the DRC gain for a signal is obtained from another signal. This reduces the power of the HOA signal. Ambient sounds in the HOA mix that share the same spatial sound area as the AO foreground sound can obtain a stronger attenuation gain than spatially distant sounds.
[0029] The gain values are transmitted to the receiver or decoder side.
[0030] Variables 1 through L related to a block of τ samples L =(N + 1) 2 gain values are transmitted. The gain values can be assigned to channel groups for transmission. In one embodiment, to minimize the transmission data, all equal gains are combined into one channel group. If a single gain is transmitted, it relates to all L L channels. What is transmitted is the channel group gain value glg and its number. The use of the channel group is signaled so that the receiver or decoder can correctly apply the gain value.
[0031] The gain value is applied as follows.
[0032] The receiver / decoder determines the number of the transmitted encoded gain values, decodes the related information (51), and assigns their gains to L L =(N + 1) 2 individual channels (52 - 55). If only one gain value (one channel group) is transmitted, it can be directly applied to the HOA signal as shown in a) of FIG. 5 (52) (B DRC = g 1 B). This has the advantage that decoding becomes much simpler and the required processing is significantly reduced. The reason is that matrix operations are not required and instead the gain value can be directly applied 52, for example, multiplied by the HOA coefficients.
[0033] When two or more gains are transmitted, the channel group gains are each assigned to L individual channel gains g = [g 1 , …, g L .
[0034] For a virtual regular loudspeaker lattice, the loudspeaker signal to which the DRC gain is applied is
Number
Number
Number
Number
[0035] This is more efficient in terms of the computational operations required for (N + 1) 2 <τ. That is, this solution has an advantage over the conventional solution because the decoding becomes much simpler and the required processing is significantly reduced. The reason is that matrix operations are not required; instead, the gain values can be directly applied, i.e., multiplied by the HOA coefficients, in the gain assignment block 54.
[0036] In one embodiment, a more efficient way to apply the gain matrix is to operate on the renderer matrix in the renderer matrix modification block 57
Number
Number
[0037] In summary, FIG. 5 shows various embodiments of applying the DRC to the HOA signal. In a) of FIG. 5, a single channel group gain is transmitted, decoded (51), and directly applied to the HOA coefficients (52). Then, the HOA coefficients are rendered using a normal renderer matrix (56).
[0038] In b) of FIG. 5, two or more channel group gains are transmitted and decoded (51). As a result of the decoding, (N + 1) 2A gain vector g of individual gain values is obtained. A gain matrix G is generated and applied to the blocks of HOA samples (54). Then, these are rendered using a normal rendering matrix.
[0039] In c) of FIG. 5, instead of directly applying the decoded gain matrix / gain values to the HOA signal, they are directly applied to the renderer's matrix. This is performed in the renderer matrix correction block 57. This is computationally beneficial if the DRC block size τ is larger than the number L of output channels. In this case, the HOA samples are rendered using the corrected rendering matrix (57).
[0040] The calculation of an ideal DSHT (Discrete Spherical Harmonic Transform) matrix for DRC is described below. Such a DSHT matrix is specifically optimized for use in DRC and is different from the DSHT matrices used for other purposes, such as data rate compression.
[0041] Ideal rendering and encoding matrices D related to an ideal spherical layout L and D L -1 The requirements for are derived below. Finally, these requirements are as follows: (1) The rendering matrix D L must be invertible. That is, there must exist D L -1 ; (2) The sum of the amplitudes in the spatial domain should be reflected as the zero - order HOA coefficient after conversion from the spatial domain to the HOA domain and should be preserved after subsequent conversion back to the spatial domain (amplitude requirement); (3) When converting to the HOA domain and then back to the spatial domain, the energy of the spatial signal should be preserved (energy conservation requirement).
[0042] Even for an ideal rendering layout, requirements 2 and 3 would seem to be in conflict with each other. When using a simple approach to derive a DSHT transform matrix as known from the prior art, only one or the other of requirements (2) and (3) can be satisfied without error. Satisfying one of requirements (2) and (3) without error leads to an error of more than 3 dB for the other. This typically leads to audible artifacts. A way to overcome this problem is described below.
[0043] First, L = (N + 1) 2 An ideal spherical layout at is selected. L directions of (virtual) speaker positions are given by Ω l and the associated mode matrix is
Number
Number
[0044] The first prototype rendering matrix [tilde D L is derived by the following formula.
[0045]
Number
[0046] Second, a compact singular value decomposition is performed:
Number
[0047]
Number
Number
[0048] Fourthly, at the last stage, the amplitude error for satisfying Requirement 2 is substituted: The row vector e is
Number
[0049]
Number
Number
Number
Number
[0050] Below, the detailed requirements for DRC will be described.
[0051] First, applying the same gain g 1 to L L elements in the spatial domain is equivalent to applying the gain g 1 to the HOA coefficients:
Number
[0052] Second, analyzing the sum signal in the spatial domain is equivalent to analyzing the zero - order HOA component. The DRC analyzer uses the energy and amplitude of the signal. Thus, the sum signal is related to the amplitude and energy.
[0053] HOA signal model
Number
[0054]
Number
Number
[0055] The zero-order component HOA signal needs to be the sum of the directional signals in order to reflect the correct amplitude of the sum signal.
[0056]
Number
[0057] In this mix,
Number
Number
[0058] The sum of the amplitudes in the spatial domain is given by the HOA pan matrix M L = D L Ψ e using
Number
[0059] This is
Number
Number
[0060] This also guarantees that the energy requirement for the sum signal can be satisfied. The sum of the energy in the spatial domain is [Number] given by [something], which is a good approximation [Number] resulting in [something], and the existence of an ideal symmetric speaker setup is required.
[0061] This leads to [Number] the requirement of [something]. Additionally, from the signal model, it can be concluded that for the re-encoded zero-order signal to maintain its amplitude and energy, the top row of D L -1 needs to be [1,1,1,1,…] (i.e., a vector of length L with elements of '1').
[0062] Thirdly, energy conservation is an essential requirement. The energy of the signal [Number] should be conserved after conversion to HOA and spatial rendering to a loudspeaker independent of the signal direction Ω S This leads to [Number] This leads to L from D being a rotation matrix and a diagonal gain matrix, D L = UV T diag(a)(direction (Ω S) The dependence on ) was removed because it was clear and can be achieved by modeling:
Number
Number
Number
[0063] Requirement VV T = 1 can be achieved for L≧(N + 1) 2 and can only be approximated for L <(N + 1) 2 is the case.
[0064] As an example, the case of having an ideal spherical position (HOA degree N = 1 to N = 3) is described below (Tables 1 to 3). The ideal spherical positions for further HOA degrees (N = 4 to N = 6) are described last (Tables 4 to 6). All the following positions are derived from the corrected positions published in Non-Patent Document 1. These positions and the related quadrature / cubature gains were published in Non-Patent Document 2. In these tables, the azimuth angle is measured counterclockwise from the front direction related to the listening position, and the tilt angle is measured from the z-axis with the tilt above the listening position set to 0.
[0065]
Table 1
[0066]
Table 2
[0067]
Table 3
[0068] The term numerical quadrature is often abbreviated to quadrature and is exactly synonymous with numerical integration, especially when applied to one-dimensional integration. Numerical integration in two or more dimensions is called cubature in this paper.
[0069] A typical application scenario of applying DRC gain to HOA signals is shown in FIG. 5 described above. For example, for mixed content applications such as HOA and audio objects, DRC gain application can be realized in at least two ways for flexible rendering.
[0070] FIG. 6 illustratively shows the dynamic range compression (DRC) process on the decoder side. In FIG. 6 a), DRC is applied before rendering and mixing. In FIG. 6 b), DRC is applied to the loudspeaker signal, that is, after rendering and mixing.
[0071] In FIG. 6a), the DRC gain is applied separately to the audio object and the HOA. The DRC gain is applied to the audio object in the audio object DRC block 610, and the DRC gain is applied to the HOA in the HOA DRC block 615. Here, the implementation of the block HOA DRC block 615 corresponds to one of those in FIG. 5. In FIG. 6b), a single gain is applied to all channels of the mixed signal of the rendered HOA and the rendered audio object signal. Here, spatial enhancement and attenuation are not possible. Since the speaker layout at the consumer site is not known at the time of generation at the broadcast or content production site, the relevant DRC gain cannot be generated by analyzing the rendered mixed sum signal. The DRC gain can be derived by analyzing [Number] This can be derived by analyzing. Here, y m is the sum of the zero-order HOA signal b w and the mono-downmix of S audio objects x s .
[0072] [Number] The following describes further details of the disclosed solution.
[0073] DRC for HOA Content The DRC may be applied to the HOA signal before rendering or in combination with rendering. The DRC for the HOA can be applied in the time domain or in the QMF filter bank domain.
[0074] For DRC in the time domain, the DRC decoder has (N + 1) 2 gain values [Number] gives
[0075] The DRC gain is
Number
Number
Number
[0076] In one embodiment, in order to reduce the computational load by only (N + 1) 4 operations per sample, including the rendering stage, the loudspeaker signal is
Number
[0077] When all gains g 1 , …, g (N+1)2 have the same value g drc as in the simplification mode, a single gain group is used to transmit the encoder DRC gain. In this case, it can be flagged by the DRC decoder. In this case, since no computation in the spatial filter is required, the computation is
Number
[0078] The above describes how to obtain and apply the DRC gain value. Below, the calculation of the DSHT matrix for DRC will be described.
[0079] Below, D L is renamed D DSHT and the spatial filter D DSHT and its inverse D DSHT -1 The matrices for determining are calculated as follows.
[0080] Indexed by the HOA order N from Tables 1 to 4, a set of spherical positions
Number
Number
Number
Number
Number
Number
Number
Number
[0081]
Number
Number
Number
[0082] Regarding DRC in the QMF filter bank region, the following applies.
[0083] The DRC decoder gives the gain value g 2 for all time-frequency tiles n, m for (N + 1) ch (n, m) for the spatial channels. The gain for time slot n and frequency band m is
Number
[0084] In the QMF filter bank region, multi-band DRC is applied. The processing stages are shown in Figure 7. The reconstructed HOA signal is (inverse DSHT): W DSHT = D DSHT is converted to the spatial domain by C. Here, [Number] is a block of τ HOA samples, [Number] is a block of spatial samples that matches the input time granularity of the QMF filter bank. Then, the QMF decomposition filter bank is applied.
[0085] [Number] is assumed to represent the vector of spatial channels for each time-frequency tile (n, m). Then, the DRC gain is applied: [Number] To minimize computational complexity, the DSHT and rendering to the loudspeaker channels are combined: [Number] Here, D represents the HOA rendering matrix. Then, the QMF signal can be input to the mixer for further processing.
[0086] Figure 7 shows the DRC for HOA in the QMF region combined with the rendering stage. If only a single gain group for the DRC is used, this should be flagged by the DRC decoder. This is because computational simplification is still possible. In this case, the gains in the vector g(n, m) all share the same value g DRC (n, m). The QMF filter bank can be applied directly to the HOA signal, with the gain g DRC(n,m) can be multiplied in the filter bank region.
[0087] FIG. 8 shows a DRC for HOA in a QMF region (filter region of a Quadrature Mirror Filter) combined with a rendering stage, having a computational simplification for the simple case of a single DRC gain group.
[0088] As will be apparent in view of the above, in one embodiment, the present invention relates to a method of applying a dynamic range compression gain factor to an HOA signal. The method includes receiving an HOA signal and one or more gain factors, converting the HOA signal into a spatial domain at 40 using an iDSHT with a transformation matrix obtained from the spherical positions of virtual loudspeakers and the quadrature gain q to obtain a transformed HOA signal, multiplying the gain factors with the transformed HOA signal to obtain a dynamically range-compressed transformed HOA signal, and converting the dynamically range-compressed transformed HOA signal into the original HOA domain which is a coefficient domain using a discrete spherical harmonic transform (DSHT) to obtain a dynamically range-compressed HOA signal.
[0089] Furthermore, the transformation matrix is
Number
Number
Number
Number
[0090] Furthermore, in one embodiment, the present invention relates to an apparatus for applying a DRC gain factor to an HOA signal. The apparatus includes a stage for receiving an HOA signal and one or more gain factors, a stage for converting the HOA signal into a spatial domain 40, where an iDSHT using a conversion matrix obtained from the spherical positions of the virtual loudspeakers and the quadrature gain q is used to obtain a converted HOA signal, a stage for multiplying the gain factor by the converted HOA signal to obtain a dynamically range-compressed converted HOA signal, and a stage for converting the dynamically range-compressed converted HOA signal back to the original HOA domain, which is the coefficient domain, using a discrete spherical harmonic transform (DSHT) to obtain a dynamically range-compressed HOA signal, and is adapted to execute using a processor or one or more processing elements. Furthermore, the conversion matrix is [Number] calculated according to, [Number] is [Number] is the normalized version of, where U, V are [Number] obtained from, Ψ DSHT is the transposed mode matrix of the spherical harmonic function related to the spherical positions where the virtual loudspeakers are used, e T is [Number] is the transposed version of
[0091] Furthermore, in certain embodiments, the present invention relates to a computer-readable storage medium having computer-executable instructions that, when executed on a computer, cause the computer to execute a method of applying a dynamic range compression gain factor to a higher-order ambisonics (HOA) signal. The method includes receiving the HOA signal and one or more gain factors; converting the HOA signal into the spatial domain 40 using an iDSHT with a transformation matrix obtained from the spherical positions of virtual loudspeakers and the quadrature gain q, resulting in a converted HOA signal; multiplying the gain factors by the converted HOA signal, resulting in a dynamically range-compressed converted HOA signal; and converting the dynamically range-compressed converted HOA signal back into the original HOA domain, which is the coefficient domain, using a discrete spherical harmonic transform (DSHT), resulting in a dynamically range-compressed HOA signal. Furthermore, the transformation matrix is
Number
Number
Number
Number
Number
[0092] Furthermore, in certain embodiments, the present invention relates to a method of performing DRC on an HOA signal. The method includes setting or determining a mode, which is either a simplified mode or a non-simplified mode; in the non-simplified mode, converting the HOA signal into the spatial domain using an inverse DSHT; in the non-simplified mode, analyzing the converted HOA signal and in the simplified mode, analyzing the HOA signal; obtaining from the result of the analysis one or more gain factors that can be used for dynamic range compression, where only one gain factor is obtained in the simplified mode and two or more different gain factors are obtained in the non-simplified mode; in the simplified mode, multiplying the obtained gain factor by the HOA signal to obtain a gain-compressed HOA signal, and in the non-simplified mode, multiplying the obtained gain factors by the converted HOA signal to obtain a gain-compressed converted HOA signal, and converting the gain-compressed converted HOA signal back to the original HOA domain to obtain a gain-compressed HOA signal.
[0093] In certain embodiments, the method further includes receiving an indicator indicating either the simplified mode or the non-simplified mode, and selecting the non-simplified mode if the indicator indicates the non-simplified mode and selecting the simplified mode if the indicator indicates the simplified mode. The steps of converting the HOA signal into the spatial domain and converting the dynamically range-compressed converted HOA signal back to the original HOA domain are performed only in the non-simplified mode, and in the simplified mode, the HOA signal is multiplied by only one gain factor.
[0094] In one embodiment, in the simplification mode, the method further comprises analyzing the HOA signal, and in the non-simplification mode, analyzing the converted HOA signal, and then obtaining one or more gain factors from the result of the analysis for dynamic range compression, wherein in the non-simplification mode, two or more different gain factors are obtained, and in the simplification mode, only one gain factor is obtained, and multiplying the obtained gain factor by the HOA signal to obtain a gain-compressed HOA signal in the simplification mode, and in the non-simplification mode, multiplying the two or more obtained gain factors by the converted HOA signal to obtain the gain-compressed converted HOA signal, and in the non-simplification mode, converting the HOA signal into the spatial domain uses the inverse DSHT.
[0095] In one embodiment, the HOA signal is divided into frequency sub-bands, and the gain factor(s) are obtained separately for each frequency sub-band and applied using the individual gain for each sub-band. In one embodiment, the steps of analyzing the HOA signal (or the converted HOA signal), obtaining one or more gain factors, multiplying the obtained gain factors by the HOA signal (or the converted HOA signal), and converting the gain-compressed converted HOA signal back to the original HOA domain are applied separately for each frequency sub-band using the individual gain for each sub-band. Note that the sequential order of dividing the HOA signal into frequency sub-bands and converting the HOA signal into the spatial domain can be interchanged and / or the sequential order of combining the sub-bands and converting the gain-compressed converted HOA signal back to the original HOA domain can be interchanged. These interchanges can be done independently of each other.
[0096] In one embodiment, the method further comprises transmitting the converted HOA signal together with the obtained gain factors and the number of these gain factors before the step of multiplying the gain factors.
[0097] In one embodiment, the transformation matrix is the mode matrix ΨDSHT and calculated from the corresponding quadrature gain, the mode matrix Ψ DSHT is
Number
Number
[0098] In one embodiment, the HOA signal B is converted into a spatial domain to obtain a converted HOA signal W DSHT and the converted HOA signal W DSHT is multiplied by a gain value diag(g)
Number
Number
Number
[0099] In one embodiment, assuming that N is the HOA order and τ is the DRC block size, if at least (N + 1) 2 < τ, then the method further
Number
Number
[0100] In an embodiment, assuming that L is the number of output channels and τ is the DRC block size, when at least L < τ, the method further includes applying the gain matrix G
Number
[0101] In an embodiment, the present invention relates to a method of applying a DRC gain factor to an HOA signal. The method includes receiving an HOA signal together with an index and one or more gain factors, where the index indicates either a simplified mode or a non-simplified mode, and only one gain factor is received when the index indicates the simplified mode; selecting either the simplified mode or the non-simplified mode according to the index; in the simplified mode, multiplying the gain factor by the HOA signal to obtain a dynamically range-compressed HOA signal; in the non-simplified mode, converting the HOA signal to the spatial domain to obtain a converted HOA signal, multiplying the gain factor by the converted HOA signal to obtain a dynamically range-compressed converted HOA signal, and converting the dynamically range-compressed converted HOA signal back to the original HOA domain to obtain a dynamically range-compressed HOA signal, including the step and
[0102] Furthermore, in certain embodiments, the present invention relates to an apparatus for performing DRC on an HOA signal. The apparatus includes steps of setting or determining a mode which is either a simplification mode or a non-simplification mode, in the non-simplification mode, converting the HOA signal into a spatial domain using an inverse DSHT, in the non-simplification mode, analyzing the converted HOA signal and in the simplification mode, analyzing the HOA signal, obtaining one or more gain factors usable for dynamic range compression from the result of the analysis, where only one gain factor is obtained in the simplification mode and two or more different gain factors are obtained in the non-simplification mode, in the simplification mode, multiplying the obtained gain factor by the HOA signal to obtain a gain-compressed HOA signal, and in the non-simplification mode, multiplying the obtained gain factors by the converted HOA signal to obtain a gain-compressed converted HOA signal, and converting the gain-compressed converted HOA signal back to the original HOA domain to obtain a gain-compressed HOA signal, and has a processor or one or more processing elements adapted to perform these steps.
[0103] In certain embodiments for the non-simplification mode only, an apparatus for performing DRC on an HOA signal includes a processor or one or more processing elements adapted to perform steps of analyzing the converted HOA signal, obtaining one or more gain factors usable for dynamic range compression from the result of the analysis, multiplying the obtained gain factors by the converted HOA signal to obtain a gain-compressed converted HOA signal, and converting the gain-compressed converted HOA signal back to the original HOA domain to obtain a gain-compressed HOA signal. In certain embodiments, the apparatus further includes a transmission unit for transmitting the HOA signal together with the obtained gain factor(s) before multiplying the obtained gain factor(s).
[0104] Furthermore, in certain embodiments, the present invention relates to an apparatus for applying a DRC gain factor to an HOA signal. The apparatus receives the HOA signal together with an indicator and one or more gain factors, where the indicator indicates either a simplified mode or a non-simplified mode, and only one gain factor is received if the indicator indicates the simplified mode, selects either the simplified mode or the non-simplified mode according to the indicator, in the simplified mode multiplies the gain factor with the HOA signal to obtain a dynamically range-compressed HOA signal, and in the non-simplified mode, converts the HOA signal into a spatial domain to obtain a converted HOA signal, multiplies the gain factor with the converted HOA signal to obtain a dynamically range-compressed converted HOA signal, and converts the dynamically range-compressed converted HOA signal back to the original HOA domain to obtain a dynamically range-compressed HOA signal, and has a processor or one or more processing elements adapted to perform the steps.
[0105] Furthermore, in certain embodiments, the apparatus further has a transmission unit for transmitting the HOA signal together with the obtained gain factor before multiplying the obtained gain factor. In certain embodiments, the HOA signal is divided into frequency sub-bands, and analyzing the converted HOA signal, obtaining a gain factor, multiplying the obtained gain factor with the converted HOA signal, and converting the gain-compressed converted HOA signal back to the original HOA domain are each applied separately to each frequency sub-band using the individual gain for each sub-band.
[0106] In certain embodiments of an apparatus for applying a DRC gain factor to an HOA signal, the HOA signal is divided into a plurality of frequency sub-bands, obtaining one or more gain factors, multiplying the obtained gain factor with the HOA signal or the converted HOA signal, and in the non-simplified mode, converting the gain-compressed converted HOA signal back to the original HOA signal are each applied separately to each frequency sub-band using the individual gain for each sub-band.
[0107] Furthermore, in certain embodiments where only the non-simplified mode is used, the present invention relates to an apparatus for applying a DRC gain factor to an HOA signal. The apparatus receives an HOA signal together with a gain factor, converts the HOA signal into a spatial domain (using iDSHT) to obtain a converted HOA signal, multiplies the gain factor with the converted HOA signal to obtain a dynamically range-compressed converted HOA signal, and converts the dynamically range-compressed converted HOA signal back into the original HOA domain (i.e., coefficient domain) (using DSHT) to obtain a dynamically range-compressed HOA signal, and has a processor or one or more processing elements adapted to perform these steps.
[0108] The following Tables 4 to 6 list the spherical positions of the virtual loudspeakers for HOA of order N where N = 4, 5 or 6.
[0109] Although the basic novel features of the present invention have been illustrated, described, and pointed out in its preferred embodiments, it will be understood that various omissions, substitutions, and changes in the described apparatus and method may be made by those skilled in the art without departing from the spirit of the present invention. It is clearly intended that any combination of elements that perform substantially the same function in substantially the same way to achieve the same result is within the scope of the present invention. Substitutions of elements from one described embodiment for those of another described embodiment are also fully intended and contemplated.
[0110] The present invention has been described purely by way of example, and it will be understood that detailed modifications can be made without departing from the scope of the present invention. Each feature disclosed in this description and (where appropriate) the claims and drawings can be provided independently or in any suitable combination. The features can be implemented, as appropriate, in hardware, software, or a combination of both.
[0111]
Table 4
[0112]
Table 5
[0113]
Table 6
Claims
1. 1. A method for dynamic range compression (DRC) comprising: receiving a reconstructed Higher Order Ambisonics (HOA) audio signal representation; The reconstructed HOA audio signal [0010] Transforming to the spatial domain based on D DSHT is the inverse discrete spherical harmonic transform (DSHT) matrix, C is a block of τ HOA samples, and W is a block of spatial samples that match the input time granularity of the quadrature mirror filter (QMF) bank; The DRC gain value g(n,m) corresponding to the time-frequency tile (n,m) is [0025] applying the method according to [0030] is the vector of spatial channels for time-frequency tile (n,m), and To the loudspeaker channel [0045] Rendering based on D DSHT -1 The matrix is D DSHT where D is the inverse of the matrix D, and D is the HOA rendering matrix. D DSHT -1 Matrices and D DSHT The queue is [0050] and [006] and [1,0,0,..,0] is the element (N+1) whose first element has value 1 and whose other elements are all 0. 2 row vector, N is the HOA order, [0070] and the compact singular value decomposition is [0080] and the new prototype matrix is [0097] It is calculated by [0089] and the set of spherical positions ##EQU00011## and related quadrature (higher order quadrature) gains ##EQU00012## is selected and the mode matrix Ψ DSHT relates to the spherical position, method.
2. 1. An apparatus for dynamic range compression (DRC), comprising: a receiver for receiving a reconstructed Higher Order Ambisonics (HOA) audio signal representation; and The reconstructed HOA audio signal ##EQU00013## Transforming to the spatial domain based on D DSHT is the inverse discrete spherical harmonic transform (DSHT) matrix, C is a block of τ HOA samples, and W is a block of spatial samples that match the input time granularity of the quadrature mirror filter (QMF) bank; The DRC gain value g(n,m) corresponding to the time-frequency tile (n,m) is ##EQU00014## applying the method according to ##EQU00015## is the vector of spatial channels for time-frequency tile (n,m), and To the loudspeaker channel ##EQU00016## Rendering based on D DSHT -1 The matrix is D DSHT is the inverse matrix of matrix D, and D is the HOA rendering matrix. an audio decoder configured to perform D DSHT -1 Matrices and D DSHT The queue is ##EQU00017## and [0018] and [1,0,0,..,0] is the element (N+1) whose first element has value 1 and whose other elements are all 0. 2 row vector, N is the HOA order, [0019] and the compact singular value decomposition is [0020] and the new prototype matrix is ##EQU00021## It is calculated by [0022] and the set of spherical positions [0023] and related quadrature (higher order quadrature) gains [0024] is selected and the mode matrix Ψ DSHT relates to the spherical position, Device.
3. A non-transitory computer readable storage medium having computer executable instructions that, when executed on a computer, cause the computer to perform the method of claim 1.
Citation Information
Patent Citations
Method and apparatus for encoding multi-channel HOA audio signals for noise reduction, and method and apparatus for decoding multi-channel HOA audio signals for noise reduction
EP2688066A1
Method and device for rendering an audio soundfield representation for audio playback
WO2014012945A1
PD130040
Method for rendering multi-channel audio signals for l1 channels to a different number l2 of loudspeaker channels and apparatus for rendering multi-channel audio signals for l1 channels to a different number l2 of loudspeaker channels
WO2015007889A2