Method and apparatus for applying dynamic range compression to higher-order ambisonic signals
The method efficiently applies dynamic range compression to higher-order ambisonics signals by analyzing and transforming HOA signals with gain coefficients, optimizing computation and rendering, addressing the challenge of flexible speaker setups in existing technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2026-02-06
- Publication Date
- 2026-04-28
AI Technical Summary
Dynamic range compression (DRC) methods are not straightforward for higher-order ambisonics (HOA) signals due to their non-sound field description, requiring flexible rendering that varies with speaker placement and number, which existing technologies fail to address effectively.
The method involves analyzing the HOA signal to obtain gain coefficients, transforming it into the spatial domain, applying the gain factors, and then transforming back to the HOA domain, with options for simplified and non-simplified modes to optimize computation, and includes transmitting these coefficients with the signal for decoder application.
This approach allows efficient dynamic range compression of HOA signals with reduced computational complexity and flexibility in rendering, maintaining audio quality by directly applying gain factors without matrix operations, suitable for various speaker setups.
Smart Images

Figure 2026071353000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and apparatus for performing dynamic range compression (DRC) on ambisonic signals, particularly on higher-order ambisonics (HOA) signals. [Background technology]
[0002] The purpose of dynamic range compression (DRC) is to reduce the dynamic range of an audio signal. A time-varying gain factor is applied to the audio signal. Typically, this gain factor depends on the amplitude envelope of the signal used to control the gain. Its mapping is generally nonlinear; large amplitudes are mapped to smaller amplitudes, while quiet sounds are often amplified. Scenarios include noisy environments, listening at night, and listening through small speakers or mobile headphones.
[0003] A common concept for streaming or broadcasting audio is to generate DRC gain before transmission and apply these gains after reception and decoding. The principle of DRC use, i.e., how DRC is typically applied to an audio signal, is shown in Figure 1a). The signal level, typically the signal envelope, is detected and the time-varying gain g involved is applied. DRC The gain is calculated. The gain is used to change the amplitude of the audio signal. Figure 1b) shows the principle of using DRC for encoding / decoding, where the gain factor is transmitted along with the encoded audio signal. On the decoder side, the gain is applied to the decoded audio signal to reduce its dynamic range.
[0004] For 3D audio, different gains can be applied to loudspeaker channels representing different spatial positions. In this case, these positions must be known at the source in order to generate a matching set of gains. This is usually only possible under idealized conditions. In practical cases, the number and placement of speakers can vary in many ways, influenced more by practical considerations than by specifications. Higher-order ambisonics is an audio format that allows for flexible rendering. HOA signals consist of coefficient channels that do not directly represent sound levels. Therefore, DRC cannot be simply applied to HOA-based signals. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] International Publication No. 2015 / 007889A (PD130040) [Non-patent literature]
[0006] [Non-Patent Document 1] Jörg Fliege, “Integration nodes for the sphere,” 2010, online, accessed 2010-10-05, http: / / www.mathematik.uni-dortmund.de / lsx / research / projects / fliege / nodes / nodes.html [Non-Patent Document 2] J¨org Fliege and Ulrike Maier, “A two-stage approach for computing cubature formula for the sphere”, Technical report, Fachbereich Mathematik, Universitat Dortmund, 1999 [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] The present invention solves at least the problem of how DRC can be applied to the HOA signal. [Means for solving the problem]
[0008] The HOA signal is analyzed to obtain one or more gain coefficients. In one embodiment, at least two gain coefficients are obtained, and the analysis of the HOA signal includes a transformation to the spatial domain (iDSHT). The one or more gain coefficients are transmitted along with the original HOA signal. A special index may be transmitted to indicate whether all gain coefficients are equal. This is the so-called simplified mode. On the other hand, in the non-simplified mode, at least two different gain coefficients are used. In the decoder, the one or more gains can be applied to the HOA signal (this is not mandatory). The user has the option of whether or not to apply the one or more gains. The advantage of the simplified mode is that only one gain factor is used, and since the gain factor can be applied directly to the coefficient channel of the HOA signal in the HOA domain, the transformation to the spatial domain and the subsequent transformation back to the HOA domain can be skipped, thus significantly reducing the computation required. In the simplified mode, the gain factor is obtained by analyzing only the zero-order coefficients of the HOA signal.
[0009] According to one embodiment of the present invention, a method for performing DRC on an HOA signal includes transforming the HOA signal into the spatial domain (by inverse DSHT), analyzing the transformed HOA signal, and obtaining a gain factor from the results of the analysis that can be used for dynamic range compression. In a further step, the obtained gain factor is multiplied by the transformed HOA signal (in the spatial domain) to obtain a gain-compressed transformed HOA signal. Finally, the gain-compressed transformed HOA signal is transformed back into the HOA domain, i.e., the coefficient domain (by DSHT), to obtain a gain-compressed HOA signal.
[0010] Furthermore, according to one embodiment of the present invention, a method for performing DRC in simplified mode on an HOA signal includes analyzing the HOA signal and obtaining a gain factor usable for dynamic range compression from the results of the analysis. In a further step, during the evaluation of the index, the obtained gain factor is multiplied by the coefficient channel of the HOA signal (in the HOA region) to obtain a gain-compressed transformed HOA signal. During the evaluation of the index, it may also be determined that the transformation of the HOA signal can be skipped. An index indicating simplified mode, i.e., that only one gain factor is used, can be set implicitly, for example, when only simplified mode is available due to hardware or other constraints, or explicitly, for example, when the user selects either simplified mode or non-simplified mode.
[0011] Furthermore, a method for applying a DRC gain factor to an HOA signal includes receiving the HOA signal, an index, and a gain factor, determining that the index indicates a non-simplification mode, converting the HOA signal to the spatial domain (using an inverse DSHT) to obtain a converted HOA signal, multiplying the converted HOA signal by the gain factor to obtain a dynamically range-compressed converted HOA signal, and converting the dynamically range-compressed converted HOA signal back to the HOA domain (i.e., the coefficient domain) (using a DSHT) to obtain a dynamically range-compressed HOA signal. The gain factor can be received together with the HOA signal or separately. Furthermore, according to one embodiment of the present invention, a method for applying a DRC gain factor to an HOA signal includes receiving the HOA signal, an index, and a gain factor, determining that the index indicates a simplification mode, and in making the determination, multiplying the HOA signal by the gain factor to obtain a dynamically range-compressed HOA signal. The gain factor can be received together with the HOA signal or separately.
[0012] An apparatus for applying a DRC gain factor to a HOA signal is disclosed in claim 11.
[0013] In one embodiment, the present invention provides a computer-readable medium having executable instructions for causing a computer to perform a method of applying a DRC gain factor to an HOA signal, the method including the steps described above.
[0014] In one embodiment, the present invention provides a computer-readable medium having executable instructions for causing a computer to perform a method of performing DRC on an HOA signal, the method comprising the steps described above.
[0015] Advantageous embodiments of the present invention are disclosed in the dependent claims, the following description and drawings. [Brief explanation of the drawing]
[0016] Exemplary embodiments of the present invention are described with reference to the accompanying drawings. [Figure 1] This diagram shows the general principles of DRC applied to audio. [Figure 2] This figure shows a general method for applying DRC to HOA-based signals according to the present invention. [Figure 3] This figure shows a spherical speaker grid for N=1 to N=6. [Figure 4] This figure shows the generation of the DRC gain for HOA. [Figure 5] This figure shows the application of DRC to the HOA signal. [Figure 6] This figure shows the dynamic range compression process performed on the decoder side. [Figure 7] This figure shows the DRC for HOA in the QMF region, combined with the rendering stage. [Figure 8] This figure shows the DRC for HOA in the QMF region, combined with the rendering stage, for a simple case of a single DRC gain group. [Modes for carrying out the invention]
[0017] This invention describes how DRC can be applied to HOA. This is not straightforward because HOA is not a sound field description. Figure 2 illustrates the principle of this approach. On the encoding or transmitting side, as shown in Figure 2a), the HOA signal is analyzed, the DRC gain g is calculated from the analysis of the HOA signal, the DRC gain is encoded and transmitted together with the encoded representation of the HOA content. This can be a multiplexed bitstream or two or more separate bitstreams.
[0018] At the decoding or receiving end, the gain g is extracted from such a bitstream(one or more) as shown in Figure 2b). After decoding the bitstream in the decoder, the gain g is applied to the HOA signal as described later. This applies the gain to the HOA signal, i.e., generally results in an HOA signal with a reduced dynamic range. Finally, the dynamic-range-adjusted HOA signal is rendered in the HOA renderer.
[0019] The following describes the assumptions and definitions used.
[0020] The assumption is that the HOA renderer conserves energy. That is, the energy of a unidirectional signal, encoded within the HOA representation using N3D normalized spherical harmonics, is preserved after rendering. For example, Patent Document 1 describes how to achieve this energy-conserving HOA rendering.
[0021] The definitions of the terms used are as follows:
[0022]
number
Number
[0023]
Number
[0024]
Number
[0025] For all HOA truncation orders N, the ideal L L =(N+1) 2 The number of virtual speaker grids and the associated rendering matrix D L The following is defined. The virtual speaker position samples the spatial area surrounding the virtual listener. The grid for N=1 to 6 is shown in Figure 3. Here, the area related to a particular speaker is the shaded cell. One sampling position is always related to the central speaker position (azimuth angle = 0, tilt angle = π / 2; note that the azimuth angle is measured from the front direction related to the listening position). Sampling position, D L , D L -1 The DRC gain is known on the encoder side when it is generated. On the decoder side, the gain value is applied to D L and D L -1 This needs to be known.
[0026] The generation of the DRC gain for HOA works as follows:
[0027] HOA signal is W L =D L These signals are transformed into spatial domains by B. L =(N+1) 2 DRC gain g up to individuall This is generated. If the content is a combination of HOA and audio objects (AO), the AO signal, such as a dialogue track, may be used for side chaining. This is shown in Figure 4b). When generating different DRC gain values related to different spatial areas, care must be taken to ensure that these gains do not affect the spatial image stability on the decoder side. To avoid this, in the simplest case (the so-called simplified mode), a single gain may be assigned to all L channels. This can be done by analyzing all spatial signals W, or by using zero-order HOA coefficient sample blocks.
number
[0028] Figure 4 shows the generation of DRC gains for HOA. Figure 4a) shows how a single gain g1 (for a single gain group) is used for the zero-order HOA component.
number
[0029] The gain value is transmitted to the receiver or decoder.
[0030] A variable number of 1 to L related to a block of τ samples L =(N+1) 2 Multiple gain values are transmitted. Gain values can be assigned to groups of channels for transmission. In one embodiment, all equal gains are combined into one group of channels to minimize the transmitted data. When a single gain is transmitted, it is L L It relates to all individual channels. What is transmitted is the channel group gain value g. lg and the number thereof. The use of channel groups is signaled so that the receiver or decoder can correctly apply the gain value.
[0031] The gain value is applied as follows:
[0032] The receiver / decoder determines the number of encoded gain values transmitted, decodes the relevant information (51), and sets their gains to L L=(N+1) 2 Assign to individual channels (52-55). If only one gain value (one channel group) is transmitted, it can be applied directly to the HOA signal as shown in Figure 5a) (52)(B DRC =g1B). This has the advantage of making decoding much simpler and requiring significantly less processing. The reason is that matrix operations are not required, and instead the gain value can be directly applied and multiplied by, for example, the HOA coefficient.
[0033] When two or more gains are transmitted, the channel group gains are each L channel gains g = [g1, ..., g L It will be assigned to ].
[0034] For a virtual regular loudspeaker grid, the loudspeaker signal to which DRC gain is applied is
number
number
number
number
[0035] This is (N+1) 2Because of <τ, this solution is more efficient in terms of the computational work required. In other words, decoding becomes much simpler and the processing required is significantly less, so this solution has advantages over conventional solutions. The reason is that matrix operations are not required, and instead the gain value can be applied directly in the gain allocation block 54, i.e., multiplied by the HOA coefficient.
[0036] In one embodiment, a more efficient way of applying the gain matrix is to modify the renderer matrix in the renderer matrix modification block 57.
number
number
[0037] In summary, Figure 5 shows various embodiments of applying DRC to the HOA signal. In Figure 5a), a single channel group gain is transmitted and decoded (51) and applied directly to the HOA coefficients (52). The HOA coefficients are then rendered using a standard rendering matrix (56).
[0038] In Figure 5b), two or more channel group gains are transmitted and decoded (51). As a result of decoding, (N+1) 2 A gain vector g of n gain values is obtained. A gain matrix G is generated and applied to the block of HOA samples (54). These are then rendered using the usual rendering matrix.
[0039] In Figure 5c), instead of directly applying the decoded gain matrix / gain value to the HOA signal, it is applied directly to the renderer matrix. This is done in renderer matrix modification block 57. This is computationally beneficial if the DRC block size τ is greater than the number of output channels L. In this case, the HOA samples are rendered using the modified rendering matrix (57).
[0040] The following describes the calculation of an ideal DSHT (Discrete Spherical Harmonic Transform) matrix for DRC. Such a DSHT matrix is specifically optimized for use in DRC and differs from DSHT matrices used for other purposes, such as data rate compression.
[0041] The ideal rendering and encoding matrix D related to the ideal spherical layout. L and D L -1 The requirements for this are derived below. Ultimately, these requirements are as follows: (1) Rendering matrix D L It must be reversible. That is, D L -1 It must exist; (2) The sum of amplitudes in the spatial domain should be reflected as zero-order HOA coefficients after the transformation from the spatial domain to the HOA domain, and should be preserved after the subsequent transformation back to the spatial domain (amplitude requirement); (3) When converting to the HOA domain and then back to the spatial domain, the energy of the spatial signal should be conserved (energy conservation requirement).
[0042] Even for an ideal rendering layout, requirements 2 and 3 seem contradictory. When using a simple approach to derive the DSHT transformation matrix as known from prior art, only one or the other of requirements (2) and (3) can be satisfied without error. Satisfying one of requirements (2) and (3) without error leads to an error of more than 3 dB for the other. This typically leads to audible artifacts. A method for overcoming this problem is described below.
[0043] First, L = (N+1) 2 An ideal spherical layout is selected. The L directions of the (virtual) speaker positions are Ω l The mode matrix is given by and the related mode matrix is
number
number
[0044] First prototype rendering matrix [D with tilde] L The formula is derived by the following equation.
[0045]
number
[0046] Secondly, a compact singular value decomposition is performed:
number
[0047]
number
number
[0048] Fourth, in the final stage, the amplitude error necessary to satisfy requirement 2 is substituted: row vector e is
number
[0049]
number
number
number
number
[0050] The following describes the detailed requirements for the DRC.
[0051] Firstly, L has a value g1 in the spatial domain. L Applying the same gain to multiple instances is equivalent to applying the gain g1 to the HOA coefficient:
number
[0052] Secondly, analyzing a sum signal in the spatial domain is equivalent to analyzing the zero-order HOA component. The DRC analyzer uses the signal's energy and amplitude. Thus, the sum signal is related to amplitude and energy.
[0053] HOA signal model
number
[0054]
number
number
[0055] The zero-order component HOA signal needs to be the sum of the directional signals in order to reflect the correct amplitude of the sum signal.
[0056]
number
[0057] In this mix,
number
number
[0058] The sum of amplitudes in the spatial domain is the HOA pan matrix M L =D L Ψ e Using
number
[0059] this is,
number
number
number
[0060] This also guarantees that the energy requirement for the sum signal can be satisfied. The sum of the energies in the spatial domain is
Number
Number
[0061] This is
Number
[0062] Thirdly, energy conservation is an essential requirement. The energy of the signal
Number
Number
[0063] Requirement VV T = 1 can be achieved for L ≧ (N + 1) 2 and can only be approximated for L < (N + 1) 2 is only possible.
[0064] As an example, the case of having an ideal position on the sphere (HOA degrees N = 1 to N = 3) is described below (Tables 1 - 3). The ideal spherical positions for further HOA degrees (N = 4 to N = 6) are described last (Tables 4 - 6). The following positions are all derived from the corrected positions published in Non - Patent Document 1. These positions and the related quadrature / cubature gains were published in Non - Patent Document 2. In these tables, the azimuth angle is measured counterclockwise from the front direction related to the listening position, and the tilt angle is measured from the z - axis with the tilt above the listening position set to 0.
[0065] [Table 1] Table 1: a) Spherical positions of virtual loudspeakers for HOA degree N = 1, b) Rendering matrix obtained as a result of spatial transformation (DSHT).
[0066] [Table 2] Table 2: a) Spherical position of the virtual loudspeaker for HOA order N=2, b) Rendering matrix obtained as a result of spatial transformation (DSHT).
[0067] [Table 3] Table 3: a) Spherical position of the virtual loudspeaker for HOA order N=2, b) Rendering matrix obtained as a result of spatial transformation (DSHT).
[0068] The term numerical quadrature is often abbreviated to quadrature and is synonymous with numerical integration, especially when applied to one-dimensional integrals. Numerical integration in two or more dimensions is referred to as higher-order quadrature (cubature) in this paper.
[0069] A typical application scenario for applying DRC gain to an HOA signal is shown in Figure 5, mentioned above. For mixed content applications such as HOA and audio objects, DRC gain application can be implemented in at least two ways for flexible rendering.
[0070] Figure 6 illustrates the dynamic range compression (DRC) process on the decoder side. In Figure 6a), DRC is applied before rendering and mixing. In Figure 6b), DRC is applied to the loudspeaker signal, i.e., after rendering and mixing.
[0071] In Figure 6a), DRC gain is applied separately to the audio object and the HOA. The DRC gain is applied to the audio object in the audio object DRC block 610, and the DRC gain is applied to the HOA in the HOA DRC block 615. Here, the implementation of block HOA DRC block 615 corresponds to one of those in Figure 5. In Figure 6b), a single gain is applied to all channels of the mixed signal of the rendered HOA and the rendered audio object signals. Here, spatial emphasis and attenuation are not possible. The DRC gains involved cannot be generated by analyzing the sum signal of the rendered mix, since the speaker layout of the consumer site is unknown at the time of generation at the broadcast or content creation site. The DRC gains are,
number
[0072]
number
[0073] DRC for HOA content DRC may be applied to the HOA signal before rendering, or combined with rendering. DRC for HOA can be applied in the time domain or the QMF filter bank domain.
[0074] For DRC in the time domain, the DRC decoder is (N+1) depending on the number of HOA coefficient channels of the HOA signal c. 2 Individual gain values
number
[0075] The DRC gain is
number
number
number
[0076] In one embodiment, (N+1) per sample 4 To reduce the computational load by only one operation, the loudspeaker signal is processed, including the rendering stage.
number
[0077] Like in the simplified mode, all gains g1, ..., g (N+1)2 The same value g drc If this is the case, a single gain set was used to transmit the encoder DRC gain. In this case, it can be flagged by the DRC decoder. In this case, no calculation in the spatial filter is required, so the calculation is
number
[0078] The above describes how to obtain and apply the DRC gain value. Below, we will discuss the calculation of the DSHT matrix for DRC.
[0079] In the following, D L is D DSHT The name will be changed to Spatial Filter D. DSHT And its inverse D DSHT -1 The matrix for determining this is calculated as follows:
[0080] The set of spherical positions is indexed by the HOA degree N from Tables 1 to 4.
number
number
number
number
number
number
number
number
[0081]
number
number
number
[0082] The following applies to DRC in the QMF filter bank region:
[0083] The DRC decoder is (N+1) 2 Gain value g for all time-frequency tiles n,m for each spatial channel. ch (n,m) is given. The gain for time slot n and frequency band m is
number
[0084] Multiband DRC is applied in the QMF filter bank region. The processing steps are shown in Figure 7. The reconstructed HOA signal is (inverse DSHT):W DSHT =D DSHT It is transformed into a spatial domain by C. Here,
number
number
[0085]
number
number
number
[0086] Figure 7 shows the DRC for HOA in the QMF domain combined with the rendering stage. If only a single gain group for the DRC is used, this should be flagged by the DRC decoder, as this also allows for computational simplification. In this case, the gains in the vector g(n,m) are all the same value g. DRC (n,m) is shared. The QMF filter bank can be applied directly to the HOA signal, with gain g DRC(n,m) can be multiplied in the filter bank region.
[0087] Figure 8 shows the DRC for HOA in the QMF domain (the filter domain of a Quadrature Mirror Filter) combined with the rendering stage, with a computational simplification for the simple case of a single DRC gain group.
[0088] As has become clear in view of the above, in one embodiment, the present invention relates to a method for applying a dynamic range compression gain factor to an HOA signal. The method includes the steps of: receiving an HOA signal and one or more gain factors; converting the HOA signal to a spatial domain 40, wherein an iDSHT using a transformation matrix obtained from the spherical position and quadrature gain q of a virtual loudspeaker is used to obtain a converted HOA signal; multiplying the gain factor with the converted HOA signal to obtain a dynamically range compressed converted HOA signal; and converting the dynamically range compressed converted HOA signal to the original HOA domain which is the coefficient domain using a discrete spherical harmonic function transform (DSHT) to obtain a dynamically range compressed HOA signal.
[0089] Furthermore, the transformation matrix is
number
number
number
number
number
[0090] Furthermore, in one embodiment, the present invention relates to an apparatus for applying a DRC gain factor to an HOA signal. The apparatus has a processor or one or more processing elements adapted to perform the following steps: receiving an HOA signal and one or more gain factors; transforming the HOA signal into a spatial domain 40, wherein an iDSHT using a transformation matrix obtained from the spherical position and quadrature gain q of a virtual loudspeaker is used to obtain a transformed HOA signal; multiplying the gain factor with the transformed HOA signal to obtain a dynamically range compressed transformed HOA signal; and transforming the dynamically range compressed transformed HOA signal into the original HOA domain which is the coefficient domain using a discrete spherical harmonic function transform (DSHT) to obtain a dynamically range compressed HOA signal. Furthermore, the transformation matrix is
number
number
number
number
number
[0091] Furthermore, in one embodiment, the present invention relates to a computer-readable storage medium having a computer-executable instruction that, when executed on a computer, causes the computer to perform a method of applying a dynamic range compression gain factor to a higher-order ambisonics (HOA) signal. The method includes the steps of: receiving an HOA signal and one or more gain factors; transforming the HOA signal into a spatial domain 40, wherein an iDSHT using a transformation matrix obtained from the spherical position and quadrature gain q of a virtual loudspeaker is used to obtain a transformed HOA signal; multiplying the gain factor by the transformed HOA signal to obtain a dynamically range compressed transformed HOA signal; and transforming the dynamically range compressed transformed HOA signal into the original HOA domain which is the coefficient domain using a discrete spherical harmonic function transform (DSHT) to obtain a dynamically range compressed HOA signal. Furthermore, the transformation matrix is
number
number
number
number
number
[0092] Furthermore, in one embodiment, the present invention relates to a method for performing DRC on an HOA signal. The method includes the steps of: setting or determining a mode, which is either a simplified mode or a non-simplified mode; converting the HOA signal to the spatial domain using an inverse DSHT in the non-simplified mode; analyzing the converted HOA signal in the non-simplified mode, and analyzing the HOA signal in the simplified mode; obtaining one or more gain factors from the results of the analysis that can be used for dynamic range compression, wherein only one gain factor is obtained in the simplified mode, and two or more different gain factors are obtained in the non-simplified mode; multiplying the obtained gain factor with the HOA signal in the simplified mode to obtain a gain-compressed HOA signal, multiplying the obtained gain factor with the converted HOA signal in the non-simplified mode to obtain a gain-compressed converted HOA signal; and converting the gain-compressed converted HOA signal back to the original HOA domain to obtain a gain-compressed HOA signal.
[0093] In one embodiment, the method further includes the steps of receiving an index indicating either a simplified mode or a non-simplified mode, selecting the non-simplified mode if the index indicates a non-simplified mode, and selecting the simplified mode if the index indicates a simplified mode, wherein the steps of converting the HOA signal to the spatial domain and converting the dynamically range-compressed converted HOA signal to the original HOA domain are performed only in the non-simplified mode, and in the simplified mode, the HOA signal is multiplied by a single gain factor.
[0094] In one embodiment, in the simplification mode, the method further comprises analyzing the HOA signal, and in the non-simplification mode, analyzing the converted HOA signal, and then obtaining one or more gain factors that can be used for dynamic range compression from the result of the analysis, wherein in the non-simplification mode, two or more different gain factors are obtained, and in the simplification mode, only one gain factor is obtained, and multiplying the obtained gain factor by the HOA signal to obtain a gain-compressed HOA signal in the simplification mode, and in the non-simplification mode, multiplying the obtained two or more gain factors by the converted HOA signal to obtain the gain-compressed converted HOA signal, and in the non-simplification mode, converting the HOA signal into the spatial domain uses an inverse DSHT.
[0095] In one embodiment, the HOA signal is divided into frequency sub-bands, and the gain factor(s) are obtained separately for each frequency sub-band and applied using the individual gain for each sub-band. In one embodiment, the steps of analyzing the HOA signal (or the converted HOA signal), obtaining one or more gain factors, multiplying the obtained gain factor by the HOA signal (or the converted HOA signal), and converting the gain-compressed converted HOA signal into the original HOA domain are applied separately for each frequency sub-band using the individual gain for each sub-band. Note that the sequential order of dividing the HOA signal into frequency sub-bands and converting the HOA signal into the spatial domain can be interchanged and / or the sequential order of synthesizing the sub-bands and converting the gain-compressed converted HOA signal into the original HOA domain can be interchanged. These interchanges can be done independently of each other.
[0096] In one embodiment, the method further comprises transmitting the converted HOA signal together with the obtained gain factors and the number of these gain factors before the step of multiplying the gain factors.
[0097] In one embodiment, the transformation matrix is a mode matrix ΨDSHT and calculated from the corresponding quadrature gain, the mode matrix Ψ DSHT is
Number
Number
[0098] In one embodiment, the HOA signal B is converted into a spatial domain to obtain a converted HOA signal W DSHT and the converted HOA signal W DSHT is multiplied by a gain value diag(g)
Number
Number
Number
[0099] In one embodiment, assuming that N is the HOA order and τ is the DRC block size, if at least (N + 1) 2 < τ, then the method further includes
Number
number
[0100] In one embodiment, where L is the number of output channels and τ is the DRC block size, the method further extends the gain matrix G to the case where L < τ.
number
[0101] In one embodiment, the present invention relates to a method for applying a DRC gain factor to an HOA signal. The method includes the steps of: receiving an HOA signal together with an index and one or more gain factors, wherein the index indicates either a simplified mode or a non-simplified mode, and if the index indicates a simplified mode, only one gain factor is received; selecting either a simplified mode or a non-simplified mode according to the index; in the simplified mode, multiplying the gain factor with the HOA signal to obtain a dynamic range compressed HOA signal; in the non-simplified mode, converting the HOA signal to the spatial domain to obtain a converted HOA signal; multiplying the gain factor with the converted HOA signal to obtain a dynamic range compressed converted HOA signal; and converting the dynamic range compressed converted HOA signal back to the original HOA domain to obtain a dynamic range compressed HOA signal.
[0102] Furthermore, in one embodiment, the present invention relates to an apparatus for performing DRC on an HOA signal. The apparatus has a processor or one or more processing elements adapted to perform the following steps: setting or determining a mode which is either a simplified mode or a non-simplified mode; in the non-simplified mode, converting the HOA signal to the spatial domain using an inverse DSHT; in the non-simplified mode, analyzing the converted HOA signal, and in the simplified mode, analyzing the HOA signal; obtaining one or more gain factors from the results of the analysis that can be used for dynamic range compression, wherein in the simplified mode, only one gain factor is obtained, and in the non-simplified mode, two or more different gain factors are obtained; and in the simplified mode, multiplying the obtained gain factor with the HOA signal to obtain a gain-compressed HOA signal, and in the non-simplified mode, multiplying the obtained gain factor with the converted HOA signal to obtain a gain-compressed converted HOA signal, and converting the gain-compressed converted HOA signal back to the original HOA domain to obtain a gain-compressed HOA signal.
[0103] In one embodiment for non-simplification mode only, the apparatus performing DRC on an HOA signal has a processor or one or more processing elements adapted to perform steps such as analyzing the converted HOA signal, obtaining one or more gain factors available for dynamic range compression from the results of the analysis, multiplying the obtained gain factors by the converted HOA signal to obtain a gain-compressed converted HOA signal, and converting the gain-compressed converted HOA signal back to the original HOA region to obtain a gain-compressed HOA signal. In one embodiment, the apparatus further has a transmitting unit for transmitting the HOA signal together with the obtained gain factors(s) before multiplying by the obtained gain factors(s).
[0104] Furthermore, in certain embodiments, the present invention relates to an apparatus for applying a DRC gain factor to an HOA signal. The apparatus receives the HOA signal together with an indicator and one or more gain factors, where the indicator indicates either a simplified mode or a non-simplified mode, and only one gain factor is received when the indicator indicates the simplified mode. It then selects either the simplified mode or the non-simplified mode according to the indicator. In the simplified mode, it multiplies the gain factor with the HOA signal to obtain a dynamically range-compressed HOA signal. In the non-simplified mode, it converts the HOA signal into the spatial domain to obtain a converted HOA signal, multiplies the gain factor with the converted HOA signal to obtain a dynamically range-compressed converted HOA signal, and then converts the dynamically range-compressed converted HOA signal back to the original HOA domain to obtain a dynamically range-compressed HOA signal. It has a processor or one or more processing elements adapted to perform these steps.
[0105] Furthermore, in certain embodiments, the apparatus further has a transmission unit for transmitting the obtained HOA signal together with the obtained gain factor before multiplying the obtained gain factor. In certain embodiments, the HOA signal is divided into frequency sub-bands, and analyzing the converted HOA signal, obtaining the gain factor, multiplying the obtained gain factor with the converted HOA signal, and converting the gain-compressed converted HOA signal back to the original HOA domain are each applied separately to each frequency sub-band using the individual gain for each sub-band.
[0106] In certain embodiments of an apparatus for applying a DRC gain factor to an HOA signal, the HOA signal is divided into a plurality of frequency sub-bands, and obtaining one or more gain factors, multiplying the obtained gain factor with the HOA signal or the converted HOA signal, and in the non-simplified mode, converting the gain-compressed converted HOA signal back to the original HOA signal are each applied separately to each frequency sub-band using the individual gain for each sub-band.
[0107] Furthermore, in a certain embodiment in which only the non-simplification mode is used, the present invention relates to an apparatus for applying a DRC gain factor to an HOA signal. The apparatus relates to an apparatus for applying a DRC gain factor to an HOA signal. The apparatus has a processor or one or more processing elements adapted to perform the steps of: receiving an HOA signal together with a gain factor; converting the HOA signal to a spatial domain (using an iDSHT) to obtain a converted HOA signal; multiplying the gain factor with the converted HOA signal to obtain a dynamically range-compressed converted HOA signal; and converting the dynamically range-compressed converted HOA signal to the original HOA domain (i.e., coefficient domain) (using a DSHT) to obtain a dynamically range-compressed HOA signal.
[0108] Tables 4 through 6 below list the spherical positions of virtual loudspeakers for HOA of order N at N=4, 5, or 6.
[0109] While the fundamental novel features of the present invention have been illustrated, described, and pointed out in their application to preferred embodiments, it will be understood that various omissions, substitutions, and modifications in the described apparatus and methods may be made by those skilled in the art in terms of the form and details of the disclosed devices and their operation, without deviating from the spirit of the invention. It is clearly intended that any combination of elements that perform substantially the same function and achieve the same results in substantially the same manner is within the scope of the invention. Substitution of elements from one described embodiment in another described embodiment is also fully intended and considered.
[0110] This invention is described purely as an example, and it will be understood that modifications to the details can be made without departing from the scope of the invention. Each feature disclosed in this description and (where appropriate) in the claims and drawings may be provided independently or in any suitable combination. Features may be implemented in hardware, software, or a combination thereof, as appropriate.
[0111] [Table 4] Table 4: Spherical position of virtual loudspeakers for HOA order N=4.
[0112] [Table 5] Table 5: Spherical position of virtual loudspeakers for HOA order N=5.
[0113] [Table 6] Table 6: Spherical position of virtual loudspeakers for HOA order N=6.
Claims
1. A method for applying dynamic range compression (DRC) to a higher-order ambisonic (HOA) signal in the time domain, the method being: The HOA signal is given one or more DRC gains g drc of [Math 1] This is the step of applying based on the above, where c is the vector of one time sample of the HOA coefficient of the HOA signal. [Math 2] And, [Math 3] And its inverse D L -1 This is a matrix related to the Discrete Spherical Harmonic Transform (DSHT) optimized for the DRC purpose, with steps; The steps include rendering the HOA signal to obtain a loudspeaker signal for audio playback by one or more loudspeakers. including, method.
2. A non-temporary computer-readable storage medium having a computer-executable instruction that, when executed by a computer, causes the computer to perform the method described in claim 1.
3. A device for applying dynamic range compression (DRC) to a higher-order ambisonic (HOA) signal in the time domain, wherein the device: The HOA signal is given one or more DRC gains g drc of [Math 4] This is the step of applying based on the above, where c is the vector of one time sample of the HOA coefficient of the HOA signal. [Math 5] And, [Math 6] And its inverse D L -1 This is a matrix related to the Discrete Spherical Harmonic Transform (DSHT) optimized for the DRC purpose, with steps; The steps include rendering the HOA signal to obtain a loudspeaker signal for audio playback by one or more loudspeakers. Having a processor configured to perform Device.
Citation Information
Patent Citations
PD130040
Method for rendering multi-channel audio signals for l1 channels to a different number l2 of loudspeaker channels and apparatus for rendering multi-channel audio signals for l1 channels to a different number l2 of loudspeaker channels
WO2015007889A2