Reverberation level adjustment

JP2026027262A5Pending Publication Date: 2026-07-30TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Filing Date
2025-10-15
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing audio rendering technologies face challenges in maintaining the correct balance between direct and reverberant sound components when the configuration of reverberation units changes, rendering non-omnidirectional and non-point sound sources, and varying definitions of direct and reverberant components, leading to suboptimal audio rendering results.

Method used

A method for rendering sound sources that involves deriving relative gains based on reverberation parameters, directivity patterns, and time limits, and adjusting audio signals to maintain the desired balance between direct and reverberant components, even with changes in reverberation unit configurations or source types.

Benefits of technology

Ensures accurate balance between direct and reverberant sound components across different configurations and source types, improving audio rendering quality in augmented and virtual reality systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for rendering a sound source.SOLUTION: The method 1000 comprises: Step s1002 of receiving an input audio signal corresponding to a sound source; Step s1004 of receiving a reverberation parameter indicating a target energy ratio of the audio reverberation component of the sound source; Deriving one or more of: (i) a relative gain associated with a first directivity pattern of the sound source; (ii) a relative gain associated with a first reference distance of the sound source; (iii) a relative gain associated with a first configuration of the reverberation unit; (iv) a relative gain related to a first time limit of the reverberation component; Further Step s1008 of generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and any one or more of the derived relative gains (I) - (iv).SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a method and apparatus for adjusting reverberation levels. [Background technology]

[0002] The MPEG-I Audio Standard for Augmented Reality (e.g., Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR)) Audio defines parameters for an acoustic environment (virtual or real) that specify the relative levels of (late) reverberation in the acoustic environment.

[0003] The parameters may comprise forming a desired ratio between the level (e.g., energy level) of the direct sound component (or the emitted energy of the sound source) and the level of (late) reverberation of the sound source when it is rendered in the acoustic environment. An audio renderer receiving the parameters may be able to render sound sources located in the acoustic environment such that a listener receives, for each sound source, the correct balance between the direct and reverberant sound components of the rendered audio at all possible listening positions in the acoustic environment. The audio renderer may achieve this by appropriately setting the (relative) levels of the processing units that generate the (late) reverberation.

[0004] An example of this parameter is the so-called direct-to-reverberant energy ratio (DRR), or its inverse, the reverberant-to-direct energy ratio (RDR), determined at a predetermined fixed distance (e.g., 1 m) from a sound source of a given type (e.g., an omnidirectional point source).

[0005] A general description of these measurements (i.e., parameters) and the conceptual method for determining them and using them to calibrate the audio renderer and set up the reverberation unit is provided in the Additional Information section below (ISO / IEC JTC1 / SC29 / WG6 Document Number: M57352, July 2021). Summary of the Invention [Problem to be solved by the invention]

[0006] Currently, the following issues exist:

[0007] Changing the reverberation unit configuration - The conceptual procedure disclosed in the Additional Information section below for obtaining measurements (RDRs obtained at specific fixed locations from specific types of sound sources) and using the measurements to calibrate the audio renderer should, in principle, only be performed once to correctly set the gain of the renderer's reverberation unit (i.e., so that the calibration results in the desired balance between the direct sound component and the reverberant sound component as specified by the received values ​​of the measurements). Therefore, there is no need to recalibrate the gain of the reverberation unit for each individual sound source, scene, and / or acoustic environment. However, this does not apply when the configuration of the reverberation unit itself is changed so that its input / output relationship changes, e.g., by loading a different room impulse response or setting a new reverberation time. In such a situation, the gain of the reverberation unit needs to be recalibrated for the new configuration.

[0008] Rendering of non-omnidirectional sources - Since the measurements are defined to be determined for a specific fixed source type, i.e., an omnidirectional point source, and the gain calibration of the renderer's reverberation units is usually done for sources of the same type, it may not be possible to achieve the correct balance between the direct sound component and the (late) reverberant part at all listening positions when rendering sources that have non-omnidirectional and / or non-point-like behavior (i.e., do not have the 1 / r law of distance attenuation for point sources).

[0009] Varying definitions of direct and reverberant components - According to the measurements defined in the Additional Information section below, different choices are possible for setting the temporal boundaries between direct and reverberant components. Making different choices on the authoring side (which determines the values ​​of the measurements) and the renderer side may lead to suboptimal results.

[0010] Rendering a sound source using a reference distance with a different distance attenuation function - The MPEG-I Audio standard allows for the setting of a so-called "reference distance". The "reference distance" is the distance from the sound source at which the distance attenuation of the sound source is defined as 1. Rendering a sound source with a "reference distance" that has a value different from the default value specified in the standard may result in an inaccurate balance between the direct and reverberant components of the sound source's audio. [Means for solving the problem]

[0011] Thus, in one aspect, there is provided a method for rendering a sound source, the method comprising receiving an input audio signal corresponding to the sound source and receiving reverberation parameters indicative of a target energy ratio between a direct sound component of the rendered audio of the sound source and a reverberant sound component of the rendered audio of the sound source, the method further comprising deriving a relative gain associated with a first configuration of reverberation units, where the relative gain is relative to a reference configuration of reverberation units, and generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and the derived relative gain.

[0012] In another aspect, a method for rendering a sound source is provided. The method includes receiving an input audio signal corresponding to the sound source and receiving reverberation parameters indicating a target energy ratio between a direct sound component of the rendered audio of the sound source and a reverberant sound component of the rendered audio of the sound source. The method further includes obtaining a directivity pattern of the sound source and deriving a relative power level of the sound source based on the obtained directivity pattern, where the relative power level is relative to a power level of an omnidirectional sound source. The method further includes generating an adjusted audio signal using the received input audio signal, the reverberation parameters, and the derived relative power level.

[0013] In another aspect, a method for rendering a sound source is provided, the method including receiving an input audio signal corresponding to the sound source and receiving reverberation parameters indicating a target energy ratio between a direct sound component of rendered audio of the sound source and a reverberant sound component of the rendered audio of the sound source, the method further including obtaining a first variable indicating an upper time limit for the direct sound component and a second variable indicating a lower time limit for the reverberant sound component, and generating an adjusted audio signal using the received input audio signal, the reverberation parameters, the obtained first variable, and the obtained second variable.

[0014] In another aspect, a method for rendering a sound source is provided, the method including receiving an input audio signal corresponding to the sound source and receiving reverberation parameters indicating a target energy ratio between a direct sound component of rendered audio for the sound source and a reverberant sound component of the rendered audio for the sound source, the method further including deriving a relative gain corresponding to a first associated reference distance of the sound source, where the relative gain is relative to a default reference distance, and generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and the derived relative gain.

[0015] In another aspect, there is provided a computer program comprising instructions that, when executed by a processing circuit, cause the processing circuit to perform the method described above.

[0016] In another aspect, an apparatus is provided that includes a memory and processing circuitry coupled to the memory, the apparatus configured to perform the above-described method.

[0017] In another aspect, a method for rendering a sound source is provided, comprising receiving an input audio signal corresponding to the sound source, receiving reverberation parameters indicating a target energy ratio for reverberant components of audio from the sound source, and deriving one or more of (i) a relative gain associated with a first directivity pattern of the sound source, (ii) a relative gain associated with a first reference distance of the sound source, (iii) a relative gain associated with a first configuration of reverberation units, and (iv) a relative gain associated with a first time limit of the reverberant components. The method further comprises generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and one or more of the derived relative gains (i)-(iv), wherein the relative gain associated with the first directivity pattern is associated with a reference directivity pattern, the relative gain associated with the first reference distance is associated with a default reference distance, the relative gain associated with the first configuration is associated with a reference configuration of reverberation units, and the relative gain associated with the first time limit is associated with a second time limit of the reverberant components.

[0018] In another aspect, a method for rendering a sound source is provided, the method including receiving an input audio signal corresponding to the sound source and receiving reverberation parameters indicating a target energy ratio for reverberant components of the audio of the sound source, obtaining a directivity pattern of the sound source, deriving a relative power level of the sound source based on the obtained directivity pattern, where the relative power level is relative to a power level of an omnidirectional sound source, and generating an adjusted audio signal using the received input audio signal, the reverberation parameters, and the derived relative power level.

[0019] In another aspect, a method for rendering a sound source is provided, the method comprising receiving an input audio signal corresponding to the sound source, receiving reverberation parameters indicative of a target energy ratio for reverberant components of audio of the sound source, deriving a relative gain corresponding to a first associated reference distance of the sound source, and generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and the derived relative gain.

[0020] In another aspect, a method for rendering a sound source is provided, the method comprising receiving an input audio signal corresponding to the sound source, receiving reverberation parameters indicative of a target energy ratio for reverberant components in rendered audio of the sound source, obtaining a variable indicative of a lower time limit for the reverberant components, and generating an adjusted audio signal using the received input audio signal, the reverberation parameters, and the obtained variable.

[0021] In another aspect, a method for rendering a sound source is provided, the method comprising receiving an input audio signal corresponding to the sound source and receiving reverberation parameters indicating a target energy ratio for reverberant components of rendered audio of the sound source, the method further comprising deriving a relative gain associated with a first configuration of reverberation units, where the relative gain is relative to a reference configuration of reverberation units, and generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and the derived relative gain.

[0022] In a different aspect, there is provided a computer program comprising instructions that, when executed by a processing circuit, cause the processing circuit to perform a method according to at least one of the above-described embodiments.

[0023] In a different aspect, there is provided an apparatus comprising a processing circuit and a memory, the memory including instructions executable by the processing circuit, the apparatus operable to perform a method according to at least one of the above-described embodiments.

[0024] advantage

[0025] According to an embodiment of the present disclosure, an efficient method is provided for providing a desired balance between the direct sound component of the rendered audio of a sound source and the (late) reverberant sound component of the rendered audio of the sound source.

[0026] To achieve the above advantages, some embodiments of the present disclosure provide a convenient way to determine and set the correct relative gains for reverberation units in an audio renderer when there are changes to the configuration of the reverberation units that change the input-output relationships of the units.

[0027] In other embodiments, methods are provided for adapting the rendering of reverberation of sound sources, such as non-omnidirectional and / or non-point sound sources, so that the obtained balance between direct and reverberant sound components is correct at all listening positions for such sound sources.

[0028] In a further embodiment, a method is provided for obtaining reverberation parameters and associated variables indicating lower limits of the reverberant components from the authoring side, such that the obtained variables are used to modify the reverberation parameters, thereby generating an adjusted audio signal in the audio renderer. [Brief explanation of the drawings]

[0029] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate various embodiments.

[0030] [Figure 1a] A diagram showing the components of audio that are rendered to the user.

[0031] [Figure 1b] FIG. 10 is a diagram showing the end time of a direct sound component and the start time of a reverberation sound component.

[0032] [Figure 2] FIG. 1 illustrates a system according to some embodiments.

[0033] [Figure 3] FIG. 1 illustrates a process according to some embodiments.

[0034] [Figure 4] FIG. 1 illustrates a process according to some embodiments.

[0035] [Figure 5] FIG. 1 illustrates a process according to some embodiments.

[0036] [Figure 6] FIG. 1 illustrates a process according to some embodiments.

[0037] [Figure 7] FIG. 1 illustrates a process according to some embodiments.

[0038] [Figure 8] 1 illustrates an apparatus according to some embodiments.

[0039] [Figure 9] FIG.

[0040] [Figure 10] FIG. 1 illustrates a process according to some embodiments.

[0041] [Figure 11] FIG. 1 illustrates a process according to some embodiments.

[0042] [Figure 12] FIG. 1 illustrates a process according to some embodiments.

[0043] [Figure 13] FIG. 1 illustrates a process according to some embodiments.

[0044] [Figure 14] FIG. 1 illustrates a process according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0045] 1a illustrates how a sound source 102 may be rendered to a user 104. As shown in FIG. 1a, the sound source 102 may be rendered to the user 104 via a direct sound component 112, an early reflection sound component 114, and a late reflection sound component 116.

[0046] As discussed above and in the Additional Information section below, the definition of the RDR (or DRR) measurement allows for different selections of the temporal boundaries for the direct and reverberant components used to calculate the measurement. The measurement obtained with any such selection of the temporal boundaries may generally be referred to as the “direct-to-reverberant energy ratio.” For example, with some selections of the temporal boundaries, the direct-to-reverberant energy ratio may refer to the ratio of the energy of the direct component 112 to the energy of all reflected components 114 and 116, while with some other selections of the temporal boundaries, the direct-to-reverberant energy ratio may refer to the ratio of the energy of the direct component 112 to the energy of only the diffuse component of the room impulse response (i.e., only the late reflected component 116). In the latter case, the resulting RDR (or DRR) measurement may be referred to as the “direct-to-diffuse” energy ratio. For yet other selections of the temporal boundaries, the reverberant components used in the RDR (or DRR) measurement may include some, but not all, reflected components that arrive before the diffuse portion of the room impulse response begins. For simplicity, the terms "direct-to-reverberant energy ratio," "direct-to-diffuse energy ratio," RDR, and DRR are used synonymously in this disclosure and are commonly referred to as "energy ratios" unless otherwise stated.

[0047] Thus, any variation of the energy ratio, including those described above and their reciprocals, is applicable to embodiments of the present disclosure. For example, equation (1) in the Additional Information section below accommodates any use of variations in the "energy ratio" by having separate parameters t1 and t2, where t1 indicates the end of the direct sound component of the room impulse response and t2 indicates the start of the "reverberant" portion of the room impulse response. Parameter t2 may be selected such that it marks either the start of the non-direct sound component of the room impulse response, the start of the diffuse component of the room impulse response (i.e., the "direct-to-diffuse" ratio definition), or some other selected start time of the "reverberant" component (e.g., any time between the start of the non-direct sound component and the start of the diffuse component of the room impulse response). Parameters t1 and t2 are shown in FIG. 1b.

[0048] The termination of the direct sound component, denoted t1, may be selected so that the direct sound component includes only the direct sound peak of the impulse response. Alternatively, the termination of the direct sound component may be selected so that the direct sound component includes not only the direct sound peak but also very early reflections. The termination of the direct sound component may also be selected to include some very early reflections, as these very early reflections perceptually merge with the direct sound component. The direct sound component and its very early reflections are perceived as one acoustic event from the direction of the direct sound, where the very early reflections result in a higher perceived level of the direct sound.

[0049] Although embodiments of the present disclosure are described using an RDR determined at a predetermined distance from a predetermined type of sound source, the embodiments are equally applicable to alternative related measurements resulting from different choices for parameters t1 and t2 (such as the diffuse-to-direct ratio). Similarly, embodiments of the present disclosure are applicable to measurements that use an alternative measure to the direct sound energy measure in the denominator of the RDR measure (provided in Equation 1 in the Additional Information section below), such as total emitted source energy (i.e., resulting in a reverberant / diffuse-to-emitted source energy ratio). In other words, in some embodiments, instead of the RDR (or DRR, direct-to-reverberant energy ratio, or direct-to-diffuse energy ratio), a different energy ratio for the reverberant component of the sound source's audio may be used, such as the ratio between the total energy emitted by the sound source and the energy corresponding to the reverberant component of the sound source's audio (e.g., the energy corresponding to the late reflections 116, or the energy corresponding to a combination of the early and late reflections 114 and 116). In summary, the embodiment applies to a "family" of metrics that indicate the energy ratio between the reverberant components of a sound source and the non-reverberant components of the rendered audio.

[0050] 2 illustrates a system 200 according to some embodiments. The system 200 may be used to render a sound source. The system 200 may include a direct sound component unit 202, an early reflection sound component unit 204, a late reflection sound component unit 206, and a combiner 208. In this disclosure, either the late reflection sound component unit 206 or the combination of the early reflection sound component unit 204 and the late reflection sound component unit 206 is referred to as a reverberation unit. Similarly, the late reflection sound component or the combination of the early reflection sound component and the late reflection sound component is referred to as a reverberation sound component.

[0051] The direct sound component unit 202 may receive an input audio signal and generate a direct sound component signal based on the received input audio signal. Similarly, the early reflection sound component unit 204 may generate an early reflection sound component signal based on the received input audio signal, and the late reflection sound component unit 206 may generate a late reflection sound component signal based on the received input audio signal. The combiner 208 may be configured to combine the three generated signals to obtain an output audio signal. In some embodiments, the combination of the three generated signals may be a weighted combination of the three generated signals. For example, s output =w direct s direct +w early-reflected s early-reflected +w late-reflected s late-reflected .

[0052] Different methods and / or systems can be implemented in the reverberation unit for reverberation generation processing. Examples of such methods and / or systems include delay networks (simulating reverberation processing using delay lines, filters, and feedback connections), convolution algorithms (convolving the dry input signal with a recorded, approximated, or simulated room impulse response (RIR)), computational acoustics (simulating sound propagation in a specified geometry), and virtual analog models (simulating electromechanical or electrical devices (tapes, plates, springs) previously used to generate reverberation effects).

[0053] Some embodiments described below are directed to a method for rendering a sound source using a relative gain, where the relative gain is a gain relative to a reverberation unit. For example, the relative gain may be a gain applied within the reverberation unit. In such an example, by applying the gain within the reverberation unit, the reverberation unit generates an adjusted reverberated audio signal.

[0054] In another example, the relative gain may be a gain applied to an input audio signal provided to a reverberation unit. By applying the gain to the input audio signal, a modified input audio signal is provided to the reverberation unit, which then generates an adjusted reverberated audio signal. In a different example, the relative gain may be a gain applied to an output audio signal output from the reverberation unit. In such an example, applying the gain to the output audio signal generates an adjusted reverberated audio signal.

[0055] Below we provide a detailed explanation of how the relative gains are derived in each of the different embodiments.

[0056] 1. Adaptation to changes in reverberation unit configuration

[0057] The method for calibrating an audio renderer, described in Section 4 of the Additional Information section below, need only be performed once for an audio renderer, as long as there are no changes to the configuration of the reverberation units included in the audio renderer. That is, one calibration procedure performed for one (e.g., any) acoustic environment may be sufficient to set the relative gains (e.g., energy gains) of the reverberation units to provide a desired balance between direct and reverberant sound components for a sound source rendered at any position in any acoustic environment with the reverberation units, as specified by given values ​​of the RDR parameters.

[0058] Even in scenarios where the value of the RDR parameters is changed, it is simple to adjust the relative gain of the reverberation units (included in the audio renderer) to achieve the desired balance between the direct and reverberant components as specified by the given RDR parameters. In particular, changing the relative gain of the reverberation units by the same amount as the value of RDR changes will achieve the desired result.

[0059] However, when the configuration of the reverberation unit is changed so that the input-output relationship of the reverberation unit changes, it is not easy to achieve the desired balance between the direct sound component and the reverberant sound component. Examples of such changes that change the input-output relationship of the reverberation unit are any one or a combination of loading a different room impulse response (RIR), setting a different reverberation time (RT, RT60), setting different absorption characteristics, etc.

[0060] In such cases, the output of the reverberation unit in response to a given input audio signal will generally differ in terms of temporal and / or spectral aspects (e.g., the temporal length of the reverberation response, the temporal density of reflections, the temporal spectral shape of the response, etc.) Also, the change in the input-output relationship of the reverberation unit will result in a different output level of the reverberation unit in response to a given input signal, which will result in a different ratio of direct to reverberant sound components at the listener position than before the change (assuming the rendering of the direct sound component remains unchanged).

[0061] One way to solve the above problem is to calibrate the reverberation units for the new configuration of the reverberation units, as described in Section 4 of the Additional Information section below. However, this has the drawback that the calibration procedure takes time to perform, which may be undesirable or even unacceptable in real-time rendering applications.

[0062] If the reverberation unit has a finite set of configurations to use, calibration can be done offline in advance for each of the configurations, and the corresponding resulting relative gains of the reverberation unit for each configuration can be stored, retrieved, and applied as needed.

[0063] By calibrating each of the available configurations, the derived relative gain of the reverberation unit for each of the reverberation unit configurations indicates both the input-output level relationship with the specific configuration of the reverberation unit and the level of the reverberation unit relative to other rendering units (particularly the direct sound rendering unit).

[0064] However, calibrating the reverberation units for each configuration is not strictly necessary and may result in the relative gains of the reverberation units containing redundant information. In fact, the only information that may be needed to derive the necessary corrections to the gains of the reverberation units may be the change in output level of the reverberation units in response to a given input signal.

[0065] Since this relationship is purely a property of the reverberation unit, this change in output level can be determined completely independently of other components of the renderer system.

[0066] Therefore, according to an embodiment of the present disclosure, a method 300 shown in FIG. 3 may be used to derive the change in output level.

[0067] The method begins at step s302, which includes determining a first output level (eg, energy level) of the reverberation unit corresponding to a reference input signal when the reverberation unit is operating in a reference configuration.

[0068] The reference input audio signal can be any audio signal suitable for determining the input-output energy level relationship of a reverberation unit, such as a stationary white noise signal, a Dirac pulse, a sine wave sweep signal, a pseudo-random noise (Maximum-Length Sequence (MLS)) signal, etc.

[0069] Referring again to FIG. 3, step s304 includes determining a second output level (e.g., energy level) of the reverberation unit corresponding to the same reference input signal when the reverberation unit is operating in a modified configuration that differs from the reference configuration.

[0070] Step s306 includes determining the difference between the first output level and the second output level.

[0071] Step s308 includes obtaining a desired balance between the direct sound component and the reverberant sound component based on the determined difference, for example, the desired balance can be obtained by applying the determined difference as an additional gain or attenuation to the original gain of the reverberation unit.

[0072] If the audio renderer has been calibrated to a reference configuration of reverberation units, as described in Section 4 of the Additional Information section below, and the output level differences are obtained using the above method, the desired balance between the direct and reverberation sound components can be obtained (achieved) for each of the different configurations by simply compensating the relative gains of the reverberation units by the differences.

[0073] The following scenario illustrates how the method 300 shown in FIG. 3 may be used.

[0074] Assume that an audio renderer including a reverberation unit has been calibrated with a reverberation unit configuration A including a room impulse response RIR_A and RDR parameters having a value of -6 dB (i.e., a reverberation-to-direct energy ratio of 0.25 on the linear energy ratio scale).

[0075] Also assume that information about a new scene is received, which includes an acoustic environment with a specified desired RDR value of -10 dB (i.e., 0.1 on the linear energy ratio scale). To achieve the desired ratio between direct and reverberant sound components, the relative gain of the reverberation unit needs to be reduced by 4 dB (2.5 times in terms of energy, or 1.6 times in terms of linear gain).

[0076] Further assume that a configuration change of the reverberation units from configuration A to configuration B is triggered (by metadata contained in the scene information or by some other trigger), i.e. a new room impulse response RIR_B is loaded.

[0077] Here, the change in the output level of the reverberation units due to the configuration change is determined using the process described above, and the change is determined to be, for example, +3 dB (i.e., the new configuration with RIR_B results in an output level 3 dB higher than the old configuration A with RIR_A). This means that the relative gain of the reverberation units needs to be reduced by 3 dB in order to maintain the correct ratio between the direct sound part and the reverberant sound part.

[0078] Alternatively, in the specific case of changing between two RIRs, the change in input-output relationship may be derived directly from calculating the energy of the RIRs, for example by integrating the square of the RIR and comparing the resulting energies for the two RIRs. In some situations, this method may be more efficient than the process described above, as the energy of each RIR may be stored with the RIRs as metadata that can be directly used by the audio renderer to make any necessary adjustments to the relative gains of the reverberation units.

[0079] Referring back to the example above, combining the two changes (a change to the scene with an RDR parameter value of -10 dB, and a change in the reverberation unit configuration from RIR_A to RIR_B), the relative gain of the reverberation units should be reduced by 4 + 3 = 7 dB to obtain the desired ratio of direct sound components (specified by the RDR parameter values) to reverberant sound components (specified by the RDR parameter values) in the overall rendered output.

[0080] Another example of a configuration change is when the reverberation time (RT60) of a reverberation unit is changed. If a shorter RT60 is set for a reverberation unit, this typically results in a lower output level for the reverberation unit (because the impulse response contains less energy). In this case, the same process as above can be used to determine the change in output level of the reverberation unit.

[0081] However, in this particular case, there is also a simpler and more direct way to determine the change. From statistical acoustic formulas, it can be derived that on a linear scale, the reverberant-to-direct energy ratio is proportional to the RT60. More specifically, if the RT60 is halved compared to the RT60 value used to calibrate the audio renderer, the resulting reverberant-to-direct energy ratio will also be halved on a linear scale (i.e., reduced by 3 dB on the decibel scale).

[0082] This means that there is a direct approximation between the change in RT60 and the change in the relative gain of the reverberation units required to obtain the correct ratio between the direct and reverberant parts. Thus, in this example, the relative gain of the reverberation units must be increased by 3 dB (10 log(2)) to achieve the correct balance.

[0083] Similar direct relationships exist between changes in other parameters of the reverberation unit and their corresponding output levels and can be advantageously used, so that instead of using the method 300 shown in FIG. 3 to determine changes in output level, a deterministic relationship can be used instead.

[0084] Although RIR and RT60 are used as examples, the method 300 shown in FIG. 3 for determining the change in output level due to a change in configuration applies equally to any other type of change to the configuration of the reverberation unit.

[0085] Also, while the above example discusses two separate changes to the reverberation unit configuration (RIR and RT60), there may be several simultaneous changes when switching between configurations. In such cases, the method 300 shown in Figure 3 can conveniently combine the effects of all those changes together to yield a single value for the overall change in the reverberation unit output level.

[0086] 2. Rendering non-omnidirectional and / or non-point sound sources

[0087] The RDR parameters and methods described in the Additional Information section below for deriving the RDR parameters and using them to calibrate an audio renderer are determined based on the assumption that the sound source is an omnidirectional point source, i.e., a source that radiates sound equally in all directions and has a distance attenuation function that is inversely proportional to distance (e.g., f(r)=1 / r).

[0088] If the sound source being rendered has these characteristics of an omnidirectional point source, then the calibrated audio renderer will produce audio with the correct balance between the direct sound portion of the audio and the reverberant sound portion of the audio for any combination of source position and listener position in any acoustic environment (unless the configuration of the reverberation units is changed as described above).

[0089] However, many sound sources rendered in VR and AR systems do not have the characteristics of an omnidirectional point source. More specifically, many sound sources rendered in VR and AR systems have a defined non-omnidirectional radiation pattern (e.g., specified in metadata accompanying the sound source). Rendering such sound sources without special measurements may not provide the correct balance between the direct portion of the audio and the reverberant portion of the audio at some (or all) listening positions.

[0090] In a real-world situation, when there are omnidirectional and directional sound sources with equal source amplitudes (or equal source signal levels with respect to the audio rendering system) and located at the same position in the same room, the level of the direct sound portion perceived at the listener position for each of the sound sources can be easily determined by simply looking at the value of the respective directional pattern of each of the sound sources in the direction of the listener position.

[0091] For an omnidirectional sound source, the value of the sound source's directional pattern is assumed to be 1 in all directions, while for a directional sound source, the value of the sound source's directional pattern may be assumed to have a value between 0 and 1 in any direction (different conventions for defining and / or normalizing the directional pattern may be used in embodiments of the present disclosure).

[0092] The level of the (late) reverberant components of sound for any sound source is essentially constant throughout the room in which the source is located and is determined by the total radiated power of the source. This means that a directional source with a directivity pattern normalized to 1 will have a lower total radiated power and less reverberant sound in the room compared to the total radiated power of a normalized omnidirectional source with the same source amplitude. More generally (i.e., not limited to directivity patterns normalized to 1), directional sources will generally generate different levels of reverberation in a room compared to an omnidirectional source with the same source amplitude.

[0093] When a VR audio rendering system has omnidirectional and directional sources with equal source signal levels (e.g., the same audio input signal is sent to each of the two sources), rendering the direct sound component of the audio is relatively straightforward for both types of sources. Furthermore, rendering the direct sound component of a directional source can be achieved using a simple scaling of the amplitude of the audio input signal by the value of the directional pattern in the direction of the listener.

[0094] However, accurate rendering of the reverberant components of directional sound sources requires more careful processing: the RDR parameters, methods for determining the RDR parameters, and methods for calibrating the audio renderer described in the Additional Information section below are all based on the assumption that the sound source is omnidirectional (and / or point source), so the reverberant components of directional (and / or non-point source) sound sources may not be rendered at the appropriate relative level (to the direct sound component).

[0095] More specifically, the relative gains of the reverberation units are set according to a calibration procedure that assumes the sound sources are omnidirectional (and / or point sources), so that when the reverberation unit is fed a first input signal that is an omnidirectional (and / or point source) and a second input signal that is a directional (and / or non-point source) both having the same signal level, the reverberation unit will produce the same reverberation output level for both, while the output levels should be different to reflect the different source powers of the two sources.

[0096] Therefore, in some embodiments of the present disclosure, to compensate for the relative levels of reverberation of directional sound sources (and / or non-point sound sources), the input gain of the reverberation unit for the directional sound source (and / or non-point sound source) signal may be modified (i.e., the signal level of the input audio signal entering the reverberation unit may be changed) to take into account the fact that the input signal level corresponds to a directional sound source (and / or non-point sound source) having a lower or different source power.

[0097] The relative source power of a directional source (and / or a non-point source) relative to an omnidirectional source (and / or a point source) may be determined from the directivity pattern of the directional source (and / or the non-point source). For example, the relative source power may be determined by integrating the directivity pattern (expressed in units of power) of the directional source (e.g., as specified in its associated directivity metadata) over a unit sphere and normalizing the obtained source power by the source power of the omnidirectional source determined in the same way.

[0098] Thus, if the directivity pattern of a directional source is specified in terms of a linear source amplitude p, the relative source power of the directional source (with respect to an omnidirectional source) may be equal to or proportional to the square of the amplitude p averaged over the unit sphere, i.e., relative source power = p 2 , is.

[0099] The obtained relative source power of the directional sound source (and / or non-point sound source) (compared to the omnidirectional sound source and / or point sound source) may be used to generate a relative gain for generating the adjusted audio signal 260 (shown in FIG. 2 ). There may be various ways to generate the adjusted audio signal 260 using the relative gain. In one embodiment, the relative gain may be used to adjust the level of the input audio signal 256 corresponding to the directional sound source that is provided to the late reflections unit 206, thereby generating the adjusted audio signal 260. In another embodiment, the relative gain may be used to adjust the level of the signal output from the late reflections unit 206, thereby generating the adjusted audio signal 260. In various embodiments, the relative gain may be used to adjust the configuration of the late reflections unit 206 such that the unit generates the adjusted audio signal 260.

[0100] For example, if averaging the directivity patterns of a directional source shows that the relative power of the directional source is half that of an omnidirectional point source, the level of the input signal of the directional source entering the reverberation unit should be reduced by 3 dB.

[0101] The above-described correction method for non-omnidirectional sound sources can also be used for sound sources that have non-point-like distance attenuation behavior, i.e., that do not follow the 1 / r distance law of point sources. Examples of such sources are line sources (with a 1 / sqrt distance attenuation curve), surface sources, or, in general, volumetric sources. U.S. Patent Application No. 17 / 344,632 discloses a model for deriving the distance attenuation behavior of all these types of sound sources as a function of the source's size in various dimensions. These documents are incorporated herein by reference.

[0102] In augmented reality scenarios where values ​​of RDR parameters may have to be derived in real-world environments using real-world (i.e., non-ideal) sound sources that are neither omnidirectional nor point sources, the corrections described above for directional patterns and / or non-point source behavior can also be applied to correct RDR values ​​derived from measurements (a footnote in Section 3 of the Additional Information section below already suggests similar corrections for measurement distances different from the default). The same applies to use cases where, on the content authoring side, values ​​of RDR parameters are derived from room impulse responses that were not obtained using an omnidirectional point source at a specified distance. As long as the directional pattern, measurement distance, and distance attenuation function are known, all of these can be corrected.

[0103] 3. Providing a time variable to the audio renderer

[0104] Equation 1 in the Additional Information section below includes two time variables t1 and t2 (shown in Figure lb). t1 represents the upper integration limit (end time) of the direct sound energy component 112 (denominator), and t2 represents the lower integration limit (start time) of the reverberant sound energy component 114 or a combination of 114 and 116 (numerator). Although Figure lb shows t1 and t2 to have different values, they could have the same value.

[0105] Different renderer implementations may distribute the generation and rendering of the reverberant sound component of diffuse reverberation (or late reverberation) and the early reflection sound component of reverberation differently. For example, in one implementation, a reverberation unit generates only the diffuse sound component of reverberation, while another unit generates the early reflection sound component of reverberation. Meanwhile, in another implementation, the reverberation unit may generate both of these components. In still other implementations, the generation and rendering of the reverberant part of the sound of the rendered sound source (i.e., everything except the direct sound component) may be divided differently, for example, across a different number of processing units or with different "handover" times between units. These different implementations may be accommodated by selecting the parameters t1 and t2 in Equation 1 in the Additional Information section below.

[0106] For example, if the reverberation unit generates all the reverberant sound, the value of t1 may be the time when the direct sound component ends, and t2 may be equal to t1 (so they are connected in time). However, in another case where the direct sound component also contains some very early reflections (since these are perceptually integrated with the actual direct sound component), the value of t1 may be a little larger than in the first case, while t2 is still equal to t1. Furthermore, in another case where the reverberation unit generates only the diffuse part of the reverberation, the value of t2 may be larger than t1, so that the two intervals are not connected in time.

[0107] If the values ​​of the parameters t1 and t2 selected on the authoring side (which determine the value of the RDR measure) differ from either or both of the values ​​of the parameters t1 and t2 selected on the renderer side (where the balance between the direct and reverberant sound components is set to match the received values ​​of the RDR), the balance between the direct and reverberant sound components generated by the renderer may not perfectly match the balance intended by the creator (e.g., a scene creator who created an augmented reality (XR) scene containing a sound source).

[0108] Therefore, it is desirable to select the values ​​of the parameters t1 and t2 on the renderer side to be the same as the values ​​of the parameters t1 and t2 selected on the authoring side. Therefore, in some embodiments of the present disclosure, the values ​​of the parameters t1 and / or t2 selected on the authoring side are transmitted to the renderer along with the RDR parameters. By receiving the parameters t1 and / or t2 selected on the authoring side, the renderer can set its parameters t1 and / or t2 to be the same as the received parameters, thereby producing the balance intended by the creator.

[0109] Alternatively, the renderer may modify the received RDR (or any other energy ratio) parameters to take into account different values ​​of the parameters t1 and / or t2 selected on the authoring side and the values ​​of the parameters t1 and / or t2 selected on the renderer side, and use the modified RDR parameters to generate the balance between the direct sound component and the reverberant sound component intended by the author. For example, the relative gain of the reverberation units may be changed by the same amount as the value of the RDR parameters change.

[0110] Specifically, a modification to account for different values ​​of the parameter t2 between the authoring side and the renderer side can be based on the diffuse field approximation of reverberant sound. On a logarithmic (dB) scale, the so-called energy decay curve of a perfectly diffuse sound field is a straight line with a slope of -60 / RT60 (dB / s) (see Figure 9). This means that in a diffuse sound field, the difference between the RDR values ​​(expressed in dB) determined by the values ​​of the parameter t2 of t2_1 and t2_2, respectively, is given by:

number

number

[0111] As will be explained, the value t2_2 of parameter t2 corresponding to the received RDR parameters may be received by the renderer as additional metadata for the XR scene. Alternatively, it may be obtained in any other way, for example, implicitly from the fact that the received RDR values ​​are known to have been determined according to a particular definition (e.g., because the XR scene is in a particular, known, e.g., standardized, format). As an example of this, the MPEG-I encoder input format [ISO / IEC JTC1 / SC29 / WG6 Document No. N0054: “MPEG-I Immersive Audio Encoder input Format”] specifies that the value of parameter t2 is equal to four times the acoustic time-of-flight associated with the longest dimension of the acoustic environment. Thus, in the latter example, the renderer can determine by itself the value of parameter t2 associated with the received RDR parameters from the fact that the received XR content was encoded according to the MPEG-I standard. If the parameter t2 associated with the received RDR parameter has a known fixed value (e.g., because the standard in which the XR scene is formatted specifies a fixed value for the t2 parameter), the value of the parameter t2 associated with the received RDR parameter may even be "baked in" in the formula used by the renderer to calculate modifications to the received RDR parameter.

[0112] As mentioned above, in some embodiments, the relative gain of the reverberation unit may be changed by the same amount as the change in the value of the RDR parameter. Once the changed relative gain is obtained, the changed relative gain (hereinafter “relative gain”) may be used to generate the adjusted audio signal 260. In one embodiment, the relative gain is used to adjust the level of the input audio signal 256 corresponding to the sound source having the received RDR value, thereby generating the adjusted audio signal 260. In another embodiment, the relative gain may be used to adjust the level of the signal output from the late reflections unit 206, thereby generating the adjusted audio signal 260. In various embodiments, the relative gain may be used to adjust the configuration of the late reflections unit 206 such that the unit generates the adjusted audio signal 260.

[0113] 4. Rendering sound sources with different reference distances for distance attenuation functions

[0114] The MPEG-I audio standard as currently developed allows setting a so-called "reference distance" attribute ("refDistance") for sound sources, which specifies the distance from the sound source at which the sound source's distance attenuation should be 1. This can be seen as a normalization of the sound source's distance attenuation function, which, among other things, allows a degree of level alignment between different renderers rendering a scene. The default value of this attribute is 1m, but content creators are free to choose a different value for sound sources if they need to.

[0115] Setting the reference distance attribute of a sound source to a value different from the default value results in the fact that the distance attenuation function of the sound source will be 1 at a different distance than when the default value is used. This will also generally result in the fact that the rendered direct sound level of the sound source will be different compared to when the default value is used (although this may actually be the intention of the content creator when setting a different value for the sound source).

[0116] In general, setting the reference distance attribute of a sound source to a value RD rather than the default RD_def will modify the rendered level of the direct sound of the sound source compared to the corresponding level of the default value by a factor that is a function of RD and RD_def. As mentioned above, this may in fact be the intention of the content creator.

[0117] Specifically, setting the reference distance attribute of a point sound source to a value RD instead of the default RD_def changes the rendering level of the direct sound of the sound source by a factor of RD / RD_def compared to the corresponding level of the default value.

[0118] For example, if a point sound source has a reference distance value of 2m instead of the default 1m, the rendered direct sound level of the sound source is effectively increased everywhere by a factor of 2(2m / 1m) (i.e., 6dB) compared to the same sound source with the same source signal level but with the default reference distance value. Similarly, if a sound source has a reference distance of 0.5, its rendered direct sound level will be everywhere 6dB lower than if the default value of 1m were used.

[0119] However, since the RDR measures and calibration of the renderer are likely determined for an omnidirectional point sound source with a default value for the reference distance, rendering a sound source with a reference distance value different from the default value may result in an inaccurate balance between the direct and reverberant sound components of that sound source, because while the rendered level of the direct sound is changed for the sound source for the non-default reference distance, the signal level of the sound source (i.e., the level of the signal provided to the sound source) is not changed, and this signal is provided to the reverberation unit to generate the reverberant component for the sound source.

[0120] Thus, in the example where the reference distance of a point sound source is 2m instead of the default 1m, the rendered level of the direct sound component of the sound source is increased by a factor of 2 (6dB), and the rendered level of the reverberant component is the same as it would be if the sound source had the default value for the reference distance. The level of reverberation is therefore too low, i.e. it does not have the correct balance to the direct sound level as specified by the RDR parameters.

[0121] Thus, in some embodiments, a relative gain for generating adjusted audio signal 260 (shown in FIG. 2 ) may be generated based on a function of the non-default reference distance and the default reference distance (e.g., the ratio between the non-default reference distance and the default reference distance). There may be various methods for using the relative gain to generate adjusted audio signal 260. In one embodiment, the relative gain may be used to adjust the level of input audio signal 256 corresponding to a sound source having a reference distance value different from the default value provided to late reflections unit 206, thereby generating adjusted audio signal 260. In another embodiment, the relative gain may be used to adjust the level of a signal output from late reflections unit 206, thereby generating adjusted audio signal 260. In various embodiments, the relative gain may be used to adjust the configuration of late reflections unit 206 such that the unit generates adjusted audio signal 260.

[0122] Specifically, the input signal level of the sound source going into the reverberation unit can be adjusted so that the reverberant sound components generated by the reverberation unit of the sound source have the correct energy balance with the direct sound component again. For example, in the case of a point sound source, the input signal level of the sound source going into the reverberation unit can be adjusted by RD / RD_def (on a linear scale) or 20log(dB) on a dB scale. 10 It can be adjusted by a factor of (RD / RD_def).

[0123] More generally, the relative gain compensates for the change in level of the direct sound component relative to the sound source that results from using a particular value of the reference distance instead of the default value. This change in level of the direct sound component may be determined by evaluating the distance attenuation functions associated with the sound source at the default and particular reference distances and calculating the difference (if the gain is expressed on a logarithmic (dB) scale) or ratio (if the gain is expressed on a linear scale).

[0124] For example, for a point sound source with a distance attenuation function proportional to 1 / r (where r is distance), the value of the distance attenuation function is 1 / RD_def when using the default reference distance value, and 1 / RD when using a specific reference distance value. The ratio of these two values ​​therefore results in the relative gain factor RD / RD_def to be applied to the input signal level for the point sound source entering the reverberation unit, as described above.

[0125] For non-point sources, i.e., sources that are not point sources and / or have an associated distance attenuation function that differs from the 1 / r function of a point source, the same approach for determining the relative gain factors may be used, using the particular distance attenuation function associated with the non-point source.

[0126] For example, if the sound source is an infinitely long line source with an associated distance attenuation function proportional to 1 / sqrt(r), the relative gain factor may be determined as sqrt(RD / RD_def). In another example, if the sound source is an infinitely large surface source with a constant distance attenuation function (i.e., the level of the direct sound component does not change with distance), the relative gain factor will be 1, regardless of the values ​​of RD and RD_def.

[0127] In the most general case, if the distance attenuation function of a sound source is DAF, then the relative gain factor is given by DAF(RD_def) / DAF(RD). U.S. patent application Ser. No. 17 / 344,632 discloses a model for deriving the distance attenuation function of a sound source as a function of the source's size in different dimensions. These documents are incorporated herein by reference.

[0128] 4 shows a process 400 for rendering a sound source according to some embodiments of the present disclosure. Process 400 may begin at step s402.

[0129] Step s402 includes receiving an input audio signal corresponding to a sound source.

[0130] Step s404 includes receiving reverberation parameters indicating a target energy ratio between the direct sound component of the rendered audio of the sound source and the reverberant sound component of the rendered audio of the sound source.

[0131] Step s406 includes deriving a relative gain associated with the first configuration of reverberation units, the relative gain being relative to a reference configuration of reverberation units.

[0132] Step s408 includes generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and the derived relative gain.

[0133] In some embodiments, the relative gain corresponds to the difference between (i) a reference output of a reverberation unit associated with a reference configuration and (ii) a first output of a reverberation unit associated with a first configuration.

[0134] In some embodiments, deriving the relative gain includes determining a reference output level of the reverberation unit for a reference input audio signal when the reverberation unit is configured in a reference configuration, and determining a first output level of the reverberation unit for a reference input audio signal when the reverberation unit is configured in a first configuration. Furthermore, deriving the relative gain includes calculating a difference between the reference output level and the first output level, and deriving the relative gain based on the calculated difference between the reference output level and the first output level. In some embodiments, the difference may be expressed using a logarithmic (dB) scale. However, in other embodiments, the difference may be expressed using a linear scale. In such embodiments, the difference may be equal to the ratio between the reference output level and the first output level.

[0135] In some embodiments, the reverberation parameters indicate a target energy ratio at a particular distance from the sound source.

[0136] In some embodiments, the reverberation parameters indicate a target energy ratio for a particular type of sound source, the particular type of sound source being a point source, an omnidirectional source.

[0137] In some embodiments, the reference configuration is a configuration used to calibrate the output levels of the reverberation units to achieve a target energy ratio at a particular distance from the sound source.

[0138] In some embodiments, the reference configuration is any one or combination of a reference room impulse response, a reference reverberation time setting, reference frequency response data, and reference absorption data.

[0139] 5 shows a process 500 for rendering a sound source according to some embodiments of the present disclosure. Process 500 may begin at step s502.

[0140] Step s502 includes receiving an input audio signal corresponding to a sound source.

[0141] Step s504 includes receiving reverberation parameters indicating a target energy ratio between the direct sound component of the rendered audio of the sound source and the reverberant sound component of the rendered audio of the sound source.

[0142] Step s506 includes obtaining the directivity pattern of the sound source.

[0143] Step s508 includes deriving relative power levels of the sound sources based on the obtained directional patterns, the relative power levels being relative to the power levels of omnidirectional sound sources.

[0144] Step s510 involves generating an adjusted audio signal using the received input audio signal, the reverberation parameters, and the derived relative power levels.

[0145] In some embodiments, the sound source is a non-omnidirectional sound source and / or a non-point sound source.

[0146] In some embodiments, the directivity pattern indicates the amplitude or power of the sound radiated by the sound source in each of a number of directions around the sound source.

[0147] In some embodiments, P i Let σ be the power of the sound emitted by a sound source towards a particular direction among a plurality of directions, and m be the number of the plurality of directions, then the relative power level is given by

number

[0148] 6 shows a process 600 for rendering a sound source according to some embodiments of the present disclosure. Process 600 may begin at step s602.

[0149] Step s602 includes receiving an input audio signal corresponding to a sound source.

[0150] Step s604 includes receiving reverberation parameters indicating a target energy ratio between the direct sound component of the rendered audio of the sound source and the reverberant sound component of the rendered audio of the sound source.

[0151] Step s606 includes obtaining at least one of a first variable indicating an upper limit time of the direct sound component and a second variable indicating a lower limit time of the reverberation sound component.

[0152] Step s608 includes generating an adjusted audio signal using the received input audio signal, the reverberation parameters, the obtained first variable, and / or the obtained second variable.

[0153] In some embodiments, where t1 is a first variable, t2 is a second variable, and p(t) is the amplitude of the room impulse response of the acoustic environment at time t, the reverberation parameter may be expressed as:

number

[0154] In some embodiments, the method further includes calculating modified reverberation parameters using the received reverberation parameters and the obtained first and / or second variables, and an adjusted audio signal is generated using the received input audio signal and the modified reverberation parameters.

[0155] 7 shows a process 700 for rendering a sound source according to some embodiments of the present disclosure. Process 700 may begin at step s702.

[0156] Step s702 includes receiving an input audio signal corresponding to a sound source.

[0157] Step s704 includes receiving reverberation parameters indicating a target energy ratio between the direct sound component of the rendered audio of the sound source and the reverberant sound component of the rendered audio of the sound source.

[0158] Step s706 includes deriving a relative gain corresponding to a first relevant reference distance of the sound source, the relative gain being relative to a default reference distance.

[0159] Step s708 includes generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and the derived relative gain.

[0160] In some embodiments, the method further includes obtaining a first associated reference distance for the sound source, the first associated reference distance indicating a distance from the sound source at which a distance attenuation function associated with the sound source has a value of 1. The relative gain is derived based on a function of the first associated reference distance and a default reference distance.

[0161] In some embodiments, the relative gain is derived based on a ratio of the first relevant reference distance to a default reference distance.

[0162] FIG. 10 shows a process 1000 for rendering a sound source. Process 1000 may start with step s1002. Step s1002 includes receiving an input audio signal corresponding to the sound source. Step s1004 includes receiving reverberation parameters indicating a target energy ratio for reverberant components of audio from the sound source. Step s1006 includes deriving one or more of: (i) a relative gain associated with a first directivity pattern of the sound source; (ii) a relative gain associated with a first reference distance of the sound source; (iii) a relative gain associated with a first configuration of reverberation units; and (iv) a relative gain associated with a first time limit for the reverberant components. Step s1008 includes generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and any one or more of the derived relative gains (i) to (iv). The relative gain associated with the first directional pattern is relative to a reference directional pattern, the relative gain associated with the first reference distance is relative to a default reference distance, the relative gain associated with the first configuration is relative to a reference configuration of the reverberation units, and the relative gain associated with the first time limit is relative to a second time limit for the reverberation sound component.

[0163] In some embodiments, the target energy ratio is a target energy ratio between the direct sound component of the audio of the sound source and the reverberant sound component of the audio of the sound source.

[0164] In some embodiments, the target energy ratio is a target energy ratio between the total energy emitted by the sound source and the energy corresponding to the reverberant component of the sound source's audio.

[0165] In some embodiments, generating the adjusted audio signal includes any one or combination of: modifying the input audio signal based on one or more of the derived relative gains (i)-(iv); modifying one or more configurations of the reverberation unit such that the reverberation unit generates the adjusted audio signal based on one or more of the derived relative gains (i)-(iv); and modifying an output signal from the reverberation unit based on one or more of the derived relative gains (i)-(iv).

[0166] In some embodiments, the first directivity pattern is at least one of a directivity pattern of a non-omnidirectional sound source and a directivity pattern of a non-point sound source, and the reference directivity pattern is at least one of a directivity pattern of an omnidirectional sound source and a directivity pattern of a point sound source.

[0167] In some embodiments, the first directivity pattern indicates the amplitude or power of sound radiated by the sound source in each of a plurality of directions around the sound source.

[0168] In some embodiments, P i is the power of sound radiated by the sound source towards a particular direction within the plurality of directions, and m is the number of directions, then the relative gain associated with a first directivity pattern of the sound source is

number

[0169] In some embodiments, the first reference distance indicates a distance from the sound source at which the distance attenuation function of the sound source has a value of 1, and a relative gain associated with the first reference distance of the sound source is derived based on a function of the first reference distance and a default reference distance.

[0170] In some embodiments, the relative gain associated with a first reference distance of the sound source is derived based on a ratio of the first reference distance to a default reference distance.

[0171] In some embodiments, the sound source is a non-omnidirectional and / or non-point sound source.

[0172] In some embodiments, the first time limit is associated with the received reverberation parameters.

[0173] In some embodiments, the relative gain associated with the first time limit is determined based on at least one of: (i) a reverberation time associated with the reverberant sound component; and (ii) a difference or ratio between the first time limit and the second time limit.

[0174] In some embodiments, the method further comprises calculating updated reverberation parameters based on the received reverberation parameters and a relative gain associated with the first time limit, and the adjusted audio signal is generated based on the updated reverberation parameters.

[0175] In some embodiments, the relative gain associated with the first configuration corresponds to the difference or ratio between a first output of a reverberation unit associated with the first configuration and a reference output of a reverberation unit associated with the reference configuration.

[0176] In some embodiments, the reverberation parameters indicate a target energy ratio at a particular distance from the sound source.

[0177] In some embodiments, the reverberation parameters indicate a target energy ratio for a particular type of sound source, the particular type of sound source being a point source, an omnidirectional source.

[0178] In some embodiments, the reference configuration is a configuration used to calibrate the output levels of the reverberation units to achieve a target energy ratio at a particular distance from the sound source.

[0179] In some embodiments, the reference configuration is a configuration associated with any one or combination of a reference room impulse response, a reference reverberation time setting, reference frequency response data, and reference absorption data.

[0180] 11 shows a process 1100 for rendering a sound source. Process 1100 may begin at step s1102. Step s1102 includes receiving an input audio signal corresponding to the sound source. Step s1104 includes receiving reverberation parameters indicating a target energy ratio for the reverberant components of the audio of the sound source. Step s1106 includes obtaining a directivity pattern for the sound source. Step s1108 includes deriving a relative power level of the sound source based on the obtained directivity pattern, the relative power level being relative to the power level of an omnidirectional sound source. Step s1110 includes generating an adjusted audio signal using the received input audio signal, the reverberation parameters, and the derived relative power level.

[0181] In some embodiments, the target energy ratio is between the direct sound component of the audio of the sound source and the reverberant sound component of the audio of the sound source.

[0182] In some embodiments, the sound source is a non-omnidirectional and / or non-point sound source.

[0183] In some embodiments, the directivity pattern indicates the amplitude or magnitude of the sound radiated by the sound source in each of a number of directions around the sound source.

[0184] In some embodiments, P i Let σ be the magnitude of the sound radiated by a sound source in a particular direction, and m be the number of directions. The relative power level is then

number

[0185] In some embodiments, A i Let be the magnitude of the sound emitted by the source in the specified direction. i =A i 2 is.

[0186] 12 shows a process 1200 for rendering a sound source. Process 1200 may begin at step s1202, which includes receiving an input audio signal corresponding to the sound source. Step s1204 includes receiving reverberation parameters indicating a target energy ratio for reverberant components of the audio of the sound source. Step s1206 includes deriving a relative gain corresponding to a first associated reference distance of the sound source. Step s1208 includes generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and the derived relative gain.

[0187] In some embodiments, the relative gain is relative to a default reference distance.

[0188] In some embodiments, the method further comprises obtaining a first associated reference distance for the sound source, the first associated reference distance indicating a distance from the sound source when a distance attenuation function associated with the sound source has a value of 1, and the relative gain is derived based on a function of the first associated reference distance and a default reference distance.

[0189] In some embodiments, the relative gain is derived based on a ratio of the first relevant reference distance to a default reference distance.

[0190] 13 shows a process 1300 for rendering a sound source. Process 1300 may begin at step s1302, which includes receiving an input audio signal corresponding to the sound source. Step s1304 includes receiving reverberation parameters indicating a target energy ratio for reverberant components in rendered audio of the sound source. Step s1306 includes obtaining a variable indicating a lower time limit for the reverberant components. Step s1308 includes generating an adjusted audio signal using the received input audio signal, the reverberation parameters, and the obtained variable.

[0191] In some embodiments, the method further comprises calculating modified reverberation parameters using the received reverberation parameters and the obtained variables, and an adjusted audio signal is generated using the received input audio signal and the modified reverberation parameters.

[0192] 14 shows a process 1400 for rendering a sound source. Process 1400 may begin at step s1402. Step s1402 includes receiving an input audio signal corresponding to the sound source. Step s1404 includes receiving reverberation parameters indicating a target energy ratio for reverberant components of rendered audio of the sound source. Step s1406 includes deriving a relative gain associated with a first configuration of reverberation units, the relative gain being relative to a reference configuration of reverberation units. Step s1408 includes generating an adjusted audio signal using the received input audio signal, the received reverberation parameters, and the derived relative gain.

[0193] In some embodiments, the relative gain corresponds to the difference between (i) a reference output of a reverberation unit associated with a reference configuration and (ii) a first output of a reverberation unit associated with a first configuration.

[0194] In some embodiments, deriving the relative gain includes determining a reference output level of the reverberation unit for a reference input audio signal when the reverberation unit is configured in a reference configuration, and determining a first output level of the reverberation unit for the reference input audio signal when the reverberation unit is configured in a first configuration, calculating a difference between the reference output level and the first output level, and deriving the relative gain based on the calculated difference between the reference output level and the first output level.

[0195] In some embodiments, the reverberation parameters indicate a target energy ratio at a particular distance from the sound source.

[0196] In some embodiments, the reverberation parameters indicate a target energy ratio for a particular type of sound source, the particular type of sound source being a point source, an omnidirectional source.

[0197] In some embodiments, the reference configuration is a configuration used to calibrate the output levels of the reverberation units to achieve a target energy ratio at a particular distance from the sound source.

[0198] In some embodiments, the reference configuration is a configuration associated with any one or combination of a reference room impulse response, a reference reverberation time setting, reference frequency response data, and reference absorption data.

[0199] 8 is a block diagram of an apparatus 800 according to some embodiments for implementing the audio renderer 200 shown in FIG. 2. As shown in FIG. 8, the apparatus 800 may include a processing circuit (PC) 802, which includes one or more processors (P) 855 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.), which may be co-located within a single enclosure or a single data center, or geographically distributed (i.e., the apparatus 800 may be a distributed computing device). The apparatus 800 includes at least one network interface 848. Each network interface 848 includes a transmitter (Tx) 845 and a receiver (Rx) 847 to enable the device 800 to transmit and receive data to and from other nodes connected to the network 110 (e.g., an Internet Protocol (IP) network) to which the network interface 848 is connected (directly or indirectly). (For example, the network interface 848 is wirelessly connected to the network 110, in which case the network interface 848 is connected in an antenna configuration.) The device 800 includes one or more storage units (also known as "data storage systems") 808, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments in which the PC 802 includes a programmable processor, a computer program product (CPP) 841 may be provided. The CPP 841 includes a computer-readable medium (CRM) 842 that stores a computer program (CP) 843, which includes computer-readable instructions (CRI) 844. CRM 842 may be a non-transitory computer-readable medium, such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., random access memory, flash memory), etc. In some embodiments, CRI 844 of computer program 843 is configured to cause device 800, when executed by PC 802, to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts).In other embodiments, the device 800 may be configured to perform the operations described herein without the need for code. That is, for example, the PC 802 may consist solely of one or more ASICs. Thus, features of the embodiments described herein may be implemented in hardware and / or software.

[0200] Additional Information Section

[0201] 1. The purpose and desired properties of the measurements in the discussion

[0202] The purpose of the measurement under discussion (here called DDR) is to make it possible to set the relative level of late reverberation when rendering a scene so that it has the correct balance ("correct" in the sense intended by the scene creator) with the level of direct sound at all positions in the room.

[0203] With this objective in mind, preferred measurements preferably have the following characteristics: - A measurement should be easily understandable by scene authors and renderer implementers, i.e. it should be intuitively clear what it represents. Preferably, the value of the measurement should also be easily interpretable in terms of the associated acoustic measurement. - Measurements should be derivable in different scene authoring scenarios and it should be clear to the scene author how this can be done. - As with RT60, the measurement should leave implementers a lot of freedom as to how to use the measurement to achieve good rendering results, without requiring the use of specific algorithms to understand the measurement. The measurement should also support the use of different types of reverberation generation techniques. - The measurements should be characteristic of the acoustic environment only, meaning, for example, that the characteristics of any particular sound source should not play a role in determining the measurements. - Absolute levels are irrelevant, it's all about the balance between the two levels / energies.

[0204] The measurements proposed herein are believed to possess all the properties listed above. They are easy to understand and use, and are applicable both to authoring scenarios where RIR (measured or simulated) is used as the basis for determining the acoustic parameters of a scene, as well as to authoring scenarios where a more "sound designer" procedure is used, where the balance between direct sound and late reverberation is adjusted by the scene creator on an artistic basis. Likewise, the proposed measurements support a wide variety of rendering scenarios.

[0205] 2. Description of proposed measurements

[0206] It is proposed to use the common direct-to-reverberant energy ratio (DRR) or its inverse (RDR) as an indicator of the relative level of late reverberation.

[0207] RDR is generally defined as the ratio of late reverberant energy to direct sound energy at a given position. This general definition has intuitive and straightforward meaning in both audio systems that generate different components of the sound field separately, as well as those that use room impulse response (RIR)-based methods. In the latter case, RDR can be expressed as:

number

[0208] As the general definition of RDR given above depends on the distance to the sound source and the radiation characteristics of the source, we propose to define the measurement used more specifically as follows: the ratio of late reverberation energy to direct sound energy (RDR) at a distance of 1 m from an omnidirectional point sound source.

[0209] By specifying the distance to the source and specifying it as an omnidirectional point source, the proposed measure defined above is unambiguous and, in addition, allows for the estimation of the ratio of direct sound to late reverberation at any position in the room, since the level of late reverberation is (by definition) the same throughout the room and the direct sound level of an omnidirectional point source follows the well-known 1 / r distance law.

[0210] It should be noted that for an omnidirectional point source as specified in the proposed measurement, the energy of the direct sound is directly proportional to the total emitted energy of the source, so the proposed measurement essentially coincides with the general definition of DDR currently found in the EIF.

[0211] Also, note that additional predefined information about source-receiver distance was missing in previous proposals for using direct-to-reverberant sound ratio as a measure in EIF.

[0212] 3. Conceptual method for determining the proposed measurements

[0213] On the scene authoring side, the proposed measurements for the acoustic environment can be evaluated by a simple conceptual procedure consisting of the following. - Place an omnidirectional point sound source (test sound source) at a reasonable location within the acoustic environment. - Place an omnidirectional receiver at a reasonable position at a predetermined distance (1 m) from the sound source. - Render a sound source and measure the direct and late reverberant energy at the receiver.

[0214] "Reasonable" above means, for example, that there are no obstructions between the source and receiver positions or immediately adjacent to reflecting surfaces used to determine the measurements. Other than that, there are no special requirements.

[0215] This conceptual method is believed to make it possible to obtain values ​​for the proposed measurements in all relevant scene authoring scenarios. Specifically, it allows: - Derive the proposed measurement (using Equation 1) from the (measured or simulated) RIR of the acoustic environment (note that it is possible to correct for RIR measurements made at different source-receiver distances, as long as the distance is known or can be estimated (e.g. from the TOA of the direct sound)). - From the respective rendered levels of the direct sound rendering module and the late reverberation rendering module, a proposed measurement is obtained, adjusted by ear to the appropriate balance (sound designer approach).

[0216] 4. Conceptual Methods for Using the Proposed Measurements for Renderer Calibration

[0217] On the rendering side, essentially the same conceptual method described above for deriving the proposed measurements is used to calibrate the renderer so that the desired balance between direct sound and late reverberation is obtained anywhere in the acoustic environment. - Place and render an omnidirectional point source in an acoustic environment. - Measure the energy ratio between the direct sound and the late reverberation using an omnidirectional receiver placed at a given distance from the sound source.

[0218] The energy ratio thus obtained can be directly compared with the desired value (provided for the acoustic environment) and the output level of the late reverberation rendering module can be adjusted accordingly.

[0219] On the rendering side, this conceptual approach also works for all relevant scenarios, especially: - RIR-based rendering (e.g., by performing real-time calculation of RIR), - Separate rendering modules for direct sound and late reverberation, whose outputs are mixed together; It is believed that this method can be applied to

[0220] It is important to note that the conceptual renderer calibration procedure described above only needs to be performed once in principle to accurately set the "master" level of the renderer's late reverberation rendering stage. Once this is done for any acoustic test environment and proposed measurements, there is a direct relationship between the received values ​​for any acoustic environment to be rendered and the adjustments that need to be made to the level of the late reverberation stage to achieve the desired balance. In other words, there is no need for scene-by-scene calibration.

[0221] 5. Alternative formulation as critical distance

[0222] An alternative, equivalent way to convey the same information contained in the RDR-based measurements described above is to instead specify a Critical Distance (CD) for the acoustic environment.

[0223] CD is defined as the distance at which the direct sound and the late reverberation are of equal intensity. Therefore, CD essentially conveys the same information as RDR, but in a different form: instead of specifying the RDR at a given distance, it specifies the distance at which the RDR is 0 dB. In fact, for an omnidirectional point source, there is a very simple relationship between the two measurements.

number

number

[0224] To determine CD on the authoring side, one can use essentially the same conceptual method as described for RDR above, with the difference that one must find the distance at which the direct sound and late reverberation are of equal strength. However, as noted above, there is a trivial relationship between CD and RDR for a monopole point source, and so CD can also be obtained using the method described for RDR (and vice versa).

[0225] Similarly, the conceptual method for renderer calibration described above for RDR can also be used for CD, with similar explicit modifications (i.e., measuring the balance at a specified critical distance and adjusting the level of late reverberation so that the direct and late reverberant energies are equal).

Claims

1. A method for rendering a sound source (1000), The process of receiving an input audio signal corresponding to the sound source (s1002), The process of receiving reverberation parameters (s1004) that indicate a target energy ratio between the direct sound component and the reverberation component of the audio of the sound source, or a target energy ratio between the total energy emitted by the sound source and the energy corresponding to the reverberation component, A step (s1006) to derive a relative gain related to the first limit time of the reverberation component, The process includes (s1008) generating an adjusted audio signal using the received input audio signal, the received reverberation parameter, and the derived relative gain, The relative gain related to the first time limit is with respect to the second time limit of the reverberation component. A method characterized by the following:

2. The method according to claim 1, characterized in that the first time limit is associated with the received reverberation parameter.

3. The method according to claim 2, wherein the relative gain related to the first limit time is determined based on at least one of (i) the reverberation time related to the reverberation component and (ii) the difference between the first limit time and the second limit time.

4. The claim further comprises the step of calculating updated reverberation parameters based on the received reverberation parameters and the relative gain associated with the first time limit, The method according to 2, characterized in that the adjusted audio signal is generated based on the updated reverberation parameters.

5. A computer program (843) characterized in that, when executed by a processing circuit (802), it includes an instruction (844) that causes the processing circuit to perform the method described in any one of Claims 1 to 4.

6. A computer-readable storage medium storing the computer program described in Claim 5.

7. Apparatus (800), Processing circuit (802), Memory (841), The apparatus (800) is characterized in that it has a memory which includes instructions that can be executed by the processing circuit, and the apparatus is operable to perform the method according to any one of claims 1 to 4.