Deriving parameter for reverberation processor
XR systems derive reverberation parameters using relationships and correction factors to address incomplete metadata, ensuring accurate reverberation generation in XR environments.
Patent Information
- Application Number
- JP2025122621
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-29
- Filing Date
- 2025-07-22
- Publication Date
- 2025-12-03
AI Technical Summary
XR systems face challenges in generating acoustically plausible reverberation characteristics when control information for reverberation time and level is incomplete or absent, due to limited data availability or inconsistent standards.
Deriving reverberation parameters using relationships between reverberation time and level, based on critical distance and Sabine's formula, to configure the reverberation processor when metadata is incomplete, and applying correction factors for improved accuracy.
Ensures generation of acoustically plausible reverberation signals in XR environments by deriving missing parameters, enhancing audio realism in XR systems.
Smart Images

Figure 2025175997000001_ABST
Abstract
Description
[Technical Field]
[0001] Embodiments related to deriving parameters of a reverberation processor are disclosed. [Background technology]
[0002] Extended reality (XR) (e.g., virtual reality (VR), augmented reality (AR), mixed reality (MR), etc.) systems typically include an audio renderer for rendering audio to a user of the XR system. The audio renderer typically includes a reverberation processor that generates late and / or diffuse reverberations that are rendered to a user of the XR system to provide the auditory sensation of being in the XR scene being rendered. The generated reverberations should provide the user with the auditory sensation of being in an acoustic environment corresponding to the XR scene (e.g., a church, a living room, a gym, an outdoor environment, etc.).
[0003] Reverberation is one of the most important acoustical properties of a room. Sound generated in a room bounces repeatedly off reflective surfaces such as floors, walls, ceilings, windows, and tables, gradually losing energy. When these reflections mix with each other, a phenomenon known as "reverberation" occurs. Reverberation is therefore a collection of many reflections of sound.
[0004] Two of the most fundamental characteristics of reverberation in an acoustical environment are 1) reverberation time and 2) reverberation level, i.e., how strong or loud the reverberation is (e.g., relative to the power or direct sound level of the sound sources in the space). Both of these are properties of the acoustical environment only, i.e., they do not depend on the individual sound sources.
[0005] Reverberation time is a measure of the time required for reflected sound to "fade out" within an enclosed space after the source sound has stopped. It is important to define how a room responds to acoustics. Reverberation time depends on the amount of sound absorption within the space, being lower in spaces with many absorbent surfaces such as curtains, padded chairs, and even people, and higher in spaces containing primarily hard, reflective surfaces.
[0006] Traditionally, reverberation time is defined as the time it takes for the sound pressure level to decrease by 60 dB after the sound source is suddenly turned off. The abbreviated form of this time is "RT60" (or sometimes T60).
[0007] Typically, for reverberation processors used in audio renderers, these two (and other) characteristics of the generated reverberation can be controlled separately and independently. For example, it is typically possible to configure a reverberation processor to generate reverberation with a particular desired reverberation time and a particular desired reverberation level.
[0008] In XR systems, the characteristics of the generated reverberation are typically controlled by control information, e.g., dedicated metadata contained in an XR scene description, e.g., as specified by a scene creator, that describes many aspects of the XR scene, including its acoustic characteristics. An audio renderer receives this control information, e.g., from a bitstream or file, and uses it to configure a reverberation processor to generate reverberation with the desired characteristics. The exact manner in which the reverberation processor obtains the desired reverberation time and reverberation level in the generated reverberation may vary depending on the type of reverberation algorithm the reverberation processor uses to generate the reverberation. Summary of the Invention [Problem to be solved by the invention]
[0009] Currently, several challenges exist. For example, as mentioned above, it is typically possible to control various characteristics of the generated reverberation (e.g., reverberation time and reverberation level) separately and independently of one another, which provides great flexibility in generating reverberation, but also introduces potential problems. In practice, the XR control information received by the audio renderer may not include control data for all characteristics of the generated reverberation that can be controlled. This can have many reasons. For example, the authoring software used to create the XR scene can only generate a limited set of acoustical characteristics for an acoustical environment. Alternatively, the scene may only correspond to a real location (e.g., a particular famous church) for which only a limited set of acoustical data is available. When the XR scene corresponds to a user's real physical space, the acoustical characteristics of that space typically need to be determined on the fly using the limited technical means available in the user's XR device.
[0010] As explained above, the two most important characteristics of the generated reverberation are the reverberation time, typically expressed in terms of RT60, and the reverberation level, typically expressed as the reverberant-to-direct sound (RDR) energy ratio. If either the reverberation time or the reverberation level is not specified in the control information provided to the audio renderer, it is unclear how the reverberation processor should be configured.
[0011] In terms of XR audio standards, even if a standard in principle supports the specification of many reverberation parameters for an acoustic environment, only some of them may be mandatory to provide an XR scene, while others are optional. For example, in the currently developed ISO / IEC MPEG-I Immersive Audio standard, the RT60 value is the only mandatory reverberation-related parameter for the acoustic environment, and the reverberation level parameter (e.g., expressed as an RDR energy ratio) is optional.
[0012] Therefore, what is needed is a solution for configuring the reverberation processor of an XR audio renderer so that when either the reverberation time and / or the reverberation level is not specified for the acoustic environment being rendered, a reverberation signal with acoustically plausible characteristics is generated for the XR scene. [Means for solving the problem]
[0013] Thus, in one aspect, a method is provided that is performed by an audio renderer. In one embodiment, the method includes obtaining (e.g., receiving or retrieving) metadata for the XR scene. The method also includes obtaining or deriving a first reverberation parameter from the metadata, where the first reverberation parameter is a reverberation time parameter or a reverberation level parameter. The method also includes deriving a second reverberation parameter using the first reverberation parameter. If the first reverberation parameter is a reverberation time parameter, the second reverberation parameter is a reverberation level parameter, and if the first reverberation parameter is a reverberation level parameter, the second reverberation parameter is a reverberation time parameter.
[0014] In one embodiment, a method performed by an audio renderer includes obtaining a set of reverberation parameters from metadata of the extended reality scene, the set including at least a first reverberation parameter and a second reverberation parameter. The method also includes determining whether the first reverberation parameter matches the second reverberation parameter. The determining step includes calculating a first value using the second reverberation parameter and comparing a difference between the first value and the first reverberation parameter to a threshold.
[0015] In another aspect, a computer program is provided that includes instructions that, when executed by a processing circuit of an audio renderer, cause the audio renderer to perform any of the above-described methods. In one embodiment, a carrier is provided that includes the computer program, the carrier being one of an electrical signal, an optical signal, a wireless signal, or a computer-readable storage medium. In another aspect, a rendering device is provided that is configured to perform any of the above-described methods. The rendering device may include a memory and a processing circuit coupled to the memory.
[0016] An advantage of the embodiments disclosed herein is that they allow an audio renderer to provide both reverberation time and reverberation level values to a reverberation processor (which may be part of the audio renderer itself or external to it), thereby enabling the reverberation processor to generate an appropriate reverberation signal for an XR scene. [Brief explanation of the drawings]
[0017] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate various embodiments.
[0018] [Figure 1A] FIG. 1 illustrates a system according to some embodiments.
[0019] [Figure 1B] FIG. 1 illustrates a system according to some embodiments.
[0020] [Figure 2] FIG. 1 illustrates a system according to some embodiments.
[0021] [Figure 3A] 1 is a flowchart showing a process according to an embodiment.
[0022] [Figure 3B] 1 is a flowchart showing a process according to an embodiment.
[0023] [Figure 4] 1 is a block diagram of an apparatus according to some embodiments.
[0024] [Figure 5] FIG.
[0025] [Figure 6] FIG. DETAILED DESCRIPTION OF THE INVENTION
[0026] 1A shows an XR system 100 to which embodiments disclosed herein can be applied. The XR system 100 includes speakers 104 and 105 (which may be speakers of headphones worn by a user) and an XR device 110, which includes a display for displaying images to a user and, in some embodiments, may be configured to be worn by a listener. In the illustrated XR system 100, the XR device 110 has a display and is designed to be worn on the user's head, commonly referred to as a head-mounted display (HMD).
[0027] As shown in FIG. 1B, the XR device 110 may include an orientation detection unit 101, a position detection unit 102, and a processing unit 103, and may be coupled (directly or indirectly) to an audio renderer 151 for generating output audio signals (e.g., a left audio signal 181 for a left speaker and a right audio signal 182 for a right speaker, as shown).
[0028] The orientation sensing unit 101 is configured to detect changes in the listener's orientation and provide information about the detected changes to the processing unit 103. In some embodiments, the processing unit 103 determines an absolute orientation (with respect to some coordinate system) taking into account the detected changes in orientation detected by the orientation sensing unit 101. Different systems for determining orientation and position may also exist, such as systems that use lighthouse trackers (lidars). In one embodiment, the orientation sensing unit 101 may determine an absolute orientation (with respect to a coordinate system) given the detected changes in orientation. In this case, the processing unit 103 may simply multiplex the absolute orientation data from the orientation sensing unit 101 with the position data from the position sensing unit 102. In some embodiments, the orientation sensing unit 101 may comprise one or more accelerometers and / or one or more gyroscopes.
[0029] Audio renderer 151 generates an audio output signal based on an input audio signal 161, metadata 162 about the XR scene the listener is experiencing, and information 163 about the listener's position and orientation. The metadata 162 for the XR scene includes metadata for each object and audio element included in the XR scene, and the metadata for an object may include information about the object's dimensions and occlusion factors for the object (e.g., the metadata may specify a set of occlusion factors, where each occlusion factor is applicable to a different frequency or frequency range). The metadata 162 may also include control information such as reverberation time values, reverberation level values, absorption parameters, etc.
[0030] The audio renderer 151 may be a component of the XR device 110 or may be remote to the XR device 110 (e.g., the audio renderer 151 or a component thereof may be implemented in the cloud).
[0031] 2 shows an example implementation of an audio renderer 151 for generating sound for an XR scene. The audio renderer 151 includes a controller 201 and an audio signal generator 202 for generating an output audio signal (e.g., an audio signal of a multi-channel audio element) based on control information 210 from the controller 201 and input audio 161. In this embodiment, the audio signal generator 202 includes a reverberation processor 204 for generating a reverberation signal.
[0032] In some embodiments, the controller 201 may be configured to receive one or more parameters and trigger the audio signal generator 202 to perform a modification to the audio signal 161 (e.g., increase or decrease the volume level) based on the received parameters. The received parameters include information 163 regarding the listener's position and / or orientation (e.g., orientation and distance to audio elements) and metadata 162 regarding the XR scene. For example, the metadata 162 may include metadata regarding the XR space in which the user is virtually located (e.g., dimensions of the space, information about objects in the space, and information about the acoustic characteristics of the space), as well as metadata regarding the audio elements and objects that occlude the audio elements. In some embodiments, the controller 201 itself generates at least a portion of the metadata 162. For example, the controller 201 may receive metadata about the XR scene and derive additional metadata (e.g., control parameters) based on the received metadata. For example, the controller 201 may use the metadata 162 and the position / orientation information 163 to calculate one or more gain factors (g) for audio elements in the XR scene.
[0033] With respect to generating the reverberation signal used by the signal generator 202 to generate the final output signal, the controller 201 provides reverberation parameters, such as reverberation time and reverberation level, to the reverberation processor 204 so that the reverberation processor 204 is operable to generate the reverberation signal. The reverberation time of the generated reverberation is most commonly provided to the reverberation processor 204 as an RT60 value, although other reverberation time scales exist and can be used as well. In some embodiments, the metadata 162 includes some or all of the necessary reverberation parameters (e.g., RT60 and reverberation level values). However, in embodiments where the metadata does not include reverberation time parameters (i.e., RT values such as RT60 values) or reverberation level parameters (i.e., RL values such as RDR energy ratios), the controller 201 is configured to generate these parameters. For example, as described herein, the controller 201 can generate the reverberation time parameters based on the reverberation level parameters, or vice versa.
[0034] The reverberation level may be expressed in various formats and provided to the reverberation processor 204. For example, it may be expressed as the energy ratio between the direct sound component and the reverberant sound component (DRR) at a certain distance from the sound source being rendered in the XR environment, or vice versa (i.e., RDR energy ratio). Alternatively, the reverberation level may be expressed as the energy ratio between the reverberant sound and the total emitted energy of the sound source. In still other cases, the reverberation level may be expressed directly as the level / gain of the reverberation processor.
[0035] In this context, the term "reverberation" may refer only to sound field components that typically correspond to the diffuse part of the acoustic room impulse response of the acoustic environment, but in some embodiments may also include sound field components that correspond to earlier parts of the room impulse response, e.g., including some late non-diffuse reflections, or even all reflected sounds.
[0036] Other metadata describing reverberation-related characteristics of the acoustic environment that may be included in metadata 162 include parameters describing the acoustic properties of the surface materials of the environment (e.g., describing the absorption, reflection, transmission, and / or diffusion properties of the materials), or specific points in time in the room impulse response associated with the acoustic environment, such as the time after source emission at which the room impulse response diffuses (sometimes called "pre-delay").
[0037] All the above mentioned reverberation related properties are typically frequency dependent, and therefore their associated metadata parameters are also typically provided and processed separately for several frequency bands.
[0038] When authoring a virtual reality sound scene, it is in principle possible to specify reverberation time and reverberation level separately and independently for the virtual acoustic environment. However, in real acoustic environments, reverberation time and reverberation level are not independent characteristics. While there is no one-to-one relationship between the two, it is possible to derive a relationship between them that, while not completely accurate in all cases, makes it at least one possible to derive a reasonable estimate for reverberation level when only information about reverberation time is available, and vice versa.
[0039] One derivation of such a relationship begins with the definition of the "critical distance (CD)," the distance (in meters) at which the sound pressure levels in the direct and reverberant fields are equal. Assuming the reverberant field is perfectly diffuse, CD can be quantified as follows:
number
[0040] We use Sabine's well-known statistical approximation formula for RT60.
number
number
[0041] Therefore, for a particular source directivity type (e.g., an omnidirectional source with γ=1), the critical distance CD is purely a property of the acoustic environment.
[0042] The reverberation level of an acoustic environment can be expressed as the ratio of reverberant to direct sound energy at a distance d from an omnidirectional point sound source (i.e., the RDR energy ratio). There is then a simple relationship between the RDR energy ratio (denoted RDR in the equation) and the critical distance (denoted CD in the equation).
number
[0043] This relationship arises because the energy of the direct sound of an omnidirectional point source varies with the square of the distance, and the RDR energy ratio should be equal to 1 at the critical distance.
[0044] Combining equations (3) and (4) gives an approximate relationship between the RDR energy ratio and RT60.
number
number
[0045] Equation (6) shows that an estimate of the RDR energy ratio can be obtained from RT60 and the volume V of the acoustic environment, and that the approximate relationship between the RDR energy ratio and RT60 is very simple and linear.
[0046] Similarly, equation (6) also allows RT60 to be estimated from a known value of the RDR energy ratio.
[0047] Combining equations (1) and (4), an approximation of the RDR energy ratio for sound absorption in an acoustic environment can be obtained as follows:
number
[0048] The equivalent absorption area A of the acoustic environment may be provided directly in the scene metadata or may be derived from other parameters contained within the scene metadata, for example from specifications of materials or material properties (e.g., absorption coefficients) specified for individual parts of the acoustic environment (e.g., individual walls, floors, ceilings, etc.).
[0049] The above derived equations allow the controller 201 to configure the reverberation processor 204 when either the reverberation time or the reverberation level, or both, are not specified for the acoustic environment being rendered, so that a reverberation signal with acoustically plausible characteristics is generated for the scene.
[0050] As mentioned above, the exact manner in which the reverberation processor 204 obtains the desired reverberation time and level in the generated reverberation may vary depending on the type of reverberation algorithm the reverberation processor uses to generate the reverberation. Common examples of such algorithms include feedback delay networks (FDNs) (which simulate reverberation processing using delay lines, filters, and feedback connections) and convolution algorithms (which convolve the dry input signal with a measured, approximated, or simulated room impulse response (RIR)).
[0051] As an example, in the case of an FDN-based reverberation processor, a desired reverberation time can be obtained by controlling the amount of feedback used. In the case of a convolution-based reverberation processor, the desired reverberation time can be obtained either by loading a specific RIR with that reverberation time, or by adapting the effective length of a generic RIR (e.g., by filtering and time-windowing the generic RIR).
[0052] For both FDN-based and convolution-based reverberation processors, the reverberation level can be controlled by applying an appropriate gain either to the input signal going into the reverberation processor, to the output of the reverberation processor, or internal to the reverberation processor (e.g., applying an overall gain to the FDN structure or RIR, respectively).
[0053] Examples of how this gain can be set to obtain a desired reverberation level (e.g., a desired RDR energy ratio) expressed as an RDR energy ratio at 1 meter from an omnidirectional point sound source are described, for example, in U.S. Provisional Patent Application No. 63 / 217,076, filed June 30, 2021, and International Patent Application No. PCT / EP2022 / 068015, filed June 30, 2022, both of which are incorporated herein by reference. The renderer performs a calibration procedure that adjusts the gain of the reverberation processor so that the rendered direct sound and reverberant components for an omnidirectional point sound source have a desired energy ratio at a distance of 1 meter from the source.
[0054] The renderer then generates an output signal for the user by combining (e.g., adding) the generated reverberation signal with other signal components for the sound source, such as a direct sound component and an early reflection component (both generated in other parts of the renderer).
[0055] As noted above, the relationships between RT60, room geometry, and RDR energy ratio used above to derive RDR energy ratio from RT60, or vice versa, are approximations that assume a diffuse reverberant sound field. This assumption is usually not perfectly valid in real acoustic spaces, and the more the real sound field deviates from a perfectly diffuse field, the less accurate the derived relationships become. However, although the diffuse field assumption is usually not perfectly valid, using the derived relationships in generating reverberation for a given virtual acoustic space usually results in a perceptually plausible reverberation for that space.
[0056] Typically, deviations from the diffuse-field assumption are greater in smaller rooms and in rooms with higher absorption, and therefore, in smaller and more absorbent rooms, the relationship derived above predicts the actual relationship between reverberation time and reverberation level less accurately. For rendering the acoustics of a virtual space, this may not be a problem because, as noted above, the results from using the relationship are typically still valid; there is no real-world reference for comparison. However, in augmented reality (AR) use cases, where virtual sound sources are rendered to appear to be in the same physical space as the user, it is desirable to achieve as close a perceptual match as possible between the reverberation of the real physical space and the generated reverberation. In that case (and other cases where an optimal match between actual and generated reverberation is desired), it is possible to improve the accuracy of the derived relationship by adding correction factors that depend on the room geometry (e.g., room volume, one or more room dimensions, the ratio between the largest and smallest dimensions, etc.), RT60, and / or the absorption characteristics of the acoustic environment (if available), and / or frequency. For example, Equation (6) can be expanded to:
number
[0057] Optionally, equation (6) may be extended by expressing the RDR energy ratio as a power of the ratio of RT60 to V, i.e.,
number
[0058] As a further example, the RDR energy ratio may be expressed as:
number
[0059] In a further embodiment, equation (6) can be generalized to express that the RDR energy ratio is a function of the ratio of RT60 to V as follows:
number
[0060] In further embodiments, equation (6) can be further generalized to express the RDR energy ratio as a function of RT60 and V, i.e., RDR=h(RT60,V) where h() is a function (Equation 10a), or as a function of RT60, i.e., RDR=j(RT60) where j() is a function (Equation 10b).
[0061] In addition to correcting the relationships between the various reverberation parameters when the reverberant sound field is not completely diffuse, the correction factors C and C in Equations 8 and 9 (as well as the correction parameters f and f in Equation 9a and the functional relationships in Equations (10), (10a) and (10b)) can correct the derived relationships for other factors as well.
[0062] One example is when a renderer (implicitly) uses a definition (or convention for measurement) of the RDR energy ratio that differs (in one or more respects) from the definition assumed in the derivation of equations (1)-(7) above.
[0063] Specifically, in the derivation of equations (1)-(7) above, which assumes a completely diffuse reverberant field, it is implicitly assumed that the energy of the reverberant field used to calculate the RDR energy ratio is determined over the entire length of the room impulse response, since in a theoretical diffuse field the room impulse response is diffuse from the onset (i.e., immediately after the direct sound is emitted by the sound source).
[0064] On the other hand, certain renderers instead (implicitly) use a slightly different definition of the RDR energy ratio, where the reverberant energy component of the RDR energy ratio only includes the energy contained in the part of the room impulse response starting from a particular instant in time indicated by the value t1.
[0065] One reason for this design choice is that in real-world spaces, after a direct sound is emitted by a sound source, the reverberation field actually begins to diffuse for a certain amount of time. This amount of time may depend on various factors, such as the room's geometry—e.g., its volume, the size of one or more of its dimensions (e.g., the longest), or the ratio of its dimensions—as well as acoustic parameters such as absorption and RT60. A definition of the RDR energy ratio that considers only reverberation energy after a time identified by t1 may be used to reflect that physical reality. Another reason is that the output response of a reverberation processor that is part of (or used by) the renderer itself begins to diffuse some time after feeding a direct sound signal to the reverberation processor. Therefore, for either of these or other reasons, the renderer may use a definition of the RDR energy ratio in which the reverberation energy component of the RDR energy ratio only includes energy after a certain time.
[0066] As a result of this choice, the resulting value of the RDR Energy Ratio will be smaller than both the value predicted from equations (1)-(7) above, as well as the value obtained if the reverberant energy of the complete room response were included in the reverberant energy component of the RDR Energy Ratio (i.e., t1=0).
[0067] Another example is when the renderer starts rendering reverberation only a certain time t1 after the emission of direct sound by the sound source, due to the fact that in real-world spaces, for example, the reverberation field starts to diffuse only a certain amount of time after emission by the sound source, as described above. This has the same effect on the value of the RDR energy ratio as described in the example above.
[0068] Equation (6) can be modified to include in the reverberation energy component of the RDR energy ratio the effect of including only reverberation energy from a certain time, identified by the value t1 onwards. As an example of this, one can look at the energy decay curve of a fully diffuse field and determine the amount of energy that would be "missed" by including only reverberation energy after the time identified by t1. On a logarithmic (dB) scale, the energy decay curve of a fully diffuse field is a straight line (see Figure 5) with a slope of -60 / RT60 (dB / s). This means that if the portion of the diffuse response before time t1 is excluded, this will result in a reduction in the calculated reverberation energy compared to using the full length of the diffuse response below.
number
number
[0069] In use cases where the received RDR values are determined using (or implicitly assuming) a specific start time t2 of the reverberation energy component that is different from the start time t1 used (implicitly) by the renderer itself, essentially the same correction method as above can also be used to correct the RDR energy ratio values (or "RDR values" for short) received by the renderer. In this case, the renderer-defined RDR values are derived by modifying the received RDR values by the correction factor in equation (11), where t1 is replaced by (t1 - t2) as follows (see Figure 6):
number
number
[0070] If the time parameter t2 of the received RDR value is greater than the renderer's own time parameter t1, the result of the modification is that the received RDR value is increased, and if t2 is less than t1 it is decreased.
[0071] The start time t2 corresponding to the received RDR values may be received by the renderer as additional metadata for the XR scene, or it may be obtained in any other way, for example implicitly from the fact that the received RDR values are known to have been determined according to a certain definition (e.g., because the XR scene is in a particular, known, e.g., standardized, format). As an example, the MPEG-I Immersive Audio Encoder Input Format (ISO / IEC JTC1 / SC29 / WG6, Document No. N0083, "MPEG-I Immersive Audio CfP Supplemental Information, Recommendations and Clarifications, Version 1", July 2021) specifies that t2 is equal to four times the acoustic time-of-flight associated with the longest dimension of the acoustic environment.
[0072] Reverberation time (e.g., RT60) and reverberation level (e.g., RDR value) are typically frequency dependent and therefore specified for various frequency bands, which means that it should be understood that all the formulas and processing steps described above may also be evaluated and performed for different frequency bands, respectively.
[0073] Although the above equations were derived for RDR energy ratios expressed on a linear energy scale, the RDR energy ratios are equally well expressed on a logarithmic (dB) scale, and equivalent logarithmic versions of the equations are readily derived.
[0074] Specifically, the logarithmic version of equation (6) is given by
number
number
number
[0075] In addition to providing a solution for configuring a reverberation processor when either or both of the reverberation time and reverberation level are not specified for the acoustic environment of an XR scene, the derived equations also make it possible to check whether the provided values are consistent with each other when at least two of the following information are provided: reverberation time, reverberation level, and absorption information. Of course, as noted above, the derived relationships are only approximate, and therefore strict consistency cannot be expected from their use, but they at least provide a means to perform a "sanity check" on the provided data, i.e., to check whether the combination of values is reasonable. (Note that "reasonable" here refers to what would occur in a real-world acoustic environment; of course, there is no reason why a virtual environment cannot have acoustic characteristics that do not exist in the real world.)
[0076] An audio renderer can use such checks in a number of ways. In one embodiment, the renderer uses derived formulas to check the provided parameters for consistency with each other, and if the consistency is worse than a threshold, it can reject the value of at least one of the parameters and replace it with a value derived from the formula provided above. If two of the three parameters (reverberation time, reverberation level, and absorption information) are consistent and one is inconsistent, a formula can be derived from the inconsistent formula that indicates the inconsistent formula and can replace its value. If only two parameters are provided, or if all three are provided and they are all mutually inconsistent, hierarchical rules can be used to determine which parameter should be replaced. For example, if reverberation time is the highest hierarchy, reverberation level is second, and absorption information is third, and so, for example, reverberation time and reverberation level are provided and found to be inconsistent, the reverberation level value is rejected and replaced, while the reverberation time value is maintained.
[0077] 3A is a flowchart illustrating a process 300 according to some embodiments. Process 300 may begin at step s302, which includes obtaining metadata for the extended reality scene. Step s304 includes obtaining or deriving a first reverberation parameter from the metadata, where the first reverberation parameter is a reverberation time (RT) parameter (e.g., RT60) or a reverberation level (RL) parameter (e.g., an RDR value). Step s306 includes deriving a second reverberation parameter using the first reverberation parameter. If the first reverberation parameter is a reverberation time parameter, the second reverberation parameter is a reverberation level parameter, and if the first reverberation parameter is a reverberation level parameter, the second reverberation parameter is a reverberation time parameter.
[0078] In some embodiments, the metadata includes an acoustic absorption parameter (denoted "A") indicating an amount of acoustic absorption, and the first reverberation parameter is derived using the acoustic absorption parameter. In some embodiments, the first reverberation parameter is an RDR value, and deriving the RDR value includes calculating RDR=Y / A, where Y is a predetermined constant. In one embodiment, Y=16×π.
[0079] In some embodiments, the first reverberation parameter is a reverberation time parameter (RT) (e.g., RT), and deriving the second reverberation parameter includes calculating X×RT or RT / X, where X is a number. In some embodiments, the first reverberation parameter and the second reverberation parameter are associated with an acoustic environment having a volume, and deriving the second reverberation parameter includes calculating f×(RT / V) f2 In some embodiments, deriving the second reverberation parameter includes calculating f(RT / V) using a function f(). In some embodiments, deriving the second reverberation parameter includes calculating h(RT,V) using a function h(). In some embodiments, deriving the second reverberation parameter includes calculating j(RT) using a function j().
[0080] In some embodiments, the first reverberation parameter is a reverberation level parameter (RL) (e.g., an RDR value), and deriving the second reverberation parameter (i.e., a reverberation time parameter) comprises calculating X×RL or RL / X, where X is a number. In some embodiments, the first reverberation parameter and the second reverberation parameter are associated with an acoustic environment having a volume, and deriving the second reverberation parameter comprises calculating i) V×RL / f1, or ii) V×(RL / f1). 1 / f2 In some embodiments, deriving the second reverberation parameter includes calculating V×g(RL) using a function g(), which is the inverse of the function f(), i.e., g()=f -1In some embodiments, deriving the second reverberation parameter includes calculating k(RL,V) using a function k(). The function k() may be the inverse of the function h(). In some embodiments, deriving the second reverberation parameter includes calculating l(RL) using a function l(). The function l() may be the inverse of the function j().
[0081] In some embodiments, the processing also includes generating a reverberant signal using the first reverberant parameter and the second reverberant parameter, and generating an output audio signal using the reverberant signal.
[0082] 3B is a flowchart illustrating a process 350 according to some embodiments. Process 350 may begin at step s352. Step s352 includes obtaining a set of reverberation parameters, including at least a first reverberation parameter and a second reverberation parameter, from metadata of the extended reality scene. Step s354 includes determining whether the first reverberation parameter matches the second reverberation parameter (step s354). Determining includes calculating a first value using the second reverberation parameter (step s356) and comparing a difference between the first value and the first reverberation parameter to a threshold (step s358).
[0083] In some embodiments, the processing also includes generating a reverberant signal using the first value instead of the first reverberant parameter as a result of determining that the difference exceeds a threshold.
[0084] In some embodiments, i) the first reverberation parameter is a reverberation level parameter and the second reverberation parameter is either a reverberation time parameter or an absorption parameter A; ii) the first reverberation parameter is a reverberation time parameter and the second reverberation parameter is either a reverberation level parameter or an absorption parameter A; or iii) the first reverberation parameter is an absorption parameter and the second reverberation parameter is either a reverberation level parameter or a reverberation time parameter.
[0085] In some embodiments, the set of reverberation parameters further includes a third reverberation parameter, and the process further includes, as a result of determining that the first reverberation parameter does not match the second reverberation parameter, determining whether the first reverberation parameter matches the third reverberation parameter, wherein determining whether the first reverberation parameter matches the third reverberation parameter includes: i) calculating a second value using the third reverberation parameter; and ii) comparing a difference between the second value and the first reverberation parameter to a threshold. In some embodiments, the process further includes, as a result of determining that the first reverberation parameter does not match either the second reverberation parameter or the third reverberation parameter, generating a reverberation signal using either the first value or the second value instead of the first reverberation parameter.
[0086] 4 is a block diagram of an audio rendering device 400 according to some embodiments for performing the methods disclosed herein (e.g., audio renderer 151 may be implemented using audio rendering device 400). As shown in FIG. 4, audio rendering device 400 includes a processing circuit (PC) 402. Processing circuit (PC) 402 may include one or more processors (P) 455 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.). Multiple processors (P) 455 may be co-located within a single housing or a single data center, or may be geographically distributed (i.e., device 400 may be a distributed computing device). Audio rendering device 400 includes a network interface 448. The network interface 448 includes a transmitter (Tx) 445 and a receiver 447 that enable the device 400 to transmit and receive data to and from other nodes connected to the network 110 (e.g., an Internet Protocol (IP) network) to which the network interface 448 is connected (directly or indirectly) (e.g., the network interface 448 is connected to an antenna device). The audio-rendering device 400 includes a storage unit (also referred to as a “data storage system”) 408, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments in which the PC 402 includes a programmable processor, a computer program product (CPP) 441 may be provided. The CPP 441 includes a computer-readable storage medium (CRM) 442 that stores a computer program (CP) 443 that includes computer-readable instructions (CRI) 444. The CRM 442 may be a non-transitory computer-readable storage medium, such as a magnetic medium (e.g., a hard disk), an optical medium, or a memory device (e.g., a random access memory, a flash memory).In some embodiments, the CRI 444 of the computer program 443 is configured such that, when executed by the PC 402, the CRI causes the audio-rendering device 400 to perform the steps described herein (e.g., steps described herein with reference to flowcharts). In other embodiments, the audio-rendering device 400 may be configured to perform the steps described herein without the need for code; that is, for example, the PC 402 may consist solely of one or more ASICs. Thus, features of the embodiments described herein may be implemented in hardware and / or software.
[0087] Overview of Various Embodiments
[0088] A1. A method (300) performed by an audio renderer (151), comprising: a step (s302) of obtaining metadata of an extended reality scene; a step (s304) of obtaining or deriving a first reverberation parameter from the metadata, wherein the first reverberation parameter is a reverberation time parameter or a reverberation level parameter; and a step (s306) of deriving a second reverberation parameter using the first reverberation parameter, wherein if the first reverberation parameter is the reverberation time parameter, the second reverberation parameter is a reverberation level parameter, and if the first reverberation parameter is the reverberation level parameter, the second reverberation parameter is a reverberation time parameter.
[0089] A2. The method of embodiment A1, wherein the metadata includes an acoustic absorption parameter indicating an amount of acoustic absorption (A), and the first reverberation parameter is derived using the acoustic absorption parameter.
[0090] A3. The method of embodiment A2, wherein the first reverberation parameter is a reverberation-to-direct sound energy ratio (RDR) value, and deriving the RDR value includes calculating 16×(π / A).
[0091] A4. The method of embodiment A1 or A2, wherein the first reverberation parameter is the reverberation time parameter (RT) (e.g., an RT60 value), and deriving the second reverberation parameter includes calculating X×RT or RT / X, where X is a number.
[0092] A5. The first reverberation parameter is the reverberation time parameter (RT) (e.g., RT60 value), and the first reverberation parameter and the second reverberation parameter are associated with an acoustic environment having a volume, and deriving the second reverberation parameter includes calculating f1x(RT / V) or (f1x(RT / V)), where f1 is a predetermined coefficient, f2 is a predetermined value (in some embodiments, f2=1), and V is a volume value indicating the volume of the acoustic environment. f2 ) In one embodiment, f1 is a function of the distance d from the omnidirectional point sound source. For example, by substituting c for a predetermined coefficient (e.g., c=3.1×10 2 ), then f1 is c×d 2 In another embodiment, f1 is equal to 3.1×10 2 In another embodiment, c is set to a predetermined factor (e.g., c=3.1×10 2 ), where C is a given coefficient, then f1 = C × c.
[0093] A6. The method according to any one of embodiments A1 to A3, wherein the first reverberation parameter is the reverberation level parameter (RL), and deriving the second reverberation parameter comprises calculating X×RL or RL / X, where X is a number.
[0094] A7. The first reverberation parameter and the second reverberation parameter are associated with an acoustic environment having a volume, and deriving the second reverberation parameter includes: deriving VxRL / f1 or (Vx(RL / f1)) where f1 is a predetermined coefficient, V is a volume value indicating the volume of the acoustic environment, and f2 is a predetermined value. 1 / f2 The apparatus of embodiment A6, comprising calculating
[0095] A8. The first reverberation parameter and the second reverberation parameter are associated with an acoustic environment having a volume, the first reverberation parameter being the reverberation time parameter (RT), and deriving the second reverberation parameter comprises:
number
[0096] A9. A method according to any one of embodiments A1, A2, A4 and A5, characterized in that the second reverberation parameter is the reverberation level parameter, and the second reverberation parameter is derived using the first reverberation parameter and a predetermined time value t1.
[0097] A10. The method of embodiment A5, wherein f1 is equal to C×c, where C is a correction coefficient that depends on the first reverberation parameter and the time value t1, and c is a predetermined value.
[0098] A11. C is
number
[0099] The method of any one of embodiments A8 to A11, wherein A12.t1 is derived based on at least one dimension of the acoustic environment.
[0100] A13. The method of any one of embodiments A8 to A11, wherein t1 is proportional to an acoustic time of flight related to the dimensions of the acoustic environment.
[0101] A14. The method of embodiment A13, wherein t1=4xL / s, where L is the size of the longest dimension of the acoustic environment and s is the speed of sound.
[0102] The method of any one of embodiments A8-A11, wherein A15.t1 indicates a pre-delay time associated with the acoustic environment.
[0103] A16. The method of any one of embodiments A8-A11, wherein t1 is a time value indicating a portion of a room impulse response associated with the acoustic environment.
[0104] A17. The method of any one of embodiments A1 to A16, wherein the reverberation level parameter is expressed as an energy ratio between the reverberant sound and the total radiant energy of the sound source.
[0105] A18. The method of any one of embodiments A1 to A17, further comprising: generating a reverberation signal using the first reverberation parameters and the second reverberation parameters; and generating an output audio signal using the reverberation signal.
[0106] B1. A method (350) performed by an audio renderer (151), comprising the steps of: obtaining (s352) a set of reverberation parameters from metadata of an extended reality scene, the set including at least a first reverberation parameter and a second reverberation parameter; and determining (s354) whether the first reverberation parameter matches the second reverberation parameter, wherein the determining step includes the steps of: calculating (s356) a first value using the second reverberation parameter; and comparing (s358) a difference between the first value and the first reverberation parameter to a threshold.
[0107] B2. The method of embodiment B1, further comprising, upon determining that the difference exceeds the threshold, generating a reverberation signal using the first value instead of the first reverberation parameter.
[0108] B3. The method of embodiment B1 or B2, wherein the first reverberation parameter is a reverberation level parameter and the second reverberation parameter is either a reverberation time parameter or an absorption parameter (A), or the first reverberation parameter is a reverberation time parameter and the second reverberation parameter is either the reverberation level parameter or the absorption parameter (A), or the first reverberation parameter is the absorption parameter and the second reverberation parameter is either the reverberation level parameter or the reverberation time parameter.
[0109] B4. The method of embodiment B1, wherein the set of reverberation parameters further includes a third reverberation parameter, and the method further comprises, upon determining that the first reverberation parameter does not match the second reverberation parameter, determining whether the first reverberation parameter matches the third reverberation parameter, wherein determining whether the first reverberation parameter matches the third reverberation parameter includes calculating a second value using the third reverberation parameter.
[0110] Comparing the difference between the second value and the first reverberation parameter to the threshold value.
[0111] B5. The method of embodiment B4, further comprising, as a result of determining that the first reverberation parameter does not match either the second reverberation parameter or the third reverberation parameter, generating a reverberation signal using either the first value or the second value instead of the first reverberation parameter.
[0112] C1. A computer program comprising instructions that, when executed by processing circuitry of an audio renderer, cause the audio renderer to perform the method of any one of the preceding embodiments.
[0113] C2. A carrier containing the computer program of embodiment C1, the carrier being one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0114] D1. An audio rendering device configured to perform the method of any one of the preceding embodiments.
[0115] D2. The audio rendering device of embodiment D1, wherein the audio rendering device comprises a memory and a processing circuit coupled to the memory.
[0116] E1. A method performed by an audio renderer, comprising the steps of: obtaining (s302) metadata for an extended reality scene; obtaining or deriving a first reverberation parameter from the metadata; and deriving a second reverberation level parameter using the first reverberation parameter.
[0117] E2. Further comprising the step of obtaining a reverberation time parameter (RT); RDR received the first reverberation level parameter, t1 is the start time used by the audio renderer, Let t2 be the start time associated with the first reverberation level parameter (e.g., the start time included in the metadata), The second reverberation level parameter is
number
[0118] While various embodiments have been described herein, it should be understood that they are presented by way of example only, and not limitation. Thus, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments. Moreover, any combination of the above-described objects in all possible variations thereof is encompassed by the present disclosure unless otherwise indicated herein or clearly contradicted by context.
[0119] Additionally, while the processes described above and illustrated in the figures are shown as a series of steps, this is done for illustrative purposes only, and as such, steps may be added, steps may be omitted, the order of steps may be rearranged, and steps may be performed in parallel.
Claims
1. A method (300) performed by an audio renderer (151), comprising: obtaining metadata for the extended reality scene (s302); obtaining or deriving a first reverberation parameter from the metadata (s304), wherein the first reverberation parameter is a reverberation time parameter or a reverberation level parameter; deriving second reverberation parameters using the first reverberation parameters (s306); and when the first reverberation parameter is the reverberation time parameter, the second reverberation parameter is a reverberation level parameter; When the first reverberation parameter is the reverberation level parameter, the second reverberation parameter is a reverberation time parameter. A method characterized by:
2. The metadata includes an acoustic absorption parameter indicating an amount of acoustic absorption (A), the first reverberation parameter is derived using the acoustic absorption parameter.
2. The method of claim 1 .
3. the first reverberation parameter is a reverberant-to-direct sound (RDR) energy ratio value; deriving the RDR energy ratio value includes calculating 16×(π / A); 3. The method of claim 2.
4. the first reverberation parameter is the reverberation time parameter (RT) (e.g., an RT60 value); deriving the second reverberation parameter includes calculating X×RT or RT / X, where X is a number; 3. The method according to claim 1 or 2.
5. the first reverberation parameter is the reverberation time parameter (RT) (e.g., an RT60 value); the first reverberation parameter and the second reverberation parameter are associated with an acoustic environment having a volume; Deriving the second reverberation parameters includes: When f1 is a predetermined coefficient, f2 is a predetermined value, and V is a volume value indicating the volume of the acoustic environment, f1 x (RT / V) or (f1 x (RT / V) f2 ), 5. The method according to claim 1, 2 or 4.
6. the first reverberation parameter is the reverberation level parameter (RL); deriving the second reverberation parameter includes calculating X×RL or RL / X, where X is a number; 4. The method according to claim 1, wherein the first and second electrodes are connected to a first electrode.
7. the first reverberation parameter is the reverberation level parameter (RL); the first reverberation parameter and the second reverberation parameter are associated with an acoustic environment having a volume; Deriving the second reverberation parameters includes: When f1 is a predetermined coefficient, V is a volume value indicating the volume of the acoustic environment, and f2 is a predetermined value, VxRL / f1 or (Vx(RL / f1) 1/f2 ), 7. The method according to claim 1, 2 or 6.
8. the first reverberation parameter and the second reverberation parameter are associated with an acoustic environment having a volume; the first reverberation parameter is the reverberation time parameter (RT); Deriving the second reverberation parameters includes: Let V be the volume of the acoustic environment and t be a time value. [Equation 22] , including calculating 3. The method according to claim 1 or 2.
9. the second reverberation parameter is the reverberation level parameter, the second reverberation parameter is derived using the first reverberation parameter and a predetermined time value t1.
6. The method according to claim 1, 2, 4 or 5.
10. Let C be a correction coefficient that depends on the first reverberation parameter and the time value t1, and c be a predetermined value. f1 is equal to C × c, 6. The method of claim 5.
11. C is [Equation 23] 11. The method of claim 10, wherein the .times. ...
12. 12. The method of claim 8, wherein t1 is derived based on at least one dimension of the acoustic environment.
13. 12. The method of any one of claims 8 to 11, wherein t1 is proportional to an acoustic time of flight related to the dimensions of the acoustic environment.
14. 14. The method of claim 13, wherein t1 = 4 x L / s, where L is the size of the longest dimension of the acoustic environment and s is the speed of sound.
15. 12. The method of any one of claims 8 to 11, wherein t1 denotes a pre-delay time associated with the acoustic environment.
16. 12. The method of any one of claims 8 to 11, wherein t1 is a time value indicative of a portion of a room impulse response associated with the acoustic environment.
17. 17. The method according to claim 1, wherein the reverberation level parameter is expressed as an energy ratio between the reverberant sound and the total radiant energy of the sound source.
18. generating a reverberation signal using the first reverberation parameters and the second reverberation parameters; generating an output audio signal using the reverberant signal; 18. The method of any one of claims 1 to 17, further comprising:
19. A method (350) performed by an audio renderer (151), comprising: obtaining (s352) a set of reverberation parameters from metadata of the extended reality scene, the set including at least a first reverberation parameter and a second reverberation parameter; determining (s354) whether the first reverberation parameters match the second reverberation parameters; and the determining step comprises: Calculating (s356) a first value using the second reverberation parameters; comparing (s358) the difference between the first value and the first reverberation parameter with a threshold; A method comprising:
20. 20. The method of claim 19, further comprising, upon determining that the difference exceeds the threshold, generating a reverberation signal using the first value instead of the first reverberation parameter.
21. the first reverberation parameter is a reverberation level parameter and the second reverberation parameter is either a reverberation time parameter or an absorption parameter (A); or the first reverberation parameter is a reverberation time parameter and the second reverberation parameter is either the reverberation level parameter or the absorption parameter (A); or the first reverberation parameter is the absorption parameter, and the second reverberation parameter is either the reverberation level parameter or the reverberation time parameter; 21. The method of claim 19 or 20.
22. The set of reverberation parameters further includes a third reverberation parameter, and the method further comprises: The method further includes a step of determining whether the first reverberation parameter matches the third reverberation parameter when it is determined that the first reverberation parameter does not match the second reverberation parameter, and the step of determining whether the first reverberation parameter matches the third reverberation parameter includes: calculating a second value using the third reverberation parameter; comparing the difference between the second value and the first reverberation parameter with the threshold; 20. The method of claim 19, comprising:
23. 23. The method of claim 22, further comprising, upon determining that the first reverberation parameter does not match either the second reverberation parameter or the third reverberation parameter, generating a reverberation signal using either the first value or the second value instead of the first reverberation parameter.
24. 24. A computer program comprising instructions which, when executed by a processing circuit of an audio renderer, cause the audio renderer to perform a method according to any one of claims 1 to 23.
25. 25. A carrier containing the computer program of claim 24, wherein the carrier is one of an electrical signal, an optical signal, a radio signal, and a computer readable storage medium.
26. An audio rendering device (400), comprising: obtaining metadata for the extended reality scene (s302); obtaining or deriving a first reverberation parameter from the metadata (s304), wherein the first reverberation parameter is a reverberation time parameter or a reverberation level parameter; deriving second reverberation parameters using the first reverberation parameters (s306); configured to perform a process including when the first reverberation parameter is the reverberation time parameter, the second reverberation parameter is a reverberation level parameter; When the first reverberation parameter is the reverberation level parameter, the second reverberation parameter is a reverberation time parameter.
1. An audio rendering device comprising:
27. 27. An audio rendering device according to claim 26, further configured to perform a method according to any one of claims 2 to 18.
28. An audio rendering device (400), comprising: obtaining (s352) a set of reverberation parameters from metadata of the extended reality scene, the set including at least a first reverberation parameter and a second reverberation parameter; determining (s354) whether the first reverberation parameters match the second reverberation parameters; The determining step is configured to perform a process including: Calculating (s356) a first value using the second reverberation parameters; comparing (s358) the difference between the first value and the first reverberation parameter with a threshold; 1. An audio rendering device comprising:
29. 27. An audio rendering device according to claim 26, further configured to perform a method according to any one of claims 20 to 23.
30. 30. An audio rendering device according to any one of claims 26 to 29, comprising a memory and a processing circuit coupled to the memory.