Rendering reverberation in connected spaces
By determining the amount of source power that the sound source propagates to the connected space through the portal and using an efficient rendering architecture, the problem of unnatural reverb response and excessive computational burden in XR scenes is solved, real and efficient reverb rendering is achieved.
Patent Information
- Application Number
- CN202380085909.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-12-15
- Publication Date
- 2025-07-22
AI Technical Summary
The existing XR audio rendering architecture cannot effectively handle the reverb generated by sound sources outside the connected space, resulting in unnatural reverb response in XR scenes. The existing methods are too computationally cumbersome, making it difficult to render reverbs in multiple connected spaces in real time.
By determining the direct propagation value, indicating the amount of source power that the sound source propagates to the connection space through the portal, and using this value and audio signal to generate an input signal to render the reverb, the reverb of multiple connection spaces is processed using an efficient rendering architecture.
Realistic and efficient rendering of reverbs in XR scenes is achieved, and the problems of unnatural reverb response and excessive computational burden in the prior art are solved, and are suitable for XR scenes in multiple connected spaces.
Smart Images

Figure CN120359767A_ABST
Abstract
Description
Technical Field
[0001] Embodiments related to rendering reverberation in a first space connected to a second space via a portal are disclosed. Background Art
[0002] Extended reality (XR) systems (such as virtual reality (VR) systems, augmented reality (AR) systems, mixed reality (MR) systems, etc.) typically include an audio renderer for rendering audio to a user of the XR system. The audio renderer typically includes a reverb processor to generate late and / or diffuse reverberation, which is rendered to the user of the XR system to provide an auditory sense of being in the XR scene being rendered. The generated reverberation should provide the user with an auditory sense of being in an acoustic environment (AE) (also referred to as a "space") corresponding to the XR scene, such as a living room, a gym, an outdoor environment, etc.
[0003] Reverberation is one of the most important acoustic properties of a room. Sound generated in a room will bounce repeatedly off reflective surfaces such as the floor, walls, ceiling, windows, or tables, while gradually losing energy. When these reflections mix with each other, a phenomenon called "reverberation" is produced. Thus, reverberation is a collection of multiple reflections of sound.
[0004] The two most fundamental properties of reverberation in any space (real or virtual) are: 1) reverberation time and 2) reverberation level, i.e., the intensity or volume of the reverberation, e.g., relative to the power or direct sound level of a sound source in the space.
[0005] Reverberation time is a measure of the time it takes for the reflected sound to "die out" in an enclosed space after the source of the sound ("sound source") has stopped. This is important in defining how a room will respond to acoustic sound. Reverberation time depends on the amount of sound absorption in the space, which is lower in a space with many absorptive surfaces such as curtains, upholstered chairs, or even people, and higher in a space mainly containing hard reflective surfaces.
[0006] Traditionally, reverberation time has been defined as the amount of time it takes for the sound pressure level to decrease by 60 dB after the sound source is suddenly turned off. The shorthand for this amount of time is "RT60" (or sometimes T60).
[0007] Typically, for a reverb processor used in an audio renderer, these two (and other) properties of the generated reverberation can be controlled separately and independently. For example, the reverb processor can typically be configured to generate reverberation with a certain desired reverberation time and a certain desired reverberation level.
[0008] In an XR system, the characteristics of the generated reverberation are typically controlled by control information, such as special metadata included in the XR scene description, e.g., metadata specified by the scene creator, which describes many aspects of the XR scene, including its acoustic characteristics. An audio renderer receives such control information, e.g., from a bitstream or a file, and uses the control information to configure a reverberation processor to produce a reverberation with desired characteristics. The exact way in which the reverberation processor obtains the desired reverberation time and reverberation level in the generated reverberation may vary, depending on the type of reverberation algorithm used by the reverberation processor to generate the reverberation.
[0009] Reverberation and connected spaces
[0010] As noted above, a key aspect of the immersive rendering of audio is the realistic rendering of reverberation associated with the virtual space of the XR scene. A particular challenge is the realistic rendering of reverberation in a space (also referred to as an "acoustic environment") that is connected (or "coupled") to another space via a "portal" (e.g., an open door, an open window, a partially transmissive window, a partially transmissive wall, etc.—i.e., anything through which sound from one space can propagate into the connected space and vice versa).
[0011] One aspect of this is the realistic modeling and rendering of reverberation that is generated in a first space in response to the direct sound of a sound source located in a second connected space, where the direct sound of the sound source propagates directly into the first space via the portal between the two spaces, i.e., without first being reflected and / or reverberated in the second space. Another aspect is the realistic modeling and rendering of reverberation from the connected second space that propagates into the first space via the portal.
[0012] A solution for rendering reverberation is described in the working draft (WD1) of ISO / IEC 23090-4 (MPEG-I Immersive Audio) (i.e., the output document of the 9th meeting of MPEG WG 6 Audio Coding, hereinafter referred to as "ISO / IEC 23090-4 WD1"). Summary of the Invention
[0013] There are certain challenges currently. For example, there is currently no practical rendering model and signal processing architecture for the realistic rendering of reverberation generated in the spaces of an XR scene due to sound sources located in connected spaces. The solution described in the WD1 of ISO / IEC 23090-4 lacks the rendering of reverberation in a first space (e.g., the active space) in response to a source located outside the first space (i.e., in the second space).
[0014] If reverberation is generated in the active space without considering sound sources outside the active space (i.e., the space in which the listener is located), no reverberation is generated in response to these external sound sources, which would be unnatural for the listener(s) located in the active space. Similarly, if a sound source is present in the active space that is connected to another highly reverberant space, the sound of that source would naturally be expected to generate some reverberation in the connected space, and a portion of that reverberation would be expected to propagate back into the active space.
[0015] The current solution in WD1 of ISO / IEC 23090-4 uses a feedback delay network (FDN) reverberator to render reverberation for each space. There are methods for combining several such FDN reverberators into a grouped feedback delay network (e.g., see reference [2]), however, such solutions can easily become computationally too heavy for real-time audio rendering.
[0016] Furthermore, existing XR audio rendering architectures are not optimized for handling reverberation due to sources outside the active space, especially for XR scenarios consisting of many connected spaces, all of which may contain any number of sound sources. Therefore, an optimized rendering architecture is needed to effectively handle such situations.
[0017] Thus, in one aspect, a method is provided for generating an input signal for a first space of an XR scenario based on a first sound source in a second space of the XR scenario, wherein the first space is directly or indirectly connected to the second space via one or more portals including at least a first portal. The method includes: obtaining a direct propagation value, wherein the direct propagation value indicates an amount or portion of source power associated with the first sound source in the second space that is directly propagated from the first sound source through the one or more portals into the first space. The method further includes: using the direct propagation value and an audio signal associated with the first sound source to generate an input signal for the first space.
[0018] In another aspect, a method is provided for enabling the rendering of reverberation in a first space of an XR scenario, the first space being directly or indirectly connected to a second space of the XR scenario via one or more portals including at least a first portal, wherein a sound source is present in the second space. The method includes: determining a direct propagation value, wherein the direct propagation value indicates an amount or portion of source power associated with the sound source that is directly propagated from the sound source in the second space through the one or more portals into the first space. The method further includes: storing and / or transmitting the direct propagation value or data from which the direct propagation value can be derived. The method can be performed by an audio encoder.
[0019] In another aspect, a method for rendering first-order reverberation in a first space of an XR scene is provided. The method includes: determining whether a first sound source is located in a second space. The method further includes: determining whether there is one or more portals connecting the first space and the second space. The method further includes: as a result of determining that the first sound source is located in the second space and there is at least a first portal connecting the first space and the second space, determining a first direct propagation value of the first sound source relative to the first portal. The method further includes: using the direct propagation value and an audio signal associated with the first sound source to generate an input signal for the first space. The method further includes: using the input signal for the first space to generate a reverberation signal for the first space.
[0020] In another aspect, an apparatus for generating an input signal for a first space of an XR scene based on a first sound source in a second space of the XR scene is provided, wherein the first space is directly or indirectly connected to the second space via one or more portals including at least a first portal. The apparatus is configured to obtain a direct propagation value, wherein the direct propagation value indicates an amount or portion of source power associated with the first sound source in the second space, and the amount or portion of source power is directly propagated from the first sound source through the one or more portals into the first space. The apparatus is further configured to use the direct propagation value and an audio signal associated with the first sound source to generate an input signal for the first space.
[0021] In another aspect, an apparatus for enabling reverberation rendering in a first space of an XR scene is provided, the first space being directly or indirectly connected to a second space of the XR scene via one or more portals including at least a first portal, wherein a sound source exists in the second space. The apparatus is configured to determine a direct propagation value, wherein the direct propagation value indicates an amount or portion of source power associated with the sound source, and the amount or portion of source power is directly propagated from the sound source in the second space through the one or more portals into the first space. The apparatus is further configured to store and / or transmit the direct propagation value or data from which the direct propagation value can be derived.
[0022] In another aspect, an apparatus for rendering first-order reverberation in a first space of an XR scene is provided. The apparatus is configured to determine whether a first sound source is located in a second space. The apparatus is further configured to determine whether there is one or more portals connecting the first space and the second space. The apparatus is further configured to, as a result of determining that the first sound source is located in the second space and there is at least a first portal connecting the first space and the second space, determine a first direct propagation value of the first sound source relative to the first portal. The apparatus is further configured to use the direct propagation value and an audio signal associated with the first sound source to generate an input signal for the first space. The apparatus is further configured to use the input signal for the first space to generate a reverberation signal for the first space.
[0023] In another aspect, there is provided a computer program comprising instructions which, when executed by a processing circuit of a device, cause the device to perform the methods disclosed herein. In one embodiment, there is provided a carrier containing a computer program, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0024] The advantages of the embodiments disclosed herein are that they enable realistic and efficient rendering of reverberation in XR scenarios with connected spaces. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings incorporated herein and forming a part of the specification illustrate various embodiments.
[0026] Figure 1A A system according to some embodiments is shown.
[0027] Figure 1B A system according to some embodiments is shown.
[0028] Figure 2 A system according to some embodiments is shown.
[0029] Figure 3 A portal from a first space to a second space is shown.
[0030] Figure 4 A vector from a point to the vertex of a portal opening is shown.
[0031] Figure 5 A grid is shown.
[0032] Figure 6 The concept of an interpolation direct propagation value grid for possible object source positions is shown.
[0033] Figure 7 An effective architecture for an audio renderer is shown.
[0034] Figure 8 A flowchart showing a process according to an embodiment is shown.
[0035] Figure 9 A flowchart showing a process according to an embodiment is shown.
[0036] Figure 10 A block diagram of a device according to some embodiments is shown.
[0037] Figure 11A Two spaces connected via a portal are shown.
[0038] Figure 11B Two spaces connected via a portal are shown.
[0039] Figure 12 is a flowchart showing a process according to an embodiment.
[0040] Figure 13 is a flowchart showing a process according to an embodiment. Detailed implementation
[0041] Figure 1A shows an XR system 100 in which the embodiments disclosed herein can be applied. The XR system 100 includes speakers 104 and 105 (which can be speakers of headphones worn by a user) and an XR device 110, which may include a display for displaying images to the user, and in some embodiments, the XR device is configured to be worn by a listener. In the illustrated XR system 100, the XR device 110 has a display and is designed to be worn on the user's head and is commonly referred to as a head-mounted display (HMD).
[0042] As Figure 1B shown, the XR device 110 may include an orientation sensing unit 101, a position sensing unit 102, and a processing unit 103, which are directly or indirectly coupled to an audio renderer 151 for generating output audio signals using an input audio signal 161 (e.g., encoded audio data) and metadata 162 (which is shown in this example as being provided by an encoder 169), e.g., as shown, a left audio signal 181 for the left speaker and a right audio signal 182 for the right speaker. For example, the encoder 169 may generate a bitstream containing the encoded audio data 161 and the metadata 162.
[0043] The orientation sensing unit 101 is configured to detect changes in the orientation of the listener and provide information about the detected changes to the processing unit 103. In some embodiments, given a change in orientation detected by the orientation sensing unit 101, the processing unit 103 determines the absolute orientation (relative to a certain coordinate system). There may also be different systems for determining orientation and position, such as a system using a beacon tracker (LIDAR). In one embodiment, given the detected change in orientation, the orientation sensing unit 101 may determine the absolute orientation (relative to a certain coordinate system). In this case, the processing unit 103 may simply multiplex the absolute orientation data from the orientation sensing unit 101 and the position data from the position sensing unit 102. In some embodiments, the orientation sensing unit 101 may include one or more accelerometers and / or one or more gyroscopes.
[0044] The audio renderer 151 generates an audio output signal based on an input audio signal 161, metadata 162 about the XR scene that the listener is experiencing, and information 163 about the listener's position and orientation. The metadata 162 for the XR scene may include metadata for each object and audio element included in the XR scene, as well as metadata for the XR space ("acoustic environment") in which the listener is virtually located. The metadata for an object may include information about the object's dimensions and an occlusion factor for the object. For example, the metadata may specify a set of occlusion factors, where each occlusion factor applies to a different frequency or frequency range. The metadata 162 may also include control parameters, such as a reverberation time value, a reverberation level value, and / or one or more absorption parameters.
[0045] The audio renderer 151 may be a component of the XR device 110, or it may be remote from the XR device 110. For example, the audio renderer 151 or its components may be implemented in the cloud.
[0046] Figure 2 An example implementation of an audio renderer 151 for generating sound for an XR scene is shown. The audio renderer 151 includes a controller 201 and an audio signal generator 202 that is configured to generate one or more output audio signals (e.g., audio signals for a multi-channel audio element) based on control information 210 from the controller 201 and the input audio 161. In this embodiment, the controller 201 includes a reverb processor 204 for determining a scaling factor (also referred to as a direct propagation value) as described below.
[0047] In some embodiments, the controller 201 may be configured to receive one or more parameters and trigger the audio signal generator 202 to perform modifications to the audio signal 161 based on the received parameters, such as, for example, increasing or decreasing the volume level. The received parameters include information 163 about the listener's position and / or orientation (such as, for example, the direction and distance to an audio element) and the metadata 162 about the XR scene. As described above, the metadata 162 may include metadata for the XR space in which the user is virtually located. For example, the metadata 162 may include information about the dimensions of the space, information about the objects in the space, information about the acoustic characteristics of the space, as well as metadata about the audio elements and metadata about the objects that occlude the audio elements. In some embodiments, the controller 201 itself generates at least a portion of the metadata 162. For example, the controller 201 may receive metadata about the XR scene and derive additional metadata, such as, for example, control parameters, based on the received metadata. For example, using the metadata 162 and the position / orientation information 163, the controller 201 may calculate one or more gain factors (g) for the audio elements in the XR scene (e.g., the scaling factor described above).
[0048] Regarding the generation of the reverberation signal used by the audio signal generator 202 to produce the final output signal, the controller 201 provides the audio signal generator 201 with reverberation parameters (such as, for example, the reverberation time and reverberation level for the space in the XR scenario) and the scaling factor described above, so that the audio signal generator 202 is operable to generate the reverberation signal. The reverberation time for the generated reverberation is most commonly provided to the reverberation processor 204 as an RT60 value, typically for individual frequency bands, but there are also other reverberation time measurement methods and they can be used. In a typical embodiment, the metadata 162 includes all necessary reverberation parameters (e.g., RT60 value and reverberation level value). However, in embodiments where the metadata does not include all necessary reverberation parameters, the controller 201 can be configured to generate the missing parameters.
[0049] The reverberation level can be represented in various formats. Typically, it will be represented as a relative level. For example, it can be represented as the energy ratio (DRR) between the direct sound and the reverberant sound component at a certain distance from the sound source rendered in the XR environment or its reciprocal (i.e., the RDR energy ratio). Alternatively, the reverberation level can be represented based on the energy ratio between the reverberant sound and the total emitted energy or power of the source. In other cases, the reverberation level can be directly represented as the level / gain for the reverberation processor.
[0050] In this context, the term "reverberation" generally can just refer to those sound field components corresponding to the diffuse part of the acoustic room impulse response of the acoustic environment, but in some embodiments, it can also include sound field components corresponding to the early part of the room impulse response (e.g., including some late non-diffuse) or even all reflected sounds.
[0051] Other metadata describing the reverberation-related characteristics of the acoustic environment that can be included in the metadata 162 include parameters describing the acoustic properties of the materials of the surfaces of the environment (describing, for example, the absorption, reflection, transmission, and / or diffusion properties of the materials), or specific time points of the room impulse response associated with the acoustic environment, e.g., the time after the source emission when the room impulse response becomes diffuse (sometimes referred to as "pre-delay").
[0052] All of the reverberation-related properties described above are typically frequency-dependent, and thus, their relevant metadata parameters are typically provided and processed separately for multiple frequency bands.
[0053] Embodiment
[0054] In one embodiment, a computationally efficient method is provided for rendering reverberation in response to a source located in a connected space. Embodiments may optionally also include rendering reverberation propagated from the connected space, and / or generating and rendering further (“second-order”) reverberation in a virtual acoustic system in response to reverberation propagated from the connected space.
[0055] In real-life situations, the reverberation present in a given space is typically generated not only in response to sound sources present within that same space, but also (partially) in response to sound sources located outside of that space.
[0056] Consider Figure 3 the case where a first space (Space 1) is connected to a second space (Space 2) via a portal 300. Initially, assume that the portal is acoustically completely transparent, i.e., it is fully “open”. For example, it may be assumed that the portal is an open door, an open window, or some other open connection between the two spaces. A sound source S is located somewhere in Space 2.
[0057] The energy radiated by the source S will generate reverberation in Space 2 (“Space 2 reverberation”), and some of this Space 2 reverberation will propagate through the portal into Space 1, where it may in turn generate further reverberation (“second-order reverberation”).
[0058] However, a portion of the energy radiated by the source S can propagate directly through the portal, i.e., not as Space 2 reverberation, but as direct sound from the source S. This direct sound entering Space 1 through the portal generates reverberation in Space 1, which is a direct response to the sound emitted by the source S.
[0059] The challenges addressed in the present disclosure include: i) determining the amount or portion of direct sound energy / power radiated by the sound source S that propagates directly through the portal into Space 1; ii) generating (and optionally rendering) reverberation in Space 1 that is consistent with the determined amount or portion of direct sound energy entering Space 1 through the portal; and iii) defining a suitable rendering architecture for efficiently handling the above processes, also in use cases with many connected spaces.
[0060] Further challenges include: i) the realistic rendering in Space 1 of the Space 2 reverberation that propagates through the portal into Space 1, and vice versa, the realistic rendering in Space 2 of the Space 1 reverberation that propagates back through the portal into Space 2 (which is essentially the same problem); and ii) the realistic rendering in Space 1 of “second-order” reverberation that is generated in Space 1 in response to the Space 2 reverberation that propagates through the portal.
[0061] The situations and challenges described above are in principle independent of the virtual location of the user of the XR system (also referred to as the "user" or "listener") within the generated XR scene. However, depending on the space in which the listener is located, the relevance and / or order in which different challenges occur may vary, and thus, the renderer may be configured to perform any of the processes described above in any combination or order depending on the spaces in which the listener and the source are located.
[0062] To illustrate this, different processes are separated into "first-order" and "second-order" processes. In Figure 3 the case of, the "first-order processes" that occur include: (1) generating first-order spatial 2 reverberation in space 2; and (2) generating first-order reverberation in space 1 due to the direct sound of source S propagating through the portal; the "second-order processes" in this case include: (3) generating second-order reverberation in space 1 in response to the first-order spatial 2 reverberation entering space 1 through the portal; and (4) generating second-order spatial 2 reverberation due to the first-order spatial 1 reverberation entering space 2 through the portal.
[0063] If the listener is located in space 1, i.e., the sound source S is in a different space from the space in which the listener is located, then processes (2) and (3) result in the sound being rendered to the listener, while process (1) requires the ability to perform process (3). In this case, process (4) can be considered irrelevant to the execution, since there is no listener in space 2 to hear the second-order reverberation in space 2, and its result is not needed to generate the sound rendered to the listener in space 1.
[0064] In principle, second-order processes can be followed by third-order and higher processes, e.g., propagating the second-order reverberation in space 2 back to space 1, but the perceptual relevance of these higher-order processes may be so small that they can be considered irrelevant to the execution and / or rendering.
[0065] On the other hand, if the listener is located in space 2 (i.e., the listener is in the same space as the sound source), then the first and fourth processes result in the sound being rendered to the listener, while the second process requires the ability to perform the fourth process. In this case, the third process can be considered irrelevant to the execution and / or rendering.
[0066] Determining the amount or portion of the direct sound source power propagating through the portal into the connected space
[0067] The amount or portion of the power radiated by the sound source propagating through the direct path through the portal depends on several factors, including: i) the location and size / shape of the portal, ii) the location of the sound source, iii) the directional pattern of the sound source, and iv) the orientation of the sound source.
[0068] One method for determining the amount or fraction of the source power that propagates directly through a portal is to use the "line-of-sight" method, which, in one embodiment, includes the following steps:
[0069] (1) Project the (edges of the) portal geometry onto an imaginary sphere surrounding the sound source;
[0070] (2) Taking into account the orientation of the sound source, integrate the directivity pattern of the sound source over the segment of the sphere covered by the projection of the portal; and
[0071] (3) Normalize the value obtained by the surface area of the sphere to produce a value for direct propagation of the source S in space 2 relative to the portal to space 1, where the direct propagation value indicates the amount or fraction of the power radiated by the sound source that propagates directly through the portal.
[0072] The latter step normalizes the amount of power propagating directly through the portal relative to the total amount of power radiated by an omnidirectional source driven with the same input signal. In many embodiments, the sphere is a unit sphere, and the surface area of the unit sphere is 4π.
[0073] If the directivity pattern of the source is normalized, i.e., it has a value of 1 in the highest output direction of the source, the process will produce a value between 0 and 1. In the case where the portal is a flat surface, the maximum value is 0.5, since an omnidirectional source radiates at most half of its power directly through a flat portal. The assumption here is that the directivity pattern is expressed in terms of the intensity / power radiated into individual directions, i.e., it is proportional to the square of the rms signal amplitude / pressure of the direct sound of the source measured for individual directions at equal distances from the source.
[0074] The directivity pattern of a sound source indicates the amount of acoustic power that the sound source radiates into individual directions. A simple omnidirectional sound source radiates sound equally in all directions, such that the directivity pattern has the same value (1, in the case where the directivity pattern is normalized) in all directions, but typically, a sound source will radiate different amounts of power into different directions.
[0075] In an XR system, the directivity pattern of a sound source will typically be directly available from the metadata 162 corresponding to the sound source. It can typically be provided as (usually normalized) dB values for individual directions, corresponding to the (normalized) sound pressure level (SPL) measured at equal distances around the sound source in each direction, or as (usually normalized) linear gain values for individual directions.
[0076] Alternatively, the directivity pattern can be provided in some other suitable format, such as a spherical harmonic representation, from which the directivity pattern can be derived for any desired set of directions.
[0077] In other cases, the directional pattern can be provided based on a "directional factor" for an individual direction, where the directional factor for a particular direction quantifies the ratio between the sound intensity radiated in that direction and the intensity averaged over all directions.
[0078] It should be clear that all these representations for the directional pattern are substantially equivalent and have a direct relationship with each other, and any one of them can be used in the context of the present disclosure by converting it into the representation assumed in the methods described above.
[0079] The initial position and orientation of the sound source will also typically be available as metadata for the sound source. In dynamic and / or interactive XR scenarios, the source position and orientation may change, but typically they are readily available to the audio renderer.
[0080] Information about the position / size / shape of the portal can also be directly available or can be readily derived from the scene description metadata or can be derived in some other way as described in more detail below.
[0081] As a simple example, consider the case where an omnidirectional source is accurately positioned at the middle of a (flat) open portal between space 1 and space 2. Figure 3 In this case, the portal occupies a solid angle of 2π (i.e., a hemisphere) relative to the sound source. Thus, since the source is omnidirectional, we find that half (2π / 4π) of the total power amount radiated by the omnidirectional source propagates directly through the portal, and thus, the direct propagation value we obtain is 0.5.
[0082] Conversely, if the sound source has a non-omnidirectional directional pattern and the orientation of the sound source is such that the "loudest side" of the source is away from the portal, the direct propagation value will be less than 0.5. In the extreme case where the sound source has a directional pattern such that all of its energy is radiated into a hemisphere and the source is oriented such that its hemispherical directional pattern exactly points away from the portal, the direct propagation value will be 0.
[0083] In a more general case as depicted in Figure 3 determining the direct propagation value for a sound source S relative to a portal between space 1 and space 2 first requires the step of determining the projection of the edges of the portal geometry on a hypothetical unit sphere surrounding the sound source S while taking into account the position of the source.
[0084] Assume that the projection of the portal covers a solid angle Ω on the unit sphere surrounding source S portal , and the source S has a normalized directional pattern D(ω) (expressed in power), where ω is the spatial angle and the orientation R is relative to some reference orientation. Then, the direct propagation value can be obtained from the integral over the solid angle Ω portalThe direct propagation value is determined by integrating the rotational directional pattern and normalizing it by the surface area 4 of the unit sphere:
[0085]
[0086] where D rot,R is the directional pattern D suitably rotated according to the source orientation R.
[0087] Note that if the source is omnidirectional, the above process simply reduces to determining the normalized surface area of the projection on the unit sphere.
[0088] To determine the direct propagation value based on the segment covered by the projection of the portal on the source directional pattern, a process based on the scene geometry needs to be performed to determine the projection and its size (area) on the unit sphere.
[0089] In one exemplary embodiment, the analysis is performed on the encoder device based on the scene geometry and the object source location. The analysis can also be performed by the renderer. In this exemplary embodiment, the portal opening is typically a box with eight vertices and represents, for example, the opening of a door between two spaces. In the case where the portal opening is not a box, the portal geometry can be transformed into a box, for example, by surrounding the portal geometry with a box.
[0090] There are different ways to obtain the portal opening from the scene geometry, which can be represented as a mesh or voxels. In one example, the content creator provides explicit metadata indicating the location of the portal, for example, as a set of vertices. Then, this portal metadata can be carried as metadata. In another example, the processing on the encoder / renderer device can be used to automatically analyze the portals between the acoustic environments. In the exemplary method, each space is surrounded by a geometry (such as a box). Whenever there is a portal between the spaces, there is a line of sight between the spaces. This line of sight can be determined, for example, by emitting rays in all possible directions from each space. Whenever a ray emitted in this way hits another space, it means that the ray has passed through the portal. Combining all the ray hits from the first space to the second space will give a rough shape of the projection of the portal between the spaces on the surface of the second space. This can be considered as one face of the portal. The starting points of the rays that hit the surface of the first space will define the other face of the portal geometry. The complete portal geometry can be obtained by forming a hull to combine these two faces.
[0091] In an example embodiment, for each object source in a scene, four vectors are aligned from the position of the object source ("objsrc_pos") to the four vertices of the portal closest to objsrc_pos to generate the geometry of the sound emitted from the sound source towards the portal. Then, four 3D points are translated 1m along each vector, and a mesh is constructed using these points as vertices. The area of the formed mesh is divided by the area of a unit sphere (4π). This yields the direct propagation value to be used in the renderer. It should be noted that in this example embodiment, the directional pattern is assumed to be omnidirectional. In some other examples, the directional pattern can be non-omnidirectional, and in these examples, an integration is performed over the directional pattern, as previously described.
[0092] The vertices for the portal opening are represented as [vpopen0, vpopen1, vpopen2, vpopen3]. These vertices are chosen as the corners of the face of the portal that is at least partially parallel to the wall of the space in which the portal is located and whose center point is closest to the center of the space. It should be noted that if the face is not rectangular, it can be bounded by a rectangle, and the boundary rectangle corners can be used as vertices. Alternatively, if the face is not rectangular, four almost equidistant points can be selected from the face perimeter to be used instead of vertices. In some other examples, the face of the portal representing the opening of the portal from the space can be indicated in other ways, such as manually by the content creator. In some other other examples, the vertices can be selected from a vertical cross-section of the portal in the middle of the portal geometry. In some other other examples, the vertices can be selected from the outermost face of the portal, and thus, towards the space where the power is transferred.
[0093] For each vertex, a 3D vector Vobjsrc_vi between objsrc_pos and vpopen0, vpopen1, vpopen2, vpopen3 is formed, where i = 0, 1, 2, 3.
[0094] Figure 4 [Vobjsrc_v0, Vobjsrc_v1, Vobjsrc_v2, Vobjscrc_v3] are respectively visualized as being towards the portal opening vertices [vpopen0, vpopen1, vpopen2, vpopen3].
[0095] To obtain the area at a distance of one meter from the sound source, four 3D points (P0, P1, P2, P3) are defined for each vertex of the portal opening [vpopen0, vpopen1, vpopen2, vpopen3]. This is achieved by translating objsrc_pos successively along each vector.
[0096] When determining the points Pi (i = 0, 1, 2, 3) along the vector Vobjsrc_vi towards the vertex vpopeni, first check the distance d between objsrc_pos and vpopeni to avoid translating P i by more than a specific threshold (basically, beyond the point vpopeni). To maintain the shape of the mesh, if plane_distance > d, subtract the value 0.1 from d. Otherwise, objsrc_pos is translated by plane_distance (in one example, plane_distance = 1), as shown in the following code:
[0097]
[0098] obj_src is translated by the amount d along each vector and assigned new x, y, z positions in 3D Cartesian space, as shown in the following code:
[0099]
[0100] the mesh objsrc is constructed with Pi (i = 0, 1, 2, 3) as vertices, as Figure 5 visualized. That is, Figure 5 shows the mesh constructed by Pi (i = 0, 1, 2, 3).
[0101] the constructed mesh objsrc has its area Amesh_objsrc divided by the area of a sphere with a diameter of one meter to obtain the direct propagation value for the object source. That is:
[0102] sphere_radius = 1
[0103] Asphere = 4π * sphere_radius 2
[0104] directPropagationValue = A mesh_objsrc / A sphere
[0105] Algorithms for calculating the area of a mesh are readily available in the literature and software libraries. For example, the mesh can be triangulated and iterations performed on the triangles of the mesh. For each triangle, vectors representing two sides are formed. The area of the triangle is obtained as half the magnitude of the cross product of the side vectors. The triangle areas are summed to obtain the complete mesh surface area.
[0106] The value directPropagationValue is used in the renderer as a coefficient for the direct sound power leaking through the portal opening into the connected room. When running the above process for determining directPropagationValue on the encoder device, the following data can be written into the payload defined in RevPortalOpeningData (see below).
[0107] Syntax for transmitting the direct propagation value in the bitstream
[0108] The following table shows the metadata syntax for providing the direct propagation value from the encoder to the renderer.
[0109]
[0110]
[0111] revNumSpaces is the number of spaces.
[0112] revNumPortalOpenings is the number of portal openings per space.
[0113] spaceBsId is the bitstream identifier of the space.
[0114] portalOpeningPositionX is the x element of the center position of the portal opening in the x,y,z space.
[0115] portalOpeningPositionY is the y element of the center position of the portal opening in the x,y,z space.
[0116] portalOpeningPositionZ is the z element of the center position of the portal opening in the x,y,z space.
[0117] objSrcBsId is the bitstream identifier of the object source. Carried in the ScenePayload.
[0118] revNumObjsrcSpaces is the number of spaces for which analysis is performed. This is greater than 1 if the objsrc is not located in any space, otherwise equal to 1.
[0119] revNumObjsrcPortalOpenings is the number of portal openings in the space being iterated.
[0120] directPropagationValue is the direct propagation value for each object source towards each portal opening.
[0121] The openingConnectionBsId is an identifier for a portal between two spaces.
[0122] It should be noted that what is listed in the above examples is for the space in which the audio object is located. For example, if the sound source is static, one space is sufficient. If the object is not static, several such lists and thus direct propagation values can be listed for more than one possible space and / or objsrc location. For a source that is not in any space, a list can be provided for all possible spaces and portal openings. Alternatively or additionally, for an object source that is not in any space, a list of the closest or otherwise most relevant space and / or portal opening can be provided.
[0123] In the example embodiment described above, the direct propagation value is derived by directly determining the surface area of the projection of the portal on the unit sphere (or rather an approximation of that projection), and does not include the step of integrating the directional pattern of the source over the segment of the unit sphere covered by the projection of the portal. As explained above, this effectively means that in the method of the example embodiment, the source is assumed to be omnidirectional.
[0124] For a static sound source, the example embodiment described above can be extended to apply to non-omnidirectional sources by replacing the step of directly determining the surface area of the projection of the portal with the step of integrating the directional pattern of the source over the segment of the unit sphere covered by the projection, taking into account the orientation of the source as previously described.
[0125] Basically, the same process as described in the above example embodiment or an extended version thereof that includes the effect of the directional pattern can be performed in real time by the renderer. In this case, when calculating the direct propagation value for the source relative to the portal, the dynamic changes in the source position and orientation can also be taken into account. In this case, the direct propagation value will be a dynamic function of the source position and orientation.
[0126] In another example embodiment for determining the projection of the portal on the unit sphere surrounding the sound source and the size of the projection, a method similar to that described in Article 6.5.5 ("DiscoverSESS") and Article 6.5.16 ("Homogeneous Range") of the working draft (WD1) of ISO / IEC 23090-4 MPEG-I Immersive Audio [1] can be used. Similar to the method described above, the method described there uses ray tracing to find the projection (and thus the size) of the extended sound source or the portal between spaces on the unit sphere, except that there it is the projection on the unit sphere surrounding the listener, rather than the source. However, with appropriate adaptation, the same method can be applied to find the projection of the portal relative to the unit sphere surrounding the sound source and the size of the projection.
[0127] Rendering of first-order reverberation in a connected space
[0128] Once the direct propagation value of the normalized radiant source power amount that propagates directly from the source S in space 2 to the connected space 1 through the portal has been determined, this information can be used to generate and render reverberation in space 1 at a level consistent with this power amount.
[0129] The principle here is to generate reverberation in space 1 corresponding to a conceptual (also known as "imaginary") omnidirectional sound source located at an arbitrary position within space 1, where the source power is equal to the amount of direct sound source power that enters space 1 through the portal.
[0130] Now it will be explained how this can be done by the renderer.
[0131] The power P of an omnidirectional point source radiating spherical sound waves is proportional to the square of the direct sound pressure p of the source:
[0132]
[0133] where r is the distance from the source, and ρ0 and c are the mass density and speed of sound in air, respectively.
[0134] Comparing the acoustic (physical) signal processing domain and the audio signal processing domain, the rms acoustic sound pressure p is proportional to the linear rms signal level of the audio input signal of the audio source, while the sound source power P is proportional to the square of this rms signal level of the input audio signal.
[0135] Therefore, if we denote the normalized source power amount that directly enters space 1 through the portal, i.e., the direct propagation value, as X, it can be concluded from the above that the audio signal for the conceptual omnidirectional source used to generate reverberation in space 1 is the audio signal corresponding to the source S scaled by the linear gain sqrt(X). That is, if we denote the signal of the sound source S in space 2 and the signal of the conceptual sound source in space 1 as s2 and s1, respectively, then: s1 = sqrt(X) * s2.
[0136] Using s1 as the input signal, reverberation for space 1 can be generated according to the provided reverberation characteristics of space 1 (such as, for example, reverberation time, reverberation energy ratio, etc.).
[0137] In the above, it is assumed that the audio signal s2 of the source S includes all factors that affect the direct sound level of the source when it is rendered, except for the directivity pattern and the distance from the source. Examples of such factors include the source gain ("volume control") that may be associated with the source, the so-called "reference distance" (which may specify the distance from the source at which its distance attenuation should be normalized to 1 (0 dB)), or muting the source S (which is equivalent to applying a source gain equal to zero). If this is not the case, i.e., if the source signal s2 associated with the source S is the "raw" input signal to which direct sound level influencing factors (such as the associated source gain, or reference distance, or muting of the source) are applied when rendering the source, then these factors should also be applied to the signal s2 in the above process. That is, if the combined effect of the direct sound level influencing factors is a direct sound gain G, then the above equation can be modified to: s1 = sqrt(X) * G * s2 (G can be equal to 1).
[0138] Rendering of first-order reverberation for all sources in all connected spaces
[0139] In an actual audio scenario, there may be several spaces connected by portals. Therefore, the method needs to iterate over all sound sources S to determine whether it should be input into the reverberation for space 1. This can be done as follows:
[0140] Determine the space in which the source S is located;
[0141] If S is located within space 1, then render the reverberation for the source S according to the reverberation characteristics for space 1;
[0142] If S is not located within space 1 but within space 2, then determine whether there is a portal from space 2 to space 1, and
[0143] If such a portal exists, then determine the direct propagation value for the source S with respect to the portal from space 2 to space 1, and use the concept (also called the imaginary) sound source in space 1 as described above to render the reverberation for the source S.
[0144] Occlusion detection and processing
[0145] If the source is not "visible", i.e., there is no direct line of sight through the portal, the direct propagation value for the source propagating through the portal can optionally be reduced. This can be the result of the geometric characteristics of the space (e.g., a wall between the source S and the portal), an object that can move around and block (e.g., temporarily block) the direct line of sight between the source and the portal (e.g., an interactive object), or a change in the state of the portal itself (e.g., a door that is open between two spaces suddenly closes). Such occlusion detection can be performed, for example, by emitting one or more rays from the position of the portal towards the sound source. If one or more rays hit the sound source, it can be determined that the source is visible and the direct propagation value does not need to be reduced. However, if no rays hit the sound source, it can be determined that the source is not visible and the direct propagation value can be reduced.
[0146] When the sound source becomes visible or stops being visible, optionally, further smoothing of the direct propagation value can be applied to prevent sudden sound level transitions. This smoothing of the direct propagation value can be performed by gradually increasing the direct propagation value (during the period when the source becomes visible) or gradually decreasing the value (during the period when the source becomes invisible). The ramp of the value can be performed, for example, by increasing or decreasing the value by a predetermined value (such as 0.05) in each audio frame until the desired value is reached.
[0147] It should be noted that the above cross-fading can be equivalently performed on the direct propagation value or its square root ("sqrt"), which is the gain value applied to the audio signal.
[0148] In another exemplary embodiment for determining the amount of occlusion of the portal relative to the sound source, a method similar to the method described in Article 6.5.6 ("Occlusion") of the working draft (WD1) of ISO / IEC 23090-4 "MPEG-I Immersive Audio [1]" can be used. The method described there uses ray tracing to project rays from the listener's position towards the extended sound source to find the amount of occlusion for the direct sound path between the source and the listener. With appropriate adaptation, the same method can be applied to find the amount of occlusion of the portal relative to the sound source.
[0149] Partially transparent portal
[0150] Above, for simplicity, it was assumed that the portal between the two spaces is acoustically completely transparent, i.e., it is fully open. Typically, examples are an open door, a window, or an open portal.
[0151] However, the same concepts and principles also apply to portals that are not completely open but only partially transmissive, i.e., general interfaces between spaces through which acoustic energy can propagate. The only modification needed for the generalization of this concept is that the propagated power and thus the direct propagation value should be scaled by a value equal to the fraction of the energy transmitted by the portal compared to the energy transmitted by a fully open portal of the same geometric size.
[0152] For example, if the interface between two spaces (e.g., a thin wall) is made of a material that transmits 30% of the power impinging on it to the adjacent space, the direct propagation value for a source relative to this interface will be 30% of the value calculated assuming the interface is an open portal.
[0153] In the general case where the energy transmission coefficient can vary with the portal, the scaling is obtained by integrating the energy transmission coefficient over the portal area and normalizing by the geometric area of the portal.
[0154] It should also be noted that the transmission coefficient of the portal can be frequency-dependent, such that the direct propagation value can also be frequency-dependent. This will be reflected, for example, in generating relatively more low-frequency reverberation than high-frequency reverberation in the connected spaces.
[0155] Cascading of connected spaces
[0156] If more than two spaces are connected to each other in a cascaded manner via multiple portals, the method described above can be extended. In this case, the main challenge is to determine the amount or fraction of the sound source power that propagates directly through the multiple portals, rather than the case of propagation through a single portal as described above.
[0157] This can be achieved by projecting the portals between successively connected spaces onto a unit sphere surrounding the sound source and determining the overlapping portions of the projections. These overlapping portions represent the angular regions where there is a direct "line of sight" between the source and the "farthest" connected space.
[0158] In the above bitstream description, the portals can be defined with the help of the starting and ending spaces: the portal connecting spaces 1, 2, and 3 such that there is a path from space 1 through space 2 into space 3 can index spaces 1 and 3.
[0159] Pre-computation of direct propagation values for any source position within a space
[0160] In another method, an interpolation grid of pre-computed values can be used to approximate the direct propagation value for an object source based on the position of the object source.
[0161] Each space has a boundary geometry (“spacebounding”) that defines the extent of the space. The object source location is interpolated from the x and y plane extents of the spacebounding at a resolution defined by interpResision. In this example, as shown in the code below, the value 0.5 is used.
[0162]
[0163]
[0164] The method distanceTo is further described as (pow(x,y) represents x to the power of y).
[0165] def distanceTo(self,another):
[0166] if not isinstance(another,Point3D):
[0167] raise ValueError("Only Point3D type is supported")
[0168] return sqrt(pow(self.x-another.x,2)+pow(self.y-another.y,2)+pow(self.z-another.z,2))
[0169] Then, the interpolated object source locations are used, and an object source-portal opening analysis is performed for each location using the above process. A grid of possible object source locations with corresponding direct propagation values is obtained.
[0170] Figure 6 The concept of a grid of interpolated direct propagation values for possible object source locations is shown. As Figure 6 depicted, all possible x, y locations interpolated with interpResolution = 0.5 are assigned direct propagation values that depend on the angle and distance. If the object source moves around in the space, this information can be used to approximate the behavior of the direct propagation values.
[0171] In some examples of the implementation, the direct propagation values within the space for different possible object locations are modeled with a function of two variables. An example is a function of two variables (x, y), such as a polynomial function of two variables. The coefficients of such a polynomial can be signaled to the renderer device in a bitstream and used there to obtain the direct propagation value for any source location within the space.
[0172] Reverberation Propagation through the Portal and Generation of "Second-Order" Reverberation
[0173] The embodiments described above focus on the reverberation generated in the first space as a result of the direct sound from a source in the second space propagating through the portal from the second space to the first space. However, further challenges have been found in the propagation of reverberation from one space through the portal to another space and the subsequent generation of second-order reverberation in response to the propagated reverberation.
[0174] Return to reference Figure 3 In the scenario of , these further challenges can be summarized as: (1) the realistic rendering in space 1 of the space 2 reverberation propagated through the portal into space 1, or vice versa, the realistic rendering in space 2 of the space 1 reverberation propagated back through the portal to space 2 (which is essentially the same problem), and (2) the generation and realistic rendering in space 1 of the "second-order" reverberation generated in space 1 in response to the space 2 reverberation propagated through the portal.
[0175] Similar to the propagation of direct sound through the portal, the key component in solving the problem of reverberation propagation through the portal and the subsequent generation of second-order reverberation is to determine the amount of power transmitted through the portal (in this case: diffuse sound field power).
[0176] It can be shown that the amount of diffuse power P transmitted from the first space to the second space through the portal 1->2 has the following relationship with the diffuse sound pressure p1 in the first space:
[0177]
[0178] where S portal is the size of the portal in m 2 . If the portal is fully open, then S portal is equal to the geometric size of the portal, otherwise it is equal to the size of a fully open portal representing the same total sound transmission, ρ0 is the mass density of air, and c is the speed of sound in air. The size S of the portal portal can be obtained directly from the scene description metadata or derived from it. For example, the size can be derived from the vertices describing the portal or from the mesh constructed from the vertices, as explained in detail above. Also as described above, ray tracing techniques can be used to find the edges of the portal, from which the size of the portal can be determined.
[0179] Therefore, if the level of reverberation in the first space and the (equivalent) size of the portal are known, the amount of diffuse power transmitted to the second space can be determined.
[0180] Based on the amount of diffused power transmitted, a second-order reverberation can be generated in the second space, similar to that described in the paragraph "Rendering of First-Order Reverberation in Connected Spaces" above. That is: the transmitted diffused power P 1->2 is assigned to a conceptual point source located at an arbitrary position in the second space.
[0181] It can be shown that this results in a relationship between the diffused pressure p1 in the first space and the direct sound pressure p at a distance of 1 m from the conceptual source in the second space, given by the following equation: 2,1m between, is given by:
[0182]
[0183] If the reverberation intensity of the second space is expressed as the reverberation-to-direct energy ratio (or its reciprocal) at a distance of 1 m from an omnidirectional point source, then this equation provides all the information needed to generate the correct level of second-order reverberation in the second space. Specifically, if the reverberation-to-direct energy ratio for the second space is denoted as RDR2, then the desired pressure of the diffused reverberation in the second space can be given by the following equation:
[0184]
[0185] In one embodiment, the conceptual process described above can be implemented by the following steps:
[0186] (1) Derive one or more reverberation input signals for generating second-order reverberation in the second space;
[0187] (2) Obtain the size of the portal through which sound is transmitted between the first space and the second space;
[0188] (3) Use the portal size to determine a reverberation scaling factor that simulates the transmission of reverberant sound from the first space to the second space; and
[0189] (4) Use the (one or more) reverberation input signals and the reverberation scaling factor to render a reverberation signal in the second space. In some embodiments, the reverberation signal in the second space is also rendered using information indicating the intensity of the reverberation in the first space.
[0190] More details and additional embodiments for generating and rendering second-order reverberation are described in the provisional application US 63 / 429643, the relevant portions of which are included in the "Additional Disclosure" section.
[0191] The generation and rendering of second-order reverberation in response to reverberation propagating from a first reverberant space to a second reverberant space through a portal has been described above. However, if the second space is a free field, i.e., has no reverberation of its own, the reverberation propagating from the first space to the second space should still be heard as reverberation from the portal in the second space. An example of this is when a listener stands in a large open outdoor space in front of an open door of a cathedral where music is being played.
[0192] In this case, the portal acts as an extended acoustic "source" whose source power is equal to the transmitted diffusion power P as described above. 1->2 . Thus, the steps for determining the correct power for this portal source are essentially the same as those for second-order reverberation described above, but the main difference now is that the portal is not a (conceptual) omnidirectional point source, but an extended source that effectively radiates all of its source power P 1->2 only into a half-space (i.e., the second space).
[0193] In one embodiment, the radiation from the portal can be modeled as hemispherical (i.e., radiating only spherical waves into the second space), in which case the relationship between the diffusion pressure p1 in the first space and the pressure p at a distance of 1 m from the portal source can be given by: 2,1m as follows:
[0194]
[0195] or
[0196]
[0197] where S portal is as previously defined.
[0198] In another embodiment, the pressure p2 at a specific location in the second space can be determined from the diffusion pressure p1 in the first space as:
[0199]
[0200] where Ω portal is the solid angle by which the portal is "seen" from a specific location in the second space, i.e., the angular size of the projection of the portal on a sphere surrounding the specific location.
[0201] In other embodiments, a specific spatial radiation model for the portal source can be used to determine the desired relationship between the intensity of the reverberation in the first space (which is proportional to the diffusion pressure p1) and the rendered sound level in the second space.
[0202] Although the above describes the rendering of reverberation of propagation for an example where the second space is a free field (i.e., non-reverberant), this also applies when the second space is reverberant, i.e., in this case, the reverberation propagating from the first space into the second space can also be rendered as a "portal source" as described above, plus generating and rendering "second-order" reverberation.
[0203] Additional details and additional embodiments for rendering reverberation propagating from a portal are described in provisional application US 63 / 429643, the relevant portions of which are included in the "Additional Disclosure" section.
[0204] Architecture
[0205] Figure 7 Shows an efficient architecture for use in implementing some of the embodiments described above.
[0206] N audio sources S1…Sn are located in a plurality of acoustic environments (e.g., rooms) AE1, AE2, AE3…AEn. For some AEs, there are acoustic connections (i.e., portals) that acoustically connect the AE to one or more other AEs.
[0207] Each AE has an associated feedback delay network (FDN) (or other kind of reverberator) that models the late reverberation characteristics of that AE considering relevant parameters such as frequency-dependent reverberation time (RT60), reverberation energy ratio, and pre-delay.
[0208] For all sources, the rendering of their direct signal parts and their early reflections is performed using well-known methods and is based on the position / orientation of the source and the position / orientation of the listener.
[0209] For the rendering of the late reverberation parts of these sources, the source signals are fed into a scaling factor matrix, which weights each contribution from a particular source into a particular target AE / FDN, and then these contributions are summed into the FDN of that AE. The scaling factors in this matrix are the direct propagation values described above, which determine how much direct acoustic energy propagates into adjacent rooms / AEs to excite the late reverberation in those rooms / AEs. These factors can also include coupling factors (such as the reverberation scaling factors described above), which reflect how much of the late reverberation energy from the space / room / AE propagates back to the listener's position. Reverberation scaling factors / coupling factors in energy representation need to be converted into amplitude factors using a square root function in order to produce linear signal scaling factors.
[0210] Then, the summed direct source contributions are fed into each AE / FDN. The output of the AE / FDN can be rendered using virtual speakers, if so desired.
[0211] Finally, all AE / FDN contributions are added together with the rendered direct sound and early reflection sounds to form the fully rendered output.
[0212] Figure 8 FIG. 8 is a flowchart showing a process 800 for generating an input signal for a first space (e.g., space 301) of an XR scene based on a first sound source (e.g., sound source 391) in a second space (e.g., space 302) of the XR scene (e.g., XR scene 390), where the first space is directly or indirectly connected to the second space via one or more portals including at least a first portal (e.g., portal 300). Process 800 may start at step s802. Step s802 includes: obtaining a direct propagation value, where the direct propagation value indicates an amount or portion of source power associated with the first sound source in the second space that is directly propagated from the first sound source through the one or more portals into the first space. Step s804 includes: using the direct propagation value and an audio signal associated with the first sound source to generate an input signal for the first space.
[0213] In some embodiments, the method further includes: using the input signal for the first space to generate a reverberation signal for the first space.
[0214] In some embodiments, the reverberation signal for the first space is generated using reverberation control information associated with the first space.
[0215] In some embodiments, the reverberation control information includes a reverberation level or a reverberation energy ratio parameter associated with the first space.
[0216] In some embodiments, the method further includes: rendering the reverberation signal.
[0217] In some embodiments, s1 = sqrt(X) * G * s2, where s1 is the input signal for the first space, X is the direct propagation value, G is a gain factor, and s2 is the audio signal associated with the first sound source.
[0218] In some embodiments, the method further includes: receiving metadata associated with the XR scene, the metadata including the direct propagation value or data from which the direct propagation value can be derived, and obtaining the direct propagation value includes: obtaining the direct propagation value from the metadata, or obtaining the direct propagation value using the data.
[0219] In some embodiments, the metadata includes coefficients of a polynomial, and obtaining the direct propagation value includes: obtaining the direct propagation value using the polynomial coefficients.
[0220] In some embodiments, obtaining the direct propagation value includes: using information indicating the position of a first sound source in a second space and one or more of the following items to derive the direct propagation value: i) information indicating the position of a first portal, ii) information indicating the size of the first portal, and iii) information indicating the shape of the first portal.
[0221] In some embodiments, one or more of the following are also used to derive the direct propagation value: information indicating the directional pattern of the first sound source or information indicating the orientation of the first sound source.
[0222] In some embodiments, information indicating the directional pattern of the first sound source and / or information indicating the orientation of the first sound source are also used to derive the direct propagation value.
[0223] In some embodiments, the process further includes: receiving metadata associated with an XR scene, the metadata including portal location information indicating the positions of one or more portals, and obtaining the direct propagation value includes: using the portal location information to derive the direct propagation value.
[0224] In some embodiments, for each portal, the metadata includes information indicating the geometry of the portal.
[0225] In some embodiments, the first portal is associated with a transmission coefficient, and the direct propagation value is obtained using the transmission coefficient.
[0226] In some embodiments, obtaining the direct propagation value includes: using an interpolation grid of pre-computed values and information indicating the position of the first sound source to obtain the direct propagation value.
[0227] In some embodiments, obtaining the direct propagation value includes: forming a grid representing the first portal; determining the area of the grid; and using the area of the grid to calculate the direct propagation value.
[0228] In some embodiments, obtaining the direct propagation value includes: projecting the edges of the first portal onto a sphere surrounding the sound source; determining the surface area of the spherical segment covered by the projection of the first portal to obtain a value; and normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
[0229] In some embodiments, obtaining the direct propagation value includes: projecting the edges of the first portal onto a sphere surrounding the sound source; integrating the directional pattern associated with the sound source over the spherical segment covered by the projection of the first portal to obtain a value; and normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
[0230] In some embodiments, the sphere is a unit sphere, and the surface area of the unit sphere is 4π.
[0231] In some embodiments, the first portal has eight vertices.
[0232] In some embodiments, the geometry of the first portal is transformed into a bounding box by surrounding the portal geometry with a box.
[0233] In some embodiments, the process further comprises: generating a first reverberation signal for a first space using an input signal for the first space; obtaining a first scaling factor, wherein the first scaling factor indicates the amount or portion of reverberation energy associated with the first reverberation signal propagated into a second space through one or more portals; generating a second input signal for the second space using the first scaling factor and the first reverberation signal; and generating a second reverberation signal for the second space using the second input signal for the second space.
[0234] Figure 9 FIG. 11 is a flow chart illustrating a process 900 for enabling rendering of reverberation in a first space (e.g., space 301) of an XR scene (e.g., XR scene 390), the first space being directly or indirectly connected to a second space (e.g., space 302) of the XR scene via one or more portals including at least a first portal (e.g., portal 300), wherein a sound source (e.g., sound source 391) is present in the second space. Process 900 may begin at step s902. Step s902 includes: determining a direct propagation value, wherein the direct propagation value indicates the amount or portion of source power associated with the sound source that is directly propagated from the sound source in the second space through one or more portals into the first space. Step s904 includes: storing and / or transmitting the direct propagation value or data from which the direct propagation value can be derived.
[0235] In some embodiments, the process further comprises: sending metadata to an audio renderer, wherein the metadata includes the direct propagation value or data from which the direct propagation value can be derived.
[0236] In some embodiments, determining the direct propagation value includes: forming a mesh representing the first portal; determining the area of the mesh; and using the area of the mesh to calculate the direct propagation value.
[0237] In some embodiments, determining the direct propagation value includes: projecting the edges of the first portal onto a sphere surrounding the sound source; determining the surface area of the spherical segment covered by the projection of the first portal to obtain a value; and normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
[0238] In some embodiments, determining the direct propagation value includes: projecting the edges of a first portal onto a sphere surrounding the sound source; integrating a directional pattern associated with the sound source over the segment of the sphere covered by the projection of the first portal to obtain a value; and normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value. In some embodiments, the sphere is a unit sphere, and the surface area of the unit sphere is 4π.
[0239] Figure 13 is a flowchart showing a process 1300 for rendering first-order reverberation in a first space (e.g., space 301) of an XR scene (e.g., XR scene 390). The process 1300 may start at step S1302. Step S1302 includes determining whether a first sound source (e.g., sound source 391) is located in a second space (e.g., space 302). Step S1304 includes determining whether there is one or more portals connecting the first space and the second space. Step S1306 includes determining a first direct propagation value of the first sound source relative to the first portal as a result of determining that the first sound source is located in the second space and that there is at least a first portal (e.g., portal 300) connecting the first space and the second space. Step S1308 includes generating an input signal for the first space using the direct propagation value and an audio signal associated with the first sound source. Step S1310 includes generating a reverberation signal for the first space using the input signal for the first space.
[0240] Figure 10 is a block diagram of an apparatus 1000 for performing the methods disclosed herein according to some embodiments. That is, the apparatus 1000 may implement an audio renderer 151 or an encoder 169. When the apparatus 1000 implements an audio renderer, the apparatus 1000 may be referred to as an audio rendering apparatus, and when the apparatus 1000 implements an encoder, the apparatus 1000 may be referred to as an encoding apparatus. As Figure 10As shown, apparatus 1000 may include: a processing circuit (PC) 1002, which may include one or more processors (P) 1055 (e.g., general microprocessors and / or one or more other processors such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.); these processors may be co-located in a single housing or a single data center, or may be geographically distributed (i.e., apparatus 1000 may be a distributed computing device); at least one network interface 1048, which includes a transmitter (Tx) 1045 and a receiver (Rx) 1047 for enabling apparatus 1000 to send data to and receive data from other nodes connected to a network 100 (e.g., an Internet Protocol (IP) network) to which network interface 1048 is (directly or indirectly) connected (e.g., network interface 1048 may be wirelessly connected to network 100, in which case network interface 1048 is connected to an antenna arrangement); and a storage unit (also referred to as a “data storage system”) 1008, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 1002 includes a programmable processor, a computer program product (CPP) 1041 may be provided. CPP 1041 includes a computer readable medium (CRM) 1042 storing a computer program (CP) 1043, and computer program (CP) 1043 includes computer readable instructions (CRI) 1044. CRSM 1042 may be a non-transitory computer readable medium such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., random access memory, flash memory), etc. In some embodiments, CRI 1044 of computer program 1043 is configured such that when executed by PC 1002, CRI causes apparatus 1000 to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, apparatus 1000 may be configured to perform the steps described herein without code. That is, for example, PC 1002 may consist of only one or more ASICs. Thus, the features of the embodiments described herein may be implemented in hardware and / or software.
[0241] Overview of additional embodiments
[0242] A1. A method for generating an input signal for a first space (space 1) of an XR scene based on a sound source S in a second space (space 2) of the XR scene, where space 1 is connected to space 2 via a portal, the method comprising: obtaining a direct propagation value X, where the direct propagation value quantifies an estimate (e.g., a normalized amount) of the amount of radiant source power directly propagated from the sound source S in space 2 through the portal to the connected space 1; and using X and a signal s2 associated with S to generate an input signal s1 for space 1.
[0243] A2. The method according to embodiment A1 further includes: using s1 to generate a reverberation signal for Space 1.
[0244] A3. The method according to embodiment A2 further includes: rendering the reverberation signal.
[0245] A4. The method according to any one of embodiments A1 to A3, wherein s1 = sqrt(X) * G * s2, where G is a gain factor (e.g., G ≥ 0).
[0246] A5. The method according to any one of embodiments A1 to A4, wherein the method further includes: receiving metadata associated with an XR scene, the metadata including a direct propagation value, and the direct propagation value is obtained from the metadata.
[0247] B1. A method for enabling reverberation to be rendered in a first space (Space 1) of an XR scene, the first space being connected to a second space (Space 2) of the XR scene via a portal, wherein a sound source S exists in Space 2, the method includes: determining a direct propagation value X, where the direct propagation value quantifies an estimate (e.g., a normalized amount) of the radiant source power amount directly propagated from the sound source S in Space 2 through the portal to the connected Space 1; and storing and / or transmitting the direct propagation value.
[0248] B2. The method according to embodiment B1 further includes: sending the metadata to an audio renderer, where the metadata includes the direct propagation value.
[0249] B3. The method according to embodiment B1 or B2, wherein determining X includes: forming a grid; determining the area of the grid; and using the area of the grid to calculate X.
[0250] B4. The method according to embodiment B1 or B2, wherein determining X includes: projecting the edges of the portal onto a unit sphere surrounding the sound source; integrating the directivity pattern of the sound source over the spherical segment covered by the projection of the portal to obtain a value; and normalizing the obtained value by the surface area (4π) of the unit sphere, where X is the normalized value.
[0251] C1. A computer program including instructions that, when executed by the processing circuit of device 1000, cause the device to execute the method according to any one of the above embodiments.
[0252] C2. A carrier containing the computer program according to embodiment C1, where the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0253] D1. A device configured to perform the method of any of the above embodiments.
[0254] D2. The device according to embodiment D1, wherein the device includes a memory and a processing circuit coupled to the memory.
[0255] Additional disclosure
[0256] As described above, the present disclosure provides embodiments for generating a plausible rendering of reverberation in a complex XR scene with connected acoustic spaces. For example, the present disclosure provides a device for determining an acoustic coupling factor that indicates the amount of reverberation propagated from a first space ("acoustic environment") to a second space through a portal (e.g., an opening or a partially transmissive surface) connecting the two spaces. In one embodiment, the determined acoustic coupling factor is determined using information indicating the size of the portal.
[0257] In some embodiments, based on the acoustic coupling factor, an appropriate signal level is set for rendering one or more audio signals into the second space, where the one or more audio signals are derived from one or more reverberation signals corresponding to the first space. For example, in some embodiments, the determined acoustic coupling factor is used to derive a scaling factor, where the scaling factor is used to scale one or more audio signals derived from a set of one or more audio signals representing the reverberation in the first space. In some embodiments, the scaling factor is further determined based on the amplitude, power, or energy of the total reverberation signal received at a position in the first space.
[0258] In some embodiments, the signal level of the signal to be rendered into the second space is determined based on one or more acoustic parameters for the first space, and more specifically, based on a reverberation level parameter or a reverberation energy ratio parameter associated with the first space.
[0259] Theoretical framework
[0260] Figure 11A A scene (real life or VR) consisting of two spaces: Space X and Space Y is shown. The spaces are connected to each other via a portal 1100 (which may alternatively be referred to as an "opening", "hole", "interface", etc.). A sound source S1 is located somewhere in Space X. The sound source S1 generates a reverberant sound field in Space X.
[0261] The portal represents the interface between Space X and Space Y through which a certain portion of the reverberant sound energy can be exchanged between the two spaces. The portal has an "acoustic" size (also referred to as the "associated" size) (area) of S portal m 2 of. In the case where the portal is an acoustically fully transparent opening, such as an open door or window, the acoustic size S of the portalportal is only equal to the geometric size of the portal (e.g., the geometric area of the portal). More generally, if the portal is not fully open but only partially acoustically transparent (e.g., a thin wall or a thick curtain separating two spaces), the acoustic size S of the portal portal is the equivalent size of a fully transparent opening representing the same amount of energy "leakage". In other words, if the portal is not fully acoustically transparent, the acoustic size S of the portal portal will be smaller than its geometric size. Hereinafter, unless otherwise explicitly stated, whenever the "size" or "area" of the portal is mentioned, it refers to the "acoustic" size.
[0262] A portion of the reverberant sound energy generated by a sound source in space X propagates through the portal into space Y, where a listener L located in space Y hears the reverberation from space X through the portal.
[0263] To determine the level of reverberation in space X perceived by a listener L in space Y, it is necessary to determine the amount of reverberant sound energy transmitted from space X to space Y through the portal.
[0264] For simplicity, space Y is first considered to be a "free field", meaning it is a very large open (e.g., outdoor) space or a space with very high sound absorption such that no reverberant energy propagates back from space Y to space X, and thus, this is a one-way problem.
[0265] Assuming a steady-state diffuse sound field in space X, which means that the sound energy leaving space X per unit time through absorption and the portal connecting to the adjoining space is equal to the power of the sound source S1 in space X, it can be shown that due to the source S1 in space X, the so-called average reverberant energy density E in space X 1,1 is equal to:
[0266] E 1,1 =(4 / c)*(P1 / A ,tot ), (1)
[0267] where P1 is the sound power of the sound source S1 in space X (expressed in watts), A 1,tot is the total absorption in space X, including the absorption represented by the portal, expressed in equivalent absorption area (m 2 ), and c is the speed of sound in air (in m / s).
[0268] A 1,tot can be further specified in terms of the absorption amount A in space X excluding the portal 1,0 and the size S of the portal portal (in m 2 ) as:
[0269] A 1,tot= A 1,0 + S portal (2)
[0270] It can be further demonstrated that, under steady-state diffusion conditions, the power P transmitted from space X to space Y through the portal 1->2 (expressed in watts) is typically equal to:
[0271] P 1->2 = (c / 4) * E 1,1 * S portal (W) (3)
[0272] Combining equations 1, 2, and 3, we obtain the power transmitted from space X to space Y through the portal:
[0273] P 1->2 = (S portal / A 1,tot ) * P1 = (S portal / (A 1,0 + S portal )) * P1 (W) (4)
[0274] The factor (S portal / A 1,tot ) in equation 4 is called the "acoustic coupling factor" from space X to space Y, which indicates the fraction of the source power in space X that is transmitted to space Y under steady-state conditions.
[0275] It can be seen that the acoustic coupling factor is equal to the fraction of the total absorption in space X due to the portal. In other words, the power transmitted from space X to space Y is determined by the fraction of the total absorption A in space X represented by the size S portal of the portal 1,tot .
[0276] If the absorption A in space X due to the portal is very small compared to S 1,0 (e.g., if the walls of space X are highly reflective and / or the portal is a very large opening), then the acoustic coupling factor (S portal / A portal / A 1,tot ) is essentially equal to 1, and the amount of power transmitted through the portal is essentially equal to the power radiated by the source.
[0277] On the other hand, if the absorption A in space X due to the portal is very large compared to S 1,0 (e.g., if the walls of space X are highly absorptive and / or the portal is very small), then the acoustic coupling factor (S portal / A portal / A 1,tot ) is equal to (S portal / A 1,0), i.e., the ratio of the size of the portal to the absorption in the space X excluding the portal, in which case it will be a very small number, i.e., only a very small fraction of the source power is transmitted through the portal.
[0278] The average steady-state reverberant energy density E in the space is directly related to the root-mean-square steady-state reverberant sound pressure p in the space by the following relationship:
[0279] p 2 = E * ρ0 * c 2 (5)
[0280] where ρ0 is the mass density of air.
[0281] Therefore, using Equation 1, the steady-state reverberant sound pressure p1 in the space X due to the source S1 can be written as:
[0282]
[0283] Combining Equations 4 and 6, we find the relationship between the steady-state reverberant sound pressure p1 in the space X and the power transmitted to the space Y:
[0284]
[0285] Therefore, Equation 7 provides an expression for the amount of power transmitted from the space X to the space Y in terms of the diffuse sound pressure p1 in the space X and the size S of the portal portal Importantly, Equation 7 shows that if we know the diffuse sound pressure in the space X and the size of the portal, this directly gives the amount of power transmitted through the portal.
[0286] Now, to obtain an expression for the relationship between the diffuse sound pressure p1 in the space X and the sound pressure p2 in the space Y associated with the radiation through the portal, the power P 1->2 transmitted through the portal is required in relation to the resulting pressure p2 in the space Y.
[0287] If the portal is relatively small, we can assume that the reverberant energy transmitted through the portal radiates equally (i.e., spherically) in all directions of the space Y. Using this assumption, this results in a pressure p at a distance of 1 m from the portal 2,1m equal to:
[0288]
[0289] This is derived from the relationship between the sound source power P and the pressure p at a distance of 1 m from the source radiating a spherical wave:
[0290] p 2 = P * ρ0 * c / 4π (9)
[0291] Among them, in Equation 8, since the power P 1->4 is only radiated into the hemisphere on the spatial Y side of the portal, the power has been increased by a factor of 2 (therefore, the resulting pressure should be the pressure corresponding to a global radiation source of twice that power).
[0292] Combining Equation 7 and 8, the following relationship between the diffusion pressure p1 in spatial X and the rms pressure p at 1 m from the portal in spatial Y can be found: 2,1m between:
[0293]
[0294] According to the audio signal level in the audio rendering system, the rms sound pressure p is proportional to the rms signal level of the corresponding audio signal, such that Equation 10 provides a direct way to relate the desired rms audio signal level in spatial Y to the rms reverberant audio signal level in spatial X.
[0295] Therefore, according to the linear sound pressure or linear rms audio signal level, Equation 10 means that the reverberation in spatial X should be rendered from the portal into spatial Y with scaling, such that at 1 m from the portal, the resulting sound pressure or rms audio signal level is scaled relative to the diffusion sound pressure or rms audio signal level of the reverberation in spatial X times.
[0296] According to the logarithmic sound pressure level or logarithmic rms audio signal level (in dB), this means that the reverberation in spatial X is rendered in spatial Y such that the resulting level at 1 m from the portal is 10*log 10 (S portal / 8π) = 10*log 10 (S portal ) - 14 (dB).
[0297] As mentioned, this result is valid for small enough portals such that the assumption of spherical radiation from the portal is reasonable. Obviously, there are limitations to the applicability of Equation 10 because if the size S of the portal portal exceeds 8π m 2 , then the rms pressure p in spatial Y 2,1m will be greater than the rms pressure p1 in spatial X, which is physically impossible. In fact, since the sound power associated with the pressure p 2,1m corresponds only to a part of the diffusion sound field incident on the opening in spatial X, that is, it corresponds at most to the diffusion sound power from the hemisphere on the spatial X side of the opening, therefore, should not exceed
[0298] Equation 10 for large S portal The reason why the value can give physically implausible results is that the use of Equation 9 implies that the total power transmitted through the portal is effectively radiated from a single point, which, if true, would indeed result in a much higher pressure near that single point than if the power were evenly distributed and radiated from the entire portal (which is actually the case in reality).
[0299] One way to solve this problem is to modify Equation 10 to:
[0300]
[0301] This modification prevents the level obtained at 1 m from the portal in space Y from exceeding the diffuse reverberation level in space X minus 3 dB (= 10 log 10 (0.5)), because this is physically required (as explained above). Note that depending on the actual rendering method used to render the reverberation from the portal, additional metrics at distances within 1 m from the portal may be required to ensure that the level there also does not exceed the diffuse reverberation level in space X minus 3 dB. In other words, it should be ensured that the level of the signal rendered from the portal anywhere in space Y does not exceed the diffuse reverberation level in space X minus 3 dB. This will also ensure a smooth transition of the reverberation level when the listener moves between spaces through the portal.
[0302] Although it is a rather simple measure for solving the problem that if S portal exceeds 4π m 2 (= 12.6 m 2 ) then p 2,1m might become too large, the solution of Equation 11 can actually provide very reasonable results in many use cases.
[0303] If S portal is substantially less than 4π m 2 , then the assumption of spherical radiation from the opening is reasonable, and using S portal / 8π as a scaling factor gives reasonable results.
[0304] When the size of the portal increases and approaches 4π m 2 , the sound pressure level at 1 m from the portal approaches the sound pressure level of the reverberation in space X minus 3 dB, and ultimately, for portals larger than 4π m 2 , it reaches and remains at that level. The latter seems quite reasonable because the reverberation experienced when standing 1 m from the center of an opening of size, for example, 4 x 3 m (= 12 m2) is very similar to that when standing at the opening itself.
[0305] According to the derived rms pressure p at 1 m from the portal 2,1m, directly derive the rms sound pressure p2 at any position in space Y at an arbitrary distance d from the portal from p2(d) = p 2,1m / d.
[0306] Alternatively, instead of the spherical radiation models of equations 8 and 9, other models for modeling the radiation from the portal can be used, which results in alternative equations to equation 8 for the relationship between the transmitted power P 1->2 and the resulting pressure p2 in space Y, and the resulting relationships between the pressures p1 and p2 of equations 10 and / or 11 (and thus ultimately being applied to the scaling of the audio signal rendered from the portal in order to obtain the correct rendered audio signal level in space Y).
[0307] For example, a reverberant portal source can be modeled as a spatially diffuse extended sound source, e.g., a spatially diffuse line source or a plane source, whose size is equal to the geometric size of the portal. Acoustic radiation models for such spatially diffuse extended sources (relating the source power to the resulting sound pressure at a given distance from the source) are available in the literature.
[0308] Thus, in some embodiments, alternative equations to equation 10 or 11 of the following form can be used:
[0309]
[0310] where C1 is a constant, or even more generally:
[0311]
[0312] where f(S portal ) is a function of the portal size S portal .
[0313] Alternative theoretical frameworks
[0314] will present alternative but largely equivalent theoretical views on the transmission of reverberation from space X to space Y.
[0315] Since the theoretical diffuse reverberant sound field at each point in space consists of uncorrelated plane waves of equal intensity arriving from all directions, the amount of reverberant energy propagating from space X through the portal to a specific point in space Y can be determined by geometric considerations.
[0316] If the average (rms) pressure of the individual reverberant plane waves arriving at an arbitrary point in space X from a single solid angle dΩ is denoted as p d , then the rms diffuse reverberant pressure p1 at an arbitrary point in space X is obtained by integrating p d over all solid angles dΩ:
[0317]
[0318] Figure 11B Shows the point L in space Y and indicates the opening angle when space X is "seen" through the portal from point L. From the perspective of point L, the portal represents a solid angle Ω portal (where 0 ≤ Ω portal ≤ 2π).
[0319] Since, by definition, the diffuse field in space X consists of uncorrelated plane waves of equal intensity from all directions, and since the pressure of a plane wave is constant along its path (i.e., it is independent of the distance traveled), therefore, each individual plane wave arriving at point L in space Y through the portal contributes the same uncorrelated pressure component p d , and thus, the resulting pressure p2 at point L in space Y can be determined by integrating p portal over the solid angle Ω d represented by the portal at point L (where 0 ≤ Ω portal ≤ 2π):
[0320]
[0321] Combining equations 14 and 15, it can be concluded that the pressure p2 at point L in space Y can be directly determined from the solid angle Ω portal and the diffuse pressure p1 in space X:[[]]
[0322]
[0323] If the portal is not completely acoustically transparent but transmits a fraction T of the power incident on it, then p2 is scaled accordingly.[[]]
[0324] This result can be compared with equations 10 - 12, and like these equations, it shows that the square of the rms pressure in space Y is proportional to the square of the rms pressure in space X, where the proportionality factor depends linearly on the size of the portal, where the size of the portal is represented by (equivalent) area (m 2 ) in equations 10 - 12 and by solid angle in equation 16.[[]]
[0325] According to equation 6, the diffuse reverberation pressure p1 in space X can be determined from the power of the sound source and the sound absorption in space X.[[]]
[0326] One thing to note when comparing the two proposed theoretical frameworks is that while the first framework models the portal as a "secondary" sound source radiating into space Y, the second framework directly considers the reverberant energy received from space X at a specific point in space Y. Since the solid angle represented by the portal depends on the position relative to the portal, the pressure p2 derived from Equation 16 also depends on the relative position of the specific location.
[0327] Specifically, for a position directly in front of the portal and a position on the side (or above / below) of the portal, the pressure derived from Equation 16 will be very different.
[0328] If the portal is small enough, the solid angle represented by a planar portal with a geometric area of S portal m 2 at a distance r from the observation point can be approximated by:
[0329]
[0330] where and are the position vector from the observation point to the portal and the normal vector of the portal, respectively.
[0331] At a distance r = 1m, this becomes:
[0332] Ω portal ≈S portal *cosθ (18)
[0333] where θ is the observation angle relative to the normal vector of the portal.
[0334] Combining Equation 18 with Equation 16, it can be seen that for a position directly in front of the portal, Ω portal ≈S portal , and Comparing it with Equation 10, it can be seen that the rms pressure at 1m in Equation 16 is larger than the rms pressure at 1m in Equation 10 by a factor of sqrt(2). On the other hand, at a position on the side completely close to the portal, Ω portal ≈0, and thus, also This can be interpreted as that while Equation 16 represents the pressure at a specific point in space Y, Equation 10 derived based on the assumption of spherical radiation from the portal represents the average value of Equation 16 at all angles at a distance of 1m (i.e., the average value of the solid angle Ω portal at all angles at a distance of 1m from the portal is equal to S portal / 2).
[0335] In the above, it has been assumed that space Y is a free field, i.e., a space that does not generate any diffuse reverberation by itself (e.g., a large outdoor space).
[0336] Rendering
[0337] The XR audio renderer can be configured to perform plausible reverberation rendering in the connected space of the XR environment using the model described above.
[0338] In one embodiment, in a second connected space Y, the reverberation rendering associated with the reverberation generated in space X is divided into two stages: (1) rendering of the reverberation from space X that reaches a listener in space Y directly through a portal between the two spaces, and (2) generating reverberation in space Y and rendering the generated reverberation in response to the reverberation from space X entering space Y through the portal.
[0339] In some embodiments, only the first rendering stage may be performed. In other embodiments, only the second rendering stage may be performed. In other embodiments, both rendering stages may be performed.
[0340] The first rendering stage
[0341] The first rendering stage is substantially independent of the acoustics of space Y. For example, a listener standing in a large open space would hear the reverberant sound coming from the open door (i.e., the portal) of a cathedral where music is being played.
[0342] In one embodiment, the level of the sound to be rendered in space Y due to the reverberant sound field in space X may include the following steps:
[0343] (1) Determine a space X reverberation intensity value representing the intensity of the reverberation in space X from one or more reverberation signals (also referred to as "space X reverberation signals") representing the reverberation in space X;
[0344] (2) Derive one or more space Y reverberation signals from the one or more space X reverberation signals for rendering in space Y (e.g., downmixing signals);
[0345] (3) Obtain (e.g., determine, derive, receive) the size of the portal through which the sound is transmitted between space X and space Y;
[0346] (4) Use the portal size to determine a scaling factor that simulates the transmission of the reverberant sound from space X to space Y; and
[0347] (5) Use the space Y reverberation signals, the space X reverberation intensity value, and the scaling factor to render one or more of the transmitted reverberation signals in space Y.
[0348] Regarding the first step, the space X reverberation intensity value can be determined in various ways.
[0349] In a simple scenario where reverberation in space X is rendered from a single (i.e., non-directional, monophonic) audio signal, the space X reverberation intensity value can be simply determined as the rms amplitude or rms power of that signal, or, in the case where reverberation is rendered based on an impulse response, as the total amount of energy contained in the impulse response.
[0350] If reverberation in space X is rendered using multiple audio signals, e.g., as multiple uncorrelated signals rendered from multiple directions around the user, the space X reverberation intensity value can be determined as the rms amplitude or power of the resulting combined signal. As an example, assume that reverberation in space X is rendered to a listener in space X as N uncorrelated reverberant signals from N corresponding directions, each reverberant signal having an rms amplitude of 1 / N (or an rms power of 1 / N 2 ), then the resulting combined reverberant signal has an rms power of 1 / N and an rms amplitude of 1 / sqrt(N).
[0351] In some cases, the space X reverberation intensity value need not be determined from the actual space X reverberant audio signal, but can be more efficiently derived from reverberation intensity metadata for space X. For example, the scene description metadata for an XR scene can include a reverberation level parameter or a reverberation energy ratio parameter for space X, which describes the desired reverberation level in space X, either absolute or relative to the direct sound level or the emitted source energy of the source in space X that generates the reverberation. In such a case, the (relative) reverberation level in space X is known a priori (and it is the job of the renderer to generate space X reverberant audio signals such that they produce the specified reverberation level in space X). For example, assume that space X has associated metadata that includes a value for the reverberation to direct energy ratio (RDR) in space X, which specifies the desired ratio of the energy of the reverberation to the energy of the direct sound at a distance of 1 m from an omnidirectional audio source located somewhere in space X. Now, if the omnidirectional audio source in space X has an associated audio signal with a linear rms signal amplitude of s and an associated linear source gain (“volume control”) of g, then the linear rms signal amplitude of the rendered direct sound at a distance of 1 m from the audio source is given by g*s, such that the rms power / energy of the direct sound signal is (proportional to) (g*s) 2 . It can then be concluded that the rms energy / power of the reverberation associated with the audio source should be equal to RDR*(g*s) 2 , such that the linear rms signal amplitude of the reverberation is sqrt(RDR)*g*s. Thus, the space X reverberation intensity value can be directly derived from the reverberation energy ratio (RDR) parameter provided for space X, as well as the source gain and the audio signal level of the audio source.
[0352] If the source is not omnidirectional but has any directional pattern associated with it, and that directional pattern causes the source to radiate a fraction X of the power of an omnidirectional source (for the same source signal), then this results in the power of the resulting reverberation also being fraction X of the omnidirectional source. Accordingly, the derived spatial X reverberation intensity values should be scaled accordingly, i.e., if a linear rms signal amplitude representation is used, the factor is sqrt(X), and if an rms signal energy / power representation is used, the factor is X.
[0353] When calculating the spatial X reverberation intensity values, other source rendering aspects that affect the rendered direct sound level or the gain of the rendered reverberant sound level can be considered in a similar manner to the source gain g, signal level s, and directional pattern discussed above.
[0354] Regarding step 2, the step of deriving one or more spatial Y reverberation signals for rendering in space Y can be accomplished in various ways. In one embodiment, the spatial Y reverberation signal can be a mono downmix from one or more spatial X reverberation audio signals.
[0355] In another embodiment, one or more spatial Y reverberation signals can be directly derived from the source signal and spatial X reverberation metadata parameters (e.g., reverberation time RT60 and reverberation energy ratio parameter), i.e., without the intermediate step of first generating the actual spatial X reverberation signal. This can be more efficient because the spatial X reverberation signal is not actually rendered to the listener (located in space Y), but is only generated as an intermediate step for generating one or more spatial Y reverberation signals.
[0356] Regarding step 3, the size of the portal can be obtained in various ways. In some embodiments, the size of the portal can be directly available in the scene description data, which can explicitly specify the location and / or size of the portal in space and which other space it is connected to. In other embodiments, the size can be derived from such scene description data, e.g., from geometric information. In other embodiments, the size of the portal can be heuristically detected, e.g., using some form of ray tracing algorithm.
[0357] In some embodiments, the size of the portal represents the area of the portal in m 2 units. In some embodiments, the area is the equivalent area of an acoustically perfectly transparent opening that has the same amount of "acoustic power leakage" as the portal.
[0358] In other embodiments, the size of the portal represents the solid angle corresponding to the portal starting from a specific location in space Y. Methods for deriving the solid angle are readily available in the literature.
[0359] The scaling factor derived in step 4 represents the desired relationship between the intensity of reverberation in space X (e.g., rms diffuse pressure, rms signal amplitude, or rms signal power) and the intensity of the rendered spatial Y reverberation signal in space Y (e.g., rms diffuse pressure, rms signal amplitude, or rms signal power).
[0360] In many embodiments, the basis for deriving the scaling factor can be given by any of equations 10 - 13 or 16, from which it can be derived as the factor that relates p1 to p2, or alternatively, the factor that relates p1 2 to p2 2 related factor.
[0361] Thus, for example, the scaling factor can be derived from equation 10 to be equal to (S portal / 8 (or its square root), while from equation 16, it can be derived as (Ω portal / 4π) (or its square root).
[0362] Finally, using the space X reverberation intensity value and the scaling factor, one or more derived spatial Y reverberation signals are rendered to a listener in space Y.
[0363] The scaling factor and the space X reverberation intensity value together determine the desired intensity of the rendered (one or more) spatial Y reverberation signals through the following relationship: Desired intensity of the rendered (one or more) spatial X reverberation signals = Scaling factor x Space X reverberation intensity value.
[0364] After determining the desired intensity for the rendered (one or more) spatial Y reverberation signals, an appropriate scaling gain for achieving the desired intensity of the rendered spatial Y reverberation signals can be determined for the (one or more) spatial Y reverberation signals.
[0365] In some embodiments, the scaling gain for the (one or more) spatial Y reverberation signals is simply equal to the scaling factor.
[0366] In some embodiments, in addition to the scaling factor, the scaling gain for the spatial Y reverberation signals can also take into account gain effects due to the specific way of deriving the (one or more) spatial Y reverberation signals from the space X reverberation signals, as well as gain effects due to different signal representations and rendering methods for the space X and spatial Y reverberation signals.
[0367] As already explained, the spatial X reverberation can be represented (and rendered) by a combination of multiple signals, and the spatial Y reverberation signal is derived from it using some signal transformation (e.g., downmixing) process, which may introduce some transformation gain effects, i.e., the difference in signal strength before and after the transformation. The scaling gain for the (one or more) spatial Y reverberation signals can compensate for this gain effect.
[0368] The scaling gain for the spatial Y reverberation signal can also compensate for the gain effect resulting from the specific way of combining the reverberation signals in the particular spatial X and spatial Y rendering methods used.
[0369] As a simple example, referring to the previous example where the spatial X reverberation is represented by N uncorrelated signals that are rendered from different directions around a listener in spatial X, and each signal has an rms amplitude of 1 / N. In this case, the spatial X reverberation intensity value is the rms amplitude of the sum of N uncorrelated signals, which is equal to 1 / sqrt(N). Now assume that the spatial Y reverberation is derived from the spatial X reverberation signal by selecting only one of the N signals (which has an rms amplitude of 1 / N). If this spatial Y reverberation signal is now rendered as a point source located at a certain position within the portal and according to the scaling factor (S portal / 8π) of Equation 10, then an additional gain of sqrt(N) must be applied to the spatial Y reverberation signal in order to obtain the correct balance between the intensities of the reverberation in spatial X and spatial Y.
[0370] Thus, the basic idea is to scale the spatial Y reverberation signal such that the intensity of the resulting rendered (one or more) spatial Y reverberation signals has the desired relationship with the intensity of the spatial X reverberation as represented by the scaling factor.
[0371] As previously discussed, different rendering methods can be used to render the derived spatial Y reverberation signal.
[0372] In one embodiment, the sound transmitted through the portal is rendered to the listener as a sound source located within the portal, i.e., the portal sound source. In one embodiment, the portal sound source is an extended sound source having a size corresponding to the geometric size of the portal. The extended sound source can be a uniformly extended sound source - radiating the same signal from each point within the range, a diffused extended sound source - radiating spatially diffused signals from different points within the range, or a non-uniformly extended sound source - radiating partially correlated signals from different points within the range.
[0373] In another embodiment, the portal sound source is a point source. In one embodiment, the point source is located at a fixed position, e.g., the central position within the portal. In another embodiment, the point source can be dynamically positioned within the portal depending on the listener's position. For example, the point source can be located at the point within the portal closest to the listener's position.
[0374] Second rendering stage
[0375] In the second rendering stage, reverberation is generated in space Y in response to the reverberation of space X entering space Y through the portal, based on the acoustic properties of space Y, such as, for example, the reverberation time, absorption, and / or reverberation level or reverberation energy ratio of space Y. Here, the rendering can be based on the amount of power transmitted from space X to space Y, for example, according to Equation 7. Then, the reverberation can be generated as the reverberation of a point source located in space Y, which has a source power equal to the transmitted power.
[0376] More specifically, the second rendering stage may include the following steps:
[0377] (1) Determine a space X reverberation intensity value representing the reverberation intensity in space X from one or more space X reverberation signals representing the reverberation in space X;
[0378] (2) Derive one or more reverberation input signals from the one or more space X reverberation signals for generating reverberation in space Y (e.g., downmixed signals);
[0379] (3) Obtain (e.g., determine, derive, receive) the size of the portal through which sound is transmitted between space X and space Y;
[0380] (4) Determine a scaling factor using the portal size, which simulates the transmission of reverberant sound from space X to space Y; and
[0381] (5) Render a reverberation signal in space Y using the (one or more) space Y reverberation signals, the space X reverberation intensity value, and the scaling factor.
[0382] Thus, the steps for the second rendering stage are basically the same as those for the first rendering stage, but with some differences in details, as will be explained below.
[0383] Steps 1 and 3 are the same as those for the first rendering stage. Thus, if both the first and second rendering stages are performed, steps 1 and 3 need to be performed only once.
[0384] In step 2, a signal for generating reverberation in space Y is derived. Generally, only a single reverberation input signal may be required. Thus, if the step 2 in the first rendering stage produces a single (e.g., mono downmixed) signal, this signal can also be used as the reverberation input signal for the second rendering stage. In principle, any signal having the general characteristics of the reverberation in space X can be used as the reverberation input signal in the second rendering stage, such as, for example, a single one of the multiple space X reverberation signals, or a single reverberation signal from which the multiple space X reverberation signals are generated.
[0385] In step 4, the scaling factor can be equal to (S portal / 16π), i.e., when using the model of Equation 10, it is 2 times smaller than in the first rendering stage. The reason is that in the second rendering stage, the reasoning that led to the addition of the factor 2 in Equation 8 does not apply here, and the "normal" relationship between the source power and pressure of the omnidirectional point source of Equation 9 should be used.
[0386] Finally, in step 5, a scaled version of the derived reverberant input signal is used as the source signal to generate the reverberation of space Y according to the reverberation characteristics corresponding to space Y (e.g., reverberation time, reverberation energy ratio). The scaling factor and the reverberation intensity value of space X are used to scale the gain of the reverberant input signal used to generate the reverberation of space Y. The scaling is such that when the scaled signal is to be rendered as a point source, it will have the desired level at a distance of 1 meter from the point source, i.e., p 2,1m 2 = scaling factor x p1 2 . Now, reverberation is generated based on the scaled reverberant input signal, which results in reverberation with the desired intensity.
[0387] Further rendering aspects
[0388] If the reverberation from space X is rendered into space Y through a portal as an extended sound source (also known as a "stereo" or "scaled" sound source) located at the portal and having the same geometric size as the portal, as in a typical implementation, then using the result of Equation 11 will be more realistic than the case where the sound from the portal is rendered as a point source located at a fixed point inside the portal. In such an implementation using an extended portal sound source, as used, for example, in the MPEG-I immersive audio standard, the distance to the extended sound source (i.e., the portal) is typically not measured relative to a reference point (e.g., the center point) inside the portal, but rather relative to its closest point. This means that if a user walks (virtually) along a path parallel to a large portal, the distance to the portal (i.e., the distance used to render the extended sound source to the user) remains constant, which means that the user's sound level experience also remains constant along that path, as would be expected. (Conversely, if the sound from the portal is to be rendered as a point source at a fixed position inside the portal, the distance and thus the rendered sound level will change as the user moves along the portal).
[0389] If the sound from the portal is rendered to a user in space Y as a point source at a dynamic position inside the portal that moves with the user rather than at a fixed position inside the portal, a similar effect can be achieved. In this case, the portal point source is dynamically positioned at the location inside the portal closest to the user.
[0390] In addition, in an implementation of using an extended sound source to render sound from a portal as described above, a distance attenuation function can typically be applied to the sound rendered from the extended portal sound source, which takes into account the geometric size of the extended sound source when viewed from the listening position, which can make the perceived effect more realistic. For example, if the listening position is initially in front of the portal and relatively close to the portal, the extended portal source can appear as a diffuse planar sound source, and if the distance from the portal is increased along a trajectory perpendicular to the portal, the sound level it renders may only decrease relatively slowly. As the distance is further increased, the rate of decrease of the level with increasing distance becomes faster, eventually approaching the rate of decrease of a point source.
[0391] On the other hand, if the listening position is initially on one side of the portal, the "perceived" geometric size of the extent of the stereo source (i.e., the geometric size when "viewed" from the listener position) is much smaller than when standing directly in front of it. If the distance is now increased while keeping the angle to the portal constant, the rendered sound level decreases more rapidly with increasing distance than in the case of the listening trajectory in front of the portal.
[0392] Cascaded-connected spaces
[0393] If more than two spaces are connected to each other, the propagation of reverberation from one space through the corresponding portals to all other spaces can be simulated by repeatedly applying Equation 7 and / or Equation 4, which simulate the amount of reverberation power transmitted from one space to the next through the portal. For example, if three spaces 1, 2, and 3 are connected via a first portal between space X and space Y and a second portal between space Y and space X, the amount of reverberation power transmitted to space X due to a sound source in space X can be determined by first applying Equation 7 to determine the power transmitted to space Y via the first portal based on the diffuse reverberation pressure in space X. Using this determined transmission power level, reverberation can be generated in space Y based on the acoustic parameters of space Y (e.g., RT60 and reverberation energy ratio), thereby providing a diffuse reverberation pressure in space Y. Then, applying Equation 7 to this space Y diffuse reverberation pressure, the amount of power transmitted to space X via the second portal can be calculated.
[0394] As an alternative to the step of rendering reverberation in space Y based on the determined amount of power from space X to space Y and thereby determining the diffuse reverberation pressure in space Y, the amount of power transmitted to space X can also be directly determined by applying Equation 4 to the result of the first step, i.e., using the amount of power transmitted from space X to space Y obtained in the first step as P1 in Equation 4. The only problem here is that applying Equation 4 requires the absorption amount A in space Y 1,tot (or A 1,0) This may not be directly available as metadata. In this case, the absorption can be estimated based on available parameters (specifically, the reverberation energy ratio, or a combination of the reverberation time RT60 and the volume of space Y). Patent application P103111 describes a method for deriving the absorption from these other parameters.
[0395] Figure 12 is a flowchart showing a process 1200 for rendering reverberation in a space Y connected to a space X via a portal according to some embodiments. The process 1200 may be executed by an audio renderer 151. The process 1200 may start from step s1202.
[0396] Step s1202 includes: determining a reverberation intensity value associated with the reverberation associated with space X.
[0397] Step s1204 includes: obtaining (e.g., deriving) information indicating the size of the portal.
[0398] Step s1206 includes: using the information indicating the size of the portal to determine a scaling factor.
[0399] Step s1208 includes: using the scaling factor and the reverberation intensity value to render a set of one or more space Y reverberation signals in space Y.
[0400] Overview of various further additional embodiments
[0401] A1. A method for rendering reverberation in a space Y connected to a space X via a portal, the method being executed by an audio renderer, the method including: determining a reverberation intensity value associated with the reverberation associated with space X; obtaining (e.g., deriving) information indicating the size of the portal; using the information indicating the size of the portal to determine a scaling factor; and using the scaling factor and the reverberation intensity value to render a set of one or more space Y reverberation signals in space Y.
[0402] A2. The method according to embodiment A1, wherein a set of one or more reverberation signals represents a reverberant sound field in space X (the set of one or more signals is referred to as "space X reverberation signals"), and the method further includes: before rendering the (one or more) space Y reverberation signals, deriving a set of one or more space Y reverberation signals from the space X reverberation signals.
[0403] A3. The method according to embodiment A2, wherein determining the reverberation intensity value includes: determining the reverberation intensity value based on a set of one or more space X reverberation signals.
[0404] A4. The method according to embodiment A2 or A3, wherein deriving a set of one or more space Y reverberation signals for rendering in space Y includes: downmixing a set of one or more space X reverberation signals.
[0405] A5. The method according to any one of embodiments A1 to A4, wherein the information indicating the size of the portal is a size value S portal , and determining the scaling factor includes: calculating C1 * S portal , where C1 is a predetermined value.
[0406] A6. The method according to embodiment A5, wherein C1 is approximately 1 / 8π.
[0407] A7. The method according to embodiment A5 or A6, wherein determining the scaling factor further includes: calculating the square root of C1 * S portal .
[0408] A8. The method according to embodiment A5 or A6, wherein determining the scaling factor further includes: determining whether C1 * S portal is less than C2, where C2 is a predetermined number (e.g., 0.5).
[0409] A9. The method according to any one of embodiments A1 to A4, wherein the information indicating the size of the portal is a solid angle value Ω portal .
[0410] A10. The method according to embodiment A9, wherein determining the scaling factor includes: calculating C1 * Ω portal , where C1 is a predetermined value (e.g., C1 = 1 / 4π).
[0411] A11. The method according to embodiment A10, wherein determining the scaling factor further includes: calculating the square root of C1 * Ω portal .
[0412] Conclusion
[0413] Although various embodiments are described herein, it should be understood that they are presented by way of example and not limitation. Accordingly, the breadth and scope of the present disclosure should not be limited by any of the exemplary embodiments described above. Moreover, unless otherwise specified herein or clearly contradicted by the context, any combination of the elements described above in all possible variations is covered by the present disclosure.
[0414] In addition, although the processes described above and illustrated in the figures are shown as a series of steps, this is done for illustrative purposes only. Accordingly, it should be anticipated that some steps may be added, some steps may be omitted, the order of the steps may be rearranged, and some steps may be performed in parallel.
[0415] References
[0416] [1]WD1 of ISO / IEC 23090-4 "MPEG-I Immersive Audio", Output document of the 9th meeting of MPEG WG 6 Audio Coding; and
[0417] [2]O. Das and J. S. Abel, "Grouped Feedback Delay Networks for Modeling of Coupled Spaces", J. Audio Eng. Soc., Vol. 69, No. 7 / 8, pp. 486-496, (July / August 2021). DOI: https: / / doi.org / 10.17743 / jaes.2021.0026.
Claims
1. A method (800) for generating an input signal for a first space (301) of the XR scene based on a first sound source (391) in a second space (302) of an extended reality XR scene (390), wherein, The first space (301) is directly or indirectly connected to the second space (302) via one or more portals including at least a first portal (300), and the method includes: Obtaining (s802) a direct propagation value, where the direct propagation value indicates an amount or portion of source power associated with the first sound source in the second space, and the amount or portion of source power is directly propagated from the first sound source through the one or more portals into the first space; and Using the direct propagation value and an audio signal associated with the first sound source to generate (s804) an input signal for the first space.
2. The method according to claim 1, wherein, The method further includes: using the input signal for the first space to generate a reverberation signal for the first space.
3. The method according to claim 2, wherein, The reverberation signal for the first space is generated using reverberation control information associated with the first space.
4. The method according to claim 3, wherein, The reverberation control information includes a reverberation level or a reverberation energy ratio parameter associated with the first space.
5. The method according to any one of claims 2 to 4, wherein The method further includes: rendering the reverberation signal.
6. The method according to any one of claims 1 to 5, wherein s1 = sqrt(X) * G * s2, where s1 is the input signal for the first space, X is the direct propagation value, G is a gain factor, and s2 is the audio signal associated with the first sound source.
7. The method according to any one of claims 1 to 6, wherein The method further includes: receiving metadata associated with the XR scene, The metadata includes the direct propagation value or data from which the direct propagation value can be derived, and Obtaining the direct propagation value includes: obtaining the direct propagation value from the metadata, or obtaining the direct propagation value using the data.
8. The method according to claim 7, wherein The metadata includes coefficients of a polynomial, and Obtaining the direct propagation value includes: obtaining the direct propagation value using the polynomial coefficients.
9. The method according to any one of claims 1 to 6, wherein Obtaining the direct propagation value includes: using information indicating the position of the first sound source in the second space and one or more of the following to derive the direct propagation value: i) Information indicating the position of the first portal, ii) Information indicating the size of the first portal, and iii) Information indicating the shape of the first portal.
10. The method according to claim 9, wherein One or more of the following are also used to derive the direct propagation value: information indicating the directivity pattern of the first sound source or information indicating the orientation of the first sound source, or Information indicating the directivity pattern of the first sound source and / or information indicating the orientation of the first sound source are also used to derive the direct propagation value.
11. The method according to any one of claims 1 to 6, wherein The method further includes: receiving metadata associated with the XR scene, The metadata includes portal position information indicating the position of the one or more portals, and Obtaining the direct propagation value includes: deriving the direct propagation value using the portal position information.
12. The method according to claim 11, wherein, For each portal, the metadata includes information indicating the geometry of the portal.
13. The method according to any one of claims 1 to 6, wherein, the first portal is associated with a transmission coefficient, and the direct propagation value is obtained using the transmission coefficient.
14. The method according to any one of claims 1 to 6, wherein, Obtaining the direct propagation value includes: obtaining the direct propagation value using an interpolation grid of pre-computed values and information indicating the position of the first sound source.
15. The method according to any one of claims 1 to 6, wherein Obtaining the direct propagation value includes: forming a grid representing the first portal; determining the area of the grid; and using the area of the grid to calculate the direct propagation value.
16. The method according to any one of claims 1 to 6, wherein Obtaining the direct propagation value includes: projecting the edges of the first portal onto a sphere surrounding the sound source; determining the surface area of the spherical segment covered by the projection of the first portal to obtain a value; and normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
17. The method according to any one of claims 1 to 6, wherein, Obtaining the direct propagation value includes: projecting the edges of the first portal onto a sphere surrounding the sound source; integrating a directional pattern associated with the sound source over the spherical segment covered by the projection of the first portal to obtain a value; and normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
18. The method according to claim 16 or 17, wherein, The sphere is a unit sphere, and the surface area of the unit sphere is 4π.
19. The method according to any one of claims 1 to 18, wherein The first portal has eight vertices.
20. The method according to any one of claims 1 to 19, wherein The geometry of the first portal is converted into a box by enclosing the portal geometry with a box.
21. The method according to any one of claims 1 to 20, wherein The method further includes: using the input signal for the first space to generate a first reverberation signal for the first space; obtaining a first scaling factor, wherein the first scaling factor indicates the amount or portion of reverberation energy associated with the first reverberation signal that is propagated into the second space through the one or more portals; using the first scaling factor and the first reverberation signal to generate a second input signal for the second space; and using the second input signal for the second space to generate a second reverberation signal for the second space.
22. A method for enabling reverberation to be rendered in a first space (301) of an extended reality XR scene (390), the first space being directly or indirectly connected to a second space (302) of the XR scene via one or more portals including at least a first portal (300), wherein, A sound source (391) is present in the second space, and the method includes: determining (s902) a direct propagation value, wherein the direct propagation value indicates the amount or portion of source power associated with the sound source that is directly propagated from the sound source in the second space through the one or more portals to the first space; and storing (s904) and / or transmitting the direct propagation value or data from which the direct propagation value can be derived.
23. The method according to claim 22, wherein, The method further includes: sending metadata to an audio renderer, wherein the metadata includes the direct propagation value or the data from which the direct propagation value can be derived.
24. The method according to claim 22 or 23, wherein Determining the direct propagation value includes: forming a grid representing the first portal; determining the area of the grid; and using the area of the grid to calculate the direct propagation value.
25. The method according to claim 22 or 23, wherein, Determining the direct propagation value includes: projecting the edges of the first portal onto a sphere surrounding the sound source; determining the surface area of the spherical segment covered by the projection of the first portal to obtain a value; and Normalize the obtained value by the surface area of the sphere, where the direct propagation value is the normalized value.
26. The method according to claim 22 or 23, wherein, Determining the direct propagation value includes: Projecting the edges of the first portal onto a sphere surrounding the sound source; Integrating the directional pattern associated with the sound source over the spherical segment covered by the projection of the first portal to obtain a value; and Normalize the obtained value by the surface area of the sphere, where the direct propagation value is the normalized value.
27. The method according to claim 25 or 26, wherein, The sphere is a unit sphere, and the surface area of the unit sphere is 4π.
28. A method (1300) for rendering first-order reverberation in a first space (301) of an extended reality XR scene (390), the method comprising: Determine (s1302) whether a first sound source (391) is located in a second space (302); Determine (s1304) whether there is one or more portals (300) connecting the first space and the second space; As a result of determining that the first sound source is located in the second space and there is at least a first portal (300) connecting the first space and the second space, determine (s1306) a first direct propagation value of the first sound source relative to the first portal; Use the direct propagation value and an audio signal associated with the first sound source to generate (s1308) an input signal for the first space; And Use (s1310) the input signal for the first space to generate a reverberation signal for the first space.
29. The method according to claim 28, wherein, The reverberation signal for the first space is generated using reverberation control information associated with the first space.
30. The method according to claim 29, wherein, The reverberation control information includes a reverberation level or a reverberation energy ratio parameter associated with the first space.
31. The method according to any one of claims 28 to 30, wherein The method further comprises: rendering the reverberation signal.
32. The method according to any one of claims 28 to 31, wherein, s1 = sqrt(X) * G * s2, where s1 is the input signal for the first space, X is the direct propagation value, G is a gain factor, and s2 is the audio signal associated with the first sound source.
33. The method according to any one of claims 28 to 32, wherein The method further comprises: receiving metadata associated with the XR scene, The metadata includes the direct propagation value or data from which the direct propagation value can be derived, and Obtaining the direct propagation value includes: obtaining the direct propagation value from the metadata, or obtaining the direct propagation value using the data.
34. The method according to claim 33, wherein The metadata includes coefficients of a polynomial, and Obtaining the direct propagation value includes: obtaining the direct propagation value using the polynomial coefficients.
35. The method according to any one of claims 28 to 32, wherein Obtaining the direct propagation value includes: using information indicating the position of the first sound source in the second space and one or more of the following to derive the direct propagation value: i) Information indicating the position of the first portal, ii) Information indicating the size of the first portal, and iii) Information indicating the shape of the first portal.
36. The method according to claim 35, wherein One or more of the following are also used to derive the direct propagation value: information indicating the directivity pattern of the first sound source or information indicating the orientation of the first sound source, or Information indicating the directivity pattern of the first sound source and / or information indicating the orientation of the first sound source are also used to derive the direct propagation value.
37. The method according to any one of claims 28 to 32, wherein The method further includes: receiving metadata associated with the XR scene, The metadata includes portal location information indicating the location of the one or more portals, and Obtaining the direct propagation value includes: using the portal location information to derive the direct propagation value.
38. The method according to claim 37, wherein For each portal, the metadata includes information indicating the geometry of the portal.
39. The method according to any one of claims 28 to 32, wherein The first portal is associated with a transmission coefficient, and The direct propagation value is obtained using the transmission coefficient.
40. The method according to any one of claims 28 to 32, wherein, Obtaining the direct propagation value includes: using an interpolation grid of pre-computed values and information indicating the location of the first sound source to obtain the direct propagation value.
41. The method according to any one of claims 28 to 32, wherein, Obtaining the direct propagation value includes: Forming a grid representing the first portal; Determining the area of the grid; and Using the area of the grid to calculate the direct propagation value.
42. The method according to any one of claims 28 to 32, wherein Obtaining the direct propagation value includes: Projecting the edges of the first portal onto a sphere surrounding the sound source; Determining the surface area of the spherical segment covered by the projection of the first portal to obtain a value; and Normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
43. The method according to any one of claims 28 to 32, wherein, Obtaining the direct propagation value includes: Projecting the edges of the first portal onto a sphere surrounding the sound source; Integrating the directivity pattern associated with the sound source over the spherical segment covered by the projection of the first portal to obtain a value; and Normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
44. The method according to claim 42 or 43, wherein, The sphere is a unit sphere, and the surface area of the unit sphere is 4π.
45. The method according to any one of claims 28 to 44, wherein The first portal has eight vertices.
46. The method according to any one of claims 28 to 45, wherein The geometry of the first portal is transformed into a box by enclosing the portal geometry with a box.
47. The method according to any one of claims 28 to 46, wherein, The method further includes: Using the input signal for the first space to generate a first reverberation signal for the first space; Obtaining a first scaling factor, wherein the first scaling factor indicates the amount or portion of the reverberation energy associated with the first reverberation signal that is propagated into the second space through the one or more portals; Using the first scaling factor and the first reverberation signal to generate a second input signal for the second space; and Using the second input signal for the second space to generate a second reverberation signal for the second space.
48. A computer program (1043) comprising instructions (1044) which, when executed by a processing circuit (1002) of a device (1000), cause the device to perform the method according to any one of the above claims.
49. A carrier, comprising the computer program according to claim 50, wherein, The carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium (1042).
50. An apparatus (1000) for generating an input signal for a first space (301) of the XR scene based on a first sound source (391) in a second space (302) of an extended reality XR scene (390), wherein, The first space (301) is directly or indirectly connected to the second space (302) via one or more portals including at least a first portal (300), and the apparatus is configured to: Obtain (s802) a direct propagation value, where the direct propagation value indicates an amount or portion of source power associated with the first sound source in the second space, and the amount or portion of source power is directly propagated from the first sound source through the one or more portals into the first space; and Use the direct propagation value and an audio signal associated with the first sound source to generate (s804) an input signal for the first space.
51. The apparatus according to claim 50, wherein, The apparatus is further configured to perform the method according to any one of claims 2 to 21.
52. An apparatus (1000) for enabling reverberation to be rendered in a first space (301) of an extended reality XR scene (390), the first space being directly or indirectly connected to a second space (302) of the XR scene via one or more portals including at least a first portal (300), wherein, A sound source (391) exists in the second space, and the apparatus is configured to: Determine (s902) a direct propagation value, where the direct propagation value indicates an amount or portion of source power associated with the sound source, and the amount or portion of source power is directly propagated from the sound source in the second space through the one or more portals into the first space; and Store (s904) and / or transmit the direct propagation value or data from which the direct propagation value can be derived.
53. The apparatus according to claim 52, wherein, The apparatus is further configured to perform the method according to any one of claims 22 to 27.
54. An apparatus (1000) for rendering first-order reverberation in a first space (301) of an extended reality XR scene (390), the apparatus being configured to: Determine (s1302) whether a first sound source (391) is located in a second space (302); Determine (s1304) whether there are one or more portals (300) connecting the first space and the second space; As a result of determining that the first sound source is located in the second space and there is at least a first portal (300) connecting the first space and the second space, determine (s1306) a first direct propagation value of the first sound source relative to the first portal; Use the direct propagation value and an audio signal associated with the first sound source to generate (s1308) an input signal for the first space; And Use (s1310) the input signal for the first space to generate a reverberation signal for the first space.
55. The apparatus according to claim 54, wherein, The apparatus is further configured to perform the method according to any one of claims 29 to 47.