Rendering reverberation in connected spaces
The method addresses the challenge of rendering reverberation in connected XR spaces by using a direct propagation value to efficiently generate reverberation in the first space from sound sources in the second space, ensuring realistic and efficient audio rendering.
Patent Information
- Application Number
- JP2025533566
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-12-15
- Publication Date
- 2025-12-16
AI Technical Summary
Existing XR systems lack practical rendering models and architectures for realistic rendering of reverberation in connected spaces due to sound sources located outside the active space, leading to unnatural audio experiences and computational inefficiencies, particularly in XR scenes with multiple connected spaces.
A method and apparatus for generating an input signal in a first space based on a sound source in a second space, using a direct propagation value to determine the amount of source power propagating through a portal, and efficiently rendering reverberation in the first space, including first-order and second-order reverberation processes.
Enables realistic and efficient rendering of reverberation in XR scenes with connected spaces, providing a computationally efficient solution for handling reverberation from external sound sources and maintaining acoustic integrity.
Smart Images

Figure 2025540822000001_ABST
Abstract
Description
[Technical Field]
[0001] Embodiments are disclosed that relate to rendering reverberation in a first space connected to a second space via a portal. [Background technology]
[0002] Augmented reality (XR) systems, such as virtual reality (VR) systems, augmented reality (AR) systems, mixed reality (MR) systems, etc., generally include an audio renderer for rendering audio to a user of the XR system. The audio renderer typically includes a reverberation processor for generating late and / or diffuse reverberation that is rendered to a user of the XR system to provide the auditory sensation of being in the XR scene being rendered. The generated reverberation should provide the user with the auditory sensation of being in an acoustic environment (AE), also known as a "space," corresponding to the XR scene, e.g., a living room, a gym, an outdoor environment, etc.
[0003] Reverberation is one of the most significant acoustic properties of a room. Sound created in a room will bounce repeatedly off reflective surfaces such as the floor, walls, ceiling, windows, or tables, gradually losing energy. When these reflections mix with each other, the phenomenon known as "reverberation" results. Reverberation is thus the collection of many reflections of sound.
[0004] Two of the most fundamental characteristics of reverberation in any space, real or virtual, are 1) reverberation time, and 2) reverberation level, i.e., how strong or loud the reverberation is relative to, for example, the power of the sound source or the direct sound level in the space.
[0005] Reverberation time is a measure of the time required for reflected sound to "die away" in an enclosed space after the sound source ("sound source") has stopped. Reverberation time is important in defining how a room will respond to acoustic sound. Reverberation time depends on the amount of sound absorption in a space, which is lower in spaces with many absorbent surfaces such as curtains, padded chairs, and even people, and higher in spaces that contain mostly hard, reflective surfaces.
[0006] Conventionally, reverberation time is defined as the amount of time it takes for the sound pressure level to decrease by 60 dB after the sound source is abruptly switched off. The abbreviation for this amount of time is "RT60" (or sometimes T60).
[0007] Typically, in a reverberation processor used in an audio renderer, these two (and other) characteristics of the generated reverberation can be controlled individually and independently. For example, it is typically possible to set the reverberation processor to generate reverberation with a certain desired reverberation time and a certain desired reverberation level.
[0008] In XR systems, the characteristics of the generated reverberation are typically controlled by control information, e.g., special metadata included in an XR scene description, e.g., as specified by a scene creator, that describes many aspects of the XR scene, including its acoustic characteristics. The audio renderer receives this control information, e.g., from a bitstream or file, and uses it to configure the reverberation processor to generate reverberation with the desired characteristics. The exact manner in which the reverberation processor obtains the desired reverberation time and reverberation level in the generated reverberation may vary depending on the type of reverberation algorithm the reverberation processor uses to generate the reverberation.
[0009] Reverberation and connected spaces
[0010] As noted above, one of the key aspects of immersive rendering of audio is the realistic rendering of reverberation associated with the virtual space of an XR scene. A particular challenge is the realistic rendering of reverberation in spaces (a.k.a., "acoustic environments") that are connected (or "coupled") to another space through a "portal," e.g., an open door, an open window, a partially transparent window, a partially transparent wall, etc., i.e., something through which sound from one space can propagate into the connected space and vice versa.
[0011] One aspect of this is the realistic modeling and rendering of reverberation generated in a first space in response to direct sound from a sound source located in a second space propagating into the first space directly through a portal between the two spaces, i.e., without first reflecting and / or reverberating in the second, connected space. Another aspect is the realistic modeling and rendering of reverberation from a connected second space propagating into the first space through a portal.
[0012] One solution for rendering reverberation is described in ISO / IEC23090-4, Working Draft (WD1) of MPEG-I Immersive Audio, an output document of the 9th meeting of MPEG WG 6 Audio Coding (hereafter "WD1 of ISO / IEC23090-4"). Summary of the Invention
[0013] Currently, several challenges exist. For example, currently, there are no practical rendering models and signal processing architectures for realistic rendering of reverberation generated in the space of an XR scene due to sound sources located in connected spaces. The solution described in WD1 of ISO / IEC 23090-4 lacks rendering of reverberation in a first space (e.g., active space) in response to sources located outside the first space (i.e., in a second space).
[0014] If sound sources outside the active space, i.e., the space in which the listener is located, are not taken into account when generating reverberation in the active space, no reverberation will be generated in response to these external sound sources, which will be unnatural to the listener(s) located in the active space. Similarly, if a sound source resides in an active space connected to another highly reverberant space, the source sound would naturally be expected to generate some reverberation in the connected space, and some of this reverberation would be expected to propagate back into the active space.
[0015] Current solutions in WD1 of ISO / IEC 23090-4 render the reverberation for each space using a feedback delay network (FDN) reverberator. Although there are methods for combining several such FDN reverberators into a grouped feedback delay network (see, for example, reference [2]), such solutions easily become too computationally intensive for real-time audio rendering.
[0016] Furthermore, existing XR audio rendering architectures are not optimized for handling reverberation due to sources outside the active space, especially for XR scenes consisting of many connected spaces, all of which may contain any number of sound sources. Therefore, an optimized rendering architecture is needed for efficient handling of such cases.
[0017] Thus, in one aspect, there is provided a method for generating an input signal for a first space of an XR scene based on a first sound source in a second space of the XR scene, where the first space is directly or indirectly connected to the second space through one or more portals, including at least a first portal. The method includes obtaining a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the first sound source in the second space that is directly propagated from the first sound source through the one or more portals into the first space. The method further includes generating the input signal for the first space using the direct propagation value and an audio signal associated with the first sound source.
[0018] In another aspect, a method is provided for enabling rendering of reverberation in a first space of an XR scene connected directly or indirectly to a second space of the XR scene through one or more portals, including at least a first portal, where a sound source is present in the second space. The method includes determining a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the sound source in the second space that is directly propagated from the sound source in the second space through one or more portals to the first space. The method further includes storing and / or transmitting the direct propagation value or data from which the direct propagation value can be derived. The method may be performed by an audio encoder.
[0019] In another aspect, a method for rendering first-order reverberation in a first space of an XR scene is provided. The method includes determining whether a first sound source is located in a second space. The method also includes determining whether there are one or more portals connecting the first space with the second space. The method also includes, as a result of determining that the first sound source is located in the second space and that there is at least a first portal connecting the first space with the second space, determining a first direct propagation value for the first sound source with respect to the first portal. The method also includes generating an input signal for the first space using the direct propagation value and an audio signal associated with the first sound source. The method further includes using the input signal for the first space to generate a reverberation signal for the first space.
[0020] In another aspect, an apparatus is provided for generating an input signal for a first space of an XR scene based on a first sound source in a second space of the XR scene, the first space being directly or indirectly connected to the second space through one or more portals, including at least a first portal. The apparatus is configured to obtain a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the first sound source in the second space that is directly propagated from the first sound source through the one or more portals into the first space. The apparatus is also configured to generate the input signal for the first space using the direct propagation value and an audio signal associated with the first sound source.
[0021] In another aspect, an apparatus is provided for enabling rendering of reverberation in a first space of an XR scene connected directly or indirectly to a second space of the XR scene through one or more portals, including at least a first portal, where a sound source is present in the second space. The apparatus is configured to determine a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the sound source in the second space that is directly propagated from the sound source in the second space through one or more portals to the first space. The apparatus is also configured to store and / or transmit the direct propagation value or data from which the direct propagation value can be derived.
[0022] In another aspect, an apparatus for rendering first-order reverberation in a first space of an XR scene is provided. The apparatus is configured to determine whether a first sound source is located in a second space. The apparatus is also configured to determine whether there are one or more portals connecting the first space with the second space. The apparatus is also configured to, as a result of determining that the first sound source is located in the second space and that there is at least a first portal connecting the first space with the second space, determine a first direct propagation value for the first sound source with respect to the first portal. The apparatus is also configured to generate an input signal for the first space using the direct propagation value and an audio signal associated with the first sound source. The apparatus is also configured to use the input signal for the first space to generate a reverberation signal for the first space.
[0023] In another aspect, a computer program is provided comprising instructions that, when executed by a processing circuit of the device, cause the device to perform the methods disclosed herein. In one embodiment, a carrier is provided that contains the computer program, the carrier being one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0024] An advantage of the embodiments disclosed herein is that they enable realistic and efficient rendering of reverberation in XR scenes with connected spaces.
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various embodiments. [Brief explanation of the drawings]
[0026] [Figure 1A] FIG. 1 illustrates a system, according to some embodiments. [Figure 1B] FIG. 1 illustrates a system, according to some embodiments. [Figure 2] FIG. 1 illustrates a system, according to some embodiments. [Figure 3] FIG. 1 shows a portal from a first space to a second space. [Figure 4] FIG. 10 shows a vector going from a point to the portal opening vertex. [Figure 5] FIG. [Figure 6] FIG. 1 illustrates the concept of an interpolated direct propagation value grid for possible object source locations. [Figure 7] FIG. 1 illustrates an efficient architecture for an audio renderer. [Figure 8] 1 is a flowchart illustrating a process according to one embodiment. [Figure 9] 1 is a flowchart illustrating a process according to one embodiment. [Figure 10] FIG. 1 is a block diagram of an apparatus, according to some embodiments. [Figure 11A] A diagram showing two spaces connected through a portal. [Figure 11B] A diagram showing two spaces connected through a portal. [Figure 12] 1 is a flowchart illustrating a process according to one embodiment. [Figure 13]1 is a flowchart illustrating a process according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0027] 1A shows an XR system 100 to which embodiments disclosed herein may be applied. The XR system 100 includes speakers 104 and 105 (which may be speakers of headphones worn by a user) and an XR device 110, which may include a display for displaying images to a user and, in some embodiments, is configured to be worn by a listener. In the illustrated XR system 100, the XR device 110 has a display and is designed to be worn on the user's head, and is commonly referred to as a head-mounted display (HMD).
[0028] 1B , XR device 110 may include orientation detection unit 101, position detection unit 102, and processing unit 103 coupled directly or indirectly to audio render 151 for generating output audio signals, e.g., left audio signal 181 for a left speaker and right audio signal 182 for a right speaker, using input audio signal 161, e.g., encoded audio data, and metadata 162, shown in this example as provided by encoder 169. For example, encoder 169 may generate a bitstream including encoded audio data 161 and metadata 162.
[0029] The orientation sensing unit 101 is configured to detect changes in the listener's orientation and provide information about the detected changes to the processing unit 103. In some embodiments, the processing unit 103 determines an absolute orientation (with respect to some coordinate system) given the detected change in orientation detected by the orientation sensing unit 101. There may also be different systems for determining orientation and position, for example, systems that use lighthouse trackers (LIDAR). In one embodiment, the orientation sensing unit 101 may determine an absolute orientation (with respect to some coordinate system) given the detected change in orientation. In this case, the processing unit 103 may simply multiplex the absolute orientation data from the orientation sensing unit 101 and the position data from the position sensing unit 102. In some embodiments, the orientation sensing unit 101 may comprise one or more accelerometers and / or one or more gyroscopes.
[0030] Audio renderer 151 generates an audio output signal based on input audio signal 161, metadata 162 about the XR scene the listener is experiencing, and information 163 about the listener's location and orientation. Metadata 162 about the XR scene may include metadata about each object and audio element included in the XR scene, as well as metadata about the XR space (the "acoustic environment") in which the listener is virtually located. Metadata about objects may include information about the object's dimensions and occlusion factors for the object. For example, the metadata may specify a set of occlusion factors, each applicable for a different frequency or frequency range. Metadata 162 may also include control parameters, such as a reverberation time value, a reverberation level value, and / or absorption parameter(s).
[0031] The audio renderer 151 may be a component of the XR device 110, or the audio renderer 151 may be remote from the XR device 110. For example, the audio renderer 151, or components of the audio renderer 151, may be implemented in the cloud.
[0032] 2 shows one exemplary implementation of an audio renderer 151 for generating sound for an XR scene. The audio renderer 151 includes a controller 201 and an audio signal generator 202 for generating output audio signal(s) (e.g., audio signals of multi-channel audio elements) based on control information 210 from the controller 201 and input audio 161. In this embodiment, the controller 201 includes a reverberation processor 204 for determining scaling factors (also known as direct propagation values) described below.
[0033] In some embodiments, the controller 201 may be configured to receive one or more parameters and trigger the audio signal generator 202 to perform a modification to the audio signal 161, such as increasing or decreasing a volume level, based on the received parameters. The received parameters include information 163 regarding the listener's position and / or orientation, such as the direction and distance to an audio element, and metadata 162 regarding the XR scene. As described above, the metadata 162 may include metadata regarding the XR space in which the user is virtually located. For example, the metadata 162 may include information regarding the dimensions of the space, information regarding objects in the space, and information regarding the acoustic properties of the space, as well as metadata regarding the audio element and metadata regarding objects that occlude the audio element. In some embodiments, the controller 201 itself generates at least a portion of the metadata 162. For example, the controller 201 may receive metadata regarding the XR scene and derive additional metadata, such as control parameters, based on the received metadata. For example, using the metadata 162 and position / orientation information 163, the controller 201 may calculate one or more gain factors (g) (e.g., the scaling factors described above) for audio elements in an XR scene.
[0034] With respect to generating the reverberation signal used by the audio signal generator 202 to generate the final output signal, the controller 201 provides reverberation parameters, such as the reverberation time and reverberation level for a space in an XR scene, and the scaling factors described above to the audio signal generator 202 so that the audio signal generator 202 is operable to generate the reverberation signal. The reverberation time for the generated reverberation is most commonly provided to the reverberation processor 204 as an RT60 value, generally for individual frequency bands, although other reverberation time measures may also exist and be used. In a typical embodiment, the metadata 162 includes all of the necessary reverberation parameters (e.g., RT60 and reverberation level values). However, in embodiments where the metadata does not include all the necessary reverberation parameters, the controller 201 may be configured to generate the missing parameters.
[0035] Reverberation levels can be expressed in various formats. Typically, reverberation levels will be expressed as relative levels. For example, reverberation levels can be expressed as the energy ratio (DRR) between the direct sound component and the reverberant sound component at a certain distance from the sound source being rendered in an XR environment, or vice versa, i.e., the RDR energy ratio. Alternatively, reverberation levels can be expressed as the energy ratio between the reverberant sound and the total emitted energy or power of the source. In other cases, reverberation levels can be expressed directly as the level / gain for a reverberation processor.
[0036] In the present context, the term "reverberant" may generally refer only to sound field components corresponding to the diffuse part of the acoustic room impulse response of an acoustic environment, but in some embodiments the term may also include sound field components corresponding to earlier parts of the room impulse response, including, for example, some late non-diffuse reflections and even all reflected sounds.
[0037] Other metadata describing reverberation-related characteristics of an acoustic environment that may be included in metadata 162 include parameters describing the acoustic properties of surface materials in the environment (e.g., describing the absorption, reflection, transmission, and / or diffusion properties of the materials), or parameters describing specific points in time of a room impulse response associated with the acoustic environment, such as the time after source emission after which the room impulse response becomes diffuse (this is sometimes called a “pre-delay”).
[0038] All the reverberation related properties described above are generally frequency dependent, and therefore their related metadata parameters are also generally provided and processed separately for several frequency bands.
[0039] Embodiment
[0040] In one embodiment, a computationally efficient method for rendering reverberation in response to sources located in connected spaces is provided. The embodiment may also optionally include rendering reverberation propagating from connected spaces and / or generating and rendering further ("second-order") reverberation in response to reverberation propagating from connected spaces in a virtual sound system.
[0041] In real-life situations, the reverberation present in a given space may generally be generated in response to sound sources present within that same space, but may also be generated (in part) in response to sound sources located outside the space.
[0042] Consider the situation in Fig. 3, where a first space (Space 1) is connected to a second space (Space 2) via a portal 300. Initially, the assumption is that the portal is completely acoustically transparent, i.e., that it is completely "open." For example, the assumption could be that the portal is an open door, an open window, or other open connection between the two spaces. A sound source S is located somewhere in Space 2.
[0043] The energy radiated by Source S will generate reverberation in Space 2 ("Space 2 reverberation"), and some of that Space 2 reverberation will propagate through the portal into Space 1, which may generate further reverberation ("secondary reverberation").
[0044] However, some of the energy radiated by Source S may propagate directly through the portal, i.e., as direct sound from Source S, and not as Space 2 reverberation. This direct sound entering Space 1 through the portal generates reverberation in Space 1 that is a direct response to the sound emitted by Source S.
[0045] The problems addressed in this disclosure include: i) determining the amount or portion of direct sound energy / power radiated by a sound source S that propagates directly into space 1 through a portal; ii) generating (and optionally rendering) reverberation in space 1 that is consistent with the determined amount or portion of direct sound energy entering space 1 through the portal; and iii) defining a suitable rendering architecture for efficient handling of the above processes, even in use cases involving many connected spaces.
[0046] Further challenges include i) realistic rendering in Space 1 of Space 2 reverberations propagating into Space 1 through a portal, and conversely, realistic rendering in Space 2 of Space 1 reverberations propagating through a portal back into Space 2 (which is essentially the same problem), and ii) realistic rendering in Space 1 of "second order" reverberations generated in Space 1 in response to Space 2 reverberations propagating through a portal.
[0047] The situations and challenges described above are, in principle, independent of where a user of an XR system (a.k.a., "user" or "listener") might virtually be located within the generated XR scene. However, depending on the space in which the listener is located, the relevance and / or order in which different challenges occur may differ, and thus a renderer may be configured to perform any of the above-described processes in any combination or order, depending on the space in which the listener and source are located.
[0048] To illustrate this, the different processes are separated into "first-order" and "second-order" processes. In the situation of Figure 3, the "first-order" processes taking place include (1) the generation of first-order Space 2 reverberation in Space 2 and (2) the generation of first-order reverberation in Space 1 due to the direct sound of Source S propagating through the portal, while the "second-order" processes in this case include (3) the generation of second-order reverberation in Space 1 in response to the first-order Space 2 reverberation entering Space 1 through the portal and (4) the generation of second-order Space 2 reverberation due to the first-order Space 1 reverberation entering Space 2 through the portal.
[0049] If the listener is located in Space 1, i.e., if sound source S is located in a space other than the space in which the listener is located, it is processes (2) and (3) that result in the sound rendered to the listener, and process (1) is required in order to be able to perform process (3). Process (4) may not be considered relevant to perform in this case, because there is no listener in Space 2 who will hear this second-order reverberation in Space 2, and the result of process (4) is not required to produce the sound rendered to the listener in Space 1.
[0050] In principle, second-order processes could be followed by third-order processes, for example, propagating second-order reverberation in Space 2 back into Space 1, but the perceptual relevance of these higher-order processes may be so small that they may be considered irrelevant to perform and / or render.
[0051] On the other hand, if the listener is located in space 2 (i.e., the listener is in the same space as the sound source), it is processes 1 and 4 that result in the sound being rendered to the listener, and process 2 is needed to be able to perform process 4. In this case, it is process 3 that may be considered irrelevant to perform and / or render.
[0052] Determining the amount or fraction of direct sound source power propagating through a portal into a connected space
[0053] The amount or portion of power radiated by a sound source that propagates through a portal via a direct path depends on several factors, including: i) the location and size / shape of the portal, ii) the location of the sound source, iii) the directivity pattern of the sound source, and iv) the orientation of the sound source.
[0054] One method for determining the amount or portion of the radiated source power that propagates directly through the portal is to use a "line of sight" approach, which in one embodiment includes the following steps:
[0055] (1) Projecting the portal geometry (edge) onto an imaginary sphere around the sound source;
[0056] (2) integrating the directivity pattern of the source over the spherical segment covered by the projection of the portal, taking into account the orientation of the sound source; and
[0057] (3) Normalizing the obtained values by the surface area of the sphere to generate a direct propagation value for source S in space 2 relative to the portal to space 1. The direct propagation value indicates the amount or portion of power radiated by the sound source that propagates directly through the portal.
[0058] The latter step normalizes the amount of power propagating directly through the portal to the total amount of power radiated by an omnidirectional source driven by the same input signal. In many embodiments, the sphere is a unit sphere, and the surface area of the unit sphere is 4π.
[0059] If the source's directivity pattern is normalized, i.e., if the source's directivity pattern has a value of 1 in the direction of the source's highest power, this process will yield a value between 0 and 1. If the portal is a flat surface, the maximum value is 0.5, since an omnidirectional source will radiate up to half of its power directly through a flat portal. The assumption here is that the directivity pattern is expressed in terms of the intensity / power radiated in each direction, i.e., that the directivity pattern is proportional to the square of the rms signal amplitude / pressure of the source's direct sound, measured in each direction at equal distances from the source.
[0060] The directivity pattern of a sound source indicates the amount of sound power that the sound source radiates in each direction. A simple omnidirectional source radiates sound equally in all directions, so the directivity pattern has the same value (1 if the directivity pattern is normalized) in all directions, but in general sound sources will radiate different amounts of power in different directions.
[0061] In an XR system, the directional pattern of a sound source will generally be available directly from the metadata 162 corresponding to the sound source. The directional pattern may be provided either as (generally normalized) dB values for each direction, corresponding to (normalized) sound pressure levels (SPLs) measured in each direction around the sound source at equal distances, or as (generally normalized) linear gain values for each direction.
[0062] Alternatively, the directional pattern may be provided in some other suitable format, for example, a spherical harmonics representation from which the directional pattern may be derived for any desired set of directions.
[0063] In other cases, the directivity pattern may be provided with "directivity factors" for individual directions, where the directivity factor for a particular direction quantifies the ratio between the sound intensity radiated in that direction and the intensity averaged over all directions.
[0064] It should be clear that all of these expressions for the directional pattern are essentially equivalent and bear simple relationships to one another, and any of them can be used in the context of the present disclosure by converting it to the expression assumed in the method described above.
[0065] Also, the initial position and orientation of the sound source will generally be available as metadata about the sound source. In dynamic and / or interactive XR scenes, the source position and orientation may vary, but they will generally be readily available to the audio renderer.
[0066] Position / size / shape information about the portal may also be directly available or easily derivable from the scene description metadata, or may be derived in some other manner, as described in more detail below.
[0067] As a simple example, consider the situation where an omnidirectional source is placed exactly in the center of the (flat) open portal between Space 1 and Space 2 in Figure 3. In this case, the portal occupies a solid angle (i.e., a hemisphere) of 2π relative to the sound source. Therefore, since the source is omnidirectional, we can see that half (2π / 4π) of the total power radiated by the omnidirectional source propagates directly through the portal, and therefore the direct propagation value we obtain is 0.5.
[0068] Alternatively, if the sound source has a non-omnidirectional directional pattern and the source orientation is such that the "loudest side" of the source is pointed away from the Portal, the direct propagation value will be less than 0.5. In the extreme case where the sound source has a directional pattern such that all of the source's energy is radiated into one hemisphere, and the source is oriented such that the source's hemispherical directional pattern is pointing exactly away from the Portal, the direct propagation value will be 0.
[0069] In the more general situation shown in Figure 3, determining the direct propagation value for source S for a portal between Space 1 and Space 2 requires first determining the projection of the (edges of) the portal geometry onto an imaginary unit sphere around that source, taking into account the position of the sound source S.
[0070] The projection of the portal into the solid angle Ω on the unit sphere around the source S portal and assume that a source S has a normalized directivity pattern D(ω) (expressed in power) with a spatial angle ω and an orientation R relative to some reference orientation. Then the direct propagation value is given by the solid angle Ω of the rotated directivity pattern. portal and normalizing by the surface area of the unit sphere, 4π. TIFF2025540822000002.tif8170D rot,R is the directional pattern D preferably rotated according to the source orientation R.
[0071] Note that if the source is omnidirectional, the above procedure simply reduces to determining the normalized surface area of the projection onto the unit sphere.
[0072] To determine the direct propagation value based on the segment covered by the portal's projection on the source directivity pattern, processing based on the scene geometry needs to be performed to determine the projection onto the unit sphere and the size (area) of that projection.
[0073] In an exemplary embodiment, an analysis based on the scene geometry and object source positions is performed on the encoder device. The analysis may also be performed by a renderer. In this exemplary embodiment, the portal opening is generally a box with eight vertices, representing, for example, a door opening between two spaces. If the portal opening is not a box, the portal geometry may be converted to a box, for example, by enclosing the portal geometry in a box.
[0074] There are various ways to obtain portal openings from scene geometry, which may be represented as meshes or voxels. In one example, a content creator provides explicit metadata indicating the location of portals, for example, as a set of vertices. Such portal metadata may then be conveyed as metadata. In another example, processing on an encoder / renderer device may be used to automatically analyze portals between acoustic environments. In an exemplary method, each space is enclosed by a geometry such as a box. Whenever there is a portal between spaces, there is a line of sight between the spaces. Such a line of sight may be determined, for example, by casting rays from each space in all possible directions. Whenever a ray cast in this way hits another space, it means that the ray traveled through a portal. Combining all ray hits from a first space to a second space provides a rough shape of the projection of the portal between the spaces onto the surface of the second space. This may be considered one face of the portal. The starting point of the ray that hits the surface of the first space defines another face of the portal geometry. By forming an outer hull to combine these two faces, the complete portal geometry can be obtained.
[0075] In one exemplary embodiment, for each object source in the scene, four vectors are directed from the object source's position ("objsrc_pos") to the four vertices of the portal closest to objsrc_pos to generate the geometry of the sound emitted from the sound source toward the portal. Four 3D points are then translated 1 m along each vector, and a mesh is constructed using these points as vertices. The area of the formed mesh is divided by the area of the unit sphere (4π). This results in a direct propagation value to be used in the renderer. Note that in this exemplary embodiment, the directional pattern is assumed to be omnidirectional. In some other examples, the directional pattern may be non-omnidirectional, in which case integration over the directional pattern is performed as previously described.
[0076] The vertices for the portal opening are denoted as [vpopen0,vpopen1,vpopen2,vpopen3]. These vertices are selected as corners of the portal's face, which is most parallel to at least a portion of the wall of the space in which the portal is located and whose center point is closest to the center of the space. Note that if the face is not rectangular, the face may be bounded by a rectangle, and the corners of the bounding rectangle may be used as the vertices. Alternatively, if the face is not rectangular, four approximately equally spaced points may be selected from the perimeter of the face and used instead of the vertices. In some other examples, the portal's face indicating the portal's opening from the space may be indicated differently, for example, manually by a content creator. In still other examples, the vertices may be selected from a portal profile in the center of the portal geometry. In still other examples, the vertices may be selected from the portal's outermost face, and therefore toward the space from which power is transmitted.
[0077] For each vertex, a 3D vector Vobjsrc_vi,i=0,1,2,3 is formed between objsrc_pos and vpopen0, vpopen1, vpopen2, vpopen3.
[0078] Figure 4 visualizes [Vobjsrc_v0, Vobjsrc_v1, Vobjsrc_v2, Vobjsrc_v3] towards the portal opening vertices [vpopen0, vpopen1, vpopen2, vpopen3], respectively.
[0079] In that case, to get the area one meter away from the sound source, four 3D points (P0, P1, P2, P3) are defined for each vertex of the portal opening [vpopen0, vpopen1, vpopen2, vpopen3]. This is done by translating objsrc_pos along each of the vectors in turn.
[0080] When determining points Pi,i=0,1,2,3 along vector Vobjsrc_vi pointing to vertex vpopeni, P i To avoid translating objsrc_pos more than a certain threshold (essentially, beyond the point vpopeni), the distance d between objsrc_pos and vpopeni is first checked. To maintain the shape of the mesh, if plane_distance>d, a value of 0.1 is subtracted from d. Otherwise, objsrc_pos is translated by plane_distance (in one example, plane_distance=1), as shown in the code below. TIFF2025540822000003.tif56170
[0081] As shown in the code below, obj_src is translated along each vector by an amount d, and assigned a new x,y,z position in 3D Cartesian space. TIFF2025540822000004.tif64170
[0082] As visualized in Figure 5, the mesh objsrc is constructed from Pi,i=0,1,2,3 as vertices. That is, Figure 5 shows a mesh constructed from Pi,i=0,1,2,3.
[0083] To obtain the direct propagation value for the object source, we use the constructed mesh objsrc The area of Amesh_objsrc is divided by the area of a sphere with a diameter of 1 meter, i.e. sphere_radius=1 Asphere=4π*sphere_radius 2 directPropagationValue=A mesh_objsrc / A sphere
[0084] Algorithms for calculating the area of a mesh are readily available in the literature and software libraries. For example, the mesh may be triangulated and an iteration is performed over the triangles of the mesh. For each triangle, vectors are formed representing the two edges. The area of the triangle is taken as half the magnitude of the cross product of the edge vectors. The areas of the triangles are summed to obtain the complete mesh surface area.
[0085] The value directPropagationValue is used in the renderer as a factor for the direct sound power leaked through the portal opening to the connected room. When the above process for determining directPropagationValue is run on an encoder device, the following data can be written into the payload, as specified in RevPortalOpeningData (see below):
[0086] Syntax for transmitting direct propagation values in the bitstream
[0087] The following table shows the metadata syntax for providing propagation values directly from the encoder to the renderer. TIFF2025540822000005.tif152170
[0088] revNumSpaces is the number of spaces.
[0089] revNumPortalOpenings is the number of portal openings per space.
[0090] spaceBsId is the spatial bitstream identifier.
[0091] portalOpeningPositionX is the x element of the portal opening center position in x,y,z space.
[0092] portalOpeningPositionY is the y element of the portal opening center position in x,y,z space.
[0093] portalOpeningPositionZ is the z element of the portal opening center position in x, y, z space.
[0094] objSrcBsId is the bitstream identifier of the object source. Carried in the ScenePayload.
[0095] revNumObjsrcSpaces is the number of spaces over which the analysis is performed. It is greater than 1 if objsrc does not lie in a space, and equal to 1 otherwise.
[0096] revNumObjsrcPortalOpenings is the number of portal openings in space that will be iterated over
[0097] directPropagationValue is the direct propagation value for each object source towards each portal opening.
[0098] openingConnectionBsId is the identifier of the portal between the two spaces.
[0099] Note that above, examples are listed for the spaces in which the audio objects are located. For example, if the sound source is static, one space is sufficient. If the object is not static, several such listings, and therefore direct propagation values, may be listed for two or more possible spaces and / or objsrc locations. For sources that are not in a space, listings may be provided for all possible spaces and portal openings. Alternatively or additionally, for object sources that are not in a space, listings may be provided for the nearest or otherwise most relevant space and / or portal opening.
[0100] In the exemplary embodiment described above, the direct propagation value is derived by directly determining the surface area of the portal's projection onto the unit sphere, or more precisely, by approximating that projection, and does not involve integrating the source's directivity pattern over the unit sphere segment covered by the portal's projection. As explained above, this effectively means that the source is assumed to be omnidirectional in the method of the exemplary embodiment.
[0101] The exemplary embodiments described above for static sound sources can be extended to be applicable to non-omnidirectional sources by replacing the step of directly determining the surface area of the portal's projection with a step of integrating the source's directivity pattern over the unit spherical segment covered by the projection, taking into account the source's orientation, as previously described.
[0102] Essentially the same procedure as described in the exemplary embodiment above, or an extended version of that procedure that includes the effect of directional patterns, may be performed in real time by the renderer, where dynamic changes in source position and orientation may also be taken into account when calculating the direct propagation value for the source with respect to the portal. In this case, the direct propagation value will be a dynamic function of the source position and orientation.
[0103] In another exemplary embodiment for determining the projection of a portal onto a unit sphere around a sound source and the size of that projection, a method similar to that described in sections 6.5.5 ("DiscoverSESS") and 6.5.16 ("Homogeneous extent") of ISO / IEC 23090-4, Working Draft (WD1) of MPEG-I Immersive Audio [1], can be used. Similar to the method described above, the method described there uses ray tracing to find the projection (and, from that projection, the size) of the spatially extended sound source or portal onto the unit sphere, only that it is a projection onto the unit sphere around the listener, not around the source. However, the same method can be applied, with appropriate adaptations, to find the projection of a portal onto the unit sphere around a sound source and the size of that projection.
[0104] Rendering first-order reverberation in connected spaces
[0105] Once a direct propagation value is determined that quantifies the normalized amount of radiated source power propagated directly from source S in space 2 through the portal to connected space 1, this information can be used to generate and render reverberation in space 1 with a level consistent with this amount of power.
[0106] The principle here is to generate reverberation in Space 1 corresponding to a conceptual (a.k.a. "imaginary") omnidirectional sound source placed anywhere within Space 1, with source power equal to the amount of direct sound source power entering Space 1 through the portal.
[0107] Next, it is explained how this can be done by the renderer.
[0108] The power P of an omnidirectional point sound source radiating spherical sound waves is proportional to the square of the direct sound pressure p of the source. TIFF2025540822000006.tif7170r is the distance from the source, and ρ and c are the mass density and speed of sound in air, respectively.
[0109] Comparing the acoustic (physical) domain with the audio signal processing domain, the rms acoustic sound pressure p is proportional to the linear rms signal level of the audio input signal of the audio source, and the acoustic source power P is proportional to the square of this rms signal level of the input audio signal.
[0110] Thus, if we denote by X the normalized amount of source power that enters Space 1 directly through the portal, i.e., the direct propagation value, it follows from the above that the audio signal for the notional omnidirectional source used to generate reverberation in Space 1 is the audio signal corresponding to source S scaled by a linear gain of sqrt(X). That is, if we denote by s2 and s1 the signal of sound source S in Space 2 and the signal of the notional sound source in Space 1, respectively, then s1=sqrt(X)*s2.
[0111] Using s1 as an input signal, the reverberation for Space 1 can be generated according to the provided reverberation characteristics of Space 1, such as, for example, reverberation time, reverberation energy ratio, and so on.
[0112] In the above, it was assumed that the audio signal s2 of a source S includes all factors that affect the direct sound level of the source when it is rendered, other than the directional pattern and the distance to the source. Examples of such factors are a source gain ("volume control") that may be associated with the source, a so-called "reference distance" that may specify a distance from the source whose distance attenuation should be normalized to 1 (0 dB), or muting the source S, which is equivalent to applying a source gain equal to 0. If this is not the case, i.e., if the source signal s2 associated with the source S is a "raw" input signal to which factors that affect the direct sound level are applied when the source is rendered, such as an associated source gain, or a reference distance, or muting the source, these factors should also be applied to the signal s2 in the above procedure; i.e., if the combined effect of factors that affect the direct sound level is a direct sound gain G, the above formula can be modified as follows: s1 = sqrt(X) * G * s2 (G may be equal to 1).
[0113] Rendering of first-order reverberation for all sources in all connected spaces
[0114] In a realistic audio scene, there may be several spaces connected by portals, so the method needs to iterate over all sound sources S to determine whether they should be input to the reverberation for space 1. This can be done as follows: Determine the space in which the source S is located, If S is located in Space 1, then render the reverberation for source S according to the reverberation characteristics for Space 1; If S is not located in space 1 but is located in space 2, determine whether there is a portal from space 2 to space 1; If there is such a portal, determine the direct propagation value for source S with respect to the portal from space 2 to space 1 as described above, and render the reverberation for source S using a conceptual (a.k.a. imaginary) sound source in space 1.
[0115] Occlusion Detection and Handling
[0116] The direct propagation value for a source propagating through a portal may be optionally ramped down if the source is not "visible," i.e., does not have a direct line of sight through the portal. This may be the result of a geometric feature of the space (e.g., a wall between the source S and the portal), an object (e.g., an interactive object) that moves around and may obstruct (e.g., temporarily obstruct) the direct line of sight between the source and the portal, or a change in the state of the portal itself (e.g., an open door between two spaces that suddenly closes). Such occlusion detection may be performed, for example, by casting one or more rays from the portal's location toward the sound source. If one or more rays strike the sound source, it may be determined that the source is visible and the direct propagation value does not need to be ramped down. However, if no rays strike the sound source, it may be determined that the source is not visible and the direct propagation value may be ramped down.
[0117] Optionally, when a sound source becomes visible or ceases to be visible, further direct propagation value smoothing may be applied to prevent abrupt sound level transitions. Such direct propagation value smoothing may be performed by gradually increasing the direct propagation value (while the source becomes visible) or gradually decreasing its value (while the source becomes invisible). Ramping the value may be done, for example, by increasing or decreasing it by a predetermined value (such as 0.05) in each audio frame until the desired value is reached.
[0118] It should be noted that the above cross-fading can be equivalently performed on the directly propagated value or on the square root ("sqrt") of that directly propagated value, which is the gain value applied on the audio signal.
[0119] In another exemplary embodiment for determining the amount of occlusion of a Portal with respect to a sound source, a method similar to that described in Section 6.5.6 ("Occlusion") of ISO / IEC 23090-4, Working Draft (WD1) of MPEG-I Immersive Audio [1] may be used. The method described there uses ray tracing to find the amount of occlusion for a direct sound path between the source and the listener by casting rays from the listener position toward the extended sound source. The same method, with appropriate adaptations, may be applied to find the amount of occlusion of a Portal with respect to a sound source.
[0120] Portals that are not fully transparent
[0121] Above, for simplicity, it was assumed that the portal between the two spaces is completely acoustically transparent, i.e., the portal is completely open. Examples are generally an open door, a window, or an open portal.
[0122] However, the same concepts and principles also apply to portals that are not fully open and are only partially transparent, i.e., a general interface between spaces through which sound energy can be propagated. The only modification required for this generalization of the concept is that the propagated power, and therefore the direct propagation value, should be scaled by a number equal to the fraction of energy transmitted by the portal compared to a fully open portal of the same geometric size.
[0123] For example, if an interface between two spaces, say a thin wall, is made from a material that transmits 30% of the power that strikes it into a neighboring space, the direct propagation value for a source relative to that interface will be 30% of the value calculated by assuming the interface is an open portal.
[0124] In the general case where the energy transmission coefficient may vary across the portal, the scaling will be obtained by integrating the energy transmission coefficient over the portal area and normalizing by the geometric area of the portal.
[0125] It should also be noted that the transmission coefficients of the portals may be frequency dependent, and therefore the direct propagation values may also be frequency dependent, which will be reflected, for example, in the generation of relatively lower frequency reverberations in the connected spaces than higher frequency reverberations.
[0126] Cascading connected spaces
[0127] The method described above can be extended when three or more spaces are connected to each other in a cascade via multiple portals, in which case the main challenge is to determine the amount or fraction of the sound source power that propagates directly through multiple portals, rather than through a single portal as in the above description.
[0128] This can be done by projecting the portals between the continuously connected spaces onto a unit sphere around the sound source and determining the overlaps of the projections, which represent the angular regions where there is a direct "line of sight" between the source and the "farthest" connected space.
[0129] In the bitstream description above, a portal can be defined with a start space and an end space, i.e., the portal connects spaces 1, 2, and 3, and therefore spaces 1 and 3 can be indexed by there being a path from space 1 through space 2 into space 3.
[0130] Pre-calculating direct propagation values for any source position in space
[0131] In another approach, an interpolated grid of pre-computed values may be used to approximate the direct propagation value for an object source based on the position of the object source.
[0132] Each space has a bounding geometry ("spacebounding") that defines the extent of the space. Object source positions are interpolated from the x and y plane extent of the spacebounding with a resolution specified by interpResolution. In this example, a value of 0.5 is used, as shown in the code below. TIFF2025540822000007.tif123170The method distanceTo is further explained as (pow(x,y) denotes x raised to the yth power). TIFF2025540822000008.tif35170
[0133] The interpolated object source locations are then used and an object source-portal aperture analysis is performed for each location using the procedure described above. A grid of possible object source locations with corresponding direct propagation values is obtained.
[0134] Figure 6 shows the concept of an interpolated direct propagation value grid for possible object source positions. As shown in Figure 6, all possible x,y positions interpolated with interpResolution=0.5 are assigned a direct propagation value that depends on the angle and distance. This information can be used to approximate the behavior of the direct propagation values as the object source moves around in space.
[0135] In some example implementations, the direct propagation values in space for different possible object positions are modeled using a function in two variables. One example is to use a function of two variables (x, y), such as a polynomial function of two variables. The coefficients of such a polynomial can be signaled in the bitstream to a renderer device, where they can be used to obtain the direct propagation value for any source position in space.
[0136] Reverberation propagation through portals and the generation of "second-order" reverberation
[0137] The embodiments described above focused on reverberation generated in a first space as a result of direct sound from a source in a second space propagating from the second space through a portal to the first space. However, additional challenges were found in the propagation of reverberation from one space through a portal to another space and the subsequent generation of second-order reverberation in response to the propagated reverberation.
[0138] Referring again to the scenario of Figure 3, these additional challenges can be summarized as (1) realistic rendering in Space 1 of Space 2 reverberation propagating into Space 1 through a portal, or vice versa, realistic rendering in Space 2 of Space 1 reverberation propagating through a portal back into Space 2 (which is essentially the same problem), and (2) generation and realistic rendering in Space 1 of "second-order" reverberation generated in Space 1 in response to Space 2 reverberation propagating through a portal.
[0139] As with the propagation of direct sound through a portal, a key component in solving the problem of reverberant propagation through a portal and the subsequent generation of second-order reverberation is determining the amount of power (in this case, diffuse field power) transmitted through the portal.
[0140] Diffusion power P transmitted from the first space through the portal to the second space 1->2 It can be shown that the quantity has the following relationship with the diffuse sound pressure p1 in the first space: TIFF2025540822000009.tif8170
[0141] where S portal is m 2 is the size of the portal in units. When the portal is fully open, it is portal is equal to the geometric size of the portal, otherwise S portal is equal to the size of the fully opened portal that would transmit the same amount of sound, ρ is the mass density of air, and c is the speed of sound in air. portal may be obtained directly from or derived from the scene description metadata. For example, its size may be derived from the vertices that describe the portal, or from a mesh constructed from those vertices, as described in detail above. Also as described above, ray tracing techniques may be used to find the edges of the portal, from which the size of the portal may be determined.
[0142] Thus, if the level of reverberation in the first space and the (equivalent) size of the portal are known, the amount of diffuse power transferred to the second space can be determined.
[0143] From this amount of transmitted diffuse power, it is possible to generate second-order reverberation in the second space in a similar manner as described in the paragraph "Rendering first-order reverberation in connected spaces" above. That is, the transmitted diffuse power P 1->2is assigned to a conceptual point source located at an arbitrary position in the second space.
[0144] This is given by the diffuse pressure p1 in the first space and the direct sound pressure p1 at 1 m from the notional source in the second space 2,1m It can be shown that this leads to a relationship between TIFF2025540822000010.tif8170
[0145] If the reverberation strength of the second space is expressed as the reverberation-to-direct energy ratio (or vice versa) at 1 m from an omnidirectional point source, then this formula provides all the information needed to generate the correct level of second-order reverberation in the second space. In particular, if the reverberation-to-direct energy ratio for the second space is denoted as RDR2, then the desired pressure of the diffuse reverberation in the second space can be given by: TIFF2025540822000011.tif8170
[0146] In one embodiment, the conceptual process described above may be implemented by the following steps.
[0147] (1) deriving one or more reverberant input signals for generating second-order reverberation in a second space;
[0148] (2) obtaining the size of the portal through which sound is transmitted between the first space and the second space;
[0149] (3) determining a reverberation scaling factor that models the transmission of reverberant sound from the first space to the second space using the portal size; and
[0150] (4) Rendering a reverberant signal in the second space using the reverberant input signal(s) and the reverberation scaling factor. In some embodiments, the reverberant signal in the second space is also rendered using information indicative of the intensity of reverberation in the first space.
[0151] Further details and additional embodiments for generating and rendering second-order reverberation are described in U.S. Provisional Application No. 63 / 429,643, relevant portions of which are included in the section "Additional Disclosure."
[0152] The above described the generation and rendering of second-order reverberation in response to reverberation propagating from a first reverberant space through a portal to a second reverberant space. However, if the second space is a free field, i.e., has no reverberation of its own, the reverberation propagating from the first space to the second space will still be heard in the second space as reverberation coming from the portal. An example of this is when a listener is standing in a large open outdoor space in front of the open doors of a large cathedral where music is being played.
[0153] In this case, the portal acts as an extended sound "source", which, as explained above, transmits a diffused power P 1->2 Therefore, the steps in determining the correct power for this portal source are essentially the same as for the second-order reverberation case described above, with the key difference here being that the portal is not a (conceptual) omnidirectional point source, but rather has a source power P 1->2 The main difference is that it is an extended source that radiates all of the energy into only half space (i.e., the second space).
[0154] In one embodiment, the radiation from the portal may be modeled as hemispherical (i.e., radiating a spherical wave only into the second space), where the diffusion pressure p in the first space and the pressure p at 1 m from the portal source are 2,1mThe relationship between may be given by: TIFF2025540822000012.tif8170 or, TIFF2025540822000013.tif8170S portal is as previously defined.
[0155] In another embodiment, the pressure p2 at a particular location in the second space can be determined from the diffusion pressure p1 in the first space as follows: TIFF2025540822000014.tif8170where, Ω portal is the solid angle that the portal represents as "seen" from a particular position in second space, i.e., the angular size of the portal's projection onto the sphere around the particular position.
[0156] In yet other embodiments, a specific spatial radiation model for the portal source may be used to determine the desired relationship between the intensity of reverberation in the first space (proportional to the diffusion pressure p1) and the rendered sound level in the second space.
[0157] Although the rendering of propagated reverb has been described above for an example where the second space is free-field, i.e., non-reverberant, the same applies when the second space is reverberant; i.e., in that case too, reverberation propagating from the first space into the second space can be rendered as a "portal source," as described above, in addition to generating and rendering "second-order" reverberation.
[0158] Additional details and embodiments for rendering propagated reverberation from portals are described in U.S. Provisional Application No. 63 / 429,643, relevant portions of which are included in the "Additional Disclosure" section.
[0159] architecture
[0160] FIG. 7 shows an efficient architecture for use in an audio renderer that implements some of the embodiments described above.
[0161] N audio sources S1...Sn are located in several acoustic environments (e.g., rooms) AE1, AE2, AE3,... AEn. In some AEs, there are acoustic connections (i.e., portals) that acoustically connect the AE to one or more other AEs.
[0162] Each AE has an associated feedback delay network (FDN) (or other kind of reverberator) that models the late reverberation characteristics of this AE, taking into account relevant parameters such as frequency-dependent reverberation time (RT60), reverberation energy ratio, and pre-delay.
[0163] For all sources, the rendering of their direct signal portion and their early reflections is done using well-known methods and is based on the position / orientation of the sources and the position / orientation of the listener.
[0164] To render the late reverberant portions of these sources, the source signals are fed into a matrix of scaling factors, which weight each contribution from a particular source to a particular target AE / FDN and then sum the contributions of this AE to the FDN. The scaling factors in the matrix are the direct propagation values described above that determine how much direct sound energy propagates into adjacent rooms / AEs and causes late reverberation in these rooms / AEs. These factors may also include coupling factors (the reverberation scaling factors described above) that reflect how much of the late reverberant energy from the space / room / AE propagates back to the listener's location. The reverberation scaling / coupling factors, which are expressed in energy, must be converted to amplitude factors using a square root function to yield linear signal scaling factors.
[0165] The summed direct source contributions are then fed into each of the AE / FDNs, whose outputs can be rendered using virtual loudspeakers if so desired.
[0166] Finally, all AE / FDN contributions are summed together with the rendered direct sound and early reflections to form the fully rendered output.
[0167] 8 is a flowchart showing a process 800 for generating an input signal for a first space (e.g., space 301) of an XR scene (e.g., XR scene 390) based on a first sound source (e.g., sound source 391) in a second space (e.g., space 302) of the XR scene, where the first space is directly or indirectly connected to the second space through one or more portals, including at least a first portal (e.g., portal 300). Process 800 may begin at step s802. Step s802 includes obtaining a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the first sound source in the second space that is directly propagated from the first sound source through one or more portals into the first space. Step s804 includes generating the input signal for the first space using the direct propagation value and an audio signal associated with the first sound source.
[0168] In some embodiments, the method further includes using the input signal for the first space to generate a reverberation signal for the first space.
[0169] In some embodiments, the reverberation signal for the first space is generated using reverberation control information associated with the first space.
[0170] In some embodiments, the reverberation control information includes a reverberation level or reverberation energy ratio parameter associated with the first space.
[0171] In some embodiments, the method further comprises rendering the reverberant signal.
[0172] In some embodiments, s1=sqrt(X)*G*s2, where s1 is the input signal for the first space, X is the direct propagation value, G is a gain factor, and s2 is the audio signal associated with the first sound source.
[0173] In some embodiments, the method further includes receiving metadata associated with the XR scene, the metadata including a direct propagation value or data from which the direct propagation value can be derived, and obtaining the direct propagation value includes obtaining the direct propagation value from the metadata or obtaining the direct propagation value using the data.
[0174] In some embodiments, the metadata includes coefficients of a polynomial, and obtaining the direct propagation value includes using the polynomial coefficients to obtain the direct propagation value.
[0175] In some embodiments, obtaining the direct propagation value includes deriving the direct propagation value using information indicating the position of the first sound source in the second space and one or more of: i) information indicating the position of the first portal; ii) information indicating the size of the first portal; and iii) information indicating the shape of the first portal.
[0176] In some embodiments, one or more of information indicative of the directivity pattern of the first sound source or information indicative of the orientation of the first sound source are also used to derive the direct propagation value.
[0177] In some embodiments, information indicative of the directivity pattern of the first sound source and / or information indicative of the orientation of the first sound source is also used to derive the direct propagation value.
[0178] In some embodiments, the process further includes receiving metadata associated with the XR scene, the metadata including portal location information indicating locations of one or more portals, and obtaining the direct propagation value includes deriving the direct propagation value using the portal location information.
[0179] In some embodiments, for each portal, the metadata includes information that indicates the geometry of the portal.
[0180] In some embodiments, the first portal is associated with a transmission coefficient, and the direct propagation value is obtained using the transmission coefficient.
[0181] In some embodiments, obtaining the direct propagation value includes obtaining the direct propagation value using an interpolated grid of pre-calculated values and information indicative of a position of the first sound source.
[0182] In some embodiments, obtaining the direct propagation value includes forming a mesh representing the first portal, determining an area of the mesh, and using the area of the mesh to calculate the direct propagation value.
[0183] In some embodiments, obtaining the direct propagation value includes projecting an edge of the first portal onto a sphere around the sound source, determining a surface area of the sphere segment covered by the projection of the first portal to obtain a value, and normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
[0184] In some embodiments, obtaining the direct propagation value includes projecting an edge of the first portal onto a sphere around the sound source, integrating a directivity pattern associated with the sound source over the sphere segment covered by the projection of the first portal to obtain a value, and normalizing the obtained value by a surface area of the sphere, wherein the direct propagation value is the normalized value.
[0185] In some embodiments, the sphere is a unit sphere, and the surface area of the unit sphere is 4π.
[0186] In some embodiments, the first portal has eight vertices.
[0187] In some embodiments, the geometry of the first portal is converted to a box by enclosing the portal geometry in a box.
[0188] In some embodiments, the process further includes using an input signal for the first space to generate a first reverberant signal for the first space; obtaining a first scaling factor, the first scaling factor indicating an amount or portion of reverberant energy associated with the first reverberant signal that is propagated into the second space through one or more portals; generating a second input signal for the second space using the first scaling factor and the first reverberant signal; and using the second input signal for the second space to generate the second reverberant signal for the second space.
[0189] 9 is a flowchart illustrating a process 900 for enabling rendering of reverberation in a first space (e.g., space 301) of an XR scene (e.g., XR scene 390) that is directly or indirectly connected to a second space (e.g., space 302) of the XR scene through one or more portals, including at least a first portal (e.g., portal 300), where a sound source (e.g., sound source 391) resides in the second space. Process 900 may begin at step s902. Step s902 includes determining a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the sound source that is directly propagated from the sound source in the second space through one or more portals to the first space. Step s904 includes storing and / or transmitting the direct propagation value or data from which the direct propagation value can be derived.
[0190] In some embodiments, the process further includes sending metadata to an audio renderer, the metadata including the direct propagation value or data from which the direct propagation value can be derived.
[0191] In some embodiments, determining the direct propagation value includes forming a mesh representing the first portal, determining an area of the mesh, and using the area of the mesh to calculate the direct propagation value.
[0192] In some embodiments, determining the direct propagation value includes projecting an edge of the first portal onto a sphere around the sound source, determining a surface area of the sphere segment covered by the projection of the first portal to obtain a value, and normalizing the obtained value by the surface area of the sphere, wherein the direct propagation value is the normalized value.
[0193] In some embodiments, determining the direct propagation value includes projecting an edge of the first portal onto a sphere around the sound source, integrating a directivity pattern associated with the sound source over a sphere segment covered by the projection of the first portal to obtain a value, and normalizing the obtained value by a surface area of the sphere, wherein the direct propagation value is the normalized value. In some embodiments, the sphere is a unit sphere, and the surface area of the unit sphere is 4π.
[0194] FIG. 13 is a flowchart illustrating a process 1300 for rendering first-order reverberation in a first space (e.g., space 301) of an XR scene (e.g., XR scene 390). Process 1300 may begin at step s1302. Step s1302 includes determining whether a first sound source (e.g., sound source 391) is located in a second space (e.g., space 302). Step s1304 includes determining whether there are one or more portals connecting the first space with the second space. Step s1306 includes, as a result of determining that the first sound source is located in the second space and that there is at least a first portal (e.g., portal 300) connecting the first space with the second space, determining a first direct propagation value for the first sound source with respect to the first portal. Step s1308 includes generating an input signal for the first space using the direct propagation value and an audio signal associated with the first sound source. Step s1310 includes using the input signal for the first space to generate a reverberation signal for the first space.
[0195] 10 is a block diagram of an apparatus 1000 according to some embodiments for performing the methods disclosed herein. That is, the apparatus 1000 may implement the audio renderer 151 or the encoder 169. The apparatus 1000 may be referred to as an audio rendering apparatus when the apparatus 1000 implements an audio renderer, and the apparatus 1000 may be referred to as an encoding apparatus when the apparatus 1000 implements an encoder. As shown in FIG. 10 , device 1000 includes a processing circuit (PC) 1002 that may include one or more processors (P) 1055 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.), which may be co-located in a single housing or in a single data center or may be geographically distributed (i.e., device 1000 may be a distributed computing device), and at least one network interface 1048, such that device 1000 is The PC 1002 may comprise at least one network interface 1048, and a storage unit (a.k.a., a "data storage system") 1008, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments in which the PC 1002 includes a programmable processor, a computer program product (CPP) 1041 may be provided. The CPP 1041 includes a computer-readable medium (CRM) 1042, which stores a computer program (CP) 1043 comprising computer-readable instructions (CRI) 1044.CRM 1042 may be a non-transitory computer-readable medium, such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., a random access memory, a flash memory), or the like. In some embodiments, CRI 1044 of computer program 1043, when executed by PC 1002, configures CRI to cause device 1000 to perform steps described herein (e.g., steps described herein with reference to flowcharts). In other embodiments, device 1000 may be configured to perform steps described herein without the need for code. That is, for example, PC 1002 may simply consist of one or more ASICs. Thus, features of the embodiments described herein may be implemented in hardware and / or software.
[0196] Summary of Additional Embodiments
[0197] A1. A method for generating an input signal for a first space, Space 1, of an XR scene based on a sound source, S, in a second space, Space 2, of the XR scene, where Space 1 is connected to Space 2 through a portal, the method including obtaining a direct propagation value, X, that quantifies an estimate of an amount (e.g., a normalized amount) of radiated source power that is directly propagated from the sound source, S, in Space 2, through the portal to the connected Space 1; and generating an input signal, s1, for Space 1 using X and a signal, s2, related to S.
[0198] A2. The method of embodiment A1, further comprising using s1 to generate a reverberation signal for Space 1.
[0199] A3. The method of embodiment A2, further comprising rendering the reverberant signal.
[0200] A4. The method of any one of embodiments A1 to A3, wherein s1=sqrt(X)*G*s2, where G is a gain factor (eg, G≧0).
[0201] A5. A method according to any one of embodiments A1 to A4, wherein the method further includes receiving metadata associated with the XR scene, the metadata including a direct propagation value, and the direct propagation value is obtained from the metadata.
[0202] B1. A method for enabling rendering of reverberation in a first space, Space 1, of an XR scene connected to a second space, Space 2, of the XR scene via a portal, wherein a sound source, S, is present in Space 2, the method including determining a direct propagation value, X, that quantifies an estimate of an amount (e.g., a normalized amount) of radiated source power that is directly propagated from the sound source, S, in Space 2, through the portal to the connected Space 1, and storing and / or transmitting the direct propagation value.
[0203] B2. The method of embodiment B1, further comprising sending metadata to an audio renderer, the metadata comprising a direct propagation value.
[0204] B3. The method of embodiment B1 or B2, wherein determining X comprises forming a mesh, determining an area of the mesh, and using the area of the mesh to calculate X.
[0205] B4. The method of embodiment B1 or B2, wherein determining X includes projecting the edge of the portal onto a unit sphere around the sound source, integrating the directivity pattern of the sound source over the spherical segment covered by the projection of the portal to obtain a value, and normalizing the obtained value by the surface of the unit sphere (4π), where X is the normalized value.
[0206] C1. A computer program comprising instructions which, when executed by a processing circuit of the device 1000, cause the device to perform the method according to any one of the above embodiments.
[0207] C2. A carrier containing the computer program of embodiment C1, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0208] D1. An apparatus configured to perform the method according to any one of the above embodiments.
[0209] D2. The apparatus of embodiment D1, wherein the apparatus comprises a memory and a processing circuit coupled to the memory.
[0210] Additional Disclosures
[0211] As described above, this disclosure provides embodiments for generating plausible renderings of reverberation in a composite XR scene with connected acoustic spaces. For example, this disclosure provides means for determining an acoustic coupling factor that indicates the amount of reverberation propagating from a first space (the "acoustic environment") into a second space through a portal (e.g., an aperture or a partially transparent surface) connecting the two spaces. In one embodiment, the determined acoustic coupling factor is determined using information indicative of the size of the portal.
[0212] In some embodiments, based on the acoustic coupling factor, appropriate signal levels are set for rendering one or more audio signals in the second space, the one or more audio signals being derived from one or more reverberant signals corresponding to the first space. For example, in some embodiments, the determined acoustic coupling factor is used to derive a scaling factor, which is used to scale one or more audio signals derived from the set of one or more audio signals representing reverberation in the first space. In some embodiments, the scaling factor is further determined based on the amplitude, power, or energy of the total reverberant signal received at a position in the first space.
[0213] In some embodiments, the signal level for the signal rendered in the second space is determined based on one or more acoustic parameters for the first space, more particularly a reverberation level parameter or a reverberation energy ratio parameter associated with the first space.
[0214] Theoretical Framework
[0215] 11A shows a scene (real life or VR) consisting of two spaces: Space X and Space Y. The spaces are connected to each other via a portal 1100 (which may alternatively be called an "opening," "aperture," "interface," or similar). A sound source S1 is located somewhere in Space X. Sound source S1 generates a reverberant sound field in Space X.
[0216] A portal represents the interface between space X and space Y, through which a portion of the reverberant energy can be exchanged between those two spaces. portal m 2If the portal is an opening that is completely acoustically transparent, e.g., an open door or window, then the portal has an acoustic size S portal is simply equal to the geometric size of the portal (e.g., the geometric area of the portal). More generally, if the portal is not fully open and is only partially acoustically transparent (e.g., a thin wall or thick curtain separating two spaces), then the acoustic size of the portal, S portal is the equivalent size of a perfectly transparent aperture that represents the same amount of energy "leakage." In other words, if the portal is not perfectly acoustically transparent, the acoustic size of the portal, S portal will be smaller than the geometric size of the Portal. In the following, whenever there is reference to the "size" or "area" of a Portal, it is the "acoustic" size that is meant, unless otherwise specified.
[0217] A portion of the reverberant sound energy generated by a sound source in space X propagates through the portal into space Y, and a listener L located in space Y hears the reverberation coming from space X through the portal.
[0218] To determine the level of Space X reverberation perceived by listener L in Space Y, the amount of reverberant sound energy transferred from Space X through the portal to Space Y needs to be determined.
[0219] For simplicity, space Y is initially considered to be a "free field", which means that space Y is either a very large open (e.g., outdoor) space or a space with very high acoustic absorption, and therefore reverberation energy does not propagate from space Y back into space X, and therefore it is a one-way problem.
[0220] Assuming a steady-state diffuse sound field in space X, meaning that the amount of sound energy leaving space X per time unit, through absorption and portals, and into the connecting space is equal to the power of sound source S1 in space X, the so-called average reverberant energy density, E, in space X due to source S1 in space X is 1,1 but, E 1,1 =(4 / c)*(P1 / A 1,tot ) (1) It can be shown that where P1 is the acoustic power (expressed in watts) of sound source S1 in space X, and A 1,tot is the equivalent absorption area (m 2 ), where c is the speed of sound in air (in m / s).
[0221] A 1,tot is further expressed as follows: the amount of absorption A in space X excluding the portal 1,0 and (m 2 Portal size (in units) portal and A 1,tot =A 1,0 +S portal (2)
[0222] Furthermore, under diffusion steady-state conditions, the power P transferred from space X through the portal to space Y, expressed in watts, 1->2 But generally, P 1->2 =(c / 4)*E 1,1 *S portal (W) (3) It can be shown that it is equal to
[0223] Combining Equation 1, Equation 2, and Equation 3, we obtain the following for the power transferred from Space X through the portal to Space Y: P 1->2 =(S portal / A 1,tot )*P1=(S portal / (A 1,0+S portal ))*P1(W) (4)
[0224] The factor (S portal / A 1,tot ) is known as the "acoustic coupling factor" from space X to space Y, which indicates the fraction of source power in space X that is transferred to space Y under steady-state conditions.
[0225] As can be seen, the acoustic coupling factor is equal to the fraction of the total absorption in space X that is due to the portal. In other words, the power transferred from space X to space Y is multiplied by the size of the portal, S. portal The total amount of absorption in space X, A, is expressed by 1,tot is determined by the ratio of
[0226] Amount of absorption in space X, excluding portals, A 1,0 But, S portal If the acoustic coupling factor (S portal / A 1,tot ) is essentially equal to 1, so the amount of power transmitted through the portal is essentially equal to the power radiated by the source.
[0227] On the other hand, the amount of absorption in space X, excluding the portal, is A 1,0 But, S portal If the acoustic coupling factor (S portal / A 1,tot ) is (S portal / A 1,0 ), i.e., equal to the ratio of the size of the portal to the amount of absorption in space X excluding the portal, which in that case will be an extremely small number, i.e., only a very small percentage of the source power will be transmitted through the portal.
[0228] The average steady-state reverberation energy density in a space, E, is directly related to the root-mean-square steady-state reverberation sound pressure in the space, p, by the following relationship: p 2 =E*ρ0*c 2 (5) ρ0 is the mass density of air.
[0229] Thus, using Equation 1, for the steady-state reverberant acoustic pressure, p1, in space X due to source S1, we can write: TIFF2025540822000015.tif6170
[0230] Combining Equation 4 and Equation 6, the relationship between the steady state reverberation pressure p1 in space X and the power transmitted to space Y is found. TIFF2025540822000016.tif6170
[0231] Therefore, Equation 7 is the diffuse acoustic pressure p1 in space X and the size of the portal S portal provides an expression for the amount of power transferred from space X to space Y. Importantly, Equation 7 shows that if we know the diffuse acoustic pressure in space X and the size of the portal, this directly gives the amount of power transferred through the portal.
[0232] Next, to arrive at an expression for the relationship between the diffuse acoustic pressure p1 in space X and the acoustic pressure p2 in space Y associated with radiation through the portal, we define the power P transmitted through the portal as 1->2 and the resulting pressure p2 in space Y is required.
[0233] If the portal is relatively small, it can be assumed that the reverberation energy transmitted through the portal radiates equally in all directions (i.e., spherically) from the portal into space Y. With this assumption, this results in a pressure p at a distance of 1 m from the portal, which is equal to 2,1m This results in: TIFF2025540822000017.tif6170
[0234] This is due to the relationship between the acoustic source power P and the pressure p at a distance of 1 m of a source radiating a spherical wave: p 2 =P*ρ0*c / 4π (9) where in Equation 8 a factor of 2 is added to the power because the power P 1->2 is emitted only into the hemisphere on the space Y side of the portal (hence the resulting pressure should be that corresponding to a full spherical radiating source of twice that power).
[0235] Combining Equation 7 and Equation 8, we obtain the diffusion pressure p1 in space X and the rms pressure p1 at a distance of 1 m from the portal in space Y. 2,1m The following relationship between and is found: TIFF2025540822000018.tif6170
[0236] With respect to audio signal levels in an audio rendering system, the rms acoustic pressure p is directly proportional to the rms signal level of the corresponding audio signal; therefore, Equation 10 provides a direct way to relate the desired rms audio signal level in space Y to the rms reverberant audio signal level in space X.
[0237] Thus, in terms of linear acoustic pressure or linear rms audio signal level, Equation 10 states that at 1 m from the portal, the resulting acoustic pressure or rms audio signal level is expressed as: This means that the reverberation in space X should be rendered from the portal into space Y, with scaling such that it is scaled by a factor of TIFF2025540822000019.tif7170.
[0238] In terms of logarithmic sound pressure level (in dB) or logarithmic rms audio signal level, this means that the resulting level at 1 m from the portal is multiplied by 10*log relative to the diffuse sound pressure level or rms audio signal level in space X. 10 (S portal / 8π)=10*log 10 (S portal This means that the reverberation of space X is rendered in space Y so that the reverberation is -14 (dB).
[0239] As stated, this result is valid for small enough portals that the assumption of spherical radiation from the portal is valid. The size of the portal, S, portal is 8πm 2 If the rms pressure p in space Y exceeds 2,1m is greater than the rms pressure p in space X, which can never be the case physically, so it is clear that there is a limit to the applicability of Equation 10. In fact, the pressure p 2,1m Since the sound power associated with corresponds to only the portion of the diffuse sound field in space X that is incident on the aperture, i.e., it corresponds at most to the diffuse sound power from the hemisphere on the space X side of the aperture, TIFF2025540822000020.tif6170 is never TIFF2025540822000021.tif should not exceed 5170.
[0240] Equation 10 is S portal The reason that this may give physically implausible results for large values of θ is that the use of Equation 9 implies that the total power transmitted through the portal is in fact radiated from a single point, which, if this were actually the case, would indeed result in a much higher pressure near this single point than if the power were to be uniformly distributed and radiated from the entire portal (as would actually be the case in reality).
[0241] One way to address this issue is by modifying Equation 10 as follows: TIFF2025540822000022.tif6170
[0242] This correction ensures that the resulting level in space Y at 1 m from the portal is -3 dB (= 10 log ) above the diffuse reverberation level in space X, as physics dictates (as explained above). 10 (0.5)). Depending on the actual rendering method used to render the reverberation coming from the Portal, additional measures may be needed to ensure that, within a distance of 1m from the Portal, the level there also does not exceed -3dB the diffuse reverberation level in Space X. In other words, it should be guaranteed that the resulting level of the signal rendered from the Portal does not exceed -3dB the diffuse reverberation level in Space X anywhere in Space Y. This will also ensure a smooth transition in reverberation levels as the listener moves between spaces through the Portal.
[0243] S portal is 4πm 2 (=12.6m 2 ) that can potentially become too large when 2,1m Although a somewhat simple approach to solving the problem of , the solution in Equation 11 can actually provide quite plausible results in many use cases.
[0244] S portal is effectively 4πm 2 is smaller than the assumption of spherical radiation from the aperture is valid and S portal Using / 8π as a scaling factor gives plausible results.
[0245] The portal size is increased to 4πm 2 As the sound pressure level approaches 4πm, the sound pressure level at 1m from the portal approaches the reverberation sound pressure level of -3dB in space X, and finally 2For portals larger than 1 m, the reverberation reaches that level and remains there. The latter seems quite plausible, since the reverberation experienced when standing at a distance of 1 m from the center of an opening with dimensions of, say, 4 x 3 m (= 12 m²) is very similar to that experienced when standing within the opening itself.
[0246] Derived rms pressure at 1 m from the portal, p 2,1m From the above, the rms acoustic pressure p2 at any position in space Y at any distance d from the portal is p2(d)=p 2,1m It is easy to derive from / d.
[0247] Alternatively, other models for modeling radiation from a portal can be used in place of the spherical radiation model of Equations 8 and 9, which would result in a transmitted power P 1->2 and the resulting pressure p2 in space Y, and the resulting relationship between pressure p1 and pressure p2 in equations 10 and / or 11 (and consequently the scaling that is ultimately applied to the audio signal rendered from the portal to obtain the correct rendered audio signal level in space Y).
[0248] For example, a reverberant portal source can be modeled as a spatially diffuse, extended sound source, e.g., a spatially diffuse line or planar source, with a size equal to the geometric size of the portal. Acoustic radiation models for such spatially diffuse, extended sources are available in the literature that relate the acoustic source power to the resulting acoustic pressure at a given distance from the source.
[0249] Thus, in some embodiments, an alternative to Equation 10 or 11 may be used, of the form: TIFF2025540822000023.tif6170C1 is a constant, and even more generally, TIFF2025540822000024.tif7170 where f(Sportal ) is portal size S portal is a function of
[0250] Alternative Theoretical Frameworks
[0251] An alternative, but to a large extent equivalent, theoretical view of the transfer of reverberation from space X to space Y is presented.
[0252] Since the theoretical diffuse reverberant sound field at each point in space is created from equally strong uncorrelated plane waves arriving from all directions, the amount of reverberant energy propagating from space X through a portal to a particular point in space Y can be determined by geometric considerations.
[0253] The mean (rms) pressure of an individual reverberant plane wave arriving at any point in space X from a single solid angle dΩ is p d If the rms diffuse reverberation pressure, p1, at any point in space X is expressed as p d Integrating over all solid angles dΩ gives TIFF2025540822000025.tif6170
[0254] FIG. 11B shows a point L in space Y and the opening angle when "looking" from point L through the portal into space X. From the perspective of point L, the portal subtends a solid angle Ω portal (0≦Ω portal ≦2π).
[0255] Since a diffuse field in space X consists, by definition, of equally strong uncorrelated plane waves from all directions, and since the pressure of a plane wave is constant along its path (i.e., its pressure does not depend on the distance traveled), each individual plane wave that reaches point L in space Y through a portal will have the same uncorrelated pressure component p d and therefore the resulting pressure p2 at point L in space Y is p dthe solid angle Ω that the portal subtends at point L portal (where 0≦Ω portal ≤ 2π).
[0256] TIFF2025540822000026.tif8170
[0257] Combining Equation 14 and Equation 15, the result is that the pressure p2 at point L in space Y is expressed as a function of the solid angle Ω portal and the diffusion pressure in space X, p1. TIFF2025540822000027.tif7170
[0258] If the portal is not perfectly acoustically transparent and transmits a fraction T of the power incident on the portal, then p2 is scaled accordingly.
[0259] This result can be compared with Equations 10-12, which, like those equations, show that the square of the rms pressure in space Y is directly proportional to the square of the rms pressure in space X, with the proportionality factor being linearly dependent on the size of the portal, which in Equations 10-12 is expressed as the (equivalent) area (m 2 ) and in Equation 16 it is expressed in solid angles.
[0260] The diffuse reverberation pressure, p1, in space X can be determined from the power of the sound source and the amount of acoustic absorption in space X according to Equation 6:
[0261] One thing to note when comparing the two theoretical frameworks presented is that the first framework models the portal as a "secondary" sound source radiating into space Y, while the second framework directly considers the reverberant energy received from space X at a particular point in space Y. Since the solid angle subtended by the portal depends on the position relative to the portal, the pressure p2 obtained from Equation 16 also depends on the relative position of a particular location.
[0262] In particular, the pressure obtained from Equation 16 will be quite different for a position directly in front of the portal and a position to the side (or above / below) the portal.
[0263] If the portal is small enough, then at distance r from the observation point, S portal m 2 The solid angle subtended by a flat surface portal with a geometric area of It can be approximated by TIFF2025540822000028.tif7170, TIFF2025540822000029.tif4170 and TIFF2025540822000030.tif4170 are the position vector from the observation point to the portal and the normal vector of the portal, respectively.
[0264] At a distance of r=1m, this becomes Ω portal ≒S portal *cosθ (18) and θ is the observation angle relative to the portal normal vector.
[0265] Combining Equation 18 with Equation 16, we obtain Ω for a position just in front of the portal. portal ≒S portal and TIFF2025540822000031.tif7170. Comparing this to Eq. 10, we see that the rms pressure at 1 m obtained from Eq. 16 is a factor sqrt(2) greater than the rms pressure at 1 m from Eq. 10. On the other hand, at a position completely to the side of the portal, Ω portal ≒0, and as a result, TIFF2025540822000032.tif5170. This means that Equation 16 expresses the pressure at a specific point in space Y, while Equation 10, derived from the assumption of spherical radiation from the portal, expresses the average of Equation 16 at a distance of 1 m over all angles (i.e., the solid angle Ω at 1 m from the portal over all angles).portal The average value of S portal / 2).
[0266] Above it was assumed that the space Y was a free field, i.e. a space that does not generate diffuse reverberation itself (eg a large outdoor space).
[0267] rendering
[0268] An XR audio renderer can be configured to utilize the above model to provide a plausible rendering of reverberation in the connected space of an XR environment.
[0269] In one embodiment, the rendering of reverberation in a second, connected space Y that is related to reverberation generated in space X is split into two stages: (1) rendering of reverberation from space X that reaches the listener in space Y directly through a portal between the two spaces, and (2) generating and rendering of reverberation generated in space Y in response to reverberation from space X that enters space Y through the portal.
[0270] In some embodiments, only the first rendering stage may be performed, in other embodiments, only the second rendering stage may be performed, and in still other embodiments, both rendering stages may be performed.
[0271] First rendering stage
[0272] The first rendering stage does not inherently rely on the acoustics of space Y. For example, it is the reverberant sound that a listener standing in a large outdoor space would hear coming from the open door (i.e., portal) of a cathedral where music is being played.
[0273] In one embodiment, the sound level to be rendered in space Y as a result of a reverberant sound field in space X may include the following steps:
[0274] (1) determining a Space X reverberation intensity value representing the intensity of reverberation in Space X from one or more reverberation signals representing reverberation in Space X (a.k.a., "Space X reverberation signals");
[0275] (2) deriving one or more spatial Y reverberation signals (e.g., downmix signals) from the one or more spatial X reverberation signals for rendering in space Y;
[0276] (3) obtaining (e.g., determining, deriving, receiving) the size of the portal through which sound is transmitted between space X and space Y;
[0277] (4) determining a scaling factor that uses portal size to model the transmission of reverberant sound from space X to space Y; and
[0278] (5) Rendering one or more transmitted reverberation signals in space Y using the space Y reverberation signal, the space X reverberation intensity value, and the scaling factor.
[0279] Regarding the first step, the spatial X reverberation intensity value can be determined in various ways.
[0280] In a simple scenario where reverberation in Space X is rendered from a single (i.e., non-directional, monophonic) audio signal, the Space X reverberation intensity value may be determined simply as the rms amplitude or rms power of that signal, or, if the reverberation is rendered based on an impulse response, as the total amount of energy contained in the impulse response.
[0281] If the reverberation in space X is rendered using multiple audio signals, for example as multiple uncorrelated signals rendered from several directions around the user, the space X reverberation intensity value may be determined as the rms amplitude or power of the resulting combined signal. As an example, if the reverberation in space X has an rms amplitude of 1 / N (or 1 / N2 If we assume that the reverberant frequencies are rendered to a listener in space X as N uncorrelated reverberant signals from N corresponding directions, each with an rms power of 1 / N and an rms amplitude of 1 / sqrt(N), then the resulting combined reverberant signal has an rms power of 1 / N and an rms amplitude of 1 / sqrt(N).
[0282] In some cases, the Space X reverberation intensity value need not be determined from the actual Space X reverberant audio signal, but may be more efficiently derived from reverberation intensity metadata for Space X. For example, scene description metadata for an XR scene may include a reverb level or reverb energy ratio parameter for Space X that describes the desired reverberation level in Space X, either absolute or relative to the direct sound level or emitted source energy of a source in Space X that generates the reverberation. In such cases, the (relative) reverberation level in Space X is known a priori (and it is the renderer's job to generate Space X reverberant audio signals such that they result in the specified reverberation level in Space X). For example, assume Space X has associated metadata that includes a value for the reverberation-to-direct energy ratio (RDR) in Space X, which specifies the desired ratio of reverberant energy to direct sound energy at a distance of 1 meter from an omnidirectional audio source located somewhere in Space X. Now, if an omnidirectional audio source in space X has an associated audio signal with a linear rms signal amplitude s and an associated linear source gain ("volume control") g, then the rendered linear rms signal amplitude of the direct sound at 1m from that audio source is given by g*s, and therefore the rms power / energy of the direct sound signal is (g*s) 2 This results in the rms energy / power of the reverberation associated with an audio source being RDR*(g*s) 2and so the linear rms signal amplitude of the reverberation is sqrt(RDR)*g*s. Thus, the Space X reverberation intensity value can be derived directly from the provided Reverberation Energy Ratio (RDR) parameter for Space X and the source gain and audio signal level of the audio source.
[0283] If a source is not omnidirectional but has an arbitrary directional pattern associated with it, which causes the source to radiate a fraction X of the power of an omnidirectional source (for the same source signal), this causes the resulting reverberation power to also be a fraction X of the power for the omnidirectional source. Therefore, the derived spatial X reverberation strength value should be scaled accordingly, i.e., by a factor sqrt(X) if expressed in linear rms signal amplitude, or by a factor X if expressed in rms signal energy / power.
[0284] In addition to the source gain g, signal level s, and directivity pattern described above, other source rendering aspects that affect gain, either the rendered direct sound level or the rendered reverberation level, may be taken into account in calculating the Space X Reverberation Intensity value in a similar manner.
[0285] Regarding step 2, the step of deriving one or more spatial Y reverberation signals for rendering in space Y may be done in various ways. In one embodiment, the spatial Y reverberation signals may be a monophonic downmix from one or more spatial X reverberation audio signals.
[0286] In another embodiment, one or more spatial Y reverberation signals may be derived directly from the source signal and spatial X reverberation metadata parameters, such as the reverberation time RT60 and the reverberation energy ratio parameters, i.e., without the intermediate step of first generating an actual spatial X reverberation signal. This may be more efficient because the spatial X reverberation signal is not actually rendered to the listener (located in space Y), but is only generated as an intermediate step in generating one or more spatial Y reverberation signals.
[0287] With respect to step 3, the size of the portal may be obtained in various ways. In some embodiments, the size of the portal may be directly available in the scene description data, which may explicitly specify the portal's position and / or size in space and which other spaces it connects to. In other embodiments, the size may be derived from such scene description data, e.g., from geometry information. In still other embodiments, the size of the portal may be detected heuristically, e.g., using some form of ray tracing algorithm.
[0288] In some embodiments, the size of the portal is m 2 represents the area of the portal in units, which in some embodiments is the equivalent area of a perfectly acoustically transparent opening that has the same amount of "acoustic power leakage" as the portal.
[0289] In other embodiments, the size of a portal represents the solid angle that corresponds to the portal from a particular position in space Y. Methods for deriving that solid angle are readily available in the literature.
[0290] The scaling factor derived in step 4 represents the desired relationship between the intensity of the reverberation in space X (e.g., rms diffusion pressure, rms signal amplitude or rms signal power) and the intensity of the rendered space Y reverberation signal in space Y (e.g., rms diffusion pressure, rms signal amplitude or rms signal power).
[0291] In many embodiments, the basis for deriving the scaling factor may be given by any one of Equations 10-13 or 16, from which the scaling factor is calculated by multiplying p1 by p2, or alternatively, p1 by p2. 2 p2 2 can be derived as a factor relating
[0292] So, for example, the scaling factor is (S portal 10 as being equal to (Ω / 8π) (or its square root), while from Equation 16 the scaling factor is portal / 4π) (or its square root).
[0293] Finally, the derived spatial Y reverberation signal(s) are rendered to the listener in space Y using the spatial X reverberation intensity values and scaling factors.
[0294] The scaling factor and the Spatial X reverberation intensity value together determined the desired intensity of the rendered Spatial Y reverberation signal(s) through the relationship: desired intensity of the rendered Spatial Y reverberation signal(s) = scaling factor × Spatial X reverberation intensity value.
[0295] Having determined the desired intensity for the rendered spatial Y reverberation signal(s), an appropriate scaling gain can be determined for the spatial Y reverberation signal(s) that achieves this desired intensity for those rendered spatial Y reverberation signals.
[0296] In some embodiments, the scaling gain for the spatial Y reverberation signal(s) is simply equal to the scaling factor.
[0297] In some embodiments, the scaling gain for the spatial Y reverberation signal may take into account, in addition to the scaling factor, gain effects due to the particular way in which the spatial Y reverberation signal(s) are derived from the spatial X reverberation signal, as well as gain effects that arise due to different signal representation and rendering methods used for the spatial X and spatial Y reverberation signals.
[0298] As already explained, the spatial X reverberation may be represented by (and rendered as) a combination of multiple signals from which the spatial Y reverberation signal is derived using some signal transformation (e.g., downmixing) process that may result in some transformation gain effect, i.e., a difference in signal strength before and after the transformation. A scaling gain for the spatial Y reverberation signal(s) may compensate for this gain effect.
[0299] The scaling gain for the spatial Y reverberation signal may also compensate for gain effects that result from the particular way in which the reverberation signal is combined in the particular spatial X and spatial Y rendering methods used.
[0300] As a simple example, see the previous example where the space X reverberation is represented by N uncorrelated signals rendered from different directions around the listener in space X, each signal having an rms amplitude of 1 / N. The space X reverberation intensity value is then the rms amplitude of the sum of the N uncorrelated signals, which is equal to 1 / sqrt(N). Now assume that the space Y reverberation is derived from the space X reverberation signal by simply selecting one of the N signals with an rms amplitude of 1 / N. This space Y reverberation signal is then rendered as a point source located at a position within the portal and (S portal , a scaling factor according to Equation 10 of sqrt(N) / 8π, then an additional gain of sqrt(N) has to be applied to the space Y reverberation signal to obtain the correct balance between the reverberation intensity in space X and the reverberation intensity in space Y.
[0301] Thus, the basic concept is that the spatial Y reverberation signal is scaled so that the resulting intensity of the rendered spatial Y reverberation signal(s) has a desired relationship to the intensity of the spatial X reverberation as expressed by the scaling factor.
[0302] As previously explained, different rendering methods can be used to render the derived spatial Y reverberation signal.
[0303] In one embodiment, sound transmitted through a Portal is rendered to the listener as a sound source located within the Portal, i.e., a Portal Sound Source. In one embodiment, the Portal Sound Source is an extended sound source having a size corresponding to the geometric size of the Portal. The extended sound source may be a homogenous extended sound source that radiates the same signal from every point within a range, a diffuse extended sound source that radiates spatially spread signals from different points within a range, or a heterogeneous extended sound source that radiates partially correlated signals from different points within a range.
[0304] In another embodiment, the portal sound source is a point source. In one embodiment, the point source is located at a fixed location, for example, a central location within the portal. In another embodiment, the point source may be dynamically located within the portal depending on the listener location. For example, the point source may be located at a point within the portal that is closest to the listener location.
[0305] Second Rendering Stage
[0306] In a second rendering stage, reverberation is generated in space Y in response to the space X reverberation that entered space Y through the portal, according to the acoustic properties of space Y, e.g., space Y reverberation time, absorption, and / or reverberation level or reverberation energy ratio. Here, the rendering may be based on the amount of power transmitted from space X to space Y, e.g., according to Equation 7. The reverberation may then be generated as the reverberation of a point source located in space Y, with source power equal to the transmitted power.
[0307] More specifically, the second rendering stage may include the following steps:
[0308] (1) determining a Space X reverberation intensity value representing an intensity of reverberation in Space X from one or more Space X reverberation signals representing reverberation in Space X;
[0309] (2) deriving one or more reverberant input signals (e.g., downmix signals) for generating reverberation in space Y from one or more space X reverberant signals;
[0310] (3) obtaining (e.g., determining, deriving, receiving) the size of the portal through which sound is transmitted between space X and space Y;
[0311] (4) determining a scaling factor that uses portal size to model the transmission of reverberant sound from space X to space Y; and
[0312] (5) Rendering the reverberant signal in space Y using the space Y reverberant signal(s), the space X reverberant intensity value, and the scaling factor.
[0313] Thus, the steps for the second rendering stage are essentially the same as for the first rendering stage, as described below, but some of the details differ.
[0314] Steps 1 and 3 are the same as for the first rendering phase. Therefore, if both a first and a second rendering phase are performed, steps 1 and 3 only need to be performed once.
[0315] In step 2, the signal used to generate reverberation in space Y is derived. Generally, only a single reverberant input signal may be needed. Thus, if step 2 in the first rendering stage produced a single (e.g., mono downmix) signal, that signal may also be used as the reverberant input signal for the second rendering stage. In principle, any signal having the general characteristics of reverberation in space X may be used as the reverberant input signal in the second rendering stage, for example, a single Space X reverberant signal of multiple Space X reverberant signals, or a single reverberant signal from which multiple Space X reverberant signals were generated.
[0316] In step 4, the scaling factor is (S portal / 16π), i.e., a factor of 2 less than in the first rendering stage when using the model of Equation 10. The reason for this is that in the second rendering stage, the reasoning that led to the addition of the factor of 2 in Equation 8 does not apply here, and it is the "normal" relationship between source power and pressure for an omnidirectional point source in Equation 9 that should be used.
[0317] Finally, in step 5, the spatial Y reverberation is generated using a scaled version of the derived reverberant input signal as a source signal according to the reverberant characteristics (e.g., reverberation time, reverberant energy ratio) corresponding to the spatial Y. The scaling factor and the spatial X reverberant intensity value are used to scale the gain of the reverberant input signal used to generate the spatial Y reverberation. The scaling is such that when the scaled signal is to be rendered as a point source, the signal will have the desired level at a distance of 1 m from the point source, i.e., p2,1m 2 = scaling factor × p1 2 Reverberation is then generated from the scaled reverberant input signal to produce a reverberation with a desired intensity.
[0318] Further rendering aspects
[0319] If, as in a typical implementation, reverberation from space X is rendered into space Y from the portal as an extended sound source (also known as a "volumetric" or "sized" sound source) located at and having the same geometric size as the portal, the result, using Equation 11, will be even more realistic than if the sound from the portal were rendered as a point source located at a fixed point within the portal. In such implementations using extended portal sound sources, such as those used in the MPEG-I Immersive Audio standard, the distance to the extended sound source, i.e., the portal, is generally measured relative to the nearest point on the portal, rather than relative to some reference point on the portal (e.g., a center point). This means that if a user were to (virtually) walk along a path parallel to a large portal, the distance to the portal, i.e., the distance used in rendering the extended sound source to the user, would remain constant, which means that the sound level experienced by the user would also remain constant along this path, just as would be expected. (In contrast, if sound coming from a Portal were to be rendered as a point source at a fixed location within the Portal, the distance, and therefore the rendered sound level, would change as the user moved along the Portal.)
[0320] A similar effect can be achieved if the sound from the Portal is rendered to the user in space Y as a point source located at a dynamic position within the Portal that moves with the user, rather than at a fixed position within the Portal. In this case, the Portal point source is dynamically located at the position within the Portal closest to the user.
[0321] Furthermore, in implementations using an extended sound source to render sound from a portal, as described above, a distance attenuation function may generally be applied to the sound rendered from the extended portal sound source, the distance attenuation function taking into account the geometric size of the extended sound source as seen from the listening position, which may make the perceived effect even more realistic. For example, if the listening position is initially in front of and relatively close to the portal, the extended portal source may behave as a diffuse, planar sound source, and its rendered sound level may only decrease relatively slowly as the distance from the portal increases along a trajectory perpendicular to the portal. As the distance is further increased, the rate of level decrease with increasing distance becomes more rapid, eventually approaching that of a point source.
[0322] On the other hand, when the listening position is initially to the side of the portal, the "perceived" geometric size of the volumetric source's extent, i.e., the geometric size of the volumetric source as "seen" from the listener position, is much smaller than when standing directly in front of the volumetric source. If the distance is then increased while keeping the angle to the portal the same, the rendered sound level decreases more rapidly with increasing distance than the sound level that applied for a listening trajectory in front of the portal.
[0323] Cascading connected spaces
[0324] When three or more spaces are connected to each other, the propagation of reverberation from one space to all other spaces through their respective portals may be modeled by iterative application of Equation 7 and / or Equation 4, which models the amount of reverberation power transferred from one space through a portal to the next. For example, if three spaces 1, 2, and 3 are connected through a first portal between space X and space Y and a second portal between space Y and space X, the amount of reverberation power transferred to space X due to a sound source in space X may be determined from first applying Equation 7 to determine the power transferred from the diffuse reverberation pressure in space X to space Y through the first portal. This determined transferred power level may be used to generate reverberation in space Y according to space Y acoustic parameters (e.g., RT60 and reverberation energy ratio) to provide the diffuse reverberation pressure in space Y. Equation 7 may then be applied to this space Y diffuse reverberation pressure to calculate the amount of power transferred to space X through the second portal.
[0325] As an alternative to the step of rendering the reverberation in space Y based on the determined amount of power from space X to space Y and determining the diffuse reverberation pressure in space Y therefrom, the amount of power transmitted to space X can also be determined directly by applying Equation 4 to the result of the first step, i.e., using the amount of power transmitted from space X to space Y obtained in the first step as P1 in Equation 4. The only problem here is that applying Equation 4 requires the amount of absorption in space Y, A, which may not be directly available as metadata. 1,tot (or A 1,0 ) In this case, the amount of absorption can be estimated from the available parameters, in particular the reverberation energy ratio or a combination of the reverberation time RT60 and the volume of the space Y. Patent application P103111 describes a method for deriving the amount of absorption from these other parameters.
[0326] 12 is a flowchart illustrating a process 1200, according to some embodiments, for rendering reverberation in a space Y connected to a space X through a portal. The process 1200 may be performed by the audio renderer 151. The process 1200 may begin at step s1202.
[0327] Step s1202 includes determining a reverberation intensity value associated with the reverberation associated with the space X.
[0328] Step s1204 includes obtaining (eg, deriving) information indicative of the size of the portal.
[0329] Step s1206 includes determining a scaling factor using information indicative of the size of the portal.
[0330] Step s1208 includes rendering a set of one or more spatial Y reverberation signals in space Y using the scaling factors and the reverberation intensity values.
[0331] Summary of Various Further and Additional Embodiments
[0332] A1. A method performed by an audio renderer for rendering reverberation in a space Y connected to a space X via a portal, the method including: determining a reverberation intensity value associated with reverberation associated with the space X; obtaining (e.g., deriving) information indicative of a size of the portal; using the information indicative of the size of the portal to determine a scaling factor; and rendering a set of one or more space Y reverberation signals in the space Y using the scaling factor and the reverberation intensity value.
[0333] A2. The method of embodiment A1, wherein the set of one or more reverberation signals represents a reverberation sound field in Space X (this set of one or more signals is referred to as the "Space X reverberation signals"), and the method further includes deriving the set of one or more Space Y reverberation signals from the Space X reverberation signals prior to rendering the Space Y reverberation signal(s).
[0334] A3. The method of embodiment A2, wherein determining the reverberation intensity value includes determining the reverberation intensity value based on a set of one or more spatial X reverberation signals.
[0335] A4. The method of embodiment A2 or A3, wherein deriving a set of one or more spatial Y reverberation signals for rendering in space Y includes downmixing a set of one or more spatial X reverberation signals.
[0336] A5. The information indicating the size of the portal is the size value, S portal and determining the scaling factor is C1*S portal The method of any one of embodiments A1 to A4, comprising calculating:
[0337] A6. The method of embodiment A5, wherein C1 is about 1 / 8π.
[0338] A7. Determining the scaling factor is C1*S portal The method of embodiment A5 or A6, further comprising calculating the square root of:
[0339] A8. Determining the scaling factor is C1*S portal The method of embodiment A5 or A6, further comprising determining whether C is less than C2, where C2 is a predetermined number (eg, 0.5).
[0340] A9. The information indicating the size of the portal is the solid angle value, Ω portal The method of any one of embodiments A1 to A4, wherein
[0341] A10. The scaling factor is determined by C1*Ω portal The method of embodiment A9, comprising calculating: where C1 is a predetermined value (eg, C1=1 / 4π).
[0342] A11. The scaling factor is determined by C1*Ω portal The method of embodiment A10, further comprising calculating the square root of:
[0343] conclusion
[0344] While various embodiments have been described herein, it should be understood that these embodiments have been presented by way of example only, and not limitation. Thus, the breadth and scope of the present disclosure should not be limited by any of the exemplary embodiments described above. Moreover, unless otherwise indicated herein or clearly contradicted by context, any combination of the above-described objects in all possible variations thereof is encompassed by the present disclosure.
[0345] Additionally, while the processes described above and illustrated in the figures have been shown as a sequence of steps, this has been done for purposes of illustration only, and it is therefore contemplated that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel.
[0346] References
[0347] [1] ISO / IEC 23090-4, WD1 of MPEG-I Immersive Audio, output document of the 9th meeting of MPEG WG 6 Audio Coding, and
[0348] [2] O. Das and J.S.Abel, "Grouped Feedback Delay Networks for Modeling of Coupled Spaces," J.Audio Eng.Soc., vol. 69, no. 7 / 8, pp. 486–496, (July / August 2021). DOI: https: / / doi.org / 10.17743 / jaes.2021.0026.
Claims
1. 1. A method (800) for generating an input signal for a first space (301) of an augmented reality (XR) scene (390) based on a first sound source (391) in a second space (302) of the XR scene, wherein the first space (301) is directly or indirectly connected to the second space (302) via one or more portals, including at least a first portal (300), the method comprising: obtaining (s802) a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the first sound source in the second space that is directly propagated from the first sound source through the one or more portals into the first space; generating (s804) the input signal for the first space using the direct propagation values and an audio signal associated with the first sound source; The method (800) includes:
2. The method of claim 1 , further comprising using the input signal for the first space to generate a reverberation signal for the first space.
3. The method of claim 2 , wherein the reverberation signal for the first space is generated using reverberation control information associated with the first space.
4. The method of claim 3 , wherein the reverberation control information comprises a reverberation level or reverberation energy ratio parameter associated with the first space.
5. The method of claim 2 , further comprising rendering the reverberant signal.
6. s1=sqrt(X)*G*s2, where: s1 is the input signal for the first space, X is the direct propagation value; G is the gain factor, s2 is the audio signal associated with the first sound source; 6. The method according to any one of claims 1 to 5.
7. the method further comprising receiving metadata associated with the XR scene; the metadata includes the directly propagated value or data from which the directly propagated value can be derived; obtaining the directly propagated value includes obtaining the directly propagated value from the metadata or using the data.
7. The method according to any one of claims 1 to 6.
8. the metadata includes polynomial coefficients; obtaining the directly propagated values includes using the polynomial coefficients to obtain the directly propagated values. The method of claim 7.
9. The acquiring of the direct propagation value includes acquiring information indicating a position of the first sound source in the second space; i) information indicating the location of the first portal; ii) information indicating the size of the first portal; and iii) Information indicating the shape of the first portal and deriving the directly propagated value using one or more of:
10. one or more of information indicative of a directivity pattern of the first sound source or information indicative of an orientation of the first sound source are also used to derive the direct propagation value; or information indicative of a directivity pattern of the first sound source and / or information indicative of an orientation of the first sound source is also used to derive the direct propagation value.
10. The method of claim 9.
11. the method further comprising receiving metadata associated with the XR scene; the metadata includes portal location information indicating a location of the one or more portals; obtaining the direct propagation value includes deriving the direct propagation value using the portal location information; 7. The method according to any one of claims 1 to 6.
12. The method of claim 11 , wherein for each portal, the metadata includes information indicative of the portal's geometry.
13. the first portal is associated with a transmission coefficient; the direct propagation values are obtained using the transmission coefficients; 7. The method according to any one of claims 1 to 6.
14. 7. The method of claim 1, wherein obtaining the direct propagation values comprises using an interpolated grid of pre-calculated values and information indicative of a position of the first sound source to obtain the direct propagation values.
15. obtaining the directly propagated values, forming a mesh representing the first portal; determining the area of the mesh; using the area of the mesh to calculate the direct propagation value; 7. The method of claim 1, comprising:
16. obtaining the directly propagated value projecting an edge of the first portal onto a spherical surface around the sound source; determining a surface area of a spherical segment covered by the projection of the first portal to obtain a value; normalizing the obtained values by a surface area of the sphere, the directly propagated values being the normalized values; 7. The method of claim 1, comprising:
17. obtaining the directly propagated value projecting an edge of the first portal onto a spherical surface around the sound source; integrating a directivity pattern associated with the sound source over a spherical segment covered by the projection of the first portal to obtain a value; normalizing the obtained values by a surface area of the sphere, the directly propagated values being the normalized values; 7. The method of claim 1, comprising:
18. 18. The method of claim 16 or 17, wherein the spherical surface is a unit sphere, and the surface area of the unit sphere is 4π.
19. The method of claim 1 , wherein the first portal has eight vertices.
20. 20. The method of any one of claims 1 to 19, wherein the geometry of the first portal is converted to a box by enclosing the portal geometry in a box.
21. The method comprises: using the input signal for the first space to generate a first reverberation signal for the first space; obtaining a first scaling factor, the first scaling factor indicating an amount or portion of reverberant energy associated with the first reverberant signal that is propagated into the second space through the one or more portals; generating a second input signal for the second space using the first scaling factor and the first reverberant signal; using the second input signal for the second space to generate a second reverberation signal for the second space; 21. The method of any one of claims 1 to 20, further comprising:
22. 1. A method (900) for enabling rendering of reverberation in a first space (301) of an augmented reality (XR) scene (390) connected directly or indirectly to a second space (392) of the XR scene via one or more portals including at least a first portal (300), wherein a sound source (391) is present in the second space, the method comprising: determining (s902) a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the sound source that is directly propagated from the sound source in the second space through the one or more portals to the first space; storing and / or transmitting the directly propagated values or data from which the directly propagated values can be derived (s904); The method (900) includes:
23. 23. The method of claim 22, wherein the method further comprises sending metadata to an audio renderer, the metadata including the directly propagated value or the data from which the directly propagated value can be derived.
24. determining the direct propagation value; forming a mesh representing the first portal; determining the area of the mesh; using the area of the mesh to calculate the direct propagation value; 24. The method of claim 22 or 23, comprising:
25. Determining the direct propagation value comprises: projecting an edge of the first portal onto a spherical surface around the sound source; determining a surface area of a spherical segment covered by the projection of the first portal to obtain a value; normalizing the obtained values by a surface area of the sphere, the directly propagated values being the normalized values; 24. The method of claim 22 or 23, comprising:
26. Determining the direct propagation value comprises: projecting an edge of the first portal onto a spherical surface around the sound source; integrating a directivity pattern associated with the sound source over a spherical segment covered by the projection of the first portal to obtain a value; normalizing the obtained values by a surface area of the sphere, the directly propagated values being the normalized values; 24. The method of claim 22 or 23, comprising:
27. 27. The method of claim 25 or 26, wherein the spherical surface is a unit sphere, and the surface area of the unit sphere is 4π.
28. 1. A method (1300) for rendering first order reverberation in a first space (301) of an augmented reality (XR) scene (390), the method comprising: determining (s1302) whether a first sound source (391) is located in a second space (302); determining (s1304) whether there are one or more portals (300) connecting the first space with the second space; As a result of determining that a first sound source is located in the second space and that there is at least a first portal (300) connecting the first space with the second space, determining (s1306) a first direct propagation value for the first sound source with respect to the first portal; generating an input signal for the first space using the direct propagation values and an audio signal associated with the first sound source (s1308); using the input signal for the first space to generate a reverberation signal for the first space (s1310); The method (1300) includes:
29. 30. The method of claim 28, wherein the reverberation signal for the first space is generated using reverberation control information associated with the first space.
30. 30. The method of claim 29, wherein the reverberation control information comprises a reverberation level or reverberation energy ratio parameter associated with the first space.
31. 31. The method of any one of claims 28 to 30, wherein the method further comprises rendering the reverberant signal.
32. s1=sqrt(X)*G*s2, where: s1 is the input signal for the first space, X is the direct propagation value; G is the gain factor, s2 is the audio signal associated with the first sound source; 32. The method of any one of claims 28 to 31.
33. the method further comprising receiving metadata associated with the XR scene; the metadata includes the directly propagated value or data from which the directly propagated value can be derived; obtaining the directly propagated value includes obtaining the directly propagated value from the metadata or using the data.
33. The method of any one of claims 28 to 32.
34. the metadata includes polynomial coefficients; obtaining the directly propagated values includes using the polynomial coefficients to obtain the directly propagated values.
34. The method of claim 33.
35. The acquiring of the direct propagation value includes acquiring information indicating a position of the first sound source in the second space; i) information indicating the location of the first portal; ii) information indicating the size of the first portal; and iii) Information indicating the shape of the first portal and deriving the directly propagated value using one or more of:
36. one or more of information indicative of a directivity pattern of the first sound source or information indicative of an orientation of the first sound source are also used to derive the direct propagation value; or information indicative of a directivity pattern of the first sound source and / or information indicative of an orientation of the first sound source is also used to derive the direct propagation value.
36. The method of claim 35.
37. the method further comprising receiving metadata associated with the XR scene; the metadata includes portal location information indicating a location of the one or more portals; obtaining the direct propagation value includes deriving the direct propagation value using the portal location information; 33. The method of any one of claims 28 to 32.
38. 38. The method of claim 37, wherein for each portal, the metadata includes information indicative of the portal's geometry.
39. the first portal is associated with a transmission coefficient; the direct propagation values are obtained using the transmission coefficients; 33. The method of any one of claims 28 to 32.
40. 33. The method of any one of claims 28 to 32, wherein obtaining the direct propagation values comprises using an interpolated grid of pre-calculated values and information indicative of a position of the first sound source to obtain the direct propagation values.
41. obtaining the directly propagated values, forming a mesh representing the first portal; determining the area of the mesh; using the area of the mesh to calculate the direct propagation value; 33. The method of any one of claims 28 to 32, comprising:
42. obtaining the directly propagated value projecting an edge of the first portal onto a spherical surface around the sound source; determining a surface area of a spherical segment covered by the projection of the first portal to obtain a value; normalizing the obtained values by a surface area of the sphere, the directly propagated values being the normalized values; 33. The method of any one of claims 28 to 32, comprising:
43. obtaining the directly propagated value projecting an edge of the first portal onto a spherical surface around the sound source; integrating a directivity pattern associated with the sound source over a spherical segment covered by the projection of the first portal to obtain a value; normalizing the obtained values by a surface area of the sphere, the directly propagated values being the normalized values; 33. The method of any one of claims 28 to 32, comprising:
44. 44. The method of claim 42 or 43, wherein the spherical surface is a unit sphere, and the surface area of the unit sphere is 4π.
45. 45. The method of any one of claims 28 to 44, wherein the first portal has eight vertices.
46. 46. The method of any one of claims 28 to 45, wherein the geometry of the first portal is converted to a box by enclosing the portal geometry in a box.
47. The method comprises: using the input signal for the first space to generate a first reverberation signal for the first space; obtaining a first scaling factor, the first scaling factor indicating an amount or portion of reverberant energy associated with the first reverberant signal that is propagated into the second space through the one or more portals; generating a second input signal for the second space using the first scaling factor and the first reverberant signal; using the second input signal for the second space to generate a second reverberation signal for the second space; 47. The method of any one of claims 28 to 46, further comprising:
48. A computer program (1043) comprising instructions (1044) which, when executed by a processing circuit (1002) of an apparatus (1000), cause said apparatus to perform the method of any one of claims 1 to 47.
49. 51. A carrier containing the computer program of claim 50, said carrier being one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (1042).
50. An apparatus (1000) for generating an input signal for a first space (301) of an augmented reality (XR) scene (390) based on a first sound source (391) in a second space (302) of the XR scene, wherein the first space (301) is directly or indirectly connected to the second space (302) via one or more portals including at least a first portal (300), the apparatus comprising: obtaining (s802) a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the first sound source in the second space that is directly propagated from the first sound source through the one or more portals into the first space; generating (s804) the input signal for the first space using the direct propagation values and an audio signal associated with the first sound source; The apparatus (1000) is configured to perform the above.
51. 51. Apparatus according to claim 50, wherein the apparatus is further configured to carry out a method according to any one of claims 2 to 21.
52. 1. An apparatus (1000) for enabling rendering of reverberation in a first space (301) of an augmented reality (XR) scene (390) connected directly or indirectly to a second space (392) of the XR scene via one or more portals including at least a first portal (300), wherein a sound source (391) is present in the second space, the apparatus comprising: determining (s902) a direct propagation value, the direct propagation value indicating an amount or portion of source power associated with the sound source that is directly propagated from the sound source in the second space through the one or more portals to the first space; storing and / or transmitting the directly propagated values or data from which the directly propagated values can be derived (s904); The apparatus (1000) is configured to perform the above.
53. 53. Apparatus according to claim 52, wherein the apparatus is further configured to carry out a method according to any one of claims 22 to 27.
54. An apparatus (1000) for rendering first order reverberation in a first space (301) of an augmented reality (XR) scene (390), the apparatus comprising: determining (s1302) whether a first sound source (391) is located in a second space (302); determining (s1304) whether there are one or more portals (300) connecting the first space with the second space; As a result of determining that a first sound source is located in the second space and that there is at least a first portal (300) connecting the first space with the second space, determining (s1306) a first direct propagation value for the first sound source with respect to the first portal; generating an input signal for the first space using the direct propagation values and an audio signal associated with the first sound source (s1308); using the input signal for the first space to generate a reverberation signal for the first space (s1310); The apparatus (1000) is configured to perform the above.
55. 55. Apparatus according to claim 54, wherein the apparatus is further configured to carry out a method according to any one of claims 29 to 47.