Rendering sound transmitted through a structure
By deriving directional and distributed sound transmission coefficients from a total coefficient, the method addresses the lack of available values in XR systems, achieving realistic and efficient sound rendering in extended reality environments.
Patent Information
- Application Number
- PCT/EP2025/058034
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-12
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-16
AI Technical Summary
Existing audio rendering systems for extended reality (XR) struggle to realistically render sound transmission through structures due to the lack of available values for individual directional and distributed sound transmission coefficients, leading to sub-optimal and non-plausible rendering results.
A method to derive directional and distributed sound transmission coefficients from a single total transmission coefficient, ensuring a natural balance between these components, and optionally checking the plausibility of provided coefficients to ensure realistic sound rendering.
Enables realistic and efficient sound rendering in XR environments by deriving coefficients from a single available coefficient, maintaining a natural balance between directional and distributed sound components, and ensuring plausibility of rendering results.
Smart Images

Figure EP2025058034_16102025_PF_FP_ABST
Abstract
Description
RENDERING SOUND TRANSMITTED THROUGH A STRUCTURE TECHNICAL FIELD
[0001] This disclosure relates to rendering sound transmitted through a structure.BACKGROUND
[0002] In the physical world, sound can, to some extent, travel from one space, e.g., asource space where the sound source is located, to another space through a structure thatseparates the spaces from each other. Examples of such separating structure include but are notlimited to walls, closed doors, windows, curtains, etc.
[0003] Depending on the properties of the material of the separating structure, a larger orsmaller fraction of the sound reaching a side of the structure in the source space may be passedthrough the structure and radiated into the space on another side of the structure. Generally, thesound power (or sound energy) that falls onto the separating structure may be broken downinto three components: a component that is reflected by the structure and radiated back into thesource space, a component that is transmitted through the structure and radiated into the spaceon the other side of the structure, and a component that is absorbed by the structure anddissipated into heat.
[0004] Due to the law of conservation of energy, the sum of the three components shouldbe equal to the total sound power / energy that falls onto the separating structure ("incididentpower / energy").
[0005] The component that is reflected by the structure (“the reflected component”) maybe broken down into: a specular reflected component which is radiated back into the sourcespace along a specific direction, and a diffuse reflected component which is radiated diffuselyback into the source space.
[0006] Similarly, the component that is transmitted through the structure (“thetransmitted component”) may be broken down into: (1) a “semi-direct” or “directional”transmitted component which is passed through the structure and radiated from the structureinto the space on the other side in a way that preserves the original direction of the sound – i.e.,the direction from which the sound reaches the structure in the source space, and (2) a“distributed” transmitted component which is passed through the structure and radiated fromthe structure into the space on the other side in a distributed way, i.e., it is radiated from the whole structure so that there is in principle no direct relation anymore between the originaldirection of the sound and the direction (or directions) from which the sound is received by areceiver on the other side of the structure.
[0007] Thus, in the “semi-direct” or “directional” transmission, the perceived directionof the sound remains essentially the same as if the separating structure was not there, and onlythe level of the transmitted sound is attenuated, typically in a frequency dependent way.
[0008] On the contrary, in the “distributed” transmission, the incident sound energy isdistributed over and radiated from the entire surface / interface of the separating structure, eitherin a uniform or non-uniform way. The transmitted sound may be radiated diffusely or coherently, or in some frequency-dependent combined / intermediate radiation mode betweenthe diffused and coherent radiation.
[0009] For an audio rendering system, in particular for providing extended reality (XR)such as virtual reality (VR), mixed reality (MR), or augmented reality (AR), which has a goalto provide a realistic audio rendering of complex virtual scenes with multiple spaces, it isimportant that the system enables a realistic rendering of the transmission of sound throughstructures, preferably including a realistic rendering of the two individual components (“directional” and “distributed”) of the transmitted sound.
[0010] The draft MPEG-I (Moving Picture Experts Group – Immersive) Audio standard[1], which is cited at the end of this disclosure, allows specifying, for a certain material, thetwo components – i.e., “directional” and “distributed” – of sound transmission as well as thetwo components of sound reflection – i.e., “specular” and “diffuse” – individually andindependently using corresponding material coefficients that have values between 0 and 1 perfrequency band. In the standard, there is a boundary condition that the sum of all fourcoefficients per frequency band may not exceed 1.
[0011] The difference between the sum of all coefficients and 1 is assumed to indicatethe fraction of sound energy / power that is absorbed by the material. In the MPEG-I ImmersiveAudio standard, a “transmission coefficient” corresponds to the “directional” component of thetotal transmitted power / energy ("transmitted component") as defined above while a “couplingcoefficient” corresponds to the “distributed” component of the total transmitted power / energy("transmitted component"). SUMMARY
[0012] Certain challenges presently exist. For example, as mentioned above, an audiorendering system, in particular for providing XR, has a goal to provide a realistic audiorendering of complex virtual scenes with multiple spaces. Thus, the system should enable arealistic rendering of the transmission of sound through structures, preferably including arealistic rendering of the two individual components – “directional” and “distributed” – of thetransmitted sound.
[0013] However, textbooks and databases of acoustic material properties typically onlyspecify the total transmission coefficient which specifies the total fraction of incident energythat is passed through the material ("transmitted component"). In other words, those textbooksand databases typically provide a value for the sum of the coefficients corresponding to the“directional” and “distributed” components of the total transmitted power / energy. Generally,they do not provide values for the coefficients for these two individual components.
[0014] Thus, for a content creator authoring an XR scene, it may be difficult to findappropriate values for the two individual components for a desired material. Therefore, even ifa renderer includes the functionality that enables the separate rendering of the two individualcomponents of the transmitted sound, e.g., as described in [2] cited at the end of this disclosure, the material property data that is needed to use this functionality in a meaningful way may typically not be available.
[0015] In many cases, this will leave the content creator no other choice than to take theavailable value of the total transmission coefficient for the material and then guess how todivide this value into the two individual components or worse, assign the available value of thetotal transmission coefficient to either one of the two components. This will result in sub- optimal, and possibly non-plausible, rendering results.
[0016] Furthermore, for a content creator, it is also quite a burden to have to specify allindividual sound transmission-related coefficients separately and ensure that their combinationsatisfies the “sum≤1” criterion. This is especially true since specifying the coefficients andensuring that the combination of the coefficients satisfies the criterion must be done essentiallyfor all materials in the multi-room scene and all relevant frequency bands.
[0017] Thus, for content creators’ convenience as well as for ensuring the plausibility ofthe rendering of the transmitted sound, it would be beneficial to have a mechanism that enablesdetermining the material transmission properties in an efficient way, e.g., requiring only asingle coefficient instead of two or more coefficients while, at the same time, guaranteeing thephysical plausibility of the rendered sound.
[0018] Accordingly, in one aspect of the embodiments of this disclosure, there isprovided a method of generating an audio signal for an extended reality (XR) scene. The XRscene comprises a first virtual space, a second virtual space, and a virtual structure. A virtualsound source is located in the first virtual space and the virtual structure at least partially divides the first and second virtual spaces. The method comprises obtaining a value of a totaltransmission (TT) coefficient for determining (i.e., that can be used to determine) how much ofvirtual sound reaching the virtual structure from the first virtual space is transmitted to thesecond space through the virtual structure (i.e., that can be used to determine how much of thevirtual sound incident on the virtual structure passes through the virtual structure and radiates into the second space). The method further comprises, based on the value of the totaltransmission coefficient, deriving one or more of: (i) a value of a directional transmission (DRT)coefficient for determining how much of the virtual sound reaching the virtual structure istransmitted to the second space from a direction of the virtual sound source (i.e., that can be usedto determine how much of the virtual sound incident on the virtual structure passes through the virtual structure and radiates into the second space from a direction of the virtual sound source)and / or (ii) a value of a distributed transmission (DBT) coefficient for determining how much ofthe virtual sound reaching the virtual structure is transmitted to the second space in a distributedway (i.e., for determining how much of the virtual sound incident on the virtual structure passesthrough the virtual structure and radiates into the second space in a distributed way). The method further comprises generating the audio signal for the XR scene based on the DRT and / or DBT coefficients.
[0019] In another aspect of the embodiments of this disclosure, there is provided amethod of generating an audio signal for an extended reality (XR) scene. The XR scenecomprises a first virtual space, a second virtual space, and a virtual structure. A virtual soundsource is located in the first virtual space and the virtual structure at least partially divides thefirst and second virtual spaces. The method comprises obtaining either one of (i) a value of adirectional transmission (DRT) coefficient for determining how much of virtual sound reachingthe virtual structure from the first virtual space is transmitted to the second space from adirection of the virtual sound source and / or (ii) a value of a distributed transmission (DBT)coefficient for determining how much of virtual sound reaching the virtual structure from thefirst virtual space is transmitted to the second space in a distributed way. The method furthercomprises, based on the obtained one of the value of the DRT coefficient and the value of the DBT coefficient, deriving another one of the value of the DRT coefficient and the value of the DBT coefficient. The method further comprises generating the audio signal for the XR scene based on the DRT and / or DBT coefficients.
[0020] In a different aspect, there is provided a computer program comprisinginstructions which when executed by processing circuitry cause the processing circuitry to perform the method of any one of the above embodiments.
[0021] In a different aspect, there is provided a carrier containing the computer programof the above embodiment. The carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.
[0022] In a different aspect, there is provided an apparatus for generating an audio signalfor an extended reality (XR) scene. The XR scene comprises a first virtual space, a second virtualspace, and a virtual structure. A virtual sound source is located in the first virtual space and thevirtual structure at least partially divides the first and second virtual spaces. The apparatus is configured to obtain a value of a total transmission coefficient (TT) for determining how much of virtual sound reaching the virtual structure from the first virtual space is transmitted to the second space through the virtual structure. The apparatus is further configured to, based on thevalue of the total transmission coefficient, derive one or more of: a value of a directionaltransmission (DRT) coefficient for determining how much of the virtual sound reaching thevirtual structure from the first virtual space is transmitted to the second space from a direction ofthe virtual sound source and / or a value of a distributed transmission (DBT) coefficient fordetermining how much of the virtual sound reaching the virtual structure from the first virtualspace is transmitted to the second space in a distributed way. The apparatus is further configuredto generate the audio signal for the XR scene based on the DRT and / or DBT coefficients.
[0023] In a different aspect, there is provided an apparatus for generating an audio signalfor an extended reality (XR) scene. The XR scene comprises a first virtual space, a second virtualspace, and a virtual structure. A virtual sound source is located in the first virtual space and thevirtual structure at least partially divides the first and second virtual spaces. The apparatus isconfigured to obtain either one of a value of a directional transmission (DRT) coefficient fordetermining how much of virtual sound reaching the virtual structure from the first virtual spaceis transmitted to the second space from a direction of the virtual sound source and / or a value of adistributed transmission (DBT) coefficient for determining how much of the virtual soundreaching the virtual structure from the first virtual space is transmitted to the second space in adistributed way. The apparatus is further configured to, based on the obtained one of the value ofthe DRT coefficient and the value of the DBT coefficient, derive another one of the value of theDRT coefficient and the value of the DBT coefficient. The apparatus is further configured togenerate the audio signal for the XR scene based on the DRT and / or DBT coefficients.
[0024] In a different aspect, there is provided an apparatus comprising processingcircuitry and a memory. The memory contains instructions executable by the processingcircuitry, whereby the apparatus is operative to perform the method of any one of the above embodiments.
[0025] Some embodiments of this disclosure enable the rendering of the transmittedsound, i.e., the sound transmitted through a structure of a certain material, with a plausiblebalance, i.e., a natural sounding balance between the directional and distributed transmittedsound components. In these embodiments, the natural sounding balance between thedirectional and distributed transmitted sound components can be achieved by deriving one ormore coefficients of one or more transmitted sound components from a single coefficient of a transmitted sound component in a physically plausible way.
[0026] For example, in case a total transmission coefficient is available / provided, thecoefficient of the directional transmitted sound component and the coefficient of the distributed transmitted sound component may be derived based on the total transmissioncoefficient. These embodiments solve the problem that typically only the total transmissioncoefficient is readily available for many materials in the acoustics literature, not thecoefficients of the individual directional and distributed transmitted sound components.
[0027] Also, deriving the coefficients of the directional and distributed transmitted soundcomponents based on the total transmission coefficient is more convenient to content creatorswho now only need to provide this single commonly available coefficient rather than having toprovide two separate coefficients of which values may not be directly available.
[0028] By deriving the coefficients of the directional and distributed transmitted soundcomponents from the single total transmission coefficient in a physically plausible way, theembodiment can also guarantee a natural balance between the directional and distributedtransmitted sound components.
[0029] In another example, in case only one of the coefficients of the directionaltransmitted sound component and the coefficient of the distributed sound component is available / provided, the remaining one of the two coefficients, that is the coefficient that is not provided or available, may be derived based on the available / provided coefficient.
[0030] Additionally, according to some embodiments, in case both of the twocoefficients of the directional and distributed transmitted sound components are provided, theconsistency of the two provided coefficients, i.e., the plausibility of the combination of the twocoefficients may be checked and, if necessary, one or both of the provided coefficients may bemodified to result in a plausible combination, thereby ensuring the natural sounding balancebetween the two transmitted sound components. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings, which are incorporated herein and form part of thespecification, illustrate various embodiments.
[0032] FIG. 1A and 1B show an exemplary XR scene.
[0033] FIG. 2A shows a relationship between total transmission coefficient anddirectional / distributed transmission coefficients.
[0034] FIG. 2B shows a relationship between total transmission coefficient anddirectional / distributed transmission coefficients.
[0035] FIGS. 3A and 3B show an XR system.
[0036] FIG. 4 shows components of an audio renderer.
[0037] FIG. 5 shows a process according to some embodiments.
[0038] FIG. 6 shows a process according to some embodiments.
[0039] FIG. 7 shows an apparatus according to some embodiments.DETAILED DESCRIPTION
[0040] FIG. 1A shows an exemplary XR scene 100 where some embodiments of thisdisclosure can be applied. Note that, as mentioned above, an XR scene can be any one of anAR scene, a VR scene, and an MR scene. One specific example of the XR scene is a virtual environment where a game character controlled by a real-world user resides.
[0041] As shown in FIG. 1A, in the XR scene 100, there are provided a first virtual space102, a second virtual space 104, and a virtual structure 106 separating the first and second virtual spaces 102 and 104 from each other. Examples of the virtual structure 106 include but are not limited to a concrete wall, a drywall, a brick wall, a wooden panel, a glass window, a curtain, etc.
[0042] In the XR scene 100, a virtual sound source 108 is in the first virtual space 102and a virtual listener 112 is in the second virtual space 104. Note that even though FIG. 1shows the virtual sound source 108 is speakers of a television, in the embodiments of thisdisclosure, the virtual sound source 108 is not limited to any specific type of sound source.
[0043] As shown in FIG. 1B, the sound 122 outputted from the virtual sound source 108reaches the structure 106. As the sound 122 passes through the structure 106, some of thesound 122 may be absorbed by the structure 106 while some of the sound 122 may be reflectedback to the first space 102. The remaining parts 126 and 128 of the sound 122 passes throughthe structure 106 and reaches the second space 104. In this disclosure, these remaining parts ofthe sound 122 is collectively referred to as the “transmitted component of the sound.”
[0044] When the transmitted component of the sound is radiated from the structure 106towards the second space 104, the transmitted component is broken down into two components– first and second parts 126 and 128.
[0045] The first part 126 is a part of the transmitted component radiated from a directionof the virtual sound source 108 with respect to the virtual listener 112. In other words, in thefirst part 126 of the transmitted component, the original direction of the sound source 108 withrespect to listener 112 is preserved. In this disclosure, the first part 126 of the transmittedcomponent is referred to as a “directional” transmitted component.
[0046] The second part 128 is a part of the transmitted component radiated from thestructure 106 into the second space 104 in a distributed way, meaning that the incident soundenergy is distributed over and radiated from the entire surface of structure 106, either in a uniform or non-uniform way. The distributed transmitted sound may be radiated diffusely or coherently, or in some frequency-dependent combined / intermediate radiation mode betweendiffuse and coherent radiation. In other words, in the second part 128 of the transmittedcomponent, there is no direct relation between the direction of the sound source 108 withrespect to the listener 112 and the direction from which the listener 112 receives the secondpart 128 of the transmitted component from the structure 106. In this disclosure, the secondpart 128 of the transmitted component is referred to as “distributed” transmitted component.
[0047] As explained above, the coefficient (a.k.a., “parameter” or “fraction”) describingthe total transmitted component, e.g., 126 + 128, is generally available from textbooks andmaterial property databases. Typically, this coefficient is simply called the “transmission coefficient” or "power transmission coefficient" in the literature. This coefficient may also be called a total transmission coefficient because it can be used to determine how much of the sound 122 reaching the structure 106 passes through the structure 106 and thus reaches the second space 104. In other words, this coefficient can be used to determine the “total” amount,or fraction, of the sound that is “transmitted” through the structure 106.
[0048] The total transmission coefficient of the structure 106 may be associated withand / or indicate a material property of the structure 106 and / or a structure property of thestructure 106. The material property of the structure 106 may indicate the amount oftransmitted sound power per a unit of thickness (e.g., 1 meter) of the material of the structure106, e.g., as a fraction (i.e., a value between 0 and 1) of the incidident sound power. Togetherwith the thickness of the structure 106, this amount of transmitted sound power per a unit ofthickness may be used to determine the total fraction of incident sound power that is transmitted through the structure 106. As an example, if the material property of the structureis expressed as a fraction with a value of 0.1 per meter thickness and the structure has athickness of 0.2 meter, then the total transmission coefficient may be derived as (0.1)^(0.2) =0.6. The structure property of the structure 106 may directly indicate the total fraction ofincident sound power that is transmitted by the structure 106 associated with the specificthickness and the specific material of which the structure 106 is made of. In case the totaltransmission coefficient indicates the structure property, the coefficient may be a fraction between 0 and 1. In the remainder of this text, it is assumed that the total transmission coefficient indicates the structure property, i.e., it is a power fraction with a value between 0(or 0%) and 1 (or 100%) that represents the total fraction of incident sound energy / power thatpasses through the structure. However, it should be clear that all concepts and equationsdisclosed herein can be adapted in a straightforward way in case the coefficients are provided as a material property instead.
[0049] As discussed above, in contrast to the general availability of the total transmissioncoefficient, the coefficient for quantifying the magnitude of the directional transmitted component is not generally available or provided. Note that this coefficient may be called a directional transmission (DRT) coefficient and can be used to determine how much of thesound 122 reaching the structure 106 passes through the structure 106 and radiates from thestructure 106 towards the second space 104 in a direction that is substantially similar to the direction from which the sound 122 reaches the structure 106. Like the total transmission coefficient, the directional transmission coefficient (DRT) of the structure 106 may beassociated with and / or indicate a material property of the structure 106 and / or a structureproperty of the structure 106.
[0050] Similarly, the coefficient for quantifying the magnitude of the distributedtransmitted component is not generally available or provided. Note that this coefficient may becalled a distributed transmission coefficient (DBT) and can be used to determine how much ofthe sound 122 reaching the structure 106 passes through the structure 106 and radiates from the structure 106 towards the second space 104 in a distributed manner. Like the total transmission coefficient, the distributed transmission coefficient (DBT) of the structure 106 may beassociated with and / or indicate a material property of the structure 106 and / or a structureproperty of the structure 106.
[0051] As explained above, the lack of general availability of the directional anddistributed transmission coefficients may put extra burden on the content creators and maycause a plausibility issue. In order to solve this problem, according to some embodiments ofthis disclosure, at least one of the directional and distributed transmission coefficients isderived based on the total transmission coefficient. This derivation is based on the insight that, for real-world materials of the separating structure, e.g., the structure 106, the magnitudes ofthe “directional” and “distributed” transmitted sound components and their correspondingmaterial coefficients, are typically not independent from each other.
[0052] For example, it would be very unnatural to have a material which allows 100% ofthe sound energy that reaches the structure 106 from a sound source located at a given positionon one side of the material (“incident energy”) to pass through the material and then to radiateall of that 100% transmitted sound energy on the other side of the material in a distributed, e.g.spatially diffuse, way that removes all source direction information.
[0053] A hypothetical structure made of such material would be an ideal “diffusor” inthe sense that when the hypothetical structure is placed between a source space, e.g., the firstspace 102, and a receiver space, e.g., the second space 104, the structure would uniformlyspread all incoming source energy over its surface and re-radiate it diffusely into the receiverspace without any loss of energy. However, such material does not exist in the physical reality.
[0054] Likewise, it would be equally unnatural to have a material which allows only atiny fraction of all incident energy to pass through the material and then radiate this tinyamount of transmitted energy on the other side of the material with a complete preservation ofthe source direction information, i.e., without any distribution / diffusion of the incident energyover the material surface.
[0055] A hypothetical structure made of such material would be an ideal “attenuator” inthe sense that when the hypothetical structure is placed between a source space, e.g., the firstspace 102, and a receiver space, e.g., the second space 104, the structure would attenuate thelevel of the incident sound and re-radiate it into the second space 104 without any change tothe spatial properties of the incident sound. However, such material does not exist in thephysical reality.
[0056] In contrast to the above described “ideal” structures, a structure made of a real-life material that transmits a very large fraction of all incident energy (i.e., a structure with ahigh total transmission coefficient) typically transmits most of this transmitted energyessentially in a completely directional way with only a (typically modest) frequency-dependentattenuation as the main modification to the transmitted sound.
[0057] One example of such structure is a thin curtain. If a sound source is positionedbehind the curtain, for a listener on the other side of the curtain, there is hardly any differencebetween the sound perceived with the curtain present or absent. Specifically, there will behardly any noticeable difference in the spatial perception of the sound source. In other words,the perceived direction and size of the sound source would be substantially the same with andwithout the curtain present and perhaps only a slight attenuation of high frequencies will be perceived with the curtain present.
[0058] In other words, for a structure made of a real-life material that has a very hightotal transmission coefficient, the directional transmission coefficient will typically be relatively large and may essentially account for almost all transmitted sound energy while the distributed transmission coefficient will be relatively low.
[0059] Likewise, a structure made of a real-life material that only transmits a tinyfraction of all incident energy (i.e., a structure with a very low total transmission coefficient)typically does so in a distributed way such that, for a listener on the other side, the spatialperception of the transmitted sound field has essentially no direct relation to the original sourcedirection and / or size.
[0060] For example, if a very low level sound of a neighbor’s TV is heard through thewall, it will typically not be possible to localize the exact position of the TV behind the wall. In other words, for a structure made of a real-life material that has a very low total transmissioncoefficient, the directional transmission coefficient will be relatively low while the distributedtransmission coefficient will be relatively high.
[0061] More generally, it can be said that, for structures made of real-world materials,the balance between the directional and distributed transmission coefficients depends on thetotal transmission coefficient in a monotonous way, i.e., the smaller the total transmissioncoefficient is, the more the balance shifts towards the distributed transmission coefficient andthe higher the total transmission coefficient is, the more the balance shifts towards thedirectional transmission coefficient.
[0062] Using the above concept, in some embodiments of this disclosure, each of thedirectional transmission coefficient (^^^^^^^^^^^^) and the distributed transmission coefficient(^^^^^^^^^^^^) may be determined based on the total transmission coefficient (^^^^^^ ). Morespecifically, in some embodiments,^^^^^^^^^^^^ = ^(^^^^^^) (1)^^^^^^^^^^^^ = ^(^^^^^^ ) (2)where ^ and ^ are mathematical functions.
[0063] In one embodiment, the functions f and g may be such that ^^^^^^ = ^^^^^^^^^^^^ +^^^^^^^^^^^^ . Also, the functions f and g may be such that ^^^^^^^^^^^^ is larger than ^^^^^^^^^^^^when ^^^^^^tends to 1, and vice versa when ^^^^^^tends to 0.
[0064] One example of functions f and g that satisfy these conditions is:^^^^^^^^^^^^ = ^^^^^^ × (1 − ^^^^^^) (4)
[0065] The rationale of the specific functions in equations (3) and (4) is that the ratio ofeach of the directional and distributed transmission coefficients to the total transmissioncoefficient is a linear function, where the ratio of ^^^^^^^^^^^^to ^^^^^^linearly increases from 0 to 1 when ^^^^^^goes from 0 to 1, while the ratio of ^^^^^^^^^^^^to ^^^^^^linearly decreasesfrom 1 to 0 when ^^^^^^ goes from 0 to 1.
[0066] FIG. 2A shows a plot of the directional transmission coefficient and thedistributed transmission coefficient as functions of the total transmission coefficient accordingto equations (3) and (4). As shown in FIG. 2A, the distributed transmission coefficientaccording to the equation (4) never exceeds 0.25. This maximum value of the distributedtransmission coefficient is reached when the total transmission coefficient is 0.5 in which casethe directional and distributed transmission coefficients are both 0.25 – meaning that each ofthe directional and distributed transmission coefficients accounts for the half of the totaltransmitted power / energy.
[0067] In case the total transmission coefficient is larger than 0.5, as the totaltransmission coefficient becomes larger, the relative contribution of the distributed transmission coefficient to the total transmission coefficient becomes smaller, reaching 0%when the total transmission coefficient is 1 (e.g., for a very thin curtain). In other words, whenthe material allows most of the energy to pass through the material, the transmitted energy will mostly be radiated on the other side in a directional way.
[0068] On the contrary, in case the total transmission coefficient is smaller than 0.5, asthe total transmission coefficient becomes smaller, the relative contribution of the distributedtransmission coefficient to the total transmission coefficient becomes larger, approaching100% when the total transmission coefficient approaches zero. However, as the totaltransmission coefficient approaches zero, the distributed transmission coefficient itself alsoapproaches zero in that case. In other words, if the material only allows a very small part of the energy to pass through it, the transmitted energy will mostly be radiated on the other side in a distributed way.
[0069] Another example of a possible relationship between ^^^^^^^^^^^^ , ^^^^^^^^^^^^ , and^^^^^^ that satisfies the requirements stated directly after equations (1) and (2) is given by:^^^^^^^^^^^^= ^^^^^^* sin2(pi*^^^^^^ / 2) (5)^2^^^^^^^^^^^= ^^^^^^* cos(pi*^^^^^^ / 2) (6)
[0070] A plot of this function is shown in FIG. 2B. It has very similar properties as therelationship in equations (3) and (4) as shown in FIG. 2A. A difference is that the curve for^^^^^^^^^^^^is now slightly asymmetrical and has a maximum value of 0.26 at ^^^^^^= 0.42.
[0071] It will be clear that many other relationships between ^^^^^^^^^^^^ , ^^^^^^^^^^^^ ,and ^^^^^^ exist that satisfy the conditions stated directly after the general equations (1) and (2),and any of these relationships may be used in any of the embodiments of this disclosure.
[0072] In the above embodiments, the directional and distributed transmissioncoefficients are derived based on the total transmission coefficient assuming that neither of thedirectional and distributed transmission coefficients is available or provided. However, in somecases, it is not the total transmission coefficient that is specified but either one of thedirectional and distributed transmission coefficients. In these cases, according to someembodiments, the unavailable (unspecified) one of the directional and distributed transmissioncoefficients can be derived based on the available (specified) one of the directional anddistributed transmission coefficients in a way that preserves the real-world plausible balance.
[0073] More specifically, in case only the directional transmission coefficient isavailable, the distributed transmission coefficient can be derived based on the directionaltransmission coefficient. For example, the equations (3) and (4) from above embodiments canbe rewritten as:which shows how the distributed transmission coefficient can be derived from a specifieddirectional transmission coefficient in a physically plausible way.
[0074] Similarly, in case only the distributed transmission coefficient is available, thedirectional transmission coefficient can be derived based on the distributed transmissioncoefficient. For example, the equations (3) and (4) from above embodiments can be rewrittenas:which shows how the directional transmission coefficient can be derived from a specifieddistributed transmission coefficient in a physically plausible way.
[0075] Note that the equation (8) implies that, as mentioned above, the distributedtransmission coefficient can never exceed 0.25 because if ^^^^^^^^^^^^ exceeds 0.25, then^1 − 4 × ^^^^^^^^^^^^ would become an imaginary number.
[0076] The equations (3) and (4) provided above capture a gradually changing balancebetween the two transmission coefficients as a function of the total transmission coefficient,and may provide a sufficiently accurate model of the balance between the directional anddistributed transmission coefficients for typical real-world materials in many situations.However, in some embodiments, a different model may be used to capture the balance even more accurately over a given range of values of the total transmission coefficient. For example, the directional and distributed transmission coefficients may be derived based on the total transmission coefficient as follows:
[0077] Note that these relationships between ^^^^^^ and ^^^^^^^^^^^^ and between ^^^^^^and ^^^^^^^^^^^^ are still captured by the general relationship expressed in the equations (1) and(2) and satisfy the general condition that ^^^^^^ = ^^^^^^^^^^^^ + ^^^^^^^^^^^^ . It isstraightforward to derive similar relationships using other powers of ^^^^^^.
[0078] In some embodiments, the relationship between the directional transmissioncoefficient and the distributed transmission coefficient may be set such that the balancebetween them does not completely shift to only one of the two transmission coefficients whenthe total transmission coefficient approaches one of its extremes (i.e., 0% or 100%). In otherwords, even at the extreme end(s) of the total transmission coefficient, the smaller one of thetwo transmission coefficients may still have a non-zero value.
[0079] As an example of such an embodiment, equations (3) and (4) may be modified as,e.g.:^^^^^^^^^^^^ = 0.9 ∗ ^^^^^^^^^^^^^^^^^^^ = ^^^^^^ × (1 − 0.9 ∗ ^^^^^^) (12)
[0080] In this example, ^^^^^^^^^^^^ still has a non-zero value (0.1, in the example) whenthe total transmission coefficient ^^^^^^ is 1.
[0081] Conversely, in other embodiments, the relationship may be set such that thebalance between the two transmission coefficients shifts completely to one of the twocoefficients before the total transmission coefficient reaches one of its extremes (0% or 100%).In other words, in such embodiments, the balance between the two transmission coefficientsmay shift completely from one of the coefficients to the other coefficient within a limited rangeof the total transmission coefficient.
[0082] An example of such an embodiment may be given by a modified version ofequations (3) and (4) as, e.g.:
[0083] In this example, ^^^^^^^^^^^^ reaches a value of 1 (and ^^^^^^^^^^^^ a value of zero)already when the total transmission coefficient ^^^^^^reaches a value of 0.95.
[0084] In yet other embodiments, the relationship may be asymmetrical at the extremeends of the range of the total transmission coefficient. For example, the balance between thetwo transmission coefficients may shift completely to the distributed transmission coefficientbefore the total transmission coefficient reaches 0% whereas the balance may never completelyshift to the directional transmission coefficient even when the total transmission coefficientreaches 100%.
[0085] In some embodiments, depending on the availability of the various transmissioncoefficients, a sound renderer for rendering the sound for the XR scene 100 may operatedifferently.
[0086] For example, in case only the total transmission coefficient is specified for thestructure 106 (e.g., specified for the material of which the structure 106 is made) but not thedirectional and distributed transmission coefficients, a sound renderer for rendering the sound for the XR scene 100 may derive the directional and distributed transmission coefficients based on the given total transmission coefficient and use the derived directional and distributed transmission coefficients to generate an audio signal for the XR scene 100.
[0087] On the other hand, in case only the directional and distributed transmissioncoefficients are specified for the structure 106 but not the total transmission coefficient, thesound renderer may just use the given directional and distributed transmission coefficients togenerate the audio signal for the XR scene 100.
[0088] In case only one of the directional and distributed transmission coefficients isspecified for the structure 106, the sound renderer may derive the remaining one of thedirectional and distributed transmission coefficients based on the given one of the directionaland distributed transmission coefficients and use the given and derived coefficients to generatethe audio signal for the XR scene.
[0089] In case both the total transmission coefficient and the directional and distributedtransmission coefficients are specified for the structure 106, the sound renderer may determinewhich one of the total transmission coefficient and the directional / distributed transmissioncoefficients has a higher priority. In case the total transmission coefficient has a higherpriority, then the sound renderer may derive the directional and distributed transmission coefficients based on the given total transmission coefficient and use the derived directionaland distributed transmission coefficients to generate the audio signal for the XR scene 100. Onthe other hand, in case the directional / distributed transmission coefficients have a higherpriority, the sound renderer may just use the given directional and distributed transmissioncoefficients to generate the audio signal for the XR scene 100. The information indicating thepriorities of the coefficients may be provided to the sound renderer, or may be pre-set in therenderer itself (either explicitly, as a configurable parameter, or implicitly, e.g., in a selectionalgorithm implemented in the renderer).
[0090] In case the renderer allows specifying more than one transmission coefficient(i.e., total, directional, and / or distributed transmission coefficients), the renderer may check theplausibility of the combination of the provided coefficients, e.g., by using one or more of theequations 1-14 or any similar equation. For example, in case all of the total, directional, anddistributed coefficients are specified, the renderer may derive the directional and distributedtransmission coefficients based on the provided total transmission coefficient and compare thederived directional and distributed transmission coefficients to the provided directional anddistributed transmission coefficients. If the differences between the derived and provided coefficients are not within a permissible limit, the renderer may determine that the combination of the provided coefficients are not plausible. In determining whether the combination of thecoefficients is plausible, a plausibility threshold value may be used.
[0091] If the renderer determines that the combination of the provided coefficients is notplausible, the renderer may replace one or more of the provided coefficients such that the combination of the remaining provided coefficient(s) and the replacement coefficient(s) is plausible. To choose one or more of the provided coefficients to replace, the renderer may rely on a hierarchy of the different coefficients. For example, if the total transmission coefficient isat a higher level in the hierarchy as compared to the directional and distributed transmissioncoefficients, and if the combination of them is determined to be not plausible, the renderer mayderive new directional and / or distributed transmission coefficients based on the given totaltransmission coefficients and use the derived coefficients to generate the audio signal for the XR scene. Similarly, the directional and distributed transmission coefficients may have different levels in the hierarchy, so that for example the directional transmission coefficient may be retained while the distributed transmission coefficient is replaced.
[0092] In some embodiments, the renderer may be configured to allow to be providedwith a maximum of two transmission coefficients without any indication of whether theprovided transmission coefficient(s) are the total transmission coefficient, the directional transmission coefficients, and / or the distributed transmission coefficient. In such embodiments, the renderer may determine whether the provided transmission coefficient(s) are the total transmission coefficient, the directional transmission coefficient, and / or the distributed transmission coefficient. For example, in case only one transmission coefficient is provided,the renderer may determine that this only provided transmission coefficient is the totaltransmission coefficient. In another example, in case two transmission coefficients are provided, the renderer may determine that the first one of the two provided transmission coefficients is the directional transmission coefficient and the second one of the two provided transmission coefficients is the distributed transmission coefficient. In a different example, in case two transmission coefficients are provided, the renderer may determine that the first oneof the two provided transmission coefficients is the distributed transmission coefficient and thesecond one of the two provided transmission coefficients is the directional transmissioncoefficient. These embodiments may provide maximum flexibility in specifying the varioustransmission coefficients in a bit-efficient way.
[0093] As described above, in some embodiments, depending on the availability of thecoefficients, the sound renderer for rendering the sound for the XR scene 100 may operatedifferently. However, in other embodiments, the operation of the sound renderer may alsodepend on other metadata that the sound renderer receives. For example, the metadata maycontain one or more flags indicating the desired operation / behavior of the sound renderer –e.g., “Derive the directional and distributed transmission coefficients based on the given totaltransmission coefficient!”, “Use the first one of the two provided transmission coefficients as the directional transmission coefficient and the second one as the distributed transmission coefficient!”, etc.
[0094] The metadata may also include a flag indicating whether all provided coefficientsshould be used as specified, i.e., no plausibility check is to be done by the renderer. This would enable the content creator to, for artistic reasons, deliberately create a material behavior that is physically non-plausible.
[0095] In the above, it has been consistently assumed that the transmission coefficientsare expressed in terms of power / energy, which is common practice. However, in some cases the transmission coefficients may be expressed in terms of linear amplitude. In such cases, allequations shown above may be modified accordingly, where ^^^^^^ , ^^^^^^^^^^^^ and^^^^^^^^^^^^ are replaced by their squares. Also, the condition that ^^^^^^ = ^^^^^^^^^^^^ +^^^^^^^^^^^^may be replaced by:
[0096] FIG. 3A illustrates an XR system 300, according to one embodiment, in whichthe embodiments disclosed herein may be applied. As shown in FIG.3A, XR system 300 comprises an XR headset 320 (e.g., XR goggles, XR glasses, XR head mounted display (HMD), etc.) that is configured to be worn by a user and that is operable to display to the user an XR scene, such as, for example, the XR scene 100, speakers 334 and 335 for producing sound for the user, and an input device 350 for receiving input from the user. In this example, input device 350 is in the form of a joystick.
[0097] By wearing the XR headset 320, the user may experience as if the user is in anXR scene. For example, the XR headset 320 may display to the user the XR scene 100 shown in FIG. 1A such that the user experiences as if the user is the virtual listener 112. More specifically, via the XR headset 320, the view of the virtual listener 112 may be provided to the user and, via the speakers 334 and 335, the sound that the virtual listener 112 hears may be provided to the user.
[0098] As shown in FIG. 3B, the XR device 320 may comprise an orientation sensingunit 301, a position sensing unit 302, and a processing unit 303 coupled (directly or indirectly)to an audio renderer 351 for producing output audio signals, e.g., a left audio signal 381 for theleft speaker 334 shown in FIG. 3A and a right audio signal 382 for the right speaker 335 shownin FIG. 3A.
[0099] The orientation sensing unit 301 is configured to detect a change in theorientation of the user and provide information regarding the detected change to the processing unit 303. In some embodiments, the processing unit 303 determines the absolute orientation (in relation to some coordinate system) given the detected change in orientation detected by the orientation sensing unit 301. There could also be different systems for determination of orientation and position, e.g. a system using lighthouse trackers (lidar). In one embodiment, the orientation sensing unit 301 may determine the absolute orientation (in relation to some coordinate system) given the detected change in orientation. In this case the processing unit 303 may simply multiplex the absolute orientation data from the orientation sensing unit 301 and the positional data from position sensing unit 302. In some embodiments, the orientation sensing unit 301 may comprise one or more accelerometers and / or one or more gyroscopes.
[0100] Audio renderer 351 produces the audio output signals based on input audiosignals 361, metadata 362 regarding the XR scene 100 the user is experiencing, andinformation 363 about the location and orientation of the user. The metadata 362 for the XRscene 100 may include metadata for each object and audio element included in the XR scene100. For example, the metadata for the separating structure 106 may include one or more of thetotal transmission coefficient, the directional transmission coefficient, and / or the distributed transmission coefficient. The coefficients may be the coefficients of the structure 106 or the coefficients of the material of which the structure 106 is made of.
[0101] The audio renderer 351 may be a component of the XR device 320 or it may beremote from the XR device 320. For example, the audio renderer 351 or components thereofmay be implemented in the cloud.
[0102] FIG. 4 shows a simplified exemplary implementation of the audio renderer 351for producing sound for the XR scene 100. The audio renderer 351 includes a controller 401and an audio signal generator 402 for generating the output audio signal(s) (e.g., the audio signals of a multi-channel audio element) based on control information 410 from the controller 401 and the input audio signals 361. In this embodiment, the audio signal generator 402comprises a coefficient deriving unit (DU) 404 for deriving one or more coefficients, e.g., thedirectional and distributed transmission coefficient(s), based on a given coefficient, e.g., the total transmission coefficient.
[0103] In some embodiments, the controller 401 may be configured to receive one ormore parameters and to trigger the audio signal generator 402 to perform modifications on the audio signals 361 based on the received parameters (e.g., increasing or decreasing the volume level). The received parameters include the information 363 regarding the position and / or orientation of the listener (e.g., direction and distance to an audio element), metadata 362regarding the XR scene 100.
[0104] For example, the metadata 362 may include metadata regarding the XR space inwhich the user is virtually located (e.g., dimensions of the space, information about objects inthe space) as well as metadata regarding audio elements. One example of the metadata included in the metadata 362 is the total transmission coefficient associated with the separating structure 106.
[0105] Using the audio renderer 351, an audio signal for generating sound for the XRscene 100 may be generated. More specifically, the audio renderer 315 may receive the audioinput signals 361 for generating the sound 122 generated by the virtual sound source 108. Theaudio renderer 351 may also receive the metadata 362 associated with the XR scene 100. The metadata 362 associated with the XR scene 100 may comprise the total transmission coefficient associated with the virtual structure 106. The total transmission coefficient may be a coefficient of the virtual structure 106 itself or a coefficient of the material of which the virtual structure 106 is made of.
[0106] Upon receiving the metadata 362, the controller 401 may trigger the audio signalgenerator 402 to derive the directional and distributed transmission coefficients based on the total transmission coefficient included in the metadata 362 using the coefficient deriving unit(DU) 404, and then trigger the audio signal generator 402 to generate the output audio signalsusing the derived directional and distributed transmission coefficients. For example, the audio signal generator 402 may generate, by multiplying the input audio signal 361 by the directionaltransmission coefficient, a directional transmission audio signal which indicates how much ofthe sound 122 reaching the structure 106 from the first virtual space 102 passes through the structure 106 and radiates from the structure 106 into the second space 104 in the direction thatthe sound 122 reaches the structure 106 from the virtual sound source 108. Similarly, the audiosignal generator 402 may generate, by multiplying the input audio signal by the distributed transmission coefficient, a distributed transmission audio signal indicating how much of the sound 122 reaching the structure 106 from the first virtual space 102 passes through the structure 106 and radiates from the structure 106 into the second space 104 in a distributed way. Then, the audio signal generator 402 may generate the output audio signals by combining the directional and distributed transmission audio signals.
[0107] FIG. 5 shows a process 500 of generating an audio signal for an extended reality(XR) scene. The XR scene comprises a first virtual space, a second virtual space, and a virtualstructure. A virtual sound source is located in the first virtual space and the virtual structure atleast partially divides the first and second virtual spaces. The process 500 may begin with steps502. The step s502 comprises obtaining a value of a total transmission (TT) coefficient fordetermining how much of virtual sound reaching the virtual structure from the first virtual space is transmitted to the second space through the virtual structure. Step s504 comprises, based onthe value of the total transmission coefficient, deriving one or more of: (i) a value of a directionaltransmission (DRT) coefficient for determining (i.e., that can be used to determine) how much ofthe transmitted virtual sound is transmitted to the second space from a direction of the virtualsound source and / or (ii) a value of a distributed transmission (DBT) coefficient for determining(i.e., that can be used to determine) how much of the transmitted virtual sound is transmitted tothe second space in a distributed way. In other words, the DRT coefficient specifies (or can beused to derive a value that specifies) how much of the virtual sound incident on the virtual structure passes through the virtual structure and radiates into the second space from a directionof the virtual sound source, and the DBT coefficient specifies (or can be used to derive a valuethat specifies) how much of the virtual sound incident on the virtual structure passes through thevirtual structure and radiates into the second space in a distributed way. Step s506 comprisesgenerating the audio signal for the XR scene based on the DRT and / or DBT coefficients.
[0108] In some embodiments, the audio signal for the XR scene is for generating soundfor a real-world listener associated with a virtual listener who is located in the second virtualspace; and listening to virtual sound generated from the virtual sound source located in the first virtual space.
[0109] In some embodiments, the value of the TT coefficient depends on a dimension ofthe virtual structure and / or the type of a material of the virtual structure.
[0110] In some embodiments, the value of the DRT coefficient is less than the value ofthe DBT coefficient in case the value of the TT coefficient is less than a first value, the value of the DRT coefficient is greater than the value of the DBT coefficient in case the value of the TTcoefficient is greater than a second value, and the first and second values are the same ordifferent.
[0111] In some embodiments, the value of the DRT coefficient increases as the value ofthe TT coefficient increases, and the value of the DBT coefficient increases as the value of theTT coefficient increases as long as the value of the TT coefficient is less than the first value.
[0112] In some embodiments, the value of the DBT coefficient decreases as the value ofthe TT coefficient increases as long as the value of the TT coefficient is greater than the second value.
[0113] In some embodiments, ^^^^^^ = ^^^^^^^^^^^^ + ^^^^^^^^^^^^ , where ^^^^^^ is thevalue of the TT coefficient, ^^^^^^^^^^^^ is the value of the DRT coefficient, and ^^^^^^^^^^^^ isthe value of the DBT coefficient.
[0114] In some embodiments, ^^^^^^^^^^^^ is derived based on ^ × ^^^^^^^, where^^^^^^^^^^^^ is the value of the DRT coefficient, ^^^^^^ is the value of the TT coefficient, and eachof a and b is a positive real number.
[0115] In some embodiments,
[0116] In some embodiments, ^^^^^^^^^^^^ is derived based on ^^^^^^ × (1 − ^ × ^^^^^^^),where ^^^^^^ is the value of the TT coefficient, ^^^^^^^^^^^^ is the value of the DBT coefficient,and each of c and d is a positive real number.
[0117] In some embodiments,
[0118] In some embodiments, ^^^^^^^^^^^^ is derived based on ^^^^^^ * sin2(pi*^^^^^^ / 2),where ^^^^^^^^^^^^ is the value of the DRT coefficient, and ^^^^^^ is the value of the TTcoefficient.
[0119] In some embodiments, ^^^^^^^^^^^^ is derived based on ^^^^^^ * cos2(pi*^^^^^^ / 2),where ^^^^^^^^^^^^ is the value of the DBT coefficient, and ^^^^^^ is the value of the TTcoefficient.
[0120] In some embodiments, the process 500 comprises obtaining a virtual audio sourcesignal for producing the virtual sound generated by the virtual sound source and modifying the virtual audio source signal using the DRT and / or DBT coefficients, thereby generating at least amodified virtual audio source signal. The audio signal for the XR scene is generated based atleast on the modified virtual audio source signal.
[0121] FIG. 6 shows a process 600 of generating an audio signal for an extended reality(XR) scene. The XR scene comprises a first virtual space, a second virtual space, and a virtual structure. A virtual sound source is located in the first virtual space and the virtual structure at least partially divides the first and second virtual spaces. The process 600 may begin with steps602. The step s602 comprises obtaining either one of a value of a directional transmission(DRT) coefficient for determining how much of virtual sound transmitted through the virtualstructure from the first virtual space is transmitted to the second space from a direction of thevirtual sound source and / or a value of a distributed transmission (DBT) coefficient fordetermining how much of the virtual sound transmitted through the virtual structure from thefirst virtual space is transmitted to the second space in a distributed way. Step s604 comprises,based on the obtained one of the value of the DRT coefficient and the value of the DBT coefficient, deriving another one of the value of the DRT coefficient and the value of the DBT coefficient. Step s606 comprises generating the audio signal for the XR scene based on the DRT and / or DBT coefficients.
[0122] In some embodiments, the audio signal for the XR scene is for generating soundfor a real-world listener associated with a virtual listener who is located in the second virtualspace; and listening to virtual sound generated from the virtual sound source located in the firstvirtual space.
[0123] In some embodiments, the obtained one of the value of the DRT coefficient andthe value of the DBT coefficient depends on: a dimension of the virtual structure; and / or the typeof a material of the virtual structure.
[0124] In some embodiments, the value of the DRT coefficient is less than the value of theDBT coefficient in case a value of a total transmission coefficient (TT) is less than a first value, the value of the DRT coefficient is greater than the value of the DBT coefficient in case the value of the TT coefficient is greater than a second value, the first and second values are the same ordifferent, and the value of the TT coefficient is for determining a total amount of the virtualsound transmitted through the virtual structure from the first virtual space to the second space.
[0125] In some embodiments, the value of the DRT coefficient increases as the value ofthe TT coefficient increases, and the value of the DBT coefficient increases as the value of theTT coefficient increases as long as the value of the TT coefficient is less than the first value.
[0126] In some embodiments, the value of the DBT coefficient decreases as the value ofthe TT coefficient increases as long as the value of the TT coefficient is greater than the second value.
[0127] In some embodiments, ^^^^^^ = ^^^^^^^^^^^^ + ^^^^^^^^^^^^ , where ^^^^^^ is thevalue of the TT coefficient, ^^^^^^^^^^^^ is the value of the DRT coefficient, and ^^^^^^^^^^^^ isthe value of the DBT coefficient.
[0128] In some embodiments, the process 600 comprises obtaining the value of the DRTcoefficient and based on the obtained value of the DRT coefficient, deriving the value of theDBT coefficient. ^^^^^^^^^^^^ may be derived based on ^^^^^^^^^^^^^ × (1 − ^^^^^^^^^^^^^),where ^^^^^^^^^^^^is the value of the DBT coefficient, and ^^^^^^^^^^^^is the value of the DRT coefficient.
[0129] In some embodiments, the process 600 comprises obtaining the value of the DBTcoefficient and based on the obtained value of the DBT coefficient, deriving the value of theDRT coefficient. ^ may be derived base^^^^^^^^^^^ d on^ , where ^^^^^^^^^^^^is the value of the DRT coefficient, and ^^^^^^^^^^^^is the value of the DBT coefficient.
[0130] In some embodiments, the process 600 comprises obtaining a virtual audio sourcesignal for producing the virtual sound generated by the virtual sound source and modifying the virtual audio source signal using the DRT and / or DBT coefficients, thereby generating at least amodified virtual audio source signal. The audio signal for the XR scene is generated based atleast on the modified virtual audio source signal.
[0131] FIG. 7 is a block diagram of an XR rendering device 724, according to someembodiments, for performing the methods disclosed herein. As shown in FIG. 7, XR renderingdevice 724 may comprise: processing circuitry (PC) 702, which includes one or moreprocessors (P) 755 such as, for example, one or more general purpose microprocessors and / or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like, which processors may be co-located in asingle housing or in a single data center or may be geographically distributed (e.g., XRrendering device 724 may be a distributed computing apparatus comprising two or more computers or a monolithic computing apparatus consisting of a single computer); at least one network interface 748 (e.g., a physical interface or air interface) comprising a transmitter (Tx)745 and a receiver (Rx) 747 for enabling XR rendering device 724 to transmit data to andreceive data from other nodes connected to a network 110 (e.g., an Internet Protocol (IP) network) to which network interface 748 is connected (physically or wirelessly) (e.g., network interface 748 may be coupled to an antenna arrangement comprising one or more antennas forenabling XR rendering device 724 to wirelessly transmit / receive data); and a storage unit(a.k.a., “data storage system”) 708, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 702 includes aprogrammable processor, a computer readable storage medium (CRSM) 742 may be provided.CRSM 742 may store a computer program (CP) 743 comprising computer readable instructions (CRI) 744. CRSM 742 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, flash memory), and the like. In some embodiments, the CRI 744 of computer program 743 is configured such that when executed by PC 702, the CRI causes XR renderingdevice 724 to perform steps described herein (e.g., steps described herein with reference to theflow charts). In other embodiments, XR rendering device 724 may be configured to performsteps described herein without the need for code. That is, for example, PC 702 may consist merely of one or more ASICs. Hence, the features of the embodiments described herein may beimplemented in hardware and / or software.
[0132] Summary of Embodiments
[0133] A1. A method (e.g., process 500) of generating an audio signal for an extendedreality, XR, scene (e.g., 100), the XR scene comprising a first virtual space (e.g., 102), a secondvirtual space (e.g., 104), and a virtual structure (e.g., 106), wherein a virtual sound source (e.g.,108) is located in the first virtual space and further wherein the virtual structure at least partiallydivides the first and second virtual spaces, the method comprising: obtaining a value of a totaltransmission coefficient (TT) for determining how much of virtual sound reaching the virtualstructure from the first virtual space is transmitted to the second space through the virtualstructure; based on the value of the total transmission coefficient, deriving one or more of: avalue of a directional transmission (DRT) coefficient for determining how much of thetransmitted virtual sound is transmitted to the second space from a direction of the virtual soundsource; and / or a value of a distributed transmission (DBT) coefficient for determining how muchof the transmitted virtual sound is transmitted to the second space in a distributed way; andgenerating the audio signal for the XR scene based on the DRT and / or DBT coefficients.
[0134] A2. The method of embodiment A1, wherein the audio signal for the XR scene isfor generating sound for a real-world listener associated with a virtual listener who is: located inthe second virtual space; and listening to virtual sound generated from the virtual sound sourcelocated in the first virtual space.
[0135] A3. The method of embodiment A1 or A2, wherein the value of the TT coefficientdepends on: a dimension of the virtual structure; and / or the type of a material of the virtualstructure.
[0136] A4. The method of any one of embodiments A1-A3, wherein the value of the DRTcoefficient is less than the value of the DBT coefficient in case the value of the TT coefficient isless than a first value (e.g., 0.5), the value of the DRT coefficient is greater than the value of theDBT coefficient in case the value of the TT coefficient is greater than a second value (e.g., 0.5),and the first and second values are the same or different.
[0137] A5. The method of embodiment A4, wherein the value of the DRT coefficientincreases as the value of the TT coefficient increases, and the value of the DBT coefficientincreases as the value of the TT coefficient increases as long as the value of the TT coefficient isless than the first value.
[0138] A6. The method of embodiment A4 or A5, wherein the value of the DBTcoefficient decreases as the value of the TT coefficient increases as long as the value of the TTcoefficient is greater than the second value.
[0139] A7. The method of any one of embodiments A1-A6, wherein ^^^^^^ =^^^^^^^^^^^^ + ^^^^^^^^^^^^ , where ^^^^^^ is the value of the TT coefficient, ^^^^^^^^^^^^ is thevalue of the DRT coefficient, and ^^^^^^^^^^^^ is the value of the DBT coefficient.
[0140] A8. The method of any one of embodiments A1-A7, wherein ^^^^^^^^^^^^ is derivedbased on ^ × ^ ^^^^^^ , where ^^^^^^^^^^^^ is the value of the DRT coefficient, ^^^^^^ is the valueof the TT coefficient, and each a and b is a positive real number.
[0141] A9. The method of embodiment A8, wherein ^^^^^^^^^^^^ =
[0142] A10. The method of any one of embodiments A1-A9, wherein ^^^^^^^^^^^^ isderived based on ^^^^^^ × (1 − ^ × ^ ^^^^^^ ), where ^^^^^^ is the value of the TT coefficient,^^^^^^^^^^^^ is the value of the DBT coefficient, and each of c and d is a positive real number.
[0143] A11. The method of embodiment A10, wherein ^^^^^^^^^^^^ =
[0144] A12. The method of any one of embodiments A1-A7, wherein ^^^^^^^^^^^^ isderived based on ^2^^^^^* sin(pi*^^^^^^ / 2), where ^^^^^^^^^^^^ is the value of the DRTcoefficient, and ^^^^^^ is the value of the TT coefficient.
[0145] A13. The method of any one of embodiments A1-A7 and A12, wherein ^^^^^^^^^^^^is derived based on ^2^^^^^* cos(pi*^^^^^^ / 2), where ^^^^^^^^^^^^ is the value of the DBTcoefficient, and ^^^^^^ is the value of the TT coefficient.
[0146] A14. The method of any one of embodiments A2-A13, wherein the methodcomprises: obtaining a virtual audio source signal for producing the virtual sound generated bythe virtual sound source; and modifying the virtual audio source signal using the DRT and / orDBT coefficients, thereby generating at least a modified virtual audio source signal, and theaudio signal for the XR scene is generated based at least on the modified virtual audio sourcesignal.
[0147] B1. A method (e.g., process 600) of generating an audio signal for an extendedreality, XR, scene (e.g., 100), the XR scene comprising a first virtual space (e.g., 102), a second virtual space (e.g., 104), and a virtual structure (e.g., 106), wherein a virtual sound source (e.g., 108) is located in the first virtual space and further wherein the virtual structure at least partiallydivides the first and second virtual spaces, the method comprising: obtaining either one of: avalue of a directional transmission (DRT) coefficient for determining how much of virtual soundtransmitted through the virtual structure from the first virtual space is transmitted to the secondspace from a direction of the virtual sound source; and / or a value of a distributed transmission(DBT) coefficient for determining how much of the virtual sound transmitted through the virtualstructure from the first virtual space is transmitted to the second space in a distributed way; andbased on the obtained one of the value of the DRT coefficient and the value of the DBTcoefficient, deriving another one of the value of the DRT coefficient and the value of the DBTcoefficient; and generating the audio signal for the XR scene based on the DRT and / or DBTcoefficients.
[0148] B2. The method of embodiment B1, wherein the audio signal for the XR scene isfor generating sound for a real-world listener associated with a virtual listener who is: located in the second virtual space; and listening to virtual sound generated from the virtual sound source located in the first virtual space.
[0149] B3. The method of embodiment B1 or B2, wherein the obtained one of the value ofthe DRT coefficient and the value of the DBT coefficient depends on: a dimension of the virtualstructure; and / or the type of a material of the virtual structure.
[0150] B4. The method of any one of embodiments B1-B3, wherein the value of the DRTcoefficient is less than the value of the DBT coefficient in case a value of a total transmission coefficient (TT) is less than a first value (e.g., 0.5), the value of the DRT coefficient is greater than the value of the DBT coefficient in case the value of the TT coefficient is greater than asecond value (e.g., 0.5), the first and second values are the same or different, and the value ofthe TT coefficient is for determining a total amount of the virtual sound transmitted through the virtual structure from the first virtual space to the second space.
[0151] B5. The method of embodiment B4, wherein the value of the DRT coefficientincreases as the value of the TT coefficient increases, and the value of the DBT coefficientincreases as the value of the TT coefficient increases as long as the value of the TT coefficient isless than the first value.
[0152] B6. The method of embodiment B4 or B5, wherein the value of the DBTcoefficient decreases as the value of the TT coefficient increases as long as the value of the TT coefficient is greater than the second value.
[0153] B7. The method of any one of embodiments B4-B6, wherein ^^^^^^ =^^^^^^^^^^^^ + ^^^^^^^^^^^^ , where ^^^^^^ is the value of the TT coefficient, ^^^^^^^^^^^^ is thevalue of the DRT coefficient, and ^^^^^^^^^^^^ is the value of the DBT coefficient.
[0154] B8. The method of any one of embodiments B1-B7, wherein the methodcomprises: obtaining the value of the DRT coefficient; and based on the obtained value of theDRT coefficient, deriving the value of the DBT coefficient; and ^^^^^^^^^^^^ is derived based on^^^^^^^^^^^^^ × (1 − ^^^^^^^^^^^^^), where ^^^^^^^^^^^^ is the value of the DBT coefficient, and^^^^^^^^^^^^is the value of the DRT coefficient.
[0155] B9. The method of any one of embodiments B1-B7, wherein the methodcomprises: obtaining the value of the DBT coefficient; and based on the obtained value of theDBT coefficient, deriving the value of the DRT coefficient; and ^^^^^^^^^^^^ is derived based on(^^^×^^^^^^^^^^^^)±^^^^×^^^^^^^^^^^^^ , where ^^^^^^^^^^^^is the value of the DRT coefficient, and ^^^^^^^^^^^^is the value of the DBT coefficient.
[0156] B10. The method of any one of embodiments B2-B9, wherein the methodcomprises: obtaining a virtual audio source signal for producing the virtual sound generated bythe virtual sound source; and modifying the virtual audio source signal using the DRT and / orDBT coefficients, thereby generating at least a modified virtual audio source signal, and theaudio signal for the XR scene is generated based at least on the modified virtual audio source signal.
[0157] C1. A computer program 700 comprising instructions 744 which when executed byprocessing circuitry 702 cause the processing circuitry to perform the method of any one of embodiments A1-B10.
[0158] C2. A carrier containing the computer program of embodiment C1, wherein thecarrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.
[0159] D1. An apparatus 700 for generating an audio signal for an extended reality, XR,scene (e.g., 100), the XR scene comprising a first virtual space (e.g., 102), a second virtual space (e.g., 104), and a virtual structure (e.g., 106), wherein a virtual sound source (e.g., 108) is located in the first virtual space and further wherein the virtual structure at least partially divides the firstand second virtual spaces, the apparatus being configured to: obtain a value of a totaltransmission coefficient (TT) for determining how much of virtual sound reaching the virtual structure from the first virtual space is transmitted to the second space through the virtualstructure; based on the value of the total transmission coefficient, derive one or more of: a valueof a directional transmission (DRT) coefficient for determining how much of the transmittedvirtual sound is transmitted to the second space from a direction of the virtual sound source;and / or a value of a distributed transmission (DBT) coefficient for determining how much of thetransmitted virtual sound is transmitted to the second space in a distributed way; and generate theaudio signal for the XR scene based on the DRT and / or DBT coefficients.
[0160] D2. The apparatus of embodiment D1, wherein the apparatus is further configuredto perform the method of any one of embodiments A2-A14.
[0161] E1. An apparatus 700 for generating an audio signal for an extended reality, XR,scene (e.g., 100), the XR scene comprising a first virtual space (e.g., 102), a second virtual space (e.g., 104), and a virtual structure (e.g., 106), wherein a virtual sound source (e.g., 108) is located in the first virtual space and further wherein the virtual structure at least partially divides the firstand second virtual spaces, the apparatus being configured to: obtain s602 either one of: a valueof a directional transmission (DRT) coefficient for determining how much of virtual soundtransmitted through the virtual structure from the first virtual space is transmitted to the secondspace from a direction of the virtual sound source; and / or a value of a distributed transmission(DBT) coefficient for determining how much of the virtual sound transmitted through the virtualstructure from the first virtual space is transmitted to the second space in a distributed way; andbased on the obtained one of the value of the DRT coefficient and the value of the DBT coefficient, derive s604 another one of the value of the DRT coefficient and the value of theDBT coefficient; and generate s606 the audio signal for the XR scene based on the DRT and / orDBT coefficients.
[0162] E2. The apparatus of embodiment E1, wherein the apparatus is further configuredto perform the method of any one of embodiments B2-10.
[0163] F1. An apparatus 700 comprising: processing circuitry 702; and a memory 741,said memory containing instructions executable by said processing circuitry, whereby the apparatus is operative to perform the method of any one of embodiments A1-B10.
[0164] While various embodiments are described herein, it should be understood that theyhave been presented by way of example only, and not limitation. Thus, the breadth and scope ofthis disclosure should not be limited by any of the above-described exemplary embodiments.Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.
[0165] As used herein transmitting a message “to” or “toward” an intended recipientencompasses transmitting the message directly to the intended recipient or transmitting themessage indirectly to the intended recipient (i.e., one or more other nodes are used to relay themessage from the source node to the intended recipient). Likewise, as used herein receiving a message “from” a sender encompasses receiving the message directly from the sender or indirectly from the sender (i.e., one or more nodes are used to relay the message from the sender to the receiving node). Further, as used herein “a” means “at least one” or “one or more.”
[0166] Additionally, while the processes described above and illustrated in the drawings areshown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be re-arranged, and some steps may be performed in parallel.
[0167] Reference List
[0168] [1] ISO 23090-4: “MPEG-I part 4: Audio”, Working Draft (WD) 4, Clause 6.5.3:“Materials”, 04-09-2023.
[0169] [2] ISO / IEC JTC1 / SC29 / WG6, M63129, “CE on support for coupled transmission”,April 2023.
Claims
CLAIMS 1. A method (500) of generating an audio signal for an extended reality, XR, scene (100),the XR scene (100) comprising a first virtual space (102), a second virtual space (104), and avirtual structure (106), wherein a virtual sound source (108) is located in the first virtual space(102) and further wherein the virtual structure (106) at least partially divides the first and secondvirtual spaces, the method comprising:obtaining (s502) a value of a total transmission, TT, coefficient for determining how much of virtual sound reaching the virtual structure from the first virtual space is transmitted to the second space through the virtual structure; based on the value of the TT coefficient, deriving (s504): avalue of a directional transmission, DRT, coefficient for determining how muchof the virtual sound reaching the virtual structure from the first virtual space istransmitted to the second space from a direction of the virtual sound source;and / or a value of a distributed transmission, DBT, coefficient for determining how much of the virtual sound reaching the virtual structure from the first virtual space is transmitted to the second space in a distributed way; andgenerating (s506) the audio signal for the XR scene based on the value of the DRT coefficient and / or the value of the DBT coefficient.
2. The method of claim 1, wherein the audio signal for the XR scene is for generatingsound for a real-world listener associated with a virtual listener located in the second virtualspace.
3. The method of claim 1 or 2, wherein the value of the TT coefficient depends on:a dimension of the virtual structure; and / or the type of a material of the virtual structure.
4. The method of any one of claims 1-3, whereinthe value of the DRT coefficient is less than the value of the DBT coefficient in case thevalue of the TT coefficient is less than a first value,the value of the DRT coefficient is greater than the value of the DBT coefficient in casethe value of the TT coefficient is greater than a second value, andthe first and second values are the same or different.
5. The method of claim 4, whereinthe value of the DRT coefficient increases as the value of the TT coefficient increases,and the value of the DBT coefficient increases as the value of the TT coefficient increases as long as the value of the TT coefficient is less than the first value.
6. The method of claim 4 or 5, whereinthe value of the DBT coefficient decreases as the value of the TT coefficient increases as long as the value of the TT coefficient is greater than the second value.
7. The method of any one of claims 1-6, wherein ^^^^^^ = ^^^^^^^^^^^^ + ^^^^^^^^^^^^ ,where ^^^^^^is the value of the TT coefficient, ^^^^^^^^^^^^is the value of the DRT coefficient, and ^^^^^^^^^^^^is the value of the DBT coefficient.
8. The method of any one of claims 1-7, wherein ^^^^^^^^^^^^ is derived based on^ × ^^^^^^^, where ^^^^^^^^^^^^is the value of the DRT coefficient, ^^^^^^is the value of the TT coefficient, and aand b are positive real numbers.
9. The method of claim 8, wherein10. The method of any one of claims 1-9, wherein ^^^^^^^^^^^^ is derived based on^^^^^^ × (1 − ^ × ^^^^^^^), where ^^^^^^is the value of the TT coefficient, ^^^^^^^^^^^^ is the value of the DBT coefficient, andc and d are positive real numbers.
11. The method of claim 10, wherein12. The method of any one of claims 1-7, wherein ^^^^^^^^^^^^ is derived based on ^^^^^^* sin2(pi*^^^^^^ / 2), where ^^^^^^^^^^^^is the value of the DRT coefficient, and ^^^^^^is the value of the TT coefficient.
13. The method of any one of claims 1-7 and 12, wherein ^^^^^^^^^^^^ is derived based on^^^^^^* cos2(pi*^^^^^^ / 2), where ^^^^^^^^^^^^ is the value of the DBT coefficient, and^^^^^^is the value of the TT coefficient.
14. The method of any one of claims 1-13, whereingenerating the audio signal for the XR scene based on the value of the DRT coefficient and / or the value of the DBT coefficient comprises: obtaining a virtual audio source signal; modifying the virtual audio source signal using the DRT and / or DBT coefficients, thereby generating at least a modified virtual audio source signal; andgenerating the audio signal for the XR scene based at least on the modified virtual audio source signal.
15. A method (600) of generating an audio signal for an extended reality, XR, scene(100), the XR scene comprising a first virtual space (102), a second virtual space (104), and avirtual structure (106), wherein a virtual sound source (108) is located in the first virtual space(102) and further wherein the virtual structure (106) at least partially divides the first and secondvirtual spaces, the method comprising:obtaining (s602): avalue of a directional transmission, DRT, coefficient for determining how much ofvirtual sound reaching the virtual structure from the first virtual space istransmitted to the second space from a direction of the virtual sound source;and / or avalue of a distributed transmission, DBT, coefficient for determining how much ofthe virtual sound reaching the virtual structure from the first virtual space istransmitted to the second space in a distributed way;if the value of the DRT coefficient is obtained, then deriving (s604) the value of the DBT coefficient based on the obtained value of the DRT coefficient, otherwise deriving (s604) the value of the DRT coefficient based on the obtained value of the DBT coefficient; and generating (s606) the audio signal for the XR scene based on the value of the DRT coefficient and / or the value of the DBT coefficient.
16. The method of claim 15, wherein the audio signal for the XR scene is for generatingsound for a real-world listener associated with a virtual listener located in the second virtualspace.
17. The method of claim 15 or 16, wherein the obtained value of the DRT coefficientand / or the obtained value of the DBT coefficient depends on: a dimension of the virtual structure; and / or the type of a material of the virtual structure.
18. The method of any one of claims 15-17, whereinthe value of the DRT coefficient is less than the value of the DBT coefficient in case avalue of a total transmission coefficient, TT, is less than a first value, the value of the DRT coefficient is greater than the value of the DBT coefficient in casethe value of the TT coefficient is greater than a second value,the first and second values are the same or different, and the value of the TT coefficient is for determining a total amount of the virtual sound transmitted through the virtual structure from the first virtual space to the second space.
19. The method of claim 18, whereinthe value of the DRT coefficient increases as the value of the TT coefficient increases,and the value of the DBT coefficient increases as the value of the TT coefficient increases as long as the value of the TT coefficient is less than the first value.
20. The method of claim 18 or 19, whereinthe value of the DBT coefficient decreases as the value of the TT coefficient increases as long as the value of the TT coefficient is greater than the second value.
21. The method of any one of claims 18-20, wherein ^^^^^^ = ^^^^^^^^^^^^ + ^^^^^^^^^^^^ ,where ^^^^^^is the value of the TT coefficient, ^^^^^^^^^^^^is the value of the DRT coefficient, and ^^^^^^^^^^^^is the value of the DBT coefficient.
22. The method of any one of claims 15-21, whereinthe method comprises: obtaining the value of the DRT coefficient; andbased on the obtained value of the DRT coefficient, deriving the value of the DBT coefficient; and ^^^^^^^^^^^^ is derived based on ^^^^^^^^^^^^^ × (1 − ^^^^^^^^^^^^^), where^^^^^^^^^^^^is the value of the DBT coefficient, and ^^^^^^^^^^^^is the value of the DRT coefficient.
23. The method of any one of claims 15-21, whereinthe method comprises: obtaining the value of the DBT coefficient; andbased on the obtained value of the DBT coefficient, deriving the value of the DRT coefficient; and ^^^^^^^^^^^^is derived based^^^^^^^^^^^^is the value of the DRT coefficient, and ^^^^^^^^^^^^is the value of the DBT coefficient.
24. The method of any one of claims 16-23, whereingenerating the audio signal for the XR scene based on the value of the DRT coefficient and / or the value of the DBT coefficient comprises: obtaining a virtual audio source signal; modifying the virtual audio source signal using the DRT and / or DBT coefficients, thereby generating at least a modified virtual audio source signal; andgenerating the audio signal for the XR scene based at least on the modified virtual audio source signal.
25. A computer program (700) comprising instructions (744) which when executed byprocessing circuitry (702) cause the processing circuitry to perform the method of at least one ofclaims 1-24.
26. A carrier containing the computer program of claim 25, wherein the carrier is one ofan electronic signal, an optical signal, a radio signal, and a computer readable storage medium.
27. An apparatus (700) for generating an audio signal for an extended reality, XR, scene(100), the XR scene (100) comprising a first virtual space (102), a second virtual space (104),and a virtual structure (106), wherein a virtual sound source (108) is located in the first virtualspace (102) and further wherein the virtual structure (106) at least partially divides the first andsecond virtual spaces, the apparatus being configured to:obtain (s502) a value of a total transmission, TT, coefficient for determining how much of virtual sound reaching the virtual structure from the first virtual space is transmitted to the second space through the virtual structure; based on the value of the TT coefficient, derive (s504): avalue of a directional transmission, DRT, coefficient for determining how muchof the virtual sound reaching the virtual structure from the first virtual space istransmitted to the second space from a direction of the virtual sound source;and / or a value of a distributed transmission, DBT, coefficient for determining how much of the virtual sound reaching the virtual structure from the first virtual space istransmitted to the second space in a distributed way; andgenerate (s506) the audio signal for the XR scene based on the value of the DRT coefficient and / or the value of the DBT coefficient.
28. The apparatus of claim 27, wherein the apparatus is further configured to perform themethod of any one of claims 2-14.
29. An apparatus (700) for generating an audio signal for an extended reality, XR, scene(100), the XR scene comprising a first virtual space (102), a second virtual space (104), and avirtual structure (106), wherein a virtual sound source (108) is located in the first virtual space(102) and further wherein the virtual structure (106) at least partially divides the first and secondvirtual spaces, the apparatus being configured to:obtain (s602): avalue of a directional transmission, DRT, coefficient for determining how much ofvirtual sound reaching the virtual structure from the first virtual space istransmitted to the second space from a direction of the virtual sound source;and / ora value of a distributed transmission, DBT, coefficient for determining how much ofvirtual sound reaching the virtual structure from the first virtual space istransmitted to the second space in a distributed way;if the value of the DRT coefficient is obtained, then derive (s604) the value of the DBT coefficient based on the obtained value of the DRT coefficient, otherwise derive (s604) the value of the DRT coefficient based on the obtained value of the DBT coefficient; and generate (s606) the audio signal for the XR scene based on the value of the DRT coefficient and / or the value of the DBT coefficient.
30. The apparatus of claim 29, wherein the apparatus is further configured to perform themethod of any one of claims 16-24.