Audio device and rendering method thereof
The audio rendering system addresses the challenge of accurately representing audio from multiple acoustic environments by using energy transfer indicators to adjust audio levels, resulting in a more natural and efficient rendering of audio signals across different rooms.
Patent Information
- Application Number
- CN202380084156.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-06
- Filing Date
- 2023-12-04
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, when rendering audio from multi-room scenes, it is difficult to accurately represent and render audio in different acoustic environments, resulting in unnatural audio experience and excessive computing resources consumption.
By using audio devices and methods to receive audio data and metadata, the renderer adjusts the level of the audio component according to the energy transfer indication to accurately represent audio in multiple acoustic environments, including energy transfer and attenuation indications of the transfer area, reducing computational complexity.
It realizes more natural and accurate audio perception in multi-room scenes, reduces computing resource requirements, and improves audio experience quality and rendering efficiency.
Smart Images

Figure CN120323039A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus and method for rendering audio signals and, in particular but not exclusively, to an apparatus and method for rendering audio for a multi-room scenario as part of, for example, an extended reality experience. Background Art
[0002] In recent years, the variety and scope of experiences based on audiovisual content have increased significantly with the continuous development and introduction of new services and ways of exploiting and consuming such content. In particular, many spatial and interactive services, applications, and experiences are being developed to give users a more engaging and immersive experience.
[0003] An example of such an application is extended reality (XR), which is a common term encompassing virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications that are rapidly becoming mainstream, with multiple solutions targeting the consumer market. Multiple standardization bodies are also developing multiple standards. Such standardization activities are actively developing standards for various aspects of VR / AR / MR systems, including, for example, streaming, broadcasting, rendering, etc.
[0004] VR applications tend to provide a user experience corresponding to the user being in a different world / environment / scenario, while AR (including mixed reality MR) applications tend to provide a user experience corresponding to the user being in the current environment but with additional information or virtual objects or information added. Thus, VR applications tend to provide a fully immersive synthetically generated world / scenario, while AR applications tend to provide a partially synthetic world / scenario that is superimposed on the real scenario in which the user physically exists. However, these terms are often used interchangeably and have a high degree of overlap. Hereinafter, the term extended reality / XR will be used to denote both virtual reality and augmented / mixed reality.
[0005] As an example, a service that is becoming increasingly popular is to provide images and audio in such a way that the user can actively and dynamically interact with the system to change the parameters of the rendering, such that this will adapt to the movement and changes in the position and orientation of the user. A feature that is very appealing in many applications is the ability to change the effective viewing position and viewing direction of the viewer, for example, allowing the viewer to move and "look around" within the presented scenario.
[0006] Such a feature can specifically allow a virtual reality experience to be provided to the user. This can allow the user to move (relatively) freely within the virtual scene and dynamically change his position and the position he is looking at. Generally, such virtual reality applications are based on a three-dimensional model of the scene, where the model is dynamically evaluated to provide a particular requested view. This approach is well known from, for example, game applications for computers and consoles, such as in the category of first-person shooters.
[0007] Especially for virtual reality applications, it is also desirable that the presented images are three-dimensional images that are typically presented using a stereoscopic display. In fact, in order to optimize the viewer's immersion, users generally prefer to experience the presented scene as a three-dimensional scene. In fact, a virtual reality experience should preferably allow the user to select his / her own position, viewpoint, and moment relative to the virtual world.
[0008] In addition to visual rendering, most XR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience, in which the audio source is perceived to arrive from a position corresponding to the position of the corresponding object in the visual scene. Therefore, the audio and video scenes are preferably perceived as consistent, and both provide a complete spatial experience.
[0009] For example, many immersive experiences are provided through a virtual audio scene, which is generated by headphone reproduction using binaural audio rendering techniques. In many cases, such headphone reproduction can be based on head tracking, so that the rendering can be performed in response to the user's head movement. This greatly increases the immersion.
[0010] An important feature of many applications is how to generate and / or distribute audio that can provide a natural and realistic perception of the audio scene. For example, when generating audio for a virtual reality application, it is important not only to generate the desired audio sources, but also to generate these audio sources to provide a realistic perception of the audio environment, including damping, reflection, coloring, etc.
[0011] For room / environment acoustics, the reflection of sound waves from walls, floors, ceilings, objects, etc. results in a delayed and attenuated (usually frequency-dependent) version of the sound source signal arriving at the listener (i.e., the user of the XR system) via different paths. The combined effect can be modeled by an impulse response, which can be referred to as the room impulse response (RIR).
[0012] As Figure 1 illustrated, the RIR typically includes a direct sound that depends on the distance from the sound source to the listener, followed by a reflected portion that characterizes the acoustic properties of the room. The size and shape of the room, the positions of the sound source and the listener in the room, and the reflective properties of the surfaces of the room all play a role in the characteristics of this reverberant portion.
[0013] The reflective portion can be decomposed into two time regions that typically overlap. The first region contains the so-called early reflections, which represent the isolated reflections of the sound source on the walls or obstacles inside the room before reaching the listener. As the time lag / (propagation) delay increases, the number of reflections present in a fixed time interval increases, and the paths can include secondary or higher-order reflections (e.g., the reflection can be off several walls or two walls and the ceiling, etc.).
[0014] The second region, known as the reverberation part, is the part where the density of these reflections increases to the point where they can no longer be isolated by the human brain. This region is commonly referred to as diffuse reverberation, late reverberation, or reverberation tail, or simply reverberation.
[0015] The RIR contains cues that give information to the auditory system about the distance of the source and the size and acoustic properties of the room. The energy of the reverberation part relative to the energy of the anechoic part largely determines the perceived distance of the sound source. The level and delay of the earliest reflections can provide cues about how close the sound source is to the walls, and the evaluation of specific walls, floors, or ceilings can be enhanced by anthropometric filtering.
[0016] (Early) The density of the reflections contributes to the perceived size of the room. The time it takes for the energy level of the reflections to drop by 60 dB (indicated by the reverberation time T 60 is a frequently used measure of how quickly the reflections dissipate in the room. The reverberation time provides information about the acoustic properties of the room, such as specifically whether the walls are very reflective (e.g., a bathroom) or there is a lot of sound absorption (e.g., a bedroom with furniture, carpets, and curtains).
[0017] In addition, when the RIR is part of a binaural room impulse response (BRIR) (since the RIR is filtered by the head, ears, and shoulders; i.e., the head-related impulse response (HRIR)), the RIR may depend on the anthropometric properties of the user.
[0018] Since the reflections in the late reverberation cannot be distinguished and isolated by the listener, they are often simulated and parametrically represented using, for example, a parametric reverberator using a feedback delay network, as in the well-known Jot reverberator.
[0019] For early reflections, the incident direction and distance-related delays are important cues for humans to extract information about the relative positions of the room and the sound source. Therefore, the simulation of early reflections must be more distinct than that of late reverberation. In effective acoustic rendering algorithms, early reflections are thus simulated differently and separately from late reverberation. A well-known method for early reflections is to mirror the sound source in each of the room boundaries to generate virtual sound sources representing the reflections.
[0020] For early reflections, the position of the user and / or the sound source relative to the boundaries (walls, ceiling, floor) of the room is relevant, while for late reverberation, the acoustic response of the room is diffuse and thus tends to be uniform throughout the room. This allows the simulation of late reverberation to generally be computationally more efficient than that of early reflections.
[0021] Two main properties of late reverberation are the slope and amplitude of the impulse response for times above a given threshold. These properties tend to strongly depend on frequency in natural rooms. Parameters characterizing these properties are often used to describe reverberation.
[0022] Figure 2 Examples of parameters characterizing reverberation are illustrated in the figure. Examples of parameters traditionally used to indicate the slope and amplitude of the impulse response corresponding to diffuse reverberation include the known T 60 value and the reverberation level / energy. More recently, other indications of the amplitude level have been proposed, such as a parameter specifically indicating the ratio between the diffuse reverberation energy and the total emitted source energy.
[0023] Specifically, the diffuse-to-source ratio DSR can be used to express the amount of diffuse reverberation energy or level of a source received by a user as a ratio of the total emitted energy of that source. The DSR can represent the ratio between the emitted source energy and the diffuse reverberation properties, such as specifically the energy or (initial) level of the diffuse reverberation signal:
[0024]
[0025] Hereafter, this will be referred to as the DSR (diffuse-to-source ratio).
[0026] Such known methods tend to provide an effective description of audio propagation in a room and tend to result in the rendering of audio for which the room in which the listener (virtually) exists is perceived as natural.
[0027] However, while conventional methods for representing and rendering sound in a room or individual acoustic environment can provide a suitable perception in many embodiments, they tend not to be suitable for all possible scenarios. In particular, for audio scenes that may include different acoustic environments / regions / rooms, the audio signals generated using, for example, the reverberation methods described may not result in an optimal experience or perception. This can typically lead to situations where audio from other rooms is not adequately or accurately represented by the rendered audio, resulting in a perception that may not fully reflect the acoustic situation and scene.
[0028] In fact, typically, reverberation is modeled for a listener within a room, taking into account the properties of the room. When the listener is outside the room or in a different room, the reverberator can be turned off or reconfigured for the properties of the other room. Even when multiple reverberators can run in parallel, the output of the reverberator is typically intended as a diffuse binaural (or multi-speaker) signal to be presented to the listener within the room. However, such methods tend to result in the generation of audio that is often not perceived as an accurate representation of the actual environment. This can, for example, lead to a perceived disconnect or even conflict between the visual perception of a scene and the associated audio being rendered.
[0029] Thus, while typical methods for rendering audio may be suitable for rendering audio of an environment in many embodiments, they tend to be sub-optimal in some cases, including especially when rendering audio for a scene that includes different acoustic rooms or environments.
[0030] Examples of methods for rendering audio representing sounds from sources in different rooms are disclosed in the following documents: Dirk et al.: "Hybrid Method for Room Acoustic Simulation in Realtime" (Volume 7, September 2, 2007, pages 4521 - 4526, XP093046083, ISBN: 978 - 1 - 61567 - 707 - 8); Dirk "Physically Based Real - Time Auralization Of Interactive Virtual Environments" (February 4, 2011, XP055593422, Berlin); and Stavrakis Efstathios et al.: "Topological Sound Propagation with Reverberation Graphs" (Acustica United With Acta Acustica, Volume 94, Number 6, November 2008, pages 921 - 932, XP093046042, ISSN: 1610 - 1928, DOI: 10.3813 / AAA, 918109).
[0031] In particular, methods for representing and rendering audio from other acoustic environments in one acoustic environment tend to be sub - optimal and / or relatively impractical, including perhaps requiring excessive computational resources or being relatively complex. Additionally, in applications, representing audio of multiple acoustic environments (especially multi - room scenes) tends to be sub - optimal in terms of not providing easy - to - use and low - data - rate information that allows representing and rendering multiple acoustic environments.
[0032] Accordingly, improved methods for rendering audio of a scene would be advantageous. In particular, methods that allow for improved operation, increased flexibility, reduced complexity, facilitated implementation, improved audio experience, improved audio quality, reduced computational burden, improved representation of multiple acoustic environments, simplified rendering, improved audio from multiple acoustic environments, improved performance of virtual / hybrid / augmented reality applications, increased processing flexibility, improved representation and rendering of audio and audio properties of multiple rooms or other acoustic environments, more natural sounding audio presentation, improved audio rendering of multi - room scenes, and / or improved performance and / or operation would be advantageous. SUMMARY OF THE INVENTION
[0033] Accordingly, the present invention seeks to alleviate, mitigate or eliminate one or more of the above disadvantages, preferably individually or in any combination.
[0034] An audio device, comprising: a first receiver arranged to receive audio data of an audio source for a scene comprising a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries; a second receiver arranged to receive metadata for the audio data, the metadata comprising a position indication of at least a first transmission region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region of the first acoustic attenuation boundary having a lower attenuation than an average attenuation of the first acoustic attenuation boundary outside the transmission region; a renderer arranged to render an audio signal for a listening position in the first acoustic environment, the rendering comprising generating a first audio component by rendering the audio data of a first audio source for the second acoustic environment and adjusting a level of the first audio component according to an energy transfer indication (wherein the energy transfer indication is for the first transmission region), the energy transfer indication indicating a proportion of energy at the first transmission region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transmission region, and wherein the energy transfer indication is provided in the metadata or provided by / through a process performed elsewhere where computational complexity is available.
[0035] The method may allow for generating an audio signal that provides an improved user experience for an audio scene having multiple acoustic environments and often provides a more realistic and natural sounding audio experience. The method may allow for improved audio rendering for, e.g., a multi-room scenario. A more natural and / or accurate audio perception of the scene may be achieved in many cases.
[0036] The method may provide improved and / or facilitated rendering of audio representing an audio source in other acoustic environments or rooms. Rendering of the audio signal may often be achieved with reduced complexity and reduced computational resource requirements.
[0037] The method may provide improved, increased and / or facilitated flexibility and / or adaptation in processing and / or rendering the audio.
[0038] The method may also allow for improved and / or facilitated representation of multi-acoustic environment sound propagation data or properties. It may provide improved and / or facilitated representation of the sound propagation characteristics of a transmission region (such as a portal) in an acoustic attenuation boundary.
[0039] In many embodiments and scenarios, the method can provide an effective and low-complexity approach for accurately representing the acoustic properties of a transmission region in an acoustic attenuation boundary and for determining and rendering appropriate audio propagating from another room into a given room via such a transmission region.
[0040] An energy transfer indication indicating the energy transfer from a source to the transmission region is equivalent to an energy attenuation indication indicating the attenuation of the energy of the source signal / audio at the transmission region. An increased attenuation indicates a decreased proportion of the audio energy reaching the transmission region from the source, which corresponds to a decreased energy transfer. A decreased attenuation indicates an increased proportion of the audio energy reaching the transmission region from the source, which corresponds to an increased energy transfer. Thus, the terms energy attenuation and energy transfer can be used interchangeably, and it should be understood that one is a monotonically decreasing function of the other.
[0041] The audio energy (or just energy) can be specifically represented by a level, amplitude, power, or time-averaged energy metric.
[0042] The acoustic attenuation boundary can attenuate the sound propagation from one acoustic environment to another through the acoustic attenuation boundary. In many embodiments and scenarios, the attenuation of the acoustic attenuation boundary outside the transmission region can be not less than 3 dB, 6 dB, 10 dB, or even 20 dB. In many embodiments, the attenuation of the transmission region in the acoustic attenuation boundary can be lower than the (average) attenuation of the acoustic attenuation boundary outside the (one or more) transmission regions by not less than 3 dB, 6 dB, 10 dB, or even 20 dB.
[0043] The first acoustic environment and the second acoustic environment are different acoustic environments. The first audio source can be an audio source of the second acoustic environment and can be, for example, an audio source corresponding to a diffuse reverberant sound or a point source. The energy transfer indication can be frequency-dependent. The energy transfer indication can be a nominal energy transfer indication. The reference position can be a normalized / standardized / predetermined position for the first transmission region.
[0044] In many embodiments, the renderer can be arranged to determine the energy transfer indication based on the position indication. In many embodiments, the renderer can be arranged to render the audio signal based on the position indication.
[0045] According to an optional feature of the present invention, the renderer is arranged to adjust the level of the first audio component in response to the position of the first audio source relative to the reference position.
[0046] This can provide improved performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help provide an improved user experience. In many embodiments, it can allow for a low-complexity but accurate determination of the energy level at a first transfer region from a first audio source.
[0047] In some embodiments, the energy transfer attenuation for a first pair of transfer regions indicates the proportion of audio energy incident on a second transfer region that propagates to leave the first transfer region (into a first acoustic environment).
[0048] According to an optional feature of the present invention, the renderer is arranged to adjust the level of the first audio component based on the difference between a reference distance from the reference position to the first transfer region and the distance from the position of the first audio source to the first transfer region.
[0049] This can provide improved performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help provide an improved user experience. In many embodiments, it can allow for a low-complexity but accurate determination of the energy level at a first transfer region from a first audio source.
[0050] According to an optional feature of the present invention, the renderer is arranged to adjust the level of the first audio component based on the difference between the direction from the reference position to the first transfer region and the direction from the first audio source to the first transfer region.
[0051] This can provide improved performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help provide an improved user experience. In many embodiments, it can allow for a low-complexity but accurate determination of the energy level at a first transfer region from a first audio source.
[0052] According to an optional feature of the present invention, the metadata further includes data describing the directivity of the sound radiation from the first audio source, and the renderer is arranged to adjust the level of the first audio component based on the directivity.
[0053] This can provide improved performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help provide an improved user experience. In many embodiments, it can allow for a low-complexity but accurate determination of the energy level at a first transfer region from a first audio source.
[0054] In some embodiments, the renderer is arranged to: generate a combined energy transfer attenuation by combining the energy transfer attenuation for a first pair of transfer regions and the energy transfer attenuation for a second pair of transfer regions, the second pair of transfer regions including a third transfer region at a boundary of a third acoustic environment and the second transfer region; and generate a first audio component by rendering a second audio source of the third acoustic environment according to the combined energy transfer attenuation.
[0055] According to an optional feature of the present invention, the renderer is arranged to scale the level of the first audio component according to a relative directivity gain of the first audio source in a direction from the first audio source to the first transfer region, the relative directivity gain indicating a gain relative to an omnidirectional source.
[0056] This can provide improved performance and / or facilitated implementation in many cases. When rendering audio for a multi-acoustic environment scenario, it can help provide an improved user experience. In many embodiments, it can allow for a low-complexity but accurate determination of the energy level at the first transfer region from the first audio source.
[0057] According to an optional feature of the present invention, the first audio source represents audio arriving at the second acoustic environment from a third acoustic environment via a second transfer region of a second boundary that separates the third acoustic environment from the second acoustic environment.
[0058] This can provide improved performance and / or facilitated implementation in many cases. When rendering audio for a multi-acoustic environment scenario, it can help provide an improved user experience. The method can particularly allow for an effective representation of the audio propagation through multiple intermediate acoustic environments.
[0059] According to an optional feature of the present invention, the metadata further includes: energy transfer parameters, each energy transfer parameter indicating an energy attenuation between a pair of transfer regions, the energy attenuation for a pair of transfer regions indicating the proportion of audio energy propagating from one transfer region of the pair of transfer regions to the other transfer region of the pair of transfer regions; and the renderer is arranged to render a second audio component of the audio signal by rendering a second audio source of the third acoustic environment according to the energy attenuation for a pair of transfer regions, the pair of transfer regions including a transfer region at an acoustic attenuation boundary of the first acoustic environment and a transfer region at a second acoustic attenuation boundary, the second acoustic attenuation boundary being a boundary of the second acoustic environment.
[0060] This can provide particularly advantageous performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help to provide an improved user experience. The method can in particular allow for improved rendering of audio that reaches the listener via a path including multiple acoustic environments and transfer regions.
[0061] According to an optional feature of the invention, the metadata includes a coupling coefficient for the first transfer region, and the renderer is arranged to render the first audio component as an audio source originating from a location near the first transfer region according to the coupling coefficient.
[0062] This can provide improved performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help to provide an improved user experience. The method can in particular allow for effective rendering of audio from other acoustic environments that reaches the listening acoustic environment via a coupling region (such as a window or the like).
[0063] According to an optional feature of the invention, the renderer is arranged to render the first audio component as a reverberant audio component of the first acoustic environment.
[0064] This can provide improved performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help to provide an improved user experience. The method is particularly advantageous for generating reverberant / diffuse / ambient sound reflection audio sources in other acoustic environments.
[0065] According to an optional feature of the invention, the renderer is arranged to generate a reverberant audio signal for the second environment to include a first component from the first audio source, the renderer is further arranged to determine an energy loss estimate as a proportion of the energy in the first audio source that reaches the first transfer region, the proportion of the energy being determined according to the energy transfer indication and the position of the first audio source relative to the reference position; and the renderer is arranged to reduce the level of the reverberant audio signal by an amount depending on the energy loss estimate.
[0066] This can provide particularly advantageous performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help to provide an improved user experience.
[0067] According to an optional feature of the invention, the energy transfer indication reflects the proportion of a sphere covered by the first transfer region, the sphere being centered at the reference position and having a radius corresponding to the distance from the reference position to the first transfer region.
[0068] This can provide improved performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help provide an improved user experience.
[0069] According to an optional feature of the present invention, the renderer is arranged to render the first audio component as a non-direct audio component.
[0070] This can provide improved performance and / or facilitated implementation in many scenarios. When rendering audio for a multi-acoustic environment scenario, it can help provide an improved user experience.
[0071] According to one aspect of the present invention, there is provided a method of rendering an audio signal, the method comprising: receiving audio data of an audio source for a scenario including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries; receiving metadata for the audio data, the metadata including: an indication of the position of at least a first transmission region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region of the first acoustic attenuation boundary having a lower attenuation than the average attenuation of the first acoustic attenuation boundary outside the transmission region; rendering an audio signal for a listening position in the first acoustic environment, the rendering including generating a first audio component by rendering the audio data of a first audio source for the second acoustic environment and adjusting the level of the first audio component according to an energy transfer indication, wherein the energy transfer indication is for the first transmission region, the energy transfer indication indicating the proportion of the energy at the first transmission region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transmission region, and wherein the energy transfer indication is provided in the metadata or provided by a process executed elsewhere where computational complexity is available.
[0072] An audio data signal may be provided, comprising: audio data of an audio source for a scenario including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries; metadata for the audio data, the metadata including: an indication of the position of at least a first transmission region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region of the first acoustic attenuation boundary having a lower attenuation than the average attenuation of the first acoustic attenuation boundary outside the transmission region; and an energy transfer indication for the first transmission region, the energy transfer indication indicating the proportion of the energy at the first transmission region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transmission region.
[0073] These and other aspects, features, and advantages of the present invention will become apparent from and will be elucidated with reference to the (one or more) embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Embodiments of the present invention will be described by way of example only with reference to the accompanying drawings, in which
[0075] Figure 1 an example of a room impulse response is illustrated;
[0076] Figure 2 an example of a room impulse response is illustrated;
[0077] Figure 3 an example of elements of a virtual reality system is illustrated;
[0078] Figure 4 an example of a scene with three rooms is illustrated;
[0079] Figure 5 an example of an audio device for generating an audio signal according to some embodiments of the present invention is illustrated;
[0080] Figure 6 an example of a scene with multiple rooms separated by a wall having a sound portal is illustrated;
[0081] Figure 7 an example of sound propagation from an audio source towards a wall having a sound portal is illustrated;
[0082] Figure 8 an example of a scene with multiple rooms separated by a wall having a sound portal is illustrated;
[0083] Figure 9 an example of sound propagation from an audio source towards a wall having a sound portal is illustrated;
[0084] Figure 10 an example of sound propagation from an audio source towards a wall having a sound portal is illustrated;
[0085] Figure 11 an example of a scene with multiple rooms separated by a wall having a sound portal is illustrated;
[0086] Figure 12 an example of a scene with multiple rooms separated by a wall having a sound portal is illustrated;
[0087] Figure 13 an example of a scene with multiple rooms separated by a wall having a sound portal is illustrated;
[0088] Figure 14An example of a scenario with multiple rooms separated by a wall with a sound portal is illustrated;
[0089] Figure 15 An example of a scenario with multiple rooms separated by a wall with a sound portal is illustrated;
[0090] Figure 16 An example of a scenario with multiple rooms separated by a wall with a sound portal is illustrated; and
[0091] Figure 17 An illustration of Figures 13 - 16 an example of a room connection diagram of an example;
[0092] Figure 18 Some elements of a possible arrangement of a processor for elements of an implementation device according to some embodiments of the present invention are illustrated. Detailed Description
[0093] The following description will focus on audio processing and rendering for extended reality applications, but it should be understood that the principles and concepts described can be used in many other applications and embodiments.
[0094] Virtual experiences that allow users to move around in a virtual world are becoming increasingly popular, and services are being developed to meet such needs.
[0095] In some systems, VR applications can be provided locally to viewers by, for example, a stand-alone device that does not use or even have any access to any remote VR data or processing. For example, a device such as a game console may include storage for storing scene data, an input for receiving / generating viewer poses, and a processor for generating corresponding images based on the scene data.
[0096] In other systems, VR applications can be implemented and executed remotely from the viewer. For example, a device local to the user can detect / receive movement / pose data, which is sent to a remote device that processes the data to generate a viewer pose. Then, the remote device can generate a suitable view image and corresponding audio signal for the user pose based on scene data describing the scene. Then, the view image and corresponding audio signal are sent to a device local to the viewer, where they are presented. For example, the remote device can directly generate a video stream (usually a stereoscopic / 3D video stream) and corresponding audio stream, which are directly presented by the local device. Thus, in such an example, the local device may not perform any VR processing other than sending movement data and presenting the received video data.
[0097] In many systems, functionality can be distributed across local and remote devices. For example, a local device can process received input and sensor data to generate a user pose that is continuously sent to a remote VR device. The remote VR device can then generate a corresponding view image and corresponding audio signals and send these to the local device for presentation. In other systems, instead of directly generating the view image and corresponding audio signals, the remote VR device can select relevant scene data and send it to the local device, which can then generate the presented view image and corresponding audio signals. For example, the remote VR device can identify a nearest capture point and extract the corresponding scene data (e.g., a set of object sources and their location metadata) and send it to the local device. The local device can then process the received scene data to generate images and audio signals for a particular current user pose. The user pose will typically correspond to a head pose, and a reference to the user pose can generally be equivalently considered to correspond to a reference to the head pose.
[0098] In many applications, especially for broadcast services, the source can send or stream scene data in the form of an image (including video) and audio representation of the scene independent of the user pose. For example, signals and metadata corresponding to audio sources within the scope of a certain virtual room can be sent or streamed to multiple clients. Individual clients can then locally synthesize the audio signals corresponding to the current user pose. Similarly, the source can send a general description of the audio environment, which includes a description of the audio sources in the environment and the acoustic characteristics of the environment. An audio representation can then be generated locally and presented to the user, for example using binaural rendering and processing.
[0099] Figure 3 An example of such a VR system is illustrated, where a remote VR client device 301 is connected to a VR server 303 via a network 305 such as the Internet. The server 303 can be arranged to support potentially a large number of client devices 301 simultaneously.
[0100] The VR server 303 can support a broadcast experience, for example, by sending an image signal that includes an image representation in the form of image data that the client device can use to locally synthesize a view image corresponding to an appropriate user pose (a pose refers to a position and / or orientation). Similarly, the VR server 303 can send an audio representation of the scene, allowing the audio to be locally synthesized for the user pose. Specifically, as the user moves around in the virtual environment, the images and audio synthesized and presented to the user are updated to reflect the user's current (virtual) position and orientation in the (virtual) environment.
[0101] Thus, in many applications such as Figure 3In the application), it is possible to expect to model a scene and generate efficient image and audio representations that can be efficiently included in a data signal, which can then be sent or streamed to various devices that can locally synthesize views and audio for poses different from the captured pose.
[0102] In some embodiments, a model representing a scene can be stored locally, for example, and can be used locally to synthesize appropriate images and audio. For example, an audio model of a room can include an indication of the nature of audio sources that can be heard in the room and the acoustic properties of the room. The model data can then be used to synthesize appropriate audio for a specific location.
[0103] In many cases, a scene can include multiple different acoustic environments or regions having different acoustic properties and specifically different reverberation properties. Specifically, a scene can include or be divided into different acoustic environments / regions, each acoustic environment / region having uniform reverberation, but the reverberation is different between them. For all positions within an acoustic environment / region, the reverberation component of the audio received at these positions can be uniform and specifically can be substantially the same (potentially except for gain differences). An acoustic environment / region can be a set of positions where the reverberation component of the audio is uniform. An acoustic environment / region can be a set of positions for which the reverberation component of the audio propagation impulse response of audio sources in the acoustic environment is uniform. Specifically, an acoustic environment / region can be a set of positions for which, except for possible gain differences, the reverberation component of the audio propagation impulse response of audio sources in the acoustic environment has the same frequency-dependent slope and / or amplitude properties. Specifically, an acoustic environment / region can be a set of positions for which, except for possible gain differences, the reverberation component of the audio propagation impulse response of audio sources in the acoustic environment is the same.
[0104] An acoustic environment / region can generally be a set of positions (usually a 2D or 3D region) having the same rendering reverberation parameters. The reverberation parameters used to render the reverberation component can be the same for all positions within the acoustic environment / region. Specifically, the same reverberation decay parameter (e.g., T 60 ) or diffusion-to-source ratio DSR can be applied to all positions within the acoustic environment / region.
[0105] The impulse response can be different between different positions in a room / acoustic environment / region due to the "noisy" characteristics resulting from many different orders of reflections that cause reverberation. However, even in such cases, the frequency-dependent slope and / or amplitude properties can be the same (except for possible gain differences), especially when represented by, for example, reverberation time (T60) or reverberation coloring.
[0106] In many cases, an acoustic environment can be separated by an acoustic attenuation boundary. In fact, in many cases, different acoustic environments can be determined by the presence of an acoustic attenuation boundary. An acoustic attenuation boundary can divide a region into different acoustic environments, and different acoustic environments can be formed by the presence of one or more acoustic attenuation boundaries. Two acoustic environments can be created by an acoustic attenuation boundary, where the two acoustic environments are on opposite sides of the acoustic attenuation boundary. Such an acoustic attenuation boundary can be formed, for example, by a wall or by any other structure that provides acoustic attenuation for dividing a space into multiple acoustic environments.
[0107] An acoustic environment / region can also be referred to as an acoustic room or simply a room. A room can be considered as the environment / region as described above.
[0108] In many embodiments, a scenario can be provided where the acoustic rooms correspond to different virtual or real rooms in which a user can move (e.g., virtually). Figure 4 An example of a scenario with three rooms A, B, and C is illustrated. In this example, a user can move between the three rooms or outside any of the rooms through doorways and openings.
[0109] In order for a room to have significant reverberation properties, it tends to represent a spatial region sufficiently bounded by geometric surfaces with fully or partially reflective properties such that most of the reflections in the room remain reflected back into the region to generate a diffuse field of reflections that do not have significant directional properties in the region. The geometric surfaces do not need to be aligned with any visual elements.
[0110] Audio rendering aimed at providing natural and realistic effects to a listener typically includes the rendering of an acoustic scene. For many environments, this includes the representation and rendering of the diffuse reverberation present in the environment (such as in the room where the listener is located). It has been found that the presentation and representation of such diffuse reverberation have a significant impact on the perception of the environment, such as having a significant impact on whether the audio is perceived as representing a natural and realistic environment.
[0111] In cases where the scenario includes multiple rooms, the method typically is to render the audio and reverberation only for the room in which the listener is present and to ignore any audio from other rooms. However, this often results in the audio experience not being perceived as optimal and often does not provide the best natural experience, especially when the user switches between rooms. Although some applications have been implemented to include the rendering of audio from adjacent rooms, it has been found that they are sub - optimal. In some embodiments, the audio from other rooms can have a substantial impact on the perceived audio scene. In particular, the audio from other rooms can, in many cases, provide a significant contribution to the reverberation or diffuse (background) sound in a room, and the sub - optimal rendering of such audio can lead to a degraded user experience.
[0112] In the following, advantageous methods for rendering an audio scene including a plurality of rooms will be described.
[0113] Figure 5 An example of an audio device arranged to render an audio scene is illustrated. The audio device may receive audio data describing audio and audio sources in a scene (such as, for example, Figure 4 the scene). Based on the received audio data, the audio device may render an audio signal representing the scene for a given listening position. The rendered audio may include contributions from audio generated in the room in which the listener is present and contributions from other neighboring (and typically adjacent) rooms.
[0114] The audio device is arranged to generate an audio output signal representing the audio in the scene. Specifically, the audio device may generate an audio representing that perceived by a user moving around in a scene having a plurality of audio sources and given acoustic properties. Each audio source is represented by an audio signal representing the sound from the audio source and metadata that may describe characteristics of the audio source (such as providing a level indication for the audio signal). In addition, metadata is provided to characterize the scene.
[0115] In the example section of the audio device, the renderer is arranged to receive the audio data and metadata of the scene and to render an audio representing at least part of the environment based on the received data.
[0116] Figure 5 The audio device of includes a first receiver 501 arranged to receive the audio data of the audio sources in the scene and thus may receive the audio data of a plurality of acoustic environments / rooms divided by an acoustic attenuation boundary. The audio data may include audio data describing a plurality of audio signals from different audio sources in the scene. Typically, audio data reflecting the sound to be rendered from those audio (point) sources may be provided to a plurality of, for example, point sources. In some embodiments, audio data may also be provided for more diffuse audio sources, such as background or ambient sound sources, or sound sources having a spatial extent.
[0117] The audio device includes a second receiver 503 arranged to receive the metadata for the audio data and specifically may receive the metadata of the audio sources represented by the audio data. As will be described in more detail later, the metadata may contain various information about the scene, including specifically related to different acoustic environments and the boundaries therebetween.
[0118] The apparatus further includes a location circuitry 505 arranged to determine a listening position in a scene. The listening position typically reflects the (virtual) position of the user in the scene. For example, the location circuitry 505 may be coupled to a user tracking device such as a VR headset, an eye tracking device, a motion capture camera, etc., and may thereby receive user movement (including or possibly limited to head movement and / or eye movement) data. The location circuitry 505 may continuously determine the current listening position based on this data.
[0119] The listening position may alternatively be represented or augmented by a controller input, by which the user may move or teleport the listening position in the scene.
[0120] It should be understood that many methods and techniques are known and used to determine a listening position in a scene for various applications, and any suitable method may be used without departing from the present invention.
[0121] The audio apparatus includes a renderer 507 arranged to generate an audio output signal representing the audio of the scene at the listening position. Generally, the audio signal may be generated to include audio components for a series of different audio sources in the scene. For example, a point audio source in the same room may be rendered as a point audio source with a direct acoustic path, a reverberation component may be rendered or generated, etc.
[0122] In the following, a method will be described in which rendering the audio signal includes an audio signal / component representing audio from other rooms than the room including the listening position. The description will focus on the generation of this audio component, but it should be understood that the rendered audio signal presented to the user may include many other components and audio sources. These may be generated and processed according to any suitable algorithm or method, and it should be understood that a person skilled in the art will know a large number of such methods.
[0123] The renderer (507) is arranged to render an audio signal at a listening position in an acoustic environment, hereinafter referred to as a first acoustic environment, based on received audio data and metadata. The rendering also causes it to include at least one audio component generated by rendering audio sources of another acoustic environment, i.e., the audio signal generated for the listening position in the first acoustic environment is generated to include a component from an audio source in a second acoustic environment (different from the first acoustic environment). Specifically, in the case / embodiment where the different acoustic environments are different rooms, the rendering of the audio signal for the listening position includes rendering the contribution from audio sources in other rooms.
[0124] In many cases, the rendering of audio and audio sources for acoustic environments / rooms other than the first acoustic environment can be at least partly as diffuse or reverberant audio. In some cases, the rendering can be the same reverberant diffuse audio for all positions in the first acoustic environment, i.e., the audio can be substantially independent of the exact listening position in the first acoustic environment. In such cases, rendering audio for a listening position can be simply achieved by rendering diffuse audio, which does not specifically depend on the listening position.
[0125] It should be understood that in many cases, audio data and metadata can be received as part of the same bitstream, and the first and second receivers 501, 503 can be implemented by the same functionality, and in fact the same receiver functionality can implement both the first and second receivers. Figure 5 The audio device can specifically correspond to Figure 3 the client device 301 or a part thereof and can receive audio data and metadata in a single bitstream sent from the server 303.
[0126] Metadata can describe the acoustic elements and properties of a scene and is specifically for different acoustic environments. For example, it can include data describing room dimensions, the acoustic properties of the room (e.g., T60, DSR, material properties), the relationship between rooms, etc. Metadata can also describe the position and orientation of some or all of the audio sources.
[0127] Metadata includes data reflecting how sound propagates or spreads between different acoustic environments (such as between different rooms). It can specifically include metadata related to the transmission regions of the acoustic attenuation boundaries.
[0128] Specifically, it can include data describing at least one and typically more or even all of the acoustic attenuation boundaries for the scene and one or more transmission regions. A transmission region can specifically be a region where the acoustic transmission level of sound from one acoustic environment to an adjacent acoustic environment (specifically from one room to an adjacent room) exceeds a threshold. Specifically, a transmission region can be a region (usually a zone) of the acoustic attenuation boundary between two acoustic environments, for which the attenuation through the boundary / attenuation across the boundary is less than a given threshold, while it may be higher outside this region. A transmission region is a region of the acoustic attenuation boundary that has a lower attenuation than the average attenuation of the acoustic attenuation boundary outside the transmission region.
[0129] Thus, the transmission region can define a region at the boundary between two acoustic environments / rooms, for which acoustic propagation / transmission / transparency / coupling exceeds a threshold. The portions of the boundary not included in the transmission region can have acoustic propagation / transmission / transparency / coupling below the threshold. Accordingly, the transmission region can define a region at the boundary between two acoustic environments / rooms where the acoustic attenuation is below a threshold. The portions of the boundary not included in the transmission region can have acoustic attenuation above the threshold. The transmission region can also be referred to as a portal (in the acoustic attenuation boundary).
[0130] The portal is associated with at least two acoustic environments, such as specifically two rooms. It can provide an acoustic link between the two acoustic environments / rooms. In addition to indicating the link between the acoustic environments, it can also include or refer to the acoustic properties of the link.
[0131] The following description will focus on an example where the acoustic environments are rooms and the acoustic attenuation boundary is the wall of the room. However, it should be understood that this is merely exemplary, and the acoustic environments can be other acoustic environments at least partially separated by the acoustic attenuation boundary.
[0132] Thus, the transmission region can indicate a region of the boundary with relatively high acoustic transparency, while the acoustic transparency outside this region may be very low. The transmission region can correspond, for example, to an opening in the boundary. For example, for a conventional room formed by an acoustic attenuation boundary in the form of a wall, the transmission region can correspond, for example, to a doorway, a window opening, or a hole in the wall separating two rooms, etc.
[0133] The transmission region can be a three-dimensional or two-dimensional region. In many embodiments, the boundary between the rooms is represented as a two-dimensional object (e.g., a wall considered to have no thickness), and in such cases, the transmission region can be a two-dimensional shape or area of the boundary with low acoustic attenuation.
[0134] Acoustic transparency can be expressed proportionally. Complete transparency means there is no acoustic suppression (e.g., an open doorway). Partial transparency can introduce attenuation to the energy transitioning from one room to another (e.g., thick curtains in a doorway, or a single-pane window). At the other end of the scale is room-separating material that does not allow any (significant) acoustic leakage between rooms (e.g., a thick concrete wall).
[0135] Thus, in some embodiments, the method can provide acoustic link metadata (in the form of transfer regions) describing how two rooms are acoustically linked. This data can be exported locally or can be obtained, for example, from a received bitstream. The data can be provided manually by the content author or can be derived indirectly from a geometric description of the rooms (e.g., box, grid, voxelized representation, etc.), including acoustic properties such as material properties indicating how much audio energy passes through the material or couples into the vibrations of the material, resulting in an acoustic link from one room to another. In many cases, the transfer regions can be considered to indicate room leakage, where acoustic energy can be exchanged between the two rooms.
[0136] Figure 6 An example of a scenario where the described method can be applied is shown. Figure 6 An example of a scenario including a building with multiple rooms A - H is shown. In the building, some audio sources are present in different rooms (indicated by circles 601). In this case, Figure 5 the audio device can determine the listener position 603 in room E and render audio for that listening position. Rendering the audio signal includes different audio components from other rooms. Sound from such sources can specifically reach room E through multiple transfer regions 605, which, for example, correspond to (open) doors or windows in the walls forming the rooms.
[0137] Rendering of audio sources within the same room as the listening position is well - established, and many algorithms are known and can be used by the renderer without departing from the invention. Rendering of audio from audio sources located in other rooms can be performed, for example, by representing the audio from other rooms as audio sources that have no position (especially for diffuse reverberation) or that have, for example, been assigned a position near the portal. For example, the sound component from audio source 605 can be considered to reach a given first room E including the listening position 603 via a first portal 4 in the first room E. The signal level reduction resulting from propagation to the first portal 4 can be determined and used to determine the level of the corresponding sound component at the first portal. The audio source can then be rendered as an audio signal component having a level corresponding to the determined level at the first portal E. As mentioned, in some embodiments, the sound source can be rendered as a spatially - defined audio source, for example, even as a point source located at the position of the first portal, or as a source having a spatial extent similar to and near the portal. In other embodiments, the sound component can be considered diffuse sound and can be rendered as diffuse reverberation in the first room E.
[0138] Such a method can be used, for example, to render an audio signal component audio / sound from room C as heard from a listening position in room E. For example, if the resulting signal level after propagation through multiple portals is determined, it can also be used to render an audio source that is more than one room away, such as an audio source from room A, for example.
[0139] Rendering an audio source as a point source, a spatially extended source, a distributed source, or a diffuse source with a given signal level is known in the art and will not be described in detail for the sake of brevity.
[0140] In some embodiments, the metadata can specifically include data describing the position of at least one of the transmission regions of the acoustic attenuation boundary. The position can be described, for example, relative to a room, for example, or as a relative position on the acoustic attenuation boundary where the transmission region is formed (which can be defined, for example, by a position within the room).
[0141] In many embodiments, the metadata can describe the scene topologically and / or geometrically, including describing the rooms, acoustic attenuation boundaries, and transmission regions among these. In some embodiments, a geometric description can be included that describes, for example, the size of all rooms (forming the acoustic environment), the extent and position of the walls (forming the acoustic attenuation boundaries), and the size, shape, and position of the portals (forming the transmission regions).
[0142] However, in other embodiments, the metadata can additionally or alternatively include a topological description of the scene. Such data can, for example, list the number of rooms and, for each room, provide some acoustic properties (such as BRIR or parameters describing reverberation). It can also define multiple portals / transmission regions, and for each transmission region, it can describe which two rooms the transmission region is connecting.
[0143] In many embodiments, the metadata includes an indication of energy transfer (which can also be referred to as a nominal energy transfer indication based on a nominal reference audio source) for at least a first transmission region formed in the acoustic attenuation boundary separating two acoustic environments / rooms. The nominal energy transfer indication indicates the proportion of the energy in an omnidirectional point audio source at a reference position that will propagate into the first transmission region, where the reference position is a relative position with respect to the first transmission region. Thus, the nominal energy transfer indication can indicate the amount of energy that will be radiated from the omnidirectional source at the reference position and will be at the portal / transmission region for which the nominal energy transfer indication is provided. Thus, the nominal energy transfer indication provides a description of the acoustic properties of the transmission region based on the reference omnidirectional source. The acoustic properties can specifically be an indication of the transfer of audio between two acoustic environments.
[0144] In some embodiments, the nominal energy transfer indication for the first transfer region can be an indication of the proportion of the energy arriving at the first transfer region, i.e., it can indicate the proportion of the energy incident on the first transfer region from the reference audio source. In some embodiments, the nominal energy transfer indication can be an indication of the proportion of the energy leaving the first transfer region, i.e., it can indicate the proportion of the energy leaving the first transfer region from the reference audio source and entering the first room. It should be understood that in the case where the first transfer region does not include any attenuation or have any other acoustic effects, such as if the first transfer region is an empty opening in the first acoustic attenuation boundary, such a measure can be the same. In other embodiments, the measure can be different, for example, due to the acoustic effects or attenuation of the first transfer region, such as if the first transfer region is formed of a material that may have some acoustic effects but allows some sound to propagate through. It should also be understood that such indications can be equivalent, i.e., an indication of the energy incident on the transfer region can equivalently be considered an indication of the energy leaving the transfer region, and vice versa. Generally, one value / property can be directly determined from another value / property by considering the acoustic effects of the transfer region (e.g., by compensating for the attenuation of the transfer region on sound).
[0145] The representation of the acoustic information of the transfer region using the nominal reference audio source as described can provide particularly advantageous operation in many embodiments. It can generally allow for a low-complexity and efficient (e.g., low data rate) description of the acoustic properties resulting from the presence of the transfer region in the acoustic environment that divides the acoustic environment. It can also be provided in a manner that allows for easy processing to provide data suitable for rendering the specific audio sources present in the scene.
[0146] This method is very advantageous for including the contribution of sources from one acoustic environment when rendering audio for another acoustic environment, where the environment is divided by an acoustic attenuation boundary including the transfer region.
[0147] In this method, the specific transfer region / portal geometry may not be described by metadata or used in rendering, but the transfer region / portal can be described by acoustic transfer properties, as expressed with reference to the reference audio source.
[0148] In many embodiments and for many transfer regions, the reference position can be within the second acoustic environment, but it should be understood that this is not required, and in fact the reference position can be outside the second acoustic environment. For example, the reference audio source position for a portal between two rooms can be within the room, but in some cases can also be outside the room.
[0149] As a specific example, the metadata for each portal / transfer region can include the following data / indications:
[0150] An indication of the normalized energy transfer from a reference source location to the portal. Thus, the portalFactor can be an example of a nominal energy transfer indication.
[0151] Optionally, it may also include one or more of the following:
[0152] Portal location (for distance and angle effects)
[0153] Portal orientation (e.g., a normal vector for angle effects)
[0154] An indication of the portal size (e.g., width and height, for determining angle effects)
[0155] An indication of which acoustic environments the portal is associated with
[0156] An indication of the acoustic environments separated by the acoustic attenuation boundary in which the portal is formed.
[0157] Thus, in some embodiments, the nominal energy transfer indication can be represented by a data field / value that can be referred to as the portalFactor. The portalFactor can indicate the proportion of an omnidirectional source that reaches the transfer region / portal, where the omnidirectional source is located at a reference position relative to the transfer region / portal. The position and orientation of the reference source relative to the transfer region are at a reference distance and a reference angle. Typically, the reference angle is advantageously selected to be substantially perpendicular to the orientation of the portal, but it can also be at a different angle (e.g., within the range of ±10°, 20°, 30°, or 45° relative to the direction perpendicular to the portal (or perpendicular to the acoustic attenuation boundary forming the portal)).
[0158] As Figure 7 Illustrated in two dimensions (while most embodiments will be in three dimensions), it shows a first transfer region 701 and a corresponding reference source 703. The audio energy radiating omnidirectionally from the reference source location causes the energy to spread over a sphere, and only a portion of the radiated energy is transferred through the portal to another acoustic environment. Thus, the proportion of the energy from a reference source at a given location (especially perpendicular to the plane extending for the transfer region) that reaches the transfer region can be determined (for negligible distance variations over the transfer region). The nominal energy transfer indication can specifically reflect the proportion of the sphere covered by the first transfer region, where the sphere is centered at the reference position and has a radius corresponding to the distance from the reference position to the transfer region. In many embodiments, the distance from the reference position to the transfer region varies only negligibly over the transfer region. In cases where the distance variation is significant over the transfer region, the maximum distance can often be advantageously used, although in other embodiments, for example, the minimum or average distance can be used.
[0159] In the following, examples for determining a nominal energy transfer indication / portalFactor representing such a value will be described.
[0160] The opening of the portal covers a certain angular fraction. It can be assumed that the portal is rectangular (or a rectangular equivalent can be derived based on surface area and aspect ratio), and can cover different width and height angles. The angles can be derived from a reference distance and the portal dimensions (width and height). From this, the ratio of the patch to the sphere surface can be derived, which is the portalFactor.
[0161] The width w of the portal gives the azimuth angle according to the following relationship:
[0162]
[0163] yielding:
[0164]
[0165] where r is the radius of the sphere. The radius is typically between the minimum distance from the reference source to the transfer region and the maximum distance from the reference source to the transfer region.
[0166] In many embodiments, the reference position is chosen to be perpendicular to the portal / transfer region and in the middle of its (rectangular equivalent) width and height. Under these conditions, the optimal radius is equal to the maximum distance, corresponding to the distance to any one of the four corner points on the (rectangular equivalent) of the transfer region.
[0167] Thus, the radius can be calculated as a function of the (rectangular equivalent) width and height and the reference distance of the reference source as:
[0168]
[0169] Similarly, the height h of the portal gives the elevation angle:
[0170]
[0171] Using these angles, the nominal energy transfer indication can be estimated as:
[0172]
[0173] Or more precisely, based on a surface integral:
[0174]
[0175] Metadata providing such a description of the transmission region can be highly suitable for representing the transmission region and can form an efficient basis for determining sound propagation through the transmission region for audio sources of a scene and specifically for audio sources that are in the same room as the transmission region and propagate into a neighboring room. For example, in Figure 8 the case of, a nominal energy transfer indication can be provided for the first transmission region 1 based on the reference position 801. Based on this nominal energy transfer indication, the sound energy arriving at the first transmission region 1 from the scene audio source 803 and propagating through this region into room A can be determined. Then, based on the propagated energy metric, determining the rendering of the audio signal for the listening position in room A includes a contribution from the audio source 803 in room B. Specifically, the renderer can generate an audio component based on the audio for the scene audio source 803 and adjust its level according to the determined propagated energy metric.
[0176] As will be described in more detail later, the renderer 203 can determine an energy reduction factor for a given transmission region / portal formed in an acoustic environment / portal that is formed in an acoustic attenuation boundary / wall that separates a first acoustic environment / room including the listening position from a second acoustic environment / room including the audio source generating the audio, and generate an audio signal for the listening position. For clarity and brevity, the following description will focus on the case of a building where audio from different rooms is rendered in other rooms and the corresponding terms will be used. However, it should be understood that these terms can be substituted for the alternative terms as indicated above.
[0177] The renderer 203 can, when rendering an audio signal component from an audio source in a source room via a (target) portal in a target room, proceed to determine the energy reduction factor for the target portal / room and can use the energy reduction factor to perform the rendering. The renderer 507 specifically adjusts the level of the rendered audio component to reflect the energy attenuation, and specifically, the higher the given energy attenuation, the lower the level of the corresponding rendered audio signal component.
[0178] Energy reduction factor F tgt can be applied to the source signal, for example. For example:
[0179]
[0180] where S in can be the input contribution of the source represented by the signal S src to the rendering algorithm (e.g., reverb, coupled source rendering).
[0181] In many embodiments, a renderer may render an immersive reverberation signal for an acoustic environment in which a listener is located, represented as in-room reverberation. Generally, all the energy emitted by a source in the room contributes to this reverberation. A nominal energy transfer indication may be used to determine how much energy reaches a transfer area of the room. These fractions of source energy may additionally or alternatively be used to reduce the source energy contributing to the in-room reverberation of the room.
[0182] In many cases, the reduction is obtained by subtracting a fraction of the source energy that contributes to the in-room reverberation. This may also depend on the material properties associated with the transfer area. That is, when the reflective properties of the transfer area are non-zero, the reduction of the source energy may be limited. For example, where F tgt indicates the total energy of the source that reaches (only) the transfer area of the room, the reduction of the in-room reverberation of the source may be determined as:
[0183]
[0184] where S rev is the input signal for the in-room reverberation, S src is the source signal, c sig2nrg is a conversion coefficient indicating the ratio between the emitted source energy and the signal energy, and c refl is the reflection coefficient associated with the transfer area. The coefficients and F tgt may be frequency-dependent.
[0185] Thus, in some embodiments, often when the listening position is in a second acoustic environment, the renderer may be arranged to render a diffuse audio signal component for the second acoustic environment (where there is an audio source). In such a case, the renderer 507 may be arranged to adjust the level of the diffuse audio signal component according to a nominal energy transfer indication. The renderer may determine an energy estimate (which may be relative) of the amount of energy reaching the transfer area from the audio source, and reduce the level of the diffuse audio signal component by a corresponding amount.
[0186] The renderer 507 may be arranged to adjust the level of an audio component generated for a given portal based on an audio source in another room, based on / according to the position of the audio source relative to a reference position for a nominal audio source. In many embodiments, it may be arranged to adjust the signal level of the audio component based on the difference in distance between the reference and actual audio source positions and the portal, based on the angular difference between the directions between these sources and the portal, and / or based on the directivity (gain) of the audio source (in the direction towards the portal).
[0187] For example, if the nominal energy transfer indication represents a portalFactor that indicates the fraction of source energy lost through the portal for an omnidirectional source at the reference source position, then the fraction of source energy lost for a target source at different positions in the room can be calculated as:
[0188] F tgt = portalFactor · G dist · G 角度 · G dir
[0189] where G dist compensates for the difference in distance, G 角度 compensates for the difference in angle, and G dir compensates for the directivity pattern. Some embodiments may use a subset of these compensations or may employ additional compensations.
[0190] In many embodiments, the renderer is arranged to adjust the level of the audio component of the audio source in the second room based on the difference between the reference distance from the reference position to the first transfer region and the source distance from the scene audio source to the first transfer region.
[0191] Reference may be made to Figure 9 the situation in which the reference position 901 is positioned perpendicular to the portal 903, while the scene audio source is located at a source position 905 different from the reference position.
[0192] The effect of distance is related to the physical phenomenon that a more distant source sounds quieter. To this end, the 1 / r law is typically used, meaning that the root mean square (RMS) amplitude level (i.e., non-energy) is inversely proportional to the distance (r). The reasoning is that the source energy is represented by the surface of a sphere, and at a greater distance, as the radius of the sphere is larger, the energy drops by approximately 6 dB because the surface of the sphere is four times larger.
[0193] Using this effect as a basis for adjusting the distance effect gives:
[0194]
[0195] A variant of the 1 / r law may be used, or alternatively a decay curve may be used, where the decay curve may be represented as an equation, a function, or a look-up table that indicates the distance attenuation gain for a given distance from the source. In such a case, the adjustment factor may be:
[0196]
[0197] where f() represents the decay curve as an equation, a function, or a look-up table.
[0198] When the portal size is relatively small and uniform, or when a lower complexity method is required, in 3D space, the distance from the scene audio source to the portal can be calculated once as the distance between the center point of the portal and the source. The distance d between two points P1 = (x1, y1, z1) and P2 = (x2, y2, z2) can be calculated by the following formula:
[0199]
[0200] In other embodiments, the method may include determining an average distance based on first calculating the distances across a plurality of uniformly distributed positions on the portal (such as the corners of the bounding box or the nodes of the mesh describing it) and taking the mean of those calculated distances.
[0201] Other methods, especially when the size of the portal is relatively large with respect to the distance from the source to the portal (e.g., maximum portal size ≥ 0.9 * d tgt ) can use the shortest distance between the target source and a point in the portal.
[0202] In many embodiments, the renderer is arranged to adjust the level of the audio component of the audio source in the second room based on the difference between the direction from the reference position to the transfer area and the direction from the scene audio source position to the transfer area.
[0203] Similar to the distance, the angle between the scene audio source position and the portal affects how much energy is lost through the portal. At a narrower angle, as seen from the source, the effective surface of the portal is smaller, so less energy will be lost.
[0204] Some embodiments may apply a simple linear relationship:
[0205]
[0206] Or more typically, in 3D:
[0207]
[0208] where the a and e subscripts indicate the azimuth and elevation angles between the source position and the portal plane.
[0209] Other embodiments may consider that when the angle is further away from the nominal angle, the energy reduction is stronger. For example, by:
[0210]
[0211] Or:
[0212]
[0213] Or:
[0214]
[0215] Calculating the angle of a scene audio source can be done in various ways, depending on which point on the transmission area is used. A simple approach could be to use the middle of the transmission area as a reference for angle calculation. In many cases, the angle to the nearest point on the transmission area may be particularly beneficial for sources close to the portal. Other embodiments may interpolate between these two angles based on the distance of the source to the transmission area.
[0216] In a more complex approach, the angle can be an average angle based on multiple points in the transmission area. For example, the four corners of the (rectangular equivalent of the) transmission area, or the nodes of a grid describing the transmission area. This may be particularly beneficial for estimating the true energy proportion for a wide range of source positions.
[0217] Determining the angle based on two coordinates means that the two coordinates form a line, and the angle of that line relative to a reference orientation is calculated. The reference orientation can be defined as part of the coordinate system used to define the scene. For example, the negative z-axis. Alternatively, the angle can be calculated relative to the normal vector of the portal. Calculating the angle between two vectors is well known in the art and will not be described further.
[0218] When the nominal source is not at a 90° (0.5π radians) angle to the transmission area, a simple method is to perform an inverse compensation of the reference angle to 90°, followed by a compensation from 90° to the target angle. For example:
[0219]
[0220] In many embodiments, the metadata may include data describing the directivity of the scene audio source, and the renderer 507 may be arranged to adjust the level of the audio component generated for the scene audio source according to / depending on the directivity. The directivity typically indicates the variation of the gain / signal level in different directions from the scene audio source.
[0221] The renderer 507 may specifically be arranged to scale the level of the audio component representing the scene audio source according to the relative directivity gain of the first audio source in the direction from the scene audio source to the transmission area, where the relative directivity gain indicates the gain relative to an omnidirectional source.
[0222] The directivity module also affects the amount of energy leaking through the portal, and it can be frequency-dependent.
[0223] The directivity can be given as a directivity pattern, which represents the amount of energy radiated in the range of azimuth and elevation angles relative to the omnidirectional pattern and the nominal front direction. In a low-complexity method, the effect of the directivity pattern can be regarded as the average energy level in the azimuth and elevation ranges covered by the portal.
[0224]
[0225] where a and e respectively represent the azimuth and elevation angles covered by the ranges from a min to a max and e min to e max which are defined by the relative position of the portal with respect to the nominal front direction (and directivity pattern) of the source, and q and n represent the number of azimuth and elevation angles considered. L a,e is the directivity gain associated with the azimuth angle a and elevation angle e as specified in the directivity pattern. Figure 10 An exemplary scenario is shown in Figure 10 which illustrates the scenario audio source 1001 with respect to the portal 1003.
[0226] In many embodiments, G dir can be frequency-dependent and can be calculated per frequency band in which the directivity pattern is specified.
[0227] In some embodiments, the rendered audio source can specifically be an audio source representing audio from a third acoustic environment, and specifically, it can represent audio arriving at a second acoustic environment via a portal between the third acoustic environment and the second acoustic environment. For example, for a scenario audio source in the third acoustic environment, the described method can be used to determine the level at the portal between the third acoustic environment and the second acoustic environment. The resulting audio signal (i.e., the audio signal from the scenario audio source after level compensation) can thus represent the audio from the scenario audio source that will propagate into the second acoustic environment via the second portal. This sound can further propagate into the first acoustic environment through the first portal. This effect can be simulated by positioning the audio source at the second portal, where the signal corresponds to the signal entering the second acoustic environment from the first acoustic environment. The audio source can then be processed as described above, thus allowing the determination and rendering of the audio entering the first acoustic environment.
[0228] The method can be used in this way to represent sound / audio propagation through multiple rooms.
[0229] Alternatively or additionally, for the metadata described previously, in some embodiments, the metadata received by the second receiver 503 may include transfer region data that describes a transfer region in an acoustic attenuation boundary of a scene, and may also include energy transfer parameters. Each energy transfer parameter indicates at least one energy attenuation between a pair of transfer regions, and specifically indicates the energy attenuation between two transfer regions of different acoustic attenuation boundaries. The energy attenuation for a pair of transfer regions indicates the proportion of audio energy propagating from one transfer region in the pair to the other transfer region in the pair. Thus, each energy transfer parameter may include one energy attenuation indication (or, as will be described later, two energy attenuation indications) for the pair of transfer regions.
[0230] Thus, although the nominal energy transfer indication may indicate the proportion of energy arriving at a transfer region from a given nominal omnidirectional audio source, the energy transfer parameter, and in particular the energy attenuation indication, reflects the proportion of audio energy arriving at a second transfer region from a first transfer region. Similar to the nominal energy transfer indication, the energy attenuation indication may reflect the proportion of energy incident on the first transfer region and / or the proportion of energy radiated from / leaving the first transfer region (into the first acoustic environment). Generally, the remarks provided for the nominal energy transfer indication apply equally to the energy attenuation indication, mutatis mutandis.
[0231] The renderer 507 may render an audio source from a second acoustic environment in the first acoustic environment based on the energy attenuation indications for a pair of transfer regions that are part of the acoustic attenuation boundary between the first acoustic environment and the second acoustic environment. Specifically, the renderer 507 may determine the level in the first acoustic environment of the signal component of the audio source in the second acoustic environment based on the energy attenuation indication. This signal component may thus represent the audio that propagates from the second acoustic environment to the first acoustic environment through the first transfer region and the second transfer region.
[0232] Renderer 507 can, for example, determine the signal energy of a given audio source incident on the second transfer region. For example, in some embodiments, the level / energy of reverberant audio in the second acoustic environment can be determined and converted into the energy / signal level of the reverberant audio that is considered to reach the second transfer region. As another example, the energy / signal level at the second transfer region can be determined for a given specific audio source (e.g., a point audio source). Specifically, the energy / signal level at the second transfer region from an audio source in the second acoustic environment can be determined based on the nominal energy transfer indication for the audio source. In practice, a method based on the nominal energy transfer indication as described above can be used to determine the energy / signal level in the audio source that reaches the second transfer region. The resulting energy / signal level for the signal at the first transfer region can be determined by directly applying the energy attenuation indication for the pair of transfer regions, and renderer 507 can adjust the signal level of the rendered signal component to reflect this attenuation. The previously described rendering method can be used, for example, as described, but with the attenuation determined by the energy attenuation indication introduced.
[0233] As a specific example, in Figure 6 the example of, in a specific example, the audio source 607 in the second acoustic environment of room A can reach the listening room E first via the transfer region / portal 1 between rooms A and C and then via the transfer region / portal 4 between rooms C and E. In this example, a nominal energy transfer indication can be provided for portal 1, for example, and based on this, the energy at portal 1 from audio source 607 can be determined as described above. This can provide a first attenuation factor for the energy / signal level from the audio source. The attenuation can then be increased based on the energy attenuation indication provided in the metadata for portals 1 and 4, or equivalently, the energy / signal level at transfer region 1 can be reduced by the amount given by the energy attenuation indication. The audio from the audio source in room A can then be rendered for the listening position in room E, but the audio has a reduced level that reflects the attenuation associated with propagation through the two portals.
[0234] The acoustic environments of the two transfer regions / portals of a given pair of transfer regions for which energy transfer parameters are provided (and thus the acoustic attenuation boundaries in the acoustic attenuation boundaries that form the transfer region / portal) can have a shared acoustic environment, i.e., the two acoustic environments can be partitioned by a single acoustic environment, and thus the two acoustic attenuation boundaries that form the portal can both be boundaries of a single shared acoustic environment. Specifically, as in Figure 6 the example of, the two portals 1 and 4 can be used for the boundaries of rooms (i.e., rooms A and E) with different acoustic attenuations, but are also two boundaries of the same room (i.e., room C).
[0235] Energy transfer parameters and energy attenuation indications can be used to describe sound propagation between different rooms via a portal to an interconnected room. In some embodiments, the metadata includes energy attenuation parameters for only the transfer area pairs that share the boundary of the acoustic environment, i.e., for which the portal / transfer area is formed in an acoustic attenuation boundary that is the boundary of the same acoustic environment. This can provide a reduced data rate for the metadata and can limit the data representation to the most likely audio propagation between acoustic environments. Additionally, in some embodiments, if it is desired to determine sound propagation for further separated acoustic environments, then such energy transfer parameters / energy attenuation indications can be combined, as described in more detail later.
[0236] A particular advantage of the method is that it can be adapted to and applied to many different topologies and connections between different acoustic environments, including providing information about sound propagation between acoustic environments that do not have a shared acoustic environment. In fact, in many embodiments, one or more of the energy transfer parameters / energy attenuation indications are provided for transfer areas that are acoustic attenuation boundaries that do not share any acoustic environment. For example, as Figure 11 illustrated, energy attenuation indications can be provided for two portals 1 and 3 separated by two acoustic environments / rooms B and C and thus having no shared adjacent acoustic environment. This can allow for the facilitated rendering of audio in room A generated by an audio source 1101 in room D, since the nature of the complete path of sound propagation through the different acoustic environments can be combined and represented by a single energy attenuation indication.
[0237] In fact, energy transfer parameters that provide energy attenuation indications can be provided for any pair of transfer areas to indicate the sound propagation that may occur between them, and in some embodiments, energy attenuation indications can be provided for every pair of possible transfer areas between any two rooms / acoustic environments in a scene.
[0238] In many typical applications, sound propagation can be symmetric, so the energy attenuation indication for propagation from transfer area x to transfer area y is the same as the propagation from transfer area y to transfer area x. In such cases, the same energy attenuation indication can be used to render an audio signal from an audio source in a second acoustic environment in a first acoustic environment and to render an audio signal from an audio source in the first acoustic environment in the second acoustic environment.
[0239] Such symmetry typically exists in many physical or virtual scenarios and particularly for diffuse or reverberant audio that tends not to be associated with a specific location. The symmetry can be used to reduce the amount of data included in the metadata to describe sound propagation from transfer area to transfer area. For example, in such cases, the energy attenuation indications for all transfer area pairs can be represented by a symmetric matrix, such as
[0240]
[0241] wherein, t xy = t yx indicates the energy attenuation indication from transfer region x to transfer region y and from transfer region y to transfer region x.
[0242] The energy attenuation data can be effectively represented as a matrix as described above, but can also be represented, for example, by direct indication as a set of portal pairs and corresponding transfer region-to-transfer region energy attenuation indications, or in other suitable ways. A matrix such as the above may be sparsely populated, or the set of portal pairs may not be a complete set of possible pairs. This is often beneficial for scenarios with many acoustic environments. For example, entries with high energy attenuation values can be excluded. For example, when 10*log10(energy attenuation[i,j]) < -60 dB.
[0243] Each energy attenuation indication is provided for two transfer regions / portals, and the metadata provides the energy attenuation indication and the identification of the transfer regions. The energy attenuation indication can also be considered an inverse energy transfer indication, i.e., the higher the energy attenuation, the lower the energy transfer. The energy attenuation indication between two transfer regions can generally indicate an increasing attenuation for an increasing distance between the transfer regions, and depends on how many intermediate acoustic environments and transfer regions the sound has to pass through to reach the destination transfer region. Additionally, if two transfer regions are misaligned (around a corner or blocked by an obstacle), the corresponding energy attenuation indication can indicate a higher attenuation to reflect the higher loss of sound attenuation.
[0244] Furthermore, in some embodiments, the energy attenuation indication can indicate a time-varying value or a value that depends, for example, on the nature of the dynamic changes of the scenario. For example, if a door-like portal is opened, closed, or moved, the energy attenuation indication can change.
[0245] The method can include the following considerations: It can be assumed that the portal radiates sound uniformly on its surface into the receiving room. When the receiving room has other portals, a portion of the sound from the first portal will reach such a second portal and may leak into the next receiving room. The amount of sound transmitted can be associated with the relative position and size of the other portals with respect to the first portal and the total room surface area.
[0246] This information can be used to effectively determine how much energy from a source in one room contributes to other rooms, and this information can be captured by the energy attenuation indication. For example, each row in the above matrix can indicate for the portals of the associated room how much contribution it has to all other rooms.
[0247] In many embodiments, it may be assumed that the transfer region locations (and the acoustic attenuation boundaries) are fixed and do not move, and thus the energy attenuation indication can be pre-computed for their specific locations. A simple approach is to calculate the visible region of the receiving portal relative to the center point of the source portal and compare this region with a hemispherical region having a radius equal to the distance between the portals. Assuming the portals are sub-parts of larger planes, it can often be assumed that they radiate hemispherically rather than omnidirectionally as in the nominal energy transfer indication.
[0248] In a more complex approach, rather than calculating only the visible area relative to the center point of the source, the area of the source portal can be considered. This can be done by calculating the visible area across multiple locations bounded by the source portal and taking the average visible area, or by other means.
[0249] In many embodiments, the energy attenuation indication can be calculated on the encoder side or using an offline process, where the computational complexity is more fully available (e.g., it can be calculated at the VR server 303 or even actually on the client device itself). In such cases, acoustic models of various levels of complexity can be used to determine how much energy from the first transfer region reaches the second transfer region. This can include occlusion and / or diffraction modeling.
[0250] Some embodiments can focus on calculating the energy transfer / attenuation from all transfer regions of a room to all other transfer regions of the same room. These transfers can then be combined to represent higher-order room-to-room transfers (i.e., including more than one shared / intermediate room). For example, when room A is associated with transfer regions 4 and 5 and room B is associated with transfer regions 5 and 2, the transfer from transfer region 4 to 2 can be obtained by combining the transfer from transfer region 4 to 5 calculated for room A with the transfer from transfer region 5 to 2 calculated for room B.
[0251] Some embodiments can also include the transfer or material properties of transfer region 5.
[0252] The energy attenuation indication can be directly used to determine the energy reduction factor for sound to reach from one acoustic environment to another acoustic environment, and the energy reduction factor can be used to perform rendering.
[0253] Specifically, for a given audio source in the source room, the energy incident on the transfer region can be determined. This can be done, for example, using the methods described previously or by other means. For example, the data can come from an audio source defined as a low-complexity replacement for several sources in the source room, can be calculated using another method, or can be generated by reverberation rendering in the source room.
[0254] The resulting energy reduction factor F tgtcan be applied to the signal by the renderer 507, for example as:
[0255]
[0256] where S in can be the input contribution of a source represented by the signal S src to a rendering algorithm (e.g., reverberation, coupled source rendering).
[0257] A particular advantage of this method is that it does not require detailed geometric information of the scene, and in particular, detailed geometric information of a room, an acoustic attenuation boundary, a transmission region, etc., or in fact, detailed geometric information of a specific acoustic property of the scene. In fact, information about the exact connection between rooms or the acoustic properties of these is not required. Instead, the energy transfer parameter can be considered a topological property that simply connects two transmission regions and provides information on the sound propagation between them. This can allow for greatly facilitated operations and rendering with significantly reduced complexity and possible resource usage.
[0258] In many embodiments, the energy attenuation indication for a pair of portals can indicate the proportion of audio energy incident on one transmission region that will propagate to be incident on the other transmission region. This can be beneficial in allowing the energy attenuation indication to be symmetric, thus allowing the use of one indication in both directions and therefore reducing the amount of metadata. It can also allow for adjustment of the rendering based on the specific acoustic properties of the transmission region. For example, if the transmission region is dynamically covered by a fabric (e.g., a curtain), this can be reflected by introducing an additional attenuation factor that can be omitted when the transmission region is not covered.
[0259] In other embodiments, the energy attenuation indication can indicate the energy attenuation of the output of the receiving transmission region, i.e., it can represent the energy leaving / radiating from a given transmission region for a given energy incident on another transmission region. This can allow for rendering with reduced complexity in many cases.
[0260] In many embodiments, the renderer 507 can be arranged to generate an audio source by combining two, more, or all audio sources in an acoustic environment into a single audio source. Such an audio source can be generated, for example, by determining the relative sound levels at given audio source locations and generating the audio as a weighted sum of audio signals from the individual audio sources, where the weights reflect the relative sound levels at the source locations. The source locations can be specifically generated to correspond to the locations of the transmission regions.
[0261] For example, the previously described method of determining the acoustic level at the transfer region based on the nominal energy transfer indication and the actual position of the individual audio sources can be performed for all audio sources in the acoustic environment. The audio signals can then be weighted and summed accordingly such that all audio of the acoustic environment is represented by a single audio source located at the transfer region.
[0262] The sound propagation to the listening acoustic environment can then be determined based on the energy transfer parameters as described above, and the renderer 507 can render the resulting signal, for example, as reverberant and diffuse sound.
[0263] Thus, in some embodiments, the sources of each acoustic environment can be combined into a single source at the associated transfer region (e.g., one source for each transfer region associated with the environment). The renderer 507 can then determine the sound in the listening room for each transfer region of the acoustic environment based on the energy attenuation indication applying the energy transfer parameters, and subsequently proceed to render all these audio signal components. This can provide a lower complexity for rendering audio originating from another acoustic environment in one acoustic environment and taking into account the sound propagation through multiple and possibly all acoustic paths between the acoustic environments.
[0264] In some embodiments, for all transfer regions for which energy transfer parameters are provided and for which some sound transfer / propagation is possible, and for the rendering of the signal components representing the inter-room propagation through the transfer region, the renderer 507 can simply extract and use the appropriate energy attenuation indication for that transfer region.
[0265] However, in some embodiments, the energy transfer parameters can be provided only for a subset of the transfer regions, such as, for example, only for the transfer regions sharing a common acoustic environment. This can allow for a reduced data rate and / or can substantially alleviate the requirements for determining accurate energy attenuation indications. For example, if these are based on measurements in a real building, the number of measurement operations required can be significantly reduced.
[0266] In such embodiments, the energy attenuation indication for other transfer region pairs can be determined, for example, in some cases by combining the energy attenuation indications for other transfer region pairs. Thus, in some embodiments, the renderer 507 is arranged to generate a combined energy transfer attenuation by combining the energy transfer attenuation for a first pair of transfer regions and for a second pair of transfer regions, where the two pairs include a common transfer region. For example, as Figure 12As illustrated by the example of, the first pair of transfer regions can be used for the transfer region between the first and second transfer regions, thereby providing an indication of the energy transfer / attenuation between the first and second environments. The second pair of transfer regions can be between the third transfer region and the second transfer region, thereby providing an indication of the energy transfer / attenuation between the third transfer region and the second transfer region, and thus providing an indication of the energy transfer / attenuation between the second transfer region and the third acoustic environment. The energy attenuation indications of the two pairs of transfer regions can be combined, for example, simply by combining the attenuations (e.g., by multiplying the two energy attenuations in the linear domain or adding them in the logarithmic domain for the attenuation values). Thus, the resulting combined value indicates the energy attenuation from the third transfer region to the first transfer region, and thus indicates the sound propagation from the third acoustic environment to the first acoustic environment. Thus, the combined energy attenuation can be used to render audio for a listening position in the first acoustic environment from an audio source in the third acoustic environment in the same manner as when a direct energy attenuation indication is provided for the pair of the first transfer region and the third transfer region. Some embodiments may also include the transfer or material properties of the second transfer region.
[0267] In some embodiments, the energy transfer parameter for a given pair of transfer regions may include a plurality of energy attenuation indications, where different energy attenuation indications are provided for different acoustic environments separated by an acoustic attenuation boundary, in which one of the transfer regions of the pair is provided.
[0268] For example, the energy transfer parameter for the first transfer region in a pair of transfer regions may include energy attenuation indications for two acoustic environments separated by a given transfer region / acoustic attenuation boundary. Thus, instead of simply providing an energy attenuation indication for the pair of transfer regions, different energy attenuation indications can be provided for the sound arriving at the source transfer region from one acoustic environment and the sound arriving at the source transfer region from the other acoustic environment. For example, for a portal in a wall dividing two rooms, separate energy attenuation indications can be provided for each of the rooms. The renderer 507 can then render the sounds from the two acoustic environments differently.
[0269] This can provide improved performance in many cases, and can specifically reflect that the sound from other acoustic environments / rooms to different acoustic environments / rooms can depend on the direction of the incident sound. In fact, in many embodiments, the sound from a given room to another given room may only be possible / suitable for the sound passing through a given portal in one direction and not in the other direction. In many embodiments, one of the directional energy attenuation indications for a given pair of transfer regions may be zero.
[0270] Such an approach of directional, individual energy decay indication can be particularly suitable for situations where the energy decays for multiple pairs of transfer regions are combined to provide a path from a source acoustic environment to a destination listening acoustic environment.
[0271] In fact, portal-to-portal transfer (transfer region to transfer region) often depends on the direction of sound incident on the portal (transfer region). For example, for a building comprising multiple rooms, a rendering algorithm can be arranged to proceed to determine the room in which each audio source is located and then, for each source, determine all the portals in that room. It can then determine the audio source energy at each portal (especially that incident thereon) and then continue to apply the energy decay of each portal in the source room to each portal in the listening room.
[0272] However, for some situations, some such methods may lead to undesirable behavior, and this can be addressed by making the energy transfer parameter indication depend on the directional energy decay of the sound incident on the (source) transfer region.
[0273] For example, consider Figure 13 the room layout. In the case where the listener is in room A and the source s1 is in room C, it can be seen that transfer p 21 (from portal 2 to portal 1) is relevant, but transfer p 31 is independent of the listening position in room A. For source s2 in room D, transfer p 31 is relevant. The topology of the room contributes to determining which transfers are important.
[0274] Relevance can be pre-determined and represented in the received metadata, which reflects different energy decays for different acoustic environments for at least one of at least a pair of transfer regions. Different acoustic environments of a transfer region are acoustic environments separated by acoustic attenuation boundaries where the transfer region exists.
[0275] Since the portal / transfer region is a region in the acoustic attenuation boundary connecting / separating two rooms, the source will be on only one of the two sides of the portal. Thus, the metadata can provide two different energy decay values, where the first value corresponds to the first room connected to the portal and the second value corresponds to the second room connected to the portal.
[0276] This can be represented by two values or two matrices for each pair of portals (or a 3D matrix with one dimension having size 2). The relationship with the rooms can be predefined. For example, when a portal is defined as having IDs to two environments, a first energy attenuation value can correspond to the first environment and a second energy attenuation value can correspond to the second environment. It should be understood that any means of metadata indicating different / directional energy attenuation can be used.
[0277] In Figure 13 the example of, for the value of p 31 can specifically indicate infinite attenuation (zero energy transfer) for room C and typically a non-infinite value (indicating some energy transfer) for room D.
[0278] In Figure 14 the example of, there may always be an acoustic path from any acoustic environment to any other acoustic environment, but this can still mean that some pairs of portals represent irrelevant / invalid paths. For example, when source s3 is in room M, p 65 is relevant, while when the source will be in room L, p 65 is not relevant. This is the case because both portal 1 and portal 2 are related to the source room L, but also because there is a path through the listening room between the entrances outside room L.
[0279] Thus, in some embodiments, four values can be provided, depending on which side of the first transfer region the source is on and which side of the second transfer region the listener is on.
[0280] Figure 15 The second layout example shown shows an example where, for source s4, only p 10,7 is relevant, while for source s5, p 10,7 is not relevant, but p 87 and p 11,7 are relevant. The energy attenuation values for non-relevant transfers can be indicated as infinite.
[0281] In Figure 16 the fourth example (which is a variant of the third example) shown, it can be seen that p 10,7 is relevant to both s4 and s5, but these values may be different.
[0282] Instead of relying on additional and explicit metadata that provides different directional energy attenuation values, some embodiments can determine relevant transfers based on metadata that is already available and used for other purposes. Specifically, based on information about which portals connect which environments, a connectivity graph can be made. This graph indicates how different environments are connected by portals and can be used to determine relevance.
[0283] Figure 17is shown from Figures 13 to 16 graphs of four examples. Each node represents a room and each vertex represents a portal. Known graph techniques can be used to determine whether there is a connection from a particular room through a particular first portal, where each vertex can be crossed only once.
[0284] In this regard, it can be seen that, for example, for the third example, there is no path from T to P through portal 10, but there is such a path in the fourth example.
[0285] Such graphs can also be used when providing energy transfer parameters for first-order transfer (i.e., through only 1 room) using a path-finding algorithm that collects relevant transfer factors on one or more paths it finds.
[0286] In many embodiments, the audio signal components determined as described above can be rendered as non-direct audio components, i.e., they can be rendered as audio sources that are propagated by means other than (only) direct line-of-sight propagation.
[0287] Specifically, the rendering can be as a reverberant audio component in a first acoustic environment. Specifically, the audio signal can be horizontally compensated and the resulting signal can be rendered using a suitable rendering method for generating reverberant audio. It should be understood that a large number of algorithms for rendering an audio signal as reverberant audio / sound are known and can be used.
[0288] Thus, in many embodiments, the method can be used to generate reverberant audio in a room / acoustic environment produced by an audio source in another room. This can provide a particularly advantageous method in many situations and can reflect a more natural experience in many cases.
[0289] In many embodiments, the renderer 507 can be arranged to render the signal components to reflect all the sound energy arriving at the corresponding transfer area. Specifically, the rendering can be such that the transfer area is considered to have no other effect on the rendered audio, and in fact, other than the extension of the transfer area in the acoustic attenuation boundary, the transfer area does not have other acoustic properties or characteristics that need to be considered. In fact, the transfer area / portal can simply be considered to correspond to an opening in the acoustic attenuation boundary / wall and can be considered to have no acoustic effect.
[0290] In such a case, the energy arriving at the transfer area can be considered equal to the energy leaving the transfer area. The determined energy attenuation (from another transfer area or from a specific sound source) can be considered the same for the incident energy and the radiated energy entering the first acoustic environment. For example, the nominal energy transfer indication or energy attenuation can inherently indicate both (since they can be the same).
[0291] However, in other embodiments, the transmission region itself may be considered to have acoustic properties that affect the amount of sound energy passing through the transmission region. For example, in some embodiments, the transmission region may not be a complete opening and may have some attenuation, although the attenuation is less than the surrounding acoustic attenuation boundary. For example, a wall may include a door covered with a drape that provides some acoustic attenuation. The renderer 507 may be arranged to take such attenuation into account and, in particular, may reduce the signal level accordingly. In some cases, the acoustic effect of the transmission region may vary, and the rendering may be adjusted dynamically to reflect this.
[0292] Portals may represent features such as windows or doors, and as such whether they are open or closed can change during runtime. When a user or other element of the rendering system interacts with a portal, such as a partially closed door, an additional weighting function may be applied to the calculated total gain such that 1 is for a fully open portal and 0 or a factor related to the portal material properties is for a fully closed portal. For example, the transmission coefficient of the scene element covering the portal may be used.
[0293] When a portal or other surface has a non-zero coupling coefficient, a similar approach can be used. The energy reaching a closed portal may not be completely blocked by the portal, but rather a portion of the energy may couple to the surface and be re-radiated. As an extreme example, a single-pane window will vibrate when a loud noise is generated on the opposite side, thereby reproducing some portion of the noise even though there is no direct path for sound to travel. Thus, in some embodiments, sound propagation through the transmission region / portal can occur fully or partially via an acoustic coupling effect. In some such embodiments, the renderer 507 may be arranged to render the corresponding sound from a source in a neighboring room by rendering a sound source at the location of the portal, and the sound source has an energy level that depends on and is, for example, proportional to the signal energy reaching the transmission region, and the signal energy is compensated by the attenuation that occurs due to the coupled propagation effect.
[0294] Often, the acoustic properties of the transmission region may take the form of material-related properties, such as reflectivity, absorptivity, transmissivity, or related effects. Reflectivity may indicate the proportion of incident sound that is reflected in a specular and / or diffuse manner. Absorption may involve dissipation or conversion into material vibrations (which may be re-emitted as a source of coupling) in the material. Transmissivity generally indicates how much energy passes through.
[0295] Thus, in some embodiments, metadata for a delivery region (e.g., a nominal energy delivery indication and energy delivery parameters) can indicate the energy arriving at the delivery region and thus the incident energy on the delivery region. Then, when determining an appropriate signal level for the resulting audio signal component, the energy can be reduced / modified based on the acoustic properties of the delivery region. This can provide, for example, improved flexibility and, for example, allow for easy adaptation to dynamic changes in the delivery region.
[0296] However, in other embodiments, metadata for a delivery region (e.g., a nominal energy delivery indication and / or energy delivery parameters) can indicate the energy leaving the delivery region. Thus, in some embodiments, the metadata (e.g., a nominal energy delivery indication and / or energy delivery parameters) can reflect / include a contribution from the acoustic properties of the delivery region itself. Thus, in some embodiments, it may not be necessary to explicitly consider or account for different acoustic properties of the delivery region during rendering, but rather it can be implicitly specified by the received metadata and may not require specific adjustment of the rendering itself.
[0297] It should be understood that the use of the (nominal) energy delivery indication as described above can be combined with or separated from the use of the energy delivery parameters as described above. Similarly, it should be understood that the use of the energy delivery parameters as described above can be combined with or separated from the use of the (nominal) energy delivery indication as described above. The principles, methods, functions, uses, etc. described above for the nominal energy delivery indication and energy delivery parameters respectively thus (as appropriate) apply to individual uses and do not imply or require that functions related to the nominal energy delivery indication must be combined with functions related to the energy delivery parameters. Different metadata and applications are independent and separate. However, it should also be understood that for embodiments using two functions related to the nominal energy delivery indication and energy delivery parameters, particularly advantageous and synergistic operations can be achieved.
[0298] (One or more) devices can be implemented specifically in one or more appropriately programmed processors. In particular, an artificial neural network can be implemented in one or more such appropriately programmed processors. Different functional blocks can be implemented in separate processors and / or can be implemented, for example, in the same processor. Examples of suitable processors are provided below.
[0299] Figure 18FIG. is a block diagram of an example processor 1800 in accordance with an embodiment of the present disclosure. The processor 1800 may be used to implement one or more processors that implement the apparatus or its elements as described above (specifically including one or more artificial neural networks). The processor 1800 may be any suitable type of processor, including but not limited to a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) (where the FPGA has been programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (where the ASIC has been designed to form a processor), or a combination thereof.
[0300] The processor 1800 may include one or more cores 1802. The cores 1802 may include one or more arithmetic logic units (ALUs) 1804. In some embodiments, in addition to or instead of the ALU 1804, the cores 1802 may include a floating point logic unit (FPLU) 1806 and / or a digital signal processing unit (DSPU) 1808.
[0301] The processor 1800 may include one or more registers 1812 communicatively coupled to the cores 1802. The registers 1812 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 1812 may be implemented using static memory. The registers may provide data, instructions, and addresses to the cores 1802.
[0302] In some embodiments, the processor 1800 may include one or more levels of cache memory 1810 communicatively coupled to the cores 1802. The cache memory 1810 may provide computer-readable instructions for execution to the cores 1802. The cache memory 1810 may provide data to be processed by the cores 1802. In some embodiments, the computer-readable instructions may have been provided to the cache memory 1810 by local memory (e.g., local memory attached to an external bus 1816). The cache memory 1810 may be implemented using any suitable type of cache memory, such as metal oxide semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.
[0303] Processor 1800 may include a controller 1814 that may control the input to processor 1800 from other processors and / or components included in the system and / or the output from processor 1800 to other processors and / or components included in the system. Controller 1814 may control the data paths in ALU 1804, FPLU 1806, and / or DSPU 1808. Controller 1814 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of controller 1814 may be implemented as discrete gates, FPGAs, ASICs, or any other suitable technology.
[0304] Registers 1812 and cache 1810 may communicate with controller 1814 and core 1802 via internal connections 1820A, 1820B, 1820C, and 1820D. The internal connections may be implemented as a bus, multiplexer, crossbar switch, and / or any other suitable connection technology.
[0305] The input and output of processor 1800 may be provided via a bus 1816 that may include one or more wires. Bus 1816 may be communicatively coupled to one or more components of processor 1800, such as controller 1814, cache 1810, and / or registers 1812. Bus 1816 may be coupled to one or more components of the system.
[0306] Bus 1816 may be coupled to one or more external memories. The external memories may include a read-only memory (ROM) 1832. ROM 1832 may be a mask ROM, an electrically programmable read-only memory (EPROM), or any other suitable technology. The external memories may include a random access memory (RAM) 1833. RAM 1833 may be a static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memories may include an electrically erasable programmable read-only memory (EEPROM) 1835. The external memories may include a flash memory 1834. The external memories may include a magnetic storage device such as a disk 1836. In some embodiments, the external memories may be included in the system.
[0307] Technical voice and sound may be considered equivalent and interchangeable and may respectively refer to physical sound pressure and / or electrical signal representations, as appropriate, in context.
[0308] It should be understood that, for clarity, the embodiments of the present invention have been described above with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable functional distribution between different functional circuits, units, or processors can be used without departing from the present invention. For example, functions illustrated as being performed by separate processors or controllers can be performed by the same processor or controller. Thus, the reference to a particular functional unit or circuit is only considered as a reference to a suitable module for providing the said function, rather than indicating a strict logical or physical structure or organization.
[0309] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The present invention can optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In fact, the functions can be implemented in a single unit, in multiple units, or as part of other functional units. In this way, the present invention can be implemented in a single unit, or can be physically and functionally distributed among different units, circuits, and processors.
[0310] Although the present invention has been described in conjunction with some embodiments, the present invention is not intended to be limited to the specific forms set forth herein. Instead, the scope of the present invention is limited only by the appended claims. Additionally, although features may seem to be described in conjunction with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term includes the presence of other elements or steps without excluding them.
[0311] Furthermore, although multiple modules, elements, circuits, or method steps are listed individually, they can be implemented by, for example, a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features may be advantageously combined, and including the features in different claims does not mean that the combination of the features is infeasible and / or advantageous. Moreover, including a feature in a category of claims does not mean a limitation to that category, but rather indicates that the feature is equally applicable to other categories of claims. Additionally, the order of the features in the claims does not imply any particular order in which the features must operate, and in particular, the order of the individual steps in method claims does not imply that the steps must be performed in that order. Instead, these steps can be performed in any suitable order. Additionally, singular references do not exclude plural. Thus, references to "a", "an", "first", "second", etc. do not exclude plural. The reference numerals in the claims are provided only as clarifying examples and should not be construed as limiting the scope of the claims in any way.
[0312] Typically, examples of an image synthesis system, an image synthesis method, and a computer program for implementing the method are indicated by the following embodiments.
[0313] Embodiment:
[0314] Embodiment 1. An audio device:
[0315] A first receiver (501) arranged to receive audio data of an audio source for a scene including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0316] A second receiver (501) arranged to receive metadata for the audio data, the metadata including:
[0317] An indication of the position of at least a first transmission region for a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region in the first acoustic attenuation boundary having a lower attenuation than the average attenuation of the first acoustic attenuation boundary outside the transmission region; and
[0318] An energy transfer indication for the first transmission region, the energy transfer indication indicating the proportion of the energy at the first transmission region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transmission region;
[0319] A renderer (507) arranged to render an audio signal for a listening position in the first acoustic environment, the rendering including generating a first audio component by rendering the audio data of a first audio source for the second acoustic environment and adjusting the level of the first audio component according to the energy transfer indication.
[0320] Embodiment 2. The audio device according to Embodiment 1, wherein the renderer (507) is arranged to adjust the level of the first audio component in response to the position of the first audio source relative to the reference position.
[0321] Embodiment 3. The audio device according to Embodiment 1 or 2, wherein the renderer (507) is arranged to adjust the level of the first audio component according to the difference between a reference distance from the reference position to the first transmission region and the distance from the position of the first audio source to the first transmission region.
[0322] Embodiment 4. The audio device according to any one of the preceding embodiments, wherein the renderer (507) is arranged to adjust the level of the first audio component according to the difference between the direction from the reference position to the first transmission region and the direction from the first audio source to the first transmission region.
[0323] Example 5. The audio device according to any one of the preceding embodiments, wherein the metadata further includes data describing the directivity of the sound radiation from the first audio source, and the renderer (507) is arranged to adjust the level of the first audio component according to the directivity.
[0324] Example 6. The audio device according to any one of the preceding embodiments, wherein the renderer (507) is arranged to scale the level of the first audio component according to the relative directivity gain of the first audio source in the direction from the first audio source to the first transfer region, and the relative directivity gain indicates the gain relative to an omnidirectional source.
[0325] Example 7. The audio device according to any one of the preceding embodiments, wherein the first audio source represents audio arriving at the second acoustic environment from a third acoustic environment via a second transfer region of a second boundary that separates the third acoustic environment from the second acoustic environment.
[0326] Example 8. The audio device according to any one of the preceding embodiments, wherein the metadata further includes:
[0327] Energy transfer parameters, each energy transfer parameter indicating the energy attenuation between a pair of transfer regions, and the energy attenuation for a pair of transfer regions indicating the proportion of the audio energy propagating from one transfer region of the pair to the other transfer region of the pair;
[0328] And the renderer (507) is arranged to render the second audio component of the audio signal by rendering a second audio source of the third acoustic environment according to the energy attenuation for a pair of transfer regions, the pair of transfer regions including the transfer region of the acoustic attenuation boundary of the first acoustic environment and the transfer region of the second acoustic attenuation boundary, and the second acoustic attenuation boundary is a boundary of the second acoustic environment.
[0329] Example 9. The audio device according to any one of the preceding embodiments, wherein the metadata includes a coupling coefficient for the first transfer region, and the renderer (507) is arranged to render the first audio component as an audio source originating from a position near the first transfer region according to the coupling factor.
[0330] Example 10. The audio device according to any one of the preceding embodiments, wherein the renderer (507) is arranged to render the first audio component as a reverberant audio component of the first acoustic environment.
[0331] Example 11. The audio device according to any one of the preceding embodiments, wherein the renderer (507) is arranged to generate a reverberant audio signal for the second environment to include a first component from the first audio source, the renderer (507) is further arranged to determine an energy loss estimate as a proportion of the energy in the first audio source that reaches the first transfer region, the proportion of the energy being determined based on the energy transfer indication and the position of the first audio source relative to the reference position; and the renderer (507) is arranged to reduce the level of the reverberant audio signal by an amount depending on the energy loss estimate.
[0332] Example 12. The audio device according to any one of the preceding embodiments, wherein the energy transfer indication reflects the proportion of a sphere covered by the first transfer region, the sphere being centered at the reference position and having a radius corresponding to the distance from the reference position to the first transfer region.
[0333] Example 13. The audio device according to any one of the preceding embodiments, wherein the renderer (507) is arranged to render the first audio component as a non-direct audio component.
[0334] Example 14. A method of rendering an audio signal, the method comprising:
[0335] Receiving audio data of an audio source for a scene including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0336] Receiving metadata for the audio data, the metadata including:
[0337] A position indication of at least a first transfer region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transfer region being a region in the first acoustic attenuation boundary having a lower attenuation than the average attenuation of the first acoustic attenuation boundary outside the transfer region; and
[0338] An energy transfer indication for the first transfer region, the energy transfer indication indicating the proportion of the energy from an omnidirectional point audio source at a reference position at the first transfer region, the reference position being a relative position with respect to the first transfer region;
[0339] Rendering an audio signal for a listening position in the first acoustic environment, the rendering including generating a first audio component by rendering the audio data of a first audio source for the second acoustic environment, and adjusting the level of the first audio component according to the energy transfer indication.
[0340] Example 15. An audio data signal, comprising:
[0341] audio data of an audio source for a scenario including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0342] metadata for the audio data, the metadata including:
[0343] a position indication for at least a first transmission region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region in the first acoustic attenuation boundary having a lower attenuation than an average attenuation of the first acoustic attenuation boundary outside the transmission region; and
[0344] an energy transfer indication for the first transmission region, the energy transfer indication indicating a proportion of energy at the first transmission region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transmission region.
[0345] Example 16. An audio device arranged to generate an audio data signal according to Example 15.
[0346] Example 17. A computer program product including a computer program code module, the computer program code module being adapted to perform all steps of Example 14 when the program runs on a computer.
[0347] Example: An audio device:
[0348] a first receiver (501) arranged to receive audio data of an audio source for a scenario including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0349] a second receiver (501) arranged to receive metadata for the audio data, the metadata including:
[0350] a position indication for at least a first transmission region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region in the first acoustic attenuation boundary having a lower attenuation than an average attenuation of the first acoustic attenuation boundary outside the transmission region;
[0351] A renderer (507) arranged to render an audio signal for a listening position in the first acoustic environment, the rendering including generating a first audio component by rendering audio data of a first audio source for the second acoustic environment and adjusting a level of the first audio component according to an energy transfer indication (wherein the energy transfer indication is for the first transfer region), the energy transfer indication indicating a proportion of energy at the first transfer region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transfer region.
[0352] Example: An audio device:
[0353] A first receiver (501) arranged to receive audio data of an audio source for a scene including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0354] A second receiver (501) arranged to receive metadata for the audio data, the metadata including:
[0355] A position indication of at least a first transfer region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transfer region being a region in the first acoustic attenuation boundary having a lower attenuation than an average attenuation of the first acoustic attenuation boundary outside the transfer region;
[0356] A renderer (507) arranged to render an audio signal for a listening position in the first acoustic environment, the rendering including generating a first audio component by rendering audio data of a first audio source for the second acoustic environment and adjusting a level of the first audio component according to an energy transfer indication (wherein the energy transfer indication is for the first transfer region), the energy transfer indication indicating a proportion of energy at the first transfer region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transfer region, and wherein the energy transfer indication is provided in the metadata or provided by / through a process performed elsewhere.
[0357] Example: An audio device:
[0358] A first receiver (501) arranged to receive audio data of an audio source for a scene including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0359] A second receiver (501) arranged to receive metadata for the audio data, the metadata including:
[0360] An indication of the location of at least a first transmission region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region of the first acoustic attenuation boundary having a lower attenuation than the average attenuation of the first acoustic attenuation boundary outside the transmission region;
[0361] A renderer (507) arranged to render an audio signal for a listening position in the first acoustic environment, the rendering including generating a first audio component by rendering audio data of a first audio source for the second acoustic environment and adjusting the level of the first audio component according to an energy transfer indication (where the energy transfer indication is for the first transmission region), the energy transfer indication indicating the proportion of the energy at the first transmission region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transmission region, and wherein the energy transfer indication is provided in the metadata or provided by / through a process executed elsewhere where sufficient computational complexity is available.
[0362] Example A. An audio device, comprising:
[0363] A first receiver (501) arranged to receive audio data of an audio source for a scene comprising a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0364] A second receiver (503) arranged to receive metadata for the audio data, the metadata including:
[0365] Transmission region data that describes transmission regions in the acoustic attenuation boundary, each transmission region being a region of the acoustic attenuation boundary having a lower attenuation than the average attenuation of the acoustic attenuation boundary outside the transmission region; and
[0366] Energy transfer parameters, each energy transfer parameter indicating the energy attenuation between a pair of transmission regions, the energy attenuation for a pair of transmission regions indicating the proportion of the audio energy propagating from one transmission region of the pair to the other transmission region of the pair;
[0367] A renderer (507) arranged to render an audio signal for a listening position in a first acoustic environment, the rendering including generating a first audio component by rendering a first audio source of a second acoustic environment according to the energy attenuation of a first pair of transmission regions, the first pair of transmission regions including a first transmission region of a first acoustic attenuation boundary that is a boundary of the first acoustic environment and a second transmission region of a second acoustic attenuation boundary that is a boundary of the second acoustic environment.
[0368] Embodiment B. The audio device according to Embodiment A, wherein the energy transfer attenuation for the first pair of transfer regions indicates the proportion of the audio energy incident on the second transfer region that propagates to be incident on the first transfer region.
[0369] Embodiment C. The audio device according to Embodiment A or B, wherein both the first acoustic attenuation boundary and the second acoustic attenuation boundary are boundaries of a third acoustic environment.
[0370] Embodiment D. The audio device according to Embodiment A or B, wherein the first acoustic attenuation boundary and the second first acoustic attenuation boundary are not boundaries of a common acoustic environment.
[0371] Embodiment E. The audio device according to any one of the preceding Embodiments A to D, wherein the first audio source represents the audio in the second audio source of the third acoustic environment that reaches the second transfer region, and the renderer (507) is arranged to: generate a combined energy transfer attenuation by combining the energy transfer attenuation for the first pair of transfer regions and the energy transfer attenuation for a second pair of transfer regions, the second pair of transfer regions including a third transfer region that is a boundary of the third acoustic environment and the second transfer region; and generate a first audio component by rendering the second audio source according to the combined energy transfer attenuation.
[0372] Embodiment F. The audio device according to any one of the preceding Embodiments A to E, wherein the second acoustic attenuation boundary is the boundary between the second acoustic environment and the third acoustic environment, the energy attenuation indicates the attenuation of the audio in the second acoustic environment between the first transfer region and the second transfer region, and the energy transfer parameter for the first pair of transfer regions further includes a second energy attenuation, the second energy attenuation indicating the attenuation of the audio in the third acoustic environment between the first transfer region and the second transfer region; and the renderer (507) is arranged to generate a second audio component by rendering the second audio source of the third acoustic environment according to the second energy attenuation.
[0373] Embodiment G. The audio device according to any one of the preceding Embodiments A to F, wherein the metadata includes energy attenuation parameters only for pairs of transfer regions that are boundaries of a shared acoustic environment.
[0374] Embodiment H. The audio device according to any one of the preceding Embodiments A to G, wherein the metadata further includes an acoustic property indication for at least the first transfer region, the acoustic property indication indicating the acoustic influence of the first transfer region on the sound passing through the first transfer region; and wherein the renderer (507) is arranged to generate a first audio component according to the acoustic property indication.
[0375] Embodiment I. An audio device according to any one of the foregoing embodiments A to H, wherein the energy attenuation for the first pair of transmission regions also indicates the proportion of the audio energy in the second transmission region that propagates into the first transmission region; and the renderer (507) is arranged to render an audio signal for a second listening position in a second acoustic environment, the rendering including generating a second audio component by rendering a second audio source in the first acoustic environment according to the energy transfer attenuation for the first pair of transmission regions.
[0376] Embodiment J. An audio device according to any one of the foregoing embodiments A to I, wherein the renderer (507) is arranged to generate a first audio source by combining audio from a plurality of audio sources in a second acoustic environment.
[0377] Embodiment K. An audio device according to any one of the foregoing embodiments A to L, wherein the metadata includes:
[0378] data describing the position of at least the second transmission region; and
[0379] an energy transfer indication for the second transmission region, the energy transfer indication indicating the proportion of the energy in an omnidirectional point audio source at a reference position that will reach the second transmission region, the reference position being a relative position with respect to the second transmission region;
[0380] and wherein the renderer (507) is arranged to determine the audio energy level of the second transmission region for audio from the first audio source in response to the position of the first audio source relative to the reference position; and to adjust the level of the first signal component according to the energy transfer attenuation and the audio energy level for a pair of transmission regions, the pair of transmission regions including the first transmission region and the second transmission region.
[0381] Embodiment L. An audio device according to any one of the foregoing embodiments A to K, wherein the metadata includes a coupling coefficient for the first transmission region, and the renderer (507) is arranged to render the first audio component as an audio source originating from a position near the first transmission region according to the coupling factor.
[0382] Embodiment M. An audio device according to any one of the foregoing embodiments A to M, wherein the renderer (507) is arranged to render the first audio component as a reverberant audio component of the first acoustic environment.
[0383] Embodiment N. An audio device according to any one of the foregoing embodiments A to M, wherein the renderer (507) is arranged to render the first audio component as a non-direct audio component.
[0384] Embodiment O. A method of rendering an audio signal, the method comprising:
[0385] Receiving audio data for an audio source for a scenario including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0386] Receiving metadata for the audio data, the metadata including:
[0387] Transmission region data that describes transmission regions in the acoustic attenuation boundary, each transmission region being a region of the acoustic attenuation boundary that has a lower attenuation than an average attenuation of the acoustic attenuation boundary outside the transmission region; and
[0388] Energy transfer parameters, each energy transfer parameter indicating an energy attenuation between a pair of transmission regions, the energy attenuation for a pair of transmission regions indicating a proportion of audio energy propagating from one transmission region of the pair of transmission regions to the other transmission region of the pair of transmission regions;
[0389] Rendering an audio signal for a listening position in a first acoustic environment, the rendering including generating a first audio component by rendering a first audio source of a second acoustic environment according to an energy attenuation of a first pair of transmission regions, the first pair of transmission regions including a first transmission region of a first acoustic attenuation boundary that is a boundary of the first acoustic environment and a second transmission region of a second acoustic attenuation boundary that is a boundary of the second acoustic environment.
[0390] Example P. An audio data signal, comprising:
[0391] Audio data for an audio source for a scenario including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries;
[0392] Metadata for the audio data, the metadata including:
[0393] Transmission region data that describes transmission regions in the acoustic attenuation boundary, each transmission region being a region of the acoustic attenuation boundary that has a lower attenuation than an average attenuation of the acoustic attenuation boundary outside the transmission region; and
[0394] Energy transfer parameters, each energy transfer parameter indicating an energy attenuation between a pair of transmission regions, the energy attenuation for a pair of transmission regions indicating a proportion of audio energy propagating from one transmission region of the pair of transmission regions to the other transmission region of the pair of transmission regions.
[0395] Example Q. An audio device arranged to generate an audio data signal according to Example P.
[0396] Embodiment R, a computer program product comprising computer program code modules which, when the program is run on a computer, are adapted to carry out all the steps of Embodiment O.
[0397] More specifically, the invention is defined by the claims.
Claims
1. An audio device, comprising: A first receiver (501) arranged to receive audio data of an audio source for a scene comprising a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries; A second receiver (501) arranged to receive metadata for the audio data, the metadata comprising: An indication of the position of at least a first transmission region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region of the first acoustic attenuation boundary having a lower attenuation than the average attenuation of the first acoustic attenuation boundary outside the transmission region; A renderer (507) arranged to render an audio signal for a listening position in the first acoustic environment, the rendering comprising generating a first audio component by rendering the audio data of a first audio source for the second acoustic environment and adjusting the level of the first audio component according to an energy transfer indication, wherein the energy transfer indication is for the first transmission region, the energy transfer indication indicating the proportion of the energy at the first transmission region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transmission region, and wherein the energy transfer indication is provided in the metadata or provided by a process performed elsewhere where computational complexity is available.
2. The audio device according to claim 1, wherein, The renderer (507) is arranged to adjust the level of the first audio component in response to the position of the first audio source relative to the reference position.
3. The audio device according to claim 1 or 2, wherein, The renderer (507) is arranged to adjust the level of the first audio component according to the difference between a reference distance from the reference position to the first transmission region and a distance from the position of the first audio source to the first transmission region.
4. The audio device according to any one of the preceding claims, wherein, The renderer (507) is arranged to adjust the level of the first audio component according to the difference between the direction from the reference position to the first transmission region and the direction from the first audio source to the first transmission region.
5. The audio device according to any one of the preceding claims, wherein, The metadata further comprises data describing the directivity of the sound radiation from the first audio source, and the renderer (507) is arranged to adjust the level of the first audio component according to the directivity.
6. The audio device according to any one of the preceding claims, wherein, The renderer (507) is arranged to scale the level of the first audio component according to a relative directivity gain of the first audio source in the direction from the first audio source to the first transmission region, the relative directivity gain indicating a gain relative to an omnidirectional source.
7. The audio device according to any one of the preceding claims, wherein, The first audio source represents audio arriving at the second acoustic environment from a third acoustic environment via a second transmission region of a second boundary that separates the third acoustic environment from the second acoustic environment.
8. The audio device according to any one of the preceding claims, wherein, The metadata further comprises: Energy transfer parameters, each energy transfer parameter indicating the energy attenuation between a pair of transmission regions, the energy attenuation for a pair of transmission regions indicating the proportion of the audio energy at one transmission region of the pair that propagates to the other transmission region of the pair; And the renderer (507) is arranged to render a second audio component of the audio signal by rendering a second audio source of a third acoustic environment according to an energy attenuation for a pair of transmission regions, the pair of transmission regions including a transmission region of an acoustic attenuation boundary of the first acoustic environment and a transmission region of a second acoustic attenuation boundary, the second acoustic attenuation boundary being a boundary of the second acoustic environment.
9. The audio device according to any one of the preceding claims, wherein, The metadata includes a coupling coefficient for the first transmission region, and the renderer (507) is arranged to render the first audio component as an audio source originating from a position near the first transmission region according to the coupling coefficient.
10. The audio device according to any one of the preceding claims, wherein, The renderer (507) is arranged to render the first audio component as a reverberant audio component of the first acoustic environment.
11. The audio device according to any one of the preceding claims, wherein, The renderer (507) is arranged to generate a reverberant audio signal for the second environment to include a first component from the first audio source, the renderer (507) is further arranged to determine an energy loss estimate as a proportion of the energy in the first audio source that reaches the first transmission region, the proportion of the energy being determined according to the energy transfer indication and the position of the first audio source relative to the reference position; and the renderer (507) is arranged to reduce the level of the reverberant audio signal by an amount depending on the energy loss estimate.
12. The audio device according to any one of the preceding claims, wherein, The energy transfer indication reflects a proportion of a sphere covered by the first transmission region, the sphere being centered at the reference position and having a radius corresponding to the distance from the reference position to the first transmission region.
13. The audio device according to any one of the preceding claims, wherein, The renderer (507) is arranged to render the first audio component as a non-direct audio component.
14. A method of rendering an audio signal, the method comprising: Receiving audio data for an audio source of a scene including a plurality of acoustic environments, the acoustic environments being delimited by acoustic attenuation boundaries; Receiving metadata for the audio data, the metadata including: A position indication for at least a first transmission region of a first acoustic attenuation boundary between a first acoustic environment and a second acoustic environment, the first transmission region being a region of the first acoustic attenuation boundary having a lower attenuation than an average attenuation of the first acoustic attenuation boundary outside the transmission region; Rendering an audio signal for a listening position in the first acoustic environment, the rendering including generating a first audio component by rendering audio data of a first audio source for the second acoustic environment, and adjusting a level of the first audio component according to an energy transfer indication, wherein the energy transfer indication is for the first transmission region, the energy transfer indication indicates a proportion of the energy at the first transmission region from an omnidirectional point audio source at a reference position, the reference position being a relative position with respect to the first transmission region, and wherein the energy transfer indication is provided in the metadata or provided by a process performed elsewhere where computational complexity is available.
15. A computer program product comprising computer program code modules which, when the program is run on a computer, are adapted to carry out all the steps according to claim 14.