Audio device and method of operating the same
The audio device and method improve audio rendering in multi-room environments by using reverberation parameter sets and energy transfer parameters to enhance realism and reduce complexity, addressing the limitations of existing technologies in handling multiple acoustic environments.
Patent Information
- Application Number
- JP2025540079
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-09
- Filing Date
- 2024-01-02
- Publication Date
- 2026-01-14
AI Technical Summary
Existing audio rendering technologies struggle to provide an optimal and realistic audio experience in scenarios involving multiple acoustic environments, leading to misalignment between visual and audio perceptions and high computational complexity.
An audio device and method that utilizes a first receiver for audio data, a second receiver for metadata including reverberation parameter sets and energy transfer parameters, and a selector to determine the appropriate reverberation parameters for each acoustic environment, reducing complexity and improving audio quality.
Enhances the realism and naturalness of audio rendering in multi-room scenes with reduced computational resources and complexity, providing a more accurate and efficient audio experience.
Smart Images

Figure 2026501417000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to audio devices and methods of operation thereof, particularly but not exclusively to audio devices and methods of operation thereof for rendering audio for multi-room scenes, for example as part of an extended reality experience. [Background technology]
[0002] In recent years, the variety and range of experiences based on audiovisual content has expanded significantly, and new services and methods for using and consuming such content are continually being developed and introduced. In particular, many spatial and interactive services, applications, and experiences are being developed to provide users with more complex and immersive experiences.
[0003] An example of such an application is Extended Reality (XR), a common term for Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) applications that are rapidly becoming mainstream, with numerous solutions for the consumer market. There are also several standards under development by a number of standards bodies, which are actively working to standardize various aspects of VR / AR / MR systems, including streaming, broadcasting, and rendering.
[0004] While VR applications tend to provide a user experience corresponding to the user being in a different world / environment / scene, AR (including mixed reality) applications tend to provide a user experience corresponding to the user being in a current environment but with additional information or virtual objects or information added. Thus, VR applications tend to provide synthetically generated worlds / scenes that are fully immersive, while AR applications tend to provide partially synthetic worlds / scenes that are overlaid on the real scene in which the user is physically present. However, these terms are often used interchangeably and there is a high degree of overlap. In the following, the term extended reality (XR) is used to refer to both virtual reality and augmented / mixed reality.
[0005] As an example, an increasingly popular service is the presentation of images and audio in a manner that allows the user to actively and dynamically interact with the system to change the parameters of the rendering, which in turn adapts to movements and changes in the user's position and orientation. A very attractive feature in many applications is the ability to change the viewer's effective viewing position and direction, which allows, for example, the viewer to move or "look around" within the scene being presented.
[0006] Such features can in particular provide a virtual reality experience for the user, allowing the user to move around (relatively) freely in a virtual scene and dynamically change their position and their point of view. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide the specific view requested. This approach is well known from gaming applications for computers and consoles, for example in the category of first-person shooters.
[0007] Also, particularly in virtual reality applications, it is desirable that the images presented are three-dimensional, typically presented using a stereoscopic display. Indeed, to optimize viewer immersion, it is generally preferred that the user experience the presented scene as a three-dimensional scene. Indeed, a virtual reality experience should preferably allow the user to select their own position, viewpoint, and moment relative to the virtual world.
[0008] In addition to the visual rendering, most XR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience, where the audio sources are perceived as arriving from positions corresponding to the positions of corresponding objects in the visual scene. Thus, the audio and video scenes are preferably perceived as coherent, providing a complete spatial experience together.
[0009] Many immersive experiences are provided by virtual audio scenes generated by headphone playback using, for example, binaural audio rendering techniques. In many scenarios, such headphone playback is based on head tracking so that the rendering can react to the user's head movements, greatly enhancing the sense of immersion.
[0010] An important aspect for many applications is how to generate and deliver audio that provides a natural and realistic perception of the audio scene. For example, when generating audio for virtual reality applications, it is important that not only are desirable audio sources generated, but that the audio sources are generated in a way that provides a realistic perception of the audio environment, including attenuation, reflections, coloration, etc.
[0011] In room / ambient acoustics, reflections of sound waves from walls, floors, ceilings, objects, etc. cause delayed and attenuated versions of the source signal (usually frequency dependent) to reach the listener (i.e. the user of an XR system) via different paths. The combined effect is modeled by an impulse response called the Room Impulse Response (RIR).
[0012] RIR typically consists of a direct sound component, which depends on the distance from the source to the listener, followed by a reflected component that characterizes the room's acoustics, as shown in Figure 1. The size and shape of the room, the location of the source and listener within the room, and the reflective properties of the room's surfaces all affect the characteristics of this reverberant component.
[0013] The reflection part can be decomposed into two time domains that usually overlap. The first domain contains the so-called early reflections, which represent the separate reflections of the sound source off the walls and obstacles of the room before reaching the listener. As the time lag / (propagation) delay increases, the number of reflections that exist within a fixed time interval increases, and the path may contain second or higher order reflections (for example, it may reflect off multiple walls or both a wall and the ceiling).
[0014] The second region, called the reverberant part, is where the density of these reflections increases to the point where they can no longer be separated by the human brain. This region is usually called diffuse reverberation, late reverberation, or reverberant tail, or simply reverberation.
[0015] RIRs contain clues that provide the auditory system with information about the distance of a sound source, the size and acoustic properties of a room. The energy in the reverberant part, relative to the energy in the anechoic part, strongly influences the perceived distance of a sound source. The level and delay of the earliest reflections can provide clues about how close a sound source is to a wall, and anthropometric filtering can enhance the assessment of specific walls, floors, or ceilings.
[0016] The density of (early) reflections contributes to the perceived size of a room. The time it takes for the reflections to drop in energy level by 60 dB is called the reverberation time, T 60Reverberation time is often used as a measure of how quickly reflections die away in a room. Reverberation time provides information about the acoustic properties of a room, specifically whether the walls are highly reflective (e.g., a bathroom) or highly sound absorbing (e.g., a bedroom with furniture, carpets, and curtains).
[0017] Furthermore, the RIR depends on the anthropometric characteristics of the user, as when it is part of the binaural room impulse response (BRIR), the RIR is filtered by the head, ears, and shoulders, i.e., the head-related impulse response (HRIR).
[0018] Because late reverberation reflections cannot be distinguished or separated by the listener, they are often simulated or represented parametrically using parametric reverberators that use feedback delay networks, such as the well-known Jot reverberator.
[0019] In early reflections, the incidence direction and distance-dependent delays are important cues that allow humans to extract information about the relative position of rooms and sound sources. Therefore, the simulation of early reflections must be clearer than late reverberation. Therefore, in efficient acoustic rendering algorithms, early reflections are simulated separately from late reverberation. A well-known method for early reflections is to mirror the sound sources at each room boundary to generate virtual sources that represent the reflections.
[0020] While early reflections involve the position of the user and sound source relative to the room boundaries (walls, ceiling, floor), late reverberation is a diffuse acoustic response of the room that tends to be more homogeneous throughout the room, which makes simulating late reverberation often more computationally efficient than early reflections.
[0021] Two main properties of late reverberation are the slope and amplitude of the impulse response for a time above a given threshold. These properties tend to be strongly frequency dependent in natural rooms. Reverberation is often described using parameters that characterize these properties.
[0022] An example of parameters characterizing reverberation is shown in Figure 2. An example of parameters conventionally used to describe the slope and amplitude of the impulse response corresponding to diffuse reverberation is the well-known T 60 These include amplitude level and reverberation level / energy. Recently, other indicators of amplitude level have been proposed, specifically parameters that indicate the ratio between diffuse reverberation energy and total source energy.
[0023] Specifically, the Diffuse to Source Ratio (DSR) can be used to express the amount of diffuse reverberant energy or the level of a sound source received by a user as a ratio of the total emitted energy of that sound source. DSR can specifically express the ratio between emitted sound source energy and diffuse reverberant characteristics, such as the energy or (initial) level of the diffuse reverberant signal:
number
[0024] Therefore, this is called DSR (Diffusion to Source Ratio).
[0025] Such known approaches tend to provide an efficient description of audio propagation in a room and also tend to lead to a rendering of audio that is perceived as natural for the room in which the listener is (virtually) present. Summary of the Invention [Problem to be solved by the invention]
[0026] However, while traditional approaches to representing and rendering sound in a room or individual acoustic environment may provide an appropriate perception in many embodiments, they tend not to be fully suited to all possible scenarios. In particular, in audio scenes that may include different acoustic environments / areas / rooms, audio signals generated using, for example, the described reverberation approaches may not provide an optimal experience or perception. This typically leads to situations where audio from other rooms is not adequately or accurately represented by the rendered audio, resulting in a perception that may not fully reflect the acoustic scenario and scene.
[0027] In practice, reverberation is typically modeled for a listener within a room, taking into account the characteristics of the room. When the listener is outside the room or in a different room, the reverberator may be turned off or reconfigured to suit the characteristics of the other room. Even if multiple reverberators can be operated in parallel, the output of the reverberators is typically a diffuse binaural (or multi-speaker) signal intended to be presented to the listener as if they were in the room. However, such an approach tends to produce audio that is often not perceived as an accurate representation of the real environment. This can result, for example, in a perceived misalignment or inconsistency between the visual perception of a scene and the associated audio being rendered.
[0028] In particular, in many scenarios, a scene contains multiple environments, and sounds from one environment may be heard in other environments. For example, in a multi-room scene, audio from other rooms can often be heard in the room where the listener is located (referred to as the listener room / acoustic environment or listening room / acoustic environment). Therefore, it is desirable to render audio to the listener that includes a representation of audio from other acoustic environments. In particular, it is often desirable to render audio that represents diffuse / reverberant audio from other acoustic environments / rooms, as this tends to result in a more realistic representation and perception. However, to provide a more realistic user perception, such rendering should reflect the acoustic characteristics of the source room rather than (or sometimes in addition to) the acoustic characteristics of the listening room.
[0029] Typical approaches to rendering audio tend to be suboptimal in some scenarios, especially when rendering audio for scenes that include different acoustic rooms or environments.
[0030] A particular problem in many scenarios is that audio rendering tends to be complex and time- and resource-intensive. In particular, when audio from multiple acoustic environments needs to be rendered and the acoustic characteristics of these acoustic environments need to be represented, multiple rendering passes are often required, resulting in high complexity and resource usage. For example, including reverberant audio from rooms other than the listening room typically requires additional reverberators to generate the relevant reverberant sounds. The trade-off between audio quality / listener perception and complexity / resource requirements is an important trade-off in many systems.
[0031] Therefore, improved approaches for rendering audio for scenes with multiple acoustic environments would be advantageous, particularly approaches that enable improved usability, increased flexibility, reduced complexity, easier implementation, improved audio experience, improved audio quality, reduced computational load, improved representation of multiple acoustic environments, improved resource usage / allocation, easier rendering, improved rendering of audio from multiple acoustic environments, improved performance for virtual / mixed / enhanced reality applications, increased processing flexibility, more natural-sounding audio rendering, improved audio rendering for multi-room scenes, improved trade-off between perceived audio quality / realism and computational load / complexity, and / or improved performance and / or usability.
[0032] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]
[0033] According to one aspect of the present invention, there is provided an audio device comprising: a first receiver for receiving audio data of an audio source of a scene including a plurality of acoustic environments separated by acoustic attenuation boundaries; a second receiver for receiving metadata of the audio data, the metadata including a plurality of reverberation parameter sets and at least one energy transfer parameter, each reverberation parameter set for one associated acoustic environment including at least one reverberation parameter indicative of a relationship between a level of reverberation in the one associated acoustic environment and a level of an audio source in the one associated acoustic environment, each energy transfer parameter indicative of an energy attenuation of audio propagation between a source acoustic environment and a destination acoustic environment; and a second receiver for receiving, for the first acoustic environment, at least a first reverberation parameter set of the plurality of acoustic environments. the audio source for the first acoustic environment based on the audio data, the audio signal including an audio component derived from the reverberant audio signal.
[0034] This approach allows for the generation of audio signals that improve the user experience of audio scenes with multiple acoustic environments, often resulting in a more realistic and natural-sounding audio experience. This approach can improve the audio rendering of, for example, multi-room scenes. In many scenarios, the audio of a scene can be perceived as more natural and / or accurate.
[0035] This approach can enhance and / or facilitate the rendering of audio representative of audio sources in other acoustic environments or rooms, and this rendering of audio signals often has reduced complexity and reduced computational resource requirements.
[0036] This approach can enhance, increase, and / or facilitate flexibility and / or adaptation of processed and / or rendered audio.
[0037] An energy transfer parameter indicating energy attenuation between a pair of transfer regions may be equivalent to an energy transfer parameter indicating energy transfer between a pair of transfer regions. An increase in attenuation indicates a decrease in the rate of audio energy reaching the destination acoustic environment from the source acoustic environment, which corresponds to a decrease in energy transfer. A decrease in attenuation indicates an increase in the rate of audio energy reaching the destination acoustic environment from the source acoustic environment, which corresponds to an increase in energy transfer. Thus, the terms "energy attenuation" and "energy transfer" can be used interchangeably, with the understanding that one is a monotonically decreasing function of the other.
[0038] Audio energy / level (or simply energy / level) can be specifically represented by level, amplitude, power, or time-averaged energy measures.
[0039] An acoustically attenuating boundary attenuates the propagation of sound through the boundary from one acoustic environment to another. In many embodiments and scenarios, the attenuation of the acoustically attenuating boundary outside the transmission region is 3 dB, 6 dB, 10 dB or more, or even 20 dB or more.
[0040] In some embodiments, the energy transfer parameter indicates the energy attenuation of audio propagation between the transmission area of the acoustic attenuation boundary of the source acoustic environment and the transmission area of the acoustic attenuation boundary of the destination acoustic environment. In many embodiments, the attenuation of the transmission area at the acoustic attenuation boundary is 3 dB, 6 dB, 10 dB or more, or even 20 dB or more lower than the attenuation (on average) of the acoustic attenuation boundary outside the transmission area.
[0041] The selector selects multiple parameter sets, and in many scenarios / embodiments, selects for multiple acoustic environments other than the first acoustic environment. In some embodiments, the selector generates rank values for the multiple reverberation parameter sets and selects at least the first reverberation parameter set in response to the rank values. The rank value for a given reverberation parameter set is determined in response to parameter values for the given reverberation parameter set and an energy transfer parameter indicative of energy attenuation between the acoustic environment in which the reverberation parameter set is provided and the first acoustic environment.
[0042] A reverberation parameter set for a given acoustic environment includes one or more parameters that describe the characteristics of reverberation in the given acoustic environment resulting from one or more audio sources in the given acoustic environment. The reverberation parameter set may include parameters that describe the level / energy attenuation / transformation of diffuse / reverberant audio from the audio source to level / energy, the delay of the diffuse / reverberant audio, and / or the decay time of the diffuse / reverberant audio.
[0043] According to an optional feature of the invention, the selector selects between the reverberation parameter sets based on a composite attenuation for each reverberation parameter set, the composite attenuation for a given reverberation parameter set indicating a level of audio in the first acoustic environment resulting from reverberation in the given acoustic environment of the given reverberation parameter set resulting from reverberation in the given acoustic environment caused by an audio source in the given acoustic environment, the composite attenuation for the given reverberation parameter set including a contribution from the attenuation indicated by at least one reverberation parameter of the given reverberation parameter set and an energy transfer parameter indicating an energy attenuation of audio propagation between the given acoustic environment and the first acoustic environment.
[0044] This provides improved performance or ease of implementation in many scenarios; helps improve the user experience when rendering audio in scenes with multiple acoustic environments; provides practical and / or improved selection of which reverberations to render in many embodiments; and provides a better trade-off between complexity and resource usage versus audio quality and user experience in many scenarios.
[0045] In accordance with an optional feature of the invention, the selector selects the first reverberation parameter set in preference to the second reverberation parameter set if the composite attenuation of the second reverberation parameter set exceeds the composite attenuation of the first reverberation parameter set.
[0046] This results in improved performance or ease of implementation in many scenarios, helps improve the user experience when rendering audio in scenes with multiple acoustic environments, and often improves utilization of limited computational resources.
[0047] In accordance with an optional feature of the invention, the selector discards a reverberation parameter set from the selection if the combined attenuation exceeds a threshold.
[0048] This results in improved performance or ease of implementation in many scenarios, helps improve the user experience when rendering audio in scenes with multiple acoustic environments, and often improves utilization of limited computational resources.
[0049] According to an optional feature of the invention, the metadata further includes audio level data indicative of a signal level of the audio source, and the selector determines a resultant signal level in a first acoustic environment of the audio source in the acoustic environment of the reverberation parameter set based on the signal level of the audio source and the combined attenuation of the reverberation parameter set, and selects the at least one reverberation parameter set in response to the resultant signal level in the first acoustic environment.
[0050] This provides improved selection and particularly advantageous performance and / or operation in many embodiments and scenarios.
[0051] The selector determines a resulting signal level in the first acoustic environment for a given reverberation parameter set in response to an audio signal level of at least one audio source in the given reverberation environment for a given reverberation parameter set and a composite attenuation for the given reverberation parameter set. The selector may determine resulting signal levels for multiple reverberation parameter sets, and possibly all reverberation parameter sets.
[0052] According to an optional feature of the invention, the selector determines a resulting signal level for the reverberation parameter set for the at least one acoustic environment by combining levels of multiple sound sources in the at least one acoustic environment into a composite audio source level and applying a composite attenuation to the composite audio source level.
[0053] This provides better performance and ease of implementation in many scenarios, and helps improve the user experience when rendering audio in scenes with multiple acoustic environments.
[0054] According to an optional feature of the invention, the selector determines a number of reverberators available for generating a reverberant signal of an audio source in an acoustic environment other than the first acoustic environment, and adapts the number of reverberation parameter sets selected to match the number of reverberators available.
[0055] This improves performance and eases implementation in many scenarios, helps improve the user experience when rendering audio in scenes with multiple acoustic environments, and allows for flexible and dynamic resource adaptation.
[0056] In accordance with an optional feature of the invention, the selector determines a first number of reverberant signal sources in the first acoustic environment and subtracts the first number from the number of reverberators of the audio device to determine the number of available reverberators.
[0057] This improves performance and eases implementation in many scenarios, helps improve the user experience when rendering audio in scenes with multiple acoustic environments, and allows for flexible and dynamic resource adaptation.
[0058] In accordance with an optional feature of the invention, at least one of the at least one reverberation parameter and the energy transfer parameter is frequency dependent.
[0059] This provides better performance and easier implementation in many scenarios.
[0060] According to an optional feature of the invention, the selector performs an iterative selection of at least one reverberation parameter set.
[0061] This improves performance and eases implementation in many scenarios, and helps improve the user experience when rendering audio in scenes with multiple acoustic environments. Reselection is performed with intervals between selections that do not exceed (on average) 1, 2, 5, 10, or 20 seconds in different embodiments.
[0062] In accordance with an optional feature of the invention, at least a first reverberator of the plurality of reverberators continues to generate the reverberated audio signal of the first reverberation parameter set for a period of time after deselection of the first reverberation parameter set.
[0063] This offers particular performance advantages and ease of implementation in many scenarios, and helps improve the user experience when rendering audio in scenes with multiple acoustic environments, in particular by reducing or mitigating the perceptual impact of switching / reselection effects.
[0064] The level of the reverberant audio signal is gradually decreased during said period. The reverberant audio signal fades out during said period.
[0065] In accordance with an optional feature of the invention, the first reverberator uses less computational resources than a reverberator of the plurality of reverberators that was used to generate the reverberated audio signal of the first reverberation parameter set before deselection of the first reverberation parameter set.
[0066] This provides particularly advantageous operation and / or performance in many embodiments.
[0067] According to an optional feature of the invention, the at least one reverberation parameter includes a diffuse-to-source ratio parameter indicative of the ratio between emitted audio source energy and diffuse reverberation energy.
[0068] This provides particularly advantageous operation and / or performance in many embodiments.
[0069] According to one aspect of the present invention, there is provided a method of operating an audio device, the method comprising the steps of receiving audio data for an audio source of a scene including a plurality of acoustic environments separated by an acoustic attenuation boundary, receiving metadata for the audio data, the metadata including a plurality of reverberation parameter sets and at least one energy transfer parameter, each reverberation parameter set for one associated acoustic environment including at least one reverberation parameter indicative of a relationship between a level of reverberation in the one associated acoustic environment and a level of an audio source in the one associated acoustic environment, each energy transfer parameter indicative of an energy attenuation of audio propagation between a source acoustic environment and a destination acoustic environment, and selecting at least a first reverberation parameter set of the plurality of acoustic environments for a first acoustic environment. the first reverberation parameter set is for a different acoustic environment than the first acoustic environment, and the selection is made in response to parameter values of the first reverberation parameter set and an energy transfer parameter indicative of energy attenuation of audio propagation between the different acoustic environment and the first acoustic environment; generating reverberant audio signals with a plurality of reverberators (509, 511), where at least one reverberator generates the reverberant audio signal based on the first reverberation parameter set and an audio source of the different acoustic environment of the first reverberation parameter set; and generating an audio signal for the first acoustic environment, where the audio signal includes audio components obtained from the reverberant audio signal.
[0070] According to one aspect of the present invention, there is provided an audio data signal comprising audio data of an audio source of a scene comprising a plurality of acoustic environments separated by acoustic attenuation boundaries, and audio data metadata, the metadata comprising a plurality of reverberation parameter sets and at least one energy transfer parameter, each reverberation parameter set for one associated acoustic environment comprising at least one reverberation parameter indicative of a relationship between a level of reverberation in the one associated acoustic environment and a level of an audio source in the one associated acoustic environment, and each energy transfer parameter indicative of an energy attenuation of audio propagation between a source acoustic environment and a destination acoustic environment.
[0071] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]
[0072] Embodiments of the invention will now be described, by way of example only, with reference to the drawings in which:
[0073] [Figure 1] FIG. 1 shows an example of a room impulse response. [Figure 2] FIG. 2 shows an example of a room impulse response. [Figure 3] FIG. 3 shows an example of elements of a virtual reality system. [Figure 4] FIG. 4 shows an example of a scene with three rooms. [Figure 5] FIG. 5 illustrates an example of an audio device for generating an audio signal in accordance with some embodiments of the present invention. [Figure 6] FIG. 6 shows an example of the Jot reverberator. [Figure 7] FIG. 7 shows an example of a feedback delay network reverberator. [Figure 8] FIG. 8 shows an example of a scene with multiple rooms separated by walls with sound doorways. [Figure 9] FIG. 9 shows an example of sound attenuation for sound propagation between different acoustic environments. [Figure 10] FIG. 10 is an example of sound attenuation for sound propagation between different acoustic environments. [Figure 11] FIG. 11 illustrates some elements of a possible arrangement of processors for implementing elements of apparatus according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0074] Although the following description focuses on audio processing and rendering for augmented reality applications, it will be appreciated that the principles and concepts described may be used in many other applications and embodiments.
[0075] Virtual experiences that allow users to move around in virtual worlds are becoming increasingly popular, and services are being developed to meet this demand.
[0076] In some systems, the VR application is provided locally to the viewer, such as by a standalone device that does not use or even have access to any remote VR data or processing. For example, a device such as a game console may include a store for storing scene data, an input for receiving / generating viewer poses, and a processor for generating corresponding images from the scene data.
[0077] In other systems, the VR application can be implemented and executed remotely by the viewer. For example, a device local to the user detects / receives movement / pose data, which is transmitted to a remote device that processes the data and generates a viewer pose. The remote device then generates a view image and corresponding audio signal appropriate for the user pose based on scene data describing the scene. The view image and corresponding audio signal are then transmitted to the viewer's local device for presentation there. For example, the remote device directly generates a video stream (typically a stereoscopic / 3D video stream) and corresponding audio stream, which are directly presented by the local device. Thus, in such an example, the local device does not perform any VR processing other than transmitting movement data and presenting the received video data.
[0078] In many systems, functionality is distributed between a local device and a remote device. For example, the local device processes received input data and sensor data to generate a user pose, which is continuously transmitted to the remote VR device. The remote VR device then generates a corresponding view image and a corresponding audio signal and transmits them to the local device for presentation. In other systems, the remote VR device may not directly generate the view image and the corresponding audio signal, but may instead select and transmit relevant scene data to the local device, which then generates the view image and the corresponding audio signal to be presented to the local device. For example, the remote VR device may identify the nearest capture point and extract the corresponding scene data (e.g., a set of object sources and their position metadata) and transmit it to the local device. The local device then processes the received scene data to generate an image and an audio signal for a particular current user pose. A user pose typically corresponds to a head pose, and a reference to a user pose is usually considered equivalent to a reference to a head pose.
[0079] In many applications, particularly broadcast services, a source transmits or streams scene data in the form of images (including video) and audio representations of a scene that are independent of the user pose. For example, signals and metadata corresponding to audio sources in a particular virtual room are transmitted or streamed to multiple clients. Each client then locally synthesizes an audio signal corresponding to the current user pose. Similarly, the source transmits a general description of the audio environment. This general description describes the audio sources in the environment and the acoustic properties of the environment. An audio representation is then generated locally and presented to the user, for example using binaural rendering and processing.
[0080] 3 shows an example of a VR system in which a remote VR client device 301 interacts with a VR server 303 over a network 305, such as the Internet. The server 303 can be configured to support multiple client devices 301 simultaneously.
[0081] For example, the VR Server 303 supports a broadcast experience by transmitting an image signal containing an image representation in the form of image data that a client device can use to locally synthesize a view image corresponding to an appropriate user pose (pose refers to position and / or orientation). Similarly, the VR Server 303 can transmit an audio representation of a scene and locally synthesize the audio to match the user pose. Specifically, as the user moves through the virtual environment, the synthesized images and audio presented to the user are updated to reflect the user's current (virtual) position and orientation in the (virtual) environment.
[0082] Therefore, in many applications, such as that of Figure 3, it is desirable to model a scene and generate efficient image and audio representations that can be efficiently included in a data signal that is transmitted or streamed to various devices that can locally synthesize views and audio for poses different from the capture pose.
[0083] In some embodiments, for example, models representing a scene are stored locally and used locally to synthesize appropriate images and audio. For example, an audio model of a room includes indications of the acoustic properties of the room as well as the properties of the audio sources heard in the room. The model data can then be used to synthesize audio appropriate for a particular location.
[0084] In many scenarios, a scene includes multiple different acoustic environments or regions with different acoustic characteristics, particularly different reverberation characteristics. Specifically, a scene includes or is divided into different acoustic environments / regions, each with homogeneous reverberation but with different reverberation characteristics among them. For all locations within an acoustic environment / region, the reverberation components of the audio received at that location are homogeneous, specifically substantially the same (except for possible gain differences). An acoustic environment / region may be a set of locations where the reverberation components of the audio are homogeneous. An acoustic environment / region may be a set of locations where the reverberation components of the audio propagation impulse responses for audio sources in that acoustic environment are homogeneous. Specifically, an acoustic environment / region may be a set of locations where the reverberation components of the audio propagation impulse responses for audio sources in that acoustic environment have the same frequency-dependent slope and / or amplitude characteristics, possibly except for possible gain differences. Specifically, an acoustic environment / region may be a collection of locations where the reverberant components of the audio propagation impulse responses for audio sources in that acoustic environment are the same, possibly except for gain differences.
[0085] An acoustic environment / area may typically be a collection of locations (usually a 2D or 3D area) with the same rendering reverberation parameters. The reverberation parameters used to render the reverberant components are the same for all locations within one acoustic environment / area. In particular, the same reverberation decay parameters (e.g., T 60 ) and diffuse-to-source ratio (DSR) are applied to all positions within an acoustic environment / area.
[0086] Impulse responses can differ between different locations within one room / acoustic environment / area due to their "noisy" nature caused by the many different reflections of different orders that cause reverberation. However, even in such cases, the frequency-dependent slope and / or amplitude characteristics are the same (except for possible gain differences), especially when expressed e.g. as reverberation time (T60) or reverberation coloration.
[0087] In many scenarios, acoustic environments are separated by acoustically attenuating boundaries. Indeed, in many scenarios, different acoustic environments are determined by the presence of acoustically attenuating boundaries. An acoustically attenuating boundary divides an area into different acoustic environments, which are formed by the presence of one or more acoustically attenuating boundaries. An acoustically attenuating boundary creates two acoustic environments, and these two acoustic environments are on either side of the acoustically attenuating boundary. Such acoustically attenuating boundaries can be formed, for example, by walls that divide a space into multiple acoustic environments, or by any other structure that provides sound attenuation.
[0088] An acoustic environment / area may also be referred to as an acoustic room or simply as a room. A room is considered an environment / area as described above.
[0089] In many embodiments, scenes are provided in which the acoustic rooms correspond to different virtual or real rooms that a user can move through (e.g., virtually). Figure 4 shows an example of a scene with three rooms, A, B, and C. In this example, the user can move between the three rooms or out of the rooms through doorways or openings.
[0090] For a room to have substantial reverberant properties, it tends to represent a spatial region that is sufficiently surrounded by geometric surfaces with fully or partially reflective properties such that a large proportion of the reflections within the room continue to reflect back into this region, producing in this region a diffuse field of reflections with no significant directional properties. The geometric surfaces do not need to be aligned with any visual elements.
[0091] Audio rendering that aims to provide a natural and realistic effect to the listener typically includes the rendering of the acoustic scene. In many environments, this includes the representation and rendering of the diffuse reverberation present in the environment, such as the room the listener is in. The rendering and representation of such diffuse reverberation has been found to have a significant impact on the perception of the environment, including whether the audio is perceived as representing a natural and realistic environment.
[0092] In situations where a scene includes multiple rooms, the approach is typically to render only the audio and reverberation of the room in which the listener is present and ignore any audio from other rooms. However, this tends to result in a suboptimal perceived audio experience and not provide an optimally natural experience, especially as the user moves between rooms. Some applications have been implemented to include rendering of audio from adjacent rooms, but these have proven to be suboptimal. In some embodiments, audio from other rooms can significantly impact the perceived audio scene. In particular, in many scenarios, audio from other rooms can significantly contribute to reverberation and diffuse (background) sound within a room, and suboptimal rendering of such audio can result in a poor user experience.
[0093] In the following, an advantageous approach for rendering an audio scene containing multiple rooms / acoustic environments is described.
[0094] 5 shows an example of an audio device rendering audio for an audio scene. The audio device receives audio data describing the audio and audio sources within a scene (such as the scene of FIG. 4). Based on the received audio data, the audio device renders an audio signal representing the scene for a given listening position. The rendered audio includes both contributions from audio generated in the room in which the listener is located and contributions from other nearby rooms (usually adjacent rooms).
[0095] The audio device generates an audio output signal that represents audio in a scene. Specifically, the audio device generates audio that represents the audio perceived by a user moving through a scene having several audio sources and given acoustic characteristics. Each audio source is represented by an audio signal that represents the sound from the audio source and metadata that describes the characteristics of the audio source (e.g., providing a level indication of the audio signal). Further metadata is provided to characterize the scene.
[0096] The renderer in this example is a portion of an audio device that receives audio data and metadata for a scene and renders audio representing at least a portion of the environment based on the received data.
[0097] The audio device of Figure 5 includes a first receiver 501 that receives audio data for audio sources in a scene, and thus for multiple acoustic environments / rooms separated by acoustic attenuation boundaries. The audio data may include audio data describing multiple audio signals from different audio sources in the scene. Typically, several, e.g., point sources, are provided with audio data reflecting the sound rendered from these audio (point) sources. In some embodiments, audio data may also be provided for more diffuse audio sources, e.g., background or ambient sound sources, or sound sources with spatial extent.
[0098] The audio device includes a second receiver 503 for receiving metadata of the audio data, in particular metadata of the audio source represented by the audio data. As will be explained in more detail below, the metadata contains various information about the scene, in particular information related to different acoustic environments and the boundaries between them.
[0099] The device further includes a position circuit 505 that determines a listening position within a scene. The listening position typically reflects the user's (virtual) position within the scene. For example, the position circuit 505 may be coupled to a user tracking device, such as a VR headset, an eye tracking device, or a motion capture camera, from which it receives user movement data (including, or possibly limited to, head movement and eye movement). The position circuit 505 continuously determines a current listening position from this data.
[0100] This listening position may alternatively be represented or emphasized by controller inputs that allow the user to move or teleport the listening position within the scene.
[0101] It will be appreciated that many approaches and techniques are known and used to determine listening positions within a scene for various applications, and any suitable approach may be used without interfering with the present invention.
[0102] In this example, much of the discussion regarding reverberant and diffuse audio depends primarily on which acoustic environment / room is the listening acoustic environment / room, rather than the specific location of the listener within the listening acoustic environment / room. Thus, in some embodiments, the listening position simply refers to the listening acoustic environment / room, i.e., the acoustic environment / room in which the listener is located.
[0103] The audio device includes a renderer 507 that generates an audio output signal that represents the audio of the scene at the listening position. Typically, the audio signal is generated to include audio components from various audio sources in the scene. For example, point audio sources in the same room are rendered as point audio sources with direct audio paths, and reverberant components are rendered as diffuse signals with no specific location.
[0104] The audio device of Figure 5 includes multiple reverberators. While Figure 5 shows a first reverberator 509 and a second reverberator 511, it will be understood that this is for illustrative purposes only and that in other embodiments, the audio device may include more than two reverberators. For brevity and clarity, much of the following description focuses on embodiments in which the audio device includes only two reverberators.
[0105] Thus, the audio device includes a plurality of reverberators / a bank of reverberators, each generating a diffuse / reverberant signal. Each reverberator receives audio signals corresponding to one or more audio sources and generates a reverberant / diffuse audio signal. The reverberators also receive a set of reverberation parameters, which are used to adapt the reverberators to generate reverberant / diffuse signals with characteristics that reflect the acoustic environment / room in which the signals are generated. For example, parameters can be provided that indicate the proportion of input audio signal energy or emitted source energy to convert to reverberant energy (e.g., a DSR value), a parameter that indicates the onset / delay time of the reverberant / diffuse energy (e.g., a delay of the reverberant tail of a room impulse response), and / or a parameter that indicates the amplitude decay rate (e.g., a T60 value).
[0106] In many other embodiments, the audio signals processed by the reverberators are a downmix of the audio signals / sources of the associated room / acoustic environment. For example, in the example of Figure 5, each reverberator includes a downmixer and / or delay that generates the audio signal that the reverberator processes from the signals of the audio sources in the audio data. The audio signals from the first receiver (still separated into signals per source) are fed to the reverberators. In this case, each reverberator generates an appropriate weighted downmix from these signals depending on its acoustic environment, and also has a (reverberator-specific) delay to suit the direct path rendering.
[0107] The reverberators 509, 511 are coupled to a renderer 507, which provides the reverberant signals to the renderer 507. The renderer 507 then combines the reverberant signals with path signals representing the individual paths of audio sources in the listening room to generate a composite audio signal representing the composite sound in the environment as perceived by the listener.
[0108] Each reverberator in the reverberator bank includes (or is) a parametric reverberator such as a Feedback Delay Network (FDN) reverberator, and in particular a Jot reverberator.
[0109] An example of a suitable reverberator is the Jot reverberator shown in Figure 6. This reverberator includes a loop input vector b and a loop extraction matrix C, which control how the input samples are distributed in the reverberator's feedback loop and how the output signal is generated from the loop.
[0110] An example of a reverberator based on a feedback delay network of the input signal is shown in Figure 7, where three feedback loops are shown. The Jot reverberator can be considered an example of a feedback delay network reverberator.
[0111] The feedback delay network includes a plurality of feedback loops, each (or at least one) of which has an input that receives an input audio signal, where each feedback loop implements a loop transfer function (which may be specifically a delay); a feedback network that returns the output signal of the feedback loop to the loop's input for combination with the input audio signal; and an output circuit that generates the output signal of the feedback delay network as a combination of the output signals of the feedback loops. For each feedback loop, the feedback network may implement a feedback path for the feedback loop's output signal to the feedback loop's input and typically also to one or more inputs of the other feedback loops. In many embodiments, the feedback network may implement a feedback path from the output of each feedback loop to each input of all of the feedback loops. Each feedback path typically implements an attenuation factor (or equivalently, a gain factor), although in some embodiments, more complex feedback paths may be provided, for example, implementing a frequency-dependent gain (e.g., implementing a filter function). In some embodiments, the loop transfer function is a filter that implements both the desired frequency response and the gain factor, and the feedback bus may simply be flat unity-gain feedback (e.g., corresponding to a feedback matrix representing feedback with a coefficient of 1 on the diagonal). In many embodiments, the feedback network is represented by a feedback matrix with coefficients for each feedback loop pair combination.
[0112] Feedback delay networks are typically based on feedback loops with different delays. An input signal is inserted into the loop and, with an appropriate feedback gain, the signal is fed back into the loop. The output signal is derived by combining the signals in the loop. Thus, the input signal is successively repeated with different delays. By using disjoint delays and having a feedback matrix that mixes the signals between the loops, it is possible to create patterns that resemble reverberation in real spaces, and they are particularly well-suited to generating diffuse reverberation, as in the Jot reverberator and other examples of parametric reverberators.
[0113] The absolute values of the elements in the feedback matrix are designed to be less than 1 to achieve a stable, decaying impulse response. The coefficients can be set in combination with delays to achieve the desired reverberation time (T60). In many implementations, additional gains and filters are included in the loop. These filters can control the decay instead of the matrix. The use of filters has the advantage that the decaying response can be made to vary with frequency. Thus, the gains, delays, and path transfer functions can be set based on the reverberation parameters that characterize the desired characteristics of the reverberation (and usually the characteristics of the room).
[0114] The renderer (507) renders an audio signal for the listening position in an acoustic environment, hereinafter referred to as the first acoustic environment, based on the received audio data and metadata. The rendering is further performed to include at least one audio component generated by rendering an audio source of another acoustic environment. That is, the audio signal generated for the listening position in the first acoustic environment is generated to include a component from an audio source of a second acoustic environment (different from the first acoustic environment). Specifically, in situations / embodiments where the different acoustic environments are different rooms, the rendering of the audio signal for the listening position includes rendering contributions from audio sources in the other rooms.
[0115] Specifically, the renderer 507 can generate audio from the current listening acoustic environment / room that represents the reverberant / diffuse audio of other acoustic environments / rooms, where such reverberant / diffuse audio from different acoustic environments / rooms can be rendered as localized sound sources (e.g., located at a door or window in a wall of the listening room) or as the reverberant / diffuse audio of the listening acoustic environment / room.
[0116] Rendering of audio and audio sources in other acoustic environments / rooms other than the first acoustic environment may be rendered at least partially as diffuse or reverberant audio. In some cases, the rendering may be as the same reverberant diffuse audio at all locations in the first acoustic environment. That is, the audio may be substantially independent of the exact listening position in the first acoustic environment. In such cases, rendering of the audio at a listening position may simply be rendering the diffuse audio independent of the listening position.
[0117] The following describes an approach in which the rendered audio signal includes an audio signal / component representing audio from rooms other than the one containing the listening position. While the description focuses on the generation of this audio component, it will be understood that the rendered audio signal presented to the user may include many other components and audio sources. These may be generated / processed according to any suitable algorithm or approach, and it will be understood that those skilled in the art will be aware of a variety of such approaches.
[0118] In particular, audio devices render an audio signal for a listener in a given listener room, which audio signal includes different audio components originating from different audio sources and / or acoustic environments.
[0119] In particular, the renderer 507 receives audio data and metadata for audio sources present in the listening room and renders these audio sources using conventional approaches, for example some audio sources in the room are rendered as point audio sources with positions indicated in the metadata, it will be appreciated that many such different rendering algorithms are known and any suitable approach may be used without prejudice to the present invention.
[0120] Additionally, in some scenarios and embodiments, reverberant signal components are also generated for audio sources present in the room. For example, a set of reverberant parameters may be provided for the listening room and used to adapt a first reverberator 509 of a bank of reverberators 509, 511. The reverberators 509, 511 are also fed with a composite audio signal representing all audio sources in (or contributing to) the listening room. The first reverberator 509 then generates reverberant signal components for the listening room, which are fed to the renderer 507, which combines them with the point source audio rendering.
[0121] Additionally, the audio device can include signal components representing audio sources in other rooms, specifically signal components representing reverberant audio from other rooms. Specifically, the audio device can use one or more of the reverberators 509, 511 to generate a reverberant signal representing the reverberation of another room. For example, a second reverberator 511 is used to generate a reverberant signal for the other room by providing an audio signal representing one or more audio sources in the second room and adapting the second reverberator 511 based on a set of reverberation parameters representing the reverberant characteristics of the second room. The resulting reverberant signal is then adapted, specifically attenuated, to reflect the reverberant audio from the second room propagating into the listening room. The resulting reverberant audio signal is then included in the audio signal output by the renderer 507.
[0122] In many embodiments, a reverberant signal from another acoustic environment can be included in the output signal by rendering the signal as a locatable audio source, such as a point source, or as an audio source with (limited) range (e.g., located in the transmission region of an acoustically attenuating boundary, such as an opening in a room wall). In some embodiments and scenarios, the reverberant signal may be included as a reverberant signal, or indeed directly (e.g., after appropriate scaling / amplification). In some embodiments, the reverberant signal generated for the second room is used as the input signal for a reverberator adapted based on the reverberation parameters of the listening room. This represents a scenario in which reverberant sounds from another room propagate into the listening room and are affected by additional reverberation there. However, in most embodiments, realistic, high-quality audio rendering and perception is achieved without requiring a second reverberation process for the listening room, thereby reducing complexity and resource usage.
[0123] It will be appreciated that in many cases the audio data and metadata are received as part of the same bitstream, the first receiver 501 and the second receiver 503 are implemented by the same functionality, and substantially the same receiver functionality implements both the first receiver and the second receiver. The audio device of Figure 5 specifically corresponds to or is part of the client device 301 of Figure 3, which receives the audio data and metadata in a single bitstream transmitted from the server 303.
[0124] The metadata contains various data that allows the rendering of the audio for the scene.
[0125] In particular, the metadata includes multiple reverberation parameter sets, each set provided for one acoustic environment. For example, in some embodiments, one reverberation parameter set is provided for each acoustic environment of the scene (e.g., for each room in a building). However, it will be appreciated that in some embodiments, no reverberation parameter sets are provided for some acoustic environments, while multiple reverberation parameter sets are provided for other acoustic environments (e.g., different reverberation parameter sets are provided depending on the position of the listener in the acoustic environment).
[0126] A reverberation parameter set includes parameters that characterize how an audio source / signal in the acoustic environment to which the reverberation parameter set is provided becomes / is transformed into reverberation. Thus, the parameters of a reverberation parameter set provide information about the reverberation resulting from an audio source in a given acoustic environment. A reverberation parameter set includes, for example, parameters that describe the dimensions of a room and the acoustic characteristics of the room.
[0127] Each reverberation parameter set specifically contains parameters that describe the relationship between the level / energy of the reverberation and the level / energy of the audio source in the acoustic environment for which the reverberation parameter set is provided, which for brevity are also referred to as level parameters.
[0128] Thus, the level parameters of a given acoustic environment indicate the level of reverberation in the acoustic environment / room resulting from a given level of audio sources in the acoustic environment. The level parameters are therefore useful for determining the level of reverberation present in a given acoustic environment taking into account the audio sources in that acoustic environment.
[0129] The reverberation parameter set may specifically include a DSR value that directly indicates the ratio between the energy / level of the diffuse audio and the (emitted) energy / level of the audio source. In many embodiments, using a DSR value is particularly advantageous, and the following description will focus on examples where the level parameter is a DSR parameter.
[0130] In many embodiments, each reverberation parameter set includes, for example, a parameter indicating the proportion of input audio signal energy that is converted to reverberant energy (e.g., a DSR value), and often a parameter indicating the onset / delay time of the reverberant / diffuse energy (e.g., a delay / lag of the estimated onset of the reverberant tail of the room impulse response), and / or a parameter indicating the rate of amplitude decay (e.g., a T60 value). In many embodiments, the reverberation parameter set provides all the information necessary to adapt the reverberator to generate a reverberant signal in a given acoustical environment that is representative of reverberation resulting from an audio source in that acoustical environment. Thus, by configuring the reverberator based on the reverberation parameter set and applying an input signal that represents an audio source in the acoustical environment of the reverberation parameter set, the reverberator can generate a reverberant signal for that acoustical environment. The reverberant signal is generated to have a level (energy / amplitude), delay, and decay time that reflects the audio source and acoustical environment characteristics.
[0131] The metadata further includes one or typically more energy transfer parameters, each energy transfer parameter indicating the energy attenuation of audio propagation between the source acoustic environment and the destination acoustic environment. The energy transfer parameters indicate the level attenuation from audio in the source acoustic environment to audio in the destination acoustic environment. Thus, the energy transfer parameters for a pair of acoustic environments reflect the sound level in the destination acoustic environment from the sound in the source acoustic environment, and thus indicate how loud the sound in the source acoustic environment is in the destination acoustic environment.
[0132] The energy transfer parameters may directly reflect all parameters of the propagation of sound from the source acoustic environment to the destination acoustic environment, and therefore include how sound propagates in free space, through an acoustically attenuating boundary, through an opening in an acoustically attenuating boundary, etc.
[0133] In many embodiments, audio from one acoustic environment to another propagates primarily through the transfer region of the acoustically attenuating boundary between the acoustic environments. For example, in a building, sound propagates primarily between rooms through wall openings, such as door and window openings. In this case, reverberation from different rooms is rendered in the listening room as an audio signal with the audio characteristics of reverberation but with a location corresponding to the wall opening. For example, reverberation from a second room is heard as reaching the user through a wall opening. In some embodiments, the energy transfer parameter directly indicates the energy attenuation of audio propagation between the transfer region of the source acoustic environment and the transfer region of the destination acoustic environment.
[0134] Thus, an energy transfer parameter may indicate the energy attenuation between a pair of acoustic environments. The energy attenuation of a pair of acoustic environments may indicate the rate at which audio energy in one acoustic environment of the pair is propagated to the other acoustic environment of the pair.
[0135] The audio device of Figure 5 renders audio of a position in a first / listening acoustic environment, as described above. The rendered audio includes renderings of audio sources having specific positions in the listening acoustic environment (including renderings of direct and reflected path audio components). Additionally, the audio device optionally generates reverberant signal components for reverberation caused by audio sources in the listening acoustic environment. Additionally, the audio device also generates signal components reflecting reverberant audio in acoustic environments other than the listening acoustic environment. Specifically, the audio device generates reverberant signals for one or more other acoustic environments and includes them as reverberant signal components in the rendered output audio signal and / or renders them as positioned audio sources (possibly as augmented audio sources).
[0136] The audio device includes a bank of reverberators 509, 511 that are used to generate reverberant signal components based on a set of reverberation parameters for the corresponding acoustic environment. The reverberators are used to generate reverberant signals for the listening acoustic environment as well as for other acoustic environments. For acoustic environments other than the listening acoustic environment, the levels of the generated reverberant signal components are determined based on the energy transfer parameters of the given acoustic environment and the listening acoustic environment. Specifically, the reverberant signals for the acoustic environment are attenuated by an amount determined from the energy transfer parameters.
[0137] In the ideal case, a reverberant signal would be generated for each reverberant parameter set, i.e., for every acoustic environment for which a reverberant parameter set is provided (including generating multiple reverberant signals for acoustic environments with multiple associated parameter sets). However, such an approach typically requires a large number of highly complex and resource-intensive reverberators, making it impractical for all but the simplest scenes and most powerful processing platforms.
[0138] 5 includes a selector 513 for selecting a number of reverberation parameter sets for generating reverberant signal components. The selector 513 selects a reverberation parameter set from the reverberation parameter sets for which the listening acoustic environment is the target acoustic environment, i.e., the reverberation parameter set includes energy transfer parameters that describe attenuation from another acoustic environment to the listening acoustic environment.
[0139] The selector 513 selects a reverberation parameter set based on its own parameters. For example, the selector 513 includes criteria or equations for generating a rank value (or equivalently, a cost or intensity value) based on one or more parameter values of the reverberation parameter set. Specifically, the rank value is generated to provide an indication of the perceptual importance of reverberant sounds to the listening acoustic environment from corresponding other acoustic environments. This may include, for example, weighting of cost values reflecting the energy (e.g., represented by DSR) and decay time (e.g., represented by T60) of the diffuse signal, or a propagation ratio (e.g., represented by an energy transfer parameter) such that, for example, low-level, rapidly decaying reverberations are rejected in favor of higher-level, slower-decaying reverberations. It will be understood that the particular equation or algorithm used to determine the rank value (or cost value) will depend on the particular preferences and requirements of each individual embodiment.
[0140] The selector 513 then selects a given number of reverberation parameter sets as those having the highest rank values, and then assigns a reverberator to each of the selected reverberation parameter sets to generate a corresponding reverberation signal component for inclusion in the rendered output audio signal.
[0141] Therefore, the selector 513 is used to select a subset of the reverberation parameter sets to be rendered, the selection being based on a ranking of the reverberation parameter sets that reflects the estimated perceptual importance of the corresponding reverberation in the listening acoustic environment.
[0142] The number of selected reverberation parameter sets may vary in different embodiments.
[0143] In many embodiments, the selector 513 can determine the number of reverberators available to generate reverberant signals for other acoustic environments. The number of reverberation parameter sets selected is then adapted accordingly, and in particular in many embodiments, the same number of reverberation parameter sets are selected as there are reverberators available.
[0144] In some embodiments, the audio device includes a fixed number of reverberators assigned to generate reverberant signals for other acoustic environments, and the selector 513 selects a corresponding fixed number of reverberation parameter sets.
[0145] In other embodiments, the number of available reverberators varies and the selector 513 dynamically determines the number of available reverberators and selects a corresponding number of reverberation parameter sets.
[0146] For example, in some embodiments, the reverberators are implemented as software, and the number of reverberators that can be implemented at a given time depends on the amount of computational resources available at a given time. For example, if a large number of complex operations are being performed, there may be sufficient resources to implement only a relatively small number of reverberators, and therefore only a small number of reverberators are available to the selector 513. Alternatively, if no other complex operations are currently being performed, the number of reverberators that can be implemented increases, and therefore the number of reverberators available to the selector 513 increases. Thus, the selector 513 determines the number of available reverberators depending on the (current) computational load of the audio processing device. The selector 513 then selects a corresponding number of reverberation parameter sets and assigns one selected reverberation parameter set to each available reverberator.
[0147] In some embodiments, the number of available reverberators depends on the number of reverberators used to render reverberation in the listening acoustic environment. For example, for the current listening acoustic environment, there may be zero, one, or multiple sets of reverberation parameters, depending on how reverberant the acoustic environment is (and whether this differs in different areas of the acoustic environment, etc.).
[0148] The audio device prioritizes local reverberation over reverberation from other acoustic environments and allocates reverberators for each reverberation parameter set accordingly. Thus, the audio device allocates zero, one, or more reverberators to generate a reverberation signal for an audio source in the listening acoustic environment. The selector 513 then determines how many reverberators remain or are available to render reverberation from the other acoustic environments. The selector 513 then selects that number of reverberation parameter sets from the other acoustic environments and allocates a reverberator for each selected reverberation parameter set. Thus, in such an embodiment, local reverberation generation for a local sound source in the listening acoustic environment is prioritized over reverberation in the other acoustic environments. However, reverberators not used to generate a local reverberation signal are allocated to generate reverberation in the other acoustic environments.
[0149] Thus, in some embodiments, the selector 513 determines the number of reverberation parameter sets provided for the listening acoustic environment and subtracts this number from the total number of reverberators in the audio device to determine the number of available reverberators.
[0150] In some embodiments, the selector 513 ranks all reverberation parameter sets that have the reverberation parameter set's acoustic environment as the source acoustic environment and that include the listening acoustic environment as the target acoustic environment for the energy transfer parameters. That is, it ranks all reverberation parameter sets that contribute audio to the listening acoustic environment. The selector 513 then selects the highest-ranked reverberation parameter set and assigns it to a reverberator until all reverberators have been assigned. Such an approach is advantageous in many embodiments, as it enhances perceptual realism by generating reverberation for signals that are most likely to have a perceptual impact.
[0151] In many embodiments, selector 513 determines a (composite) attenuation that represents at least a portion of the reduction in the level of an audio source in the source acoustic environment to the level of reverberant audio in the listening acoustic environment, particularly the reduction in the level of reverberant sound in the source acoustic environment relative to the level of the audio source and the reduction in the level of reverberant sound resulting from propagation from the source acoustic environment to the listening acoustic environment.
[0152] The level reduction resulting from the conversion of audio source energy to reverberant energy in the source acoustic environment is indicated by the reverberant parameters, and the composite attenuation is determined depending on the reverberant parameters. As a specific example, the reverberant parameter set includes a DSR value indicating the ratio of reverberant / diffuse energy resulting from a given sound source energy.
[0153] The composite attenuation is therefore determined based on reverberation parameters, ie, level parameters, which indicate the relationship between the level of reverberation in the source acoustic environment and the level of the audio source in the source acoustic environment.
[0154] The attenuation from the source acoustic environment to the listening acoustic environment is expressed in terms of energy transfer parameters of the source acoustic environment and the listening acoustic environment, and the composite attenuation is determined in response to the energy transfer parameters.
[0155] Thus, composite attenuation indicates the combined effect of attenuation in the source acoustic environment in converting audio source energy to reverberant energy, and the effect of propagation attenuation from the source acoustic environment to the destination / listening acoustic environment.
[0156] In some embodiments, the composite attenuation is determined directly as the composite attenuation of, for example, the attenuation represented by the level parameter and the attenuation represented by the energy transfer parameter between the source acoustic environment and the listening acoustic environment, or in some embodiments, the composite attenuation is determined simply as the combination of these two attenuations by addition (in the log / dB domain) or multiplication (in the linear domain).
[0157] In some cases, the selector 513 may consider other aspects when determining the composite attenuation, for example, if there are multiple sets of reverberation parameters for the source acoustic environment, the selector 513 may interpolate between these to determine the appropriate attenuation to include.
[0158] The selector 513 selects between the reverberation parameter sets based on the composite decay, and in fact may be based solely on the composite decay, and in particular, in some embodiments, the selector 513 may make the selection without taking into account the signal level of the audio source.
[0159] For example, in some embodiments, selector 513 determines, for a given listening acoustic environment, a composite attenuation for each reverberation parameter set of an acoustic environment that is a source acoustic environment of energy transfer parameters that also has the listening acoustic environment as a destination acoustic environment. Selector 513 then ranks all determined composite attenuations and selects from among them in order of increasing attenuation. Thus, selector 513 selects a first reverberation parameter set in preference to a second reverberation parameter set if the first set has a lower composite attenuation than the second set.
[0160] For example, the number of available reverberators is determined and the corresponding number of reverberation parameter sets are selected that have the smallest composite decay.
[0161] In some embodiments, the selector 513 discards a reverberation parameter set from selection if its composite attenuation exceeds a threshold. For example, if the composite attenuation is simply too high, it may be determined that the corresponding audio in the listening acoustic environment is not loud enough to have a significant perceptual effect, no matter how high the signal source level. For example, if the composite attenuation exceeds, say, 60 dB, even if the audio source is loud, the resulting sound in the listening acoustic environment is likely not loud, and therefore the reverberation parameter set can be ignored.
[0162] In some embodiments, the threshold is a dynamic threshold that can be varied. For example, in many embodiments, the selector 513 determines the threshold based on the (e.g., composite) level / energy of audio sources in the listening acoustic environment (thus, whether reverberant audio from other acoustic environments is rendered depends on how loud the local sound sources are). Thus, such an approach may take into account the masking effect of local audio sources.
[0163] This approach is highly advantageous in many embodiments. For example, resources allocated to reverberation processing can be adapted to reflect perceptual importance. For example, in some embodiments, the number of reverberators used is limited to only those reverberation parameter sets that are likely to have a significant perceptual impact. In some embodiments, computational resources are dynamically reallocated to other functions when they are not needed for generating reverberant components.
[0164] In some embodiments, the selector 513 determines the resulting signal levels in the listening acoustic environment and the selection is based on these signal levels.
[0165] Specifically, in some embodiments, the metadata includes audio level data indicating the signal levels of the audio sources. The selector 513 determines the resulting signal level in the listening acoustic environment based on these signal levels and the composite attenuation. For example, in some embodiments, the resulting signal level is simply determined as the signal level of the audio source attenuated by a value corresponding to the composite attenuation.
[0166] In many embodiments, the selector 513 generates a signal level resulting from all reverberation in the source acoustic environment, rather than from individual signal sources. The selector 513 combines the levels of multiple sources in the source acoustic environment into a composite audio source level for the source acoustic environment. For example, in some cases, the levels can be simply added together to generate the composite audio source level. A composite attenuation is applied to the composite audio source level to generate a resulting (composite) signal level in the listening acoustic environment.
[0167] In some embodiments, the selector 513 selects the reverberation parameter sets based on the resulting signal level in the listening acoustic environment, for example, selecting the given number of reverberation parameter sets as the number of reverberation parameter sets having the highest signal level.
[0168] This approach, in many embodiments and scenarios, provides improved rendering of scenes containing multiple environments. Reverberation from environments where the listener is not present is rendered, for example, as a locatable sound source. Typically, this reflects the reverberation heard through a doorway (an opening through which audio propagates between environments). However, this approach takes into account the recognition that in many cases, it may not be necessary to render audio from all other environments because they do not significantly contribute to the listener's perception of the scene. This approach can be used to reduce the computational complexity for rendering such scenes or to scale the computational complexity to available or allocated capacity.
[0169] In this approach, the reverberation of other acoustic environments is represented by a reverberation parameter set. Some individual acoustic environments may be represented by multiple parameter sets. This often means that a reverberator is instantiated for each parameter set, and its output is interpolated depending on the listener's position. Also, here, some parameter sets contribute very little to the listener's perception of the (virtual) environment.
[0170] The described approach provides an efficient and high-performance means of deciding which set of reverberation parameters / acoustic environments to render. The choice is influenced by several factors: the presence of audio sources in the environment; reverberant audio levels, Transfer function from the doorway in the environment to the listener's position, The level of reverberant audio combined with the transfer function relative to the level of other audio presented to the listener.
[0171] As an example, a room floor plan is shown in Figure 8. In this example, some environments / rooms (F, G) have no sound sources of their own and therefore do not emit reverberation, but they do receive reverberation from external sources (e.g. the reverberation in room G is noticeable from the sound sources in room H).
[0172] If a listener is in room E, the transmission of reverberant energy over long distances and / or through obstacles (e.g., from room H) may be too weak to be perceptible.
[0173] Some nearby room sources have reverberant energy that is so weak that it is barely audible at the listening position (e.g., room A), while much stronger reverberation in another environment (e.g., room B) produces significant energy at the listening position for similar source levels in rooms A and B.
[0174] The audio device of FIG. 5 dynamically adapts its rendering to reflect such characteristics and features.
[0175] The room selection is constantly updated as sound sources become active, doors open and close, listeners move around, etc. The deactivation and (re)activation of the reverberators is also controlled to avoid switching artifacts.
[0176] As a specific example, the selector 513 generates a prioritized list of reverberation parameter sets for assignment to the reverberators. Such a prioritized list may include: Reverberation (room reverberation) of the current acoustic environment. Reverberation in all other acoustic environments. The signals are prioritized according to a rank value that reflects the combined attenuation and / or resulting signal level.
[0177] The prioritization of reverberation sounds / reverberation parameter sets is based on a single intensity metric based on the reverberant energy at the listener position / listening acoustic environment, with some modifications possible, such as applying a bonus weighting factor (e.g., 1.5) to all room reverberation sounds.
[0178] The reverberator rank value or "intensity weight" can be selected as the resulting reverberation energy / level perceived in the listening acoustic environment. Specifically, this is the amount of reverberator energy received at the listening position (or acoustic environment containing the listening position). This perceived / received reverberator energy is derived for each reverberation parameter set from the metadata, as shown in Figure 9.
[0179] In the first step, the (normalized) emitted source energy is calculated for all sources that contribute to the reverberation of the source acoustic environment.
[0180] All source energies contributing to a given source acoustic environment are now combined (summed, averaged, maximized) into the total source acoustic environment energy, which DSR then converts into the corresponding reverberant energy / level in the source acoustic environment.
[0181] In the case of multiple parameter sets per room, an interpolation factor is applied in this step to isolate the contribution of a particular reverberation parameter set. For example, each parameter set of a room is considered as a separate reverberator, similar to other parameter sets of the acoustic environment represented by a single reverberation parameter set. If the source acoustic environment has only one reverberation parameter set, this step can be omitted or the interpolation factor can be assumed to be 1.
[0182] Next, energy transfer parameters describing the acoustic transfer / propagation from the source acoustic environment to the listening acoustic environment are applied to determine the perceived reverberator energy in the listening acoustic environment. This step involves combining doorway transfer parameters from all doorways in the source acoustic environment to all doorways in the listening acoustic environment. For example, in the example of Figure 8, for listening room E, and taking into account the perceived reverberator energy from room D, the energy transfer parameter for doorway 7 and the energy transfer parameter from doorways 3 to 4 are combined into a composite energy transfer parameter (also called an inter-room transfer parameter).
[0183] In the previous examples, the source energies were nominal and were considered to be identical. Therefore, in effect, the selection is based directly on the composite attenuation. However, in some embodiments, the signal level of the sound source is available. This can provide a much better estimate of the perceived reverberant energy. The signal level is the signal energy or signal loudness, either calculated directly from the signal or received as metadata.
[0184] In such cases, the approach of Figure 9 can be modified as shown in Figure 10. In this case, rather than assuming a nominal signal energy / level, the actual signal source energy / level can be used.
[0185] The signal level can be represented as loudness (e.g., according to ITU-R BS.1770) within a medium-term window of the signal, but it can also be RMS level or signal energy. Level indicators are provided per signal or per source, where a source is represented by multiple signals. For example, for HOA (Higher Order Ambisonics) sources, a single value is calculated for the omnidirectional representation. Based on loudness standards, a range of, e.g., -62 to 0 dBFS is appropriate, with an additional value to indicate -∞ dBFS. As a result, approximately 6 bits per single value per signal / source are required per update. Updates are performed, e.g., once per second. Under these assumptions, the required bitrate is 6 bits / s per source. Or, for example, for a scene with 30 signals / sources, this would be 180 bps. Even for scenes with more signals / sources, the bitrate does not exceed 1 kbps until a loudness value of 171 is reached. Further optimizations such as dedicated Huffman coding, non-uniform quantization, and differential coding further improve transmission efficiency.
[0186] This value can represent perceptual loudness according to a model, and in most cases is a single value representing the entire bandwidth. However, in such cases, the individual loudness measures may not be well suited to simply adding together to obtain an accurate composite loudness, and a more conservative audibility threshold may be required. In other, more accurate embodiments, it may be desirable to convey signal energy at different frequencies instead. This can be done with different frequency resolutions. There is a trade-off between bitrate and performance. For example, selecting octave bands requires a maximum of 11 values per signal / source. This results in a maximum bitrate of 2 kbps for the example 30 values without entropy coding.
[0187] If no audio source signal level is received or calculated, the selection is based on normalized source energy, or equivalently on composite attenuation. Since this actual reverberant energy in the listening acoustic environment depends on the unknown signal level, in some embodiments, the selection is performed only when there is significant composite attenuation.
[0188] In many embodiments, one or more reverberation and / or energy transfer parameters are frequency dependent. For example, multiple different values are provided for a given parameter, each value being associated with a given frequency interval. In such cases, the above approach can be performed for each frequency interval to determine the frequency-dependent signal level and / or composite attenuation. A single value is then determined by combining the values for the individual frequency ranges. For example, a weighted averaging / summing of the values for different frequency intervals is performed. The averaging / summing is weighted by a coefficient that reflects the magnitude of the different frequency intervals.
[0189] In some embodiments, the combination is a perceptually weighted combination. For example, the coefficients of the weighted summation are set to reflect how perceptually salient each frequency interval is. For example, a frequency interval in the range 500 Hz to 1000 Hz is weighted more highly than a similarly sized frequency interval in the range 2 kHz to 3 kHz.
[0190] Thus, in some cases, some or all of the parameters used to derive the perceived reverberator energy in a listening acoustic environment (e.g., as shown in Figures 9 or 10) may be frequency dependent. In such cases, it is beneficial to calculate the perceived energy as frequency dependent. Furthermore, values at different frequencies can be summed to obtain a single metric. While doing so, perceptual weighting (e.g., A-weighting) can be applied to different frequencies to obtain a more perceptually relevant metric.
[0191] In many embodiments, the selector 513 repeatedly selects the set of reverberation parameters for rendering the reverberant signal components, for example, at intervals not exceeding 1, 2, 5, 10, 20 seconds in different embodiments.
[0192] Such an approach offers advantageous operation in many scenarios where the rendered audio is dynamically adapted and modified to provide a more perceptually accurate and realistic sounding audio scene.
[0193] To mitigate and reduce the perceptual impact of the continuous reselection and change of the rendered reverberation parameter set, the audio device includes the ability to perform a gradual transition.
[0194] For example, after a given reverberation parameter set is deselected, the audio device does not immediately stop generating the corresponding reverberation signal, but instead generates a reverberant audio signal for that reverberation parameter set for a given period of time after deselection. Furthermore, during that period, the audio device generates a reverberant signal for a newly selected reverberation parameter set, during which time the former signal is gradually attenuated and the latter signal is gradually de-attenuated, thereby achieving a gradual transition. Thus, both the old and new reverberant signals are generated during the transition period, during which a gradual cross-fade is also performed.
[0195] In some embodiments, the two reverberators have different complexities and resource usage (or equivalently, the complexity and resource usage of the individual reverberators are changed, e.g., by each reverberator operating in or switching between different modes of operation).
[0196] In such an embodiment, the deselected reverberation parameter set continues to be rendered for a period of time, but using a less complex and less resource-intensive reverberator. For example, during the transient period when a given reverberation parameter set fades out after deselection, the rendering of the reverberation signal not only reduces the level but also uses less resources. The quality of the generated reverberation signal is reduced due to the reduced resource usage. This may allow other functions to be performed.
[0197] For example, in some embodiments, the generation of reverberation signals for the newly selected reverberation parameter set is also performed using a low-resource reverberator. For example, during a transient period, one reverberator is replaced by two low-complexity reverberators that generate both the previously selected reverberation signal and the newly selected reverberation signal during the transient period. After the cross-fade, the reverberators are reconfigured to generate a more complex, resource-intensive reverberation signal for the newly selected reverberation parameter set, while no resources are allocated to the deselected reverberation parameter set. Thus, during normal operation, a high-quality reverberation signal is generated, but during the transient period, two low-quality reverberation signals are generated and cross-faded between them.
[0198] In some embodiments, the selector 513 includes a hysteresis effect in the selection, for example introduced in the intensity metric / energy decision threshold.
[0199] In the previous example, the energy transfer parameters were primarily considered to directly reflect the propagation attenuation between acoustic environments, specifically how much audio energy from the source acoustic environment propagates to the destination acoustic environment. In this example, the energy transfer parameters indicate the losses associated with propagation, but do not take into account how (through which medium) the propagation occurs.
[0200] However, in many embodiments, sound propagation between environments occurs primarily through transfer regions at acoustically attenuating boundaries between acoustic environments, for example, in the case of rooms in a building, sound propagation occurs primarily through openings in the walls (such as doors and windows).
[0201] Thus, in some embodiments, the metadata includes data related to the transmission area of the acoustic attenuation boundary, and some or all of the energy transfer parameters are indicative of the propagation loss / attenuation between the transmission area of the source acoustic environment and the transmission area of the destination acoustic environment.
[0202] In such an embodiment, the metadata includes data describing one or more transmission regions of at least one, and typically several, or all, of the sound attenuation boundaries of the scene.
[0203] Specifically, a transmission region is a region where the acoustic transmission level of sound from one acoustic environment to an adjacent acoustic environment (e.g., from one room to an adjacent room) exceeds a threshold. Specifically, a transmission region is a region (typically an area) of an acoustically attenuating boundary between two acoustic environments, where the attenuation by / across the boundary is less than a given threshold and higher outside the region. A transmission region is a region of an acoustically attenuating boundary that has attenuation lower than the average attenuation of the acoustically attenuating boundary outside the transmission region.
[0204] Thus, a transmission zone defines an area of the boundary between two acoustic environments / rooms where sound propagation / transmission / transparency / coupling exceeds a threshold. Parts of the boundary not included in the transmission zone have sound propagation / transmission / transparency / coupling below the threshold. Correspondingly, a transmission zone defines an area of the boundary between two acoustic environments / rooms where sound attenuation is below a threshold. Parts of the boundary not included in the transmission zone have sound attenuation above the threshold. A transmission zone is also called a doorway (at an acoustically attenuating boundary).
[0205] A Doorway is associated with at least two acoustic environments (e.g., two rooms). The Doorway provides an acoustic connection between the two acoustic environments / rooms. Apart from indicating the connection between the acoustic environments, the Doorway may also contain or refer to the acoustic properties of this connection.
[0206] The following description focuses on an example in which the acoustic environment is a room and the acoustically attenuating boundaries are the walls of the room, however, it will be understood that this is merely exemplary and that the acoustic environment may be other acoustic environments that are at least partially bounded by acoustically attenuating boundaries.
[0207] Thus, the transmission zone refers to a region of the boundary where acoustic transparency is relatively high, outside of which the acoustic transparency is low. For example, the transmission zone corresponds to an opening in the boundary. For example, in a conventional room formed by an acoustically attenuating boundary in the form of a wall, the transmission zone corresponds to, for example, a doorway, an open window, or a hole in the wall separating two rooms.
[0208] The transmission region may be a three-dimensional region or a two-dimensional region. In many embodiments, the boundaries between rooms are represented as two-dimensional objects (e.g., walls are considered to have no thickness), in which case the transmission region will be a two-dimensional shape or region of the boundary with low acoustic attenuation.
[0209] Acoustic transparency can be expressed on a scale: complete transparency means no acoustic suppression (e.g., an open doorway). partial transparency provides attenuation of energy transferring from one room to another (e.g., heavy curtains in a doorway or single-pane windows). At the other end of the scale are room-separating materials (e.g., thick concrete walls) that do not allow any (significant) sound leakage between rooms.
[0210] Thus, in some embodiments, the energy transfer parameters (in the form of a transfer area) provide acoustic coupling metadata that describes how two rooms are acoustically coupled via the transfer area. This data can be obtained locally or, for example, from a received bitstream. The data can be provided manually by a content creator or indirectly from a geometric description of the rooms (e.g., a box, mesh, voxelized representation, etc.) that includes acoustic properties such as material properties that indicate the amount of audio energy that is transferred through a material or coupled into material vibrations that cause acoustic coupling from one room to another. The transfer area is often considered to indicate the room leakage through which acoustic energy is exchanged between two rooms.
[0211] FIG. 8 shows an example of a scene in which the described approach can be applied. FIG. 8 shows an example of a scene consisting of a building with several rooms A-H. Within the building, several audio sources are present in different rooms (indicated by circles). In this case, the audio device of FIG. 5 determines a listening position in room E and renders audio for this listening position. The rendered audio signal includes different audio components in the other rooms, specifically reverberant audio from the other rooms. Specifically, sound from such sources reaches room E via several transmission areas, corresponding for example to (open) doors and windows in the walls forming the room.
[0212] Rendering audio sources in the same room as the listening position is well established, and many algorithms are known and can be used by the renderer without interfering with the present invention. Rendering audio from audio sources located in other rooms can be performed, for example, by representing the audio from the other rooms as audio sources that are assigned a location, e.g., no location (especially in the case of diffuse reverberation) or a location near a doorway. For example, reverberant sound components from another acoustic environment are considered to reach a given first room E where the listener is located via a first doorway 4 of the first room E. The signal level reduction caused by propagation to the first doorway 4 is determined and used to determine the corresponding sound component level at the first doorway. The audio source is then rendered as an audio signal component having a level corresponding to the determined level at the first doorway 4. As mentioned above, in some embodiments, the sound source is rendered as a spatially defined audio source, e.g., a point source located at the location of the first doorway or a sound source with a spatial extent similar to or near the doorway. In other embodiments, the sound components are considered as diffuse sound and are rendered as diffuse reverberation in the first room E. In a specific example, the sound components from the other rooms are reverberant sound and the corresponding signal components are generated by one of the banks of reverberators 509, 511.
[0213] Such an approach is used, for example, to render reverberant audio / sound from room C as heard from a listening position in room E. This approach is also used to render audio sources separated by more than one room, such as an audio source from room A, where, for example, the resulting signal levels after propagation through multiple doorways are determined.
[0214] In some embodiments, the metadata includes data describing, among other things, the location of at least one of the transmission regions of the acoustic attenuation boundary, which may be described relative to the room, for example, or as a relative location on the acoustic attenuation boundary where the transmission region is formed (e.g., defined by a location within the room).
[0215] In many embodiments, the metadata describes the scene topologically and / or geometrically, including, for example, describing rooms, sound-attenuating boundaries, and transmission areas within these. In some embodiments, a geometric description may be included, describing, for example, the size of all rooms (which form the acoustic environment), the extent and location of walls (which form sound-attenuating boundaries), and the size, shape, and location of doorways (which form transmission areas).
[0216] However, in other embodiments, the metadata additionally or alternatively includes a topological description of the scene. Such data may for example list several rooms and provide for each room some acoustic characteristics (such as a set of reverberation parameters describing reverberation), and may further define several doorways / transmission areas, each describing the two rooms it connects.
[0217] In some embodiments, the energy transfer parameter indicates at least one energy attenuation between a pair of transmission areas, one of which is a transmission area of the source acoustic environment and the other of which is a transmission area of the destination acoustic environment, in particular the energy attenuation between two transmission areas of different acoustic attenuation boundaries. The energy attenuation of a pair of transmission areas indicates the rate at which audio energy in one transmission area of the pair is propagated to the other transmission area of the pair. Thus, each energy transfer parameter includes one energy attenuation index (or, as will be explained later, two energy attenuation indexes) for a pair of transmission areas.
[0218] Energy transfer parameters provide energy attenuation metrics, which are particularly useful for describing sound propagation between different rooms through doorways to interconnected rooms. In some embodiments, the metadata only includes energy attenuation parameters for pairs of boundary transfer areas that share acoustic environments. That is, doorways / transfer areas are formed within an acoustic attenuation boundary that is a boundary of the same (intermediate) acoustic environment (e.g., energy transfer parameters are provided for Doorway 1 and Doorway 4 in FIG. 8). The energy transfer parameters indicate propagation from the doorway in Room A to Room E through the doorway to the shared intermediate Room C). This reduces the data rate of the metadata and limits the data representation to the most likely audio propagation between acoustic environments. Furthermore, in some embodiments, if sound propagation is desired to be determined for acoustic environments far apart, such energy transfer parameters / energy attenuation metrics can be combined, as described in more detail below.
[0219] A particular advantage of this approach is that it is suitable for and applicable to many different topologies and connections between different acoustic environments, including providing information about sound propagation between acoustic environments that do not share a common acoustic environment. Indeed, in many embodiments, one or more energy transfer parameters / energy attenuation metrics are provided for transfer regions / sound attenuation boundaries that do not share any common acoustic environment.
[0220] In fact, energy transfer parameters providing an energy attenuation index are provided for any pair of transfer areas or acoustic environments to indicate the possible sound propagation between them, and indeed in some embodiments an energy attenuation index is provided for each possible pair of transfer areas between any two rooms / acoustic environments in a scene.
[0221] In many typical applications, sound propagation is symmetric, and therefore the energy attenuation metric for propagation from transmission area x to transmission area y is the same as for propagation from transmission area y to transmission area x. In such cases, the same energy attenuation metric is used for rendering an audio signal in a first acoustic environment from an audio source in a second acoustic environment and for rendering an audio signal in a second acoustic environment from an audio source in the first acoustic environment.
[0222] Such symmetries are typically present in many physical or virtual scenes, especially in diffuse or reverberant audio, which tends not to be associated with a specific location. This symmetry is used to reduce the amount of data contained in the metadata describing sound propagation between transmission areas (or more generally, sound propagation between acoustic environments). In such cases, for example, the energy decay index for all pairs of transmission areas is represented by a symmetric matrix such as:
number
[0223] The energy attenuation data is efficiently represented as a matrix as above, but may also be represented in any other suitable way, for example, as a direct index as a set of doorway pairs and corresponding transmission area inter-energy attenuation index. Such a matrix may be sparsely populated, and the set of doorway pairs may not be the complete set of possible pairs. This is often useful in scenes with many acoustic environments. Entries with high energy attenuation values may be filtered out, for example, if 10*log10(energy attenuation[i,j])<-60dB.
[0224] Each energy transfer parameter provides an energy attenuation index for two transmission areas / doorways / acoustic environments, and the metadata specifically includes the energy attenuation index and the identification of the associated transmission area / acoustic environment. The energy attenuation index can also be considered an inverse energy transfer index; i.e., the higher the energy attenuation, the lower the energy transfer. The energy attenuation index between two transmission areas / acoustic environments typically indicates that attenuation increases as the distance between the transmission areas increases, and also depends on the number of intermediate acoustic environments and transmission areas that the sound must traverse to reach the desired transmission area. Furthermore, if the two transmission areas are not aligned (around a corner or blocked by an obstacle), the corresponding energy attenuation index will show greater attenuation, reflecting a greater loss in sound attenuation.
[0225] Furthermore, in some embodiments, the energy decay index exhibits a time-varying value, e.g., a value that depends on dynamically changing characteristics of the scene, e.g., the energy decay index changes as a doorway, such as a door, opens, closes, or moves.
[0226] This approach includes the consideration that a doorway is assumed to radiate sound uniformly across its surface into a receiving room. When there are other doorways in the receiving room, some of the sound from the first doorway will reach those second doorways and leak into the next receiving room. The amount of sound transmitted is related to the relative position and size of the other doorways with respect to the first doorway and the total surface area of the room.
[0227] Using this information, we can efficiently determine how much of a sound source's energy in one room contributes to other rooms, and this information is given by the energy attenuation index. For example, each row in the above matrix indicates, for a given doorway of the associated room, how much it contributes to all other rooms.
[0228] In many embodiments, the locations of the transmission areas (and acoustic attenuation boundaries) are assumed to be fixed and unmoving, so the energy attenuation index can be pre-calculated for those specific locations. A simple method is to calculate the visibility area of the receiving doorway relative to the center point of the source doorway and compare that area to that of a hemisphere with a radius equal to the distance between the doorways. Because doorways are assumed to be subsections of a larger plane, they are often assumed to radiate in a hemispherical shape rather than omnidirectionally for the nominal energy transfer index.
[0229] More complex methods take into account the area of the source's doorway rather than just calculating the visibility relative to the source's center point, by calculating the visibility over several locations bounded by the source's doorway and taking the average visibility, or by other techniques.
[0230] In many embodiments, the energy decay metric is computed at the encoder side, or in an offline process where computational resources are more fully available (e.g., computed on the VR Server 303). In such cases, acoustic models of various levels of complexity are used to determine how much energy reaches the second transmission region from the first transmission region, including occlusion and / or diffraction modeling.
[0231] Some embodiments focus on calculating the energy transfer / attenuation from all transfer areas in a room to all other transfer areas in the same room (e.g., for all doorway pairs in room C). These transfers can then be combined to represent higher-order room-to-room transfers (i.e., including multiple shared / intermediate rooms).
[0232] The energy decay measure is used directly to determine the energy reduction factor for a sound in one acoustic environment to reach another acoustic environment, and rendering is performed using this energy reduction factor.
[0233] Specifically, for a given audio source in a source room, the energy incident on the transmission volume is determined, for example, using the approaches described above or other means, for example, resulting from the described reverberant rendering of the source room.
[0234] The resulting energy reduction factor F tgt is applied to the signal by the renderer 507:
number
[0235] A particular advantage of this approach is that it does not require detailed geometric information of the scene, in particular rooms, acoustic attenuation boundaries, transmission regions, etc., or indeed any specific acoustic properties of the scene. Indeed, information about the exact connections between rooms or their acoustic properties is not required. Rather, the energy transfer parameters can be seen as simply topological properties that connect two transmission regions and provide information about the sound propagation between them. This makes it much easier to manipulate and render, and significantly reduces complexity and resource usage.
[0236] In many embodiments, the energy attenuation metric for a pair of doorways indicates the proportion of audio energy incident on one transmission area that propagates to be incident on the other transmission area. This is advantageous in that the energy attenuation metric can be symmetric, allowing one metric to be used for both directions, thus reducing the amount of metadata. It also allows for the rendering to be adapted based on the particular acoustic characteristics of the transmission areas. For example, if the transmission area is dynamically covered by a fabric (e.g., a curtain), this is reflected by introducing an additional attenuation coefficient that can be removed when the transmission area is uncovered.
[0237] In other embodiments, the energy attenuation index indicates the energy attenuation of the output of a receiving transmission area, i.e., the energy exiting / radiating from a given transmission area for a given energy incident on another transmission area, which can reduce rendering complexity in many situations.
[0238] In some embodiments, energy transfer parameters are provided for all pairs of transfer areas through which any sound can be transferred / propagated, and rendering signal components representing inter-room propagation through a transfer area simply involves extracting and using the appropriate energy attenuation metric for that transfer area.
[0239] However, in some embodiments, energy transfer parameters are provided only for a subset of the transfer areas, e.g., only for transfer areas that share a common acoustic environment. This allows for lower data rates and significantly reduces the requirements for determining accurate energy attenuation metrics. For example, if these are based on measurements in real buildings, the number of measurements required can be significantly reduced.
[0240] In such embodiments, the energy attenuation index for other pairs of transmission areas may be determined, for example, by combining the energy attenuation index for other pairs of transmission areas. Thus, in some embodiments, the renderer 507 generates a composite energy transfer attenuation by combining the energy transfer attenuation of a first pair of transmission areas with the energy transfer attenuation of a second pair of transmission areas. These two pairs include a common transfer area. For example, the first pair of transmission areas is for the transfer area between the first and second transfer areas, thereby providing an index of energy transfer / attenuation between the first and second environments. The second pair of transmission areas is between a third and second transfer area, thereby providing an index of energy transfer / attenuation between the third and second transfer areas, and therefore providing an index of energy transfer / attenuation between the second and third acoustic environments. The energy attenuation indexes of two pairs of transmission areas are combined, for example, by simply combining the attenuation (e.g., by multiplying the two energy attenuations in the linear domain or by adding them in the logarithmic domain of the attenuation values). The resulting composite value thus indicates the energy attenuation from the third transmission area to the first transmission area, and thus the propagation of sound from the third acoustic environment to the first acoustic environment. The composite energy attenuation is then used to render audio from an audio source in the third acoustic environment to a listening position in the first acoustic environment, just as if a direct energy attenuation measure had been provided for the pair of first and third transmission areas. Some embodiments further include the transmission characteristics or material properties of the second transmission area.
[0241] In particular, the apparatus may be implemented in one or more suitably programmed processors. In particular, the artificial neural network may be implemented in one or more such suitably programmed processors. The various functional blocks may be implemented in separate processors and / or may be implemented in, for example, the same processor. An example of a suitable processor is given below:
[0242] 11 is a block diagram illustrating an exemplary processor 1100 according to an embodiment of the present disclosure. Processor 1100 may be used to implement one or more processors that implement the aforementioned devices or elements thereof (including, inter alia, one or more artificial neural networks). Processor 1100 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) (wherein the FPGA is programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (wherein the ASIC is designed to form a processor), or a combination thereof.
[0243] Processor 1100 may include one or more cores 1102. Core 1102 may include one or more arithmetic logic units (ALUs) 1104. In some embodiments, core 1102 includes a floating point logic unit (FPLU) 1106 and / or a digital signal processing unit (DSPU) 1108 in addition to or instead of the ALUs 1104.
[0244] The processor 1100 may include one or more registers 1112 communicatively coupled to the core 1102. The registers 1112 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 1112 are implemented using static memory. The registers may provide data, instructions, and addresses to the core 1102.
[0245] In some embodiments, processor 1100 may include one or more levels of cache memory 1110 communicatively coupled to cores 1102. Cache memory 1110 may provide computer-readable instructions to cores 1102 for execution. Cache memory 1110 may provide data for processing by cores 1102. In some embodiments, computer-readable instructions may be provided to cache memory 1110 by local memory (e.g., local memory attached to external bus 1116). Cache memory 1110 may be implemented with any suitable cache memory type, such as, for example, static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology, e.g., metal-oxide semiconductor (MOS) memory.
[0246] The processor 1100 may include a controller 1114. The controller 1114 may control input to the processor 1100 from other processors and / or components included in the system and / or output from the processor 1100 to other processors and / or components included in the system. The controller 1114 controls data paths within the ALU 1104, the FPLU 1106, and / or the DSPU 1108. The controller 1114 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 1114 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.
[0247] The registers 1112 and the cache 1110 may communicate with the controller 1114 and the core 1102 via internal connections 1120A, 1120B, 1120C, and 1120D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.
[0248] Input and output for processor 1100 is provided via bus 1116, which may include one or more conductive lines. Bus 1116 may be communicatively coupled to one or more components of processor 1100, such as controller 1114, cache 1110, and / or registers 1112. Bus 1116 may be coupled to one or more components of the system.
[0249] The bus 1116 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 1132. The ROM 1132 may be masked ROM, electronically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory (RAM) 1133. The RAM 1133 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 1135. The external memory may include flash memory 1134. The external memory may include a magnetic storage device such as a disk 1136. In some embodiments, external memory may be included in the system.
[0250] The terms "audio" and "sound" are considered equivalent and interchangeable and both may refer to physical sound pressure and / or electrical signal representation, as appropriate in the context.
[0251] It will be appreciated that, for clarity, the above description describes embodiments of the invention with reference to various functional circuits, units, and processors. However, it will be apparent that functionality may be distributed among various functional circuits, units, or processors as appropriate without detracting from the invention. For example, functionality described as being performed by separate processors or controllers may be performed by the same processor or controller. Accordingly, references to specific functional units or circuits are to be regarded merely as references to suitable means for providing the described functionality, and not as indicative of a strict logical or physical structure or organization.
[0252] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the invention may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in several units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits, and processors.
[0253] While the present invention has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, while certain features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0254] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Moreover, although individual features may be included in different claims, these features may also be advantageously combined, and the inclusion of features in various claims does not imply that such combinations are not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply a limitation of this category, but rather indicates that the feature may be applied to other claim categories as well, where appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must function, and in particular the order of individual steps in method claims does not imply that the steps must be performed in this order. Rather, steps may be performed in any suitable order. Furthermore, a reference in the singular does not exclude a reference in the plural; thus, a reference to "first," "second," etc. does not exclude a plural. Reference signs in the claims are provided solely as examples for clarity, and these examples should not be construed as limiting the scope of the claims in any way.
Claims
1. a first receiver for receiving audio data of an audio source of a scene including a plurality of acoustic environments separated by acoustic attenuation boundaries; a second receiver for receiving metadata of the audio data, the metadata comprising a plurality of reverberation parameter sets and at least one energy transfer parameter, each reverberation parameter set for one associated acoustic environment comprising at least one reverberation parameter indicative of a relationship between a level of reverberation in the one associated acoustic environment and a level of an audio source in the one associated acoustic environment, each energy transfer parameter comprising at least one energy transfer parameter indicative of an energy attenuation of audio propagation between a source acoustic environment and a destination acoustic environment; a selector for selecting at least a first reverberation parameter set of the plurality of acoustic environments for a first acoustic environment, the first reverberation parameter set being for an acoustic environment different from the first acoustic environment, the selection being made in response to parameter values of the first reverberation parameter set and an energy transfer parameter indicative of energy attenuation of audio propagation between the different acoustic environment and the first acoustic environment; a plurality of reverberators for generating reverberant audio signals, at least one of the reverberators generating the reverberant audio signal based on the first reverberation parameter set and audio data of audio sources in the different acoustic environments of the first reverberation parameter set; a renderer for generating an audio signal for the first acoustic environment based on the audio data, the audio signal including an audio component derived from the reverberant audio signal; and An audio device comprising:
2. 2. The audio device of claim 1, wherein the selector selects between reverberation parameter sets based on a composite attenuation for each reverberation parameter set, the composite attenuation for a given reverberation parameter set indicating a level attenuation for audio in the first acoustic environment resulting from reverberation in the given acoustic environment caused by an audio source in the given acoustic environment, and the composite attenuation for the given reverberation parameter set includes a contribution from attenuation indicated by the at least one reverberation parameter of the given reverberation parameter set and an energy transfer parameter indicating an energy attenuation of audio propagation between the given acoustic environment and the first acoustic environment.
3. 3. The audio device of claim 2, wherein the selector selects the first reverberation parameter set in preference to the second reverberation parameter set if the composite decay of the second reverberation parameter set exceeds the composite decay of the first reverberation parameter set.
4. 4. An audio device according to claim 2 or 3, wherein the selector discards a reverberation parameter set from the selection if the combined decay exceeds a threshold.
5. 5. The audio device of claim 2, wherein the metadata further comprises audio level data indicating a signal level of an audio source, and wherein the selector determines a resulting signal level in the first acoustic environment of the audio source in the acoustic environment of the reverberation parameter set based on the signal level of the audio source and the combined attenuation of the reverberation parameter set, and selects the at least one reverberation parameter set in response to the resulting signal level in the first acoustic environment.
6. 6. The audio device of claim 5, wherein the selector determines the resulting signal level of the reverberation parameter set for the at least one acoustic environment by combining levels of multiple sound sources in at least one acoustic environment into a composite audio source level and applying the composite attenuation of the at least one acoustic environment to the composite audio source level.
7. 7. The audio device of claim 1, wherein the selector determines a number of the plurality of reverberators available for generating a reverberant signal of an audio source in an acoustic environment other than the first acoustic environment, and adapts a number of selected reverberation parameter sets to match the number of the plurality of reverberators available.
8. 8. The audio device of claim 7, wherein the selector determines a first number of reverberant signal sources in the first acoustic environment and subtracts the first number from a number of reverberators of the audio device to determine the number of available reverberators of the plurality of reverberators.
9. 9. An audio device according to claim 1, wherein at least one of the at least one reverberation parameter and the energy transfer parameter is frequency dependent.
10. 10. Audio device according to claim 1, wherein the selector performs the selection of the at least one reverberation parameter set iteratively.
11. 11. The audio device of claim 10, wherein at least a first reverberator of the plurality of reverberators continues to generate the reverberated audio signal of the first reverberation parameter set for a period of time after deselection of the first reverberation parameter set.
12. 12. The audio device of claim 11, wherein the first reverberator uses less computational resources than a reverberator of the plurality of reverberators that was used to generate the reverberated audio signal of the first reverberation parameter set before deselection of the first reverberation parameter set.
13. 13. An audio device according to any one of claims 1 to 12, wherein the at least one reverberation parameter comprises a diffuse-to-source ratio parameter indicative of a ratio between emitted audio source energy and diffuse reverberation energy.
14. receiving audio data for an audio source of a scene including a plurality of acoustic environments separated by acoustic attenuation boundaries; receiving metadata of the audio data, the metadata comprising a plurality of reverberation parameter sets and at least one energy transfer parameter, each reverberation parameter set for one associated acoustic environment and including at least one reverberation parameter indicative of a relationship between a level of reverberation in the one associated acoustic environment and a level of an audio source in the one associated acoustic environment, each energy transfer parameter indicative of an energy attenuation of audio propagation between a source acoustic environment and a destination acoustic environment; selecting at least a first reverberation parameter set of the plurality of acoustic environments for a first acoustic environment, the first reverberation parameter set being for a different acoustic environment from the first acoustic environment, the selection being made in response to parameter values of the first reverberation parameter set and an energy transfer parameter indicative of energy attenuation of audio propagation between the different acoustic environment and the first acoustic environment; generating reverberant audio signals with a plurality of reverberators, at least one of the reverberators generating the reverberant audio signal based on the first reverberation parameter set and audio data of audio sources in the different acoustic environments of the first reverberation parameter set; generating an audio signal for the first acoustic environment based on the audio data, the audio signal including an audio component derived from the reverberant audio signal; 10. A method for operating an audio device, comprising:
15. A computer program comprising computer program code means adapted to perform all the steps of the method according to claim 14 when said program is run on a computer.