Audio device and method

JP2024513889A5Active Publication Date: 2025-09-01KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023561329
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-04-08
Filing Date
2022-03-10
Publication Date
2025-09-01
Estimated Expiration
2042-03-10

AI Technical Summary

Technical Problem

Existing methods for generating audio in virtual reality applications do not accurately reflect all acoustic phenomena, leading to suboptimal rendering of room acoustics and a less immersive user experience.

Method used

An audio device that generates a flutter echo audio signal using a feedback delay network with multiple feedback loops, adapting parameters based on room metadata to simulate flutter echoes and diffuse reverberation, enhancing the perception of the acoustic environment.

Benefits of technology

Improves the accuracy and naturalness of audio rendering, reducing computational complexity while providing a more immersive and realistic acoustic experience in virtual reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The audio device for generating a flutter echo audio signal comprises a receiver 401 configured to receive room metadata indicative of characteristics of a room. The room metadata may for example be indicative of room dimensions and / or acoustic reflection data of room boundaries. An estimator 405 determines a room flutter echo estimate in response to the room metadata, the flutter echo estimate indicative of a level of flutter echo in the room. A signal generator 403 includes a feedback delay network with a plurality of feedback loops. The signal generator 403 is configured to generate the flutter echo audio signal from output signals of a group of feedback loops of a plurality of feedback loops to which the audio source signal is fed. An adapter 407 adapts a first parameter for a first feedback loop in the group of feedback loops in response to the flutter echo estimate.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an apparatus and method for generating a flutter echo audio signal, particularly, but not exclusively, in combination with the generation of a diffuse reverberation signal. [Background technology]

[0002] In recent years, the variety and breadth of experiences based on audiovisual content has increased significantly with the continued development and introduction of new services and methods for using and consuming such content. In particular, many spatial and interactive services, applications and experiences have been developed to provide users with more engaging and immersive experiences.

[0003] Examples of such applications are Virtual Reality (VR), Augmented Reality (AR) and Mixed Reality (MR) applications, which are rapidly becoming mainstream, with many solutions targeting the consumer market. Numerous standards are also under development by a number of standards bodies. Such standardization activities are actively developing standards for various aspects of VR / AR / MR systems, including, for example, streaming, broadcasting, rendering, etc.

[0004] VR applications tend to provide a user experience that corresponds to the user being in another world / environment / scene, whereas AR (including mixed reality MR) applications tend to provide a user experience that corresponds to the user being in the current environment but with additional information or virtual objects or information added. Thus, VR applications tend to provide fully immersive synthetically generated worlds / scenes, whereas AR applications tend to provide partially synthetic worlds / scenes that are overlaid on the real scene in which the user is physically present. However, these terms are often used interchangeably and are highly overlapping. In the following, the term virtual reality / VR is used to refer to both virtual reality and augmented / mixed reality.

[0005] As an example, an increasingly popular service is the ability for a user to actively and dynamically interact with the system to modify the parameters of the rendering, providing images and audio in a manner that adapts to movements and changes in the user's position and orientation. A very attractive feature in many applications is the ability to change the viewer's effective viewing position and direction, for example to enable the viewer to move and "look around" within the scene being presented.

[0006] Such features may allow a virtual reality experience to be provided to the user in particular, allowing the user to move around (relatively) freely in the virtual environment and dynamically change his position and where he is looking. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known, for example, from gaming applications, such as in the class of first-person shooter games for computers and consoles.

[0007] In addition to the visual rendering, most VR / AR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience, such that the audio sources are perceived to arrive from positions corresponding to the positions of corresponding objects in the visual scene. In this way, the audio scene and the video scene are preferably perceived to be consistent and both provide a complete spatial experience.

[0008] For example, many immersive experiences are provided where the virtual audio scene is generated through headphone playback using binaural audio rendering techniques. In many scenarios, such headphone playback can be based on head tracking so that rendering can be made responsive to the user's head movements, which greatly increases the sense of immersion.

[0009] An important feature for many applications is how to generate and / or deliver audio that can provide a natural and realistic perception of the audio environment. For example, when generating audio for virtual reality applications, it is important not only that the desired audio sources are generated, but that these audio sources are modified to provide a realistic perception of the audio environment, including attenuation, reflections, coloration, etc.

[0010] In the case of room acoustics, or more generally environmental acoustics, reflections of sound waves from the walls, floor, ceiling, objects, etc. of the environment cause delayed and attenuated (usually frequency dependent) versions of the source signal to reach the listener (i.e. the user of the VR / AR system) via different channels. This combined effect can be modeled by an impulse response, hereafter called Room Impulse Response (RIR) (although this term implies a specific use for an acoustic environment in the form of a room, the term tends to be used more generally in relation to acoustic environments, whether this applies to a room or not).

[0011] As shown in Figure 1, a room impulse response typically consists of a direct sound, which depends on the distance from the sound source to the listener, followed by a reverberant part that characterizes the acoustic properties of the room. The size and shape of the room, the position of the sound source and the listener within the room, and the reflective properties of the room's surfaces all affect the characteristics of this reverberant part.

[0012] The reverberant part can be decomposed into two time domains that usually overlap. The first domain contains the so-called early reflections, which represent isolated reflections of the sound source off walls or obstacles in the room before reaching the listener. As the time delay increases, the number of reflections present within a certain time interval increases, and their paths may include second or higher order reflections (e.g. reflections may be from multiple walls, or both walls and ceiling, etc.).

[0013] The second region in the reverberant section is where the density of these reflections increases to the point where they can no longer be separated by the human brain. This region is usually called the diffuse reverberation, late reverberation, or reverberation tail.

[0014] The reverberant part contains cues that give the hearing system information about the distance of a sound source and the size and acoustic properties of the room. The energy of the reverberant part relative to the energy of the anechoic part largely determines the perceived distance of a sound source. The level and delay of early reflections can provide clues about how close a sound source is to a wall, and anthropometric filtering can enhance the assessment of a particular wall, floor or ceiling.

[0015] The density of (early) reflections affects the perceived size of a room. The time it takes for the energy level of the reflections to drop to 60 dB (reverberation time T 60 Reverberation time, denoted by (R) is often used as a measure of how quickly reflections dissipate in a room. Reverberation time provides information about the acoustic properties of a room, especially whether the walls are very reflective (such as a bathroom) or very sound absorbing (such as a bedroom with furniture, carpets and curtains).

[0016] Furthermore, the RIR, if part of the Binaural Room Impulse Response (BRIR), may depend on the user's anthropometric characteristics since the RIR is filtered by the head, ears and shoulders (i.e., it is a Head Related Impulse Response (HRIR)).

[0017] Since reflections in the late reverberation cannot be distinguished and separated by a listener, they are often simulated and represented parametrically using, for example, parametric reverberators that use feedback delay networks, as in the well-known Jot reverberator.

[0018] For early reflections, the incidence direction and distance-dependent delay are important cues for humans to extract information about the relative position of the room and the sound source. Therefore, the simulation of early reflections must be more explicit than late reverberation. Therefore, in efficient acoustic rendering algorithms, early reflections are simulated differently from late reverberation. A well-known method for early reflections is to mirror the sound source at each room boundary and generate virtual sources that represent the reflections.

[0019] For early reflections, the position of the user and / or sound source relative to the room boundaries (walls, ceiling, floor) is relevant, whereas for late reverberation, the acoustic response of the room is diffuse and therefore tends to be more uniform across the board, which often makes simulating late reverberation more computationally efficient than early reflections.

[0020] The two main characteristics of the late reverberation defined by a room are the T60 value and the reverberation level. For a diffuse reverberation impulse response, these values ​​represent the slope and amplitude of the impulse response. Both are usually strongly frequency dependent in natural rooms.

[0021] While the T60 parameter is important to give an impression of the reflectivity and size of a room, the reverberation level indicates the combined effect of multiple reflections at the room boundaries. The reverberation level and its frequency behavior depend on the pre-delay, which indicates where the distinction between early and late reflections is made (see Figure 2).

[0022] The reverberation level has the main psychoacoustic relevance with respect to the direct sound. The level difference between the two is an indication of the distance between the sound source and the user (or the RIR measurement point). As the distance increases, the direct sound attenuates more, while the level of late reverberation remains the same (it is the same throughout the room). Similarly, for sound sources with directionality that depends on where the user is relative to the source, the directivity affects the direct response as the user moves around the source, but not the level of reverberation.

[0023] To render a realistic audio experience and provide a perception of the audio environment, in particular the acoustic properties of a virtual room in which a listener is considered to be located, one or more audio signals and objects may be rendered through a rendering process that reflects the room impulse response, which typically involves generating the direct path, early reflections, and diffuse late reverberation components separately and then combining them in the rendered output.

[0024] Typically, different approaches are used to generate the different components: direct sound and early reflections are often generated by simple filtering (e.g. using binaural processing and head-related transfer function filters), whereas diffuse late reverberation is often generated using a parametric reverberator such as the Jotto reverberator.

[0025] Such approaches can generate advantageous and natural-sounding audio in many situations and applications. However, known approaches may not be optimal in some situations and for some applications. For example, many embodiments may result in rendered audio that is not a perfect representation of the intended room acoustics. In many situations, generating a more accurate acoustic environment may require additional complexity and / or computational resources. Current approaches and proposals for how to represent and generate audio representative of an acoustic environment may tend to be suboptimal and / or insufficient and / or incomplete. This may be particularly true in, for example, virtual reality applications, where the rendered acoustic environment may have a significant impact on the immersion and overall user experience. Summary of the Invention [Problem to be solved by the invention]

[0026] Accordingly, improved approaches would be advantageous, particularly approaches that allow for improved operation, increased flexibility, reduced complexity, easier implementation, improved audio experience, improved audio quality, reduced computational burden, improved suitability and / or performance for virtual / mixed / augmented reality applications, improved perceptual cues, improved representation and rendering of different acoustic environments, and / or improved performance and / or operation.

[0027] SUMMARY OF THE DISCLOSURE Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0028] According to one aspect of the present invention, there is provided an audio device for generating a flutter echo audio signal, the audio device comprising: a receiver configured to receive room metadata indicative of characteristics of a room; an estimator configured to determine a flutter echo estimate for the room in response to the room metadata, the flutter echo estimate being indicative of a level of flutter echo in the room; a signal generator including a feedback delay network with a plurality of feedback loops, the signal generator configured to generate the flutter echo audio signal from output signals of a group (set) of feedback loops of a plurality of feedback loops to which an audio source signal is supplied; and an adapter configured to adapt a first parameter for a first feedback loop in the group of feedback loops in response to the flutter echo estimate.

[0029] The present invention can provide an improved user experience in many embodiments and in many scenarios, and in particular can provide an improved user perception of the acoustic environment. The approach can further enable efficient communication of data enabling such improvements, and in particular can be based on environmental data (in particular room data) that can be communicated for other purposes without the need for additional data in many scenarios.

[0030] In particular, the inventors have recognized that existing approaches do not accurately reflect all acoustic phenomena, and that significant improvements can be achieved by generating and rendering a flutter echo audio signal that can provide the perception of flutter echo effects in an acoustic environment. Furthermore, generating such a flutter echo audio signal using a signal generator with a feedback delay network with multiple feedback loops can provide a very efficient implementation while allowing accurate rendering of flutter echo effects in many embodiments. Furthermore, this allows commonality with functionality for generating diffuse reverberation, for example allowing a highly efficient combined reverberator functionality to be provided that can dynamically adapt resources allocated to different types of reverberation and echo.

[0031] The approach can provide adaptation that allows a more natural sounding echo to be perceived. The approach can, in many embodiments, allow a flutter echo effect to be created without requiring dedicated data to be transmitted to control the echo. The device can, among other things, determine whether to generate a flutter echo or not depending on the room metadata, or can adapt, for example, the parameters of the flutter echo (such as delay, frequency response and / or level) to provide a signal that more accurately reflects the natural acoustic environment.

[0032] The flutter echo estimate may be indicative of the level / degree / amount / prevalence of flutter echo in a room, in particular the level / degree / amount / prevalence of flutter echo relative to the diffuse reverberation in the room. The flutter echo may be a flutter echo between two opposing walls / boundaries / sides of a room, in particular between two parallel walls / boundaries / sides of a room.

[0033] The feedback delay network may include a network configured to couple at least an audio source signal to a feedback loop in at least one group of feedback loops, and an output circuit configured to generate a flutter echo audio signal by combining output signals of the feedback loops in the at least one group of feedback loops.

[0034] The group of feedback loops may include one or more feedback loops.

[0035] The first parameter may be, for example, a feedback coefficient for the first feedback loop, a transfer function parameter for the first feedback loop, a frequency dependency of the first feedback loop, a loop gain of the first feedback loop, a delay of the feedback loop, a weight / gain / level for the output signal of the feedback loop and / or a flutter echo audio signal.

[0036] In some embodiments, the adapter may be configured to vary a number of feedback loops in a group of feedback loops in response to a flutter echo estimate.

[0037] In some embodiments, the signal generator is configured to further generate a diffuse reverberation signal from an output of a feedback loop not included in the group of feedback loops. A diffuse reverberation signal may be generated when an audio source signal and / or another audio source signal is fed into a feedback loop not in the group of feedback loops.

[0038] In many embodiments, the apparatus may be configured to determine a flutter echo estimate in response to a room impulse response.

[0039] The receiver may be configured to receive an audio source signal.

[0040] According to an optional feature of the invention, the room metadata includes dimensional data relating to a room, and the flutter echo estimate is determined in response to a room dimension in a first direction relative to a room dimension in a second direction.

[0041] This configuration can provide particularly advantageous operation and improved adaptive flutter echo simulation in many embodiments.

[0042] The dimensional data may provide an indication of the distance between one or more opposing walls / sides / boundaries of a room.

[0043] The flutter echo estimate may indicate that a level of flutter echo increases as the difference between a room dimension in a first direction and a room dimension in a second direction increases.

[0044] According to an optional feature of the invention, the room metadata includes acoustic reflection data relating to sides of the room, and the flutter echo estimate is determined in response to the acoustic reflection attenuation of a first boundary of the room relative to the acoustic reflection attenuation of a second boundary of the room.

[0045] This configuration can provide particularly advantageous operation and improved adaptive flutter echo simulation in many embodiments.

[0046] The first and second boundaries may be walls or sides of a room.

[0047] According to an optional feature of the invention, the adapter is configured to increase a feedback coefficient from the first feedback loop to itself in response to the flutter echo estimate indicating an increase in the level of flutter echo.

[0048] This configuration can provide particularly advantageous operation, resulting in an improved user experience and a more natural perception of the acoustic environment.

[0049] The adapter may be configured to decrease a feedback coefficient from a first feedback loop to a second feedback loop of the plurality of feedback loops in response to the flutter echo estimate indicating an increased level of flutter echo, the second feedback loop being a feedback loop not included in the group of feedback loops.

[0050] In accordance with an optional feature of the invention, feedback coefficients of at least some of the plurality of feedback loops relative to other feedback loops of the plurality of feedback loops are dependent on a room dimension of the room.

[0051] In some embodiments, feedback coefficients of at least some of the group of feedback loops to other feedback loops of the plurality of feedback loops depend on a room dimension of a room.

[0052] According to an optional feature of the invention, the signal generator is configured to further generate a diffuse reverberation signal from outputs of feedback loops not included in the group of feedback loops, and the adapter is configured to vary a number of feedback loops included in the group of feedback loops in response to a flutter echo estimate.

[0053] The approach may allow for a very efficient audio emulation of a room, and may allow for a low complexity implementation, since for example feedback loops can be used for different purposes (generation of diffuse reverberation and feedback loops) with the allocation of feedback loops between them being dynamically adapted.

[0054] A diffuse reverberation signal may be generated when an audio source signal and / or another audio source signal is fed into a feedback loop that is not within the group of feedback loops.

[0055] According to an optional feature of the invention, the signal generator has a delay with respect to the audio source signal before it is provided to a feedback loop of the group of feedback loops, and the adapter is configured to adapt the delay in response to a position of at least one of an audio source of the audio source signal, a listener, and a room boundary.

[0056] This configuration may provide particularly advantageous operation, resulting in an improved user experience and a more natural perception of the acoustic environment.

[0057] According to an optional feature of the invention, the group of feedback loops comprises at least two feedback loops, and the signal generator has a delay with respect to the audio source signal before being provided to the at least two feedback loops, the delay being different for the at least two feedback loops.

[0058] This configuration can provide particularly advantageous operation and / or performance in many embodiments.

[0059] In accordance with an optional feature of the invention, the group of feedback loops comprises two or less loops.

[0060] This configuration can provide particularly advantageous operation and / or performance in many embodiments.

[0061] According to an optional feature of the invention, the adapter is configured to adapt feedback coefficients for a plurality of feedback loops such that there is no feedback from a feedback loop in the group of feedback loops to any feedback loop not included in the group of feedback loops.

[0062] In some embodiments, the adapter is configured to adapt feedback coefficients of a plurality of feedback loops such that there is no feedback from any feedback loop not included in the group of feedback loops to any feedback loop of the group of feedback loops.

[0063] According to an optional feature of the invention, the signal generator is configured to further generate a diffuse reverberation signal, the audio device further comprising: a spatial processor for applying spatial processing to the flutter echo audio signal, the spatial processing being dependent on a position of at least one of a source of the audio source signal and a room boundary; and a combiner for combining the diffuse reverberation signal and the spatially processed flutter echo audio signal.

[0064] In accordance with an optional feature of the invention, the audio device further comprises: a spatial processor that applies spatial processing to the flutter echo audio signal, the spatial processing depending on a position of a source of the audio source signal and at least one of a side of the room.

[0065] In accordance with an optional feature of the invention, the audio device further comprises circuitry for supplying a plurality of audio source signals to a plurality of feedback loops, at least one audio source signal being supplied to only a feedback loop of the group of feedback loops.

[0066] According to an optional feature of the invention, the signal generator has a gain for the audio source signal before it is fed to a feedback loop of the group of feedback loops, and the adapter is configured to adapt the gain in response to at least one of a position of an audio source of the audio source signal, a position of a listener, a position of the room boundary and a reflection order of a start of a flutter echo audio signal.

[0067] According to an optional feature of the invention, the flutter echo audio signal represents a flutter echo between a pair of opposing boundaries of a room, the signal generator has a frequency-dependent gain for the audio source signal before being fed to a feedback loop of the group of feedback loops, and the adapter is configured to adapt the gain in response to acoustic reflection data of room metadata relating to room boundaries, the acoustic reflection data indicating frequency-dependent acoustic characteristics for at least one room boundary that is not one of the pair of opposing room boundaries.

[0068] In accordance with an optional feature of the invention, the group of feedback loops includes at least two feedback loops having different loop gains.

[0069] According to another aspect of the present invention, there is provided a method for generating a flutter echo audio signal, the method comprising: receiving room metadata indicative of characteristics of a room; determining a flutter echo estimate for the room in response to the room metadata, the flutter echo estimate indicative of a level of flutter echo in the room; generating the flutter echo audio signal from an output signal of a group of feedback loops to which an audio source signal is supplied, the group of feedback loops comprising a feedback loop of a plurality of feedback loops in a feedback delay network; and adapting a first parameter for a first feedback loop of the group of feedback loops in response to the flutter echo estimate.

[0070] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0071] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which: [Brief description of the drawings]

[0072] [Figure 1] FIG. 1 shows an example of a room impulse response. [Diagram 2] FIG. 2 shows an example of a room impulse response. [Diagram 3] FIG. 3 shows an example of elements of a virtual reality system. [Figure 4] FIG. 4 illustrates an example of an audio device for generating a flutter echo audio signal according to some embodiments of the present invention. [Diagram 5] FIG. 5 illustrates an example of a signal generator for generating an audio signal according to some embodiments of the present invention. [Figure 6] FIG. 6 shows an example of a Jot reverberator. [Figure 7] FIG. 7 illustrates an example of a flutter echo signal generator according to some embodiments of the present invention. [Figure 8] FIG. 8 shows an example of a room impulse response. [Figure 9] FIG. 9 shows an example of a flutter echo. [Figure 10] FIG. 10 shows an example of a room impulse response characteristic. [Figure 11] FIG. 11 shows an example of a flutter echo between two opposing walls. [Figure 12] FIG. 12 illustrates an example of a signal generator circuit for generating an audio signal according to some embodiments of the present invention. [Figure 13] FIG. 13 illustrates an example of a signal generator circuit for generating an audio signal according to some embodiments of the present invention. [Figure 14] FIG. 14 shows an example of a flutter echo between two opposing walls. [Figure 15] FIG. 15 shows an example of distance gain as a function of the number of reflections between walls. [Figure 16] FIG. 16 illustrates an example of a circuit of a signal generator for generating an audio signal according to some embodiments of the present invention. [Figure 17]FIG. 17 illustrates an example of a circuit of a signal generator for generating an audio signal according to some embodiments of the present invention. [Figure 18] FIG. 18 illustrates an example of a signal generator circuit for generating an audio signal according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0073] Although the following description focuses on audio processing and generation for virtual reality applications, it will be appreciated that the principles and concepts described can be used in many other applications and embodiments.

[0074] Virtual experiences that allow users to move through virtual worlds are becoming increasingly popular, and services are being developed to meet such demand.

[0075] In some systems, VR applications may be provided locally to a user, for example by a standalone device that does not use or even have access to any remote VR data or processing. For example, a device such as a game console may include a memory for storing scene data, an input for receiving / generating user poses, and a processor for generating corresponding images from the scene data.

[0076] In other systems, the VR application may be implemented and executed remotely from the user. For example, the user's local device may detect / receive motion / pose data, which is transmitted to a remote device, which processes the data to generate a user pose. The remote device may then generate a view image and corresponding audio signal appropriate for the user's pose based on scene data describing the scene. The view image and corresponding audio signal are then transmitted to the user's local device for presentation therein. For example, the remote device may directly generate a video stream (usually a stereo / 3D video stream) and a corresponding audio stream, which are directly presented by the local device. Thus, in such an example, the local device would not perform any VR processing other than transmitting motion data and presenting the received video data.

[0077] In many systems, functionality may be distributed between local and remote devices. For example, the local device may process received input and sensor data to generate a user pose, which is continually transmitted to the remote VR device. The remote VR device may then generate a corresponding view image and a corresponding audio signal, which it may transmit to the local device for presentation. In other systems, the remote VR device does not directly generate a view image and a corresponding audio signal, but may select relevant scene data and transmit this to the local device, in which case the local device may generate the presented view image and corresponding audio signal. For example, the remote VR device may identify the closest capture point and extract the corresponding scene data (e.g., a set of object sources and their position metadata), which it may transmit to the local device. In this case, the local device may process the received scene data to generate images and audio signals for a particular current user pose. The user pose typically corresponds to a head pose, and a reference to the user pose may be equivalently considered to typically correspond to a reference to the head pose.

[0078] In many applications, particularly for broadcast services, a source may transmit or stream scene data in the form of images (including video) and audio representations of a scene independent of the user's pose. For example, signals and metadata corresponding to audio sources within a particular virtual room may be transmitted or streamed to multiple clients. In this case, each client may locally synthesize an audio signal corresponding to the current user's pose. Similarly, a source may transmit a schematic description of the audio environment, including a description of the audio sources within the environment and the acoustic properties of the environment. In this case, an audio representation may be generated locally and presented to the user, for example using binaural rendering and processing.

[0079] 3 shows such an example of a VR system in which remote VR client devices 301 communicate with a VR server 303 over a network 305, such as the Internet. The server 303 can be configured to potentially support multiple client devices 301 simultaneously.

[0080] The VR Server 303 can support the broadcasted experience by, for example, transmitting an image signal including an image representation in the form of image data that can be used by a client device to locally synthesize a view image corresponding to an appropriate user pose (pose refers to a position and / or orientation). Similarly, the VR Server 303 can transmit an audio representation of the scene, allowing the audio to be locally synthesized for the user's pose. In particular, as the user moves around in the virtual environment, the images and audio synthesized and presented to the user are updated to reflect the user's current (virtual) position and orientation in the (virtual) environment.

[0081] Thus, in many applications, such as that of FIG. 3, it will be desirable to generate efficient image and audio representations that can be modeled on a scene and efficiently included in a data signal that can be transmitted or streamed to various devices, allowing these devices to locally synthesize views and audio for poses different from the capture pose.

[0082] In some embodiments, a model representing a scene can be stored, for example, locally and used locally to synthesize appropriate images and audio. For example, an audio model of a room can include an indication of the characteristics of audio sources that can be heard in the room, as well as the acoustic characteristics of the room. The model can then be used to synthesize appropriate audio for a particular location.

[0083] How an audio scene is represented and how this representation is used to generate audio are important issues. Audio rendering, which aims to provide a natural and realistic effect to the listener, usually involves a rendering of the acoustic environment. For many environments, this involves a representation and rendering of the diffuse reverberation present in the environment, such as a room. It has been found that the rendering and representation of such diffuse reverberation has a significant impact on the perception of the environment, such as whether the audio is perceived to represent a natural and realistic environment. In the following, an advantageous approach for representing an audio scene and for rendering audio based on this representation, and in particular for enhancing diffuse reverberant audio, will be described.

[0084] The approach is described with reference to an audio device shown in Figure 4. The audio device is configured to generate an audio output signal representative of audio in an acoustic environment. In particular, the audio device is capable of generating audio representative of the audio perceived by a user moving around in a virtual environment having multiple audio sources and given acoustic characteristics. Each audio source is represented by an audio signal representative of the sound from that audio source, and metadata capable of describing characteristics of that audio source (such as providing an indication of the level of the audio signal and / or the location of the audio source). Furthermore, metadata is provided to characterize the acoustic environment.

[0085] In this example, the audio device is specifically configured to generate an audio signal representative of how an audio source is perceived in the current listening environment. The device has the capability to generate direct and early reflection audio signal components as well as diffuse reverberation audio signal components. In this way, the audio device can receive one or more audio source signals and process one, some or all of them to generate a corresponding output signal that includes different components reflecting the behavior of the acoustic environment.

[0086] Furthermore, the device is configured to generate a flutter echo audio signal dependent on room metadata indicative of room characteristics. If the acoustic environment is a room, the room is characterized by room metadata and the audio device may be configured to generate a flutter echo audio signal emulating a flutter echo that may occur in such a room. The flutter echo audio signal may be an additional audio component that is combined with the direct sound, early reflections and / or diffuse reverberation audio components to provide a more accurate and natural perceived acoustic environment (however, in some embodiments only the flutter echo audio signal is generated). Furthermore, since flutter echo is typically very specific to an individual room and indeed tends to be significant or noticeable only for some room types / characteristics, the audio device may specifically provide a flutter echo audio signal when appropriate for a particular room, and typically the flutter echo audio signal is adapted to reflect these particular conditions. In particular, in many embodiments the generation of the flutter echo audio signal may be conditional on the room metadata and the flutter echo audio signal may be generated only if the room metadata meets certain criteria.

[0087] For some room types and characteristics, opposing (especially parallel) boundaries / walls of the room, in addition to helping generate possible early reflections and diffuse reverberation, may also produce a certain rate of repeating echoes. Such effects may be perceived as flutter echoes reflecting sound bouncing between opposing walls, with energy decaying as the order of reflection increases. Flutter echoes may include many frequencies (specifically, for example, all audio frequencies) and are not limited to standing wave frequencies, such as those known from room modes. They tend to be most noticeable for mid and high frequencies.

[0088] In the case of flutter echoes, the reflected sound essentially returns from the reflecting walls at a constant rate but at a slightly lower level. The rate of echoes depends on the distance (i.e., the time of flight) between the walls that give rise to the echo. The reduction in level depends on the distance attenuation and the reflective properties of the walls involved. These parameters are usually frequency dependent.

[0089] Flutter echo is an acoustic property that can occur in many rooms where the characteristics of the particular room allow for proper reflections, for example, hallways, stairwells, or rooms with very different material properties on different boundaries. Including an emulation of this acoustic effect can provide a compelling experience and generate a greater sense of immersion for the user. Nevertheless, commonly used methods cannot and do not perform such emulation.

[0090] The audio device of Fig. 4 comprises a receiver (RX) 401 configured to receive room metadata indicative of characteristics of a room, among other things. A flutter echo audio signal is generated to represent flutter echoes in the room, and the generated output signal may include, among other things, a flutter echo audio signal reflecting a particular flutter echo characteristic of the room.

[0091] The device specifically uses a feedback delay network to generate the flutter echo audio signal. Such a feedback delay network can also be used by a parametric reverberator to generate diffuse reverberation, so that the function can reuse the same functionality. Such an approach can provide reduced complexity and / or ease of operation, for example in some embodiments can enable dynamic and flexible allocation of resources between diffuse reverberation and flutter echo simulation depending on the characteristics of a particular room. In contrast to the existing configuration of a feedback delay network in a parametric reverberator, the approach of FIG. 4 can add further characteristic features of room acoustics to the set of simulation tools used for audio rendering, thereby providing a more realistic modeling of a typical room in the virtual rendering.

[0092] The apparatus of FIG. 4 is configured to generate a flutter echo audio signal, while comprising a receiver 401, which is configured to receive room metadata indicative of characteristics of the room.

[0093] The room metadata may include data characterizing the dimensions of a room, such as, in particular, the three-dimensional dimensions of a rectangular room. In some embodiments, only one or two dimensions of a room may be represented by the room metadata. The remaining dimension(s) may be, for example, pre-defined or assumed dimensions, e.g., the room metadata may indicate the width and length of a room and the audio device may assume a standard height. In some embodiments, absolute dimension data may be provided, while other embodiments may instead or in addition use relative dimension data information. In some embodiments, a room outline may be provided indicating, for example, the distances between the sides / borders / walls of a room, as well as the layout of the room.

[0094] Dimensional data may be provided in different ways in different embodiments, for example room metadata may include distance in meters, room volume including dimensional ratios, time of flight for each dimension, 2D or 3D data as a mesh, etc.

[0095] In some embodiments, the room metadata may include acoustic reflection data, such as reflection or absorption coefficients, for one or more walls of the room, and in many cases for all walls / boundaries of the room.

[0096] Such information may be provided as acoustic absorption, transmission, coupling and diffusion coefficients for each wall in the room.

[0097] In addition to the room metadata, the receiver 401 may also receive one or more audio source signals representing the audio of the audio sources in the room to be rendered. In many embodiments, the audio sources may be represented by audio objects, although it will be appreciated that the particular audio source signals will depend on the particular embodiment and may be, for example, channel sources or Higher Order Ambisonics (HOA) sources. The audio device is configured to generate an output signal relating to one or more of the received audio source signals / objects, typically generating an output signal including all audio sources. In many cases, the output signal will be generated from a subset of all audio sources that have position metadata indicating that they are in the room. The audio device may, among other things, process all received audio source signals to generate an output signal that reflects the acoustic characteristics of the room, including the direct sound path, early reflections, diffuse reverberation and flutter echoes. This processing may, for example, be applied to each audio source signal sequentially or in parallel. The resulting output signals may be combined to generate a single rendering signal. For example, a binaural stereo signal may be generated by binaurally processing (at least a portion of) the output signals generated for each source and then combining the binaural signals into a single output stereo signal.

[0098] It will be appreciated that the described approach may be applied to audio devices that generate only flutter echo audio signals and, for example, do not generate any direct, early reflection and / or diffuse reverberation signal components, however, the following description will focus on embodiments in which the audio device is configured to simulate a range of acoustic effects of a typical acoustic environment.

[0099] The audio device comprises a signal generator 403 arranged to generate one or more output signals from one or more (usually all) received audio source signals. The signal generator 403 in this example generates the output signals to reflect the intended acoustic environment.

[0100] 5 shows an example of a signal generator 403. The audio device comprises a path renderer 501 for each audio source. Each path renderer 501 is configured to generate a direct path signal component representing a direct path from the audio source to a listener. The direct path signal component is generated based on the position of the listener and the audio source, in particular by scaling the audio signal in a distance-dependent and potentially frequency-dependent manner with respect to the audio source, and the relative gain for audio sources in a particular direction with respect to the user (e.g. in the case of non-omnidirectional sources).

[0101] In many embodiments, the renderer 501 can also generate direct path signals based on occlusion or diffraction (virtual) elements that lie between the source position and the user position.

[0102] In many embodiments, the path renderer 501 may also generate further signal components for the individual paths if they include one or more reflections. This may be done, for example, by evaluating reflections off walls, ceilings, etc., as known by those skilled in the art. In this manner, the path renderer 501 may also generate an early reflection component. The direct path and reflected path components may be combined into a single output signal for each path renderer, such that a single signal representing the direct path and early / discrete reflections may be generated for each audio source.

[0103] In some embodiments, the output audio signal for each audio source may be a binaural signal generated, for example, by applying HRTF or HRIR filters based on the relative (angular) positions of the audio source and the listener, and thus each output signal may include both left-ear and left-right-ear (sub-) signals.

[0104] The output signals from the path renderers 501 are provided to a combiner 503, which combines the signals from the different path renderers 501 to generate a single combined signal. In many embodiments, a binaural output signal may be generated and the combiner may perform a combination, such as a weighted combination, of the individual signals from the path renderers 501. That is, all right ear signals from the path renderers 501 may be added together to generate a combined right ear signal, while all left ear signals from the path renderers 501 may be added together to generate a combined left ear signal.

[0105] It will be appreciated that binaural rendering can be replaced by rendering to a speaker configuration (e.g. 2.0, 5.1, 7.1, 9.1.4, 22.2) using a panning algorithm such as VBAP to generate two or more speaker signals. The combiner 503 will in most such embodiments combine all the contributions to each speaker signal in the speaker configuration.

[0106] The path renderers and combiners may be implemented in any suitable manner, typically comprising executable code for processing on a suitable computational resource such as a microcontroller, microprocessor, digital signal processor, or central processing unit including supporting circuitry such as memory. It will be appreciated that the multiple path renderers may be implemented as parallel functional units, for example a bank of dedicated processing units, or as iterative operations for each audio source. Typically the same algorithm / code is executed for each audio source / signal.

[0107] In addition to the individual path audio components, the audio device is further configured to generate a signal component representative of the diffuse reverberation in the environment. The diffuse reverberation signal is (effectively) generated by combining the source signals into a downmix signal and then applying a reverberation algorithm to the downmix signal to generate the diffuse reverberation signal.

[0108] The audio device of Fig. 5 comprises a downmixer (DMX) 505, which receives audio signals of several sound sources (typically all sound sources in the acoustic environment for which the reverberator is simulating diffuse reverberation) and combines these audio signals into a downmix. The downmix thus reflects all sounds generated in the environment. The downmix is ​​fed to a reverberator (REVB) 507, which is configured to generate a diffuse reverberation signal based on the downmix. The reverberator 507 may in particular be a parametric reverberator such as a Jotto reverberator. The reverberator 507 is coupled to a combiner 503, to which the diffuse reverberation signal is fed. In this case, the combiner 503 combines said diffuse reverberation signal with path signals representative of the individual paths to generate a combined audio signal representative of the combined sound in the environment as perceived by the listener.

[0109] An example of a suitable reverberator is the Jot reverberator shown in Figure 6. This reverberator contains a loop input vector b and a loop extraction matrix C to control how the input samples are distributed to the feedback loop of the reverberator and how the output signal is generated from the loop.

[0110] The audio device further comprises an echo signal generator (FLTECHGEN) 509 configured to generate a flutter echo audio signal (in many embodiments multiple flutter echo audio signals may be generated). The echo signal generator 509 receives the input audio source signal and generates one or more flutter echo audio signals, which are fed to a combiner 503 where they are combined with other generated signal components to provide an output signal reflecting the acoustic characteristics of the room being simulated.

[0111] The echo signal generator 509, and therefore the signal generator 403, includes a feedback delay network with multiple feedback loops.

[0112] An example of such a feedback delay network of the echo signal generator 509 is shown in FIG. 7 where three feedback loops are illustrated. The feedback delay network may have multiple feedback loops, each (or at least one) feedback loop having an input receiving an input audio signal, each feedback loop configuring a loop transfer function (which may in particular be a delay), a feedback network that returns an output signal of the feedback loop to an input of the loop to be combined with the input audio signal, and an output circuit configured to generate the output signal of the feedback delay network as a combination of the output signals of the feedback loops. The feedback network may configure, for each feedback loop, a feedback path for the output signal of the feedback loop to the input of the feedback loop, and typically also a feedback path to one or more inputs of the other feedback loops. In many embodiments, the feedback network may configure a feedback path from the output of each feedback loop to each input of all feedback loops. Each feedback path typically may configure an attenuation factor (or equivalently a gain factor), but in some embodiments may provide a more complex feedback path (e.g. provide a filter function), such as for example configuring a frequency dependent gain. In some embodiments, the loop transfer function may be a filter that achieves both the desired frequency response and gain factor, while the feedback path may simply be a flat unity gain feedback (e.g., corresponding to a feedback matrix that represents feedback with a coefficient of 1 on the diagonal). In many embodiments, the feedback network may be represented by a feedback matrix with a coefficient for each feedback loop pair combination.

[0113] Feedback delay networks are usually based on feedback loops with different delays. Input signals are injected into the loops and these signals are fed back to the loops with appropriate feedback gains. The output signal is extracted by combining the signals in the loops. Thus, the input signal is repeated successively with different delays. Using relatively prime delays and a feedback matrix to mix the signals between the loops can create patterns similar to reverberation in real spaces and is particularly suitable for generating diffuse reverberation as in the example of Jot or other parametric reverberators.

[0114] The absolute values ​​of the elements of the feedback matrix are designed to be less than 1 to achieve a stable decaying impulse response. The coefficients can be set in combination with delays to achieve the desired reverberation time (T60). In many implementations, additional gains or filters are included in the loop. These filters can control the decay instead of the matrix. Using filters has the advantage that the decay response can be different for different frequencies.

[0115] In the audio device, such a feedback delay network can be used to generate a flutter echo audio signal, and in many embodiments, a feedback delay network can be used to generate both the flutter echo audio signal and the diffuse reverberation. In particular, the same feedback delay network can be used for both, with parameter values ​​being determined to provide the desired effect. In particular, if a flutter echo is to be generated, all feedback loops of the feedback delay network can be used to generate the diffuse reverberation component, and the parameters can be set accordingly. If a flutter echo audio signal is to be generated, one or more (usually only a small number, such as 2 or 3 or less) feedback loops can be used to generate the flutter echo audio signal, and the remaining feedback loops are used to generate the diffuse reverberation signal. The reassigned feedback loops are then set with appropriate parameters for generating the flutter echo audio signal. In many embodiments, a total of, for example, 8 to 20 feedback loops can be provided, and if appropriate, 3 or less of these are used to generate the flutter echo audio signal.

[0116] As a specific example, the approach can provide a way to include flutter echo simulation using existing configurations of feedback delay networks within a parametric reverberator that generates diffuse reverberation, thereby adding further characteristic features of room acoustics to the suite of simulation tools and providing a more realistic modeling of typical rooms in virtual renderings.

[0117] Thus, the feedback delay network may be common to the echo signal generator 509 and the reverberator 507 .

[0118] In the example of Figure 7, an input signal is fed to each feedback loop via an input circuit comprising a pre-gain section 701. The inputs of these feedback loops comprise a combiner 703 which combines the input audio source signal with the signal fed back to the feedback loop. Each loop comprises a loop filter 705 (which may include a delay) whose output is fed to a feedback network / matrix 707 which provides feedback to the loop inputs. Furthermore, an output circuit combines the output signals from each loop into an output signal. The output circuit specifically comprises a group of gain sections 709 and a combiner 711 configured to generate the output signal of the feedback delay network as a weighted combination of the output signals from each feedback loop.

[0119] The audio device is configured to adapt the generation of the flutter echo audio signal. In particular, in many embodiments, the audio device can be configured to adapt the degree or level of the flutter echo depending on the room characteristics of the simulated room, and indeed in many embodiments, the audio device can adapt whether or not the flutter echo audio signal is generated depending on the room characteristics. In this way, the simulation of the flutter echo is not merely a static generation of a flutter echo audio signal providing a flutter echo effect, but rather a dynamically adapted generation of flutter echo depending on the room characteristics, and in particular, rather than always producing the flutter echo effect, in many embodiments, such generation can be performed only if it is determined that the flutter echo is likely to be significant in a particular room.

[0120] The audio device comprises an estimator (EST) 405, which is configured to determine a flutter echo estimate of the room based on the received room metadata, the flutter echo estimate being indicative of the level / degree / amount / prevalence of the flutter echo in the room.

[0121] The exact approach and algorithm or function for determining the flutter echo estimate may vary between different embodiments and may depend on the exact performance and behavior desired for a particular application. In many embodiments, the flutter echo estimate may be generated to indicate an increased level of flutter echo when the room metadata indicates that the reflection between one pair of opposing boundaries / walls is higher than for other pairs of boundaries / walls. This may be the case, for example, when the opposing walls of the pair are farther apart from each other than the opposing walls of other pairs and / or when the combined reflection attenuation of the opposing walls of the pair is lower than for other pairs of walls. In such cases, the echo occurring between the opposing walls of the pair may be significantly stronger than other reflection paths occurring between the walls, which may lead to a larger flutter echo (generated by the opposing walls of the pair) relative to other reflections that produce, for example, diffuse reverberation, etc. That is, these flutter echoes may decay slower than other reflections that produce, for example, diffuse reverberation. This may result in a more noticeable flutter echo after a certain time, for example 30 milliseconds, after emission by the source.

[0122] The estimator 405 is coupled to an adaptor (ADP) 407, which is configured to adapt parameters of at least one of the feedback loops of the feedback delay network in response to the flutter echo estimate. In many embodiments, the parameters can be a feedback coefficient of a loop to itself (which may be frequency dependent), a feedback coefficient from a loop of the feedback delay network to another loop (which may be frequency dependent), a feedback coefficient from another loop to this loop (which may be frequency dependent), a loop gain / weight, a loop delay, a loop transfer function, and / or an extraction coefficient / weight for generating an output signal.

[0123] In many embodiments, a common feedback delay network can be used for generating the diffuse reverberation and for generating the flutter echo signal. In such a case, a feedback loop can be dynamically assigned to be used either for the diffuse reverberation generation or for the flutter echo audio signal generation, which can be done by adapting the parameters of the loop to be appropriate for the diffuse reverberation or the flutter echo audio signal. Thus, in many embodiments, the adapter 407 can be configured to switch, for at least one feedback loop, between parameter values ​​for generating the diffuse reverberation signal and parameters for generating the flutter echo audio signal in response to the flutter echo estimate.

[0124] The audio device can thus be configured to determine the degree of flutter echo believed to be present in a room, and the feedback loop of the feedback delay network can be set to generate a flutter echo audio signal corresponding to the flutter echo.

[0125] This approach can provide improved acoustic simulation in many embodiments, particularly providing more natural sounding audio when simulating rooms with particular characteristics, allowing certain flutter echoes to be more pronounced without sacrificing performance for rooms where the flutter echoes are not loud or even noticeable.

[0126] The main driving factor that defines the reverberation response is the travel distance of the sound waves. The travel distance causes attenuation and delay. However, each reflection off a surface causes additional attenuation without adding any delay. Thus, repeated reflections in small room sizes will decay faster than for larger room sizes. Flutter echoes will decay faster in short room sizes than in large room sizes.

[0127] The decay rate of flutter echoes is almost always consistent with the room's reverberation time T60, since the various dimensions of the room are roughly similar. This means that flutter echoes are mixed with other reflections that take different paths across multiple dimensions. These result in less regular reflection behavior. Due to similar decay characteristics, flutter echoes are not particularly noticeable in many situations and are not considered in typical current approaches.

[0128] However, if one room dimension deviates significantly from the other dimensions, there will be a flutter echo in this dimension that deviates significantly from most of the room's reflectivities. It will decay slower than other reflection paths because it has fewer reflection interactions with the room boundaries. This makes the flutter echo stand out from the rest of the reverberation, as fewer reflections result in less decay over time and the flutter echo will be correspondingly more audible. An example of a room impulse response showing a flutter echo is shown in Figure 8 (the example is for a hallway with dimensions 40x2x2.5m).

[0129] Similarly, flutter echoes can be noticeable as reverberant responses when two parallel walls in a room are significantly more reflective than the other walls. This causes flutter echoes in this dimension to decay more slowly, as each interaction with the wall is less destructive than in flutter echoes in other dimensions and reflection paths that span multiple dimensions.

[0130] As mentioned above, flutter echoes can result from the repeated bouncing of sound waves between two parallel surfaces. Such echoes tend to be present in all rooms, but may be more noticeable in some rooms depending on their shape or the relative material properties of their boundaries.

[0131] In this example, the estimator 405 can generate a flutter echo audio signal to reflect the difference in room dimensions. The room metadata can include room dimension data, and the adapter 407 can determine a flutter echo estimate based on the room dimension in a first direction relative to the room dimension in a second direction. For example, the horizontal dimension between two parallel wall pairs in a rectangular room can be determined from the room size information indicated by the room metadata. The ratio of the longest dimension to the shortest dimension (or the second longest dimension) can then be determined and used as an indication of how strong the flutter echo is. That is, the ratio can be used directly as a flutter echo estimate.

[0132] The adapter 407 may then, for example, compare the flutter echo estimate in the form of a ratio to a threshold and configure some of the feedback loops of the feedback delay network to generate a flutter echo audio signal if the threshold is exceeded, while configuring the loops to instead contribute to the generation of diffuse reverberation (and thus no flutter echo audio signal is generated) if the ratio is below the threshold. In other embodiments, a more step-wise approach is used, for example by permanently using one or more feedback loops to generate a flutter echo audio signal, but having it have an amplitude that is a monotonically increasing function of the ratio / flutter echo estimate.

[0133] Alternatively or additionally, the adaptor 407 may in some embodiments determine the flutter echo audio signal in response to changes in the acoustic return loss of the sides / boundaries / walls of the room. The room metadata may include the acoustic return loss of the walls of the room, and the flutter echo estimate may be generated to reflect these changes. In particular, the flutter echo estimate may be generated as a function of the difference between the composite acoustic return loss of a pair of opposing sides of the room and the composite acoustic return loss of another pair of opposing sides of the room. For example, a ratio between such composite acoustic return losses may be determined, and the flutter echo estimate may be generated directly as this ratio. The larger the difference, the larger the flutter echo estimate. As described for the dimension example, the adaptor 407 may adapt its operation based on the ratio.

[0134] It will be appreciated that in many embodiments the flutter echo estimate may be generated as a combination of different considerations, and in particular in many embodiments both the dimensions of the room and the acoustic return loss of the walls / sides of the room may be taken into account when generating the flutter echo estimate.

[0135] As mentioned above, one possible cause of noticeable flutter echoes is a room in which one deviating dimension is significantly longer than the other, such as a hallway. In such a case, the echoes of two opposing walls in the deviating dimension will have longer path lengths that result in a noticeable flutter echo from the rest of the room impulse response (RIR). However, the reflection path that is completely perpendicular to the wall will involve additional reflections at other boundaries in the shorter dimension, but the lateral spread may be captured by a relatively small reflection path.

[0136] As a result, the path length of a significant portion of the early reflections is dominated by the distance in the dimension they diverge in. This effect is stronger the higher the reflection order. If the mapped source is spread out in one dimension by, say, about 40 meters, a spread in the other dimension by, say, about 4 meters, does not increase the distance very much. Thus, multiple reflections of different orders will be grouped close to each other in the RIR with slightly different delays.

[0137] This means that flutter echoes are not caused purely by sound waves bouncing back and forth between two parallel surfaces. The effect is simply to produce the first strongest reflection in a series of reflections. Many subsequent reflections may follow, representing one or more shallow additional reflections on one of the longer boundaries. These produce clearly visible repeated bursts of concentrated energy in the RIR. This can result in flutter echoes, where each echo essentially contains a series of compound reflections, rather than just a single echo reflection.

[0138] As the order of the main flutter echo becomes higher, other reflections of similar distances become more tightly compressed in time. That is, the path length that bounces once or twice off a long room boundary will be closer to the path length without reflections off that long boundary than for lower orders. An example of such a compound flutter echo is shown in Figure 9, which also shows the possible time compression.

[0139] Operating in the digital domain, this means that at any given time, multiple reflections contribute to the same (discrete) filter delay: these contributions add up, making the impulse response amplitude of these bursts larger than it would be with an infinite sample rate.

[0140] In a particular example, the audio device implements an approach that adds the simulation of flutter echoes by utilizing the existing framework of a parametric reverberator, thus leaving the overall complexity of the audio device substantially unchanged.

[0141] The audio device performs the following actions: · Room dimensions; the location and orientation of room boundaries; and / or · Material properties related to room boundaries; It can be based on.

[0142] Based on this metadata, the estimator 405 can first determine whether flutter echo is a likely audible acoustic feature of the room the user is in. For example, this may be the case if one dimensional dimension is significantly larger than the other two, or if the reflective properties of the material on the walls in one dimension are significantly larger than in the others. A flutter echo estimate reflecting this may be generated.

[0143] The adapter 407 can adapt the operation of the signal generator 403 in response to the flutter echo estimate: if this indicates that the flutter echo is significant, the configuration parameters of the feedback delay network of the parametric reverberator are altered so that one or more of its feedback loops models the flutter echo.

[0144] The adapter 407 can then set the loop delay to be proportional to the dimensions of the room in which the flutter echo occurs, the loop filter can be set to correspond to the (combined) material properties of the walls involved in the flutter echo, and the feedback matrix can be adapted to isolate the loop from the rest of the normal feedback loop. In this way, multiple parameters of the feedback loop can be set to emulate the flutter echo.

[0145] Thus, in some embodiments, a flutter echo estimate can be generated and evaluated to determine whether or not to simulate a flutter echo. This is only necessary if the flutter echo is audible. Typically, there are two possible main root causes for an audible flutter echo: One room's dimensions are significantly larger than the other two; and The reflective properties of the material on one wall dimension are significantly stronger than in the other two.

[0146] Combinations of the above can also cause flutter, for example when two dimensions are significantly larger than the third, but one of these is much less reflective than the other.

[0147] A dimension of a room may be considered significantly larger than the other two if, for example, it is twice as large as the maximum of the other two dimensions. An alternative criterion may be if one dimension of a room is at least 3.1 times longer than the average dimension of the other two dimensions. In some embodiments, this may be if one dimension of a room is at least 50% longer than the average of all three room dimensions.

[0148] If the room is not a rectangular prism (shoebox), the dimensional dimensions can be set to the outer limits of the geometric structure in all three dimensions.

[0149] As another example, a room may qualify for flutter echo simulation if the material properties of the room boundaries in one dimension are significantly different from those in other dimensions. Reflection can be represented by a parameter reflecting the acoustic reflection attenuation, such as a reflection coefficient or an absorption coefficient. For example, the average reflection coefficient (value between 0 (non-reflective) and 1 (fully reflective)) of both walls in one room dimension is at least 0.2 higher than the maximum average reflection coefficient of both walls in two other directions. Similarly, the average reflection coefficient of each pair of walls can be compared to the average of all walls or the average of two other wall pairs. For example, the average reflection coefficient is at least 20% higher than the overall average. Furthermore, a minimum required reflection coefficient can be introduced. For example, the average reflection coefficient must be at least 0.67.

[0150] In other embodiments, absorption coefficients can be used to reflect acoustic return loss, and these coefficients may be required to be smaller in candidate flutter dimensions than in other dimensions, for example, an average absorption coefficient less than 85% of the average absorption coefficient of wall pairs in other dimensions.

[0151] Reflection (or absorption) coefficients are often frequency dependent. They can be averaged over all frequencies, or over a subset of frequencies. Furthermore, averaging can be done over wall sections with different material properties.

[0152] In this manner, a flutter echo estimate can be generated to reflect such parameters, and the adaptor 407 can determine whether or not to simulate a flutter echo based on whether the flutter echo estimate meets appropriate criteria.

[0153] Estimating flutter echoes, and in particular determining whether to simulate flutter echoes, may include consideration of a combination of room dimensions and material properties. For example, any of the separate evaluation criteria may be met to cause flutter echoes to be simulated. Other embodiments may simulate flutter echoes only if the room dimensions are significantly larger and the corresponding average material properties are significantly different. Optionally or alternatively, a minimum reflection coefficient for the candidate flutter dimension may be additionally required.

[0154] In some embodiments, dimensions and material properties are combined with an estimated decay time (e.g., T60). If the estimated 1D decay time in a dimension is at least 30% longer than the maximum of the 1D decay times in the other two dimensions, a flutter echo can be simulated in that dimension. In other embodiments, the decay time can be required to be at least 0.5 seconds longer than in the other dimensions.

[0155] The decay time can be estimated from the dimensions and the average reflection coefficient of the corresponding walls. In the time it takes a sound wave to travel round trip through a room in that dimension, it is attenuated by the distance traveled and the two reflections on the walls. As an example, the estimated T60 decay time is:

number

[0156] This formula determines the attenuation in one round trip path through a room of size D. ref and the average reflection coefficient

number

[0157] In other embodiments, the estimated one-dimensional decay time can be compared to the whole-room decay time, for example if the one-dimensional T60 is 10% longer than that estimated for the whole room. The whole-room T60 can be estimated with equations such as the Sabine or Norris-Eyring equations.

[0158] The decision whether a flutter echo should be simulated can also be a soft decision: for example, by choosing a low threshold when the flutter echo is unlikely to be audible and a high threshold when the flutter echo is likely to be audible, any case between these thresholds will yield a confidence level between 0 and 1. A weight w=0 corresponds to no audible flutter echo, while w=1 corresponds to full confidence that a flutter echo is audible.

[0159] For example, the one-dimensional decay time in dimension 1

number

number

[0160] In some embodiments, the room characteristics may not be directly available, for example, they may be characterized by a room impulse response. In some embodiments, the room metadata may include the RIR, and the estimator 405 may be configured to generate a flutter echo estimate in response to the RIR. In this example, the parameters of the feedback delay network may be determined from the flutter echo estimate generated from the RIR. Measuring the impulse response is more amenable to rooms of arbitrary shapes that deviate from the rectangular shoebox model.

[0161] In such an embodiment, the presence of flutter echoes can be detected by a smoothed version of the magnitude squared IR(e smooth (n)) can be measured using IR(e min By applying minimum tracking to (n), the flutter echo components can be isolated because the noticeable flutter echoes decay more slowly than the remaining reverberation reflections, and tracking the minimum approximates the reverberation decay envelope. An example of this is shown in Figure 10.

[0162] Subtracting the two signals isolates the flutter echo component, if present. If the energy of this signal exceeds a certain threshold, it can be determined that the flutter echo is present. This determination is made based on the percentage of reverberation, i.e.:

number

[0163] As another example, the difference between two echograms, e smooth (n)-e min (n) can be used to derive properties related to the delay and decay of the flutter echo, which can be used to construct a feedback delay network.

[0164] In some embodiments, a peak extraction algorithm can be used to extract the local maxima and their timestamps. The decay rates of these echoes can be determined by fitting an exponential decay model to the peaks. The decay rates and timestamps together can be used to determine the parameters of a feedback loop.

[0165] The adapter 407 can be configured to adapt the parameters in different ways in different embodiments depending on the desired performance. The parameters for generating a flutter echo audio signal can be significantly different from the parameters used by the feedback loop when generating a diffuse reverberation.

[0166] The delay in the feedback delay network for generating the reverberation is usually chosen to be relatively small, allowing the reflection density to increase rapidly - for example, an average of 12 ms is often used, but for high bandwidth signals (e.g. 48 kHz) this will usually be much smaller.

[0167] The choice of delay often depends on the reverberation time (T60), which is usually positively correlated to the room dimensions, but the material properties of the room boundaries also have a large effect on T60; that is, material properties introduce additional attenuations (beyond those caused by distance attenuation) into the RIR without adding latency, and the room dimensions determine the proportion of these attenuations in the RIR. Thus, the configuration of a parametric reverberator is primarily driven by the overall reverberation characteristics T60, and the need to quickly reach a minimum reflection density to accurately model the room (e.g. 1,000-10,000 reflections per second).

[0168] In contrast, if the feedback loop is configured to generate a flutter echo audio signal, the adapter 407 can select a loop delay corresponding to the dimensions of the room to simulate the proportion of flutter echo. The loop filter, which normally simulates the overall reverberation slope T60, can instead be selected to correspond to the average material properties of the walls involved in the flutter echo, to simulate the effect of the walls on each reflection.

[0169] The feedback matrix can in many embodiments be adjusted to keep the flutter echoes separated from the diffuse reverberation products, allowing a consistent reproduction of the flutter echo to be simulated. If there are multiple different flutter echoes in a room, multiple feedback loops can be reused in a similar manner.

[0170] In many embodiments, the adaptor 407 may be configured to increase a feedback coefficient / gain from a first feedback loop to itself in response to the flutter echo estimate indicating an increasing level of flutter echo. If the extent of flutter echo increases, the feedback from a given feedback loop to itself may be increased. Alternatively, or typically in addition, the adaptor 407 may be configured to decrease a feedback coefficient from a first feedback loop to a second feedback loop of the multiple feedback loops in response to the flutter echo estimate indicating an increasing level of flutter echo. The second feedback loop may not be configured to be used to generate flutter echo, but instead may be used to generate diffuse reverberation.

[0171] In some examples, a feedback loop used to generate a flutter echo may only feed back to itself. In some examples, a feedback loop used to generate a flutter echo will not feed back to any other feedback loops configured to generate a flutter echo. In some examples, a feedback loop used to generate a flutter echo may receive a feedback signal only from itself (out of the group of feedback loops used to generate the flutter echo, or perhaps out of all feedback loops of a feedback delay network).

[0172] The adaptation may be, for example, gradual, but in other embodiments, the adaptation may be, for example, a step function. For example, if the flutter echo estimate indicates that the flutter echo is not significant, the appropriate feedback coefficient may be relatively small since the feedback loop may be used primarily to contribute to the diffuse reverberation, and the feedback from a given loop may then be increasingly distributed among different loops to reflect the many different reflections that make up the diffuse echo. However, if the flutter echo estimate indicates that the flutter echo is significant, the feedback coefficient of that loop may be increased while the feedback coefficient for the other loops may be decreased to reflect an increased amount of periodic reflections that correspond to typical flutter echoes.

[0173] Some examples of such adaptations can be explained below with reference to the example of Fig. 11, which shows an example of flutter echo as a function of time and space. In this example, flutter echoes occurring at a source 1101 are transmitted to a listener 1103 at a constant rate (τ γ = 2D / c, where D is the distance between the walls and c is the speed of sound), but arrives at four different offsets depending on where the user and source are between the walls.

[0174] In some low-complexity embodiments, the audio device may be configured to detect a single path length between walls (τ r This can be simplified by reusing only a single reverberator loop, with a delay corresponding to (D'=D / c), which corresponds to the listener and source being halfway between the rooms, in which case both the dotted and solid reflections reach the listener at the same time with a constant rate.

[0175] The feedback matrix for a reverberator with N loops, with the first loop used for flutter echoes, is:

number

number

number

number

number

[0176] Function G d(x) gives the distance attenuation for a sound wave propagating x meters. This can be a simple attenuation based on an omnidirectional source with energy spread over a sphere of radius x. It is well known that every doubling of distance (i.e. radius) produces a 6 dB attenuation. In many embodiments, a reference distance can be used as the distance at which the source signal is defined, in which case the distance attenuation is considered to be included in the signal, and G d The additional distance attenuation from (x) is equal to 0 dB.

[0177] Furthermore, the effect of air absorption G abs (x) and other aspects of G d It can also be added to (x). This effect is usually more pronounced at larger distances and tends to be frequency dependent. Usually the effect of air absorption is very small, especially when considered for realistic room dimensions D.

[0178]

number

[0179] The described embodiment uses the average reflection coefficient

number

[0180]

number

[0181] The reflection coefficients do not have to be averaged, but may be adapted, for example, to the lateral location of the source in front of a wall, i.e., where most of the flutter echoes will occur. In such an embodiment, multiple sources may be grouped into separate loops according to their associated reflection coefficients.

[0182] As another example, the audio device may be configured to emulate the flutter echo of FIG. r )2. The four loops may have separate inputs that are pre-delayed to reflect the offset between the listener / source and the wall. For example, a pre-delay circuit as shown in Figure 12 may be used.

[0183] The feedback matrix in this case is:

number

[0184] The loop filters of the reverberators in the flutter loop will together simulate the average reflection characteristics of the walls. For example:

number

[0185] In this way, each loop filter simulates the attenuation that results from a sound wave propagating through the medium (eg, air) of interest twice over the distance between the walls, and reflections on both walls.

[0186] The advantage of this embodiment is that it simulates the asymmetry of the two loops, similar to how it would be in a real room, and adjustment of the pre-delay can be used to adapt the asymmetry to the user's position in the room, without having to update the parameters of the feedback delay network itself.

[0187] For example, if the listener is 30% of the wall distance from wall 1105 and the source is 15% of the wall distance from wall 1107, the pre-delay can be set based on the first four path lengths from the source to the listener as follows:

number

number

[0188] The previous embodiment can possibly be simplified by combining the pre-delayed signals before feeding the resultant signal into a single feedback loop, for example as shown in the example of FIG.

[0189] The feedback matrix is:

number

[0190] And the loop filter would be the same as in the previous embodiment:

number

[0191] The loop simulates the path length attenuation and reflections on the two walls, while the pre-delay configuration is responsible for creating an offset in the signal. The delay will be the same as in the previous example.

[0192] The pre-delay configuration can also be extended to include gains or filters that simulate wall reflections and distance attenuation in these first paths. Such filters can also include additional filtering and / or attenuation to simulate early propagation and reflections of flutter echoes, since the simulation in the feedback loop does not represent the first few reflection orders. However, such effects are typically already built into the normal reverberation pre-mix and its coloration filters.

[0193] The separate input signals can also come from a single tapped delay line. Parametric reverberators, usually used in combination with direct path and early reflection rendering, include a pre-delay for their normal operation to control where the reverberation starts relative to the direct path and early reflections. If this pre-delay is long enough, a delay buffer can be used as a tapped delay line. In this case, flutter echoes will start early, but this can be compensated for by early reflection modeling.

[0194] As another example, one could use a set of feedback loops with two interacting loops, with a signal alternating between the loops on each iteration. This could be achieved for the first two loops using the following feedback matrices:

number

[0195] The delays in this embodiment can be set for any listener position to create a regular but asymmetric pattern that better matches realistic scenarios. As another example, the delays can be adjusted depending on the user's position between walls. For example, if the user is 30% of the wall distance from wall 1105, the first delay can be τ r1 = (2·0.3·D) / c, while the second delay is τ r2 = (2 (1-0.3) D) / c.

[0196] As with the previous embodiment, a pre-delay configuration can be used to create the missing offset due to the signal bouncing in two directions. This can be done with two delays (Delay1 and Delay2 above) corresponding to the first two paths.

[0197] A particular advantage of this approach is that the two loop filters can simulate each wall separately, i.e. the first delay τ r1 The first filter related to:

number

[0198] Similarly, the second delay τ r2 The second filter related to:

number

number

[0199] A possibility with such an embodiment is that if the flutter loops are excluded from the regular extraction matrix to generate the diffuse reverberation tails, these can be extracted to separate the outputs for rendering with dedicated HRTF pairs.

[0200] In some embodiments, the signal generator 403 has a gain for the audio source signal before it is fed into the feedback loop of the feedback delay network, and the adapter 407 is configured to adapt the gain depending on the position of the audio source relative to the audio source signal. This can in particular, but not necessarily, be combined with the pre-delays discussed above, in particular each delay of the circuits shown in Figures 12 and 13 can include an adaptive gain, which can be adjusted by the adapter 407 based on the position of the audio source, the listener and / or the wall.

[0201] The received data may include audio signals representative of an audio source and its position, which may be used to adapt the gain. In particular, the gain may be adapted based on the position of the audio source relative to the walls / boundaries / sides of the room. Typically, the gain may be adapted based on the distance from the audio source to the wall that is the flutter echo reflector (typically the nearest wall). The pre-gain may be used to adjust the relative strength / level of the overall flutter echo effect, in particular to adjust the level to reflect the strength of the signal when it is first reflected.

[0202] In some embodiments, the pre-gain may be adapted based on the distance of the listener / user, specifically the relative distance from the listener to the source, or the distance from the source to the listener through at least one reflection off a reflecting wall for flutter echoes.

[0203] Furthermore, in many embodiments, the first reflection can be represented by an early reflection simulation, and the flutter echo signal generator 403 can be used only to represent further reflections of the flutter echo. For example, the flutter echo signal generator 403 can be used to generate flutter echo components corresponding to the fourth or subsequent reflections. In such cases, the reflected sound is already attenuated by previous reflections, including both distance attenuation and reflection attenuation. Such effects can alternatively or additionally be represented by the pre-gain.

[0204] In some embodiments, the adapter 407 may be configured to adapt the gain depending on the distance between two walls / sides / boundaries of the room (particularly the wall / boundary / side that produces the flutter echo). In some embodiments, the adapter 407 may be configured to adapt the gain in response to the acoustic reflection attenuation of at least one wall / side / boundary of the room (particularly the wall / boundary / side that produces the flutter echo). In some embodiments, the adapter 407 may be configured to adapt the gain in response to a number of early flutter echo reflections that are not emulated by the set of feedback loops of the feedback delay network assigned to the flutter echo simulation.

[0205] In particular, the (distance) gain component of the loop filter can represent the attenuation for adjustments in the previous loop path (reflection), and the pre-gain can be used to adapt the input signal level, i.e., the level at the start of the reflection being simulated.

[0206] Signals can often be represented at a level that corresponds to a particular reference distance. Compensation / pre-gain can be used specifically to match the level of the signal to the distance it has already traveled before it is injected into the loop, i.e. to represent the initial distance gain. For example, a simulation based on a feedback delay network can be configured to represent flutter echoes from their fourth order (since the first three are represented by early reflection modeling by other algorithms). In this particular example, and with reference to FIG. 14, the input gain is:

number

number

number

[0207] The feedback loop may have an overall loop gain set to reflect the attenuation of the reflection path (which may include one or more reflections depending on the particular approach). The loop gain may be set by a loop filter and / or a feedback coefficient (feedback matrix). In the example described, the feedback coefficient of the loop to itself is set to 1 and the loop gain (less than 1) is determined by the loop filter. The loop gain / attenuation is typically frequency dependent, and the frequency dependence is typically implemented by the use of an appropriate loop filter.

[0208] In different embodiments, the loop gain / attenuation G d Different approaches for determining

[0209] Typically, loop filters contain two main components: material properties (e.g., reflection coefficients) and distance-related gains. Each loop filter may represent one or more reflection coefficients corresponding to reflections on one or two walls, and a distance gain corresponding to the propagation distance consistent with the reflections represented by the average reflection coefficients.

[0210] Since the loop-related distance relative to the reference distance keeps increasing, the required distance decay component must get weaker with each iteration. For example:

number

[0211] This means that successive reflections decay faster than exponentially, which may not be accurately simulated by a single feedback loop. Filters in feedback loops may be constant due to their recursive nature. Any processed sample may contain components of many different iterations.

[0212] If we isolate the energy dispersion component (the most important component), the distance attenuation (with respect to the signal amplitude) is d ref / d. Let every iteration correspond to a progression distance D. With each iteration, the distance d increases by this constant distance D, so that the additional attenuation with respect to the previous iteration is:

number

[0213] The problem is that d is growing with every iteration, and a fixed gain is required with every iteration. If we express the distance in another way, such that d is a multiple of D, we get a similar result, which shows a further simplification:

number

[0214] It can be seen that the effect of distance attenuation at each iteration does not depend much on the actual propagation distance, but on the increase in distance already propagated (following the rule of thumb that the attenuation is 6 dB for every doubling of distance). Figure 15 shows how the distance gain changes with each iteration.

[0215] The distance gain in the first iteration has a very large effect because the distance corresponding to one iteration is relatively small compared to the total propagated distance. Rapidly, the dynamic effect of the distance decay in each iteration decreases (i.e., there is less change between iterations). As a result, the decay approaches an exponential shape.

[0216] As the distance gain approaches unity, the gain per iteration stabilizes towards the average reflection coefficient of the flutter boundary material properties. When simulating flutter echoes, different embodiments can choose different approaches. For example, the average reflection coefficient can be chosen to simulate the decay in the higher orders. Alternatively, a steeper decay can be used to simulate the decay in the lower orders. Or, in most implementations, an intermediate value would be beneficial, so as not to have too steep or too shallow decays. An exact simulation of the slope in the higher orders may not be necessary in many cases, as they will not be heard by the listener. A good trade-off can be made by choosing a slope that corresponds to, for example, the 5th iteration.

[0217] As mentioned above, many embodiments can adjust the input level of the signal injected into the flutter loop. In addition to compensating for the reference distance and reflection orders that are otherwise simulated, the input gain can be adjusted by adjusting the attenuation gain G d It may also be advantageous to adjust for the trade-off selected for : choosing a relatively slow decay may cause flutter echoes to be overly noticeable, while choosing a relatively steep decay may cause the flutter to go inaudible in an accurate simulation.

[0218] Thus, if a relatively slow decay is set, the initial level can be further reduced to prevent it from being too noticeable. The additional decay basically compensates for the faster decay in early iterations that are not accurately modeled in the recursive process. As a result, stronger first reflections may not be accurately modeled. In many cases, these would be (largely) masked by the reverberation anyway.

[0219] As an example, based on the model in Figure 14, a feedback delay network can be used to simulate second order or higher flutter echoes with a decay slope corresponding to the 10th iteration, resulting in a loop filter of:

number

[0220] The initial input gain can be configured to represent the first reflection order:

number

[0221] Compensation for the attenuation lost in the first I=9 iterations is:

number

number

number

[0222] This compensation ensures consistency of both slope and level at the 10th iteration, which can be stored for different attenuation reference iterations J in a lookup table:

number

number

[0223] In some applications it is preferable to simulate both higher levels of low order reflections, as well as lower levels of mid and high order reflections. This may be possible in rooms with relatively low diffuse reverberation energy (e.g. highly absorptive boundaries, except those involving flutter echoes). Such applications may employ embodiments in which two or more loops simulate different decay rates of the same delay.

[0224] If the delays on the input signal and in the flutter feedback loop are equal, reflections will be generated with the same delay. The first flutter loop may be configured with a steep decay and a relatively large input gain, while the second flutter loop may be configured with a slow decay and a relatively small input gain. If the two are combined by an output circuit, the combined effect may more closely resemble an accurate simulation with an iteratively dependent loop gain.

[0225] In some embodiments, the set of feedback loops assigned to generate flutter echoes may thus have at least two feedback loops with different loop gains, but the at least two feedback loops may have the same delay.

[0226] The above embodiments configure loop filters according to the material properties of the walls where the flutter echoes originate. These filters can be extended to include the effect of shallow reflections at the boundaries of long rooms.

[0227] The material properties of the boundary in the flutter dimension (short boundary) have no effect on the energy ratio of the first reflection to the total complex reflection, however it does affect how quickly successive complex reflections decay.

[0228] Conversely, the material properties of the long boundaries (i.e., not along the flutter dimension) determine how quickly each complex reflection decays, and therefore the energy ratio between the first reflection and the total complex reflection. The decay of the first response amplitude in successive complex reflections is not affected by this material.

[0229] As mentioned before, in RIRs these responses vary with the order of the flutter echo they contribute, compressing the individual reflections in time. However, essentially there is an additional contribution due to one, two or more additional material properties. The main effect is that this increases the energy of the individual flutter echoes and their coloration. The coloration is affected by adding the contribution due to the additional frequency-dependent material properties, but also, in theory, by delayed reflections which give rise to a comb filter effect. However, due to the large number of repetitions with different delays, the comb filter effect is not very large.

[0230] Multiple reflections can be modeled by a single reflection. The loop filter H τ can be set to represent a single pulse with a spectral response that matches that of the complex reflection.

[0231] The net effect of multiple reflections also constitutes more energy than a single reflection, and this total energy must also be represented in terms of a single reflection. Due to delays, the amplitudes of the individual responses usually do not add coherently.

[0232] The energy of the complex reflection is:

number

[0233]

number

number

[0234] E c The above expression for ignoring distance attenuation, whose contribution is relatively low. This also makes the energy ratio between the initial amplitude and the complex reflected energy independent of the flutter order.

[0235] In an alternative embodiment, a separate loop with a very short delay can simulate the tail to the main flutter response. This loop is fed only by the main flutter loop and does not have a direct signal input (b i =0), but feeds back to itself. The short delay may depend on the shortest dimension of the room. The attenuation by the filter may be greater at long room boundaries (e.g.

number

[0236] Another alternative is to use a sparse IIR as the loop filter in the flutter loop to simulate the fast decay response of the complex reflections.

[0237] In many embodiments, the audio device can be configured to provide multiple audio source signals to a feedback delay network, and in particular to provide multiple audio source signals to a set of feedback loops that generate a flutter echo audio signal. The audio device can receive audio sources, for example, related to multiple audio sources in a room, and provide multiple (and possibly all) of these signals to a set of feedback loops that generate a flutter echo audio signal. The multiple signals can be, for example, combined into a composite signal, which can then be provided to the set of feedback loops. Each signal can be subject to delay and / or gain adjustment before being combined with the other signals. The gain and / or delay for each signal can be adapted, for example, to reflect the initial and / or relative signal level and / or arrival time of the individual signal (with respect to the other signals). In some embodiments, for example, the gain and / or delay can be common to some or perhaps all of the source signals provided to the set of feedback loops.

[0238] The above-described embodiments may allow accurate simulation of offsets between individual reflections. This may provide a particularly realistic rendering. The described approach focuses on the generation of flutter echoes for a single source, and the characteristics of the loop parameters etc. may depend on the specific characteristics of the source, such as its position. However, in many cases there are two or more sources generating flutter echoes in the simulated room. In such cases, each source may be simulated with its own dedicated feedback loop etc. These may be implemented, for example, with separate parallel paths to premixing and predelay sections before the feedback delay network.

[0239] However, in many applications, such a level of accuracy is not required. The parameters can be set to appropriate values ​​(e.g., arbitrarily or artificially selected values). In some embodiments, they can be selected equally for all simulated sources. For example, the approach shown in FIG. 16 uses separate gains g n can be used where the delays are applied to the input audio source signals before they are synthesized, and then one or more delays are applied to the common signal. This reduces computational and structural complexity. In such an approach, the flutter echo audio signal can still be adapted to the user's position in the room.

[0240] The input to the feedback loop of this feedback delay network is mathematically:

number

[0241] Some embodiments require or benefit from separate inputs to the feedback loops. This can be achieved by expanding the input gain vector b into a matrix B that takes into account more than one signal and maps it to the different loops.

[0242] For example, the inputs provided to a feedback delay network with five feedback loops (P=5) can be processed by an input matrix B:

number

[0243] As another example, and in keeping with FIG. 17, an exemplary matrix is:

number

[0244] The delays produce different offsets for the P different paths from the source to the listener. Typically, for a shoebox shaped room, P=4 per flutter dimension. The delays can be chosen to represent relative offsets to a minimum offset, and the common offset is ignored. In other embodiments, all delays can be set to absolute offsets, potentially dynamically adjusting to the listener's position.

[0245] The delays can also be commonly adjusted to provide an additional common delay component for the flutter echoes. Such a common delay component can be useful to control the offset of the flutter echoes simulated by the parametric reverberator relative to early reflections simulated by other means, for example to ensure proper latency between the last early reflection associated with the flutter dimension and the first simulated flutter echo response from the feedback delay network.

[0246] In some cases, it may be advantageous to start the flutter echoes earlier than the diffuse late reverberation portion. In these embodiments, the input to the flutter loop may bypass the pre-delay and pass only through a dedicated flutter delay that controls the start of the flutter echo simulation relative to the source radiation. For example, the separately generated early reflections may exclude all reflections related to the flutter dimension and instead simulate these using only a feedback delay network.

[0247] In another embodiment, an early reflection signal may be generated and fed into a flutter echo feedback loop of a feedback delay network, the early reflection signal may include only reflections in the flutter dimension.

[0248] In some embodiments, the audio device may be configured such that at least one audio source signal is only fed into the feedback loop used to generate the flutter echo audio signal.

[0249] In some embodiments, the audio device may comprise a spatial processor configured to apply spatial processing to the flutter echo signal, where the spatial processing depends on the position of the source of the audio source signal and / or the side of the room.

[0250] The spatial processing may be a process that modifies or generates spatial cues of the flutter echo audio signal. In particular, the spatial processor may be configured to perform binaural processing of the flutter echo audio signal, for example as shown in Fig. 18, where the spatial processor is represented by two HRTF blocks HRTF1, HRTF2. The spatial processor may use the HRTFs to apply binaural processing to generate a stereo signal that, when rendered by headphones, results in a spatial perception of the flutter echo originating from the appropriate position / direction. For example, the binaural processing may apply HRTF processing based on the position of one of the walls that generates the flutter echo and the position of the listener, such that the flutter echo is perceived to come from the direction of this wall.

[0251] In some embodiments, the spatially processed flutter echo audio signal can be combined with other generated audio components, in particular it can be combined with diffuse reverberation generated by other feedback loops of the feedback delay network, however this diffuse reverberation cannot be spatially processed since it is generally a distributed sound.

[0252] Thus, in some embodiments, the audio device comprises a combiner for combining the spatially processed flutter echo audio signal with the (non-spatially processed) diffuse reverberation signal. In the example of Fig. 18, the combiner MIX can generate a stereo output signal for a set of headphones by combining the spatially processed flutter echo audio signal with the non-spatially processed diffuse reverberation signal and typically other audio components such as direct and early reflection audio components.

[0253] In many embodiments, the feedback delay network can generate a flutter echo audio signal by combining output signals of feedback loops used to generate the flutter echo audio signal. Similarly, a diffuse reverberation signal can be generated by combining output signals of feedback loops used to generate reverberation.

[0254] Typically, the feedback loop of the feedback delay network is used either for the generation of reverberation or for the generation of flutter echoes.

[0255] In most embodiments, the adapter 407 may be configured to assign a set of feedback loops to the generation of the flutter echo audio signal, while the remaining feedback loops are used to generate the reverberation. In such a case, the adapter 407 may be configured to generally keep the loops separated. In particular, the adapter 407 may adapt the feedback coefficients of the feedback loops such that there is no feedback from one feedback loop to any other feedback loop of the set of feedback loops used to generate the flutter echo audio signal, and vice versa. That is, the adapter may set to zero all feedback coefficients of the feedback matrix that relate to feedback between two loops belonging to two different sets.

[0256] Similarly, when generating an output signal, a flutter echo audio signal may be generated by combining output signals of only those feedback loops of a group of feedback loops that are used to generate the flutter echo audio signal, while a reverberation signal may be generated by combining output signals of only those feedback loops of a group of feedback loops that are not used to generate the flutter echo audio signal.

[0257] In many embodiments, the output signal of the flutter feedback loop can be processed in the same way as other feedback loops, by using a weighted combination to generate an output signal that can be represented by an extraction matrix C. This can include, for example, applying correlation and / or coloration filters known from the generation of diffuse reverberation. In this case, the resulting flutter echoes will not originate from a specific direction.

[0258] However, in embodiments where it is desired that the flutter echoes are directional, the flutter echo feedback loop output signals may be extracted separately for alternative processing (as in the example of FIG. 18). The extraction matrix for the (binaural) diffuse reverberation tail is of dimension 2×(N−2), which can be expanded to become 4×N, processing all N feedback loops of the feedback delay network. In the following example where N=4, the first two rows relate to further diffuse reverberation tail processing and the last two rows relate to flutter echoes.

number

[0259] The first and second output signals generated by the extraction matrix may be processed as usual by the remainder of the parametric reverberator function. The third and fourth output signals may be processed separately, for example with different HRTF pairs corresponding to opposite directions of both walls. These may be adaptive depending on the user's orientation.

[0260] This may be particularly advantageous for embodiments in which each wall is simulated in a separate loop: a first loop simulates wall 1105, while a second loop simulates wall 1107. The HRTF pair for the third output signal may correspond to the orientation of wall 1105 relative to the listener, and similarly, for the fourth output signal, the HRTF pair may correspond to the orientation of wall 1107.

[0261] For example, in Figure 18, different HRTF pairs can be applied to the two signals (corresponding to opposing walls), and a binaural mixer can mix all three left ear signals and all three right ear signals into a single binaural output.

[0262] When a soft decision is made as to whether to generate a flutter echo audio signal, the rendering of the flutter echo can advantageously be adapted to the soft decision. For example, if the soft decision results in a flutter echo estimate that includes (or results in) a confidence value α between 0 and 1, the rendering can thereby be controlled between no flutter echo effect at a confidence value of 0 and full flutter echo effect at a confidence value of 1.

[0263] In a particularly simple implementation, the extraction matrix elements associated with the flutter echoes are multiplied by a confidence value. As a result, if the confidence is low, the level of the flutter echo is low. The confidence value can also be modified, for example to achieve a non-linear behavior with respect to the confidence. For example,

number

[0264] Similarly, the confidence values ​​can be used to modify the corresponding elements in the feedback matrix. This has the effect that flutter echoes die out more quickly, since additional damping is applied at each iteration. The confidence values ​​can also be modified, for example, to achieve a non-linear behavior with respect to the confidence. For example:

number

[0265] In another embodiment, the parametric reverberator can be cross-faded between the diffusion and flutter echo scheme described above and a normal diffusion reverberator. A simple implementation of this could be to cross-fade the feedback matrices for the two schemes controlled by a confidence value.

number

[0266] The effect is some bleeding from the diffuse reverb generation into the flutter echo generation and vice versa, which makes the flutter echo more diffuse as the confidence value decreases.

[0267] Other such embodiments may additionally cross-fade other aspects of the feedback loop, which may only affect the flutter loop: delays may be altered and / or the target spectrum of the loop filter may be cross-faded.

[0268] Note that in a room, multiple flutter echo instances can occur with different reflectivities. In some cases, there may be multiple dimensions in which strong reflections exist. Rooms with unusual shapes may have staggered surfaces in the flutter direction.

[0269] In such cases, the additional flutter echo cases can be processed using additional feedback loops as described above. In this way, the described approach can be replicated for the generation of multiple flutter echo audio signals. If too many feedback loops are required for the simulation of flutter echoes, it may be beneficial to increase the number of feedback loops in the feedback delay network configuration. Typically, if the number of loops for reverberation processing is less than eight, quality may be compromised.

[0270] It will be appreciated that the above description has, for clarity, described embodiments of the invention in terms of different functional circuits, units and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units or processors can be used without departing from the invention. For example, functionality shown to be performed by separate processors or controllers may be performed by the same processor or controller. Thus, references to specific functional units or circuits should be seen only as references to suitable means for providing the described functionality, rather than indicative of a strict logical or physical organization or organisation.

[0271] The invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, the functionality may be implemented in a single unit, in several units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.

[0272] Although the present invention has been described in connection with certain embodiments, it is not intended that the present invention be limited to the specific form described herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, while features may appear to be described in connection with certain embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0273] Moreover, although individually listed, a plurality of means, elements, circuits or method steps may be implemented by, for example, a single circuit, unit or processor. Moreover, although individual features may be included in different claims, they may be advantageously combined, and their inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. Moreover, the inclusion of a feature in one class of claims does not imply a limitation to this class, but rather indicates that the feature may be applied to other claim classes as well, as appropriate. Moreover, the order of features in the claims does not imply any particular order in which the features must be performed, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. Moreover, a reference to the singular does not exclude a plurality. Thus, reference to the singular, "first", "second", etc. does not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and are not to be construed as limiting the scope of the claims in any way.

Claims

1. An audio device having a receiver, an estimator circuit, a signal generator circuit, a feedback delay network, and an adapter circuit, the receiver receives room metadata; the room metadata indicates a plurality of characteristics of the room; the estimator circuit determines the room flutter echo estimate in response to the room metadata; the flutter echo estimate indicates a level of flutter echo in the room; the signal generator circuit includes a feedback delay network; the feedback delay network comprises a plurality of feedback loops; the signal generator circuit generates a flutter echo audio signal from the first output signal; the first output signal is from a first portion of the plurality of feedback loops when an audio source signal is provided to the plurality of feedback loops; The adaptor circuit varies a first parameter for a first feedback loop in the first portion of the plurality of feedback loops in response to the flutter echo estimate. Audio equipment.

2. the room metadata includes dimensional data about the room; 2. The audio device of claim 1, wherein the flutter echo estimate is determined in response to a room dimension in a first direction relative to a room dimension in a second direction.

3. the room metadata includes acoustic reflection data for sides of the room; 2. The audio device of claim 1, wherein the flutter echo estimate is determined in response to an acoustic return loss of a first boundary of the room relative to an acoustic return loss of a second boundary of the room.

4. the adaptor circuit increases a feedback coefficient; the feedback coefficient is for the first feedback loop; 2. The audio device of claim 1, wherein the feedback coefficient is based on the flutter echo estimate.

5. The method of claim 1, wherein at least a second portion of the plurality of feedback loops each has a second feedback coefficient; 2. The audio device of claim 1, wherein a third portion of the second portion depends on a room dimension of the room.

6. the signal generator circuit generates a diffuse reverberation signal from an output of a fourth portion of the plurality of feedback loops; the first portion of the plurality of feedback loops does not include any part of the fourth portion of the plurality of feedback loops; the adaptor circuitry varies a fifth portion of the plurality of feedback loops in response to the flutter echo estimate; 2. The audio device of claim 1, wherein the first portion of the plurality of feedback loops includes the fifth portion of the plurality of feedback loops.

7. the signal generator circuit has a delay relative to the audio source signal; the delay is applied to the audio source signal before the first feedback loop; the first portion of the plurality of feedback loops includes the first feedback loop; 2. The audio device of claim 1, wherein the adapter circuitry varies the delay in response to a position of at least one of an audio source of the audio source signal, a listener, and a boundary of the room.

8. the first portion of the plurality of feedback loops comprises at least two feedback loops; the signal generator circuit has a delay relative to the audio source signal; the delay is applied to the audio source signal before the at least two feedback loops; 2. The audio device of claim 1, wherein the delay is different for each of the at least two feedback loops.

9. An audio device as described in claim 1, wherein the first portion of the plurality of feedback loops has one loop or two loops.

10. The plurality of feedback loops include the first portion and the sixth portion, the first portion of the plurality of feedback loops does not include any part of the sixth portion of the plurality of feedback loops; 2. The audio device of claim 1, wherein the adapter circuit changes a feedback coefficient for at least one of the plurality of feedback loops such that there is no feedback from a feedback loop in the first portion of the plurality of feedback loops to any feedback loop in the sixth portion of the plurality of feedback loops.

11. The audio device further comprising a spatial processor circuit and a combiner circuit; the spatial processor circuit applies spatial processing to the flutter echo audio signal; the spatial processing is dependent on the location of at least one of the source of the audio source signal and the boundary of the room; the combiner circuit combines the diffuse reverberation signal and the spatially processed flutter echo audio signal; The audio device of claim 1 , wherein the signal generator circuit generates the diffuse reverberation signal.

12. The audio device further comprises a spatial processor circuit; the spatial processor circuit applies spatial processing to the flutter echo audio signal; The audio device of claim 1 , wherein the spatial processing depends on the location of at least one of a source of the audio source signal and a side of the room.

13. The audio device further comprising a first circuit; the first circuit supplies a plurality of audio source signals to the plurality of feedback loops; 2. The audio device of claim 1, wherein at least one audio source signal is supplied to only the feedback loops of the first portion of the plurality of feedback loops.

14. the signal generator circuit has a gain circuit, the gain circuit being applied to the audio source signal before the first feedback loop; the first portion of the plurality of feedback loops includes the first feedback loop; 2. The audio device of claim 1, wherein the adapter circuit varies the gain circuit in response to at least one of a position of an audio source of the audio source signal, a position of a listener, a position of the room boundary, and a reflection order of the beginning of the flutter echo audio signal.

15. the flutter echo audio signal represents flutter echoes between a pair of opposing boundaries of the room; the signal generator circuit having a frequency dependent gain circuit; the frequency dependent gain circuit is applied to the audio source signal before a first feedback loop; the first portion of the plurality of feedback loops includes the first feedback loop; the adaptor circuitry varies the frequency dependent gain circuitry in response to acoustic reflection data of the room metadata relating to room boundaries; the acoustic reflection data indicates frequency-dependent acoustic characteristics for at least one room boundary; The audio device of claim 1 , wherein a pair of opposing room boundaries does not include the at least one room boundary.

16. The method of claim 1, wherein the first portion of the plurality of feedback loops has at least two feedback loops; 2. The audio device of claim 1, wherein each of the at least two feedback loops has a different loop gain.

17. A method for providing a system for receiving room metadata, the room metadata indicating a plurality of characteristics of the room; determining a flutter echo estimate of the room in response to the room metadata, the flutter echo estimate indicating a level of flutter echo in the room; generating flutter echo audio signals from output signals of a first portion of a plurality of feedback loops when the plurality of feedback loops are supplied with an audio source signal, wherein a feedback delay network comprises the first portion of the plurality of feedback loops; adapting a first parameter for a first feedback loop of the first portion of the plurality of feedback loops in response to the flutter echo estimate; A method comprising:

18. A computer program stored on a non-transitory medium which, when executed on a processor, performs the method of claim 17.

19. The audio device of claim 4, wherein the feedback coefficient is increased when the flutter echo increases.

20. The room metadata includes dimensional data about the room, 18. The method of claim 17, wherein the flutter echo estimate is determined in response to a room dimension in a first direction relative to a room dimension in a second direction.

21. The room metadata includes acoustic reflection data relating to sides of the room, 18. The method of claim 17, wherein the flutter echo estimate is determined in response to the acoustic return loss of a first boundary of the room relative to the acoustic return loss of a second boundary of the room.

22. The adapter circuit increases the feedback coefficient, the feedback coefficient is for the first feedback loop; The method of claim 17 , wherein the feedback coefficient is based on the flutter echo estimate.

23. The method of claim 20, wherein at least a second portion of the plurality of feedback loops each has a second feedback coefficient; The method of claim 17 , wherein a third portion of the second portion depends on a room dimension of the room.