Audio device and method
The audio device generates flutter echo signals using a feedback delay network to improve acoustic environment representation in virtual reality, addressing computational inefficiencies and enhancing user experience.
Patent Information
- Application Number
- JP2023561329
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-08
- Filing Date
- 2022-03-10
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2042-03-10
AI Technical Summary
Existing approaches for generating audio in virtual reality applications often fail to accurately represent the acoustic environment, leading to suboptimal user experiences due to insufficient representation and computational inefficiencies.
An audio device and method that generates a flutter echo audio signal using a feedback delay network with multiple feedback loops, adapting parameters based on room metadata to simulate flutter echo effects, allowing for efficient and accurate rendering of acoustic environments.
Enhances user perception of the acoustic environment by providing a more natural-sounding echo effect with reduced computational complexity and improved flexibility, particularly suitable for virtual reality applications.
Smart Images

Figure 0007797526000048 
Figure 0007797526000049 
Figure 0007797526000050
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus and method for generating a flutter echo audio signal, particularly, but not exclusively, for generating a flutter echo audio signal in combination with generating a diffuse reverberation signal. [Background technology]
[0002] In recent years, the variety and breadth of experiences based on audiovisual content has increased significantly with the continuous development and introduction of new services and methods for using and consuming such content. In particular, many spatial and interactive services, applications and experiences have been developed to provide users with more engaging and immersive experiences.
[0003] Examples of such applications are virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications, which are rapidly becoming mainstream, with many solutions targeted at the consumer market. Numerous standards are also under development by a number of standards bodies. These standardization efforts are actively developing standards for various aspects of VR / AR / MR systems, including, for example, streaming, broadcasting, rendering, etc.
[0004] VR applications tend to provide a user experience that corresponds to the user being in another world / environment / scene, whereas AR (including mixed reality MR) applications tend to provide a user experience that corresponds to the user being in the current environment but with additional information or virtual objects or information added. Thus, VR applications tend to provide fully immersive synthetically generated worlds / scenes, whereas AR applications tend to provide partially synthetic worlds / scenes that are overlaid on the real scene in which the user is physically present. However, these terms are often used interchangeably and are highly overlapping. In the following, the term virtual reality / VR will be used to refer to both virtual reality and augmented / mixed reality.
[0005] As an example, an increasingly popular service is one in which a user can actively and dynamically interact with the system to change the parameters of the rendering, providing images and audio in a manner that adapts to movements and changes in the user's position and orientation. A very attractive feature in many applications is the ability to change the viewer's effective viewing position and direction, for example, to enable the viewer to move and "look around" within the scene being presented.
[0006] Such functionality can enable a virtual reality experience to be provided to the user in particular, allowing the user to move around (relatively) freely in the virtual environment and dynamically change their position and their point of view. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known, for example, from gaming applications, such as in the class of first-person shooter games for computers and consoles.
[0007] In addition to visual rendering, most VR / AR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience, such that audio sources are perceived as arriving from positions corresponding to the positions of corresponding objects in the visual scene. In this way, the audio and video scenes are preferably perceived as being consistent, and both provide a complete spatial experience.
[0008] For example, many immersive experiences are provided in which the virtual audio scene is generated through headphone playback using binaural audio rendering techniques. In many scenarios, such headphone playback can be based on head tracking so that rendering can be made responsive to the user's head movements, which greatly increases the sense of immersion.
[0009] An important feature for many applications is how to generate and / or deliver audio that provides a natural and realistic perception of the audio environment. For example, when generating audio for virtual reality applications, it is important not only that the desired audio sources are generated, but that these audio sources are modified to provide a realistic perception of the audio environment, including attenuation, reflections, coloration, etc.
[0010] In the case of room acoustics, or more generally environmental acoustics, reflections of sound waves from the walls, floor, ceiling, objects, etc. of the environment cause delayed and attenuated (usually frequency-dependent) versions of the source signal to reach the listener (i.e., the user of a VR / AR system) via different routes. This combined effect can be modeled by an impulse response, hereafter referred to as Room Impulse Response (RIR) (although this term implies a specific use for acoustic environments in the form of a room, the term tends to be used more generally in relation to acoustic environments, whether this applies to a room or not).
[0011] As shown in Figure 1, a room impulse response typically consists of a direct sound, which depends on the distance from the sound source to the listener, followed by a reverberant portion that characterizes the acoustic properties of the room. The size and shape of the room, the position of the sound source and listener within the room, and the reflective properties of the room's surfaces all affect the characteristics of this reverberant portion.
[0012] The reverberant part can be decomposed into two time domains that usually overlap. The first domain contains the so-called early reflections, which represent isolated reflections of the sound source off walls or obstacles in the room before reaching the listener. As the time delay increases, the number of reflections that exist within a certain time interval increases, and their paths may include second- or higher-order reflections (e.g., reflections may be from multiple walls, or both walls and ceiling, etc.).
[0013] The second region in the reverberant section is where the density of these reflections increases to the point where they can no longer be separated by the human brain. This region is usually called the diffuse reverberation, late reverberation, or reverberation tail.
[0014] The reverberant part contains clues that give the auditory system information about the distance of a sound source and the size and acoustic properties of the room. The energy in the reverberant part relative to the energy in the anechoic part largely determines the perceived distance of a sound source. The level and delay of early reflections can provide clues about how close a sound source is to a wall, and anthropometric filtering can enhance the assessment of a particular wall, floor, or ceiling.
[0015] The density of (early) reflections affects the perceived size of a room. The time it takes for the energy level of the reflections to drop to 60 dB (reverberation time T 60 Reverberation time, denoted by ( ), is often used as a measure of how quickly reflections dissipate in a room. Reverberation time provides information about the acoustic properties of a room, especially whether the walls are highly reflective (such as a bathroom) or highly sound absorbing (such as a bedroom with furniture, carpets, and curtains).
[0016] Furthermore, the RIR, when part of the binaural room impulse response (BRIR), may depend on the anthropometric characteristics of the user since the RIR is filtered by the head, ears and shoulders (i.e., it is a head-related impulse response (HRIR)).
[0017] Because reflections in late reverberation cannot be distinguished and separated by a listener, they are often simulated and represented parametrically using parametric reverberators, for example using feedback delay networks, as in the well-known Jot reverberator.
[0018] For early reflections, the incidence direction and distance-dependent delay are important cues that allow humans to extract information about the relative position of the room and the sound source. Therefore, the simulation of early reflections must be more explicit than late reverberation. Therefore, in efficient acoustic rendering algorithms, early reflections are simulated differently from late reverberation. A well-known method for early reflections is to mirror the sound source at each boundary of the room and generate virtual sources that represent the reflections.
[0019] For early reflections, the position of the user and / or sound source relative to the room boundaries (walls, ceiling, floor) is relevant, whereas for late reverberation, the acoustic response of the room is diffuse and therefore tends to be more uniform across the effect, which often makes simulating late reverberation more computationally efficient than early reflections.
[0020] Two main characteristics of the late reverberation defined by a room are the T60 value and the reverberation level. For a diffuse reverberation impulse response, these values represent the slope and amplitude of the impulse response. Both are usually strongly frequency dependent in natural rooms.
[0021] The T60 parameter is important to give an impression of the reflectivity and size of a room, while the reverberation level shows the combined effect of multiple reflections at the room boundaries. The reverberation level and its frequency behavior depend on the pre-delay, which indicates where the distinction between early and late reflections is made (see Figure 2).
[0022] The reverberation level has a primary psychoacoustic relevance with respect to the direct sound. The level difference between the two is an indication of the distance between the sound source and the user (or RIR measurement point). As the distance increases, the direct sound attenuates more, while the level of late reverberation remains the same (it is the same throughout the room). Similarly, for sound sources with directionality that depends on where the user is relative to the source, the directionality affects the direct response as the user moves around the source, but not the level of reverberation.
[0023] To render a realistic audio experience and provide a perception of the audio environment, in particular the acoustic properties of a virtual room in which the listener is considered to be located, one or more audio signals and objects can be rendered through a rendering process that reflects the room impulse response, which typically involves generating the direct path, early reflection, and diffuse late reverberation components separately and then combining them in the rendered output.
[0024] Typically, different approaches are used to generate these different components: direct sound and early reflections are often generated by simple filtering (e.g., using binaural processing and head-related transfer function filters), whereas diffuse late reverberation is often generated using a parametric reverberator such as a Jot reverberator.
[0025] Such approaches can produce advantageous, natural-sounding audio in many situations and applications. However, known approaches may not be optimal in some situations and for some applications. For example, many embodiments may result in rendered audio that is not a perfect representation of the intended room acoustics. In many situations, generating a more accurate acoustic environment may require additional complexity and / or computational resources. Current approaches and proposals for how to represent and generate audio representative of an acoustic environment may tend to be suboptimal and / or insufficient and / or incomplete. This may be particularly true in virtual reality applications, for example, where the rendered acoustic environment can have a significant impact on immersion and the overall user experience. Summary of the Invention [Problem to be solved by the invention]
[0026] Therefore, improved approaches would be advantageous, particularly approaches that allow for improved operation, increased flexibility, reduced complexity, easier implementation, improved audio experience, improved audio quality, reduced computational load, improved suitability and / or performance for virtual / mixed / augmented reality applications, improved perceptual cues, improved representation and rendering of different acoustic environments, and / or improved performance and / or operation.
[0027] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]
[0028] According to one aspect of the present invention, there is provided an audio device for generating a flutter echo audio signal, the audio device comprising: a receiver configured to receive room metadata indicative of room characteristics; an estimator configured to determine a flutter echo estimate for the room in response to the room metadata, the flutter echo estimate being indicative of a level of flutter echo in the room; a signal generator including a feedback delay network with a plurality of feedback loops, the signal generator configured to generate the flutter echo audio signal from output signals of a group (set) of feedback loops of a plurality of feedback loops to which an audio source signal is supplied; and an adapter configured to adapt a first parameter for a first feedback loop in the group of feedback loops in response to the flutter echo estimate.
[0029] The present invention can provide an improved user experience in many embodiments and in many scenarios, and in particular can provide an improved user perception of the acoustic environment. The approach can further enable efficient communication of data that enables such improvements, and in particular can be based on environmental data (especially room data) that can be communicated for other purposes without requiring additional data in many scenarios.
[0030] In particular, the inventors have recognized that existing approaches do not accurately reflect all acoustic phenomena, and that significant improvements can be achieved by generating and rendering a flutter echo audio signal that can provide the perception of flutter echo effects in an acoustic environment. Furthermore, generating such a flutter echo audio signal using a signal generator with a feedback delay network having multiple feedback loops can provide a very efficient implementation while still allowing accurate rendering of the flutter echo effect in many embodiments. Furthermore, this allows for commonality with functionality for generating diffuse reverberation, allowing, for example, a highly efficient combined reverberator function to be provided that can dynamically adapt resources allocated to different types of reverberation and echo.
[0031] The approach can provide adaptations that allow a more natural-sounding echo to be perceived. In many embodiments, the approach can allow a flutter echo effect to be created without requiring dedicated data to be transmitted to control the echo. In particular, the device can determine whether to generate a flutter echo depending on room metadata, or can adapt parameters of the flutter echo (such as delay, frequency response, and / or level) to provide a signal that more accurately reflects the natural acoustic environment.
[0032] The flutter echo estimate may indicate the level / degree / amount / prevalence of flutter echo in a room, in particular the level / degree / amount / prevalence of flutter echo relative to diffuse reverberation in a room. The flutter echo may be a flutter echo between two opposite walls / boundaries / sides of a room, in particular between two parallel walls / boundaries / sides of a room.
[0033] The feedback delay network may include a network configured to couple at least an audio source signal to a feedback loop in at least one group of feedback loops, and an output circuit configured to generate a flutter echo audio signal by combining output signals of the feedback loops in the at least one group of feedback loops.
[0034] The group of feedback loops may include one or more feedback loops.
[0035] The first parameter may be, for example, a feedback coefficient for the first feedback loop, a transfer function parameter for the first feedback loop, a frequency dependency of the first feedback loop, a loop gain of the first feedback loop, a delay of the feedback loop, a weight / gain / level for the output signal of the feedback loop and / or a flutter echo audio signal.
[0036] In some embodiments, the adapter may be configured to vary the number of feedback loops in the group of feedback loops in response to the flutter echo estimate.
[0037] In some embodiments, the signal generator is further configured to generate a diffuse reverberation signal from an output of a feedback loop that is not included in the group of feedback loops. A diffuse reverberation signal may be generated when an audio source signal and / or another audio source signal is fed into a feedback loop that is not in the group of feedback loops.
[0038] In many embodiments, the apparatus may be configured to determine a flutter echo estimate in response to a room impulse response.
[0039] The receiver may be configured to receive an audio source signal.
[0040] According to an optional feature of the invention, the room metadata includes dimensional data relating to the room, and the flutter echo estimate is determined in response to a room dimension in a first direction relative to a room dimension in a second direction.
[0041] This configuration can provide particularly advantageous operation and improved adaptive flutter echo simulation in many embodiments.
[0042] The dimensional data may provide an indication of the distance between one or more opposing walls / sides / boundaries of a room.
[0043] The flutter echo estimate may indicate that the level of flutter echo increases as the difference between a room dimension in a first direction and a room dimension in a second direction increases.
[0044] According to an optional feature of the invention, the room metadata includes acoustic reflection data relating to sides of the room, and the flutter echo estimate is determined in response to the acoustic reflection attenuation of a first boundary of the room relative to the acoustic reflection attenuation of a second boundary of the room.
[0045] This configuration can provide particularly advantageous operation and improved adaptive flutter echo simulation in many embodiments.
[0046] The first and second boundaries may be walls or sides of a room.
[0047] According to an optional feature of the invention, the adapter is configured to increase a feedback coefficient from the first feedback loop to itself in response to the flutter echo estimate indicating an increase in the level of flutter echo.
[0048] This configuration can provide particularly advantageous operation, resulting in an improved user experience and a more natural perception of the acoustic environment.
[0049] The adapter may be configured to decrease a feedback coefficient from a first feedback loop to a second feedback loop of the plurality of feedback loops in response to the flutter echo estimate indicating an increasing level of flutter echo, wherein the second feedback loop may be a feedback loop not included in the group of feedback loops.
[0050] In accordance with an optional feature of the invention, feedback coefficients of at least some of the plurality of feedback loops relative to other feedback loops of the plurality of feedback loops depend on a room dimension of the room.
[0051] In some embodiments, the feedback coefficients of at least some of the feedback loops to other feedback loops of the plurality of feedback loops depend on a room dimension of the room.
[0052] According to an optional feature of the invention, the signal generator is configured to further generate a diffuse reverberation signal from outputs of feedback loops not included in the group of feedback loops, and the adapter is configured to vary the number of feedback loops included in the group of feedback loops in response to a flutter echo estimate.
[0053] This approach may allow for a very efficient audio emulation of a room, and may allow for a low complexity implementation, since, for example, feedback loops can be used for different purposes (diffuse reverberation and feedback loop generation) with the allocation of feedback loops between them being dynamically adapted.
[0054] A diffuse reverberant signal may be generated when an audio source signal and / or another audio source signal is fed into a feedback loop that is not within the group of feedback loops.
[0055] According to an optional feature of the invention, the signal generator has a delay with respect to the audio source signal before being provided to a feedback loop of the group of feedback loops, and the adapter is configured to adapt the delay in response to a position of at least one of an audio source of the audio source signal, a listener, and a room boundary.
[0056] This configuration may provide particularly advantageous operation, resulting in an improved user experience and a more natural perception of the acoustic environment.
[0057] In accordance with an optional feature of the invention, the group of feedback loops includes at least two feedback loops, and the signal generator has a delay with respect to the audio source signal before being provided to the at least two feedback loops, the delay being different for the at least two feedback loops.
[0058] This configuration can provide particularly advantageous operation and / or performance in many embodiments.
[0059] In accordance with an optional feature of the invention, the group of feedback loops comprises two or less loops.
[0060] This configuration can provide particularly advantageous operation and / or performance in many embodiments.
[0061] According to an optional feature of the invention, the adapter is configured to adapt feedback coefficients for a plurality of feedback loops such that there is no feedback from a feedback loop in the group of feedback loops to any feedback loop not included in the group of feedback loops.
[0062] In some embodiments, the adapter is configured to adapt feedback coefficients of a plurality of feedback loops such that there is no feedback from any feedback loop not included in the group of feedback loops to any feedback loop of the group of feedback loops.
[0063] According to an optional feature of the invention, the signal generator is configured to further generate a diffuse reverberation signal, and the audio device further comprises: a spatial processor for applying spatial processing to the flutter echo audio signal, the spatial processing depending on a position of at least one of a source of the audio source signal and a room boundary; and a combiner for combining the diffuse reverberation signal and the spatially processed flutter echo audio signal.
[0064] According to an optional feature of the invention, the audio device further comprises: a spatial processor that applies spatial processing to the flutter echo audio signal, the spatial processing depending on the location of at least one of a source of the audio source signal and a side of the room.
[0065] In accordance with an optional feature of the invention, the audio device further comprises circuitry for supplying a plurality of audio source signals to a plurality of feedback loops, at least one audio source signal being supplied to only a feedback loop of the group of feedback loops.
[0066] According to an optional feature of the invention, the signal generator has a gain for the audio source signal before being provided to a feedback loop of the group of feedback loops, and the adapter is configured to adapt the gain in response to at least one of a position of an audio source of the audio source signal, a position of a listener, a position of the room boundary, and a reflection order of an onset of a flutter echo audio signal.
[0067] According to an optional feature of the invention, the flutter echo audio signal represents flutter echo between a pair of opposing room boundaries, the signal generator has a frequency-dependent gain for the audio source signal before being supplied to a feedback loop of the group of feedback loops, and the adapter is configured to adapt the gain in response to acoustic reflection data of room metadata relating to room boundaries, the acoustic reflection data indicating frequency-dependent acoustic characteristics for at least one room boundary that is not one of the pair of opposing room boundaries.
[0068] In accordance with an optional feature of the invention, the group of feedback loops comprises at least two feedback loops having different loop gains.
[0069] According to another aspect of the present invention, there is provided a method for generating a flutter echo audio signal, the method comprising: receiving room metadata indicative of a characteristic of a room; determining a flutter echo estimate for the room in response to the room metadata, the flutter echo estimate indicative of a level of flutter echo in the room; generating the flutter echo audio signal from output signals of a group of feedback loops to which an audio source signal is supplied, the group of feedback loops comprising a feedback loop of a plurality of feedback loops in a feedback delay network; and adapting a first parameter for a first feedback loop of the group of feedback loops in response to the flutter echo estimate.
[0070] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.
[0071] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which: [Brief explanation of the drawings]
[0072] [Figure 1] FIG. 1 shows an example of a room impulse response. [Figure 2] FIG. 2 shows an example of a room impulse response. [Figure 3] FIG. 3 shows an example of elements of a virtual reality system. [Figure 4] FIG. 4 illustrates an example of an audio device for generating a flutter echo audio signal according to some embodiments of the present invention. [Figure 5] FIG. 5 illustrates an example of a signal generator for generating an audio signal according to some embodiments of the present invention. [Figure 6] Figure 6 shows an example of a Jot reverberator. [Figure 7] FIG. 7 illustrates an example of a flutter echo signal generator according to some embodiments of the present invention. [Figure 8] FIG. 8 shows an example of a room impulse response. [Figure 9] Figure 9 shows an example of a flutter echo. [Figure 10] FIG. 10 shows an example of a room impulse response characteristic. [Figure 11] Figure 11 shows an example of a flutter echo between two opposing walls. [Figure 12] FIG. 12 illustrates an example circuit of a signal generator for generating an audio signal according to some embodiments of the present invention. [Figure 13] FIG. 13 illustrates an example circuit of a signal generator for generating an audio signal according to some embodiments of the present invention. [Figure 14] FIG. 14 shows an example of a flutter echo between two opposing walls. [Figure 15] FIG. 15 shows an example of distance gain as a function of the number of reflections between walls. [Figure 16] FIG. 16 illustrates an example circuit of a signal generator for generating an audio signal according to some embodiments of the present invention. [Figure 17]FIG. 17 illustrates an example circuit of a signal generator for generating an audio signal according to some embodiments of the present invention. [Figure 18] FIG. 18 illustrates an example circuit of a signal generator for generating an audio signal according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0073] Although the following description focuses on audio processing and generation for virtual reality applications, it will be appreciated that the principles and concepts described can be used in many other applications and embodiments.
[0074] Virtual experiences that allow users to move through virtual worlds are becoming increasingly popular, and services are being developed to meet such demand.
[0075] In some systems, VR applications may be provided locally to a user, e.g., by a standalone device that does not use or even access any remote VR data or processing. For example, a device such as a game console may include a memory for storing scene data, an input for receiving / generating user poses, and a processor for generating corresponding images from the scene data.
[0076] In other systems, VR applications can be implemented and executed remotely from the user. For example, a user's local device can detect / receive motion / pose data, which is transmitted to a remote device, which processes the data to generate a user pose. The remote device can then generate a view image and corresponding audio signal appropriate for the user's pose based on scene data describing the scene. The view image and corresponding audio signal are then transmitted to the user's local device for presentation there. For example, the remote device can directly generate a video stream (typically a stereo / 3D video stream) and corresponding audio stream, which are then presented directly by the local device. Thus, in such an example, the local device would not perform any VR processing other than transmitting motion data and presenting the received video data.
[0077] In many systems, functionality may be distributed between local and remote devices. For example, the local device may process received input and sensor data to generate a user pose, which is continuously transmitted to the remote VR device. The remote VR device may then generate a corresponding view image and a corresponding audio signal and transmit them to the local device for presentation. In other systems, the remote VR device does not directly generate a view image and a corresponding audio signal, but may select relevant scene data and transmit it to the local device, in which case the local device may generate the presented view image and corresponding audio signal. For example, the remote VR device may identify the nearest capture point and extract corresponding scene data (e.g., a set of object sources and their position metadata) and transmit it to the local device. In this case, the local device may process the received scene data to generate images and audio signals related to a particular current user pose. A user pose typically corresponds to a head pose, and a reference to a user pose may be equivalently considered to typically correspond to a reference to a head pose.
[0078] In many applications, particularly for broadcast services, a source may transmit or stream scene data in the form of images (including video) and audio representations of a scene that are independent of the user's pose. For example, signals and metadata corresponding to audio sources within a particular virtual room may be transmitted or streamed to multiple clients. In this case, each client may locally synthesize an audio signal corresponding to the current user's pose. Similarly, a source may transmit a schematic description of an audio environment, including a description of the audio sources within the environment and the acoustic characteristics of the environment. In this case, an audio representation may be generated locally and presented to the user, for example, using binaural rendering and processing.
[0079] 3 shows such an example of a VR system in which remote VR client devices 301 communicate with a VR server 303 over a network 305, such as the Internet. The server 303 can be configured to support a large number of client devices 301 simultaneously.
[0080] The VR Server 303 can support the broadcasted experience by, for example, transmitting an image signal containing an image representation in the form of image data that can be used by a client device to locally synthesize a view image corresponding to the appropriate user pose (pose refers to position and / or orientation). Similarly, the VR Server 303 can transmit an audio representation of the scene, allowing the audio to be synthesized locally relative to the user's pose. Specifically, as the user moves around in the virtual environment, the images and audio synthesized and presented to the user are updated to reflect the user's current (virtual) position and orientation in the (virtual) environment.
[0081] Thus, in many applications, such as that of FIG. 3, it will be desirable to generate efficient image and audio representations that model a scene and can be efficiently included in a data signal that can be transmitted or streamed to various devices, allowing these devices to locally synthesize views and audio for poses different from the capture pose.
[0082] In some embodiments, a model representing a scene can be stored locally, for example, and used locally to synthesize appropriate images and audio. For example, an audio model of a room can include an indication of the characteristics of audio sources that can be heard in the room, as well as the acoustic characteristics of the room. The model can then be used to synthesize appropriate audio for a particular location.
[0083] How an audio scene is represented and how this representation is used to generate audio are important issues. Audio rendering, which aims to provide a natural and realistic effect to the listener, usually includes a rendering of the acoustic environment. For many environments, this includes a representation and rendering of the diffuse reverberation present in the environment, such as a room. The rendering and representation of such diffuse reverberation has been found to have a significant impact on the perception of the environment, including whether the audio is perceived to represent a natural and realistic environment. In the following, advantageous approaches for representing an audio scene and for rendering audio based on this representation, and in particular for enhancing diffuse reverberant audio, will be described.
[0084] The approach will be described with reference to the audio device shown in Figure 4. The audio device is configured to generate an audio output signal representative of audio in an acoustic environment. Specifically, the audio device can generate audio representative of the audio perceived by a user moving around in a virtual environment having multiple audio sources and given acoustic characteristics. Each audio source is represented by an audio signal representing the sound from that audio source, and metadata that can describe characteristics of the audio source (e.g., providing an indication of the level of the audio signal and / or the location of the audio source). Further, metadata is provided to characterize the acoustic environment.
[0085] In this example, the audio device is specifically configured to generate an audio signal representative of how an audio source is perceived in the current listening environment. The device is capable of generating direct and early reflection audio signal components, as well as diffuse reverberation audio signal components. In this manner, the audio device can receive one or more audio source signals and process one, some, or all of them to generate a corresponding output signal that includes different components that reflect the behavior of the acoustic environment.
[0086] Furthermore, the device is configured to generate a flutter echo audio signal that depends on room metadata that describes the characteristics of the room. When the acoustic environment is a room, the room is characterized by room metadata, and the audio device can be configured to generate a flutter echo audio signal that emulates flutter echoes that may occur in such a room. The flutter echo audio signal can be an additional audio component that is combined with direct sound, early reflections, and / or diffuse reverberation audio components to provide a more accurate and natural-perceived acoustic environment (although in some embodiments, only the flutter echo audio signal is generated). Furthermore, because flutter echoes are typically very specific to individual rooms and, in fact, tend to be significant or noticeable only for some room types / characteristics, the audio device can specifically provide a flutter echo audio signal when appropriate for a particular room, and the flutter echo audio signal is typically adapted to reflect these particular conditions. In particular, in many embodiments, generation of the flutter echo audio signal can be conditional on room metadata, and the flutter echo audio signal can be generated only if the room metadata meets certain criteria.
[0087] For some room types and characteristics, opposing (especially parallel) boundaries / walls of a room, in addition to helping to generate possible early reflections and diffuse reverberation, can also produce a certain percentage of repeating echoes. Such effects can be perceived as flutter echoes, reflecting sound bouncing between opposing walls, with energy decaying as the order of reflection increases. Flutter echoes can include many frequencies (specifically, for example, all audio frequencies) and are not limited to standing wave frequencies, such as those known from room modes. They tend to be most noticeable for mid- and high-frequency sounds.
[0088] In the case of flutter echoes, the reflected sound essentially returns from the reflecting walls at a constant rate but at a slightly lower level. The rate of echoes depends on the distance (i.e., the time of flight) between the walls that give rise to the echoes. The reduction in level depends on the distance attenuation and the reflective properties of the walls involved. These parameters are usually frequency dependent.
[0089] Flutter echo is an acoustic characteristic that can occur in many rooms where the characteristics of the particular room allow for appropriate reflections, such as in a hallway, a stairwell, or a room with very different material properties on different boundaries. Including an emulation of this acoustic effect can provide a compelling experience and create a greater sense of immersion for the user. Nevertheless, commonly used methods cannot and do not perform such emulation.
[0090] 4 comprises, among other things, a receiver (RX) 401 configured to receive room metadata indicative of room characteristics. A flutter echo audio signal is generated to represent flutter echoes in the room, and the generated output signal may include, among other things, a flutter echo audio signal reflecting particular flutter echo characteristics of the room.
[0091] Specifically, the apparatus generates a flutter echo audio signal using a feedback delay network. Such a feedback delay network can also be used by a parametric reverberator to generate diffuse reverberation, thus reusing the same functionality. Such an approach can provide reduced complexity and / or ease of operation, for example, in some embodiments, allowing for dynamic and flexible allocation of resources between diffuse reverberation and flutter echo simulation depending on the characteristics of a particular room. Compared to existing configurations of feedback delay networks in parametric reverberators, the approach of FIG. 4 can add additional characteristic features of room acoustics to the suite of simulation tools used for audio rendering, thereby providing a more realistic modeling of typical rooms in virtual renderings.
[0092] The apparatus of Figure 4 is configured to generate a flutter echo audio signal, while comprising a receiver 401, which is configured to receive room metadata indicative of room characteristics.
[0093] The room metadata may include data characterizing the dimensions of a room, such as the three-dimensional dimensions of a rectangular room. In some embodiments, only one or two dimensions of the room may be represented by the room metadata. The remaining dimension(s) may be predetermined or assumed, for example; for example, the room metadata may indicate the width and length of the room, and the audio device may assume a standard height. In some embodiments, absolute dimension data may be provided, while other embodiments may instead or in addition use relative dimension data information. In some embodiments, a room outline may be provided, for example, indicating the layout of the room as well as the distances between the sides / boundaries / walls of the room.
[0094] Dimensional data may be provided in different ways in different embodiments, for example, room metadata may include distance in meters, room volume including dimensional ratios, time of flight for each dimension, 2D or 3D data as a mesh, etc.
[0095] In some embodiments, the room metadata may include acoustic reflection data, such as reflection or absorption coefficients for one or more walls of the room, and in many cases for all walls / boundaries of the room.
[0096] Such information may be provided as acoustic absorption coefficients, transmission coefficients, coupling coefficients, and diffusion coefficients for each wall in the room.
[0097] In addition to the room metadata, the receiver 401 can also receive one or more audio source signals representing audio from audio sources within the room to be rendered. In many embodiments, the audio sources may be represented by audio objects, although it will be understood that the specific audio source signals depend on the particular embodiment and may be, for example, channel sources or higher-order Ambisonics (HOA) sources. The audio device is configured to generate an output signal related to one or more of the received audio source signals / objects, typically generating an output signal including all audio sources. In many cases, the output signal will be generated from a subset of all audio sources with positional metadata indicating that they are in the room. The audio device can, among other things, process all received audio source signals to generate an output signal that reflects the acoustic characteristics of the room, including the direct sound path, early reflections, diffuse reverberation, and flutter echoes. This processing may, for example, be applied to each audio source signal sequentially or in parallel. The resulting output signals may be combined to generate a single rendering signal. For example, a binaural stereo signal may be generated by binaurally processing (at least a portion of) the output signals generated for each source and then combining the binaural signals into a single output stereo signal.
[0098] It will be appreciated that the described approach may be applied to audio devices that generate only flutter echo audio signals, and not, for example, any direct, early reflection, and / or diffuse reverberation signal components. However, the following description will focus on embodiments in which the audio device is configured to simulate a range of acoustic effects of a typical acoustic environment.
[0099] The audio device comprises a signal generator 403 configured to generate one or more output signals from one or more (typically all) received audio source signals, in this example the signal generator 403 generating the output signals to reflect the intended acoustic environment.
[0100] 5 shows an example of a signal generator 403. The audio device includes a path renderer 501 for each audio source. Each path renderer 501 is configured to generate a direct path signal component representing a direct path from the audio source to the listener. The direct path signal component is generated based on the position of the listener and the audio source. Specifically, the direct path signal component may be generated by scaling the audio signal in a distance-dependent, potentially frequency-dependent manner relative to the audio source, and by scaling the relative gain for audio sources in specific directions relative to the user (e.g., in the case of non-omnidirectional sources).
[0101] In many embodiments, the renderer 501 can also generate direct path signals based on occlusion or diffraction (virtual) elements between the source position and the user position.
[0102] In many embodiments, the path renderer 501 may also generate additional signal components for individual paths if they include one or more reflections. This may be done, for example, by evaluating reflections off walls, ceilings, etc., as known by those skilled in the art. In this manner, the path renderer 501 may also generate early reflection components. The direct path and reflected path components may be combined into a single output signal for each path renderer, and thus a single signal representing the direct path and early / discrete reflections may be generated for each audio source.
[0103] In some embodiments, the output audio signal for each audio source may be a binaural signal generated, for example, by applying HRTF or HRIR filters based on the relative (angular) positions of the audio source and the listener, and thus each output signal may include both left-ear and left-right ear (sub-) signals.
[0104] The output signals from the path renderers 501 are provided to a combiner 503, which combines the signals from the different path renderers 501 to generate a single combined signal. In many embodiments, a binaural output signal may be generated, and the combiner may perform a combination, such as a weighted combination, of the individual signals from the path renderers 501. That is, all right ear signals from the path renderers 501 may be added together to generate a combined right ear signal, while all left ear signals from the path renderers 501 may be added together to generate a combined left ear signal.
[0105] It will be appreciated that binaural rendering can be replaced by rendering to a speaker configuration (e.g., 2.0, 5.1, 7.1, 9.1.4, 22.2) using a panning algorithm such as VBAP to generate two or more speaker signals. Combiner 503 will, in most such embodiments, combine all the contributions to each speaker signal in the speaker configuration.
[0106] The path renderers and combiners may be implemented in any suitable manner, typically comprising executable code for processing on a suitable computational resource such as a microcontroller, microprocessor, digital signal processor, or central processing unit including supporting circuitry such as memory. It will be appreciated that the multiple path renderers may be implemented as parallel functional units, for example a bank of dedicated processing units, or as iterative operations for each audio source. Typically, the same algorithm / code is run for each audio source / signal.
[0107] In addition to the individual path audio components, the audio device is also configured to generate a signal component representing the diffuse reverberation in the environment. The diffuse reverberation signal is (effectively) generated by combining the source signals into a downmix signal and then applying a reverberation algorithm to the downmix signal to generate the diffuse reverberation signal.
[0108] The audio device of FIG. 5 includes a downmixer (DMX) 505 that receives audio signals from multiple sound sources (typically all sound sources in the acoustic environment for which the reverberator is simulating diffuse reverberation) and combines these audio signals into a downmix. The downmix thus reflects all sounds generated in the environment. The downmix is provided to a reverberator (REVB) 507, which is configured to generate a diffuse reverberation signal based on the downmix. The reverberator 507 may be, in particular, a parametric reverberator, such as a Jot reverberator. The reverberator 507 is coupled to a combiner 503, to which the diffuse reverberation signal is provided. The combiner 503 then combines the diffuse reverberation signal with path signals representing the individual paths to generate a combined audio signal representing the combined sound in the environment as perceived by the listener.
[0109] An example of a suitable reverberator is the Jot reverberator shown in Figure 6. This reverberator includes a loop input vector b and a loop extraction matrix C to control how input samples are distributed in the reverberator's feedback loop and how an output signal is generated from the loop.
[0110] The audio device further comprises an echo signal generator (FLTECHGEN) 509 configured to generate a flutter echo audio signal (in many embodiments, multiple flutter echo audio signals may be generated). The echo signal generator 509 receives an input audio source signal and generates one or more flutter echo audio signals, which are fed to a combiner 503 where they are combined with other generated signal components to provide an output signal that reflects the acoustic characteristics of the room being simulated.
[0111] The echo signal generator 509, and therefore the signal generator 403, has a feedback delay network with multiple feedback loops.
[0112] An example of such a feedback delay network of the echo signal generator 509 is shown in FIG. 7 , which illustrates three feedback loops. The feedback delay network can have multiple feedback loops, each (or at least one) of which has an input that receives an input audio signal. Each feedback loop includes a loop transfer function (which may be a delay), a feedback network that returns the output signal of the feedback loop to the input of the loop to be combined with the input audio signal, and an output circuit configured to generate the output signal of the feedback delay network as a combination of the output signals of the feedback loops. For each feedback loop, the feedback network can include a feedback path for the output signal of the feedback loop to the input of the feedback loop and typically also to one or more inputs of the other feedback loops. In many embodiments, the feedback network can include feedback paths from the output of each feedback loop to each input of all feedback loops. Each feedback path typically includes an attenuation factor (or equivalently, a gain factor), although some embodiments can provide more complex feedback paths (e.g., provide a filtering function), such as by including a frequency-dependent gain. In some embodiments, the loop transfer function may be a filter that achieves both the desired frequency response and gain coefficient, while the feedback path may simply be flat unity-gain feedback (e.g., corresponding to a feedback matrix that represents feedback with coefficients of 1 on the diagonal). In many embodiments, the feedback network may be represented by a feedback matrix with coefficients for each feedback loop pair combination.
[0113] Feedback delay networks are typically based on feedback loops with various delays. Input signals are injected into the loops, and these signals are fed back to the loops with appropriate feedback gains. The output signal is derived by combining the signals in the loops. Thus, the input signal is successively repeated with various delays. Using relatively prime delays and a feedback matrix to mix the signals between the loops can create patterns similar to reverberation in real spaces, and is particularly suitable for generating diffuse reverberation, as in the example of Jot or other parametric reverberators.
[0114] The absolute values of the elements of the feedback matrix are designed to be less than 1 to achieve a stable, decaying impulse response. The coefficients can be set in combination with delays to achieve the desired reverberation time (T60). In many implementations, additional gains or filters are included in the loop. These filters can control the decay instead of the matrix. Using filters has the advantage that the decay response can be different for different frequencies.
[0115] In the audio device, such a feedback delay network can be used to generate a flutter echo audio signal, and in many embodiments, a feedback delay network can be used to generate both the flutter echo audio signal and the diffuse reverberation. In particular, the same feedback delay network can be used for both, with parameter values determined to provide the desired effect. Specifically, if a flutter echo is to be generated, all feedback loops of the feedback delay network can be used to generate the diffuse reverberation component, and parameters can be set accordingly. If a flutter echo audio signal is to be generated, one or more (typically only a few, such as two or three or less) feedback loops can be used to generate the flutter echo audio signal, with the remaining feedback loops being used to generate the diffuse reverberation signal. The reassigned feedback loops are then set with appropriate parameters for generating the flutter echo audio signal. In many embodiments, a total of, for example, 8 to 20 feedback loops can be provided, with three or fewer of these being used to generate the flutter echo audio signal, if appropriate.
[0116] As a specific example, the approach can provide a way to include flutter echo simulation using existing configurations of feedback delay networks within a parametric reverberator that generates diffuse reverberation, thereby adding further characteristic features of room acoustics to the suite of simulation tools and providing a more realistic modeling of typical rooms in virtual renderings.
[0117] In this way, the feedback delay network may be common to the echo signal generator 509 and the reverberator 507 .
[0118] In the example of Figure 7, an input signal is provided to each feedback loop via an input circuit comprising a pre-gain section 701. The inputs of these feedback loops comprise a combiner 703 that combines the input audio source signal with the signal being fed back to the feedback loop. Each loop comprises a loop filter 705 (which may include a delay), the output of which is provided to a feedback network / matrix 707 that provides feedback to the loop inputs. Additionally, an output circuit combines the output signals from each loop into an output signal. The output circuit specifically includes a group of gain sections 709 and a combiner 711 configured to generate the output signal of the feedback delay network as a weighted combination of the output signals from each feedback loop.
[0119] The audio device is configured to adapt the generation of the flutter echo audio signal. In particular, in many embodiments, the audio device can be configured to adapt the degree or level of flutter echo depending on the room characteristics of the simulated room; indeed, in many embodiments, the audio device can adapt whether or not a flutter echo audio signal is generated depending on the room characteristics. In this way, the flutter echo simulation is not simply a static generation of a flutter echo audio signal that provides a flutter echo effect, but rather a dynamically adapted generation of flutter echo that is dependent on the room characteristics; in particular, rather than always producing a flutter echo effect, in many embodiments, such generation may be performed only if it is determined that flutter echo is likely to be significant in a particular room.
[0120] The audio device comprises an estimator (EST) 405, which is configured to determine a flutter echo estimate for the room based on the received room metadata, the flutter echo estimate indicating the level / degree / amount / prevalence (distribution) of the flutter echo in the room.
[0121] The exact approach and algorithm or function for determining the flutter echo estimate may vary between different embodiments and may depend on the exact performance and operation desired for a particular application. In many embodiments, a flutter echo estimate may be generated to indicate an increased level of flutter echo when room metadata indicates that reflections between one pair of opposing boundaries / walls are higher than for other pairs of boundaries / walls. This may be true, for example, when the opposing walls of the pair are farther apart than the opposing walls of other pairs and / or when the combined reflection attenuation of the opposing walls of the pair is lower than for other pairs of walls. In such cases, the echo occurring between the opposing walls of the pair may be significantly stronger than other reflection paths occurring between the walls, which may lead to a larger flutter echo (generated by the opposing walls of the pair) relative to other reflections that produce, for example, diffuse reverberation. That is, these flutter echoes may decay more slowly than other reflections that produce, for example, diffuse reverberation. This may result in a more pronounced flutter echo after a certain time, for example, 30 milliseconds, after emission by the source.
[0122] The estimator 405 is coupled to an adaptor (ADP) 407 configured to adapt parameters of at least one of the feedback loops of the feedback delay network in response to the flutter echo estimate. In many embodiments, the parameters may be a feedback coefficient of a loop to itself (which may be frequency dependent), a feedback coefficient from a loop to another loop of the feedback delay network (which may be frequency dependent), a feedback coefficient from another loop to this loop (which may be frequency dependent), a loop gain / weight, a loop delay, a loop transfer function, and / or an extraction coefficient / weight for generating an output signal.
[0123] In many embodiments, a common feedback delay network can be used for generating the diffuse reverberation signal and the flutter echo signal. In such cases, a feedback loop can be dynamically assigned to be used for either the diffuse reverberation signal generation or the flutter echo signal generation by adapting the loop's parameters to suit the diffuse reverberation signal or the flutter echo audio signal. Thus, in many embodiments, the adapter 407 can be configured to switch, for at least one feedback loop, between parameter values for generating the diffuse reverberation signal and parameters for generating the flutter echo audio signal in response to a flutter echo estimate.
[0124] The audio device can therefore be configured to determine the degree of flutter echo believed to be present in the room and set the feedback loop of the feedback delay network to generate a flutter echo audio signal corresponding to this flutter echo.
[0125] This approach can provide improved acoustic simulation in many embodiments, and can provide more natural-sounding audio, particularly when simulating rooms with particular characteristics, allowing certain flutter echoes to be more pronounced without sacrificing performance for rooms where the flutter echoes are not loud or even noticeable.
[0126] The main driving factor that defines the reverberation response is the travel distance of the sound waves. The travel distance causes attenuation and delay. However, each reflection off a surface causes additional attenuation without adding any delay. Therefore, repeated reflections in small room sizes will decay faster than for larger room sizes. Flutter echoes will decay faster in short room sizes than in large room sizes.
[0127] The decay rate of flutter echoes is often consistent with the room's reverberation time, T60, because the various dimensions of the room are roughly similar. This means that flutter echoes are mixed with other reflections that take different paths across multiple dimensions. These result in less regular reflection behavior. Due to similar decay characteristics, flutter echoes are not particularly noticeable in many situations and are not considered in typical current approaches.
[0128] However, if one room dimension deviates significantly from the other dimensions, there will be a flutter echo in this dimension that deviates significantly from most of the room's reflectivity. This will decay slower than other reflection paths because it will have fewer reflective interactions with the room boundaries. This makes the flutter echo stand out from the rest of the reverberation. Fewer reflections result in less decay over time, and the flutter echo will be correspondingly more audible. An example of a room impulse response showing a flutter echo is shown in Figure 8 (the example is for a hallway with dimensions 40 x 2 x 2.5 m).
[0129] Similarly, flutter echoes can be noticeable as reverberant responses when two parallel walls in a room are significantly more reflective than the other walls. This causes flutter echoes in this dimension to decay more slowly, as each interaction with the wall is less destructive than in flutter echoes in other dimensions and in reflection paths that span multiple dimensions.
[0130] As mentioned above, flutter echoes can result from the repeated bouncing of sound waves between two parallel surfaces. Such echoes tend to be present in all rooms, but may be more noticeable in some rooms depending on their shape or the relative material properties of their boundaries.
[0131] In this example, the estimator 405 can generate a flutter echo audio signal to reflect differences in room dimensions. The room metadata can include room dimension data, and the adapter 407 can determine a flutter echo estimate based on the room dimension in a first direction relative to the room dimension in a second direction. For example, the horizontal dimension between two pairs of parallel walls in a rectangular room can be determined from the room size information indicated by the room metadata. The ratio of the longest dimension to the shortest dimension (or the second longest dimension) can then be determined and used as an indicator of how strong the flutter echo is. That is, the ratio can be used directly as a flutter echo estimate.
[0132] The adapter 407 may then, for example, compare the flutter echo estimate in the form of a ratio to a threshold and configure some of the feedback loops of the feedback delay network to generate a flutter echo audio signal if the threshold is exceeded, while configuring the loops to instead contribute to the generation of diffuse reverberation (and therefore not generate a flutter echo audio signal) if the ratio is below the threshold. In other embodiments, a more step-wise approach is used, for example by permanently using one or more feedback loops to generate a flutter echo audio signal, but having it have an amplitude that is a monotonically increasing function of the ratio / flutter echo estimate.
[0133] Alternatively or additionally, in some embodiments, the adapter 407 can determine the flutter echo audio signal in response to changes in the acoustic return loss of the room's sides / boundaries / walls. The room metadata can include the acoustic return loss of the room's walls, and the flutter echo estimate can be generated to reflect these changes. Specifically, the flutter echo estimate can be generated as a function of the difference between the combined acoustic return loss of one pair of opposing sides of the room and the combined acoustic return loss of another pair of opposing sides of the room. For example, a ratio between such combined acoustic return losses can be determined, and the flutter echo estimate can be generated directly as this ratio. The greater the difference, the greater the flutter echo estimate. As described for the dimensional example, the adapter 407 can adapt its operation based on this ratio.
[0134] It will be appreciated that in many embodiments the flutter echo estimate may be generated as a combination of different considerations, and in particular in many embodiments both the dimensions of the room and the acoustic return loss of the walls / sides of the room may be taken into account when generating the flutter echo estimate.
[0135] As mentioned above, one possible cause of noticeable flutter echoes is a room in which one deviant dimension is significantly longer than the other, such as a hallway. In such a case, the echoes of two opposing walls in the deviant dimension will have longer path lengths that result in noticeable flutter echoes from the rest of the room impulse response (RIR). However, the reflection path that is completely perpendicular to the wall in question will involve additional reflections at other boundaries in the shorter dimension, but the lateral extent can be captured by the relatively small reflection path.
[0136] As a result, the path length of a significant portion of early reflections is dominated by the distance in the dimension they diverge. This effect is stronger the higher the reflection order. If the mapped source is spread in one dimension by, say, about 40 meters, a spread in the other dimension by, say, about 4 meters, does not increase the distance significantly. Thus, multiple reflections of different orders will be grouped close to each other in the RIR with slightly different delays.
[0137] This means that flutter echoes are not caused purely by sound waves bouncing back and forth between two parallel surfaces. The effect is simply to produce the first, strongest reflection in a series of reflections. Many subsequent reflections may follow, representing one or more shallow additional reflections on one of the longer boundaries. These produce clearly visible, repeated bursts of concentrated energy in the RIR. This can result in flutter echoes, where each echo essentially contains a series of compound reflections, rather than just a single echo reflection.
[0138] As the order of the main flutter echo increases, other reflections of similar distances become more tightly compressed in time. That is, the path length of a single or double bounce off a long room boundary will be closer to the path length without that long boundary reflection than for lower orders. An example of such a compound flutter echo is shown in Figure 9 (which also illustrates the possible time compression).
[0139] Operating in the digital domain, this means that at any given time, multiple reflections contribute to the same (discrete) filter delay: these contributions add up, causing the impulse response amplitude of these bursts to be larger than it would be with an infinite sample rate.
[0140] In a particular example, the audio device implements an approach that adds flutter echo simulation by utilizing the existing framework of a parametric reverberator, thus leaving the overall complexity of the audio device substantially unchanged.
[0141] The audio device: · Room dimensions; the location and orientation of room boundaries; and / or · Material properties related to room boundaries; It can be based on.
[0142] Based on this metadata, the estimator 405 can first determine whether flutter echo is a likely audible acoustic feature of the room the user is in. For example, this may be the case if one dimension is significantly larger than the other two, or if the reflective properties of material on a wall in one dimension are significantly larger than in the others. A flutter echo estimate reflecting this can be generated.
[0143] The adapter 407 can adapt the operation of the signal generator 403 in response to the flutter echo estimate: if this indicates that the flutter echo is significant, the configuration parameters of the feedback delay network of the parametric reverberator are changed so that one or more of its feedback loops models the flutter echo.
[0144] The adapter 407 can then set the loop delay to be proportional to the dimensions of the room in which the flutter echo occurs, the loop filter can be set to correspond to the (combined) material properties of the walls involved in the flutter echo, and the feedback matrix can be adapted to isolate the loop from the rest of the normal feedback loop. In this way, multiple parameters of the feedback loop can be set to emulate the flutter echo.
[0145] Thus, in some embodiments, a flutter echo estimate can be generated and evaluated to determine whether or not to simulate a flutter echo. This is only necessary if the flutter echo is audible. Typically, there are two possible main root causes for an audible flutter echo: One room's dimensions are significantly larger than the other two; and The reflective properties of the material on one wall dimension are significantly stronger than in the other two.
[0146] Combinations of the above can also cause flutter, for example when two dimensions are significantly larger than the third, but one of these is much less reflective than the other.
[0147] A room dimension may be considered significantly larger than the other two if, for example, it is twice as large as the largest of the other two dimensions. An alternative criterion may be if one room dimension is at least 3.1 times longer than the average of the other two dimensions. In some embodiments, this may be if one room dimension is at least 50% longer than the average of all three room dimensions.
[0148] If the room is not a rectangular prism (shoebox), the dimension can be set to the outer limits of the geometric structure in all three dimensions.
[0149] As another example, a room may qualify for flutter echo simulation if the material properties of the room's boundaries in one dimension differ significantly from those in other dimensions. Reflections can be represented by parameters reflecting acoustic reflection attenuation, such as the reflection coefficient or absorption coefficient. For example, the average reflection coefficient (a value between 0 (no reflection) and 1 (complete reflection)) of both walls in one room dimension is at least 0.2 higher than the maximum average reflection coefficient of both walls in two other directions. Similarly, the average reflection coefficient of each pair of walls can be compared to the average of all walls or the average of two other wall pairs. For example, the average reflection coefficient may be at least 20% greater than the overall average. Furthermore, a minimum required reflection coefficient can be introduced. For example, the average reflection coefficient must be at least 0.67.
[0150] In other embodiments, absorption coefficients can be used to reflect acoustic reflection loss, and these coefficients may be required to be smaller in candidate flutter dimensions than in other dimensions, for example, an average absorption coefficient less than 85% of the average absorption coefficient of wall pairs in other dimensions.
[0151] Reflection (or absorption) coefficients are often frequency dependent. They can be averaged over all frequencies or over a subset of frequencies. Furthermore, averaging can be done over wall sections with different material properties.
[0152] In this manner, a flutter echo estimate can be generated to reflect such parameters, and the adapter 407 can decide whether to simulate a flutter echo based on whether the flutter echo estimate meets appropriate criteria.
[0153] Estimating flutter echoes, and in particular determining whether to simulate flutter echoes, may involve considering a combination of room dimensions and material properties. For example, the satisfaction of any of the separate criteria may result in flutter echoes being simulated. Other embodiments may simulate flutter echoes only if the room dimensions are significantly larger and the corresponding average material properties are significantly different. Optionally or alternatively, a minimum reflection coefficient for the candidate flutter dimension may be additionally required.
[0154] In some embodiments, dimensions and material properties are combined into an estimated decay time (e.g., T60). If the estimated one-dimensional decay time of a dimension is at least 30% longer than the maximum one-dimensional decay time in the other two dimensions, a flutter echo can be simulated in that dimension. In other embodiments, the decay time may need to be at least 0.5 seconds longer than in the other dimensions.
[0155] The decay time can be estimated from the dimensions and the average reflection coefficient of the corresponding walls. In the time it takes a sound wave to travel around a room in that dimension, it is attenuated by the distance traveled and by two reflections on the walls. As an example, the estimated T60 decay time is:
number
[0156] This formula determines the attenuation for one round trip path in a room dimension of magnitude D. The reference distance d of the source ref and the average reflection coefficient
number
[0157] In other embodiments, the estimated one-dimensional decay time can be compared to the whole-room decay time, for example, if the one-dimensional T60 is 10% longer than that estimated for the whole room. The whole-room T60 can be estimated with equations such as the Sabine or Norris-Eyring equations.
[0158] The decision whether a flutter echo should be simulated can also be a soft decision: for example, by choosing a low threshold if the flutter echo is unlikely to be audible and a high threshold if the flutter echo is likely to be audible, any case between these thresholds will yield a confidence level between 0 and 1. A weight w=0 corresponds to no audible flutter echo, while w=1 corresponds to full confidence that a flutter echo is audible.
[0159] For example, the one-dimensional decay time in dimension 1
number
number
[0160] In some embodiments, room characteristics may not be directly available, for example, they may be characterized by a room impulse response. In some embodiments, the room metadata may include the RIR, and the estimator 405 may be configured to generate a flutter echo estimate in response to the RIR. In this example, parameters of the feedback delay network may be determined from the flutter echo estimate generated from the RIR. Measuring the impulse response is more amenable to rooms of arbitrary shapes that deviate from the rectangular shoebox model.
[0161] In such an embodiment, the presence of flutter echoes can be determined by a smoothed version of the magnitude squared, IR(e smooth (n)) can be measured using IR(e min By applying minimum tracking to (n), the flutter echo component can be isolated because the noticeable flutter echo decays more slowly than the remaining reverberation reflections, and tracking the minimum approximates the reverberation decay envelope. An example of this is shown in Figure 10.
[0162] Subtracting the two signals isolates the flutter echo component, if present. If the energy of this signal exceeds a certain threshold, it can be determined that a flutter echo is present. This determination is made based on the percentage of reverberation, i.e.:
number
[0163] Another example is the difference between two echograms, e smooth (n)-e min (n) can be used to derive properties related to the delay and decay of the flutter echo, which can be used to construct a feedback delay network.
[0164] In some embodiments, a peak extraction algorithm can be used to extract the local maxima and their timestamps. The decay rates of these echoes can be determined by fitting an exponential decay model to the peaks. The decay rates and timestamps can be used together to determine the parameters of a feedback loop.
[0165] Adapter 407 can be configured to adapt the parameters in different ways in different embodiments depending on the desired performance. The parameters for generating a flutter echo audio signal can be significantly different from the parameters used by the feedback loop when generating a diffuse reverberation.
[0166] The delay in the feedback delay network used to generate the reverberation is usually chosen to be relatively small, allowing the reflection density to increase rapidly: for example, an average of 12 ms is often used, but for high bandwidth signals (e.g., 48 kHz) this is usually much smaller.
[0167] The choice of delay often depends on the reverberation time (T60), which is usually positively correlated with the room dimensions, but the material properties of the room boundaries also have a significant effect on T60. That is, material properties introduce additional attenuation (over and above that caused by distance attenuation) into the RIR without adding latency, and the room dimensions determine the proportion of these attenuations in the RIR. Therefore, the configuration of a parametric reverberator is primarily determined by the overall reverberation characteristic T60 and the need to quickly reach a minimum reflection density (e.g., 1,000 to 10,000 reflections per second) to accurately model the room.
[0168] In contrast, if the feedback loop is configured to generate a flutter echo audio signal, the adapter 407 can select a loop delay corresponding to the room dimensions to simulate the proportion of flutter echo. The loop filter, which normally simulates the overall reverberation slope T60, can instead be selected to correspond to the average material properties of the walls involved in the flutter echo, to simulate the effect of the walls on each reflection.
[0169] The feedback matrix can be adjusted in many embodiments to keep flutter echoes separated from the diffuse reverberation products, so that a consistent reproduction of the flutter echo is simulated. If multiple different flutter echoes are present in a room, multiple feedback loops can be reused in a similar manner.
[0170] In many embodiments, adapter 407 may be configured to increase the feedback coefficient / gain from a first feedback loop to itself in response to the flutter echo estimate indicating an increasing level of flutter echo. If the extent of flutter echo increases, the feedback from a given feedback loop to itself may be increased. Alternatively, or typically in addition, adapter 407 may be configured to decrease the feedback coefficient from a first feedback loop to a second feedback loop of the multiple feedback loops in response to the flutter echo estimate indicating an increasing level of flutter echo. The second feedback loop may not be configured to be used to generate flutter echo, but instead may be used to generate diffuse reverberation.
[0171] In some examples, a feedback loop used to generate flutter echoes may only feed back to itself. In some examples, a feedback loop used to generate flutter echoes will not feed back to any other feedback loops configured to generate flutter echoes. In some examples, a feedback loop used to generate flutter echoes may receive feedback signals only from itself (out of a group of feedback loops used to generate flutter echoes, or possibly out of all feedback loops in a feedback delay network).
[0172] The adaptation may be, for example, gradual, although in other embodiments the adaptation may be, for example, a step function. For example, if the flutter echo estimate indicates that the flutter echo is not significant, the appropriate feedback coefficient may be relatively small since the feedback loop may be used primarily to contribute to diffuse reverberation, and the feedback from a given loop may then be increasingly distributed among different loops to reflect the many different reflections that make up the diffuse echo. However, if the flutter echo estimate indicates that the flutter echo is significant, the feedback coefficient for that loop may be increased while the feedback coefficient for other loops may be decreased to reflect an increased amount of periodic reflections that correspond to typical flutter echoes.
[0173] Some examples of such adaptations can be explained below with reference to the example of Figure 11, which shows an example of flutter echoes as a function of time and space. In this example, flutter echoes originating at a source 1101 are transmitted to a listener 1103 at a constant rate (τ γ = 2D / c, where D is the distance between the walls and c is the speed of sound), but arrives at four different offsets depending on where the user and source are between the walls.
[0174] In some low-complexity embodiments, the audio device may have a single path length (τ r This can be simplified by reusing only a single reverberator loop, with a delay corresponding to (=D / c), which corresponds to the listener and source being halfway between the rooms, in which case both the dotted and solid reflections reach the listener at the same time, with a constant rate.
[0175] The feedback matrix for a reverberator with N loops, with the first loop used for flutter echoes, is:
number
number
number
number
number
[0176] Function G d(x) gives the distance attenuation for a sound wave propagating x meters. This can be simple attenuation based on an omnidirectional source whose energy is spread over a sphere of radius x. It is well known that every doubling of distance (i.e., radius) results in a 6 dB attenuation. In many embodiments, a reference distance can be used as the distance at which the source signal is defined, in which case the distance attenuation is considered to be included in the signal, and G d The additional distance attenuation from (x) is equal to 0 dB.
[0177] Furthermore, the effect of air absorption G abs Other aspects such as (x) can be d can be added to (x). This effect is usually more pronounced at larger distances and tends to be frequency dependent. Usually the effect of air absorption is very small, especially when considered for realistic room dimensions D.
[0178]
number
[0179] The described embodiment uses the average reflection coefficient
number
[0180]
number
[0181] The reflection coefficients may not be averaged, but may be adapted to the lateral position of the source, for example, in front of a wall, i.e., where most of the flutter echoes will occur. In such an embodiment, multiple sources may be grouped into separate loops according to their associated reflection coefficients.
[0182] As another example, the audio device may be configured to emulate the flutter echo of FIG. 11 by doubling the path length (τ r )2. The four loops may have separate inputs that are pre-delayed to reflect the offset between the listener / source and the wall. For example, a pre-delay circuit such as that shown in Figure 12 may be used.
[0183] The feedback matrix in this case is:
number
[0184] The loop filters of the reverberators in the flutter loop will together simulate the average reflection characteristics of the walls, for example:
number
[0185] In this way, each loop filter simulates the attenuation and reflections on both walls that result from a sound wave propagating twice through the medium (eg, air) over the distance between the walls.
[0186] The advantage of this embodiment is that it simulates the asymmetry of the two loops, similar to how it would be in a real room, and adjustment of the pre-delay can be used to adapt the asymmetry to the user's position in the room, without having to update the parameters of the feedback delay network itself.
[0187] For example, if the listener is 30% of the wall distance from wall 1105 and the source is 15% of the wall distance from wall 1107, the pre-delay can be set based on the first four path lengths from the source to the listener as follows:
number
number
[0188] The previous embodiment can optionally be simplified by combining pre-delay signals before feeding the combined signal into a single feedback loop, for example as shown in the example of FIG.
[0189] The feedback matrix is:
number
[0190] And the loop filter would be the same as in the previous embodiment:
number
[0191] The loop simulates the path length attenuation and reflections on the two walls, while the pre-delay configuration is responsible for creating an offset in the signal. The delay will be the same as in the previous example.
[0192] The pre-delay configuration can also be extended to include gains or filters that simulate wall reflections and distance attenuation in these first paths. Such filters can also include additional filtering and / or attenuation to simulate early propagation and reflections of flutter echoes, since the simulation in the feedback loop does not represent the first few reflection orders. However, such effects are typically already incorporated into conventional reverberation pre-mixing and its coloration filters.
[0193] The separate input signals can also be derived from a single tapped delay line. Parametric reverberators, typically used in combination with direct path and early reflection rendering, include a pre-delay for their normal operation to control where reverberation starts relative to the direct path and early reflections. If this pre-delay is long enough, a delay buffer can be used as a tapped delay line. In this case, flutter echoes will start earlier, but this can be compensated for by early reflection modeling.
[0194] As another example, one could use a set of feedback loops with two interacting loops, with a signal alternating between the loops on each iteration. This could be achieved for the first two loops using the following feedback matrix:
number
[0195] The delays in this embodiment can be set for any listener position to create a regular but asymmetric pattern that better matches realistic scenarios. As another example, the delays can be adjusted depending on the user's position between walls. For example, if the user is 30% of the wall distance from wall 1105, the first delay is τ r1 = (2 0.3 D) / c, while the second delay is τ r2 = (2·(1 − 0.3)·D) / c.
[0196] As with the previous embodiment, a pre-delay configuration can be used to create the missing offset due to the signal bouncing in two directions. This can be done with two delays (Delay1 and Delay2 above) corresponding to the first two paths.
[0197] A particular advantage of this approach is that two loop filters can simulate each wall independently, i.e., the first delay τ r1 The first filter related to:
number
[0198] Similarly, the second delay τ r2 The second filter related to:
number
number
[0199] A possibility with such an embodiment is that if the flutter loops are excluded from the regular extraction matrix to generate the diffuse reverberation tails, these can be extracted to separate the outputs for rendering with dedicated HRTF pairs.
[0200] In some embodiments, signal generator 403 has a gain for the audio source signal before it is fed into the feedback loop of the feedback delay network, and adapter 407 is configured to adapt the gain depending on the position of the audio source relative to the audio source signal. This can particularly, but not necessarily, be combined with the pre-delays discussed above, and in particular each delay in the circuits shown in Figures 12 and 13 can include an adaptive gain, which can be adjusted by adapter 407 based on the position of the audio source, the listener and / or the wall.
[0201] The received data may include an audio signal representing an audio source and its position, which may be used to adapt the gain. Specifically, the gain may be adapted based on the position of the audio source relative to the walls / boundaries / sides of the room. Typically, the gain may be adapted based on the distance from the audio source to the wall (typically the nearest wall) that is the flutter echo reflector. The pre-gain may be used to adjust the relative strength / level of the overall flutter echo effect, and in particular to adjust the level to reflect the strength of the signal when first reflected.
[0202] In some embodiments, the pre-gain may be adapted based on the distance of the listener / user, specifically the relative distance from the listener to the source, or the distance from the source to the listener via at least one reflection off a reflecting wall for flutter echoes.
[0203] Furthermore, in many embodiments, the first reflection can be represented by an early reflection simulation, and the flutter echo signal generator 403 can be used only to represent further reflections of the flutter echo. For example, the flutter echo signal generator 403 can be used to generate flutter echo components corresponding to the fourth or subsequent reflections. In such cases, the reflected sound has already been attenuated by previous reflections, including both distance attenuation and reflection attenuation. Such effects can alternatively or additionally be represented by the pre-gain.
[0204] In some embodiments, the adapter 407 may be configured to adapt the gain depending on the distance between two walls / sides / boundaries of the room (particularly, a wall / boundary / side that produces a flutter echo). In some embodiments, the adapter 407 may be configured to adapt the gain in response to the acoustic reflection attenuation of at least one wall / side / boundary of the room (particularly, a wall / boundary / side that produces a flutter echo). In some embodiments, the adapter 407 may be configured to adapt the gain in response to a plurality of early flutter echo reflections that are not emulated by the set of feedback loops of the feedback delay network assigned to the flutter echo simulation.
[0205] Specifically, the (distance) gain component of the loop filter can represent the attenuation for adjustments in the previous loop path (reflection), and the pre-gain can be used to adapt the input signal level, i.e., the level at the start of the reflection being simulated.
[0206] Signals can often be represented at a level corresponding to a particular reference distance. Compensation / pre-gain can be used specifically to match the level of the signal to the distance it has already traveled before injecting it into the loop, i.e., to represent the initial distance gain. For example, a simulation based on a feedback delay network can be configured to represent flutter echoes from their fourth order (since the first three are represented by early reflection modeling by other algorithms). In this particular example, and with reference to Figure 14, the input gain is:
number
number
number
[0207] The feedback loop may have an overall loop gain set to reflect the attenuation of the reflection path (which may include one or more reflections depending on the particular approach). The loop gain can be set by a loop filter and / or feedback coefficients (feedback matrix). In the example described, the feedback coefficient of the loop to itself is set to 1, and the loop gain (less than 1) is determined by the loop filter. Loop gain / attenuation is typically frequency dependent, and the frequency dependence is typically implemented through the use of an appropriate loop filter.
[0208] In different embodiments, the loop gain / attenuation G d Different approaches for determining
[0209] Typically, a loop filter contains two main components: material properties (e.g., reflection coefficients) and distance-related gains. Each loop filter may represent one or more reflection coefficients corresponding to reflections on one or two walls, and a distance gain corresponding to the propagation distance consistent with the reflections represented by the average reflection coefficients.
[0210] Since the loop-related distance relative to the reference distance continues to increase, the required distance decay component must become weaker with each iteration. For example:
number
[0211] This means that successive reflections decay faster than exponentially, which may not be accurately simulated by a single feedback loop. Filters in feedback loops may be constant due to their recursive nature. Any processed sample may contain components of many different iterations.
[0212] If we isolate the energy dispersion component (the most important component), the distance attenuation (with respect to the signal amplitude) is d ref / d. Let every iteration correspond to a progression distance D. With each iteration, the distance d increases by this constant distance D, so that the additional attenuation relative to the previous iteration is:
number
[0213] The problem is that d is increasing with each iteration, and a fixed gain is required with each iteration. If we express the distance in another way, such that d is a multiple of D, we get a similar result that shows further simplification:
number
[0214] It can be seen that the effect of distance attenuation at each iteration does not depend much on the actual propagation distance, but on the increase in distance already propagated (using the rule of thumb that the attenuation is 6 dB for every doubling of distance). Figure 15 shows how the distance gain varies with each iteration.
[0215] The distance gain in the first iteration has a very large effect because the distance corresponding to one iteration is relatively small compared to the total propagated distance. Rapidly, the dynamic effect of the distance decay in each iteration decreases (i.e., there is less change between iterations). As a result, the decay approaches an exponential shape.
[0216] As the distance gain approaches unity, the gain per iteration stabilizes toward the average reflection coefficient of the flutter boundary material properties. When simulating flutter echoes, different embodiments can choose different approaches. For example, the average reflection coefficient can be chosen to simulate the decay in higher orders. Alternatively, a steeper decay can be used to simulate the decay in lower orders. Alternatively, an intermediate value will be beneficial in most implementations, so as not to have an overly steep or overly shallow decay. Accurate simulation of the slope in higher orders may not be necessary in many cases, as they will not be audible to the listener. A good trade-off can be made, for example, by choosing a slope corresponding to the fifth iteration.
[0217] As mentioned above, many embodiments can adjust the input level of the signal injected into the flutter loop. In addition to compensating for the reference distance and reflection orders that are otherwise simulated, the input gain can be adjusted by adjusting the attenuation gain G d It can also be advantageous to adjust for the trade-off selected for : choosing a relatively slow decay may cause flutter echoes to be overly noticeable, while choosing a relatively steep decay may make flutter inaudible in an accurate simulation.
[0218] Therefore, if a relatively slow decay is set, the initial level can be further reduced to prevent it from being too noticeable. The additional decay essentially compensates for the faster decay in early iterations that are not accurately modeled in the recursive process. As a result, stronger initial reflections may not be accurately modeled. In many cases, these would be (largely) masked by reverberation anyway.
[0219] As an example, based on the model in Figure 14, a feedback delay network can be used to simulate second order or higher flutter echoes with a decay slope corresponding to the 10th iteration, resulting in a loop filter:
number
[0220] The initial input gain can be configured to represent the first reflection order:
number
[0221] Compensation for the attenuation lost in the first I=9 iterations is:
number
number
number
[0222] This compensation ensures matching of both slope and level at the 10th iteration, which can be stored for different decay reference iterations J in a lookup table:
number
number
[0223] In some applications, it may be desirable to simulate both higher levels of low-order reflections and lower levels of mid- and high-order reflections. This may be possible in rooms with relatively low diffuse reverberation energy (e.g., highly absorptive boundaries, except for those involving flutter echoes). Such applications may employ embodiments in which two or more loops simulate different decay rates of the same delay.
[0224] If the delays on the input signal and in the flutter feedback loop are equal, reflections will be generated with the same delay. The first flutter loop may be configured with a steep decay and a relatively large input gain, while the second flutter loop may be configured with a slow decay and a relatively small input gain. When the two are combined by an output circuit, the combined effect may more closely resemble an accurate simulation with an iteratively dependent loop gain.
[0225] In some embodiments, the set of feedback loops assigned to generate flutter echoes may therefore have at least two feedback loops with different loop gains, but the at least two feedback loops may have the same delay.
[0226] The above embodiments configure loop filters according to the material properties of the walls where the flutter echoes originate. These filters can be extended to include the effect of shallow reflections at long room boundaries.
[0227] The material properties of the boundary in the flutter dimension (short boundary) have no effect on the energy ratio of the first reflection to the total complex reflection, however, it does affect how quickly successive complex reflections decay.
[0228] Conversely, the material properties of the long boundary (i.e., not along the flutter dimension) determine how quickly each complex reflection decays, and therefore the energy ratio between the first reflection and the total complex reflection. The decay of the first response amplitude in successive complex reflections is not affected by this material.
[0229] As mentioned above, within an RIR, these responses vary with the order of the flutter echo to which they contribute, compressing the individual reflections in time. However, there are essentially additional contributions from one, two, or more additional material properties. The main effect is that this increases the energy of the individual flutter echoes and their coloration. Coloration is affected by adding the contributions from the additional frequency-dependent material properties, but also, in theory, by delayed reflections that create a comb-filtering effect. However, with a large number of repetitions at different delays, the comb-filtering effect is not very significant.
[0230] Multiple reflections can be modeled by a single reflection. Loop filter H τ can be set to represent a single pulse with a spectral response that matches that of the complex reflection.
[0231] The net effect of multiple reflections also constitutes more energy than a single reflection, and this total energy must also be represented by a single reflection. Due to delays, the amplitudes of the individual responses usually do not add coherently.
[0232] The energy of the complex reflection is:
number
[0233]
number
number
[0234] E c The above expression for ignoring distance attenuation, whose contribution is relatively low, also makes the energy ratio between the initial amplitude and the complex reflected energy independent of the flutter order.
[0235] In an alternative embodiment, a separate loop with a very short delay can simulate the tail to the main flutter response. This loop is fed only by the main flutter loop and does not have a direct signal input (b i =0), but feeds back to itself. The short delay may depend on the shortest dimension of the room. The attenuation by the filter may be greater at long room boundaries (e.g.
number
[0236] Another alternative is to use a sparse IIR as the loop filter in the flutter loop to simulate the fast decay response of the complex reflections.
[0237] In many embodiments, the audio device can be configured to provide multiple audio source signals to a feedback delay network, and in particular, to provide multiple audio source signals to a set of feedback loops that generate a flutter echo audio signal. The audio device can receive audio signals, e.g., for multiple audio sources in a room, and provide multiple (and possibly all) of these signals to a set of feedback loops that generate a flutter echo audio signal. The multiple signals can be combined, for example, into a composite signal, which can then be provided to the set of feedback loops. Each signal can be subject to delay and / or gain adjustment before being combined with the other signals. The gain and / or delay for each signal can be adapted, for example, to reflect the initial and / or relative signal level and / or arrival time of the individual signal (relative to the other signals). In some embodiments, for example, the gain and / or delay can be common to some or perhaps all of the source signals provided to the set of feedback loops.
[0238] The above-described embodiments may enable accurate simulation of offsets between individual reflections, which may provide particularly realistic renderings. The described approach focuses on generating flutter echoes for a single source, and the characteristics of the loop parameters may depend on the specific characteristics of the source, such as its location. However, in many cases, there are two or more sources generating flutter echoes in the simulated room. In such cases, each source may be simulated with its own dedicated feedback loop, etc. These may be implemented, for example, with separate parallel paths to pre-mixing and pre-delay sections before the feedback delay network.
[0239] However, in many applications, such a level of accuracy is not required. The parameters can be set to appropriate values (e.g., arbitrarily or artificially selected values). In some embodiments, they can be selected equally for all simulated sources. For example, the approach shown in Figure 16 uses individual gains g n can be used where a delay is applied to the input audio source signals before they are combined, and then one or more delays are applied to the common signal. This reduces computational and structural complexity. In such an approach, the flutter echo audio signal can still be adapted to the user's position in the room.
[0240] The input to the feedback loop of this feedback delay network is mathematically:
number
[0241] Some embodiments require or benefit from separate inputs to the feedback loops, which can be achieved by expanding the input gain vector b into a matrix B that takes into account more than one signal and maps it to the different loops.
[0242] For example, the inputs provided to a feedback delay network with five feedback loops (P=5) can be processed by an input matrix B:
number
[0243] As another example, in line with FIG. 17, an exemplary matrix is:
number
[0244] The delays create different offsets for the P different paths from the source to the listener. Typically, for a shoebox-shaped room, P=4 per flutter dimension. The delays can be chosen to represent relative offsets to the smallest offset, and the common offset is ignored. In other embodiments, all delays can be set to absolute offsets, potentially dynamically adjusting to the listener's position.
[0245] The delay can also be commonly adjusted to provide an additional common delay component for the flutter echoes. Such a common delay component can be useful to control the offset of the flutter echoes simulated by the parametric reverberator relative to early reflections simulated by other means, for example, to ensure appropriate latency between the last early reflection associated with the flutter dimension and the first simulated flutter echo response from the feedback delay network.
[0246] In some cases, it may be advantageous to start the flutter echo earlier than the diffuse late reverberation portion. In these embodiments, the input to the flutter loop may bypass the pre-delay and pass only through a dedicated flutter delay that controls the start of the flutter echo simulation relative to the source radiation. For example, separately generated early reflections may exclude all reflections associated with the flutter dimension and instead simulate these using only a feedback delay network.
[0247] In another embodiment, an early reflection signal may be generated and fed into a flutter echo feedback loop of a feedback delay network, the early reflection signal may include only reflections in the flutter dimension.
[0248] In some embodiments, the audio device may be configured such that at least one audio source signal is only fed into the feedback loop used to generate the flutter echo audio signal.
[0249] In some embodiments, the audio device may comprise a spatial processor configured to apply spatial processing to the flutter echo signal, where the spatial processing depends on the position of the source of the audio source signal and / or the side of the room.
[0250] The spatial processing may be a process that modifies or generates spatial cues for the flutter echo audio signal. In particular, the spatial processor may be configured to perform binaural processing of the flutter echo audio signal, for example, as shown in FIG. 18, where the spatial processor is represented by two HRTF blocks HRTF1 and HRTF2. The spatial processor may use the HRTFs to apply binaural processing to generate a stereo signal that, when rendered by headphones, results in a spatial perception of the flutter echo originating from the appropriate position / direction. For example, the binaural processing may apply HRTF processing based on the position of one of the walls that generates the flutter echo and the listener's position, so that the flutter echo is perceived as coming from the direction of this wall.
[0251] In some embodiments, the spatially processed flutter echo audio signal can be combined with other generated audio components, in particular, it can be combined with diffuse reverberation generated by other feedback loops of the feedback delay network, however, this diffuse reverberation is generally a distributed sound and therefore cannot be subjected to spatial processing.
[0252] Thus, in some embodiments, the audio device comprises a combiner for combining the spatially processed flutter echo audio signal with the (non-spatially processed) diffuse reverberation signal. In the example of Figure 18, the combiner MIX may generate a stereo output signal for a set of headphones by combining the spatially processed flutter echo audio signal with the non-spatially processed diffuse reverberation signal and typically other audio components such as direct and early reflection audio components.
[0253] In many embodiments, the feedback delay network can generate a flutter echo audio signal by combining the output signals of feedback loops used to generate the flutter echo audio signal. Similarly, a diffuse reverberation signal can be generated by combining the output signals of feedback loops used to generate reverberation.
[0254] Typically, the feedback loop of the feedback delay network is used either for the generation of reverberation or for the generation of flutter echoes.
[0255] In most embodiments, the adapter 407 may be configured to assign a group of feedback loops to generating the flutter echo audio signal, while the remaining feedback loops are used to generate reverberation. In such cases, the adapter 407 may be configured to generally keep the loops separate. Specifically, the adapter 407 may adapt the feedback coefficients of the feedback loops so that there is no feedback from one feedback loop to any other feedback loop in the group of feedback loops used to generate the flutter echo audio signal, and vice versa. That is, the adapter may set to zero all feedback coefficients in the feedback matrix that relate to feedback between two loops belonging to two different groups.
[0256] Similarly, when generating an output signal, a flutter echo audio signal may be generated by combining output signals from only those feedback loops in a group of feedback loops that are used to generate the flutter echo audio signal, while a reverberation signal may be generated by combining output signals from only those feedback loops in a group of feedback loops that are not used to generate the flutter echo audio signal.
[0257] In many embodiments, the output signal of the flutter feedback loop can be processed in the same way as other feedback loops, by using a weighted combination to generate the output signal, which can be represented by an extraction matrix C. This can involve, for example, applying correlation and / or coloration filters known from the generation of diffuse reverberation, in which case the resulting flutter echoes will not originate from any particular direction.
[0258] However, in embodiments where it is desired that the flutter echoes be directional, the flutter echo feedback loop output signal can be extracted separately for alternative processing (as in the example of Figure 18). The extraction matrix for the (binaural) diffuse reverberation tail is of dimension 2x(N-2), which can be expanded to become 4xN to process all N feedback loops of the feedback delay network. In the following example where N=4, the first two rows relate to the processing of the further diffuse reverberation tail, and the last two rows relate to the flutter echoes.
number
[0259] The first and second output signals generated by the extraction matrix may be processed as usual by the rest of the parametric reverberator function. The third and fourth output signals may be processed separately, for example with different HRTF pairs corresponding to opposite directions of both walls. These may be adaptive depending on the user's orientation.
[0260] This may be particularly advantageous for embodiments in which each wall is simulated in a separate loop: a first loop simulates wall 1105, while a second loop simulates wall 1107. The HRTF pair for the third output signal may correspond to the direction of wall 1105 relative to the listener, and similarly, for the fourth output signal, the HRTF pair may correspond to the direction of wall 1107.
[0261] For example, in Figure 18, different HRTF pairs can be applied to the two signals (corresponding to opposite walls), and a binaural mixer can mix all three left-ear signals and all three right-ear signals into a single binaural output.
[0262] If a soft decision is made as to whether to generate a flutter echo audio signal, the rendering of the flutter echo may advantageously be adapted to the soft decision. For example, if the soft decision results in a flutter echo estimate that includes (or results in) a confidence value α between 0 and 1, then the rendering may be controlled between no flutter echo effect at a confidence value of 0 and full flutter echo effect at a confidence value of 1.
[0263] In a particularly simple implementation, the extraction matrix element associated with the flutter echo is multiplied by a confidence value. As a result, if the confidence is low, the level of the flutter echo will be low. The confidence value can also be modified, for example, to achieve a non-linear behavior with respect to the confidence. For example,
number
[0264] Similarly, the confidence values can also be used to modify the corresponding elements in the feedback matrix. This has the effect that flutter echoes die out more quickly, since additional damping is applied at each iteration. The confidence values can also be modified, for example, to achieve non-linear behavior with respect to confidence, for example:
number
[0265] In another embodiment, the parametric reverberator can be cross-faded between the diffusion and flutter echo scheme described above and a regular diffusion reverberator. A simple implementation of this could be to cross-fade the feedback matrices for the two schemes controlled by a confidence value.
number
[0266] The effect is a bleed from diffuse reverb generation to flutter echo generation and vice versa, which makes flutter echoes more diffuse as the confidence value decreases.
[0267] Other such embodiments may additionally cross-fade other aspects of the feedback loop, which may only affect the flutter loop: delays may be changed and / or the target spectrum of the loop filter may be cross-faded.
[0268] Note that in a room, multiple flutter echo instances can occur with different reflectivities. In some cases, there may be multiple dimensions in which strong reflections exist. Rooms with unusual shapes may have staggered surfaces in the flutter direction.
[0269] In such cases, additional flutter echo instances can be processed using additional feedback loops as described above. In this manner, the described approach can be replicated for generating multiple flutter echo audio signals. If too many feedback loops are required for flutter echo simulation, it may be beneficial to increase the number of feedback loops in the feedback delay network configuration. Typically, if the number of loops for reverberation processing is less than eight, quality may suffer.
[0270] It will be understood that the above description has, for clarity, described embodiments of the invention in terms of different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units, or processors may be used without departing from the invention. For example, functionality shown to be performed by separate processors or controllers may be performed by the same processor or controller. References to specific functional units or circuits should therefore be seen merely as references to suitable means for providing the described functionality, rather than as indicating a strict logical or physical organization or organisation.
[0271] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally, and logically implemented in any suitable way. Indeed, its functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and processors.
[0272] While the present invention has been described in connection with several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, while features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0273] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual features may be included in different claims, they may be advantageously combined, and their inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one class of claims does not imply limitation to that class, but rather indicates that the feature may be applied to other claim classes as well, as appropriate. Furthermore, the order of features in the claims does not imply any particular order in which the features must be performed, and in particular the order of individual steps in method claims does not imply that the steps must be performed in that order. Rather, the steps may be performed in any suitable order. Furthermore, singular references do not exclude pluralities. Thus, references to the singular, "first," "second," etc. do not exclude pluralities. Reference signs in the claims are provided merely as a clarifying example and are not to be construed as limiting the scope of the claims in any way.
Claims
1. An audio device having a receiver, an estimator circuit, a signal generator circuit, a feedback delay network, and an adapter circuit, the receiver receives room metadata; the room metadata indicates a plurality of characteristics of the room; the estimator circuit determines the room flutter echo estimate in response to the room metadata; the flutter echo estimate indicates a level of flutter echo in the room; the signal generator circuit includes a feedback delay network; the feedback delay network comprises a plurality of feedback loops; the signal generator circuit generates a flutter echo audio signal from the first output signal; the first output signal is from a first portion of the plurality of feedback loops when an audio source signal is provided to the plurality of feedback loops; The adaptor circuit varies a first parameter for a first feedback loop in the first portion of the plurality of feedback loops in response to the flutter echo estimate. Audio equipment.
2. the room metadata includes dimensional data about the room; 2. The audio device of claim 1, wherein the flutter echo estimate is determined in response to a room dimension in a first direction relative to a room dimension in a second direction.
3. the room metadata includes acoustic reflection data for sides of the room; 2. The audio device of claim 1, wherein the flutter echo estimate is determined in response to an acoustic return loss of a first boundary of the room relative to an acoustic return loss of a second boundary of the room.
4. the adaptor circuit increases a feedback coefficient; the feedback coefficient is for the first feedback loop; 2. The audio device of claim 1, wherein the feedback coefficient is based on the flutter echo estimate.
5. The method of claim 1, wherein at least a second portion of the plurality of feedback loops each has a second feedback coefficient; 2. The audio device of claim 1, wherein a third portion of the second portion depends on a room dimension of the room.
6. the signal generator circuit generates a diffuse reverberation signal from an output of a fourth portion of the plurality of feedback loops; the first portion of the plurality of feedback loops does not include any part of the fourth portion of the plurality of feedback loops; the adaptor circuitry varies a fifth portion of the plurality of feedback loops in response to the flutter echo estimate; 2. The audio device of claim 1, wherein the first portion of the plurality of feedback loops includes the fifth portion of the plurality of feedback loops.
7. the signal generator circuit has a delay relative to the audio source signal; the delay is applied to the audio source signal before the first feedback loop; the first portion of the plurality of feedback loops includes the first feedback loop; 2. The audio device of claim 1, wherein the adapter circuitry varies the delay in response to a position of at least one of an audio source of the audio source signal, a listener, and a boundary of the room.
8. the first portion of the plurality of feedback loops comprises at least two feedback loops; the signal generator circuit has a delay relative to the audio source signal; the delay is applied to the audio source signal before the at least two feedback loops; 2. The audio device of claim 1, wherein the delay is different for each of the at least two feedback loops.
9. An audio device as described in claim 1, wherein the first portion of the plurality of feedback loops has one loop or two loops.
10. The plurality of feedback loops include the first portion and the sixth portion, the first portion of the plurality of feedback loops does not include any part of the sixth portion of the plurality of feedback loops; 2. The audio device of claim 1, wherein the adapter circuit changes a feedback coefficient for at least one of the plurality of feedback loops such that there is no feedback from a feedback loop in the first portion of the plurality of feedback loops to any feedback loop in the sixth portion of the plurality of feedback loops.
11. The audio device further comprising a spatial processor circuit and a combiner circuit; the spatial processor circuit applies spatial processing to the flutter echo audio signal; the spatial processing is dependent on the location of at least one of the source of the audio source signal and the boundary of the room; the combiner circuit combines the diffuse reverberation signal and the spatially processed flutter echo audio signal; The audio device of claim 1 , wherein the signal generator circuit generates the diffuse reverberation signal.
12. The audio device further comprises a spatial processor circuit; the spatial processor circuit applies spatial processing to the flutter echo audio signal; The audio device of claim 1 , wherein the spatial processing depends on the location of at least one of a source of the audio source signal and a side of the room.
13. The audio device further comprising a first circuit; the first circuit supplies a plurality of audio source signals to the plurality of feedback loops; 2. The audio device of claim 1, wherein at least one audio source signal is supplied to only the feedback loops of the first portion of the plurality of feedback loops.
14. the signal generator circuit has a gain circuit, the gain circuit being applied to the audio source signal before the first feedback loop; the first portion of the plurality of feedback loops includes the first feedback loop; 2. The audio device of claim 1, wherein the adapter circuit varies the gain circuit in response to at least one of a position of an audio source of the audio source signal, a position of a listener, a position of the room boundary, and a reflection order of the beginning of the flutter echo audio signal.
15. the flutter echo audio signal represents flutter echoes between a pair of opposing boundaries of the room; the signal generator circuit having a frequency dependent gain circuit; the frequency dependent gain circuit is applied to the audio source signal before a first feedback loop; the first portion of the plurality of feedback loops includes the first feedback loop; the adaptor circuitry varies the frequency dependent gain circuitry in response to acoustic reflection data of the room metadata relating to room boundaries; the acoustic reflection data indicates frequency-dependent acoustic characteristics for at least one room boundary; The audio device of claim 1 , wherein a pair of opposing room boundaries does not include the at least one room boundary.
16. The method of claim 1, wherein the first portion of the plurality of feedback loops has at least two feedback loops; 2. The audio device of claim 1, wherein each of the at least two feedback loops has a different loop gain.
17. A method for providing a system for receiving room metadata, the room metadata indicating a plurality of characteristics of the room; determining a flutter echo estimate of the room in response to the room metadata, the flutter echo estimate indicating a level of flutter echo in the room; generating flutter echo audio signals from output signals of a first portion of a plurality of feedback loops when the plurality of feedback loops are supplied with an audio source signal, wherein a feedback delay network comprises the first portion of the plurality of feedback loops; adapting a first parameter for a first feedback loop of the first portion of the plurality of feedback loops in response to the flutter echo estimate; A method comprising:
18. A computer program stored on a non-transitory medium which, when executed on a processor, performs the method of claim 17.
19. The audio device of claim 4, wherein the feedback coefficient is increased when the flutter echo increases.
20. The room metadata includes dimensional data about the room, 18. The method of claim 17, wherein the flutter echo estimate is determined in response to a room dimension in a first direction relative to a room dimension in a second direction.
21. The room metadata includes acoustic reflection data relating to sides of the room, 18. The method of claim 17, wherein the flutter echo estimate is determined in response to the acoustic return loss of a first boundary of the room relative to the acoustic return loss of a second boundary of the room.
22. The adapter circuit increases the feedback coefficient, the feedback coefficient is for the first feedback loop; The method of claim 17 , wherein the feedback coefficient is based on the flutter echo estimate.
23. The method of claim 20, wherein each of at least a second portion of the plurality of feedback loops has a second feedback coefficient; The method of claim 17 , wherein a third portion of the second portion depends on a room dimension of the room.
Citation Information
Patent Citations
Artificial reverberator room size control
US10559295B1
Electronic device with digital reverberator and method
US20130202125A1