Audio device and its operating method
The audio device dynamically generates airflow sounds based on listener pose and airflow characteristics, improving immersion and reducing resource demands in virtual reality and augmented reality applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2026-03-16
AI Technical Summary
Existing audio rendering technologies in virtual reality and augmented reality applications fail to provide a realistic and immersive representation of airflow sounds, such as wind noise, which often results in an artificial and restrictive audio experience.
An audio device that generates an audio signal by receiving airflow audio frequency profile data, determining listener pose characteristics, and using a frequency response generator to adapt airflow sound components based on airflow velocity and direction relative to the listener, allowing for dynamic and flexible audio generation.
The solution provides a more immersive and natural-sounding airflow experience, reducing computational resources and complexity while efficiently adapting to changes in listener posture and head orientation.
Smart Images

Figure 2026509002000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an audio device and an operating method thereof, and more particularly to an approach for generating an audio signal that provides an improved representation of audio corresponding to, for example, airflow and wind noise in a virtual environment, but is not limited thereto.
Background Art
[0002] In recent years, the diversity and scope of experiences based on audiovisual content have increased significantly, and new services and their usage and consumption methods have been continuously developed and introduced. In particular, many spatial and interactive services, applications, and experiences have been developed to provide more complex and immersive experiences.
[0003] Examples of such applications include virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications (often collectively referred to as extended reality (XR)), which are rapidly becoming mainstream, and many solutions are targeted at the consumer market. Some standards are under development by several standardization bodies. In such standardization activities, standard development for various aspects of VR / AR / MR systems, including streaming, broadcasting, rendering, etc., is actively underway.
[0004] VR applications tend to provide a user experience that corresponds to the user being in a different world / environment / scene, while AR (including Mixed Reality MR) applications tend to provide a user experience that corresponds to the user being in the current environment but with additional information or virtual objects or information added. Therefore, VR applications tend to provide a fully immersive, synthetically generated world / scene, while AR applications tend to provide a partially synthetic world / scene that is overlaid on the user's physically existing real-world scene. However, these terms are often used interchangeably and have a high degree of overlap. Below, the term eXtended Reality / XR will be used to refer to both virtual reality and augmented reality / mixed reality.
[0005] One example of an increasingly popular service is one that provides images and audio in a way that allows users to actively and dynamically interact with the system to change rendering parameters and adapt to changes in the user's position and orientation. A particularly appealing feature in many applications is the ability to change the viewer's effective viewing position and direction, for example, allowing the observer to move around and "look around" within the presented scene.
[0006] Such features, in particular, can enable the provision of virtual reality experiences to users. This allows users to move around (relatively) freely within the virtual environment and dynamically change their position and what they are looking at. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known from gaming applications, such as in the category of first-person shooter games for computers and consoles.
[0007] Furthermore, especially in virtual reality applications, it is desirable that the presented images are typically three-dimensional images, such as those displayed using a stereoscopic display. In fact, to optimize the observer's immersion, it is preferable that the user typically experiences the presented scene as a three-dimensional scene. Indeed, in virtual reality experiences, it is desirable that the user be able to select their position, viewpoint, and point in time relative to the virtual world.
[0008] In addition to visual rendering, most XR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience in which the audio source is perceived as arriving from a position corresponding to the position of a corresponding object in the visual scene. Thus, the audio scene and video scene are preferably perceived as cohesive, and both provide a complete spatial experience.
[0009] For example, virtual audio scenes generated by headphone playback using binaural audio rendering technology offer many immersive experiences. In many scenarios, such headphone playback can be based on head tracking, allowing rendering to respond to the user's head movements, significantly enhancing immersion.
[0010] A key feature for many applications is the ability to generate and / or deliver audio that can perceive the audio environment naturally and realistically.
[0011] To create an immersive experience, it is desirable to render a complete audio scene that is as close as possible to a realistic environment. Therefore, it is desirable to render not only specific active audio sources such as speakers and active sound generators, but also more subtle and general audio sources such as various environmental audio sources and background audio sources. Specific examples of such audio components are airflow and wind noise in an environment. By rendering recorded wind noise, various applications such as games have been developed that include sound components corresponding to wind noise. However, while this can provide users with the perception of such sounds, it is usually not the optimal experience and may be perceived as relatively artificial. This typically reduces user / listener immersion.
[0012] Therefore, an improved approach to rendering audio is advantageous, and in particular, an improved approach to rendering audio that responds to airflow, such as wind noise. Specifically, an approach that enables improved performance and / or behavior is advantageous because it allows for improved operation, increased flexibility, reduced complexity, faster implementation, improved audio experience, improved audio quality, reduced computational load, improved performance in virtual / mixed / augmented reality applications, improved performance and user experience in game applications, improved adaptability to listener pose variations, improved immersion, improved and / or facilitated adaptability, and / or improved performance and / or operation. [Overview of the project] [Problems that the invention aims to solve]
[0013] Therefore, the present invention preferably seeks to mitigate, reduce, or eliminate one or more of the above-mentioned drawbacks, either individually or in any combination. [Means for solving the problem]
[0014] According to one aspect of the present invention, an audio device for generating an audio signal is provided, the device comprising: a receiver configured to receive airflow audio frequency profile data showing the dependence of the airflow audio frequency profile on airflow velocity parameters; a pause determiner configured to determine the listener pause characteristics of a listener; a frequency response generator configured to determine the airflow frequency response depending on the airflow audio frequency profile data, the user's pause characteristics, and the airflow velocity characteristics of the airflow; an audio source configured to provide a first audio signal; an audio component generator configured to generate airflow audio signal components, wherein the generation includes filtering the first audio signal using airflow frequency characteristics; and an output unit configured to generate an audio signal including airflow audio signal components.
[0015] This approach enables an improved audio experience in many embodiments, providing a more immersive experience in numerous applications and scenarios. In many scenarios, it can improve the representation of airflow sounds, such as wind noise, as perceived by the listener. Furthermore, it can dynamically and flexibly adapt to reflect changes in the listener's posture, thereby providing a more immersive and realistic effect. Moreover, this approach enables the efficient generation of airflow-representing audio that adapts to changes in the user's posture. In many scenarios, it can reduce the requirements for computational resources.
[0016] Furthermore, in embodiments where airflow audio frequency profile data is received from, for example, a remote source, this approach can be configured to provide advantageous audio under remote control / guidance while maintaining low communication overhead.
[0017] This approach enables efficient content-side control and support in the generation of airflow audio on the renderer / user side.
[0018] The airflow velocity parameter can include at least one of an airflow direction parameter and an airflow speed parameter. The airflow velocity parameter can be an airflow velocity parameter with respect to a listener. The listener pose characteristic can include at least one of a listener's position parameter, a listener's orientation parameter, a listener's position change parameter (especially a listener's speed parameter, etc.), and a listener's orientation change parameter. The pose of the listener can be a pose in the (coordinate system) of the rendered audio scene.
[0019] The pose can be a position and / or an orientation.
[0020] According to an optional feature of the present invention, the frequency response generator is configured to generate an airflow frequency response depending on the airflow speed of the airflow with respect to the listener.
[0021] Thereby, the performance is improved and the complexity and resource requirements are reduced. It can provide an airflow sound that is particularly immersive and natural-sounding in many embodiments.
[0022] The frequency response generator can be configured to determine the airflow speed according to the speed value of the airflow velocity parameter of the airflow and the speed value of the listener pose.
[0023] According to an optional feature of the present invention, the frequency response generator is configured to generate an airflow frequency response depending on the airflow direction of the airflow with respect to the listener.
[0024] Thereby, the performance is improved and the complexity and resource requirements are reduced. It can provide an airflow sound that is particularly immersive and natural-sounding in many embodiments.
[0025] The frequency response generator can be configured to determine the airflow direction according to the airflow direction parameter value of the airflow and the value of the listener's pose direction.
[0026] According to an optional feature of the present invention, the first audio signal is a noise audio signal.
[0027] Thereby, the performance is improved, and the complexity and resource requirements are reduced. Thereby, the complexity is low, the implementation is easy, and the resource usage is reduced. The noise signal can be generated using operations with low complexity, and several different algorithms with low resource usage are known and can be used. The noise audio signal can be dynamically generated during operation.
[0028] The audio source can consist of, or include, a pseudo-noise generator that generates a pseudo-noise audio signal. The noise audio signal can be a probabilistic signal. The noise audio signal can be, for example, a white noise or pink noise audio signal.
[0029] According to an optional feature of the present invention, the audio component generator is configured to generate an airflow audio signal component as a stereo airflow audio signal component having a first channel and a second channel, and the output unit is configured to generate an audio signal as a stereo audio signal having a first channel and a second channel.
[0030] In many embodiments, this can provide an improved user experience, and in particular, it can provide an audio signal that delivers a more natural-sounding and immersive airflow sound. This can improve the adaptability of the airflow sound to user movements, including changes in head orientation. The first channel of the stereo airflow audio signal component and the first channel of the stereo audio signal may be the left channel, and the second channel of the stereo airflow audio signal component and the second channel of the stereo audio signal may be the right channel. The output section may be configured to include the first channel signal component of the airflow audio signal component in the first channel of the stereo audio signal, and the second channel signal component of the airflow audio signal component in the second channel of the stereo audio signal.
[0031] According to an optional feature of the present invention, the audio source is configured to generate a first audio signal as a stereo audio signal having different signals for the first channel and the second channel.
[0032] In many embodiments, this can provide an improved user experience, particularly by providing audio signals that offer a more natural-sounding and immersive airflow sound. This can improve the adaptability of the airflow sound to listener movements, such as when the listener turns their head. This enables a computationally and / or functionally efficient approach to generating airflow noise / audio with appropriate adaptability to head orientation and / or appropriate externalization.
[0033] According to an optional feature of the present invention, the frequency response generator is configured to generate an airflow frequency response including a first airflow frequency response for a first channel and a second airflow frequency response for a second channel, and the audio component generator is configured to generate a first channel signal component of the airflow audio signal component using the first airflow frequency response for filtering and a second channel signal component of the airflow audio signal component using the second airflow frequency response for filtering.
[0034] This can provide an improved user experience in many embodiments, and in particular, it can provide an audio signal that delivers a more natural-sounding and immersive airflow sound. This can improve the adaptability of the airflow sound to user movements, including changes in head orientation. This enables a computationally and / or functionally efficient approach to generating airflow noise with appropriate adaptability to head orientation and / or appropriate externalization. The first and second airflow frequency responses may differ at at least some values of the airflow characteristics.
[0035] According to an optional feature of the present invention, the audio device is configured to generate airflow audio signal components such that they have signals that are at least partially uncorrelated with respect to the first channel and the second channel.
[0036] In many embodiments, this can provide an improved user experience, and in particular, it can provide an audio signal that delivers a more natural-sounding and immersive airflow sound. This can improve the adaptability of the airflow sound to user movements, including changes in head orientation. This approach can provide a degree of externalization of the airflow sound.
[0037] According to an optional feature of the present invention, the audio device is configured to adapt the degree of decorrelation between the first and second channels of the stereo airflow audio signal components depending on the direction of the airflow toward the listener.
[0038] This can provide an improved user experience in many embodiments, and in particular, can provide an audio signal that delivers immersive airflow sounds that sound more natural. This can improve the adaptability of the airflow sounds to user movements, including changes in head orientation. This approach can provide varying degrees of externalization of the airflow sounds.
[0039] According to an optional feature of the present invention, the airflow audio frequency profile data includes an index of the first dependence of the first airflow audio frequency profile on an airflow direction parameter, and an index of the second dependence of the second airflow audio frequency profile on an airflow velocity parameter, and the frequency response generator is configured to generate a first frequency response according to the first dependence and the airflow direction of the airflow to the listener, and a second frequency response according to the second dependence and the airflow velocity of the airflow to the listener, and is configured to generate a frequency response as a combination of the first and second frequency responses.
[0040] This improves performance and reduces complexity and resource demands.
[0041] According to an optional feature of the present invention, the audio signal is a stored audio signal.
[0042] This improves performance and reduces complexity and resource demands.
[0043] According to an optional feature of the present invention, the airflow audio frequency profile data includes an index of relative airflow audio frequency response values for each of several airflow velocity parameter values, and the frequency response generator is configured to determine other relative airflow audio frequency response values for other values of the airflow velocity parameter by interpolating from a plurality of airflow velocity parameter values.
[0044] This improves performance and reduces complexity and resource demands.
[0045] According to an optional feature of the present invention, the receiver is configured to receive an indicator of the properties of the airflow source of the airflow, and the frequency response generator is configured to determine the airflow velocity characteristics according to the characteristics of the airflow source.
[0046] This improves performance and reduces complexity and resource demands. In many embodiments, it enables improved adaptation, for example, allowing for lower complexity and enabling characterization / adaptation of airflow characteristics with lower communication overhead.
[0047] According to an optional feature of the present invention, the indicator of the characteristics of the airflow source is configured to indicate that the airflow source is at least one of the following: a global airflow source, an omnidirectional airflow source, a point airflow source, and a conical airflow source.
[0048] This improves performance and reduces complexity and resource demands.
[0049] According to an optional feature of the present invention, the receiver is configured to receive an indicator of the characteristics of the airflow source as part of the metadata of the audio bitstream received from the remote source.
[0050] This enables, for example, the characterization and adaptation of airflow characteristics using remote sources with low communication overhead.
[0051] According to one aspect of the present invention, a method for generating an audio signal is provided, the method comprising: receiving airflow audio frequency profile data showing the dependence of the airflow audio frequency profile on airflow velocity parameters; determining the listener pause characteristics of a listener; determining the airflow frequency response depending on the airflow audio frequency profile data, user pause characteristics and airflow velocity characteristics of the airflow; generating airflow audio signal components, the method comprising filtering a first audio signal using the airflow frequency response; and generating an audio signal to include airflow audio signal components.
[0052] These and other aspects, features and advantages of the present invention will become apparent from and be described with reference to the embodiments described below. [Brief explanation of the drawing]
[0053] Embodiments of the present invention will be described with reference to the drawings, merely as examples. [Figure 1] A diagram illustrating an example of elements in an augmented reality system. [Figure 2] A figure showing an example of an audio device according to several embodiments of the present invention. [Figure 3] A diagram showing some elements of a possible configuration of a processor for implementing elements of an audio device according to some embodiments of the present invention. [Modes for carrying out the invention]
[0054] The following description focuses on audio processing and rendering in augmented reality (XR) applications, such as virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications. The approach described focuses on applications where audio rendering is adapted to reflect acoustic changes as the (possibly virtual) user / listener's pose changes, or changes in the audio perception of airflow audio. However, it will be understood that the principles and concepts described can be used in many other applications and embodiments, including, for example, game applications where a virtual game world is presented as a spatial audio signal on a two-dimensional display.
[0055] Semi-virtual or fully virtual experiences, where users can (perhaps partially) navigate a virtual world, are becoming increasingly popular, and services are being developed to meet such demand.
[0056] In some systems, XR applications can be provided locally to the observer by a standalone device that does not use, or even have access to, remote XR data or processing. For example, a device such as a game console may have a storage device for storing scene data, an input for receiving / generating observer poses, and a processor for generating corresponding images from the scene data.
[0057] In other systems, XR applications can be implemented and executed remotely from the observer. For example, a device local to the user can detect / receive motion / pose data that is sent to a remote device that processes the data to generate observer poses. The remote device can then generate a view image and corresponding audio signal appropriate for the user pose, based on scene data describing the scene. The view image and corresponding audio signal are then sent to the device local to the observer and presented there. For example, the remote device can directly generate a video stream (usually a stereo / 3D video stream) and a corresponding audio stream that are presented directly by the local device. Therefore, in such an example, the local device does not need to perform any XR processing other than sending motion data and presenting the received video data.
[0058] In many systems, functionality can be distributed across local and remote devices. For example, a local device can process received input and sensor data to generate user poses that are continuously transmitted to a remote XR device. The remote XR device can generate corresponding view images and audio signals and transmit them to the local device for presentation. In other systems, the remote XR device may not directly generate view images and corresponding audio signals, but instead select relevant scene data and transmit it to the local device, which then generates the presented view images and corresponding audio signals. For example, a remote XR device may identify the nearest capture point, extract the corresponding scene data (e.g., a set of object sources and their position metadata), and transmit this to the local device. The local device can process the received scene data to generate images and audio signals for a specific current user pose. User poses typically correspond to head poses, and references to user poses can usually be considered equivalent to references to head poses.
[0059] In many applications, particularly broadcast services, a source can transmit or stream scene data in the form of image (including video) and audio representations of the scene, independently of user pauses. For example, signals and metadata corresponding to audio sources within a certain virtual room can be transmitted or streamed to multiple clients. Each client can locally synthesize the audio signal corresponding to the current user pause. Similarly, a source can transmit a general description of an audio environment, including audio sources within the environment and a description of the environment's acoustic properties. The audio representation can then be generated locally and presented to the user, for example, using binaural rendering and processing.
[0060] Figure 1 shows an example of an XR system in which a remote XR client device 101 exchanges information with an XR server 103 via a network 105, such as the Internet. The server 103 can potentially be configured to support a large number of client devices 101 simultaneously.
[0061] The XR server 103 can support the broadcast experience by, for example, transmitting an image signal containing an image representation in the form of image data that can be used by a client device to locally synthesize a view image corresponding to an appropriate user pose (where pose refers to position and / or orientation). Similarly, the XR server 103 transmits an audio representation of the scene, enabling the audio to be synthesized locally in response to the user pose. Specifically, as the user moves around in the virtual environment, the synthesized and presented images and audio are updated to reflect the user's current (virtual) position and orientation in the (virtual) environment.
[0062] Therefore, in many applications like the one shown in Figure 1, it is desirable to model the scene and generate efficient image and audio representations that can be efficiently included in the data signal, and then the data signal can be sent to or streamed to various devices that can locally synthesize views and audio for poses different from the captured pose.
[0063] In computer games and (fully or partially) virtual environments such as augmented reality (AR) and virtual reality (VR), content creators typically aim to provide users with an immersive experience. Part of this immersion involves creating realistic sound effects that match visual elements, mimicking the weather and other factors of the user's local environment.
[0064] A good example of this is the sound of wind that corresponds to the movement of trees, leaves, and other objects in the environment. One problem is that wind itself has no sound; rather, it is the interaction of wind with other objects that produces sound, such as the movement of tree branches or the whistling sound (Aeolian sound) of wind blowing on ship's rigging. One of the key elements of how a listener experiences wind is the sound produced as the air passes through the listener's ears.
[0065] When air passes through the ear, turbulence is created within the ear's structure, which then causes the eardrum to vibrate, resulting in the sound we hear. The level and timbre of the sound are influenced by the speed of the air passing through the ear and the angle at which the air passes through the listener.
[0066] For example, if you stand in an open area facing a strong wind, you can hear the sound of the wind, and the sound changes as the wind strength changes. However, if you turn your head away from the wind, the level and spectrum of the sound change. Another example is when cycling; you hear sounds as they pass through the air, but if you turn your head, that is, check over your shoulder before turning, the sound changes.
[0067] Current approaches generate environmental wind noise by adding pre-recorded or synthesized sound effects. However, such approaches tend to provide relatively static sounds, which often result in an audio experience that is perceived as restrictive and relatively unrealistic.
[0068] The following describes a specific approach to generating an audio signal that includes an improved airflow audio component representing airflow audio in an audio scene. This approach can provide an improved and more realistic audio experience, and in particular, it can provide an audio signal that can more realistically represent how the sound of airflow changes dynamically in different scenarios.
[0069] The audio device in Figure 2 is configured to generate an audio signal that includes audio components representing airflow, such as wind noise, for a given listener pose. The audio signal can represent the audio of an audio scene, and this audio is generated for the listener pose within the audio scene. The airflow audio component is generated to reflect, in particular, the specific position and / or direction and / or velocity of the listener (more specifically the listener's head), as well as the characteristics of the airflow itself. In many embodiments, in addition to the airflow audio component, the generated audio signal may include several other audio components in the output audio signal, such as audio components representing other audio sources in the audio scene.
[0070] The audio device includes a pause determiner 201 configured to determine the listener's pause generated in response to an audio signal.
[0071] In this field, the terms position and pose are used as general terms relating to location and / or orientation. For example, a combination of the position and orientation of an object, camera, head, or view may be called a pose or position. Thus, a position or pose index can include six values / components / degrees of freedom, each value / component typically describing an individual characteristic of the corresponding object's position or orientation. Of course, in many situations, for example, when one or more components are considered fixed or irrelevant (e.g., if all objects are considered to be at the same height and have a horizontal orientation, four components may provide a complete representation of the object's pose), the position or pose may be considered or represented with fewer components. Hereinafter, the term pose is used to refer to a position and / or orientation that can be represented by 1 to 6 values (and more if the orientation is represented by a quaternion or rotation matrix) (corresponding to the maximum possible degrees of freedom). The term pose can be replaced with the term position. The term pose can be replaced with the term location and / or orientation. The term "pose" can be replaced by the terms "position and orientation" (if the pose provides both positional and orientational information), by the term "position" (if the pose provides positional information, if applicable), or by "orientation" (if the pose provides orientational information, if applicable).
[0072] The pose determiner 201 can determine listener pose characteristics that reflect the position and / or orientation characteristics of the (nominal / virtual) listener to which the audio signal is generated. The listener pose characteristics are typically provided with reference to the presented audio scene (e.g., with reference to the scene coordinate system of the audio scene), and especially in the case of rendering a virtual scene, the listener pose and characteristics are provided with reference to the scene coordinate system of the virtual scene. The characteristics are typically pose values or rate of change values of the pose. Specifically, in many embodiments, the pose determiner 201 can be configured to determine the orientation and / or speed / velocity of the listener, or in particular the listener's head. In many embodiments, the listener pose characteristics can also (or perhaps instead) indicate the position of the listener (in particular the listener's head).
[0073] Many different approaches are known for determining and providing listener / user / observer poses in a scene / environment, and it will be understood that the appropriate approach can be used. For example, the second receiver 203 can be configured to receive pose data from a VR headset, eye tracker, etc., worn by the user. In other embodiments and applications, a controller or joystick can be used to control, for example, a virtual person / avatar / character in a virtual environment. Such control is well known, for example, in computer game applications. For example, in a game application, a player can use a joystick or other game controller to control an avatar in a virtual environment. The corresponding pose (typically both position and orientation) in the virtual environment is determined, and the game application can generate a view of the virtual scene presented to the player on, for example, a monitor or other suitable 2D display. The audio device in Figure 2 can further use this pose as a listener pose to generate an audio signal presented to the user, and this audio signal and airflow sound components are generated by the audio device in Figure 2 based on the determined pose (i.e., poses controlled by the controller are also used as listener poses).
[0074] The audio device further includes a receiver 203 configured to receive airflow audio frequency profile data that shows the dependence of the airflow audio frequency profile on an airflow velocity parameter. The airflow velocity parameter is, in particular, the speed and / or direction of the airflow (usually relative to the listener pause), and the airflow audio frequency profile data can provide information on the frequency distribution / profile of airflow audio components for different values of the airflow velocity parameter. The airflow audio components can represent the audio perceived by the listener for a particular airflow velocity parameter. The airflow audio frequency profile data can usually be provided for a constant nominal / reference listener pause. Similarly, the airflow velocity parameter can usually be a relative airflow velocity parameter that shows the relative characteristics of the airflow velocity with respect to the listener's pause.
[0075] Therefore, airflow audio frequency profile data can provide indicators / information showing how the frequency distribution of airflow sound / audio changes with changes in the values of airflow velocity parameters, particularly changes in the values of airflow velocity and / or direction (relative to a reference listener pause).
[0076] The audio device further includes a frequency response generator 205 coupled to a receiver 203 and a pause determiner 201, which receives airflow audio frequency profile data and listener pause data. The frequency response generator 205 is configured to determine the airflow frequency response depending on the airflow audio frequency profile data, listener pause characteristics, and airflow velocity characteristics of the airflow. Specifically, airflow velocity characteristics can be airflow speed and / or airflow direction. For example, airflow characteristics can indicate the speed and direction of wind in the audio scene. Listener pause characteristics can be, for example, the listener's position and / or orientation, and / or their derivatives, such as the speed and direction of the listener's movement.
[0077] For example, airflow audio frequency profile data can provide an index of the frequency response of different relative airflow velocities to the listener's pause (corresponding to the listener's head). The airflow velocity characteristics indicate the direction and speed of the airflow, and the listener pause characteristics indicate the direction and speed of the listener pause / listener's head. From these, the frequency response generator 205 can determine the relative speed and direction of the airflow to the listener's pause / head. Then, by accessing the airflow audio frequency profile data, the frequency response provided for the corresponding airflow velocity parameter can be extracted, that is, the frequency profile provided for the airflow velocity parameter that matches the relative speed and direction of the airflow can be extracted.
[0078] The frequency response generator 205 is coupled to the audio component generator 207, which in turn is coupled to the audio source 209. The audio source 209 provides an audio signal to the audio component generator 207, which is configured to filter the audio signal based on the airflow frequency response. The audio component generator 207 can, in particular, generate a filter having a frequency response corresponding to / matching the determined airflow frequency response and apply it to the received audio signal. The filtering can adapt the audio of the audio signal accordingly, allowing it to more closely reflect the characteristics of the airflow audio perceived by the listener in listener pauses (and in the direction / velocity represented by the listener pauses). The audio component generator 207 thus generates an airflow audio signal component by filtering the audio signal from the audio source 209 using the airflow frequency response. In some embodiments, it will be understood that the audio source 209 can also perform other operations to generate the airflow audio component, such as amplitude level setting, other filtering, etc.
[0079] In many embodiments, the audio component generator 207 is configured to adapt the level of the airflow audio component in accordance with (typically relative) airflow velocity parameters. In many embodiments, the audio component generator 207 is configured to adapt the level of the airflow audio component for at least one other audio component in the output audio signal, and typically for all other audio components. For example, the audio component generator 207 can be configured to set the level of the airflow audio component as a monotonically increasing function of the airflow velocity (at the listener pause) relative to the listener pause. Thus, for example, the stronger / faster the wind, the louder the wind noise.
[0080] The filtered audio signal is fed to an output generator 211 that generates an audio signal containing an airflow audio signal component. In many embodiments, the output generator 211 includes a mixer / combiner configured to mix / combine different audio components into a single audio signal. For example, audio signal components can be generated for individual audio sources within an audio scene, such as ambient background audio and individual audio point sources. The various audio components are combined into a single output audio signal that provides a complete rendering of the audio scene, with the airflow noise component contributing to the overall proximity of the audio sources.
[0081] The audio device in Figure 2 is configured to receive data from a local or remote source showing how the frequency response of the airflow audio changes with changes in the listener pause.
[0082] In some embodiments, the receiver 203 can be coupled to an internal memory of an auxiliary power supply from which airflow audio frequency profile data is stored and from which the receiver 101 retrieves appropriate airflow audio frequency profile data.
[0083] For example, the internal memory may contain frequency responses to various different values of one or more pause parameters, such as frequency responses to each of several different airflow directions relative to the listener pause, and / or frequency responses to each of several different airflow velocities relative to the listener. Each frequency response may be represented, for example, by several different gain values for different frequencies, or by parameter values of a given gain function as a function of frequency.
[0084] In such cases, the frequency response generator 205 can be configured to determine and extract, for example, the stored frequency response for the speed and direction that best match the determined relative airflow velocity.
[0085] In many embodiments, airflow audio frequency profile data can be received from a remote source. For example, an audio device may be part of a client device 103, and the airflow audio frequency profile data may be received from a server 103. The airflow audio frequency profile data may be received in a bitstream that includes, in particular, audio data for individual audio sources, location information for such audio sources, background audio data, and other data describing the audio scene. Thus, the bitstream can provide a representation of the audio scene that enables the client device 103 to render the audio scene. This rendering may include rendering airflow audio / wind noise based on the airflow audio frequency profile data provided in the bitstream.
[0086] Therefore, this approach can be adapted locally to changes in listener pause, for example, while providing an efficient approach on the content source side or assisting with how to render airflow audio on the client side.
[0087] In this approach, the filter's frequency response can be adapted to reflect changes in airflow audio resulting from changes in airflow relative to the listener. Typically, the frequency distribution of locally generated audio signals is modified accordingly, and the audio signal is shaped in response to fluctuations in relative airflow velocity.
[0088] In different embodiments, audio signals from different sources can be used and provided by the audio source 209.
[0089] In many embodiments, the audio source 209 can generate audio signals as probabilistic / pseudorandom signals. In many embodiments, the audio source can be a noise generator that generates noise signals, and in many embodiments, the noise signals can be white noise signals or color noise signals such as pink noise signals. In fact, it has been found that using such noise signals as the basis for characterizing relative airflow velocity dependence by a determined frequency response can produce airflow sounds that sound very realistic in many scenarios and applications.
[0090] In some embodiments, the audio source 209 can specifically be an audio signal that is a non-uniform mix of frequencies over a given range, or a noise source that provides pink noise having equal energy per octave.
[0091] In some embodiments, the audio source 209 may be configured to dynamically generate an audio signal during operation, and in particular, to generate it as a noise signal. However, in other embodiments, the audio signal may be a dedicated audio signal stored locally, for example. For example, the audio source 209 may include a recorded and stored wind noise audio signal that is acquired and provided to the audio component generator 207.
[0092] In many embodiments, the receiver 203 can receive the audio signal extracted by the audio source 209 and provided to the audio component generator 207 as an audio signal for frequency shaping to generate airflow audio components, as part of a bitstream that provides, for example, audio for an audio scene (and also, for example, airflow audio frequency profile data).
[0093] In some embodiments, such audio signals may be recorded and stored for different relative velocity values, for example, and the audio source 209 may be configured to extract the one that best matches the determined relative velocity characteristics.
[0094] Therefore, in some embodiments, pre-rendered or recorded audio fragments are provided for one or more known speeds, and the frequency response can be determined to reflect the relative variation of these signals to the frequency distribution. For example, the determined frequency response is designed to have a flat response for known speeds, so that filtering is applied only to other speeds.
[0095] In some embodiments, the generated airflow audio component and audio signal may be single-channel signals; that is, the device can generate a monaural audio signal that includes a monaural representation of airflow noise. However, to improve the user experience through enhanced spatial awareness and advanced externalization, the airflow audio component and output audio signal are generated as multi-channel signals, specifically as stereo signals. The output stereo signal can be a binaural signal, which is suitable for rendering to the user using headphones, for example.
[0096] Therefore, in many embodiments, the audio component generator 207 is configured to generate airflow audio components such that they become stereo airflow audio signal components having two channels. Similarly, the output unit 211 can generate an audio signal as a stereo audio signal having two channels. In particular, the output unit 211 can include one channel of the airflow audio component in one channel of the output stereo audio signal and the other channel of the airflow audio component in the other channel of the output stereo audio signal.
[0097] Therefore, the airflow audio component is generated to have different signal components for the two channels (hereinafter referred to as left and right channels for convenience, but please understand that this does not impose any limitations). Consequently, the output audio signal is also generated as a stereo signal with different signals for the left and right channels.
[0098] In some embodiments, the audio source 209 generates a stereo audio signal with different signals in two channels, and these signals are filtered by the same filter / frequency response generated by the frequency response generator 205 to produce an airflow audio component with different signals in two channels.
[0099] In other embodiments, the audio source 209 may generate a monaural signal, and the audio component generator 207 may be configured to apply different filters to the two channels. Thus, in this example, the audio component generator 207 generates stereo components through the different filters applied.
[0100] In many embodiments, the audio source 209 can generate stereo signals having different channel signals, which can then be filtered by the audio component generator 207 using different frequency responses for the two channels. Therefore, in this example, the difference between channels can be caused by both the audio signals used and the different filters employed.
[0101] In many embodiments, the airflow audio frequency profile data can include indices of a stereo frequency profile. For example, for each airflow velocity parameter for which data is provided, two frequency responses may be shown, one for the left channel and one for the right channel. Thus, in many embodiments, the frequency response generator 205 can generate two airflow frequency responses, which are applied to the left and right channels respectively by the audio component generator 207.
[0102] In many embodiments, the device can generate airflow audio signal components such that the signals for the two channels are at least partially uncorrelated. This can be specifically achieved, for example, by having the audio source 209 generate a pseudo-noise signal with a predetermined amount of uncorrelatedness. Uncorrelatedness can specifically help provide increased externalization so that the airflow noise is increasingly perceived as being outside the listener's head.
[0103] In some embodiments, the degree of decorrelation can be adapted according to the relative airflow velocity. Specifically, the device can be configured to adapt the degree of decorrelation depending on the airflow direction relative to the listener. This is achieved, for example, by the audio source 209 generating a decorrelation signal to which the amount of decorrelation is adapted.
[0104] Such a variable degree of decorrelation can be achieved, for example, by introducing a small time offset between the two channels, where the amount of difference controls the degree of decorrelation, or by filtering the two signals with independent, fully transparent filters constructed to have different phase responses.
[0105] In some embodiments, the device can be configured to vary the degree of decorrelation as a monotonically increasing function of the angular difference between the relative airflow direction and the direction directly in front of the listener. Thus, the further the relative airflow direction deviates from directly in front of the listener, the higher the level of decorrelation.
[0106] Such variable discorrelation can provide a more realistic and immersive experience, particularly by reflecting how wind tends to be perceived more externally, as the difference in amplitude levels between the left and right ears is greater when the wind is coming from the side to the listener. Varying the correlation level can enhance this effect while maintaining a lower level of sound presentation, which may be more comfortable for the listener.
[0107] In an exemplary embodiment, the airflow audio frequency profile data can parameterize the induced airflow noise frequency profile, for example, at two or more known velocities and / or in two or more relative directions (e.g., the angle of incidence of the airflow relative to the listener pause). The frequency response generator 205 can specifically determine the airflow velocity vector relative to the listener pause in the virtual environment.
[0108] The frequency response generator 205 can further determine the noise frequency profile with respect to the relative velocity parameter and generate airflow audio components by filtering, for example, a white noise or pink noise audio signal.
[0109] To generate a sound source that emulates noise generated in the ear by airflow over the ear at an arbitrary velocity, the frequency response generator 205 generates a target spectrum and can filter random noise, such as white noise or pink noise, through that target spectrum. The random noise can typically be a binaural signal, and the interaural correlation can be controlled so that the perceived externalization of the reproduced signal is altered. However, in some embodiments, the device can generate only a monaural signal without considering the direction of pause (e.g., head rotation) and present the same signal to both ears.
[0110] This approach provides an efficient way to characterize the relationship between perceived sound and the user's relative motion in the air, and / or allows content creators to specify a desired sound and identify how it changes with relative velocity and angle. Since the actual sound effects played to the user can be generated in the renderer, there is no need to store or send additional audio from the content creator to the user.
[0111] This approach can improve performance, for example, by generating airflow noise audio in real time in the audio renderer while adapting to the relative airflow velocity and relative direction of the airflow relative to the listener.
[0112] This approach uses an efficient method to parameterize frequency-dependent amplitude changes, yielding different angle and velocity combinations. This is particularly useful for streaming immersive experiences where using the smallest possible amount of data is beneficial.
[0113] In many embodiments, airflow audio frequency profile data can include data showing how the frequency response depends on airflow direction and airflow velocity parameters. Thus, airflow audio frequency profile data can reflect dependencies on both relative direction and relative velocity. In many embodiments, airflow audio frequency profile data can include frequency responses for each of several different combinations of velocity and direction. In this case, the frequency response generator 205 can extract the frequency response provided for the velocity and direction that best matches the determined relative airflow velocity and direction.
[0114] However, in some embodiments, the airflow audio frequency profile data may include separate data for airflow velocity and airflow direction. For example, the airflow audio frequency profile data may include frequency responses for each of a plurality of relative airflow directions, and further include separate frequency responses for each of a plurality of relative airflow velocities.
[0115] In such a case, the frequency response generator 205 can be configured to extract one frequency response for the relative airflow direction that best matches the relative airflow direction determined relative to the listener pause. Furthermore, it can extract one frequency response for the relative airflow direction that is closest to the relative airflow velocity determined relative to the listener pause. The frequency response generator 205 can then combine the two extracted filters to generate a frequency response used to filter the audio signal from the audio source 209. For example, the frequency responses can be considered to correspond to separate sequential filters. Therefore, for example, the combined frequency response can be determined by multiplying it by a normalized frequency gain coefficient.
[0116] Therefore, in some embodiments, separate frequency responses are provided for different relative airflow velocities and different relative airflow directions, and the separate air frequency responses can then be coupled to the frequency response for use in filtering an audio signal to generate an airflow audio component.
[0117] In the previous example, the airflow audio frequency profile data provides frequency responses for different values of the airflow velocity parameter, and the frequency response generator 205 can be configured to extract and use the frequency response of the airflow velocity parameter closest to the determined relative airflow velocity parameter. However, in some embodiments, the frequency response generator 205 can be configured to interpolate between the frequency responses for the airflow velocity parameter values.
[0118] In situations where airflow audio frequency profile data provides two or more frequency responses / spectrums, improved interpolation of the frequency response is provided, leading to a more immersive experience and a more realistic perception of the audio scene.
[0119] In some such cases, the interpolation method can be specified, for example, in the airflow audio frequency profile data and / or as part of the metadata of the received bitstream characterizing the audio scene. For example, the airflow audio frequency profile data may define that, for example, cubic interpolation should be used. This allows content creators to have more granular control over how intermediate target spectra are constructed while minimizing the number of spectra required.
[0120] In some embodiments, minimum and maximum speeds may also be given, beyond which no interpolation occurs, or only interpolation of specific parameters occurs. When the target spectrum is given as filter parameters, it is desirable that only the gain changes beyond the upper speed threshold, or that only the center frequency and Q change, rather than the gain. Similarly, when frequency-gain pairs are specified, beyond a certain threshold, it may be desirable to continue changing the relative levels between frequencies rather than increasing the maximum gain further. (i.e., modifying the shape of the spectrum so as not to exceed the maximum level the system can handle, but without changing the overall level).
[0121] In some embodiments, the frequency response / target spectrum can be represented by multiple frequency-gain pairs. However, in some embodiments, it is advantageous to specify the frequency response in terms of a set of bandpass filters. For example, each known target spectrum for a given velocity and / or angle of incidence can be constructed from a set of parametric bandpass filters that can be defined using a center frequency (CF), bandwidth (or Q, Q = center frequency / -3dB bandwidth), and gain (g), and optionally the order of the filter. Typically, three bandpass filters are sufficient to represent the target spectrum, which means nine parameters per spectrum to be preserved / transmitted, compared to 17 when using frequency-gain pairs spaced 1 / 3 octave apart in the range of 20–1000 Hz as shown in Figure 1. Filters for velocities other than the known velocity can be derived in the same way as direct frequency-gain pairs using linear interpolation of CF, Q, and gain values.
[0122] The filter can be applied directly to the signal, for example, as a Butterworth bandpass filter, or frequency-gain values can be derived for multiple specified frequencies. One way to do this is to solve the quadratic function of each filter, and once three known frequency-gain values are obtained from combinations of center frequency, Q, and gain, the polynomial can be computed. any frequency f i For this, use the following quadratic function to gain g i It is possible to calculate this. g i = af i 2 + bf i + c
[0123] Here,
number
[0124] In some embodiments, airflow audio can be generated based on the general characteristics of the airflow present in the audio scene. However, in some embodiments, the receiver 203 can be configured to receive data describing the characteristics of the airflow source from which the audio is being generated. In such cases, the frequency response generator 205 can determine the airflow velocity characteristics depending on the indicated characteristics of the airflow source.
[0125] Airflow source characteristics can, in particular, indicate the spatial characteristics of the airflow source / origin that generates / generates the airflow. In many embodiments, airflow source characteristics can, in particular, indicate the characteristics of the origin of the airflow.
[0126] For example, airflow source characteristics indicate the spatial extent of the airflow source, and specifically, they can indicate the spatial extent and / or location and / or direction of the airflow source, such as whether it is a global source without a specific starting point, an airflow point source where airflow originates from a specific starting point, an omnidirectional airflow source where airflow originates in all directions, or a conical airflow source where airflow originates and spreads according to a conical shape, for example.
[0127] The frequency response generator 205 can determine the airflow velocity value at the listener's position from the airflow source characteristics. Such calculations / determinations can be based on known physical properties, such as those known from airflow physics / dynamics. For example, the velocity value is assumed to decrease according to the inverse square law, where the decrease in velocity is inversely proportional to the square of the distance to the airflow source, if the airflow source is a point source, or in proportion to the cross-sectional area of a conical source.
[0128] In some embodiments, data describing the characteristics of an airflow source can be stored or generated locally. For example, in a game application, the game can locally generate characteristics reflecting an airflow source and provide this to the receiver 203. For example, if the game environment includes wind noise, the game application can generate data indicating the presence of a global airflow source with a specific airflow / wind speed. If a specific airflow source is present, data can be provided indicating, for example, that an airflow of a specific direction and speed originates from a specific point in the audio scene.
[0129] However, in many embodiments, the receiver 203 can be configured to receive indicators of the characteristics of the airflow source as part of the metadata of the audio bitstream received from the remote source. In this example, the client device 101 can receive airflow source data from the server 103, in particular, as metadata of the bitstream describing the audio scene.
[0130] In this case, the receiver 203 can extract the airflow source characteristics, and the frequency response generator 205 can determine the airflow velocity at the current listener position. It then determines the relative airflow velocity (with respect to the user pause) and uses this to determine the appropriate frequency response used to determine the airflow audio component as described above.
[0131] This approach can provide a highly efficient and advantageous approach for generating airflow audio, particularly for remote content-side control of such audio. This approach enables this while keeping complexity and communication overhead low.
[0132] In many common applications and audio scenes, airflow can be generated by various sources, such as ambient wind, fans, and HVAC systems. For visual objects and audio scenes to match well, it's desirable to be able to describe the source generating the airflow and how it moves within the virtual environment.
[0133] One example is characterizing an airflow source as a global source or global airflow source. Such a global source is the simplest form and usually best represents a source of airflow, such as atmospheric wind. A global source may not have a location, but it does have a flow vector (or direction) and velocity. The velocity received by the user is not affected by the location in the virtual environment.
[0134] An optional region can be specified where the global source is active or inactive. The inactive region can be used to represent buildings where the global source does not affect users but does affect those outside of it.
[0135] As another example, an omnidirectional source may have position, velocity, and an optional distance fade parameter. The flow vector points from the omnidirectional source's position to the user, regardless of the user's position. The optional distance fade parameter reduces the velocity as the distance from the source increases, for example, using an inverse square law.
[0136] As another example, a point source has a position, orientation, azimuth range, elevation range, an optional edge fade parameter, and an optional distance fade parameter. The flow vector points from the point source's position to the user's position, but is only active if the vector is within the azimuth and elevation ranges. The azimuth and elevation ranges are centered on the front direction, and the orientation parameter rotates the point source relative to the virtual environment.
[0137] The edge fade parameter reduces the intensity of the fade as the user approaches the limits of the azimuth or elevation range. This is expressed as a percentage (or a value between 0 and 1), where 0% means no fade, and 100% means the fade starts at 0° azimuth, 0° elevation and fades linearly (or logarithmically, or according to another curve) towards the edge. Other values represent the point at which the fade begins.
[0138] For example, if we consider only the azimuth angle and assume the same applies to the elevation angle, the azimuth range is set to 90° and the edge fade parameter is set to 50%. A listener positioned at 30° relative to the source will not have edge fade applied, a user positioned at 60° relative to the source will have a 33% edge fade reduction applied, and a user at 90° will have a 100% reduction applied.
[0139] As yet another example, an airflow audio source can be characterized as a cone / cylinder airflow source. Such a source has position, direction, length, end radius, optional start radius, and optional edge fade and distance fade parameters. Essentially, they can behave similarly to point sources, except that the azimuthal and elevation edges are calculated based on the cone length and end radius. The optional start radius forms a frustum, and the flow vector points to the user from the theoretical tip of the cone. If the start and end radii are the same, a cylinder is created, and the flow vector is perpendicular to the axis of the cylinder defined by the direction parameter.
[0140] This approach allows for changes to the position, direction, and velocity of the airflow source. These changes can be performed by the user, an external source such as a physics engine controlling the visual rendering, or a random sequence generator.
[0141] Content creators can also specify animations for changes, such as time-based animations, to toggle sources on and off, or move them in a predetermined order.
[0142] For velocity or direction, a pseudo-random number sequence can be specified. Alternatively, given a desired range and distribution of random values, a random number generator can be used to construct a random number sequence that fits that range and distribution.
[0143] Audio devices can be implemented, in particular, as one or more appropriately programmed processors. Different functional blocks can be implemented in separate processors, and / or, for example, in the same processor. Examples of suitable processors are shown below.
[0144] Figure 3 is a block diagram showing an exemplary processor 300 according to an embodiment of the disclosure. The processor 300 can be used to implement one or more processors that implement the aforementioned device or its elements (in particular, including one or more artificial neural networks). The processor 300 can be any suitable processor type, including but not limited to a microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA) / (the FPGA is programmed to form a processor), graphics processing unit (GPU), application-specific integrated circuit (ASIC) / (the ASIC is designed to form a processor), or a combination thereof.
[0145] The processor 300 may include one or more cores 302. A core 302 may include one or more arithmetic logic units (ALUs) 304. In some embodiments, a core 302 may include a floating-point logic unit (FPLU) 306 and / or a digital signal processing unit (DSPU) 308 in addition to or instead of the ALU 304.
[0146] The processor 300 may include one or more registers 312 that are communicatively coupled to the core 302. The registers 312 can be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 312 can be implemented using static memory. The registers can provide data, instructions, and addresses to the core 302.
[0147] In some embodiments, the processor 300 may include one or more levels of cache memory 310 that are communicatively coupled to the core 302. The cache memory 310 can provide computer-readable instructions to the core 302 for execution. The cache memory 310 can provide data for processing by the core 302. In some embodiments, computer-readable instructions may be provided to the cache memory 310 by local memory, for example, local memory attached to an external bus 316. The cache memory 310 can be implemented using any suitable cache memory type, such as static random access memory, dynamic random access memory, and / or any other suitable memory technology.
[0148] The processor 300 may include a controller 314 that can control inputs to the processor 300 from other processors and / or components included in the system, and / or outputs from the processor 300 to other processors and / or components included in the system. The controller 314 can control data paths in the ALU 304, FPLU 306, and / or DSPU 308. The controller 314 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 314 can be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.
[0149] The registers 312 and cache memory 310 can communicate with the controller 314 and core 302 via internal connections 320A, 320B, 320C, and 320D. The internal connections can be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection techniques.
[0150] Inputs and outputs for the processor 300 may be provided via a bus 316 which may include one or more conductive wires. The bus 316 may be communicatively coupled to one or more components of the processor 300, such as a controller 314, a cache memory 310, and / or registers 312. The bus 316 may be coupled to one or more components of the system.
[0151] Bus 316 can be coupled to one or more external memories. The external memory may have read-only memory 332. ROM 332 can be mask ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may have random access memory 333. RAM 333 can be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may have electrically erasable programmable read-only memory (EEPROM) 335. The external memory may have flash memory 334. The external memory may have magnetic storage devices such as disks 336. In some embodiments, the external memory may be included in the system.
[0152] For clarification, the above description will be understood to have illustrated embodiments of the invention with reference to different functional circuits, units, and processors. However, it will be apparent that any appropriate distribution of functions between different functional circuits, units, or processors can be used without departing from the invention. For example, functions shown to be performed by separate processors or controllers can also be performed by the same processor or controller. Thus, references to specific functional units or circuits should be considered only as references to appropriate means for providing the described functions, and not as indicating a strict logical or physical structure or organization.
[0153] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, the present invention can be implemented at least partially as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the present invention can be implemented physically, functionally, and logically in any suitable manner. Indeed, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention can be implemented in a single unit or physically and functionally distributed among different units, circuits, and processors.
[0154] Although the present invention has been described in relation to several embodiments, it is not intended to be limited to any particular form described herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term “comprising” does not preclude the existence of other elements or steps.
[0155] Furthermore, although listed individually, multiple means, elements, circuits, or method steps can be implemented, for example, by a single circuit, unit, or processor. Additionally, individual features may be included in different claims, but these can be advantageously combined in some cases, and inclusion in different claims does not mean that the combination of features is unfeasible and / or unfavorable. Also, including a feature in one category of claims does not imply limitation to that category, but rather indicates that the feature is equally applicable to other claim categories as needed. Furthermore, the order of features in a claim does not imply a specific order in which the features must operate, and in particular, the order of individual steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Furthermore, singular references do not exclude plurals; therefore, references to "a," "an," "first," "second," etc., do not exclude plurals. Reference numerals in a claim are provided merely as clear examples and should not be construed as limiting the scope of the claim in any way.
Claims
1. An audio device for generating audio signals, A receiver configured to receive airflow audio frequency profile data that shows the dependence of the airflow audio frequency profile on airflow velocity parameters, A pose determiner configured to determine the listener pose characteristics of a listener, A frequency response generator configured to determine the airflow frequency response depending on the aforementioned airflow audio frequency profile data, the user's pause characteristics, and the airflow velocity characteristics of the airflow, An audio source configured to provide a first audio signal, An audio component generator configured to generate airflow audio signal components, comprising filtering the first audio signal using the airflow frequency response, An output unit configured to generate the audio signal so as to include the aforementioned airflow audio signal component, An audio device having the following features.
2. The audio device according to claim 1, wherein the frequency response generator is configured to generate the airflow frequency response depending on the airflow velocity of the airflow to the listener.
3. The audio apparatus according to claim 1, wherein the frequency response generator is configured to generate the airflow frequency response depending on the airflow direction of the airflow to the listener.
4. The audio device according to any one of claims 1 to 3, wherein the first audio signal is a noise audio signal.
5. The audio device according to any one of claims 1 to 4, wherein the audio component generator is configured to generate the airflow audio signal component as a stereo airflow audio signal component having a first channel and a second channel, and the output unit is configured to generate the audio signal as a stereo audio signal having a first channel and a second channel.
6. The audio apparatus according to claim 5, wherein the audio source is configured to generate the first audio signal as a stereo audio signal having different signals for the first channel and the second channel.
7. The audio apparatus according to claim 5 or 6, wherein the frequency response generator is configured to generate the airflow frequency response such that it has a first airflow frequency response for the first channel and a second airflow frequency response for the second channel, and the audio component generator is configured to generate a first channel signal component of the airflow audio signal component using the first airflow frequency response for filtering, and generate a second channel signal component of the airflow audio signal component using the second airflow frequency response for filtering.
8. The audio apparatus according to any one of claims 5 to 7, configured to generate the airflow audio signal components such that they have at least partially uncorrelated signals for the first channel and the second channel.
9. The audio device according to any one of claims 5 to 8, configured to adapt the degree of decorrelation between the first channel and the second channel of the stereo airflow audio signal components depending on the airflow direction of the airflow to the listener.
10. The airflow audio frequency profile data includes an index of the first dependence of the first airflow audio frequency profile on the airflow direction parameter and an index of the second dependence of the second airflow audio frequency profile on the airflow velocity parameter. The audio device according to any one of claims 1 to 9, wherein the frequency response generator is configured to generate a first frequency response according to the first dependency and the airflow direction of the airflow to the listener, generate a second frequency response according to the second dependency and the airflow velocity of the airflow to the listener, and generate the frequency response as a combination of the first frequency response and the second frequency response.
11. The audio device according to any one of claims 1 to 10, wherein the audio signal is an audio signal that has been stored.
12. The audio device according to any one of claims 1 to 11, wherein the airflow audio frequency profile data includes an index of relative airflow audio frequency response values for each of several airflow velocity parameter values, and the frequency response generator is configured to determine other relative airflow audio frequency response values for other values of the airflow velocity parameter by interpolation from the several airflow velocity parameter values.
13. The audio device according to any one of claims 1 to 13, wherein the receiver is configured to receive an index of the characteristics of the airflow source of the airflow, and the frequency response generator is configured to determine the airflow velocity characteristics according to the characteristics of the airflow source.
14. The index of the characteristics of the airflow source is that the airflow source is Global airflow source, Omnidirectional airflow source, Point airflow source, and Conical airflow source, The audio device of claim 13, configured to indicate that it is at least one of the following.
15. The receiver is configured to receive the indicator of the characteristics of the airflow source as part of the metadata of the audio bitstream received from the remote source. The audio apparatus according to claim 13 or 14.
16. A method for generating an audio signal, A step of receiving airflow audio frequency profile data that shows the dependence of the airflow audio frequency profile on the airflow velocity parameter, Steps to determine the listener's listener pose characteristics, The steps include determining the airflow frequency response according to the airflow audio frequency profile data, the user pause characteristics, and the airflow velocity characteristics of the airflow, The steps include providing a first audio signal, A step of generating an airflow audio signal component, wherein the generation includes filtering the first audio signal using the airflow frequency response, The steps of generating the audio signal so as to include the airflow audio signal component, A method of having.
17. A computer program that runs on a computer and causes the computer to perform the method described in claim 16.