Generation of audio data signals

The apparatus and method address suboptimal audio quality in VR/AR/MR by correcting audio playback distortions using a binaural filter, enhancing audio quality and efficiency while adapting to listener changes.

JP2026525325APending Publication Date: 2026-07-29KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2024-07-24
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing virtual, augmented, and mixed reality applications often provide suboptimal audio quality due to distortion in rendered audio, failing to accurately reflect variations in signal processing or audio playback, especially on the source and rendering side.

Method used

An apparatus and method for generating an output audio signal that includes a receiver, listener attitude processor, binaural transfer function processor, adapter, and renderer, which determine and apply a binaural filter to correct for audio playback frequency responses and distortions, using metadata and frequency equalization data to enhance audio quality and efficiency.

Benefits of technology

This approach improves audio quality by reducing frequency distortion, enables efficient implementation with low computational resources, and adapts to changes in listener position/orientation, providing enhanced virtual reality experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026525325000001_ABST
    Figure 2026525325000001_ABST
Patent Text Reader

Abstract

The device has a receiver 201 that receives a data signal having an audio signal and metadata including an audio signal, a sound source attitude index for the first audio signal, and frequency equalization data indicating a reference audio playback frequency response for the audio signal. A listener attitude processor 205 determines the listening attitude, and a binaural transfer function processor 207 determines the binaural transfer function according to the listening attitude and the sound source attitude. A playback processor 211 determines the audio rendering frequency response indicating the audio playback frequency response of the audio playback path for the output audio signal. An adapter 101 generates a binaural filter having a frequency response that depends on the combination of the audio rendering frequency response, the reference audio playback frequency response, and the binaural transfer function. A renderer 203 generates the output audio signal using the binaural filter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to generating an audio signal and / or an audio data signal, and more particularly, but not exclusively, to generating such a signal to support, for example, an extended reality application.

Background Art

[0002] The diversity and scope of experiences based on audiovisual content have increased significantly in recent years as new services and methods for utilizing and consuming such content are continuously developed and introduced. In particular, many spatial and interactive services, applications, and experiences have been developed to provide users with a deeper immersive experience.

[0003] Examples of such applications include virtual reality (VR), extended reality (AR), and mixed reality (MR) applications (commonly referred to as extended reality XR applications), which are rapidly becoming mainstream and many solutions are targeted at the consumer market. Some standards are being developed by several standardization bodies. Such standardization activities are actively developing standards for various aspects of VR / AR / MR / XR systems, including streaming, broadcasting, rendering, etc. VR applications tend to provide a user experience corresponding to the user being in a different world / environment / scene, while AR (including decoded reality MR) applications tend to provide a user experience corresponding to the user being in the current environment but with additional information or virtual objects or information added. Thus, VR applications tend to provide a fully immersive synthetically generated world / scene, while AR applications tend to provide a partially synthetic world / scene superimposed on the actual scene where the user physically exists. However, these terms are often used interchangeably and have a high degree of overlap. Hereinafter, the term virtual reality / VR is used to denote both virtual reality and extended reality. VR applications typically provide users with a virtual reality experience, allowing them to move (relatively) freely within the virtual environment and dynamically change their position and what they are looking at. In this field, the term posture is used to refer to position and / or orientation. For example, user posture is used to refer to the user's position and / or orientation. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known from gaming applications, such as in the category of first-person shooter (FPS) games for computers and consoles.

[0004] In addition to visual rendering, most XR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience in which the sound source is perceived to arrive from a position corresponding to the position of the corresponding object in the visual scene (including both currently visible objects and currently invisible objects, such as those behind the user). Thus, the audio scene and the video scene are preferably perceived as consistent and both provide a complete spatial experience.

[0005] Regarding audio, headphone playback using binaural audio rendering technology is widely used. In many scenarios, headphone playback enables users to have a highly immersive and personalized experience. Using head tracking, rendering can be performed in response to the user's head movements, which greatly increases immersion.

[0006] However, while such applications can provide a satisfactory user experience in many embodiments, they tend not to provide an optimal user experience in all situations. In particular, in many situations, suboptimal audio quality may be provided, typically resulting in distortion in the rendered audio compared to the original or desired content audio. Especially in many applications, this approach may not accurately reflect or compensate for variations in signal processing or audio playback on the source or rendering side. [Overview of the project] [Problems that the invention aims to solve]

[0007] Therefore, improved approaches to audio signal distribution and / or rendering and / or processing are advantageous, especially for virtual / augmented / mixed / eXtended reality experiences / applications. In particular, approaches that enable improved behavior, increased flexibility, reduced complexity, easier implementation, improved user experience, improved audio quality, improved adaptation to audio playback capabilities, easier and / or improved adaptation to changes in listener position / orientation (e.g., virtual listener position / orientation), improved virtual reality experience, and / or improved performance and / or behavior are advantageous.

[0008] Therefore, the present invention preferably seeks to mitigate, reduce, or eliminate one or more of the above-mentioned drawbacks, either individually or in any combination. [Means for solving the problem]

[0009] According to one aspect of the present invention, an apparatus for generating an output audio signal is provided, comprising: a receiver configured to receive a data signal having metadata including at least a first audio signal, an attitude index of the first audio signal relative to a sound source, and frequency equalization data indicating a reference audio playback frequency response for the first audio signal; a listener attitude processor configured to determine a listening attitude; a binaural transfer function processor configured to determine a binaural transfer function dependent on the listening attitude and the attitude index relative to a sound source; a processor configured to determine an audio rendering frequency response indicating an audio playback frequency response of an audio playback path for an output audio signal; an adapter configured to generate a binaural filter having a frequency response dependent on a combination of the audio rendering frequency response, the reference audio playback frequency response, and the binaural transfer function; and a renderer configured to generate an output audio signal from a first audio signal, including applying filtering by the binaural filter.

[0010] This approach can provide an improved output audio signal. In many embodiments and scenarios, this approach can provide improved audio quality, for example, by reducing frequency distortion or degradation caused by the processing path. This approach can, for example, enable rendering-side correction of distortion caused by audio playback equipment (e.g., headphones) used by the content-side to adapt / generate the audio signal on the encoding / creation side. For example, a content creator or sound engineer may adjust or generate an audio signal to a desired sound quality and content, and since this is based on the sound heard by the content creator / sound engineer, the frequency response of the headphones (or other audio playback means) used may affect the generated audio. This approach can enable rendering-side correction of such effects while simultaneously correcting the local rendering-side audio playback means.

[0011] Furthermore, this approach can enable highly efficient implementation and operation, particularly low complexity and / or resource usage. This approach can enable single filtering of the received audio signal, specifically to provide correction for both audio playback effects on the content creator side and rendering side, as well as effective binaural filtering. Indeed, this approach can provide such improved effects with virtually no increase in computational resources compared to conventional binaural signal generation. For example, the binaural filter may only be updated slowly to reflect user movement (or sound source movement), and therefore the updates may require significantly fewer computational resources compared to continuous filtering of the audio signal.

[0012] The posture, also referred to as the arrangement, may be position and / or orientation. The listening posture may be the posture in which the (spatial) output audio signal is generated. The output audio signal may be a stereo audio signal.

[0013] The binaural transfer function processor may be configured to determine the frequency response of the binaural transfer function based on the listening posture and the posture index relative to the sound source. Alternatively, the binaural transfer function processor may be configured to determine the binaural transfer function based on the difference between the listening posture and the posture index relative to the sound source.

[0014] The audio playback path for an output audio signal may have or consist of a function that converts the output audio signal from an electrical signal into an audio / sound / acoustic signal. The audio playback path for an output audio signal may have or consist of an audio transducer, such as headphones and / or speakers.

[0015] The audio playback frequency response to the output audio signal may have or consist of the frequency response of an audio transducer, such as a pair of headphones or speakers.

[0016] The reference audio playback frequency response for the first audio signal may have or consist of the frequency response of an audio transducer, such as a pair of headphones or speakers.

[0017] The adapter may be configured to generate a binaural filter having a frequency response that is a combination (particularly in series) of a filter having a frequency response that matches the audio rendering frequency response, a filter having a frequency response that matches the reference audio playback frequency response, and a filter having a frequency response that matches the frequency response of the binaural transfer function.

[0018] The renderer may be configured to generate an output audio signal from a first audio signal, and this generation includes applying binaural filtering to the signal generated from the first audio signal. In some embodiments, the renderer may be configured to generate an output audio signal by applying a binaural filter to the first audio signal. The renderer may be configured to generate an output audio signal based on the filtering of the first audio signal by a binaural filter. The renderer (203) may be configured to generate an output audio signal from a first audio signal, and the generation of the output audio signal includes applying binaural filtering.

[0019] According to an optional feature of the present invention, the reference audio playback frequency response is the target frequency response for the first audio signal.

[0020] This may provide improved and / or facilitated operation or performance in many embodiments. The target frequency response may represent a (desired / target) frequency response for the entire audio path or for example only for the rendering audio path.

[0021] According to an optional feature of the invention, the reference audio playback frequency response is the source frequency response applied in the generation of the first audio signal.

[0022] This may provide improved and / or facilitated operation or performance in many embodiments.

[0023] According to an optional feature of the invention, the frequency equalization data has an indicator of the audio playback device, and the adapter is configured to determine the reference audio playback frequency response as a predetermined frequency response for the audio playback device.

[0024] This may provide improved and / or facilitated operation or performance in many embodiments. The audio playback device may be (or may include) headphones and / or speakers. The predetermined frequency response may be, for example, a stored frequency response for the audio playback device. The predetermined frequency response may be stored, for example, inside or remotely from the audio device. For example, the adapter may be configured to obtain the frequency response from a remote server.

[0025] According to an optional feature of the invention, the frequency equalization data has data describing a finite impulse response (FIR) filter having a frequency response that matches the reference audio playback frequency response.

[0026] This may provide improved and / or facilitated operation or performance in many embodiments. This may enable particularly advantageous representations that can be effectively shown at relatively low data rates. This may further provide improved audio quality by enabling accurate representation.

[0027] According to an optional feature of the invention, the frequency equalization data has data describing a set of sections, each section being a primary section or a secondary section, and the combination having a frequency response that matches a reference audio playback frequency response.

[0028] This may provide improved and / or facilitated operation or performance in many embodiments. This may enable particularly advantageous representations that can be effectively shown at relatively low data rates. This may further provide improved audio quality by enabling accurate representation.

[0029] According to an optional feature of the invention, the frequency equalization data has data describing a set of parallel filters, the set of parallel filters having a frequency response that matches a reference audio playback frequency response.

[0030] This may provide improved and / or facilitated operation or performance in many embodiments. This may enable particularly advantageous representations that can be effectively shown at relatively low data rates. This may further provide improved audio quality by enabling accurate representation. For example, this approach may facilitate processing that reflects perceptual importance, such as a reference audio playback frequency response that reflects, for example, perceptual frequency bands. <x

[0031] According to an optional feature of the present invention, the frequency equalization data includes frequency response identification data, and the adapter is configured to access a remote server to obtain a reference audio playback frequency response based on the frequency response identification data.

[0032] This may provide improved and / or simplified operation or performance in many embodiments.

[0033] According to an optional feature of the present invention, the metadata further includes a reference playback level, and the audio device is configured to adapt the frequency response of a binaural filter according to a first playback level for playback of an output audio signal relative to the reference playback level.

[0034] This may provide improved and / or simplified operation or performance in many embodiments.

[0035] According to an optional feature of the present invention, the metadata includes data indicating the dependency of a reference audio playback frequency response on the playback level, and the adapter is configured to adapt the reference audio playback frequency response according to the playback level.

[0036] This may provide improved and / or simplified operation or performance in many embodiments.

[0037] An apparatus for generating audio data signals may be provided, comprising: a receiver configured to receive at least a first audio signal from a first sound source having a posture in a scene; a playback processor configured to determine a reference audio playback frequency response for the first audio signal; and a data signal generator configured to generate an audio data signal including audio data for the first audio signal, a posture index indicating the posture of the sound source, and frequency equalization data indicating the reference audio playback frequency response.

[0038] According to one aspect of the present invention, a method for generating an output audio signal is provided, the method comprising the steps of: receiving a data signal having metadata including at least a first audio signal, an attitude index of the first audio signal relative to a sound source, and frequency equalization data indicating a reference audio playback frequency response for the first audio signal; determining a listening attitude; determining a binaural transfer function that depends on the listening attitude and the attitude index relative to the sound source; determining an audio rendering frequency response indicating the audio playback frequency response of an audio playback path for the output audio signal; generating a binaural filter having a frequency response that depends on a combination of the audio rendering frequency response, the reference audio playback frequency response, and the binaural transfer function; and generating an output audio signal from a first audio signal, the method comprising applying filtering by the binaural filter.

[0039] A method for generating an audio data signal may be provided, the method comprising: receiving at least a first audio signal from a first sound source having a posture in a scene; determining a reference audio playback frequency response for the first audio signal; and a data signal generator configured to generate an audio data signal having audio data for the first audio signal, a posture index indicating the posture of the sound source, and frequency equalization data indicating the reference audio playback frequency response.

[0040] An audio data signal may be provided that includes at least a first audio signal, metadata including an attitude index of the first audio signal relative to a sound source, and frequency equalization data indicating a reference audio playback frequency response for the first audio signal.

[0041] These and other aspects, features and advantages of the present invention will become apparent from and be described with reference to the embodiments described below.

[0042] Embodiments of the present invention will be described with reference to the drawings, merely as examples. [Brief explanation of the drawing]

[0043] [Figure 1] This is an example of a client-server based virtual reality system. [Figure 2] Examples of elements of an audio rendering device according to several embodiments of the present invention are shown. [Figure 3] Examples of elements of an audio data signal generation device according to several embodiments of the present invention are shown. [Figure 4] Some elements of possible processor configurations for implementing elements of the apparatus according to some embodiments of the present invention are shown. [Modes for carrying out the invention]

[0044] The following explanation focuses on augmented reality applications where audio is rendered according to the user's position within an audio scene to provide an immersive user experience. Typically, audio rendering may accompany image rendering, providing the user with a complete audiovisual experience. However, it will be understood that the approach described may be used in many other applications as well.

[0045] Augmented reality (including virtual augmented reality and mixed reality) experiences, which allow users to move around in a virtual or augmented world, are becoming increasingly popular, and services are being developed to improve such applications. In many such approaches, visual and audio data may be dynamically generated to reflect the user's (or viewer's) current posture.

[0046] In this field, the terms position and orientation are used as general terms for position and / or orientation. For example, a combination of position and orientation of an object, camera, head, or view may be called orientation or position. Thus, an index of position or orientation may have up to six values / components / degrees of freedom, each value / component typically describing an individual characteristic of the position or orientation of the corresponding object. Of course, in many situations, position or orientation may be represented by fewer components, for example, when one or more components are considered fixed or irrelevant (for example, if all objects are considered to be at the same height and horizontal, four components may provide a complete representation of the object's orientation). Hereafter, the term orientation will be used to refer to position and / or orientation that can be represented by 1 to 6 values ​​(corresponding to the maximum possible degrees of freedom).

[0047] Many XR applications are based on pose, which has the most degrees of freedom; that is, the three degrees of freedom for position and orientation each result in a total of six degrees of freedom. The pose may therefore be represented by a set of six values ​​or a vector representing the six degrees of freedom, and thus the pose vector may provide an index of the three-dimensional position and / or three-dimensional orientation. However, it will be understood that in other embodiments, the pose may be represented by fewer values.

[0048] A system or entity that provides the maximum number of degrees of freedom to the observer is typically said to have 6 degrees of freedom (6DoF). Many systems and entities provide only orientation or position, and these are typically known as having 3 degrees of freedom (3DoF).

[0049] Typically, a virtual reality application generates a three-dimensional output in the form of separate view images for the left and right eyes. These can then be presented to the user by appropriate means, typically such as the individual left-eye and right-eye displays of a VR headset. In other embodiments, one or more view images may be displayed, for example, on an automated stereoscopic display, or in fact, in some embodiments, only a single two-dimensional image may be generated (for example, using a conventional two-dimensional display).

[0050] Similarly, an audio representation of a scene may be provided for a given viewer / user / listener posture. The audio scene is typically rendered to provide a spatial experience in which the sound source is perceived to originate from a desired location. Since the sound source can be static within the scene, a change in user posture results in a change in the relative position of the sound source to the user's posture. Accordingly, the spatial perception of the sound source may change to reflect its new position relative to the user. The audio rendering is adjusted as appropriate according to the user's posture.

[0051] Listener pose input may be determined in different ways in different applications. In many embodiments, the user's physical movement may be directly tracked. For example, a camera monitoring the user area may detect and track the user's head (or eyes (gaze tracking)). In many embodiments, the user may wear a VR headset that can be tracked by external and / or internal means. For example, the headset may have accelerometers and gyroscopes that provide information about the movement and rotation of the headset and therefore the head. In some examples, the VR headset may have (e.g., visual) identifiers that transmit signals enabling external sensors to determine the position and orientation of the VR headset.

[0052] In many systems, VR / scene data may be provided from a remote device or server. For example, a remote server may generate audio data representing an audio scene, transmitting audio signals corresponding to audio components / objects / channels, or other audio elements corresponding to different sound sources within the audio scene, along with positional information indicating their locations (which may change dynamically for moving objects, for example). The audio signals / elements may include elements associated with specific locations, but may also include elements for more dispersed or diffuse sound sources. For example, audio elements representing generic (unlocalized) background sounds, ambient sounds, diffuse reverberation, etc., may be provided.

[0053] A local VR device can properly render audio elements by applying appropriate binaural processing that specifically reflects the relative position of the sound source of the audio component.

[0054] Similarly, a remote device may generate visual / video data representing a visual audio scene, and transmit visual scene components / objects / signals, or other visual elements corresponding to different objects within the visual scene, along with positional information indicating their locations (which may change dynamically for moving objects, for example). Visual items may include elements associated with specific locations, or they may include video items for more distributed sources.

[0055] In some embodiments, visual items may be provided as individual and separate items, such as descriptions of individual scene objects (e.g., dimensions, texture, opacity, reflectivity, etc.). Alternatively or additionally, visual items may be represented as part of an overall model of the scene, including, for example, descriptions of different objects and their relationships to one another. For VR services, the central server may, in some embodiments, generate audiovisual data representing a three-dimensional scene, specifically representing audio with multiple audio signals representing sound sources within the scene, which can be rendered by local clients / devices.

[0056] Figure 1 shows an example of a VR / XR system in which a central server 101 interacts with a large number of remote clients 103 via a network 105, such as the internet. The central server 101 may be configured to support a potentially large number of remote clients 103 simultaneously. Such an approach can offer improved trade-offs in many scenarios, for example, between complexity and resource requirements for different devices, communication requirements, etc. For instance, scene data may only need to be transmitted once or relatively rarely, and the local rendering device (remote client 103) processes the scene data locally to receive the viewer's pose and render audio and / or video to reflect changes in the viewer's pose. This approach can provide an efficient system and a compelling user experience. For example, it can allow scene data to be centrally stored, generated, and maintained while significantly reducing the required communication bandwidth and providing a low-latency real-time experience. This could be suitable, for example, for applications where a VR experience is delivered to multiple remote devices.

[0057] Figure 2 shows elements of an output audio signal generating device, also referred to below as an audio rendering device, which can generate improved audio signals in many applications and scenarios. In particular, the audio rendering device may provide improved rendering for many VR applications, and the audio rendering device may be specifically configured to perform audio processing and rendering for the VR client 103 in Figure 1. Figure 3 shows an example of a device, also referred to below as an audio encoding device, which generates audio data signals representing audio for one or more sound sources in a scene. In particular, the audio encoding device may provide improved representation of sound sources or scenes for many VR applications, and the audio encoding device may be specifically configured to perform audio processing and functions for the VR server 101 in Figure 1.

[0058] The audio device in Figure 2 is configured to render audio for a three-dimensional scene to provide a three-dimensional perception of the scene. While the specific description will focus on audio rendering, it will be understood that in many embodiments this can be complemented by visual rendering of the scene. Specifically, view images may be generated and presented to the user in many embodiments.

[0059] The audio rendering device has a first receiver 201 configured to receive audio items from a local or remote source. In a specific example, the first receiver 201 receives data describing audio items from a server 101. The first receiver 201 may be configured to receive data describing a virtual scene, specifically an audio scene. The data may include data providing a visual description of the scene and data providing an audio description of the scene. Thus, audio scene descriptions and visual scene descriptions may be provided by the received data.

[0060] An audio signal / item can be encoded audio data, such as an encoded audio signal. An audio signal may also be different types of audio elements, including different types of elements and components. In fact, in many embodiments, the first receiver 201 may receive audio data that defines different types / formats of audio. For example, audio data may include audio represented by audio channel signals, individual audio objects, scene-based audio such as higher-order ambisonics (HOA), etc. For example, audio may be represented as encoded audio for a given audio component to be rendered.

[0061] The received data signal further includes data indicating the orientation (position and / or orientation) of the sound source from which the audio signal is provided. In many embodiments, at least a portion of the audio signal is linked to data describing the location of the sound source in the scene. In some cases, the data may alternatively, or typically additionally, include data defining the orientation of the sound source (e.g., to allow for non-omnidirectional sound sources).

[0062] The received signal may have metadata that specifically includes the position and / or orientation of an audio item, specifically the position and / or orientation of a sound source and / or pose data indicating a visual scene object or element. The pose data may include, for example, absolute position and / or orientation defining the position of each or at least some of the sound sources.

[0063] The first receiver 201 is coupled to a renderer 203 which proceeds to render an audio scene based on the received data describing the audio items. In the case of encoded data, the renderer 203 may be configured to decode the audio data (or, in some embodiments, the decoding may be performed by the first receiver 201).

[0064] Renderer 203 is configured to render an audio scene by generating an audio signal based on received audio data for a sound source. In this example, renderer 203 is a binaural audio renderer that generates binaural audio signals for the user's left and right ears. The binaural audio signals are generated to provide a desired spatial experience and are typically played back by headphones or earphones, which may be part of a headset worn by the user (the headset typically also has left and right eye displays).

[0065] Therefore, in many embodiments, audio rendering by renderer 203 is a binaural rendering process that uses an appropriate binaural transfer function to provide the desired spatial effect to a user wearing headphones. For example, renderer 203 may be configured to use binaural processing to generate audio components that are perceived to be arriving from a particular location.

[0066] Binaural processing is known to be used to provide a spatial experience by using individual signals to the listener's ears to create a virtual arrangement of sound sources. With proper binaural rendering, the signals required by the eardrum for the listener to perceive sound from any desired direction can be calculated, and these signals can be rendered to produce the desired effect. These signals are reproduced at the eardrum using either headphones or a crosstalk cancellation method (suitable for rendering to closely spaced speakers). Binaural rendering can be considered an approach that generates signals to the listener's ears, thereby deceiving the human auditory system into perceiving that the sound is coming from a desired location.

[0067] Binaural rendering is based on binaural transfer functions, which differ from person to person due to the acoustic properties of reflective surfaces such as the head, ears, and shoulders. Therefore, binaural transfer functions may be personalized for the optimal binaural experience. For example, binaural filters can be used to create binaural recordings that simulate multiple sound sources located in different places. This can be achieved by convolving each sound source with, for example, a pair of head-related impulse responses (HRIRs) corresponding to the location of the sound source.

[0068] A well-known method for determining binaural transfer functions is binaural recording. This method uses a dedicated microphone placement to record sounds intended for playback using headphones. Recording is done either by placing microphones in the subject's ear canal or by using a dummy head with a built-in microphone and a bust that includes the auricle (outer ear). The use of such a dummy head with an auricle provides a very similar spatial impression as if the person listening to the recording had been present during the recording.

[0069] For example, an appropriate binaural filter can be determined by measuring the response from a sound source located at a specific location in 2D or 3D space to a microphone placed inside or near a human ear. Based on such measurements, a binaural filter can be generated that reflects the acoustic transfer function to the user's ear. Binaural filters can be used to create binaural recordings that simulate multiple sound sources at various locations. This can be achieved, for example, by convolving each sound source with a pair of measured impulse responses for a desired location of the sound source. To create the illusion that the sound source is moving around the listener, multiple binaural filters are typically required at a specific spatial resolution, e.g., 10 degrees.

[0070] The head-related binaural transfer function may be expressed, for example, as a head-related impulse response (HRIR), or equivalently as a head-related transfer function (HRTF), binaural room impulse response (BRIR), or binaural room transfer function (BRTF). The (estimated or assumed) transfer function from a given position to the listener's ear (or eardrum) may be expressed, for example, in the frequency domain, in which case it is typically referred to as HRTF or BRTF, or in the time domain, in which case it is typically referred to as HRIR or BRIR. In some scenarios, the head-related binaural transfer function is determined to include the acoustic environment, particularly the manner or characteristics of the room in which the measurement is taken, while in other examples, only user characteristics are considered. Examples of the first type of function are BRIR and BRTF.

[0071] The renderer 203 may be configured to apply binaural processing to multiple audio signals / sources individually, and the results may be combined into a single binaural output audio signal representing an audio scene with multiple sound sources positioned appropriately within a soundstage.

[0072] The renderer 203 is configured to generate an output audio signal in reliance on a binaural filter, specifically by generating a (first / input / received) audio signal, the generation of which includes applying filtering by / using the binaural filter. In many embodiments, the binaural filter may be applied directly to the (first / input / received) audio signal, and / or the binaural output signal may be generated directly by the binaural filter. However, it will be understood that in many embodiments, other signal processing and calculations may be performed as part of the rendering / generation of the binaural output audio signal.

[0073] For example, in many embodiments, the input audio signal may undergo signal level normalization, and / or an equalization filter may be applied to the input signal level, or before the audio signal is fed to the binaural filter. Often, gain adjustment or equalization filters may be applied alternatively or additionally to the binaural audio signal produced by the binaural filter.

[0074] The generation of the output audio signal may include applying a binaural filter to the signal generated from the (first / input / received) audio signal. This may often be the audio signal itself, or a signal that has been filtered, amplified, attenuated, digitized, resampled, compressed, expanded, etc., before being fed to the binaural filter. Similarly, the binaural output audio signal may be the output of the binaural filter itself, which may be a signal that has been filtered, amplified, attenuated, digitized, resampled, compressed, expanded, etc.

[0075] Specifically, the renderer 203 may have a signal path that receives a (first input / received) audio signal and generates a binaural output audio signal, the signal path including a binaural filter. In many embodiments, the renderer 203 may include multiple such signal paths that receive different input / received audio signals and apply different binaural filters. The resulting binaural audio signals can be combined into a single output audio signal.

[0076] As an example, in some embodiments, the input audio signal may first be resampled to a different sample rate, and then the resampled signal may be subjected to dynamic range compression. The resulting signal may be filtered by a variable frequency filter (e.g., a low-pass, band-pass, and / or high-pass filter), and the resulting signal may have scaling / gain compensation. The resulting signal may then be applied to a binaural filter to generate a binaural audio signal. For example, it is understood that any subset of these operations may be applied, and may be applied in different / any order. In other embodiments, other signal processing operations may be applied additionally or alternatively before filtering by the binaural filter. Alternatively or additionally, signal processing operations may be applied to the output of the binaural filter as part of generating a binaural output audio signal. As an example, in some embodiments, the binaural filter output signal may first have scaling / gain compensation, followed by filtering by a variable frequency filter (e.g., a low-pass, band-pass, and / or high-pass filter), and the resulting signal may have dynamic range expansion before being resampled, for example. The resulting signal may then be output as a binaural output audio signal. For example, any subset of these operations may be applied, and may be applied in different / any order. In other embodiments, other signal processing operations may be applied additionally or alternatively after filtering by the binaural filter. It is also understood that different operations may be applied before and after filtering, or in fact, operations may be applied only before or after binaural filtering (signal operations / processing do not need to be symmetrical or complementary before and after binaural filtering).

[0077] The audio rendering device further includes a listener posture processor 205 configured to determine the listening posture from which the output audio signal is generated. The listener posture may, accordingly, correspond to the user / listener's position in the scene.

[0078] The listener's posture may be determined specifically in response to sensor input, for example, from a suitable sensor that is part of a headset. It is understood that many suitable algorithms are known to those skilled in the art, and for the sake of brevity, this will not be described in further detail.

[0079] The renderer 203 of the apparatus in Figure 2 is configured to generate an output audio signal by a binaural rendering process, for example, so that spatial cues are provided from rendering using headphones. The renderer may filter the audio signal with appropriate binaural filters corresponding to the left and right ears, respectively, to generate a binaural output signal that provides a spatial audio experience when presented to the user's ears.

[0080] The listener attitude processor 205 receives the listener attitude from the listener attitude processor 205 and is coupled to a binaural transfer function processor 207 configured to determine a binaural transfer function for this listener attitude and a given sound source attitude. The binaural transfer function processor 207 is configured to determine an appropriate binaural transfer function for a given audio signal and sound source, depending on the relative position / attitude of the sound source with respect to the listener attitude.

[0081] The binaural transfer function processor 207 may have a store / service / database containing a number of binaural transfer functions for several different locations and possibly orientations, each binaural transfer function providing information on how the audio signal should be processed / filtered so that it is perceived to originate from that location.

[0082] The binaural transfer function processor 207 may select and retrieve a stored binaural transfer function that best matches a given audio signal perceived to originate from a given position relative to the user's head, i.e., best matches the sound source position relative to the listener's position. In some embodiments, the binaural transfer function processor 207 may be configured to generate a binaural transfer function by interpolating between a plurality of close stored binaural transfer functions. The selected binaural transfer function can then be applied to the audio signal of the audio element to generate audio signals for the left and right ears.

[0083] However, in the apparatus shown in Figure 2, the extracted binaural transfer function is not used directly by the renderer 203, but is adapted before rendering. The binaural transfer function processor 207 is coupled to an adapter 209 configured to adapt the binaural transfer function, specifically, to generate a binaural filter that has a frequency response corresponding to the frequency response of the binaural transfer function, but which is adapted / modified. The binaural filter is then supplied to the renderer 203 and used to perform binaural rendering. Typically, the binaural filter is a stereo filter that has both a transfer function / impulse response / frequency response for generating the left ear signal and a transfer function / impulse response / frequency response for generating the right ear signal. (Similarly, the adapter 207 may generate two filters used by the renderer 203 to generate the left ear signal and the right ear signal, respectively. In such a case, the process described can be applied to both binaural filters.)

[0084] The renderer can generate an output stereo signal by specifically filtering the audio signal with a binaural filter. The generated output stereo signal, in the form of left and right ear signals, is suitable for headphone rendering and may be amplified to generate a drive signal supplied to the user's headset. The user perceives the audio signal / source as if it were originating from a desired location.

[0085] In some embodiments, it will be understood that the audio signal may be processed to add, for example, acoustic environmental effects. For example, the audio signal may be processed to add reverberation or, for example, decorrelation / diffusivity. In many embodiments, this processing may be performed on the generated binaural signal rather than directly on the audio element signal.

[0086] Therefore, the renderer 203 may be configured to generate an output audio (stereo) signal such that a given sound source is rendered and a user wearing headphones perceives the audio as being received from a desired position. This approach typically applies to multiple sound sources and signals that are coupled by the renderer 203 to generate an output signal.

[0087] For example, many algorithms and approaches for rendering spatial audio, particularly binaural rendering, using headphones, are known to those skilled in the art, and it will be understood that any suitable approach may be used without impairing the present invention.

[0088] Specifically, the adapter 209 may be configured to correct / adapt the binaural transfer function to the audio playback characteristics on both the rendering side and the source side.

[0089] The apparatus in Figure 2 specifically includes a playback processor 211 configured to determine an audio rendering frequency response that reflects the audio playback frequency response of the audio playback path to the output audio signal. Specifically, the audio rendering frequency response may reflect the frequency response of a pair of headphones used to render the generated output audio signal to the user. More generally, the audio rendering frequency response is typically the frequency response of an audio transducer configured to convert an electrical signal into an acoustic signal. The audio transducer may be, for example, headphones and / or speakers.

[0090] In this approach, the received data signal further includes frequency equalization data that indicates the reference audio playback frequency response for the generated audio. The received data may also include frequency equalization data that indicates the desired frequency spectrum correction, target frequency response, and / or frequency response of the audio playback device used, for example, in the generation or post-processing of the audio of the sound source on the content side.

[0091] The playback processor 211 is supplied with an audio rendering frequency response and is coupled to an adapter 209 which is further supplied with a reference audio playback frequency response from the receiver 201. The adapter 209 is supplied to the renderer 203 and is configured to adapt a binaural transfer function according to both the audio rendering frequency response and the reference audio playback frequency response to generate a binaural filter applied to the audio signal to generate a binaural audio signal.

[0092] The precise adaptation depends on the specific preferences and requirements of each individual embodiment, for example, on the format and specific characteristics of the provided audio rendering frequency response and reference audio playback frequency response.

[0093] In many embodiments, the reference audio playback frequency response may represent the frequency response of an audio playback device, such as headphones, used by the content creator to generate (e.g., equalize) the audio signal. Similarly, the reference audio playback frequency response may characterize the frequency response of a set of headphones or speakers used to render the output stereo signal to the user. In such cases, the adapter 209 may be configured to compensate for the frequency response of the extracted binaural transfer function to compensate for the frequency response of the headphones / speakers. For example, the gain of the binaural transfer function for a first frequency may be multiplied by the (normalized) gain of the reference audio playback frequency response and divided by the (normalized) gain of the audio rendering frequency response. This is repeated for all frequency values, resulting in the frequency response of the binaural filter. This frequency response therefore represents an audio playback compensation binaural transfer function / filter that can provide improved audio quality when applied to an audio signal.

[0094] In this approach, the rendering side may adapt / correct frequency distortion on both the encoder and decoder sides accordingly. Correction and audio quality may be supported or controlled by the content side, thereby allowing content creators to influence rendering characteristics. Furthermore, adaptation of characteristics on both the content creator and rendering sides is combined synergistically with binaural processing. In fact, pre-processing with binaural filters (rather than applying filters as post-processing of the output stereo signal, for example) provides highly efficient processing and computation that can deliver improved audio quality for low implementation complexity and resource usage. In particular, in this approach, only a single filtering of the audio signal is required to achieve a beneficial effect.

[0095] The encoder in Figure 3 has an encoder receiver 301 configured to receive at least a first audio signal from a first sound source having a given orientation in an (audio) scene. Typically, the receiver receives multiple sound sources that are in different locations and / or have different orientations, and / or sound sources that do not have a specific location and / or orientation, such as a sound source representing diffuse reverberation audio in an audio scene.

[0096] The encoder further includes an encoder playback processor 303 configured to determine a reference audio playback frequency response for a first audio signal. The reference audio playback frequency response may depend on the frequency response of an audio playback path or device used on the encoder / content side. For example, the reference audio playback frequency response may be determined based on the frequency response of a pair of headphones or speakers used by the content creator to generate the audio signal. In many embodiments, the reference audio playback frequency response may be, or include, the frequency response of an audio transducer configured to convert an electrical signal into an acoustic signal. The audio transducer may be, for example, headphones and / or speakers.

[0097] The encoder playback processor 303 may include a user interface that allows, for example, a user or content creator to manually input data describing the frequency response of, for example, headphones or speakers. In another example, in some embodiments, the encoder playback processor 303 may be configured to control a test system to measure the frequency response of a speaker. For example, the encoder playback processor 303 may have a variable frequency tone generator that can be coupled to an external audio playback device, such as headphones or speakers. A frequency sweep may be performed, and a microphone appropriately positioned for a given audio playback device may generate a microphone signal that captures the generated audio. The levels of the captured audio for different frequencies may be used to determine the frequency response of the playback device / headphones. In another example, the playback device may have a function to identify, for example, a brand, model, type, etc., and the encoder playback processor 303 may be configured to retrieve, for example, a stored nominal frequency response for that device from a remote database. For example, Bluetooth® headphones may provide identification data that allows the encoder playback processor 303 to access a remote database to retrieve, for example, the frequency response for the headphones.

[0098] The encoder playback processor 303 and encoder receiver 301 are coupled to a data signal generator 305 configured to generate an output audio data signal that can be transmitted to or distributed to the decoder (and optionally other decoder devices) shown in Figure 1. The data signal generator 305 may encode the audio signal to generate encoded audio data to be included in the audio data signal. Typically, other sound sources may be encoded to provide a combined audio representation of an audio scene. The data signal generator 305 may further receive sound source attitude data of the first audio signal, and typically other sound source attitude data, and may be configured to include sound source attitude data in the output audio signal.

[0099] The data signal generator 305 may further receive frequency equalization data that represents a reference audio playback frequency response, such as the frequency response of a playback device used by the content creator. The data signal generator 305 may then include data that represents and describes the reference audio playback frequency response in the output audio data signal.

[0100] In some embodiments, the reference audio playback frequency response may be a target frequency response for the playback of the first audio signal on the decoder side. For example, the reference audio playback frequency response may represent a desired frequency response for the entire audio path, including the frequency response of the playback path in the rendering device.

[0101] For example, a content creator may manually process an audio signal to adapt it to the desired sound, and such processing may typically include frequency equalization and filtering; that is, the content creator may adapt the frequency spectrum of the audio signal to produce the desired sound. Furthermore, it may be desirable for the overall frequency response of the entire signal processing to be flat so that the rendered sound corresponds more directly to the audio manually produced by the content creator. However, the content creator uses an audio playback device such as headphones or speakers, which inherently has a frequency response that tends not to be flat. Therefore, the audio heard by the content creator may be colored by the playback. In the approach described, a reference audio playback frequency response may be used to compensate for such distortion by indicating a target playback frequency response that corrects the frequency response of the content-side playback path.

[0102] Specifically, the target reference audio playback frequency response may be expressed as a frequency response that corrects the determined content-side playback device frequency response, for example, by having an inverse frequency response. For example, if the content-side playback frequency response has increased attenuation, the target reference audio playback frequency response may be determined to have a corresponding gain. In particular, the reference audio playback frequency response may be determined as the inverse / reciprocal of the determined playback frequency response for the content-side playback device. Upon receiving such a target reference audio playback frequency response, the rendering device in Figure 2 may proceed to perform frequency equalization, which includes applying the corresponding reference audio playback frequency response. The rendering device may further include corrections to the audio rendering frequency response. Thus, specifically, the binaural transfer function / HRTF may be modified to provide the desired target response while correcting the local audio playback. As a specific example, the frequency response of the binaural filter used by renderer 203 may be determined as follows: H RND (f) = (H T (f) / H R (f))H B (f)

[0103] Here, H T (f) shows the received target reference audio playback frequency response, H R (f) shows the local audio rendering frequency response, H B (f) shows the determined binaural transfer function (specifically HRTF) with respect to the sound source position relative to the listening position.

[0104] In some embodiments, the reference audio playback frequency response is the source frequency response applied in the generation of the first audio signal. Specifically, the reference audio playback frequency response may provide an index of the frequency response of, for example, headphones or speakers used by, for example, a sound engineer or content creator on the content side. In such cases, the adapter 209 may be configured to adapt the frequency response of a binaural transfer function to compensate for the filter response. For example, the frequency response of a binaural filter used by the renderer 203 may be determined using the same approach as described above (keeping in mind that the effect of the headphone / speaker frequency response on the transmitted signal is the inverse of the frequency response, and therefore the headphone frequency response is also the desired target frequency response for compensating the headphones).

[0105] Frequency equalization data may use different approaches to represent the reference audio playback frequency response.

[0106] In some embodiments, the frequency equalization data may include data describing a finite impulse response (FIR) filter having a frequency response that matches a reference audio playback frequency response. The reference audio playback frequency response may be specifically described as an FIR filter having a given frequency response. The FIR filter may be described in the frequency domain and / or time domain. For example, the frequency equalization data may include a set of coefficients for an FIR filter, specifically, the coefficients may correspond to a suitable sample of impulse responses for an FIR filter having a frequency response that represents / matches the reference audio playback frequency response.

[0107] Such an approach can provide a highly efficient representation that is easy for the renderer to process. It may also be generated with relatively low complexity on the content side. Furthermore, FIR filters can effectively provide both amplitude and phase equalization, and these can be performed independently of each other, which is an advantage.

[0108] In some embodiments, the frequency equalization data may include data describing the reference audio playback frequency response as an infinite impulse response IIR filter. This may provide a more compressed representation and / or a less complex implementation in some scenarios.

[0109] In some embodiments, the frequency equalization data may include data describing a combination of sets of sections, the combination having a frequency response that matches a reference audio playback frequency response. Each section may be a primary section (FOS) or a secondary section (SOS). Specifically, the combination may be a series of FOS and SOS sections. Thus, each section may correspond to a primary or secondary filter, and the combination may correspond to a filter formed by a series of these primary and secondary filters. The combined frequency response of the series of filters may be considered by the renderer 203 as the reference audio playback frequency response.

[0110] Such an approach can be particularly advantageous in many embodiments and scenarios, and can enable computationally efficient implementations and functions.

[0111] In some embodiments, the frequency equalization data may include data describing a set of parallel filters, and the reference audio playback frequency response may be given as the combined frequency response of these parallel filters. For example, a set of parallel bandpass filters may be provided, each bandpass filter providing a frequency response for a given frequency band, and the other bandpass bands being added together to the overall frequency response. Specifically, the parallel filters may be provided, for example, one octave apart. In some embodiments, parallel filters with different bandwidths may be provided, and the advantage of such an approach is that the filter set may be designed, for example, on a perceptually motivated logarithmic frequency scale, which is an advantage over FIR filters designed on a linear frequency scale. For example, the filters and filter bandwidths may be determined to correspond to, for example, logarithmic, octave, 1 / 3 octave, equivalent square bandwidth (ERB), Burk scale, etc.

[0112] As a specific example, an audio data signal may have frequency equalization data in the form of an EQ (Equalization) field specifier that describes, for example, the type and method of frequency equalization. Such an EQ field may indicate the method of frequency equalization and the equalization filters that can or may be used. The frequency equalization data may include, for example, data indices for one or more of the following examples of frequency equalization indices.

[0113] A reference audio playback frequency response expressed as the coefficients of a specified FIR filter at a particular sample rate.

[0114] A reference audio playback frequency response, expressed as one or more series of first- and second-order sections (FOS & SOS) specified either by coefficients and sample rate or by descriptive parameters. For example, a second-order peak EQ filter with center frequency, gain, and quality coefficient may be provided, followed by a first-order shelving filter with a given gain and center frequency, etc.

[0115] The reference audio playback frequency response, expressed as the sum of parallel filters.

[0116] It will be understood that other approaches may be used depending on the requirements and preferences of each individual embodiment.

[0117] In some embodiments, the audio rendering apparatus may receive frequency equalization data having frequency response identification data. This identification data may not fully describe the frequency response, but may provide identification information of the frequency response, such as a specific name and / or number.

[0118] In some embodiments, the audio rendering device may be configured to extract a corresponding reference audio playback frequency response from a locally stored set of reference audio playback frequency responses. For example, during manufacturing, multiple reference audio playback frequency responses may be stored in local memory along with an associated identifier for each reference audio playback frequency response. During operation, the rendering device may receive an audio data signal along with the identifier of the reference audio playback frequency response, which may then proceed to access memory to retrieve the reference audio playback frequency response stored for the received identifier. It may then proceed to use this reference audio playback frequency response to modify the binaural transfer function.

[0119] In some embodiments, the adapter 209 may be configured to access a remote server to obtain a reference audio playback frequency response that matches a frequency response identifier received in the data audio signal. For example, the adapter 209 may have a network interface that can connect to the internet to access a remote server with a request containing a frequency response identifier. The server may extract the corresponding reference audio playback frequency response and send it back to the rendering device.

[0120] Such an approach can be highly advantageous in many embodiments because it allows a single (or a small number of) central servers to serve a large number of rendering devices. This can greatly facilitate the distribution and updating of the reference audio playback frequency response.

[0121] In many embodiments, the identification information may specifically be the identification information of an audio playback device, and more specifically, the identification information of an audio playback device that can be used in content-side operation. For example, this identification information may identify headphones, speakers, amplifiers, etc., that can be used as part of the content-side setup. In many examples, the identification information may simply be data describing the manufacturer and model of the device. The adapter 209 may access a remote database / server to obtain the frequency response for that device, and the binaural transfer function may be modified to compensate for this response.

[0122] As a specific example, if the brand and type of headphones are specified, the adapter 209 may access a (third-party) database containing preferred equalization filter frequency responses or headphone frequency responses, which may be used in this case to provide a desired target frequency response, specifically such as a flat frequency response.

[0123] In some embodiments, the encoder may be further configured to generate an audio data signal that includes metadata providing an index of a reference playback level. The reference playback level may be a nominal value or may be dynamically determined to directly indicate the playback level applied by, for example, a content creator or sound engineer manually adjusting the frequency spectrum of a first audio signal. For example, a sound engineer may manually perform frequency equalization while listening to a stereo audio signal. The playback level of the stereo signal may be determined by measurement or (for example, from a volume setting by a sound engineer), and data indicating this playback level may be included in the audio data signal.

[0124] The playback level may be provided, for example, as a relative or absolute value. For example, in some embodiments, a general audio level within a predetermined range (e.g., 0 to 11) may be provided. In other embodiments, the reference playback level may be provided as an indicator of sound pressure level, or SPL.

[0125] The renderer may receive such audio data signals and extract a reference playback level from the metadata of the audio data signals, and the adapter 209 may be configured to apply a binaural filter according to the reference playback level and the playback level of the generated output audio signal.

[0126] For example, the current level setting (e.g., volume setting) for rendering / playback of the output stereo signal may be provided to the adapter 209, and depending on this and the reference playback level, the adapter 209 may modify, for example, the frequency response adaptation applied to the binaural transfer function. For example, if the reference playback level indicates a low level and the current volume setting is set to a low playback sound level, the adapter 209 does not provide additional sound level frequency adaptation. However, if the volume setting is currently set to correspond to a high playback sound level, the adapter 209 may proceed to perform additional frequency response adjustments to attenuate low and high frequencies (e.g., to prevent hearing damage). Alternatively, if the reference playback level indicates a high sound level, the adapter 209 may be configured not to include additional frequency response adaptations for high volume settings, but may increase the high and low frequency gains for low volume settings (e.g., to provide a loudness effect).

[0127] Such an approach can, in many embodiments, provide an improved listening experience.

[0128] In some embodiments, the metadata may include data indicating the dependency of the reference audio playback frequency response on the playback level. Specifically, the metadata may not only provide data describing the reference audio playback frequency response, but also data describing how it may change for different playback sound levels.

[0129] As an example of low complexity, the metadata may include different reference audio playback frequency responses for different sound playback levels. For example, the metadata may include a first reference audio playback frequency response for a first (e.g., low) reference playback level and a second reference audio playback frequency response for a second (e.g., high) reference playback level. As another example, the metadata may include a reference audio playback frequency response for which the gain for one or more frequencies is provided as a function of the sound / playback level.

[0130] The adapter 209 may adapt a reference audio playback frequency response used to modify the binaural transfer function depending on the current playback level of the output audio signal. For example, depending on the current volume setting (e.g., whether it is "low-frequency" or "high-frequency," e.g., below or above a setting of 5), the adapter 209 may select from among different reference audio playback frequency responses provided in the metadata. In another example, in some embodiments, the actual current acoustic playback level may be measured (e.g., by a microphone), and the adapter 209 may proceed to determine various gains of the reference audio playback frequency response by evaluating a function provided in the metadata for a given measured acoustic playback level.

[0131] Such an approach can provide a highly advantageous user experience in many embodiments, and in particular, it can enable content-side control or assistance when adapting the rendering-side system to provide desired audio playback.

[0132] Figure 4 is a block diagram showing an exemplary processor 400 according to an embodiment of the present disclosure. The processor 400 may be used to implement one or more processors that implement the aforementioned devices or elements thereof. The processor 400 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA) programmed to form a processor, a graphics processing unit (GPU), an application-specific integrated circuit (ASIC) designed to form a processor, or a combination thereof.

[0133] The processor 400 may include one or more cores 402. A core 402 may include one or more arithmetic logic units (ALUs) 404. In some embodiments, a core 402 may include a digital signal processing unit (DSPU) 408 in addition to or instead of a floating-point logic unit (FPLU) 406 and / or ALU 404.

[0134] The processor 400 may include one or more registers 312 that are communicatively coupled to the core 402. The registers 412 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 412 may be implemented using static memory. The registers may provide data, instructions, and addresses to the core 402.

[0135] In some embodiments, the processor 400 may include one or more levels of cache memory 410 communicably coupled to the core 402. The cache memory 410 may provide computer-readable instructions to the core 402 for execution. The cache memory 410 may provide data for processing by the core 402. In some embodiments, computer-readable instructions may be provided to the cache memory 410 by local memory, for example, local memory attached to an external bus 416. The cache memory 410 may be implemented in metal-oxide-semiconductor (MOS) memory of any suitable cache memory type, such as static random-access memory (SRAM), dynamic random-access memory (DRAM), and / or any other suitable memory technology.

[0136] The processor 400 may include a controller 414 that can control inputs to the processor 400 from other processors and / or components included in the system and / or outputs from the processor 400 to other processors and / or components included in the system. The controller 414 can control data paths in the ALU 404, FPLU 406, and / or DSPU 408. The controller 414 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 414 can be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.

[0137] The registers 412 and cache 410 can communicate with the controller 414 and core 402 via internal connections 420A, 420B, 420C, and 420D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection techniques.

[0138] Inputs and outputs to the processor 400 may be provided via a bus 416 which may include one or more conductive wires. The bus 416 may be communicatively coupled to one or more components of the processor 400, such as a controller 414, a cache 410, and / or registers 412. The bus 416 may be coupled to one or more components of the system.

[0139] The bus 416 may be coupled to one or more external memories. The external memory may have read-only memory 432. ROM 432 may be a mask ROM, an electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may have random access memory 433. RAM 433 may be static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memory may have electrically erasable programmable read-only memory (EEPROM) 435. The external memory may have flash memory 434. The external memory may have a magnetic storage device such as a disk 436. In some embodiments, the external memory may be included in the system.

[0140] For clarification, the above description will be understood to have illustrated embodiments of the invention with reference to different functional circuits, units, and processors. However, it will be apparent that any appropriate distribution of functions between different functional circuits, units, or processors can be used without departing from the invention. For example, functions shown to be performed by separate processors or controllers can also be performed by the same processor or controller. Thus, references to specific functional units or circuits should be considered only as references to appropriate means for providing the described functions, and not as indicating a strict logical or physical structure or organization.

[0141] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, the present invention can be implemented at least partially as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the present invention can be implemented physically, functionally, and logically in any suitable manner. In fact, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention can be implemented in a single unit, or physically and functionally distributed among different units, circuits, and processors.

[0142] Although the present invention has been described in relation to several embodiments, it is not intended to be limited to any particular form described herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term “comprising” does not preclude the existence of other elements or steps.

[0143] Furthermore, although listed individually, multiple means, elements, circuits, or method steps may be implemented, for example, by a single circuit, unit, or processor. In addition, individual features may be included in different claims, but these may be advantageously combined as they may be, and inclusion in different claims does not mean that the combination of features is unfeasible and / or unfavorable. Also, the inclusion of a feature in one category of claims does not mean that it is limited to that category, but rather that the feature is equally applicable to other claim categories as needed. Furthermore, the order of features in a claim does not mean that the feature must operate in a specific order, and in particular, the order of individual steps in a method claim does not mean that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Furthermore, a singular reference does not exclude the plural. Thus, references to "a," "an," "first," "second," etc., do not exclude the plural. Reference numerals in a claim are provided merely as clear examples and should not be construed as limiting the scope of the claim in any way.

[0144] Generally, an apparatus / method for generating an output audio signal, an apparatus / method for generating an audio data signal, a data signal, and a computer program are shown by the following embodiments. Embodiments:

[0145] 1. A device for generating an output audio signal, At least a first audio signal, and, The attitude indicator of the sound source of the first audio signal, and Frequency equalization data showing the reference audio playback frequency response for the first audio signal, A receiver (201) configured to receive a data signal having metadata including, A listener posture processor (205) configured to determine listening posture, A binaural transfer function processor (207) configured to determine a binaural transfer function depending on the listening posture and the posture index of the sound source, A processor (211) configured to determine an audio rendering frequency response that shows the audio playback frequency response of the audio playback path for the output audio signal, An adapter (209) configured to generate a binaural filter having a frequency response that depends on a combination of the audio rendering frequency response, the reference audio playback frequency response, and the binaural transfer function, A renderer (203) configured to generate the output audio signal using the binaural filter, A device having.

[0146] 2. The audio apparatus according to claim 1, wherein the reference audio playback frequency response is the target frequency response to the first audio signal.

[0147] 3. The audio apparatus according to claim 1, wherein the reference audio playback frequency response is the source frequency response applied in the generation of the first audio signal.

[0148] 4. Any of the above-described audio devices, wherein the frequency equalization data has an index for the audio playback device, and the adapter (209) is configured to determine the reference playback frequency response as a predetermined frequency response for the audio playback device.

[0149] 5. Any of the above-mentioned audio devices, wherein the frequency equalization data includes data describing a finite impulse response (FIR) filter having a frequency response that matches the reference audio playback frequency response.

[0150] 6. The frequency equalization data has data describing a combination of sets of sections, each section being either a primary or secondary section, and the combination has a frequency response that matches the reference audio playback frequency response, any of the above audio devices.

[0151] 7. Any of the above-described audio devices, wherein the frequency equalization data has data describing a set of parallel filters, and the set of parallel filters has a frequency response that matches the reference audio playback frequency response.

[0152] 8. Any of the above audio devices, wherein the frequency equalization data includes frequency response identification data, and the adapter (209) is configured to access a remote server to obtain the reference playback frequency response based on the frequency response identification data.

[0153] 9. Any of the above audio devices, wherein the metadata further includes a reference playback level, and the audio device is configured to adapt the frequency response of the binaural filter in accordance with a first playback level for playback of the output audio signal relative to the reference playback level.

[0154] 10. The audio device according to 9, wherein the metadata includes data indicating the dependency of the reference audio playback frequency response on the playback level, and the adapter (209) is configured to adapt the reference audio playback frequency response according to the playback level.

[0155] 11. A device for generating audio data signals, A receiver (301) configured to receive at least a first audio signal from a first sound source having a posture within the scene, A playback processor (303) configured to determine a reference audio playback frequency response for the first audio signal, Audio data for the first audio signal, A posture indicator showing the posture of the sound source, and Frequency equalization data showing the aforementioned reference playback frequency response, A data signal generator (305) configured to generate the audio data signal having the following characteristics: A device having.

[0156] 12. A method for generating an output audio signal, At least a first audio signal, and, The attitude indicator of the sound source of the first audio signal, and Frequency equalization data showing the reference audio playback frequency response for the first audio signal, The steps include receiving a data signal having metadata including, Steps to determine your listening posture, The steps include determining a binaural transfer function that depends on the listening posture and the posture indicator of the sound source, The steps include determining the audio rendering frequency response, which shows the audio playback frequency response of the audio playback path for the output audio signal, The steps include generating a binaural filter having a frequency response that depends on a combination of the audio rendering frequency response, the reference audio playback frequency response, and the binaural transfer function, The steps include generating the output audio signal using the binaural filter, A method of having.

[0157] 13. A method for generating an audio data signal, The steps include receiving at least a first audio signal from a first sound source having a posture within the scene, The steps include determining a reference playback frequency response for the first audio signal, Audio data for the first audio signal, A posture indicator showing the posture of the sound source, and Frequency equalization data showing the aforementioned reference playback frequency response, A data signal generator configured to generate the audio data signal having the following characteristics: A method of having.

[0158] 14. At least a first audio signal and The attitude indicator of the sound source of the first audio signal, and Frequency equalization data showing the reference playback frequency response for the first audio signal, An audio data signal having metadata including the above.

[0159] 15. A computer program product having computer program code means configured to perform any of the above steps when executed on a computer.

[0160] More specifically, the present invention is defined by the appended claims.

Claims

1. A device for generating an output audio signal, At least a first audio signal, and, The attitude indicator of the sound source of the first audio signal, and Frequency equalization data showing the reference audio playback frequency response for the first audio signal, A receiver configured to receive a data signal having metadata including, A listener posture processor configured to determine listening posture, A binaural transfer function processor configured to determine a binaural transfer function depending on the listening posture and the posture index of the sound source, A processor configured to determine an audio rendering frequency response that shows the audio playback frequency response of an audio playback path for the output audio signal, An adapter configured to generate a binaural filter having a frequency response that depends on a combination of the audio rendering frequency response, the reference audio playback frequency response, and the binaural transfer function, A renderer configured to generate the output audio signal from the first audio signal, including applying the filtering by the binaural filter, A device having.

2. The audio apparatus according to claim 1, wherein the reference audio playback frequency response is a target frequency response for the first audio signal.

3. The audio apparatus according to claim 1, wherein the reference audio playback frequency response is a source frequency response applied in the generation of the first audio signal.

4. The audio device according to any one of claims 1 to 3, wherein the frequency equalization data has an index for an audio playback device, and the adapter is configured to determine the reference audio playback frequency response as a predetermined frequency response for the audio playback device.

5. The audio apparatus according to any one of claims 1 to 4, wherein the frequency equalization data includes data describing a finite impulse response filter having a frequency response that matches the reference audio playback frequency response.

6. The audio apparatus according to any one of claims 1 to 5, wherein the frequency equalization data has data describing a combination of sets of sections, each section being a primary or secondary section, and the combination has a frequency response that matches the reference audio playback frequency response.

7. The audio apparatus according to any one of claims 1 to 6, wherein the frequency equalization data has data describing a set of parallel filters, and the set of parallel filters has a frequency response that matches the reference audio playback frequency response.

8. The audio apparatus according to any one of claims 1 to 7, wherein the frequency equalization data includes frequency response identification data, and the adapter is configured to access a remote server to obtain the reference audio playback frequency response based on the frequency response identification data.

9. The audio device according to any one of claims 1 to 8, wherein the metadata further includes a reference playback level, and the audio device is configured to adapt the frequency response of the binaural filter in accordance with a first playback level for playback of the output audio signal relative to the reference playback level.

10. The audio device according to claim 9, wherein the metadata includes data indicating the dependency of the reference audio playback frequency response on the playback level, and the adapter is configured to adapt the reference audio playback frequency response according to the playback level.

11. A method for generating an output audio signal, At least a first audio signal, and, The attitude indicator of the sound source of the first audio signal, and Frequency equalization data showing the reference audio playback frequency response for the first audio signal, The steps include receiving a data signal having metadata including, Steps to determine your listening posture, The steps include determining a binaural transfer function that depends on the listening posture and the posture indicator of the sound source, The steps include determining the audio rendering frequency response, which shows the audio playback frequency response of the audio playback path for the output audio signal, The steps include generating a binaural filter having a frequency response that depends on a combination of the audio rendering frequency response, the reference audio playback frequency response, and the binaural transfer function, A step of generating the output audio signal from the first audio signal, which includes applying the filtering by the binaural filter, A method of having.

12. A computer program having computer program code means configured to perform all the steps of the method described in claim 11 when executed on a computer.