Generation of audio data signals
Patent Information
- Application Number
- JP2026517860
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-03
- Filing Date
- 2024-10-01
- Publication Date
- 2026-09-17
AI Technical Summary
【0040】 本発明のこれら及び他の態様、特徴及び利点は以下に記載される実施形態から明らかになり、それを参照して説明される。
Smart Images

Figure 2026531705000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to generating and optionally rendering audio data signals, and more particularly, though not exclusively, to generating such signals, for example, to support split rendering applications. Background Art
[0002] The diversity and scope of experiences based on audiovisual content have increased significantly in recent years as new services and methods for utilizing and consuming such content are continuously developed and introduced. In particular, many spatial and interactive services, applications and experiences are being developed to provide users with deeper immersive experiences.
[0003] Examples of such applications include virtual reality (VR), augmented reality (AR), mixed reality (MR) applications (generally referred to as extended reality XR applications), which are rapidly becoming mainstream, with many solutions targeting the consumer market. In addition, many standards have been developed by various standardization organizations. Such standardization activities are actively developing standards for various aspects of VR / AR / MR / XR systems, including streaming, broadcasting, rendering and so on.
[0004] VR applications tend to provide a user experience that corresponds to the user being in a different world / environment / scene, while AR (including decoded reality / MR) applications tend to provide a user experience that corresponds to the user being in the current environment but with additional information or virtual objects or information added. Therefore, VR applications tend to provide a fully immersive, synthetically generated world / scene, while AR applications tend to provide a partially synthesized world / scene that is superimposed on the user's physically existing real-world scene. However, these terms are often used interchangeably and have a high degree of overlap. Below, the term virtual reality / VR will be used to refer to both virtual reality and augmented reality.
[0005] VR applications typically provide users with a virtual reality experience, allowing them to move (relatively) freely within the virtual environment and dynamically change their posture (position and / or orientation). Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well-known from gaming applications, such as in the category of first-person shooter (FPS) games for computers and consoles.
[0006] In addition to visual rendering, most VR (and more generally XR) applications further provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience in which the sound source is perceived to arrive from a position corresponding to the position of the corresponding object in the visual scene (including both currently visible objects and currently or partially invisible objects, such as those behind the user). Thus, the audio and video scenes are preferably perceived to provide a consistent and complete spatial experience.
[0007] Regarding audio, headphone playback using binaural audio rendering technology is widely used. In many scenarios, headphone playback enables a highly immersive and personalized user experience. Using head tracking, rendering can be performed in response to the user's head movements, which significantly increases immersion.
[0008] To generate appropriate audio, rendering devices provide spatial audio data representing the audio scene. However, sound sources can be represented in many different ways, including audio channel signals, audio objects, diffuse non-spatial specific sound sources (e.g., background noise), and ambisonic audio. In fact, it is becoming increasingly common to use multiple different audio representations to describe an audio scene in order to reflect the different types of sound sources that can be represented. However, such an approach increases complexity and resource requirements, often resulting in computationally intensive rendering algorithms.
[0009] These points are problematic and, in particular, contradict the desire to enable devices that can render audiovisual content with very low complexity and resources. In particular, providing low-cost, small, lightweight, low-complexity, and low-computational-resource end-user devices is increasingly desirable. For example, wearable devices with such characteristics are becoming increasingly widespread. In particular, to support a variety of XR applications, it is desirable to be able to provide small and lightweight XR headsets, such as relatively small and lightweight XR glasses.
[0010] To enable or facilitate such devices, it has been proposed to use a partitioned rendering approach in which the (final) part of the rendering process is performed by an end device, and other, more computationally intensive parts of the rendering are performed by another device called an edge device. The edge device may perform potentially critical parts of the rendering process and generate audio signals that are sent to the end device for final adaptation.
[0011] Edge devices are typically more complex and resource-rich than end devices, and accordingly, can implement more complex rendering algorithms and functions. For example, edge devices may be mobile phones, game consoles, computers, or remote servers, while end devices may be user-worn rendering and audio playback devices such as XR headsets / glasses.
[0012] As an example, it has been proposed that an edge device renders received audio data to generate a rendered audio signal, which is then sent to an end device, where the end device can perform additional simple operations (e.g., simple panning) on these audio signals to adapt to the user's movements.
[0013] However, while this approach may provide desirable applications and operations in many scenarios, the current approach tends to be suboptimal. In particular, in many situations, the current approach may provide suboptimal audio quality, suboptimal response to user movements, suboptimal resource usage and distribution, and so on. [Overview of the Initiative] [Problems that the invention aims to solve]
[0014] Therefore, improved approaches to audio signal distribution and / or rendering and / or processing are advantageous, especially for virtual / augmented / hybrid / eXtended reality experiences / applications. In particular, approaches that enable improved behavior, increased flexibility, reduced complexity, easier implementation, improved user experience, improved audio quality, improved adaptation to audio playback capabilities, easier and / or improved adaptation to changes in listener position / orientation (e.g., virtual listener position / orientation), improved resource demand / processing distribution for split rendering approaches, improved eXtended reality experiences, and / or improved performance and / or behavior are advantageous.
[0015] Therefore, the present invention aims to mitigate, reduce, or eliminate one or more of the above-mentioned drawbacks, either individually or in combination. [Means for solving the problem]
[0016] According to one aspect of the present invention, an audio device is provided, the audio device comprising: a receiver configured to receive a plurality of audio elements representing a three-dimensional audio scene, wherein the audio elements comprise a plurality of audio objects, each audio object linked to a position in the three-dimensional audio scene; a listener attitude receiver configured to receive listener attitude indications; and a designator configured to designate at least a first audio object among the plurality of audio objects as a close audio object or a far audio object in response to a comparison between a distance scale indicating the distance between the attitude of a first audio object and the listener attitude and a threshold, wherein the distance scale for the first audio object is far from the threshold. A designator is designated as an audio object, and the distance scale to a first audio object is designated as a close audio object that does not exceed a threshold; an audio mix generator is configured to generate an ambisonic audio mix from a first plurality of audio elements, wherein the first plurality of audio elements include the first audio object when the first audio object is designated as a far audio object; and a data generator is configured to generate an audio data signal having an ambisonic audio mix, wherein the data generator is configured to include the first audio object in the audio data signal when the first audio object is designated as a close audio object.
[0017] This approach may enable the generation of an improved output audio signal. In many embodiments and scenarios, this approach can provide improved audio quality and, for example, reduce complexity for rendering devices that render audio based on audio data signals.
[0018] This approach could, in particular, enable improvements to split rendering, where the rendering of an audio scene can be divided among multiple devices. The audio device could provide an audio data signal that allows for improved distribution of functionality, complexity, and computational resources. The audio device could generate an audio data signal that offers an improved trade-off between the additional requirements for rendering individual audio objects and rendering multiple sources contained within a single ambisonic audio mix.
[0019] Posture can be position and / or orientation. Listener posture can be the listener's position, orientation, or position and orientation.
[0020] The designator may be configured to designate at least one of a plurality of audio objects as a close audio object or a far audio object depending on whether a distance scale indicating the distance between the posture of the first audio object and the listener posture exceeds a threshold. The designator may be configured to designate at least one of a plurality of audio objects as a close audio object if the distance scale indicating the distance between the posture of the first audio object and the listener posture is below a threshold, and as a far audio object if the distance scale exceeds the threshold (typically, depending on the individual embodiment, it may designate it as a far object or a close object if the distance scale is equal to the threshold).
[0021] Listener attitude can represent attitude within the audio scene. The audio scene can be a three-dimensional audio scene.
[0022] The audio mixer may be configured to generate an ambisonic audio mix comprising the first audio object only when the first audio object is designated as a distant audio object. The data generator may be configured to not include the first audio object in the audio data signal when the first audio object is designated as a distant audio object.
[0023] According to an optional feature of the present invention, the designator is configured to designate the first audio object as a close audio object or a distant audio object in accordance with a volume scale of the first audio object.
[0024] According to an optional feature of the present invention, the designator is configured to designate the first audio object as a close audio object or a distant audio object in accordance with a direction from a listener posture to a position of the first audio object.
[0025] According to an optional feature of the present invention, the designator is configured to designate the first audio object as a close audio object or a distant audio object in accordance with a trajectory of a position of the first audio object within an audio scene.
[0026] According to an optional feature of the present invention, the designator is configured to designate the first audio object as a close audio object or a distant audio object in accordance with an order of the ambisonic audio mix.
[0027] According to an optional feature of the present invention, the designator is configured to designate the first audio object as a close audio object or a distant audio object in accordance with a distance between a position of the first audio object and a position of a second audio object.
[0028] According to an optional feature of the present invention, the data generator is configured to transmit an audio data signal to a remote device via a communication link, and the specifier is configured to specify the first audio object as a near audio object or a far audio object in accordance with a data transfer characteristic of the communication link.
[0029] These features may provide improved and / or facilitated operation and / or performance in many embodiments. This feature may result in generation of an audio data signal that provides an improved trade-off between additional requirements for rendering of individual audio objects and rendering of multiple sources included in a single Ambisonic audio mix.
[0030] According to an optional feature of the present invention, the audio mix generator is configured to change a data rate for the first audio object when transitioning between a state specified as a near object and a state specified as a far object.
[0031] This may provide improved and / or facilitated operation and / or performance in many embodiments.
[0032] According to an optional feature of the present invention, the listener posture receiver is configured to receive a plurality of listener postures, the specifier is configured to determine a distance measure in accordance with the plurality of listener postures, and the data generator is configured to transmit the audio data signal to a plurality of remote devices.
[0033] This may provide improved and / or facilitated operation and / or performance in many embodiments.
[0034] According to an optional feature of the present invention, an audio system is provided comprising the above-described audio device, a receiver configured to receive audio data signals from the audio device, and a rendering device having a renderer configured to generate binaural audio signals representing an audio scene, the renderer including contributions from a first audio mix and a first audio object when the binaural audio signals are present in the audio data signals.
[0035] According to an optional feature of the present invention, the rendering device is a headset having an audio transducer configured to reproduce a binaural rendering signal.
[0036] According to an optional feature of the present invention, the designator is configured to designate the first audio object as a nearby audio object or a distant audio object, depending on the characteristics of the rendering device.
[0037] This may provide improved and / or facilitated operation and / or performance in many embodiments.
[0038] According to one aspect of the present invention, a method for operating an audio device is provided. This method includes the steps of: receiving a plurality of audio elements representing a three-dimensional audio scene, wherein the audio elements comprise a plurality of audio objects, each audio object being linked to a position in the three-dimensional audio scene; receiving an indication of a listener's posture; designating at least one of the plurality of audio objects as a close audio object or a far audio object based on a comparison between a distance scale indicating the distance between the posture of the first audio object and the listener's posture and a threshold, wherein the distance scale for the first audio object exceeding the threshold is designated as a far audio object, and the distance scale for the first audio object not exceeding the threshold is designated as a close audio object; generating an ambisonic audio mix from the first plurality of audio elements, wherein the first plurality of audio elements include the first audio object when the first audio object is designated as a far audio object; generating an audio data signal including the ambisonic audio mix; and including the first audio object in the audio data signal when the first audio object is designated as a close audio object.
[0039] According to an optional feature of the present invention, the audio device performs the method of claim 13 and transmits an audio data signal to a rendering device, the rendering device performs the steps of receiving the audio data signal from the audio device and generating a binaural audio signal representing an audio scene, the binaural audio signal including contributions from a first audio mix and a first audio object, if present in the audio data signal.
[0040] These and other aspects, features and advantages of the present invention will become apparent from and be described with reference to the embodiments described below.
[0041] Embodiments of the present invention will be described with reference to the drawings, merely as examples. [Brief explanation of the drawing]
[0042] [Figure 1] This shows an example of a client-server based augmented reality system. [Figure 2] Examples of elements of an audio rendering configuration according to several embodiments of the present invention are shown. [Figure 3] Examples of elements of an audio device according to several embodiments of the present invention are shown. [Figure 4] Examples of elements of an audio rendering device according to several embodiments of the present invention are shown. [Figure 5] This shows an example of an audio object within an audio scene. [Figure 6] This shows an example of an audio object within an audio scene. [Figure 7] Some elements of possible processor configurations for implementing elements of the apparatus according to some embodiments of the present invention are shown. [Modes for carrying out the invention]
[0043] The following description focuses on augmented reality (eXtended Reality) applications where audio is rendered to reflect the user's posture within the audio scene, providing an immersive user experience. Typically, audio rendering may accompany image rendering, providing the user with a complete audiovisual experience. However, it will be understood that the approach described may be used in many other applications as well.
[0044] Augmented reality (including virtual augmented reality and mixed reality) experiences, which allow users to move around in a virtual or augmented world, are becoming increasingly popular, and services are being developed to improve such applications. In many such approaches, visual and audio data may be dynamically generated to reflect the user's (or observer's) current posture.
[0045] In this field, the terms position and orientation are used as general terms for position and / or orientation. For example, a combination of position and orientation of an object, camera, head, or view may be called orientation or position. Thus, an index of position or orientation may have up to six values / components / degrees of freedom, each value / component typically describing an individual characteristic of the position or orientation of the corresponding object. Of course, in many situations, position or orientation may be represented by fewer components, for example, when one or more components are considered fixed or irrelevant (for example, if all objects are considered to be at the same height and horizontal, four components may provide a complete representation of the object's orientation). Hereafter, the term orientation is used to refer to position and / or orientation that can be represented by 1 to 6 values (corresponding to the maximum possible degrees of freedom). The term "position" may be replaced with the terms "at least one of position and orientation" or "position and / or orientation". Listener orientation may be the listener's position and / or listener orientation.
[0046] Many XR applications are based on orientation, which has the maximum degrees of freedom: 3 degrees of freedom for position and 3 degrees of freedom for orientation, resulting in a total of 6 degrees of freedom. Orientation may therefore be represented by a set of 6 values or a vector representing the 6 degrees of freedom, and thus the orientation vector may provide indications for 3D position and / or 3D orientation. However, in other embodiments, it will be understood that orientation may be represented by fewer values (rotation may be in the form of a quaternion or rotation matrix, and therefore more than 6 values are possible).
[0047] A system or entity that provides the maximum number of degrees of freedom to the observer is typically said to have 6 degrees of freedom (6DoF). Many systems and entities provide only orientation or position, and these are typically known to have 3 degrees of freedom (3DoF – typically used to represent an approach that uses orientation with variable orientation and fixed position).
[0048] Typically, a virtual reality application generates a three-dimensional output in the form of separate view images for the left and right eyes. These can then be presented to the user by appropriate means, typically such as the individual left-eye and right-eye displays of a VR headset. In other embodiments, one or more view images may be displayed, for example, on an automated stereoscopic display, or in fact, in some embodiments, only a single two-dimensional image may be generated (for example, using a conventional two-dimensional display).
[0049] Similarly, an audio representation of a scene may be provided for a given observer / user / listener posture. The audio scene is typically rendered to provide a spatial experience in which the sound source is perceived to originate from a desired location. Since the sound source can be static within the scene, a change in listener posture results in a change in the relative position of the sound source to the user's posture. Accordingly, the spatial perception of the sound source may change to reflect its new position relative to the user. Audio rendering may be adjusted in response to listener posture.
[0050] Listener pose input can be determined in different ways in different applications. In many embodiments, the user's physical movement can be directly tracked. For example, a camera monitoring the user area may detect and track the user's head (or eyes (gaze tracking)). In many embodiments, the user may wear a VR headset that can be tracked by external and / or internal means. For example, the headset may have an accelerometer and a gyroscope that provide information about the movement and rotation of the headset and therefore the head. In some examples, the VR headset may have an (e.g., visual) identifier that transmits signals enabling external sensors to determine the position and orientation of the VR headset.
[0051] In many systems, VR / scene data, particularly audio data representing audio scenes, may be provided by a remote device or server. For example, a remote server may generate audio data representing an audio scene and transmit audio signals corresponding to audio components / objects / channels, or other audio elements corresponding to different sound sources within the audio scene, along with positional information indicating their locations (which may change dynamically for moving objects, for example). The audio signals / elements may include elements associated with specific locations, but may also include elements for more dispersed or diffuse sound sources. For example, audio elements representing generic (unlocalized) background sounds, ambient sounds, diffuse reverberation, etc., may be provided.
[0052] A local VR device can properly render audio elements by applying appropriate binaural processing that specifically reflects the relative position of the sound source of the audio component.
[0053] Similarly, a remote device may generate visual / video data representing a visual audio scene, and transmit visual scene components / objects / signals, or other visual elements corresponding to different objects within the visual scene, along with positional information indicating their locations (which may change dynamically for moving objects, for example). Visual items may include elements associated with specific locations, or they may include video items for more distributed sources.
[0054] In some embodiments, visual items may be provided as individual and separate items, such as descriptions of individual scene objects (e.g., dimensions, texture, opacity, reflectivity, etc.). Alternatively or additionally, visual items may be represented as part of an overall model of the scene, including, for example, descriptions of different objects and their relationships to one another.
[0055] For VR services, the central server may, in some embodiments, generate audiovisual data representing a three-dimensional scene, specifically representing audio with multiple audio signals representing sound sources within the scene, which can be rendered by local clients / devices.
[0056] Figure 1 shows an example of a VR / XR system in which a central server 101 interacts with several remote clients 103 via a network 105, such as the internet. The central server 101 may be configured to support a potentially large number of remote clients 103 simultaneously.
[0057] Such an approach can offer improved trade-offs in many scenarios, for example, between complexity and resource requirements for different devices, communication requirements, etc. For instance, scene data may only need to be transmitted once or relatively rarely, and the local rendering device (remote client 103) processes the scene data locally to receive the observer's pose and render audio and / or video to reflect changes in the observer's pose. This approach can provide an efficient system and an engaging user experience. For example, it can allow scene data to be centrally stored, generated, and maintained while significantly reducing the required communication bandwidth while providing a low-latency real-time experience. This could be suitable, for example, for applications where a VR experience is delivered to multiple remote devices.
[0058] In some cases, a remote client may include multiple devices configured to interact with each other to provide rendering of audiovisual data. In particular, as shown in Figure 2, the rendering configuration (which may specifically correspond to the remote client) may include a first device 201, also called edge device 201, which receives audiovisual data describing the scene. To generate an audio representation, the first / edge device 201 receives audio data describing the audio scene accordingly.
[0059] The edge device 201 is configured to process the received audio data to generate intermediate audio data, which is then transmitted to the end device 203, which is configured to process the intermediate audio data to generate rendered output audio signals. These output signals may specifically be binaural signals for the listener's left and right ears, respectively. Thus, this approach employs a split rendering approach in which the rendering of the audio (scene) is divided among more than one device. In many embodiments, the end device 203 may specifically include audio playback means such as audio transducers (often having one audio transducer for the user's right ear and one audio transducer for the user's left ear).
[0060] The edge device 201 may specifically be a mobile phone, computer, game console, laptop, tablet, etc., and the end device 203 may typically be a user-worn playback device such as XR glasses and / or headphones.
[0061] Figure 3 shows an example of several elements of an audio device configured to generate audio data signals for a typical three-dimensional audio scene from audio data describing the audio scene. This device may specifically be edge device 201 in Figure 2, and will be described with reference to it. A key issue for an approach like that in Figure 2 is how to distribute functions and processing across different devices and what data to send from the edge device to the end device. Such considerations include not only the computational load on different devices but also other parameters such as the impact of communication between devices, including the effects of communication errors and delays. For example, the responsiveness of rendered audio to rapid changes in listener posture typically depends heavily on communication delay. Therefore, from the standpoint of speed and responsiveness, it is desirable that processing be performed mainly on the end device. However, from the standpoint of resources, size, complexity, etc., it is typically desirable that processing be performed mainly on the edge device. Therefore, trade-offs between conflicting requirements tend to be important for the performance of the approach.
[0062] The edge device 201 in Figure 3 has a receiver configured to receive audio data describing a three-dimensional audio scene. The audio data includes multiple audio elements, each providing a representation of audio / sound in the audio scene. The audio elements may be audio signals representing sound sources, diffuse noise, ambient noise, multiple sources, etc. Often, the audio elements may include different types of audio elements, such as audio objects, audio channel signals, and ambisonic audio signals / mixes.
[0063] In some cases, one or more audio elements may include one or more audio channel signals. Such channel signals may be provided for a given and / or nominal location, such as a nominal speaker position. In such cases, each channel may contain information about a particular portion of the mix for a given speaker configuration.
[0064] An audio element specifically includes several audio objects. Each audio object provides a description / representation of the audio, and typically a description / representation of the audio from a sound source. Each audio object is linked to a position within the audio scene, and the audio object provides audio data and spatial data. The position / orientation may be absolute (global) relative to the scene, relative to another object in the scene, or relative to the listener. Typically, an audio object provides audio data and positional data relative to a sound source within the audio scene. Object-based immersive audio may add audio streams and metadata to a decoder to give instructions on how each stream should be positioned within the 3D sound field.
[0065] In many scenarios, the audio element may also include one or more additional ambisonic signals / mixes.
[0066] Ambisonics, particularly higher-order ambisonics (HOA), is a method for recording and reproducing a three-dimensional sound field (scene-based audio). Unlike common channel-oriented transmission methods, this system focuses on reproducing the entire sound field at the listening position and does not require a predetermined number of speaker positions for sound reproduction. The relevant loudspeaker signals are calculated for each loudspeaker position, used by mathematical derivation of transmitted values for sound pressure and velocity. In its basic version, known as first-order ambisonics, ambisonics can be understood as a three-dimensional extension of M / S (mid / side) stereo, adding additional difference channels for height and depth. The resulting signal set is called the B format. Sound information is transmitted in four channels: W, X, Y, and Z. Component W contains only sound pressure, typically recorded with an omnidirectional or omnidirectional microphone. Signals X, Y, and Z are directional components in their corresponding spatial axes. They may be recorded with a figure-eight oriented microphone in their corresponding axes. A simple ambisonic approach is the A format, which uses four cardioid microphones, and this can typically be converted to the more practical B format.
[0067] The purpose of this process is to reconstruct the sound pressure and associated sound direction vectors recorded from these signals at the listener's listening position.
[0068] For audio objects that are part of an ambisonic mix, the ambisonic order determines the amount of spatial smearing of the point source. The higher the order, the less spatial smearing there is and the greater the spatial directionality.
[0069] The edge device 301 further includes a listener attitude receiver 303 configured to receive listener attitude indications. The listener attitude receiver 303 is configured to determine the listener attitude from data that may be provided from the end device 203 to the edge device 201. Thus, the edge device 201 and the end device 203 can establish a communication link that enables listener attitude data to be transmitted from the end device 203 to the edge device 201. The processing of the edge device 201 may depend on the attitude of the end device 203; for example, a VR headset / glasses may track the user's movements and generate sensor data to be transmitted to the edge device 201, and the edge device 201 may adapt its processing according to the user's movements.
[0070] The listener's posture may be determined specifically in response to sensor input, for example, from a suitable sensor that is part of a headset. It is understood that many suitable algorithms are known to those skilled in the art, and for the sake of brevity, this will not be described in more detail here.
[0071] In some embodiments, the determination of the listener pose in the audio scene is performed in the end device 203, which may then transmit the determined listener pose to the edge device. In some embodiments, other data, such as sensor data, may be transmitted to the edge device 201, which may be configured to determine the listener pose based on this data.
[0072] Listener posture is typically the user's position within the audio scene in which the audio presented is perceived, i.e., it represents the listener / user's position within the audio scene. In the approach shown in Figure 2, the edge device 201 and the end device 203 work together to process the received audio data to generate an output audio signal, particularly a binaural audio output signal, for the user, in order to provide a spatial audio experience / perception of the audio scene based on the listener posture. Listener posture information is provided to the edge device 201 so that the processing of both the edge device 201 and the end device 203 can be adapted to the current listener posture.
[0073] A key issue in such scenarios is which processing should be performed on the edge device 201 and the end device 203, respectively, and which specific information should be sent from the edge device 201 to the end device 203. Typically, the edge device 201 has significantly more computing resources than the end device 203, and therefore it is desirable to perform most of the processing, especially complex and resource-intensive processing, on the edge device 201. However, the communication of data regarding listener posture and the corresponding audio data to the end device 203 is typically significant and introduces communication delays that can result in a reduced user experience. This introduces round-trip delays that can create a perceptually large delay between user movement and the corresponding perceived audio. Therefore, it is often desirable that the end device 203 be able to perform some, preferably low-complexity, processing, such as adapting the rendered audio locally to changes in listener posture, while keeping most of the processing on the edge device 201. However, to achieve an efficient trade-off, the distribution of processing and data communicated from the edge device 201 to the end device 203 is important.
[0074] Figure 4 shows an example of the elements of the end device 203 in Figure 2. The end device 203 has a receiver 401 that receives audio data signals from the edge device 201. As will be described in more detail below, the audio data signals have audio signals generated by the edge device 201 to represent an audio scene. The receiver 401 is coupled to a renderer 403 configured to render an audio signal for the listener from the received audio signal. Specifically, the renderer 403 generates a binaural audio signal representing the audio scene from the received audio data. As will be described later, the audio data signal includes one or more ambisonic audio mixes and typically one or more audio objects, and the renderer 403 is configured to generate a binaural audio signal from these signals.
[0075] In this example, the binaural output signal is supplied to output circuit 405, which outputs the binaural signal to headphones, which are typically part of a VR headset. Output circuit 405 may include, for example, a digital-to-analog converter, an amplifier, and so on, as is well known to those skilled in the art.
[0076] The renderer 403 is coupled to a listener pose processor 409 configured to determine the listener pose. In this example, the listener pose is determined based on sensor input from headphones / VR headset 407. The listener pose is provided to the renderer 403, which is configured to generate a binaural signal to represent the audio scene from the listener pose, specifically, to act dynamically on it to track changes in the listener pose. The listener pose may, in turn, correspond to the user / listener's position in the scene.
[0077] The listener attitude processor 409 is further coupled to a transmitter 411 configured to transmit the listener attitude to the edge device 201.
[0078] In this example, edge device 201 is configured to generate an audio data signal and send it to end device 203, the audio data signal containing an ambisonic audio mix and several audio objects. Edge device 201 is configured to determine whether several audio elements are included in the audio data signal as part of the ambisonic audio mix or as listener-position-dependent audio objects.
[0079] The edge device 201 has a designator 305 configured to designate at least one of the received audio objects as a near or far audio object, depending on the listener's orientation. Specifically, for a given audio object, the designator 305 may determine a distance scale indicating the distance between the orientation of the first audio object and the listener's orientation. The distance scale is then compared to a threshold, and if the distance scale exceeds the threshold (indicating that the distance is greater than a given value), the audio object is designated as a far object; otherwise, it is designated as a near object. In some embodiments, it will be understood that the audio object may include designations to other possible categories, which include subcategories of the designation of an audio object as a near or far object.
[0080] In this approach, a distance scale for a first audio object designated as a far audio object exceeds a given threshold, while a distance scale for a first audio object designated as a close audio object does not exceed the threshold. Therefore, designating a first audio object as a far audio object requires that the distance scale indicates a distance exceeding the threshold. Designating a first audio object as a close audio object requires that the distance scale indicates a distance not exceeding the threshold.
[0081] The algorithms / criteria used for specific implementation and precise specification may depend on the particular preferences and requirements of each individual embodiment. In particular, it should be noted that the thresholds used will clearly depend on many characteristics and design choices of the individual application. It should also be noted that the specification may depend on several other parameters and considerations. However, fundamentally, if the first audio object is designated as a far audio object, it is at a distance above the threshold; if it is designated as a close audio object, it is at a distance below the threshold.
[0082] The distance may be any appropriate distance measure that indicates the distance between the listener's pose and the poses of audio objects within the audio scene, such as the Euclidean distance or the sum of absolute coordinate differences.
[0083] The distance scale may have an increasing value for increasing distance between the pose of the first audio object and the pose of the listener, and the explanation will focus on such examples (for example, when referring to comparison with an appropriate threshold).
[0084] The edge device 201 further includes an audio mix generator 307 configured to generate an ambisonic audio mix from multiple audio elements. The audio mix generator 307 is configured to include or exclude audio objects from the ambisonic audio mix depending on whether the audio objects are designated as nearby or distant audio objects. In particular, if an audio object is designated as a distant audio object, it is included in the ambisonic audio mix, but if it is designated as a nearby audio object, it is not included in the ambisonic audio mix.
[0085] In many embodiments, the audio mix generator 307 may be configured to include multiple audio objects designated as distant audio objects, often all audio objects. In many embodiments, it may be configured not to include audio objects designated as close audio objects.
[0086] Therefore, the audio mix generator 307 is configured to generate an ambisonic audio mix that includes audio objects designated as distant audio objects.
[0087] Furthermore, in many embodiments, the audio mix generator 307 is configured to include other types of audio elements in an ambisonic audio mix, such as a received ambisonic audio mix. In fact, in some embodiments, the audio mix generator 307 is configured to add audio objects designated as distant audio objects to an existing (received) ambisonic audio mix. In some embodiments, the ambisonic audio mix may be generated to include, for example, channel-based audio signals.
[0088] In some embodiments, the audio mix generator 307 may be coupled to a first audio signal generator 309, which may be configured to generate a first audio signal for transmission to an end device 203. The first audio signal represents an ambisonic audio mix, and in many cases, the generated ambisonic audio mix may be transmitted directly without modification. However, in other embodiments, the ambisonic audio mix may be processed to generate a different representation, such as providing a binaural representation.
[0089] The ambisonic audio mix / first audio signal is supplied to a data signal generator 311 configured to generate an audio data signal that includes the ambisonic audio mix (which may be directly represented as an ambisonic audio mix or, in some cases, represented by a set of binaural signals, etc.).
[0090] Audio objects designated as nearby audio objects are supplied to the generator 311 and included in the audio data signal as audio objects. In some cases, the edge device 201 further has a second audio signal generator 313 which can generate a second audio signal that provides a proper representation of the audio objects, for example by encoding a binaural representation.
[0091] The generator 311 may generate an audio data signal that includes an ambisonic audio mix and any audio objects designated as nearby audio objects. Furthermore, the audio data signal may be generated without including any audio objects designated as distant audio objects, but these audio objects may instead be included in the ambisonic audio mix.
[0092] Next, the end device 203 may receive an audio data signal and process the received ambisonic audio mix and audio objects to generate a binaural output signal for the current listener posture.
[0093] In many embodiments, it will be understood that the audio data signal may include audio elements other than the ambisonic audio mix and audio objects. For example, it may include channel-based audio, diffuse background audio signals, other ambisonic audio mixes, etc. In such cases, the end device 203 may include the ability to render such signals and combine them with the audio signals generated from the ambisonic audio mix and audio objects.
[0094] In many embodiments, this approach offers particularly favorable distribution of processing and favorable selection of audio data to send from edge device 201 to end device 203. This typically requires relatively few computing resources at edge device 201, while enabling a favorable user experience with high-quality audio and fast adaptation.
[0095] Ambisonics is an efficient form for representing a (large) set of audio objects. In particular, if these audio objects are diffuse, a low-order ambisonic mix is sufficient. The achievable spatial resolution of the audio objects within an ambisonic mix is directly coupled to the order of the ambisonic mix. Note that even if there is only one audio object in the ambisonic mix, a higher spatial resolution applies to the ambisonic mix as a whole. Therefore, it is not efficient to represent a limited set of point sources as an ambisonic mix with a sufficient order to achieve a specific spatial resolution. Instead, it is more efficient to represent a limited set of point sources as separate audio objects, optionally combined with an ambisonic mix.
[0096] A scene rendered from an Ambisonics mix can be considered to "exist" on a sphere. As a result, when a user approaches a particular audio object on that sphere, the audio object remains on the sphere, even if rendered with a different gain. That is, when moving towards the sphere, the sphere essentially moves with the user. Therefore, audio objects rendered using an Ambisonics mix are not "approachable." Note that instead of the user approaching a point source audio object, the point source audio object may "approach" the user, for example, determined by the trajectory of a particular object. When audio objects become approachable, they typically cease to be properly represented using an Ambisonics mix, independently of their order.
[0097] The ambisonic mix can be rendered for a speaker setup or efficiently converted into a binaural signal for playback in headphones. Essentially, the ambisonic mix is easily converted to account for the user's (3D) rotation. That is, if the user rotates their head by a specific angle in any (3D) direction, the ambisonic mix can be efficiently updated to account for the rotation. As a result, the user will experience an updated ambisonic mix that reflects the rotation of their head.
[0098] However, unlike rotation, translation (change of position) cannot be efficiently compensated for in an ambisonic mix. In particular, there is no efficient way to accommodate translation when an audio object is approachable in response to user translation, such as when the user moves towards a point source or when an audio object in the ambisonic mix moves towards the user. As shown in Figure 5, for an audio object close to the user (object 1), the user's translation (x) has a greater impact on the perceived angle to the audio object than for an audio object farther away from the user (object 2). For object 3, which is positioned more laterally, the user's translation (x) has only a slight impact on the perceived angle to the audio object. However, since the distance between object 3 and the (translated) user is below a threshold, the translation may render the audio object approachable. A new shape (a circle in this example) can be defined for the approachable object relative to the user's new position. Object 1 has a greater impact on the perceived angle, but in this example, it may move outside the (new) circle.
[0099] Since an ambisonic mix is an efficient representation of (preferably diffuse) audio objects, one approach can be to ignore the effects of (small) user translation and not update the ambisonic mix accordingly.
[0100] However, when audio objects are point sources, especially if they are accessible, and / or if there are associated visual components, there is a clear advantage to an accurate representation of the audio object. In the case of corresponding visual components, even a match between the audio and visual components is highly desirable for a proper user experience. This requires a dynamic trade-off between representing audio objects separately and representing them as part of an ambisonic mix.
[0101] One approach to split rendering can be to pre-render all content to an ambisonic mix based on the predicted listener pose, and then finally render from ambisonic to binaural on the end device. Since the user's three degrees of freedom in the directions of translation (X, Y, Z) are not represented in the ambisonic rendering, this can result in increased round-trip delay (from the user's translation until receiving the updated pre-rendered ambisonic representation from the edge device). Even if the round-trip delay is acceptable, such translational changes can result in audible artifacts, especially for sound sources relatively close to the user, as they are likely to result in significant changes in the angle of incidence of the sound source, which is an important perceptual cue for estimating the location of the sound source.
[0102] Therefore, the inventors recognized that the approach in which the scene is appropriately pre-rendered in ambisonics format based on the most recent known listener posture is not suitable for scenarios in which the listener can move closer to the sound source or can rapidly change the distance to the sound source.
[0103] The approach of pre-rendering multiple binaural signal pairs on edge devices and interpolating the final binaural pair based on those signals on end devices tends to be suboptimal in terms of both the resulting audio quality and the computational complexity required for pre- and post-rendering. In such cases, the output sound quality depends on the final pose offset. For example, using a variation rotated by only 15 degrees to interpolate the output in the final pose, serious artifacts begin to appear when the rotation exceeds 20 degrees. The artifacts from interpolation are proportional to the minimum angle between the available pose and the actual pose. Furthermore, the proposed approach has artifacts when the listener's pose changes in directions other than yaw, such as the pitch and roll of the user's head.
[0104] The aforementioned approach of dynamic adaptation between representing different audio objects, and therefore corresponding sound sources, as part of an ambisonic audio mix or as separate audio objects can address many of the problems described above and, in many embodiments, may offer an improved and particularly advantageous approach.
[0105] This approach can be illustrated by Figure 6, which shows an example of a scene composed of multiple audio objects (1, 2, 3, 4, 5). Some audio objects are included as part of an ambisonic audio mix (3, 4, 5), while other audio objects (1, 2) are sent separately as separate audio objects for independent binaural rendering. The ambisonic audio mix and the separate audio objects are included in alternative audio data and sent to the end device. The ambisonic audio mix may be provided as a binauralized mix generated by a second audio signal generator. The audio objects are rendered (binauralized or rendered to speakers) at the end device and then combined with the pre-rendered ambisonic audio mix to produce a rendered output for consumption on one or more transducers, typically headphones or an AR / VR set. Specifically, the ambisonic audio mix may be binauralized in combination with the audio objects or rendered to speakers.
[0106] In this example, the set of audio objects included in the ambisonic audio mix, each individually represented, can change dynamically. For example, in Figure 7, the trajectory may move the audio object / sound source 4 closer to the listener's position, specifically moving within a threshold and changing from a state designated as a distant object to a state designated as a close audio object. Thus, it may be removed from the ambisonic audio mix and introduced as another audio object in the audio data signal, which may allow for more accurate and dedicated rendering by the end device 203.
[0107] Changes in relative position and distance can similarly occur due to movement of the user / listening position, specifically translation. In the case of audio objects (3, 4, 5) included in the ambisonics mix, user translation (X, Y, Z) does not result in the expected change in the direction of the audio objects until after the round-trip delay, i.e., until the user translation is reflected in the update of the ambisonics audio mix (the translation is reported in the listener posture transmitted to the edge device 201, where the ambisonics audio mix is modified to reflect the new position, resulting in a transmitted ambisonics audio mix corresponding to the new position). As previously shown, such translation and parallax between the listener posture from which the ambisonics audio mix was generated and the new listener posture are far less perceptible for distant sources than for nearby or approachable sources. Both spatial smearing and round-trip delay due to the ambisonics representation make it more difficult to accurately locate nearby sources by attempting to approach them. The distance threshold for including sound sources and audio objects in an ambisonic audio mix can be determined as an appropriate threshold at the point where the audio object begins to become accessible, or where, in the worst case, user translation (during round-trip delay) is considered unacceptable due to the sound source position error.
[0108] The distance from the user to individual audio objects is influenced by the movement of both the user and the audio objects themselves. For example, if the user moves physically or virtually towards object 4 in Figure 6, this audio object may be translated into a circle representing a distance threshold, while other audio objects may be translated outside the circle (e.g., audio object 1). Alternatively, a similar effect can be achieved if the trajectory of object 4 enters the circle. In that case, the audio object may become accessible without the user's translation. Thus, the representation of audio objects, either as components within an ambisonic audio mix or as separate audio objects, can be dynamically updated. In many embodiments, such switching may involve an element of hysteresis.
[0109] In this approach, as soon as an audio object crosses or approaches a threshold, the audio object may transition from the ambisonic audio mix to the audio object, or vice versa. Smaller distances mean that the object will transition from ambisonics to the audio object if it approaches. These transitions can be seamless, with the audio object crossfading from the ambisonic mix to separately transmitted audio objects, and vice versa.
[0110] During these transitions, the resolution of separately transmitted audio objects can be advantageously reduced due to the perceptual masking of audio objects within the ambisonic audio mix; that is, a lower bitrate can be used to encode audio objects during such transitions.
[0111] At points where an audio object transitions outside or inside an ambisonic audio mix, portions of the audio object represented as separate audio objects may be assigned a lower bitrate. As the audio object transitions further outside the ambisonic audio mix, the bitrate may increase. Additionally, masking (by the ambisonic audio mix and / or other separate objects) may be considered when determining the bitrate required to represent an audio object transitioning from the ambisonic audio mix.
[0112] Therefore, in some embodiments, the audio mix generator is configured to change the data rate for a first audio object when transitioning between a state designated as a close object and a state designated as a far object, and thus between a state that is part of an ambisonic audio mix and a state that is not part of an ambisonic audio mix. The transition may be gradual.
[0113] In different embodiments, the designation of audio objects as near or far objects may favorably take into account different parameters and characteristics.
[0114] In many embodiments, the designation of an audio object may depend on the volume scale of the audio object. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, as a function of the volume scale for the first audio object. Thus, in some embodiments, at least one of the distance scale and the threshold depends on the volume scale for the audio object.
[0115] For example, the louder a given audio object is, the more likely it is to be considered a nearby audio object than a quieter one. Therefore, the threshold may be a function of the volume of the audio object, specifically a monotonically increasing function of volume (or, conversely, the distance measure may be a monotonically decreasing function of volume). Thus, in many embodiments, the distance at which an audio object is considered a nearby audio object may increase with increasing volume.
[0116] In many embodiments, such an approach can offer advantageous performance and allow for adaptations to provide a more perceptually consistent experience.
[0117] In some embodiments, the specification may further consider the volume of one or more other audio objects, specifically their relative volume. This may allow for consideration of, for example, masking effects from other sound sources.
[0118] Specifically, a threshold in the distance scale can be adapted to move an audio object from being included in an ambisonic audio mix to being represented as a separate audio object (and vice versa), depending on the volume and / or temporal / spatial masking of the audio object. For example, for a loud (and therefore more prominent) audio object, the threshold may be modified to increase the likelihood that the audio object will be designated as a nearby audio object, while for an audio object masked by another object, the threshold for the masked (less important) object may be modified to decrease the likelihood that it will be designated as a nearby audio object (thus, the threshold may decrease for increasing volume and increase for decreasing volume).
[0119] In many embodiments, the designation of an audio object may depend on a diffusion scale for the audio object. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, as a function of the diffusion scale for the first audio object. Thus, in some embodiments, at least one of the distance scale and the threshold depends on the diffusion scale for the audio object.
[0120] In particular, because diffuse audio objects have lower localization than point source audio objects, the threshold may depend on a measure of the degree of diffusion of the audio object. For example, if the threshold distance (radius) of a point source audio object (for a given angle of incidence) is d p In this case, the threshold distance for the diffuse audio object is < d p (For example) it could be.
[0121] In many embodiments, the designation of an audio object may depend on the direction from the listener's posture to the position of the audio object. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, as a function of the direction from the listener's posture to the position of the audio object. Thus, in some embodiments, at least one of the distance scale and the threshold depends on the direction from the listener's posture to the position of the audio object.
[0122] Audio objects represented by the angle of incidence to the listener can be localized more effectively than, for example, audio objects represented by the angle of incidence to the listener. Therefore, thresholds, which were previously represented by circles (i.e., direction-independent distances) in diagrams, may be represented by asymmetric shapes. For example, if the threshold distance (radius) from the angle of incidence to the listener is d, then the threshold distance for an audio object with a back incidence angle may be < d (e.g., ). For all other directions, the threshold distance may transition seamlessly between the angle of incidence to the listener and the back incidence angle. Typically, the threshold pattern may correspond to localization accuracy as a function of the angle of incidence.
[0123] In many embodiments, the designation of an audio object may depend on the trajectory of the audio object's position within the audio scene. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, as a function of the trajectory for the first audio object. Thus, in some embodiments, at least one of the distance scale and the threshold depends on the trajectory of the audio object.
[0124] The trajectory may be a trajectory relative to the listener's posture, and therefore may be a relative trajectory caused by the movement of the audio object and / or the listener's posture within the audio scene.
[0125] The designator 305 may, for example, track the position of an audio object and determine whether it is moving toward or away from the listener's pose. In particular, based on its trajectory, it may estimate whether the audio object is likely to move toward the listener's pose and, as a result, likely to be designated as a nearby object. If so, the distance threshold may be increased to result in, for example, earlier designation of the audio object, earlier removal from the ambisonic audio mix, and inclusion as another audio object.
[0126] The velocity of an audio object's trajectory can also affect the threshold at which an object moves from a state designated as a nearby object to a state designated as a farther object, or vice versa. For example, for a substantially stationary object, the threshold may remain unchanged, while for a fast-moving object, the threshold may be increased so that the object is represented as a different object at the appropriate time. Different criteria may be used to determine the object's velocity, either based on metadata or prediction, for example, as follows: • Velocity vector from listener's posture to audio object position (prediction, potential change in angular distance) • Animated (moving) point source audio object
[0127] In many embodiments, the designation of an audio object may depend on the order of the ambisonic audio mix. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, as a function of the order of the ambisonic audio mix. Thus, in some embodiments, at least one of the distance scale and the threshold depends on the order of the ambisonic audio mix.
[0128] In some embodiments, distance may take into account the order of the ambisonic audio mix, specifically, the higher the order of the ambisonic audio mix, the more likely audio objects are to be considered distant audio objects and therefore more likely to be included in the ambisonic audio mix. The higher the order of the ambisonic audio mix, the better the ambisonic audio mix can represent the audio objects, and accordingly, it may be more appropriate to include an increasing number of audio objects in the ambisonic audio mix.
[0129] In many embodiments, the designator 305 may also consider the number of audio objects included in the ambisonic audio mix, specifically the number of audio objects designated as distant audio objects. For example, a maximum or preferred number of audio objects for a given order may be determined, and the number of distant audio objects included in the ambisonic audio mix may be limited to this maximum number.
[0130] In some embodiments, the order of the ambisonic audio mix may be adapted depending on the number of audio objects designated as distant audio objects (and audio objects that should be included in the ambisonic audio mix).
[0131] The threshold may include a balance between the ambisonics order and the number of audio objects. As previously shown, increasing the ambisonics order increases the object identifiability. Therefore, to represent a large number of objects, it may be beneficial to increase the ambisonics order so that fewer objects are represented separately.
[0132] In many embodiments, the designation of an audio object may depend on an importance or priority indication for the audio object. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, as a function of the importance or priority indication for the first audio object. Thus, in some embodiments, at least one of the distance scale and the threshold depends on the importance or priority indication for the audio object.
[0133] Importance / priority indications may be determined, for example, based on the type of audio provided, may be assigned manually, and / or may be received, for example, along with the received data for the input audio element.
[0134] For example, an object representing a speaker, such as a dialogue object, is likely to have higher importance in the scene than a (background) object that is emitting some sound. Audio metadata may be used to classify the importance of objects and thereby influence the threshold at which an object moves out of the ambisonic audio mix and is represented as a separate audio object. Examples of embodiments that may indicate higher importance include: (If an audio object represents a user, the need to render it as a separate audio object increases.) • Whether there is a visual component connected to the audio object • Size of audio object • Flags or "importance" scales indicated in the metadata.
[0135] In many embodiments, the designation of an audio object may depend on the distance between the location of one audio object and the location of another audio object. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, as a function of the distance between the location of one audio object and the location of another audio object. Thus, in some embodiments, at least one of the distance scale or threshold depends on the distance between the location of one audio object and the location of another audio object.
[0136] In some embodiments, the designation may depend on the position of one or more other audio objects. For audio objects that are close to each other (located substantially in the same place) or have a small angular distance from the user, the ability to identify them as individual objects is reduced. Therefore, instead of representing all these same-location objects in an ambisonic audio mix, the same-location objects may be rendered as one or more combined audio objects. This improves the accessibility of such groups of same-location objects, reduces the bitrate required to represent these objects jointly, and decreases the computational complexity of rendering them as individual audio objects.
[0137] For audio objects that are close to each other, or for one or more dominant audio objects within a group of nearby audio objects, it may not be necessary to transmit all audio objects in the group, or to use the same transmission rate for all of these individual audio objects. Audio objects within a group of nearby audio objects with similar dominance may be grouped into a single audio object.
[0138] From a computational complexity standpoint, it may be desirable to limit the number of separately transmitted audio objects to a maximum value. For example, there may be separate maximum values for the number of separately transmitted audio objects and for the number of audio objects that are expected to approach a threshold.
[0139] Separate audio objects can be represented as pre-rendered binaural signals. This can be done as part of a trade-off between complexity and quality, depending on the capabilities of the end device.
[0140] For example, with low available computing resources, only a single representation of the audio object may be pre-rendered. With medium available computing resources, additional representations of the audio object may be pre-rendered.
[0141] In many embodiments, the designation of audio objects may depend on the characteristics of the rendering / end device. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, depending on the characteristics of the rendering / end device. Thus, in some embodiments, at least one of the distance scale or threshold depends on the characteristics of the rendering / end device.
[0142] The characteristics may specifically refer to the capabilities of the end device, such as the available rendering algorithms and processing power.
[0143] For example, during setup or initialization, the end device 203 may send a data message to the edge device 201 indicating characteristics such as the functionality of the end device 203. For example, the edge device 201 may indicate the type or processing capacity of the edge device. The specifier 305 may take such information into consideration. For example, the maximum number of audio objects that can be processed by the end device 203 may be determined according to the processing capacity of the end device 203. The distance threshold may be set accordingly, specifically, so that the number of isolated audio objects represented in the audio data signal does not exceed the determined maximum number.
[0144] In some embodiments, the distance threshold may depend on the capabilities of the end device. For example, for end devices with limited processing power, it may be beneficial to modify the overall threshold to reduce the number of audio objects represented as separate audio objects. Other trade-offs, including combinations of the parameters described above, may also be possible.
[0145] In some embodiments, the threshold may be dynamically changed to have a specific number of audio objects that are represented separately to optimize audio quality within the functionality of the end device.
[0146] Therefore, the processing power of the end device 203 may be taken into consideration. For example, in the case of a low-cost device, a limited set of pre-rendered audio objects may be provided.
[0147] In many embodiments, the designation of an audio object may depend on the data transfer characteristics of the communication link for transmitting audio data signals to the edge device 201. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance measure, as a function of the data transfer characteristics. Thus, in some embodiments, at least one of the distance measure or the threshold depends on the data transfer characteristics. Specifically, the data transfer characteristics may be the data rate of the communication.
[0148] For example, edge device 201 may estimate the current bandwidth or throughput for transmitting audio data to end device 203. It may then determine the maximum number of separate audio objects that can be transmitted (e.g., in addition to one ambisonic audio mix) and proceed to adapt a distance threshold to ensure that the maximum number is not exceeded.
[0149] As another example, if bandwidth is limited and the number of objects designated as nearby is relatively large, the distance threshold may be reduced so that fewer objects are designated as nearby. In addition, the bandwidth freed up by moving previously designated nearby objects may increase the Ambisonic (HOA) order so that previously designated nearby objects are represented "better" in the Ambisonic audio mix.
[0150] In many embodiments, the designation of an audio object may depend on multiple listener poses, for example, in a game scenario. Specifically, the designator 305 may be configured to determine a threshold, or equivalently a distance scale, depending on the multiple listener poses. Thus, in some embodiments, at least one of the distance scale or the threshold depends on the multiple listener poses.
[0151] As an example, in some embodiments, the edge device 201 may receive listener poses from two or more end devices 203 and proceed to generate a single audio data signal that can be sent back to multiple end devices 203. In such cases, the designation of an audio object may take into account multiple listener poses. For example, a distance threshold may be determined for each listener pose, and an audio object may be designated as a nearby audio object if its distance to any of the listener poses is less than the corresponding threshold. Thus, in such embodiments, an audio object may be designated as a nearby object if it is close to any of the listener poses, and may only be included in the ambisonic audio mix if it is far from all of the listener poses.
[0152] The distance threshold (or equivalent distance measure) for an audio object can be determined as a function of many different parameters, including one or more of the following: • Distance of audio object relative to listener's posture • The (3D) angle (angle of incidence) of the audio object relative to the listener's posture. • Diffusivity of audio objects (objects with a range relative to a point source) • Position of other audio objects Audio objects placed at the same location (small angular distance) can be rendered as a single audio object, rather than remaining in the spatial audio mix. • Volume / temporal / spatial masking • Speed of audio objects • Velocity vector from user to object (prediction, potential change in angular distance) • Animated (moving) point source audio object • Number of input objects • Metadata for audio objects including the following: • The type of audio object (i.e., if an audio object represents a user, it is highly likely that it should be rendered as a separate audio object). • Whether a visual component is connected to the audio object • Size of audio objects • Dynamic balance between ambisonics order and the number of audio objects • Bitrate allocation to an audio object during the transition from an ambisonic mix to another audio object. • Use of masking • End device capabilities (e.g., processing power)
[0153] Figure 7 is a block diagram showing an exemplary processor 700 according to an embodiment of the present disclosure. The processor 700 may be used to implement one or more processors that implement the aforementioned devices or elements thereof. The processor 700 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA) programmed to form a processor, a graphics processing unit (GPU), an application-specific integrated circuit (ASIC) designed to form a processor, or a combination thereof.
[0154] The processor 700 may include one or more cores 702. A core 702 may include one or more arithmetic logic units (ALUs) 704. In some embodiments, the core 702 may include, in addition to or instead of, a floating-point logic unit (FPLU) 706 and / or a digital signal processing unit (DSPU) 708.
[0155] The processor 700 may include one or more registers 812 that are communicatively coupled to the core 702. The registers 712 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 712 may be implemented using static memory. The registers may provide data, instructions, and addresses to the core 702.
[0156] In some embodiments, the processor 700 may include one or more levels of cache memory 710 communicatively coupled to the core 702. The cache memory 710 may provide computer-readable instructions to the core 702 for execution. The cache memory 710 may provide data for processing by the core 702. In some embodiments, computer-readable instructions may be provided to the cache memory 710 by local memory, for example, local memory attached to an external bus 716. The cache memory 710 may be implemented in metal-oxide-semiconductor (MOS) memory of any suitable cache memory type, such as static random-access memory (SRAM), dynamic random-access memory (DRAM), and / or any other suitable memory technology.
[0157] The processor 700 may include a controller 714 that can control inputs to the processor 700 from other processors and / or components included in the system and / or outputs from the processor 700 to other processors and / or components included in the system. The controller 714 can control data paths in the ALU 704, FPLU 706, and / or DSPU 708. The controller 714 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 714 can be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.
[0158] The registers 712 and cache 710 can communicate with the controller 714 and core 702 via internal connections 720A, 720B, 720C, and 720D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection techniques.
[0159] Inputs and outputs to the processor 700 may be provided via a bus 716, which may include one or more conductive wires. The bus 716 may be communicatively coupled to one or more components of the processor 700, such as a controller 714, a cache 710, and / or registers 712. The bus 716 may be coupled to one or more components of the system.
[0160] The bus 716 may be coupled to one or more external memories. The external memory may have read-only memory 732. ROM 732 may be a mask ROM, an electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may have random access memory 733. RAM 733 may be static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memory may have electrically erasable programmable read-only memory (EEPROM) 735. The external memory may have flash memory 734. The external memory may have a magnetic storage device such as a disk 736. In some embodiments, the external memory may be included in the system.
[0161] For clarification, the above description will be understood to have illustrated embodiments of the invention with reference to different functional circuits, units, and processors. However, it will be apparent that any appropriate distribution of functions between different functional circuits, units, or processors can be used without departing from the invention. For example, functions shown to be performed by separate processors or controllers can also be performed by the same processor or controller. Thus, references to specific functional units or circuits should be considered only as references to appropriate means for providing the described functions, and not as indicating a strict logical or physical structure or organization.
[0162] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, the present invention can be implemented at least partially as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the present invention can be implemented physically, functionally, and logically in any suitable manner. In fact, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention can be implemented in a single unit, or physically and functionally distributed among different units, circuits, and processors.
[0163] Although the present invention has been described in relation to several embodiments, it is not intended to be limited to any particular form described herein. Rather, the scope of the present invention is limited only by the appended claims. In addition, while some features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term “comprising” does not preclude the existence of other elements or steps.
[0164] Furthermore, although listed individually, multiple means, elements, circuits, or method steps may be implemented, for example, by a single circuit, unit, or processor. In addition, individual features may be included in different claims, but these may be advantageously combined as they may be, and inclusion in different claims does not mean that the combination of features is unfeasible and / or unfavorable. Also, the inclusion of a feature in one category of claims does not mean that it is limited to that category, but rather that the feature is equally applicable to other claim categories as needed. Furthermore, the order of features in a claim does not mean that the feature must operate in a specific order, and in particular, the order of individual steps in a method claim does not mean that the steps must be performed in a specific order. Rather, the steps can be performed in any suitable order. In addition, singular references do not exclude plurals. Accordingly, references such as "a," "an," "first," "second," etc., do not exclude plurals. Reference numerals in a claim are provided merely as clear examples and should not be construed as limiting the scope of the claim in any way.
[0165] Generally, examples of approaches are shown by the following embodiments.
[0166] Embodiments:
[0167] Embodiment 1. A receiver (301) configured to receive multiple audio elements representing an audio scene, wherein the audio elements include several audio objects, and each audio object is linked to a location within the audio scene, and the receiver and A listener attitude receiver (303) configured to receive listener attitude indications, A designator (305) configured to designate at least one of a plurality of audio objects as a close audio object or a far audio object in response to a comparison between a distance scale indicating the distance between the posture of a first audio object and the listener posture and a threshold, wherein the distance scale for the first audio object is such that it is designated as a far audio object exceeding the threshold, and the distance scale for the first audio object is such that it is designated as a close audio object not exceeding the threshold. An audio mix generator (307) configured to generate an ambisonic audio mix from a first plurality of audio elements, wherein the first plurality of audio elements include the first audio object when the first audio object is specified as a distant audio object, and the audio mix generator A data generator (311) configured to generate an audio data signal including an ambisonic audio mix, and configured to include a first audio object in the audio data signal when the first audio object is designated as a nearby audio object, An audio device having the following features.
[0168] Embodiment 2. The audio device of Embodiment 1, wherein the designator (305) is configured to designate the first audio object as a nearby audio object or a distant audio object according to the volume scale of the first audio object.
[0169] Embodiment 3. An audio device according to any of the above embodiments, wherein the designator (305) is configured to designate the first audio object as a near audio object or a far audio object depending on the direction from the listener's posture to the position of the first audio object.
[0170] Embodiment 4. An audio device according to any of the above embodiments, wherein the designator (305) is configured to designate the first audio object as a nearby audio object or a distant audio object depending on the trajectory of the position of the first audio object in the audio scene.
[0171] Embodiment 5. An audio device according to any of the above embodiments, wherein the designator (305) is configured to designate a first audio object as a close audio object or a far audio object depending on the order of the ambisonic audio mix.
[0172] Embodiment 6. An audio device according to any of the above embodiments, wherein the designator (305) is configured to designate the first audio object as a close audio object or a far audio object depending on the distance between the position of the first audio object and the position of the second audio object.
[0173] Embodiment 7. An audio device according to any of the above embodiments, wherein a data generator (311) is configured to transmit audio data signals to a remote device over a communication link, and a designator (305) is configured to designate a first audio object as a nearby audio object or a distant audio object depending on the data transfer characteristics of the communication link.
[0174] Embodiment 8. An audio device according to any of the above embodiments, wherein the audio mix generator (307) is configured to change the data rate for a first audio object when transitioning between a state designated as a nearby object and a state designated as a far-away object.
[0175] Embodiment 9. An audio device according to any of the above embodiments, wherein a listener attitude receiver (303) is configured to receive a plurality of listener attitudes, a designator (305) is configured to determine a distance scale according to the plurality of listener attitudes, and a data generator (311) is configured to transmit an audio data signal to a plurality of remote devices.
[0176] Embodiment 10. An audio system having an audio device (201) and a rendering device (203) according to any of the embodiments described above, wherein a data generator (311) is configured to transmit audio data signals to the rendering device (203), The rendering device (203) is, A receiver (401) configured to receive audio data signals from an audio device (201), A renderer (403) configured to generate a binaural audio signal representing an audio scene, wherein the binaural audio signal includes a first audio mix and contributions from a first audio object if present in the audio data signal, An audio system that has the following features.
[0177] Embodiment 11. The audio system of Embodiment 10, wherein the rendering device (203) is a headset having an audio transducer configured to play back a binaural rendering signal.
[0178] Embodiment 12. An audio system according to Embodiment 10 or 11, wherein the designator (305) is configured to designate a first audio object as a nearby audio object or a distant audio object depending on the characteristics of the rendering device.
[0179] Embodiment 13. A step of receiving a plurality of audio elements representing an audio scene, wherein the audio elements include a plurality of audio objects, and each audio object is linked to a position within the audio scene. Steps to receive listener attitude indications, A step of designating at least one first audio object from among multiple audio objects as a close audio object or a far audio object based on a comparison between a distance scale indicating the distance between the posture of the first audio object and the listener posture and a threshold, wherein the distance scale for the first audio object is greater than the threshold and designated as a far audio object, and the distance scale for the first audio object is less than the threshold and designated as a close audio object. A step of generating an ambisonic audio mix from a first set of audio elements, wherein the first set of audio elements includes a first audio object when the first audio object is designated as a distant audio object. A step of generating an audio data signal that includes an ambisonic audio mix, The steps include including the first audio object in the audio data signal when the first audio object is designated as a nearby audio object, A method for operating an audio device having the following characteristics.
[0180] Embodiment 14. A method for operating an audio device (201), wherein the audio device performs the method described in claim 13, transmits an audio data signal to a rendering device (203), and the rendering device (203) The steps include receiving an audio data signal from an audio device (201), A step of generating a binaural audio signal representing an audio scene, wherein the binaural audio signal includes a first audio mix and contributions from a first audio object if present in the audio data signal. How to do it.
[0181] More specifically, the present invention is defined by the appended claims.
Claims
1. A receiver configured to receive multiple audio elements representing an audio scene, wherein the audio elements include several audio objects, each audio object being linked to a location within the audio scene, and the receiver A listener attitude receiver configured to receive listener attitude indications, A designator configured to designate at least one of several audio objects as a close audio object or a far audio object based on a comparison between a distance scale indicating the distance between the posture of the first audio object and the listener posture and a threshold, wherein the distance scale for the first audio object is such that it is designated as a far audio object exceeding the threshold, and the distance scale for the first audio object is such that it is designated as a close audio object not exceeding the threshold. An audio mix generator configured to generate an ambisonic audio mix from a first plurality of audio elements, wherein the first plurality of audio elements include the first audio object when the first audio object is designated as a distant audio object, and the audio mix generator A data generator configured to generate an audio data signal including the ambisonic audio mix, wherein the data generator is configured to include the first audio object in the audio data signal when the first audio object is designated as a nearby audio object, An audio device having the following features.
2. The audio device according to claim 1, wherein the designator is configured to designate the first audio object as a nearby audio object or a distant audio object according to the volume scale for the first audio object.
3. The audio apparatus according to claim 1 or 2, wherein the designator is configured to designate the first audio object as a nearby audio object or a distant audio object depending on the direction from the listener's posture to the position of the first audio object.
4. The audio apparatus according to claims 1 to 3, wherein the designator is configured to designate the first audio object as a nearby audio object or a distant audio object according to the trajectory of the position of the first audio object in the audio scene.
5. The audio apparatus according to any one of claims 1 to 4, wherein the designator is configured to designate the first audio object as a close audio object or a far audio object according to the order of the ambisonic audio mix.
6. The audio apparatus according to claims 1 to 5, wherein the designator is configured to designate the first audio object as a close audio object or a far audio object depending on the distance between the position of the first audio object and the position of the second audio object.
7. The audio device according to claims 1 to 6, wherein the data generator is configured to transmit the audio data signal to a remote device over a communication link, and the designator is configured to designate the first audio object as a nearby audio object or a distant audio object according to the data transfer characteristics of the communication link.
8. The audio device according to any one of claims 1 to 7, wherein the audio mix generator is configured to change the data rate for the first audio object when transitioning between a state designated as a nearby object and a state designated as a distant object.
9. The audio apparatus according to any one of claims 1 to 8, wherein the listener attitude receiver is configured to receive a plurality of listener attitudes, the designator is configured to determine the distance scale according to the plurality of listener attitudes, and the data generator is configured to transmit the audio data signal to a plurality of remote devices.
10. An audio system having an audio device and a rendering device according to any one of claims 1 to 4, wherein the data generator is configured to transmit the audio data signal to the rendering device, The rendering device, A receiver configured to receive the audio data signal from the aforementioned audio device, A renderer configured to generate a binaural audio signal representing the aforementioned audio scene, wherein the binaural audio signal includes contributions from the first audio object if present in the first audio mix and the audio data signal, An audio system that has the following features.
11. The audio system according to claim 10, wherein the rendering device is a headset having an audio transducer configured to reproduce the binaural rendering signal.
12. The audio system according to claim 10 or 11, wherein the designator is configured to designate the first audio object as a nearby audio object or a distant audio object depending on the characteristics of the rendering device.
13. A step of receiving multiple audio elements representing an audio scene, wherein the audio elements include several audio objects, and each audio object is linked to a position within the audio scene. Steps to receive listener attitude indications, A step of designating at least one of several audio objects as a close audio object or a far audio object based on a comparison between a distance scale indicating the distance between the posture of the first audio object and the listener posture and a threshold, wherein the distance scale for the first audio object is greater than the threshold and designated as a far audio object, and the distance scale for the first audio object is less than the threshold and designated as a close audio object. A step of generating an ambisonic audio mix from a first plurality of audio elements, wherein the first plurality of audio elements include the first audio object when the first audio object is designated as a distant audio object; The steps include generating an audio data signal that includes the aforementioned ambisonic audio mix, The steps include including the first audio object in the audio data signal when the first audio object is designated as a nearby audio object, A method for operating an audio device having the following characteristics.
14. A method for operating an audio device, wherein the audio device performs the method described in claim 13, transmits the audio data signal to a rendering device, and the rendering device The steps include receiving the audio data signal from the audio device, A step of generating a binaural audio signal representing the audio scene, wherein the binaural audio signal includes contributions from the first audio mix and, if present in the audio data signal, from the first audio object; How to do it.
15. A computer program having computer program code means configured to perform all steps of the method according to claim 13 or 14 when executed on a computer.