Generating audio data signal

By specifying audio objects as near or far objects in the audio device and generating adapted audio data signals, the suboptimal problem of audio rendering on low-complexity devices is solved, achieving a high-quality and flexible audio experience.

CN121970375APending Publication Date: 2026-05-01KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2024-10-01
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing audio signal rendering methods struggle to provide a high-quality, flexible audio experience on low-complexity and resource-constrained terminal devices, especially in virtual reality and augmented reality applications, resulting in suboptimal audio quality and responsiveness.

Method used

An audio device and method are used to receive multiple audio elements of a 3D audio scene through a receiver. An assignor is used to assign audio objects as near or far objects based on the listener's posture and distance threshold. An audio mixer generates surround sound audio mixes and an audio data signal is generated through a data generator to adapt to the rendering requirements of the terminal device.

Benefits of technology

It improves audio quality, reduces rendering complexity, enhances responsiveness to user movement and resource utilization, and provides better audio compatibility and an immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970375A_ABST
    Figure CN121970375A_ABST
Patent Text Reader

Abstract

An audio device comprises a receiver (301) that receives an audio element comprising a number of audio objects linked with locations in an audio scene. A listener gesture receiver (303) to receive an indication of a listener gesture, and a designator (305) to designate the audio object as a short range audio object or a remote audio object depending on a comparison of a distance metric to a threshold, the distance metric being indicative of a distance between a gesture of the first audio object and the listener gesture. An audio mix generator (307) generates a surround sound audio mix from a first plurality of audio elements, the first plurality of audio elements including one or more audio objects designated as remote audio objects. A data generator (311) generating an audio data signal comprising surround sound audio mixing, and further comprising an audio object designated as a short range audio object. The audio device may be an edge device that operates with an audio end device to provide split rendering of audio across the device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to generating audio data signals, and in some cases to rendering audio data signals, and more specifically, but not exclusively, to generating such signals, for example, to support segmented rendering applications. Background Technology

[0002] The variety and scope of experiences based on audiovisual content have increased significantly in recent years, as new services and ways of utilizing and consuming such content are constantly being developed and introduced. Specifically, many spatial and interactive services, applications, and experiences are being developed to provide users with more immersive and engaging experiences.

[0003] Examples of such applications are Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) applications (often referred to as Extended Reality (XR) applications), which are rapidly becoming mainstream, with numerous solutions targeting the consumer market. Standards are also being developed by various standardization bodies. These standardization activities are actively developing standards for all aspects of VR / AR / MR / XR systems, including, for example, streaming, broadcasting, rendering, etc.

[0004] VR applications tend to provide a user experience that corresponds to the user in a different world / environment / scene, while AR (including Mixed Reality) applications tend to provide a user experience that corresponds to the user in the current environment, but with added supplementary information or virtual objects or information. Therefore, VR applications tend to provide a fully immersive, synthetically generated world / scene, while AR applications tend to provide a partially synthetic world / scene that overlays the user's actual real-world environment. However, these terms are often used interchangeably and have a high degree of overlap. In the following text, the term Virtual Reality / VR will be used to refer to both Virtual Reality and Augmented Reality.

[0005] VR applications typically provide users with a virtual reality experience, allowing them to move (relatively) freely within a virtual environment and dynamically change their posture (position and / or orientation). Typically, such VR applications are based on a 3D model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well-known, for example, in gaming applications such as first-person shooters for computers and consoles.

[0006] Beyond visual rendering, most VR (and more generally XR) applications further provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience, in which the perceived audio source originates from a location corresponding to the location of a corresponding object in the visual scene (including both currently visible objects and currently invisible or partially visible objects (e.g., behind the user)). Therefore, the audio and video scenes are preferably perceived as consistent and provide a complete spatial experience.

[0007] For audio, headphone reproduction using binaural audio rendering technology is widely adopted. In many scenarios, headphone reproduction enables a highly immersive and personalized user experience. Using head tracking allows the rendering to respond to the user's head movements, which greatly enhances the sense of immersion.

[0008] To generate appropriate audio, rendering devices are provided with spatial audio data representing the audio scene. However, audio sources can be represented in many different ways, including as audio channel signals, audio objects, diffuse non-spatial-specific audio sources (such as background noise), ambisonic audio, and so on. In practice, it is becoming increasingly common to use multiple different audio representations to provide an audio scene description, reflecting the different types of audio sources that can be represented. However, such an approach can increase complexity and resource requirements, and often results in processing-intensive rendering algorithms.

[0009] These issues are problematic, particularly in conflict with the desire to render audiovisual content with very low complexity and resources. Specifically, there is a growing expectation for end-user devices that are low-cost, small in size, lightweight, low-complexity, and have low computational resources. For example, wearable devices with these characteristics are becoming increasingly common. Specifically, to support various XR applications, there is a need to provide small and lightweight XR headsets, such as relatively small and lightweight XR glasses.

[0010] To enable or facilitate the use of such a device, a segmented rendering approach has been proposed, in which a portion of the (final) rendering process is performed by the terminal device, while other, more computationally demanding portions of the rendering are performed by another device (called an edge device). The edge device may perform a significant portion of the rendering process and may generate an audio signal, which is then sent to the terminal device for final adaptation.

[0011] Edge devices are typically more complex and resource-rich than end devices, and therefore can implement more sophisticated rendering algorithms and functionalities. For example, edge devices can be mobile phones, game consoles, computers, remote servers, etc., while end devices can be rendering and audio reproduction devices worn by users, such as XR headsets / glasses.

[0012] As an example, it has been proposed that an edge device renders the received audio data to generate rendered audio signals, which are then sent to a terminal device that can perform additional simple operations (such as simple translation) on these audio signals to adapt to user movement.

[0013] However, while these methods can provide the desired application and operation in many situations, current methods are often not optimal. Specifically, in many cases, current methods can provide suboptimal audio quality, suboptimal response to user movement, suboptimal resource usage and distribution, and so on.

[0014] Therefore, improved methods for the distribution and / or rendering and / or processing of audio signals are particularly advantageous for virtual reality / augmented reality / mixed reality / extended reality experiences / applications. Specifically, methods that allow for improved operation, increased flexibility, reduced complexity, ease of implementation, improved user experience, improved audio quality, improved adaptability to audio reproduction capabilities, ease and / or improved adaptability to changes in listener position / or orientation (e.g., virtual listener position / or orientation), improved resource requirements / processing distribution for split rendering methods, improved extended reality experiences, and / or improved performance and / or operation will be advantageous. Summary of the Invention

[0015] Therefore, the present invention preferably mitigates, alleviates, or eliminates one or more of the above-mentioned disadvantages, either alone or in any combination.

[0016] According to one aspect of the present invention, an audio device is provided, comprising: a receiver arranged to receive a plurality of audio elements representing a three-dimensional audio scene, the audio elements including a plurality of audio objects, each audio object being associated with a position in the three-dimensional audio scene; a listener posture receiver arranged to receive an indication of a listener posture; and a designator arranged to designate at least a first audio object among the plurality of audio objects as a near-range audio object or a far-range audio object based on a comparison of a distance metric and a threshold, the distance metric indicating the distance between the posture of the first audio object and the listener posture, for which the far-range audio object is designated... The distance metric of the first audio object exceeds a threshold, and the distance metric for the first audio object designated as a near-range audio object does not exceed a threshold; an audio mix generator is arranged to generate a surround sound audio mix from a first plurality of audio elements, wherein the first plurality of audio elements include the first audio object if the first audio object is designated as a remote audio object; a data generator is arranged to generate an audio data signal including the surround sound audio mix; and wherein the data generator is arranged to include the first audio object in the audio data signal if the first audio object is designated as a near-range audio object.

[0017] This method allows for the generation of improved output audio signals. It can provide improved audio quality in many embodiments and scenarios, and, for example, can reduce the complexity of rendering audio using rendering devices based on audio data signals.

[0018] This method particularly allows for improved split rendering, where the rendering of an audio scene can be split across multiple devices. The audio device can provide audio data signals that allow for improved distribution of functionality, complexity, and computational resource usage. The audio device can generate audio data signals that offer an improved trade-off between the additional requirements of rendering a single audio object and rendering multiple sources contained within a single surround sound audio mix.

[0019] Posture can be position and / or orientation. A listener's posture can be listener position, orientation, or both.

[0020] The designator can be arranged to designate at least a first audio object among a plurality of audio objects as a near-field audio object or a far-field audio object, depending on whether a distance metric used to indicate the distance between the pose of the first audio object and the pose of the listener exceeds a threshold. The designator can be arranged to designate at least a first audio object as a near-field audio object if the distance metric used to indicate the distance between the pose of the first audio object and the pose of the listener is below a threshold, and to designate at least a first audio object as a far-field audio object if the distance metric exceeds a threshold (and if the distance metric is the same as the threshold, it is typically designated as a far-field or near-field object depending on the specific embodiment).

[0021] The listener's posture can indicate the posture within the audio scene. The audio scene can be a three-dimensional audio scene.

[0022] The audio mix generator can be configured to generate a surround sound audio mix that includes the first audio object only if the first audio object is designated as a remote audio object. The data generator can be configured to not include the first audio object in the audio data signal if the first audio object is designated as a remote audio object.

[0023] According to any optional feature of the invention, the designator is arranged to designate the first audio object as a near-field audio object or a far-field audio object depending on the loudness metric for the first audio object.

[0024] According to any optional feature of the invention, the designator is arranged to designate the first audio object as a near-field audio object or a far-field audio object, depending on the direction from the listener's posture to the posture of the first audio object.

[0025] According to any optional feature of the invention, the designator is arranged to designate the first audio object as a near-field audio object or a far-field audio object, depending on the trajectory of the first audio object's position in the audio scene.

[0026] According to any optional feature of the invention, the designator is arranged to designate the first audio object as a near-field audio object or a far-field audio object, depending on the order of the surround sound audio mix.

[0027] According to any optional feature of the invention, the designator is arranged to designate the first audio object as a near-field audio object or a far-field audio object, depending on the distance between the position of the first audio object and the position of the second audio object.

[0028] According to any optional feature of the invention, the data generator is arranged to transmit audio data signals to a remote device via a communication link; and the designator is arranged to designate the first audio object as a near-field audio object or a remote audio object, depending on the data transmission properties of the communication link.

[0029] In many embodiments, these features can provide improved and / or enhanced operation and / or performance. Features can result in the generation of audio data signals that offer an improved trade-off between the additional requirements of rendering a single audio object and rendering multiple sources contained within a single surround sound audio mix.

[0030] According to an optional feature of the invention, the audio mixer is arranged to change the data rate of the first audio object when switching between being designated as a near object and being designated as a far object.

[0031] In many embodiments, this can provide improved and / or enhanced operation and / or performance.

[0032] According to an optional feature of the invention, a listener gesture receiver is arranged to receive multiple listener gestures, a designator is arranged to determine a distance measurement based on the multiple listener gestures, and a data generator is arranged to transmit audio data signals to multiple remote devices.

[0033] In many embodiments, this can provide improved and / or enhanced operation and / or performance.

[0034] According to an optional feature of the invention, an audio system is provided, comprising an audio device and a rendering device as described above, the rendering device comprising: a receiver arranged to receive an audio data signal from the audio device; and a renderer arranged to generate a binaural audio signal representing an audio scene, wherein the binaural audio signal includes contributions from a first audio mix and the first audio object if a first audio object is present in the audio data signal.

[0035] According to an optional feature of the invention, the rendering device is a head-mounted device that includes an audio transducer arranged to reproduce binaural rendering signals.

[0036] According to any optional feature of the invention, the designator is arranged to designate the first audio object as a near-field audio object or a far-field audio object, depending on the properties of the rendering device.

[0037] In many embodiments, this can provide improved and / or enhanced operation and / or performance.

[0038] According to one aspect of the present invention, an operating method for an audio device is provided, the method comprising: receiving a plurality of audio elements representing a three-dimensional audio scene, the audio elements including a plurality of audio objects, each audio object being linked to a position in the three-dimensional audio scene; receiving an indication of a listener's posture; designating at least a first audio object among the plurality of audio objects as a near-field audio object or a far-field audio object based on a comparison of a distance metric indicating a distance between the posture of a first audio object and the listener's posture, wherein the distance metric for the first audio object designated as a far-field audio object exceeds the threshold, and the distance metric for the first audio object designated as a near-field audio object does not exceed the threshold; generating a surround sound audio mix from the first plurality of audio elements, wherein if the first audio object is designated as a far-field audio object, the first plurality of audio elements include the first audio object; and generating an audio data signal including the surround sound audio mix; wherein if the first audio object is designated as a near-field audio object, the first audio object is included in the audio data signal.

[0039] According to an optional feature of the invention, the audio device performs the method of claim 13 and sends an audio data signal to the rendering device; and the rendering device performs the following steps: receiving the audio data signal from the audio device; and generating a binaural audio signal for representing an audio scene, wherein if a first audio object is present in the audio data signal, the binaural audio signal includes contributions from a first audio mix and the first audio object.

[0040] These and other aspects, features, and advantages of the present invention will become apparent from the embodiments described below, and will be explained with reference to the embodiments described. Attached Figure Description

[0041] Embodiments of the present invention are described by way of example only, with reference to the accompanying drawings, in which: Figure 1 An example of a client-server based extended reality system is shown; Figure 2 Examples of elements of an audio rendering arrangement according to some embodiments of the present invention are shown; Figure 3 Examples of elements of an audio device according to some embodiments of the present invention are shown; Figure 4 Examples of elements of an audio rendering apparatus according to some embodiments of the present invention are shown; Figure 5 An example of an audio object in an audio scene is shown; Figure 6 An example of an audio object in an audio scene is shown; Figure 7 Some elements of a processor for implementing an apparatus are shown in some embodiments of the invention. Detailed Implementation

[0042] The following description will focus on extended reality applications where audio is rendered to reflect the user's posture within the audio scene to provide an immersive user experience. Typically, audio rendering can accompany image rendering to provide the user with a complete audiovisual experience. However, it is understood that the described approach can be used for many other applications.

[0043] Extended reality (including virtual reality, augmented reality, and mixed reality) experiences that allow users to move around in virtual or augmented worlds are becoming increasingly popular, and services are being developed to improve such applications. In many of these schemes, visual and audio data can be dynamically generated to reflect the user's (or viewer's) current posture.

[0044] In this art, the terms “placement” and “pose” are used as general terms for position and / or orientation / direction. For example, a combination of position and orientation / direction of an object, camera, head, or view can be referred to as a pose or placement. Therefore, a placement or pose indication can include up to six values / components / degrees of freedom, where each value / component typically describes a separate attribute of the position / location or orientation / direction of the corresponding object. Of course, in many cases, placement or pose can be represented with fewer components, for example, if one or more components are considered fixed or unrelated (e.g., if all objects are considered to be at the same height and have a horizontal orientation, four components can provide a complete representation of the object's pose). In the following, the term “pose” is used to refer to a position and / or orientation that can be represented by one to six values ​​(corresponding to the maximum possible degrees of freedom). The term “pose” may be replaced by the terms “at least one of position and orientation” or “position and / or orientation.” A listener’s pose can be a listener position and / or listener orientation.

[0045] Many XR applications are based on pose, which has the maximum degree of freedom—three degrees of freedom for position and three degrees of freedom for orientation—resulting in a total of six degrees of freedom. Therefore, pose can be represented by a set or vector of six values ​​representing the six degrees of freedom, and thus, the pose vector can provide a three-dimensional position and / or three-dimensional orientation indication. However, it should be understood that in other embodiments, pose can be represented with fewer values ​​(rotation can also be in the form of quaternions or rotation matrices, and therefore may have more than six values).

[0046] Systems or entities that offer the maximum degrees of freedom to the viewer are often referred to as having six degrees of freedom (6DoF). Many systems and entities offer only orientation or position, and these are often referred to as having three degrees of freedom (3DoF – which is often used to represent a method of using poses with variable orientation and fixed position).

[0047] Typically, virtual reality applications generate 3D output in the form of separate view images for the left and right eyes. This information can then be presented to the user through appropriate units (e.g., separate left and right eye displays, typically for a VR headset). In other embodiments, one or more view images may be presented, for example, on an autostereoscopic display, or more precisely, in some embodiments, only a single 2D image is generated (e.g., using a conventional 2D display).

[0048] Similarly, for a given viewer / user / listener pose, an audio representation of the scene can be provided. Audio scenes are typically rendered to provide a spatial experience, where the audio source is perceived as originating from a desired location. Since the audio source is static within the scene, changes in the listener's pose cause a change in the audio source's position relative to the user's pose. Therefore, the spatial perception of the audio source can be altered to reflect its new position relative to the user. Audio rendering can thus adapt to the listener's pose.

[0049] Listener gesture input can be determined in different ways in different applications. In many embodiments, the user's physical movement can be tracked directly. For example, a camera used to survey the user's area can detect and track the user's head (or even eyes (eye tracking)). In many embodiments, the user can wear a VR headset that can be tracked by external and / or internal units. For example, the headset can include accelerometers and gyroscopes for providing information about the movement and rotation of the headset (and therefore the head). In some instances, the VR headset can send signals or include (e.g., visual) identifiers that enable external sensors to determine the VR headset's position and orientation.

[0050] In many systems, VR / scene data, and particularly audio data representing audio scenes, can be provided from remote devices or servers. For example, a remote server can generate audio data representing an audio scene and can send audio signals corresponding to, and location information indicating the location of, for example, dynamically changing locations of moving objects: audio components / objects / channels, or other audio elements corresponding to different audio sources in the audio scene. Audio signals / elements can include elements associated with a specific location, but can also include elements for more dispersed or diffused audio sources. For example, audio elements representing general (non-localized) background sounds, ambient sounds, diffuse reverberation, etc., can be provided.

[0051] The local VR device can then render audio elements appropriately, specifically by applying appropriate binaural processing to reflect the relative position of the audio source for the audio components.

[0052] Similarly, remote devices can generate visual / video data to represent a visual-audio scene, and can send visual scene components / objects / signals, or other visual elements corresponding to different objects in the visual scene, as well as location information indicating the location of these items (e.g., dynamically changing for moving objects). Visual items can include elements associated with a specific location, but can also include video items for more distributed sources.

[0053] In some embodiments, visual items may be provided as separate and distinct items, such as descriptions of individual scene objects (e.g., size, texture, opacity, reflectivity, etc.). Alternatively or additionally, visual items may be represented as part of an overall model of the scene, such as including descriptions of different objects and their relationships.

[0054] For VR services, in some embodiments, the central server can generate audiovisual data to represent a three-dimensional scene, and in particular, audio can be represented by several audio signals that represent audio sources in the scene, and then the local client / device can render these audio signals.

[0055] Figure 1 An example of a VR / XR system is shown, in which a central server 101 communicates with several remote clients 103, for example, via a network 105 (such as the Internet). The central server 101 can be configured to simultaneously support a potentially large number of remote clients 103.

[0056] In many situations, this approach offers improved trade-offs between factors such as the complexity and resource requirements of different devices, and communication requirements. For example, scene data can be sent only once or relatively infrequently, and the local rendering device (remote client 103) receives the viewer's posture and processes the scene data locally to render audio and / or video to reflect changes in the viewer's posture. This approach provides an efficient system and an engaging user experience. For example, it can substantially reduce the required communication bandwidth while providing a low-latency, real-time experience, and allows scene data to be centrally stored, generated, and maintained. For example, it is suitable for applications that deliver VR experiences to multiple remote devices.

[0057] In some cases, remote clients may include multiple devices arranged to interoperate to provide rendering of audiovisual data. Specifically, such as Figure 2 As shown, the rendering setup (which may specifically correspond to a remote client) may include a first device 201, also referred to as an edge device 201, which receives audiovisual data for describing the scene. In order to generate an audio representation, the first / edge device 201 accordingly receives audio data for describing the audio scene.

[0058] Edge device 201 is arranged to process received audio data to generate intermediate audio data, which is then sent to end device 203. End device 203 is arranged to process the intermediate audio data to generate rendered output audio signals. These output signals may specifically be binaural signals for the listener's left and right ears, respectively. This method therefore employs a split rendering approach, where the rendering of the audio (scene) is split across more than one device. In many embodiments, end device 203 may include an audio reproduction unit, specifically an audio transducer (and typically has one audio transducer for the user's right ear and one audio transducer for the user's left ear).

[0059] Edge device 201 may specifically be a mobile phone, computer, game console, laptop, tablet, etc., and terminal device 203 may typically be a playback device worn by the user, such as XR glasses and / or headphones.

[0060] Figure 3 Examples of some elements of an audio device are shown, the device being arranged to generate audio data signals from audio data used to describe an audio scene (and typically for a three-dimensional audio scene). Specifically, the device may be... Figure 2 The edge device 201 will be described with reference to it. For example, Figure 2 A key question regarding this approach is how to distribute functionality and processing across different devices, and what data to send from edge devices to end devices. These considerations include not only the computational load across different devices but also other parameters such as the impact on communication between devices (including the effects of communication errors and latency). For example, the responsiveness of rendered audio to rapid changes in listener posture is often highly dependent on communication latency. Therefore, for speed and responsiveness, it is desirable to perform processing primarily at the end device. However, from the perspective of resources, size, complexity, etc., it is often desirable to perform processing primarily on edge devices. Therefore, trade-offs between conflicting requirements are often crucial to the performance of the approach.

[0061] Figure 3 The edge device 201 includes a receiver arranged to receive audio data used to describe a three-dimensional audio scene. The audio data includes multiple audio elements, each providing a representation of audio / sound in the audio scene. Audio elements can be audio signals representing audio sources, diffuse noise, ambient sound, multiple sources, etc. In many cases, audio elements can include a range of different types of audio elements, including audio objects, audio channel signals, surround sound audio signals / mixes, etc.

[0062] In some cases, one or more audio elements may include one or more audio channel signals. For example, such channel signals may be provided for a predetermined and / or nominal location (e.g., nominal speaker location). In these cases, each channel has information about a specific part of the mix configured relative to a specified speaker.

[0063] An audio element specifically comprises several audio objects. Each audio object provides a description / representation of the audio, and typically provides a description / representation of the audio from an audio source. Each audio object is linked to a location within the audio scene, and thus provides both audio and spatial data. Location / orientation can be, for example, absolute (global) with respect to the scene, relative to another object within the scene, or relative to the listener. Typically, an audio object provides audio and location data for an audio source within the audio scene. Object-based immersive audio can add audio streams and metadata to the decoder, giving instructions on how to place each stream within the 3D sound field.

[0064] In many scenarios, audio elements may further include one or more surround sound signals / mixes.

[0065] Surround sound, and especially high-order surround sound (HOA), is a method for recording and reproducing a three-dimensional sound field (scene-based audio). Unlike typical channel-oriented transmission methods, this system focuses on the reproduction of the entire sound field at the listening position and does not require a predetermined number of speaker positions for sound reproduction. The relevant speaker signal for each speaker position is calculated by mathematically deriving the transmission values ​​of sound pressure and sound velocity. In the basic version known as first-order surround sound, surround sound can be understood as a three-dimensional extension of M / S (center / side) stereo, adding additional difference channels for height and depth. The resulting signal set is called the B-format. The sound information is transmitted in four channels: W, X, Y, and Z. The component W contains only sound pressure, which is typically recorded using non-directional or omnidirectional microphones. The signals X, Y, and Z are directional components on the corresponding spatial axes. They can be recorded using microphones whose figure-eight shape is aligned with the corresponding axes. A simpler surround sound method uses the A-format with four cardioid microphones, which can then be converted to the more practical B-format.

[0066] The purpose of this process is to reconstruct the recorded sound pressure and associated sound direction vectors based on these signals at the listener's listening position.

[0067] For audio objects that are part of a surround sound mix, the surround sound order determines the amount of spatial blur of a point source. The higher the order, the lower the spatial blur and the higher the spatial directionality.

[0068] Edge device 301 further includes a listener posture receiver 303 arranged to receive indications of a listener's posture. Listener posture receiver 303 is arranged to determine the listener's posture based on data that can be provided from end device 203 to edge device 201. Therefore, edge device 201 and end device 203 can establish a communication link that allows listener posture data to be transmitted from end device 203 to edge device 201. Thus, processing by edge device 201 can depend on the posture of end device 203; for example, a VR headset / glasses may track user movement and generate sensor data that is sent to edge device 201, and edge device 201 can adapt its processing based on the user's movement.

[0069] The listener's posture can be determined specifically in response to sensor inputs (e.g., from appropriate sensors that are part of the head-mounted device). It will be understood that many suitable algorithms are well known to those skilled in the art, and for the sake of brevity, they will not be described in further detail here.

[0070] In some embodiments, the determination of a listener's posture in an audio scene can be performed in the end device 203, and the determined listener posture can then be sent to an edge device. In some embodiments, other data, such as sensor data, can be sent to the edge device 201, which can be configured to determine the listener's posture based on this data.

[0071] A listener's posture is typically their position within an audio context, from which they can perceive the audio being presented to them; that is, it indicates the listener's / user's location within the audio scene. Figure 2 In this method, edge device 201 and terminal device 203 collaborate to process received audio data to generate output audio signals (e.g., binaural audio output signals) to provide users with a spatial audio experience / perception of the audio scene based on the listener's posture. Information about the listener's posture is provided to edge device 201 so that the processing by both edge device 201 and terminal device 203 can be adapted to the current listener posture.

[0072] In such a scenario, a key issue is what processing should be performed at the edge device 201 and the terminal device 203, and what specific information should be sent from the edge device 201 to the terminal device 203. Typically, the edge device 201 has significantly more computing resources than the terminal device 203, and therefore it is expected that most processing, especially complex resource-intensive processing, should be performed at the edge device 201. However, transmitting data about the listener's posture to the terminal device 203, and transmitting the corresponding audio data, introduces communication latency, which is often significant and can lead to a degraded user experience. It introduces round-trip latency, which can result in a perceptually significant delay between the user's movement and the corresponding perceived audio. Therefore, it is generally necessary for the terminal device 203 to be able to perform some, preferably low-complexity, processing to locally adapt the rendered audio to changes in the listener's posture, while still maintaining a significant amount of processing at the edge device 201. However, to achieve an efficient trade-off, the distribution of processing and data transmitted from the edge device 201 to the terminal device 203 is crucial.

[0073] Figure 4 It shows Figure 2 Examples of elements of the edge device 203 are provided. The edge device 203 includes a receiver 401 that receives audio data signals from the edge device 201. As described in more detail below, the audio data signals include audio signals generated by the edge device 201 to represent an audio scene. The receiver 401 is coupled to a renderer 403, which is arranged to render audio signals to a listener based on the received audio signals. Specifically, the renderer 403 generates binaural audio signals to represent an audio scene based on the received audio data. As will be described below, the audio data signals include one or more surround sound audio mixes and generally one or more audio objects, and the renderer 403 is arranged to generate binaural audio signals based on these signals.

[0074] In this example, the binaural output signals are fed to an output circuit 405, which outputs the binaural signals to headphones, which are typically part of a VR headset. The output circuit 405 may include, for example, a digital-to-analog converter, an amplifier function, etc., as are well known to those skilled in the art.

[0075] Renderer 403 is coupled to listener pose processor 409, which is arranged to determine the listener's pose. In this example, the listener pose is determined based on sensor input from headphones / VR headset 407. The listener pose is provided to renderer 403, which is arranged to generate binaural signals to represent the audio scene from the listener pose, and specifically dynamically influences the binaural signals to follow changes in the listener pose. Thus, the listener pose can correspond to the user / listener's position in the scene.

[0076] The listener gesture processor 409 is also coupled to a transmitter 411, which is arranged to send listener gestures to the edge device 201.

[0077] In this example, edge device 201 is arranged to generate an audio data signal and send the audio data signal to end device 203, wherein the audio data signal includes a surround sound audio mix and several audio objects. Edge device 201 is arranged to determine whether some audio elements are included in the audio data signal as part of the surround sound audio mix or as audio objects depending on the listener's posture.

[0078] Edge device 201 includes a designator 305 arranged to designate at least one audio object among received audio objects as a near-range audio object or a far-range audio object, depending on the listener's posture. Specifically, for a given audio object, designator 305 may determine a distance metric indicating the distance between the posture of the first audio object and the listener's posture. The distance metric is then compared to a threshold, and if the distance metric exceeds the threshold (which indicates a distance greater than a given value), the audio object is designated as a far-range object; otherwise, it is designated as a near-range object. It should be understood that in some embodiments, the audio object may include designation of other possible categories, including designation of subcategories of audio objects that are either near-range or far-range objects.

[0079] In this method, the distance metric for a first audio object designated as a remote audio object exceeds a given threshold, while the distance metric for a first audio object designated as a near-range audio object does not exceed the threshold. Therefore, designating the first audio object as a remote audio object requires a distance metric indicating a distance exceeding the threshold. Designating the first audio object as a near-range audio object requires a distance metric indicating a distance not exceeding the threshold.

[0080] It is understood that the specific implementation and algorithm / criteria used for precise specification may depend on the specific preferences and requirements of each embodiment. Specifically, it should be noted that the threshold used is clearly dependent on many attributes and design choices of each application. It should also be noted that the specification may depend on several other parameters and considerations. However, fundamentally, if the first audio object is specified as a remote audio object, its distance will exceed the threshold; while if it is specified as a near-range audio object, its distance will be below the threshold.

[0081] Distance can be any suitable distance measure used to represent the distance between the listener's pose and the pose of the audio object in an audio scene, such as Euclidean distance, sum of absolute coordinate differences, etc.

[0082] For the increased distance between the pose of the first audio object and the pose of the listener, the distance metric may have an increased value, and the description will focus on such examples (e.g., when the reference is compared with an appropriate threshold).

[0083] Edge device 201 also includes an audio mix generator 307, which is arranged to generate a surround sound audio mix based on multiple audio elements. The audio mix generator 307 is arranged to include or exclude an audio object in the surround sound audio mix based on whether the audio object is designated as a near-field audio object or a far-field audio object. Specifically, if an audio object is designated as a far-field audio object, it is included in the surround sound audio mix; however, if an audio object is designated as a near-field audio object, it is not included in the surround sound audio mix.

[0084] In many embodiments, the audio mixer 307 may be arranged to include multiple, and in many cases, all, audio objects designated as remote audio objects. In many embodiments, it may be arranged to exclude any audio objects designated as near-range audio objects.

[0085] Therefore, the audio mix generator 307 is arranged to generate a surround sound audio mix that includes an audio object designated as a remote audio object.

[0086] Furthermore, in many embodiments, the audio mix generator 307 is arranged to include other types of audio elements into the surround sound audio mix, such as, for example, a received surround sound audio mix, and in fact, in some embodiments, the audio mix generator 307 is arranged to add audio objects designated as remote audio objects to an existing (received) surround sound audio mix. In some embodiments, the surround sound audio mix may also be generated to include, for example, channel-based audio signals.

[0087] In some embodiments, audio mix generator 307 may be coupled to a first audio generator 309, which is arranged to generate a first audio signal to be sent to end device 203. The first audio signal represents a surround sound audio mix, and in many cases the generated surround sound audio mix can be sent directly without modification. However, in other embodiments, the surround sound audio mix may be processed to generate different representations, such as providing a binaural representation.

[0088] The surround sound audio mix / first audio signal is fed to a data signal generator 311, which is arranged to generate an audio data signal including the surround sound audio mix (directly represented as surround sound audio mix, or possibly represented by a set of binaural signals, etc.).

[0089] An audio object designated as a near-field audio object is fed to generator 311 and included as an audio object in the audio data signal. In some cases, edge device 201 also includes a second audio signal generator 313, which can generate a second audio signal for providing an appropriate representation of the audio object, for example, by encoding a binaural representation.

[0090] Generator 311 can then generate an audio data signal to include the surround sound audio mix as well as any audio objects designated as near-field audio objects. Alternatively, an audio data signal can be generated to exclude any audio objects designated as far-field audio objects, but these audio objects can be included in the surround sound audio mix.

[0091] Then, the end device 203 can receive audio data signals and process the received surround sound audio mix and audio objects to generate binaural output signals for the current listener's posture.

[0092] It is understood that in many embodiments, the audio data signal may include other audio elements besides surround sound audio mixes and audio objects. For example, it may include channel-based audio, diffuse background audio signals, other surround sound audio mixes, etc. In such cases, the end device 203 may include functionality for rendering these signals and combining them with audio signals generated based on surround sound audio mixes and audio objects.

[0093] This method, in many embodiments, provides a particularly advantageous distribution of processing from edge device 201 to end device 203 and an advantageous selection of audio data from edge device 201 to end device 203. It can generally allow for a favorable user experience with high audio quality and fast adaptation, while requiring relatively few computational resources in edge device 201.

[0094] Surround sound is an efficient format for representing a (large) set of audio objects. Specifically, when these audio objects are diffuse, a low-order surround sound mix is ​​sufficient. The achievable spatial resolution of audio objects in a surround sound mix is ​​directly related to the surround sound order. Note that a higher achievable spatial resolution will be applied to the entire surround sound mix, even if only a single audio object exists in the surround sound mix. Therefore, representing a finite set of point sources with a surround sound mix of a sufficient order to achieve a specific spatial resolution is inefficient. Instead, representing a finite set of point sources as individual audio objects (optionally combined with the surround sound mix) is more efficient.

[0095] A scene rendered from a surround sound mix can be considered to "exist" on a sphere. Therefore, when a user approaches a specific audio object on the sphere, the audio object will remain on the sphere, although its rendering gain will differ. In other words, as you move toward the sphere, the sphere essentially moves with you. Therefore, audio objects rendered using a surround sound mix are "inaccessible." Note that instead of the user approaching a point source audio object, the point source audio object can also be "approachable," for example, indicated by the trajectory of a specific object. Once audio objects become accessible, they are generally not well represented using surround sound mixes, regardless of their order.

[0096] Surround sound mixes can be rendered onto a set of speakers or efficiently converted into binaural signals for reproduction on headphones. Inherently, surround sound mixes are readily converted to account for the user's (3D) rotation. That is, if a user turns their head at a certain angle in any (3D) direction, the surround sound mix can be efficiently updated to account for that rotation. Therefore, the user will experience an updated surround sound mix that reflects their head rotation.

[0097] However, very different from rotation, translation (changes in position) cannot be efficiently compensated for in surround sound mixes. There is no efficient way to adapt to translation, especially when audio objects become accessible in response to, for example, the user's translation (such as the user moving closer to a point source in the surround sound mix or the audio object moving). Figure 5 As shown, compared to an audio object farther from the user (object 2), the user's translation (x) has a greater impact on the perceived angle of the audio object (object 1) that is closer to the user. For object 3, which is located further to the edge, the user's translation (x) has only a smaller impact on the perceived angle of the audio object. However, since the distance between object 3 and the (translated) user is below a threshold, this translation can make the audio object more accessible. For the user's new position, a new shape (a circle in this example) can be defined for the accessible object. Although it has a greater impact on the perceived angle of object 1, in this example, object 1 may move out of the (new) circle.

[0098] Since surround sound mixes are an efficient representation of (preferably diffused) audio objects, one approach is to ignore the effects of (small) user pans and not update the surround sound mix accordingly.

[0099] However, when the audio object is a point source—specifically, when the audio object is accessible and / or also has a visual component associated with the source—there is a clear benefit to accurate representation of the audio object. In the case of a corresponding visual component, even matching between the audio and visual components is highly desirable for a proper user experience. This necessitates a dynamic trade-off between representing the audio object separately or as part of a surround sound mix.

[0100] One approach to split rendering could be to pre-render all content into the surround sound mix based on predicted listener pose, and finally render it to both ears on the end device according to the surround sound. Since the user's three translational degrees of freedom (X, Y, Z) are not represented in the surround sound rendering, this can lead to increased round-trip latency (from the user translating to receiving an updated pre-rendered surround sound representation from the edge device). Even if the round-trip latency is acceptable, this translational change can still cause audible artifacts, specifically for sources relatively close to the user, as they are more likely to cause significant changes in the source's angle of incidence, which is an important perceptual cue for assessing the location of a sound source.

[0101] Therefore, the inventors have recognized that the method of appropriately pre-rendering a scene into a surround sound format based on the last known listener posture is not suitable for scenes where the listener can quickly move closer to the sound source or significantly change their distance from the sound source.

[0102] The method of pre-rendering multiple binaural signal pairs at the edge device and interpolating the final binaural pair based on those signals at the end device tends to be suboptimal in terms of both the resulting audio quality and the computational complexity required for pre-rendering and post-rendering. In these cases, the output sound quality depends on the final pose offset. For example, using a variant that interpolates the output at the final pose by a 15-degree rotation, severe artifacts begin to appear when the rotation exceeds 20 degrees. The artifacts from the interpolation are proportional to the minimum angle between the available pose and the actual pose. Furthermore, if the listener's pose varies in directions other than yaw angle (e.g., the pitch and roll angles of the user's head), the proposed method will have artifacts.

[0103] The methods described above, namely the methods for dynamically adapting different audio objects (and therefore corresponding audio sources) as part of a surround sound audio mix or as separate audio objects, can provide an improved and particularly advantageous approach in many embodiments because it addresses many of the problems described above.

[0104] This method can be used Figure 6 To explain, among them Figure 6An example of a scene consisting of several audio objects (1, 2, 3, 4, 5) is shown. Some audio objects are included as part of a surround sound audio mix (3, 4, 5), while other audio objects (1, 2) are sent separately as audio objects for independent binaural rendering. The surround sound audio mix and the individual audio objects are included in alternative audio data and sent to the end device. The surround sound audio mix can be provided as a binaural mix generated by a second audio signal generator. The audio objects are rendered (binauralized or rendered to speakers) at the end device, and then the audio objects are combined with the pre-rendered surround sound audio mix to produce a rendered output for consumption by one or more converters (typically headphones or AR / VR kits). Specifically, the surround sound audio mix can be binauralized or rendered to speakers for combination with the audio objects.

[0105] In this example, the collection of audio objects included in the surround sound audio mix, as well as the collection of audio objects represented separately, can change dynamically. For example, in Figure 6 In this process, the trajectory allows the audio object / source 4 to approach the listener's posture, and specifically, the audio object / source 4 can move within a threshold and change from being designated as a remote object to being designated as a near-range audio object. Therefore, the audio object / source 4 can be removed from the surround sound audio mix and introduced as a separate audio object into the audio data signal, thereby allowing the end device 203 to perform more accurate and dedicated rendering.

[0106] Changes in relative position and distance can occur equivalently through movement (specifically, translation) of the user / listening position. For audio objects included in surround sound mixes (3, 4, 5), user translation (X, Y, Z) does not result in a change in the intended orientation of the audio object until after the round-trip delay, i.e., after the user translation has been incorporated into the update of the surround sound audio mix (the translation is reported in the listener pose sent to edge device 201, where the surround sound audio mix is ​​modified to reflect the new position, which results in a sent surround sound audio mix corresponding to the new position). As mentioned above, for audio sources farther away from the sound source, this translation and gap between the listener pose for which the surround sound audio mix is ​​generated and the new listener pose is more difficult to perceive than for nearby or accessible audio sources. Both the spatial ambiguity of the surround sound representation and the round-trip delay make it more difficult to accurately locate a nearby audio source by attempting to approach it. A distance threshold for audio sources and objects included in a surround sound audio mix can be determined as an appropriate threshold at which an audio object begins to become accessible, or at which worst-case user translation (during round-trip delay) results in an unacceptable source location error.

[0107] The distance from a user to an individual audio object is affected by the movement of both the user and the audio object itself. For example, when the user is... Figure 6 When an audio object moves physically or virtually along the direction of object 4, it will transition into the circle representing a distance threshold, and other audio objects (e.g., audio object 1) can transition out of that circle. Alternatively, a similar effect is achieved when the trajectory of object 4 enters the circle. In this case, the audio object can become accessible without any translation by the user. Therefore, the audio object can be dynamically updated as a representation of a component in the surround sound audio mix or as a separate audio object. In many embodiments, this transition may include a hysteresis element.

[0108] In this method, an audio object can be transitioned from surround sound audio mix to audio object or vice versa once it passes or approaches a threshold. Smaller distances mean that the object will transition from surround sound to audio object upon approach. These transitions can be seamless, allowing the audio object to crossfade from the surround sound mix into separately sent audio objects, and vice versa.

[0109] During these transitions, the resolution of the separately transmitted audio objects can be advantageously reduced due to the perceptual masking of audio objects in the surround sound audio mix; that is, a lower bit rate can be used to decode the audio objects during this transition.

[0110] When an audio object transitions out of or into a surround sound audio mix, a lower bit rate can be assigned to portions of the audio object represented as a separate audio object. The bit rate can be increased as the audio object transitions further out of the surround sound audio mix. Masking (by the surround sound audio mix and / or other separate objects) can also be considered when determining the bit rate required to represent the audio object transitioning from the surround sound audio mix.

[0111] Therefore, in some embodiments, the audio mixer is arranged to change the data rate for the first audio object when switching between being designated as a near object and being designated as a far object, and thus, to change the data rate for the first audio object when switching between being part of a surround sound audio mix and not being part of a surround sound audio mix. The transition can be gradual.

[0112] In different embodiments, different parameters and attributes can be advantageously considered to designate an audio object as a near object or a remote object.

[0113] In many embodiments, the designation of an audio object may depend on a loudness metric for that audio object. Specifically, the designator 305 may be arranged to determine a threshold or equivalently a distance metric based on the loudness metric for the first audio object. Thus, in some embodiments, at least one of the distance metric and the threshold depends on the loudness metric for the audio object.

[0114] For example, a given audio object is more likely to be considered a near-range audio object if it is louder than a quieter audio object. Therefore, the threshold can be a function of the loudness of the audio object, and specifically, a monotonically increasing function of loudness (or equivalently, a distance measure can be a monotonically decreasing function of loudness). Thus, in many embodiments, the distance at which an audio object is considered a near-range audio object can increase for an increase in loudness.

[0115] In many embodiments, this approach can provide advantageous performance and can allow for adaptation to deliver a more perceptually consistent experience.

[0116] In some embodiments, this specification may further consider the loudness of one or more other audio objects, and specifically, relative loudness may be considered. This may, for example, allow for consideration of masking effects from other audio sources.

[0117] Specifically, depending on the loudness and / or temporal / spatial masking of an audio object, a threshold can be adapted to measure the distance at which an audio object is transformed from being included in a surround sound audio mix to being represented as a separate audio object (or vice versa). For example, for a loud (and therefore potentially prominent) audio object, the threshold can be changed to increase the likelihood that the audio object is designated as a near-range audio object, while for an audio object masked by another object, the threshold of the masked (less important) object can be changed to decrease the likelihood that it is designated as a near-range audio object (thus, the threshold decreases as loudness increases, and the threshold increases as loudness decreases).

[0118] In many embodiments, the designation of an audio object may depend on a diffusion metric for that audio object. Specifically, the designator 305 may be arranged to determine a threshold or equivalently a distance metric based on the diffusion metric for the first audio object. Thus, in some embodiments, at least one of the distance metric and the threshold depends on the diffusion metric for the audio object.

[0119] Specifically, since the localization of a diffuse audio object is worse than that of a point-source audio object, the threshold can depend on the measure of the audio object's diffusion. For example, (for a specific angle of incidence) the threshold distance (radius) for a point-source audio object is equal to... d p In this case, the threshold distance for diffuse audio objects can be<d p (For example, ).

[0120] In many embodiments, the designation of an audio object may depend on the direction from the listener's posture to the location of the audio object. Specifically, the designator 305 may be arranged to determine a threshold or equivalently a distance measure based on the direction from the listener's posture to the location of the audio object. Thus, in some embodiments, at least one of the distance measure and the threshold depends on the direction from the listener's posture to the location of the audio object.

[0121] Audio objects represented by a forward angle of incidence can be located better than those represented by a backward angle of incidence, for example. Therefore, the threshold (represented by a circle in the diagram, i.e., the direction-independent distance) can be represented by an asymmetrical shape. For example, if the threshold distance (radius) for a forward angle of incidence is equal to d, the threshold distance for an audio object with a backward angle of incidence can be <d (e.g., For all other directions, the threshold distance can seamlessly transition between the front and rear incident angles. Typically, the threshold mode corresponds to the positioning accuracy as a function of the incident angle.

[0122] In many embodiments, the designation of an audio object may depend on the trajectory of the audio object's position within the audio scene. Specifically, the designator 305 may be arranged to determine a threshold or equivalent distance metric based on the trajectory for the first audio object. Thus, in some embodiments, at least one of the distance metric and the threshold depends on the trajectory for the audio object.

[0123] The trajectory can be a trajectory relative to the listener's posture, and therefore can be a relative trajectory caused by the movement of the audio object and / or the listener's posture in the audio scene.

[0124] For example, designator 305 can track the position of an audio object to determine whether the audio object is moving toward or away from the listener's posture. Specifically, based on the trajectory, it can be estimated whether the audio object is likely to move toward the listener's posture, and thus potentially be designated as a near object. If so, a distance threshold can be increased, for example, resulting in earlier designation of the audio object and earlier removal from the surround sound audio mix and inclusion as a separate audio object.

[0125] The velocity of an audio object's trajectory also affects the threshold at which the object moves from being designated a near object to being designated a far object (or vice versa). For example, the threshold may remain constant for a essentially stationary object, while for a rapidly moving object, the threshold may increase so that the object is represented as a separate object in a timely manner. Different criteria can be used to determine the velocity of an object, based on metadata or predictions, such as the following: • Listener's posture to the velocity vector of the audio object's position (expected, probability of angular distance change). • Active (moving) point source audio objects In many embodiments, the designation of an audio object may depend on the order of the surround sound audio mix. Specifically, the designator 305 may be arranged to determine a threshold or, equivalently, a distance metric based on the order of the surround sound audio mix. Thus, in some embodiments, at least one of the distance metric and the threshold depends on the order of the surround sound audio mix.

[0126] In some embodiments, distance can be considered in relation to the order of the surround sound audio mix, and specifically, the higher the order of the surround sound audio mix, the more likely an audio object is to be considered a distant audio object and thus included in the surround sound audio mix. A higher order of the surround sound audio mix allows for better representation of audio objects, and therefore, it is more appropriate to include a greater number of audio objects in the surround sound audio mix.

[0127] In many embodiments, the designator 305 may also consider the number of audio objects included in the surround sound audio mix, and specifically, the number of audio objects designated as remote audio objects. For example, a maximum or preferred number of audio objects for a given order may be determined, and the number of remote audio objects included in the surround sound audio mix may be limited to that maximum number.

[0128] In some embodiments, the number of surround sound audio mixes may be adapted depending on the number of audio objects designated as remote audio objects (and to be included in the surround sound audio mix).

[0129] A threshold can be defined as a balance between the surround sound level and the number of audio objects. As mentioned earlier, increasing the surround sound level improves the discriminability of objects. To represent a large number of objects, it may be advantageous to increase the surround sound level so that fewer objects need to be represented individually.

[0130] In many embodiments, the designation of an audio object may depend on an indication of the importance or priority of that audio object. Specifically, the designator 305 may be arranged to determine a threshold or equivalent distance metric based on an indication of the importance or priority of a first audio object. Thus, in some embodiments, at least one of the distance metric and the threshold depends on an indication of the importance or priority of the audio object.

[0131] For example, an importance / priority indicator can be determined based on the provided audio type. The importance / priority indicator can be assigned manually, for example, and / or the importance / priority indicator can be received, for example, along with the received data used for the input audio element.

[0132] For example, an object representing a person speaking (e.g., a dialogue object) might have higher importance in a scene than a (background) object making some sounds. Audio metadata can be used to categorize the importance of objects and thus influence the threshold at which objects are removed from the surround sound audio mix to be represented as separate audio objects. Examples of aspects that can indicate higher importance include: When an audio object represents a user, there is a high demand for rendering it as separate audio objects. Does a visual component exist that is connected to the audio object? • Audio object size • Flags or “importance” measures indicated in the metadata.

[0133] In many embodiments, the designation of an audio object may depend on the distance between the location of one audio object and the location of another audio object. Specifically, the designator 305 may be arranged to determine a threshold or equivalent distance metric based on the distance between the location of one audio object and the location of another audio object. Thus, in some embodiments, at least one of the distance metric or threshold depends on the distance between the location of one audio object and the location of another audio object.

[0134] In some embodiments, this designation may depend on the location of one or more other audio objects. For audio objects that are close together (essentially co-located) or have a small angular distance relative to the user, the ability to distinguish them as individual objects is reduced. Therefore, instead of representing all these co-located objects in the surround sound audio mix, the co-located objects can be rendered as one or more combined audio objects. This improves the accessibility of the group of co-located objects, reduces the bit rate required to jointly represent these objects, and reduces the computational complexity of rendering them as individual audio objects.

[0135] For audio objects that are close to each other, or for one or more dominant audio objects in a group of close audio objects, it may not be necessary to send all audio objects in the group or to use the same transmission rate for all these individual audio objects. Audio objects in a group of close audio objects with similar dominance can be grouped into individual audio objects.

[0136] From a computational complexity perspective, it may be necessary to limit the number of audio objects sent individually to a maximum. For example, there could be maximum values ​​for the number of audio objects sent individually and for the number of audio objects expected to be close to a threshold.

[0137] Individual audio objects can be represented as pre-rendered binaural signals. This can be done as part of a trade-off between complexity and quality, depending on the capabilities of the end device.

[0138] For example, in cases with low available computing resources, only a single representation of the audio object can be pre-rendered. In cases with moderate available computing resources, additional representations of the audio object can be pre-rendered.

[0139] In many embodiments, the designation of an audio object may depend on the attributes of the rendering device / end device. Specifically, the designator 305 may be arranged to determine a threshold or equivalent distance metric based on the attributes of the rendering device / end device. Thus, in some embodiments, at least one of the distance metric or threshold depends on the attributes of the rendering device / end device.

[0140] This attribute can be specifically the capabilities of the end device, such as available rendering algorithms, processing power, etc.

[0141] For example, during setup or initialization, the end device 203 may send a data message to the edge device 201 indicating the attributes (e.g., capabilities) of the end device 203. For example, the edge device 201 may indicate its type or processing capabilities. The designator 305 may take this information into account. For example, the maximum number of audio objects that the end device 203 can process may be determined based on the processing capabilities of the end device 203. A distance threshold may be set accordingly, specifically such that the number of individual audio objects represented in the audio data signal does not exceed the determined maximum number.

[0142] In some embodiments, the distance threshold may depend on the capabilities of the end device. For example, for end devices with limited processing power, it is beneficial to change the overall threshold to reduce the number of audio objects represented as separate audio objects. Other trade-offs involving combinations of the above parameters are also possible.

[0143] In some embodiments, the threshold can be dynamically changed to have a certain number of separately represented audio objects to optimize audio quality within the capabilities of the end device.

[0144] Therefore, the processing power of the end device 203 can be considered. For example, for low-cost devices, a limited set of pre-rendered audio objects can be provided.

[0145] In many embodiments, the designation of an audio object may depend on the data transmission attributes of the communication link used to send audio data signals to the edge device 201. Specifically, the designator 305 may be arranged to determine a threshold or equivalent distance metric based on the data transmission attributes. Thus, in some embodiments, at least one of the distance metric or threshold depends on the data transmission attributes. The data transmission attributes may specifically be the data rate used for communication.

[0146] For example, edge device 201 can estimate the current bandwidth or throughput used to transmit audio data to end device 203. It can then go on to determine (in addition to, for example, a surround sound audio mix) the maximum number of individual audio objects that can be sent, and continue to adapt distance thresholds to ensure that the maximum number is not exceeded.

[0147] As another example, when bandwidth is constrained and the number of objects designated as near objects is relatively high, the distance threshold can be reduced so that fewer objects are designated as near objects. Additionally, the surround sound (HOA) order can be increased by moving previously designated near objects, allowing them to be represented "better" in the surround sound audio mix.

[0148] In many embodiments, the designation of an audio object may depend on multiple listener poses, for example, in a game scene. Specifically, the designator 305 may be arranged to determine a threshold or equivalent distance metric based on multiple listener poses. Thus, in some embodiments, at least one of the distance metric or threshold depends on multiple listener poses.

[0149] As an example, in some implementations, edge device 201 may receive listener poses from two or more end devices 203, and edge device 201 may continue to generate a single audio data signal that can be sent back to the multiple end devices 203. In such a case, the designation of an audio object can take multiple listener poses into account. For example, a distance threshold can be determined for each listener pose, and an audio object can be designated as a near-range audio object if the distance to any listener pose is less than the corresponding threshold. Thus, in this embodiment, an acoustic object can be designated as a near-range object if it is close to any listener pose, and it can only be included in the surround sound audio mix if it is far away from all listener poses.

[0150] The distance threshold (or equivalent, distance metric) for an audio object can be determined based on a number of different parameters, including one or more of the following: The distance of the audio object relative to the listener's posture o The (3D) angle (angle of incidence) of the audio object relative to the listener's posture oDiffusion of audio objects (point source vs. objects with a range) o Location of other audio objects • The co-location (small angular distance) of audio objects can be rendered as a single audio object instead of being preserved in the spatial audio mix. Loudness / Time / Spatial Masking speed of audio objects • User-to-object velocity vector (expected, probability of angular distance change) • Active (moving) point source audio objects o Number of input objects The audio object metadata includes: • Audio object type (i.e., when the audio object represents the user, there is a greater need to render it as a separate audio object). Does a visual component exist that is connected to the audio object? • Audio object size o Dynamic balance between surround sound level and number of audio objects o During the transition from surround sound mix to individual audio objects, the bitrate allocation of the audio objects. o Use cover o The capabilities of the end device (e.g., processing power).

[0151] Figure 7This is a block diagram illustrating an example processor 700 according to an embodiment of the present disclosure. Processor 700 can be used to implement one or more processors that implement the previously described means or elements thereof. Processor 700 can be any suitable processor type, including but not limited to microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable arrays (FPGAs) (which have been programmed to form processors), graphics processing units (GPUs), application-specific integrated circuits (ASICs) (which have been designed to form processors), or combinations thereof.

[0152] Processor 700 may include one or more cores 702. Core 702 may include one or more arithmetic logic units (ALUs) 704. In some embodiments, in addition to or in place of ALU 704, core 702 may include a floating-point logic unit (FPLU) 706 and / or a digital signal processing unit (DSPU) 708.

[0153] Processor 700 may include one or more registers 812 communicatively coupled to core 702. Registers 712 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 712 may be implemented using static memory. Registers may provide data, instructions, and addresses to core 702.

[0154] In some embodiments, processor 700 may include cache memory 710 communicatively coupled to one or more stages of core 702. Cache memory 710 may provide computer-readable instructions for execution to core 702. Cache memory 710 may provide data for processing by core 702. In some embodiments, computer-readable instructions may be provided to cache memory 710 from local memory (e.g., local memory connected to external bus 716). Cache memory 710 may be implemented using any suitable cache memory type, such as metal-oxide-semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0155] Processor 700 may include controller 714, which controls inputs to and / or outputs from other processors and / or components included in the system to other processors and / or components included in the system. Controller 714 may control data paths in ALU 704, FPLU 706, and / or DSPU 708. Controller 714 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of controller 714 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.

[0156] Register 712 and cache memory 710 can communicate with controller 714 and core 702 via internal connections 720A, 720B, 720C and 720D. The internal connections can be implemented as a bus, multiplexer, crossbar switch and / or any other suitable connection technology.

[0157] Inputs and outputs to processor 700 may be provided via bus 716, which may include one or more conductive lines. Bus 716 may be communicatively coupled to one or more components of processor 700, such as controller 714, cache 710, and / or register 712. Bus 716 may be coupled to one or more components of the system.

[0158] Bus 716 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 732. ROM 732 may be a mask ROM, electronically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory (RAM) 733. RAM 733 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 735. The external memory may include flash memory 734. The external memory may include a magnetic storage device such as a disk 736. In some embodiments, the external memory may be included in the system.

[0159] It should be understood that, for clarity, embodiments of the invention have been described above with reference to various functional circuits, units, and processors. However, it will be apparent that any suitable functional distribution among different functional circuits, units, or processors may be used without prejudice to the invention. For example, functionality illustrated as being performed by separate processors or controllers may be performed by the same processor or controller. Therefore, references to specific functional units or circuits should be considered merely as references to suitable means for providing said functionality, and not as indications of a strict logical or physical structure or organization.

[0160] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, this invention can be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of this invention can be implemented physically, functionally, and logically in any suitable manner. In practice, the function can be implemented as a single unit, multiple units, or as part of other functional units. Therefore, this invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.

[0161] While the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is defined only by the appended claims. Furthermore, although it may appear that the features are described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0162] Furthermore, although listed separately, multiple means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Additionally, while different claims may include individual features, these features may be advantageously combined, and inclusion in different claims does not imply that such combinations are infeasible and / or disadvantageous. Moreover, the inclusion of a feature in one class of claims does not imply a limitation on that class, but rather indicates that the feature is equally applicable to other appropriate classes of claims. Furthermore, the order of features in a claim does not imply that the features must be performed in a particular order, and specifically, the order of steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. Furthermore, the singular form does not exclude the plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude multiple instances. Reference numerals in the claims are provided only as clarifying examples and should not be construed as limiting the scope of the claims in any way.

[0163] Typically, instances of this method are indicated by the following embodiments.

[0164] Example: Example 1. An audio device, comprising: A receiver (301) is arranged to receive a plurality of audio elements representing an audio scene, the audio elements comprising several audio objects, each audio object being linked to a position in the audio scene; A listener gesture receiver (303) is arranged to receive instructions on the listener's gesture; The designator (305) is arranged to designate at least a first audio object among the plurality of audio objects as a near-field audio object or a far-field audio object based on a comparison of a distance metric with a threshold, the distance metric indicating the distance between the pose of the first audio object and the pose of the listener, wherein the distance metric for the first audio object designated as a far-field audio object exceeds the threshold, and the distance metric for the first audio object designated as a near-field audio object does not exceed the threshold; An audio mix generator (307) is arranged to generate a surround sound audio mix based on a first plurality of audio elements, wherein the first plurality of audio elements include the first audio object if the first audio object is specified as a remote audio object. A data generator (311) is arranged to generate an audio data signal including the surround sound audio mix; and wherein, if the first audio object is designated as a near-field audio object, the data generator is arranged to include the first audio object in the audio data signal.

[0165] Example 2. The audio device according to Example 1, wherein the designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object depending on the loudness metric for the first audio object.

[0166] Example 3. An audio device according to any of the foregoing embodiments, wherein the designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object depending on the direction from the listener's posture to the position of the first audio object.

[0167] Example 4. An audio device according to any of the foregoing embodiments, wherein the designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object depending on the trajectory of the position of the first audio object in the audio scene.

[0168] Example 5. An audio device according to any of the foregoing embodiments, wherein the designator (305) is arranged to designate the first audio object as a near-field audio object or as a far-field audio object depending on the order of the surround sound audio mix.

[0169] Example 6. An audio device according to any of the foregoing embodiments, wherein the designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object depending on the distance between the position of the first audio object and the position of the second audio object.

[0170] Example 7. An audio device according to any of the foregoing embodiments, wherein the data generator (311) is arranged to send an audio data signal to a remote device via a communication link; and the designator (305) is arranged to designate the first audio object as a near-field audio object or a remote audio object depending on the data transmission attributes of the communication link.

[0171] Example 8. An audio device according to any of the foregoing embodiments, wherein the audio mixer (307) is arranged to change the data rate for the first audio object when switching between being designated as a near object and being designated as a far object.

[0172] Example 9. An audio device according to any of the foregoing embodiments, wherein the listener gesture receiver (303) is arranged to receive a plurality of listener gestures, the designator (305) is arranged to determine the distance measure depending on the plurality of listener gestures, and the data generator (311) is arranged to transmit the audio data signal to a plurality of remote devices.

[0173] Example 10. An audio system comprising an audio device (201) according to any of the foregoing embodiments and a rendering device (203), wherein the data generator (311) is arranged to send the audio data signal to the rendering device (203); and The rendering device (203) includes: A receiver (401) is arranged to receive the audio data signal from the audio device (201); and A renderer (403) is configured to generate a binaural audio signal representing the audio scene, wherein the binaural audio signal includes contributions from a first audio mix and the first audio object if a first audio object is present in the audio data signal.

[0174] Example 11. The audio system according to Example 10, wherein the rendering device (203) is a head-mounted device including an audio transducer arranged to reproduce the binaural rendering signal.

[0175] Example 12. An audio system according to Example 10 or 11, wherein the designator (305) is arranged to designate the first audio object as a near-field audio object or as a far-field audio object depending on the attributes of the rendering device.

[0176] Example 13. A method for operating an audio device, the method comprising: Receive multiple audio elements for representing an audio scene, the audio elements including several audio objects, each audio object being linked to a position in the audio scene; Receive instructions on the listener's posture; Depending on the comparison between the distance metric and a threshold, at least a first audio object among the plurality of audio objects is designated as a near-field audio object or a far-field audio object, the distance metric indicating the distance between the pose of the first audio object and the pose of the listener, wherein the distance metric for the first audio object designated as a far-field audio object exceeds the threshold, and the distance metric for the first audio object designated as a near-field audio object does not exceed the threshold; A surround sound audio mix is ​​generated based on a first plurality of audio elements, wherein if the first audio object is specified as a remote audio object, the first plurality of audio elements include the first audio object; and Generates audio data signals including surround sound audio mixes; and If the first audio object is designated as a near-field audio object, then the first audio object is included in the audio data signal.

[0177] Example 14. A method for operating an audio device (201), the audio device performing the method of claim 13 and sending the audio data signal to a rendering device (203); and the rendering device (203) performing the following steps: Receive the audio data signal (201) from the audio device (201); and Generate a binaural audio signal to represent the audio scene, wherein if a first audio object is present in the audio data signal, the binaural audio signal includes contributions from the first audio mix and the first audio object.

[0178] More specifically, the invention is defined by the appended claims.

Claims

1. An audio device, comprising: A receiver (301) is arranged to receive a plurality of audio elements representing an audio scene, the audio elements comprising several audio objects, each audio object being linked to a position in the audio scene; A listener gesture receiver (303) is arranged to receive instructions on the listener's gesture; The designator (305) is arranged to designate at least a first audio object among the plurality of audio objects as a near-field audio object or a far-field audio object based on a comparison of a distance metric with a threshold, the distance metric indicating the distance between the pose of the first audio object and the pose of the listener, wherein the distance metric for the first audio object designated as a far-field audio object exceeds the threshold, and the distance metric for the first audio object designated as a near-field audio object does not exceed the threshold; An audio mix generator (307) is arranged to generate a surround sound audio mix based on a first plurality of audio elements, wherein the first plurality of audio elements include the first audio object if the first audio object is specified as a remote audio object. A data generator (311) is arranged to generate an audio data signal including the surround sound audio mix; and wherein, if the first audio object is designated as a near-field audio object, the data generator is arranged to include the first audio object in the audio data signal.

2. The audio device according to claim 1, wherein, The designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object, depending on the loudness metric for the first audio object.

3. The audio device according to any of the preceding claims, wherein, The designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object, depending on the direction from the listener's posture to the location of the first audio object.

4. The audio device according to any of the preceding claims, wherein, The designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object based on the trajectory of the position of the first audio object in the audio scene.

5. The audio device according to any of the preceding claims, wherein, The designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object, depending on the order of the surround sound audio mix.

6. The audio device according to any of the preceding claims, wherein, The designator (305) is arranged to designate the first audio object as a near-field audio object or a far-field audio object depending on the distance between the location of the first audio object and the location of the second audio object.

7. The audio device according to any of the preceding claims, wherein, The data generator (311) is arranged to send the audio data signal to a remote device via a communication link; and the designator (305) is arranged to designate the first audio object as a near-field audio object or a remote audio object, depending on the data transmission attributes of the communication link.

8. The audio device according to any of the preceding claims, wherein, The audio mixer (307) is configured to change the data rate for the first audio object when switching between being designated as a near object and being designated as a remote object.

9. The audio device according to any of the preceding claims, wherein, The listener gesture receiver (303) is arranged to receive multiple listener gestures, the designator (305) is arranged to determine the distance measurement based on the multiple listener gestures, and the data generator (311) is arranged to send the audio data signal to multiple remote devices.

10. An audio system comprising an audio device (201) according to any of the preceding claims and a rendering device (203); wherein, The data generator (311) is arranged to send the audio data signal to the rendering device (203); as well as The rendering device (203) includes: A receiver (401) is arranged to receive the audio data signal from the audio device (201); as well as A renderer (403) is configured to generate a binaural audio signal representing the audio scene, wherein the binaural audio signal includes contributions from the first audio mix and the first audio object if the first audio object is present in the audio data signal.

11. The audio system according to claim 10, wherein, The rendering device (203) is a head-mounted device including an audio transducer arranged to reproduce the binaural rendering signal.

12. The audio system according to claim 10 or 11, wherein, The designator (305) is configured to designate the first audio object as a near-field audio object or a far-field audio object, depending on the attributes of the rendering device.

13. A method for operating an audio device, the method comprising: Receive multiple audio elements for representing an audio scene, the audio elements including several audio objects, each audio object being linked to a position in the audio scene; Receive instructions on the listener's posture; Depending on the comparison between the distance metric and a threshold, at least a first audio object among the plurality of audio objects is designated as a near-field audio object or a far-field audio object, the distance metric indicating the distance between the pose of the first audio object and the pose of the listener, wherein the distance metric for the first audio object designated as a far-field audio object exceeds the threshold, and the distance metric for the first audio object designated as a near-field audio object does not exceed the threshold; A surround sound audio mix is ​​generated based on a first plurality of audio elements, wherein if the first audio object is specified as a remote audio object, the first plurality of audio elements include the first audio object. as well as Generate an audio data signal including the surround sound audio mix; as well as If the first audio object is designated as a near-field audio object, then the first audio object is included in the audio data signal.

14. A method for operating an audio device (201), the audio device performing the method of claim 13 and sending the audio data signal to a rendering device (203); and the rendering device (203) performing the following steps: Receive the audio data signal (201) from the audio device (201); and Generate a binaural audio signal to represent the audio scene, wherein if the first audio object is present in the audio data signal, the binaural audio signal includes contributions from the first audio mix and the first audio object.

15. A computer program product comprising computer program code units adapted to perform all the steps of claim 13 or 14 when the program is run on a computer.