Method and apparatus for processing communication audio in immersive audio scene rendering
The apparatus and method address the challenge of rendering communication audio in AR by determining rendering parameters and insertion points to adapt to network dynamics and user preferences, ensuring low-latency and immersive audio integration.
Patent Information
- Application Number
- JP2022145812
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-17
- Filing Date
- 2022-09-14
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2042-09-14
AI Technical Summary
Existing augmented reality (AR) systems face challenges in rendering communication audio within immersive audio scenes due to dynamic network conditions and unknown delay parameters during content consumption, which affect the acoustic properties and interaction quality.
An apparatus and method for determining rendering process parameters, including insertion points and selection of rendering elements based on communication audio signal characteristics, network conditions, and user preferences to ensure low-latency and immersive audio integration.
Enables reliable and immersive rendering of communication audio within AR environments, adapting to dynamic network conditions and user interactions, maintaining desired audio quality and latency.
Smart Images

Figure 0007744311000005 
Figure 0007744311000006 
Figure 0007744311000007
Abstract
Description
[Technical Field]
[0001] This application relates to a method and apparatus for processing communication audio within an augmented reality rendering, but is not limited to a method and apparatus for processing communication audio within an augmented reality six degrees of freedom rendering. [Background technology]
[0002] Augmented reality (AR) applications (and other similar virtual scene creation applications, such as mixed reality (MR) and virtual reality (VR)), which present virtual scenes to a user wearing a head-mounted device (HMD), have become increasingly complex and sophisticated over time. Applications may contain data that includes visual and audio components (or overlays) that are presented to the user. These components are provided to the user according to the user's position and orientation (in the case of six-degree-of-freedom applications) within the augmented reality (AR) scene.
[0003] Scene information for rendering an AR scene typically includes two parts. One part is virtual scene information, which can be described during content creation (or by an appropriate capture device) and represents the captured (or initially generated) scene. The virtual scene may be provided in Encoder Input Format (EIF) data format. The EIF and (captured or generated) audio data are used by an encoder to generate a scene description and spatial audio metadata (and audio signals), which can be delivered via a bitstream to a rendering (playback) device or equipment. Thus, the scene description for an AR or VR scene is specified by the content creator during the content creation phase. In the case of VR, the entire scene is specified and rendered as specified in the content creator's bitstream.
[0004] The second part of the rendering of an AR audio scene is related to the listener's (or end user's) physical listening space (or physical space). Spatial information about the scene or listener can be obtained during the AR rendering (when the listener is consuming the content). This means that a fundamental aspect of AR that differs from VR is that the acoustic properties of the audio scene are only known during content consumption (in the case of AR), and cannot be known or optimized during content creation.
[0005] FIG. 1 illustrates an example of an AR scene in which a virtual scene is located within a physical listening space. In this example, there is a user 107 located within the physical listening space 101. Furthermore, in this example, the user 109 is experiencing a six-degree-of-freedom (6DOF) virtual scene 113 with virtual scene elements. In this example, the elements of the virtual scene 113 are represented by two audio objects, a first object 103 (a guitar player) and a second object 105 (a drummer), a virtual occlusion element (e.g., represented as a virtual partition 117), and a virtual room 115 (e.g., a wall with a size, position, and acoustic material defined within the virtual scene description). A renderer (which in this example is a handheld electronic device or apparatus 111) is configured to perform rendering such that the auralization is plausible relative to the user's physical listening space (e.g., wall locations and acoustic material properties of the walls). The rendering is presented to the user 107, in this example, via suitable headphones or a headset 109.
[0006] Thus, for an AR scene, the content creator bitstream carries information about which audio and scene geometry elements correspond to which anchors in the listening space. As a result, the positions of audio elements, reflective elements, occlusion elements, etc. are known only during rendering. Furthermore, acoustic modeling parameters are known only during rendering.
[0007] Social VR / AR is a further development of such systems, which are envisioned to support the rendering of voices and audio from other users in the virtual environment. Furthermore, it has been proposed to render the received voice and audio communications as an immersive audio signal. Summary of the Invention [Problem to be solved by the invention]
[0008] The embodiments of the present application aim to address the problems of the prior art. [Means for solving the problem]
[0009] According to a first aspect, there is provided an apparatus comprising means configured to: acquire at least one spatial audio signal for rendering within an immersive audio scene; acquire a communication audio signal and position information related to the communication audio signal; acquire rendering process parameters related to the communication audio signal; determine a rendering method based on the rendering process parameters; and determine an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0010] The means may be further configured to generate at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method, an insertion point in the rendering process for the determined rendering method, and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0011] The means may be further configured to determine at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal.
[0012] The means configured to determine an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters may further be configured to at least one of: determine the insertion point in the rendering process based on at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal; and determine the rendering method and / or a selection of rendering elements for the determined rendering method based on at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal.
[0013] The allowable delay value may be the amount of delay allowed for utilizing the communication audio signal, and the communication audio signal delay may be a delay value determined based on the end-to-end delivery delay and the delay in rendering the communication audio.
[0014] The audio format associated with the communication audio signal may include one of a one-way communication audio signal and a conversational communication audio signal between users in an immersive scene.
[0015] The means configured to generate at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters may be configured to represent the communication audio signal as a high-order Ambisonic audio signal.
[0016] The means may be further configured to obtain user input, and the means configured to generate at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the rendering method determined based on the rendering process parameters and / or a selection of rendering elements for the determined rendering method may be configured to generate the at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal further based on the user input, and the user input may be configured to define at least one of allowed communication audio signal types, allowed audio formats, allowed delay values, and at least one acoustic modeling preference parameter.
[0017] The means may be further configured to obtain a communication audio signal type associated with the at least one spatial audio signal, and the means configured to generate at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the rendering method determined based on the rendering process parameters may be configured to generate the at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal further based on the at least one communication audio signal type associated with the at least one spatial audio signal.
[0018] The rendering processes and / or rendering elements may include one or more of Doppler processing, direct sound processing, material filtering processing, early reflection processing, diffuse late reverberation processing, sound source extension processing, occlusion processing, diffraction processing, sound source transformation processing, externalization rendering, and in-head rendering.
[0019] The means for determining an insertion point in a rendering process of the determined rendering method and / or a selection of rendering elements of the determined rendering method based on the rendering process parameters may be configured to determine a rendering mode, and the rendering mode may include a value indicating an insertion point of the communication audio signal.
[0020] The value indicating the insertion point may include one of a first mode value indicating that the communication audio signal and the at least one spatial audio signal are inserted at the start of the rendering processing method, a second mode value indicating that the communication audio signal is to bypass the rendering processing and be mixed directly with the output of the rendering processing applied to the at least one spatial audio signal, and a third mode value indicating that the rendering processing is to be fully applied to the at least one spatial audio signal while the communication audio signal is partially rendered.
[0021] A third mode value indicating that the communication audio signal has been partially rendered may be a value indicating that the communication audio signal is a direct sound rendering of a point sound source and a binaural rendering relative to the user position.
[0022] The means may be further configured to determine an audio format type of the communication audio signal based on the rendering process parameters.
[0023] Furthermore, the means for determining an insertion point in a rendering process of the determined rendering method and / or a selection of rendering elements of the determined rendering method based on the rendering process parameters may be configured to determine an insertion point in a rendering process of the communication audio signal in the determined rendering method based on the audio format type.
[0024] The means configured to determine an insertion point in a rendering process for a communication audio signal in a rendering method determined based on an audio format type may be configured to determine that, if the communication audio signal has an audio format type of a pre-rendered spatial audio format, the insertion point in the rendering method is in direct mixing with the output of the rendering process applied to the at least one spatial audio signal.
[0025] According to a second aspect, there is provided a method for an apparatus for rendering a communication audio signal within an immersive audio scene, the method comprising: obtaining at least one spatial audio signal for rendering within the immersive audio scene; obtaining the communication audio signal and position information associated with the communication audio signal; obtaining rendering process parameters associated with the communication audio signal; determining a rendering method based on the rendering process parameters; determining an insertion point in the rendering process for the determined rendering method; and / or selecting a rendering element for the determined rendering method based on the rendering process parameters.
[0026] The method may further include generating at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method, and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0027] The method may further include determining at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal.
[0028] Determining an insertion point in the rendering process for the determined rendering method and / or selecting a rendering element for the determined rendering method based on the rendering process parameters may further include at least one of determining an insertion point in the rendering process further based on the determined at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal, and selecting a rendering method and / or a rendering element for the determined rendering method based on the determined at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal.
[0029] The allowable delay value may be the amount of delay allowed for utilizing the communication audio signal, and the communication audio signal delay may be a delay value determined based on the end-to-end delivery delay and the delay in rendering the communication audio.
[0030] The audio format associated with the communication audio signal may include one of a one-way communication audio signal and a conversational communication audio signal between users in an immersive scene.
[0031] Generating at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method, and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters may include representing the communication audio signal as a high-order Ambisonic audio signal.
[0032] The method may further include obtaining user input and generating at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the rendering method determined based on the rendering process parameters, and may further comprise generating the at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal further based on the user input, wherein the user input may include defining at least one of allowed communication audio signal types, allowed audio formats, allowed delay values, and at least one acoustic modeling preference parameter.
[0033] The method may further include obtaining a communication audio signal type associated with the at least one spatial audio signal, and generating at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the rendering method determined based on the rendering process parameters, and may include generating the at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal further based on the at least one communication audio signal type associated with the at least one spatial audio signal.
[0034] The rendering processes and / or rendering elements may include one or more of Doppler processing, direct sound processing, material filtering processing, early reflection processing, diffuse late reverberation processing, sound source extension processing, occlusion processing, diffraction processing, sound source transformation processing, externalization rendering, and in-head rendering.
[0035] Determining an insertion point in a rendering process for the determined rendering method and / or selecting a rendering element for the determined rendering method based on the rendering process parameters may include determining a rendering mode, the rendering mode including a value indicating an insertion point of the communication audio signal.
[0036] The value indicating the insertion point may include one of a first mode value indicating that the communication audio signal and the at least one spatial audio signal are inserted at the start of the rendering processing method, a second mode value indicating that the communication audio signal bypasses the rendering processing and is mixed directly with the output of the rendering processing applied to the at least one spatial audio signal, and a third mode value indicating that the rendering processing is fully applied to the at least one spatial audio signal while the communication audio signal is partially rendered.
[0037] A third mode value indicating that the communication audio signal has been partially rendered may be a value indicating that the communication audio signal is a direct sound rendering of a point sound source and a binaural rendering relative to the user position.
[0038] The method may further include determining an audio format type of the communication audio signal based on the rendering process parameters.
[0039] Determining an insertion point in a rendering process for the determined rendering method and / or selecting a rendering element for the determined rendering method based on the rendering process parameters may include determining an insertion point in a rendering process for the communication audio signal in the determined rendering method based on the audio format type.
[0040] Determining an insertion point in a rendering process of the communication audio signal within the determined rendering method based on the audio format type may include determining that if the communication audio signal has an audio format type of a pre-rendered spatial audio format, the insertion point in the rendering method is to be in direct mixing with the output of the rendering process applied to the at least one spatial audio signal.
[0041] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured by the at least one processor to cause the apparatus to at least: acquire at least one spatial audio signal for rendering within an immersive audio scene; acquire a communication audio signal and position information associated with the communication audio signal; acquire rendering process parameters associated with the communication audio signal; determine a rendering method based on the rendering process parameters; and determine an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0042] The device may further be configured to generate at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method, and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0043] The apparatus may further be adapted to determine at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal.
[0044] An apparatus adapted to determine an insertion point in a rendering process for a determined rendering method and / or a selection of rendering elements for the determined rendering method based on rendering process parameters may further be adapted to perform at least one of: determining an insertion point in the rendering process further based on determined at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal; and determining a rendering method and / or a selection of rendering elements for the determined rendering method based on determined at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal.
[0045] The allowable delay value may be the amount of delay allowed for utilizing the communication audio signal, and the communication audio signal delay may be a delay value determined based on the end-to-end delivery delay and the delay in rendering the communication audio.
[0046] The audio format associated with the communication audio signal may include one of a one-way communication audio signal and a conversational communication audio signal between users in an immersive scene.
[0047] An apparatus for generating at least one output spatial audio signal from at least one spatial audio signal and a communication audio signal based on a determined rendering method and an insertion point in a rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on rendering process parameters may be configured to represent the communication audio signal as a high-order Ambisonic audio signal.
[0048] The apparatus may be further configured to obtain user input, and the apparatus is configured to generate at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the rendering method determined based on the rendering process parameters. The apparatus may further be configured to generate the at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the user input, and the user input may be configured to define at least one of allowed communication audio signal types, allowed audio formats, allowed delay values, and at least one acoustic modeling preference parameter.
[0049] The apparatus may be configured to obtain a communication audio signal type associated with the at least one spatial audio signal, and the apparatus configured to generate at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the rendering method determined based on the rendering process parameters may further be configured to generate the at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the at least one communication audio signal type associated with the at least one spatial audio signal.
[0050] The rendering processes and / or rendering elements may include one or more of Doppler processing, direct sound processing, material filtering processing, early reflection processing, diffuse late reverberation processing, sound source extension processing, occlusion processing, diffraction processing, sound source transformation processing, externalization rendering, and in-head rendering.
[0051] An apparatus configured to determine an insertion point in a rendering process for a determined rendering method and / or a selection of rendering elements for the determined rendering method based on rendering process parameters may be configured to determine a rendering mode, the rendering mode including a value indicating an insertion point of the communication audio signal.
[0052] The value indicating the insertion point may include one of a first mode value indicating that the communication audio signal and the at least one spatial audio signal are inserted at the start of the rendering processing method, a second mode value indicating that the communication audio signal is to bypass the rendering processing and be mixed directly with the output of the rendering processing applied to the at least one spatial audio signal, and a third mode value indicating that the rendering processing is to be fully applied to the at least one spatial audio signal while the communication audio signal is partially rendered.
[0053] A third mode value indicating that the communication audio signal is partially rendered may be a value indicating that the communication audio signal is a direct sound rendering of a point sound source and a binaural rendering relative to the user position.
[0054] The device may be further characterized by determining an audio format type of the communication audio signal based on the rendering process parameters.
[0055] The device for selecting an insertion point in the rendering process of the determined rendering method and / or a rendering element of the determined rendering method based on the rendering process parameters may be configured to determine an insertion point in the rendering process of the communication audio signal in the determined rendering method based on the audio format type.
[0056] An apparatus for determining an insertion point in a rendering process of a communication audio signal within a rendering method determined based on an audio format type may be configured to determine that, if the communication audio signal has an audio format type of a pre-rendered spatial audio format, the insertion point in the rendering method is in direct mixing with the output of the rendering process applied to the at least one spatial audio signal.
[0057] According to a fourth aspect, there is provided an apparatus comprising: means for obtaining at least one spatial audio signal for rendering in an immersive audio scene; means for obtaining a communication audio signal and position information related to the communication audio signal; means for obtaining rendering process parameters related to the communication audio signal; means for determining a rendering method based on the rendering process parameters; and means for determining an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0058] According to a fifth aspect, there is provided a computer program comprising instructions (or a computer readable medium comprising program instructions) to cause an apparatus to at least: acquire at least one spatial audio signal for rendering within an immersive audio scene; acquire a communication audio signal and position information associated with the communication audio signal; acquire rendering process parameters associated with the communication audio signal; determine a rendering method based on the rendering process parameters; and determine an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0059] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus to at least acquire at least one spatial audio signal for rendering within an immersive audio scene; acquire a communication audio signal and position information associated with the communication audio signal; acquire rendering process parameters associated with the communication audio signal; determine a rendering method based on the rendering process parameters; and determine an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0060] According to a seventh aspect, there is provided an apparatus comprising: an acquisition circuit configured to acquire at least one spatial audio signal for rendering within an immersive audio scene; an acquisition circuit configured to acquire a communication audio signal and position information associated with the communication audio signal; an acquisition circuit configured to acquire rendering process parameters associated with the communication audio signal; a decision circuit configured to determine a rendering method based on the rendering process parameters; and a decision circuit configured to determine an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0061] According to an eighth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus to at least: acquire at least one spatial audio signal for rendering within an immersive audio scene; acquire a communication audio signal and position information associated with the communication audio signal; acquire rendering process parameters associated with the communication audio signal; determine a rendering method based on the rendering process parameters; and determine an insertion point in the rendering process for the determined rendering method and / or a selection of rendering elements for the determined rendering method based on the rendering process parameters.
[0062] An apparatus comprising means for performing the operations of the above method.
[0063] An apparatus configured to perform the operations of the above method.
[0064] A computer program comprising program instructions for causing a computer to carry out the above method.
[0065] A computer program product stored on the medium may cause an apparatus to perform the methods described herein.
[0066] The electronic device may include the apparatus described herein.
[0067] The chipset may include the devices described herein. [Brief explanation of the drawings]
[0068] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings, in which: [Figure 1] FIG. 1 is a schematic diagram of a suitable environment showing an example of the combination of virtual scene elements within a physical listening space. [Figure 2]FIG. 2 is a diagram that schematically illustrates a system of devices for implementing an exemplary capture-to-rendering of an augmented reality scene, according to some embodiments. [Figure 3] FIG. 3 shows a flow diagram of the operation of a system of devices such as that shown in FIG. 2, according to some embodiments. [Figure 4] 3A and 3B are diagrams illustrating examples of the renderer shown in FIG. 2 according to some embodiments. [Figure 5] 5A and 5B are schematic diagrams illustrating an example of a scene state derivation audio processor shown in FIG. 4, according to some embodiments. [Figure 6] 6 illustrates a flow diagram of operations within the exemplary scene state derivation audio processor shown in FIG. 5, according to some embodiments. [Figure 7] 1 illustrates a flow diagram of operations for implementing communication audio within a system according to some embodiments. [Figure 8] FIG. 8 shows a schematic diagram of an exemplary device suitable for implementing the illustrated apparatus. DETAILED DESCRIPTION OF THE INVENTION
[0069] The following describes in more detail suitable devices and possible mechanisms for rendering augmented reality (AR) scene experiences and providing immersive audio communication processing capabilities.
[0070] As mentioned above, it is envisioned that "social VR" will be specified as a requirement of the MPEG-I6DoF audio standard. This requirement is envisioned to require that a system support the rendering of voice and audio from other users within a virtual environment. The voice and audio may be immersive. Furthermore, in some embodiments, the apparatus and method are envisioned to support low-latency conversation between users within a given virtual environment. Furthermore, the apparatus and method must support low-latency conversation between users within the given virtual environment and users outside the given virtual environment.
[0071] Additionally, the apparatus and method should enable audio and video synchronization of users and scenes, and further support metadata that specifies restrictions and recommendations on the rendering of voice / audio from other users (e.g., regarding placement and sound levels).
[0072] Thus, the embodiments described herein may implement low latency communication solutions such as, for example, 3GPP EVS / IVAS, and interface with or implement MPEG-I6DoF rendering.
[0073] The differences between immersive audio signals used in 6DoF scene rendering and telecommunications audio signals may lie in decoding and rendering delays, different delivery mechanisms, and different usage constraints. Streaming or content delivery delays (such as those obtained with DASH) are expected or tolerated in 6DoF immersive audio delivery, as these are not low-latency delivery methods, whereas telecommunications audio typically requires interactivity or low latency.
[0074] Because the rendering of communication audio in an immersive audio scene (MPEG Immersive Audio) is dynamic in nature, pre-determined rendering characteristics may not always be sufficient. This is because the delay budget for communication audio depends on several factors. One important factor is the use case (whether the communication is one-way commentary or two-way conversation, etc.), which requires different delay constraints depending on the use case. Another factor is the network conditions for communication audio, which may differ for different use cases within a single virtual immersive audio usage session or between different usage sessions. Changes in network conditions may result in different rendering delay budgets.
[0075] This can be considered similar to the situation in AR, where the usage scenario is unknown, but communication audio processing is more dynamic and can change significantly over time. Therefore, for communication audio, the delay budget (e.g., specified in the EIF or known via any higher-level module) can be known or estimated during content creation. However, communication audio delay parameters almost always vary (e.g., end-to-end delay, jitter, etc.) and are unknown during content creation. In an embodiment of the present invention, the content creator delay budget parameter can implicitly signal the budget by indicating the rendering mode of the communication audio. For example, a value of 0 indicates that the communication audio passes through the start of the entire immersive audio rendering pipeline; a value of 1 indicates that the communication audio bypasses the entire rendering pipeline and is directly mixed at the final stage; and a value of 2 indicates that the communication audio passes through minimal rendering (e.g., direct sound rendering of a point sound source and binaural rendering of a specified position).
[0076] The concepts described in further detail in the following embodiments relate to apparatus and methods configured to reliably render audio signals to a listener or user in the presence of various implementation-specific delay differences that are not significantly perceptible during streaming-based consumption of 6DoF scene data. These embodiments are configured to enable interactive usage scenarios to be implemented (as streaming-based consumption is less sensitive to delays than interactive consumption).
[0077] In the following description, the tolerable delay budget is the amount of delay that is tolerated for utilizing received communication audio. Typically, this is composed of the end-to-end delivery delay and the delay for rendering the communication audio into the 6DoF immersive audio scene. In the following, this term is used interchangeably with the tolerable delay value. Furthermore, in the following, network variations refer to changes in end-to-end network delivery, which can be based on delay, jitter, etc.
[0078] The embodiments described below relate to rendering communication audio within an immersive audio scene with six degrees of freedom (i.e., the listener can move within the scene and the listener's position and pose are tracked). Apparatuses and methods are described that are configured to ensure desired (e.g., player, user, or content creator specified) rendering of received communication audio within an immersive audio scene, despite the dynamic nature of the communication audio due to network fluctuations and immersive audio scene-dependent characteristics. In some embodiments, this can be achieved by deriving position information associated with the received communication audio (which can be, for example, an EIF placeholder indicating where to "place" the communication audio) from a bitstream associated with at least one spatial audio signal. In another embodiment, a communication audio placeholder can be added to any audio element, such as an audio object source, channel source, or HOA source. Audio element properties for each signal type (HOA, object, channel) can be added to such communication audio placeholders. In some embodiments, the apparatus and method may be further configured to obtain an allowable delay budget associated with the communication audio (e.g., obtained from the system, one-way or two-way, or from the EIF). Additionally, in some embodiments, the apparatus and method may be configured to obtain an audio format associated with the communication audio based on the allowable delay. Embodiments may be further configured to determine an appropriate insertion point in the rendering pipeline (in other words, derive where in the audio signal pipeline the communication audio is received for processing) based on the allowable delay, the communication audio format, the communication audio delay, and the audio rendering pipeline declaration.The apparatus and method according to some embodiments are further configured to determine at least one processing method and spatial rendering parameters for the communication audio based on at least one of the determined allowable delay, the communication audio delay, and the audio format.
[0079] Thus, in some embodiments, the processing method relates to selecting and processing communication audio depending on the signal type (e.g., whether the signal type is an object format signal, an HOA signal, or a channel format signal), and the rendering parameters may then refer to indicate one or more of the rendering approaches to apply. For example, the rendering parameters may indicate that the rendering stage employs point-source rendering of direct sound and binaural rendering of audio objects.
[0080] In some further embodiments, the rendering parameters may be used to indicate or control other elements of the rendering process. For example, in some embodiments, the rendering parameters may be employed to indicate the selection or skipping of rendering stages. Thus, in some embodiments, based on a certain insertion point within the rendering process operation, the start of rendering follows a certain sequence that is dependent on the audio rendering pipeline stage sequence. However, the rendering parameters may be employed to indicate the skipping of some intermediate stages within the rendering process sequence.
[0081] In some embodiments where the communication audio signal is represented as a Higher Order Ambisonics (HOA) format, whether it is rendered as a single-point HOA with internal transformation or as a 3DoF HOA depends on the allowable rendering delay. In some embodiments, the communication audio is implemented as audio objects or channels, and the amount of acoustic modeling performed on the audio objects depends on the network delay and the allowable usage delay.
[0082] In some embodiments, the communication audio rendering is adapted according to preferences of the immersive audio scene player, which preferences include at least one of the following aspects: The allowed communication audio types (e.g., one-way, conversational, between users using the same 6DoF scene, etc.) Allowed communication audio formats (e.g., audio objects, channels, 3DoF HOA, single HOA with transform) Allowable delay for rendering communications audio Acoustic Modeling Preferences
[0083] In some embodiments, depending on latency preferences, communication audio can be input and processed through the 6DoF audio rendering pipeline or mixed separately.
[0084] For example, if an immersive audio scene is configured to place a communications audio signal with maximum acoustic merging, but the delay budget dictates that certain features such as diffraction, occlusion, etc. are not possible, the renderer will determine the audio object path to minimize the occurrence of occlusion or diffraction.
[0085] 2 shows an overview of an end-to-end AR / XR 6DoF audio system. The example shows three parts of the system: a capture / generator device 201 configured to capture / generate and store / transmit audio information and associated metadata; and an augmented reality (AR) device 207 configured to output an appropriately processed audio signal based on the audio information and associated metadata. The example AR device 207 shown in FIG. 2 includes a 6DoF audio player 205 that retrieves and renders a 6DoF bitstream from a storage / distribution device 203.
[0086] In some embodiments, such as shown in FIG. 2 , the capture / generator device 201 includes an encoder input format (EIF) generator 211. The encoder input format (EIF) generator 211 (or, more generally, a scene definer) is configured to define a 6DoF audio scene. In some embodiments, the scene may be described by EIF (Encoder Input Format) or any other suitable 6DoF scene description format. EIF also refers to audio data that constitutes the audio scene. The encoder input format (EIF) generator 211 is configured to create EIF (Encoder Input Format) data, which is a content creator's scene description. The scene description information includes geometric information of the virtual scene, such as the positions of audio elements. Additionally, the scene description information may include other relevant metadata, such as directivity, size, and other acoustically relevant elements. For example, the relevant metadata may include the positions of virtual walls and their acoustic properties, as well as other acoustically relevant objects, such as occlusions. Examples of acoustic properties are acoustic material properties, such as (frequency-dependent) absorption or reflection coefficients, the amount of scattered energy, or transmission characteristics. In some embodiments, the virtual acoustic environment may be described according to its (frequency-dependent) reverberation time or diffuse-to-direct sound ratio. The EIF generator 211 in some embodiments is more commonly known as a virtual scene information generator. The EIF parameters 214 may, in some embodiments, be provided to a suitable (MPEG-I) encoder 217.
[0087] In some embodiments, the capture / generator device 201 includes an audio content generator 213. The audio content generator 213 is configured to generate audio content corresponding to an audio scene. In some embodiments, the audio content generator 213 is configured to generate and / or acquire audio signals associated with the virtual scene. For example, in some embodiments, these audio signals may be acquired or captured using a suitable microphone or microphone array, may be based on processed captured audio signals, or may be synthesized. In some embodiments, the audio content generator 213 is further configured to generate or acquire audio parameters associated with the audio signals, such as their position within the virtual scene, signal directionality, etc. The audio signals and / or parameters 212 may, in some embodiments, be provided to a suitable (MPEG-I) encoder 217.
[0088] In some embodiments, the capture / generator device 201 includes a communication audio processing data generator 215. The communication audio processing data generator 215 is configured to generate information carried in the content creator bitstream to indicate what type of communication audio (e.g., interactive, one-way, etc.) is allowed for this particular immersive audio scene. For example, a content creator may allow incoming communication audio from any caller, allow communication audio only from other users using the same 6DoF audio content, or allow communication audio between any two users using any 6DoF audio content. Additionally, the content creator bitstream carries information about which rendering stages are allowed and which are not.
[0089] In some implementations, communication audio processing parameters may depend on device profile preferences, application settings, or user preference settings.
[0090] For example, in some embodiments, the parameters may be implemented in a structure such as ObjectSourceCAStruct(). The ObjectSourceCAStruct() structure is an extension of the audio object metadata. In some embodiments, this structure may appear as a structure within the audio object metadata. The following examples describe audio objects, but can similarly be extended to communication audio structures for HOAs and channels.
[0091] aligned(8) ObjectSourceCAStruct(){ unsigned int(16) object_audio_identifier; / / object audio index unsigned int(1) ca_prototype_flag; / / commmunication audio prototype flag unsigned int(1) active; / / active or inactive flag unsigned int(1) hasExtent; unsigned int(32) gainDB; unsigned int(32) referenceDistance; bit(5) reserved = 0; if(ca_prototype_flag){ unsigned int(1) exclude_clustering_flag; / / communication audio is excluded from clustering bit(7) reserved = 0; CommunicationAudioIngestionStruct(); DynamicIndexStruct(); } else { MPEGHDecodedAudioIndex; / / index to obtain MPEG-H encoded audio stream Location(); if(hasExtent) ExtentStruct(); } } aligned(8) DynamicIndexStruct(){ unsigned int(16) stream_identifier; / / dynamic ID allocated by renderer / player for the communication audio } aligned(8) Location(){ signed int(32) pos_x; signed int(32) pos_y; signed int(32) pos_z; signed int(32) orient_yaw; signed int(32) orient_pitch; signed int(32) orient_roll; unsigned int(1) cspace; / / with respect to listening space origin if 1 with respect to user if 0 bit(7) reserved = 0; }
[0092] A ca_prototype_flag of 1 indicates to the player that it should prepare to receive communication audio. Communication audio ingestion-related information is described by CommunicationAudioIngestionStruct(), which also contains information about the type(s) of communication audio that are allowed or permitted for the 6DoF audio scene. Furthermore, there is a flag, exclude_clustering_flag, that indicates whether clustering is possible for communication audio. If this flag is absent, clustering is disabled by default. These communication audio types can be two-way, one-way, two-way, between users using the same content, or between users using different content. The ingestion structure also carries information about required (RequiredRenderingStagesStruct()) and disallowed (DisallowedRenderingStagesStruct()) rendering stages. Furthermore, rendering_modes can also compactly indicate whether communication audio needs to be processed through the immersive audio rendering pipeline or be completely bypassed. In the latter case, the communication audio is rendered outside the immersive rendering pipeline and mixed with the output of the immersive audio content rendering pipeline. If the rendering_modes_present flag value is absent or equal to 0, rendering is performed according to the communication audio signal element properties. Typically, if the ca_rendering_modes_present value is 1, other data structures such as RequiredRenderingStagesStruct(), DisallowedRenderingStagesStruct(), and ca_rendering_max_latency do not need to be present.
[0093] aligned(8) CommunicationAudioIngestionStruct(){ unsigned int(1) ca_co_conversational_allowed; / / bidirectional call with another user in the same 6DoF immersive audio scene unsigned int(1) ca_co_oneway_allowed; / / commentary from another user in the same 6DoF immersive audio scene unsigned int(1) ca_conversational_allowed; / / bidirectional call unsigned int(1) ca_oneway_allowed; / / commentary if(ca_co_conversational_allowed){ unsigned int(1) ca_delay_threshold_present; unsigned int(1) required_stages_present; unsigned int(1) disallowed_rendering_stages_present; unsigned int(1) ca_rendering_modes_present; bit(4) reserved = 0; if(ca_delay_threshold_present) unsigned int(32) ca_rendering_maxlatency; if(required_stages_present) RequiredRenderingStagesStruct(); if(disallowed_stages_present) DisallowedRenderingStagesStruct(); if(ca_rendering_modes_present) unsigned int(8) rendering_modes_type; } if(ca_co_oneway_allowed){ unsigned int(1) ca_delay_threshold_present; unsigned int(1) required_stages_present; unsigned int(1) disallowed_rendering_stages_present; unsigned int(1) ca_rendering_modes_present; bit(4) reserved = 0; if(ca_delay_threshold_present) unsigned int(32) ca_rendering_maxlatency; if(required_stages_present) RequiredRenderingStagesStruct(); if(disallowed_stages_present) DisallowedRenderingStagesStruct(); if(ca_rendering_modes_present) unsigned int(8) rendering_modes_type; } if(ca_oneway_allowed){ unsigned int(1) ca_delay_threshold_present; unsigned int(1) required_stages_present; unsigned int(1) disallowed_rendering_stages_present; unsigned int(1) ca_rendering_modes_present; bit(4) reserved = 0; if(ca_delay_threshold_present) unsigned int(32) ca_rendering_maxlatency; if(required_stages_present) RequiredRenderingStagesStruct(); if(disallowed_stages_present) DisallowedRenderingStagesStruct(); if(ca_rendering_modes_present) unsigned int(8) rendering_modes_type; } if(ca_conversational_allowed){ unsigned int(1) ca_delay_threshold_present; unsigned int(1) required_stages_present; unsigned int(1) disallowed_rendering_stages_present; unsigned int(1) ca_rendering_modes_present; bit(4) reserved = 0; if(ca_delay_threshold_present) unsigned int(32) ca_rendering_maxlatency; if(required_stages_present) RequiredRenderingStagesStruct(); if(disallowed_stages_present) DisallowedRenderingStagesStruct(); if(ca_rendering_modes_present) unsigned int(8) rendering_modes_type; } } aligned(8) RequiredRenderingStagesStruct(){ unsigned int(8) num_stages; for(i=0;i <num_stages;i++){ unsigned int(8) rendering_stage_idx; } } aligned(8) DisallowedRenderingStagesStruct(){ unsigned int(8) num_stages; for(i=0;i <num_stages;i++){ unsigned int(8) rendering_stage_idx; } }
[0094] [Table 1]
[0095] [Table 2]
[0096] In an embodiment, in addition to the rendering stage, the MPEG-H decoding delay is also taken into consideration to determine the tolerable delay threshold for communication audio rendering in an MPEG-I immersive audio scene. The decoding delay may depend on the audio format of the audio elements in the MPEG-I immersive audio scene.
[0097] In some embodiments, the capture / generator device 201 comprises an encoder 217. The encoder is configured to receive the EIF parameters 212, the communication audio processing parameters 216, and the audio signal / audio parameters 214 and decode them to generate an appropriate bitstream.
[0098] The encoder 217 can use, for example, the EIF parameters 212, the communication audio processing parameters 216, and the audio signal / audio parameters 214 to generate MPEG-I 6DoF audio scene content that is stored in a format suitable for streaming over a network. The distribution can be in any suitable format, such as MPEG-DASH (Dynamic Adaptive Streaming Over HTTP), HLS (HTTP Live Streaming), etc. The 6DoF bitstream carries the MPEG-H encoded audio content and the MPEG-I 6DoF bitstream. The content creator bitstream generated by the encoder based on the EIF and audio data can be formatted and encapsulated in a manner similar to an MHAS packet (MPEG-H 3D audio stream). In some embodiments, the encoded bitstream is passed to an appropriate content storage module 219. For example, as shown in FIG. 2, the encoded bitstream is passed to the MPEG-I 6DoF content storage module 219. In this example, the encoder 217 is located within the capture / generator device 201 , but it will be appreciated that the encoder 217 can be separate from the capture / generator device 201 .
[0099] In some embodiments, the capture / generator device 201 includes a content storage module. For example, as shown in FIG. 2, the encoded bitstream is passed to an MPEG-16DoF content storage 219 module. In such an embodiment, the audio signal is transmitted in a separate data stream from the encoded parameters. In some embodiments, the audio signal and parameters are stored / transmitted as a single data stream or format.
[0100] The content storage 219 is configured to store and provide content (including EIF-derived content creator bitstreams with communication audio processing parameters) to the AR device 207.
[0101] In some embodiments, the AR device 207 having a head-mounted device (HMD) is a playback device for AR use of 6DoF audio scenes.
[0102] In some embodiments, the AR device 207 has at least one AR sensor 221. The at least one AR sensor 221 may include a multimodal sensor such as a visual camera array, a depth sensor, LiDAR, etc. The multimodal sensor is used by the AR-utilizing device to generate information about the listening space. This information may include material information, objects of interest, etc. This sensor information, in some embodiments, can be passed to an AR processor 223.
[0103] In some embodiments, the AR device 207 includes at least one position / orientation sensor 227. The at least one position / orientation sensor 227 may include any suitable sensor or sensors configured to determine the position and / or orientation of the listener within the physical listening space. For example, the sensors may include a digital compass / gyroscope, a positioning beacon, etc. In some embodiments, the sensors employed in the AR sensor 221 are further used to determine the orientation and / or orientation of the listener. This sensor information may be passed to the renderer 235 in some embodiments.
[0104] In some embodiments, the AR device 207 includes an input configured to receive the communication audio 200 .
[0105] Additionally, in some embodiments, the AR device 207 includes a communication audio controller 253. The communication audio controller 253 is configured to output control information to the renderer 235 and control the integration of the communication audio 200 and the interactive audio (which may be part of the bitstream 220). In some embodiments, the communication audio controller 253 is configured to generate information in a format described below or any other suitable format. In some embodiments, this information is in the form of a rendering stage declaration that can signal a desired order so that any of the processing stages can be skipped without causing repetitive processing or undesirable output for subsequent rendering operations. The rendering pipeline declaration in some embodiments is available to the scene management processor. An exemplary structure for the information may be as follows:
[0106] aligned(8) RenderingStagesInfoStruct(){ unsigned int(8) num_stages; for(i=0;i <num_stages;i++){ unsigned int(8) render_stage_idx; unsigned int(32) mean_delay_value; unsigned int(32) sd_delay_value; unsigned int(8) input_audio_type; } }
[0107] In some embodiments, the rendering stages of the sonification pipeline are listed, with the values for mean_delay_value and sd_delay_value being -1 if not available.
[0108] In some embodiments, the communications audio controller 253 is configured to generate control information such as the format, delivery delay, and jitter of the communications audio. In some embodiments, this information can be pre-defined values stored in the renderer, so that if the controller 253 is not present or does not provide this information, the renderer is configured to use default values. For example, the delay value of a class type can be a pre-defined default value used by the renderer. In some embodiments, the control information can be passed in the following structure:
[0109] aligned(8) CommunicationAudioInfoStruct(){ unsigned int(4) ca_class_type; unsigned int(4) ca_format_type; unsigned int(32) ca_delivery_latency; }
[0110] [Table 3]
[0111] [Table 4]
[0112] If the communication audio format is 0 or 1, it is input as communication audio by the immersive audio rendering pipeline. However, in the case of pre-rendered spatial audio, the communication audio is input directly into the mixer block.
[0113] In another embodiment, communication audio processing is embedded within audio element metadata specified for objects, channels, and HOA sources. As a result, communication audio can be rendered by an MPEG-I renderer as other audio elements with communication audio-specific properties. These communication audio-specific rendering preferences are specified as control data for selecting and / or rejecting one or more rendering stages.
[0114] aligned(8) ObjectSourceStruct(){ unsigned int(16) index; / / object audio index unsigned int(1) ca_flag; / / placeholder for rendering communication audio as audio object unsigned int(1) active; / / active or inactive flag unsigned int(1) hasExtent; unsigned int(32) gainDB; unsigned int(32) referenceDistance; bit(5) reserved = 0; if(ca_flag){ CommunicationAudioRenderingStruct(); unsigned int(16) CommunicationAudioIndex; / / Identifier to receive the communication audio stream } else { unsigned int(16) MPEGHDecodedAudioIndex; / / index to obtain MPEG-H encoded audio stream Location(); / / index to obtain MPEG-H encoded audio stream } if(hasExtent) ExtentStruct(); } aligned(8) CommunicationAudioRenderingStruct(){ unsigned int(1) rendering_modes_present; unsigned int(1) dynamic_modes_; if(rendering_modes_present){ unsigned int(8) rendering_modes; } bit(7) reserved = 0; }
[0115] The ObjectSourceStruct(), ObjectSourceCAStruct(), or any of their configuration parameters or structures may change over the duration of an audio scene. As a result, communication audio in some embodiments may be allowed or disallowed depending on the prevailing metadata information. Additionally, the rendering mode or insertion point may change over the duration of an audio scene.
[0116] In yet another embodiment, the communication audio metadata, eg, the audio element metadata mentioned above that is a placeholder for communication audio, carries an indication flag communicationAudioRenderImmediateFlag.
[0117] If communicationAudioRenderImmediateFlag==0, communication audio is rendered or mixed immediately into the rendered immersive audio scene without additional delay.
[0118] If communicationAudioRenderImmediateFlag==1, communication audio is rendered according to the rendering metadata and properties specified on the audio element.
[0119] Communications audio metadata can also be delivered as dynamic updates to the renderer. Communications audio can be delivered as a new MHAS packet with PACTYP_CAAUDIODATA, and labels can be used to indicate the corresponding metadata and the dynamic update metadata that applies to PACTYP_CAAUDIODATA. The PACTYP_CAAUDIODATA packet carries a payload of ObjectSourceStruct(), HOASourceStruct(), ChannelSourceStruct(), or a subset of these structures, along with communications audio rendering or capture parameters.
[0120] In one implementation, PACTYP_CAAUDIODATA carries audio data in the form of PCM. As a result, PACTYP_CAAUDIODATA is followed by PACTYP_PCMCONFIG and PACTYP_PCMDATA. The preceding PACTYP_CAAUDIODATA packet allows the renderer to identify the PCM data that corresponds to the transmitted audio data.
[0121] In some embodiments, the AR device 207 has a suitable output device, which in the example shown in Figure 2 is shown as headphones 241 configured to receive spatial audio output 240 generated by the renderer 235, although any suitable output transducer may be arranged.
[0122] In some embodiments, the AR device 207 includes a player / renderer apparatus 205. The player / renderer apparatus 205 is configured to receive a bitstream including the EIF-derived content creator bitstream 220, AR sensor information, user position and / or posture information, communication audio 220, and control information from a communication audio controller, and to determine from this information an appropriate spatial audio output 240 (which may be incorporated within the AR device 207) that can be passed to an appropriate output device, shown in FIG. 2 as headphones 241.
[0123] In some embodiments, the player / renderer device 205 includes an AR processor 223. The AR processor 223 is configured to receive sensor information from at least one AR sensor 221 and generate appropriate AR information that can be passed to the LSDF generator 225. For example, in some embodiments, the AR processor is configured to perform fusion of the sensor information from each of the sensor types.
[0124] In some embodiments, the player / renderer device 205 includes a listening space description file (LSDF) generator 225. The listening space description file (LSDF) generator 225 is configured to receive the output of the AR processor 223 and generate a listening space description for AR use from information obtained from the AR sensing interface. The format of the listening space can be any appropriate format. The LSDF format can be used to create the LSDF. This description conveys listening space or room information, including acoustic properties (e.g., a mesh enclosing the listening space, including the materials of the mesh surfaces) and spatially variable elements of the scene, called anchors in the listening space description. The LSDF generator is configured to output this listening scene description information to the renderer 235.
[0125] In some embodiments, the player / renderer device 205 comprises a receive buffer 231 configured to receive the content creator bitstream (including EIF information) 220. The buffer 231 is configured to pass the received data and pass the data to a decoder 233.
[0126] In some embodiments, the player / renderer device 205 has a decoder 233 configured to obtain the encoded bitstream from the buffer 231 and output the decoded EIF information and communication audio processing parameters (together with the decoded audio data if in the same data stream) to the renderer 235.
[0127] In some embodiments, the player / renderer device 205 includes a communication receiver buffer and decoder 251 configured to receive the communication audio 200 and decode the encoded audio data to pass it to the renderer 235.
[0128] In some embodiments, the player / renderer device 205 comprises a renderer 235. The renderer 235 is configured to receive the decoded EIF information (including the decoded immersive audio data if in the same data stream), the listening scene description information, the listener position and / or attitude information, the decoded communication audio, and the communication audio control information. The renderer 235 is configured to generate spatial audio output signals and pass these to an output device, such as shown in FIG. 2 by spatial audio output 240 to headphones 241.
[0129] With reference to FIG. 3, an example of the operation of the system shown in FIG. 2 is shown.
[0130] As shown in FIG. 3, in step 301, communication audio processing data is obtained (or generated).
[0131] The method may include generating or alternatively obtaining EIF information, as shown in FIG. 3, via step 303.
[0132] Further, as shown in FIG. 3, audio data is obtained (or generated) by step 305 .
[0133] Then, as shown in FIG. 3, step 307 encodes the EIF information, the communication audio processing data, and the audio data.
[0134] The encoded data is then stored / retrieved or transmitted / received, per step 309, as shown in FIG.
[0135] Further, as shown in FIG. 3, AR scene data is acquired by step 311.
[0136] As shown in FIG. 3, step 313 generates listening space description (file) information from the detected AR scene data.
[0137] Further, as shown in FIG. 3, communication audio control information may be obtained via step 312 .
[0138] Also shown in FIG. 3, step 314 obtains communication audio data.
[0139] Additionally, as shown in FIG. 3, position and / or attitude data of the listener / user may be obtained by step 315 .
[0140] A spatial audio signal may then be rendered based on the audio data, communication audio control information, communication audio, EIF information, LSDF data, and position and / or attitude data. Specifically, rendering involves combining the audio signals, as shown in FIG. 3 at step 317.
[0141] After rendering the spatial audio signals, they may be output to an appropriate output device, such as headphones, by step 319, as shown in FIG.
[0142] FIG. 4 illustrates an exemplary renderer 235 suitable for implementing some embodiments and may be configured to provide a fused audio signal.
[0143] For example, Figure 4 shows that before the renderer 235 there is a bitstream parser 401 configured to receive the decoded 6DoF bitstream. The parsed EIF data can then be passed to a scene manager / processor 403.
[0144] The renderer 235 in some embodiments includes a scene manager / processor 403 configured to receive the parsed EIF from the bitstream parser 401, communication control information such as delay, jitter, and parameters defining the format of the communication audio data 402.
[0145] In some embodiments, the scene manager / processor 403 includes a communications audio adaptation processor 411 configured to control the auralization pipeline (or DSP processing for rendering) according to information such as content creator preferences (obtained from the bitstream) regarding communications audio processing, communications audio tolerance budget for delay, etc.
[0146] The scene management information may then be passed to the scene state derivation audio processor 405.
[0147] The scene manager / processor 403 may further be configured to take the decoded 6DoF audio signal, the processed scene information, and the position and / or pose of the listener, and generate therefrom a spatial audio signal output. As indicated above, the effect of the scene manager / processor 403 is such that any known or suitable spatial audio processing implementation can be employed (the auralization pipeline is not involved in prior scene processing).
[0148] In some embodiments, the renderer 235 includes a scene state derived audio processor (DSP processing and auralization) 405. The scene state derived audio processor (DSP processing and auralization) 405 is configured to receive configuration information from the scene manager / processor 403, the decoded immersive audio (MPEG-I audio / decoded MPEG-H audio) 400, and the decoded communication audio signal 450, and generate a spatial audio signal.
[0149] With reference to FIG. 5, an exemplary scene state derivation audio processor (DSP processing and auralization) 405 is shown. The scene state derivation audio processor is configured to obtain decoded immersive audio (IA) and communication audio (CA) flows through different rendering modules in the auralization pipeline. In some embodiments, the processor 405 is configured to employ multiple processing paths based on the IA format type (IA1 and IA2). Similarly, the communication audio (CA) may have multiple paths (in this example, there are three candidate paths, CA1, CA2, and CA3, according to the CA format). Different rendering modules are annotated with index numbers indicating possible insertion points for IA or CA. Furthermore, the exemplary processor 405 shown herein has two output options: a first option (O2) with a mixer and another option (O1) without a mixer.
[0150] The first path IA1-CA1 is configured to apply the following processing operations to the audio signal, and to apply audio object modeling operations:
[0151] In some embodiments, the scene state derived audio processor (DSP processing and auralization) 405 has a Doppler processor 501 configured to control Doppler processing of immersive audio (IA) and communications audio (CA) based on control information from configuration information from the scene manager / processor 403.
[0152] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has a direct sound processor 503 configured to control direct sound processing of immersive audio (IA) and communications audio (CA) based on control information of the configuration information from the scene manager / processor 403.
[0153] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has a material filter processor 505 configured to control material filtering of immersive audio (IA) and communication audio (CA) based on control information of the setting information from the scene manager / processor 403.
[0154] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has an early reflection processor 507 configured to control early reflection sound processing of immersive audio (IA) and communications audio (CA) based on control information from configuration information from the scene manager / processor 403.
[0155] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has a late reverberation processor 509 configured to control late reverberation processing of immersive audio (IA) and communications audio (CA) based on control information from configuration information from the scene manager / processor 403.
[0156] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has an extension processor configured to control extension sound processing of immersive audio (IA) and communications audio (CA) based on control information from configuration information from the scene manager / processor 403.
[0157] It will be appreciated that the order of the above processes may be any suitable order. The output of the object processing processor may be passed to a higher order Ambisonics or spatializer processor 541.
[0158] The second path IA2-CA2 applies the following processing operations to the audio signal, which may apply higher order Ambisonics processing:
[0159] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has an SP High Order Ambisonics Processor 521 configured to control SP High Order Ambisonics processing of immersive audio (IA) and communications audio (CA) based on control information from configuration information from the scene manager / processor 403.
[0160] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has an MP High Order Ambisonic Processor 523 configured to control MP High Order Ambisonic processing of immersive audio (IA) and communications audio (CA) based on control information from configuration information from the scene manager / processor 403.
[0161] The output of the higher order Ambisonics processing may, in some embodiments, be passed to a higher order Ambisonics spatializer processor 541.
[0162] The first and second paths go to the high order Ambisonics of the spatializer processor 541 and are configured to select either one of the two paths to output as output O1 or to the mixer 551 based on the audio format, such as control information from the setting information from the scene manager / processor 403.
[0163] The third path CA3 may apply the following processing operations to the audio signal, and may also apply rendering processing.
[0164] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has a communications audio output rendering audio processor 531 configured to control the processing of communications audio (CA) based on control information from configuration information from the scene manager / processor 403.
[0165] The output of the communication audio output rendering audio processor 531 may, in some embodiments, be passed to a mixer 551 .
[0166] In some embodiments, the scene state derivation audio processor (DSP processing and auralization) 405 has a mixer configured to receive the outputs of the communication audio output rendering audio processor 531 and the high-order Ambisonics or spatializer processor 541 and mix them based on control information from the setting information from the scene manager / processor 403 to generate a mixed output O2.
[0167] An exemplary flow diagram illustrating the operation of the scene state derivation audio processor (DSP processing and auralization) 405 shown in some embodiments is shown in FIG.
[0168] The decoded immersive audio input is received or obtained from the listening space description by step 601, as shown in FIG.
[0169] Further, communication audio input is received or acquired by step 603 as shown in FIG.
[0170] Additionally, a processing pipeline is shown.
[0171] The spatialization processing pipeline is shown by steps 605 to 615. These are Doppler processing shown in step 605 of Figure 6, direct sound processing shown in step 607 of Figure 6, material filtering processing shown in step 609 of Figure 6, early reflection processing shown in step 611 of Figure 6, late reverberation processing shown in step 613 of Figure 6, and extension processing shown in step 615 of Figure 6.
[0172] The Ambisonics processing pipeline is shown by steps 621 and 623. These are SP higher order Ambisonics processing, shown in Figure 6 as step 621, and MP higher order Ambisonics processing, shown in Figure 6 as step 623.
[0173] The pre-rendered communications audio output pipeline is indicated by the communications audio output rendering audio process by step 631 as shown in FIG.
[0174] Further, as shown in FIG. 6, step 625 presents the selected spatialized or Ambisonics pipeline output.
[0175] In some embodiments, the output of the selected spatialized or Ambisonics processing pipeline, in other words the output of the immersive audio rendering pipeline, is then output as an unmixed audio signal by step 627, as shown in FIG. 6.
[0176] In some embodiments, the selected spatialized or Ambisonics processing pipeline output and the pre-rendered communications audio output pipeline are mixed and output by step 641, as shown in FIG. 6.
[0177] FIG. 7 is a flow diagram illustrating a procedure for determining appropriate audio insertion points and selection of processing steps for rendering communication audio.
[0178] As shown in Figure 7, the method comprises receiving a content creator bitstream for communication audio processing in step 701, which may be obtained as part of an MPEG-I6DoF content.
[0179] A check is then made to determine whether the currently active immersive audio scene allows the use of communications audio, which is shown in Figure 7 as step 703.
[0180] If the usage is permitted and supported, the next action is configured to retrieve communication audio information, as shown in FIG. 7, by step 705 .
[0181] 7, a further check may be performed to determine whether the permitted communication audio type (e.g., description, interactivity, etc.) is supported, per step 707. In some embodiments (not shown) where the check determines that the audio type is not permitted, the method should pause the immersive audio scene and continue or switch the communication audio.
[0182] If the type is supported, then a rendering stage declaration and associated delay information is received or otherwise obtained, as shown in Figure 7, by step 709. This provides information about available rendering stages and potential insertion points for the communication audio.
[0183] In some embodiments, as shown in FIG. 7, a further checking operation may follow to determine whether delay threshold information is present, per step 711.
[0184] If delay threshold information is present, the method may be configured to utilize the delay of the rendering stage to select a stage from the declared rendering pipeline, as shown in FIG. 7, by step 713 .
[0185] In some embodiments, this is subtracting the communication audio delay (ca_delivery_latency) from the delay threshold to obtain an updated delay requirement; Using the communication audio format type, determine candidate rendering stages applicable to the format type (ca_format_type); If DisallowedRenderingStagesStruct() is present, its information can be used to discard stages from the associated declared rendering pipeline, Prioritize the inclusion of rendering stages represented in RequiredRenderingStagesStruct() while complying with the latest latency requirements; and Controlling the insertion or input of communication audio in the first rendering stage obtained by the above operation; It can be implemented by:
[0186] If the delay threshold information is not present, the method can be configured to select a stage from the declared rendering pipeline using the required and disallowed stage information from the content creator bitstream. Furthermore, if a rendering mode is present, the rendering is configured to be performed based on the rendering mode. The selection of a stage and, if present, the configuration of the rendering based on the rendering mode is shown in FIG. 7 at step 715. This can, for example, use the following operations: Obtaining the latest latency requirement based on the requirement derived by taking the difference between ca_delivery_latency and ca_class_type; determining, based on the communication audio format type, candidate rendering stages applicable to the ca_format_type; DisallowedRenderingStagesStruct(), if present, is used to discard stages from the associated declared rendering pipeline, Prioritize including rendering stages indicated by RequiredRenderingStagesStruct() while complying with the latest latency requirements, and Inserting communication audio into the first rendering stage obtained by the above operation.
[0187] The communication audio is then rendered, as shown in FIG. 7, via step 717.
[0188] In some embodiments, a new communication audio delay is then obtained. If the difference from the current estimated delay requirement changes by more than a predetermined threshold, the audio rendering pipeline is modified. This is shown in Figure 7, where step 719 performs a check step to determine whether there has been a change in the communication audio information; if the answer is yes, operation returns to step 705.
[0189] An example of a communications audio scenario for audio object formats is treating a communications audio signal whose format type is mono as an audio object if the content creator's bitstream specifies similar acoustic modeling as an audio object. The audio object is rendered with the specified acoustic processing steps in the rendering pipeline as long as the latency requirements are met. However, in another example, if a specific rendering processing step (e.g., a sound source extension, or such a step that adds significant rendering latency) is observed, that specific rendering step is omitted unless it is part of the required stage metadata.
[0190] A further example is HOA format communication audio, where communication audio is delivered as an HOA source with a content creator bitstream indicating the HOA source with translation support within the HOA source. It has been determined that if the rendering pipeline is selected to include single-HOA source rendering, it can accommodate this processing based on delay constraints. In another example of the same scene, it has been determined that the communication audio delay is too large to allow single-point HOA rendering with translation. As a result, the communication audio from the HOA source is rendered without translation processing, but is mixed directly with the immersive audio output in the mixing block.
[0191] 8 is an exemplary electronic device that can represent any of the devices described above. The device may be any suitable electronic device or device. For example, in some embodiments, device 1400 may be a mobile device, user equipment, tablet computer, computer, audio playback device, etc.
[0192] In some embodiments, the device 1400 includes at least one processor or central processing unit 1407. The processor 1407 may be configured to execute various program code, such as methods as described herein.
[0193] In some embodiments, the apparatus 1400 comprises a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 may be any suitable storage means. In some embodiments, the memory 1411 comprises a program code section for storing program code implementable by the processor 1407. Furthermore, in some embodiments, the memory 1411 may further comprise a storage data section for storing data, e.g., data that has been processed or is to be processed according to embodiments as described herein. The implemented program code stored in the program code section and the data stored in the storage data section may be retrieved by the processor 1407 whenever needed via the memory-processor coupling.
[0194] In some embodiments, device 1400 comprises a user interface 1405. User interface 1405, in some embodiments, may be coupled to processor 1407. In some embodiments, processor 1407 may control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 may allow a user to input instructions to device 1400, for example, via a keypad. In some embodiments, user interface 1405 may allow a user to obtain information from device 1400. For example, user interface 1405 may have a display configured to display information from device 1400 to a user. User interface 1405, in some embodiments, may be configured with a touchscreen or touch interface that can both allow information to be input into device 1400 and also display information to a user of device 1400. In some embodiments, user interface 1405 may be a user interface for communicating with a position-determining device as described herein.
[0195] In some embodiments, apparatus 1400 comprises an input / output port 1409. The input / output port 1409 in some embodiments comprises a transceiver. The transceiver in such embodiments may be coupled to processor 1407 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitting and / or receiving means may, in some embodiments, be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.
[0196] The transceiver may communicate with the further device by any suitable known communication protocol, for example, in some embodiments the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an infrared data channel (IRDA).
[0197] The transceiver input / output port 1409 may be configured to receive signals and, in some embodiments, determine parameters as described herein, using a processor 1407 executing appropriate code.
[0198] Furthermore, although exemplary embodiments have been described above, it is pointed out herein that several variations and modifications are possible to the disclosed solution without departing from the scope of the present invention.
[0199] In general, various embodiments may be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. Some aspects of the present disclosure may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, although the present disclosure is not limited thereto. While various aspects of the present disclosure may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it will be appreciated that these blocks, devices, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controller or other computing device, or combinations thereof.
[0200] As used herein, the term "circuitry" may refer to one or more or all of the following: (a) hardware-only circuit implementations (e.g., implementations in analog and / or digital circuits only); and (b) a combination of hardware circuitry and software, e.g., (where applicable); (i) a combination of analog and / or digital hardware circuitry and software / firmware; and (ii) software (including digital signal processors), hardware processor portions having software and memory, that work together to cause devices such as mobile phones and servers to perform various functions; and (c) Hardware circuits and processors, such as microprocessors or portions of microprocessors, that require software (e.g., firmware) to operate, but the software may be absent when not necessary for operation.
[0201] This definition of circuitry applies to all uses of the term in this application, including any claims. As a further example, as used in this application, the term circuitry also covers simply a hardware circuit or processor(s) or portion of a hardware circuit or processor and its(their) accompanying software and / or firmware implementation.
[0202] The term circuitry also covers, for example, baseband or processor integrated circuits for mobile terminals, or similar integrated circuits in servers, cellular network devices, or other computing or network devices, where applicable to particular claim elements.
[0203] Embodiments of the present disclosure may be implemented by computer software executable by a data processor of a mobile device, such as a processor entity, or by hardware, or a combination of software and hardware. Computer software or programs, also referred to as program products, including software routines, applets, and / or macros, may be stored on any device-readable data storage medium and consist of program instructions that perform specific tasks. A computer program product may consist of one or more computer-executable components configured to perform embodiments when the program is executed. The one or more computer-executable components may be at least one software code or portion thereof.
[0204] Further, in this regard, it should be noted that any block of logic flow as shown may represent a program step, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. Software may be stored on physical media such as memory chips or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs. Physical media are non-transitory media.
[0205] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memory, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), an FPGA, a gate level circuit, and a processor based on a multi-core processor architecture.
[0206] Embodiments of the present disclosure can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs ready to be etched onto semiconductor substrates.
[0207] The scope of protection sought for various embodiments of the present disclosure is defined by the independent claims. The embodiments and features described herein (if any) that do not fall within the scope of the independent claims shall be interpreted as examples useful for understanding various embodiments of the present disclosure.
[0208] The foregoing description has provided, by way of non-limiting example, a complete and informative description of the exemplary embodiments of the present disclosure. However, various modifications and adaptations will become apparent to those skilled in the relevant art in light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of the present invention as defined by the appended claims. Indeed, further embodiments exist that consist of a combination of one or more of the embodiments with any of the other embodiments previously described.
Claims
1. 1. An apparatus for rendering a communication audio signal within a six degrees of freedom immersive audio scene, the apparatus having at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured to, using the at least one processor, cause the apparatus to at least: obtaining at least one spatial audio signal for rendering within the six degrees of freedom immersive audio scene; obtaining the communication audio signal and location information associated with the communication audio signal; obtaining rendering process parameters associated with the communication audio signal and the location information; determining a rendering method based on the rendering process parameters; determining an insertion point for at least the communication audio signal in a rendering process for the determined rendering method and / or determining a selection of rendering elements for the determined rendering method; generating at least one output spatial audio signal for rendering the communication audio signal within the six-degree-of-freedom immersive audio scene according to a listener's position / pose, the at least one output spatial audio signal comprising: said at least one spatial audio signal; the communication audio signal; the determined insertion point; the position / attitude of the listener; is generated based on An apparatus configured to perform the following.
2. The device of claim 1, further configured to generate at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and the insertion point.
3. The device comprises: an audio format associated with said communication audio signal; the allowable delay value, and Communication audio signal delay, 2. The apparatus of claim 1, adapted to determine at least one of:
4. The determined insertion point in the rendering process may be determining the insertion point in the rendering process further based on the determined at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal; determining the rendering method and / or the selection of rendering elements for the determined rendering method based on the determined at least one of an audio format, a tolerable delay value, and a communication audio signal delay associated with the communication audio signal; 4. The apparatus of claim 3, further comprising:
5. 5. The device of claim 4, wherein the allowable delay value is an amount of delay allowed for utilizing the communication audio signal, and the communication audio signal delay is a delay value determined based on an end-to-end delivery delay and a delay for rendering the communication audio.
6. The device of claim 2 , wherein the generated at least one output spatial audio signal causes the device to represent the communication audio signal as a high-order Ambisonic audio signal.
7. The device is further adapted to obtain a user input, the device being further adapted to generate the at least one output spatial audio signal based on the user input, the user input comprising: allowed communication audio signal types, Allowed audio formats, Tolerable delay value, at least one acoustic modeling preference parameter; The apparatus of claim 2 , configured to define at least one of:
8. 3. The apparatus of claim 2, wherein the apparatus is further adapted to obtain a communication audio signal type associated with the at least one spatial audio signal, and wherein the apparatus is further adapted to generate the at least one output spatial audio signal based on the at least one communication audio signal type associated with the at least one spatial audio signal.
9. The rendering process and / or rendering element may include: Doppler processing, Direct sound processing, Material filtering, Early reflection processing, Diffuse post-reverberation processing, Sound source expansion processing, Shielding treatment, Diffraction Processing Sound source conversion processing, Externalized rendering, In-head rendering, The apparatus of claim 1 , comprising one or more of:
10. The device of claim 1 , wherein the determined insertion point causes the device to determine a rendering mode, the rendering mode comprising a value indicative of the insertion point of the communication audio signal.
11. The value indicating the insertion point is a first mode value indicating that the communication audio signal and the at least one spatial audio signal are inserted at the start of the rendering process; a second mode value indicating that the communication audio signal bypasses the rendering process and is mixed directly with the output of the rendering process applied to the at least one spatial audio signal; a third mode value indicating that the rendering process is fully applied to the at least one spatial audio signal, while the communication audio signal is partially rendered; 11. The apparatus of claim 10, comprising one of:
12. 12. The device of claim 11, wherein the third mode value indicating that the communication audio signal has been partially rendered is a value indicating that the communication audio signal is a direct sound rendering relative to a point sound source and a binaural rendering relative to a user position.
13. The device comprises: an audio format type of the communication audio signal based on the rendering process parameters; the insertion point in the rendering process of the communication audio signal within the rendering method determined based on the audio format type; 2. The apparatus of claim 1, adapted to determine at least one of:
14. 14. The apparatus of claim 13, wherein the determined insertion point in the rendering process within the determined rendering method based on the audio format type causes the apparatus to determine that if the communication audio signal has an audio format type of a pre-rendered spatial audio format, the insertion point in the rendering method is in direct mixing with an output of the rendering process applied to the at least one spatial audio signal.
15. 1. A method for an apparatus for rendering a communication audio signal within a six degrees of freedom immersive audio scene, said method comprising: obtaining at least one spatial audio signal for rendering within the six degrees of freedom immersive audio scene; obtaining the communication audio signal and location information associated with the communication audio signal; obtaining rendering process parameters and the position information associated with the communication audio signal; determining a rendering method based on the rendering process parameters; determining an insertion point for at least the communication audio signal in a rendering process for the determined rendering method and / or selecting a rendering element for the determined rendering method; generating at least one output spatial audio signal for rendering the communication audio signal within the six-degree-of-freedom immersive audio scene according to a listener's position / pose, the at least one output spatial audio signal comprising: said at least one spatial audio signal; the communication audio signal; the determined insertion point; the position / attitude of the listener; is generated based on The method includes:
16. The method of claim 15 , further comprising generating the at least one output spatial audio signal from the at least one spatial audio signal and the communication audio signal based on the determined rendering method and insertion point.
17. an audio format associated with the communication audio signal; Tolerable delay value, Communication audio signal delay, The method of claim 15 , further comprising determining at least one of:
18. Determining the insertion point in the rendering process includes: determining the insertion point in the rendering process further based on determining at least one of the audio format associated with the communication audio signal, the allowable delay value, and the communication audio signal delay; determining the selection of rendering elements for the determined rendering method based on determining at least one of the audio format, the allowed delay value, and the communication audio signal delay associated with the rendering method and / or the communication audio signal; 20. The method of claim 17, comprising at least one of:
19. The method of claim 15 , wherein determining the insertion point comprises determining a rendering mode, the rendering mode comprising a value indicative of the insertion point of the communication audio signal.
20. an audio format type of the communication audio signal based on the rendering process parameters; the insertion point in the rendering process of the communication audio signal within the rendering method determined based on the audio format type; The method of claim 15 , further comprising determining at least one of:
Citation Information
Patent Citations
Immersive audio communication
JP2008547290A
Stereophonic sound IP telephone using binaural recording
JP2015119248A
Receiver, content transfer system, and program
JP2021136465A