Audio representations and associated rendering

By embedding the spatial audio stream into the object stream and updating the metadata in the immersive audio codec, the problem of high complexity in audio format mixing and forwarding in the existing technology is solved, low-latency and high-quality audio processing is achieved, and the user experience is improved.

CN120751329APending Publication Date: 2025-10-03NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510935866.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-02-28
Filing Date
2021-02-10
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively handling the mixing and forwarding of multiple audio formats in immersive audio codecs, resulting in high complexity and increased latency, which particularly affects user experience in virtual reality and augmented reality applications.

Method used

By embedding the spatial audio stream into the object stream and processing it at the receiving end, and utilizing object metadata updates and scene description parameters, flexible embedding and rendering of the audio data stream can be achieved, avoiding decoding and encoding operations at intermediate nodes, and reducing complexity and latency.

Benefits of technology

It improves the flexibility of audio source mixing and forwarding, reduces latency and complexity, maintains the characteristics and quality of the original input, and optimizes the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751329A_ABST
    Figure CN120751329A_ABST
Patent Text Reader

Abstract

The invention relates to audio representations and associated rendering. There is provided an apparatus configured to: receive at least one first audio channel and at least one second audio channel, where at least one of the at least one first audio channel or the at least one second audio channel comprises spatial audio configured to enable immersive audio communication; determining a format of at least one of the at least one first audio channel or the at least one second audio channel to identify which of the received at least one first audio channel and at least one second audio channel includes the spatial audio; processing the identified at least one audio channel with at least one parameter according to the determined format; and rendering the processed at least one audio channel and the other of the at least one first audio channel or the at least one second audio channel.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese invention patent application with the invention name “Audio Representation and Associated Rendering” (application number 202180016741.7, application date February 10, 2021). Technical Field

[0002] The present application relates to apparatus and methods for soundfield-related audio representation and associated rendering, but not exclusively to apparatus and methods for audio representation for audio encoders and decoders. Background Art

[0003] Immersive audio codecs are being implemented to support a large number of operating points ranging from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed to be suitable for use on communication networks such as 3GPP 4G / 5G networks. Such immersive services include, for example, the use of immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR) and mixed reality (MR). The audio codec is expected to handle the encoding, decoding and rendering of speech, music and general audio. It is also expected that the audio codec supports channel-based audio and scene-based audio input, including spatial information about the sound field and sound sources. It is also expected that the codec operates with low latency to enable conversational services and supports high error robustness under various transmission conditions.

[0004] Furthermore, parametric spatial audio processing is an area of ​​audio signal processing that uses a set of parameters to describe the spatial aspects of sound. For example, when performing parametric spatial audio capture from a microphone array, estimating a set of parameters from the microphone array signal, such as the direction of the sound in a frequency band and the ratio of the directional to non-directional portion of the captured sound in the frequency band, is a typical and effective choice. It is well known that these parameters describe the perceived spatial characteristics of the captured sound at the position of the microphone array. These parameters can be used accordingly in the synthesis of spatial sound for use with binaural headphones, speakers, or other formats such as Ambisonics. Summary of the Invention

[0005] According to a first aspect, a device for immersive audio communication is provided, comprising components configured to: receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; determine a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and the second audio data stream includes the spatial audio stream; process the second audio data stream with at least one parameter according to the determined type; and render the first audio data stream and the processed second audio data stream.

[0006] The second audio data stream may be configured to include at least one other audio data stream, and wherein the at least one other audio data stream may include the determined type, and the at least one other audio data stream may be an embedded-level audio data stream relative to the second audio data stream.

[0007] The at least one further audio data stream may comprise at least one further embedding level, wherein each embedding level may comprise at least one additional audio data stream of the determined type.

[0008] The second audio data stream may be a primary audio data stream.

[0009] Each audio data stream may be further associated with at least one of: a stream identifier configured to uniquely identify the audio data stream; and a stream descriptor configured to describe the type of the audio data stream.

[0010] The type can be one of the following: mono audio signal type; immersive voice and audio service audio signal.

[0011] At least one parameter may be configured to define a room characteristic or scene description.

[0012] The at least one parameter defining the indoor characteristic or scene description may include at least one of: direction; direction azimuth; direction elevation; distance; gain; spatial range; energy ratio; and position.

[0013] The component may be further configured to: receive an additional audio data stream; and embed the additional audio data stream within one or the other of the first audio data stream and the second audio data stream.

[0014] According to a second aspect, a method for a device is provided, the method comprising: receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; determining a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and the second audio data stream includes the spatial audio stream; processing the second audio data stream with at least one parameter according to the determined type; and rendering the first audio data stream and the processed second audio data stream.

[0015] The second audio data stream may be configured to include at least one other audio data stream, and wherein the at least one other audio data stream may include the determined type, and the at least one other audio data stream may be an embedded-level audio data stream relative to the second audio data stream.

[0016] The at least one further audio data stream may comprise at least one further embedding level, wherein each embedding level may comprise at least one additional audio data stream of the determined type.

[0017] The second audio data stream may be a primary audio data stream.

[0018] Each audio data stream may be further associated with at least one of: a stream identifier configured to uniquely identify the audio data stream; and a stream descriptor configured to describe the type of the audio data stream.

[0019] The type can be one of the following: mono audio signal type; immersive voice and audio service audio signal.

[0020] At least one parameter may be configured to define a room characteristic or scene description.

[0021] The at least one parameter defining the indoor characteristic or scene description may include at least one of: direction; direction azimuth; direction elevation; distance; gain; spatial range; energy ratio; and position.

[0022] The method may further comprise: receiving an additional audio data stream; and embedding the additional audio data stream in one or the other of the first audio data stream and the second audio data stream.

[0023] According to a third aspect, a device is provided, comprising at least one processor and at least one memory comprising computer program code, the at least one memory and the computer program code being configured to, together with the at least one processor, cause the device to at least: receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream comprises a spatial audio stream to enable immersive audio during communication; determine a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and the second audio data stream comprises the spatial audio stream; process the second audio data stream with at least one parameter based on the determined type; and render the first audio data stream and the processed second audio data stream.

[0024] The second audio data stream may be configured to include at least one other audio data stream, and wherein the at least one other audio data stream may include the determined type, and the at least one other audio data stream may be an embedded-level audio data stream relative to the second audio data stream.

[0025] The at least one further audio data stream may comprise at least one further embedding level, wherein each embedding level may comprise at least one additional audio data stream of the determined type.

[0026] The second audio data stream may be a primary audio data stream.

[0027] Each audio data stream may be further associated with at least one of: a stream identifier configured to uniquely identify the audio data stream; and a stream descriptor configured to describe the type of the audio data stream.

[0028] The type can be one of the following: mono audio signal type; immersive voice and audio service audio signal.

[0029] At least one parameter may be configured to define a room characteristic or scene description.

[0030] The at least one parameter defining the indoor characteristic or scene description may include at least one of: direction; direction azimuth; direction elevation; distance; gain; spatial range; energy ratio; and position.

[0031] The apparatus may be further caused to: receive an additional audio data stream; and embed the additional audio data stream within one or the other of the first audio data stream and the second audio data stream.

[0032] According to a fourth aspect, a device is provided, comprising: a receiving circuit configured to receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; a determining circuit configured to determine a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and the second audio data stream includes the spatial audio stream; a processing circuit configured to process the second audio data stream with at least one parameter according to the determined type; and a rendering circuit configured to render the first audio data stream and the processed second audio data stream.

[0033] According to a fifth aspect, there is provided a computer program [or a computer-readable medium comprising program instructions], the instructions / program instructions being configured to cause an apparatus to at least perform the following operations: receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream comprises a spatial audio stream to enable immersive audio during communication; determine a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and the second audio data stream comprises the spatial audio stream; process the second audio data stream using at least one parameter based on the determined type; and render the first audio data stream and the processed second audio data stream.

[0034] According to a sixth aspect, a non-transitory computer-readable medium comprising program instructions is provided, which are used to cause an apparatus to at least perform the following operations: receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream comprises a spatial audio stream to enable immersive audio during communication; determine a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and the second audio data stream comprises the spatial audio stream; process the second audio data stream with at least one parameter according to the determined type; and render the first audio data stream and the processed second audio data stream.

[0035] According to a seventh aspect, a device is provided, comprising: a component for receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; a component for determining a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and the second audio data stream includes the spatial audio stream; a component for processing the second audio data stream with at least one parameter according to the determined type; and a component for rendering the first audio data stream and the processed second audio data stream.

[0036] According to an eighth aspect, a computer-readable medium comprising program instructions is provided, which are used to cause an apparatus to at least perform the following operations: receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream comprises a spatial audio stream to enable immersive audio during communication; determine a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and the second audio data stream comprises the spatial audio stream; process the second audio data stream with at least one parameter according to the determined type; and render the first audio data stream and the processed second audio data stream.

[0037] An apparatus comprises means for performing the actions of the method as described above.

[0038] An apparatus is configured to perform the actions of the method described above.

[0039] A computer program includes program instructions for causing a computer to execute the method described above.

[0040] A computer program product stored on a medium may cause an apparatus to perform the method described herein.

[0041] An electronic device may include an apparatus as described herein.

[0042] A chipset may include the apparatus as described herein.

[0043] The embodiments of the present application are intended to solve the problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings, in which:

[0045] Figure 1 schematically illustrates an example conferencing system suitable for employing some embodiments;

[0046] Figures 2a to 2d schematically illustrate systems suitable for implementing the apparatus of some embodiments;

[0047] Figure 3 schematically illustrates a bitstream-to-object-to-bitstream converter according to some embodiments;

[0048] Figure 4 Schematically illustrating how Figure 3 A flowchart of the operation of the bitstream-object-bitstream converter shown in;

[0049] Figures 5a-5d illustrate example object formats according to some embodiments;

[0050] Figure 6 illustrates example object nesting according to some embodiments;

[0051] Figure 7 illustrates example operating scenarios according to some embodiments;

[0052] Figures 8a-8c illustrate example object grouping according to some embodiments; and

[0053] Figure 9 An example device suitable for implementing the illustrated means is shown. DETAILED DESCRIPTION

[0054] The following describes in more detail suitable means and possible mechanisms for embedding spatial streams into object streams and transmitting the spatial streams intact as objects to receiving participants. Object metadata is updated based on the spatial scene. In other words, the object stream type itself is another audio stream with corresponding object metadata generated by the processing unit. This operation can be performed by a suitable device that receives more than one input format (e.g., a mobile device, user equipment (UE)) or, for example, a conference bridge (e.g., a multipoint control unit (MCU)).

[0055] The present invention relates to an immersive audio codec capable of supporting multiple input audio formats, immersive audio scene representations, and services where the input encoded audio can, for example, be mixed, re-encoded and / or forwarded to a listener.

[0056] The IVAS codec discussed above is an extension of the 3GPP EVS codec and is intended for new real-time immersive voice and audio services on 4G / 5G. Such immersive services include, for example, immersive voice and audio for virtual reality (VR) and augmented reality (AR). The multifunctional audio codec is expected to handle the encoding, decoding, and rendering of voice, music, and general audio. The audio codec is expected to support channel-based audio and scene-based audio input, including spatial information about sound fields and sound sources. The audio codec is also expected to operate with low latency to enable conversational services and support high error robustness under various transmission conditions.

[0057] The IVAS encoder is configured to receive inputs in a variety of supported formats (as well as some permitted combinations of formats). Similarly, the decoder is expected to output audio in a variety of supported formats. A pass-through mode has been proposed in which audio is provided in its original format after transmission (encoding / decoding).

[0058] A method for describing an acceptable format for object-based audio implemented as an IVAS codec has been proposed, which is configured to process spatial metadata combined with a suitable (mono) audio signal and which can be rendered to a user. Metadata parameters can, for example, be captured from a real environment with the help of any visual or auditory tracking method or any other sensory channel / modality. In some embodiments, radio-based technology can be used to generate metadata, for example, Bluetooth, Wifi or GPS locator technology can be used to obtain object coordinates. In some embodiments, sensors such as magnetometers, accelerometers and / or gyroscopes can be used to receive directional data. In addition, other sensors such as proximity sensors can also be used to generate scene-related metadata from a real environment.

[0059] Alternatively, metadata can be created artificially according to a defined virtual scene, for example, by a teleconference bridge or by a user device (e.g., a smart phone). For example, a user can set or indicate some desired acoustic features via a suitable UI.

[0060] In some embodiments, object-based audio spatial metadata may be defined as one or more objects, where each object may be defined by parameters such as azimuth, elevation, distance, gain, and spatial extent.

[0061] Furthermore, Metadata Assisted Spatial Audio (MASA) is a parameterized spatial audio format and representation. At a high level, it can be thought of as a representation consisting of "N channels + spatial metadata". It is a scene-based audio format that is particularly well suited for spatial audio capture on real devices such as smartphones, where spherical arrays for FOA / HOA capture are impractical or inconvenient. The idea is to describe the sound scene in terms of the direction of the sound sources that vary in time and frequency. If no directional sound source is detected, the audio is described as diffuse. In MASA (as currently proposed for IVAS), there can be one or two directions for each time-frequency (TF) tile. The spatial metadata is described relative to the directions and may include, for example, spatial metadata for each direction and common spatial metadata that is independent of the directions.

[0062] For example, direction-dependent spatial metadata may include parameters such as direction index, direct-to-total energy ratio, spread coherence, and distance. Direction-independent spatial metadata may include parameters such as diffuse-to-total energy ratio, surround coherence, and remainder-to-total energy ratio.

[0063] An example use case for IVAS is AR / VR teleconferencing. In an AR / VR teleconferencing, each participant can have his / her own objects that can be translated arbitrarily in 3D space. In a teleconferencing scenario, a conference bridge can, for example, receive several IVAS streams from multiple participants. These streams are then combined into a common stream, for example using an object for at least each active participant. Alternatively, a pre-rendered spatial scene can be created and represented, for example, in MASA or FOA / HOA audio format. If objects are used, the incoming objects or other mono streams (e.g., EVS streams) can be copied directly as object streams to the outgoing common conference stream by attaching appropriate metadata representations to the waveforms. This may or may not include re-encoding of the audio waveforms. However, if the participants are sending spatial audio streams such as MASA or HOA, the conference bridge must decode all incoming IVAS streams and reduce these streams to mono before sending them downstream as (mono) audio objects.

[0064] Another use case is where a user is capturing a scene (e.g., making a live podcast video) using a mobile device that enables spatial audio capture on a fixed stand. In addition, a headset or some other form of close-range microphone can be used to enhance the voice recording. The close-range capture device is also capable of capturing spatial audio from a headset, for example, using binaural capture, or capturing MASA from a lavalier microphone that supports spatial audio. Furthermore, the close-range captured voice can be added to the IVAS spatial audio stream captured by the device as an object stream. Object position and distance can be conveniently captured, for example, using a suitable location beacon attached to the close-range capture device. When only mono objects are allowed in IVAS, the device must downmix the spatial stream from the close-range capture into mono and then embed it in the IVAS stream. The embodiments described herein attempt to avoid or minimize the added delay and complexity, and further attempt to improve the maximum achievable quality.

[0065] Thus, some embodiments as described herein increase flexibility in audio source mixing and forwarding of various IVAS audio inputs, for example, in AR / VR teleconferencing and other immersive use cases.

[0066] Additionally, in some embodiments, latency and complexity are significantly reduced, thereby avoiding the need to generate a down-converted spatial stream at the AR / VR conference bridge or capture device. Additionally, there is no loss of original input characteristics and / or quality loss for the converted audio format.

[0067] In some embodiments, the decoder is configured to have an interface output format (a so-called pass-through mode to enable an external renderer with more capabilities than a normal integrated renderer) to act as an output mode.

[0068] about Figure 1 , shows an example system in which some embodiments may be implemented. System 200 shows a conference scenario where some participants are sending mono and some spatial streams, and some participants have mono, some spatial, and even some 6DoF rendering and playback capabilities. For example, Figure 1 As shown in FIG, in room A 209, user 202 is using mono capture and fixed spatial playback, in room B 213, user 206 is using spatial capture and 6DoF (degrees of freedom) playback, in room C 211, user 204 is using mono capture and playback, and in room D 215, users 208 and 210 are using spatial capture and mono object capture and spatial playback, but without head tracking. The conferencing service 201 connects all users.

[0069] like Figure 1The system shown in FIG has user-operated devices with different capabilities, and the embodiments described herein attempt to optimize the user experience without requiring conference service 201 to decode, mix, and encode various inputs separately. In the embodiments described herein, any decision is related to the level of immersion. For example, in some embodiments, the apparatus can be implemented at the receiving UE.

[0070] Thus, in some embodiments, an (IVAS) object stream can be configured to include another "objectified" (IVAS) data stream. Furthermore, the object metadata is configured to contain information whether the object is an audio representation based on a (mono) object (e.g., an EVS stream with spatial metadata) or a complete IVAS spatial stream (e.g., MASA or stereo or even containing IVAS objects) that can be given object-like metadata (e.g., position metadata). In such embodiments, any "objectified" (IVAS) data stream can contain another (IVAS) object. These (IVAS) objects can be moved around to become part of any other (IVAS) object or "main" (IVAS) data stream. Furthermore, any one object metadata is updated so that it remains meaningful for the entire newly formed IVAS stream. Furthermore, in some embodiments, the remaining object metadata fields are updated based on the spatial scene description.

[0071] In such embodiments, higher quality and lower latency are expected for conference bridge use cases where the input audio stream is spatially captured / created. Furthermore, some embodiments can be implemented in use cases where there is primary spatial audio captured by, for example, a mobile phone (UE) and additional spatial audio objects are captured by wireless microphones to similarly enhance voice capture benefits and allow (IVAS) encoding to be performed on a new class of devices (wireless microphones) without the need to decode the audio at the UE to allow further encoding. Instead, the streams can simply be embedded as is.

[0072] Before further discussing embodiments, we first discuss a system for obtaining and rendering spatial audio signals that may be used in some embodiments.

[0073] 2, it is shown that Figure 1 Example apparatus used within the system shown in and suitable for implementing some embodiments as described herein.

[0074] Figure 2AAs an example, an apparatus suitable for implementing some embodiments with respect to a user in room A is shown. In this example, the apparatus includes a single microphone 101 configured to generate a mono audio signal that is passed to an encoder 103. The apparatus also includes an encoder 103 configured to receive the mono audio signal and encode the mono audio signal before sending it to a suitable conferencing network.

[0075] Figure 2A Also shown is a decoder / renderer 105 configured to receive encoded spatial / mono audio signals, which are decoded and rendered into suitable audio signal outputs, which are passed to a plurality of speakers 107 for outputting the spatial audio signals to a user.

[0076] also, Figure 2B An example apparatus suitable for implementing some embodiments with respect to a user in room B is shown. In this example, the apparatus includes a multi-microphone 111 audio input configured to generate a plurality of audio signals that can be used to generate a spatial audio signal that is passed to an encoder 113. The apparatus also includes an encoder 113 configured to receive the spatial audio signal and encode the spatial audio signal before transmitting it to a suitable conferencing network.

[0077] Figure 2B Also shown is a decoder / renderer 115 configured to receive encoded spatial / mono audio signals, which are decoded and rendered into suitable audio signal outputs, which are passed to headphones equipped with a head tracker / localizer 117 to output these spatial audio signals to the user and pass the user position to the decoder / renderer 115 to control the rendering.

[0078] Figure 2C An example apparatus suitable for implementing some embodiments with respect to a user in room C is shown. In this example, the apparatus includes a mono microphone 121 audio input configured to generate a mono audio signal, which can be used to generate a mono audio signal that is passed to encoder 113. The apparatus also includes encoder 123, which is configured to receive the mono audio signal and encode the mono audio signal into a spatial audio signal before sending it to a suitable conferencing network.

[0079] Figure 2C Also shown is a decoder / renderer 125 configured to receive encoded spatial / mono audio signals, which are decoded and rendered into suitable audio signal outputs, which are passed to a mono speaker 127 for outputting these audio signals to a user.

[0080] also, Figure 2D An example apparatus suitable for implementing some embodiments with respect to a user in room D is shown. In this example, the apparatus includes a multi-microphone 131 audio input configured to generate multiple audio signals, and an external microphone (e.g., a mono microphone or multiple microphones) that can be used to generate a spatial audio signal and an external mono / spatial audio signal that is passed to an encoder 113. The apparatus also includes an encoder 133 that is configured to receive the spatial / mono audio signal and encode the spatial / mono audio signal before sending it to a suitable conferencing network.

[0081] Figure 2D Also shown is a decoder / renderer 135 configured to receive encoded spatial / mono audio signals, which are decoded and rendered into suitable audio signal outputs, which are passed to headphones 137 for outputting these spatial audio signals to the user.

[0082] about Figure 3 , shows a high-level view of an example (IVAS) encoder 103 / 113 / 123 / 133 including various inputs (as non-exclusive examples, which may be expected for a codec).

[0083] In some embodiments, encoder 103 / 113 / 123 / 133 includes an audio (IVAS) input 301. Audio input 301 is configured to receive one or more spatial data (IVAS) streams from multiple sources, either local or remote. These sources can be local (e.g., more than one spatial capture device configured in a known space at the encoder's location) and / or multiple remote participants sending spatial IVAS streams. Audio input 301 is configured to pass the audio data streams to an object header creator 303 and an (IVAS) decoder 311, which is part of an IVAS data stream processor 313.

[0084] In some embodiments, the encoder 103 / 113 / 123 / 133 includes a scene controller 305 configured to control processing of the received audio input 301 .

[0085] For example, in some embodiments, the encoder 103 / 113 / 123 / 133 includes an object header creator 303. The object header creator 303, controlled by the scene controller 305, is configured to insert each data stream as an object into the "main" data stream. In some embodiments, the object header creator 305 can also be configured to add missing object parameters, such as distance and direction, based on the real space configuration or the virtual defined scene.

[0086] In some embodiments, the object header creator 303 is configured to determine whether any inserted data stream contains objects, and optionally move those audio objects to become part of the "main" IVAS stream directly and update their metadata, or move the objects below any other IVAS objects. Additionally, the object header creator 303 is configured to update the object metadata so that it is correct for the entire spatial configuration.

[0087] In some embodiments, the encoder 103 / 113 / 123 / 133 includes an IVAS data stream processor 313. The IVAS data stream processor 313 may include an (IVAS) decoder 311. The (IVAS) decoder 311 is configured to receive one or more sets of spatial audio data streams, decode the spatial audio signals, and pass them to the audio scene renderer 231.

[0088] The IVAS data stream processor 313 may include an audio scene renderer 231 configured to receive an audio signal and generate an audio scene rendering based on the decoded (IVAS) spatial audio signal. The audio scene rendering may constitute a downmix of various inputs from the (IVAS) decoder 311, for example. In turn, the rendered audio scene audio signal may be passed to the encoder 315.

[0089] The IVAS data stream processor 313 may include an encoder 315 that receives and encodes the rendered spatial audio signal. In other words, the IVAS data stream processor 313 is configured to decode all or at least some of the input data streams and generate a common spatial scene using, for example, IVAS MASA, IVAS HOA / FOA, or IVAS mono objects.

[0090] In some embodiments where there are multiple embedded objects, these embedded objects can then be sent to those recipients that have high-performance rendering available. The remaining recipients receive only the pre-rendered spatial scene. Alternatively, a combination of at least one "IVAS stream object" and a pre-rendered "spatial scene IVAS stream object" can be used to reduce the bit rate.

[0091] Additionally, the encoder comprises an audio object multiplexer 309 configured to combine the objects and output a combined object data stream.

[0092] Figure 4 The operation of the encoder is further illustrated by the flowchart in .

[0093] exist Figure 4 At step 401 in the embodiment, an audio (IVAS) data stream is received.

[0094] In addition, Figure 4 At step 411 in the process, the spatial scene configuration and control are determined.

[0095] like Figure 4 As shown in step 403, an object header for the audio data stream is created based on the determined spatial scene configuration and control and the input audio data stream.

[0096] In addition, if Figure 4 As shown in step 404, optionally, the data stream is decoded based on the determined spatial scene configuration and control and the input audio data stream.

[0097] Furthermore, if Figure 4 As shown in step 406, the decoded data stream may be rendered.

[0098] Furthermore, if Figure 4 As shown in step 408, the rendered audio scene is encoded using a suitable (IVAS) encoder.

[0099] Furthermore, if Figure 4 As shown in step 409, these data streams can be multiplexed and output.

[0100] IVAS object stream metadata can use any suitable acoustic / spatial metadata. The following table provides an example.

[0101]

[0102] However, in some embodiments, other position information such as xyz or Cartesian coordinates may be used instead of azimuth-elevation-distance. For example, the following table may provide another configuration:

[0103] Field Bit describe Location x The xyz position of the audio object in 3D space

[0104] However, some minimal stream description metadata is still needed to signal (IVAS) the object data stream configuration information. For example, the following format can be used to signal this information.

[0105]

[0106] In this embodiment, the "stream ID" parameter is used to uniquely identify each IVAS object stream in the current session. Therefore, each original and mixed audio component (input stream) can be transmitted by signaling. For example, signaling allows identification of components in the system or on the user interface. The "stream type" parameter defines the meaning of each "audio object". In some embodiments, audio objects are therefore not just object-based audio inputs. Instead, object data streams can be object-based audio (inputs), which can also be any IVAS scene. For example, as shown in Figure 5, three types of objects are shown.

[0107] For example, in Figure 5A In FIG, a simple conventional (mono) audio object 501 is shown. The audio object 501 is defined in terms of a PCM audio signal portion 505 and an acoustic (spatial) metadata portion 503. It will be appreciated that additional metadata may be present.

[0108] about Figure 5B , showing the Figure 5A An encoded representation 507 of the same audio object shown in FIG.

[0109] Figure 5C Shown with Figure 5A and Figure 5B , but processed according to some embodiments as discussed herein. The processed audio objects are depicted as object data streams 509, which are defined by a "stream type = 0" parameter 513. In other words, the object data stream 509 includes a data stream identifier that identifies it as an object-based audio IVAS object stream. In addition, the object data stream 509 includes an object audio bitstream portion 515 (an encoded representation of the audio object) and a stream identifier 511 that uniquely identifies the object data stream.

[0110] Figure 5D Another (IVAS) object data stream 517 is shown. Another object data stream 517 includes an identifier portion 521 with "stream type = 1". In some embodiments, stream type = 0 corresponds to a "simple" object type, such as a mono signal. In addition, in some embodiments, stream type = 1 corresponds to a potentially "complex" stream. For example, in this example, stream type = 1 corresponds to a complete IVAS stream, in this case, it contains a MASA spatial stream. Since IVAS can contain one or more object streams, this allows for nested objects. If stream type = 0, then it is known that there are no other objects and the stream is of a simple type (effectively a mono object).

[0111] Another object data stream 517 may also include an explicit stream description portion 523, or the stream content may be determined by starting to decode the object stream. In this case, it is explicitly described as a MASA-based scenario (e.g., "stream description = MASA").

[0112] In addition, another object data stream 517 includes a MASA format bitstream portion 525 (an encoded representation of the audio object) and a stream identifier 519 "stream ID=000002" that uniquely identifies the object data stream.

[0113] The first advantage of the method discussed herein is that IVAS input can often be conveniently forwarded without the need for decoding / encoding operations. For example, in the presence of a mixer device, a conference call bridge (e.g., an AR / VR conference server), or other entity for combining and / or forwarding audio input in an IVAS end-to-end service, no decoding / encoding operations are required. Thus, by reallocating the received (encoded) input as an IVAS object stream, the complexity and latency of the operation are reduced. For example, if the playback capabilities of the recipient are unknown, the server can optimize the complexity by simply providing the received scene as is. Any IVAS stream can be decoded and rendered as mono to support even the simplest IVAS device. Skipping any decoding / encoding operations at an intermediate point (e.g., a conference server) also reduces the end-to-end latency for that audio component. Thus, the user experience is improved.

[0114] In addition, the embodiment is configured to make there be only shallow embedded " objectized " IVAS stream. In other words, if there is an object stream that also contains objects (and therefore can include multiple object levels), then avoid deep data structure, and therefore reduce the complexity of decoder. Therefore, in some embodiments, the proposed embedding allows one IVAS object to include another IVAS object. In other words, although the IVAS object can be two or more levels deep (two or more levelsdeep), in some embodiments, any "deep" object can be moved to a "higher level" object closer to the "main" IVAS stream, and its metadata can be updated so that its representation remains meaningful for the newly formed scene. In some embodiments, the IVAS object can be moved to become a part of another IVAS object. Therefore, the object is moved "deeper". This can allow, for example, audio objects (for example, monophonic objects) to be encoded or decoded together, so as to save complexity or bit rate. If the formats of the same type are at different levels in the structure, then they usually need to be encoded / decoded at different times or using different instances. This can introduce additional complexity.

[0115] In addition, the embodiment as discussed herein may have a second advantage, i.e., IVAS object streams may be conveniently nested, for example, for content distribution purposes. In such an embodiment, more complex scenes may be processed as a single (mono) audio object. Figure 6 Provides an example of nested packetization. This can be used, for example, to distribute decoding complexity. This is useful, for example, for edge cloud services.

[0116] So, for example, Figure 6 6. The entire scene object data stream 601 is shown. The entire scene object data stream 601 includes a plurality of object data streams 602, 604, 606, and 608. For example, the first object data stream 602 includes a stream ID 621 (stream ID = 000001) that uniquely identifies the object data stream, a stream type identifier 623 (stream type = 0), and a data portion 625. The second object data stream 604 includes a stream ID 631 (stream ID = 000006) that uniquely identifies the object data stream, a stream type identifier 633 (stream type = 1), and a data portion 635. The third object data stream 606 includes a stream ID 641 (stream ID = 000007) that uniquely identifies the object data stream, a stream type identifier 643 (stream type = 1), and a data portion 645. The fourth object data stream 608 includes a stream ID 651 (stream ID=000008) that uniquely identifies the object data stream, a stream type identifier 653 (stream type=0), and a data portion 655 .

[0117] In addition, if Figure 6 As shown in FIG, the second object data stream 604 also includes nested object data streams 612 and 614. These nested object data streams may be, for example, object data streams associated with sub-portions of the entire scene. The fifth object data stream 612 includes a stream ID 661 (stream ID = 000004) that uniquely identifies the object data stream, a stream type identifier 663 (stream type = 0), and a data portion 665. The sixth object data stream 614 includes a stream ID 671 (stream ID = 000005) that uniquely identifies the object data stream, a stream type identifier 673 (stream type = 1), and a data portion 675.

[0118] In addition, the nested sixth object data stream 614 also includes further nested object data streams 622 and 624. These further nested object data streams may be, for example, object data streams associated with sub-sub-sections of sub-sections of the entire scene. The seventh object data stream 622 includes a stream ID 681 (stream ID = 000002) that uniquely identifies the object data stream, a stream type identifier 683 (stream type = 1), and a data portion 685. The eighth object data stream 624 includes a stream ID 691 (stream ID = 000003) that uniquely identifies the object data stream, a stream type identifier 693 (stream type = 1), and a data portion 695.

[0119] Another advantage of implementing some embodiments is that any IVAS input or IVAS scene that does already include spatial parameters (e.g., positional characteristics) can determine such characteristics. For example, this can be achieved by adding acoustic spatial metadata (e.g., one of the parameters in the previous table) to the IVAS object stream ("stream type = 1"). This enables, for example, enhanced experiences in AR / VR teleconferencing use cases.

[0120] For example, Figure 7 A capture scenario 701 is shown where there is a UE or similar capture device 707 enabling spatial capture at a first (UE) location, and a second capture device 705 enabling a second spatial capture (or object capture) at a second (user) location.

[0121] exist Figure 1 The conventional method shown in the upper right portion of the shows the audio object rendering 713 position and the first spatial capture scene 711. Thus, while the user can use a multi-microphone UE to capture the spatial scene (e.g., in MASA format), and the user can use a close-range microphone or, for example, a second UE that can be connected to a "master" device to capture audio objects. These two inputs will be combined and provided to the IVAS encoder. In terms of the listening experience, one can then listen to a combined rendering of spatial audio (e.g., background audio) and audio objects (e.g., user voice).

[0122] By implementing the embodiments described herein, a listener can switch 730 between a second spatially captured audio object rendering 723 and a first option for the first spatially captured scene 721, or between a first spatially captured audio object rendering 733 and a second option for the second spatially captured scene 731. Thus, the IVAS codec can introduce a second spatial audio representation as an IVAS object stream. Thus, while a user is using their UE to capture a spatial audio scene, a wireless multi-microphone device, or indeed a second UE connected to the "primary" UE, can capture a complete spatial representation of the sound scene at a second location. This sound scene can now be encoded by the second device into an IVAS bitstream and provided to a second UE, which can "act as a conference bridge," acquiring the IVAS bitstream and embedding it as an IVAS object stream. This, in turn, will deliver two spatial audio scenes to the listener. For example, the user can switch between them so that a mono downmix of each scene is provided as an audio object rendering for the other scene being rendered to the user.

[0123] Although Figure 6 An example of object stream nesting is shown, but it should be understood that this is not the only mechanism of IVAS streaming / packetization as enabled by the present invention. Figure 8 shows two examples of IVAS stream packetization according to some embodiments.

[0124] In some embodiments, a lookup table specifying the contents of a packet may be used. This lookup table may be defined as a "payload header," and it may be, for example, an RTP payload header. This may include, for example, the sizes of various blocks, etc. Following the header is the payload.

[0125] For example, as shown in FIG8 , a data stream may include various IVAS object streams and IVAS content. Thus, the entire scene object stream 801 includes a payload header 811 or a lookup table that may specify the contents of the packet. For example, Figure 8A As shown in FIG, a first object data stream 813 and a second object data stream 819 and payloads such as a first payload 815 (MASA and objects) and a second payload 817 (5.1 channel audio data) are specified.

[0126] In such Figure 8C In some embodiments shown in , the data stream may include only IVAS object streams. Thus, the entire scene object stream 831 includes a payload header or lookup table that can specify the packet content including object data stream 833, which in turn may include nested object data streams 835, which in turn may include further nested object data streams.

[0127] Figure 8BA "hybrid" embodiment is provided with payload and nested object data streams 813 throughout the scene.

[0128] There is a nesting cost associated in the generation of the additional "payload header" information and its parsing.

[0129] Regarding the decoders / renderers 105, 115, 125, 135. The decoders / renderers 105, 115, 125, 135 are configured to receive various (IVAS) object data streams, and to decode and render these data streams in parallel.

[0130] In some embodiments, processing of the nested audio object data streams may be performed separately for each sub-scene level and then combined at a higher level.

[0131] For example, about Figure 6 . Here, decoding can start from "stream ID = 000002" and "stream ID = 000003". Therefore, we have decoded "stream ID = 000005" (because it is a container for a subscene). In turn, the decoder can be configured to decode the next "stream ID = 000004". After this, the other streams are decoded. This approach can have advantages in terms of memory consumption, for example, where some memory can be freed between subscene levels and the total memory footprint is therefore not defined by all combined streams.

[0132] In such embodiments, rendering can be performed on a sub-scene level and summed in the rendering domain, or combined rendering can be performed at the end of decoding.

[0133] In some embodiments, the decoder is configured to start a separate decoder instance for each sub-scene.Thus, for each "stream type=1", a separate IVAS decoder instance is initialized.

[0134] about Figure 9 , illustrates an example electronic device that can be used as an analysis or synthesis device. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 1400 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.

[0135] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 may be configured to execute various program codes, such as the methods described herein.

[0136] In some embodiments, the device 1400 includes a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage component. In some embodiments, the memory 1411 includes a program code portion for storing program codes that can be implemented on the processor 1407. In addition, in some embodiments, the memory 1411 can also include a storage data portion for storing data (e.g., data that has been processed or is to be processed according to the embodiments described herein). Whenever necessary, the implemented program code stored in the program code portion and the data stored in the storage data portion can be retrieved by the processor 1407 via the memory-processor coupling.

[0137] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, user interface 1405 can be coupled to processor 1407. In some embodiments, processor 1407 can control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 can enable a user to enter commands to device 1400, for example, via a keypad. In some embodiments, user interface 1405 can enable a user to obtain information from device 1400. For example, user interface 1405 can include a display configured to display information from device 1400 to the user. In some embodiments, user interface 1405 can include a touch screen or touch interface that enables information to be entered into device 1400 and to display information to a user of device 1400. In some embodiments, user interface 1405 can be a user interface for communicating with a location determiner as described herein.

[0138] In some embodiments, device 1400 includes input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. In such embodiments, the transceiver can be coupled to processor 1407 and configured to communicate with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver components can be configured to communicate with other electronic devices or devices via a wired or wired coupling.

[0139] The transceiver can communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.x, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication path (IRDA).

[0140] The transceiver input / output port 1409 may be configured to receive signals and, in some embodiments, determine parameters as described herein using the processor 1407 executing appropriate code. Additionally, the device may generate appropriate downmix signals and parameter outputs for transmission to a synthesis device.

[0141] In some embodiments, the device 1400 can be used as at least part of a synthesis device. Thus, the input / output port 1409 can be configured to receive the downmix signal and, in some embodiments, receive parameters determined at a capture device or processing device as described herein, and generate a suitable audio signal format output using the processor 1407 executing appropriate code. The input / output port 1409 can be coupled to any suitable audio output, such as a multi-channel speaker system and / or headphones (which can be head tracking or non-tracking headphones), etc.

[0142] In general, various embodiments of the present invention may be implemented using hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented using hardware, while other aspects may be implemented using firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flow charts, or using some other graphical representation, it is understood that the blocks, devices, systems, techniques, or methods described herein may be implemented using hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof, as non-limiting examples.

[0143] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of the logic flow as in the accompanying drawings may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on a physical medium such as a memory chip or a memory block implemented within a processor, on a magnetic medium such as a hard disk or floppy disk, and on an optical medium such as a DVD and its data variants, a CD.

[0144] The memory may be of any type suitable for the local technical environment and may be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include, by way of non-limiting example, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit based on a multi-core processor architecture, and a processor.

[0145] Embodiments of the present invention may be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0146] Programs, such as those offered by Synopsys, Inc. of Mountain View, Calif., and Cadence Design, Inc. of San Jose, Calif., use well-established design rules and a library of pre-stored design blocks to automatically route conductors and position components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the resulting design, in a standardized electronic format (e.g., Opus, GDSII, etc.), can be transferred to a semiconductor fabrication facility, or "fab," for fabrication.

[0147] The foregoing description has provided by way of exemplary and non-limiting examples a complete and informative description of the exemplary embodiments of the present invention. However, various modifications and adaptations will become apparent to those skilled in the relevant arts in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such and similar modifications of the teachings of this invention will still fall within the scope of the invention as defined by the appended claims.

Claims

1. A device comprising: at least one processor; as well as at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: receiving at least at least one first audio channel and at least one second audio channel, wherein at least one of the at least one first audio channel or the at least one second audio channel comprises spatial audio configured to enable immersive audio communication; determining a format of at least one of the at least one first audio channel or the at least one second audio channel to identify which of the received at least one first audio channel and at least one second audio channel includes the spatial audio; processing the identified at least one audio channel using at least one parameter according to the determined format; and The processed at least one audio channel and the other of the at least one first audio channel or the at least one second audio channel are rendered.

2. The device according to claim 1, wherein The identified at least one audio channel comprising the spatial audio is aided with spatial metadata.

3. The device according to claim 1, wherein The at least one second audio channel is configured to include at least one other audio channel, and wherein the at least one other audio channel includes the determined format and is an embedded-level audio channel relative to the at least one second audio channel.

4. The device according to claim 3, wherein The at least one further audio channel comprises at least one further embedding level, wherein a respective one of the at least one further embedding level comprises at least one additional audio channel having the determined format.

5. The device according to claim 1, wherein The at least one second audio channel comprises a primary audio channel.

6. The device according to claim 1, wherein The at least one first audio channel and the at least one second audio channel are each further associated with at least one of: a channel identifier configured to uniquely identify an audio channel; or A channel descriptor configured to describe the format of the audio channel.

7. The device according to claim 1, wherein The format of at least one of the at least one first audio channel or the at least one second audio channel is one of: Mono audio signal format; or Immersive voice and audio services audio signals.

8. The device according to claim 1, wherein The at least one parameter is configured to define a room characteristic or a scene description.

9. The device according to claim 8, wherein The at least one parameter defining the indoor characteristic or scene description includes at least one of the following: direction; Direction azimuth; Direction elevation angle; distance; gain; spatial extent; Energy ratio; or Location.

10. The device according to claim 8, wherein The at least one parameter is generated based on the spatial audio capture.

11. The device according to claim 1, wherein The at least one memory stores instructions that, when executed by the at least one processor, cause the apparatus to at least: Receive additional audio channels; as well as The additional audio channel is embedded within one or the other of the at least one first audio channel or the at least one second audio channel.

12. A method comprising: receiving at least at least one first audio channel and at least one second audio channel, wherein at least one of the at least one first audio channel or the at least one second audio channel comprises spatial audio configured to enable immersive audio communication; determining a format of at least one of the at least one first audio channel or the at least one second audio channel to identify which of the received at least one first audio channel and at least one second audio channel includes the spatial audio; processing the identified at least one audio channel using at least one parameter according to the determined format; and The processed at least one audio channel and the other of the at least one first audio channel or the at least one second audio channel are rendered.

13. The method according to claim 12, wherein: The identified at least one audio channel comprising the spatial audio is aided with spatial metadata.

14. The method according to claim 12, wherein: The at least one second audio channel is configured to include at least one other audio channel, and wherein the at least one other audio channel includes the determined format and is an embedded-level audio channel relative to the at least one second audio channel.

15. The method according to claim 14, wherein The at least one further audio channel comprises at least one further embedding level, wherein a respective one of the at least one further embedding level comprises at least one additional audio channel having the determined format.

16. The method according to claim 12, wherein The at least one second audio channel comprises a primary audio channel.

17. The method according to claim 12, wherein: The at least one first audio channel and the at least one second audio channel are each further associated with at least one of: a channel identifier configured to uniquely identify an audio channel; or A channel descriptor configured to describe the format of the audio channel.

18. The method according to claim 12, wherein: The format of at least one of the at least one first audio channel or the at least one second audio channel is one of: Mono audio signal type; or Immersive voice and audio services audio signals.

19. The method according to claim 12, wherein: The at least one parameter is configured to define a room characteristic or a scene description.

20. The method of claim 12, further comprising: Receive additional audio channels; as well as The additional audio channel is embedded within one or the other of the at least one first audio channel or the at least one second audio channel.