Audio representation and related rendering

By embedding spatial audio streams as object streams and optimizing processing within IVAS codecs, the solution addresses complexity and latency issues in immersive audio systems, ensuring high-quality rendering across diverse audio formats and playback capabilities.

JP2025169455APending Publication Date: 2025-11-12NOKIA TECHNOLOGIES OY
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2025143100
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-02-28
Filing Date
2025-08-29
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Existing immersive audio codecs face challenges in efficiently handling and rendering diverse audio formats with varying levels of immersion, leading to increased complexity, latency, and loss of quality in conferencing and teleconferencing scenarios.

Method used

The proposed solution involves embedding spatial audio streams as object streams, allowing for flexible processing and rendering of immersive audio data using IVAS codecs, which includes determining the type of audio streams, applying parameters like room characteristics and scene descriptions, and embedding additional audio data within existing streams to optimize rendering without decoding and encoding at intermediate points.

Benefits of technology

This approach reduces operational complexity and latency while maintaining high quality, enabling seamless integration of diverse audio inputs and outputs across varying playback capabilities, enhancing user experience in immersive audio environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169455000001_ABST
    Figure 2025169455000001_ABST
Patent Text Reader

Abstract

To provide an apparatus for immersive audio communication.SOLUTION: In a conference system 200, a device implemented in a receiving UE receives at least a first audio data stream and a second audio data stream. At least one of the first audio stream and the second audio stream includes a spatial audio stream to enable immersive audio during communication. The device further determines a type of each audio stream to identify which of the received first and second audio data streams includes the spatial audio stream, processes the second audio data stream with at least one parameter dependent on the determined type, and renders the first audio data stream and the processed second audio data stream.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to apparatus and methods for audio representation and associated rendering in the field of sound technology, and relates to apparatus and methods for audio representation for audio encoders and decoders, but is not limited thereto. [Background technology]

[0002] Immersive audio codecs are implemented to support multiple operating points, ranging from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed for use over communications networks such as 3GPP 4G / 5G networks. Such immersive services include immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR), and mixed reality (MR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general-purpose audio. Furthermore, it is expected to support channel-based and scene-based audio inputs, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services and support high error robustness under various transmission conditions.

[0003] Furthermore, parametric spatial audio processing is a field of audio signal processing in which the spatial aspects of sound are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a typical and effective choice to signal a set of parameters from the microphone array, such as the direction of the sound in a frequency band and the ratio between the directional and non-directional parts of the captured sound in a frequency band. These parameters are known to sufficiently describe the perceptual spatial characteristics of the captured sound at the position of the microphone array. These parameters can be accordingly utilized in spatial sound synthesis, for binaural over headphones, loudspeakers, or other formats such as Ambisonics. Summary of the Invention

[0004] According to a first aspect, there is provided an apparatus comprising: means configured to receive at least a first audio data stream and a second audio data stream, where at least one of the first and second audio streams constitutes a spatial audio stream for enabling immersive audio during communication; determine a type of each of the first audio stream and the second audio stream to identify which of the received first audio data stream and second audio data stream comprises the spatial audio stream; process the second audio data stream using at least one parameter depending on the determined type; and render the first audio data stream and the processed second audio data stream.

[0005] The second audio data stream may be configured to comprise at least one further audio data stream, wherein the at least one further audio data stream may be of a determined type, and wherein the at least one further audio data stream may be an embedded-level audio data stream relative to the second audio data stream.

[0006] The at least one further audio data stream may comprise at least one further embedding level, each embedding level may comprise at least one additional audio data stream having the determined type.

[0007] The second audio data stream may be a master level audio data stream.

[0008] Each audio data stream may be further associated with at least one of a stream identifier configured to uniquely identify the audio data stream and a stream descriptor configured to describe the type of the audio data stream.

[0009] The type can be one of the following: a mono audio signal type, an immersive voice and an audio service audio signal.

[0010] The at least one parameter may be configured to define a room characteristic or a scene description.

[0011] The at least one parameter defining the room characteristic or scene description may include at least one of direction, azimuth, azimuth elevation, distance, gain, spatial extent, energy ratio, and position.

[0012] The means may be further configured to receive an additional audio data stream and to embed the additional audio data stream within one or other of the first audio data stream and the second audio data stream.

[0013] According to a second aspect, there is provided a method for an apparatus, the method comprising: receiving at least a first audio data stream and a second audio data stream, at least one of the first audio stream and the second audio stream comprising a spatial audio stream for enabling immersive audio during communication; determining a type of each of the first audio stream and the second audio stream to identify which of the received first and second audio data streams comprises the spatial audio stream; processing the second audio data stream using at least one parameter depending on the determined type; and rendering the first audio data stream and the processed second audio data stream.

[0014] The second audio data stream may be configured to comprise at least one further audio data stream, wherein the at least one further audio data stream may be of a determined type, and wherein the at least one further audio data stream may be an embedded-level audio data stream relative to the second audio data stream.

[0015] The at least one further audio data stream may comprise at least one further embedding level, each embedding level may comprise at least one additional audio data stream having the determined type.

[0016] The second audio data stream may be a master level audio data stream.

[0017] Each audio data stream may further be associated with at least one of a stream identifier configured to uniquely identify the audio data stream and a stream descriptor configured to describe the type of the audio data stream.

[0018] The type can be one of the following: a mono audio signal type, an immersive voice and an audio service audio signal.

[0019] The at least one parameter may be configured to define a room characteristic or a scene description.

[0020] The at least one parameter defining the room characteristic or scene description may include at least one of a direction, an azimuth, a direction elevation, a distance, a gain, a spatial extent, an energy ratio, and a position. The method may further include receiving an additional audio data stream and embedding the additional audio data stream within one or other of the first audio data stream and the second audio data stream.

[0021] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory containing computer program code configured, using the at least one processor, to cause the apparatus to receive at least a first audio data stream and a second audio data stream comprising at least a spatial audio stream for enabling immersive audio during communication; determine a type of each of the first audio stream and the second audio stream to identify which of the received first audio data stream and second audio data stream comprises a spatial audio stream; process the second audio data stream with at least one parameter depending on the determined type; and render the first audio data stream and the processed second audio data stream.

[0022] The second audio data stream can be configured to comprise at least one further audio data stream, the at least one further audio data stream can be of a determined type, and the at least one further audio data stream can be an embedded-level audio data stream relative to the second audio data stream.

[0023] The at least one further audio data stream may comprise at least one further embedding level, each embedding level may comprise at least one additional audio data stream having the determined type.

[0024] The second audio data stream may be a master level audio data stream.

[0025] Each audio data stream may be further associated with at least one of a stream identifier configured to uniquely identify the audio data stream and a stream descriptor configured to describe the type of the audio data stream.

[0026] The type may be one of a mono audio signal type, an immersive voice, and an audio service audio signal. The at least one parameter may be configured to define a room characteristic or a scene description. The at least one parameter defining the room characteristic or the scene description may include at least one of a direction, an azimuth angle, a direction elevation angle, a distance, a gain, a spatial extent, an energy ratio, and a position.

[0027] The apparatus is further capable of receiving an additional audio data stream and embedding the additional audio data stream within one or other of the first audio data stream and the second audio data stream.

[0028] According to a fourth aspect, there is provided an apparatus for receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first and second audio streams comprises a spatial audio stream enabling immersive audio during communication, the apparatus comprising: a receiving circuit configured to determine a type of each of the first and second audio data streams to identify which of the received first and second audio data streams constitutes the spatial audio stream; a processing circuit configured to process the second audio data stream with at least one parameter depending on the determined type; and a rendering circuit configured to render the first audio data stream and the processed second audio data stream.

[0029] According to a fifth aspect, there is provided a computer program comprising instructions (or a computer-readable medium comprising program instructions) to cause an apparatus to perform at least the steps of: receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio stream and the second audio stream comprises a spatial audio stream for enabling immersive audio during communication; determining a type of each of the first audio stream and the second audio stream to identify which of the received first and second audio data streams comprises a spatial audio stream; processing the second audio data stream using at least one parameter depending on the determined type; and rendering the first audio data stream and the processed second audio data stream.

[0030] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus to at least receive a first audio data stream and a second audio data stream, the first audio data stream and the second audio data stream comprising spatial audio streams that enable immersive audio during communication; determine a type of each of the received first and second audio data streams to identify which of the received first and second audio data streams comprises a spatial audio stream; process the second audio data stream using at least one parameter that depends on the determined type; and render the first audio data stream and the processed second audio data stream.

[0031] According to a seventh aspect, there is provided an apparatus comprising: means for receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio stream and the second audio stream comprises a spatial audio stream for enabling immersive audio during communication; means for determining a type of each of the first audio stream and the second audio stream to identify which of the received first and second audio data streams comprises the second audio data stream; means for processing the second audio data stream with at least one parameter depending on the determined type; and means for rendering the first audio data stream and the processed second audio data stream.

[0032] According to an eighth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus to at least receive a first audio stream and a second audio stream, wherein at least one of the first audio stream and the second audio stream comprises a spatial audio stream for enabling immersive audio during communication, determine a type of each of the first audio stream and the second audio stream to identify which of the received first and second audio data streams comprises the spatial audio stream, process the second audio data stream using at least one parameter depending on the determined type, and render the first audio data stream and the processed second audio data stream.

[0033] The apparatus includes means for performing the operations described above.

[0034] The apparatus is configured to perform the operations of the method as described above.

[0035] The computer program comprises program instructions for causing a computer to carry out the above-described method.

[0036] A computer program product stored on a medium can cause an apparatus to perform the methods described herein.

[0037] The electronic device may comprise an apparatus as described herein.

[0038] A chipset may comprise the apparatus described herein.Embodiments of the present application aim to address challenges associated with the state of the art. [Brief explanation of the drawings]

[0039] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1] FIG. 1 illustrates a schematic diagram of an exemplary conferencing system suitable for employing some embodiments. [Figure 2] 2a-2d show a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Figure 3] FIG. 3 illustrates a schematic diagram of a bitstream-to-object-to-bitstream converter according to some embodiments. [Figure 4] FIG. 4 shows a schematic flow diagram of the operation of a bitstream-to-object-to-bitstream converter such as that shown in FIG. 3, according to some embodiments. [Figure 5] 5a-5d illustrate exemplary object formats according to some embodiments. [Figure 6] FIG. 6 illustrates an example nesting of objects according to some embodiments. [Figure 7] FIG. 7 illustrates an exemplary operating scenario according to some embodiments. [Figure 8] 8a-8c illustrate exemplary object packetization according to some embodiments. [Figure 9] FIG. 9 shows an exemplary device suitable for implementing the apparatus shown. DETAILED DESCRIPTION OF THE INVENTION

[0040] The following describes in more detail a suitable apparatus and possible mechanisms for embedding the spatial stream as an object stream and transmitting the spatial stream as an object to the receiving participants. The object metadata is updated based on the spatial scene. In other words, the object stream type is itself another audio stream with respective object metadata generated by the processing elements. This operation can be performed by an appropriate device (e.g., a mobile user equipment (UE)) that receives two or more input formats, or, for example, a conference bridge (e.g., a multipoint control unit - MCU).

[0041] The present invention relates to an immersive audio codec capable of supporting many input audio formats, immersive audio scene representations, and services where the incoming encoded audio can be, for example, mixed, re-encoded, and / or forwarded to a listener.

[0042] The IVAS codec described above is an extension of the 3GPP EVS codec and is intended for new real-time immersive voice and audio services beyond 4G / 5G. Such immersive services include, for example, immersive voice and audio for virtual reality (VR) and augmented reality (AR). A versatile audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general-purpose audio. It is expected to support channel-based audio and scene-based audio inputs, including spatial information about the sound field and sound sources. It is also expected to operate with low latency to enable conversational services and support high error robustness under various transmission conditions.

[0043] An IVAS encoder is configured to be able to receive input in any supported format (and in some allowed combinations of formats). Similarly, a decoder is expected to be able to output audio in some supported format. A pass-through mode is proposed that can provide audio in its original format after transmission (encoding / decoding).

[0044] A method for describing object-based audio has been proposed, which is implemented as an acceptable format for an IVAS codec configured to process spatial metadata combined with an appropriate (mono) audio signal and render it to a user. Metadata parameters can be captured from the real-world environment, for example, with the aid of any visual or auditory tracking method or any other modality. In some embodiments, wireless-based technologies can be used to generate the metadata; for example, object coordinates can be obtained using Bluetooth, Wi-Fi, or GPS locator technologies. Orientation data can be received in some embodiments using sensors such as magnetometers, accelerometers, and / or gyrometers. Other sensors, such as proximity sensors, can also be used to generate scene-related metadata from the real-world environment.

[0045] Alternatively, it may be artificially created, for example, by a video conference bridge or by a user device (e.g., a smartphone) according to a metadata-defined virtual scene. For example, a user may set or indicate some desired acoustic features via an appropriate UI.

[0046] In some embodiments, object-based audio spatial metadata can be defined as one or more objects, each of which can be defined by parameters such as azimuth, elevation, distance, gain, and spatial extent.

[0047] Furthermore, Metadata-Assisted Spatial Audio (MASA) is a parametric spatial audio format and representation. At a high level, it can be considered a representation consisting of "N channels + spatial metadata." It is a scene-based audio format particularly suited for spatial audio capture on practical devices such as smartphones, where spherical arrays for FOA / HOA capture are neither realistic nor convenient. The idea is to describe a sound scene in terms of source directions that vary in time and frequency. If no directional sound sources are detected, the audio is described as diffuse. In MASA (currently proposed for IVAS), there can be one or two directions for each time-frequency (TF) tile. Spatial metadata is described in terms of directions and can include, for example, spatial metadata for each direction and common spatial metadata that is direction independent.

[0048] For example, spatial metadata for direction may comprise parameters such as direction index, direct-to-total energy ratio, diffuse coherence, and distance. Direction-independent spatial metadata may include parameters such as diffuse-to-total energy ratio, surround coherence, and residual-to-total energy ratio.

[0049] An exemplary use case for IVAS is for AR / VR teleconferencing. Each participant can have their own object that they can freely pan around in 3D space. In a teleconferencing scenario, a conference bridge can, for example, receive several IVAS streams from multiple participants. These streams are then combined into a common stream, for example, using an object for at least each active participant. Alternatively, a pre-rendered spatial scene may be created and represented, for example, as MASA or FOA / HOA audio format. Incoming object or other mono streams (e.g., EVS streams) can be directly copied to become object streams in the outgoing common conference stream by attaching the appropriate metadata representation to the waveform. This may or may not involve re-encoding of the audio waveform. However, if a participant is sending a spatial audio stream such as MASA or HOA, the conference bridge must decode all incoming IVAS streams and reduce the stream to mono before sending them downstream as (mono) audio objects.

[0050] A further use case is when a user is capturing a scene using a mobile device on a fixed stand that is enabled for spatial audio capture (e.g., creating a live podcast video). In addition, a headset or some other form of close-up microphone can be used to enhance the audio recording. The close-up capture device can also capture spatial audio, for example, using a headset from a spatial audio-enabled lavalier microphone or binaural capture from a MASA. The close-up captured audio can then be added as an object stream to the device that captured the IVAS spatial audio stream. The object's location and distance can be conveniently captured, for example, using an appropriate position beacon attached to the close-up capture device. If IVAS only allows mono objects, the device must downmix the incoming spatial stream from the close-up capture to mono before embedding it in the IVAS stream. The embodiments described herein attempt to avoid or minimize added latency and complexity, while also increasing the maximum achievable quality.

[0051] Thus, some embodiments described herein provide increased flexibility for various IVAS audio inputs in audio source mixing and forwarding, for example, for AR / VR teleconferencing and other immersive use cases.

[0052] Additionally, some embodiments have substantially less delay and complexity, avoiding the need to generate downmixed spatial streams at the AR / VR conference bridge or capture device, and there is no loss of quality or original input properties in the converted audio format.

[0053] In some embodiments, the decoder is configured with an interface output format, a so-called pass-through mode, and has an external renderer with more capabilities than the normal integrated renderer operating as output mode.

[0054] With reference to FIG. 1, an exemplary system in which some embodiments may be implemented is shown. System 200 illustrates a conferencing scenario in which some participants transmit mono and some spatial streams, and some participants have mono, some spatial, and some 6DoF rendering and playback capabilities. For example, as shown in room A 209 of FIG. 1, user 202 is using mono capture and fixed spatial playback; in room B 213, user 206 is using spatial capture and 6DoF (degrees of freedom) playback; in room C 211, user 204 is using mono capture and playback; and in room D 215, users 208 and 210 are using spatial capture and mono object capture and spatial playback, but not head tracking. A conferencing service 201 connects all users.

[0055] The system shown in FIG. 1 has user-operated devices with different capabilities, and the embodiments described herein attempt to optimize the user's experience without requiring the conferencing service 201 to separately decode, mix, and encode the various inputs. In the embodiments described herein, any determination related to the level of immersion is made. For example, in some embodiments, the device may be implemented in the receiving UE. Thus, in some embodiments, an (IVAS) object stream can be configured to comprise a separate “objectified” (IVAS) data stream. Furthermore, the object metadata can be configured to include information on whether the object is a (mono) object-based audio representation (e.g., an EVS stream with spatial metadata) or a full IVAS spatial stream (e.g., MASA or stereo, or an object containing IVAS) that can provide object-like metadata (e.g., positional metadata). In such embodiments, any “objectified” (IVAS) data stream can include other (IVAS) objects. These (IVAS) objects can be moved to become part of other (IVAS) objects or the “main” (IVAS) data stream. The object metadata is then updated to remain meaningful across the newly formed IVAS stream. Additionally, in some embodiments, the remainder of the object metadata field is updated according to the spatial scene description.

[0056] In such embodiments, higher quality and lower latency are expected for conference bridge use cases where the input audio stream is spatially captured / created. Furthermore, some embodiments may be implemented in use cases where, for example, there is primary spatial audio captured by a mobile phone (UE), and additional spatial audio objects are captured by a wireless microphone, for example, similarly enhancing the voice capture benefits and enabling (IVAS) encoding on a new class of device (wireless microphone) without the need to decode the audio at the UE to enable further encoding. Instead, the stream can simply be embedded as is.

[0057] Before describing the embodiments further, we first describe a system for acquiring and rendering spatial audio signals that may be used in some embodiments.

[0058] With reference to FIG. 2, an exemplary apparatus suitable for use in a system such as that shown in FIG. 1 and for implementing some embodiments as described herein is shown.

[0059] 2A illustrates an apparatus suitable for implementing some embodiments, for example, with respect to a user in Room A. In this example, the apparatus comprises a single microphone 101 configured to generate a mono audio signal that is passed to an encoder 103. The apparatus further comprises an encoder 103 configured to receive the mono audio signal and encode the mono audio signal before transmitting it to an appropriate conference network.

[0060] FIG. 2A further shows a decoder / renderer 105 configured to receive the encoded spatial / mono audio signals which are decoded and rendered into appropriate audio signal outputs, which are passed to a number of speakers 107 for outputting the spatial audio signals to a user.

[0061] 2B further illustrates an exemplary apparatus suitable for implementing some embodiments with respect to a user in Room B. In this example, the apparatus comprises a plurality of microphone 111 audio inputs configured to generate a plurality of audio signals that can be used to generate a spatial audio signal that is passed to an encoder 113. The apparatus further comprises an encoder 113 configured to receive the spatial audio signals and encode the spatial audio signals before transmitting them to an appropriate conference network.

[0062] FIG. 2B further shows a decoder / renderer 115 configured to receive the encoded spatial / mono audio signal which is decoded and rendered into an appropriate audio signal output which is passed to headphones with a head tracker / locator 117 which outputs the spatial audio signal to the user and passes the user position to the decoder / renderer 115 to control the rendering.

[0063] 2C shows an exemplary device suitable for implementing some embodiments with respect to a user in Room C. In this example, the device comprises a mono microphone 121 audio input configured to generate a mono audio signal, which may be used to generate a mono audio signal that is passed to an encoder 123. The device further comprises an encoder 123 configured to receive the mono audio signal and encode the mono audio signal as a spatial audio signal before transmitting it to an appropriate conference network.

[0064] FIG. 2C further shows a decoder / renderer 125 configured to receive the encoded spatial / mono audio signal, which is decoded and rendered into an appropriate audio signal output, which is passed to a mono speaker 127 to output the audio signal to the user.

[0065] 2D further illustrates an exemplary device suitable for implementing some embodiments with respect to a user in room D. In this example, the device comprises multiple microphone 131 audio inputs configured to generate multiple audio signals, and an external microphone (e.g., a mono microphone or a multi-microphone) that may be used to generate a spatial audio signal and an external mono / spatial audio signal that are passed to encoder 133. The device further comprises encoder 133 configured to receive the spatial / mono audio signal and encode the spatial / mono audio signal before transmitting it to an appropriate conference network.

[0066] FIG. 2D further shows a decoder / renderer 135 configured to receive the encoded spatial / mono audio signals which are decoded and rendered into appropriate audio signal outputs, which are passed to headphones 137 to output the spatial audio signals to the user.

[0067] With reference to FIG. 3, a high-level view of an exemplary (IVAS) encoder 103 / 113 / 123 / 133 is shown, including, by way of non-exclusive example, various inputs that may be expected for the codec.

[0068] In some embodiments, the encoder 103 / 113 / 123 / 133 includes an audio (IVAS) input 301. The audio input 301 is configured to receive one or more configurations of spatial data (IVAS) streams from multiple sources, either local or remote. The source(s) may be local, such as multiple spatial capture devices with known spatial configurations at the encoder's location, and / or multiple remote participants transmitting spatial IVAS streams. The audio input 301 is configured to pass the audio data stream to an object header creator 303 and to an (IVAS) decoder 311 as part of an IVAS data stream processor 313.

[0069] In some embodiments, the encoder 103 / 113 / 123 / 133 comprises a scene control 305 configured to control the processing of the received audio input 301 .

[0070] For example, in some embodiments, the encoder 103 / 113 / 123 / 133 comprises an object header creator 303. Controlled by a scene control unit 305, the object header creator 303 is configured to insert each data stream as an object into a "master" data stream. In some embodiments, the object header creator 305 can be further configured to add missing object parameters, such as distance and direction, based on either a true spatial configuration or a virtually defined scene.

[0071] In some embodiments, the object header creator 303 is configured to determine if the inserted data stream contains objects, and to either move those audio objects freely and update their metadata so that they are directly part of the "master" IVAS stream, or move the objects under any other IVAS object. Additionally, the object header creator 303 is configured to update the object metadata so that it is correct for the overall spatial configuration.

[0072] The encoder 103 / 113 / 123 / 133 in some embodiments comprises an IVAS data stream processor 313. The IVAS data stream processor 313 may comprise an (IVAS) decoder 311. The (IVAS) decoder 311 is configured to receive one or more settings of the spatial audio data stream, decode the spatial audio signals, and pass them to the audio scene renderer 231.

[0073] The IVAS data stream processor 313 may include an audio scene renderer 231 configured to receive the audio signal and generate an audio scene rendering based on the decoded (IVAS) spatial audio signal. The audio scene rendering may, for example, comprise a downmix of various inputs from the (IVAS) decoder 311. The rendered audio scene audio signal may then be passed to the encoder 315.

[0074] The IVAS data stream processor 313 may comprise an encoder 315 for receiving the rendered spatial audio signals and encoding them. In other words, the IVAS data stream processor 313 is configured to decode all or at least some of the incoming data streams and generate a common spatial scene using, for example, IVAS MASA, IVAS HOA / FOA or IVAS mono objects.

[0075] In some embodiments where there are multiple embedded objects, these can be transmitted to receivers with high rendering capabilities available, while the remaining receivers receive only the pre-rendered spatial scene. Alternatively, a combination of at least one "IVAS Stream Object" and a pre-rendered "Spatial Scene IVAS Stream Object" can be used to reduce the bitrate.

[0076] Furthermore, the encoder comprises an audio object multiplexer 309 configured to combine the objects and output a combined object data stream.

[0077] The operation of the encoder is further illustrated by the flow diagram of FIG.

[0078] In step 401, an audio (IVAS) data stream is received in FIG.

[0079] Additionally, spatial scene configuration and control is determined in FIG.

[0080] based on the determined spatial scene configuration and control and the input audio data stream; An object header for the audio data stream is created by step 403 as shown in FIG.

[0081] Further, optionally, the data stream is decoded based on the determined spatial scene configuration and control and the input audio data stream, as shown in FIG. 4, by step 404 .

[0082] The decoded data stream can then be rendered, as shown in FIG. 4, via step 406 .

[0083] The rendered audio scene is then encoded using a suitable (IVAS) encoder, as shown in FIG. 4, per step 408 .

[0084] The data stream can then be multiplexed and output as shown in FIG. 4, via step 409 .

[0085] IVAS object stream metadata can utilize any suitable acoustic / spatial metadata, an example of which is shown in the table below. [Table 1]

[0086] However, in some embodiments, other location information such as xyz or Cartesian coordinates may be used instead of azimuth-elevation-distance. For example, further configurations may be provided by tables. [Table 2]

[0087] However, some minimal stream description metadata is additionally required to signal (IVAS) object data stream configuration information. For example, this information may be signaled using the following format: [Table 3]

[0088] In such an embodiment, the "Stream ID" parameter is used to uniquely identify each IVAS object stream in the current session. It can therefore signal each original and mixed audio component (input stream). For example, the signal allows for identification of the component within the system or on a user interface. The "Stream Type" parameter defines the meaning of each "audio object." Thus, in some embodiments, an audio object is not just an object-based audio input. Rather, the object data stream can be object-based audio (input) or any IVAS scene. An example of this is shown in Figure 5, where three types of objects are shown:

[0089] For example, Figure 5A shows a simple conventional (mono) audio object 501. Audio object 501 is defined by a PCM audio signal portion 505 and an acoustic (spatial) metadata portion 503. It will be understood that additional metadata may be present.

[0090] With reference to Figure 5B, a coded representation 507 of the same audio object as shown in Figure 5A is shown.

[0091] Figure 5C shows the same audio object as shown in Figures 5A and 5B, but processed according to some embodiments as discussed herein. The processed audio object is described as an object data stream 509 defined by a "stream type = 0" parameter 513. In other words, the object data stream 509 includes a data stream identifier that identifies it as an object-based audio IVAS object stream. Furthermore, the object data stream 509 includes an object audio bitstream portion 515 (a coded representation of the audio object) and a stream identifier 511 that uniquely identifies the object data stream.

[0092] FIG. 5D shows a further (IVAS) object data stream 517. The further object data stream 517 includes an identifier portion 521 with "Stream Type=1." In some embodiments, Stream Type=0 corresponds to a "simple" object type, e.g., a mono signal. Furthermore, in some embodiments, Stream Type=1 corresponds to a potentially "complex" stream. For example, in this example, Stream Type=1 corresponds to a complete IVAS stream, which in this case includes a MASA spatial stream. IVAS may contain one or more object streams, allowing for nested objects. When Stream Type=0, it is known that there are no further objects and the stream is a simple type (actually a mono object).

[0093] The further object data stream 517 may further comprise an explicit stream description 523, or the stream content may be determined by starting the decoding of the object stream, in which case it is explicitly described as a MASA-based scene (e.g., "stream description=MASA").

[0094] Furthermore, the object data stream 517 comprises a MASA format bitstream portion 525 (a coded representation of the audio object) and a stream identifier 519 "Stream ID=000002" that uniquely identifies the object data stream.

[0095] The first advantage of the approach discussed herein is that IVAS inputs can often be conveniently forwarded without decoding / encoding operations. For example, if a mixer device, a remote conference bridge (e.g., an AR / VR conference server), or other entity used to combine and / or forward audio inputs exists in the IVAS end-to-end service, decoding / encoding operations are not necessary. Therefore, by reassigning the received (encoded) input as an IVAS object stream, operational complexity and latency are reduced. For example, if the playback capabilities of the receiver are unknown, the server can optimize complexity by simply providing the received scene as is. Any IVAS stream can be decoded and rendered as mono to support even the simplest IVAS devices. Also, skipping the decoding / encoding operations at an intermediate point (e.g., the conference server) reduces the end-to-end latency of that audio component. Therefore, the user experience is improved.

[0096] Furthermore, embodiments are configured such that only shallowly embedded "objective" IVAS streams exist. In other words, if there is an object stream that also contains objects (and thus can contain multiple levels of objects), deep data structures are avoided, thus reducing decoder complexity. Thus, embedding as proposed in some embodiments allows IVAS objects to contain other IVAS objects, but in other words, any "deep" object can be moved to a "higher" object, in some embodiments closer to the "master" IVAS stream, and its metadata can be updated so that its representation remains meaningful for the newly formed scene. In some embodiments, an IVAS object can be moved to become part of another IVAS object. Thus, the object is moved "deeper." This may allow, for example, audio objects (e.g., mono objects) to be coded or decoded together to save complexity or bitrate. When formats of the same type are at different levels in the structure, they generally need to be coded / decoded at different times or using different instances. This can introduce additional complexity.

[0097] Furthermore, the embodiments discussed herein may have a second advantage in that it allows for convenient nesting of IVAS object streams, e.g., for content distribution purposes. In such embodiments, a more complex scene can be treated as a single (mono) audio object. An example of nested packetization is shown in Figure 6. This can be used, for example, to distribute decoding complexity. This can be very useful, for example, for edge cloud services.

[0098] 6 shows an overall scene object data stream 601. The overall scene object data stream 601 includes a plurality of object data streams 602, 604, 606, and 608. For example, the first object data stream 602 comprises a stream ID 621 (stream ID=000001) that uniquely identifies the object data stream, a stream type identifier 623 (stream type=0), and a data portion 625. The second object data stream 604 comprises a stream ID 631 (stream ID=000006) that uniquely identifies the object data stream, a stream type identifier 633 (stream type=1), and a data portion 635. The third object data stream 606 comprises a stream ID 641 (stream ID=000007) that uniquely identifies the object data stream, a stream type identifier 643 (stream type=1), and a data portion 645. The fourth object data stream 608 comprises a stream ID 651 (stream ID=000008) that uniquely identifies the object data stream, a stream type identifier 653 (stream type=0), and a data portion 655.

[0099] 6, the second object data stream 604 further comprises nested object data streams 612 and 614. These may be, for example, object data streams associated with subsections of the overall scene. The fifth object data stream 612 comprises a stream ID 661 (stream ID=000004) that uniquely identifies the object data stream, a stream type identifier 663 (stream type=0), and a data portion 665. The sixth object data stream 614 comprises a stream ID 671 (stream ID=000005) that uniquely identifies the object data stream, a stream type identifier 673 (stream type=1), and a data portion 675.

[0100] Furthermore, the sixth nested object data stream 614 further includes nested object data streams 622 and 624. These may, for example, be object data streams associated with subsections of subsections of the overall scene. The seventh object data stream 622 comprises a stream ID 681 (stream ID=000002) that uniquely identifies the object data stream, a stream type identifier 683 (stream type=1), and a data portion 685. The eighth object data stream 624 comprises a stream ID 691 (stream ID=000003) that uniquely identifies the object data stream, a stream type identifier 693 (stream type=1), and a data portion 695.

[0101] A further advantage of implementing some embodiments is that any IVAS input or IVAS scene that already contains spatial parameters, such as positional characteristics, can determine such characteristics. For example, this can be implemented by adding acoustic spatial metadata (e.g., one of the parameters from the previous table) to the IVAS object stream ("Stream Type=1"). This enables an enhanced experience, for example, in AR / VR teleconferencing use cases.

[0102] For example, FIG. 7 shows a capture scene 701 with a UE or similar capture device 707 implementing spatial capture at a first (UE) location and a second capture device 705 implementing second spatial capture (or object capture) at a second (user) location.

[0103] The conventional approach shown in the upper right of FIG. 1 shows where the audio object rendering 713 is located and where the first spatial capture scene 711 is located. Thus, a user can use a multi-microphone UE to capture a spatial scene (e.g., in MASA format), but the user can use a close-up microphone or, for example, a second UE that can be connected to a "master" device to capture audio objects. These two inputs are combined and provided to an IVAS encoder. In terms of the listening experience, it is possible to listen to a combined rendering of spatial audio (e.g., background audio) and audio objects (e.g., user voice).

[0104] By implementing embodiments as described herein, a listener can switch (730) between a first option of audio object rendering of the second spatial capture 723 and a second option of audio object rendering of the first spatial capture scene 721 or the first spatial capture 733 and the second spatial capture scene 731. Thus, the IVAS codec can import the second spatial audio representation as an IVAS object stream. Thus, when a user captures a spatial audio scene using the user's UE, a second UE connected to a wireless multi-microphone device, or indeed a "master" UE, can capture a complete spatial representation of the sound scene at a second location. This sound scene can be encoded as an IVAS bitstream by a second device and provided to a second UE that can "act as a conference bridge," ingesting the IVAS bitstream and embedding it as an IVAS object stream, which is then delivered to the listener in two spatial audio scenes. For example, a user can switch between them, with a mono downmix of each scene provided as an audio object rendering of the other scene being rendered for the user.

[0105] While Figure 6 shows an example of object stream nesting, it should be understood that this is not the only mechanism for IVAS stream transport / packetization enabled by the present invention. Figure 8 shows two examples of IVAS stream packetization according to some embodiments.

[0106] In some embodiments, a lookup table may be used that specifies the packet contents. The lookup table may be defined as a "payload header," which may be, for example, an RTP payload header. This may include, for example, the sizes of the various blocks. Following the header is the payload.

[0107] For example, as shown in Figure 8, a data stream can include various IVAS object streams and IVAS content. Thus, the entire scene object stream 801 includes a payload header 811 or lookup table that can specify packet contents. For example, as shown in Figure 8A, it specifies payloads such as a first object data stream 813 and a second object data stream 819, and a first payload 815 (MASA and objects) and a second payload 817 (5.1 channel audio data).

[0108] In some embodiments shown in Figure 8C, the data stream may include only IVAS object streams. Thus, the entire scene object stream 831 comprises a payload header or lookup table that can specify packet contents including object data stream 833, which may comprise nested object data stream 835, which may further comprise nested object data streams.

[0109] FIG. 8B shows a "hybrid" embodiment with payload and nested object data streams 813 across the scene.

[0110] There is an associated cost of nesting in generating additional "payload header" information and parsing it.

[0111] Regarding the decoders / renderers 105, 115, 125, 135: The decoders / renderers 105, 115, 125, 135 are configured to receive various (IVAS) object data streams and to decode and render the data streams in parallel.

[0112] In some embodiments, processing of nested audio object data streams may be performed separately for each sub-scene level and then combined at a higher level.

[0113] For example, with respect to the example shown in FIG. 6, here decoding may begin with "Stream ID=000002" and "Stream ID=000003." Therefore, "Stream ID=000005" is decoded (as a container for sub-scenes). The decoder can then be configured to decode the next "Stream ID=000004." After this, other streams are decoded. This approach can have advantages, for example, in memory consumption, where certain memory can be freed between sub-scene levels and therefore the overall memory footprint is not defined by all the streams combined.

[0114] In such an embodiment, rendering may be performed at a sub-scene level using summation within the rendered domain, or composite rendering may be performed at the beginning of the decoding.

[0115] In some embodiments, the decoder is configured to launch a separate decoder instance for each sub-scene, so for each "stream type=1" a separate IVAS decoder instance is initialized.

[0116] 9, an exemplary electronic device that may be used as an analysis or synthesis device is shown. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.

[0117] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 can be configured to execute various program code, such as the methods described herein.

[0118] In some embodiments, device 1400 comprises memory 1411. In some embodiments, at least one processor 1407 is coupled to memory 1411. Memory 1411 can be any suitable storage means. In some embodiments, memory 1411 comprises program code sections for storing program code executable on processor 1407. Additionally, in some embodiments, memory 1411 can further comprise a stored data section for storing data, e.g., data that has been processed or is to be processed according to embodiments described herein. The executed program code stored in the program code sections and the data stored in the stored data sections can be retrieved by processor 1407 via the memory-processor coupling as needed.

[0119] In some embodiments, the device 1400 comprises a user interface 1405 . The user interface 1405 may be coupled to the processor 1407 in some embodiments. In some embodiments, the processor 1407 may control the operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1405 may allow a user to input commands into the device 1400, for example, via a keypad. In some embodiments, the user interface 1405 may allow a user to obtain information from the device 1400. For example, the user interface 1405 may comprise a display configured to display information from the device 1400 to the user. The user interface 1405 may, in some embodiments, comprise a touchscreen or touch interface that can both allow information to be input into the device 1400 and further display information to the user of the device 1400. In some embodiments, the user interface 1405 may be a user interface for communicating with a position determiner, as described herein.

[0120] In some embodiments, device 1400 comprises an input / output port 1409. In some embodiments, input / output port 1409 comprises a transceiver. The transceiver in such embodiments may be coupled to processor 1407 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may in some embodiments be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.

[0121] The transceiver may communicate with the further device by any suitable known communication protocol, for example, in some embodiments the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an infrared data channel (IRDA).

[0122] The transceiver input / output port 1409 can be configured to receive the signal, and in some embodiments, determine the parameters as described herein by using the processor 1407 executing appropriate code. Additionally, the device may generate an appropriate downmix signal and parameter output to be sent to the combining device.

[0123] In some embodiments, device 1400 may be used as at least part of a synthesis device. Thus, input / output port 1409 may be configured to receive the downmix signal and, in some embodiments, receive parameters determined in a capture device or processing device as described herein and generate an appropriate audio signal format output by using processor 1407 to execute appropriate code. Input / output port 1409 may be coupled to any appropriate audio output, for example, a multi-channel speaker system and / or headphones (which may be head-tracked or non-head-tracked headphones) or the like.

[0124] In general, various embodiments of the present invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. While various aspects of the present invention may be illustrated and depicted as block diagrams, flowcharts, or using some other graphical representation, it is fully understood that these blocks, apparatus, systems, techniques, or methods contemplated herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller, or other computing device, or some combination thereof.

[0125] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, such as within a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of a logic flow, such as that shown in the figures, may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. Software may be stored on object-oriented media, such as memory chips, or memory blocks implemented within a processor, magnetic media, such as hard disks or floppy disks, and optical media, such as DVDs and their data variants, CDs.

[0126] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memory, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0127] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is a large-scale, highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.

[0128] Programs such as those offered by Synopsys, Inc. of Mountain View, California, and Cadence Design, Inc. of San Jose, California, automatically route conductors and locate components on semiconductor chips using well-established design rules and pre-stored libraries of design modules. Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor manufacturing facility or "fab" for fabrication.

[0129] The foregoing description has provided a full and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting example. However, various modifications and adaptations will become apparent to those skilled in the art in view of the foregoing description upon perusal of the accompanying drawings and the appended claims. However, all such similar modifications of the teachings of this invention will still fall within the scope of the present invention, as defined in the appended claims.

Claims

1. 1. An apparatus comprising: at least one processor; When executed by the at least one processor, the device includes at least: receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream configured to enable immersive audio during communication; determining a type of each of the received first and second audio data streams to identify which of the received first and second audio data streams includes the spatial audio stream; processing the second audio data stream using at least one parameter dependent on the determined type of the second audio data stream; rendering the first audio data stream and the processed second audio data stream; at least one memory storing instructions for executing the An apparatus comprising:

2. the second audio data stream is configured to comprise at least one further audio data stream; the at least one further audio data stream comprises a determined type; the at least one further audio data stream is an embedded-level audio data stream relative to the second audio data stream.

10. The apparatus of claim 1.

3. said at least one further audio data stream comprising at least one further embedding level; each of the at least one further embedding level includes at least one additional audio data stream having a determined type; 3. The apparatus of claim 2.

4. 10. The apparatus of claim 1, wherein the second audio data stream is a master level audio data stream.

5. The first audio data stream and the second audio data stream further include: a stream identifier configured to uniquely identify the audio data stream; and a stream descriptor configured to describe a type of said audio data stream; The device of claim 1 , wherein the device is associated with at least one of:

6. The type of at least one of the first audio data stream or the second audio data stream is: Mono audio signal type, or Immersive Voice and Audio Services audio signals, 10. The device of claim 1, wherein the device is one of:

7. The apparatus of claim 1 , wherein the at least one parameter is configured to define a room characteristic or a scene description.

8. The at least one memory, when executed by the at least one processor, causes the apparatus to further: receiving an additional audio data stream; embedding the additional audio data stream into either the first audio data stream or the second audio data stream; 10. The apparatus of claim 1, further comprising instructions for causing the apparatus to:

9. 1. A method for an apparatus for immersive audio communication, comprising: receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream or the second audio data stream includes a spatial audio stream configured to enable immersive audio during communication; determining a type of each of the received first and second audio data streams to identify which of the received first and second audio data streams includes the spatial audio stream; processing the second audio data stream using at least one parameter dependent on the determined type of the second audio data stream; rendering the first audio data stream and the processed second audio data stream; A method comprising:

10. the second audio data stream is configured to include at least one further audio data stream; the at least one further audio data stream comprises a determined type; the at least one further audio data stream is an embedded-level audio data stream relative to the second audio data stream.

10. The method of claim 9.

11. the at least one further audio data stream comprises at least one further embedding level; each of the at least one further embedding level includes at least one additional audio data stream having a determined type; The method of claim 10.

12. 10. The method of claim 9, wherein the second audio data stream is a master level audio data stream.

13. The first audio data stream and the second audio data stream each further comprise: a stream identifier configured to uniquely identify the audio data stream; a stream descriptor configured to describe a type of said audio data stream; The method of claim 9 , wherein the method is associated with at least one of:

14. The type of at least one of the first audio data stream or the second audio data stream is: Mono audio signal type, or Immersive Voice and Audio Services audio signals, 10. The method of claim 9, wherein the

15. The method of claim 9 , wherein the at least one parameter is configured to define a room characteristic or a scene description.

16. The at least one parameter defining the room characteristic or scene description is: direction, Azimuth, azimuth elevation, distance, gain, spatial extent, Energy ratio, or position, 16. The method of claim 15, comprising at least one of:

17. receiving an additional audio data stream; embedding the additional audio data stream into either the first audio data stream or the second audio data stream; 10. The method of claim 9, further comprising:

18. The method of claim 9 , wherein the at least one parameter is generated based on spatial audio capture.

19. The at least one parameter defining the room characteristic or scene description is: direction, Azimuth, azimuth elevation, distance, gain, spatial extent, Energy ratio, or position, The apparatus of claim 7 , comprising at least one of:

20. The apparatus of claim 7 , wherein the at least one parameter is generated based on spatial audio capture.

Citation Information

Patent Citations

  • Spatial audio capture, transmission and reproduction

    GB2575509A

  • multimedia systems based on mpeg-4, service providers for such systems and telecommunications equipment based on content

    JP2004538727A

  • Decoder and method for multi-instance spatial acoustic object coding employing a parametric concept for multi-channel downmix / upmix configurations

    JP2015527611A

  • Audio representation and related rendering

    JP2023516303A

  • Multi-stream audio coding

    US20190103118A1