Audio representation and associated rendering
By identifying and processing types, embedding and rendering technologies of multiple audio data streams, the complexity and latency of audio formats in immersive audio communications are solved, and low-latency and high-quality audio transmission is achieved, improving the user experience.
Patent Information
- Application Number
- CN202180016741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-28
- Filing Date
- 2021-02-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-02-10
AI Technical Summary
The prior art is difficult to effectively deal with the complexity and latency problems of multiple audio data streams in immersive audio communications, especially under low bit rate conditions, especially in virtual reality, augmented reality and mixed reality applications, where audio codecs have difficulty supporting flexible mixing and low latency transmission of multiple audio formats.
By receiving and identifying the types of different audio data streams, using parameter processing and rendering techniques, spatial audio streams are embedded in other audio streams, leveraging metadata updates and object stream conversions, reducing decoding and encoding operations, and optimizing audio mixing and transmission processes.
It enables flexible mixing and transmission of multiple audio formats at low latency and low complexity, improves audio quality and user experience, and reduces processing delays and device complexity of audio data streams.
Smart Images

Figure CN115211146B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to apparatuses and methods for audio representation and associated rendering related to sound fields, but non-exclusively to apparatuses and methods for audio representation for audio encoders and decoders. Background Art
[0002] Immersive audio codecs are being implemented to support a large number of operating points ranging from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Service (IVAS) codec, which is designed to be suitable for use on communication networks such as 3GPP 4G / 5G networks. Such immersive services include, for example, the use of immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR), and mixed reality (MR). It is expected that this audio codec processes the encoding, decoding, and rendering of voice, music, and general audio. It is also expected that this audio codec supports channel-based audio and scene-based audio inputs, including spatial information about the sound field and sound sources. It is further expected that the codec operates with low latency to enable session services and supports high error robustness under various transmission conditions.
[0003] In addition, parametric spatial audio processing is an area of audio signal processing that uses a set of parameters to describe the spatial aspects of sound. For example, when performing parametric spatial audio capture from a microphone array, estimating a set of parameters from the microphone array signals is a typical and effective option, such as the direction of the sound in a frequency band, and the ratio of the directional to non-directional parts of the captured sound in the frequency band. It is well known that these parameters well describe the perceived spatial characteristics of the captured sound at the location of the microphone array. These parameters can accordingly be used in the synthesis of spatial sound for use with binaural headphones, speakers, or other formats such as Ambisonics. Summary of the Invention
[0004] According to a first aspect, there is provided an apparatus for immersive audio communication, comprising components configured to perform the following operations: receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; determining the type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; processing the second audio data stream with at least one parameter according to the determined type; and rendering the first audio data stream and the processed second audio data stream.
[0005] The second audio data stream may be configured to include at least one other audio data stream, and wherein, the at least one other audio data stream may include the determined type, and the at least one other audio data stream may be an embedded-level audio data stream relative to the second audio data stream.
[0006] The at least one other audio data stream may include at least one other embedded level, wherein each embedded level may include at least one additional audio data stream having the determined type.
[0007] The second audio data stream may be a primary-level audio data stream.
[0008] Each audio data stream may be further associated with at least one of the following: a stream identifier configured to uniquely identify the audio data stream; and a stream descriptor configured to describe the type of the audio data stream.
[0009] The type may be one of the following: a mono audio signal type; an immersive voice and audio service audio signal.
[0010] At least one parameter may be configured to define an indoor characteristic or a scene description.
[0011] The at least one parameter defining the indoor characteristic or the scene description may include at least one of the following: direction; direction azimuth; direction elevation; distance; gain; spatial extent; energy ratio; and position.
[0012] The component may be further configured to: receive an additional audio data stream; and embed the additional audio data stream within one or the other of the first audio data stream and the second audio data stream.
[0013] According to a second aspect, there is provided a method for a device, the method comprising: receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; determining the type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; processing the second audio data stream with at least one parameter according to the determined type; and rendering the first audio data stream and the processed second audio data stream.
[0014] The second audio data stream may be configured to include at least one other audio data stream, and wherein, the at least one other audio data stream may include the determined type, and the at least one other audio data stream may be an embedded-level audio data stream relative to the second audio data stream.
[0015] At least one other audio data stream may include at least one other embedding level, wherein each embedding level may include at least one additional audio data stream having a determined type.
[0016] The second audio data stream may be a primary level audio data stream.
[0017] Each audio data stream may be further associated with at least one of the following: a stream identifier configured to uniquely identify the audio data stream; and a stream descriptor configured to describe the type of the audio data stream.
[0018] The type may be one of the following: a mono audio signal type; an immersive voice and audio service audio signal.
[0019] At least one parameter may be configured to define indoor characteristics or a scene description.
[0020] The at least one parameter defining indoor characteristics or a scene description may include at least one of the following: direction; direction azimuth; direction elevation; distance; gain; spatial extent; energy ratio; and position.
[0021] The method may further include: receiving an additional audio data stream; embedding the additional audio data stream in one or the other of the first audio data stream and the second audio data stream.
[0022] According to a third aspect, there is provided an apparatus including at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; determine the type of each audio data stream of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; process the second audio data stream with at least one parameter according to the determined type; and render the first audio data stream and the processed second audio data stream.
[0023] The second audio data stream may be configured to include at least one other audio data stream, and wherein the at least one other audio data stream may include a determined type, and the at least one other audio data stream may be an embedded level audio data stream relative to the second audio data stream.
[0024] At least one other audio data stream may include at least one other embedding level, wherein each embedding level may include at least one additional audio data stream having a determined type.
[0025] The second audio data stream may be a primary audio data stream.
[0026] Each audio data stream may be further associated with at least one of the following: a stream identifier configured to uniquely identify the audio data stream; and a stream descriptor configured to describe the type of the audio data stream.
[0027] The type may be one of the following: a mono audio signal type; an immersive voice and audio service audio signal.
[0028] At least one parameter may be configured to define an indoor characteristic or a scene description.
[0029] The at least one parameter defining the indoor characteristic or the scene description may include at least one of the following: orientation; azimuth of orientation; elevation of orientation; distance; gain; spatial extent; energy ratio; and position.
[0030] The apparatus may be further caused to: receive an additional audio data stream; and embed the additional audio data stream within one or the other of the first audio data stream and the second audio data stream.
[0031] According to a fourth aspect, there is provided an apparatus comprising: a receiving circuit configured to receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; a determining circuit configured to determine the type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; a processing circuit configured to process the second audio data stream with at least one parameter according to the determined type; and a rendering circuit configured to render the first audio data stream and the processed second audio data stream.
[0032] According to a fifth aspect, there is provided a computer program [or a computer-readable medium comprising program instructions] comprising instructions for causing an apparatus to at least perform the following operations: receive at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; determine the type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; process the second audio data stream with at least one parameter according to the determined type; and render the first audio data stream and the processed second audio data stream.
[0033] According to a sixth aspect, there is provided a non-transitory computer-readable medium including program instructions for causing an apparatus to perform at least the following operations: receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; determining a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; processing the second audio data stream with at least one parameter according to the determined type; and rendering the first audio data stream and the processed second audio data stream.
[0034] According to a seventh aspect, there is provided an apparatus including: means for receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; means for determining a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; means for processing the second audio data stream with at least one parameter according to the determined type; and means for rendering the first audio data stream and the processed second audio data stream.
[0035] According to an eighth aspect, there is provided a computer-readable medium including program instructions for causing an apparatus to perform at least the following operations: receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream to enable immersive audio during communication; determining a type of each of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; processing the second audio data stream with at least one parameter according to the determined type; and rendering the first audio data stream and the processed second audio data stream.
[0036] An apparatus includes means for performing the actions of the method as described above.
[0037] An apparatus is configured to perform the actions of the method as described above.
[0038] A computer program includes program instructions for causing a computer to perform the method as described above.
[0039] A computer program product stored on a medium can cause an apparatus to perform the method described herein.
[0040] An electronic device may include the apparatus as described herein.
[0041] A chipset may include the apparatus as described herein.
[0042] Embodiments of the present application are intended to solve problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] To better understand the present application, reference will now be made, by way of example, to the accompanying drawings, in which:
[0044] Figure 1 Schematically shows an example conference system suitable for adopting some embodiments;
[0045] Figures 2a to 2d schematically show a system of an apparatus suitable for implementing some embodiments;
[0046] Figure 3 Schematically shows a bitstream-object-bitstream converter according to some embodiments;
[0047] Figure 4 Schematically shows, according to some embodiments, as Figure 3 shown in the flowchart of the operation of the bitstream-object-bitstream converter;
[0048] Figures 5a to 5d show an example object format according to some embodiments;
[0049] Figure 6 Shows an example object nesting according to some embodiments;
[0050] Figure 7 Shows an example operation scenario according to some embodiments;
[0051] Figures 8a to 8c show an example object grouping according to some embodiments; and
[0052] Figure 9 Shows an example device suitable for implementing the shown apparatus. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The suitable apparatus and possible mechanisms for embedding a spatial stream as an object stream and sending the spatial stream as an object as it is to a receiving participant will be described in more detail below. The object metadata is updated based on the spatial scenario. In other words, the object stream type itself is another audio stream having corresponding object metadata generated by a processing unit. This operation can be performed by a suitable device (e.g., a mobile device, a user equipment UE) that receives more than one input format or, for example, a conference bridge (e.g., a multipoint control unit MCU).
[0054] The present invention relates to an immersive audio codec capable of supporting multiple input audio formats, immersive audio scene representation, and services in which the input encoded audio can be mixed, re-encoded, and / or forwarded to a listener.
[0055] The IVAS codec discussed above is an extension of the 3GPP EVS codec and is intended for new real-time immersive voice and audio services over 4G / 5G. Such immersive services include, for example, immersive voice and audio for virtual reality (VR) and augmented reality (AR). The multi-functional audio codec is expected to handle the encoding, decoding, and rendering of voice, music, and general audio. The audio codec is expected to support channel-based audio and scene-based audio inputs, including spatial information about the sound field and sound sources. The audio codec is also expected to operate with low latency to enable session services and support high error robustness under various transmission conditions.
[0056] The IVAS encoder is configured to be able to receive inputs in the supported formats (and certain allowed combinations of some formats). Similarly, the decoder is expected to be able to output audio in multiple supported formats. A pass-through mode has been proposed in which the audio can be provided in its original format after transmission (encoding / decoding).
[0057] Methods have been proposed to describe an acceptable format in which object-based audio is implemented as an IVAS codec, which is configured to process spatial metadata combined with a suitable (mono) audio signal and which can be rendered to the user. The metadata parameters can, for example, be captured from the real environment with the help of any visual or auditory tracking method or any other sensory channel / modality. In some embodiments, radio-based techniques can be used to generate the metadata, for example, Bluetooth, Wifi, or GPS locator technology can be used to obtain object coordinates. In some embodiments, sensors such as magnetometers, accelerometers, and / or gyroscopes can be used to receive orientation data. In addition, other sensors such as proximity sensors can also be used to generate scene-related metadata from the real environment.
[0058] Alternatively, the metadata can be artificially created according to a defined virtual scene, for example, via a conference bridge or via a user device (e.g., a smart phone). For example, the user can set or indicate some desired acoustic characteristics via a suitable UI.
[0059] In some embodiments, the object-based audio spatial metadata can be defined as one or more objects, where each object can be defined by parameters such as azimuth, elevation, distance, gain, and spatial extent.
[0060] In addition, Metadata-Assisted Spatial Audio (MASA) is a parametric spatial audio format and representation. At a high level, it can be considered a representation consisting of "N channels + spatial metadata". It is a scene-based audio format that is particularly suitable for spatial audio capture on practical devices such as smart phones, where a spherical array for FOA / HOA capture is unrealistic or inconvenient. The idea is to describe the sound scene based on the direction of the sound source that changes over time and frequency. If no directional sound source is detected, the audio is described as diffuse. In MASA (as currently proposed for IVAS), there can be one or two directions for each time-frequency (TF) tile. The spatial metadata is described relative to the direction and can include, for example, spatial metadata for each direction and common spatial metadata independent of the direction.
[0061] For example, the spatial metadata related to the direction can include parameters such as direction index, direct-to-total energy ratio, spread coherence, and distance. The spatial metadata independent of the direction can include parameters such as diffuse-to-total energy ratio, surround coherence, and remainder-to-total energy ratio.
[0062] An example use case of IVAS is AR / VR teleconferencing. In an AR / VR teleconference, each participant can have his / her own objects, which can be translated arbitrarily in 3D space. In the teleconference scenario, the conference bridge can, for example, receive a number of IVAS streams from multiple participants. Then, using, for example, the objects for at least each active participant, these streams are combined into a common stream. Alternatively, a pre-rendered spatial scene can be created and represented, for example, as a MASA or FOA / HOA audio format. If objects are used, the input objects or other mono streams (e.g., EVS streams) can be directly copied as the object stream of the output common conference stream by attaching the appropriate metadata representation to the waveform. This may or may not include re-encoding of the audio waveform. However, if the participants are sending spatial audio streams such as MASA or HOA, the conference bridge must decode all the input IVAS streams and reduce these streams to mono and then send them downstream as (mono) audio objects.
[0063] Another use case is where the user is capturing a scene (e.g., making a live podcast video) with a mobile device on a fixed mount that enables spatial audio capture. Additionally, a headset or some other form of close - up microphone can be used to enhance the voice recording. The close - up capture device is also capable of capturing spatial audio from a headset, for example, using binaural capture, or MASA from a lapel microphone that supports spatial audio. Further, the voice captured up close can be added to the IVAS spatial audio stream captured by the device as an object stream. The object position and distance can be conveniently captured, for example, using a suitable position beacon attached to the close - up capture device. When only mono objects are allowed in the IVAS, the device must down - mix the spatial stream from the close - up capture into mono and then embed it in the IVAS stream. The embodiments described herein attempt to avoid or minimize the added latency and complexity and also attempt to improve the maximum achievable quality.
[0064] Accordingly, some embodiments described herein increase the flexibility of various IVAS audio inputs in audio source mixing and forwarding. For example, in AR / VR conference calls and other immersive use cases.
[0065] Additionally, in some embodiments, the latency and complexity are significantly reduced, thus avoiding the generation of down - scaled spatial streams at the AR / VR conference bridge or capture device. Also, there is no loss of the original input characteristics and / or quality loss for the converted audio format.
[0066] In some embodiments, the decoder is configured to have an interface output format (so - called pass - through mode to give an external renderer more capabilities than a normal integrated renderer) to act as an output mode.
[0067] Regarding Figure 1 , an example system in which some embodiments can be implemented is shown. System 200 shows a conference scenario where some participants are sending mono and some spatial streams, and some participants have mono, some spatial, and even some 6DoF rendering and playback capabilities. For example, as Figure 1 shown, in room A 209, user 202 is using mono capture and fixed - spatial playback, in room B 213, user 206 is using spatial capture and 6DoF (degrees of freedom) playback, in room C 211, user 204 is using mono capture and playback, and in room D 215, users 208 and 210 are using spatial capture as well as mono - object capture and spatial playback, but without head tracking. Conference service 201 connects all users.
[0068] As Figure 1The system shown in [description] has a user operating device that has different capabilities, and the embodiments described herein attempt to optimize the user experience without requiring the conferencing service 201 to separately decode, mix, and encode various inputs. In the embodiments described herein, any decisions are related to the level of immersion. For example, in some embodiments, the device may be implemented at the receiving UE.
[0069] Thus, in some embodiments, an (IVAS) object stream may be configured to include another “objectified” (IVAS) data stream. Additionally, object metadata is configured to contain information whether the object is a (mono) object-based audio representation (e.g., an EVS stream with spatial metadata) or a full IVAS spatial stream that can be metadata of a given class of object (object-like metadata, e.g., location metadata) (e.g., MASA or stereo or even an object containing IVAS). In such an embodiment, any “objectified” (IVAS) data stream may contain another (IVAS) object. These (IVAS) objects can be moved around to become part of any other (IVAS) object or the “main” (IVAS) data stream. Subsequently, any one object metadata is updated such that it remains meaningful for the newly formed IVAS stream as a whole. Additionally, in some embodiments, the remaining object metadata fields are updated according to the spatial scene description.
[0070] In such an embodiment, for the conferencing bridge use case where the input audio stream is spatially captured / created, higher quality and lower latency are expected. Additionally, some embodiments may be implemented in use cases where there is primary spatial audio captured by, for example, a mobile phone (UE) and additional spatial audio objects are captured by a wireless microphone to, for example, similarly enhance the voice capture benefit and allow (IVAS) encoding on a new class of devices (wireless microphones) without decoding the audio at the UE to allow further encoding. Instead, the stream can simply be embedded as is.
[0071] Before discussing the embodiments further, we first discuss a system for obtaining and rendering a spatial audio signal that may be used in some embodiments.
[0072] Regarding FIG. 2, an example device is shown that is used within the system shown in [description] and is suitable for implementing some embodiments described herein. Figure 1 and is suitable for implementing some embodiments described herein.
[0073] Figure 2AFor example, a device suitable for implementing some embodiments with respect to a user in Room A is shown. In this example, the device includes a single microphone 101 configured to generate a mono audio signal that is passed to an encoder 103. The device also includes an encoder 103 configured to receive the mono audio signal and encode the mono audio signal before sending it to a suitable conferencing network.
[0074] Figure 2A A decoder / renderer 105 is also shown, configured to receive the encoded spatial / mono audio signal, which is decoded and rendered into suitable audio signal outputs that are passed to a plurality of speakers 107 to output the spatial audio signals to the user.
[0075] In addition, Figure 2B An example device suitable for implementing some embodiments with respect to a user in Room B is shown. In this example, the device includes a multi-microphone 111 audio input configured to generate a plurality of audio signals that can be used to generate a spatial audio signal that is passed to an encoder 113. The device also includes an encoder 113 configured to receive the spatial audio signal and encode the spatial audio signal before sending it to a suitable conferencing network.
[0076] Figure 2B A decoder / renderer 115 is also shown, configured to receive the encoded spatial / mono audio signal, which is decoded and rendered into suitable audio signal outputs that are passed to a headset equipped with a head tracker / locator 117 to output the spatial audio signals to the user and pass the user position to the decoder / renderer 115 to control the rendering.
[0077] Figure 2C An example device suitable for implementing some embodiments with respect to a user in Room C is shown. In this example, the device includes a mono microphone 121 audio input configured to generate a mono audio signal that can be used to generate a mono audio signal that is passed to an encoder 113. The device also includes an encoder 123 configured to receive the mono audio signal and encode the mono audio signal as a spatial audio signal before sending it to a suitable conferencing network.
[0078] Figure 2C A decoder 125 is also shown, configured to receive the encoded spatial / mono audio signal, which is decoded and rendered into suitable audio signal outputs that are passed to a mono speaker 127 to output the audio signals to the user.
[0079] In addition, Figure 2D An example device suitable for implementing some embodiments with respect to a user in room D is shown. In this example, the device includes a multi-microphone 131 audio input configured to generate a plurality of audio signals, and an external microphone (e.g., a mono microphone or a multi-microphone), which can be used to generate a spatial audio signal and an external mono / spatial audio signal that are passed to the encoder 113. The device also includes an encoder 133, which is configured to receive the spatial / mono audio signal and encode the spatial / mono audio signal before sending it to a suitable conferencing network.
[0080] Figure 2D A decoder / renderer 135 is also shown, which is configured to receive the encoded spatial / mono audio signal, and the encoded spatial / mono audio signal is decoded and rendered into a suitable audio signal output, and these audio signal outputs are passed to the headset 137 to output these spatial audio signals to the user.
[0081] Regarding Figure 3 , a high-level view of an example (IVAS) encoder 103 / 113 / 123 / 133 including various inputs (as non-exclusive examples, which are expected for a codec) is shown.
[0082] In some embodiments, the encoder 103 / 113 / 123 / 133 includes an audio (IVAS) input 301. The audio input 301 is configured to be able to receive one or more sets of spatial data (IVAS) streams from multiple sources, either locally or remotely. These sources can be local (e.g., more than one spatial capture device with a known spatial configuration at the encoder location) and / or multiple remote participants sending spatial IVAS streams. The audio input 301 is configured to pass the audio data stream to an object header creator 303 and a (IVAS) decoder 311 that is part of an IVAS data stream processor 313.
[0083] In some embodiments, the encoder 103 / 113 / 123 / 133 includes a scene controller 305, which is configured to control the processing of the received audio input 301.
[0084] For example, in some embodiments, the encoder 103 / 113 / 123 / 133 includes an object header creator 303. The object header creator 303 controlled by the scene controller 305 is configured to insert each data stream as an object into the "main" data stream. In some embodiments, the object header creator 305 can also be configured to add missing object parameters, such as distance and direction, based on a real spatial configuration or a virtual defined scene.
[0085] In some embodiments, the object header creator 303 is configured to determine whether any inserted data stream contains objects, optionally move those audio objects to directly become part of the "main" IVAS stream and update their metadata, or move the objects under any other IVAS objects. Additionally, the object header creator 303 is configured to update the object metadata to be correct for the entire spatial configuration.
[0086] In some embodiments, the encoders 103 / 113 / 123 / 133 include an IVAS data stream processor 313. The IVAS data stream processor 313 may include an (IVAS) decoder 311. The (IVAS) decoder 311 is configured to receive one or more sets of spatial audio data streams, decode the spatial audio signals, and pass them to the audio scene renderer 231.
[0087] The IVAS data stream processor 313 may include an audio scene renderer 231, which is configured to receive the audio signals and generate an audio scene rendering based on the decoded (IVAS) spatial audio signals. This audio scene rendering may constitute, for example, a downmix of the various inputs from the (IVAS) decoder 311. Further, the rendered audio scene audio signals may be passed to the encoder 315.
[0088] The IVAS data stream processor 313 may include an encoder 315, which receives the rendered spatial audio signals and encodes them. In other words, the IVAS data stream processor 313 is configured to decode all or at least some of the input data streams and generate a common spatial scene, for example, using IVAS MASA, IVAS HOA / FOA, or IVAS mono objects.
[0089] In some embodiments where there are multiple embedded objects, these embedded objects may subsequently be sent to recipients with available high-performance rendering. The remaining recipients only receive the pre-rendered spatial scene. Alternatively, a combination of at least one "IVAS stream object" and a pre-rendered "spatial scene IVAS stream object" may be used to reduce the bit rate.
[0090] Additionally, the encoder includes an audio object multiplexer 309, which is configured to combine these objects and output a combined object data stream.
[0091] Figure 4 The flowchart in [ ] further illustrates the operation of the encoder.
[0092] In Figure 4 At step 401 in [ ], an audio (IVAS) data stream is received.
[0093] Furthermore, inFigure 4 At step 411 therein, determine the spatial scene configuration and control.
[0094] As Figure 4 shown in step 403, based on the determined spatial scene configuration and control and the input audio data stream, create an object header for the audio data stream.
[0095] In addition, as Figure 4 shown in step 404, optionally, based on the determined spatial scene configuration and control and the input audio data stream, decode the data stream.
[0096] Furthermore, as Figure 4 shown in step 406, the decoded data stream can be rendered.
[0097] Furthermore, as Figure 4 shown in step 408, use a suitable (IVAS) encoder to encode the rendered audio scene.
[0098] Furthermore, as Figure 4 shown in step 409, these data streams can be multiplexed and output.
[0099] IVAS object stream metadata can use any suitable acoustic / spatial metadata. The following table provides an example.
[0100]
[0101] However, in some embodiments, other position information such as x-y-z or Cartesian coordinates can be used instead of azimuth-elevation-distance. For example, the following table can provide another configuration:
[0102] Field Bit Description Position x Audio object x-y-z position in 3D space
[0103] However, some minimum stream description metadata is also required to signal the (IVAS) object data stream configuration information. For example, the following format can be used to signal this information.
[0104]
[0105] In this embodiment, the "stream ID" parameter is used to uniquely identify each IVAS object stream within the current session. Thus, each original and mixed audio component (input stream) can be signaled. For example, the signaling allows identification of components in the system or on the user interface. The "stream type" parameter defines the meaning of each "audio object". In some embodiments, the audio object is thus not just an object-based audio input. Instead, the object data stream can be object-based audio (input), or it can be any IVAS scenario. For example, as shown in FIG. 5, three types of objects are shown.
[0106] For example, in Figure 5A , a simple traditional (mono) audio object 501 is shown. The audio object 501 is defined in accordance with a PCM audio signal portion 505 and an acoustic (spatial) metadata portion 503. It will be understood that additional metadata may be present.
[0107] Regarding Figure 5B , an encoded representation 507 of the same audio object as shown in Figure 5A is shown.
[0108] Figure 5C Shown is the same audio object as shown in Figure 5A and Figure 5B but processed according to some embodiments as discussed herein. The processed audio object is described as an object data stream 509, which is defined by a "stream type = 0" parameter 513. In other words, the object data stream 509 includes a data stream identifier that identifies it as an object-based audio IVAS object stream. Additionally, the object data stream 509 includes an object audio bitstream portion 515 (the encoded representation of the audio object) and a stream identifier 511 that uniquely identifies the object data stream.
[0109] Figure 5D Shown is another (IVAS) object data stream 517. The other object data stream 517 includes an identifier portion 521 with a "stream type = 1". In some embodiments, stream type = 0 corresponds to a "simple" object type, e.g., a mono signal. Additionally, in some embodiments, stream type = 1 corresponds to a potentially "complex" stream. For example, in this example, stream type = 1 corresponds to a complete IVAS stream, in which case it contains a MASA spatial stream. Since IVAS can contain one or more object streams, this allows nested objects. If the stream type = 0, then no other objects are known and the stream has a simple type (effectively a mono object).
[0110] Another object data stream 517 may also include an explicit stream description part 523, or the stream content may be determined by starting to decode the object stream. In this case, it is explicitly described as a MASA-based scenario (e.g., "stream description = MASA").
[0111] In addition, another object data stream 517 includes a MASA format bitstream part 525 (the encoded representation of the audio object) and a stream identifier 519 that uniquely identifies the object data stream "stream ID = 000002".
[0112] The first advantage of the method discussed herein is that IVAS inputs can often be conveniently forwarded without decoding / encoding operations. For example, in the case where there is a mixer device, a conference bridge (e.g., an AR / VR conference server), or other entities for combining and / or forwarding audio inputs in an IVAS end-to-end service, no decoding / encoding operations are required. Thus, by reassigning the received (encoded) input as an IVAS object stream, the operational complexity and latency are reduced. For example, if the playback capabilities of the receiver are unknown, the server can optimize the complexity by simply providing the received scenario as is. Any IVAS stream can be decoded and rendered as mono to support even the simplest IVAS devices. Skipping any decoding / encoding operations at an intermediate point (e.g., a conference server) also reduces the end-to-end latency for that audio component. Thus, the user experience is improved.
[0113] Furthermore, the embodiments are configured such that there is only a shallowly embedded "objectified" IVAS stream. In other words, if there is an object stream that also contains objects (and thus may include multiple object levels), deep data structures are avoided, and thus the decoder complexity is reduced. Thus, in some embodiments, the proposed embedding allows an IVAS object to include another IVAS object. In other words, although an IVAS object may be two or more levels deep, in some embodiments, any "deep" object can be moved into a "higher-level" object closer to the "main" IVAS stream, and its metadata can be updated such that its representation remains meaningful for the newly formed scenario. In some embodiments, an IVAS object can be moved to become part of another IVAS object. Thus, the object is moved "deeper". This can allow, for example, audio objects (e.g., mono objects) to be encoded or decoded together to save complexity or bitrate. If the same type of format is at different levels in the structure, they typically need to be encoded / decoded at different times or using different instances. This can introduce additional complexity.
[0114] In addition, embodiments as discussed herein can have a second advantage, namely, IVAS object streams can be conveniently nested, e.g., for content distribution purposes. In such embodiments, more complex scenarios can be treated as a single (mono) audio object. Figure 6 Example nested packetization is provided. This can be used, e.g., for distributing decoding complexity. This is very useful, e.g., for edge cloud services.
[0115] Thus, for example, Figure 6 The entire scene object data stream 601 is shown. The entire scene object data stream 601 includes a plurality of object data streams 602, 604, 606, and 608. For example, the first object data stream 602 includes a stream ID 621 (stream ID = 000001) that uniquely identifies the object data stream, a stream type identifier 623 (stream type = 0), and a data portion 625. The second object data stream 604 includes a stream ID 631 (stream ID = 000006) that uniquely identifies the object data stream, a stream type identifier 633 (stream type = 1), and a data portion 635. The third object data stream 606 includes a stream ID 641 (stream ID = 000007) that uniquely identifies the object data stream, a stream type identifier 643 (stream type = 1), and a data portion 645. The fourth object data stream 608 includes a stream ID 651 (stream ID = 000008) that uniquely identifies the object data stream, a stream type identifier 653 (stream type = 0), and a data portion 655.
[0116] In addition, as Figure 6 shown, the second object data stream 604 also includes nested object data streams 612 and 614. These nested object data streams can be, e.g., object data streams associated with sub - parts of the entire scene. The fifth object data stream 612 includes a stream ID 661 (stream ID = 000004) that uniquely identifies the object data stream, a stream type identifier 663 (stream type = 0), and a data portion 665. The sixth object data stream 614 includes a stream ID 671 (stream ID = 000005) that uniquely identifies the object data stream, a stream type identifier 673 (stream type = 1), and a data portion 675.
[0117] Additionally, the nested sixth object data stream 614 further includes further nested object data streams 622 and 624. These further nested object data streams can be, for example, object data streams associated with a sub-sub-section of a sub-section of the entire scene. The seventh object data stream 622 includes a stream ID 681 (stream ID = 000002) that uniquely identifies the object data stream, a stream type identifier 683 (stream type = 1), and a data portion 685. The eighth object data stream 624 includes a stream ID 691 (stream ID = 000003) that uniquely identifies the object data stream, a stream type identifier 693 (stream type = 1), and a data portion 695.
[0118] Another advantage of implementing some embodiments is that any IVAS input or IVAS scene that already includes spatial parameters (e.g., location characteristics) can determine such characteristics. For example, this can be achieved by adding acoustic spatial metadata (e.g., one of the parameters in the previous table) to the IVAS object stream ("stream type = 1"). This enables an enhanced experience, for example, in AR / VR conference call use cases.
[0119] For example, Figure 7 A captured scene 701 is shown, where there is a UE or similar capture device 707 that performs a spatial capture at a first (UE) location, and a second capture device 705 that performs a second spatial capture (or object capture) at a second (user) location.
[0120] In Figure 1 The conventional method shown in the upper right shows the audio object rendering 713 position and the first spatial capture scene 711. Thus, although the user can use a multi-microphone UE to capture a spatial scene (e.g., in the MASA format), and the user can use a close-range microphone or, for example, a second UE that can be connected to the "main" device to capture audio objects. These two inputs will be combined and provided to the IVAS encoder. In terms of the listening experience, a combined rendering of spatial audio (e.g., background audio) and audio objects (e.g., user speech) can then be listened to.
[0121] By implementing the embodiments described herein, a listener can switch 730 between a first option of audio object rendering 723 captured in a second space and a first space captured scene 721 or a second option of audio object rendering 733 captured in a first space and a second space captured scene 731. Thus, the IVAS codec can introduce a second space audio representation as an IVAS object stream. Accordingly, when a user uses their UE to capture a spatial audio scene, a wireless multi-microphone device or, in fact, a second UE connected to the "primary" UE can capture a complete spatial representation of the sound scene at a second location. This sound scene can now be encoded by the second device as an IVAS bitstream and provided to a second UE that can "act as a conference bridge", obtain the IVAS bitstream, and embed it as an IVAS object stream. In turn, it will be delivered to the listener two spatial audio scenes. For example, the user can switch between them such that the mono downmix of each scene is provided as the audio object rendering for the other scene for forward rendering to the user.
[0122] While Figure 6 an example of object stream nesting is shown, it should be understood that this is not the only mechanism for IVAS stream transport / packeting enabled by the present invention. FIG. 8 shows two examples of IVAS stream packeting according to some embodiments.
[0123] In some embodiments, a look-up table specifying the packet content can be used. The look-up table can be defined as a "payload header", and it can be, for example, an RTP payload header. This can include, for example, the sizes of various blocks, etc. Behind the header is the payload.
[0124] For example, as shown in FIG. 8, the data stream can include various IVAS object streams and IVAS content. Thus, the entire scene object stream 801 includes a payload header 811 or a look-up table that can specify the packet content. For example, as Figure 8A shown, a first object data stream 813 and a second object data stream 819 are specified, as well as payloads such as a first payload 815 (MASA and object) and a second payload 817 (5.1 channel audio data).
[0125] In some embodiments, as Figure 8C shown, the data stream can include only IVAS object streams. Thus, the entire scene object stream 831 includes a payload header or a look-up table that can specify the packet content including an object data stream 833, which in turn can include a nested object data stream 835, which in turn includes a further nested object data stream.
[0126] Figure 8BA "hybrid" embodiment is provided with a payload and a nested object data stream 813 across the scene.
[0127] There are associated nested costs in generating additional "payload header" information and its parsing.
[0128] Regarding decoders / renderers 105, 115, 125, 135. The decoders / renderers 105, 115, 125, 135 are configured to receive various (IVAS) object data streams and decode and render these data streams in parallel.
[0129] In some embodiments, the processing of the nested audio object data stream can be performed separately for each sub-scene level and then combined at a higher level.
[0130] For example, regarding Figure 6 the example shown in. Here, decoding can start from "stream ID = 000002" and "stream ID = 000003". Thus, we have decoded "stream ID = 000005" (since it is a container for the sub-scene). Further, the decoder can be configured to decode the next "stream ID = 000004". After that, other streams are decoded. This method can have advantages, for example, in terms of memory consumption, where some memory can be released between sub-scene levels and thus the total memory footprint is not defined by all combined streams.
[0131] In such an embodiment, rendering can be performed at the sub-scene level and summed in the rendering domain, or combined rendering can be performed at the end of decoding.
[0132] In some embodiments, the decoder is configured to start a separate decoder instance for each sub-scene. Thus, for each "stream type = 1", a separate IVAS decoder instance is initialized.
[0133] Regarding Figure 9 , an example electronic device that can be used as an analysis or synthesis device is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.
[0134] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program codes, such as the methods described herein.
[0135] In some embodiments, device 1400 includes a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage component. In some embodiments, the memory 1411 includes a program code portion for storing program code that can be implemented on the processor 1407. Additionally, in some embodiments, the memory 1411 may also include a stored data portion for storing data (e.g., data that has been processed or is to be processed according to the embodiments described herein). As needed, the implemented program code stored within the program code portion and the data stored within the stored data portion can be retrieved by the processor 1407 via the memory-processor coupling.
[0136] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, the user interface 1405 can be coupled to the processor 1407. In some embodiments, the processor 1707 can control the operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1705 can enable a user to input commands to the device 1400, for example, via a keypad. In some embodiments, the user interface 1405 can enable a user to obtain information from the device 1400. For example, the user interface 1405 can include a display configured to display information from the device 1400 to the user. In some embodiments, the user interface 1405 can include a touch screen or touch interface that can both enable information to be input into the device 1400 and display information to the user of the device 1400. In some embodiments, the user interface 1405 can be a user interface for communicating with a location determiner as described herein.
[0137] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, the input / output port 1409 includes a transceiver. In such embodiments, the transceiver can be coupled to the processor 1407 and is configured to communicate with other devices or electronic apparatuses, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver components can be configured to communicate with other electronic devices or apparatuses via a wired or wired coupling.
[0138] The transceiver can communicate with other devices through any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an Infrared Data Association (IrDA) communication path.
[0139] The transceiver input / output port 1409 can be configured to receive signals and, in some embodiments, determine parameters as described herein by using a processor 1407 that executes appropriate code. Additionally, the device can generate appropriate down-mixed signals and parameter outputs to send to a synthesis device.
[0140] In some embodiments, the device 1400 can be used as at least a part of a synthesis device. Thus, the input / output port 1409 can be configured to receive down-mixed signals and, in some embodiments, receive parameters determined at a capture device or a processing device as described herein, and generate appropriate audio signal format outputs by using a processor 1407 that executes appropriate code. The input / output port 1409 can be coupled to any suitable audio output, such as, for example, a multi-channel speaker system and / or a headset (which can be a head-tracked or non-tracked headset), etc.
[0141] In general, various embodiments of the present invention can be implemented using hardware or a special-purpose circuit, software, logic, or any combination thereof. For example, some aspects can be implemented using hardware, while other aspects can be implemented using firmware or software executable by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although the various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof.
[0142] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Additionally, in this regard, it should be noted that any block of the logical flow in the drawings can represent a program step, or an interconnected logical circuit, block, and function, or a combination of program steps and logical circuit, block, and function. The software can be stored on a physical medium such as a memory chip or a memory block implemented within a processor, on a magnetic medium such as a hard disk or a floppy disk, and on an optical medium such as a DVD and its data variants, a CD, etc.
[0143] The memory can be of any type suitable for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, by way of non-limiting example, can include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), gate-level circuitry, and processors based on multi-core processor architectures.
[0144] Embodiments of the present invention can be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Sophisticated and powerful software tools can be used to transform a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0145] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design of San Jose, California, use well-established design rules and a library of pre-stored design modules to automatically route conductors and position components on a semiconductor chip. Once the design of the semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transferred to a semiconductor manufacturing facility or “fab” for fabrication.
[0146] The foregoing description has provided a complete and beneficial description of exemplary embodiments of the present invention by way of example and not limitation. However, various modifications and adaptations will become apparent to those skilled in the relevant art in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the present invention will still fall within the scope of the present invention as defined by the appended claims.
Claims
1. An apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: Receiving at least a first audio data stream and a second audio data stream, wherein, At least one of the first audio data stream and the second audio data stream includes a spatial audio stream, the spatial audio stream being configured to enable immersive audio during communication; Determine the type of each audio data stream of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; Process the second audio data stream with at least one parameter according to the determined type of the second audio data stream; And Render the first audio data stream and the processed second audio data stream.
2. The device according to claim 1, wherein, The second audio data stream is configured to include at least one other audio data stream, and wherein the at least one other audio data stream includes the determined type and the at least one other audio data stream is an embedded-level audio data stream relative to the second audio data stream.
3. The device according to claim 2, wherein, The at least one other audio data stream includes at least one other embedded level, wherein each embedded level of the at least one other embedded level includes at least one additional audio data stream having the determined type.
4. The apparatus according to claim 1, wherein, The second audio data stream is a primary-level audio data stream.
5. The device according to claim 1, wherein Each of the first audio data stream and the second audio data stream is further associated with at least one of the following: A stream identifier configured to uniquely identify the audio data stream; And A stream descriptor configured to describe the type of the audio data stream.
6. The device according to claim 1, wherein The type of at least one of the first audio data stream and the second audio data stream is one of the following: Mono audio signal type; or Immersive voice and audio service audio signal.
7. The apparatus according to claim 1, wherein The at least one parameter is configured to define indoor characteristics or a scene description.
8. The apparatus according to claim 7, wherein, The at least one parameter defining the indoor characteristics or the scene description includes at least one of the following: Direction; Direction azimuth; Direction elevation; Distance; Gain; Spatial range; Energy ratio; and Position.
9. The apparatus according to claim 7, wherein The at least one parameter is generated based on spatial audio capture.
10. The device according to claim 1, wherein, The apparatus is further caused to: Receive an additional audio data stream; and Embed the additional audio data stream within one or the other of the first audio data stream and the second audio data stream.
11. A method for an apparatus for immersive audio communication, the method comprising: Receiving at least a first audio data stream and a second audio data stream, wherein at least one of the first audio data stream and the second audio data stream includes a spatial audio stream, the spatial audio stream being configured to enable immersive audio during communication; Determine the type of each audio data stream of the first audio data stream and the second audio data stream to identify which of the received first audio data stream and second audio data stream includes the spatial audio stream; Process the second audio data stream with at least one parameter according to the determined type of the second audio data stream; and Render the first audio data stream and the processed second audio data stream.
12. The method according to claim 11, wherein, The second audio data stream is configured to include at least one other audio data stream, and wherein the at least one other audio data stream includes the determined type, and the at least one other audio data stream is an embedded-level audio data stream relative to the second audio data stream.
13. The method according to claim 12, wherein, The at least one other audio data stream includes at least one other embedded level, wherein each embedded level in the at least one other embedded level includes at least one additional audio data stream having the determined type.
14. The method according to claim 11, wherein, The second audio data stream is a primary-level audio data stream.
15. The method according to claim 11, wherein Each of the first audio data stream and the second audio data stream is further associated with at least one of the following: A stream identifier, which is configured to uniquely identify the audio data stream; And A stream descriptor, which is configured to describe the type of the audio data stream.
16. The method according to claim 11, wherein, The type of at least one of the first audio data stream and the second audio data stream is one of the following: Mono audio signal type; or Immersive voice and audio service audio signal.
17. The method according to claim 11, wherein, The at least one parameter is configured to define indoor characteristics or a scene description.
18. The method according to claim 17, wherein, The at least one parameter defining the indoor characteristics or the scene description includes at least one of the following: Direction; Direction azimuth; Direction elevation; Distance; Gain; Spatial range; Energy ratio; and Position.
19. The method according to claim 11, further comprising: Receiving an additional audio data stream; And Embedding the additional audio data stream into one or the other of the first audio data stream and the second audio data stream.
20. The method according to claim 11, wherein Generate the at least one parameter based on spatial audio capture.
Citation Information
Patent Citations
Synchronizing enhanced audio transports with backward compatible audio transports
US20200013426A1