Masa with embedded near-far stereo for mobile devices
By using Metadata-Assisted Spatial Audio Format (MASA) and parametric analysis on mobile devices to generate encoded multi-channel audio signals, the problem of multi-format audio signal processing and spatial audio rendering in existing technologies is solved, and efficient, low-latency independent spatial audio presentation is achieved.
Patent Information
- Application Number
- CN202080055573.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-02
- Filing Date
- 2020-07-21
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2040-07-21
AI Technical Summary
Existing immersive audio codecs struggle to effectively process multi-format audio signals and perform spatial audio rendering on mobile devices, especially at low bit rates, and lack the ability to accurately describe the direction of sound sources and independently present ambient audio.
By receiving and generating encoded multichannel audio signals, using the Metadata-Assisted Spatial Audio Format (MASA), combined with parametric analysis and microphone signal processing, spatial representation of speech and object audio is achieved independently of ambient audio signals, and the encoding level and bit rate are adjusted according to the rendering device and transmission channel capabilities.
It enables efficient, low-latency multi-format audio signal processing and independent spatial rendering on mobile devices, adapting to different rendering devices and transmission conditions, and improving the accuracy of sound source direction description and the ability to render ambient audio independently.
Smart Images

Figure CN114207714B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to apparatuses and methods for spatial audio capture and associated rendering for mobile devices, but not exclusively to apparatuses and methods for Immersive Voice and Audio Services (IVAS) codec and Metadata Assisted Spatial Audio (MASA) with embedded near-far stereo for mobile devices. BACKGROUND
[0002] Immersive audio codecs are being implemented to support a large number of operating points ranging from low bit rate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed to be suitable for use over communication networks such as 3GPP 4G / 5G networks. Such immersive services include use in, for example, immersive voice and audio in applications such as immersive communications, virtual reality (VR), augmented reality (AR) and mixed reality (MR). The audio codec is expected to handle the encoding, decoding and rendering of speech, music and general audio. In addition it is also expected to support channel-based audio and scene-based audio input, including spatial information about the soundfield and sound sources. The codec is also expected to operate with low delay to enable conversational services and support high error robustness under a variety of transmission conditions.
[0003] Input signals are presented to the IVAS encoder in one of the supported number of formats (and in some allowed combinations of formats). Similarly, the decoder is expected to be able to output audio in a number of supported formats.
[0004] Some input formats of interest are Metadata Assisted Spatial Audio (MASA), object-based audio, in particular a combination of MASA and at least one object. Metadata Assisted Spatial Audio (MASA) is a parametric spatial audio format and representation. It can be considered as a representation consisting of “N channels + spatial metadata”. It is a scene-based audio format, particularly well suited for spatial audio capture on practical devices such as smartphones. The idea is to describe the sound scene in terms of time-varying and frequency-varying sound source directions. If no directional sound sources are detected, the audio is described as diffuse. The spatial metadata is described with respect to at least one direction indicated for each time-frequency (TF) tile, and can include, for example, spatial metadata for each direction and spatial metadata independent of the number of directions. SUMMARY
[0005] According to a first aspect, an apparatus is provided comprising components configured to perform the following operations: receiving at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, the at least one channel speech audio signal and the metadata being generated from at least one microphone audio signal; receiving at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and the metadata are generated based on parametric analysis of the at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel speech audio signal; and generating an encoded multichannel audio signal based on the at least one channel speech audio signal and the metadata, and further based on the at least one channel ambient audio signal and the metadata, such that the encoded multichannel audio signal enables the at least one channel speech audio signal to be spatially represented spatially independently of the at least one channel ambient audio signal.
[0006] The component can be further configured to receive at least one other audio object audio signal, wherein the component configured to generate an encoded multi-channel audio signal is configured to further generate an encoded multi-channel audio signal based on the at least one other audio object audio signal, such that the encoded multi-channel audio signal enables the at least one other audio object audio signal to be spatially presented independently of at least one channel of speech audio signal and at least one channel of ambient audio signal.
[0007] At least one microphone audio signal from which at least one channel of speech audio signal and metadata is generated, and at least one microphone audio signal from which at least one channel of ambient audio signal and metadata is generated, may include: a separate microphone group without a shared microphone; or a microphone group having at least one shared microphone.
[0008] The component can be further configured to receive an input configured to control the generation of encoded multichannel audio signals.
[0009] The component can be further configured to: modify the location parameters of the metadata associated with the at least one channel audio signal, or change the allocation of the near channel rendering channel associated with the at least one channel audio signal, based on the mismatch between the determined location parameters of the metadata associated with the at least one channel audio signal and the allocated near channel rendering channel.
[0010] A component configured to generate an encoded multi-channel audio signal based on at least one channel of speech audio signal and metadata, and further based on at least one channel of ambient audio signal and metadata, can be configured to: obtain an encoder bit rate; select embedded encoding levels and assign a bit rate to each selected embedded encoding level, wherein a first level is associated with the at least one channel of speech audio signal and metadata, a second level is associated with the at least one channel of ambient audio signal, and a third level is associated with the metadata associated with the at least one channel of ambient audio signal; and encode the at least one channel of speech audio signal and metadata, the at least one channel of ambient audio signal, and the metadata associated with the at least one channel of ambient audio signal based on the assigned bit rate.
[0011] The component can be further configured to: determine capability parameters based on at least one of the following: transmission channel capability; rendering device capability, wherein the component configured to generate encoded multi-channel audio signals can be further configured to: generate encoded multi-channel audio signals based on the capability parameters.
[0012] The component configured to further generate encoded multichannel audio signals based on capability parameters can be configured to: select an embedded coding level based on at least one of transmission channel capability and rendering device capability, and assign a bit rate to each selected embedded coding level.
[0013] At least one microphone audio signal used to generate at least one channel of environmental audio signal and metadata based on parametric analysis may include at least two microphone audio signals.
[0014] This component can be further configured to output encoded multi-channel audio signals.
[0015] According to a second aspect, an apparatus is provided comprising components configured to perform the following operations: receiving an embedded encoded audio signal, the embedded encoded audio signal comprising at least one of the following levels of embedded audio signals: at least one channel of speech audio signal and associated metadata to be rendered as a spatial speech scene; at least one channel of speech audio signal and associated metadata and at least one channel of ambient audio signal to be rendered as a near-far stereo scene; at least one channel of speech audio signal and associated metadata and at least one channel of ambient audio signal and associated spatial metadata to be rendered as a spatial audio scene; and decoding the embedded encoded audio signal and outputting a multi-channel audio signal representing the scene, such that the encoded multi-channel audio signal enables spatial representation of at least one channel of speech audio signal independent of at least one channel of ambient audio signal.
[0016] The embedded audio signal of the above level can further comprise at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered into a spatial audio scene, and at least one further audio object audio signal and associated metadata, and wherein the means configured to decode the embedded encoded audio signal and output a multi-channel audio signal representing the scene can be configured to decode and output the multi-channel audio signal such that the spatial rendering of the at least one further audio object audio signal is spatially independent from the at least one channel speech audio signal and the at least one channel ambient audio signal.
[0017] The means can be further configured to receive an input configured to control the decoding of the embedded encoded audio signal and the output of the multi-channel audio signal.
[0018] The input can comprise a switching of capabilities, wherein the means configured to decode the embedded encoded audio signal and output the multi-channel audio signal can be configured to update the decoding and output based on the switching of capabilities.
[0019] The switching of capabilities can comprise at least one of: a determination of an earbud / earphone configuration; a determination of a headphone configuration; and a determination of a loudspeaker output configuration.
[0020] The input can comprise a determination of a change of embedded level, wherein the means configured to decode the embedded encoded audio signal and output the multi-channel audio signal can be configured to update the decoding and output based on the change of embedded level.
[0021] The input can comprise a determination of a change of bitrate for an embedded level, wherein the means configured to decode the embedded encoded audio signal and output the multi-channel audio signal can be configured to update the decoding and output based on the change of bitrate for the embedded level.
[0022] The means can be configured to control the decoding of the embedded encoded audio signal and the output of the multi-channel audio signal to modify at least one channel speech audio signal position or change a near-channel rendering channel assignment associated with at least one channel speech audio signal based on a determined at least one speech audio signal detection position and / or a mismatch between assigned near-channel rendering channels.
[0023] The input can comprise a determination of a correlation between the at least one channel speech audio signal and the at least one channel ambient audio signal, and the component configured to decode and output the multi-channel audio signal is configured to: when the correlation is less than a determined threshold: control a position associated with the at least one channel speech audio signal, and control an ambient spatial scene formed by the at least one channel ambient audio signal by rotating the ambient spatial scene based on the at least one channel ambient audio signal according to the obtained rotation parameter, or compensating for a rotation of another device by applying a corresponding inverse rotation to the ambient spatial scene; and when the correlation is greater than or equal to the determined threshold: control a position associated with the at least one channel speech audio signal, and control an ambient spatial scene formed by the at least one channel ambient audio signal by rotating the ambient spatial scene based on the at least one channel ambient audio signal according to the obtained rotation parameter, or compensating for a rotation of another device by applying a corresponding inverse rotation to the ambient spatial scene while leaving the rest of the scene rotated.
[0024] According to a third aspect, a method is provided, comprising: receiving at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, the at least one channel speech audio signal and metadata being generated from at least one microphone audio signal; receiving at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and metadata are generated based on a parametric analysis of the at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel speech audio signal; and generating an encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata, such that the encoded multi-channel audio signal enables a spatial rendering of the at least one channel speech audio signal spatially independently of the at least one channel ambient audio signal.
[0025] The method can further comprise receiving at least one further audio object audio signal, wherein generating the encoded multi-channel audio signal comprises generating the encoded multi-channel audio signal further based on the at least one further audio object audio signal, such that the at least one further audio object audio signal is enabled to be spatially rendered spatially independently of the at least one channel speech audio signal and the at least one channel ambient audio signal.
[0026] The at least one microphone audio signal from which the at least one channel speech audio signal and metadata are generated and the at least one microphone audio signal from which the at least one channel ambient audio signal and metadata are generated can comprise separate microphone groups without a common microphone or microphone groups with at least one common microphone.
[0027] The method can further comprise receiving an input configured to control the generation of the encoded multi-channel audio signal.
[0028] The method can further comprise modifying a position parameter of metadata associated with the at least one channel speech audio signal or changing a near-channel rendering channel assignment associated with the at least one channel speech audio signal based on a mismatch between the determined position parameter of metadata associated with the at least one channel speech audio signal and the assigned near-channel rendering channel.
[0029] Generating the encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata can comprise obtaining an encoder bitrate, selecting embedded encoding levels and assigning a bitrate to each selected embedded encoding level, wherein a first level is associated with the at least one channel speech audio signal and metadata, a second level is associated with the at least one channel ambient audio signal and a third level is associated with metadata associated with the at least one channel ambient audio signal, encoding the at least one channel speech audio signal and metadata, the at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal based on the assigned bitrates.
[0030] The method can further comprise determining a capability parameter, the capability parameter being determined based on at least one of a transport channel capability and a rendering device capability, wherein generating the encoded multi-channel audio signal can comprise generating the encoded multi-channel audio signal further based on the capability parameter.
[0031] Generating the encoded multi-channel audio signal further based on the capability parameter can comprise selecting embedded encoding levels and assigning a bitrate to each selected embedded encoding level based on at least one of a transport channel capability and a rendering device capability.
[0032] The at least one microphone audio signal for generating the at least one channel ambient audio signal and metadata based on a parametric analysis can comprise at least two microphone audio signals.
[0033] The method can further comprise outputting the encoded multi-channel audio signal.
[0034] According to a fourth aspect, there is provided a method comprising: receiving an embedded encoded audio signal comprising at least one of the following levels of embedded audio signals: at least one channel speech audio signal and associated metadata to be rendered into a spatial voice scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal to be rendered into a near-far stereo scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered into a spatial audio scene; and decoding the embedded encoded audio signal and outputting a multi-channel audio signal representing the scene, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal independently of the at least one channel ambient audio signal.
[0035] The above-mentioned levels of embedded audio signals can further comprise: at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered into a spatial audio scene, and at least one other audio object audio signal and associated metadata, and wherein decoding the embedded encoded audio signal and outputting a multi-channel audio signal representing the scene can comprise decoding and outputting the multi-channel audio signal such that spatial rendering of the at least one other audio object audio signal is spatially independent of the at least one channel speech audio signal and the at least one channel ambient audio signal.
[0036] The method can further comprise receiving an input configured to control decoding of the embedded encoded audio signal and output of the multi-channel audio signal.
[0037] The input can comprise a switching of capabilities, wherein decoding the embedded encoded audio signal and outputting the multi-channel audio signal can comprise updating the decoding and output based on the switching of capabilities.
[0038] The switching of capabilities can comprise at least one of: a determination of an earbud / headphone configuration; a determination of a headphone configuration; and a determination of a loudspeaker output configuration.
[0039] The input can comprise a determination of a change of embedded level, wherein decoding the embedded encoded audio signal and outputting the multi-channel audio signal can comprise updating the decoding and output based on the change of embedded level.
[0040] The input can comprise a determination of a change of bit rate for an embedded level, wherein decoding the embedded encoded audio signal and outputting the multi-channel audio signal can comprise updating the decoding and output based on the change of bit rate for the embedded level.
[0041] The method can comprise controlling decoding of the embedded encoded audio signal and outputting of the multi-channel audio signal to modify at least one of a position of a channel speech audio signal based on the determined at least one of a speech audio signal detection position and / or a mismatch between assigned proximity channel rendering channels, or a proximity channel rendering channel assignment associated with the at least one channel speech audio signal.
[0042] The input can comprise a determination of a correlation between the at least one channel speech audio signal and the at least one channel ambient audio signal, and decoding and outputting the multi-channel audio signal can comprise: when the correlation is less than a determined threshold: controlling a position associated with the at least one channel speech audio signal, and controlling an ambient spatial scene formed by the at least one channel ambient audio signal by rotating the ambient spatial scene based on the at least one channel ambient audio signal according to the obtained rotation parameter, or compensating for a rotation of another device by applying a corresponding inverse rotation to the ambient spatial scene; and when the correlation is greater than or equal to the determined threshold: controlling a position associated with the at least one channel speech audio signal, and controlling an ambient spatial scene formed by the at least one channel ambient audio signal by rotating the ambient spatial scene based on the at least one channel ambient audio signal according to the obtained rotation parameter, or compensating for a rotation of another device by applying a corresponding inverse rotation to the ambient spatial scene while leaving the rest of the scene rotated.
[0043] According to a fifth aspect, there is provided an apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, the at least one channel speech audio signal and metadata being generated from at least one microphone audio signal; receive at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and metadata are generated based on a parametric analysis of the at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel speech audio signal; and generate an encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal spatially independently of the at least one channel ambient audio signal.
[0044] The apparatus can be further caused to receive at least one further audio object audio signal, wherein the apparatus caused to generate the encoded multi-channel audio signal is caused to generate the encoded multi-channel audio signal further based on the at least one further audio object audio signal such that the at least one further audio object audio signal can be spatially rendered independently from the at least one channel speech audio signal and the at least one channel ambient audio signal.
[0045] The at least one microphone audio signal from which the at least one channel speech audio signal and metadata is generated and the at least one microphone audio signal from which the at least one channel ambient audio signal and metadata is generated can comprise separate groups of microphones without a common microphone or a group of microphones with at least one common microphone.
[0046] The apparatus can be further caused to receive an input configured to control the generation of the encoded multi-channel audio signal.
[0047] The apparatus can be further caused to modify a position parameter of metadata associated with the at least one channel speech audio signal or change a near-channel rendering channel assignment associated with the at least one channel speech audio signal based on a determined mismatch between the position parameter of metadata associated with the at least one channel speech audio signal and the assigned near-channel rendering channel.
[0048] The apparatus caused to generate the encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata can be caused to obtain an encoder bit rate, select embedded encoding levels and assign a bit rate to each selected embedded encoding level, wherein a first level is associated with the at least one channel speech audio signal and metadata, a second level is associated with the at least one channel ambient audio signal and a third level is associated with metadata associated with the at least one channel ambient audio signal, encode the at least one channel speech audio signal and metadata, the at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal based on the assigned bit rates.
[0049] The apparatus can be further caused to determine a capability parameter, the capability parameter being determined based on at least one of: a transmission channel capability; a rendering apparatus capability, wherein the apparatus caused to generate the encoded multi-channel audio signal can be caused to generate the encoded multi-channel audio signal further based on the capability parameter.
[0050] The apparatus caused to generate the encoded multi-channel audio signal based on the capability parameter can be caused to: select an embedded encoding level based on at least one of the transport channel capability and the rendering device capability, and assign a bit rate to each selected embedded encoding level.
[0051] The at least one microphone audio signal used for generating the at least one channel ambient audio signal and the metadata based on the parametric analysis can comprise at least two microphone audio signals.
[0052] The apparatus can be further caused to: output the encoded multi-channel audio signal.
[0053] According to a sixth aspect, there is provided an apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive an embedded encoded audio signal, the embedded encoded audio signal comprising at least one of the following levels of embedded audio signals: at least one channel speech audio signal and associated metadata to be rendered into a spatial voice scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal to be rendered into a near-far stereo scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered into a spatial audio scene; and decode the embedded encoded audio signal and output a multi-channel audio signal representing the scene, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal independently of the at least one channel ambient audio signal.
[0054] The above-mentioned levels of embedded audio signals can further comprise: at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered into a spatial audio scene, and at least one other audio object audio signal and associated metadata, and wherein the apparatus caused to decode the embedded encoded audio signal and output a multi-channel audio signal representing the scene can be caused to: decode and output the multi-channel audio signal such that spatial rendering of the at least one other audio object audio signal is spatially independent of the at least one channel speech audio signal and the at least one channel ambient audio signal.
[0055] The apparatus can be further caused to: receive an input configured to control the decoding of the embedded encoded audio signal and the output of the multi-channel audio signal.
[0056] The input can comprise a determination of a change in embedded level, wherein the apparatus caused to decode an embedded encoded audio signal and output a multi-channel audio signal can be caused to update the decoding and output based on the change in embedded level.
[0057] The change in capability can comprise at least one of: a determination of an earbud / headphone configuration; a determination of a headphone configuration; and a determination of a speaker output configuration.
[0058] The input can comprise a determination of a change in embedded level, wherein the apparatus caused to decode an embedded encoded audio signal and output a multi-channel audio signal can be caused to update the decoding and output based on the change in embedded level.
[0059] The input can comprise a determination of a change in bit rate for an embedded level, wherein the apparatus caused to decode an embedded encoded audio signal and output a multi-channel audio signal can be caused to update the decoding and output based on the change in bit rate for the embedded level.
[0060] The apparatus can be caused to control the decoding of the embedded encoded audio signal and the output of the multi-channel audio signal to modify at least one channel speech audio signal position or change a near-channel rendering channel assignment associated with at least one channel speech audio signal based on the determined at least one speech audio signal detection position and / or a mismatch between assigned near-channel rendering channel channels.
[0061] The input can comprise a determination of a correlation between at least one channel speech audio signal and at least one channel ambient audio signal, and the apparatus caused to decode and output a multi-channel audio signal can be caused to: when the correlation is less than a determined threshold: control a position associated with the at least one channel speech audio signal, and control an ambient spatial scene formed by the at least one channel ambient audio signal by rotating the ambient spatial scene based on the at least one channel ambient audio signal according to an obtained rotation parameter or by applying a corresponding inverse rotation to the ambient spatial scene to compensate for a rotation of another device; and when the correlation is greater than or equal to the determined threshold: control a position associated with the at least one channel speech audio signal, and control an ambient spatial scene formed by the at least one channel ambient audio signal by rotating the ambient spatial scene based on the at least one channel ambient audio signal according to an obtained rotation parameter or by applying a corresponding inverse rotation to the ambient spatial scene to compensate for a rotation of another device while leaving the rest of the scene rotated.
[0062] According to a seventh aspect, there is provided an apparatus comprising: receiving circuitry configured to receive at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, the at least one channel speech audio signal and metadata being generated from at least one microphone audio signal; receiving circuitry configured to receive at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and metadata are generated based on a parametric analysis of the at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel speech audio signal; and encoding circuitry configured to generate an encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal spatially independently of the at least one channel ambient audio signal.
[0063] According to an eighth aspect, there is provided an apparatus comprising: receiving circuitry configured to receive an embedded encoded audio signal, the embedded encoded audio signal comprising at least one of the following levels of embedded audio signals: at least one channel speech audio signal and associated metadata to be rendered as a spatial speech scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal to be rendered as a near-far stereo scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered as a spatial audio scene; and decoding circuitry configured to decode the embedded encoded audio signal and output a multi-channel audio signal representing a scene, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal independently of the at least one channel ambient audio signal.
[0064] According to a ninth aspect, a computer program (or computer readable medium comprising program instructions) comprising instructions (or program instructions) for causing an apparatus to perform at least the following is provided: receiving at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, the at least one channel speech audio signal and metadata being generated from at least one microphone audio signal; receiving at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and metadata are generated based on a parametric analysis of the at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel speech audio signal; and generating an encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal spatially independent of the at least one channel ambient audio signal.
[0065] According to a tenth aspect, a computer program (or computer readable medium comprising program instructions) comprising instructions (or program instructions) for causing an apparatus to perform at least the following is provided: receiving an embedded encoded audio signal, the embedded encoded audio signal comprising at least one of the following levels of embedded audio signals: at least one channel speech audio signal and associated metadata to be rendered into a spatial speech scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal to be rendered into a near-far stereo scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered into a spatial audio scene; and decoding the embedded encoded audio signal and outputting a multi-channel audio signal representing the scene, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal independent of the at least one channel ambient audio signal.
[0066] According to an eleventh aspect, there is provided a non-transitory computer- readable medium comprising program instructions for causing an apparatus to perform at least the following: receiving at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, the at least one channel speech audio signal and metadata being generated from at least one microphone audio signal; receiving at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and metadata are generated based on a parametric analysis of the at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel speech audio signal; and generating an encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal spatially independent of the at least one channel ambient audio signal.
[0067] According to a twelfth aspect, there is provided a non-transitory computer- readable medium comprising program instructions for causing an apparatus to perform at least the following: receiving an embedded encoded audio signal comprising at least one of the following levels of embedded audio signals: at least one channel speech audio signal and associated metadata to be rendered as a spatial speech scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal to be rendered as a near-far stereo scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered as a spatial audio scene; and decoding the embedded encoded audio signal and outputting a multi-channel audio signal representing a scene, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal independent of the at least one channel ambient audio signal.
[0068] According to a thirteenth aspect, there is provided an apparatus comprising: means for receiving at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, wherein the at least one channel speech audio signal and metadata are generated from at least one microphone audio signal; means for receiving at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and metadata are generated based on a parametric analysis of the at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel speech audio signal; and means for generating an encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal spatially independent of the at least one channel ambient audio signal.
[0069] According to a fourteenth aspect, there is provided an apparatus comprising: means for receiving an embedded encoded audio signal, wherein the embedded encoded audio signal comprises at least one of the following levels of embedded audio signals: at least one channel speech audio signal and associated metadata to be rendered as a spatial speech scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal to be rendered as a near-far stereo scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered as a spatial audio scene; and means for decoding the embedded encoded audio signal and outputting a multi-channel audio signal representing a scene, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal independent of the at least one channel ambient audio signal.
[0070] According to a fifteenth aspect, there is provided a computer readable medium comprising program instructions causing an apparatus to perform at least the following: receiving at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, the at least one channel speech audio signal and metadata being generated from at least one microphone audio signal; receiving at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and metadata are generated based on a parametric analysis of the at least one microphone audio signal and the at least one channel ambient audio signal is associated with the at least one channel speech audio signal; and generating an encoded multi-channel audio signal based on the at least one channel speech audio signal and metadata and further based on the at least one channel ambient audio signal and metadata, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal spatially independent of the at least one channel ambient audio signal.
[0071] According to a sixteenth aspect, there is provided a computer readable medium comprising program instructions causing an apparatus to perform at least the following: receiving an embedded encoded audio signal, the embedded encoded audio signal comprising at least one of the following levels of embedded audio signals: at least one channel speech audio signal and associated metadata to be rendered as a spatial speech scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal to be rendered as a near-far stereo scene; at least one channel speech audio signal and associated metadata and at least one channel ambient audio signal and associated spatial metadata to be rendered as a spatial audio scene; and decoding the embedded encoded audio signal and outputting a multi-channel audio signal representing the scene, such that the encoded multi-channel audio signal enables spatial rendering of the at least one channel speech audio signal independent of the at least one channel ambient audio signal.
[0072] An apparatus comprising means for performing the actions of the methods as described above.
[0073] An apparatus configured to perform the actions of the methods as described above.
[0074] A computer program comprising program instructions for causing a computer to perform the methods as described above.
[0075] A computer program product stored on a medium can cause an apparatus to perform the methods described herein.
[0076] An electronic device can comprise an apparatus as described herein.
[0077] A chipset can comprise an apparatus as described herein.
[0078] Embodiments of the present application aim to address problems associated with the prior art. BRIEF DESCRIPTION OF DRAWINGS
[0079] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which:
[0080] Figure 1 and Figure 2 schematically illustrates a typical audio capture scenario that can be encountered when using a mobile device;
[0081] Figure 3a schematically illustrates example encoder and decoder architectures;
[0082] Figure 3b and Figure 3c schematically illustrates an example input device suitable for use in some embodiments;
[0083] Figure 4 schematically illustrates an example decoder / renderer apparatus suitable for receiving the output of Figure 3c
[0084] Figure 5 schematically illustrates example encoder and decoder architectures in accordance with some embodiments;
[0085] Figure 6 schematically illustrates an example input format in accordance with some embodiments;
[0086] Figure 7 schematically illustrates an example embedding scheme with three layers in accordance with some embodiments;
[0087] Figure 8a and Figure 8b schematically illustrates example input generator, codec generator, encoder, decoder and output device architectures suitable for implementing some embodiments;
[0088] Figure 9 shows a flow diagram of embedded format encoding in accordance with some embodiments;
[0089] Figure 10 shows a flow diagram of embedded format encoding and level selection and waveform and metadata encoding in accordance with some embodiments in more detail;
[0090] Figures 11a to 11d shows an example rendering / presentation apparatus in accordance with some embodiments;
[0091] Figure 12 shows an example immersive rendering of voice objects and MASA ambient audio for stereo headphone output in accordance with some embodiments;
[0092] Figure 13 An example change in rendering is shown according to some embodiments in the case of implementing bit switching of example near-far representations (wherein an embedded level is dropped);
[0093] Figure 14 An example of switching between mono and stereo capabilities is shown;
[0094] Figure 15 A flowchart of a rendering control method including presentation capability switching is shown according to some embodiments;
[0095] Figure 16 An example of rendering control during a change from immersive mode to embedded format mode for stereo output is shown according to some embodiments;
[0096] Figure 17 An example of rendering control during a change in embedded format mode caused by a capability change is shown according to some embodiments;
[0097] Figure 18 An example of rendering control during a change in embedded format mode caused by a capability change from mono to stereo is shown according to some embodiments;
[0098] Figure 19 An example of rendering control during an embedded format mode change from mono to immersive mode is shown according to some embodiments;
[0099] Figure 20 An example of rendering control during adaptation of lower spatial dimension audio based on user interaction in higher spatial dimension audio is shown according to some embodiments;
[0100] Figure 21 shows an example of user experience allowed by the method used in some embodiments;
[0101] Figure 22 An example of near-far stereo channel preference selection when answering a phone call is shown when implementing some embodiments;
[0102] Figure 23 An example device suitable for implementing the illustrated apparatus is shown. DETAILED DESCRIPTION
[0103] Suitable apparatus and possible mechanisms for spatial speech and audio environment input format definition and for the coding framework for IVAS are described in more detail below. In such embodiments, backward compatible transport and playback for stereo and mono representations can be provided with embedded structure and capabilities for corresponding rendering. In some embodiments, rendering capability switching enables decoders / renderers to allocate speech in the best channels, e.g., in a way that does not correspond to the preferences of the transmitter, without the need for blind downmixing in the rendering device. Furthermore, the input format definition and coding framework can be particularly suitable for real-world mobile device spatial audio capture, e.g., better allowing UE-on-ear immersive capture.
[0104] The concept is described with respect to the following embodiments configured to define an input format and metadata signaling with separate speech and spatial audio. In such embodiments, the complete audio scene is provided to a suitable encoder in (at least) two captured streams. The first stream is a mono speech object based on at least one microphone capture, and the second stream is a parametric spatial environment signal based on parametric analysis of signals from at least three microphones. In some embodiments, audio objects of additional sound sources can optionally be provided. The metadata signaling can include at least a speech priority indicator for the speech object stream (and optionally its spatial position). The first stream can be at least predominantly related to user speech and audio content close to the user’s mouth (near channel), and the second stream can be at least predominantly related to audio content far from the user’s mouth (far channel).
[0105] In some embodiments, the input format is generated so as to facilitate directional compensation of correlated and / or uncorrelated signals. For example, for uncorrelated first and second signals, the signals can be processed independently, while for correlated first and second signals, the parametric spatial audio (MASA) metadata can be modified according to the speech object position.
[0106] In some embodiments, spatial speech and audio encoding can be provided according to the defined input format. The encoding can be configured to allow separation of speech and spatial environment, where the speech object position is modified (if needed) prior to waveform and metadata encoding, or where updates to near-channel rendering channel assignments are applied based on a mismatch between the active part of the speech object real position and the assigned channel.
[0107] In some embodiments, rendering control and rendering can be provided based on changing (of the transmitted signal) level of immersion and rendering device capabilities. For example, in some embodiments, audio signal rendering characteristics are modified according to switched capabilities communicated to the decoder / renderer. In some embodiments, audio signal rendering characteristics and channel allocation can be modified according to changes in the level of embedding received over the transport. Furthermore, in some embodiments, audio signals can be rendered according to rendering characteristics and channel allocation to one or more output channels. In other words, near and far signals (which are a 2-channel or stereo representation different from the traditional stereo representation using left and right stereo channels) are allocated to a left and right channel stereo rendering according to some predetermined information. Similarly, near and far signals can have channel allocation or downmix information indicating how a mono presentation should be performed.
[0108] Furthermore, in some embodiments, spatial positions of mono voice signals (near signals) can be used in spatial IVAS audio streams according to UE settings (e.g. selected by the user via a UI or provided by an immersive voice service) and / or spatial analysis. In such embodiments, MASA signals (spatial environment, far signals) are configured to automatically contain spatial information obtained during MASA spatial analysis.
[0109] Thus, in some embodiments, audio signals can be decoded and rendered according to transport bitrates and rendering capabilities to provide one of the following to a receiving user or listener:
[0110] 1. Mono voice object.
[0111] 2. Stereo voice (as mono) + environment (as another mono) according to near-far dual channel configuration.
[0112] 3. Full spatial audio. Providing correct spatial placement for transmitted streams, where mono voice objects are rendered at object positions and spatial environment consists of both directional components and diffuse sound field.
[0113] In some embodiments, additional backward interoperability of transmitted streams can be achieved with EVS encoding of mono voice / audio waveforms.
[0114] Mobile handsets or devices, also referred to herein as UEs (User Equipment), represent the largest segment of the market for immersive audio services, with annual sales exceeding 1 billion. In terms of spatial playback, these devices can be connected to stereo headphones (preferably with head tracking). In terms of spatial capture, the UE itself can be considered the preferred device. To increase the popularity of multi-microphone spatial audio capture on the market and to provide an immersive audio experience to as many users as possible, it is therefore an aspect that has been considered in detail to optimize capture and codec performance for immersive communication.
[0115] For example, Figure 1 A typical audio capture scenario 100 on a mobile device is shown. In this example, there is a first user 104 without a headset. Thus, the user 104 is holding the UE 102 to their ear for a phone call. The user can call another user 106 equipped with stereo headphones and thus capable of using the headphones to hear the spatial audio captured by the first user. Based on the spatial audio capture by the first user, an immersive experience can be provided for the second user. However, considering, for example, regular MASA capture and encoding, the device can have problems with the capturing user's ear. For example, the user's speech can dominate the spatial capture, thereby reducing the level of immersion. Furthermore, all head rotations of the user and rotations of the device relative to the user's head cause the audio scene to rotate for the receiving user. In some cases, this can cause confusion for the user. Thus, for example, at time 121, the captured spatial audio scene has a first orientation, as shown by the rendered sound scene from the experience of the other user 106, which shows the first user 110 at a first position relative to the other user 106 and one of the more than one audio sources 108 in the scene at a position directly in front of the other user 106. In turn, when the user turns their head at time 123 and further at time 125, the captured spatial scene in turn rotates, as shown by the rotation of the position of the audio source 108 relative to the other user 106 and the audio position of the first user 110. This can, for example, lead to an overlap 112 of the audio sources. Furthermore, if the sensor information of the user device is used to compensate for the rotation (as shown by the arrow 118 with respect to time 125), the position of the audio source of the first user (the speech of the first user) rotates in the opposite direction, thereby making the experience worse.
[0116] Figure 2 Another audio capture scenario 200 using a mobile device is shown. The user 204 can, for example, start from the UE 202 operating in a UE-ear-capture mode (as shown on the left side 221), in turn change to a hands-free mode (which can, for example, be a handheld hands-free as shown in the center 223 of Figure 2 or a wearable hands-free as shown on the right side 225 of Figure 2the UE / device 202 shown on the right 225 is placed on a table). Another user / listener can also wear earbuds or headphones to present the captured audio signals. In this case, the listener / another user can walk around in a handheld, hands-free mode, or move around, for example, around the device placed on the table (in a hands-free capture mode). In this example use case, the rotation of the device relative to the user speech position and the overall immersive audio scene is much more complex than in the case of Figure 1
[0117] In some embodiments, various device rotations can be at least partially compensated in the capture-side spatial analysis based on any suitable sensors such as gyroscopes, magnetometers, accelerometers, and / or orientation sensors. Alternatively, in some embodiments, if rotation compensation angle information is sent as side-information (or metadata), the capture device rotation can be compensated at the playback side. Embodiments as described herein show means and methods to enable compensation of capture device rotation such that compensation or audio modification can be applied separately to speech and background sound spatial audio (in other words, for example, only the "near" or "far" audio signals can be compensated).
[0118] One key use case for IVAS is the automatic stabilization of an audio scene where an audio call is established between two participants (especially for actual audio-only mobile capture), with the participant with immersive audio capture capability starting the call, for example, with their UE on-ear, but switching to handheld hands-free and finally to the device on a table hands-free, and the spatial sound scene being sent and rendered to the other party in a way that does not result in the listener experiencing incorrect rotation of "near" or "far" audio signals.
[0119] Such an example can be where user Bob comes home from work. He walks through a park and suddenly remembers that he needs to discuss plans for the weekend. He places an immersive audio call to another user, his friend Peter, also conveying the nice environment of birds singing on the trees around him. Bob does not have a headset, so he places his smartphone to his ear to better hear Peter's voice. On his way home, he stops at a crossroad, looks left and right, and then again to the left to safely cross the street. As soon as Bob gets home, he switches to handheld hands-free operation and finally places the smartphone on the table to continue the call through the loudspeaker. The spatial capture also provides Peter with the directionality of Bob's cuckoo clock series of sounds. A stable immersive audio scene is provided to Peter regardless of the orientation and mode of operation of the capture device.
[0120] Accordingly, these embodiments attempt to optimize the capture for real mobile device use cases so that the speech dominance of the spatial audio representation is minimized and both directional speech and spatial background remain stable during rotation of the capture device.
[0121] In addition, these embodiments attempt to enable backward interoperability with 3GPP EVS, which IVAS is an extension of, and provide mono compatibility with EVS. In addition, these embodiments also allow for stereo and spatial compatibility. This is particularly applicable in spatial audio for UE close-talker use cases (e.g., as shown in Figure 1
[0122] Further, these embodiments attempt to provide the best audio rendering to the listener based on the various rendering capabilities available to the user on their device. In particular, some embodiments attempt to enable low complexity switching / downmixing from spatial audio calls to lower spatial dimensions. In some embodiments, this can be implemented in immersive audio voice calls from UEs with various rendering or rendering methods.
[0123] Figures 3a to 3c and Figure 4 An example capture and rendering apparatus is shown which provides a suitable relevant background that can describe these embodiments.
[0124] For example, Figure 3a A multi-channel capture and binaural audio rendering apparatus 300 is shown. In this example, a spatial multi-channel audio signal (2 or more channels) 301 is provided to a noise reduction and signal separation processor 303 which is configured to perform noise reduction and signal separation. For example, this can be performed by generating a noise reduced primary mono signal 306 (containing primarily near field speech) extracted from the original mono / stereo / multi-microphone signal. This uses general mono or multi-microphone noise suppression techniques used by many current devices. This stream is sent as is to the receiver as a backward compatible stream. In addition, the noise reduction and signal separation processor 303 is configured to generate an ambient signal 305. These can be obtained by removing the mono signal (speech) from all microphone signals, resulting in as many ambient streams as there are microphones present. According to this approach, therefore, there is one or more spatial signals available for further encoding.
[0125] In turn, the mono 306 and ambient 305 audio signals can be processed by a spatial image normalization processor 307 which performs suitable spatial image processing in order to produce a binaural compatible output 309 comprising left ear 310 and right ear 311 outputs.
[0126] Figure 3b A first example input generator (and encoder) 320 is shown. The input generator 320 is configured to receive spatial microphone 321 (multi-channel) audio signals, which are provided to a noise reduction and ambient extraction processor 323, which is configured to perform extraction of a mono-channel speech signal 325 (containing mainly near-field speech) from the multi-microphone signals and pass it to a legacy speech codec 327. In addition, ambient signals 326 are generated (based on the microphone audio signals from which the speech signal is extracted), which can be passed to a suitable stereo / multi-channel audio codec 328.
[0127] The legacy speech codec 327 encodes the mono-channel speech audio signal 325 and outputs a suitable speech codec bitstream 329, and the stereo / multi-channel audio codec 328 encodes the ambient audio signals 326 to output a suitable stereo / multi-channel ambient bitstream 330.
[0128] Figure 3c A second example input generator (and encoder) 340 is shown. The second input generator (and encoder) 340 is configured to receive spatial microphone 341 (multi-channel) audio signals, which are provided to a noise reduction and ambient extraction processor 343, which is configured to perform extraction of a mono-channel speech signal 345 (containing mainly near-field speech) from the multi-microphone signals and pass it to a legacy speech codec 347. In addition, ambient signals 346 are generated (based on the microphone audio signals from which the speech signal is extracted), which can be passed to a suitable parametric stereo / multi-channel audio processor 348.
[0129] The legacy speech codec 347 encodes the mono-channel speech audio signal 345 and outputs a suitable speech codec bitstream 349.
[0130] The parametric stereo / multi-channel audio processor 348 is configured to generate a mono-channel ambient audio signal 351 (which can be passed to a suitable audio codec 352) and a spatial parameter bitstream 350. The audio codec 352 receives the mono-channel ambient audio signal 251 and encodes it to generate a mono-channel ambient bitstream 353.
[0131] Figure 4 An example decoder / renderer 400 is shown, which is configured to process the bitstreams from Figure 3c The decoder / renderer 400 thus comprises a speech / ambient audio decoder and spatial audio renderer 401, which is configured to receive the mono-channel ambient bitstream 353, the spatial parameter bitstream 350 and the speech codec bitstream 349, and generate a suitable multi-channel audio output 403.
[0132] Figure 5 A high level overview of an IVAS encoder / decoder architecture suitable for implementing embodiments discussed below is shown. The system can include a range of possible input types, including a mono audio signal input type 501, a stereo and binaural audio signal input type 502, a MASA input type 503, an Ambisonics input type 504, a channel-based audio signal input type 505, and an audio object input type 506.
[0133] In addition, the system includes an IVAS encoder 511. The IVAS encoder 511 can include an enhanced voice services (EVS) encoder 513, which can be configured to receive the mono input format 501 and provide at least a portion of a bitstream 521 for at least some of the input types.
[0134] In addition, the IVAS encoder 511 can include a stereo and spatial encoder 515. The stereo and spatial encoder 515 can be configured to receive signals from any input from the stereo and binaural audio signal input type 502, the MASA input type 503, the Ambisonics input type 504, the channel-based audio signal input type 505, and the audio object input type 506 and provide at least a portion of the bitstream 521. In some embodiments, the EVS encoder 513 can be used to encode mono audio signals obtained from the input types 502, 503, 504, 505, 506.
[0135] In addition, the IVAS encoder 511 can include a metadata quantizer 517 configured to receive side information / metadata associated with input types such as the MASA input type 503, the Ambisonics input type 504, the channel-based audio signal input type 505, and the audio object input type 506 and quantize / encode them to provide at least a portion of the bitstream 521.
[0136] In addition, the system includes an IVAS decoder 531. The IVAS decoder 531 can include an enhanced voice services (EVS) decoder 533, which can be configured to receive the bitstream 521 and generate appropriate decoded mono signals for output or further processing.
[0137] In addition, the IVAS decoder 531 can include a stereo and spatial decoder 535. The stereo and spatial decoder 535 can be configured to receive the bitstream and decode them to generate appropriate output signals.
[0138] In addition, the IVAS decoder 531 can include a metadata dequantizer 537 configured to receive the bitstream 521 and regenerate metadata that can be used to assist spatial audio signal processing. For example, based on the combination of the stereo and spatial decoder 535 and the metadata dequantizer 537, at least some spatial audio can be generated in the decoder.
[0139] In the following examples, the primary input types of interest are the MASA 503 and the object 506. The embodiments described below feature a codec input that can be a MASA + at least one object input, where the object is particularly for providing user speech. In some embodiments, the MASA input can or can not be related to the user speech object. In some embodiments, signaling related to this related state can be provided as IVAS input (e.g., input metadata).
[0140] In some embodiments, the MASA 503 input is provided to the IVAS encoder as a mono or stereo audio signal and metadata. However, in some embodiments, the input can instead consist of 3 (e.g., planar first order Ambisonic - FOA) or 4 (e.g., FOA) channels. In some embodiments, the encoder is configured to encode an Ambisonics input as MASA (e.g., via modified DirAC encoding), to encode a channel-based input (e.g., 5.1 or 7.1 + 4) as MASA, or to encode one or more object tracks as MASA or as a modified MASA representation. In some embodiments, object-based audio can be defined as a mono audio signal with at least associated metadata.
[0141] Embodiments as described herein can be flexible in terms of accurate audio object input definition for user speech objects. In some embodiments, a specific metadata flag defines the user speech object as the main signal (speech) for communication. For certain input signals and codec configurations, like user generated content (UGC), this signaling can be ignored or treated differently from the main conversation mode.
[0142] In some embodiments, the UE is configured to implement spatial audio capture that provides not only a spatial signal (e.g., MASA) but also a dual component signal, where user speech is handled separately. In some embodiments, the user speech is represented by a mono object.
[0143] Thus, in some embodiments, the UE is configured to provide a combination such as “MASA + object” at the IVAS encoder 511 input. This is for example in Figure 6An example is shown in Fig. 6, where a user 601 and a UE 603 are configured to capture / provide mono speech object 605 and MASA environment 607 input to a suitable IVAS encoder. In other words, the UE spatial audio capture provides not only a spatial signal (e.g. MASA), but also a dual component signal, where the user speech is processed separately.
[0144] Thus, in some embodiments, the input is a mono speech object 609 captured using at least one microphone close to the mouth of the user. The at least one microphone can be a microphone on the UE, or, e.g. a headset microphone or so-called lavalier microphone, from which the audio stream is provided to the UE.
[0145] In some embodiments, the mono speech object has an associated speech priority signaling flag / metadata 621. The mono speech object 609 can comprise a mono waveform and metadata comprising at least a spatial position of the sound source (i.e. the user speech). The position can be a real position, e.g. a position relative to the position of the capturing device (UE), or a virtual position based on some other setting / input. In practice, the speech object can utilize other ways of utilizing the same / similar metadata as is generally known in the art for object-based audio.
[0146] The following table summarizes the minimum characteristics of a speech object according to some embodiments.
[0147] Characteristic Description (value) Audio waveform Single mono thin profile (intended for user speech) Metadata: Speech priority Indicate user speech signal with (most) high priority Metadata: Location Object location (e.g. x-y-z or azimuth / elevation / (distance))
[0148] The following table provides some signaling options that can be additionally or alternatively used for the regular object audio position of a speech object according to some embodiments.
[0149]
[0150]
[0151] In some embodiments, different signaling (metadata) can be implemented for the speech object rendering channel, where there is either mono-only or near-far stereo transmission.
[0152] For audio objects, the object placement in the scene can be arbitrary. When the scene (consisting of at least one object) is binauralized for rendering, the object position can change, e.g. over time. Thus, time-varying position information can also be provided in the speech object metadata (as shown in the table above).
[0153] The mono voice object input can be considered the “near” signal, which can always be rendered according to its signaled position in immersive rendering, or alternatively downmixed to a fixed position in downmix domain rendering. By “near” is denoted the signal capture spatial position / distance relative to the captured voice source. According to embodiments, this “near” signal is always provided to the user and can always be heard in the rendering, regardless of the exact rendering configuration and bitrate. For this purpose, voice priority metadata or equivalent signaling is provided (as shown in the table above). In some embodiments, this stream can be the default mono signal from an immersive IVAS UE, even in the absence of any MASA spatial input.
[0154] Accordingly, the spatial audio of immersive communication from an immersive UE (smartphone) is represented as two parts. The first part can be defined as the voice signal in the form of a mono voice object 609.
[0155] The second part (ambient part 623) can be defined as the spatial MASA signal (including MASA channels 611 and MASA metadata 613). In some embodiments, the spatial MASA signal includes at least substantially no or only weakly correlates with the track of the user voice object. For example, the mono voice object can be captured using a lavalier microphone or with strong beamforming.
[0156] In some embodiments, additional acoustic correlation (or separation) information for the voice object and ambient signal can be signaled. This metadata provides information about how much acoustic “leakage” or crosstalk exists between the voice object and the ambient signal. In particular, this information can be used to control a directional compensation process, as explained below.
[0157] In some embodiments, processing of the spatial MASA signal is implemented. This can be according to the following steps:
[0158] 1. If there is no or low correlation (correlation < threshold) between the voice object and the MASA ambient waveform, then
[0159] a. independently control the position of the voice object
[0160] b. control the MASA spatial scene by:
[0161] i. let the MASA spatial scene rotate according to the actual rotation, or
[0162] ii. compensate for device rotation by applying a corresponding negative rotation to the MASA spatial scene direction on a per-TF-tile basis.
[0163] 2. If there is a correlation between the voice object and the MASA environment waveforms (correlation >= threshold), then
[0164] a. Control the position of the voice object
[0165] b. Control the MASA spatial scene by:
[0166] i. Compensate for the device rotation of the TF-tile corresponding to the user's voice (TF-tile and direction) by applying the rotation used in "a." on a TF-tile-per-TF-tile basis to the MASA spatial scene (at least when VAD = 1 for the voice object), while leaving the rest of the scene to rotate according to the actual capture device rotation, or
[0167] ii. Leave the MASA spatial scene to rotate according to the actual rotation, while at least diffusing the TF-tile corresponding to the user's voice (TF-tile and direction) (at least when VAD = 1 for the voice object), where the amount of directional-to-diffuse modification can depend on the confidence value associated with the MASA TF-tile corresponding to the user's voice.
[0168] From the above, it can be understood that the correlation computation can be performed based on a long-term average (i.e., not based on correlation values computed for individual frames to make the decision), and that voice activity detection (VAD) can be used to identify the presence of the user's voice within the voice object channel / signal. In some embodiments, the correlation computation is based on encoder processing of at least two signals. In some embodiments, metadata signaling is provided, e.g., in addition to the signal correlation computation, which can be based at least in part on capture device specific information.
[0169] In some embodiments, the voice object position control can involve a (pre-) determined voice position or stability of the voice object in the spatial scene. In other words, related to the unwanted rotation compensation.
[0170] In some embodiments, the consideration of confidence values as discussed above can be a weighted function of angular distance (between directions) and signal correlation.
[0171] It is observed herein that when voice is separated from the spatial environment, if suitable signaling is implemented, rotation compensation (e.g., in UE close-talker spatial audio capture use cases) can be simplified, either done locally in capture or in rendering. The representations as discussed herein can also enable arbitrary placement of mono objects, which need not rely on scene rotation. Thus, in some embodiments, a single audio waveform and MASA metadata can be used to convey the environment, where device rotation can already be compensated in a way that there is no perceptible or annoying mismatch between voice and environment, even if they are related.
[0172] The MASA input audio can be, for example, mono-based and stereo-based audio signals. Thus, in addition to the MASA spatial metadata, there can also be mono waveforms or stereo waveforms. In some embodiments, either type of input and conveyance can be implemented. However, in some embodiments, mono-based MASA input for the environment along with the user voice object can be the preferred format.
[0173] In some embodiments, there can also optionally be other objects 625 represented as object 615 audio signals. These objects are typically related to some other audio component of the overall scene, rather than the user voice. For example, the user can add a virtual loudspeaker to the transmitted scene to play a music signal. Such an audio element would typically be provided to the encoder as an audio object.
[0174] Since the user voice is the primary communication signal, for conversational operation, a large portion of the available bit rate can be allocated to the voice object coding. For example, at 48 kbps, it can be considered that approximately 20 kbps can be allocated to encode the voice, while the remaining 28 kbps is allocated to the spatial environment representation. At such bit rates, and especially at lower bit rates, it can be beneficial to encode the mono representation of the spatial MASA waveform to achieve the highest possible quality. For these reasons, in some examples, mono-based MASA input can be the most practical.
[0175] Another consideration is embedded coding. In A suitable embedded coding scheme proposal is provided in Figure 7 where mono-based MASA (encoding) is practical. Embodiments of embedded coding can enable three levels of embedded coding.
[0176] The lowest level, as indicated by the Mono: voice column, is mono operation, which includes a mono voice object 701. This can also be designated as a "near" signal.
[0177] The middle level, as indicated by the Mono: voice + Mono: environment column, is mono operation, which includes a mono voice object 701 and a mono environment object 703. This can also be designated as a "far" signal. Figure 7Stereo: near-far column in the table. In such embodiments, the input is both the mono speech object 701 and the ambient MASA channel 703. This can be implemented as a "near-far" stereo configuration. However, the difference is that the "near" signal in these embodiments is the complete mono speech object, whose parameters are for example as shown in the table above.
[0178] In these embodiments, the previously defined approach is extended to handle immersive audio rather than just stereo and various levels of rendering. The "far" channel is the mono portion of the MASA representation based on the mono speech object. Thus, it includes the complete spatial environment, but without the actual method to render it correctly in space. What is rendered in the case of stereo delivery will also depend on the spatial rendering setup and capabilities. The following table provides some additional characteristics of the "far" channel that can be used for rendering according to some embodiments.
[0179]
[0180]
[0181] The third and highest level of the embedded structure as indicated by the Spatial audio column is the spatial MASA environment representation, which includes both the mono speech object 701 and the spatial MASA environment representation, which includes both the ambient MASA channel 703 and the ambient MASA spatial metadata 705. In these embodiments, the spatial information is provided in a way that the environment in space can be correctly rendered to the listener.
[0182] Additionally, as indicated by the Spatial audio + objects column in the table, in some embodiments, additional objects 707 can be included in case of a combined input to the encoder. Note that these additional individual objects 707 are not assumed to be part of the embedded encoding, but rather they are handled individually. Figure 7
[0183] In some embodiments, there can be priority signaling at the codec input (e.g., input metadata) that indicates whether for example a particular object 707 is more or less important than the ambient audio. Typically, such information is based on user input (e.g., via a UI) or service settings.
[0184] For example, there can be priority signaling that results in a lower embedded mode to include separate objects on that side for transmission before entering the next embedded level.
[0185] In other words, in some cases (input setup) and operating points, lower embedded modes and optional objects can be conveyed, e.g. mono speech object + separate audio objects, before considering switching to near-far stereo transmission mode.
[0186] Figure 8a An example apparatus for implementing some embodiments is presented. Figure 8a A UE 801 is shown, for example. The UE 801 comprises at least one microphone for capturing speech 803 of a user and is configured to provide a mono speech audio signal to a mono speech object (near) input 811. In other words, a dedicated microphone setup is used to capture the mono speech object. It can be a single microphone or a plurality of microphones, for example, which are used to perform suitable beamforming, for example.
[0187] In some embodiments, the UE 801 comprises a spatial capture microphone array 805 for the environment, which is configured to capture and pass on an environmental component to a MASA (far) input 813.
[0188] In some embodiments, the mono speech object (near) 811 input is configured to receive a microphone’s speech audio signal and pass on a mono audio signal as a mono speech object to an IVAS encoder 821. In some embodiments, the mono speech object (near) 811 input is configured to process the audio signal (e.g. optimize the mono speech object’s audio signal before passing it on to the IVAS encoder 821).
[0189] The MASA input 813 is configured to receive an environmental audio signal of a spatial capture microphone array and pass it on to the IVAS encoder 821. In some embodiments, a separate spatial capture microphone array is used to obtain a spatial environmental signal (MASA) and it is processed according to any suitable means to improve the quality of the captured audio signal.
[0190] In turn, the IVAS encoder 821 is configured to encode the two input format audio signals based on them, as shown by bitstream 831.
[0191] Furthermore, an IVAS decoder 841 is configured to decode the encoded audio signals and pass them on to a mono speech output 851, a near-far stereo output 853, and a spatial audio output 855.
[0192] Figure 8b Another example apparatus for implementing some embodiments is presented. Figure 8bFor example, a UE 851 is shown, which comprises a combined spatial audio capture multi-microphone arrangement 853, which is configured to provide a mono-voice audio signal to a mono-voice object (near) input 861 and to provide a MASA audio signal to a MASA input 863. In other words, Figure 8b A combined spatial audio capture is shown, which outputs a mono for voice and spatial signals for the environment. The combined analysis processing can for example suppress the user voice from the spatial capture. This can be done for the individual channels or the MASA waveform (“far” channel). It can be considered that this audio capture looks very similar to Figure 8a but at least some common processing, e.g. removing the beamformed microphone signals from the spatial capture signals.
[0193] In some embodiments, the mono-voice object (near) 861 input is configured to receive a microphone’s voice audio signal and pass a mono audio signal as a mono-voice object to the IVAS encoder 871. In some embodiments, the mono-voice object (near) 861 input is configured to process the audio signal (e.g. optimize the mono-voice object’s audio signal before passing it to the IVAS encoder 871).
[0194] The MASA input 863 is configured to receive an environment audio signal of a spatial capture microphone array and pass it to the IVAS encoder 871. In some embodiments, a separate spatial capture microphone array is used to obtain a spatial environment signal (MASA) and it is processed according to any suitable means to improve the quality of the captured audio signal.
[0195] In turn, the IVAS encoder 871 is configured to encode the two input format audio signals based on them as shown by the bitstream 881.
[0196] Furthermore, the IVAS decoder 891 is configured to decode the encoded audio signals and pass them to a mono-voice output 893, a near-far stereo output 895, and a spatial audio output 897.
[0197] In some embodiments, in addition to the possibility to suppress (remove) user speech from the individual channels or MASA waveform, other spatial audio processing can be applied to optimize the mono speech object + MASA spatial audio input. For example, during active speech, the direction can not be considered to correspond to the main microphone direction or to increase the diffuse value uniformly. For example, a full spatial analysis can be performed when the local VAD does not activate the speech microphone. In some embodiments, such additional processing can be utilized, e.g., only when the UE is on the ear or in handheld hands-free operation and the microphone used for the speech signal is close to the user’s mouth.
[0198] For multi-microphone IVAS UEs, the default capture mode can be the one that utilizes the mono speech object + MASA spatial audio input format.
[0199] In some embodiments, the UE determines the capture mode based on other (non-audio) sensor information. For example, there can be known methods to detect that the UE is in contact with the user’s ear or is substantially located near the user’s ear. In this case, the spatial audio capture can enter the mode as described above. In other embodiments, the mode can depend on some other mode selection (e.g., the user can provide input using a suitable UI to select whether the device is in a hands-free mode, a handheld hands-free mode, or a handheld mode).
[0200] Regarding Figure 9 the operation of the IVAS encoder 813, 863 is described in more detail.
[0201] For example, in some embodiments, as Figure 9 indicated at step 901, the IVAS encoder 813, 863 can be configured to obtain the negotiated settings (e.g., the mode as described above) and to initialize the encoder.
[0202] The next operation can be to obtain the input. For example, as Figure 9 indicated at step 903, the mono speech object + MASA spatial audio is obtained.
[0203] Further, in some embodiments, as Figure 9 indicated at step 905, the encoder is configured to obtain the current encoder bitrate.
[0204] In turn, as Figure 9 indicated at step 907, a codec mode request can be obtained. Such a request relates to a receiver request, e.g., for a specific encoding mode to be used by the sending device.
[0205] In turn, as Figure 9 indicated at step 909, an embedded encoding level can be selected and a bitrate can be allocated.
[0206] Furthermore, such as Figure 9 As shown in step 911, waveforms and metadata can be encoded.
[0207] In addition, such as Figure 9 As shown, the encoder's output is passed as a bitstream 913, which is then decoded by the IVAS decoder 915, for example, in an embedded output format. For instance, in some embodiments, the output format may be a mono audio object 917. This mono audio object may further be a mono audio object located near channel 919 (and another ambient far channel 921). Furthermore, the output of the IVAS decoder 915 may be a mono audio object 923, an ambient MASA channel 925, and ambient MASA channel spatial metadata 927.
[0208] also, Figure 10 The selection of the embedded coding level is shown in more detail. Figure 9 Step 909) and encoding the waveform and metadata ( Figure 9 Step 911) is the operation. The initial operation is as follows: Figure 10 Step 1001 shows the acquisition of the mono audio signal of the speech object, such as Figure 10 Step 1003 shows the acquisition of the MASA spatial environment audio signal, and as shown in step 1003. Figure 10 The total bit rate is obtained as shown in step 1005.
[0209] like Figure 10 As shown in step 1007, after obtaining the mono audio signal of the speech object, the method may further include comparing the speech object position with the near-channel rendering channel allocation.
[0210] In addition, such as Figure 10 As shown in step 1011, after obtaining the mono audio signal of the speech object and the MASA spatial environment audio signal, the method may further include determining the activity level of the input signal and pre-allocating the bit budget.
[0211] like Figure 10 As shown in step 1013, after comparing the speech object positions, determining the input signal activity level, and obtaining the total bit rate, the method may further include estimating the need for a handover and determining a speech object position modification.
[0212] like Figure 10As shown in step 1009, after estimating the need for switching and determining the modification of the speech object position, the method may include modifying the speech object position or updating the near channel rendering channel assignment when necessary. Specifically, for embedding mode switching, modification of the mono object position is considered to smooth out any potential discontinuities. Alternatively, the recommended L / R assignment for the near channel (as received at the encoder input) may be updated. This update may also include updating the far channel L / R assignment (or, for example, the mix balance). The potential need for such modification is based on the possibility of significant deviations between near-far stereo rendering channel preferences and the speech object position. Such deviations are possible because the speech object position can be understood as a continuous, time-varying location in space. On the other hand, near-far stereo channel selection (entering the L or R channel / ear presentation) is typically static, or updated only, for example, during certain pauses, to avoid creating unnecessary and annoying discontinuities (i.e., content jumping between channels). Therefore, the signal activity levels of both component waveforms are tracked.
[0213] In addition, such as Figure 10 As shown in step 1015, the method may include determining the embedded level to be used. This may be based on an understanding of the total bit rate (and, for example, a negotiated bit rate range), estimating the embedded level (mono, stereo, spatial, etc.) to be used for switching. Figure 7 and Figure 9 (as shown in the image) potential needs.
[0214] After that, such as Figure 10 As shown in step 1017, bit rates can then be allocated for speech and environment.
[0215] like Figure 10 As shown in step 1019, after the bit rate has been allocated and the speech object position has been modified or the near-channel rendering channel allocation has been updated as needed, the method may further include performing waveform and metadata encoding based on the allocated bit rate.
[0216] Furthermore, such as Figure 10 As shown in step 1021, the encoded bit stream can be output.
[0217] In some embodiments, the encoder is configured to be EVS-compatible. In this embodiment, the IVAS codec can encode the user's voice object in an EVS-compatible encoding mode (e.g., EVS 16.4kbps). By stripping away any IVAS voice object metadata and spatial audio components and decoding only EVS-compatible mono audio, compatibility with legacy EVS devices becomes very straightforward. Furthermore, this corresponds to an end-to-end experience from EVS UE to EVS UE, despite the use of an IVAS UE (with immersive capture).
[0218] Figures 11a to 11d Some typical rendering / presentation use cases for immersive conversational services using IVAS are shown, which can be used according to some embodiments. Although in some embodiments the apparatus can be configured to implement the conference application discussed in the following, the apparatus and method implement a (immersive) voice call, in other words a call between two people. Thus, for example, a user can be configured to use the UE for audio rendering as shown in Figure 11a In some embodiments where this is a legacy UE, typically a mono EVS encoding (or, for example, AMR-WB) has been negotiated. However, if this UE is an IVAS UE, it can be configured to receive an immersive bitstream. In such embodiments, typically mono audio signal playback is implemented, and thus the user can only play mono speech objects (and indeed only send mono speech objects). In some embodiments, the user can be provided with an option to select the level of output embedding. For example, the UE can be configured with a user interface (UI) that enables the user to control the level of environmental reproduction. Thus, in some embodiments, the user can also be provided with and presented with a monaural mix of the near and far channels. In some embodiments, a default mix balance value is received based on the far channel characteristics table as shown above. In some embodiments, for example, a mono speech default distance gain can be provided for the additional / alternative speech object characteristics table also shown above.
[0219] In some embodiments, the rendering apparatus comprises earbuds (such as shown by the wireless left channel earbud 1113 and right channel earbud 1111 in Figure 11b or a headset / headset microphone (such as shown by the wireless headset 1121 in Figure 11c In the case of these embodiments, the user can be presented with a stereo or full spatial audio signal. Thus, in some embodiments and depending on the bit rate, the spatial audio can comprise a mono speech object + MASA spatial audio representation (as provided at the encoder input) or just MASA (where the speaker’s speech has been downmixed into MASA) in various embodiments. According to some embodiments, the user can receive a mono speech object, near-far stereo, or a mono speech object + MASA spatial audio.
[0220] In some embodiments, the rendering apparatus comprises stereo or multi-channel loudspeakers (such as shown by the stereo loudspeakers 1123 in Figure 11dThe left channel speaker 1133 and right channel speaker 1131 are shown in the diagram, and the listener or receiving user is configured to listen to the call using the (stereo) speakers. This could be, for example, pure stereo playback. Alternatively, the speaker arrangement could be a multi-channel arrangement, such as 5.1 or 7.1+4 or any suitable configuration. Furthermore, spatial speaker presentation could be synthesized or implemented by a suitable soundbar playback device. In some embodiments where the rendering apparatus includes stereo speakers, any spatial component received by the user equipment can be ignored. Instead, playback could be a near-far stereo format. It is understood that near-far stereo can be configured to provide discrete playback of both signals, or the ability to shift at least one of these signals (e.g., according to metadata provided based on the table shown above). Therefore, in some embodiments, at least one of the playback channels could be a mixture of these two discrete (near and far) transmission channels.
[0221] Figure 12 An example of immersive rendering of “voice object + MASA environment” audio for stereo headphones / earpods is shown (visualized here as a user wearing an earpod with left channel 1113 and right channel 1111). The capture scene is shown with several orientations: front 1215, left 1211, and right 1213, and the scene includes a capture device 1201, a mono object 1203 located between the front and left, and an ambient sound source. Figure 12 In the example shown, source 1 1205 is located in front of capture device 1201, source 2 1209 is located to the right of capture device 1201, and source 3 1207 is located between the left and rear of capture device 1201. Renderer 1211 (the listener) is able to generate and render a facsimile of the capture scene using embodiments as described herein. For example, renderer 1211 is configured to generate and render a mono object 1213 located between the front and left sides, and ambient sound source 1 1215 located in front of renderer 1211, sound source 2 1219 located to the right of renderer 1211, and source 3 1217 located between the left and rear of renderer 1211. In this way, an immersive audio scene is presented to the receiving user based on the input representation. Furthermore, the speech stream can be controlled independently of the environment; for example, the speaker's volume level can be increased, or the speaker's position can be manipulated on a suitable UE UI.
[0222] Figure 13The rendering changes under bitrate switching (where the embedding level decreases) are shown according to the example near-far representation. Here, the presented speech object position 1313 is matched with the near-far rendering channel information (i.e., the user's speech is rendered on the default left channel corresponding to the general speech object position relative to the listening position). Furthermore, the environment is mixed into the right channel 1315. This can typically be achieved using methods such as... Figure 10 The system shown is used to implement this.
[0223] Figure 14 The presentation capability switching use case is illustrated. While many other examples are possible, the earpod use case is foreseeable as a general capability switch for currently available devices. The earpod form factor and size are becoming increasingly popular and are likely highly relevant to IVAS audio presentation in both user-generated content (UGC) and conversational use cases. Furthermore, head-tracking capabilities implemented in this form factor can be expected to become available in the coming years. Therefore, this form factor is a natural candidate for devices to be used across all embedded immersion levels (mono, stereo, spatial).
[0224] In this example, there is a user 1401 at the receiving end of an immersive IVAS call. User 1401 has a UE 1403 in their hand, and this UE is used for audio capture (which can be immersive audio capture). To render the incoming audio, UE 1403 is connected to smart earpods 1413, 1411, which can be operated individually or together. Therefore, the wireless earpods 1413, 1411 act as mono or stereo playback devices depending on user preferences and behavior. For example, in some embodiments, earpods 1413, 1411 may be characterized by automatically detecting whether they are placed in the user's ears. Figure 14 The image shows a user wearing a single Earpod 1411 on the left side. However, the user can add a second Earpod 1413, and... Figure 14 The right-hand illustration shows a user wearing the two Earpods, 1413 and 1411. Therefore, a user can switch back and forth between dual-channel stereo (or immersive) playback and single-channel mono playback, for example, during a call. In some embodiments, an immersive dialogue codec renderer is able to handle this use case by providing a consistently good user experience. Otherwise, the user would be distracted by incorrect rendering of the incoming communication audio and might, for example, completely lose the transmitted user voice.
[0225] about Figure 15 A flowchart illustrates a suitable method for controlling rendering based on rendering capability switching in the decoder / renderer. In some embodiments, rendering control may be implemented as part of at least one of an internal renderer and an external renderer (IVAS).
[0226] Therefore, in some embodiments, such as Figure 15 As shown in step 1501, the method includes receiving a bit stream input.
[0227] like Figure 15 As shown in step 1503, after the bitstream has been received, the method may further include obtaining the decoded audio signal and metadata, as well as determining the embedded level of the transmitted data.
[0228] In addition, such as Figure 15 As shown in step 1505, the method may further include receiving appropriate user interface input.
[0229] like Figure 15 As shown in step 1507, after obtaining appropriate user interface input, the decoded audio signal and metadata, and determining the embedded level to be sent, the method may further include obtaining presentation capability information.
[0230] like Figure 15 As shown in step 1509, the next operation is to determine whether there is a capability switch.
[0231] like Figure 15 As shown in step 1511, if the switching of capabilities is determined, the method may include updating the audio signal rendering characteristics according to the switched capabilities.
[0232] like Figure 15 As shown in step 1513, if there is no switch or update to follow, the method may include determining whether there is an embedded-level change.
[0233] like Figure 15 As shown in step 1515, if there is an embedded level change, the method may include updating the audio signal rendering characteristics and channel allocation according to the embedded level change.
[0234] like Figure 15 As shown in step 1517, if there is no change in the embedded level or an update to the audio signal rendering characteristics and channel allocation, the method may include rendering the audio signal based on the rendering characteristics (including the sent metadata) and the channel allocation for one or more output channels.
[0235] Therefore, this rendering can lead to, as by Figure 15 The presentation of the mono signal shown in step 1523, as described by... Figure 15 The presentation of the stereo signal shown in step 1521, and as by Figure 15 The presentation of the immersive signal shown in step 1519.
[0236] Thus, modifications to the speech object can be applied and rendering of the ambient signal under rendering capability switching and embedding level changes can be implemented.
[0237] Thus, for example with respect to Figure 16 rendering control during a change in the default embedding level from an immersive embedding level to a stereo embedding level or directly to a mono embedding level is shown. For example, network congestion causes a reduction in bit rate and the encoder changes the embedding level accordingly. In the example on the left-hand side, immersive rendering of “speech object + MASA ambient” audio is used for stereo headphones / earbuds / earpods presentation (visualized here as a user wearing left channel 1113 and right channel 1111 earpods). The presented scene is shown to include a renderer 1601, a mono object 1603 located between front and left, and ambient sound sources, source 1 1607 located in front of the renderer device 1601, source 2 1609 located to the right of the renderer device 1601, source 3 1605 located to the left rear of the renderer device 1601.
[0238] As shown by arrow 1611, the embedding level is reduced from immersive to stereo, which results in rendering the mono speech object 1621 through the left channel and the ambient sound sources located to the right of the renderer device 1625.
[0239] As shown by arrow 1613, the embedding level is also reduced from immersive to mono, which results in rendering the mono speech object 1631 through the left channel. In other words, when the user is listening to a stereo presentation, the stereo and mono signals are not binauralized. Instead, the signaling is considered and the presentation side of the talker speech is selected (the speech object preferred channel according to the encoder signaling).
[0240] In some embodiments where the “smart” presentation device is able to signal its current capabilities / usage to the (IVAS internal / external) renderer, the renderer can be able to the capabilities for mono or stereo presentation and can mono presentation on that channel (or which ear). If this is unknown or cannot be determined, it is up to the user to ensure their earpods / headphones / earphones etc. are placed correctly, otherwise the user can receive an incorrectly rotated immersive scene (in case of spatial presentation) or be provided with a mostly diffuse ambient presentation (depending on the renderer).
[0241] This can be shown, for example with respect to Figure 17 Figure 17 rendering control during a capability switch from a two-channel presentation to a single-channel presentation is shown. Thus, it can be understood here that the decoder / renderer is configured to receive an indication of the capability change. This can be, for example, a Figure 15 the input step 1505 shown in the flowchart of FIG. 15.
[0242] For example, Figure 17 showing a similar scenario as shown in Figure 16 the same as shown in FIG. 17, but where the user removes or turns off 1715 the left channel earpod, resulting in only the right channel earpod 1111 being worn. Thus, as shown on the left hand side, the rendered scene is shown, including the renderer 1601, the mono object 1603 located between the front and left, and the ambient sound sources, source 1 1607 located in front of the renderer device 1601, source 2 1609 located to the right of the renderer device 1601, and source 3 1605 located to the left rear of the renderer device 1601.
[0243] As shown by arrow 1711, the embedded level is lowered from immersive to stereo, which results in the mono speech object 1725 being rendered through the right channel and the ambient sound source 1721 also being located to the right of the renderer device 1725.
[0244] As shown by arrow 1713, the embedded level is also shown as being lowered from immersive to mono, which results in the mono speech object 1733 being rendered through the right channel.
[0245] In this embodiment, an immersive signal can be received (or there is a substantially simultaneous embedded level change to stereo or mono, which can be caused, for example, by a codec mode request (CMR) to the encoder based on a change in receiver presentation device capability). Thus, the audio is routed to the available channel in the renderer. Note that this is a renderer control, it is not just downmixing the two channels in the presentation device, if only the presentation capability switches and the embedded level does not change, this would be a direct downmix of the immersive signal visible on the left hand side of FIG. 18. Thus, the user experience is improved, the sound is clearer. Figure 17
[0246] With respect to Figure 18 , an example of an embedded level change from mono to stereo is shown. The left hand side shows the location of the mono speech object 1803 relative to the renderer, to the right of the renderer 1801. When the mono to near-far stereo capability is changed (based on encoder signaling or associated user preference at the decoder), the speech is panned from the right to the left, as shown by the contour 1833 (right contour) to 1831 (left contour), which results in the object renderer outputting the speech object (near) 1823 to the left and the environment (far) 1825 to the right.
[0247] With respect to Figure 19 This illustrates an example of an embedded level change from mono to immersive. The left-hand side shows the position of the mono speech object 1803 relative to the renderer, to the right of the renderer 1801. When the mono-to-immersive capability is changed (e.g., a higher bit rate is assigned), the mono speech object is smoothly moved 1928 to its correct position 1924 in the spatial scene. This is possible because it is transmitted separately. Therefore, the position metadata (in the decoder / renderer) is modified to achieve this transition. If there is no externalization / binauralization initially in the mono rendering, the speech object is first externalized from 1913 to 1923 as shown. Additionally, ambient audio sources 1922, 1925, and 1927 are rendered in their correct positions.
[0248] about Figure 20 This illustrates an example rendering adaptation of audio in a lower spatial dimension based on previous user interactions with the receiving user (in a higher spatial dimension). It illustrates the capture of a spatial scene and its default rendering in near-far stereo mode.
[0249] therefore, Figure 20 The capture and immersive rendering of audio for "voice object 2031 + MASA environment (MASA channel 2033 and MASA metadata 2035)" is illustrated. The capture scene is shown with several orientations: front 2013, left 2003, and right 2015, and the scene includes a capture device 2001, a mono object 2009 located between the front and left sides, and an ambient sound source. Figure 20 In the example shown, source 1 2013 is located in front of capture device 2001, source 2 2015 is located to the right of capture device 2001, and source 3 2011 is located between the left and rear of capture device 1201. Renderer 2041 (listener) is capable of generating and rendering a copy of the captured scene using embodiments as described herein. For example, renderer 2041 is configured to generate and render a mono object 2049 located between the front and left sides, and an ambient sound source 12043 located in front of renderer 2041, a sound source 2 2045 located to the right of renderer 2041, and a source 33042 located between the left and rear of renderer 1201.
[0250] Moreover, the receiving user can freely manipulate at least some aspects of the scene (according to any signaling that can limit the user’s freedom to do so). For example, as indicated by arrow 2050, the user can move the speech object to their preferred location. This can trigger a re-mapping of the local preference for the speech channel in the application. Thus, the Tenderer 2051 (the listener) is able to generate and render a modified facsimile of the captured scene by using embodiments as described herein. For example, the Tenderer 2051 is configured to generate and render the mono speech object 2059 between front and left, while keeping the ambient sound sources 2042, 2043, 2045 in their original orientation.
[0251] Moreover, when the embedded level changes, e.g. due to network congestion / reduced bitrate 2060, the mono speech object 2063 is kept by the rendering device 2061 on the channel it was previously mainly heard by the user, while the mono ambient audio object 2069 is kept on the other side. (Note that in binauralized rendering, the user would hear sound from both channels. The application of HRTFs and the direction of arrival of the speech object give its location in the virtual scene. In the case of non-binauralized rendering of the stereo signal, the speech can alternatively appear from only a single channel). This is different from the default situation based on the capture and transmission, as shown on the right upper side, where the mono speech object 2019 is positioned by the rendering device 2021 on the original capture side, while the mono ambient audio object 2023 is positioned on the opposite side.
[0252] With respect to Fig. 21, a comparison of three user experiences according to embodiments is shown. These user experiences are all controlled by encoder or decoder / rendering side signaling. Three states of the same scene (e.g. the scene previously described with respect to Fig. 20) are rendered, where the states are transitioned, and where differences are implemented in the rendering control. In the top Figure 12 Figure 21a of Fig. 21, a rendering according to encoder side signaling is shown. Thus, the Tenderer 2101 is initially configured with a mono speech source 2103 on the left and a mono ambient audio source 2102 on the right. A stereo to immersive transition 2104 results in the Tenderer 2105 having a mono speech source 2109 and ambient audio sources 2106, 2107, 2108 in the correct positions. A further immersive to mono transition results in the Tenderer 2131 having a mono speech source 2133 on the left.
[0253] In the middle Figure 21b of Fig. 21, a rendering according to a combination of encoder side and decoder / rendering side signaling is shown. Here, the user has a preference to have the speech object in the R channel (e.g. the right ear). Figure 21a The rendering is shown according to the encoder side signaling. Thus, the renderer 2111 is initially configured with a mono speech source 2113 on the right and a mono ambient audio source 2112 on the left. The stereo to immersive transition 2104 results in the renderer 2115 having the mono speech source 2119 in its correct position and the ambient audio sources 2116, 2117, 2118. The further immersive to mono transition results in the renderer 2141 having a mono speech source 2143 on the right.
[0254] Finally, the bottom shows how the user preference is applied to the same immersive scene. This can be the case, for example, where there is a first transition from near-far stereo transmission (and its rendering according to the user preference) to immersive scene transmission. The speech object position is adapted to maintain the user preference. Thus, the renderer 2121 is initially configured with a mono speech source 2123 on the right and a mono ambient audio source 2122 on the left. The stereo to immersive transition 2104 results in the renderer 2125 having the mono speech source 2129 moved to a new position and the ambient audio sources 2126, 2127, 2108 in the correct position. The further immersive to mono transition results in the renderer 2151 having a mono speech source 2153 on the left. Figure 21c
[0255] Figure 22 The channel preference selection (provided as decoder / renderer input) based on the receiving user's earpod activation is shown with respect to
[0256] Thus, for example, the top shows that the user 2201 has an incoming call 2203 (where the transmission uses, for example, near-far stereo configuration), the user adds the right channel earpod 2205, as shown by arrow 2207, the call is answered and speech is rendered on the right channel. Further, as shown by reference 2210, the user can in turn add the left channel earpod 2209, which in turn causes the renderer to add the ambient on the left channel.
[0257] The bottom shows that the user 2221 has an incoming call 2223 (where the transmission uses, for example, near-far stereo configuration), the user adds the left channel earpod 2225, as shown by arrow 2227, the call is answered and speech is rendered on the left channel. Further, as shown by reference 2230, the user can in turn add the right channel earpod 2229, which in turn causes the renderer to add the ambient on the right channel.
[0258] It can be appreciated that in many conversational immersive audio use cases, the receiving user does not know what is the correct scene. It is therefore important to provide a high quality and consistent experience, where at least the most important signal is always conveyed to the user. Generally, for conversational use cases, the most important signal is the speaker's voice. In this document, the voice signal presentation is maintained during capability switching. If needed, the voice signal is thus switched from one channel to another or from one direction to the rest of the channels. Based on the presentation capabilities and signaling, the environment is automatically added or removed.
[0259] Thus, according to embodiments herein, two possibly completely independent audio streams are transmitted in an embedded spatial stereo configuration, where the first stereo channel is a mono voice and the second stereo channel is the basis for a spatial audio environment scene. Thus, for rendering, it is important to understand the intended or expected spatial meaning / positioning of the two channels at least in terms of L-R channel placement. In other words, it is generally needed to know which near-far channel is L and which near-far channel is R. Alternatively, as explained in embodiments, information can be provided about the desired way to mix them together to render at least one channel. For any backward compatible playback, the channels can always be played back as L and R, regardless of how (although the choice can be arbitrary, e.g. by designating the first channel as L and the second channel as R).
[0260] Thus, in some embodiments, it can be decided based on mono / stereo capability signaling how to present any of the following received signals in the actual rendering system. These can include:
[0261] - only mono voice objects (or mono-only signals stripped of metadata)
[0262] - near-far stereo with mono objects (or mono voice stripped of metadata) and mono environment waveforms
[0263] - spatial audio scene with mono voice objects transmitted separately
[0264] In the case of mono-only playback, the default presentation can be simple:
[0265] - mono objects are rendered in the available channels
[0266] - near components (mono voice objects) are rendered in the available channels
[0267] - mono voice objects transmitted separately are rendered in the available channels
[0268] In the case of stereo playback, the default presentation is suggested as follows:
[0269] - render mono speech objects in the preferred channel, or binauralize mono objects according to the signaled direction
[0270] - render near components (mono speech objects) in the preferred channel, and far components (mono ambient waveforms) in the second channel, or binauralize near components (mono speech objects) according to the signaled direction, and binauralize far components (mono ambient waveforms) according to a default or user preferred way
[0271] o binauralization of ambient signals can be fully diffuse by default, or it can depend on the near channel direction or some previous state
[0272] - binauralize mono speech objects that are transmitted separately according to the signaled direction, and binauralize the spatial ambience according to the MASA spatial metadata description
[0273] With respect to Figure 23 , an example electronic device that can be used as an analysis or synthesis device is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 2400 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc.
[0274] In some embodiments, the device 2400 includes at least one processor or central processing unit 2407. The processor 2407 can be configured to execute various program codes, such as the methods described herein.
[0275] In some embodiments, the device 2400 includes a memory 2411. In some embodiments, the at least one processor 2407 is coupled to the memory 2411. The memory 2411 can be any suitable storage component. In some embodiments, the memory 2411 includes a program code portion for storing program codes that can be implemented on the processor 2407. Further, in some embodiments, the memory 2411 can also include a storage data portion for storing data, such as data that has been processed or will be processed according to the embodiments described herein. When needed, implemented program codes stored within the program code portion and data stored within the storage data portion can be fetched by the processor 2407 via the memory-processor coupling.
[0276] In some embodiments, the interface device 2400 includes a user interface 2405. In some embodiments, the user interface 2405 can be coupled to the processor 2407. In some embodiments, the processor 2407 can control the operation of the user interface 2405 and receive input from the user interface 2405. In some embodiments, the user interface 2405 can enable a user to enter commands to the device 2400, for example via a keypad. In some embodiments, the user interface 2405 can enable the user to obtain information from the device 1700. For example, the user interface 2405 can include a display configured to display information from the device 2400 to a user. In some embodiments, the user interface 2405 can include a touchscreen or touch interface that is capable of both enabling information to be entered into the device 2400 and displaying information to a user of the device 2400. In some embodiments, the user interface 2405 can be a user interface as described herein.
[0277] In some embodiments, the device 2400 includes an input / output port 2409. In some embodiments, the input / output port 2409 includes a transceiver. In such embodiments, the transceiver can be coupled to the processor 2407 and configured to enable communication with other apparatuses or electronic devices, for example via a wireless communication network. In some embodiments, the transceiver, or any suitable transceiver or transmitter and / or receiver components, can be configured to communicate with other electronic devices or apparatuses via a wired or wired coupling.
[0278] The transceiver can communicate with other apparatuses by any suitable known communication protocol. For example, in some embodiments, the transceiver can use suitable Universal Mobile Telecommunications System (UMTS) protocols, wireless local area network (WLAN) protocols such as IEEE 802.X, suitable short-range radio frequency communication protocols such as Bluetooth, or infrared data communication paths (IRDA).
[0279] The input / output port 2409 can be coupled to any suitable audio output, for example to a multi-channel speaker system and / or headphones (which can be head-tracked or non-tracked headphones), etc.
[0280] In general, the various embodiments of the application can be implemented in hardware or special-purpose circuits, software, logic or any combination thereof. For example, some aspects can be implemented in hardware, while other aspects can be implemented in
[0281] Embodiments of the application can be implemented by computer software executable by a data processor of the mobile device such as in the processor entity, or by hardware, or by a combination of software and hardware. Further, in this regard, it should be noted that any blocks of the logic flow as disclosed in the accompanying figures can represent program steps, or interconnected logic circuits, blocks and functions, or combinations of one or more of them. The software can be stored on such physical media as memory chips, or memory blocks implemented in the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof CD.
[0282] The memory can be of any type appropriate for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processor can be of any type appropriate for the local technical environment, and can encompass one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application- specific integrated circuits (ASICs), gate level circuits and processors based on multi-core processor architectures, as non-limiting examples.
[0283] Embodiments of the application can be practiced in a variety of components such as integrated circuit modules. The design of integrated circuits is by nature a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on semiconductor chips.
[0284] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well-established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) can be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0285] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiments of the application. However, various modifications and adaptations of the embodiments will become apparent to those skilled in the relevant arts, in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this application will still fall within the scope of the application as defined by the appended claims.
Claims
1. An apparatus comprising at least one processor and at least one memory including computer program code, said at least one memory and said computer program code being configured, together with said at least one processor, to cause the apparatus to at least: Receive at least one channel of voice audio signal and metadata associated with the at least one channel of voice audio signal, wherein, The at least one channel voice audio signal and the metadata associated with the at least one channel voice audio signal are generated from at least one microphone audio signal; Receive at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and the metadata associated with the at least one channel ambient audio signal are generated based on the analysis of at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel voice audio signal; and Based on the at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, and further based on the at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, an encoded multi-channel audio signal is generated, wherein the encoded multi-channel audio signal enables the at least one channel speech audio signal to be spatially presented independently of the at least one channel ambient audio signal, and the encoded multi-channel audio signal enables the at least one channel speech audio signal to be spatially presented with and at least substantially without the at least one channel ambient audio signal, wherein the encoded multi-channel audio signal includes at least a first portion encoded with a first encoder associated with at least one first input type and a second portion encoded with a second encoder associated with at least one second input type.
2. The apparatus according to claim 1, wherein, The apparatus is further configured to: receive at least one other audio object audio signal, wherein generating the encoded multi-channel audio signal is further based on the at least one other audio object audio signal, wherein the encoded multi-channel audio signal enables the at least one other audio object audio signal to be spatially presented independently of the at least one channel speech audio signal and the at least one channel ambient audio signal.
3. The apparatus according to claim 1, wherein, The at least one microphone audio signal from which the at least one channel voice audio signal and metadata associated with the at least one channel voice audio signal are generated, and the at least one microphone audio signal from which the at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal are generated, includes one of the following: A separate microphone group without a shared microphone; or A microphone group having at least one shared microphone.
4. The apparatus according to claim 1, wherein, The device is further configured to receive an input, the input being configured to control the generation of the encoded multichannel audio signal.
5. The apparatus according to claim 1, wherein, The apparatus is further configured to: modify the position parameters of the metadata associated with the at least one channel audio signal, or change the allocation of the near-channel rendering channel associated with the at least one channel audio signal, based on a mismatch between the determined position parameters of the metadata associated with the at least one channel audio signal and the allocated near-channel rendering channel.
6. The apparatus according to claim 1, wherein, The device generates encoded multi-channel audio signals to enable the following: Obtain the encoder bit rate; Select an embedded coding level and assign a bit rate to each selected embedded coding level, wherein a first level is associated with the at least one channel speech audio signal and the metadata associated with the at least one channel speech audio signal, a second level is associated with the at least one channel ambient audio signal, and a third level is associated with the metadata associated with the at least one channel ambient audio signal. Based on the allocated bit rate, the at least one channel speech audio signal and the metadata associated with the at least one channel speech audio signal, the at least one channel ambient audio signal and the metadata associated with the at least one channel ambient audio signal are encoded.
7. The apparatus according to claim 1, wherein, The device is further configured to: determine capability parameters, the capability parameters being determined based on at least one of the following: Transmission channel capacity; as well as The rendering device capability, wherein generating encoded multichannel audio signals is configured to further generate encoded multichannel audio signals based on the capability parameters.
8. The apparatus according to claim 7, wherein, The generation of encoded multichannel audio signals based on the capability parameters is further configured to: select an embedded encoding level based on at least one of the transmission channel capability and the rendering device capability, and allocate a bit rate to each selected embedded encoding level.
9. The apparatus according to claim 1, wherein, The at least one microphone audio signal used to generate the at least one channel ambient audio signal and the metadata associated with the at least one channel ambient audio signal based on parametric analysis includes at least two microphone audio signals.
10. The apparatus according to claim 1, wherein, The device is further configured to output the encoded multi-channel audio signal.
11. An apparatus comprising at least one processor and at least one memory including computer program code, said at least one memory and said computer program code being configured, together with said at least one processor, to cause the apparatus to at least: Receive embedded encoded audio signals, wherein the embedded encoded audio signals include at least one channel of speech audio signals and its associated metadata, and at least one channel of ambient audio signals, wherein, The embedded coded audio signal includes at least a first portion encoded with a first encoder associated with at least one first input type and a second portion encoded with a second encoder associated with at least one second input type, wherein the embedded coded audio signal is encoded according to at least one of the following levels of embedded audio signals: The first level is used to render the at least one channel voice audio signal and its associated metadata into a spatial voice scene; The second level is used to render the at least one channel of speech audio signal and its associated metadata, and the at least one channel of ambient audio signal, into a near-far stereo scene; or The third level is used to render the at least one channel speech audio signal and its associated metadata, and the at least one channel ambient audio signal and its associated spatial metadata, into a spatial audio scene; and The embedded encoded audio signal is decoded, and a multi-channel audio signal representing the scene is output, wherein the multi-channel audio signal enables the spatial presentation of the at least one channel speech audio signal independently of the at least one channel ambient channel audio signal, and the multi-channel audio signal enables the spatial presentation of the at least one channel speech audio signal in the presence of and at least substantially without the at least one channel ambient audio signal.
12. The apparatus according to claim 11, wherein, The embedded coded audio signal is encoded according to the third level, wherein the embedded coded audio signal further includes at least one other audio object audio signal and its associated metadata, and wherein the apparatus is configured to: decode and output a multi-channel audio signal representing the scene, such that the spatial presentation of the at least one other audio object audio signal is spatially independent of the at least one channel speech audio signal and the at least one channel ambient audio signal.
13. The apparatus according to claim 11, wherein, The device is further configured to receive an input configured to control the decoding of the embedded coded audio signal and the output of the multi-channel audio signal.
14. The apparatus according to claim 13, wherein, The input includes a capability switch, wherein the device that is configured to decode the embedded encoded audio signal and output the multi-channel audio signal is configured to update the decoding and output based on the capability switch.
15. The apparatus according to claim 14, wherein, The switching of the capability includes at least one of the following: Determining the earbud / headphone configuration; The configuration of the headphones was determined; and Determining the speaker output configuration.
16. The apparatus according to claim 13, wherein, The input includes at least one of the following: The determination of the embedded level change, wherein the apparatus for decoding the embedded encoded audio signal and outputting the multi-channel audio signal is configured to: update the decoding and output based on the embedded level change; and The determination of a change in bit rate at the embedded level, wherein the apparatus for decoding the embedded encoded audio signal and outputting the multi-channel audio signal is configured to update the decoding and output based on the change in bit rate at the embedded level.
17. The apparatus according to claim 13, wherein, The apparatus is configured to: control the decoding of the embedded encoded audio signal and the output of the multi-channel audio signal to modify the position of at least one channel audio signal or change the near-channel rendering channel assignment associated with the at least one channel audio signal based on the determined position of at least one voice audio signal and / or the mismatch between the assigned near-channel rendering channels.
18. The apparatus according to claim 13, wherein, The input includes determining the correlation between the at least one channel speech audio signal and the at least one channel ambient audio signal, and the means for decoding and outputting the multi-channel audio signal is configured to: When the correlation is less than the determined threshold: Controlling the position associated with the at least one channel of voice audio signal, and The ambient spatial scene formed by the at least one channel ambient audio signal is controlled by: rotating the ambient spatial scene based on the at least one channel ambient audio signal according to obtained rotation parameters, or compensating for the rotation of another device by applying a corresponding reverse rotation to the ambient spatial scene; and When the correlation is greater than or equal to the determined threshold: Controlling the position associated with the at least one channel of voice audio signal, and The ambient space scene formed by the at least one channel ambient audio signal is controlled by: compensating for the rotation of the other device by applying a corresponding reverse rotation to the ambient space scene while rotating the rest of the scene, or rotating the ambient space scene based on the at least one channel ambient audio signal according to the obtained rotation parameters.
19. A method comprising: Receive at least one channel voice audio signal and metadata associated with the at least one channel voice audio signal, wherein the at least one channel voice audio signal and the metadata associated with the at least one channel voice audio signal are generated from at least one microphone audio signal; Receive at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, wherein the at least one channel ambient audio signal and the metadata associated with the at least one channel ambient audio signal are generated based on the analysis of at least one microphone audio signal, and the at least one channel ambient audio signal is associated with the at least one channel voice audio signal; Based on the at least one channel speech audio signal and metadata associated with the at least one channel speech audio signal, and further based on the at least one channel ambient audio signal and metadata associated with the at least one channel ambient audio signal, an encoded multi-channel audio signal is generated, wherein the encoded multi-channel audio signal enables the at least one channel speech audio signal to be spatially presented independently of the at least one channel ambient audio signal, and the encoded multi-channel audio signal enables the at least one channel speech audio signal to be spatially presented with and at least substantially without the at least one channel ambient audio signal, wherein the encoded multi-channel audio signal includes at least a first portion encoded with a first encoder associated with at least one first input type and a second portion encoded with a second encoder associated with at least one second input type.
20. A method comprising: Receive an embedded encoded audio signal, the embedded encoded audio signal comprising at least one channel of speech audio signal and its associated metadata and at least one channel of ambient audio signal, wherein the embedded encoded audio signal comprises at least a first portion encoded by a first encoder associated with at least one first input type and a second portion encoded by a second encoder associated with at least one second input type, wherein the embedded encoded audio signal is encoded according to at least one of the following levels of embedded audio signals: The first level is used to render the at least one channel voice audio signal and its associated metadata into a spatial voice scene; The second level is used to render the at least one channel of speech audio signal and its associated metadata, and the at least one channel of ambient audio signal, into a near-far stereo scene; or The third level is used to render the at least one channel speech audio signal and its associated metadata, and the at least one channel ambient audio signal and its associated spatial metadata, into a spatial audio scene; and The embedded encoded audio signal is decoded, and a multi-channel audio signal representing the scene is output, wherein the multi-channel audio signal enables the spatial presentation of the at least one channel speech audio signal independently of the at least one channel ambient channel audio signal, and the multi-channel audio signal enables the spatial presentation of the at least one channel speech audio signal in the presence of and at least substantially without the at least one channel ambient audio signal.
Citation Information
Patent Citations
Speech capturing and speech rendering
US20110264450A1