Decoder and decoding method for discontinuous transmission of independent streams with parametrically encoded metadata - Patent Application 20070122997
The audio encoder and decoder system addresses DTX challenges by encoding transport channels actively and background noise in inactive phases, reducing bit rates and artifacts in immersive audio systems.
Patent Information
- Application Number
- JP2025514406
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-09
- Filing Date
- 2023-09-07
- Publication Date
- 2025-09-09
AI Technical Summary
Existing discontinuous transmission (DTX) systems for audio scenes with independent streams (ISM) and parametrically coded metadata face challenges in efficiently reducing bit rates while maintaining spatial synchronization and avoiding artifacts, particularly when applied to immersive audio with multiple objects.
An audio encoder and decoder system that determines voice activity to selectively encode transport channels during active phases and background noise during inactive phases, using metadata to generate immersive audio outputs, with separate quantization resolutions for active and inactive phases.
Achieves a dramatic reduction in bit rate for immersive audio transmission with improved spatial cues and reduced artifacts, ensuring seamless audio rendering across active and inactive frames.
Smart Images

Figure 2025529989000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to audio scenes with independent streams (ISM) with parametrically coded metadata, discontinuous transmission (DTX) modes for audio scenes with independent streams (ISM) with parametrically coded metadata and comfort noise generation (CNG), immersive voice and audio services (IVAS). In particular, the present invention relates to a coder and a method for discontinuous transmission of parametrically coded independent streams with metadata (DTX of Param-ISM). [Background technology]
[0002] In the IVAS codec, at low bit rates, audio objects or independent streams with metadata are encoded in a parametric manner. In a first step, for example, a downmix (e.g., a stereo downmix or a virtual cardioid) and metadata can be calculated from the audio objects and quantized directional information (e.g., from azimuth and elevation angles). The downmix can then be encoded, for example, to obtain one or more transport channels and, for example, transmitted to a decoder together with the metadata. The metadata can include, for example, directional information (e.g., azimuth and elevation angles), power ratios, and object indices corresponding to dominant objects that are a subset of the input objects. In the decoder, a covariance renderer can receive, for example, the metadata transmitted together with the stereo downmix / transport channels as input and, for example, render it into the required speaker layout (see [1], [2]).
[0003] Communications codecs typically employ discontinuous transmission (DTX) to significantly reduce the transmission rate when there is no voice input. In this mode, frames are first classified as "active" frames (i.e., frames containing voice) and "inactive" frames (i.e., frames containing either background noise or silence). For inactive frames, the codec then operates in DTX mode, significantly reducing the transmission rate. Most frames determined to contain background noise are dropped from transmission and replaced by some form of comfort noise generation (CNG) at the decoder. For these frames, a very low-rate parametric representation of the signal is transmitted using silence insertion descriptor (SID) frames, which are transmitted periodically rather than frame-by-frame. This allows the decoder's CNG to generate artificial noise that resembles the actual background noise.
[0004] A concept adopted according to the prior art is discontinuous transmission (DTX). Comfort noise generators are typically used for discontinuous transmission of speech. According to this concept, speech is first classified into active and inactive frames by a Voice Activity Detector (VAD). An example of VAD can be found in [3]. Based on the results of the VAD, only active speech frames are coded and transmitted at the nominal bit rate. During long pauses where only background noise or silence is present, the bit rate is reduced or even zeroed, and the background noise / silence is episodically and parametrically coded. Thus, the average bit rate is significantly reduced. Noise is generated during inactive frames by a comfort noise generator (CNG) on the decoder side. For example, the speech coders AMR-WB [3] and 3GPP EVS [4], [5] both have the possibility to operate in DTX mode. An example of an efficient CNG is given in [6]. In the IVAS codec, discontinuous transmission (DTX) systems exist for audio scenes that are parametrically coded by the Directional Audio Coding (DirAC) paradigm or transmitted in the Metadata-Assisted Spatial Audio (MASA) format (see [7]).
[0005] In discrete independent streams with metadata (discrete ISM), the encoder for discrete ISM accepts audio objects and their associated metadata. The objects are then individually encoded on a frame-by-frame basis along with metadata including object orientation information, such as azimuth and elevation angles, and the encodings are then transmitted to a decoder. The decoder then independently decodes the individual objects and renders them into a specified output layout by applying amplitude panning techniques using the quantized orientation information.
[0006] Another concept in the prior art is the Independent Stream with Parametrically Encoded Metadata (Param-ISM). Figure 4 shows an overview of the corresponding encoder, showing, among other things, the encoded audio signal 491 and the encoded parametric side information 495, 496, 497.
[0007] A parametric ISM (Param-ISM) encoder receives audio objects and associated metadata as input. The metadata may include, for example, frame-based object directions (e.g., azimuth angles with values between [-180, 180] and elevation angles with values between [-90, 90]), which are then quantized and used during the calculation of a stereo downmix (e.g., a virtual cardioid or transport channel). In addition, two dominant objects and a power ratio between the two dominant objects among the input audio objects may be determined, for example, for each time / frequency tile. The metadata may then be quantized and encoded, for example, along with the two dominant objects per time / frequency tile, i.e., the object indices of the two dominant objects.
[0008] The coded bitstream 490 may include, for example, a stereo downmix / transport channel 491 coded separately with the help of a core coder, a coded dominant object index 495, a quantized and coded power ratio 496, and quantized and coded directional information such as azimuth and elevation angles 497.
[0009] Figure 5 shows a simplified overview of a decoder. The decoder receives a bitstream 490 and obtains an encoded stereo downmix / transport channel 491, an encoded object index 495, an encoded power ratio 496, and an encoded directional information 497. The encoded stereo downmix / transport channel 491 is then decoded using a core decoder and converted to a time / frequency representation using an analysis filter bank, e.g., a complex low-delay filter bank (CLDFB). The decoded object index can be used, for example, with decoded and dequantized directional information, e.g., azimuth and elevation angles, and an output configuration, e.g., 5.1, 5.1+4, 7.1, 7.1+4, etc., to calculate a directional response. The direct response, e.g., the transport channel / stereo downmix in the time / frequency representation, along with a prototype matrix and decoded and dequantized power ratios, are provided as inputs to a covariance synthesis operating in the time / frequency domain. The output of the covariance synthesis is converted from the time / frequency representation to a time-domain representation using a synthesis filter, e.g., a CLDFB.
[0010] Figure 6 shows a detailed overview of the covariance synthesis step, without reflecting the dimensionality of the input / output data.
[0011] Covariance combining is performed by calculating a mixing matrix (M) for each time / frequency tile to combine the input transport channels. ( TIFF2025529989000002.tif566) The desired output speaker layout ( TIFF2025529989000003.tif586) (for example, 5.1 loudspeaker layout, 7.1 loudspeaker layout, 7.1+4 loudspeaker layout, etc.). TIFF2025529989000004.tif416For the mixing matrix, covariance synthesis is done using the prototype matrix, input covariance matrix TIFF2025529989000005.tif518, and the target covariance matrix TIFF2025529989000006.tif45 can be used. The target covariance matrix is calculated with the help of the transport channels / stereo downmix, power ratios and signal powers calculated from the direct response. Summary of the Invention [Problem to be solved by the invention]
[0012] It is an object of the present invention to provide an improved concept for discontinuous transmission of audio content. [Means for solving the problem]
[0013] The object of the present invention is solved by the subject matter of the independent claims.
[0014] According to one embodiment, an audio encoder is provided. The audio encoder comprises: a transport signal generator for generating two or more transport channels of a transport signal from an audio input including a plurality of audio input objects and at least one of a plurality of audio input channels. The audio encoder further comprises: a voice activity determiner for determining a voice activity decision for the transport signal indicating whether the audio input in the transport signal indicates voice activity. The audio encoder further comprises: a bitstream generator for generating a bitstream in response to the audio input. If the voice activity determiner determines that the transport signal indicates voice activity, the bitstream generator is adapted to encode the two or more transport channels in the bitstream. If the voice activity determiner determines that the transport signal does not indicate voice activity, the bitstream generator is adapted to encode information about background noise instead of the two or more transport channels, the information about background noise including information about background noise of at least one of the two or more transport channels or information about background noise of a derived signal dependent on at least one of the two or more transport channels.
[0015] For example, in one embodiment, the number of transport channels is less than or equal to the number of input channels.
[0016] Furthermore, a method for audio coding according to an embodiment is provided, the method comprising: generating two or more transport channels of a transport signal from an audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels; determining a voice activity decision for the transport signal indicating whether audio input in the transport signal indicates voice activity; determining a bitstream in response to an audio input; Includes.
[0017] If it is determined that the transport signal indicates voice activity, the method includes encoding two or more transport channels in the bitstream. If it is determined that the transport signal does not indicate voice activity, the method includes encoding, instead of the two or more transport channels, information about background noise in at least one of the two or more transport channels or information about background noise in a derived signal that depends on at least one of the two or more transport channels.
[0018] Furthermore, a computer program is provided for carrying out the above-described method when run on a computer or signal processor.
[0019] An audio decoder according to an embodiment is further provided. The audio decoder comprises an input interface for receiving a bitstream dependent on audio content, the bitstream including a plurality of audio objects and at least one of a plurality of audio channels. A transport signal including two or more transport channels is encoded in the bitstream, and the audio content is encoded in the transport signal. Alternatively, information regarding background noise is encoded in the bitstream instead of the transport signal, and the information regarding background noise includes information regarding the background noise of at least one of the two or more transport channels or information regarding the background noise of a derived signal dependent on at least one of the two or more transport channels. The audio decoder further comprises a renderer for generating one or more audio output signals according to the audio content encoded in the bitstream. If the transport signal including two or more transport channels is encoded in the bitstream, the renderer is configured to generate the one or more audio output signals according to the two or more transport channels. If the information regarding background noise is encoded in the bitstream instead of the transport signal, the renderer is configured to generate the one or more audio output signals according to the information regarding the background noise.
[0020] Further provided is a method for audio decoding, the method comprising: The method includes receiving a bitstream dependent on audio content, the bitstream including a plurality of audio objects and at least one of a plurality of audio channels, a transport signal including two or more transport channels being encoded in the bitstream, the audio content being encoded in the transport signal, or information about background noise being encoded in the bitstream instead of the transport signal, the information about background noise including information about background noise of at least one of the two or more transport channels or information about background noise of a derived signal dependent on at least one of the two or more transport channels. The method further includes generating one or more audio output signals in response to audio content encoded in the bitstream.
[0021] If a transport signal including two or more transport channels is encoded in the bitstream, the generation of the one or more audio output signals is performed in response to the two or more transport channels. If information about background noise is encoded in the bitstream instead of the transport signal, the generation of the one or more audio output signals is performed in response to the information about background noise.
[0022] Furthermore, a computer program is provided for carrying out the above-described method when run on a computer or signal processor.
[0023] Some embodiments are based on the discovery that combining existing solutions allows DTX to be applied independently to individual streams, e.g., audio objects, or individual channels of a stereo downmix / transport channel. However, this would be incompatible with DTX designed for low-bitrate communications, since in the case of two or more objects, or in the case of transport channels, or in the case of a downmix with two or more channels, the number of available bits would be insufficient to efficiently describe the inactive parts of the input signal. Furthermore, such an approach would also face problems due to the lack of synchronization of the individual VAD decisions. Spatial artifacts would result.
[0024] In an embodiment, a DTX system is provided for an audio scene described by (audio) objects and their associated metadata.
[0025] Some embodiments provide SID and CNG for DTX systems, in particular for parametrically coded audio objects (also known as ISM, i.e., independent streams with metadata) (e.g., as Param-ISM).
[0026] In some embodiments, a dramatic reduction in the bit rate required to transmit conversational immersive audio is achieved.
[0027] According to some embodiments, the DTX concept is provided extended to immersive audio with spatial cues.
[0028] In some embodiments, two most dominant objects per time / frequency unit are considered. In other embodiments, more than two most dominant objects per time / frequency unit are considered, especially when the number of input objects increases. For ease of reading the text, the following embodiments are primarily described with respect to two dominant objects per time / frequency unit, but these embodiments may also be extended to more than two dominant objects per time / frequency unit in a similar manner, for example in other embodiments.
[0029] A particular embodiment of an audio encoder is provided.
[0030] According to one embodiment, an audio encoder is provided for encoding a plurality of (audio) objects and their associated metadata.
[0031] The audio encoder may, for example, comprise a directional information determiner for extracting the directional information and a directional information quantizer for quantizing the directional information.
[0032] Further, the audio encoder may comprise a transport signal generator (downmixer) for generating a transport signal (downmix) comprising at least two transport channels (e.g., downmix channels), e.g., from the input audio object and from quantized directional information, e.g., azimuth and elevation angles associated with the input audio object.
[0033] Additionally, the audio encoder may comprise a decision logic module for combining the individual VAD decisions of the transport channels to compute an overall decision, for example, as to whether the frame is active or not.
[0034] Furthermore, the audio encoder may comprise a mono signal generator (eg, a stereo to mono converter) for outputting a mono signal from the transport channel to be encoded, for example, in the inactive phase.
[0035] Additionally, the audio encoder may comprise an inactive phase metadata generator for generating (eg, computing) inactive phase metadata, for example, to be transmitted during the inactive phase.
[0036] Additionally, the audio encoder may comprise an active metadata generator, for example for generating (eg, computing) active metadata to be transmitted during the active phase.
[0037] Furthermore, the audio encoder may comprise, for example, a transport channel encoder configured to generate the encoded data by encoding a downmix signal comprising a transport channel in an active phase.
[0038] Furthermore, the audio encoder may comprise a transport channel silence insertion description generator for generating silence insertion descriptions, for example of background noise of the mono signal during an inactive phase.
[0039] Furthermore, the audio encoder may comprise a multiplexer for, for example, combining active metadata and encoded data into a bitstream during an active phase and transmitting no data or silence insertion descriptions, or the multiplexer may be configured to, for example, combine transmission of silence insertion descriptions and inactive phase metadata during an inactive phase.
[0040] According to an embodiment, the transport signal generator / downmixer can apply, for example, a CELP coding scheme (CELP = Code Excited Linear Prediction), or can apply, for example, an MDCT-based coding scheme (MDCT = Modified Discrete Cosine Transform), or can apply, for example, a combination of the two coding schemes, with switching between them.
[0041] In one embodiment, the active and inactive phases may be determined, for example, by first running a voice activity detector on the transport / downmix channels separately and then combining the results of the transport / downmix channels to determine an overall decision.
[0042] According to one embodiment, the mono signal may be calculated, for example, from the transport / downmix channels, for example by adding the transport channels or, for example, by selecting the channel with higher long-term energy.
[0043] In one embodiment, the active and inactive metadata may differ, for example, in quantization resolution or type (nature) of parameters (used).
[0044] According to one embodiment, the quantization resolution of the transmitted directional information and the directional information used to calculate the downmix may be different, for example in the inactive phase.
[0045] In one embodiment, the spatial audio input format may be described, for example, by an object and its associated metadata (eg, by a separate stream with the metadata).
[0046] According to one embodiment, for example, two or more transport channels may be generated.
[0047] Furthermore, a specific embodiment of an audio decoder is provided.
[0048] According to an embodiment, an audio decoder is provided for (decoding and) generating a spatial audio output signal from a bitstream. The bitstream may, for example, indicate at least an active phase followed by at least an inactive phase. Furthermore, the bitstream may encode at least one silence insertion descriptor frame (SLD), which may, for example, describe background noise characteristics of a transport / downmix channel and / or spatial image information.
[0049] The audio decoder may comprise, for example, a SID decoder (Silence Insertion Descriptor Decoder) which may be configured to decode silence insertion descriptor frames of the mono signal.
[0050] Furthermore, the audio decoder may comprise a mono-to-stereo converter that may be configured, e.g., during an inactive phase / mode, to generate at least two (downmix) channels from SID information of the mono signal and control parameters, the control parameters may, e.g., describe characteristics of the stereo downmix / transport channel, e.g., scaling parameters, and / or either wideband coherence or wideband correlation calculated from the stereo downmix / transport channel on the encoder side.
[0051] Furthermore, the audio decoder may comprise a transport channel decoder which may be configured, for example, during an active phase / mode, to reconstruct transport / downmix channels from the bitstream during the active phase.
[0052] Furthermore, the audio decoder may include a (spatial) renderer that may be configured to reconstruct a spatial output signal, e.g., during an active phase / mode, from the decoded transport / downmix channels and e.g. from the transmitted active metadata and e.g. from the reconstructed background noise in the transport / downmix channels and e.g. from the transmitted inactive metadata during an inactive phase.
[0053] According to one embodiment, the mono-to-stereo converter may for example comprise a random generator, which may for example be run at least twice with different seeds to generate noise, which may for example use the decoded SID information of the mono signal and may be processed using control parameters which may for example describe the characteristics of the stereo downmix / transport channel, for example scaling parameters and / or wideband coherence or wideband correlation calculated from the stereo downmix / transport channel at the encoder side.
[0054] In one embodiment, the spatial parameters transmitted in the active phase may include, for example, an object index, a power ratio that may be transmitted, for example, in a frequency subband, and directional information (e.g., azimuth and elevation angles) that may be transmitted, for example, in a wideband.
[0055] According to one embodiment, the spatial parameters transmitted in the inactive phase may include, for example, directional information (e.g., azimuth and elevation angles) that may be transmitted in wideband, and control parameters that may describe, for example, the characteristics of the stereo downmix / transport channel, e.g., scaling parameters, and / or either the wideband coherence or the wideband correlation that may be calculated from the stereo downmix / transport channel on the encoder side.
[0056] In one embodiment, the quantization resolution of the direction information during the inactive phase is different from the quantization resolution of the direction information during the active phase.
[0057] According to one embodiment, the transmission of the control parameters may be, for example, in a wideband or, for example, in a frequency sub-band, and the decision whether to do it in a wideband or in a frequency sub-band may depend, for example, on bit rate availability.
[0058] In one embodiment, the renderer may be configured to perform covariance compositing, for example.
[0059] The renderer may for example comprise a signal power calculation unit for calculating a reference power depending on the transport / downmix channels for each time / frequency tile.
[0060] Furthermore, the renderer may comprise a direct power calculation unit for scaling the reference power, for example using the transmit power ratio in the active phase and using a fixed scaling factor in the inactive phase.
[0061] Furthermore, the renderer may comprise a direct response calculation unit for calculating the direct response, for example, depending on the quantized direction information of the dominant object during the active phase or depending on the quantized direction information of all transmitted objects during the inactive phase.
[0062] Furthermore, the renderer may comprise an input covariance matrix calculation unit for calculating an input covariance matrix based on, for example, the transport / downmix channels.
[0063] Furthermore, the renderer may comprise a target covariance matrix calculation unit for calculating a target covariance matrix based on the outputs of, for example, the direct response calculation block and the direct power calculation block.
[0064] Furthermore, the renderer may comprise a mixing matrix calculation unit for calculating a mixing matrix for rendering depending on, for example, the input covariance matrix and the target covariance matrix.
[0065] According to one embodiment, the constant scaling factor used during the inactive phase may be determined depending on, for example, the number of transmitted objects, or a control parameter may be used, for example.
[0066] In one embodiment, the dominant objects may be, for example, a subset of all transmitted objects, and the number of dominant objects may be, for example, less than / smaller than the number of transmitted objects.
[0067] According to an embodiment, the transport channel decoder may comprise, for example, a speech decoder, for example a CELP-based speech decoder, and / or may comprise, for example, a general-purpose audio decoder, for example a TCX-based decoder, and / or may comprise, for example, a bandwidth extension module.
[0068] Further particular embodiments are provided in the dependent claims.
[0069] In the following, embodiments of the invention will be explained in more detail with reference to the drawings. [Brief explanation of the drawings]
[0070] [Figure 1] 1 illustrates an audio encoder according to an embodiment; [Figure 2] FIG. 2 illustrates an audio decoder according to an embodiment. [Figure 3] FIG. 1 illustrates a system according to one embodiment. [Figure 4] FIG. 1 is a diagram illustrating an overview of a Param-ISM encoder. [Figure 5]FIG. 1 is a diagram illustrating an overview of a Param-ISM decoder. [Figure 6] FIG. 1 shows a detailed overview of the covariance synthesis step in Param-ISM without reflecting the dimensionality of the input and output data. [Figure 7] FIG. 2 is a block diagram according to one embodiment for determining whether a frame is active or inactive. [Figure 8] FIG. 2 is a block diagram of an encoder according to one embodiment. [Figure 9] FIG. 2 is a block diagram of a decoder according to one embodiment. [Figure 10] FIG. 1 illustrates a spatial renderer according to one embodiment. [Figure 11] FIG. 1 illustrates the generation of a stereo signal according to one embodiment using three random seeds Seed 1, Seed 2, and Seed 3, derived scaling factors, and control parameters. [Figure 12] FIG. 10 illustrates the generation of a stereo signal according to another embodiment, where the noise N3(k,n) generated from a third random generator for the left channel is also used to generate the right channel. DETAILED DESCRIPTION OF THE INVENTION
[0071] FIG. 1 illustrates an audio encoder 100 according to one embodiment.
[0072] The audio encoder 100 comprises a transport signal generator 110 for generating two or more transport channels of a transport signal from an audio input including at least one of a plurality of audio input objects and a plurality of audio input channels.
[0073] Furthermore, the audio encoder 100 comprises a voice activity determiner 120 for determining a voice activity decision of the transport signal indicative of whether the audio input in the transport signal exhibits voice activity or not.
[0074] Furthermore, the audio encoder 100 comprises a bitstream generator 130 for generating a bitstream depending on the audio input.
[0075] If the voice activity determiner 120 determines that the transport signal exhibits voice activity, the bitstream generator 130 is adapted to encode two or more transport channels in a bitstream.
[0076] If the voice activity determiner 120 determines that the transport signal does not exhibit voice activity, the bitstream generator 130 is adapted to encode information about background noise instead of the two or more transport channels, the information about background noise comprising information about the background noise of at least one of the two or more transport channels or information about the background noise of a derived signal dependent on at least one of the two or more transport channels.
[0077] According to an embodiment, the voice activity determiner 120 may be configured to determine an individual voice activity decision for each transport channel of one or more transport channels of the transport signal, e.g., indicating whether the audio input in the transport channel indicates voice activity or not. Furthermore, the voice activity determiner 120 may be configured to determine the voice activity decision of the transport signal in response to the individual voice activity decision for each transport channel of the one or more transport channels.
[0078] In one embodiment, the voice activity determiner 120 may be configured to determine, for example, for each of two or more transport channels of a transport signal, an individual voice activity decision indicating whether the audio input in said transport channel exhibits voice activity. Further, the voice activity determiner 120 may be configured to determine the voice activity decision of the transport signal in response to, for example, the individual voice activity decision of each of the two or more transport channels of the transport signal.
[0079] According to one embodiment, the voice activity determiner 120 may be configured to determine that a transport signal exhibits voice activity if, for example, at least one of two or more transport channels of the transport signal exhibits voice activity. Furthermore, the voice activity determiner 120 may be configured to determine that a transport signal does not exhibit voice activity if, for example, none of the two or more transport channels of the transport signal exhibits voice activity.
[0080] In one embodiment, the audio encoder 100 may be configured to determine whether to transmit a bitstream encoding information about background noise, or whether to not generate and transmit a bitstream, for example if the voice activity determiner 120 determines that the transport signal does not exhibit voice activity.
[0081] According to an embodiment, the audio encoder 100 may comprise a mono signal generator 830 (see FIG. 8) for generating a derived signal as a mono signal from at least one of the two or more transport channels, e.g., if the transport signal is determined by the voice activity determiner 120 not to exhibit voice activity. The audio encoder 100 may also comprise an information generator for generating information about background noise, e.g., as information about background noise for the mono signal.
[0082] In one embodiment, the mono signal generator 830 may be configured to generate the mono signal, for example, by adding two or more transport channels or adding two or more channels derived from two or more transport channels, or by selecting a transport channel exhibiting a higher energy among the two or more transport channels.
[0083] According to an embodiment, the information generator may be configured to generate information about the background noise of the mono signal as information about the mono signal, for example.
[0084] In one embodiment, the information generator may be configured to generate, for example, a silence insertion description of the background noise of the mono signal as information about the background noise of the mono signal.
[0085] According to an embodiment, the audio encoder 100 may, for example, comprise a direction information determiner 802 (see FIG. 8) for determining direction information in response to an audio input. The audio encoder 100 may, for example, comprise a direction information quantizer 804 (see FIG. 8) for quantizing the direction information to obtain quantized direction information. The bitstream generator 130 may, for example, be configured to encode the quantized direction information in a bitstream.
[0086] In one embodiment, the transport signal generator 110 may be configured to generate two or more transport channels of a transport signal from the audio input using, for example, directional information.
[0087] According to one embodiment, the audio input may, for example, include a plurality of audio input objects, and the directional information may, for example, include information regarding azimuth and elevation angles of audio input objects of the plurality of audio input objects of the audio input.
[0088] In one embodiment, the audio encoder 100 may comprise an active metadata generator 825 (see FIG. 8 ) for generating metadata including at least one of quantized directional information, object indices, and power ratios of a plurality of audio input objects and / or a plurality of audio input channels of the audio input when, for example, the voice activity determiner 120 determines that the transport signal indicates voice activity.
[0089] According to an embodiment, the audio input may, for example, comprise a plurality of audio input objects. The audio encoder 100 may comprise an inactivity metadata generator 826 (see Fig. 8) for generating metadata comprising quantized directional information and control parameters, such as, for example, a scaling factor depending on the number of audio input objects of the plurality of audio input objects of the audio input, or a scaling factor depending on the long-term energy of a transport channel of the transport signal and / or a scaling factor depending on the coherence or correlation between the transport channels of the transport signal, if, for example, the voice activity determiner 120 determines that the transport signal does not exhibit voice activity.
[0090] In one embodiment, the quantization resolution of the directional information that may be generated, for example, by the inactive metadata generator 826 is different from the quantization resolution of the directional information that may be generated, for example, by the active metadata generator 825 .
[0091] In one embodiment, for example, the characteristics of the metadata that may be generated by the inactive metadata generator 826 differ from the characteristics of the metadata that may be generated by the active metadata generator 825, for example.
[0092] In one embodiment, the audio input may include, for example, multiple audio input objects and metadata associated with the audio input objects.
[0093] In one embodiment, the transport signal generator 110 may be configured to generate two or more transport channels of the transport signal from the audio input, for example, by downmixing at least one of the plurality of audio input objects and the plurality of audio input channels to obtain a downmix as a transport signal, which downmix may include, for example, two or more downmix channels as the two or more transport channels.
[0094] According to one embodiment, if the audio input in the transport signal does not exhibit voice activity, the directional information quantizer 804 is configured to determine the quantized directional information such that the quantization resolution of the quantized directional information is different from the quantization resolution used, for example, to calculate the downmix.
[0095] In one embodiment, the bitstream generator 130 may be configured to encode a control parameter in the bitstream if, for example, the voice activity determiner 120 determines that the transport signal does not exhibit voice activity. The control parameter may, for example, be suitable for guiding the generation of the intermediate signal from random noise. The control parameter may, for example, include multiple parameter values for multiple subbands, or the control parameter may, for example, include a single wideband control parameter.
[0096] According to one embodiment, the audio encoder 100 may be configured to generate the control parameters by selecting, e.g., depending on the available bitrate, whether the control parameters may include, e.g., multiple parameter values for multiple subbands, or whether the control parameters may include, e.g., a single wideband control parameter.
[0097] In one embodiment, the transport signal generator 110 may be configured to encode the audio input, for example, by applying code-excited linear prediction, or by applying a modified discrete cosine transform, or by applying a combination of code-excited linear prediction and a modified discrete cosine transform.
[0098] According to one embodiment, if the audio input includes multiple audio input channels but does not include multiple audio input objects, the number of the two or more transport channels may be, for example, less than the number of the multiple audio input channels. If the audio input includes multiple audio input objects but does not include multiple audio input channels, the number of the two or more transport channels may be, for example, less than the number of the multiple audio input objects. If the audio input includes both multiple audio input objects and multiple audio input channels, the number of the two or more transport channels may be, for example, less than the sum of the number of the multiple audio input channels and the number of the multiple audio input objects.
[0099] Alternatively, according to one embodiment, if the audio input includes multiple audio input channels but does not include multiple audio input objects, the number of the two or more transport channels may be, for example, equal to or less than the number of the multiple audio input channels. If the audio input includes multiple audio input objects but does not include multiple audio input channels, the number of the two or more transport channels may be, for example, equal to or less than the number of the multiple audio input objects. If the audio input includes both multiple audio input objects and multiple audio input channels, the number of the two or more transport channels may be, for example, equal to or less than the sum of the number of the multiple audio input channels and the number of the multiple audio input objects.
[0100] FIG. 2 illustrates an audio decoder 200 according to one embodiment.
[0101] The audio decoder 200 includes an input interface 210 for receiving a bitstream dependent on audio content, the bitstream including multiple audio objects and at least one of multiple audio channels. A transport signal including two or more transport channels is encoded in the bitstream, and the audio content is encoded in the transport signal. Alternatively, information about background noise is encoded in the bitstream instead of the transport signal, and the information about background noise includes information about background noise of at least one of the two or more transport channels or information about background noise of a derived signal dependent on at least one of the two or more transport channels.
[0102] Furthermore, the audio decoder 200 comprises a renderer 220 for generating one or more audio output signals depending on the audio content encoded in the bitstream.
[0103] If a transport signal including two or more transport channels is encoded in the bitstream, the renderer 220 is configured to generate one or more audio output signals in response to the two or more transport channels.
[0104] If the information about the background noise is encoded in the bitstream instead of the transport signal, the renderer 220 is configured to generate one or more audio output signals depending on the information about the background noise.
[0105] According to one embodiment, if the audio content indicates voice activity, a transport signal including two or more transport channels may be encoded in the bitstream, for example. If the audio content does not indicate voice activity, information about background noise may be encoded in the bitstream instead of the transport signal, for example.
[0106] In one embodiment, the audio decoder 200 may comprise, for example, a demultiplexer 902, a noise information determiner 920, and a multi-channel generator 930 (see FIG. 9 ). The demultiplexer may be configured to determine whether a transmitted bitstream corresponds to an active frame or an inactive frame, for example, based on the size of the bitstream. If information about background noise is encoded in the bitstream, the noise information determiner 920 may be configured, for example, to determine the information about the background noise from the bitstream, the multi-channel generator 930 may be configured to generate a derived signal, for example, as an intermediate signal including two or more intermediate channels from the information about the background noise, and the renderer 220 may be configured, for example, to generate one or more audio output signals according to the two or more intermediate channels of the intermediate signal.
[0107] According to one embodiment, the multi-channel generator 930 may comprise, for example, a random generator for generating random noise, and may be configured to generate two or more intermediate channels in response to the random noise.
[0108] In one embodiment, the multi-channel generator 930 may be configured to, for example, shape random noise in response to information about the background noise to obtain shaped noise, and may be configured to generate two or more intermediate channels from the shaped noise.
[0109] According to one embodiment, the multi-channel generator 930 may be configured to obtain the random noise, for example by running a random generator at least twice with different seeds.
[0110] In one embodiment, the multi-channel generator 930 may be configured to generate two or more intermediate channels, for example, in response to random noise and in response to control parameters that depend on the transport channel of the transport signal, such as scaling and / or coherence or correlation, for example, which control parameters may be encoded in the bitstream, for example, as part of inactive metadata.
[0111] According to one embodiment, the control parameter may, for example, be coded in the bitstream and may, for example, comprise a plurality of parameter values for a plurality of sub-bands, and the multi-channel generator 930 may, for example, be configured to generate each sub-band of the plurality of sub-bands of the two or more intermediate channels depending on one parameter value of the plurality of parameter values of the control parameter associated with said sub-band.
[0112] In one embodiment, the control parameters may be, for example, encoded within the bitstream, and the control parameters may, for example, include a single wideband control parameter.
[0113] According to one embodiment, the multi-channel generator 930 may be configured to generate two or more intermediate channels, for example, by generating a first random noise portion of the random noise using a random generator having a first seed and generating a first intermediate channel of the two or more intermediate channels in response to the first random noise portion, by generating a second random noise portion of the random noise using a random generator having a second seed different from the first seed, and generating a second intermediate channel of the two or more intermediate channels in response to the second random noise portion.
[0114] According to one embodiment, the multi-channel generator 930 may be configured to generate a first intermediate channel of the two or more intermediate channels, for example, in response to the first random noise portion and the third noise portion, and in response to a control parameter, for example, a scaling factor and / or, for example, coherence or correlation. Furthermore, the multi-channel generator 930 may be configured to generate a second intermediate channel of the two or more intermediate channels, for example, in response to the second random noise portion and the third noise portion, and in response to a control parameter, for example, a scaling factor and / or, for example, coherence or correlation. The multi-channel generator 930 may be configured, for example, to generate a first random noise portion of the random noise using a random generator having a first seed, to generate a second random noise portion of the random noise using a random generator having a second seed, and to generate a third random noise portion of the random noise using a random generator having a third seed, where the second seed is different from the first seed, and the third seed is different from both the first seed and the second seed.
[0115] In one embodiment, the multi-channel generator 930 may be configured to generate two or more intermediate channels, for example, by generating a first intermediate channel of the two or more intermediate channels in response to random noise, and by generating a second intermediate channel of the two or more intermediate channels from the first intermediate channel of the two or more intermediate channels.
[0116] According to one embodiment, the multi-channel generator 930 may be configured to generate a second of the two or more intermediate channels, such that the second of the two or more intermediate channels may be identical to the first of the two or more intermediate channels, or the multi-channel generator 930 may be configured to generate the second of the two or more intermediate channels, such as by modifying the first of the two or more intermediate channels.
[0117] In one embodiment, the renderer 220 may be configured to generate, for example, two or more audio output signals as the one or more audio output signals.
[0118] According to one embodiment, the audio content may include, for example, multiple audio objects. If the audio content indicates voice activity, multiple audio object indices associated with the multiple audio objects, multiple power ratios associated with the multiple audio objects of the multiple subbands, and wideband directional information of the multiple audio objects may be encoded, for example, in the bitstream, and the renderer 220 may be configured to generate one or more audio output signals in response to, for example, the multiple audio object indices, the multiple power ratios, and the wideband directional information of the multiple audio objects.
[0119] In one embodiment, the audio content may include, for example, multiple audio objects. If the audio content does not exhibit voice activity, wideband directional information and control parameters of the multiple audio objects may be encoded, for example, in the bitstream, and the renderer 220 may be configured to generate one or more audio output signals, for example, depending on the wideband directional information and depending on all object indices and a fixed power ratio depending on the number of objects to be transmitted, where the fixed power ratio depends on the number of objects to be transmitted.
[0120] According to one embodiment, when the audio content indicates voice activity, the first quantization resolution of the wideband directional information encoded in the bitstream may be different from, for example, the second quantization resolution of the wideband directional information when the audio content does not indicate voice activity.
[0121] In one embodiment, the renderer 220 may comprise a signal power calculation unit 951 (see FIG. 10 ) for calculating, for example, a reference power according to two or more transport channels for each of a plurality of time-frequency tiles. Furthermore, the renderer 220 may comprise a direct power calculation unit 952 (see FIG. 10 ) for scaling the reference power to obtain a scaled reference power, for example using a transmission power ratio encoded in the bitstream if the audio content indicates voice activity, or using a scaling factor encoded in the bitstream if the audio content does not indicate voice activity. Furthermore, the renderer 220 may be configured to generate one or more audio output signals, for example depending on the scaled reference power.
[0122] According to one embodiment, the renderer 220 may comprise a direct response calculation unit 953 (see FIG. 10 ) for calculating the direct response, and may be configured to calculate the direct response according to quantized directional information of dominant objects that are a proper subset of the audio objects of the audio content if the audio content exhibits voice activity, or may be configured to calculate the direct response according to quantized directional information of all audio objects of the audio content if the audio content does not exhibit voice activity, and the quantized directional information may be encoded in the bitstream, for example. The renderer 220 may be configured to generate one or more audio output signals according to the direct response, for example.
[0123] In one embodiment, the renderer 220 may comprise an input covariance matrix calculation unit 954 (see FIG. 10 ) for calculating an input covariance matrix, for example, according to two or more transport channels. Furthermore, the renderer 220 may comprise a target covariance matrix calculation unit 955 (see FIG. 10 ) for calculating a target covariance matrix, for example, according to the direct response and the scaled reference power. Furthermore, the renderer 220 may comprise a mixing matrix calculation unit 956 (see FIG. 10 ) for calculating a mixing matrix for rendering, for example, according to the input covariance matrix and the target covariance matrix. The renderer 220 may be configured to generate one or more audio output signals, for example, according to the mixing matrix.
[0124] According to one embodiment, the renderer 220 may be configured to generate one or more of the transport channels of the transport signal, for example, by applying code-excited linear prediction, or by applying a modified discrete cosine transform or an inverse modified discrete cosine transform, or by applying a combination of code-excited linear prediction and a modified discrete cosine transform.
[0125] According to one embodiment, if the audio content includes multiple audio channels but does not include multiple audio objects, the number of the two or more transport channels may be, for example, less than the number of the multiple audio channels. If the audio content includes multiple audio objects but does not include multiple audio channels, the number of the two or more transport channels may be, for example, less than the number of the multiple audio objects. If the audio content includes both multiple audio objects and multiple audio channels, the number of the two or more transport channels may be, for example, less than the sum of the number of the multiple audio channels and the number of the multiple audio objects.
[0126] Alternatively, according to one embodiment, if the audio content includes multiple audio channels but does not include multiple audio objects, the number of the two or more transport channels may be, for example, equal to or less than the number of the multiple audio channels. If the audio content includes multiple audio objects but does not include multiple audio channels, the number of the two or more transport channels may be, for example, equal to or less than the number of the multiple audio objects. If the audio content includes both multiple audio objects and multiple audio channels, the number of the two or more transport channels may be, for example, equal to or less than the sum of the number of the multiple audio channels and the number of the multiple audio objects.
[0127] Figure 3 shows a system according to one embodiment, which comprises an audio encoder 100 according to one of the above-mentioned embodiments and an audio decoder 200 according to one of the above-mentioned embodiments.
[0128] The audio encoder 100 is configured to generate a bitstream from an audio input.
[0129] The audio decoder 200 is configured to generate one or more audio output signals from the bitstream.
[0130] The embodiments will be described in detail below.
[0131] According to one embodiment, the DTX system (e.g., the encoder) may be configured to determine an overall decision on whether a frame is inactive or active, e.g., depending on independent determination of channels of a stereo downmix and / or depending on individual audio objects.
[0132] A DTX system (eg, an encoder) may be configured to transmit a mono signal to a decoder, for example, using a silence insertion descriptor (SID) along with inactive metadata.
[0133] Furthermore, the DTX system (e.g., a decoder) may be configured to generate a transport channel / downmix comprising at least two channels using, for example, a comfort noise generator (CNG) from the SID information of a mono signal only.
[0134] Furthermore, the DTX system (e.g., a decoder) can be configured to, for example, post-process the generated transport channels / downmix with control parameters, which can, for example, be calculated from the stereo downmix / transport channels on the encoder side.
[0135] Furthermore, a DTX system (eg, a decoder) may render a multi-channel transport signal into a defined output layout, eg, using modified covariance combining.
[0136] Further specific embodiments are described below.
[0137] 7 shows a block diagram according to one embodiment for determining whether a frame is active or inactive, where the overall decision is based on the individual decisions of the transport channels / downmix channels.
[0138] In FIG. 7, a transport signal generator (eg, downmixer) 710 may be configured to receive, for example, audio objects and their associated quantized directional information (eg, azimuth and elevation angles).
[0139] First transport channel (e.g., left downmix channel) DMX L and a second transport channel (e.g., right downmix channel) DMX R A transport signal (e.g., downmix (DMX)) for TIFF2025529989000007.tif1557TIFF2025529989000008.tif1558Here, N is the total number of input objects, k is the index of the sample, and i is the index of the object. TIFF2025529989000009.tif585TIFF2025529989000010.tif586In another embodiment, two transport channels (e.g., downmix channels) may be generated using, for example, a downmix matrix D, for example, as follows: TIFF2025529989000011.tif1543 where, TIFF2025529989000012.tif58... TIFF2025529989000013.tif59 indicates audio object 1 to audio object N.
[0140] Additionally, FIG. 7 shows a decision logic module 720 comprising individual decision logic 722 and overall decision logic 725 .
[0141] 7, the individual decision logic 722 can be configured to, for example, determine whether an individual channel is active or inactive. The individual decision regarding whether each of the two (or more) transport channels is active or inactive may be indicated, for example, by a (e.g., internal) flag.
[0142] In one embodiment, each decision logic 722 may be configured to receive, for example, two (or more) transport channels as input. TIFF2025529989000014.tif412, For each transport channel of TIFF2025529989000015.tif412, it may be configured to determine whether said transport channel indicates voice activity, for example by analysing said transport channel.
[0143] In another embodiment, each decision logic 722 may, for example, determine whether two (or more) transport channels are being used. TIFF2025529989000016.tif412, All audio input channels or all audio input objects used by the transport signal generator 710 to form TIFF2025529989000017.tif412 may be analyzed. For example, if the individual determination logic 722 detects voice activity in at least one of the audio input channels or audio input objects, the individual determination logic 722 may conclude, for example, that there is voice activity in the respective transport channel, e.g., that the respective transport channel is active. For example, if the individual determination logic 722 detects no voice activity in any of the audio input channels or audio input objects used to generate the respective transport channel, the individual determination logic 722 may conclude, for example, that there is no voice activity in the respective transport channel, e.g., that the respective transport channel is inactive.
[0144] 7, the overall decision logic 725 may be configured to receive, for example, individual decisions (e.g., for transport channels) as input and may be configured to determine, for example, an overall decision in response to the individual decisions. For example, the overall decision logic 725 may indicate the decision using, for example, DTX_FLAG. The overall decision logic may determine, for example, an overall decision according to Table 1 below, which shows a frame-by-frame decision based on the individual decisions for each frame of the downmix.
[0145] [Table 1] The overall decision may be determined, for example, by using a hysteresis buffer of a predetermined size. Using a hysteresis buffer helps avoid artifacts that may be caused by frequent switching between active and inactive portions. For example, a hysteresis buffer of size 10 may require, for example, 10 frames before switching from an active decision to an inactive decision.
[0146] Exemplary pseudocode for determining the overall decision is shown below: Shift the hysteresis buffer by one step, e.g. buffer_decision[i] = buffer_decision[i+1] where i = 0, 1, 2 …. (Buff_size - 1) Buff_decision[buff_size] = Decision_Overall Here, Decision_Overall can be calculated, for example, as shown in Table 1.
[0147] The overall decision can be calculated, for example, as outlined in the following pseudocode: DTX_Flag = 1; for (i=0; i <buff_size; i++) { DTX_Flag = DTX_Flag && buffer_decision[i]; } In this pseudocode, DTX_Flag=1 means "inactive" and DTX_FLAG=0 means "active."
[0148] Figure 8 illustrates an audio encoder 800 according to one embodiment. The audio encoder of Figure 8 may, for example, implement a particular embodiment of the audio encoder 100 of Figure 1. In particular, Figure 8 illustrates a block diagram of an encoder that may, for example, be configured to receive an input audio object and its associated metadata.
[0149] Further, the audio encoder 800 may comprise a transport signal generator (e.g., downmixer) 810 (e.g., transport signal generator 710 of FIG. 7) for generating a downmix (transport channels) including at least two channels, e.g., from the input audio object and from quantized directional information, e.g., azimuth and elevation angles associated with the input audio object.
[0150] Furthermore, the audio encoder 800 may comprise a voice activity determiner, for example, implemented with a decision logic module 820 (e.g., decision logic module 720 of FIG. 7) for combining the individual VAD decisions of the transport channels to calculate an overall decision as to whether the frame is active or not.
[0151] The stereo downmix may be calculated in the transport signal generator 810, for example, using quantized directional information (eg, azimuth and elevation) from the input audio objects.
[0152] The stereo downmix may then be provided to a decision logic module 820, where a decision as to whether the frame is active or inactive may be determined based on, for example, the logic described above. For example, the decision logic module 820 may include, for example, individual decision logic 722 and overall decision logic 725, as described above.
[0153] When the decision logic module 820 makes an overall decision of "active" (for an active frame), the encoder of Figure 8 provides a more efficient approach compared to the encoder of Figure 4. In the case of an active downmix, both channels of the stereo downmix may be encoded independently in a transport channel encoder along with metadata, for example, as described in Table 2 (see below).
[0154] In contrast, if the decision logic module 820 determines "inactive" as the overall decision (for an inactive frame), the SID bitrate (e.g., either 4.4 kbps or 5.2 kbps) is too low to efficiently transmit both channels of the stereo downmix along with the active metadata. Thus, for occasional / occasionally transmitted SID frames, the metadata bitrate may be, for example, either 1.85 kbps or 2.45 kbps, and may include coarsely quantized directional information (e.g., azimuth and elevation) along with control parameters, e.g., controlling the spatiality of the background noise and derived from the stereo downmix / transport signal, such as scaling factors and / or, for example, coherence or correlation.
[0155] In an embodiment, during inactive frames, no transmission of the object may be indicated, but a power ratio, for example. The main motivation for not transmitting either the object index or the power ratio during inactive frames is the assumption that background noise has no particular direction and is naturally diffuse.
[0156] Furthermore, the audio encoder 800 may comprise a transport channel silence insertion description generator 840 for generating silence insertion descriptions, for example for background noise of the mono signal during the inactive phase. The transport channel SID generator (transport channel SID encoder) 840 may operate, for example, at 2.4 kbps and may receive, for example, a mono downmix as input.
[0157] Furthermore, the audio encoder 800 may comprise a mono signal generator (e.g., a stereo-to-mono converter) 830 for outputting a mono signal from a transport channel to be encoded in an inactive phase, for example. The conversion from a stereo downmix to a mono downmix may be performed by the mono signal generator (e.g., a stereo-to-mono converter) 830, for example.
[0158] In one embodiment, the downmix, eg, stereo to mono conversion, may be performed, eg, as the addition of two stereo transport / downmix channels, eg, as follows: In another embodiment, the downmix, e.g., stereo to mono conversion, may be implemented as a transmission of only one channel of the stereo downmix. The decision of which channel to select may depend, for example, on the (e.g., long-term) energy of the individual channels of the stereo downmix. For example, the channel with the higher long-term energy may be selected, for example, as follows: TIFF2025529989000020.tif1054 where, TIFF2025529989000021.tif47 shows the long-term energy of the first (e.g., left) channel, TIFF2025529989000022.tif47 shows the long-term energy of the second (e.g., right) channel.
[0159] Table 2 shows, for example, metadata that may be transmitted during active and inactive frames.
[0160] [Table 2] The audio encoder 800 of FIG. 8 may, for example, comprise a directional information extractor 802 for extracting directional information and a directional information quantizer 804 for quantizing the directional information.
[0161] Additionally, the audio encoder 800 may comprise an inactive metadata generator 826 for generating (eg, computing) inactive metadata, for example, to be transmitted during an inactive phase.
[0162] Additionally, the audio encoder 800 may comprise an active metadata generator 825 for generating (eg, computing) active metadata, for example, to be transmitted during the active phase.
[0163] Furthermore, the audio encoder 800 may comprise a transport channel encoder 828 configured to generate encoded data, for example, by encoding a downmix signal including a transport channel in an active phase.
[0164] Furthermore, the audio encoder 800 may comprise, for example, a bitstream generator, which may be implemented as, for example, a multiplexer 850 for combining (e.g., encoding) active metadata and encoded data (e.g., two or more transport channels) into a bitstream during an active phase and transmitting no data or silence insertion descriptions. Alternatively, the multiplexer 850 may be configured, for example, to combine transmission of silence insertion descriptions and inactive phase metadata during an inactive phase.
[0165] 9 shows an audio decoder 900 according to one embodiment. The audio decoder 900 of FIG. 9 may, for example, implement a particular embodiment of the audio decoder 200 of FIG.
[0166] The audio decoder 900 may, for example, receive a bitstream via an input interface, which may, for example, be implemented as a demultiplexer 902 .
[0167] The audio decoder 900 of FIG. 9 may include a transport channel decoder 910 that may be configured, for example, during an active phase / mode, to reconstruct transport / downmix channels from a bitstream during the active phase.
[0168] Furthermore, the audio decoder 900 may comprise a noise information determiner, for example implemented as a SID decoder (Silence Insertion Descriptor Decoder) 920, which may be configured to decode, for example, silence insertion descriptor frames of the mono signal.
[0169] Furthermore, the audio decoder 900 may comprise a multi-channel generator 930, implemented for example as a mono-to-stereo converter 930, which may be configured to generate at least two (downmix) channels from the SID information of the mono signal and from control parameters, for example during an inactive phase / mode.
[0170] Additionally, the audio decoder 900 of FIG. 9 may comprise, for example, a filter bank analysis module 940.
[0171] Furthermore, the audio decoder 900 may comprise a (e.g., spatial) renderer 950 that may be configured to reconstruct a spatial output signal from the decoded transport / downmix channels, e.g., from transmitted active metadata, e.g., from reconstructed background noise in the transport / downmix channels, e.g., during an active phase / mode, and from transmitted inactive metadata, e.g., during an inactive phase.
[0172] The audio decoder 900 of FIG. 9 may include, for example, a synthesis module for performing (eg, frequency band) synthesis of the spatial output signals of the renderer 950.
[0173] The audio decoder 900 of FIG. 9 may further comprise a voice activity information determiner 905 for determining, for example, whether the decoder is operating in an active or inactive fashion (in either an active or inactive mode), e.g., depending on VAD data in the bitstream.
[0174] In the active mode described here, the decoder described in FIG. 9 is more efficient than the decoder described in FIG.
[0175] 10 illustrates a spatial renderer, for example for covariance rendering, according to one embodiment. The renderer 950 illustrated in FIG. 9 can be implemented as the spatial renderer of FIG.
[0176] The renderer may for example comprise a signal power calculation unit 951 for calculating a reference power depending on the transport / downmix channel for each time / frequency tile.
[0177] Further, the renderer may comprise a direct power calculation unit 952 for scaling the reference power, e.g. using the power ratio transmitted in the active phase and either using a fixed scaling factor that depends on the number of transmitted objects, e.g., or a scaling factor transmitted as part of the metadata, or e.g., no scaling in the inactive phase.
[0178] Further, the renderer may comprise a direct response calculation unit 953 for calculating a direct response, for example depending on the quantized direction information of the dominant object during the active phase or depending on the quantized direction information of all transmitted objects during the inactive phase.
[0179] Furthermore, the renderer may comprise an input covariance matrix calculation unit 954 for calculating an input covariance matrix based on, for example, the transport / downmix channels.
[0180] Further, the renderer may include a target covariance matrix calculation unit 955 for calculating a target covariance matrix, for example, depending on the output of the direct power calculation block 952 and the output of the direct response calculation block 953 (or depending on a calculated covariance matrix that depends on the output of the direct response calculation block 953).
[0181] Furthermore, the renderer may comprise a mixing matrix calculation unit 956 for calculating a mixing matrix for rendering as a function of, for example, the input covariance matrix and the target covariance matrix.
[0182] For example, in the case of a mixing matrix, covariance synthesis is performed using the prototype matrix, the input covariance matrix, TIFF2025529989000024.tif518, and the target covariance matrix TIFF2025529989000025.tif45 can be used as described with reference to Figure 6.
[0183] Furthermore, the renderer may comprise an amplitude panning unit 957 for performing amplitude panning on the transport channels depending on the mixing matrix calculated by, for example, the mixing matrix calculation unit 956 .
[0184] The spatial renderer for covariance synthesis-based rendering shown in Figure 10 can use active metadata such as quantized direction information, object indices, and power ratios, etc. Therefore, covariance rendering is more efficient compared to the covariance rendering shown in Figure 3.
[0185] 9 may, for example, independently decode the two channels of a stereo downmix in the bitstream, which may be fed to a filter bank analysis module 940 before being provided as input to a covariance synthesis.
[0186] In the inactive mode described here, the SID decoder 920 and mono-to-stereo converter 930 can use, for example, the encoded SID information of the mono channel to generate a stereo signal with some degree of spatial correlation reduction.
[0187] According to one embodiment, an efficient implementation of mono to stereo conversion may be used, for example, by running a random generator twice with different seeds. In one embodiment, the generated noise may be shaped, for example, with the SID information of the mono channel. This produces a stereo signal (with zero coherence).
[0188] In another embodiment, the mono channel may for example be copied to both stereo channels (however this has the drawback of causing spatial collapse and coherence to become unity).
[0189] In a preferred embodiment, a stereo signal with similar coherence and energy as the input stereo downmix is generated. To generate TIFF2025529989000026.tif514, control parameters such as coherence and / or correlation and scaling factors, which may be transmitted as part of the inactive metadata, may be used. TIFF2025529989000027.tif8118TIFF2025529989000028.tif8119Where, TIFF2025529989000029.tif519TIFF2025529989000030.tif527where k is the frequency index, n is the sample index, and c(n) is either the coherence or correlation that is sent as part of the inactive metadata. TIFF2025529989000031.tif531 is a scaling factor derived from the scaling factor s sent as part of the inactive metadata, TIFF2025529989000032.tif534 and TIFF2025529989000033.tif516 are random noises generated by different random generators with seed 1, seed 2, and seed 3 respectively.
[0190] Since the inactive metadata does not include the power ratios and object indices, during the direct power calculation a scaling factor that may depend on, for example, the number of objects may be used instead of, for example, the power ratio, or alternatively, a scaling factor that is transmitted as part of the inactive metadata may be used instead of, for example, the power ratio.
[0191] FIG. 11 illustrates the generation of a stereo signal according to one embodiment using three random seeds Seed 1, Seed 2, and Seed 3, the derived scaling factors, and the control parameters.
[0192] Furthermore, FIG. 11 shows a random generator comprising a random generator unit 1 and a random generator unit 3 for generating the left channel, and a random generator unit 2 and another random generator unit 3 for generating the right channel.
[0193] In FIG. 11, the random generator unit 3 for generating the left channel and the random generator unit 3 for generating the right channel receive the same seed 3 and therefore, for example, the same random noise TIFF2025529989000034.tif516 can be generated.
[0194] FIG. 12 shows the generation of a stereo signal according to another embodiment, where the generated noise of the random generator unit 3 for the left channel is 12 includes a random generator unit 1, a random generator unit 2, and only one random generator unit 3.
[0195] In further embodiments, the random generator may, for example, comprise only a single random generator unit, which may, for example, generate random noises in response to receiving Seed 1, Seed 2, and Seed 3, respectively. TIFF2025529989000036.tif534 and This may be used to generate TIFF2025529989000037.tif516 in turn.
[0196] In other embodiments, the above concepts apply equally to generating multi-channel signals having more than two channels.
[0197] Furthermore, for example, the directional information of all objects can be used to compute the direct response, rather than just the dominant object.
[0198] Embodiments allow extending DTX in an efficient way to spatial audio coding with an independent stream with metadata (ISM), which maintains high perceptual fidelity with respect to background noise, even during inactive frames where transmission may be interrupted, for example, to save communication bandwidth.
[0199] Decoder-side transport channels with two or more channels can be generated from only the mono signal transmitted by a comfort noise generator (CNG) to present a spatial image from the SID information. The generated transport channels can then be fed to a covariance synthesis module, for example, along with direct responses calculated from the directional information of all audio objects, equal power ratios, and prototype matrices for rendering to the required output layout.
[0200] While some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most significant method steps may be performed by such an apparatus.
[0201] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partially in hardware, or at least partially in software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, on which electronically readable control signals are stored, which cooperates (or can cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable.
[0202] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0203] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code can be stored on, for example, a machine-readable carrier.
[0204] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0205] In other words, therefore, an embodiment of the inventive methods is a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0206] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.
[0207] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can for example be adapted to be transferred via a data communication connection, for example via the Internet.
[0208] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0209] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0210] Further embodiments according to the invention comprise an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0211] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0212] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0213] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0214] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented as descriptions and explanations of the embodiments herein.
[0215] References [1] WO 2022 / 079049 A2, A. “Apparatus and method for encoding a plurality of audio objects and apparatus and method for decoding using two or more relevant audio objects”. [2] WO 2022 / 079044 A1 “Apparatus and method for encoding a plurality of audio objects using direction information during a downmixing or apparatus and method for decoding using an optimized covariance synthesis”. [3] 3GPP TS 26.194; Voice Activity Detector (VAD); - 3GPP technical specification Retrieved on 2009-06-17. [4] 3GPP TS 26.449, "Codec for Enhanced Voice Services (EVS); Comfort Noise Generation (CNG) Aspects". [5] 3GPP TS 26.450, "Codec for Enhanced Voice Services (EVS); Discontinuous Transmission (DTX)". [6] A. Lombard, S. Wilde, E. Ravelli, S. Doehla, G. Fuchs and M. Dietz, "Frequency-domain Comfort Noise Generation for Discontinuous Transmission in EVS," 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, 2015, pp. 5893-5897, doi: 10.1109 / ICASSP.2015.7179102. [7] WO 2022 / 022876 A1 ”Apparatus, method and computer program for encoding an audio signal or for decoding an encoded audio scene”.
Claims
1. an input interface (210; 902) for receiving a bitstream dependent on audio content, the bitstream comprising at least one of a plurality of audio objects and a plurality of audio channels, wherein a transport signal comprising two or more transport channels is encoded in the bitstream, the audio content is encoded in the transport signal, or information about background noise is encoded in the bitstream instead of the transport signal, the information about background noise comprising information about background noise of at least one of the two or more transport channels or information about background noise of a derived signal dependent on at least one of the two or more transport channels; a renderer (220; 950) for generating one or more audio output signals in response to the audio content encoded in the bitstream; Equipped with if the transport signal comprising the two or more transport channels is encoded in the bitstream, the renderer (220; 950) is configured to generate the one or more audio output signals in response to the two or more transport channels, - if the information about the background noise is encoded in the bitstream instead of the transport signal, the renderer (220; 950) is configured to generate the one or more audio output signals in response to the information about the background noise. Audio decoder (200; 900).
2. if the audio content indicates voice activity, the transport signal including the two or more transport channels is encoded into the bitstream; If the audio content does not show voice activity, the information about the background noise is encoded in the bitstream instead of the transport signal. An audio decoder (200; 900) according to claim 1.
3. The audio decoder (200; 900) comprises a noise information determiner (920) and a multi-channel generator (930), wherein the noise information determiner (920) is configured to determine the information about the background noise from the bitstream if the information about the background noise is coded in the bitstream, the multi-channel generator (930) is configured to generate the derived signal as an intermediate signal comprising two or more intermediate channels from the information about the background noise, and the renderer (220; 950) is configured to generate the one or more audio output signals according to the two or more intermediate channels of the intermediate signal. An audio decoder (200; 900) according to claim 1 or 2.
4. the multi-channel generator (930) comprises a random generator for generating random noise; the multi-channel generator (930) is configured to generate the two or more intermediate channels in response to the random noise generated by the random generator. An audio decoder (200; 900) according to claim 3.
5. the multi-channel generator (930) is configured to shape the random noise in response to the information about the background noise to obtain a shaped noise; the multi-channel generator (930) is configured to generate the two or more intermediate channels from the shaped noise. An audio decoder (200; 900) according to claim 4.
6. 6. An audio decoder (200; 900) according to claim 4 or 5, wherein the multi-channel generator (930) is configured to run the random generator at least twice with different seeds to obtain the random noise.
7. the multi-channel generator (930) is configured to generate the two or more intermediate channels in response to the random noise and in response to control parameters encoded in the bitstream, e.g., the control parameters comprising e.g., a scaling factor and / or e.g., any of coherence or correlation; An audio decoder (200; 900) according to any one of claims 4 to 6.
8. at least one of the control parameters is encoded in the bitstream and includes a plurality of parameter values for a plurality of subbands, and the multi-channel generator (930) is configured to generate each subband of the plurality of subbands of the two or more intermediate channels in response to one parameter value of the plurality of parameter values of the at least one of the control parameters associated with the subband. An audio decoder (200; 900) according to claim 7.
9. the control parameter is encoded within the bitstream, the control parameter being a single wideband control parameter; An audio decoder (200; 900) according to claim 7.
10. the multi-channel generator (930) is configured to generate the two or more intermediate channels by generating a first random noise portion of the random noise using the random generator having a first seed, generating a first intermediate channel of the two or more intermediate channels in response to the first random noise portion, generating a second random noise portion of the random noise using the random generator having a second seed different from the first seed, and generating a second intermediate channel of the two or more intermediate channels in response to the second random noise portion. An audio decoder (200; 900) according to any one of claims 4 to 9.
11. the multi-channel generator (930) is configured to generate a first intermediate channel of the two or more intermediate channels in response to a first random noise portion, in response to a third noise portion, and in response to the control parameters, such as the scaling factor and the coherence and / or correlation; the multi-channel generator (930) is configured to generate a second intermediate channel of the two or more intermediate channels in response to a second random noise portion, in response to the third noise portion, and in response to the control parameter; the multi-channel generator (930) is configured to generate the first random noise portion of the random noise using the random generator having a first seed; the multi-channel generator (930) is configured to generate the second random noise portion of the random noise using the random generator having a second seed; the multi-channel generator (930) is configured to generate the third random noise portion of the random noise using the random generator having a third seed; the second seed is different from the first seed, the third seed is different from the first seed and different from the second seed; Audio decoder (200; 900) according to any one of claims 4 to 9, further dependent on claim 7.
12. the multi-channel generator (930) is configured to generate the two or more intermediate channels by generating a first intermediate channel of the two or more intermediate channels in response to the random noise, and by generating a second intermediate channel of the two or more intermediate channels from the first intermediate channel of the two or more intermediate channels. An audio decoder (200; 900) according to any one of claims 4 to 9.
13. the multi-channel generator (930) is configured to generate the second intermediate channel of the two or more intermediate channels such that the second intermediate channel of the two or more intermediate channels is identical to the first intermediate channel of the two or more intermediate channels; or the multi-channel generator (930) is configured to generate the second intermediate channel of the two or more intermediate channels by modifying the first intermediate channel of the two or more intermediate channels. An audio decoder (200; 900) according to claim 12.
14. the renderer (220; 950) is configured to generate the two or more audio output signals as the one or more audio output signals. An audio decoder (200; 900) according to any one of claims 1 to 13.
15. the audio content includes the plurality of audio objects; if the audio content indicates voice activity, a plurality of audio object indices associated with the plurality of audio objects, a plurality of power ratios associated with the plurality of audio objects for a plurality of subbands, and wideband directional information for the plurality of audio objects are coded in the bitstream, and the renderer (220; 950) is configured to generate the one or more audio output signals in response to the plurality of audio object indices, in response to the plurality of power ratios, and in response to the wideband directional information for the plurality of audio objects. An audio decoder (200; 900) according to any one of claims 1 to 14.
16. the audio content includes the plurality of audio objects; if the audio content does not indicate voice activity, wideband directional information of the plurality of audio objects and the control parameters are encoded in the bitstream, and the renderer (220; 950) is configured to generate the one or more audio output signals in response to the wideband directional information. An audio decoder (200; 900) according to any one of claims 1 to 15, further dependent on claim 7.
17. a first quantization resolution of the wideband directional information encoded in the bitstream when the audio content indicates voice activity is different from a second quantization resolution of the wideband directional information when the audio content does not indicate voice activity; An audio decoder (200; 900) according to claims 15 and 16.
18. said renderer (220; 950) comprising a signal power calculation unit (951) for calculating a reference power according to said two or more transport channels for each of a plurality of time-frequency tiles; the renderer (220; 950) comprises a direct power calculation unit (952), which is configured to scale the reference power using a transmission power ratio coded in the bitstream to obtain a scaled reference power if the audio content does not exhibit voice activity, and to use a scaling factor if the audio content exhibits voice activity, the scaling factor being coded in the bitstream or the scaling factor being a fixed scaling factor that depends for example on the number of transmitted objects, the renderer (220; 950) is configured to generate the one or more audio output signals in response to the scaled reference power. An audio decoder (200; 900) according to any one of claims 1 to 17.
19. the renderer (220; 950) comprises a direct response calculation unit (953) for calculating a direct response, the renderer (220; 950) being configured to calculate the direct response in response to quantized directional information of dominant objects that are a proper subset of the plurality of audio objects of the audio content if the audio content exhibits voice activity, and the renderer (220; 950) being configured to calculate the direct response in response to quantized directional information of all audio objects of the audio content if the audio content does not exhibit voice activity, the quantized directional information being coded in the bitstream; the renderer (220; 950) is configured to generate the one or more audio output signals in response to the direct response. An audio decoder (200; 900) according to claim 18.
20. said renderer (220; 950) comprising an input covariance matrix calculation unit (954) for calculating an input covariance matrix in response to said two or more transport channels; said renderer (220; 950) comprising a target covariance matrix calculation unit (955) for calculating a target covariance matrix as a function of said direct response and said scaled reference power; the renderer (220; 950) comprises a mixing matrix calculation unit (956) for calculating a mixing matrix for rendering according to the input covariance matrix and the target covariance matrix; the renderer (220; 950) is configured to generate the one or more audio output signals in response to the mixing matrix.
20. An audio decoder (200; 900) according to claim 19.
21. the renderer (220; 950) is configured to generate one or more of the two or more transport channels by applying code-excited linear prediction, or by applying a modified discrete cosine transform or an inverse of the modified discrete cosine transform, or by applying a combination of the code-excited linear prediction and the modified discrete cosine transform, An audio decoder (200; 900) according to any one of claims 1 to 20.
22. if the audio content includes the plurality of audio channels but does not include the plurality of audio objects, the number of the two or more transport channels is less than the number of the plurality of audio channels; if the audio content includes the plurality of audio objects but does not include the plurality of audio channels, the number of the two or more transport channels is less than the number of the plurality of audio objects; when the audio content includes both the plurality of audio objects and the plurality of audio channels, the number of the two or more transport channels is smaller than the sum of the number of the plurality of audio channels and the number of the plurality of audio objects; or when the audio content includes the plurality of audio channels but does not include the plurality of audio objects, the number of the two or more transport channels is equal to or less than the number of the plurality of audio channels; if the audio content includes the plurality of audio objects but does not include the plurality of audio channels, the number of the two or more transport channels is equal to or less than the number of the plurality of audio objects; When the audio content includes both the plurality of audio objects and the plurality of audio channels, the number of the two or more transport channels is equal to or less than the sum of the number of the plurality of audio channels and the number of the plurality of audio objects. An audio decoder (200; 900) according to any one of claims 1 to 21.
23. an audio encoder (100; 800); An audio decoder (200; 900) according to any one of claims 1 to 22, Equipped with The audio encoder (100; 800) a transport signal generator (110; 710; 810) for generating two or more transport channels of a transport signal from an audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels; a voice activity determiner (120; 820) for determining a voice activity decision of said transport signal, indicative of whether said audio input in said transport signal exhibits voice activity; and a bitstream generator (130; 850) for generating a bitstream in response to said audio input; Equipped with the bitstream generator (130; 850) is adapted to encode the two or more transport channels in the bitstream if the voice activity determiner (120; 820) determines that the transport signal exhibits voice activity, if the voice activity determiner (120; 820) determines that the transport signal does not exhibit voice activity, the bitstream generator (130; 850) is adapted to encode information about background noise instead of the two or more transport channels, the information about the background noise comprising information about the background noise of at least one of the two or more transport channels or information about the background noise of a derived signal dependent on at least one of the two or more transport channels; said audio encoder (100; 800) configured to generate a bitstream from an audio input; the audio decoder (200; 900) is configured to generate one or more audio output signals from the bitstream; system.
24. receiving a bitstream dependent on audio content comprising at least one of a plurality of audio objects and a plurality of audio channels, wherein a transport signal comprising two or more transport channels is encoded in the bitstream, the audio content is encoded in the transport signal, or information regarding background noise is encoded in the bitstream instead of the transport signal, the information regarding background noise comprising information regarding background noise of at least one of the two or more transport channels or information regarding background noise of a derived signal dependent on at least one of the two or more transport channels; generating one or more audio output signals in response to the audio content encoded in the bitstream; if the transport signal including the two or more transport channels is encoded in the bitstream, generating the one or more audio output signals is performed in response to the two or more transport channels; if the information about the background noise is encoded in the bitstream instead of the transport signal, generating the one or more audio output signals is performed in response to the information about the background noise. A method for decryption.
25. 25. A computer program for performing the method according to claim 24 when the computer program is run on a computer or signal processor.
Citation Information
Patent Citations
Voice decoding switching system, voice decoding switching method and voice decoding switching program
JP2009198652A
Multi-channel audio signal processing method, device, and system
JP2019533189A
Stereo signal encoding device, stereo signal decoding device, stereo signal encoding method, and stereo signal decoding method
WO2012066727A1
Spatial comfort noise
WO2014143582A1
Methods and devices for encoding and / or decoding spatial background noise within a multi-channel input signal
WO2021252705A1