Decoder and decoding method for discontinuous transmission of independent streams with metadata in parameter coded

By introducing a transmission signal generator, a voice activity determiner and a bitstream generator into the audio encoder and decoder, the encoding method of the transmission signal is dynamically adjusted, and the bit rate management problem of multiple audio objects or transmission channels in the prior art is solved, thereby achieving efficient audio transmission and reducing spatial falsetto effects.

CN120112995APending Publication Date: 2025-06-06FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380075680.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-09
Filing Date
2023-09-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In audio encoding and decoding, it is difficult to effectively manage the bit rates of multiple audio objects or transmission channels in the non-continuous transmission (DTX) mode, resulting in a decrease in the average bit rate and the occurrence of spatial falsetto.

Method used

An audio encoder and decoder are adopted to dynamically adjust the encoding method of the transmitted signal according to the voice activity of the audio input to ensure that the bit rate usage is optimized in the active and inactive stages respectively.

Benefits of technology

It realizes that in multiple audio objects or transmission channel scenarios, effectively reduces the bit rate requirement, reduces the appearance of spatial falsetto, and improves the efficiency and quality of audio transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112995A_ABST
    Figure CN120112995A_ABST
Patent Text Reader

Abstract

An audio decoder (200) according to an embodiment is provided. The audio decoder (200) comprises an input interface (210) for receiving a bitstream dependent on audio content comprising at least one of a plurality of audio objects and a plurality of audio channels; wherein a transmission signal comprising two or more transmission channels is encoded within the bitstream, and the audio content is encoded within the transmission signal; or wherein information on background noise rather than the transmission signal is encoded within the bitstream wherein the information on background noise comprises information on background noise of at least one of the two or more transmission channels or information on background noise of a derived signal, the derived signal is dependent on at least one of the two or more transmission channels. Furthermore, the audio decoder (200) comprises a renderer (220) for generating one or more audio output signals in dependence on the audio content encoded with the bitstream. If the transmission signal comprising the two or more transmission channels is encoded within the bitstream, the renderer (220) is configured to generate the one or more audio output signals in dependence on the two or more transmission channels. If the information about the background noise rather than the transmit signal is encoded within the bitstream, the renderer (220) is configured to generate the one or more audio output signals in dependence on the information about the background noise.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] manual

[0002] The present invention relates to parameterized coded independent stream with metadata (ISM) audio scenes, discontinuous transmission (DTX) mode and comfort noise generation (CNG) for parameterized coded independent stream with metadata (ISM) audio scenes, immersive voice and audio services (IVAS). In particular, the present invention relates to a codec and method for parameterized coded independent stream with metadata discontinuous transmission (DTX for Param-ISM).

[0003] In the IVAS codec, audio objects or separate streams with metadata are parametrically encoded and decoded at low bitrates. In a first step, a downmix (e.g., a stereo downmix or a virtual cardioid) and metadata may be calculated, for example, from the audio objects and from quantized directional information (e.g., from azimuth and elevation). The downmix is ​​then encoded, for example, to obtain one or more transport channels, and may be transmitted, for example, together with the metadata, to a decoder. The metadata may, for example, include directional information (e.g., azimuth and elevation), power ratios, and object indices corresponding to primary objects that are a subset of the input objects. At the decoder, a covariance renderer may, for example, receive as input the transmitted metadata and the stereo downmix / transmission channels and may, for example, render them to the desired loudspeaker layout (see [1], [2]).

[0004] Typically, in communications codecs, discontinuous transmission (DTX) is used to significantly reduce the transmission rate in the absence of speech input. In this mode, frames are first classified into "active" frames (i.e., frames containing speech) and "inactive" frames (i.e., frames containing background noise or silence). Later, for inactive frames, the codec runs in DTX mode to significantly reduce the transmission rate. Most frames that are determined to include background noise are stopped from transmission and replaced with some comfort noise generation (CNG) at the decoder. For these frames, a very low rate parameter representation of the signal is transmitted using a silence insertion descriptor (SID) frame that is sent periodically but not at every frame. This allows the CNG in the decoder to generate artificial noise similar to actual background noise.

[0005] The concept used according to the prior art is discontinuous transmission (DTX). A comfort noise generator is usually used for discontinuous transmission of speech. According to this concept, speech is first classified into active and inactive frames by a voice activity detector (VAD). An example of VAD can be found in [3]. Based on the VAD results, only active speech frames are encoded, decoded and transmitted at the nominal bit rate. During long pauses when only background noise or silence is present, the bit rate is reduced or zeroed and the background noise / silence is encoded and decoded in an episodic and parametric manner. As a result, the average bit rate is significantly reduced. The noise is generated at the decoder side by a comfort noise generator (CNG) during inactive frames. For example, the speech codecs AMR-WB [3] and 3GPP EVS [4], [5] can both operate in DTX mode. An example of an effective CNG is given in [6]. In the IVAS codec, the discontinuous transmission (DTX) system exists in the audio scenario transmitted via parametric codecs of the Directional Audio Codec (DirAC) paradigm or in the Metadata Assisted Spatial Audio (MASA) format (see [7]).

[0006] In Discrete Independent Streams with Metadata (Discrete ISM), the encoder of Discrete ISM accepts audio objects and their associated metadata. Then, on a frame basis, the objects are individually encoded (along with metadata including object directional information (such as azimuth and elevation)), and then the encoding is transmitted to the decoder. The decoder then decodes the individual objects independently and renders them to the specified output structure by applying an amplitude shifting technique that uses the quantized directional information.

[0007] Another concept in the prior art is parametrically coded independent stream with metadata (Param-ISM). Figure 4 An overview of a corresponding encoder is shown, wherein in particular an encoded audio signal 491 and encoded parametric side information 495 , 496 , 497 are depicted.

[0008] The encoder of the parameter ISM (Param-ISM) receives as input the audio objects and the associated metadata. The metadata may include, for example, on a frame basis, the object direction (e.g., azimuth, whose value is, for example, between [-180, 180]; and, for example, elevation, whose value is, for example, between [-90, 90]), which is then quantized and used during the calculation of the stereo downmix (e.g., virtual cardioid or transmission channel). In addition, among the input audio objects, two main objects and the power ratio between the two main objects may be determined, for example, per time / frequency block. Subsequently, the metadata may be quantized and encoded, for example, together with the object indexes of the two main objects (two main objects per time / frequency block).

[0009] The encoded bitstream 490 may, for example, include a stereo downmix / transmission channel 491 (which is separately encoded with the aid of the core codec), an encoded primary object index 495, a quantized and encoded power ratio 496, and quantized and encoded direction information 497 (e.g., azimuth and elevation).

[0010] Figure 5 A simplified overview of the decoder is shown. The decoder receives a bitstream 490 and obtains an encoded stereo downmix / transmission channel 491, an encoded object index 495, an encoded power ratio 496, and an encoded direction information 497. The encoded stereo downmix / transmission channel 491 is then decoded using a core decoder and converted into a time / frequency representation using an analysis filter bank (e.g., a composite low-delay filter bank, CLDFB). The decoded object index may be used, for example, together with decoded and dequantized direction information (e.g., azimuth and elevation and output configurations, such as 5.1, 5.1+4, 7.1, 7.1+4, etc.) to calculate the direction response. The direct response may be provided, for example, together with a transmission channel / stereo downmix represented in time / frequency, a prototype matrix, and a decoded and dequantized power ratio as an input in a covariance synthesis operating in the time / frequency domain. The output of the covariance synthesis is converted from a time / frequency representation to a time domain representation using a synthesis filter (e.g., CLDFB).

[0011] Figure 6 A detailed overview of the covariance synthesis steps is shown without reflecting the dimensionality of the input / output data.

[0012] Covariance synthesis computes a mixing matrix (M) for each time / frequency block, which renders the input transmission channels

[0013] (x=x(k,n)=[X 1 (k,n),X 2 (k,n)] T )

[0014] To the desired speaker configuration

[0015] (y=y(k,n)=[Y 1 (k,n),Y 2 (k,n),Y 3 (k,n),…] T )

[0016] (For example, 5.1 speaker architecture, 7.1 speaker architecture, 7.1+4 speaker architecture, etc.):

[0017] y=Mx

[0018] For the mixing matrix, covariance synthesis can use the prototype matrix, input covariance matrix Cx =xx T And the target covariance matrix C Y The target covariance matrix is ​​calculated with the aid of the signal powers calculated from the transmission channels / stereo downmix, the power ratio and the direct response.

[0019] It is an object of the invention to provide an improved concept for discontinuous transmission of audio content.The object of the invention is solved by the subject-matter of the independent claims.

[0020] According to an embodiment, an audio encoder is provided. The audio encoder includes a transmission signal generator for generating two or more transmission channels of a transmission signal from an audio input, the audio input including at least one of a plurality of audio input objects and a plurality of audio input channels. In addition, the audio encoder includes a voice activity determiner for determining a voice activity decision of the transmission signal, which indicates whether the audio input within the transmission signal exhibits voice activity. In addition, the audio encoder includes a bitstream generator for generating a bitstream based on the audio input. If the voice activity determiner has determined that the transmission signal exhibits voice activity, the bitstream generator is suitable for encoding the two or more transmission channels within the bitstream. If the voice activity determiner has determined that the transmission signal does not exhibit voice activity, the bitstream generator is suitable for encoding information about background noise instead of the two or more transmission channels, wherein the information about background noise includes information about background noise of at least one of the two or more transmission channels or information about background noise of a derived signal, the derived signal depending on at least one of the two or more transmission channels.

[0021] For example, according to an embodiment, the number of transmit channels is less than or equal to the number of input channels.

[0022] In addition, according to an embodiment, a method for audio encoding is provided. The method comprises:

[0023] - generating two or more transmission channels of a transmission signal from an audio input, the audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels.

[0024] - Determining a voice activity decision for the transmit signal, which indicates whether an audio input within the transmit signal exhibits voice activity.

[0025] and

[0026] -Determine the bitstream based on the audio input.

[0027] If it has been determined that the transmitted signal exhibits voice activity, the method includes encoding the two or more transmitted channels within the bitstream. If it has been determined that the transmitted signal does not exhibit voice activity, the method includes encoding information about background noise of at least one of the two or more transmitted channels or information about background noise of a derived signal instead of encoding the two or more transmitted channels, the derived signal being dependent on at least one of the two or more transmitted channels.

[0028] Furthermore, a computer program is provided, which is used to implement the above method when the computer program is executed on a computer or a signal processor.

[0029] In addition, according to an embodiment, an audio decoder is provided. The audio decoder includes an input interface for receiving a bitstream, and the bitstream depends on an audio content including at least one of a plurality of audio objects and a plurality of audio channels. A transmission signal including two or more transmission channels is encoded in the bitstream, and the audio content is encoded in the transmission signal. Alternatively, information about background noise is encoded in the bitstream instead of the transmission signal, and the information about background noise includes information about background noise of at least one of the two or more transmission channels or information about background noise of a derived signal, and the derived signal depends on at least one of the two or more transmission channels. In addition, the audio decoder includes a renderer, which is used to generate one or more audio output signals according to the audio content encoded in the bitstream. If the transmission signal including two or more transmission channels is encoded in the bitstream, the renderer is configured to generate one or more audio output signals according to the two or more transmission channels. If information about background noise is encoded in the bitstream instead of the transmission signal, the renderer is configured to generate one or more audio output signals according to the information about background noise.

[0030] In addition, a method for audio decoding is provided. The method comprises:

[0031] - receiving a bitstream depending on audio content, the audio content comprising a plurality of audio objects and at least one of a plurality of audio channels. A transmission signal comprising two or more transmission channels is encoded in the bitstream. The audio content is encoded in the transmission signal. Alternatively, information about background noise is encoded in the bitstream instead of the transmission signal, and the information about background noise comprises information about background noise of at least one of the two or more transmission channels or information about background noise of a derived signal, the derived signal being dependent on at least one of the two or more transmission channels. And:

[0032] - generating one or more audio output signals based on the audio content encoded in the bitstream.

[0033] If a transmission signal comprising two or more transmission channels is encoded in the bitstream, the generation of the one or more audio output signals is performed in dependence on the two or more transmission channels. If information about background noise is encoded in the bitstream instead of the transmission signal, the generation of the one or more audio output signals is performed in dependence on the information about the background noise.

[0034] In addition, a computer program is provided, which is used to implement the above method when the computer program is executed on a computer or a signal processor.

[0035] Some embodiments are based on the discovery that by combining existing schemes, DTX can be applied independently, for example, for individual streams (e.g. for audio objects or for individual channels, such as stereo downmix / transmission channels). However, this will be incompatible with DTX designed for low bit rate communication, because for more than one object or for a transmission channel or for a downmix with more than one channel, the available number of bits will not be sufficient to effectively describe the inactive parts of the input signal. In addition, such an approach will also face problems, because the individual VAD decisions are not synchronized. Spatial artifacts will be generated as a result.

[0036] In an embodiment, a DTX system for audio scenes described by (audio) objects and their associated metadata is provided.

[0037] Some embodiments provide a DTX system, in particular SID and CNG, for audio objects (also called ISM, ie "Independent Stream with Metadata"), which are parametrically coded, for example as Param-ISM.

[0038] In some embodiments, a drastic reduction in the bit rate requirements for transmitting conversational immersive voice is achieved.

[0039] According to some embodiments, a DTX concept is provided that is extended to immersive speech with spatial cues.

[0040] In some embodiments, two most important objects per time / frequency unit are considered. In other embodiments, more than two most important objects per time / frequency unit are considered, especially for an increased number of input objects. For the readability of text, the embodiments below are mainly described with respect to two main objects per time / frequency unit, but similarly, these embodiments may be extended to more than two main objects per time / frequency unit, for example, in other embodiments.

[0041] Certain embodiments of audio encoders are provided.

[0042] According to an embodiment, an audio encoder for encoding a plurality of (audio) objects and their associated metadata is provided.

[0043] The audio encoder may, for example, comprise a directional information determiner for extracting directional information and a directional information quantizer for quantizing the directional information.

[0044] Furthermore, the audio encoder may, for example, comprise a transmission signal generator (downmixer) for generating a transmission signal (downmix) comprising at least two transmission channels (e.g. downmix channels) from an input audio object and from quantized directional information (e.g. azimuth and elevation) associated with the input audio object.

[0045] Furthermore, the audio encoder may, for example, comprise a decision logic module for combining the individual VAD decisions of the transport channels to calculate an overall decision as to whether a frame is active or not.

[0046] Furthermore, the audio encoder may, for example, comprise a mono signal generator (eg a stereo-to-mono converter) for outputting a mono signal from the transport channel to be encoded in the inactive phase.

[0047] Furthermore, the audio encoder may, for example, comprise an inactivity metadata generator for generating (eg calculating) inactivity metadata to be transmitted during the inactivity phase.

[0048] Furthermore, the audio encoder may, for example, comprise an activity metadata generator for generating (eg calculating) activity metadata to be transmitted during the activity phase.

[0049] Furthermore, the audio encoder may, for example, include a transmission channel encoder configured to generate encoded data by encoding a downmix signal including a transmission channel in an active phase.

[0050] Furthermore, the audio encoder may for example comprise a transport channel silence insertion description generator for generating a silence insertion description of the background noise of the mono signal in inactive phases.

[0051] In addition, the audio encoder may, for example, include a multiplexer for combining active metadata and encoded data into a bitstream during an active phase, and for not sending data or for sending a silence insertion description. Alternatively, the multiplexer may, for example, be configured to combine sending a silence insertion description and inactive metadata during an inactive phase.

[0052] According to an embodiment, the transmission signal generator / downmixer may for example apply a CELP coding scheme (CELP=Code Excited Linear Prediction), or may for example apply an MDCT-based coding scheme (MDCT=Modified Discrete Cosine Transform), or may for example apply a transform combination of these two coding schemes.

[0053] In an embodiment, the active and inactive phases may be determined, for example, by first running a speech activity detector on the transmit / downmix channels individually and then combining the results of the transmit / downmix channels to determine an overall decision.

[0054] According to an embodiment, the mono signal may be calculated from the transmit / downmix channels, eg by adding the transmit channels or eg by selecting the channel with higher long term energy.

[0055] In embodiments, the active and inactive metadata may differ, for example, in quantization resolution, or in the type (nature) of parameters (used).

[0056] According to an embodiment, the quantization resolution of the transmitted directional information and the quantization resolution of the directional information used to calculate the downmix may be different, for example in an inactive phase.

[0057] In an embodiment, the spatial audio input format may be described, for example, by an object and its associated metadata (eg, by a separate stream with metadata).

[0058] Depending on the embodiment, two or more transmission channels may be generated, for example.

[0059] Furthermore, specific embodiments of an audio decoder are provided.

[0060] According to an embodiment, an audio decoder for (decoding and) generating a spatial audio output signal from a bitstream. The bitstream may, for example, exhibit at least one active phase and then at least one inactive phase. Furthermore, the bitstream may, for example, have encoded therein at least a silence insertion descriptor frame (SID), which may, for example, describe background noise characteristics of a transmission / downmix channel and / or spatial image information.

[0061] The audio decoder may, for example, comprise a SID decoder (Silence Insertion Descriptor Decoder), which may, for example, be configured to decode silence insertion descriptor frames of a mono signal.

[0062] Furthermore, the audio decoder may, for example, comprise a mono to stereo converter which may, for example, be configured to generate at least two (downmix) channels from the SID information of the mono signal and control parameters during an inactive phase / mode, which control parameters may, for example, describe characteristics of a stereo downmix / transmission channel, such as scale parameters calculated at the encoder side from the stereo downmix / transmission channel and / or for example wideband coherence or wideband correlation.

[0063] Furthermore, the audio decoder may, for example, comprise a transport channel decoder which may, for example, be configured to reconstruct the transport / downmix channel during the active phase / mode from the bitstream during the active phase.

[0064] Furthermore, the audio decoder may, for example, comprise a (spatial) renderer which may, for example, be configured to reconstruct the spatial output signal from the decoded transmit / downmix channel during an active phase / mode, for example from transmitted activity metadata, for example from reconstructed background noise in the transmit / downmix channel and from transmitted inactive metadata during an inactive phase.

[0065] According to an embodiment, the mono to stereo converter may, for example, comprise a random generator which may, for example, be executed at least twice using different seeds to generate noise and may, for example, process the generated noise using decoded SID information of the mono signal and using control parameters which may, for example, describe characteristics of a stereo downmix / transmission channel, such as scale parameters calculated from the stereo downmix / transmission channel at the encoder side and / or, for example, wideband coherence or wideband correlation.

[0066] In an embodiment, the spatial parameters transmitted in the active phase may, for example, include object index, power ratio (which may, for example, be transmitted in frequency sub-bands) and direction information (eg azimuth and elevation), which may, for example, be transmitted wideband.

[0067] According to an embodiment, the spatial parameters transmitted in the inactive phase may, for example, include directional information (e.g. azimuth and elevation), which may, for example, be a transmitted wideband, and control parameters, which may, for example, describe characteristics of the stereo downmix / transmission channel, such as scale parameters calculated from the stereo downmix / transmission channel at the encoder side and / or for example wideband coherence or wideband correlation.

[0068] In an embodiment, the quantization resolution of the direction information in the inactive phase is different from the quantization resolution of the direction information in the active phase.

[0069] According to an embodiment, the transmission of the control parameters may eg be done in a wideband or may eg be done in a frequency subband, wherein the decision whether to do so in a wideband or in a frequency subband may eg be determined depending on bitrate availability.

[0070] In an embodiment, the renderer may, for example, be configured to perform covariance synthesis.

[0071] The renderer may for example comprise a signal power calculation unit for calculating a reference power according to a transmit / downmix channel for each time / frequency block.

[0072] Furthermore, the renderer may, for example, comprise a direct power calculation unit for scaling the reference power using the transmission power ratio in the active phase and scaling the reference power using a constant scaling factor in the inactive phase.

[0073] Furthermore, the renderer may, for example, comprise a direct response calculation unit for calculating the direct response depending on the quantized direction information of the main object during the active phase or depending on the quantized direction information of all transmitted objects during the inactive phase.

[0074] Furthermore, the renderer may for example comprise an input covariance matrix calculation unit for calculating an input covariance matrix based on the transmit / downmix channels.

[0075] Furthermore, the renderer may, for example, include a target covariance matrix calculation unit for calculating a target covariance matrix based on outputs of the direct response calculation block and the direct power calculation block.

[0076] Furthermore, the renderer may, for example, comprise a mixing matrix calculation unit for calculating a mixing matrix for rendering depending on the input covariance matrix and depending on the target covariance matrix.

[0077] According to an embodiment, the constant scaling factor used during the inactive phase may be determined, for example, in dependence on the number of transmitted objects; or a control parameter may be used, for example.

[0078] In an embodiment, the primary objects may, for example, be a subset of all transmitted objects, and the number of primary objects may, for example, be less than / smaller than the number of transmitted objects.

[0079] According to embodiments, the transport channel decoder may for example include a speech decoder (eg a CELP-based speech decoder) and / or may for example include a generic audio decoder (eg a TCX-based decoder) and / or may for example include a bandwidth extension module.

[0080] Further particular embodiments are provided in the dependent claims.

[0081] In the following, embodiments of the present invention are described in more detail with reference to the accompanying drawings, in which:

[0082] Figure 1 An audio encoder according to an embodiment is shown.

[0083] Figure 2 An audio decoder according to an embodiment is shown.

[0084] Figure 3 A system according to an embodiment is shown.

[0085] Figure 4 An overview of the Param-ISM encoder is shown.

[0086] Figure 5 An overview of the Param-ISM decoder is shown.

[0087] Figure 6 A detailed overview of the covariance synthesis step in Param-ISM is shown without reflecting the dimensionality of the input / output data.

[0088] Figure 7 A block diagram for determining whether a frame is active or inactive is shown according to an embodiment.

[0089] Figure 8 A block diagram of an encoder according to an embodiment is shown.

[0090] Fig. 9 A block diagram of a decoder according to an embodiment is shown.

[0091] Fig.10 A spatial renderer according to an embodiment is shown.

[0092] Fig.11 It shows the generation of a stereo signal using three random seeds - Seed 1, Seed 2 and Seed 3, derived scale factors and control parameters according to an embodiment.

[0093] Fig.12 shows the generation of a stereo signal according to another embodiment, in which the noise N generated by the third random generator for the left channel 3 (k,n) is also used to generate the right channel.

[0094] Figure 1 An audio encoder 100 according to an embodiment is shown.

[0095] The audio encoder 100 comprises a transmit signal generator 110 for generating two or more transmit channels of a transmit signal from an audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels.

[0096] Furthermore, the audio encoder 100 comprises a voice activity determiner 120 for determining a voice activity decision for the transmit signal, which indicates whether an audio input within the transmit signal exhibits voice activity.

[0097] Furthermore, the audio encoder 100 comprises a bitstream generator 130 for generating a bitstream according to an audio input.

[0098] If the voice activity determiner 120 has determined that the transmission signal exhibits voice activity, the bitstream generator 130 is adapted to encode two or more transmission channels within the bitstream.

[0099] If the voice activity determiner 120 has determined that the transmitted signal does not exhibit voice activity, the bitstream generator 130 is adapted to encode information about the background noise instead of the two or more transmission channels, wherein the information about the background noise comprises information about the background noise of at least one of the two or more transmission channels or information about the background noise of a derived signal that depends on at least one of the two or more transmission channels.

[0100] According to an embodiment, the voice activity determiner 120 may, for example, be configured to determine a separate voice activity decision for each of the one or more transmission channels of the transmission signal, the separate voice activity decision indicating whether the audio input within the transmission channel exhibits voice activity. In addition, the voice activity determiner 120 may, for example, be configured to determine the voice activity decision of the transmission signal based on the separate voice activity decision for each of the one or more transmission channels.

[0101] In an embodiment, the voice activity determiner 120 may, for example, be configured to determine a separate voice activity decision for each of two or more transmission channels of the transmission signal, the separate voice activity decision indicating whether the audio input within the transmission channel exhibits voice activity. In addition, the voice activity determiner 120 may, for example, be configured to determine the voice activity decision of the transmission signal based on the separate voice activity decision for each of the two or more transmission channels of the transmission signal.

[0102] According to an embodiment, the voice activity determiner 120 may, for example, be configured to determine that the transmission signal exhibits voice activity if at least one of the two or more transmission channels of the transmission signal exhibits voice activity. In addition, the voice activity determiner 120 may, for example, be configured to determine that the transmission signal does not exhibit voice activity if none of the two or more transmission channels of the transmission signal exhibits voice activity.

[0103] In an embodiment, the audio encoder 100 may, for example, be configured to determine whether to transmit a bitstream in which information about background noise has been encoded, or whether to not generate and transmit the bitstream if the voice activity determiner 120 has determined that the transmitted signal exhibits no voice activity.

[0104] According to an embodiment, the audio encoder 100 may, for example, include a mono signal generator 830 (see Figure 8 ), the mono signal generator for generating a derived signal from at least one of the two or more transmission channels as a mono signal if the voice activity determiner 120 has determined that the transmission signal does not exhibit voice activity. The audio encoder 100 may, for example, include an information generator for generating the information about the background noise as the information about the background noise of the mono signal.

[0105] In an embodiment, the mono signal generator 830 may be configured to generate a mono signal by adding two or more transmission channels or by adding two or more channels derived from two or more transmission channels. Alternatively, the mono signal generator 830 may be configured to generate a mono signal by selecting a transmission channel exhibiting higher energy among two or more transmission channels.

[0106] According to an embodiment, the information generator may, for example, be configured to generate information about background noise of the mono signal as the information about the mono signal.

[0107] In an embodiment, the information generator may, for example, be configured to generate a silence insertion description of the background noise of the monophonic signal as the information on the background noise of the monophonic signal.

[0108] According to an embodiment, the audio encoder 100 may, for example, include a direction information determiner 802 (see Figure 8 ), for determining direction information based on the audio input. The audio encoder 100 may, for example, include a direction information quantizer 804 (see Figure 8 ), for quantizing the direction information to obtain quantized direction information. The bitstream generator 130 may, for example, be configured to encode the quantized direction information in a bitstream.

[0109] In an embodiment, the transmit signal generator 110 may, for example, be configured to generate two or more transmit channels of the transmit signal from the audio input using the direction information.

[0110] According to an embodiment, the audio input may, for example, include a plurality of audio input objects. The direction information may, for example, include information on an azimuth and an elevation of an audio input object among the plurality of audio input objects of the audio input.

[0111] In an embodiment, the audio encoder 100 may, for example, include an activity metadata generator 825 (see Figure 8 ), the activity metadata generator is used to generate metadata when the voice activity determiner 120 has determined that the transmitted signal exhibits voice activity, the metadata including at least one of quantized direction information, object index and power ratio of multiple audio input objects and or multiple audio input channels of the audio input.

[0112] According to an embodiment, the audio input may, for example, include a plurality of audio input objects. The audio encoder 100 may, for example, include an inactive metadata generator 826 (see Figure 8 ), which is used to generate metadata when the voice activity determiner 120 has determined that the transmission signal does not exhibit voice activity, the metadata comprising quantized directional information and control parameters, such as a scaling factor depending on the number of audio input objects in a plurality of audio input objects of the audio input, or a scaling factor depending on the long-term energy of a transmission channel of the transmission signal and / or depending on the coherence or correlation between transmission channels of the transmission signal.

[0113] In an embodiment, the quantization resolution of the directional information, which may be generated, for example, by the inactive metadata generator 826 , is different from the quantization resolution of the directional information, which may be generated, for example, by the active metadata generator 825 .

[0114] In an embodiment, characteristics of metadata that may be generated, for example, by the inactive metadata generator 826 , differ from characteristics of metadata that may be generated, for example, by the active metadata generator 825 .

[0115] According to an embodiment, the audio input may, for example, include a plurality of audio input objects and metadata associated with the audio input objects.

[0116] In an embodiment, the transmission signal generator 110 may, for example, be configured to generate two or more transmission channels of a transmission signal from an audio input, including by downmixing a plurality of audio input objects and at least one of a plurality of audio input channels to obtain a downmix as a transmission signal, which may, for example, include two or more downmix channels as two or more transmission channels.

[0117] According to an embodiment, if the audio input within the transmission signal exhibits no speech activity, the directional information quantizer 804 is configured to determine the quantized directional information such that the quantization resolution of the quantized directional information may, for example, be different from the quantization resolution used for calculating the downmix.

[0118] In an embodiment, the bitstream generator 130 may, for example, be configured to encode a control parameter into the bitstream if the voice activity determiner 120 has determined that the transmitted signal does not exhibit voice activity. The control parameter may, for example, be suitable for controlling the generation of the intermediate signal from random noise. The control parameter may, for example, comprise a plurality of parameter values ​​for a plurality of frequency sub-bands, or wherein the control parameter may, for example, comprise a single wideband control parameter.

[0119] According to an embodiment, the audio encoder 100 may for example be configured to generate the control parameter by selecting, depending on the available bit rate, whether the control parameter may for example comprise multiple parameter values ​​for multiple frequency subbands or whether the control parameter may for example comprise a single wideband control parameter.

[0120] In an embodiment, the transmit signal generator 110 may, for example, be configured to encode the audio input by applying code excited linear prediction or by applying modified discrete cosine transform or by applying a combination of code excited linear prediction and modified discrete cosine transform.

[0121] According to an embodiment, if the audio input includes multiple audio input channels but does not include multiple audio input objects, the number of two or more transmission channels may, for example, be less than the number of multiple audio input channels. If the audio input includes multiple audio input objects but does not include multiple audio input channels, the number of two or more transmission channels may, for example, be less than the number of multiple audio input objects. If the audio input includes both multiple audio input objects and multiple audio input channels, the number of two or more transmission channels may, for example, be less than the sum of the number of multiple audio input channels and the number of multiple audio input objects.

[0122] Alternatively, according to an embodiment, if the audio input includes multiple audio input channels but does not include multiple audio input objects, the number of two or more transmission channels may, for example, be less than or equal to the number of the multiple audio input channels. If the audio input includes multiple audio input objects but does not include multiple audio input channels, the number of two or more transmission channels may, for example, be less than or equal to the number of the multiple audio input objects. If the audio input includes both multiple audio input objects and multiple audio input channels, the number of two or more transmission channels may, for example, be less than or equal to the sum of the number of the multiple audio input channels and the number of the multiple audio input objects.

[0123] Figure 2 An audio decoder 200 according to an embodiment is shown.

[0124] The audio decoder 200 comprises an input interface 210 for receiving a bitstream, the bitstream depending on audio content including a plurality of audio objects and at least one of a plurality of audio channels. A transport signal including two or more transport channels is encoded in the bitstream, and the audio content is encoded in the transport signal. Alternatively, information about background noise is encoded in the bitstream instead of the transport signal, and the information about background noise comprises information about background noise of at least one of the two or more transport channels or information about background noise of a derived signal, the derived signal depending on at least one of the two or more transport channels.

[0125] Furthermore, the audio decoder 200 comprises a renderer 220 for generating one or more audio output signals depending on the audio content encoded in the bitstream.

[0126] If a transport signal comprising two or more transport channels is encoded within the bitstream, the renderer 220 is configured to generate one or more audio output signals depending on the two or more transport channels.

[0127] If information about the background noise is encoded in the bitstream instead of the transmitted signal, the renderer 220 is configured to generate one or more audio output signals in dependence on the information about the background noise.

[0128] According to an embodiment, if the audio content exhibits voice activity, a transmission signal comprising two or more transmission channels may be encoded in the bitstream, for example. If the audio content does not exhibit voice activity, information about background noise instead of the transmission signal may be encoded in the bitstream, for example.

[0129] In an embodiment, the audio decoder 200 may include, for example, a demultiplexer 902, a noise information determiner 920, and a multi-channel generator 930 (see Fig. 9 ). The demultiplexer may, for example, be configured to determine whether the transmitted bitstream corresponds to an active frame or an inactive frame based on the size of the bitstream. If the information about the background noise is encoded in the bitstream, the noise information determiner 920 may, for example, be configured to determine the information about the background noise from the bitstream, the multi-channel generator 930 may, for example, be configured to generate a derived signal from the information about the background noise as an intermediate signal including two or more intermediate channels, and the renderer 220 may, for example, be configured to generate one or more audio output signals based on the two or more intermediate channels of the intermediate signal.

[0130] According to an embodiment, the multi-channel generator 930 may, for example, include a random generator for generating random noise. The multi-channel generator 930 may, for example, be configured to generate two or more intermediate channels according to the random noise.

[0131] In an embodiment, the multi-channel generator 930 may, for example, be configured to adjust the random noise in dependence on information about the background noise to obtain the shaped noise. The multi-channel generator 930 may, for example, be configured to generate two or more intermediate channels from the shaped noise.

[0132] According to an embodiment, the multi-channel generator 930 may, for example, be configured to run the random generator at least twice using different seeds to obtain the random noise.

[0133] In an embodiment, the multi-channel generator 930 may, for example, be configured to generate two or more intermediate channels based on random noise and based on control parameters (e.g. depending on the ratio and / or the coherence or correlation of the transmission channels of the transmission signal), wherein the control parameters may, for example, be encoded in the bitstream as part of inactive metadata.

[0134] According to an embodiment, the control parameter may, for example, be encoded in a bitstream and may, for example, include multiple parameter values ​​for multiple sub-bands, and the multi-channel generator 930 may, for example, be configured to generate each of the multiple sub-bands of the two or more intermediate channels based on a parameter value from the multiple parameter values ​​of the control parameter associated with the sub-band.

[0135] In an embodiment, the control parameters may, for example, be encoded in a bitstream, wherein the control parameters may, for example, comprise a single wideband control parameter.

[0136] According to an embodiment, the multi-channel generator 930 may be configured, for example, to generate two or more intermediate channels by generating a first random noise portion of random noise using a random generator using a first seed, by generating a first intermediate channel of the two or more intermediate channels based on the first random noise portion, by generating a second random noise portion of random noise using a random generator using a second seed different from the first seed, and by generating a second intermediate channel of the two or more intermediate channels based on the second random noise portion.

[0137] According to an embodiment, the multi-channel generator 930 may be configured, for example, to generate a first intermediate channel of the two or more intermediate channels based on a first random noise portion, based on a third noise portion, and based on a control parameter, such as a scaling factor and / or, for example, coherence or correlation. In addition, the multi-channel generator 930 may be configured, for example, to generate a second intermediate channel of the two or more intermediate channels based on a second random noise portion, based on a third noise portion, and based on a control parameter, such as a scaling factor and / or, for example, coherence or correlation. The multi-channel generator 930 may be configured, for example, to generate a first random noise portion of random noise using a random generator using a first seed, to generate a second random noise portion of random noise using a random generator using a second seed, and to generate a third random noise portion of random noise using a random generator using a third seed, wherein the second seed is different from the first seed, and wherein the third seed is different from the first seed and different from the second seed.

[0138] In an embodiment, the multi-channel generator 930 may, for example, be configured to generate two or more intermediate channels by generating a first intermediate channel of the two or more intermediate channels based on random noise and by generating a second intermediate channel of the two or more intermediate channels from the first intermediate channel of the two or more intermediate channels.

[0139] According to an embodiment, the multi-channel generator 930 may be configured, for example, to generate a second intermediate channel of the two or more intermediate channels, such that the second intermediate channel of the two or more intermediate channels may be, for example, the same as the first intermediate channel of the two or more intermediate channels. Alternatively, the multi-channel generator 930 may be configured, for example, to generate a second intermediate channel of the two or more intermediate channels by modifying the first intermediate channel of the two or more intermediate channels.

[0140] In an embodiment, the renderer 220 may, for example, be configured to generate two or more audio output signals as the one or more audio output signals.

[0141] According to an embodiment, the audio content may, for example, include a plurality of audio objects. If the audio content exhibits speech activity, a plurality of audio object indexes associated with the plurality of audio objects, a plurality of power ratios associated with the plurality of audio objects of the plurality of sub-bands, and broadband directional information of the plurality of audio objects may, for example, be encoded in the bitstream, and the renderer 220 may, for example, be configured to generate one or more audio output signals according to the plurality of audio object indexes, according to the plurality of power ratios, and according to the broadband directional information of the plurality of audio objects.

[0142] In an embodiment, the audio content may, for example, include a plurality of audio objects. If the audio content does not exhibit voice activity, broadband directional information and control parameters of the plurality of audio objects may, for example, be encoded in the bitstream, and the renderer 220 may, for example, be configured to generate one or more audio output signals based on the broadband directional information and based on all object indices and a constant power ratio, wherein the constant power ratio depends on the number of transmitted objects.

[0143] According to an embodiment, a first quantization resolution of the wideband directional information encoded in the bitstream when the audio content exhibits voice activity may, for example, be different from a second quantization resolution of the wideband directional information when the audio content does not exhibit voice activity.

[0144] In an embodiment, the renderer 220 may include, for example, a signal power calculation unit 951 (see Fig.10 ), the signal power calculation unit is used to calculate the reference power according to two or more transmission channels of each time-frequency block in the plurality of time-frequency blocks. In addition, the renderer 220 may, for example, include a direct power calculation unit 952 (see Fig.10 ), the direct power calculation unit is used to proportionally adjust the reference power using the transmitted power ratio encoded in the bitstream when the audio content exhibits voice activity, and proportionally adjust the reference power using the scale factor encoded in the bitstream when the audio content does not exhibit voice activity to obtain the proportionally adjusted reference power. In addition, the renderer 220 can, for example, be configured to generate one or more audio output signals based on the proportionally adjusted reference power.

[0145] According to an embodiment, the renderer 220 may include, for example, a direct response calculation unit 953 for calculating a direct response (see Fig.10 ), wherein the renderer 220 may, for example, be configured to calculate the direct response based on the quantized direction information of the main object that is an appropriate subset of the plurality of audio objects of the audio content when the audio content exhibits voice activity, wherein the renderer 220 may, for example, be configured to calculate the direct response based on the quantized direction information of all audio objects of the audio content when the audio content does not exhibit voice activity, wherein the quantized direction information may, for example, be encoded in the bitstream. The renderer 220 may, for example, be configured to generate one or more audio output signals based on the direct response.

[0146] In an embodiment, the renderer 220 may include, for example, an input covariance matrix calculation unit 954 (see Fig.10 ), the input covariance matrix calculation unit is used to calculate the input covariance matrix according to two or more transmission channels. In addition, the renderer 220 may, for example, include a target covariance matrix calculation unit 955 (see Fig.10), the target covariance matrix calculation unit is used to calculate the target covariance matrix based on the direct response and based on the proportionally adjusted reference power. In addition, the renderer 220 may, for example, include a mixing matrix calculation unit 956 (see Fig.10 ), the mixing matrix calculation unit is used to calculate a mixing matrix for rendering according to the input covariance matrix and according to the target covariance matrix. The renderer 220 can be configured to generate one or more audio output signals according to the mixing matrix.

[0147] According to an embodiment, the renderer 220 may, for example, be configured to generate one or more transmission channels of a transmission signal by applying code excited linear prediction, by applying modified discrete cosine transform or an inverse transform of modified discrete cosine transform, or by applying a combination of code excited linear prediction and modified discrete cosine transform.

[0148] According to an embodiment, if the audio content includes multiple audio channels but does not include multiple audio objects, the number of two or more transmission channels may be, for example, less than the number of multiple audio channels. If the audio content includes multiple audio objects but does not include multiple audio channels, the number of two or more transmission channels may be, for example, less than the number of multiple audio objects. If the audio content includes both multiple audio objects and multiple audio channels, the number of two or more transmission channels may be, for example, less than the sum of the number of multiple audio channels and the number of multiple audio objects.

[0149] Alternatively, according to an embodiment, if the audio content includes multiple audio channels but does not include multiple audio objects, the number of two or more transmission channels may be, for example, less than or equal to the number of multiple audio channels. If the audio content includes multiple audio objects but does not include multiple audio channels, the number of two or more transmission channels may be, for example, less than or equal to the number of multiple audio objects. If the audio content includes both multiple audio objects and multiple audio channels, the number of two or more transmission channels may be, for example, less than or equal to the sum of the number of multiple audio channels and the number of multiple audio objects.

[0150] Figure 3 A system according to an embodiment is shown. The system comprises an audio encoder 100 according to one of the above described embodiments and an audio decoder 200 according to one of the above described embodiments.

[0151] The audio encoder 100 is configured to generate a bitstream from an audio input.

[0152] The audio decoder 200 is configured to generate one or more audio output signals from a bitstream.

[0153] Hereinafter, embodiments are described in detail.

[0154] According to an embodiment, the DTX system (eg, an encoder thereof) may be configured to determine an overall decision of whether a frame is active or inactive based on individual decisions of stereo downmix channels and / or based on individual audio objects, for example.

[0155] A DTX system (eg, an encoder thereof) may, for example, be configured to transmit a mono signal to a decoder together with inactive metadata using a silence insertion descriptor (SID).

[0156] Furthermore, the DTX system (eg, a decoder thereof) may, for example, be configured to generate a transport channel / downmix comprising at least two channels using a comfort noise generator (CNG) according to the SID information of the mono-only signal.

[0157] Furthermore, the DTX system (eg a decoder thereof) may for example be configured to post-process the generated transport channel / downmix applying control parameters, wherein the control parameters may for example be calculated at the encoder side from the stereo downmix / transmit channel.

[0158] Furthermore, the DTX system (eg, a decoder thereof) may render the multi-channel transmit signal to a defined output structure, for example using modified covariance synthesis.

[0159] Hereinafter, other specific embodiments are described.

[0160] Figure 7 A block diagram for determining whether a frame is active or inactive according to an embodiment is shown. The overall decision is based on individual decisions of the transmit channel / downmix channel.

[0161] exist Figure 7 In the embodiment of the present invention, the transmission signal generator (eg, downmixer) 710 may be configured to receive an audio object and its associated quantized direction information (eg, azimuth and elevation).

[0162] A transmission signal (eg, downmix (DMX)) for a first transmission channel (eg, a left downmix channel) DMX L and a transmission signal DMX for a second transmission channel (eg, a right downmix channel) R It can be generated, for example, as follows:

[0163]

[0164] Where N is the total number of input objects, k is the sample index and i is the object index

[0165] w L [i]=0.5+0.5*cos(azimuth[i]-π / 2)

[0166] w R[i]=0.5+0.5*cos(azimuth[i]+π / 2)

[0167] In another embodiment, two transport channels (eg, downmix channels) may be generated, for example, using the downmix matrix D as follows:

[0168]

[0169] where obj 1 …obj N Represents audio object 1 to audio object N.

[0170] also, Figure 7 A decision logic module 720 is depicted, which includes individual decision logic 722 and overall decision logic 725 .

[0171] exist Figure 7 In the embodiment, the separate decision logic 722 may be configured to determine whether the separate channel is active or inactive. The separate decision as to whether each of the two (or more) transmission channels is active or inactive may be indicated, for example, by a (eg, internal) tag.

[0172] In an embodiment, the separate decision logic 722 may, for example, be configured to receive two (or more) transmission channels as input. The separate decision logic 722 may, for example, be configured to determine a DMX for the two (or more) transmission channels, for example by analyzing the transmission channels. L 、DMX R For each delivery channel in a channel, determine whether the delivery channel exhibits voice activity.

[0173] In another embodiment, the separate decision logic 722 may, for example, analyze all audio input channels or all audio input objects used by the transmit signal generator 710 to form two (or more) transmit channel DMX L 、DMX R . For example, if the separate decision logic 722 detects voice activity in at least one of the audio input channels or audio input objects, the separate decision logic 722 may, for example, infer that there is voice activity in each transmission channel, and may, for example, infer that each transmission channel is active. For example, if the separate decision logic 722 that detects voice activity does not detect voice activity in any of the audio input channels or audio input objects used to generate each transmission channel, the separate decision logic 722 may, for example, infer that there is no voice activity in each transmission channel, and may, for example, infer that each transmission channel is inactive.

[0174] In addition, Figure 7In the embodiment of the present invention, the overall decision logic 725 may be configured to receive individual decisions (e.g., for transport channels) as input, and may be configured to determine the overall decision based on the individual decisions. For example, the overall decision logic 725 may indicate the decision using DTX_FLAG, for example. The overall decision logic may determine the overall decision, for example, according to the following Table 1, which depicts the frame-by-frame decision based on the individual downmix decisions on a frame-by-frame basis:

[0175] Table 1

[0176] Activity in the first delivery channel (D_L) Activity in the second delivery channel (D_R) Overall decision (Decision_Overall) Active Active Active Inactive Active Active Active Inactive Active Inactive Inactive Inactive

[0177] For example, the overall decision may be determined, for example, by using a hysteresis buffer with a predefined size. Using a hysteresis buffer helps to avoid artifacts caused by frequent switching between active and inactive parts. For example, a hysteresis buffer of size 10 may, for example, require 10 frames before switching from an active decision to an inactive decision.

[0178] The following is an example pseudo code to determine the overall decision:

[0179] Shift the hysteresis buffer by one step, e.g.

[0180] buffer_decision[i]=buffer_decision[i+1]

[0181] Where i = 0, 1, 2...(Buff_size-1)

[0182] Buff_decision[buff_size]=Decision_Overall

[0183] Decision_Overall may be calculated, for example, as shown in Table 1.

[0184] The overall decision may be calculated, for example, as outlined in the following pseudocode:

[0185] DTX_Flag = 1;

[0186] for(i=0;i <buff_size;i++)

[0187] {

[0188] DTX_Flag=DTX_Flag&&buffer_decision[i];

[0189] }.

[0190] In this pseudocode, DTX_Flag=1 means "inactive", and DTX_FLAG=0 means "active".

[0191] Figure 8 An audio encoder 800 according to an embodiment is shown. Figure 8 The audio encoder may be implemented for example Figure 1 A specific embodiment of the audio encoder 100. Specifically, Figure 8 A block diagram of an encoder is shown, which may, for example, be configured to receive an input audio object and its associated metadata.

[0192] Furthermore, the audio encoder 800 may, for example, include a transmission signal generator (eg, a downmixer) 810 (eg, a transmission channel) for generating a downmix. Figure 7 The downmix comprises at least two channels from an input audio object and from quantized directional information (eg, azimuth and elevation) associated with the input audio object.

[0193] Furthermore, the audio encoder 800 may, for example, include a speech activity determiner, for example, implemented with a decision logic module 820 (eg, Figure 7 A voice activity determiner of the decision logic module 720) is used to combine the individual VAD decisions of the transmission channels to calculate an overall decision as to whether a frame is an active frame.

[0194] A stereo downmix may be computed from the input audio objects in the transmit signal generator 810, for example, using quantized directional information (eg, azimuth and elevation).

[0195] The stereo downmix is ​​then fed back into decision logic module 820, where a decision as to whether a frame is active or inactive may be determined, for example, based on the above logic. Decision logic module 820 may, for example, include individual decision logic 722 and overall decision logic 725 as described above.

[0196] If the decision logic module 820 has determined "active" as the overall decision (for an active frame), then Figure 8 The encoder in Figure 4 The encoder of provides a more efficient method. For the active downmix, the two channels of the stereo downmix can be encoded, for example, independently by the transport channel encoder together with metadata, as described in Table 2 (see below).

[0197] In contrast, if the decision logic module 820 has determined "inactive" as the overall decision (for inactive frames), the SID bitrate (e.g., 4.4 kbps or 5.2 kbps) will be too low to efficiently transmit the two channels of the stereo downmix and the active metadata. Therefore, for the sporadic / intermittently transmitted SID frames, the metadata bitrate may be, for example, 1.85 kbps or 2.45 kbps, and may, for example, include coarsely quantized directional information (e.g., azimuth and elevation) and control parameters that control the spatiality of the background noise and are derived from the stereo downmix / transmit signal, such as scale factors and / or, for example, coherence or correlation.

[0198] In an embodiment, during inactive frames, transmission of object index and power ratio may not occur.The main motivation for not transmitting object index or power ratio during inactive frames is that the background noise is assumed not to have any specific direction and is diffuse in nature.

[0199] In addition, the audio encoder 800 may, for example, include a transmission channel silence insertion description generator 840 for generating a silence insertion description of the background noise of the mono signal in an inactive phase. The transmission channel SID generator (transmission channel SID encoder) 840 may, for example, operate at 2.4 kbps and may, for example, receive a mono downmix as input.

[0200] In addition, the audio encoder 800 may, for example, include a mono signal generator (e.g., a stereo to mono converter) 830 for outputting a mono signal from a transport channel to be encoded in an inactive phase. The conversion of a stereo downmix to a mono downmix may, for example, be performed by the mono signal generator (e.g., a stereo to mono converter) 830.

[0201] In an embodiment, a downmix (eg, stereo to mono conversion) may be implemented, for example, as an addition of two stereo transmit / downmix channels, for example:

[0202] M=DMZ L +DMX R

[0203] In another embodiment, the downmix (e.g., stereo to mono conversion) may be implemented, for example, as a transmission of only one channel of the stereo downmix. The decision of which channel to select may, for example, depend on the (e.g., long-term) energies of the individual channels of the stereo downmix. For example, the channel with the higher long-term energy may be selected:

[0204]

[0205] Among them, LE Lindicates the long-term energy of the first (eg, left) channel, and LE R Indicates the long-term energy of the second (eg, right) channel.

[0206] Table 2 depicts metadata that may be transmitted, for example, during active and inactive frames:

[0207] Table 2

[0208]

[0209] Figure 8 The audio encoder 800 may, for example, include a directional information extractor 802 for extracting directional information and a directional information quantizer 804 for quantizing the directional information.

[0210] Furthermore, the audio encoder 800 may, for example, include an inactive metadata generator 826 for generating (eg, calculating) inactive metadata to be transmitted during the inactive phase.

[0211] Furthermore, the audio encoder 800 may, for example, include an activity metadata generator 825 for generating (eg, calculating) activity metadata to be transmitted during the activity phase.

[0212] Furthermore, the audio encoder 800 may include, for example, a transmission channel encoder 828 configured to generate encoded data by encoding a downmix signal including a transmission channel in an active phase.

[0213] In addition, the audio encoder 800 may, for example, include a bitstream generator, which may, for example, be implemented as a multiplexer 850, for combining (e.g., encoding) the activity metadata with the encoded data (e.g., two or more transmission channels) into a bitstream during the active phase, and for sending no data or for sending silence insertion descriptions. Alternatively, the multiplexer 850 may, for example, be configured to send silence insertion descriptions and inactive metadata in combination during the inactive phase.

[0214] Fig. 9 An audio decoder 900 according to an embodiment is shown. Fig. 9 The audio decoder 900 may for example implement Figure 2 A specific embodiment of the audio decoder 200.

[0215] The audio decoder 900 may receive a bitstream, for example, via an input interface, which may be implemented, for example, as a demultiplexer 902 .

[0216] Fig. 9The audio decoder 900 may, for example, comprise a transport channel decoder 910 which may, for example, be configured to reconstruct a transport / downmix channel during an active phase / mode from the bitstream during the active phase.

[0217] Furthermore, the audio decoder 900 may, for example, comprise a noise information determiner, for example implemented as a SID decoder (Silence Insertion Descriptor Decoder) 920, which may, for example, be configured to decode silence insertion descriptor frames of a mono signal.

[0218] Furthermore, the audio decoder 900 may, for example, comprise a multi-channel generator 930, for example implemented as a mono to stereo converter 930, which may, for example, be configured to generate at least two (downmix) channels from SID information of the mono signal and from control parameters during an inactive phase / mode.

[0219] also, Fig. 9 The audio decoder 900 may, for example, include a filter bank analysis module 940 .

[0220] Furthermore, the audio decoder 900 may, for example, comprise a (spatial) renderer 950, which may, for example, be configured to: reconstruct the spatial output signal during an active phase / mode based on the decoded transmit / downmix channel and, for example, based on the transmitted activity metadata; and to reconstruct the spatial output signal, for example, based on the reconstructed background noise in the transmit / downmix channel during an inactive phase and, for example, based on the transmitted inactive metadata.

[0221] Fig. 9 The audio decoder 900 may, for example, include a synthesis module for performing (eg frequency band) synthesis of a spatial output signal of the renderer 950 .

[0222] Fig. 9 The audio decoder 900 may, for example, further comprise a voice activity information determiner 905 for determining, for example based on VAD data in the bitstream, whether the decoder shall operate in an active or inactive form (in an active mode or in an inactive mode).

[0223] In the currently described activity mode (in the activity form), Fig. 9 The decoder described in Figure 5 The decoder described in is more efficient.

[0224] Fig.10 A spatial renderer, eg for covariance rendering, according to an embodiment is shown. Fig. 9 The renderer 950 shown in FIG. 9 may be implemented, for example, as Fig.10 Space renderer.

[0225] The renderer may for example comprise a signal power calculation unit 951 for calculating a reference power according to a transmit / downmix channel of each time / frequency block.

[0226] Furthermore, the renderer may, for example, include a direct power calculation unit 952 for: scaling the reference power in an active phase using the transmitted power ratio; and scaling the reference power in an inactive phase using, for example, a constant scaling factor depending on the number of transmitted objects or a scaling factor transmitted, for example, as part of metadata, or for example no scaling.

[0227] Furthermore, the renderer may, for example, comprise a direct response calculation unit 953 for calculating a direct response depending on the quantized direction information of the main object during the active phase or depending on the quantized direction information of all transmitted objects during the inactive phase.

[0228] Furthermore, the renderer may, for example, comprise an input covariance matrix calculation unit 954 for calculating an input covariance matrix based on the transmit / downmix channels.

[0229] In addition, the renderer may, for example, include a target covariance matrix calculation unit 955, which is used to calculate the target covariance matrix based on the output of the direct power calculation block 952 and based on the output of the direct response calculation block 953 (or based on the calculated covariance matrix that depends on the output of the direct response calculation block 953).

[0230] Furthermore, the renderer may, for example, comprise a mixing matrix calculation unit 956 for calculating a mixing matrix for rendering depending on the input covariance matrix and depending on the target covariance matrix.

[0231] For example, for a mixing matrix, covariance synthesis can use the prototype matrix, the input covariance matrix C x =xx T And the target covariance matrix C Y As reference Figure 6 As described.

[0232] Furthermore, the renderer may, for example, include an amplitude shifting unit 957 for performing amplitude shifting on the transmission channel according to the mixing matrix calculated by the mixing matrix calculating unit 956 .

[0233] Fig.10 The spatial renderer for covariance synthesis based rendering depicted in FIG. 4 may use, for example, activity metadata such as quantized direction information, object index, and power ratio. The covariance rendering is thus comparable to Figure 3 The covariance rendering shown in is more efficient.

[0234] Fig. 9 The transport channel decoder 910 may, for example, decode the two channels of the stereo downmix in the bitstream independently. The stereo downmix may then be fed back into the filter bank analysis module 940, for example, before being provided as input to the covariance synthesis.

[0235] In the presently described inactive mode (in the inactive mode), the SID decoder 920 and the mono-to-stereo converter 930 may, for example, employ the encoded SID information of the mono channel to generate a stereo signal with some spatial decorrelation.

[0236] According to an embodiment, an efficient implementation of mono to stereo conversion may be used, for example, which may run the random generator twice using different seeds. In an embodiment, the generated noise may be adjusted, for example, using the SID information of the mono channel. Thus, a stereo signal (with zero coherence) is generated.

[0237] In another embodiment, a mono channel may be duplicated into two stereo channels, for example (however, this has the disadvantage of causing spatial collapse and coherence to be unity).

[0238] In a preferred embodiment, to generate a stereo signal with coherence and energy similar to the input stereo downmix Control parameters such as coherence and / or correlation and scaling factors may be used, for example, and may be transmitted, for example, as part of the inactive metadata.

[0239]

[0240] in

[0241] s L (n) = s

[0242] s R (n) = 1-s

[0243] where k is the frequency index, n is the sample index, c(n) is the coherence or relevance of the partial transmission as inactive metadata, and s L (n) and R (n) is the scale factor derived from the scale factor s transmitted as part of the inactive metadata, N 1 (k,n),N 2 (k,n) and N 3 (k,n) is the random noise generated by different random generators using seed 1, seed 2 and seed 3 respectively.

[0244] Since the inactive metadata does not include the power ratio and the object index, during direct power calculation, a scaling factor that may depend on the number of objects may be used instead of the power ratio, for example. Alternatively, a scaling factor transmitted as part of the inactive metadata may be used, for example, instead of the power ratio.

[0245] Fig.11 The generation of a stereo signal using three random seeds - Seed 1, Seed 2 and Seed 3, derived scaling factors and control parameters according to an embodiment is shown.

[0246] also, Fig.11 A random generator is shown, which comprises a random generator unit 1 and a random generator unit 3 for generating a left channel and a random generator unit 2 and a further random generator unit 3 for generating a right channel.

[0247] exist Fig.11 In the example, the random generator unit 3 for generating the left channel and the random generator unit 3 for generating the right channel receive the same seed - seed 3, and thus can generate the same random noise N 3 (k,n).

[0248] Fig.12 shows the generation of a stereo signal according to another embodiment, in which the random generator unit 3 for the left channel generates a noise N 3 (j,n) is also used for the right channel. In other words, Fig.12 The random generator includes a random generator unit 1, a random generator unit 2 and only a single random generator unit 3.

[0249] In another embodiment, the random generator may include only a single random generator unit, which may be used to sequentially generate random noise N in response to receiving seed 1, seed 2 and seed 3 respectively. 1 (k,n),N 2 (k,n) and N 3 (k,n).

[0250] In other embodiments, the above concept is similarly applied to generate a multi-channel signal having more than two channels.

[0251] Additionally, the direct response may be calculated, for example, using the direction information of all objects instead of just the primary object.

[0252] Embodiments allow extending DTX to spatial audio codecs using independent streams with metadata (ISM) in an efficient manner. Spatial audio codecs maintain high perceptual fidelity with respect to background noise even for inactive frames, whereby transmission may be interrupted, for example, to save communication bandwidth.

[0253] Decoder-side transport channels with a channel number greater than one may be generated, for example, only by a comfort noise generator (CNG) from the transport mono signal, such that it exhibits a spatial image according to the SID information. The generated transport channels may then be fed back into the covariance synthesis module, for example, together with the direct response, equal power ratios and prototype matrices calculated from the directional information of all audio objects, for rendering to the desired output structure.

[0254] Although some aspects have been described herein in the context of an apparatus, it is apparent that these aspects also represent a description of a corresponding method, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, the aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.

[0255] According to certain implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. The implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM or flash memory, on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system to perform the corresponding method. Therefore, the digital storage medium may be computer readable.

[0256] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0257] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.The program code may, for example, be stored on a machine readable carrier.

[0258] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0259] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0260] Therefore, another embodiment of the inventive method is a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory.

[0261] Therefore, another embodiment of the inventive method is a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.The data stream or the sequence of signals may, for example, be configured to be transmitted via a data communication connection (eg via the Internet).

[0262] A further embodiment comprises a processing means, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0263] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0264] Another embodiment according to the invention comprises an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may, for example, comprise a file server for transmitting the computer program to the receiver.

[0265] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may be used in conjunction with a microprocessor to perform one of the methods described herein. Typically, these methods are preferably performed by any hardware device.

[0266] The devices described herein may be implemented using hardware devices, or using computers, or using a combination of hardware devices and computers.

[0267] The methods described herein may be performed using a hardware device or using a computer or using a combination of a hardware device and a computer.

[0268] The above embodiments are only used to illustrate the principles of the present invention. It is to be understood that modifications and variations of the arrangements and details described herein will be apparent to other persons skilled in the art. Therefore, it is intended that the invention be limited only by the scope of the pending patent claims and not by the specific details presented by the description and explanation of the embodiments herein.

[0269] refer to:

[0270] [1]WO 2022 / 079049 A2,A.“Apparatus and method for encoding a pluralityof audio objects and apparatus and method for decoding using two or morerelevant audio objects”.

[0271] [2]WO 2022 / 079044 A1“Apparatus and method for encoding a plurality ofaudio objects using direction information during a downmixing or apparatusand method for decoding using an optimized covariance synthesis”.

[0272] [3]3GPP TS26.194;Voice Activity Detector(VAD);-3GPP technicalspecification Retrieved on 2009-06-17.

[0273] [4]3GPP TS26.449,"Codec for Enhanced Voice Services(EVS);ComfortNoise Generation(CNG)Aspects".

[0274] [5]3GPP TS26.450,"Codec for Enhanced Voice Services(EVS);Discontinuous Transmission(DTX)".

[0275] [6]A.Lombard,S.Wilde,E.Ravelli,S. G.Fuchs and M.Dietz,"Frequency-domain Comfort Noise Generation for Discontinuous Transmission in EVS,"2015IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP),Brisbane,QLD,2015,pp.5893-5897,doi:10.1109 / ICASSP.2015.7179102.

[0277] [7]WO 2022 / 022876 A1”Apparatus,method and computer program forencoding an audio signal or for decoding an encoded audio scene”.

Claims

1. An audio decoder (200 ; 900), including: An input interface (210; 902) for receiving a bitstream, the bitstream depending on an audio content comprising a plurality of audio objects and at least one of a plurality of audio channels; wherein a transport signal comprising two or more transport channels is encoded in the bitstream and the audio content is encoded in the transport signal; or wherein information about background noise is encoded in the bitstream instead of the transport signal, wherein the information about background noise comprises information about background noise of at least one of the two or more transport channels or comprises information about background noise of a derived signal, the derived signal depending on at least one of the two or more transport channels; and a renderer (220; 950) for generating one or more audio output signals based on the audio content encoded in the bitstream; wherein if the transmission signal comprising the two or more transmission channels is encoded in the bitstream, the renderer (220; 950) is configured to generate the one or more audio output signals in dependence on the two or more transmission channels, and Wherein, if the information about the background noise is encoded in the bitstream instead of the transmission signal, the renderer (220; 950) is configured to generate the one or more audio output signals according to the information about the background noise.

2. An audio decoder (200; 900) as claimed in claim 1, in, If the audio content exhibits voice activity, the transport signal comprising the two or more transport channels is encoded within the bitstream; and Wherein, if the audio content does not exhibit voice activity, the information about the background noise is encoded in the bitstream instead of the transmission signal.

3. An audio decoder (200; 900) as claimed in claim 1 or 2, wherein the audio decoder (200; 900) comprises a noise information determiner (920) and a multi-channel generator (930), in, If the information about the background noise is encoded in the bitstream, the noise information determiner (920) is configured to determine the information about the background noise from the bitstream, the multi-channel generator (930) is configured to generate the derived signal from the information about the background noise as an intermediate signal including two or more intermediate channels, and the renderer (220; 950) is configured to generate the one or more audio output signals based on the two or more intermediate channels of the intermediate signal.

4. An audio decoder (200; 900) as claimed in claim 3, wherein the multi-channel generator (930) comprises a random generator for generating random noise, The multi-channel generator (930) is configured to generate the two or more intermediate channels based on the random noise generated by the random generator.

5. An audio decoder (200; 900) as claimed in claim 4, wherein the multi-channel generator (930) is configured to adjust the random noise according to the information about the background noise to obtain the shaped noise, Wherein the multi-channel generator (930) is configured to generate the two or more intermediate channels from the shaped noise.

6. An audio decoder (200; 900) as claimed in claim 4 or 5, The multi-channel generator (930) is configured to run the random generator at least twice using different seeds to obtain the random noise.

7. The audio decoder (200; 900) as claimed in any one of claims 4 to 6, The multi-channel generator (930) is configured to generate the two or more intermediate channels based on the random noise and based on control parameters encoded in the bitstream, for example, wherein the control parameters include for example scaling factors and / or for example coherence or correlation.

8. An audio decoder (200; 900) as claimed in claim 7, wherein at least one of the control parameters is encoded within the bitstream and comprises a plurality of parameter values ​​for a plurality of frequency sub-bands, and wherein the multi-channel generator (930) is configured to generate each of the plurality of frequency sub-bands of the two or more intermediate channels based on a parameter value from the plurality of parameter values ​​for at least one of the control parameters associated with the frequency sub-band.

9. The audio decoder (200; 900) as claimed in claim 7, The control parameter is encoded in the bitstream, wherein the control parameter is a single wideband control parameter.

10. The audio decoder (200; 900) as claimed in any one of claims 4 to 9, The multi-channel generator (930) is configured to generate the two or more intermediate channels by: generating a first random noise portion of the random noise using the random generator using a first seed, and generating a first intermediate channel of the two or more intermediate channels based on the first random noise portion; generating a second random noise portion of the random noise using the random generator using a second seed different from the first seed, and generating a second intermediate channel of the two or more intermediate channels based on the second random noise portion.

11. The audio decoder (200; 900) according to any one of claims 4 to 9, further according to claim 7, wherein the multi-channel generator (930) is configured to generate a first intermediate channel of the two or more intermediate channels in dependence on the first random noise portion, in dependence on the third noise portion and in dependence on the control parameters, such as the scaling factor and the coherence and / or correlation, wherein the multi-channel generator (930) is configured to generate a second intermediate channel of the two or more intermediate channels based on the second random noise portion, based on the third noise portion and based on the control parameter, wherein the multi-channel generator (930) is configured to generate the first random noise portion of the random noise using the random generator using a first seed, wherein the multi-channel generator (930) is configured to generate the second random noise portion of the random noise using the random generator using a second seed, and wherein the multi-channel generator (930) is configured to generate the third random noise portion of the random noise using the random generator using a third seed, Wherein the second seed is different from the first seed, and wherein the third seed is different from the first seed and different from the second seed.

12. The audio decoder (200; 900) as claimed in any one of claims 4 to 9, Wherein the multi-channel generator (930) is configured to generate the two or more intermediate channels by: generating a first intermediate channel of the two or more intermediate channels based on the random noise, and generating a second intermediate channel of the two or more intermediate channels from the first intermediate channel of the two or more intermediate channels.

13. An audio decoder (200; 900) as claimed in claim 12, wherein the multi-channel generator (930) is configured to generate the second intermediate channel of the two or more intermediate channels such that the second intermediate channel of the two or more intermediate channels is identical to the first intermediate channel of the two or more intermediate channels, or The multi-channel generator (930) is configured to generate the second intermediate channel of the two or more intermediate channels by modifying the first intermediate channel of the two or more intermediate channels.

14. An audio decoder (200; 900) as claimed in any preceding claim, Wherein the renderer (220; 950) is configured to generate the two or more audio output signals as the one or more audio output signals.

15. An audio decoder (200; 900) as claimed in any preceding claim, wherein the audio content includes the plurality of audio objects, in, If the audio content exhibits voice activity, multiple audio object indexes associated with the multiple audio objects, multiple power ratios associated with the multiple audio objects for multiple sub-bands, and wide-band directional information of the multiple audio objects are encoded in the bitstream, and the renderer (220; 950) is configured to generate the one or more audio output signals based on the multiple audio object indexes, based on the multiple power ratios and based on the wide-band directional information of the multiple audio objects.

16. The audio decoder (200; 900) of any preceding claim, further comprising: wherein the audio content includes the plurality of audio objects, in, If the audio content does not exhibit voice activity, wideband directional information of the plurality of audio objects and the control parameters are encoded in the bitstream, and the renderer (220; 950) is configured to generate the one or more audio output signals based on the wideband directional information.

17. An audio decoder (200; 900) as claimed in claim 15 or claim 16, in, A first quantization resolution of the wideband directional information encoded in the bitstream when the audio content exhibits voice activity is different from a second quantization resolution of the wideband directional information when the audio content does not exhibit voice activity.

18. An audio decoder (200; 900) as claimed in any preceding claim, wherein the renderer (220; 950) comprises a signal power calculation unit (951) for calculating a reference power for each of the plurality of time-frequency blocks according to the two or more transmission channels, wherein the renderer (220; 950) comprises a direct power calculation unit (952), in, If the audio content does not exhibit voice activity, the direct power calculation unit (952) is configured to scale the reference power using the transmitted power ratio encoded in the bitstream to obtain a scaled reference power, if the audio content exhibits voice activity, and scale the reference power using a scaling factor to obtain a scaled reference power, wherein the scaling factor is encoded in the bitstream or wherein the scaling factor is a constant scaling factor, for example depending on the number of transmitted objects, The renderer (220; 950) is configured to generate the one or more audio output signals in dependence on the scaled reference power.

19. An audio decoder (200; 900) as claimed in claim 18, wherein the renderer (220; 950) comprises a direct response calculation unit (953) for calculating a direct response, wherein the renderer (220; 950) is configured to calculate the direct response based on quantized direction information of a primary object that is a proper subset of the plurality of audio objects of the audio content if the audio content exhibits voice activity, wherein the renderer (220; 950) is configured to calculate the direct response based on quantized direction information of all audio objects of the audio content if the audio content does not exhibit voice activity, wherein the quantized direction information is encoded in the bitstream, The renderer (220; 950) is configured to generate the one or more audio output signals based on the direct response.

20. The audio decoder (200; 900) of claim 19, The renderer (220; 950) comprises an input covariance matrix calculation unit (954), wherein the input covariance matrix calculation unit (954) is used to calculate an input covariance matrix according to the two or more transmission channels, wherein the renderer (220; 950) comprises a target covariance matrix calculation unit (955), the target covariance matrix calculation unit (955) being configured to calculate a target covariance matrix based on the direct response and based on the scaled reference power, The renderer (220; 950) comprises a mixing matrix calculation unit (956), which is used to calculate a mixing matrix for rendering based on the input covariance matrix and based on the target covariance matrix. The renderer (220; 950) is configured to generate the one or more audio output signals according to the mixing matrix.

21. The audio decoder (200; 900) as claimed in any one of the preceding claims, The renderer (220; 950) is configured to generate one or more of the two or more transmission channels by applying code-excited linear prediction, or by applying modified discrete cosine transform or an inverse of the modified discrete cosine transform, or by applying a combination of the code-excited linear prediction and the modified discrete cosine transform.

22. The audio decoder (200; 900) as claimed in any one of the preceding claims, in, If the audio content includes the plurality of audio channels but does not include the plurality of audio objects, the number of the two or more transmission channels is less than the number of the plurality of audio channels, wherein if the audio content includes the plurality of audio objects but does not include the plurality of audio channels, the number of the two or more transmission channels is less than the number of the plurality of audio objects, wherein, if the audio content includes both the plurality of audio objects and the plurality of audio channels, the number of the two or more transmission channels is less than the sum of the number of the plurality of audio channels and the number of the plurality of audio objects; or wherein, if the audio content includes the plurality of audio channels but does not include the plurality of audio objects, the number of the two or more transmission channels is less than or equal to the number of the plurality of audio channels, wherein, if the audio content includes the plurality of audio objects but does not include the plurality of audio channels, the number of the two or more transmission channels is less than or equal to the number of the plurality of audio objects, Wherein, if the audio content includes both the plurality of audio objects and the plurality of audio channels, the number of the two or more transmission channels is less than or equal to the sum of the number of the plurality of audio channels and the number of the plurality of audio objects.

23. System, include: an audio encoder (100; 800), and An audio decoder (200; 900), The audio encoder (100; 800) comprises: a transmit signal generator (110; 710; 810) for generating two or more transmit channels of a transmit signal from an audio input, the audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels, a voice activity determiner (120; 820) for determining a voice activity decision for the transmitted signal, indicating whether the audio input within the transmitted signal exhibits voice activity, and a bitstream generator (130; 850) for generating a bitstream based on the audio input, wherein the bitstream generator (130; 850) is adapted to encode the two or more transport channels in the bitstream if the voice activity determiner (120; 820) has determined that the transport signal exhibits voice activity, wherein if the voice activity determiner (120; 820) has determined that the transmission signal does not exhibit voice activity, the bitstream generator (130; 850) is adapted to encode information about background noise instead of the two or more transmission channels, wherein the information about background noise comprises information about background noise of at least one of the two or more transmission channels or comprises information about background noise of a derived signal, the derived signal being dependent on at least one of the two or more transmission channels, wherein the audio encoder (100; 800) is configured to generate a bitstream from an audio input, and The audio decoder (200; 900) is configured to generate one or more audio output signals from the bitstream.

24. A method for decoding, include: receiving a bitstream dependent on audio content, the audio content comprising at least one of a plurality of audio objects and a plurality of audio channels; in a transmission signal comprising two or more transmission channels is encoded in the bitstream, and the audio content is encoded in the transmission signal; or wherein information about background noise is encoded in the bitstream instead of the transmission signal, wherein the information about background noise comprises information about background noise of at least one of the two or more transmission channels or information about background noise of a derived signal, the derived signal being dependent on at least one of the two or more transmission channels; and generating one or more audio output signals based on the audio content encoded in the bitstream; wherein, if the transport signal comprising the two or more transport channels is encoded in the bitstream, generating the one or more audio output signals is performed in dependence on the two or more transport channels, and If the information about the background noise is encoded in the bitstream instead of the transmission signal, then generating the one or more audio output signals is performed according to the information about the background noise.

25. A computer program for implementing the method according to claim 24 when the computer program is executed on a computer or a signal processor.

Citation Information

Patent Citations

  • Apparatus, method and computer program for encoding an audio signal or for decoding an encoded audio scene

    WO2022022876A1

  • Apparatus and method for encoding a plurality of audio objects using direction information during a downmixing or apparatus and method for decoding using an optimized covariance synthesis

    WO2022079044A1