Encoder and encoding method for parameterized encoded and decoded discontinuous transmission of independent streams with metadata

By adopting specific audio encoders and decoders in audio encoding and decoders, using transmission signal generators, voice activity determiners and bitstream generators, bit rate management and spatial prosecutor problems in the prior art under multi-audio objects and transmission channels are solved, and efficient bit rate demand reduction and noise management are achieved.

CN120129940APending Publication Date: 2025-06-10FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380075691.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-09
Filing Date
2023-09-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In prior art In audio encoding and decoding, it is difficult to effectively manage multiple audio objects and transmission channels in the non-continuous transmission (DTX) mode, resulting in a decrease in average bit rate and the generation of spatial falsetto.

Method used

An audio encoder and decoder is adopted to generate a transmission signal and bitstream according to the audio input through a transmission signal generator, a speech activity determiner and a bitstream generator, and encode and decode respectively in the speech activity and inactive phases, and background noise is managed using a mute insertion descriptor (SID) and a comfort noise generator (CNG).

Benefits of technology

It realizes that in multiple audio objects and transmission channel scenarios, the bit rate requirement is effectively reduced and the generation of spatial artificial sounds is reduced, and the high perceptual fidelity of background noise is maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129940A_ABST
    Figure CN120129940A_ABST
Patent Text Reader

Abstract

An audio encoder (100) according to an embodiment is provided. The audio encoder (100) comprises a transmission signal generator (110) for generating two or more transmission channels of a transmission signal from an audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels. Furthermore, the audio encoder (100) comprises a voice activity determiner (120) for determining a voice activity decision of the transmission signal indicating whether the audio input within the transmission signal exhibits voice activity. Furthermore, the audio encoder (100) comprises a bitstream generator (130) for generating a bitstream from the audio input. If the voice activity determiner (120) has determined that the transmission signal exhibits voice activity, the bitstream generator (130) is adapted to encode two or more transmission channels within the bitstream. If the voice activity determiner (120) has determined that the transmission signal does not exhibit voice activity, the bitstream generator (130) is adapted to encode information on background noise rather than the two or more transmission channels, wherein the information about the background noise comprises information about the background noise of at least one of the two or more transmission channels or information about the background noise of a derived signal dependent on the at least one of the two or more transmission channels.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Specification

[0002] The present invention relates to audio scenarios of independently streamable media (ISM) with metadata, parameterized encoding and decoding, discontinuous transmission (DTX) mode and comfort noise generation (CNG) for audio scenarios of independently streamable media (ISM) with metadata, parameterized encoding and decoding, and immersive voice and audio services (IVAS). In particular, the present invention relates to a codec and method for discontinuous transmission (DTX for Param-ISM) of independently streamable media with metadata, parameterized encoding and decoding.

[0003] In an IVAS codec, at low bitrates, audio objects or independently streamable media with metadata are encoded and decoded parametrically. In a first step, downmixing (e.g., stereo downmixing or virtual cardioid) and metadata can be calculated, for example, from the audio objects and from quantized direction information (e.g., from azimuth and elevation). This downmix is then encoded, for example, to obtain one or more transmission channels and can be transmitted to the decoder, for example, together with the metadata. The metadata can include, for example, direction information (e.g., azimuth and elevation), power ratios, and object indices corresponding to the main objects of a subset of input objects. At the decoder, a covariance renderer can receive the transmitted metadata and the stereo downmix / transmission channels as inputs and can render them, for example, to a desired loudspeaker layout (see [1], [2]).

[0004] Typically, in a communication codec, discontinuous transmission (DTX) is used to significantly reduce the transmission rate in the absence of voice input. In this mode, frames are first classified as "active" frames (i.e., frames containing speech) and "inactive" frames (i.e., frames containing background noise or silence). Later, for inactive frames, the codec operates in DTX mode to significantly reduce the transmission rate. Most frames determined to include background noise are stopped from being transmitted and are replaced at the decoder by some comfort noise generation (CNG). For these frames, a silence insertion descriptor (SID) frame, which is sent periodically but not at every frame, is used to transmit a very low-rate parametric representation of the signal. This allows the CNG in the decoder to generate artificial noise similar to the actual background noise.

[0005] The concept used in the prior art is discontinuous transmission (DTX). A comfort noise generator is typically used for discontinuous transmission of speech. According to this concept, speech is first classified into active and inactive frames by a voice activity detector (VAD). Examples of VADs can be found in [3]. Based on the VAD results, only the active speech frames are encoded, decoded, and transmitted at a nominal bit rate. During long pauses with only background noise or silence, the bit rate is reduced or set to zero, and the background noise / silence is encoded and decoded in an occasional and parametric manner. As a result, the average bit rate is significantly reduced. This noise is generated by a comfort noise generator (CNG) at the decoder side during inactive frames. For example, both the speech codec AMR-WB [3] and 3GPP EVS [4], [5] are capable of operating in the DTX mode. Examples of effective CNGs are given in [6]. In the IVAS codec, a discontinuous transmission (DTX) system exists in audio scenarios transmitted by parametric coding through the directional audio coding (DirAC) paradigm or in the metadata-assisted spatial audio (MASA) format (see [7]).

[0006] In a discrete independent stream with metadata (discrete ISM), the encoder of the discrete ISM receives an audio object and its associated metadata. Then, based on frames, the object is encoded separately (along with the metadata including object direction information (such as azimuth and elevation angles)), and the encoded data is subsequently transmitted to the decoder. The decoder then independently decodes the individual objects and renders them to a specified output architecture by applying an amplitude translation technique using the quantized direction information.

[0007] Another concept of the prior art is an independent stream with metadata encoded parametrically (Param-ISM). Figure 4 An overview of the corresponding encoder is shown, which particularly depicts the encoded audio signal 491 and the encoded parametric side information 495, 496, 497.

[0008] The encoder of the Param-ISM receives an audio object and associated metadata as input. The metadata can include, for example, on a frame basis, the object direction (e.g., azimuth, whose value is, for example, between [-180, 180]; and elevation, whose value is, for example, between [-90, 90]), which is then quantized and used during the calculation of stereo downmixing (such as virtual cardioid or transmission channel). Additionally, among the input audio objects, the two main objects and the power ratio between the two main objects can be determined, for example, on a time / frequency block basis. Subsequently, the metadata can be quantized and encoded, for example, together with the object indices of the two main objects (the two main objects on a time / frequency block).

[0009] The encoded bitstream 490 may include, for example, a stereo downmix / transmission channel 491 (which is separately encoded by means of a core codec), an encoded primary object index 495, a quantized and encoded power ratio 496, and a quantized and encoded direction information 497 (such as azimuth and elevation).

[0010] Figure 5 A simplified overview of a decoder is shown. The decoder receives the bitstream 490 and obtains the encoded stereo downmix / transmission channel 491, the encoded object index 495, the encoded power ratio 496, and the encoded direction information 497. The encoded stereo downmix / transmission channel 491 is then decoded using a core decoder and converted into a time / frequency representation using an analysis filter bank (such as a composite low-delay filter bank, CLDFB). The decoded object index may be used, for example, together with the decoded and dequantized direction information (such as azimuth and elevation and output configuration, such as 5.1, 5.1+4, 7.1, 7.1+4, etc.) to calculate the direction response. The direct response may be provided, for example, together with the transmission channel / stereo downmix in time / frequency representation, the prototype matrix, and the decoded and dequantized power ratio as an input in covariance synthesis operating in the time / frequency domain. The output of the covariance synthesis is converted from the time / frequency representation to the time domain representation using a synthesis filter (such as CLDFB).

[0011] Figure 6 A detailed overview of the covariance synthesis step is shown without reflecting the dimensions of the input / output data.

[0012] Covariance synthesis calculates the mixing matrix (M) for each time / frequency block, which renders the input transmission channel

[0013] (x = x(k,n) = [X 1 (k,n), X 2 (k,n)] T )

[0014] to the desired output loudspeaker configuration

[0015] (y = y(k,n) = [Y 1 (k,n), Y 2 (k,n), Y 3 (k,n),…] T )

[0016] (such as, 5.1 loudspeaker configuration, 7.1 loudspeaker configuration, 7.1+4 loudspeaker configuration, etc.):

[0017] y = Mx

[0018] For the mixing matrix, covariance synthesis may use the prototype matrix, the input covariance matrix Cx = xx T and the target covariance matrix C Y . The target covariance matrix is calculated by means of the signal power calculated from the transmission channel / stereo downmix, power ratio, and direct response.

[0019] The object of the present invention is to provide an improved concept for the discontinuous transmission of audio content. The object of the present invention is solved by the subject matter of the independent claims.

[0020] According to an embodiment, an audio encoder is provided. The audio encoder includes a transmission signal generator for generating two or more transmission channels of a transmission signal from an audio input, the audio input including at least one of a plurality of audio input objects and a plurality of audio input channels. In addition, the audio encoder includes a voice activity determiner for determining a voice activity decision of the transmission signal, which indicates whether the audio input within the transmission signal exhibits voice activity. In addition, the audio encoder includes a bitstream generator for generating a bitstream based on the audio input. If the voice activity determiner has determined that the transmission signal exhibits voice activity, the bitstream generator is adapted to encode two or more transmission channels within the bitstream. If the voice activity determiner has determined that the transmission signal does not exhibit voice activity, the bitstream generator is adapted to encode information about background noise rather than two or more transmission channels, where the information about background noise includes information about the background noise of at least one of the two or more transmission channels or information about the background noise of a derived signal that depends on at least one of the two or more transmission channels.

[0021] For example, according to an embodiment, the number of transmission channels is less than or equal to the number of input channels.

[0022] In addition, according to an embodiment, a method for audio coding is provided. The method includes:

[0023] - Generating two or more transmission channels of a transmission signal from an audio input, the audio input including at least one of a plurality of audio input objects and a plurality of audio input channels.

[0024] - Determining a voice activity decision of the transmission signal, which indicates whether the audio input within the transmission signal exhibits voice activity. And

[0025] - Determining a bitstream based on the audio input.

[0026] If it is determined that the transmission signal exhibits voice activity, the method includes encoding two or more transmission channels in a bitstream. If it is determined that the transmission signal does not exhibit voice activity, the method includes encoding information about background noise for at least one of two or more transmission channels or information about background noise for a derived signal that depends on at least one of two or more transmission channels, rather than encoding the two or more transmission channels.

[0027] Furthermore, a computer program is provided that, when executed on a computer or a signal processor, is configured to implement the above method.

[0028] Additionally, according to an embodiment, an audio decoder is provided. The audio decoder includes an input interface for receiving a bitstream that depends on audio content including at least one of a plurality of audio objects and a plurality of audio channels. A transmission signal including two or more transmission channels is encoded in the bitstream, and the audio content is encoded in the transmission signal. Alternatively, information about background noise rather than the transmission signal is encoded in the bitstream, and the information about background noise includes information about background noise for at least one of two or more transmission channels or information about background noise for a derived signal that depends on at least one of two or more transmission channels. Furthermore, the audio decoder includes a renderer configured to generate one or more audio output signals based on the audio content encoded in the bitstream. If a transmission signal including two or more transmission channels is encoded in the bitstream, the renderer is configured to generate one or more audio output signals based on the two or more transmission channels. If information about background noise rather than the transmission signal is encoded in the bitstream, the renderer is configured to generate one or more audio output signals based on the information about background noise.

[0029] Furthermore, a method for audio decoding is provided. The method includes:

[0030] - Receiving a bitstream that depends on audio content including at least one of a plurality of audio objects and a plurality of audio channels. A transmission signal including two or more transmission channels is encoded in the bitstream. The audio content is encoded in the transmission signal. Alternatively, information about background noise rather than the transmission signal is encoded in the bitstream, and the information about background noise includes information about background noise for at least one of two or more transmission channels or information about background noise for a derived signal that depends on at least one of two or more transmission channels. And:

[0031] - Generating one or more audio output signals based on the audio content encoded in the bitstream.

[0032] If a transmission signal comprising two or more transmission channels is encoded in a bitstream, one or more audio output signals are generated based on the two or more transmission channels. If information about background noise rather than the transmission signal is encoded in the bitstream, one or more audio output signals are generated based on the information about background noise.

[0033] Furthermore, a computer program is provided which, when executed on a computer or a signal processor, is for implementing the above method.

[0034] Some embodiments are based on the discovery that by combining existing solutions, DTX can be applied independently, for example, to individual streams (such as to audio objects or to individual channels, like stereo downmix / transmission channels). However, this would be incompatible with DTX designed for low bitrate communication because the available number of bits would not be sufficient to effectively describe the inactive part of the input signal for more than one object or for a transmission channel or for a downmix with more than one channel. Additionally, since the individual VAD decisions are not synchronized, such methods would also face problems. Spatial artefacts would be generated thereby.

[0035] In an embodiment, a DTX system for an audio scene described by (audio) objects and their associated metadata is provided.

[0036] Some embodiments provide DTX systems for audio objects (also referred to as ISM, i.e., "Independent Streams with Metadata"), especially SID and CNG, which are parametrically coded, for example, as Param-ISM.

[0037] In some embodiments, a sharp reduction in the bitrate requirements for transmitting conversational immersive voice is achieved.

[0038] According to some embodiments, the DTX concept is provided, which is extended to immersive voice with spatial cues.

[0039] In some embodiments, the two most dominant objects per time / frequency unit are considered. In other embodiments, more than two most dominant objects per time / frequency unit are considered, especially for an increased number of input objects. For the sake of readability of the text, the embodiments below mainly describe the two dominant objects per time / frequency unit, but similarly, these embodiments can be extended, for example, to more than two dominant objects per time / frequency unit in other embodiments.

[0040] A specific embodiment of an audio encoder is provided.

[0041] According to an embodiment, an audio encoder for encoding a plurality of (audio) objects and their associated metadata is provided.

[0042] The audio encoder may for example include a direction information determiner for extracting direction information and a direction information quantizer for quantizing the direction information.

[0043] Furthermore, the audio encoder may for example include a transmission signal generator (downmixer) for generating a transmission signal (downmix) that includes at least two transmission channels (e.g., downmix channels) from an input audio object and from the quantized direction information (e.g., azimuth and elevation angle) associated with the input audio object.

[0044] Furthermore, the audio encoder may for example include a decision logic module that combines individual VAD decisions of the transmission channels to calculate an overall decision regarding whether a frame is active.

[0045] Furthermore, the audio encoder may for example include a mono signal generator (e.g., stereo-to-mono converter) that outputs a mono signal from the transmission channels to be encoded during an inactive phase.

[0046] Furthermore, the audio encoder may for example include an inactive metadata generator that generates (e.g., calculates) inactive metadata to be transmitted during an inactive phase.

[0047] Furthermore, the audio encoder may for example include an active metadata generator that generates (e.g., calculates) active metadata to be transmitted during an active phase.

[0048] Furthermore, the audio encoder may for example include a transmission channel encoder configured to generate encoded data by encoding the downmixed signal including the transmission channels in an active phase.

[0049] Furthermore, the audio encoder may for example include a transmission channel silence insertion description generator that generates a silence insertion description of the background noise of the mono signal during an inactive phase.

[0050] Furthermore, the audio encoder may for example include a multiplexer that combines the active metadata and the encoded data into a bitstream during an active phase and is configured to either not transmit data or transmit the silence insertion description. Alternatively, the multiplexer may for example be configured to combine and transmit the silence insertion description and the inactive metadata during an inactive phase.

[0051] According to an embodiment, the transmission signal generator / downmixer may, for example, apply a CELP codec scheme (CELP = Code-Excited Linear Prediction), or may, for example, apply an MDCT-based codec scheme (MDCT = Modified Discrete Cosine Transform), or may, for example, apply a conversion combination of these two codec schemes.

[0052] In an embodiment, the active and inactive phases may, for example, be determined as follows: First, a voice activity detector is run separately on the transmission / downmix channel, and then the results of the transmission / downmix channel are combined to determine an overall decision.

[0053] According to an embodiment, a mono signal may, for example, be calculated from the transmission / downmix channel by adding the transmission channels or, for example, by selecting the channel with the higher long-term energy.

[0054] In an embodiment, the active and inactive metadata may, for example, differ in terms of quantization resolution or in terms of the type (nature) of the parameters (used).

[0055] According to an embodiment, the quantization resolution of the transmitted direction information and the quantization resolution of the direction information used to calculate the downmix may, for example, be different in the inactive phase.

[0056] In an embodiment, the spatial audio input format may, for example, be described by objects and their associated metadata (e.g., by independent streams with metadata).

[0057] According to an embodiment, two or more transmission channels may, for example, be generated.

[0058] Furthermore, a specific embodiment of an audio decoder is provided.

[0059] According to an embodiment, an audio decoder for generating a spatial audio output signal from a bitstream (decoding and). The bitstream may, for example, exhibit at least one active phase followed by at least one inactive phase. Furthermore, the bitstream may, for example, have at least a silence insertion descriptor frame (SlD) encoded therein, which may, for example, describe the background noise characteristics of the transmission / downmix channel and / or the spatial image information.

[0060] The audio decoder may, for example, include a SID decoder (silence insertion descriptor decoder), which may, for example, be configured to decode the silence insertion descriptor frame of the mono signal.

[0061] In addition, the audio decoder may for example include a mono-to-stereo converter which may be configured to generate at least two (downmixed) channels from the SID information and control parameters of the mono signal during an inactive phase / mode, where the control parameters may describe the characteristics of the stereo downmix / transmission channels, such as the ratio parameters calculated from the stereo downmix / transmission channels on the encoder side and / or, for example, broadband coherence or broadband correlation.

[0062] In addition, the audio decoder may for example include a transmission channel decoder which may be configured to reconstruct the transmission / downmix channels from the bitstream during the active phase / mode according to the bitstream during the active phase.

[0063] In addition, the audio decoder may for example include a (spatial) renderer which may be configured to reconstruct a spatial output signal during the active phase / mode based on the decoded transmission / downmix channels, for example based on the transmitted active metadata, for example based on the reconstructed background noise in the transmission / downmix channels and for example based on the transmitted inactive metadata during the inactive phase.

[0064] According to an embodiment, the mono-to-stereo converter may for example include a random generator which may execute at least twice with different seeds to generate noise, and may process the generated noise using the decoded SID information of the mono signal and using control parameters which may describe the characteristics of the stereo downmix / transmission channels, such as the ratio parameters calculated from the stereo downmix / transmission channels on the encoder side and / or, for example, broadband coherence or broadband correlation.

[0065] In an embodiment, the spatial parameters transmitted during the active phase may for example include object indices, power ratios (which may be transmitted in frequency subbands) and direction information (for example, azimuth and elevation), which may be broadband transmitted.

[0066] According to an embodiment, the spatial parameters transmitted during the inactive phase may for example include direction information (for example, azimuth and elevation) (which may be broadband transmitted) and control parameters which may describe the characteristics of the stereo downmix / transmission channels, such as the ratio parameters calculated from the stereo downmix / transmission channels on the encoder side and / or, for example, broadband coherence or broadband correlation.

[0067] In an embodiment, the quantization resolution of the direction information in the inactive phase is different from the quantization resolution of the direction information in the active phase.

[0068] According to an embodiment, the transmission of control parameters can be performed, for example, in a wideband or in frequency sub-bands, where the decision of "performing in a wideband or in frequency sub-bands" can be determined, for example, based on bitrate availability.

[0069] In an embodiment, the renderer can be configured, for example, to perform covariance synthesis.

[0070] The renderer can include, for example, a signal power calculation unit that calculates a reference power based on the transmission / downmix channel for each time / frequency block.

[0071] In addition, the renderer can include, for example, a direct power calculation unit that scales the reference power using a transmission power ratio during the active phase and scales the reference power using a constant scaling factor during the inactive phase.

[0072] In addition, the renderer can include, for example, a direct response calculation unit that calculates a direct response based on the quantized direction information of the primary object during the active phase or based on the quantized direction information of all transmitted objects during the inactive phase.

[0073] In addition, the renderer can include, for example, an input covariance matrix calculation unit that calculates an input covariance matrix based on the transmission / downmix channel.

[0074] In addition, the renderer can include, for example, a target covariance matrix calculation unit that calculates a target covariance matrix based on the outputs of the direct response calculation block and the direct power calculation block.

[0075] In addition, the renderer can include, for example, a mixing matrix calculation unit that calculates a mixing matrix for rendering based on the input covariance matrix and based on the target covariance matrix.

[0076] According to an embodiment, the constant scaling factor used during the inactive phase can be determined, for example, based on the number of transmitted objects; or a control parameter can be used, for example.

[0077] In an embodiment, the primary object can be, for example, a subset of all transmitted objects, and the number of primary objects can be less than, for example, the number of transmitted objects.

[0078] According to an embodiment, the transmission channel decoder can include, for example, a voice decoder (e.g., a CELP-based voice decoder), and / or can include, for example, a general audio decoder (e.g., a TCX-based decoder), and / or can include, for example, a bandwidth expansion module.

[0079] Other specific embodiments are provided in the dependent claims.

[0080] Hereinafter, embodiments of the present invention are described in more detail with reference to the accompanying drawings, where:

[0081] Figure 1 Shows an audio encoder according to an embodiment.

[0082] Figure 2 Shows an audio decoder according to an embodiment.

[0083] Figure 3 Shows a system according to an embodiment.

[0084] Figure 4 Shows an overview of the Param-ISM encoder.

[0085] Figure 5 Shows an overview of the Param-ISM decoder.

[0086] Figure 6 Shows a detailed overview of the covariance synthesis step in Param-ISM without reflecting the dimensions of the input / output data.

[0087] Figure 7 Shows a block diagram for determining whether a frame is active or inactive according to an embodiment.

[0088] Figure 8 Shows a block diagram of an encoder according to an embodiment.

[0089] Figure 9 Shows a block diagram of a decoder according to an embodiment.

[0090] Figure 10 Shows a spatial renderer according to an embodiment.

[0091] Figure 11 Shows generating a stereo signal according to an embodiment using three random seeds - seed 1, seed 2, and seed 3, to derive scale factors and control parameters.

[0092] Figure 12 Shows the generation of a stereo signal according to another embodiment, where the noise N 3 (k,n) generated by the third random generator for the left channel is also used to generate the right channel.

[0093] Figure 1 Shows an audio encoder 100 according to an embodiment.

[0094] The audio encoder 100 includes a transmission signal generator 110 that is configured to generate two or more transmission channels of a transmission signal from an audio input that includes at least one of a plurality of audio input objects and a plurality of audio input channels.

[0095] In addition, the audio encoder 100 includes a voice activity determiner 120 that determines a voice activity decision for the transmission signal, which indicates whether the audio input within the transmission signal exhibits voice activity.

[0096] In addition, the audio encoder 100 includes a bitstream generator 130 that generates a bitstream based on the audio input.

[0097] If the voice activity determiner 120 has determined that the transmission signal exhibits voice activity, the bitstream generator 130 is adapted to encode two or more transmission channels within the bitstream.

[0098] If the voice activity determiner 120 has determined that the transmission signal does not exhibit voice activity, the bitstream generator 130 is suitable for encoding information about background noise rather than two or more transmission channels, where the information about background noise includes information about the background noise of at least one of the two or more transmission channels or information about the background noise of a derived signal that depends on at least one of the two or more transmission channels.

[0099] According to an embodiment, the voice activity determiner 120 may be configured, for example, to determine a separate voice activity decision for each transmission channel among one or more transmission channels of the transmission signal, where the separate voice activity decision indicates whether the audio input within the transmission channel exhibits voice activity. In addition, the voice activity determiner 120 may be configured, for example, to determine the voice activity decision of the transmission signal based on the separate voice activity decisions of each transmission channel among one or more transmission channels.

[0100] In an embodiment, the voice activity determiner 120 may be configured, for example, to determine a separate voice activity decision for each transmission channel among two or more transmission channels of the transmission signal, where the separate voice activity decision indicates whether the audio input within the transmission channel exhibits voice activity. Additionally, the voice activity determiner 120 may be configured, for example, to determine the voice activity decision of the transmission signal based on the separate voice activity decisions of each transmission channel among two or more transmission channels of the transmission signal.

[0101] According to an embodiment, the voice activity determiner 120 may be configured, for example, to determine that the transmission signal exhibits voice activity if at least one of two or more transmission channels of the transmission signal exhibits voice activity. In addition, the voice activity determiner 120 may be configured, for example, to determine that the transmission signal does not exhibit voice activity if none of two or more transmission channels of the transmission signal exhibits voice activity.

[0102] In an embodiment, the audio encoder 100 may be configured, for example, to determine whether to transmit a bitstream in which information about background noise has been encoded, or whether not to generate and transmit the bitstream, in a case where the voice activity determiner 120 has determined that the transmission signal does not exhibit voice activity.

[0103] According to an embodiment, the audio encoder 100 may include, for example, a mono signal generator 830 (see Figure 8 ), which is configured to generate a derived signal as a mono signal from at least one of two or more transmission channels in a case where the voice activity determiner 120 has determined that the transmission signal does not exhibit voice activity. The audio encoder 100 may include, for example, an information generator, which is configured to generate information about background noise as information about the background noise of the mono signal.

[0104] In an embodiment, the mono signal generator 830 may be configured, for example, to generate a mono signal by adding two or more transmission channels or by adding two or more channels derived from two or more transmission channels. Alternatively, the mono signal generator 830 may be configured, for example, to generate a mono signal by selecting a transmission channel that exhibits higher energy among two or more transmission channels.

[0105] According to an embodiment, the information generator may be configured, for example, to generate information about the background noise of the mono signal as information about the mono signal.

[0106] In an embodiment, the information generator may be configured, for example, to generate a silence insertion description of the background noise of the mono signal as information about the background noise of the mono signal.

[0107] According to an embodiment, the audio encoder 100 may include, for example, a direction information determiner 802 (see Figure 8 ), which is configured to determine direction information based on an audio input. The audio encoder 100 may include, for example, a direction information quantizer 804 (see Figure 8 ), which is configured to quantize the direction information to obtain quantized direction information. The bitstream generator 130 may be configured, for example, to encode the quantized direction information in the bitstream.

[0108] In an embodiment, the transmission signal generator 110 may be configured, for example, to generate two or more transmission channels of the transmission signal from the audio input using the direction information.

[0109] According to an embodiment, the audio input may include a plurality of audio input objects, for example. The direction information may include information about the azimuth angle and elevation angle of an audio input object among the plurality of audio input objects of the audio input, for example.

[0110] In an embodiment, the audio encoder 100 may include, for example, an active metadata generator 825 (see Figure 8 ), which is configured to generate metadata in the case where the voice activity determiner 120 has determined that the transmission signal exhibits voice activity, and the metadata includes at least one of quantized direction information, object index, and power ratio of a plurality of audio input objects and / or a plurality of audio input channels of the audio input.

[0111] According to an embodiment, the audio input may include, for example, a plurality of audio input objects. The audio encoder 100 may include, for example, an inactive metadata generator 826 (see Figure 8 ), which is configured to generate metadata in the case where the voice activity determiner 120 has determined that the transmission signal does not exhibit voice activity, and the metadata includes quantized direction information and control parameters, such as a scaling factor depending on the number of audio input objects among the plurality of audio input objects of the audio input, or a scaling factor depending on the long-term energy of the transmission channel of the transmission signal and / or depending on the coherence or correlation between the transmission channels of the transmission signal.

[0112] In an embodiment, the quantization resolution of the direction information that may be generated, for example, by the inactive metadata generator 826 may be different from the quantization resolution of the direction information that may be generated, for example, by the active metadata generator 825.

[0113] In an embodiment, the characteristics of the metadata that may be generated, for example, by the inactive metadata generator 826 may be different from the characteristics of the metadata that may be generated, for example, by the active metadata generator 825.

[0114] According to an embodiment, the audio input may include, for example, a plurality of audio input objects and metadata associated with the audio input objects.

[0115] In an embodiment, the transmission signal generator 110 may be configured to generate two or more transmission channels of the transmission signal from the audio input, including downmixing at least one of the plurality of audio input objects and the plurality of audio input channels to obtain a downmix as the transmission signal, which may include two or more downmix channels as two or more transmission channels.

[0116] According to an embodiment, if the audio input within the transmission signal does not exhibit voice activity, the direction information quantizer 804 is configured to determine the quantized direction information such that the quantization resolution of the quantized direction information may be different from the quantization resolution used for calculating the downmix, for example.

[0117] In an embodiment, the bitstream generator 130 may be configured, for example, to encode control parameters in a bitstream in case the voice activity determiner 120 has determined that the transmitted signal does not exhibit voice activity. The control parameters may be suitable, for example, for controlling the generation of an intermediate signal from random noise. The control parameters may include, for example, parameter values for a plurality of subbands, or wherein the control parameters may include a single wideband control parameter.

[0118] According to an embodiment, the audio encoder 100 may be configured, for example, to generate control parameters by selecting whether the control parameters may include parameter values for a plurality of subbands, or whether the control parameters may include a single wideband control parameter, depending on the available bitrate.

[0119] In an embodiment, the transmitted signal generator 110 may be configured, for example, to encode an audio input by applying code-excited linear prediction or by applying a modified discrete cosine transform or by applying a combination of code-excited linear prediction and a modified discrete cosine transform.

[0120] According to an embodiment, if the audio input includes a plurality of audio input channels but does not include a plurality of audio input objects, the number of two or more transmitted channels may be less than, for example, the number of audio input channels. If the audio input includes a plurality of audio input objects but does not include a plurality of audio input channels, the number of two or more transmitted channels may be less than, for example, the number of audio input objects. If the audio input includes both a plurality of audio input objects and a plurality of audio input channels, the number of two or more transmitted channels may be less than, for example, the sum of the number of audio input channels and the number of audio input objects.

[0121] Alternatively, according to an embodiment, if the audio input includes a plurality of audio input channels but does not include a plurality of audio input objects, the number of two or more transmitted channels may be less than or equal to, for example, the number of audio input channels. If the audio input includes a plurality of audio input objects but does not include a plurality of audio input channels, the number of two or more transmitted channels may be less than or equal to, for example, the number of audio input objects. If the audio input includes both a plurality of audio input objects and a plurality of audio input channels, the number of two or more transmitted channels may be less than or equal to, for example, the sum of the number of audio input channels and the number of audio input objects.

[0122] Figure 2 An audio decoder 200 according to an embodiment is shown.

[0123] The audio decoder 200 includes an input interface 210 for receiving a bitstream that depends on audio content including at least one of a plurality of audio objects and a plurality of audio channels. A transmission signal including two or more transmission channels is encoded in the bitstream, and the audio content is encoded in the transmission signal. Alternatively, information about background noise instead of the transmission signal is encoded in the bitstream, and the information about background noise includes information about background noise of at least one of two or more transmission channels or information about background noise of a derived signal that depends on at least one of two or more transmission channels.

[0124] In addition, the audio decoder 200 includes a renderer 220 for generating one or more audio output signals based on the audio content encoded in the bitstream.

[0125] If a transmission signal including two or more transmission channels is encoded in the bitstream, the renderer 220 is configured to generate one or more audio output signals based on the two or more transmission channels.

[0126] If information about background noise instead of the transmission signal is encoded in the bitstream, the renderer 220 is configured to generate one or more audio output signals based on the information about background noise.

[0127] According to an embodiment, if the audio content exhibits voice activity, a transmission signal including two or more transmission channels may be encoded in the bitstream, for example. If the audio content does not exhibit voice activity, information about background noise instead of the transmission signal may be encoded in the bitstream, for example.

[0128] In an embodiment, the audio decoder 200 may include, for example, a demultiplexer 902, a noise information determiner 920, and a multi-channel generator 930 (see Figure 9 ). The demultiplexer may be configured to determine whether the transmitted bitstream corresponds to an active frame or an inactive frame based on the size of the bitstream, for example. If information about background noise is encoded in the bitstream, the noise information determiner 920 may be configured to determine information about background noise from the bitstream, the multi-channel generator 930 may be configured to generate a derived signal as an intermediate signal including two or more intermediate channels from the information about background noise, and the renderer 220 may be configured to generate one or more audio output signals based on the two or more intermediate channels of the intermediate signal, for example.

[0129] According to an embodiment, the multi-channel generator 930 may include, for example, a random generator for generating random noise. The multi-channel generator 930 may be configured to generate two or more intermediate channels based on the random noise.

[0130] In an embodiment, the multi-channel generator 930 may be configured, for example, to adjust random noise based on information about background noise to obtain shaped noise. The multi-channel generator 930 may be configured, for example, to generate two or more intermediate channels from the shaped noise.

[0131] According to an embodiment, the multi-channel generator 930 may be configured, for example, to run a random generator at least twice with different seeds to obtain random noise.

[0132] In an embodiment, the multi-channel generator 930 may be configured, for example, to generate two or more intermediate channels based on the random noise and based on control parameters (e.g., depending on the proportion and / or coherence or correlation of the transmission channels of the transmitted signal), where the control parameters may be encoded, for example, in the bitstream as part of the inactive metadata.

[0133] According to an embodiment, the control parameters may be encoded, for example, in the bitstream and may include, for example, multiple parameter values for multiple subbands, and the multi-channel generator 930 may be configured, for example, to generate each subband of the two or more intermediate channels' multiple subbands based on the parameter value among the multiple parameter values of the control parameter associated with the subband.

[0134] In an embodiment, the control parameters may be encoded, for example, in the bitstream, where the control parameters may include, for example, a single wideband control parameter.

[0135] According to an embodiment, the multi-channel generator 930 may be configured to generate two or more intermediate channels by generating a first random noise portion of the random noise using a random generator with a first seed, by generating a first intermediate channel of the two or more intermediate channels based on the first random noise portion, by generating a second random noise portion of the random noise using a random generator with a second seed different from the first seed, and by generating a second intermediate channel of the two or more intermediate channels based on the second random noise portion.

[0136] According to an embodiment, the multi-channel generator 930 may be configured, for example, to generate a first intermediate channel among two or more intermediate channels based on a first random noise portion, based on a third noise portion, and based on control parameters (such as a scaling factor and / or, for example, coherence or correlation). Additionally, the multi-channel generator 930 may be configured, for example, to generate a second intermediate channel among two or more intermediate channels based on a second random noise portion, based on a third noise portion, and based on control parameters (such as a scaling factor and / or, for example, coherence or correlation). The multi-channel generator 930 may be configured, for example, to generate a first random noise portion of the random noise using a random generator operating on a first seed, generate a second random noise portion of the random noise using a random generator operating on a second seed, and generate a third random noise portion of the random noise using a random generator operating on a third seed, where the second seed is different from the first seed, and where the third seed is different from the first seed and different from the second seed.

[0137] In an embodiment, the multi-channel generator 930 may be configured, for example, to generate two or more intermediate channels by generating a first intermediate channel among two or more intermediate channels based on random noise and by generating a second intermediate channel among two or more intermediate channels from the first intermediate channel among two or more intermediate channels.

[0138] According to an embodiment, the multi-channel generator 930 may be configured, for example, to generate a second intermediate channel among two or more intermediate channels such that the second intermediate channel among two or more intermediate channels may be, for example, the same as the first intermediate channel among two or more intermediate channels. Alternatively, the multi-channel generator 930 may be configured, for example, to generate a second intermediate channel among two or more intermediate channels by modifying the first intermediate channel among two or more intermediate channels.

[0139] In an embodiment, the renderer 220 may be configured, for example, to generate two or more audio output signals as one or more audio output signals.

[0140] According to an embodiment, the audio content may include, for example, a plurality of audio objects. If the audio content exhibits voice activity, a plurality of audio object indices associated with the plurality of audio objects, a plurality of power ratios associated with the plurality of audio objects in a plurality of subbands, and wideband direction information of the plurality of audio objects may be encoded, for example, in a bitstream, and the renderer 220 may be configured, for example, to generate one or more audio output signals based on the plurality of audio object indices, based on the plurality of power ratios, and based on the wideband direction information of the plurality of audio objects.

[0141] In an embodiment, the audio content may include, for example, a plurality of audio objects. If the audio content does not exhibit voice activity, the wideband direction information and control parameters of the plurality of audio objects may be encoded, for example, in a bitstream, and the renderer 220 may be configured, for example, to generate one or more audio output signals based on the wideband direction information and based on all object indices and a constant power ratio, where the constant power ratio depends on the number of transmitted objects.

[0142] According to an embodiment, a first quantization resolution of the wideband direction information encoded in the bitstream when the audio content exhibits voice activity may be different from, for example, a second quantization resolution of the wideband direction information when the audio content does not exhibit voice activity.

[0143] In an embodiment, the renderer 220 may include, for example, a signal power calculation unit 951 (see Figure 10 ), which is used to calculate a reference power based on two or more transmission channels of each time-frequency block among a plurality of time-frequency blocks. In addition, the renderer 220 may include, for example, a direct power calculation unit 952 (see Figure 10 ), which is used to scale the reference power using the transmitted power ratio encoded in the bitstream in the case where the audio content exhibits voice activity, and to scale the reference power using a scaling factor encoded in the bitstream in the case where the audio content does not exhibit voice activity, to obtain a scaled reference power. In addition, the renderer 220 may be configured, for example, to generate one or more audio output signals based on the scaled reference power.

[0144] According to an embodiment, the renderer 220 may include, for example, a direct response calculation unit 953 (see Figure 10 ) for calculating a direct response, where the renderer 220 may be configured, for example, to calculate the direct response based on the quantized direction information of a primary object that is a suitable subset of the plurality of audio objects of the audio content in the case where the audio content exhibits voice activity, where the renderer 220 may be configured, for example, to calculate the direct response based on the quantized direction information of all audio objects of the audio content in the case where the audio content does not exhibit voice activity, where the quantized direction information may be encoded, for example, in the bitstream. The renderer 220 may be configured, for example, to generate one or more audio output signals based on the direct response.

[0145] In an embodiment, the renderer 220 may include, for example, an input covariance matrix calculation unit 954 (see Figure 10 ), which is used to calculate an input covariance matrix based on two or more transmission channels. In addition, the renderer 220 may include, for example, a target covariance matrix calculation unit 955 (see Figure 10), the target covariance matrix calculation unit is configured to calculate a target covariance matrix based on the direct response and based on the scaled reference power. Additionally, the renderer 220 may include, for example, a mixing matrix calculation unit 956 (see Figure 10 ), the mixing matrix calculation unit is configured to calculate a mixing matrix for rendering based on the input covariance matrix and based on the target covariance matrix. The renderer 220 may be configured to generate one or more audio output signals based on the mixing matrix, for example.

[0146] According to an embodiment, the renderer 220 may be configured to generate one or more transmission channels of a transmission signal by applying code-excited linear prediction, by applying an inverse modified discrete cosine transform or a modified discrete cosine transform, or by applying a combination of code-excited linear prediction and a modified discrete cosine transform.

[0147] According to an embodiment, if the audio content includes a plurality of audio channels but does not include a plurality of audio objects, the number of two or more transmission channels may be, for example, less than the number of the plurality of audio channels. If the audio content includes a plurality of audio objects but does not include a plurality of audio channels, the number of two or more transmission channels may be, for example, less than the number of the plurality of audio objects. If the audio content includes both a plurality of audio objects and a plurality of audio channels, the number of two or more transmission channels may be, for example, less than the sum of the number of the plurality of audio channels and the number of the plurality of audio objects.

[0148] Alternatively, according to an embodiment, if the audio content includes a plurality of audio channels but does not include a plurality of audio objects, the number of two or more transmission channels may be, for example, less than or equal to the number of the plurality of audio channels. If the audio content includes a plurality of audio objects but does not include a plurality of audio channels, the number of two or more transmission channels may be, for example, less than or equal to the number of the plurality of audio objects. If the audio content includes both a plurality of audio objects and a plurality of audio channels, the number of two or more transmission channels may be, for example, less than or equal to the sum of the number of the plurality of audio channels and the number of the plurality of audio objects.

[0149] Figure 3 A system according to an embodiment is shown. The system includes an audio encoder 100 according to one of the above embodiments and an audio decoder 200 according to one of the above embodiments.

[0150] The audio encoder 100 is configured to generate a bitstream from an audio input.

[0151] The audio decoder 200 is configured to generate one or more audio output signals from the bitstream.

[0152] Embodiments will be described in detail below.

[0153] According to an embodiment, a DTX system (e.g., its encoder) may be configured, for example, to determine an overall decision of "whether a frame is active or inactive" based on independent decisions for stereo downmix channels and / or based on individual audio objects.

[0154] A DTX system (e.g., its encoder) may be configured, for example, to transmit a mono signal, together with inactive metadata, to a decoder using a silence insertion descriptor (SID).

[0155] Furthermore, a DTX system (e.g., its decoder) may be configured, for example, to generate a transmission channel / downmix including at least two channels using a comfort noise generator (CNG) based on SID information of only the mono signal.

[0156] Furthermore, a DTX system (e.g., its decoder) may be configured, for example, to post-process the generated transmission channel / downmix using control parameters, where the control parameters may be calculated, for example, on the encoder side from the stereo downmix / transmission channel.

[0157] Furthermore, a DTX system (e.g., its decoder) may be configured, for example, to render a multi-channel transmission signal to a defined output architecture using modified covariance synthesis.

[0158] In the following, other specific embodiments are described.

[0159] Figure 7 A block diagram for determining whether a frame is active or inactive according to an embodiment is shown. The overall decision is based on individual decisions for the transmission channel / downmix channel.

[0160] In Figure 7 , a transmission signal generator (e.g., a downmixer) 710 may be configured, for example, to receive an audio object and its associated quantized direction information (e.g., azimuth and elevation).

[0161] The transmission signal (e.g., downmix (DMX)) DMX for a first transmission channel (e.g., left downmix channel) L and the transmission signal DMX for a second transmission channel (e.g., right downmix channel) R may be generated, for example, as follows:

[0162]

[0163] where N is the total number of input objects, k is the sample index, and i is the object index

[0164] w L [i] = 0.5 + 0.5 * cos(azimuth[i] - π / 2)

[0165] w R[i] = 0.5 + 0.5 * cos(azimuth[i] + π / 2)

[0166] In another embodiment, two transmission channels (e.g., downmix channels) can be generated as follows, for example, using a downmix matrix D:

[0167]

[0168] where obj 1 …obj N represents audio objects 1 to audio object N.

[0169] In addition, Figure 7 a decision logic module 720 is depicted, which includes individual decision logic 722 and overall decision logic 725.

[0170] In Figure 7 , the individual decision logic 722 can be configured, for example, to determine whether an individual channel is active or inactive. The individual decision as to whether each of two (or more) transmission channels is active or inactive can be indicated, for example, by a (e.g., internal) tag.

[0171] In an embodiment, the individual decision logic 722 can be configured, for example, to receive two (or more) transmission channels as inputs. The individual decision logic 722 can be configured, for example, to determine for each of the two (or more) transmission channels DMX L 、DMX R whether the transmission channel exhibits voice activity, for example, by analyzing the transmission channel.

[0172] In another embodiment, the individual decision logic 722 can analyze all audio input channels or all audio input objects used by the transmission signal generator 710 to form two (or more) transmission channels DMX L 、DMX R . For example, if the individual decision logic 722 detects voice activity in at least one of the audio input channels or audio input objects, the individual decision logic 722 can, for example, infer the presence of voice activity in each transmission channel and can, for example, infer that each transmission channel is active. For example, if the individual decision logic 722 that detects voice activity does not detect voice activity in any of the audio input channels or audio input objects used to generate each transmission channel, the individual decision logic 722 can, for example, infer the absence of voice activity in each transmission channel and can, for example, infer that each transmission channel is inactive.

[0173] In addition, in Figure 7In it, the overall decision logic 725 can be configured, for example, to receive individual decisions (e.g., for transmission channels) as input and can be configured, for example, to determine an overall decision based on the individual decisions. For example, the overall decision logic 725 can use the DTX_FLAG to indicate the decision, for example. The overall decision logic can determine the overall decision according to Table 1 below, which depicts frame-by-frame decisions based on individual downmix decisions on a frame-by-frame basis:

[0174] Table 1

[0175] Activity in the first transmission channel (D_L) Activity in the second transmission channel (D_R) Overall decision (Decision_Overall) Active Active Active Inactive Active Active Active Inactive Active Inactive Inactive Inactive

[0176] For example, the overall decision can be determined by using a hysteresis buffer with a predefined size. Using a hysteresis buffer helps avoid artifacts caused by frequent switching between active and inactive portions. For example, a hysteresis buffer of size 10 may require 10 frames before switching from an active decision to an inactive decision.

[0177] The following gives an example of pseudocode for determining the overall decision:

[0178] Shift the hysteresis buffer by one step, for example

[0179] buffer_decision[i] = buffer_decision[i + 1]

[0180] where i = 0, 1, 2...(Buff_size - 1)

[0181] Buff_decision[buff_size] = Decision_Overall

[0182] where Decision_Overall can be calculated as shown in Table 1, for example.

[0183] The overall decision can be calculated as outlined in the following pseudocode:

[0184] DTX_Flag = 1;

[0185] for(i = 0; i < buff_size; i++)

[0186] {

[0187] DTX_Flag = DTX_Flag && buffer_decision[i];

[0188] }

[0189] In this pseudocode, DTX_Flag = 1 means "inactive" and DTX_FLAG = 0 means "active".

[0190] Figure 8 FIG. 800 shows an audio encoder according to an embodiment. Figure 8 The audio encoder may implement, for example, Figure 1 a specific embodiment of the audio encoder 100. Specifically, Figure 8 FIG. shows a block diagram of the encoder, which may be configured, for example, to receive an input audio object and its associated metadata.

[0191] In addition, the audio encoder 800 may include, for example, a transmission signal generator (e.g., a downmixer) 810 for generating a downmixed (transmission channel) signal (e.g., Figure 7 the transmission signal generator 710), the downmix including at least two channels from the input audio object and from the quantized direction information (e.g., azimuth and elevation) associated with the input audio object.

[0192] In addition, the audio encoder 800 may include, for example, a voice activity determiner, such as a voice activity determiner implemented with a decision logic module 820 (e.g., Figure 7 the decision logic module 720), for combining individual VAD decisions for the transmission channel to calculate an overall decision as to whether a frame is an active frame.

[0193] Stereo downmix may be calculated from the input audio object in the transmission signal generator 810 using, for example, the quantized direction information (e.g., azimuth and elevation).

[0194] The stereo downmix is then fed back into the decision logic module 820, where the decision as to whether the frame is active or inactive may be determined, for example, based on the above logic. For example, the decision logic module 820 may include the individual decision logic 722 and the overall decision logic 725 as described above.

[0195] If the decision logic module 820 has determined "active" as the overall decision (for an active frame), then Figure 8 the encoder in Figure 4 provides a more efficient method compared to the encoder in

[0196] Conversely, if the decision logic module 820 has determined "inactive" as the overall decision (for an inactive frame), the SID bitrate (e.g., 4.4 kbps or 5.2 kbps) will be too low to efficiently transmit the two channels of the stereo downmix and the active metadata. Therefore, for the sporadically / occasionally transmitted SID frames, the metadata bitrate can be, for example, 1.85 kbps or 2.45 kbps, and can include, for example, coarsely quantized direction information (e.g., azimuth and elevation angle) and control parameters that control the spatial sense of the background noise and are derived from the stereo downmix / transmitted signal, such as a scaling factor and / or, for example, coherence or correlation.

[0197] In an embodiment, during an inactive frame, the transmission of the object index and the power ratio may not occur. The main motivation for not transmitting the object index or the power ratio during an inactive frame is the assumption that the background noise does not have any specific direction and is diffuse in nature.

[0198] In addition, the audio encoder 800 may include, for example, a transmission channel silence insertion descriptor generator 840 that is used to generate a silence insertion descriptor for the background noise of the mono signal during an inactive phase. The transmission channel SID generator (transmission channel SID encoder) 840 may operate, for example, at 2.4 kbps and may receive, for example, a mono downmix as input.

[0199] In addition, the audio encoder 800 may include, for example, a mono signal generator (e.g., a stereo-to-mono converter) 830 that is used to output a mono signal from the transmission channel to be encoded during an inactive phase. The conversion from stereo downmix to mono downmix may be performed, for example, by the mono signal generator (e.g., a stereo-to-mono converter) 830.

[0200] In an embodiment, the downmix (e.g., stereo-to-mono conversion) may be implemented, for example, as the addition of two stereo transmission / downmix channels, for example:

[0201] M = DMZ L + DMX R

[0202] In another embodiment, the downmix (e.g., stereo-to-mono conversion) may be implemented, for example, as the transmission of only one channel of the stereo downmix. The decision of which channel to select may depend, for example, on the (e.g., long-term) energy of the individual channels of the stereo downmix. For example, the channel with the higher long-term energy may be selected:

[0203]

[0204] where LE LIndicates the long-term energy of the first (e.g., left) channel, and LE R Indicates the long-term energy of the second (e.g., right) channel.

[0205] Table 2 depicts metadata that can be transmitted, for example, during active and inactive frames:

[0206] Table 2

[0207]

[0208] Figure 8 The audio encoder 800 may, for example, include a direction information extractor 802 that extracts direction information and a direction information quantizer 804 that quantizes the direction information.

[0209] In addition, the audio encoder 800 may, for example, include an inactive metadata generator 826 that is configured to generate (e.g., calculate) inactive metadata to be transmitted during the inactive phase.

[0210] In addition, the audio encoder 800 may, for example, include an active metadata generator 825 that is configured to generate (e.g., calculate) active metadata to be transmitted during the active phase.

[0211] In addition, the audio encoder 800 may, for example, include a transmission channel encoder 828 that is configured to generate encoded data by encoding the downmixed signal including the transmission channel in the active phase.

[0212] In addition, the audio encoder 800 may, for example, include a bitstream generator that may be implemented, for example, as a multiplexer 850 that combines (e.g., encodes) the active metadata with the encoded data (e.g., two or more transmission channels) into a bitstream during the active phase and is used to transmit no data or to transmit a silence insertion description. Alternatively, the multiplexer 850 may, for example, be configured to combine and transmit the silence insertion description and the inactive metadata during the inactive phase.

[0213] Figure 9 Shows an audio decoder 900 according to an embodiment. Figure 9 The audio decoder 900 may, for example, be implemented Figure 2 A specific embodiment of the audio decoder 200.

[0214] The audio decoder 900 may, for example, receive a bitstream through an input interface, which may be implemented, for example, as a demultiplexer 902.

[0215] Figure 9The audio decoder 900 may, for example, include a transmission channel decoder 910, which may be configured, for example, to reconstruct the transmission / downmix channel according to the bitstream during the active phase / mode.

[0216] In addition, the audio decoder 900 may, for example, include a noise information determiner implemented as a SID decoder (silence insertion descriptor decoder) 920, which may be configured, for example, to decode the silence insertion descriptor frames of the mono signal.

[0217] In addition, the audio decoder 900 may, for example, include a multi-channel generator 930 implemented as a mono-to-stereo converter 930, which may be configured, for example, to generate at least two (downmixed) channels from the SID information of the mono signal and from the control parameters during the inactive phase / mode.

[0218] In addition, Figure 9 the audio decoder 900 may, for example, include a filter bank analysis module 940.

[0219] In addition, the audio decoder 900 may, for example, include a (spatial) renderer 950, which may be configured, for example, to: reconstruct a spatial output signal according to the decoded transmission / downmix channel and, for example, according to the transmitted active metadata during the active phase / mode; and reconstruct a spatial output signal according to the reconstructed background noise in the transmission / downmix channel during the inactive phase and, for example, according to the transmitted inactive metadata.

[0220] Figure 9 The audio decoder 900 may, for example, include a synthesis module for (e.g., band) synthesizing the spatial output signal of the renderer 950.

[0221] Figure 9 The audio decoder 900 may, for example, further include a voice activity information determiner 905, which is used to determine, for example, whether the decoder will operate in an active or inactive form (in the active mode or in the inactive mode) based on the VAD data in the bitstream.

[0222] In the currently described active mode (in the active form), Figure 9 the decoder described in Figure 5 is more efficient than the decoder described in

[0223] Figure 10 shows a spatial renderer for covariance rendering according to an embodiment. Figure 9 The renderer 950 shown in Figure 10 may be implemented as a spatial renderer of

[0224] The renderer may include, for example, a signal power calculation unit 951 that calculates a reference power based on the transmission / downmix channel for each time / frequency block.

[0225] In addition, the renderer may include, for example, a direct power calculation unit 952 that: adjusts the reference power proportionally using the transmitted power ratio during the active phase; and adjusts the reference power proportionally using, for example, a constant scaling factor depending on the number of transmitted objects or a scaling factor transmitted as part of the metadata during the inactive phase, or without scaling.

[0226] In addition, the renderer may include, for example, a direct response calculation unit 953 that calculates a direct response based on the quantized direction information of the primary object during the active phase or based on the quantized direction information of all transmitted objects during the inactive phase.

[0227] In addition, the renderer may include, for example, an input covariance matrix calculation unit 954 that calculates an input covariance matrix based on the transmission / downmix channel.

[0228] In addition, the renderer may include, for example, a target covariance matrix calculation unit 955 that calculates a target covariance matrix based on the output of the direct power calculation block 952 and based on the output of the direct response calculation block 953 (or based on a calculated covariance matrix depending on the output of the direct response calculation block 953).

[0229] In addition, the renderer may include, for example, a mixing matrix calculation unit 956 that calculates a mixing matrix for rendering based on the input covariance matrix and based on the target covariance matrix.

[0230] For example, for the mixing matrix, covariance synthesis may use a prototype matrix, the input covariance matrix C x = xx T and the target covariance matrix C Y . As described in reference Figure 6 .

[0231] In addition, the renderer may include, for example, an amplitude translation unit 957 that performs amplitude translation on the transmission channel based on the mixing matrix calculated by the mixing matrix calculation unit 956.

[0232] Figure 10 The spatial renderer for covariance synthesis-based rendering depicted in may use, for example, active metadata such as quantized direction information, object indices, and power ratios. This covariance rendering is therefore more efficient than the covariance rendering shown in Figure 3 .

[0233] Figure 9 The transmission channel decoder 910 can, for example, independently decode two channels of a stereo downmix in a bitstream. Subsequently, the stereo downmix can, for example, be fed back into the filter bank analysis module 940 before being provided as an input to covariance synthesis.

[0234] In the currently described inactive mode (in the inactive mode), the SID decoder 920 and the mono-to-stereo converter 930 can, for example, generate a stereo signal with some spatial decorrelation using the encoded SID information of a mono channel.

[0235] According to an embodiment, an efficient implementation of mono-to-stereo conversion can, for example, be employed, which can, for example, run a random generator twice with different seeds. In an embodiment, the generated noise can, for example, be adjusted using the SID information of the mono channel. Thereby, a stereo signal (with zero coherence) is generated.

[0236] In another embodiment, the mono channel can, for example, be copied to two stereo channels (however, this has the drawback of causing spatial collapse and coherence of one).

[0237] In a preferred embodiment, to generate a stereo signal with coherence and energy similar to the input stereo downmix control parameters such as coherence and / or correlation and scale factors can, for example, be used, and the control parameters and scale factors can, for example, be transmitted as part of the inactive metadata.

[0238]

[0239] where

[0240] s L (n) = s

[0241] s R (n) = 1 - s

[0242] where k is the frequency index, n is the sample index, c(n) is the coherence or correlation transmitted as part of the inactive metadata, s L (n) and s R (n) are scale factors derived from a scale factor s transmitted as part of the inactive metadata, N 1 (k,n), N 2 (k,n) and N 3 (k,n) are random noises generated by different random generators using seeds 1, seed 2, and seed 3 respectively.

[0243] Since the inactive metadata does not include the power ratio and the object index, during direct power calculation, a scale factor that may depend, for example, on the number of objects may be used instead of the power ratio. Alternatively, a scale factor that is transmitted as part of the inactive metadata may be used, for example, to replace the power ratio.

[0244] Figure 11 Illustrated is the generation of a stereo signal according to an embodiment using three random seeds - Seed 1, Seed 2, and Seed 3, the derived scale factor, and control parameters.

[0245] In addition, Figure 11 Illustrated is a random generator that includes a random generator unit 1 and a random generator unit 3 for generating the left channel, and a random generator unit 2 and another random generator unit 3 for generating the right channel.

[0246] In Figure 11 the random generator unit 3 for generating the left channel and the random generator unit 3 for generating the right channel receive the same seed - Seed 3, and thus may generate, for example, the same random noise N 3 (k,n).

[0247] Figure 12 Illustrated is the generation of a stereo signal according to another embodiment, wherein the noise N 3 (j,n) generated by the random generator unit 3 for the left channel is also used for the right channel. In other words, Figure 12 the random generator includes a random generator unit 1, a random generator unit 2, and only a single random generator unit 3.

[0248] In another embodiment, the random generator may, for example, include only a single random generator unit that may be used to sequentially generate random noise N 1 (k,n), N 2 (k,n), and N 3 (k,n) in response to receiving Seed 1, Seed 2, and Seed 3, respectively.

[0249] In other embodiments, the above concepts are similarly applied to the generation of multi - channel signals having more than two channels.

[0250] Additionally, the direct response may be calculated using the direction information of all objects instead of only the primary objects, for example.

[0251] The embodiments allow for the efficient extension of DTX to spatial audio coding using Independent Stream with Metadata (ISM). Spatial audio coding maintains a high perceptual fidelity regarding background noise even for inactive frames, thereby allowing transmission to be interrupted, for example, to save communication bandwidth.

[0252] Decoder-side transport channels with a channel number greater than one can be generated, for example, only by a comfort noise generator (CNG) from a transmitted mono signal such that it exhibits a spatial image according to SID information. The generated transport channels can then be fed back, for example, together with the direct responses, equal power ratios, and prototype matrices calculated from the direction information of all audio objects, into a covariance synthesis module for rendering to a desired output architecture.

[0253] Although some aspects have been described herein in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.

[0254] According to certain implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. The implementation may be carried out using a digital storage medium, such as, for example, a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory, on which electronically readable control signals are stored, which cooperate (or are capable of cooperating) with a programmable computer system to perform the corresponding method. Thus, the digital storage medium may be computer-readable.

[0255] Some embodiments according to the present invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0256] Generally, embodiments of the present invention may be implemented as a computer program product having program code that is operable to perform one of the methods when the computer program product is run on a computer. The program code may be stored, for example, on a machine-readable carrier.

[0257] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the methods described herein.

[0258] In other words, thus, embodiments of the method of the present invention are a computer program having program code for performing one of the methods described herein when the computer program is run on a computer.

[0259] Accordingly, another embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transitory.

[0260] Accordingly, another embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence can be configured, for example, to be transmitted via a data communication connection (such as via the Internet).

[0261] Another embodiment includes a processing device, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.

[0262] Another embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.

[0263] Another embodiment according to the present invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system can include, for example, a file server for transmitting the computer program to the receiver.

[0264] In some embodiments, a programmable logic device (such as a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0265] The apparatus described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0266] The methods described herein can be performed using a hardware device or using a computer or using a combination of a hardware device and a computer.

[0267] The above embodiments are only used to illustrate the principles of the present invention. It is to be understood that modifications and variations of the arrangements and details described herein will be apparent to other technicians in the art. Accordingly, the intention is only limited by the scope of the pending patent claims and not by the specific details set forth in the description and explanation of the embodiments herein.

[0268] Reference:

[0269] [1]WO 2022 / 079049 A2, A. “Apparatus and method for encoding a plurality of audio objects and apparatus and method for decoding using two or more relevant audio objects”.

[0270] [2]WO 2022 / 079044 A1 “Apparatus and method for encoding a plurality of audio objects using direction information during a downmixing or apparatus and method for decoding using an optimized covariance synthesis”.

[0271] [3]3GPP TS26.194; Voice Activity Detector (VAD); - 3GPP technical specification Retrieved on 2009 - 06 - 17.

[0272] [4]3GPP TS26.449, "Codec for Enhanced Voice Services (EVS); Comfort Noise Generation (CNG) Aspects".

[0273] [5]3GPP TS26.450, "Codec for Enhanced Voice Services (EVS); Discontinuous Transmission (DTX)".

[0274] [6]A. Lombard, S. Wilde, E. Ravelli, S. G.Fuchs and M.Dietz,"Frequency-domain Comfort Noise Generation for Discontinuous Transmission in EVS,"2015IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP),Brisbane,QLD,2015,pp.5893-5897,

[0276] doi:10.1109 / ICASSP.2015.7179102.

[0277] [7]WO 2022 / 022876 A1”Apparatus,method and computer program forencoding an audio signal or for decoding an encoded audio scene”.

Claims

1. An audio encoder (100; 800), comprising: a transmission signal generator (110; 710; 810) for generating two or more transmission channels of a transmission signal from an audio input, the audio input including at least one of a plurality of audio input objects and a plurality of audio input channels, a voice activity determiner (120; 820) for determining a voice activity decision of the transmission signal, the language activity decision indicating whether the audio input within the transmission signal exhibits voice activity, and a bitstream generator (130; 850) for generating a bitstream based on the audio input, wherein if the voice activity determiner (120; 820) has determined that the transmission signal exhibits voice activity, the bitstream generator (130; 850) is adapted to encode the two or more transmission channels within the bitstream, wherein if the voice activity determiner (120; 820) has determined that the transmission signal does not exhibit voice activity, the bitstream generator (130; 850) is suitable for encoding information about background noise rather than the two or more transmission channels, wherein the information about background noise includes information about background noise of at least one of the two or more transmission channels or includes information about background noise of a derived signal that depends on at least one of the two or more transmission channels.

2. The audio encoder (100; 800) according to claim 1, wherein the voice activity determiner (120; 820) is configured to determine a separate voice activity decision for each transmission channel among one or more transmission channels of the transmission signal, the separate voice activity decision indicating whether the audio input within the transmission channel exhibits voice activity, and wherein the voice activity determiner (120; 820) is configured to determine the voice activity decision of the transmission signal based on the separate voice activity decisions of each transmission channel among the one or more transmission channels.

3. The audio encoder (100; 800) according to claim 2, wherein the voice activity determiner (120; 820) is configured to determine a separate voice activity decision for each transmission channel among the two or more transmission channels of the transmission signal, the separate voice activity decision indicating whether the audio input within the transmission channel exhibits voice activity, and wherein the voice activity determiner (120; 820) is configured to determine the voice activity decision of the transmission signal based on the separate voice activity decisions of each transmission channel among the two or more transmission channels of the transmission signal.

4. The audio encoder (100; 800) according to claim 3, wherein the voice activity determiner (120; 820) is configured to determine that the transmission signal exhibits voice activity in the case where at least one of the two or more transmission channels of the transmission signal exhibits voice activity, and wherein the voice activity determiner (120; 820) is configured to determine that the transmitted signal does not exhibit voice activity in the case where none of the two or more transmission channels of the transmitted signal exhibits voice activity.

5. The audio encoder (100; 800) according to any one of the preceding claims, wherein the audio encoder (100; 800) is configured to determine whether to transmit the bitstream in which the information about the background noise has been encoded, or whether not to generate and not to transmit the bitstream, in the case where the voice activity determiner (120; 820) has determined that the transmitted signal does not exhibit voice activity.

6. The audio encoder (100; 800) according to any one of the preceding claims, wherein the audio encoder (100; 800) includes a mono signal generator (830) configured to generate the derived signal as a mono signal from at least one of the two or more transmission channels in the case where the voice activity determiner (120; 820) has determined that the transmitted signal does not exhibit voice activity, and wherein the audio encoder (100; 800) includes an information generator configured to generate the information about the background noise as information about the background noise of the mono signal.

7. The audio encoder (100; 800) according to claim 6, wherein the mono signal generator (830) is configured to generate the mono signal by adding the two or more transmission channels or by adding two or more channels derived from the two or more transmission channels, or wherein the mono signal generator (830) is configured to generate the mono signal by selecting the transmission channel that exhibits higher energy among the two or more transmission channels.

8. The audio encoder (100; 800) according to claim 6 or 7, wherein the information generator is configured to generate the information about the background noise of the mono signal as the information about the mono signal.

9. The audio encoder (100; 800) according to claim 8, wherein the information generator is configured to generate a silence insertion description of the background noise of the mono signal as the information about the background noise of the mono signal.

10. The audio encoder (100; 800) according to any one of the preceding claims, wherein the audio encoder (100; 800) includes a direction information determiner (802) for determining direction information based on the audio input, wherein the audio encoder (100; 800) includes a direction information quantizer (804) for quantifying the direction information to obtain quantized direction information, and wherein the bitstream generator (130; 850) is configured to encode the quantized direction information in the bitstream.

11. The audio encoder (100; 800) according to claim 10, wherein the transmission signal generator (110; 710; 810) is configured to generate the two or more transmission channels of the transmission signal from the audio input using the direction information.

12. The audio encoder (100; 800) according to claim 10 or 11, wherein the audio input includes the plurality of audio input objects, wherein the direction information includes information on the azimuth and elevation angles of an audio input object among the plurality of audio input objects of the audio input.

13. The audio encoder (100; 800) according to any one of the preceding claims, wherein the audio encoder (100; 800) includes an active metadata generator (825) configured to generate metadata in the case where the voice activity determiner (120; 820) has determined that the transmission signal exhibits voice activity, the metadata including at least one of the quantized direction information, object index, and power ratio of the plurality of audio input objects and / or the plurality of audio input channels of the audio input.

14. The audio encoder (100; 800) according to any one of the preceding claims, wherein the audio input includes the plurality of audio input objects, and wherein the audio encoder (100; 800) includes an inactive metadata generator (826) configured to generate metadata in the case where the voice activity determiner (120; 820) has determined that the transmission signal does not exhibit voice activity, the metadata including quantized direction information and control parameters, the control parameters including, for example, a scale factor and / or coherence or correlation.

15. The audio encoder (100; 800) according to claim 13 and claim 14, wherein the quantization resolution of the direction information generated by the inactive metadata generator (826) is different from the metadata generated by the active metadata generator (825).

16. The audio encoder (100; 800) according to claim 14 or 15, further according to claim 13, wherein the inactive metadata generator (826) is configured to generate the control parameters such that the characteristics of the control parameters are different from the characteristics of the power ratio and object index generated by the active metadata generator (825), for example, wherein the control parameters include, for example, the scale factor and / or, for example, the coherence or the correlation.

17. The audio encoder (100; 800) according to any one of the preceding claims, wherein the audio input includes a plurality of audio input objects and metadata associated with the audio input objects.

18. The audio encoder (100; 800) according to any one of the preceding claims, The transmission signal generator (110; 710; 810) is configured to generate the two or more transmission channels of the transmission signal from the audio input, including obtaining a downmix as the transmission signal by downmixing at least one of a plurality of audio input objects and a plurality of audio input channels, the transmission signal including two or more downmix channels as the two or more transmission channels.

19. The audio encoder (100; 800) according to any one of claims 10 to 13 and as claimed in claim 18, wherein, if the audio input within the transmission signal does not exhibit voice activity, the direction information quantizer (804) is configured to determine the quantized direction information such that the quantization resolution of the quantized direction information is different from the quantization resolution used for calculating the downmix.

20. The audio encoder (100; 800) according to any one of the preceding claims, further according to claim 14, wherein the bitstream generator (130; 850) is configured to encode the control parameter within the bitstream if the voice activity determiner (120; 820) has determined that the transmission signal does not exhibit voice activity, wherein the control parameter is suitable for controlling the generation of an intermediate signal from random noise, wherein the control parameter includes a plurality of parameter values for a plurality of subbands, or wherein the control parameter is a single wideband control parameter.

21. The audio encoder (100; 800) according to claim 20, wherein the audio encoder (100; 800) is configured to generate the control parameter by selecting whether the control parameter includes the plurality of parameter values for the plurality of subbands, or whether the control parameter is the single wideband control parameter according to the available bitrate.

22. The audio encoder (100; 800) according to any one of the preceding claims, wherein the transmission signal generator (110; 710; 810) is configured to encode the audio input by applying code-excited linear prediction or by applying a modified discrete cosine transform or by applying a combination of the code-excited linear prediction and the modified discrete cosine transform.

23. The audio encoder (100; 800) according to any one of the preceding claims, wherein, if the audio input includes the plurality of audio input channels and does not include the plurality of audio input objects, the number of the two or more transmission channels is less than the number of the plurality of audio input channels, wherein, if the audio input includes the plurality of audio input objects and does not include the plurality of audio input channels, the number of the two or more transmission channels is less than the number of the plurality of audio input objects, wherein, if the audio input includes both the plurality of audio input objects and the plurality of audio input channels, the number of the two or more transmission channels is less than the sum of the number of the plurality of audio input channels and the number of the plurality of audio input objects; or Wherein, if the audio input includes the plurality of audio input channels and does not include the plurality of audio input objects, the number of the two or more transmission channels is less than or equal to the number of the plurality of audio input channels, Wherein, if the audio input includes the plurality of audio input objects and does not include the plurality of audio input channels, the number of the two or more transmission channels is less than or equal to the number of the plurality of audio input objects, Wherein, if the audio input includes both the plurality of audio input objects and the plurality of audio input channels, the number of the two or more transmission channels is less than or equal to the sum of the number of the plurality of audio input channels and the number of the plurality of audio input objects.

24. System, Comprising: An audio encoder (100; 800) as claimed in any one of the preceding claims, and An audio decoder (200; 900), Wherein the audio decoder (200; 900) comprises: An input interface (210; 902) for receiving a bitstream, the bitstream being dependent on audio content comprising at least one of a plurality of audio objects and a plurality of audio channels; wherein a transmission signal comprising two or more transmission channels is encoded in the bitstream, and the audio content is encoded in the transmission signal; or wherein information about background noise rather than the transmission signal is encoded in the bitstream, the information about background noise comprising information about background noise of at least one of the two or more transmission channels or information about background noise of a derived signal, the derived signal being dependent on at least one of the two or more transmission channels; and A renderer (220; 950) for generating one or more audio output signals based on the audio content encoded in the bitstream; Wherein, if the transmission signal comprising the two or more transmission channels is encoded in the bitstream, the renderer (220; 950) is configured to generate the one or more audio output signals based on the two or more transmission channels, and Wherein, if the information about background noise rather than the transmission signal is encoded in the bitstream, the renderer (220; 950) is configured to generate the one or more audio output signals based on the information about background noise, Wherein the audio encoder (100; 800) is configured to generate a bitstream from an audio input, and Wherein the audio decoder (200; 900) is configured to generate one or more audio output signals from the bitstream.

25. A method for audio encoding, wherein the method Comprises: Generating two or more transmission channels of a transmission signal from an audio input, the audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels, Determining a voice activity decision of the transmission signal, the voice activity decision indicating whether the audio input within the transmission signal exhibits voice activity, and Determining a bitstream based on the audio input, Wherein, if it is determined that the transmission signal exhibits voice activity, the method includes encoding the two or more transmission channels within the bitstream, Wherein, if it is determined that the transmission signal does not exhibit voice activity, the method includes encoding information on background noise for at least one of the two or more transmission channels or information on background noise for a derived signal, which depends on at least one of the two or more transmission channels, rather than encoding the two or more transmission channels.

26. A computer program for implementing the method according to claim 25 when the computer program is executed on a computer or a signal processor.

Citation Information

Patent Citations

  • Apparatus, method and computer program for encoding an audio signal or for decoding an encoded audio scene

    WO2022022876A1

  • Apparatus and method for encoding a plurality of audio objects using direction information during a downmixing or apparatus and method for decoding using an optimized covariance synthesis

    WO2022079044A1