Comfort noise generation for multi-modal spatial audio coding

CN122598665APending Publication Date: 2026-08-18TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610888576.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-07-07
Filing Date
2021-07-06
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

非活动段(例如,语音中的停顿)中的完全静音感知起来是恼人的,并且经常导致误解呼叫已中断

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598665A_ABST
    Figure CN122598665A_ABST
Patent Text Reader

Abstract

A method for generating comfort noise is provided. The method comprises providing a first set N1 of background noise parameters for at least one audio signal in a first spatial audio coding mode; and providing a second set N2 of background noise parameters for at least one audio signal in a second spatial audio coding mode. The first spatial audio coding mode is used for active segments; the second spatial audio coding mode is used for non-active segments. The method further comprises adapting the first set N1 of background noise parameters to the second spatial audio coding mode, thereby providing an adapted first set of background noise parameters. The method further comprises generating comfort noise parameters by combining N1 and N2 over a transition period. The method further comprises generating comfort noise based on the comfort noise parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Case Analysis

[0001] This application is a divisional application of the invention patent application filed on July 6, 2021, with application number 202180044437.3 and invention title "Comfort Noise Generation Based on Multi-Mode Spatial Audio Coding". Technical Field

[0002] Embodiments relating to multimodal spatial audio discontinuity transmission (DTX) and comfort noise generation are disclosed. Background Technology

[0003] Despite the increasing capacity in telecommunications networks, limiting the bandwidth required for each communication channel remains a significant concern. In mobile networks, smaller transmission bandwidth per call means that the mobile network can serve a large number of users in parallel. Reduced transmission bandwidth also results in lower power consumption in both mobile devices and base stations. This translates to energy and cost savings for mobile operators, while end users experience extended battery life and increased talk time.

[0004] One method for reducing transmission bandwidth in voice communications is to utilize natural pauses in speech. In most conversations, only one speaker is active at a time; therefore, pauses in one direction typically account for more than half of the signal. A method to reduce transmission bandwidth by utilizing this property of typical conversations is to employ a discontinuous transmission (DTX) scheme, in which the encoding of the active signal is interrupted during speech pauses. The DTX scheme is standardized for all 3GPP mobile phone standards, including 2G, 3G, and VoLTE. It is also commonly used in Voice over IP (VoIP) systems.

[0005] During speech pauses, extremely low bit-rate encoding of background noise is typically transmitted to allow a comfort noise generator (CNG) at the receiver to fill the pause with background noise that has similar characteristics to the original noise. CNG makes the sound more natural because the background noise is preserved and does not turn on and off with the speech. Complete silence in inactive segments (e.g., pauses in speech) can be annoying and often leads to the misconception that the call has been interrupted.

[0006] The DTX scheme may include a Voice Activity Detector (VAD), which instructs the system whether to use an activity signal coding method (when voice activity is detected) or low-rate background noise coding (when no voice activity is detected). This in Figure 1The diagram is schematically shown. System 100 includes a VAD 102, a speech / audio encoder 104, and a CNG encoder 106. When VAD 102 detects speech activity, it signals to use “high bit rate” encoding with the speech / audio encoder 104, and when no speech activity is detected, it signals to use “low bit rate” encoding with the CNG encoder 106. The system can be summarized as distinguishing between other source types by using a (general) sound activity detector (GSAD or SAD) that can distinguish speech not only from background noise but also detect music or other signal types (which are considered relevant).

[0007] Communication services can be further enhanced by supporting stereo or multi-channel audio transmission. For stereo transmission, one solution is to use two mono codecs that independently encode the left and right portions of the stereo signal. A more efficient and complex solution is often to combine the encoding of the left and right input signals, a process known as joint stereo coding. The terms "signal" and "channel" are often used interchangeably to refer to the signal of an audio channel, such as the left and right channels of stereo audio. Summary of the Invention

[0008] A common method for generating comfort noise (CN) (used in all 3GPP speech codecs) is to transmit information relating to the energy and spectral shape of background noise during speech pauses. This can be accomplished with significantly fewer bits than the conventional coding of a speech segment. On the receiver side, CN is generated by creating a pseudo-random signal and then using filters to shape the signal's spectrum based on information received from the transmitter. This signal generation and spectral shaping can be performed in either the time or frequency domain.

[0009] In a typical DTX system, the capacity gain comes partly from the fact that fewer bits are used to encode the CN (Channel Code), but primarily from the fact that CN parameters are typically not sent as frequently as regular encoded parameters. This usually works well because background noise characteristics do not change as rapidly as, for example, speech signals. The encoded CN parameters are sent in what are commonly referred to as “SID frames,” where SID stands for silence descriptor. Typically, CN parameters are sent every 8th speech encoder frame, where a speech encoder frame is typically 20 ms. The CN parameters are then used as the basis for CNG (Channel Code Generation) in the receiver until the next set of CN parameters is received. Figure 2This is illustrated schematically, showing that when "active coding" is enabled (also known as the active segment or active coding segment), there is no "CN coding," and when "active coding" is not enabled (also known as the inactive segment or inactive coding segment), then "CN coding" is performed intermittently every 8th frame.

[0010] One solution to avoid unwanted fluctuations in the CN is to sample the CN parameters during all 8 speech encoder frames and then send parameters based on (e.g., by averaging) all 8 frames. Figure 3 This is illustrated schematically, showing the average interval over 8 frames. While a fixed SID interval of 8 frames is typical for speech codecs, shorter or longer intervals can be used to transmit CNG parameters. The SID interval can also be based, for example, on signal characteristics varying over time, such that CN parameters are updated less frequently for stationary signals and more frequently for changing signals.

[0011] Speech / audio codecs with DTX systems incorporate low-bit-rate coding patterns used for encoding inactive segments (e.g., non-speech segments), allowing the decoder to generate comfort noise with characteristics similar to the input signal. An example is the 3GPP EVS codec. In the EVS codec, the decoder also performs the following function: analyzing the signal during an active segment and using the results of this analysis to improve the generation of comfort noise in the next inactive segment.

[0012] The EVS codec is an example of a multi-mode codec, where a combination of different coding techniques is used to create a highly flexible codec to handle, for example, different input signals and different network conditions. Future codecs will be even more flexible, supporting stereo and multi-channel audio as well as virtual reality scenarios. To cover a wide range of input signals, such a codec will use several different coding techniques that can be adaptively selected based on characteristics such as the input signal and network conditions.

[0013] Given the specific purpose of CN encoding and the expectation of keeping CN encoding low complexity, it is reasonable to have a specific mode for CN encoding, even if the encoder combines several different modes used for encoding speech, music or other signals.

[0014] Ideally, the conversion from active coding to CN coding should be audible, but this is not always achievable. The risk of audible conversion is higher when active segments are encoded using coding techniques different from CN coding. A typical example is... Figure 4 As shown, the CN level is higher than the preceding active segment. Note that although only one signal is shown, similar audible transitions can exist in all channels.

[0015] Typically, the comfort noise coding process generates CN parameters, which allow the decoder to recreate comfort noise with an energy corresponding to the energy of the input signal. In some cases, modifying the level of the comfort noise can be advantageous, for example, by slightly reducing it to achieve noise suppression in speech pauses or to better match the level of background noise being reproduced during the encoding of the active signal.

[0016] Active signal coding can have a noise suppression effect, which reduces the level of reproduced background noise to be lower than the original signal, especially when noise is mixed with speech. This is not necessarily an intentional design choice; it can be a side effect of the coding scheme used. If this level reduction is fixed, or fixed for a particular coding pattern, or otherwise known in the decoder, the level of comfort noise can be reduced by the same amount to make the transition from active coding to comfort noise smooth. However, if the level reduction (or increase) is signal-dependent, there may be a step in energy when switching from active coding to CN coding. This step change in energy will be perceived as annoying by the listener, especially when the level of comfort noise is higher than the level of noise in the active coding preceding the comfort noise.

[0017] Combining multichannel audio codecs (e.g., stereo codecs) can present additional challenges, requiring consideration not only of mono signal characteristics but also spatial properties such as inter-channel level differences and inter-channel coherence. For encoding and representing such multichannel signals, individual encoding of each channel (including DTX and CNG) is inefficient due to redundancy between channels. As an alternative, various multichannel coding techniques can be used for more efficient representation. Stereo codecs, for example, can use different encoding modes for different signal characteristics of the input channels, such as a single audio source (speaker) versus multiple audio sources, different capture techniques / microphone settings, and different stereo codec modes for DTX operation.

[0018] For CN generation, a compact parameterized stereo representation is suitable, as it can efficiently represent the signal and spatial characteristics of the CN. This parameterized representation typically uses a downmixed signal and additional parameters describing the stereo image to represent stereo channel pairs. However, for encoding active signal segments, different stereo coding techniques may offer higher performance. Note that although a single signal is shown, similar audible transitions can exist for all channels.

[0019] Figure 4An example operation of a multi-mode audio codec is illustrated. For active segments, the codec operates under two spatial coding modes (mode_1, mode_2) (e.g., stereo mode), selected, for example, depending on signal characteristics, bit rate, or similar control features. When the codec switches to inactive (SID) coding using the DTX scheme, the spatial coding mode changes to the spatial coding mode used for SID coding and CN generation (mode_CNG). It should be noted that, in terms of their spatial representation, mode_CNG can be similar to or even identical to one of the modes used for active coding (i.e., mode_1 or mode_2 in this example). However, mode_CNG typically operates at a much lower bit rate than the corresponding mode used for active signal coding.

[0020] Multi-mode mono audio codecs (e.g., 3GPP EVS codecs) efficiently handle the transitions between different codec modes and CN generation during DTX operation. These methods typically analyze signal characteristics at the end of active speech segments, for example, during the so-called VAD tail delay protection period indicating the background signal, but for safety, the regular transmission remains active to avoid speech clipping. However, for multi-channel codecs, this existing technology may be insufficient and leads to annoying transitions between active and inactive coding (DTX / CNG operation), especially when different spatial audio representations or multi-channel / stereo coding techniques are used for active and inactive (SID / CNG) coding.

[0021] Figure 4 This illustrates the troublesome transition from active coding using a first spatial coding mode to inactive (SID) coding using a second spatial coding mode and CN generation. Despite the use of existing methods for smooth active-to-inactive transitions for mono signals, a visibly noticeable transition can exist due to the change in spatial coding mode.

[0022] By converting and adapting the background noise characteristics estimated during operation in the first spatial coding mode to background noise characteristics suitable for CNG in the second spatial coding mode, the embodiment provides a solution to the problem of perceptually annoying activity-to-inactivity (CNG) conversion. The obtained background noise characteristics are further adapted based on the parameters sent to the decoder in the second spatial coding mode.

[0023] The implementation improves the conversion between active coding and comfort noise (CN) in a multi-mode spatial audio codec by making the conversion to CN smoother. This enables DTX to be used in high-quality applications, thereby reducing the bandwidth required for transmission in such services, and also improves perceived audio quality.

[0024] According to a first aspect, a method for generating comfortable noise is provided. The method includes: providing a first set N1 of background noise parameters for at least one audio signal in a first spatial audio coding mode, wherein the first spatial audio coding mode is used for active segments. The method also includes: providing a second set N2 of background noise parameters for at least one audio signal in a second spatial audio coding mode, wherein the second spatial audio coding mode is used for inactive segments. The method further includes: adapting the first set N1 of background noise parameters to the second spatial audio coding mode, thereby providing an adapted first set of background noise parameters. The method includes: combining a first set of appropriately matched background noise parameters within a conversion period. The method involves generating comfort noise parameters based on a second set N2 of background noise parameters. The method further includes generating comfort noise for at least one output audio channel based on the comfort noise parameters.

[0025] In some embodiments, generating comfort noise for at least one output audio channel includes applying the generated comfort noise parameters to at least one intermediate audio signal. In some embodiments, generating comfort noise for at least one output audio channel includes upmixing the at least one intermediate audio signal. In some embodiments, the at least one audio signal is based on signals from at least two input audio channels, and wherein a first set N1 of background noise parameters and a second set N2 of background noise parameters are each based on a single audio signal, wherein the single audio signal is based on downmixing the signals from at least two input audio channels. In some embodiments, the at least one output audio channel includes at least two output audio channels.

[0026] In some embodiments, providing a first set N1 of background noise parameters includes receiving the first set N1 of background noise parameters from a node. In some embodiments, providing a second set N2 of background noise parameters includes receiving the second set N2 of background noise parameters from a node. In some embodiments, adapting the first set N1 of background noise parameters to a second spatial audio coding mode includes applying a transform function. In some embodiments, the transform function includes a function of N1, NS1, and NS2, wherein NS1 includes a first set of spatial coding parameters indicating the undermixing and / or spatial characteristics of the background noise of the first spatial audio coding mode, and NS2 includes a second set of spatial coding parameters indicating the undermixing and / or spatial characteristics of the background noise of the second spatial audio coding mode.

[0027] In some embodiments, applying the transformation function includes calculating ,in, It is a scalar compensation factor. In some embodiments, It has the following values: , where ratio LRIt's a mixed bag. Corresponding to coherence or correlation coefficient, and c is derived from Given, among which, and This is the gain parameter. In some embodiments, It has the following values: , where ratio LR It's a mixed bag. Corresponding to coherence or correlation coefficient, and c is derived from Given, among which, , and It is the gain parameter.

[0028] In some embodiments, the conversion period is a fixed-length inactive frame. In some embodiments, the conversion period is a variable-length inactive frame. In some embodiments, a first set of appropriately matched background noise parameters is combined within the conversion period. The second set of background noise parameters N2 is used to generate comfortable noise, including: application The weighted average of N2. In some embodiments, this is achieved by combining a first set of appropriately matched background noise parameters within the conversion period. The calculation of the second set N2 of background noise parameters to generate comfort noise parameters includes:

[0029]

[0030] Where CN is the generated comfort noise parameter. It is the current inactive frame count, and k is the indicator for the application. The length of the conversion period is the weighted average of the number of inactive frames and N2. In some embodiments, this is achieved by combining a first set of suitable background noise parameters within the conversion period. The calculation of the second set N2 of background noise parameters to generate comfort noise parameters includes:

[0031]

[0032] in,

[0033]

[0034] Where CN is the generated comfort noise parameter. It is the current inactive frame count, and k is the indicator for the application. The length of the conversion period for the weighted average number of inactive frames of N2, and It is a frequency subband index. In some embodiments, generating comfort noise parameters includes targeting the frequency subband. at least one frequency coefficient calculate:

[0035]

[0036] In some embodiments, k is determined as:

[0037]

[0038] Where M is the maximum value of k, and r1 is the energy ratio of the estimated background noise level, determined as follows:

[0039]

[0040] in, There are N frequency sub-bands. This refers to the case of a given subband b. The adapted background noise parameters, and This refers to the background noise parameter adapted for N2 of a given subband b.

[0041] In some embodiments, by combining a first set of appropriately matched background noise parameters within the conversion cycle The second set N2 of background noise parameters is used to generate comfort noise parameters, including: application A nonlinear combination of N2. In some embodiments, the method further includes: determining a first set of appropriately matched background noise parameters by combining them within a conversion period. The comfort noise parameters are generated by combining a second set N2 of background noise parameters, wherein the first set of background noise parameters is appropriately matched within the conversion period. The second set N2 of background noise parameters is used to generate the comfort noise parameters as a first set of background noise parameters that are appropriately matched within the conversion period. The process is performed to generate the comfort noise parameters by combining the second set N2 of background noise parameters.

[0042] In some embodiments, a first set of appropriately matched background noise parameters is determined by combining them within the conversion cycle. The comfort noise parameters are generated by combining a second set of background noise parameters N1 and a second set of background noise parameters N2, based on the evaluation of the first energy of the primary channel and the second energy of the secondary channel. In some embodiments, the first set of background noise parameters N1, the second set of background noise parameters N2, and the first set of adapted background noise parameters are used. One or more of the parameters include one or more parameters describing signal characteristics and / or spatial characteristics, and the one or more parameters include one or more of the following: (i) linear prediction coefficients representing signal energy and spectral shape; (ii) excitation energy; (iii) inter-channel coherence; (iv) inter-channel level difference; and (v) side gain parameters.

[0043] According to a second aspect, a node is provided, comprising processing circuitry and a memory containing instructions executable by the processing circuitry. The processing circuitry is operable to: provide a first set N1 of background noise parameters for at least one audio signal in a first spatial audio coding mode, wherein the first spatial audio coding mode is used for an active segment. The processing circuitry is operable to: provide a second set N2 of background noise parameters for at least one audio signal in a second spatial audio coding mode, wherein the second spatial audio coding mode is used for an inactive segment. The processing circuitry is operable to: adapt the first set N1 of background noise parameters to the second spatial audio coding mode, thereby providing an adapted first set of background noise parameters. The processing circuitry is operable to: combine a first set of appropriately matched background noise parameters within the conversion cycle. The second set N2 of background noise parameters is used to generate comfort noise parameters. The processing circuitry is operable to generate comfort noise for at least one output audio channel based on the comfort noise parameters.

[0044] According to a third aspect, a computer program including instructions is provided that, when executed by processing circuitry, causes the processing circuitry to perform a method according to any embodiment of the first aspect.

[0045] According to the fourth aspect, a carrier containing the computer program of the third aspect is provided, wherein the carrier is one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium. Attached Figure Description

[0046] The accompanying drawings, which are included in and form part of this specification, illustrate various embodiments.

[0047] Figure 1 A system for generating comfort noise is shown.

[0048] Figure 2 The encoding of active and inactive segments is shown.

[0049] Figure 3 The encoding of inactive segments is shown.

[0050] Figure 4 The encoding of active and inactive segments using multiple encoding modes is shown.

[0051] Figure 5A system for decoding comfort noise according to an embodiment is shown.

[0052] Figure 6 An encoder according to an embodiment is shown.

[0053] Figure 7 A decoder according to an embodiment is shown.

[0054] Figure 8 This is a flowchart based on an embodiment.

[0055] Figure 9 The encoding of active and inactive segments using various encoding modes according to an embodiment is illustrated.

[0056] Figure 10 This is a schematic representation of stereo downmixing according to an embodiment.

[0057] Figure 11 This is a schematic representation of stereo upmixing according to an embodiment.

[0058] Figure 12 This is a flowchart based on an embodiment.

[0059] Figure 13 This is a block diagram of an apparatus according to an embodiment.

[0060] Figure 14 This is a block diagram of an apparatus according to an embodiment. Detailed Implementation

[0061] The following embodiments describe a stereo codec including an encoder and a decoder. The codec can utilize more than one spatial coding technique to more efficiently compress stereo audio with various characteristics, such as single-speaker speech, two-speaker speech, music, and background noise.

[0062] The codec can be used by nodes (e.g., user equipment (UE)). For example, two or more nodes can communicate with each other, for example, via a UE connected to a telecommunications network using network standards such as 3G, 4G, 5G, etc. A node can be an "encoding" node, in which speech is encoded and sent to a "decoding" node, where speech is decoded. The "encoding" node can send background noise parameters to the "decoding" node, which can use these parameters to generate comfortable noise according to any embodiment disclosed herein. For example, in two-way voice communication, nodes can also switch between "encoding" and "decoding". In this case, a given node can be both an "encoding" node and a "decoding" node, and can switch between one and the other or perform both tasks simultaneously.

[0063] Figure 5A system 500 for decoding comfort noise according to an embodiment is illustrated. System 500 may include a speech / audio decoder 502, a CNG decoder 504, a background estimator 506, a transform node 508, and a CN generator 510. A received bitstream enters system 500; this bitstream may be a "high" bitrate stream (for active segments) or a "low" bitrate stream (for inactive segments). If the bitstream is a "high" bitrate stream (for active segments), it is decoded by the speech / audio decoder 502, which generates the speech / audio output. Additionally, the output of the speech / audio decoder 502 may be passed to the background estimator 506, which can estimate background noise parameters. The estimated background noise parameters may be passed to the transform node 508, which applies a transform to the parameters, which are then sent to the CN generator 510. If the bitstream is a "low" bitrate stream (for inactive segments), it is decoded by the CNG decoder 504 and passed to the CN generator 510. The CN generator 510 can generate comfort noise based on the decoded bitstream and can additionally use information from the transform node 508 regarding the estimated background parameters during the active segment (and similarly use information from nodes 502 and / or 506). The result of the CN generator 510 is a CNG output, which can be applied to the audio output channel.

[0064] Two-channel parametric stereo coding

[0065] Joint stereo coding techniques aim to reduce the information required to represent the audio channel pair to be encoded (e.g., left and right channels). Various (down)mixing techniques can be used to form a channel pair with lower correlation than the original left and right channels, thus containing less redundant information, making coding more efficient. One well-known technique of this kind is center-side stereo, where the sum and difference of the input signals form the center and side channels. Further extensions utilize more adaptive downmixing schemes aimed at minimizing redundant information within the channels for more efficient coding. This adaptive downmixing can be based on energy compression techniques such as principal component analysis or Karhunen-Loève transform, or any other suitable technique. The adaptive downmixing process can be written as:

[0066]

[0067] Where P and S are the primary and secondary (submixer) channels, respectively; L and R are the left and right channel inputs, respectively; and ratio is... LR It's a mixed bag.

[0068] Downmixing ratio is calculated based on the characteristics of the input signal; it can be based on, for example, interchannel correlation and level difference. Fixed. This corresponds to a regular mid-transform / side-transform. Downmixing can be performed on audio samples in the time domain or on frequency ranges or subbands in the frequency domain. For clarity, sample, range, and / or subband indices have been omitted from the equations presented here.

[0069] In the decoder, the decoded parameters are used. and decoded audio channels and To perform the inverse operation (upmixing) to recreate the left and right output signals (respectively). and ):

[0070]

[0071] in

[0072]

[0073] In this case, the downmixing parameters It is typically encoded and sent to the decoder for upmixing. Additional parameters can be used to further improve compression efficiency.

[0074] Mono Parametric Stereo Coding

[0075] Depending on the signal characteristics, other stereo coding techniques can be more efficient than two-channel parametric stereo coding. Especially for CNG, the bit rate of the transmitted SID parameters needs to be reduced to achieve an efficient DTX system. In this case, only one of the sub-mix channels can be described or encoded (e.g., In this case, the additional parameters encoded and sent to the decoder can be used to estimate the other channels required for upmixing (e.g., The stereo parameters will allow the decoder to invert the encoder downmix in an approximate manner and recreate the (upmixed) stereo signal (upmixed signal pair) from the decoded mono downmixed signal.

[0076] Figure 6 and Figure 7 The diagram shows a block diagram of an encoder and decoder operating in the Discrete Fourier Transform (DFT) domain. Figure 6As shown, encoder 600 includes a DFT transform unit 602, a stereo processing and mixing unit 604, and a mono speech / audio encoder 606. A time-domain stereo input enters encoder 600, where the time-domain stereo signal undergoes a DFT transform performed by DFT transform unit 602. DFT transform unit 602 can then pass its output (DFT-transformed signal) to stereo processing and mixing unit 604. Stereo processing and mixing unit 604 can then perform stereo processing and mixing, outputting a mono mix (or downmix) and stereo parameters. The mono mix can be passed to mono speech / audio encoder 606, which produces an encoded mono signal. Figure 7 As shown, the decoder 700 includes a mono speech / audio decoder 702, a stereo processing and upmixing unit 704, and an inverse DFT transform unit 706. The encoded mono signal and stereo parameters enter the decoder 700. The encoded mono signal is passed to the mono speech / audio decoder 702, which causes the mono mix signal to be sent to the stereo processing and upmixing unit 704. The stereo processing and upmixing unit 704 also receives the stereo parameters and performs stereo processing and upmixing on the mono mix signal using these parameters. The output is then passed to the inverse DFT transform unit 706, which outputs a time-domain stereo output.

[0077] Appropriate parameters for describing the spatial characteristics of stereo signals are usually related to inter-channel level difference (ILD), inter-channel coherence (IC), inter-channel phase difference (IPD), and inter-channel time difference (ITD).

[0078] The processes of creating the downmix signal and extracting stereo parameters in the encoder can be performed in the time domain; alternatively, this process can be performed in the frequency domain by first transforming the input signal to the frequency domain, for example, through a Discrete Fourier Transform (DFT) or any other suitable filter bank. This also applies to the decoder, where processes for stereo synthesis, for example, can be performed in either the time or frequency domain. For frequency domain processing, a frequency-adaptive downmixing process can be used to optimize downmixing across different frequency bands, for example, to avoid signal cancellation in the downmixed signal. Furthermore, channel time alignment can be performed before downmixing based on the inter-channel time difference determined at the encoder.

[0079] For CNG, and The signal can be generated at the decoder based on a noise signal spectrally shaped using transmitted SID parameters that describe the spectral properties of the estimated background noise characteristics. Furthermore, the coherence, level, time, and phase difference between channels can be described to allow for good reconstruction of the spatial characteristics of the background noise, represented by CN.

[0080] In one embodiment, the side gain parameter Used to describe and coherent The amount from Estimate or predict The edge gain can be estimated as a normalized inner product (or dot product):

[0081]

[0082] in, express and The inner product of signals. This can be shown as the inner product of signals. and In the multi-dimensional space that spans arrive The projection onto the frequency domain, for example, is a vector of time-domain samples or correspondingly in the frequency domain.

[0083] Use passive blending, for example, as follows:

[0084]

[0085] The corresponding supermix can be obtained as:

[0086]

[0087] in, and Unrelated, having the same The same spectral characteristics and signal energy. Here, Uncorrelated components The gain factor can be obtained from interchannel coherence as follows:

[0088]

[0089] Given frequency vocal tract coherence It is given by the following formula:

[0090]

[0091] in, and Indicates two audio channels and The corresponding power spectrum, and It has two channels. and The cross-power spectrum. In a DFT-based solution, the spectrum can be represented by the DFT spectrum. Specifically, according to an embodiment, the frame index... and frequency range index Spatial coherence It was identified as:

[0092]

[0093] in, and Representing a frame and frequency range The left and right audio channels.

[0094] Alternatively or additionally, inter-channel cross-correlation (ICC) can be estimated. Conventional ICC estimation relies on the cross-correlation function (CCF). It consists of two waveforms. and The similarity between them is measured and is usually defined in the time domain as follows:

[0095]

[0096] in, It is a time lag, and It is the expectation operator. For length... The cross-correlation of signal frames is typically estimated as:

[0097]

[0098] ICC is then obtained as the maximum value of CCF normalized by the signal energy, as follows:

[0099]

[0100] In this case, the gain factor It can be calculated as:

[0101]

[0102] It can be noted that coherence or correlation coefficient corresponds to Figure 10 Angle shown ,in, .

[0103] Furthermore, if available parameters exist to describe those characteristics, inter-channel phase and time difference or similar spatial characteristics can be synthesized.

[0104] DTX operation with stereo mode switching

[0105] In an example embodiment, the stereo codec is based on... Figure 4 The operation uses a first stereo mode for active signal encoding and a second stereo mode for inactive (SID) encoding for CNG at the decoder.

[0106] Background noise estimation

[0107] In this embodiment, parameters for comfort noise generation (CNG) in the transition section are determined based on two different background noise estimates. Figure 9 An example of such a transition segment at the beginning of the comfort noise segment is shown. The first background noise estimate can be determined based on the background noise estimate performed by the decoder when operating in the first stereo mode (e.g., based on minimum statistics analysis of the decoded audio signal). The second background noise estimate can be determined based on the estimated background noise characteristics of the encoded audio signal determined at the encoder operating in the second stereo mode for SID encoding.

[0108] Background noise estimation can include one or more parameters describing signal characteristics, such as signal energy and spectral shape described by linear prediction coefficients and excitation energy or equivalent representations (e.g., line spectrum pairs (LSPs), line spectrum frequencies (LSFs), etc.). Background noise characteristics can also be represented in a transform domain (e.g., the Discrete Fourier Transform (DFT) or Modified Discrete Cosine Transform (MDCT) domain) as, for example, amplitude or power spectrum. Using minimum statistics to estimate the level and spectral shape of background noise during active coding is just one example of a technique that can be used; other techniques may also be employed. Furthermore, the undermixing and / or spatial characteristics of the background estimation can be estimated, encoded, and sent to the decoder in the SID frame.

[0109] In one embodiment, the first set of background noise parameters Describes the first stereo coding mode The spectral characteristics of the vocal tract. A set of spatial coding parameters. This describes the downmixing and / or spatial characteristics of the background noise in the first stereo mode. A second set of background noise parameters. Describes the second stereo coding mode The spectral characteristics of the vocal tract. A set of spatial coding parameters. The background noise downmixing and / or spatial characteristics of the second stereo mode are described.

[0110] In one embodiment, the set of spatial coding parameters This includes the undermixing parameter, such as the undermixing parameter corresponding to the mixing factor according to equation (1). .

[0111] In one embodiment, the set of spatial coding parameters Includes: corresponding to and Relevant The first gain parameter of the gain of the component ; and corresponding to Irrelevant (unrelated) The second gain parameter of the gain of the component Spatial encoding parameters This can represent the gain of a complete frame of an audio sample or the corresponding gain in a specific frequency subband. The latter implies the existence of a gain parameter that represents the gain parameter of the frame of the audio sample. and The set. In another embodiment, the second gain parameter It is based on the interchannel coherence received at the decoder ( ) or correlation coefficient ( This is determined at the decoder. Similarly, inter-channel coherence can be described in frequency subbands, resulting in a set of parameters for each audio frame.

[0112] Although various representations (e.g., frequency subband energy or linear prediction coefficients and excitation energy) can be used to describe background noise characteristics, and It can be converted to a general representation (e.g., a DFT field). This means that... and It can be obtained as a function of the confirmed parameters describing the background noise characteristics, for example, through DFT transformation. In one embodiment, the background noise parameters... and It is expressed as frequency band energy or amplitude.

[0113] Background noise estimation transformation

[0114] To smoothly transition from active signal encoding in the first stereo mode to SID encoding and CNG at the decoder, a first set of background noise parameters is used. (Originally from the first stereo mode) was adapted to a second stereo mode for SID encoding and CNG. The set of transformation parameters. It can be determined as:

[0115]

[0116] in, It is a transform function. The transform function can be frequency-dependent or constant at all frequencies.

[0117] In another embodiment, the set of parameters of the transformation It can be determined as:

[0118]

[0119] In one embodiment, the set of parameters of the transformation identified as Scaling version:

[0120]

[0121] in, It is between two stereo modes The scalar compensation factor for the energy difference.

[0122] If the downmix of the first stereo mode is

[0123]

[0124] And the downmixing in the second stereo mode is

[0125]

[0126] Scaling factor It can be determined as:

[0127]

[0128] in

[0129]

[0130] Undermixing factor Origin (First stereo mode), and gain parameters and Origin (Second stereo mode).

[0131] In another embodiment, energy differences between channels can be compensated at the encoder. The downmixing of the first stereo mode can then be determined by the following formula:

[0132]

[0133] scaling factor Then it can be determined as:

[0134]

[0135] In one embodiment, scaling factor In frequency subband It was determined in the middle.

[0136] In another embodiment, based on frequency sub-band The obtained spatial coding parameters are used to determine the scaling factor across the entire frequency band (not a frequency sub-band). In this case, the average scaling factor It can be determined, for example, as the arithmetic mean:

[0137]

[0138] in, It is for each frequency sub-band It is determined, as described above in equations (19) or (22) with subband-related parameters.

[0139] Comfort noise generation

[0140] Once the first set of background noise parameters Adapted to second stereo mode, transformed into The codec operating in the second stereo mode generates comfort noise. For smooth transitions, the parameters of CN are determined using two background noise estimates. and The weighted sum.

[0141] At the beginning of the transition segment, a larger weight is applied to the first background noise estimate (based on the estimate from the preceding active segment), and at the end of the transition segment, a larger weight is applied to the second background noise estimate (based on the received SID parameter). This smooth shift in weights between the first and second background noise estimates achieves a smooth transition between the active and inactive segments.

[0142] The length of the transition segment can be fixed or adaptively variable.

[0143] Comfort noise parameters It can be determined as:

[0144]

[0145] in:

[0146] It is the background noise parameter based on the transformation of the minimum statistics of the first stereo mode encoding;

[0147] These are the comfort noise parameters of the SID frame based on the second stereo mode encoding;

[0148] It is a counter for the number of inactive frames; and

[0149] It is the length of the crossfade.

[0150] when When increased, the conversion between the background noise level in the active encoding and the background noise level of the CN generated using the CNG parameters takes longer. In this case, [the following is obtained] and A linear cross-fade can be used between the elements, but other transformation functions with similar effects can be used. The length of the cross-fade can be fixed or adaptive, depending on the background noise parameter.

[0151] In one embodiment, the length of the adaptive crossfade is... It was identified as:

[0152]

[0153] in, This is the maximum number of frames that can have crossfade applied, for example, it can be set to 50 frames.

[0154]

[0155] It is the energy ratio of the estimated background noise level, for example and frequency subband The sum of energy.

[0156] In another embodiment, and The cross fade-out between them is obtained as

[0157]

[0158] in

[0159]

[0160] in, It is a frequency subband index, and It can be adaptive or fixed, for example... In one embodiment, frequency subband b can correspond to several frequency coefficients. This makes for frequency subbands frequency range , .

[0161] Based on the obtained comfort noise parameters The stereo channels can be synthesized in stereo mode 2 according to equation (6), that is...

[0162]

[0163] in, and Based on the obtained comfort noise parameters Uncorrelated random noise signals undergo spectral shaping. For example, uncorrelated noise signals can be generated in the frequency domain as follows:

[0164]

[0165] in, It is a pseudo-random generator that generates unit variance noise sequences by targeting frequency subbands. frequency range Obtained comfort noise parameters Scaling is performed. Figure 11 CNG upmixing is shown as a geometric representation of a multidimensional vector (e.g., a frame as an audio sample) according to equation (29). This is achieved by using vectors with the correct length (energy) and correlation (angle). The vector composition of ) is Figure 10 Encoder input channel and This yields CNG with correct inter-channel level differences and coherence. As mentioned earlier, CNG upmixing can also include control over inter-channel time and / or phase differences, or a similar representation of CN generation that is even more accurate relative to the spatial characteristics of the input channels.

[0166] In addition, control and Should the conversion between them be performed, or is CNG best based solely on... (as well as , It might be useful. If Only for If an estimate is made, then if Significant signal cancellation exists (e.g., for anticorrelated or inverted input stereo channels), which may cause inaccuracy.

[0167] In one embodiment, the decision to perform a crossfade between two background noise estimates is based on the energy relationship between the primary and secondary channels, which can be expressed in the time domain as:

[0168]

[0169] in

[0170]

[0171] Threshold SP thr The good value is already 2.0, but other values ​​are also acceptable. E P and E S It is given by the following formula:

[0172]

[0173] Low-pass filter coefficients and Should be within the range Inside. In one embodiment, and .

[0174] Figure 9 This illustrates the improved conversion from active coding in the first stereo mode to CNG in the second stereo mode. (Compared to...) Figure 4 Compared to the conversion shown, it can be seen that the conversion to CNG is smoother, which results in fewer audible conversions and increased perceptual performance for stereo codecs that utilize DTX to improve transmission efficiency.

[0175] Figure 8 This is a flowchart of process 800 according to an embodiment. The process begins with input voice / audio at box 802. Next, at box 802, VAD (or SAD) detects whether there is an active or inactive segment.

[0176] If it is an active segment, stereo encoding mode 1 is performed at box 806, followed by stereo decoding mode 1 at box 808. Next, background estimation 810 is performed at box 810, then buffered at box 812 for transforming the background estimation (from mode 1 to mode 2) at box 814, comfortable noise generation at box 816, and comfortable noise output at box 818.

[0177] If it is an inactive segment, background estimation is performed at box 820, followed by stereo coding mode 2 (SID) at box 822, and stereo decoding mode 2 at box 824. The output of stereo decoding mode 2 can be used at boxes 810 (background estimation) and 816 (CN generation). Typically, the transformation of the buffered background estimation parameters is triggered in the inactive segment, followed by the generation of comfortable noise at box 816, and the output of comfortable noise at box 818.

[0178] Figure 12 A flowchart according to an embodiment is shown. Process 1200 is a method performed by a node (e.g., a decoder). Process 1200 may begin at step s1202.

[0179] Step s1202 includes: providing a first set N1 of background noise parameters for at least one audio signal in a first spatial audio coding mode, wherein the first spatial audio coding mode is used for an active segment.

[0180] Step s1204 includes: providing a second set N2 of background noise parameters for at least one audio signal in a second spatial audio coding mode, wherein the second spatial audio coding mode is used for inactive segments.

[0181] Step s1206 includes: adapting a first set N1 of background noise parameters to a second spatial audio coding mode, thereby providing a first set of adapted background noise parameters. .

[0182] Step s1208 includes: combining a first set of appropriately matched background noise parameters within the conversion cycle. The second set N2 of background noise parameters is used to generate the comfort noise parameters.

[0183] Step s1210 includes: generating comfort noise for at least one output audio channel based on comfort noise parameters.

[0184] In some embodiments, generating comfort noise for at least one output audio channel includes applying the generated comfort noise parameters to at least one intermediate audio signal. In some embodiments, generating comfort noise for at least one output audio channel includes upmixing the at least one intermediate audio signal. In some embodiments, the at least one audio signal is based on signals from at least two input audio channels, and wherein a first set N1 of background noise parameters and a second set N2 of background noise parameters are each based on a single audio signal, wherein the single audio signal is based on downmixing the signals from at least two input audio channels. In some embodiments, the at least one output audio channel includes at least two output audio channels. In some embodiments, providing the first set N1 of background noise parameters includes receiving the first set N1 of background noise parameters from a node. In some embodiments, providing the second set N2 of background noise parameters includes receiving the second set N2 of background noise parameters from a node.

[0185] In some embodiments, adapting a first set N1 of background noise parameters to a second spatial audio coding mode includes applying a transform function. In some embodiments, the transform function includes a function of N1, NS1, and NS2, wherein NS1 includes a first set of spatial coding parameters indicating the undermixing and / or spatial characteristics of the background noise in the first spatial audio coding mode, and NS2 includes a second set of spatial coding parameters indicating the undermixing and / or spatial characteristics of the background noise in the second spatial audio coding mode. In some embodiments, applying the transform function includes calculating... ,in, It is a scalar compensation factor.

[0186] In some embodiments, It has the following values:

[0187]

[0188] Where, ratio LR It's a mixed bag. The corresponding coherence or correlation coefficient, and c, are given by the following formula.

[0189]

[0190] in, and These are gain parameters, and L and R are the left and right channel inputs, respectively.

[0191] In some embodiments, the conversion period is a fixed-length inactive frame. In some embodiments, the conversion period is a variable-length inactive frame. In some embodiments, a first set of appropriately matched background noise parameters is combined within the conversion period. The second set of background noise parameters N2 is used to generate comfortable noise, including: application The weighted average of N2.

[0192] In some embodiments, by combining a first set of appropriately matched background noise parameters within the conversion cycle The second set of background noise parameters, N2, generates comfort noise parameters, including calculations such as:

[0193]

[0194] Where CN is the generated comfort noise. It is the current inactive frame count, and k is the length of the transition cycle, which indicates the application. The number of inactive frames, weighted averaged with N2. 17. In some embodiments, a first set of appropriately matched background noise parameters is combined within the conversion period. The calculation of the second set N2 of background noise parameters to generate comfort noise parameters includes:

[0195]

[0196] in

[0197]

[0198] Where CN is the generated comfort noise parameter. It is the current inactive frame count, and k is the indicator for the application. The length of the conversion period for the weighted average number of inactive frames of N2, and It is a frequency subband index. In some embodiments, generating comfort noise parameters includes targeting the frequency subband. at least one frequency coefficient calculate:

[0199]

[0200] In some embodiments, k is determined as:

[0201]

[0202] Where M is the maximum value of k, and r1 is the energy ratio of the estimated background noise level, determined as follows:

[0203]

[0204] in, There are N frequency sub-bands. This refers to the case of a given subband b. The adapted background noise parameters, and This refers to the background noise parameter adapted for N2 of a given subband b.

[0205] In some embodiments, by combining a first set of appropriately matched background noise parameters within the conversion cycle The second set N2 of background noise parameters is used to generate comfort noise parameters, including: application A nonlinear combination of N2. In some embodiments, the method further includes: determining a first set of appropriately matched background noise parameters by combining them within a conversion period. Comfort noise parameters are generated by combining a second set N2 of background noise parameters, wherein a first set of appropriately matched background noise parameters is combined within the conversion period. The second set N2 of background noise parameters is used to generate the comfort noise parameters as a first set of background noise parameters that are appropriately matched within the conversion period. The result of generating the comfort noise parameters is performed by combining the second set N2 of background noise parameters.

[0206] In some embodiments, a first set of appropriately matched background noise parameters is determined by combining them within the conversion cycle. The second set of background noise parameters N2 is used to generate comfort noise parameters based on the evaluation of the first energy of the primary channel and the second energy of the secondary channel. In some embodiments, the first set of background noise parameters N1, the second set of background noise parameters N2, and the first set of adapted background noise parameters are used. One or more of the parameters include one or more parameters describing signal characteristics and / or spatial characteristics, and the one or more parameters include one or more of the following: (i) linear prediction coefficients representing signal energy and spectral shape; (ii) excitation energy; (iii) inter-channel coherence; (iv) inter-channel level difference; and (v) side gain parameters.

[0207] Figure 8 This is a block diagram of an apparatus according to an embodiment. As shown, node 1300 (e.g., a decoder) may include a providing unit 1302, an adapting unit 1304, a generating unit 1306, and an application unit 1308.

[0208] The providing unit 1302 is configured to provide a first set N1 of background noise parameters for at least one audio signal in a first spatial audio coding mode, wherein the first spatial audio coding mode is used for an active segment.

[0209] The providing unit 1302 is further configured to provide a second set N2 of background noise parameters for at least one audio signal in a second spatial audio coding mode, wherein the second spatial audio coding mode is used for inactive segments.

[0210] The adaptation unit 1304 is configured to adapt a first set N1 of background noise parameters to a second spatial audio coding mode, thereby providing an adapted first set of background noise parameters. .

[0211] The generation unit 1306 is configured to: combine a first set of appropriately matched background noise parameters within the conversion cycle. The second set N2 of background noise parameters is used to generate the comfort noise parameters.

[0212] The application unit 1308 is configured to apply the generated comfort noise to at least one output audio channel.

[0213] Figure 14 This is a block diagram of an apparatus 1300 (e.g., a node (e.g., a decoder)) according to some embodiments. Figure 14As shown, the device may include: a processing circuitry (PC) 1402, which may include one or more processors (P) 1455 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.); a network interface 948, including a transmitter (Tx) 1445 and a receiver (Rx) 1447, for enabling the device to send data to and receive data from other nodes connected to a network 1410 (e.g., an Internet Protocol (IP) network), the network interface 1448 being connected to the network 1410; and a local storage unit (also referred to as a “data storage system”) 1408, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where the PC 1402 includes a programmable processor, a computer program product (CPP) 1441 may be provided. CPP 1441 includes a computer-readable medium (CRM) 1442 that stores a computer program (CP) 1443 including computer-readable instructions (CRI) 1444. CRM 1442 may be a non-transitory computer-readable medium, such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., random access memory, flash memory), etc. In some embodiments, the CRI 1444 of the computer program 1443 is configured such that, when executed by PC 1402, the CRI causes the device to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, the device may be configured to perform the steps described herein without requiring code. That is, for example, PC 1402 may consist only of one or more ASICs. Therefore, the features of the embodiments described herein can be implemented in hardware and / or software.

[0214] Brief description of various embodiments

[0215] A1. A method for generating comfort noise, comprising:

[0216] A first set N1 of background noise parameters is provided for at least one audio signal in a first spatial audio coding mode, wherein the first spatial audio coding mode is used for an active segment;

[0217] A second set N2 of background noise parameters is provided for at least one audio signal in a second spatial audio coding mode, wherein the second spatial audio coding mode is used for inactive segments;

[0218] This adapts a first set N1 of background noise parameters to a second spatial audio coding mode, thereby providing a first set of adapted background noise parameters. ;

[0219] By combining a first set of appropriately matched background noise parameters within the conversion cycle The second set N2 of background noise parameters is used to generate the comfort noise parameters; and

[0220] Comfort noise is generated for at least one output audio channel based on comfort noise parameters.

[0221] A1a. The method according to embodiment A1, wherein generating comfort noise for at least one output audio channel includes: applying the generated comfort noise parameters to at least one intermediate audio signal.

[0222] A1b. The method according to embodiment A1a, wherein generating comfort noise for at least one output audio channel includes upmixing at least one intermediate audio signal.

[0223] A2. The method according to any one of embodiments A1, A1a and A1b, wherein at least one audio signal is based on signals from at least two input audio channels, and wherein a first set N1 of background noise parameters and a second set N2 of background noise parameters are each based on a single audio signal, wherein the single audio signal is based on downmixing signals from at least two input audio channels.

[0224] A3. The method according to any one of embodiments A1 to A2, wherein at least one output audio channel includes at least two output audio channels.

[0225] A4. The method according to any one of embodiments A1 to A3, wherein providing a first set N1 of background noise parameters includes: receiving a first set N1 of background noise parameters from a node.

[0226] A5. The method according to any one of embodiments A1 to A4, wherein providing the second set N2 of background noise parameters includes: receiving the second set N2 of background noise parameters from the node.

[0227] A6. The method according to embodiment A1, wherein adapting the first set N1 of background noise parameters to the second spatial audio coding mode includes applying a transformation function.

[0228] A7. The method according to embodiment A6, wherein the transformation function includes functions of N1, NS1 and NS2, wherein NS1 includes a first set of spatial coding parameters indicating the undermixing / or spatial characteristics of background noise in a first spatial audio coding mode, and NS2 includes a second set of spatial coding parameters indicating the undermixing / or spatial characteristics of background noise in a second spatial audio coding mode.

[0229] A8. The method according to any one of embodiments A6 to A7, wherein applying the transformation function includes calculating ,in, It is a scalar compensation factor.

[0230] A9. The method according to embodiment A8, wherein, It has the following values:

[0231]

[0232] Where, ratio LR It's a mixed bag. Corresponding to the coherence or correlation coefficient, c is given by the following formula.

[0233]

[0234] in, and It is the gain parameter.

[0235] A9a. The method according to embodiment A8, wherein, It has the following values:

[0236]

[0237] Where, ratio LR It's a mixed bag. Corresponding to the coherence or correlation coefficient, c is given by the following formula.

[0238]

[0239] in, , and It is the gain parameter.

[0240] A10. The method according to any one of embodiments A1 to A9a, wherein the conversion period is a fixed-length inactive frame.

[0241] A11. The method according to any one of embodiments A1 to A9a, wherein the conversion period is a variable-length inactive frame.

[0242] A12. The method according to any one of embodiments A1 to A11, wherein a first set of appropriately matched background noise parameters is assembled during the conversion period. The second set of background noise parameters N2 is used to generate comfortable noise, including: application The weighted average of N2.

[0243] A13. The method according to any one of embodiments A1 to A12, wherein a first set of appropriately matched background noise parameters is assembled during the conversion period. The calculation of the second set N2 of background noise parameters to generate comfort noise parameters includes:

[0244]

[0245] Where CN is the generated comfort noise parameter. It is the current inactive frame count, and k is the indicator for the application. The length of the conversion period is the weighted average number of inactive frames of N2.

[0246] A13a. The method according to any one of embodiments A1 to A12, wherein a first set of appropriately matched background noise parameters is assembled during the conversion period. The calculation of the second set N2 of background noise parameters to generate comfort noise parameters includes:

[0247]

[0248] in

[0249]

[0250] Where CN is the generated comfort noise parameter. It is the current inactive frame count, and k is the indicator for the application. The length of the conversion period for the weighted average number of inactive frames of N2, and It is a frequency subband index.

[0251] A13b. The method according to embodiment A13a, wherein generating comfort noise parameters includes targeting frequency sub-bands. at least one frequency coefficient calculate:

[0252]

[0253] A14. The method according to any one of embodiments A13, A13a, and A13b, wherein k is determined as

[0254]

[0255] Where M is the maximum value of k, and r1 is the energy ratio of the estimated background noise level, determined as follows:

[0256]

[0257] in, There are N frequency sub-bands. This refers to the case of a given subband b. The adapted background noise parameters, and This refers to the background noise parameter adapted for N2 of a given subband b.

[0258] A15. The method according to any one of embodiments A1 to A11, wherein a first set of appropriately matched background noise parameters is assembled during the conversion period. The second set N2 of background noise parameters is used to generate comfort noise parameters, including: application Nonlinear combination of N2 and N2.

[0259] A16. The method according to any one of embodiments A1 to A15 further includes: determining a first set of appropriately matched background noise parameters by combining them within a conversion period. The comfort noise parameters are generated by combining a second set N2 of background noise parameters, wherein the first set of background noise parameters is appropriately matched within the conversion period. The second set N2 of background noise parameters is used to generate the comfort noise parameters as a first set of background noise parameters that are appropriately matched within the conversion period. The process is performed to generate the comfort noise parameters by combining the second set N2 of background noise parameters.

[0260] A17. The method according to embodiment A16, wherein a first set of appropriately matched background noise parameters is determined by combining them within the conversion cycle. The comfort noise parameters are generated by using a second set of background noise parameters, N2, based on the evaluation of the first energy of the primary channel and the second energy of the secondary channel.

[0261] A18. The method according to any one of embodiments A1 to A17, wherein a first set N1 of background noise parameters, a second set N2 of background noise parameters, and a first set of adapted background noise parameters are defined. One or more of the parameters include one or more parameters describing signal characteristics and / or spatial characteristics, and the one or more parameters include one or more of the following: (i) linear prediction coefficients representing signal energy and spectral shape; (ii) excitation energy; (iii) inter-channel coherence; (iv) inter-channel level difference; and (v) side gain parameters.

[0262] B1. A node including processing circuitry and a memory containing instructions executable by the processing circuitry, wherein the processing circuitry is operable to:

[0263] A first set N1 of background noise parameters is provided for at least one audio signal in a first spatial audio coding mode, wherein the first spatial audio coding mode is used for an active segment;

[0264] A second set N2 of background noise parameters is provided for at least one audio signal in a second spatial audio coding mode, wherein the second spatial audio coding mode is used for inactive segments;

[0265] This adapts a first set N1 of background noise parameters to a second spatial audio coding mode, thereby providing a first set of adapted background noise parameters. ;

[0266] By combining a first set of appropriately matched background noise parameters within the conversion cycle The second set N2 of background noise parameters is used to generate the comfort noise parameters; and

[0267] Comfort noise is generated for at least one output audio channel based on comfort noise parameters.

[0268] B1a. The node according to embodiment B1, wherein generating comfort noise for at least one output audio channel includes: applying the generated comfort noise parameters to at least one intermediate audio signal.

[0269] B1b. The node according to embodiment B1a, wherein generating comfort noise for at least one output audio channel includes upmixing at least one intermediate audio signal.

[0270] B2. A node according to any one of embodiments B1, B1a and B1b, wherein at least one audio signal is based on signals from at least two input audio channels, and wherein a first set N1 of background noise parameters and a second set N2 of background noise parameters are each based on a single audio signal, wherein the single audio signal is based on downmixing signals from at least two input audio channels.

[0271] B3. The node according to any one of embodiments B1 to B2, wherein at least one output audio channel includes at least two output audio channels.

[0272] B4. The node according to any one of embodiments B1 to B3, wherein providing a first set N1 of background noise parameters includes: receiving a first set N1 of background noise parameters from the node.

[0273] B5. The node according to any one of embodiments B1 to B4, wherein the second set N2 providing background noise parameters includes: receiving a second set N2 of background noise parameters from another node.

[0274] B5a. The node according to embodiment B5, wherein another node includes an encoder.

[0275] B6. The node according to embodiment B1, wherein adapting the first set N1 of background noise parameters to the second spatial audio coding mode includes applying a transformation function.

[0276] B7. The node according to embodiment B6, wherein the transformation function includes a function of N1, NS1 and NS2, wherein NS1 includes a first set of spatial coding parameters indicating the undermixing / or spatial characteristics of background noise in a first spatial audio coding mode, and NS2 includes a second set of spatial coding parameters indicating the undermixing / or spatial characteristics of background noise in a second spatial audio coding mode.

[0277] B8. The node according to any one of embodiments B6 to B7, wherein applying the transformation function includes calculating ,in, It is a scalar compensation factor.

[0278] B9. The node according to embodiment B8, wherein, It has the following values:

[0279]

[0280] Where, ratio LR It's a mixed bag. Corresponding to the coherence or correlation coefficient, c is given by the following formula.

[0281]

[0282] in, and It is the gain parameter.

[0283] B9a. The node according to embodiment B8, wherein, It has the following values:

[0284]

[0285] Where, ratio LR It's a mixed bag. Corresponding to the coherence or correlation coefficient, c is given by the following formula.

[0286]

[0287] in, , and It is the gain parameter.

[0288] B10. A node according to any one of embodiments B1 to B9a, wherein the conversion period is a fixed-length inactive frame.

[0289] B11. A node according to any one of embodiments B1 to B9a, wherein the conversion period is a variable-length inactive frame.

[0290] B12. The node according to any one of embodiments B1 to B11, wherein a first set of appropriately matched background noise parameters is assembled during the conversion period. The second set N2 of background noise parameters is used to generate comfort noise parameters, including: application The weighted average of N2.

[0291] B13. The node according to any one of embodiments B1 to B12, wherein a first set of appropriately matched background noise parameters is assembled during the conversion period. The calculation of the second set N2 of background noise parameters to generate comfort noise parameters includes:

[0292]

[0293] Where CN is the generated comfort noise parameter. It is the current inactive frame count, and k is the indicator for the application. The length of the conversion period is the weighted average number of inactive frames of N2.

[0294] B13a. The node according to any of embodiments B1 to B12, wherein a first set of appropriately matched background noise parameters is assembled during the conversion period. The calculation of the second set N2 of background noise parameters to generate comfort noise parameters includes:

[0295]

[0296] in

[0297]

[0298] Where CN is the generated comfort noise parameter. It is the current inactive frame count, and k is the indicator for the application. The length of the conversion period for the weighted average number of inactive frames of N2, and It is a frequency subband index.

[0299] B13b. The node according to embodiment B13a, wherein generating comfort noise parameters includes targeting frequency sub-bands. at least one frequency coefficient calculate:

[0300]

[0301] B14. A node according to any one of embodiments B13, B13a, and B13b, wherein k is determined as:

[0302]

[0303] Where M is the maximum value of k, and r1 is the energy ratio of the estimated background noise level, determined as follows:

[0304]

[0305] in, There are N frequency sub-bands. This refers to the case of a given subband b. The adapted background noise parameters, and This refers to the background noise parameter adapted for N2 of a given subband b.

[0306] B15. The node according to any one of embodiments B1 to B11, wherein a first set of appropriately matched background noise parameters is assembled during the conversion period. The second set of background noise parameters N2 is used to generate comfortable noise, including: application Nonlinear combination of N2 and N2.

[0307] B16. The node according to any one of embodiments B1 to B15 further includes: determining a first set of appropriately matched background noise parameters by combining them within the conversion period. The comfort noise parameters are generated by combining a second set N2 of background noise parameters, wherein the first set of background noise parameters is appropriately matched within the conversion period. The second set N2 of background noise parameters is used to generate the comfort noise parameters as a first set of background noise parameters that are appropriately matched within the conversion period. The process is performed to generate the comfort noise parameters by combining the second set N2 of background noise parameters.

[0308] B17. The node according to embodiment B16, wherein a first set of appropriately matched background noise parameters is determined by combining them within the conversion cycle. The comfort noise parameters are generated by using a second set of background noise parameters, N2, based on the evaluation of the first energy of the primary channel and the second energy of the secondary channel.

[0309] B18. A node according to any one of embodiments B1 to B17, wherein a first set N1 of background noise parameters, a second set N2 of background noise parameters, and a first set of adapted background noise parameters. One or more of the parameters include one or more parameters describing signal characteristics and / or spatial characteristics, and the one or more parameters include one or more of the following: (i) linear prediction coefficients representing signal energy and spectral shape; (ii) excitation energy; (iii) inter-channel coherence; (iv) inter-channel level difference; and (v) side gain parameters.

[0310] B19. A node according to any one of embodiments B1 to B18, wherein the node includes a decoder.

[0311] B20. A node according to any one of embodiments B1 to B18, wherein the node includes an encoder.

[0312] C1. A computer program including instructions that, when executed by processing circuitry, cause the processing circuitry to perform the method described according to any one of embodiments A1 to A18.

[0313] C2. A carrier comprising a computer program according to embodiment C1, wherein the carrier is one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium.

[0314] Although various embodiments of this disclosure have been described herein, it should be understood that they are presented by way of example only and not by way of limitation. Therefore, the breadth and scope of this disclosure should not be limited to any of the exemplary embodiments described above. Furthermore, any combination of the foregoing elements with all possible variations is included in this disclosure unless otherwise indicated or otherwise explicitly conflicted by the context.

[0315] Additionally, although the process described above and shown in the accompanying figures is presented as a series of steps, it is for illustrative purposes only. Therefore, it is conceivable that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel.

Claims

1. A method for determining parameters of comfort noise generation CNG in an audio decoder, the method comprising: Two sets of background noise parameters (s1202, s1204) are obtained. The first set N1 of background noise parameters is determined based on the background noise estimation performed by the audio decoder when operating in the first stereo coding mode for the active segment, and The second set N2 of background noise parameters is determined based on parameters received in the silence descriptor SID frame of the second stereo coding mode used for inactive segments; The first set N1 of background noise parameters is adapted to the second stereo coding mode (s1206), thereby providing a set of adapted background noise parameters. ; By combining the set of adapted background noise parameters within the conversion cycle and the second set N2 of the background noise parameters as The (s1208) comfort noise parameter is generated by weighted summation of N2 and N2; and (s1210) Comfort noise is generated based on the aforementioned comfort noise parameters.

2. The method according to claim 1, wherein, The second set N2 of background noise parameters is determined based on the estimated background noise characteristics of the encoded audio signal at the encoder operating in the encoding mode used for SID encoding.

3. The method according to claim 1 or 2, wherein, Adapting the first set N1 of background noise parameters to the second coding mode includes calculating... ,in, It is a compensation factor.

4. The method according to any one of claims 1 to 3, wherein, The conversion period is a variable-length inactive frame.

5. The method according to any one of claims 1 to 4, wherein, By combining the set of adapted background noise parameters within the conversion cycle Generating comfort noise parameters from a second set N2 of background noise parameters includes calculations such as: in, It is the generated comfort noise parameter. is the number of inactive frames, and k is the length of the conversion period.

6. An audio decoder (1300) configured to: Two sets of background noise parameters are obtained. The first set N1 of the background noise parameters is determined based on the background noise estimation performed by the audio decoder when operating in the first stereo coding mode for the active segment, and the second set N2 of the background noise parameters is determined based on the parameters received in the silence descriptor SID frame in the second stereo coding mode for the inactive segment. The first set N1 of background noise parameters is adapted to the second stereo coding mode, thereby providing a first set of adapted background noise parameters. ; By combining the set of adapted background noise parameters within the conversion cycle and the second set N2 of the background noise parameters as The comfort noise parameter is generated by weighting N2 and N2; and Comfort noise is generated based on the aforementioned comfort noise parameters.

7. The audio decoder according to claim 6, wherein, The second set N2 of background noise parameters is determined based on the estimated background noise characteristics of the encoded audio signal at the encoder operating in the encoding mode used for SID encoding.

8. The audio decoder according to claim 6 or 7, wherein, Adapting the first set N1 of background noise parameters to the second coding mode includes calculating... ,in, It is a compensation factor.

9. The audio decoder according to any one of claims 6 to 8, wherein, The conversion period is a variable-length inactive frame.

10. The audio decoder according to any one of claims 6 to 9, wherein, By combining the set of adapted background noise parameters within the conversion cycle Generating comfort noise parameters from a second set N2 of background noise parameters includes calculations such as: in, It is the generated comfort noise parameter. is the number of inactive frames, and k is the length of the conversion period.