Coherence calculation for stereo intermittent transmission (DTX)

By initializing the cross-spectral filter state with previous inactive period coherence, the method enhances coherence estimation in DTX systems, ensuring natural comfort noise generation with reduced memory requirements.

JP2025532351APending Publication Date: 2025-09-29TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025519573
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-05
Filing Date
2023-09-20
Publication Date
2025-09-29

AI Technical Summary

Technical Problem

Existing DTX systems face challenges in accurately estimating coherence during speech pauses due to the influence of speech segments, leading to incorrect generation of comfort noise, and require significant memory for storing previous frames.

Method used

The proposed method initializes the cross-spectral low-pass filter state during inactive segments using coherence parameters from the previous inactive period, minimizing the impact of speech portions and reducing memory requirements.

Benefits of technology

This approach maintains accurate coherence estimation and natural-sounding comfort noise by avoiding abrupt spatial characteristic changes, while minimizing memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532351000001_ABST
    Figure 2025532351000001_ABST
Patent Text Reader

Abstract

Enabling generation of comfort noise in an encoder using estimated coherence parameters in a network using discontinuous transmission (DTX) includes receiving a time-domain audio input including an audio input signal; processing the input signal frame-by-frame, encoding active content of each input signal at a first bit rate until an inactive period is detected in the input signal; switching encoding from active to inactive encoding to encode background noise at a second bit rate during the inactive periods; estimating coherence parameters during the inactive periods based on cross-spectral low-pass filtering or averaging, including re-initializing the cross-spectral low-pass filter states based on the coherence parameters from a previous inactive period; and processing the input signal frame-by-frame by encoding the estimated coherence parameters; and initiating transmission of the encoded active content, background noise, and coherence parameters to a decoder.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to communications, and more particularly to communication methods and related devices and nodes that support encoding and decoding. [Background technology]

[0002] In a communication network, there can be challenges to obtain good performance and capacity for a given communication protocol, its parameters, and the physical environment in which the communication network is located.

[0003] For example, despite the continuous increase in capacity of communication networks, limiting the resource usage required per user remains a concern. In mobile communication networks, lower resource usage required per call means that the mobile communication network can serve a larger number of users in parallel. Lower resource usage also results in lower power consumption in both user-side devices (such as terminal devices) and network-side devices (such as network nodes). This translates into energy and cost savings for network operators while enabling longer battery life and increased talk time experienced at terminal devices.

[0004] One mechanism for reducing the required resource usage for voice communication applications in mobile communication networks is to exploit natural pauses in speech. More specifically, in most conversations, only one party is active at a time, and therefore speech pauses in one communication direction typically account for more than half of the signal. One way to take advantage of this property to reduce the required resource usage is to employ a discontinuous transmission (DTX) system, in which active signal coding is suspended during speech pauses.

[0005] The encoding process operates on a segment (or segments) of the audio signal, called a frame, where input audio samples during a certain time interval, typically 10-20 ms, are buffered by the encoder and used to extract parameters to be transmitted to the decoder.

[0006] During speech pauses, it is common to transmit so-called SID (silence insertion descriptor) frames in a very low bitrate encoding of background noise to enable a Comfort Noise Generator (CNG) system at the receiving end to fill said pauses with background noise having similar characteristics to the original noise. CNG makes the sound more natural compared to having silence during speech pauses, since the background noise is maintained and does not switch on and off with the speech. Complete silence during speech pauses is usually perceived as unpleasant and often leads to the mistaken belief that the call has been dropped.

[0007] The DTX system may further rely on a Voice Activity Detector (VAD) that instructs the transmitting device whether to use active signal coding or low-rate background noise coding. In this regard, the transmitting device may be configured to distinguish between other source types by using a (Generic) Sound Activity Detector (GSAD or SAD), which not only distinguishes speech from background noise but may also be configured to detect music or other signal types deemed relevant. A block diagram of a DTX system 100 is shown in FIG. 1.

[0008] In Figure 1, input audio is received by a VAD 102, a speech / audio coder 104, and a CNG coder 106. The VAD 102 indicates whether a "high" bit rate should be sent from the speech / audio coder 104 or a "low" bit rate should be sent from the CNG coder 106.

[0009] Communication services can be further enhanced by supporting stereo or multi-channel audio transmission. In these cases, the DTX / CNG system can also take into account the spatial characteristics of the signal to provide pleasant-to-hear comfort noise.

[0010] A common mechanism for generating comfort noise is to transmit information about the energy and spectral shape of the background noise during speech pauses. This can be done using a significantly lower number of bits than the normal coding of speech segments. Typically, this information is sent less frequently than in active segments, as shown in Figure 2, where the active segments are denoted as active coding and the information about the energy and spectral shape of the background noise during speech pauses is denoted as CN coding.

[0011] A common feature in DTX systems is the addition of a so-called "hangover period" to the VAD decision, as shown in Figure 3. During this period, active coding will still be used even if the VAD decision is that there should not be any active coding. This is to avoid short segments of CNG in the middle of longer active segments, for example, during breathing pauses in speech utterances. The parameters used for CNG generation can be estimated during this period.

[0012] At the receiving end, comfort noise is generated by creating a pseudorandom signal and then shaping the spectrum of the signal with a filter based on information received from the transmitting device. The signal generation and spectral shaping can be performed in the time or frequency domain.

[0013] For stereo operation, additional parameters are transmitted to the receiver. In typical stereo signals, channel pairs exhibit a high degree of similarity, or correlation. State-of-the-art stereo coding schemes exploit this correlation by employing parametric coding, in which a single channel is encoded with high quality and complemented with a parametric description that enables the reconstruction of a full stereo image. The process of reducing a channel pair to a single channel is often called downmixing, and the resulting channel is the downmix or mixdown channel. Downmix procedures generally attempt to preserve energy by aligning the inter-channel time difference (ITD) and inter-channel phase difference (IPD) before mixing the channels. To maintain energy balance in the input signal, the inter-channel level difference (ILD) is also measured. The ITD, IPD, and ILD can then be coded and used in the inverse upmix procedure when reconstructing the stereo channel pair at the decoder. Figures 4 and 5 show block diagrams of a parametric stereo encoder 400 and decoder 500.

[0014] 4, a time-domain stereo input is received by a stereo processing and mixdown module 402. The stereo processing and mixdown module 402 processes the time-domain stereo input signal to produce a mono mixdown signal and stereo parameters. The mono mixdown signal is received by a mono audio / audio encoder 404, which processes the mono mixdown signal to produce an encoded mono signal. The encoded mono signal and stereo parameters are transmitted to a decoder, such as a parametric stereo decoder 500.

[0015] 5, the encoded mono signal is received by a mono audio / audio decoder 502, which decodes the encoded mono signal to produce a mono mixdown signal. The mono mixdown signal and stereo parameters are received by a stereo processing and upmix decoder 504, which processes the mono mixdown signal and stereo parameters to produce a time-domain stereo output. The time-domain stereo output can be stored or sent to an audio player for playback.

[0016] In addition to ITD, IPD, and ILD, the coherence between the left and right channels can be calculated at the encoder and transmitted to the receiver. Coherence basically represents how well the left and right signals are correlated at different frequencies.

[0017] For DTX operation and CNG, the parametric representation of spatial characteristics (stereo image in the case of stereo audio) is particularly relevant since it is a compact representation. The same or similar parameters as those used for parametric stereo coding modes for active frames can be transmitted in silence insertion descriptor (SID) frames for comfort noise generation at the decoder. However, for SID frames, larger quantization errors can be allowed without significant perceptual degradation, which means that for CNG, fewer bits can be used to represent spatial characteristics than for active coded frames.

[0018] If a coherence parameter is used to represent the properties of spatial audio for CNG, the coherence can be reconstructed at the decoder, and a comfort noise signal with properties similar to the original sound can be created. For further details, see U.S. Patent Application Publication No. 20170047072. Note that, in general, additional parameters (e.g., ILD, IPD, ITD parameters) will be needed to capture / represent all of the most perceptually relevant spatial characteristics and will be transmitted together with the coherence in the SID frame.

[0019] A solution for efficient representation of coherence is described in PCT Publication No. WO2019193173, where the coherence is calculated at the transmitter with high frequency resolution and then divided into a small number of frequency bands, and the coherence within each band is weighted together to one value per band. A vector containing the coherence per band is then encoded and transmitted to a decoder.

[0020] A stereo coder receives as input a channel pair [l(m,n) r(m,n)], where l(m,n) and r(m,n) denote the input signals for the left and right channels, respectively, for sample index n of frame m. The audio is sampled at a sampling frequency F s The audio is processed in frames of length N samples in , where the length of the frame may include overlap (look-ahead samples and samples from past memory). Typically, 20 ms of new audio samples are buffered and included in the frame being coded.

[0021] Coding parameters such as ITD are estimated at the encoding side for each frame and transmitted to the decoder. It is also common not to transmit a parameter if there is no clear gain in the encoding process that involves using it. In the case of ITD, this will be when the left and right signals are nearly uncorrelated.

[0022] The input signal is transformed into the frequency domain, for example by a DFT (Discrete Fourier Transform) or any other suitable filter bank or transform, such as a QMF, a hybrid QMF (Quadrature Mirror Filter) or an MDCT (Modified Discrete Cosine Transform). When a DFT or an MDCT is used, the input signal is generally windowed before the transformation. The choice of window depends on various parameters, such as time and frequency resolution characteristics, algorithm delay (overlap length), reconstruction properties, etc. In the case of a DFT, the spectra of the left and right audio channels are [l win (m,n)r win (m,n)]=[l(m,n)win(n) r(m,n)win(n)], n=0,1,2,...,N-1 TIFF2025532351000002.tif13170, where win(n) is the selected window function.

[0023] A general specification of channel coherence C for frequency f gen (f) is TIFF2025532351000003.tif10170, where S x (f) and S y (f) represents the frequency spectrum of two channels x and y, and S xy (f) is the cross spectrum. Working in the DFT domain, the coherence is TIFF2025532351000004.tif9170Xspec(k,m)=SPD_L(m,k) * SPD_R(m,k) can be estimated based on the cross and power spectra according to Here, * denotes the complex conjugate.

[0024] However, this relies on good estimates of the cross- and power spectra, which may be obtained, for example, using the well-known Welch method. Another way to stabilize the coherence estimate is to low-pass filter the short-time spectra Xspec(k,m), SPD_L(m,k) and SPD_R(m,k) with a first-order low-pass filter before being used in the coherence calculation, as shown in the following equations: Xspec smooth [k,m]=(1-α)·Xspec smooth [k,m-1]+α·Xspec(k,m) SPD_L smooth [k,m]=(1-α)·SPD_L smooth [k,m-1]+α·SPD_L(m,k) SPD_R smooth [k,m]=(1-α)·SPD_R smooth [k,m-1]+α·SPD_R(m,k)

[0025] In that case, the coherence is It can be obtained as TIFF2025532351000005.tif9170.

[0026] To obtain a good, stable coherence estimate, a somewhat small value of α is required. Figure 7 shows how to estimate the coherence of two signals with a fixed coherence of 0.2 across all frequencies: Two examples are shown in TIFF2025532351000006.tif8170. It can be seen that in the early frames, where the cross- and power spectrum smoothing filters only contain information from the current frame, the coherence will be 1 for all frequencies in both cases. However, for later frames, there is a significant difference. The coherence estimate gradually approaches the true coherence value in case (b), while in case (a), there is a significant amount of noise in the estimate. In Figure 8, it can be seen that, averaging across frequency, the coherence estimate in case (b) does indeed approach the true coherence (0.2 in this case), while in case (a), there is a clear bias in the coherence estimate.

[0027] In DTX solutions where coherence is only used in generating comfort noise, low-pass filter updates can be skipped during speech segments, i.e., when the VAD indicates speech or active content, because otherwise the speech signal would reside in the low-pass filter state for some time after the speech stops and a new comfort noise segment starts, which would cause a bias in the coherence estimate for the background signal.

[0028] To reduce the number of bits for encoding the coherence values, the spectrum is scaled by N, as shown in Figure 6 and in the following equation: band The signal is divided into bands. TIFF2025532351000007.tif10170 where bandlimits(b) is a vector containing the limits between frequency bands.

[0029] The width of these bands is intended to mimic the frequency resolution of human hearing, with narrow bands for low frequencies and increasing bandwidth for higher frequencies.

[0030] Instead of using the average of the coherence within one band, a weighted average can be used for each band, where the DFT energy spectrum |LR(m,k)| for a mono signal that is a downmix of the input signal, e.g., lr(m,n)=l(m,n)+r(m,n) 2 is used as the weighting function. Details can be found in PCT Publication No. WO2019193156. With the weighting function, the formula becomes: It can be written as TIFF2025532351000008.tif12170. Summary of the Invention

[0031] Currently, there exists one or more challenges. If the smoothed left and right power spectra and cross-spectrum used in the coherence calculation contain a portion in time where a speech signal is present, they may not reflect the characteristics of the background noise, leading to incorrect generation of comfort noise. One reason for this may be that the last frame before the speech segment contains the onset of the speech segment. Although the energy of this portion may be too low and / or other features of the audio signal may not be sufficient to trigger the VAD to detect speech, the portion may still have a negative impact on the background noise coherence estimation.

[0032] One solution to this problem is to store the left and right spectra and cross-spectrum for the previous frame and remove it from the low-pass filter state if the next frame is a speech frame. However, for high-resolution DFTs, this can mean that several kilowords of memory must be spent storing the previous values.

[0033] Some aspects of the present disclosure and their embodiments may provide solutions to these and other challenges, improving coherence estimation in speech pauses by minimizing the influence of speech portions. Various embodiments described herein determine the coherence for a small number of frequency bands and use that coherence in creating filter states for a low-pass filter that reflect the previous background noise.

[0034] According to some embodiments, a method in an encoder for enabling generation of comfort noise using estimated coherence parameters in a network using discontinuous transmission (DTX) includes receiving a time-domain audio input including audio input signals, the method including: processing the audio input signals frame by frame, encoding active content of each audio input signal at a first bit rate until an inactive period is detected in the audio input signals, switching encoding from the active encoded content to inactive encoding to encode background noise at a second bit rate during the inactive periods, estimating coherence parameters during the inactive periods based on cross-spectral low-pass filtering or cross-spectral averaging, where estimating the coherence parameters includes re-initializing cross-spectral low-pass filter states based on the coherence parameters from a previous inactive period, and encoding the estimated coherence parameters; and initiating transmission of the encoded active content, the encoded background noise, and the encoded coherence parameters to a decoder.

[0035] Similar encoders, computer programs, and computer program products are also provided.

[0036] Some embodiments may provide one or more of the following technical advantages: Various embodiments make comfort noise sound more natural and avoid the unpleasant effects of abrupt changes in spatial characteristics during CNG after changing from active coding. In particular, they avoid DTX starting with a segment of comfort noise characterized by active content and then, after a while, suddenly changing to comfort noise that more closely resembles the original input noise. Various embodiments can estimate coherence and minimize the influence of speech portions in the estimate.

[0037] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate several non-limiting embodiments of the inventive concepts. [Brief explanation of the drawings]

[0038] [Figure 1] FIG. 1 is a block diagram of a DTX system. [Figure 2] 1 is a flow diagram illustrating CNG parameter encoding and transmission. [Figure 3] 1 is a flow diagram illustrating a VAD (or DTX) hangover period. [Figure 4] FIG. 1 is a block diagram of a parametric stereo encoder. [Figure 5] FIG. 2 is a block diagram of a parametric stereo decoder according to some embodiments. [Figure 6] FIG. 1 is a diagram of coherence band divisions, according to some embodiments. [Figure 7] 1 is a diagram of a coherence estimate for a stereo signal according to some embodiments. [Figure 8] 10 is a graph showing the average over frequency of coherence estimates for a stereo signal with fixed filter coefficients; [Figure 9] 10 is a graph illustrating estimating coherence compared to resetting cross and power spectra according to various embodiments. [Figure 10] 10 is a flow diagram illustrating filter state initialization according to some embodiments. [Figure 11] 1 is a flow diagram illustrating coherence memory shifting according to some embodiments. [Figure 12] FIG. 1 is a block diagram illustrating an operating environment for a parametric stereo encoder and a parametric stereo decoder according to some embodiments. [Figure 13] FIG. 2 is a block diagram of an encoder according to some embodiments. [Figure 14] FIG. 2 is a block diagram of a decoder according to some embodiments. [Figure 15] FIG. 2 is a block diagram of a host, according to some embodiments. [Figure 16] FIG. 1 is a block diagram of a virtualized environment, according to some embodiments. [Figure 17] 10 is a flowchart illustrating the operation of an encoder, according to some embodiments. [Figure 18] 10 is a flowchart illustrating the operation of an encoder, according to some embodiments. [Figure 19] 10 is a flowchart illustrating the operation of an encoder, according to some embodiments. [Figure 20] 10 is a flowchart illustrating the operation of an encoder, according to some embodiments. [Figure 21] 10 is a flowchart illustrating the operation of an encoder, according to some embodiments. [Figure 22] 10 is a flowchart illustrating the operation of an encoder, according to some embodiments. [Figure 23] 10 is a flowchart illustrating the operation of an encoder, according to some embodiments. [Figure 24] 10 is a flowchart illustrating the operation of an encoder, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0039] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Examples of embodiments of the inventive concepts are shown. The embodiments are provided as examples to convey the scope of the subject matter to those skilled in the art. However, the inventive concepts may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the inventive concepts to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. It may be implicitly assumed that a component from one embodiment is present / used in another embodiment.

[0040] As shown previously, avoiding the unpleasant effect of sudden changes in spatial characteristics during CNG after changing from active coding will make the comfort noise sound more natural.

[0041] Various embodiments calculate a set of coherence values ​​for each frame in which the VAD or SAD signals non-speech. These coherence values ​​are stored for at least two previous frames in time. Figure 10 shows the previous two frames in time. In some embodiments, more coherence values ​​may be used. In the following description, the last two frames are used to describe these embodiments.

[0042] When a new inactive segment begins, the low-pass filter state of the cross-spectrum Xspec(k,m) is initialized to: TIFF2025532351000009.tif16170∀k∈k b b=0,1,...,N band -1 rand(k) is a complex number with absolute value = 1 and random phase, where k b is the set of frequency coefficients for band b.

[0043] Smoothed left spectrum SPD_Lsmooth and the right spectrum SPD_R smooth is used as is, SPD_L smooth [k,m]=(1-α)·SPD_L smooth [k,m-1]+α·SPD_L(m,k) SPD_R smooth [k,m]=(1-α)·SPD_R smooth [k,m-1]+α·SPD_R(m,k) where α is a smoothing coefficient.

[0044] With this initialization, the coherence calculation is band This gives the result of (b,m-2). This is not important in itself, but band (b,m-2) is the updated Xspec smooth Instead of recalculating the coherence using the filter state, it can be used directly for the first frame in the inactive segment. smooth The idea is to start from a point where the filter state gives the same coherence as at the end of the previous inactive segment.

[0045] Note that in other embodiments, other frames near the end of the last inactive period may be used, e.g., C band (b,m-3) can be used, resulting in only a small increase in memory usage.

[0046] In some scenarios, Xspec smooth is set to 0 at the beginning of the VAD hangover period and then updated during the hangover period and the first inactive frame. smooth has not been updated enough times to give a reliable coherence estimate, but when computing the initialization value, Xspec smooth We show that using topological information from Xspec gives an improvement over using random numbers.smooth This is done by scaling [k,m] by its absolute value, which has absolute value 1, where Xspec smooth A complex number with topology [k,m] is given. TIFF2025532351000010.tif16170∀k∈k b b=0,1,...,N band -1 where k b is the set of frequency coefficients for band b.

[0047] For clarity, the first Xspec smooth [k,m] is Xspec smooth [k,m]=(1-α)·Xspec smooth [k,m-1]+α·Xspec(k,m) is used to calculate the phase of the resulting TIFF2025532351000011.tif8170Furthermore, TIFF2025532351000012.tif10170 is scaled and preserved.

[0048] A special case that needs to be handled is when the inactive segment is only one frame long. In that case, Xspec smooth C to initialize the [k,m] filter state band At this time point, C (b,m-2) will be used from the last frame of the previous inactive segment, i.e., a frame that may contain part of the speech onset. band (b,m-1) is taken. In normal operation, C band (b,m-2) is updated to C band (b,m-1), which in this case is C bandThis would lead to (b,m-2) sometimes containing the onset frame, which is something we want to avoid. The solution to this is to first select C in the first frame of the inactive segment, except for the second frame of the inactive segment. band (b,m-2) is not updated.

[0049] This problem is illustrated in Figure 11, where dashed frames are used in one frame-long inactive segment, and black frames containing onsets are used in the next inactive segment. band If (b,m-2) is not updated in the first frame of the inactive segment, the dashed frame will be used instead.

[0050] Figure 9 shows the advantage of the disclosed method in estimating coherence for a segment of inactive coding being transmitted in an SID frame to be used for CNG at the receiving side. Just like in Figure 8, the average over frequency is plotted. The true coherence of the signal across all frequencies is fixed at 0.2. It can be seen that the proposed method maintains a good coherence estimate while resetting the cross and power spectrum and restarting the estimation process. Restarting the estimation process means that there will be an inaccurate coherence estimate at the beginning of the second inactive segment.

[0051] One reason for resetting the cross-spectrum during active segments may be to improve ITD estimation for inactive segments, which would otherwise rely on cross-spectrum data highly influenced by the ITD of the active speech segments. It may also be advantageous to increase the filter coefficient α during hangover periods; as can be seen from Error! Reference source not found, having a large filter coefficient results in unstable coherence estimates. However, with the proposed method of reinitialized cross-spectrum, stable and reliable coherence estimates, as well as accurate ITD estimates, can be obtained while minimizing the memory footprint. The same memory may be used to store cross-spectral filter states, which may be reset and updated more rapidly during hangover periods and then copied into other cross-spectral filter states to be used for ITD estimation during CNG segments, and the same memory may be used to store cross-spectral estimates to be used to estimate coherence during CNG segments. As a result, in the case of a high-resolution DFT, two filter state vectors can be used for the cross-spectral estimate instead of three, which can imply a significant amount of memory, especially for applications in codecs targeting mobile devices, or any other device with limited memory capacity.

[0052] Before describing operations from the perspective of the encoder 400, FIG. 12 is a block diagram of an example operating environment 1200 in which the encoder 400 and decoder 500 may be implemented. In FIG. 12, the encoder 400 receives data, such as audio files, to be encoded from an entity such as a host 1204 and / or from storage 1206 over a network 1202. In some embodiments, the host 1204 may communicate directly with the encoder 400. The host 1204 may be included in various combinations of hardware and / or software, including a UE, a mobile phone, a terminal, a standalone server, a blade server, a cloud-implemented server, a distributed server, a virtual machine, a container, or processing resources in a server farm, etc. The encoder 400 either encodes audio files and stores the encoded audio files in storage 1206 and / or transmits the encoded audio files to the decoder 500 over a network 1208, as described herein. The decoder 500 decodes the audio files and transmits the decoded audio files to an audio player, such as a multi-channel audio player 1210, for playback. The decoder 500 may be in a UE, a mobile phone, a terminal, etc. The multi-channel audio player 1210 may be included in a user equipment, a terminal, a mobile phone, etc.

[0053] 13 is a block diagram illustrating elements of an encoder 400 configured to encode audio frames, in accordance with various embodiments herein. As shown, the encoder 400 may include a network interface circuit 1305 (also referred to as a network interface) configured to provide communication with other devices / entities / functions, etc. The encoder 400 may also include a processing circuit 1301 (also referred to as a processor and processor circuit) coupled to the network interface circuit 1305, and a memory circuit 1303 (also referred to as a memory) coupled to the processing circuit. The memory circuit 1303 may include computer-readable program code that, when executed by the processing circuit 1301, causes the processing circuit to perform operations in accordance with embodiments disclosed herein.

[0054] According to other embodiments, the processing circuit 1301 may be defined to include memory such that a separate memory circuit is not required. As described herein, the operations of the encoder 400 may be performed by the processing circuit 1301 and / or the network interface circuit 1305. For example, the processing circuit 1301 may control the network interface 1405 to send communications to the decoder 500 and / or receive communications through the network interface 1305 from one or more other network nodes / entities / servers, such as other encoder nodes, depository servers, etc. Moreover, modules may be stored in the memory 1303, and these modules may provide instructions such that, when the instructions of the modules are executed by the processing circuit 1301, the processing circuit 1301 performs respective operations.

[0055] 14 is a block diagram illustrating elements of a decoder 500 configured to decode audio frames, in accordance with some embodiments of the inventive concepts. As shown, the decoder 120 may include a network interface circuit 1405 (also referred to as a network interface) configured to provide communication with other devices / entities / functions, etc. The decoder 500 may also include a processing circuit 1401 (also referred to as a processor or processor circuit) coupled to the network interface circuit 1405, and a memory circuit 1403 (also referred to as a memory) coupled to the processing circuit. The memory circuit 1403 may include computer-readable program code that, when executed by the processing circuit 1401, causes the processing circuit to perform operations according to embodiments disclosed herein.

[0056] According to other embodiments, the processing circuit 1401 may be defined to include memory such that a separate memory circuit is not required. As described herein, the operations of the decoder 500 may be performed by the processor 1401 and / or the network interface 1405. For example, the processing circuit 1401 may control the network interface circuit 1405 to receive communications from the encoder 400. Moreover, modules may be stored in the memory 1403, and these modules may provide instructions such that, when the instructions of the modules are executed by the processing circuit 1401, the processing circuit 1401 performs respective operations.

[0057] 15 is a block diagram illustrating elements of a host 1204 configured to provide audio files to an encoder for encoding the audio files and sending the encoded audio files to a decoder 500, according to some embodiments. As shown, the host 1204 may include a network interface circuit 1505 (also referred to as a network interface) configured to provide communication with other devices / entities / functions, etc. The host 1204 may also include a processing circuit 1501 (also referred to as a processor or processor circuit) coupled to the network interface circuit 1505, and a memory circuit 1503 (also referred to as a memory) coupled to the processing circuit. The memory circuit 1503 may include computer-readable program code that, when executed by the processing circuit 1501, causes the processing circuit to perform operations according to embodiments disclosed herein.

[0058] According to other embodiments, the processing circuit 1501 may be defined to include memory such that a separate memory circuit is not required. As described herein, operations of the host 1204 may be performed by the processor 1501 and / or the network interface 1505. For example, the processing circuit 1501 may control the network interface circuit 1505 to send communications to the encoder 400. Moreover, modules may be stored in the memory 1503, and these modules may provide instructions such that, when the instructions of the modules are executed by the processing circuit 1501, the processing circuit 1501 performs respective operations.

[0059] The encoder 400 and decoder 500 may be virtualized in some embodiments by distributing the encoder 400 and / or decoder 500 across various components. Figure 16 is a block diagram illustrating an example of a virtualization environment 1600 in which functionality implemented by some embodiments may be virtualized. In this context, virtualizing means creating a virtual version of an apparatus or device, which may include virtualizing a hardware platform, storage devices, and networking resources. Virtualization, as used herein, may apply to any device described herein, or components thereof, and relates to implementations in which at least a portion of functionality is implemented as one or more virtual components. Some or all of the functionality described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtual environments 1600 hosted by one or more of the hardware nodes, such as a network node, a UE, a core network node, or a hardware computing device acting as a host. Furthermore, in embodiments in which the virtual node does not require wireless connectivity (e.g., to a core network node or host), the node may be fully virtualized.

[0060] An application 1602 (which may alternatively be referred to as a software instance, a virtual appliance, a network function, a virtual node, a virtual network function, etc.) is run in the virtualized environment 1600 to implement some of the features, functions, and / or benefits of some of the embodiments disclosed herein.

[0061] Hardware 1604 includes processing circuitry, memory that stores software and / or instructions executable by the hardware processing circuitry, and / or other hardware devices described herein, such as network interfaces, input / output interfaces, etc. Software is executed by the processing circuitry to instantiate one or more virtualization layers 1606 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 1608A and 1608B (one or more of which may be referred to generically as VMs 1608), and / or implement any of the functions, features, and / or benefits described with respect to some embodiments described herein. Virtualization layer 1606 may present to VMs 1608 a virtual operating platform that appears to be networking hardware.

[0062] The VMs 1608 may comprise virtual processing, virtual memory, virtual networking or interfaces, and virtual storage, and may be run by a corresponding virtualization layer 1606. Different embodiments of the virtual appliance 1602 instance may be implemented on one or more of the VMs 1608, and the implementation may be done in different ways. Hardware virtualization is referred to in some contexts as network functions virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry-standard high-volume server hardware, physical switches, and physical storage, which may be located in data centers and customer premises equipment.

[0063] In the context of NFV, a VM 1608 may be a software implementation of a physical machine that runs programs as if those programs were running on a physical, non-virtualized machine. Each VM 1608 and the portion of the hardware 1604 on which it runs, whether hardware dedicated to that VM and / or hardware shared by that VM with other VMs, form a separate virtual network element. Further, in the context of NFV, a virtual network function is responsible for handling a particular network function running in one or more VMs 1608 on the hardware 1604 and corresponds to the application 1602.

[0064] The hardware 1604 may be implemented in a standalone network node with general or specific components. The hardware 1604 may implement some functions via virtualization. Alternatively, the hardware 1604 may be part of a larger cluster of hardware (e.g., as in a data center or CPE) where many hardware nodes cooperate and are managed via a management and orchestration 1610 that, among other things, oversees the lifecycle management of the application 1602. In some embodiments, the hardware 1604 is coupled to one or more radio units, each including one or more transmitters and one or more receivers, which may be coupled to one or more antennas. The radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces and may be used in combination with virtual components to provide a virtual node with wireless capabilities, such as a wireless access node or base station. In some embodiments, some signaling may be provided using a control system 1612, which may alternatively be used for communication between the hardware nodes and the radio units.

[0065] The operation of encoder 400 (implemented using the block diagram structures of FIGS. 4 and 13) will now be described with reference to the flowchart of FIG. 17, in accordance with some embodiments of the inventive concept. For example, modules may be stored in memory 1303 of FIG. 13 that provide instructions such that, when the instructions of the modules are executed by respective encoder processing circuits 1301, encoder 400 performs the respective operations of the flowchart.

[0066] 17 illustrates operations performed by the encoder 400 in various embodiments. Referring to FIG. 17, in block 1701, the encoder 400 receives a time-domain audio input comprising an audio input signal. The audio input signal may be speech, music, or a combination thereof.

[0067] In block 1703, the encoder 400 processes the audio input signal frame by frame, as shown in blocks 1705 to 1711. The encoder 400 can perform the processing in the time domain or in the frequency domain.

[0068] In blocks 1705-1711, the encoder 400 encodes each of the audio input signals. Specifically, in block 1705, the encoder 400 encodes the active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal. To detect the periods of inactivity described above, a VAD (e.g., VAD 102) or a SAD may be used.

[0069] In block 1707, the encoder 400 switches encoding from the active encoding content to the inactive encoding to encode background noise at a second bit rate during the pause period, the second bit rate generally being less than the first bit rate described above.

[0070] In block 1709, the encoder 400 estimates coherence parameters during the inactive period based on cross-spectral low-pass filtering or cross-spectral averaging, where estimating the coherence parameters includes activating a cross-spectral low-pass filter state based on the coherence parameters from a previous inactive period.

[0071] In some embodiments, as shown in FIG. 18, in estimating the coherence parameters, the encoder 400 may, in block 1801, generate a first cross-spectral low-pass filter X in the first coded frame after active coding based on the coherence parameters from a period prior to inactive coding. corr_smooth Reinitialize the state of

[0072] In some embodiments, the encoder 400 generates a first cross-spectral low-pass filter X based on the last two frames from the previous period of inactive coding. spec_smooth Reinitialize the state of

[0073] In other embodiments, the coherence parameter may be various functions of the previous coherence value. For example, the encoder 400 may calculate the coherence parameter by choosing the penultimate one from the previous inactive period, by taking the average of the last estimated coherence parameters (and potentially excluding the last one), by taking a weighted average of the previous coherence values, by using a filtered version of an earlier coherence value, e.g., By creating TIFF2025532351000013.tif7170, we estimate band C instead of (b,m-2) band_filt Use (b,m) to create X spec_smooth can be reinitialized, and so on.

[0074] In block 1901 of FIG. 19, the encoder 400 performs a second low-pass filter X during the DTX hangover period. spec_smooth Start updating.

[0075] In block 1711, the encoder 400 encodes the estimated coherence parameters.

[0076] In block 1713 , the encoder 400 initiates sending the encoded active content, the encoded background noise, and the coherence parameters to the decoder 500 .

[0077] In some embodiments that process the audio input signal frame by frame, the encoder 400 processes the audio input signal frame by frame to produce a mono mixdown signal, and the encoder 400 encodes the active content of each audio input signal by encoding the active content of the mono mixdown signal.

[0078] In still other embodiments, encoder 400 processes the audio input signal on a frame-by-frame basis to produce a mono mixdown signal and one or more stereo parameters. In these embodiments, encoder 400 encodes the active content of the mono mixdown signal and the one or more stereo parameters.

[0079] In some embodiments, the encoder 400 spec_smooth of, TIFF2025532351000014.tif16170∀k∈k b b=0,1,...,N band -1 SPD_L smooth [k,m]=(1-α)·SPD_L smooth [k,m-1]+α·SPD_L(m,k) SPD_R smooth [k,m]=(1-α)·SPD_R smooth [k,m-1]+α·SPD_R(m,k) Determined according to TIFF2025532351000015.tif20170, where · denotes multiplication, α is the low-pass coefficient, and k b is the set of frequency coefficients for band b, bandlimits(b) is a vector containing the limits between frequency bands, and rand(k) is a complex number with modulus=1 and random phase.

[0080] In some other embodiments, the encoder 400 spec_smooth of, TIFF2025532351000016.tif16170∀k∈k b b=0,1,...,N band -1 SPD_L smooth [k,m]=(1-α)·SPD_L smooth [k,m-1]+α·SPD_L(m,k) SPD_R smooth [k,m]=(1-α)·SPD_R smooth [k,m-1]+α·SPD_R(m,k) Determined according to TIFF2025532351000017.tif20170, where · denotes multiplication, α is the low-pass coefficient, and k b is the set of frequency coefficients for band b, and bandlimits(b) is a vector containing the limits between frequency bands. As explained before, for clarity, the first Xspec smooth [k,m] is Xspec smooth [k,m]=(1-α)·Xspec smooth [k,m-1]+α·Xspec(k,m) is used to calculate the phase of the resulting TIFF2025532351000018.tif8170Furthermore, TIFF2025532351000019.tif10170 is scaled and preserved.

[0081] X spec_smooth In the above manner for determining C, the encoder 400 uses a weighting function, as shown in block 2001 of FIG. band(b,m). For example, as described above, the encoder 400 may C with weighting function according to TIFF2025532351000020.tif12170 band (b,m) may be weighted, where |LR(m,k)| 2 is the discrete Fourier transform (DFT) energy spectrum for a mono signal that is a downmix of the audio input signal.

[0082] In some embodiments, the previous inactive period may consist of only one frame. In such cases, C band Processing (b,m-2) may result in an onset frame that is part of the comfort noise, which is undesirable. To take this into account, the encoder 400 may add C band (b,m-2) is not updated in the first frame of an inactive period having multiple frames, except in the second frame of an inactive period having multiple frames.

[0083] In other embodiments, a dedicated cross-correlation estimate may be used. As shown in block 2201 of Figure 22, the encoder 400 performs a dedicated cross-correlation estimate that is updated only during quiet periods and / or during DTX hangover frames for the cross-spectrum, and uses the dedicated cross-correlation estimate for coherence estimation during inactive periods.

[0084] In a further embodiment, as shown in block 2301 of Figure 23, the encoder 400 speeds up cross-spectral smoothing by low-pass filtering by resetting the cross-spectral low-pass filter state one of before an update during a DTX hangover period and before an update during a quiescent period. Additionally or alternatively, the filter coefficient α may be increased to speed up the effect of a new frame being processed.

[0085] In yet another embodiment, as shown in block 2301 of FIG. 23, the encoder 400 accelerates cross-spectral smoothing through low-pass filtering by replacing the low-pass filter state at the start of a hangover period or at the start of an inactive period.

[0086] In yet a further embodiment, as shown in block 2401 of FIG. 24, the encoder 400 reinitializes the low-pass filtering state at the start of a hangover period or at the start of an inactivity period.

[0087] While the computing devices (e.g., encoders, decoders, hosts) described herein may include the depicted combinations of hardware components, other embodiments may comprise computing devices with different combinations of components. It should be understood that these computing devices may comprise any suitable combination of hardware and / or software required to perform the tasks, features, functions, and methods disclosed herein. The determining, calculating, obtaining, or similar operations described herein may be performed by processing circuitry, which may process information by, for example, converting obtained information to other information, comparing the obtained or converted information with information stored in a network node, and / or performing one or more operations based on the obtained or converted information and as a result of the processing making a decision. Moreover, while components are shown as a single box located within a larger box or nested within multiple boxes, in reality the computing device may comprise multiple different physical components that make up the single depicted component, and functionality may be partitioned among the separate components. For example, a communications interface may be configured to include any of the components described herein, and / or the functionality of those components may be partitioned between the processing circuitry and the communications interface. In another example, non-computationally intensive functionality of any of such components may be implemented in software or firmware, and computationally intensive functionality may be implemented in hardware.

[0088] In some embodiments, some or all of the functionality described herein may be provided by a processing circuit executing instructions stored in a memory, which in some embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by the processing circuit without executing instructions stored on a separate or discrete device-readable storage medium, such as in a hardwired manner. In any of these particular embodiments, the processing circuit may be configured to perform the described functionality, regardless of whether or not it executes instructions stored on a non-transitory computer-readable storage medium. Benefits provided by such functionality are not limited to the processing circuit alone or to other components of the computing device, but are enjoyed by the computing device as a whole and / or by end users and wireless networks generally.

[0089] Embodiment Embodiment 1. A method in an encoder (400) for enabling generation of comfort noise using estimated coherence parameters in a network using discontinuous transmission (DTX), the method comprising: Receiving a time-domain audio input (1701) comprising an audio input signal; Processing (1703) an audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to the inactive encoding to encode background noise at a second bit rate during the inactive period; estimating coherence parameters during an inactive period based on cross-spectral low-pass filtering or cross-spectral averaging (1709), where estimating the coherence parameters includes reinitializing cross-spectral low-pass filter states based on coherence parameters from a previous inactive period; Encoding the estimated coherence parameters (1711) processing (1703) the audio input signal frame by frame by Initiating (1713) the transmission of the encoded active content, the encoded background noise, and the encoded coherence parameters to the decoder (500); A method comprising: Embodiment 2. Estimating the coherence parameters comprises: In the first coding frame after active coding, a first cross-spectral low-pass filter X is generated based on the coherence parameters from the previous period of inactive coding. spec_smooth Reinitializing the state of (1801) 2. The method of embodiment 1, comprising: Embodiment 3. A first cross-spectral low-pass filter X is generated based on the coherence parameters from the previous period of inactive coding. spec_smooth reinitializing the state of the first cross-spectral low-pass filter X based on the last two frames from the previous period of inactive coding; spec_smooth 3. The method of claim 2, further comprising reinitializing a state of Embodiment 4. Low-pass filter X during DTX hangover period spec_smooth Start updating (1901) 4. The method of embodiment 2 or 3, further comprising: Embodiment 5. A method according to any one of embodiments 1 to 4, wherein processing the audio input signal frame by frame includes processing the audio input signal frame by frame to produce a mono mixdown signal, and encoding the active content of each audio input signal includes encoding the active content of the mono mixdown signal. Embodiment 6. The method of embodiment 5, wherein processing the audio input signal frame by frame to produce a mono mixdown signal includes processing the audio input signal frame by frame to produce a mono mixdown signal and one or more stereo parameters, and encoding the active content of the mono mixdown signal includes encoding the active content of the mono mixdown signal and one or more stereo parameters. Embodiment 7. X spec_smooth but, TIFF2025532351000021.tif16170∀k∈k b b=0,1,...,N band -1 SPD_L smooth [k,m]=(1-α)·SPD_L smooth [k,m-1]+α·SPD_L(k,m) SPD_R smooth [k,m]=(1-α)·SPD_R smooth [k,m-1]+α·SPD_R(k,m) Determined according to TIFF2025532351000022.tif20170, where · denotes multiplication, α is the low-pass coefficient, and k b 7. A method according to any one of claims 2 to 6, wherein k is a set of frequency coefficients for band b, bandlimits(b) is a vector containing limits between frequency bands, and rand(k) is a complex number with absolute value = 1 and a random phase. Embodiment 8. Weighting function C band 8. The method of embodiment 7, further comprising weighting (b, m) (2001). Embodiment 9. Weighting function Cband Weighting (b,m) is Weighted according to TIFF2025532351000023.tif12170, where |LR(m,k)| 2 9. The method of embodiment 8, wherein {right arrow over (x)} is a discrete Fourier transform (DFT) energy spectrum for a mono signal that is a downmix of the audio input signal. Embodiment 10. X spec_smooth but, TIFF2025532351000024.tif16170∀k∈k b b=0,1,...,N band -1 SPD_L smooth [k,m]=(1-α)·SPD_L smooth [k,m-1]+α·SPD_L(k,m) SPD_R smooth [k,m]=(1-α)·SPD_R smooth [k,m-1]+α·SPD_R(k,m) Determined according to TIFF2025532351000025.tif20170, where · denotes multiplication, α is the low-pass coefficient, and k b 7. The method according to any one of embodiments 2 to 6, wherein {circumflex over (x)} is a set of frequency coefficients for band b, and bandlimits(b) is a vector containing limits between frequency bands. Embodiment 11. Weighting function C band 11. The method of embodiment 10, further comprising weighting (b, m) (2001). Embodiment 12. Weighting function C band Weighting (b,m) is Weighted according to TIFF2025532351000026.tif12170, where |LR(m,k)| 2 12. The method of embodiment 11, wherein {right arrow over (x)} is a discrete Fourier transform (DFT) energy spectrum for a mono signal that is a downmix of the audio input signal. Embodiment 13. Cband (b, m-2) in the first frame of an inactive period having multiple frames, except in the second frame of the inactive period having multiple frames (2101). 13. The method of any one of claims 1 to 12, further comprising: Embodiment 14. performing dedicated cross-correlation estimates for the cross spectrum that are updated only during inactive periods and / or during DTX hangover frames, and using the dedicated cross-correlation estimates for coherence estimation during inactive periods (2201); 13. The method of any one of embodiments 1 to 12, further comprising: Embodiment 15. resetting the cross-spectral low-pass filter state one of before an update during a DTX hangover period and before an update during an inactive period (2301); 15. The method of any one of embodiments 1 to 14, further comprising: Embodiment 16. Reinitializing the low-pass filter state at the start of a hangover period or at the start of an inactivity period (2401) 16. The method of any one of embodiments 1 to 15, further comprising: Embodiment 17. An encoder (400) adapted to enable the generation of comfort noise using an estimated coherence parameter in a network using discontinuous transmission (DTX), the encoder comprising: Receiving a time-domain audio input (1701) comprising an audio input signal; Processing (1703) an audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to the inactive encoding to encode background noise at a second bit rate during the inactive period; estimating 1709 a coherence parameter during an inactive period based on cross-spectral low-pass filtering or cross-spectral averaging, where estimating the coherence parameter includes activating a cross-spectral low-pass filter state based on the coherence parameter from a previous inactive period; Encoding the estimated coherence parameters (1711) processing (1703) the audio input signal frame by frame by Initiating (1713) the transmission of the encoded active content, the encoded background noise, and the encoded coherence parameters to the decoder (500); an encoder (400) adapted to: Embodiment 18. The encoder (400) of embodiment 17, wherein the encoder is further adapted to perform according to any one of embodiments 2 to 16. Embodiment 19. An encoder (400) adapted to enable generation of comfort noise using estimated coherence parameters in a network using discontinuous transmission (DTX), the encoder comprising: A processing circuit (1301); a memory (1303) coupled to the processing circuit; wherein the memory contains instructions that, when executed by the processing circuit, cause the encoder to perform operations, the operations including: Receiving a time-domain audio input (1701) comprising an audio input signal; Processing (1703) an audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to the inactive encoding to encode background noise at a second bit rate during the inactive period; estimating 1709 a coherence parameter during an inactive period based on cross-spectral low-pass filtering or cross-spectral averaging, where estimating the coherence parameter includes activating a cross-spectral low-pass filter state based on the coherence parameter from a previous inactive period; Encoding the estimated coherence parameters (1711) processing (1703) the audio input signal frame by frame by Initiating (1713) the transmission of the encoded active content, the encoded background noise, and the encoded coherence parameters to the decoder (500); an encoder (400) including: Embodiment 20. The encoder (400) of embodiment 19, wherein the memory includes further instructions that, when executed by the processing circuit, cause the encoder to perform any one of embodiments 2 to 16. Embodiment 21. A computer program comprising program code to be executed by a processing circuit (1301) of an encoder (400), whereby execution of the program code causes the encoder (400) to perform operations, the operations being: Receiving a time-domain audio input (1701) comprising an audio input signal; Processing (1703) an audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to the inactive encoding to encode background noise at a second bit rate during the inactive period; estimating 1709 a coherence parameter during an inactive period based on cross-spectral low-pass filtering or cross-spectral averaging, where estimating the coherence parameter includes activating a cross-spectral low-pass filter state based on the coherence parameter from a previous inactive period; Encoding the estimated coherence parameters (1711) processing (1703) the audio input signal frame by frame by Initiating (1713) the transmission of the encoded active content, the encoded background noise, and the encoded coherence parameters to the decoder (500); a computer program comprising: Embodiment 22. The computer program of embodiment 21, comprising further program code to be executed by a processing circuit of the encoder, whereby execution of the program code causes the encoder (400) to perform the operations of any one of embodiments 2 to 16. Embodiment 23. A computer program product comprising a non-transitory storage medium containing program code to be executed by a processing circuit (1301) of an encoder (400), whereby execution of the program code causes the encoder (400) to perform operations, the operations being: Receiving a time-domain audio input (1701) comprising an audio input signal; Processing (1703) an audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to the inactive encoding to encode background noise at a second bit rate during the inactive period; estimating 1709 a coherence parameter during an inactive period based on cross-spectral low-pass filtering or cross-spectral averaging, where estimating the coherence parameter includes activating a cross-spectral low-pass filter state based on the coherence parameter from a previous inactive period; Encoding the estimated coherence parameters (1711) processing (1703) the audio input signal frame by frame by Initiating (1713) the transmission of the encoded active content, the encoded background noise, and the encoded coherence parameters to the decoder (500); a computer program product, Embodiment 24. A computer program product as described in embodiment 23, wherein the non-transitory storage medium includes further program code to be executed by the processing circuitry (1301) of the encoder (400), whereby execution of the further program code causes the encoder (400) to perform the operations described in any one of embodiments 2 to 16. Embodiment 25. A method implemented by a host (1204) configured to operate in a communication system further including an encoder, a decoder, and an audio player, the method comprising: providing user data to a decoder (500) for decoding an audio file for an audio player; Initiating transmission of an audio file to an audio player via a cellular network comprising an encoder (400), the encoder (400) performing the following operations for transmitting user data from a host to a decoder (500): Receiving a time-domain audio input (1701) comprising an audio input signal; Processing (1703) an audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to the inactive encoding to encode background noise at a second bit rate during the inactive period; estimating 1709 a coherence parameter during an inactive period based on cross-spectral low-pass filtering or cross-spectral averaging, where estimating the coherence parameter includes activating a cross-spectral low-pass filter state based on the coherence parameter from a previous inactive period; Encoding the estimated coherence parameters (1711) processing (1703) the audio input signal frame by frame by Initiating (1713) the transmission of the encoded active content, the encoded background noise, and the encoded coherence parameters to the decoder (500); Initiating a transmission carrying an audio file, A method comprising: Embodiment 26. The encoder (400) Implementing the method according to any one of embodiments 2 to 16 26. The method of embodiment 25, further configured as follows: Embodiment 27. A host (1204) configured to operate in a communication system for providing over-the-top (OTT) services, the host comprising: a processing circuit (1501) configured to provide an audio file; a network interface (1505) configured to initiate transmission of an audio file to the encoder (400) over a cellular network for transmission to the decoder (500); The encoder (400) includes a network interface (1305) and a processing circuit (1301), and the processing circuit (1301) of the encoder (400) performs the following operations to transmit an audio file from a host to a decoder (500): Receiving a time-domain audio input (1701) comprising an audio input signal; Processing (1703) an audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to the inactive encoding to encode background noise at a second bit rate during the inactive period; estimating 1709 a coherence parameter during an inactive period based on cross-spectral low-pass filtering or cross-spectral averaging, where estimating the coherence parameter includes activating a cross-spectral low-pass filter state based on the coherence parameter from a previous inactive period; Encoding the estimated coherence parameters (1711) processing (1703) the audio input signal frame by frame by Initiating (1713) the transmission of the encoded active content, the encoded background noise, and the encoded coherence parameters to the decoder (500); The host (1204) configured to implement this. Embodiment 28. The processing circuit of the encoder (400) Implementing the method according to any one of embodiments 2 to 16 28. The host of embodiment 27, further configured to:

[0090] References are identified below. U.S. Patent Application Publication No. 20200194013 U.S. Patent No. 11,417,348 U.S. Patent No. 11,404,069

Claims

1. 1. A method in an encoder (400) for enabling generation of comfort noise using estimated coherence parameters in a network using discontinuous transmission (DTX), the method comprising: Receiving a time-domain audio input (1701) comprising an audio input signal; processing (1703) the audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to inactive encoding to encode background noise at a second bit rate during the inactive period; estimating coherence parameters during the inactive period based on low-pass filtering of the cross-spectrum or averaging of the cross-spectrum (1709), wherein estimating the coherence parameters includes re-initializing the cross-spectral low-pass filter state based on coherence parameters from a previous inactive period; encoding (1711) the estimated coherence parameters; processing (1703) the audio input signal frame by frame by A method comprising:

2. Initiating (1713) the transmission of the encoded active content, the encoded background noise, and the encoded coherence parameter to a decoder (500). The method of claim 1 further comprising:

3. estimating the coherence parameter In the first coding frame after active coding, a first cross-spectral low-pass filter X is generated based on the coherence parameters from the previous period of inactive coding. spec_smooth Reinitializing the state of (1801) 3. The method of claim 1 or 2, comprising:

4. the first cross-spectral low-pass filter X based on a coherence parameter from a previous period of inactive coding; spec_smooth reinitializing the state of the first cross-spectral low-pass filter X based on the last two frames from a previous period of inactive coding; spec_smooth 4. The method of claim 3, further comprising reinitializing the state of

5. the first cross-spectral low-pass filter X based on a coherence parameter from a previous period of inactive coding; spec_smooth reinitializing the state of the first cross-spectral low-pass filter X based on the penultimate frame from a previous period of inactive coding; spec_smooth 4. The method of claim 3, further comprising reinitializing the state of

6. During the DTX hangover period, the low-pass filter X spec_smooth Start updating (1901) 6. The method of claim 3, further comprising:

7. 7. The method of claim 1, wherein processing the audio input signals frame by frame comprises processing the audio input signals frame by frame to produce a mono mixdown signal, and wherein encoding the active content of each audio input signal comprises encoding the active content of the mono mixdown signal.

8. 8. The method of claim 7, wherein processing the audio input signal frame-by-frame to produce the mono mixdown signal comprises processing the audio input signal frame-by-frame to produce the mono mixdown signal and one or more stereo parameters, and wherein encoding the active content of the mono mixdown signal comprises encoding the active content of the mono mixdown signal and the one or more stereo parameters.

9. X spec_smooth but, ∀k∈k b b=0,1,...,N band -1 SPD_L smooth [k,m]=(1-α)・SPD_L smooth [k,m-1]+α・SPD_L(k,m) SPD_R smooth [k,m]=(1-α)・SPD_R smooth [k,m-1]+α・SPD_R(k,m) is determined in accordance with where ⋅ denotes multiplication, α is the low-pass coefficient, and k b 9. The method of claim 3, wherein k is a set of frequency coefficients for band b, bandlimits(b) is a vector containing the limits between the frequency bands, and rand(k) is a complex number with absolute value = 1 and a random phase.

10. The weighting function is band The method of claim 9 further comprising weighting (2001) (b, m).

11. The weighting function band Weighting (b, m) are weighted according to Here, |LR(m, k)| 2 The method of claim 10 , wherein ∑ i = 1 i ...

12. X spec_smooth but, ∀k∈k b b=0,1,...,N band -1 SPD_L smooth [k,m]=(1-α)・SPD_L smooth [k,m-1]+α・SPD_L(k,m) SPD_R smooth [k,m]=(1-α)・SPD_R smooth [k,m-1]+α・SPD_R(k,m) is determined in accordance with where ⋅ denotes multiplication, α is the low-pass coefficient, and k b 9. The method of claim 3, wherein {circumflex over (x)} is a set of frequency coefficients for band b, and bandlimits(b) is a vector containing the limits between frequency bands.

13. The weighting function is band The method of claim 12 further comprising weighting (2001) (b, m).

14. The weighting function band Weighting (b, m) and weighted according to Here, |LR(m, k)| 2 The method of claim 13 , wherein ∑ i = 1 ⁢ ...

15. Said C band (b, m-2) in the first frame of the inactive period having a plurality of frames except in the second frame of the inactive period having the plurality of frames (2101).

15. The method of any one of claims 1 to 14, further comprising:

16. performing dedicated cross-correlation estimates for the cross-spectrum, the estimates being updated only during the inactive periods and / or during DTX hangover frames, and using the dedicated cross-correlation estimates for coherence estimation during the inactive periods (2201); 15. The method of any one of claims 1 to 14, further comprising:

17. resetting the cross-spectral low-pass filter state one of before an update during a DTX hangover period and before an update during the inactive period (2301); 17. The method of any one of claims 1 to 16, further comprising:

18. Reinitializing the low pass filter state at the beginning of a hangover period or at the beginning of said inactive period (2401).

18. The method of any one of claims 1 to 17, further comprising:

19. 1. An encoder (400) adapted to enable generation of comfort noise using an estimated coherence parameter in a network using discontinuous transmission (DTX), the encoder comprising: Receiving a time-domain audio input (1701) comprising an audio input signal; processing (1703) the audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to inactive encoding to encode background noise at a second bit rate during the inactive period; estimating coherence parameters during the inactive period based on low-pass filtering of the cross-spectrum or averaging of the cross-spectrum (1709), wherein estimating the coherence parameters includes activating a low-pass filter state of the cross-spectrum based on coherence parameters from a previous inactive period; encoding (1711) the estimated coherence parameters; processing (1703) the audio input signal frame by frame by An encoder (400) adapted to:

20. 20. The encoder (400) of claim 19, wherein said encoder is further adapted to perform the method of any one of claims 2 to 18.

21. 1. An encoder (400) adapted to enable generation of comfort noise using an estimated coherence parameter in a network using discontinuous transmission (DTX), the encoder comprising: A processing circuit (1301); a memory (1303) coupled to said processing circuit; the memory containing instructions that, when executed by the processing circuitry, cause the encoder to perform operations, the operations including: Receiving a time-domain audio input (1701) comprising an audio input signal; processing (1703) the audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to inactive encoding to encode background noise at a second bit rate during the inactive period; estimating coherence parameters during the inactive period based on low-pass filtering of the cross-spectrum or averaging of the cross-spectrum (1709), wherein estimating the coherence parameters includes activating a low-pass filter state of the cross-spectrum based on coherence parameters from a previous inactive period; encoding (1711) the estimated coherence parameters; processing (1703) the audio input signal frame by frame by an encoder (400) including:

22. 22. The encoder (400) of claim 21, wherein the memory includes further instructions that, when executed by the processing circuitry, cause the encoder to perform a method according to any one of claims 2 to 18.

23. A computer program comprising program code to be executed by a processing circuit (1301) of an encoder (400), whereby execution of said program code causes said encoder (400) to perform operations, said operations being: Receiving a time-domain audio input (1701) comprising an audio input signal; processing (1703) the audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to inactive encoding to encode background noise at a second bit rate during the inactive period; estimating coherence parameters during the inactive period based on low-pass filtering of the cross-spectrum or averaging of the cross-spectrum (1709), wherein estimating the coherence parameters includes activating a low-pass filter state of the cross-spectrum based on coherence parameters from a previous inactive period; encoding (1711) the estimated coherence parameters; processing (1703) the audio input signal frame by frame by a computer program comprising:

24. 24. The computer program of claim 23, comprising further program code to be executed by the processing circuitry of the encoder, whereby execution of the program code causes the encoder (400) to perform the operations of any one of claims 2 to 18.

25. A computer program product comprising a non-transitory storage medium containing program code to be executed by a processing circuit (1301) of an encoder (400), whereby execution of the program code causes the encoder (400) to perform operations, the operations being: Receiving a time-domain audio input (1701) comprising an audio input signal; processing (1703) the audio input signal frame by frame, encoding (1705) active content of each audio input signal at a first bit rate until a period of inactivity is detected in the audio input signal; switching (1707) the encoding from the active encoding content to inactive encoding to encode background noise at a second bit rate during the inactive period; estimating coherence parameters during the inactive period based on low-pass filtering of the cross-spectrum or averaging of the cross-spectrum (1709), wherein estimating the coherence parameters includes activating a low-pass filter state of the cross-spectrum based on coherence parameters from a previous inactive period; encoding (1711) the estimated coherence parameters; processing (1703) the audio input signal frame by frame by a computer program product,

26. 26. The computer program product of claim 25, wherein the non-transitory storage medium comprises further program code to be executed by a processing circuit (1301) of an encoder (400), whereby execution of the further program code causes the encoder (400) to perform the operations of any one of claims 2 to 18.

Citation Information

Patent Citations

  • Support for comfort noise generation

    JP2021520515A

  • Comfort noise generation

    US20170047072A1

  • Multi-channel signal generator, audio encoder and related methods relying on a mixing noise signal

    WO2022042908A1