Support for comfort noise generation

By determining and compressing spectral characteristics and spatial coherence between audio channels using perceptual importance measures, the method enhances comfort noise generation in multi-channel systems, reducing resource usage and maintaining a realistic stereo image.

JP7746431B2Active Publication Date: 2025-09-30TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024019342
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-04-05
Filing Date
2024-02-13
Publication Date
2025-09-30
Estimated Expiration
2039-04-05

AI Technical Summary

Technical Problem

Existing mechanisms for generating comfort noise in multi-channel audio systems suffer from inefficiencies and artifacts due to unsynchronized encoding and decoding of audio channels, leading to inconsistent stereo images and increased resource usage.

Method used

A method and system that determines spectral characteristics and spatial coherence between audio channels, compresses this information into frequency bands using perceptual importance measures, and signals this information to the receiving node for efficient comfort noise generation, maintaining a realistic stereo image.

Benefits of technology

This approach reduces the amount of encoded information while preserving a realistic stereo image, thereby minimizing resource usage and improving the quality of comfort noise generation for multiple audio channels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007746431000061
    Figure 0007746431000061
  • Figure 0007746431000062
    Figure 0007746431000062
  • Figure 0007746431000063
    Figure 0007746431000063
Patent Text Reader

Abstract

To provide methods, transmitting nodes, and programs that enable efficient generation of comfort noise for two or more channels, as well as receiving nodes, methods and programs for the generation of comfort noise at receiving nodes.SOLUTION: A method for supporting generation of comfort noise for at least two audio channels at a receiving node, which is performed by a transmitting node, includes determining spectral characteristics of audio signals of the at least two input audio channels, and determining spatial coherence between the audio signals. A compressed representation of the spatial coherence associated with perceptual importance measure is determined for each frequency band by weighting the spatial coherence within each frequency band according to the perceptual importance measure. And signaling information about the spectral characteristics and the compressed representation of spatial coherence for each frequency band is performed to the receiving node to enable the generation of the comfort noise at the receiving node.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The embodiments presented herein relate to a method, a transmitting node, a computer program, and a computer program product for supporting comfort noise generation for at least two audio channels at a receiving node. The embodiments presented herein further relate to a method, a receiving node, a computer program, and a computer program product for comfort noise generation at a receiving node. [Background technology]

[0002] In a communication network, there can be challenges to obtain good performance and capacity for a given communication protocol, its parameters, and the physical environment in which the communication network is deployed.

[0003] For example, capacity in telecommunications networks is continually increasing, but limiting the required resource usage per user remains a concern. In mobile telecommunications networks, lower resource usage per call means that the mobile telecommunications network can serve multiple users in parallel. Lower resource usage also leads to lower power consumption in both user-side devices (e.g., terminal devices) and network-side devices (e.g., network nodes). This translates into energy and cost savings for network operators, while enabling longer battery life and talk time to be experienced in terminal devices.

[0004] One mechanism for reducing the required resource usage of voice communication applications in mobile telecommunications networks is to exploit natural pauses in speech. More specifically, in most conversations, only one party is active at a time, and therefore speech pauses in one direction of communication typically dominate the signal. One way to take advantage of this characteristic and reduce the required resource usage is to use a discontinuous transmission (DTX) system, in which active signal encoding is suspended during speech pauses.

[0005] During speech pauses, it is common to transmit a very low bitrate encoding of background noise so that a Comfort Noise Generator (CNG) system at the receiving end can fill these pauses with background noise having similar characteristics to the original noise. Because the background noise is maintained and does not switch on and off with the speech, CNG makes the sound more natural compared to having silence during speech pauses. Complete silence during speech pauses is generally perceived as unpleasant and often leads to the mistaken belief that the call has been dropped.

[0006] DTX systems may further rely on a Voice Activity Detector (VAD) that instructs the transmitting device whether to use active signal encoding or low-rate background noise encoding. In this regard, the transmitting device may be configured to distinguish between other source types by using a (Generic) Sound Activity Detector (GSAD or SAD), which not only distinguishes between background noise and speech, but may also be configured to detect music or other signal types deemed relevant.

[0007] Communication services can be further enhanced by supporting stereo or multi-channel audio transmission. In these cases, the DTX / CNG system can also take into account the spatial characteristics of the signal to provide a pleasant-sounding comfort noise.

[0008] A common mechanism for generating comfort noise would be to transmit information about the energy and spectral shape of the background noise during pauses in speech, which can be done using significantly fewer bits than conventional coding of speech segments.

[0009] At the receiving device, comfort noise is generated by creating a pseudorandom signal and then shaping the spectrum of the signal with a filter based on information received from the transmitting device. Signal generation and spectrum shaping can be performed in the time or frequency domain. Summary of the Invention

[0010] An objective of the embodiments herein is to enable efficient generation of comfort noise for two or more channels.

[0011] According to a first aspect, a method for supporting comfort noise generation for at least two audio channels at a receiving node is presented. The method is executed by a transmitting node. The method includes determining spectral characteristics of audio signals of at least two input audio channels. The method includes determining spatial coherence between the audio signals of the respective input audio channels, the spatial coherence being associated with a perceptual importance measure. The method includes dividing the spatial coherence into frequency bands, and a compressed representation of the spatial coherence is determined for each frequency band by weighting the spatial coherence within each frequency band according to the perceptual importance measure. The method includes signaling, to the receiving node, information regarding the spectral characteristics and information regarding the compressed representation of the spatial coherence for each frequency band to enable comfort noise generation for the at least two audio channels at the receiving node.

[0012] According to a second aspect, a transmitting node for supporting generation of comfort noise for at least two audio channels at a receiving node is presented. The transmitting node includes a processing circuit. The processing circuit is configured to cause the transmitting node to determine spectral characteristics of audio signals of at least two input audio channels. The processing circuit is configured to cause the transmitting node to determine spatial coherence between the audio signals of each input audio channel, the spatial coherence being associated with a perceptual importance measure. The processing circuit is configured to cause the transmitting node to divide the spatial coherence into frequency bands, and a compressed representation of the spatial coherence is determined for each frequency band by weighting the spatial coherence within each frequency band according to the perceptual importance measure. The processing circuit is configured to cause the transmitting node to signal information regarding the spectral characteristics and information regarding the compressed representation of the spatial coherence for each frequency band to the receiving node to enable generation of comfort noise for the at least two audio channels at the receiving node.

[0013] According to a third aspect, a computer program for supporting comfort noise generation for at least two audio channels at a receiving node is presented, the computer program comprising computer program code that, when executed at a transmitting node, causes the transmitting node to perform at least the method according to the first aspect.

[0014] According to a fourth aspect, there is provided a computer program product comprising the computer program according to the third aspect and a computer-readable storage medium on which the computer program is stored. The computer-readable storage medium may be a non-transitory computer-readable storage medium.

[0015] According to a fifth aspect, a wireless transceiver device is presented, the wireless transceiver device comprising a transmitting node according to the second aspect.

[0016] Advantageously, the methods, the transmitting node, the computer program, the computer program product and the wireless transceiver device enable efficient generation of comfort noise for two or more channels.

[0017] Advantageously, these methods, these transmitting nodes, this computer program, this computer program product and this wireless transceiver device allow comfort noise to be generated for two or more channels without suffering from the aforementioned problems.

[0018] Advantageously, these methods, these transmitting nodes, this computer program, this computer program product and this wireless transceiver device make it possible to reduce the amount of information that needs to be encoded in a stereo or multi-channel DTX system while retaining the ability to recreate a realistic stereo image at the receiving node.

[0019] Other objects, features, and advantages of the included embodiments will become apparent from the following detailed disclosure, from the claims, and from the drawings.

[0020] The inventive concepts will now be described, by way of example, with reference to the accompanying drawings in which: [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a schematic diagram illustrating a communication network according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating a schematic diagram of a DTX system according to an embodiment. [Figure 3] 1 is a flow diagram of a method according to an embodiment. [Figure 4] 1 is a flow diagram of a method according to an embodiment. [Figure 5] FIG. 2 is a diagram illustrating a spectrum of channel coherence values ​​according to an embodiment; [Figure 6] FIG. 2 is a diagram illustrating a spectrum of channel coherence values ​​according to an embodiment; [Figure 7] 1 is a flow diagram illustrating an encoding process according to some embodiments. [Figure 8] FIG. 1 illustrates a truncation scheme according to some embodiments. [Figure 9] 1 is a flow diagram illustrating a decoding process according to some embodiments. [Figure 10] 1 is a flow diagram illustrating a process according to one embodiment. [Figure 11] 1 is a flow diagram illustrating a process according to one embodiment. [Figure 12] FIG. 2 is a schematic diagram illustrating functional units of a sending node according to one embodiment. [Figure 13] FIG. 2 is a schematic diagram illustrating functional modules of a sending node according to one embodiment. [Figure 14] FIG. 1 illustrates an example of a computer program product comprising a computer-readable storage medium according to one embodiment. [Figure 15] FIG. 1 illustrates a stereo encoding and decoding system according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0022] The inventive concepts are described more fully below with reference to the accompanying drawings, in which certain embodiments of the inventive concepts are shown. However, the inventive concepts may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided as examples so that this disclosure will be thorough and complete, and will fully convey the scope of the inventive concepts to those skilled in the art. Like numbers refer to like elements throughout the specification. Any step or feature indicated by a dashed line should be considered optional.

[0023] Since spatial coherence describes the coherence between audio channels, it constitutes a spatial property of a multi-channel audio representation and may also be referred to as channel coherence. In the following description, the terms channel coherence and spatial coherence are used interchangeably.

[0024] When two mono encoders are used, each with its own DTX system operating separately on the signal in each of the two stereo channels, different energies and spectral shapes in the two different signals will be transmitted.

[0025] In most realistic cases, the differences in energy and spectral shape between the signals in the left and right channels will not be large, but there may still be large differences in how wide the stereo image of the signals is perceived.

[0026] If the random sequence used to generate the comfort noise is synchronized between the signal in the left channel and the signal in the right channel, the result is a stereo signal sounding with a very narrow stereo image and giving the sensation of sound coming from the center of the listener's head. If, instead, the signals in the left channel and the right channel are not synchronized, it will give the opposite effect, i.e., a signal with a very wide stereo image.

[0027] In most cases, the initial background noise will have a stereo image that lies somewhere between these two extremes, which means that there will be annoying differences in the stereo image when the transmitting device switches between active speech encoding and inactive noise encoding, which has a good representation of stereo width, along with synchronized or unsynchronized random sequences.

[0028] For example, the perceived stereo image width of the initial background noise may also change during a call due to the user of the transmitting device moving around and / or what is happening in the background. A system with two mono encoders, each with its own DTX system, has no mechanism for tracking these changes.

[0029] One additional problem with using a dual-mono DTX system is that, for example, when the signal in the left channel is encoded with active encoding and the signal in the right channel is encoded with low-bit-rate comfort noise encoding, the VAD decisions will become unsynchronized between the two channels, which can result in audible artifacts. Random sequences will be synchronized at some time instances and unsynchronized at others, which can lead to a stereo image that toggles between being extremely wide and extremely narrow over time.

[0030] Therefore, there remains a need for improvements in comfort noise generation for two or more channels.

[0031] The following embodiment describes a DTX system for two channels (stereo audio), but the method can be generally applied for DTX and CNG for multi-channel audio.

[0032] 1 is a schematic diagram illustrating a communication network 100 to which the embodiments presented herein may be applied. The communication network 100 comprises a sending node 200a in communication with a receiving node 200b via a communication link 110.

[0033] The transmitting node 200a may communicate with the receiving node 200b via a direct communication link 110 or via an indirect communication link 110 via one or more other devices, nodes, or entities in the communication network 100, such as network nodes.

[0034] In some aspects, the transmitting node 200a is part of a wireless transceiver device 200, and the receiving node 200b is part of another wireless transceiver device 200. Additionally, in some aspects, the wireless transceiver device 200 comprises both the transmitting node 200a and the receiving node 200b. Different examples of wireless transceiver devices may exist. Examples include, but are not limited to, portable wireless devices, mobile stations, mobile phones, handsets, wireless local loop telephones, user equipment (UE), smartphones, laptop computers, and tablet computers.

[0035] As mentioned above, a DTX system can be used to transmit encoded speech / audio only when necessary. Figure 2 is a schematic block diagram of a DTX system 300 for one or more audio channels. The DTX system 300 may be part of, collocated with, or implemented in a transmitting node 200a. Input audio is provided to a VAD 310, a speech / audio encoder 320, and a CNG encoder 330. When the VAD indicates that the signal contains speech or audio, the speech / audio encoder is activated, and when the VAD indicates that the signal contains background noise, the CNG encoder is activated. The VAD selectively controls whether to transmit the output from the speech / audio encoder or the CNG encoder accordingly. Problems with existing mechanisms for comfort noise generation for two or more channels were disclosed above.

[0036] Accordingly, embodiments disclosed herein relate to a mechanism for supporting comfort noise generation for at least two audio channels at a receiving node 200b and for the generation of comfort noise for at least two audio channels at a receiving node 200b. To obtain such a mechanism, a computer program product is provided that includes a transmitting node 200a, a method executed by the transmitting node 200a, and code, e.g., in the form of a computer program, that, when executed on the transmitting node 200a, causes the transmitting node 200a to perform the method. To obtain such a mechanism, a computer program product is further provided that includes a receiving node 200b, a method executed by the receiving node 200b, and code, e.g., in the form of a computer program, that, when executed on processing circuitry of the receiving node 200b, causes the receiving node 200b to perform the method.

[0037] 3 is a flow diagram illustrating an embodiment of a method for supporting comfort noise generation for at least two audio channels in a receiving node 200b. The method is performed by a transmitting node 200a. The method is advantageously provided as a computer program 1420.

[0038] S102: The sending node 200a determines the spectral characteristics of the audio signals of at least two input audio channels.

[0039] S104: The sending node 200a determines spatial coherence between the audio signals of each input audio channel, where the spatial coherence is related to a perceptual importance measure.

[0040] Since the whole rationale behind using the DTX system 300 is to transmit the minimum information needed in the pauses between speech / audio, the spatial coherence is coded in a very efficient way before transmission.

[0041] S106: The sending node 200a separates the spatial coherence into frequency bands, and a condensed representation of the spatial coherence is determined for each frequency band by weighting the spatial coherence values ​​in each frequency band according to a perceptual importance measure.

[0042] S108: The transmitting node 200a signals to the receiving node 200b information regarding the spectral characteristics and the compressed representation of the spatial coherence for each frequency band to enable generation of comfort noise for at least two audio channels at the receiving node 200b.

[0043] According to one embodiment, the perceptual importance measure is based on the spectral characteristics of the at least two input audio channels.

[0044] According to one embodiment, the perceptual importance measure is determined based on the power spectra of the at least two input audio channels.

[0045] According to one embodiment, the perceptual importance measure is determined based on the power spectrum of a weighted sum of at least two input audio channels.

[0046] According to one embodiment, the condensed representation of the spatial coherence is one single value per frequency band.

[0047] 4 is a flow diagram illustrating an embodiment of a method for supporting comfort noise generation for at least two audio channels in a receiving node 200b. The method is performed by a transmitting node 200a. The method is advantageously provided as a computer program 1420.

[0048] S202: The sending node 200a determines spectral characteristics of the audio signals of at least two input audio channels, the spectral characteristics being associated with perceptual importance measures.

[0049] S204: The sending node 200a determines the spatial coherence between the audio signals of each input audio channel. The spatial coherence is divided into frequency bands.

[0050] Since the whole rationale behind using the DTX system 300 is to transmit the minimum information needed in the pauses between speech / audio, the spatial coherence is coded in a very efficient way before transmission. Thus, one single value of spatial coherence is determined per frequency band.

[0051] The single value of spatial coherence is determined by weighting the spatial coherence values ​​within each frequency band. One purpose of the weighting function used for weighting is to place a higher weight on spatial coherence occurring at frequencies that are perceptually more important than others. Thus, the spatial coherence values ​​within each frequency band are weighted according to the perceptual importance measure of the corresponding value of the spectral feature.

[0052] S206: The transmitting node 200a signals to the receiving node 200b information regarding the spectral characteristics and the single value of spatial coherence for each frequency band to enable generation of comfort noise for at least two audio channels at the receiving node 200b.

[0053] At the decoder at the receiving node 200b, the coherence is reconstructed and a comfort noise signal is created that has a stereo image similar to the original sound.

[0054] Embodiments are now disclosed that relate to further details of supporting comfort noise generation for at least two audio channels at a receiving node 200b as performed by a transmitting node 200a.

[0055] The embodiments disclosed herein are applicable to stereo encoder and decoder architectures as well as for multi-channel encoders and decoders where channel coherence is considered across channel pairs.

[0056] In some aspects, a stereo encoder receives as input a channel pair [l(m,n)r(m,n)], where l(m,n) and r(m,n) denote the input signals for the left and right channels, respectively, at sample index n of frame m. The signals are sampled at a sampling frequency f s are processed in samples of frame length N, which may include overlap (look-ahead and / or memory of past samples).

[0057] As shown in Figure 2, when the stereo encoder VAD indicates that the signal contains background noise, the stereo CNG encoder is activated. The signal is transformed into the frequency domain, for example, using a discrete Fourier transform (DFT) or any other suitable filter bank or transform, such as a quadrature mirror filter (QMF), a hybrid QMF, or a modified discrete cosine transform (MDCT). When a DFT or MDCT transform is used, the input signal is divided into a channel pair [l] determined according to: win (m,n)r win (m,n)] is windowed before the transformation: [l win (m,n)r win (m,n)]=[l(m,n)win(n)r(m,n)win(n)],n=0,1,2,…,N-1.

[0058] Therefore, according to one embodiment, before the spectral characteristics are determined, the audio signals l(m,n), r(m,n) for frame index m and sample index n of at least two audio channels are windowed to obtain respective windowed signals l(m,n), r(m,n), win (m,n), r win (m,n). The choice of window may generally depend on various parameters, such as time and frequency resolution characteristics, algorithm delay (length of overlap), reconstruction characteristics, etc. Thus, the windowed channel pair [l win (m,n)r win (m,n)] is then transformed according to: TIFF0007746431000001.tif17170

[0059] Channel coherence C for frequency f gen The general provisions of (f) are given below: TIFF0007746431000002.tif16170So, Sxx (f) and S yy (f) represents the power spectrum of each of the two channels x and y, and S xy (f) is the cross power spectrum of the two channels x and y. In a DFT-based solution, the spectrum may be represented by the DFT spectrum. Specifically, according to one embodiment, the spatial coherence C(m,k) for frame index m and frequency bin index k is determined as follows: TIFF0007746431000003.tif14170 where L(m,k) is the windowed audio signal l win (m,n) and R(m,k) is the spectrum of the windowed audio signal r win is the spectrum of (m,n), and * denotes the complex conjugate.

[0060] The above expressions of coherence are generally calculated with high frequency resolution. One reason for this is that the frequency resolution depends on the signal frame size, which is usually the same for CNG encoding as for active speech / audio encoding, where high resolution is desirable. Another reason is that high frequency resolution allows for perceptually motivated frequency band division. Yet another reason is that the components of the coherence calculation, namely L(m,k), R(m,k), and S, xx , S xy , S yy , but in a typical audio encoder it may be used for other purposes where higher frequency resolution is desirable. s A typical value with a frequency of 48 kHz and a frame length of 20 ms would result in 960 frequency bins of channel coherence.

[0061] For DTX applications, where keeping the bit rate for encoding inactive (i.e., non-voice) segments low is crucial, transmitting channel coherence with high frequency resolution is not feasible. To reduce the number of bits required to represent channel coherence, the spectrum can be divided into frequency bands, as shown in Figure 5, and the channel coherence within each frequency band would be represented by a single value or some other compressed representation. The number of frequency bands is typically around 2-50 for the full audible bandwidth of 20-20,000 Hz.

[0062] Although all frequency bands may have similar widths, it is more common in audio coding applications to match the width of each frequency band to human perception of audio, resulting in relatively narrow frequency bands at low frequencies and increasing widths at higher frequencies. Specifically, according to one embodiment, spatial coherence is divided into frequency bands of unequal length. For example, frequency bands can be created using a scale of ERB rates, where ERBs are short relative to the width of an equivalent rectangular frequency band.

[0063] In one embodiment, the compressed representation of the coherence is defined by the average value of the coherence within each frequency band, and this single value per frequency band is transmitted to the decoder at the receiving node 200b so that the decoder can then use this single value for all frequencies within the frequency band when generating comfort noise, possibly with some smoothing of the signal frames and / or frequency bands to avoid abrupt changes in time and / or frequency.

[0064] However, as previously described in step S204, in another embodiment, different frequencies within a frequency band are given different weights according to a perceptual importance measure in determining a single coherence value for each frequency band.

[0065] There can be different examples of perceptual importance measures.

[0066] In some aspects, the perceptual importance measure is related to spectral characteristics.

[0067] Specifically, in one embodiment, the perceptual importance measure relates to the magnitude or power spectrum of the at least two input audio signals.

[0068] In another embodiment, the perceptual importance measure relates to the magnitude or power spectrum of a weighted sum of at least two input audio channels.

[0069] In some aspects, high energy corresponds to high perceptual importance, and vice versa. Specifically, according to one embodiment, the spatial coherence values ​​within each frequency band are weighted so that spatial coherence values ​​corresponding to frequency coefficients with higher power have more influence on this single value of spatial coherence compared to spatial coherence values ​​corresponding to frequency coefficients with lower energy.

[0070] According to one embodiment, different frequencies within a frequency band are given different weights depending on the power at each frequency. One rationale behind this embodiment is that a frequency with higher energy should have more influence on the combined coherence value compared to another frequency with lower energy.

[0071] In some other aspects, the perceptual importance measure is related to coded spectral characteristics, which may more closely (i.e., more closely than uncoded spectral characteristics) reflect the signal as reconstructed at the receiving node 200b.

[0072] In some other embodiments, the perceptual importance measure is related to spatial coherence. It may be perceptually more important to represent signal components with higher spatial coherence more accurately than signal components with lower spatial coherence. In another embodiment, the perceptual importance measure may be related to spatial coherence over time, including actively encoded speech / audio segments. One reason for this is that it may be perceptually important to generate spatial coherence of similar characteristics as in actively encoded speech / audio segments.

[0073] Other perceptual importance measures are also contemplated.

[0074] According to one embodiment, a weighted average is used to represent the coherence in each frequency band, where the transformed energy spectrum |LR(m,k)| for the mono signal lr(m,n) = w1l(m,n) + w2r(m,n) 2 defines a perceptual importance measure in frame m and is used as a weighting function. That is, in some aspects, the energy spectrum |LR(m,k)| of lr(m,n)=w1l(m,n)+w2r(m,n) 2 is used to weight the spatial coherence values. The downmix weights w1 and w2 may be constant or variable over time, or constant or variable across frequency if similar operations are performed in the frequency domain. In one embodiment, the channel weights are equal, e.g., w1 = w2 = 0.5. Then, according to one embodiment, each frequency band spans between a lower frequency bin and an upper frequency bin and is weighted by one single value C of spatial coherence for frame index m and frequency band b. w (m,b) is determined as follows: TIFF0007746431000004.tif18170where m is the frame index, b is the frequency band index, and N bandis the total number of frequency bands, and limit(b) indicates the lowest frequency bin of frequency band b. Thus, the parameter limit(b) indicates the first coefficient in each frequency band and defines the boundary between frequency bands. In this embodiment, limit(b) also indicates the upper limit N of the frequency bands. band To define -1, frequency band N band There can be different ways to obtain limit(b). According to one embodiment, limit(b) is provided as a function or a look-up table.

[0075] Figure 6 illustrates weighting in frequency band b+1. For each frequency bin, the dot with a vertical solid line indicates the coherence value, and the dot with a vertical dashed line indicates the energy of the corresponding value of the spectral characteristic. The horizontal dotted line indicates the average of the four coherence values ​​in frequency band b+1, and the dashed line indicates the weighted average. In this example, the third bin in frequency band b+1 has both a high coherence value and high energy, leading to a weighted average that is higher than the unweighted average.

[0076] If we assume that the energy is the same for all bins in a frequency band, then the weighted and unweighted averages will be equivalent. Furthermore, if we assume that the energy is zero for all bins in a frequency band except for one bin, then the weighted average will be equivalent to the coherence value of that one bin.

[0077] Spatial coherence value C w (m, b) is then encoded and stored or transmitted to a decoder at the receiving node 200b, where comfort noise is generated using the decoded coherence to create a realistic stereo image.

[0078] Encoding Spatial Coherence According to One Embodiment

[0079] The coherence representative value given for each frequency band is the spatial coherence vector TIFF0007746431000005.tif8170 is formed, where N bnd is the number of frequency bands, b is the frequency band index, and m is the frame index. m The value of C b,m is the weighted spatial coherence value C for frame m and band b w Corresponds to (m,b).

[0080] In one embodiment, the coherence vectors are coded using a prediction scheme followed by variable bit rate entropy coding. The coding scheme further improves performance through adaptive inter-frame prediction. The encoding of the coherence vectors takes into account the following attributes: (1) varying per-frame bit allocation B m (2) the coherence vectors exhibit strong frame-to-frame similarity, and (3) error propagation should be kept low for lost frames.

[0081] To cope with the varying frame-by-frame bit allocation, a coarse-fine encoding strategy is implemented. More specifically, coarse encoding is first achieved at a low bit rate, and subsequent fine encoding can be truncated when the bit limit is reached.

[0082] In some embodiments, the coarse encoding is performed using a prediction scheme. In such embodiments, a predictor works along the coherence vector of increasing band b, estimating each coefficient based on the previous value of the vector. That is, an intra-frame prediction of the coherence vector is performed, given by: TIFF0007746431000006.tif20170

[0083] For each predictor set P (q) is (N bnd−1) predictors, each containing (b−1) predictor coefficients for each band b, where q=1, 2, …N q and N q denotes the total number of predictor sets. As mentioned above, when b=1, there is no previous value and the intra-frame prediction of the coherence vector is zero. As an example, when there are six coherence bands, N bnd = 6, the number of predictor sets q is given by: TIFF0007746431000007.tif12170

[0084] As another example, the total number of predictor sets may be 4, i.e., N q = 4, which indicates that the selected predictor set may be signaled using 2 bits. In some embodiments, the predictor coefficients for predictor set q may be addressed consecutively, with length The image can be stored in a single vector: TIFF0007746431000008.tif13170.

[0085] 7 is a flow diagram illustrating an encoding process 701 according to some embodiments. The encoding process 701 may be performed by an encoder according to the following steps:

[0086] In step 700, for each frame m, a bit variable (also called a bit counter) for recording the bits used for encoding is initialized to zero (B curr,m= 0). The encoding algorithm uses the coherence vector (C b,m ) and the previous reconstructed coherence vector Copy of TIFF0007746431000009.tif8170, and bit allocation B m In some embodiments, the bits used in the encoding step are B m and B curr,mIn such an embodiment, the bit allocation in the algorithm described below may be m -B curr,m can be given by

[0087] In step 710, the available predictors p (q) , q=1,2,…,N q , the predictor set p that gives the minimum prediction error (q*) The set of predictors selected is Given by TIFF0007746431000010.tif15170.

[0088] In some embodiments, b=1 is omitted from the predictor set because the prediction is zero and the contribution to the error will be the same for all predictor sets. The selected predictor set index is stored and a bit counter (B curr,m ) is augmented by the required number of bits, e.g., if two bits are required to encode the predictor set, then B curr,m =B curr,m It becomes +2.

[0089] In step 720, a prediction weighting factor α is calculated. The prediction weighting factor α is used to generate a weighted prediction as described in step 760 below. The weighting factor α determines the bit allocation B available for encoding the vector of spatial coherence values ​​in each frame m. m The judgment is based on the following.

[0090] In general, the weighting factor α can range in value from 0 to 1, i.e., from using only information from the current frame (α=1) to using only information from the previous frame (α=0) and anywhere in between (0<α<1). Since a lower weighting factor α can make the encoding more susceptible to lost frames, in some aspects it is desirable to use as high a weighting factor α as possible. However, since a lower value of the weighting factor α generally results in fewer coded bits, the selection of the weighting factor α depends on the bit allocation B per frame m. m must be balanced.

[0091] The value of the weighting factor α used in the encoding must be known, at least implicitly, to the decoder at the receiving node 200b. That is, in one embodiment, information about the weighting factor α needs to be coded and transmitted to the decoder (as in step S1016). In other embodiments, the decoder can derive the predicted weighting factor based on other parameters already available at the decoder. Further aspects of how to provide information about the weighting factor α are disclosed below.

[0092] Bit allocation B for frame m for encoding spatial coherence m is known at the decoder at the receiving node 200b without explicit signaling from the transmitting node 200a. In this regard, the bit allocation B m The value of does not need to be explicitly signaled to receiving node 200b. Because the decoder at receiving node 200b knows how to interpret the bitstream, a side effect is that it also knows how many bits have been decoded. The remaining bits are found at the decoder at receiving node 200b simply by subtracting the decoded number of bits from the total bit budget (which is also known).

[0093] In some aspects, bit allocation B mBased on , a set of candidate weighting factors is selected, and trial encoding using a combined prediction and residual encoding scheme (without implementing a rate truncation strategy as described below) is performed on all these candidate weighting factors to find the total number of coded bits given the candidate weighting factors used. Specifically, according to one embodiment, the weighting factor α is determined by selecting a set of at least two candidate weighting factors and performing trial encoding of a vector of spatial coherence values ​​for each candidate weighting factor.

[0094] In some aspects, which candidate weighting factors to use during trial encoding is determined by the bit allocation B m In this respect, the candidate weight coefficients are based on the bit allocation B m or by performing a table lookup having bit allocation B m into the function. The table lookup may be performed with table values ​​obtained through training on a set of background noise.

[0095] The trial encoding of each candidate weighting factor results in a respective total number of coded bits for the vector of spatial coherence values. The weighting factor α is then determined so that the total number of coded bits for the candidate weighting factors is the bit allocation B m Specifically, according to one embodiment, the weighting factor α may be selected depending on whether the total number of coded bits falls within the bit allocation B m According to one embodiment, the total number of coded bits is selected as the largest candidate weighting factor that fits within the bit allocation B m If neither of these conditions is satisfied, then the weighting factor α is selected as the candidate weighting factor that results in the smallest total number of coded bits.

[0096] That is, all candidate weighting factors are used to determine whether the total number of coded bits is equal to the bit allocation B mIf the result is within , the highest candidate weighting factor is selected as the weighting factor α. Similarly, the lowest candidate weighting factor is selected as the bit allocation B m or any of the candidate weighting factors leads to the total number of bits in the bit allocation B m The candidate weighting factor that leads to the lowest number of bits is selected as the weighting factor α only if it does not lead to the total number of bits in. Which of the candidate weighting factors was selected is then signaled to the decoder.

[0097] the number of bits required for encoding the vector of spatial coherence values, respectively B currlow,m and B currhigh,m Two candidate weight coefficients α low and α high An example is now disclosed in which trial encoding is performed for .

[0098] B as input curr,m Using the two candidate weight coefficients α low and α high is the input bit allocation B m or by performing a table lookup using m The trial encoding is obtained by inputting two values ​​of the number of bits required for the encoding, B currlow,m and B currhigh,m Each candidate weight coefficient α low and α high Based on this, two candidate weighting factors α low and α high One of the is selected according to the encoding: TIFF0007746431000011.tif22170

[0099] The selected weighting factor α is coded using one bit, e.g., α low "0" and α for highThe third alternative in the above expression of the weighting factor α should be interpreted as follows: low and α high Both of them are bit allocation B m If the candidate weighting factors yield a resulting number of coded bits greater than 1, then the candidate weighting factor that yields the lowest number of coded bits is selected.

[0100] Band b=1, 2, .. N in step 730 bnd For each of the following steps are performed:

[0101] In step 740, the intra-frame prediction TIFF0007746431000012.tif9170 is obtained. There is no coded coherence value for the first band (b=1). In some embodiments, the intra-frame prediction for the first band may be set to zero. In some embodiments, the intra-frame prediction of the first band is TIFF0007746431000014.tif6170, TIFF0007746431000015.tif9170.

[0102] In some alternative embodiments, the coherence values ​​of the first band may be coded separately. In such embodiments, the first values ​​are coded using a scalar quantizer and the reconstructed values ​​are coded using a scalar quantizer. TIFF0007746431000016.tif8170. In response, intra prediction of the first band yields the reconstructed value, TIFF0007746431000017.tif9170, can be set to a bit counter, B curr,mis increased by the amount of bits required to encode the coefficient. For example, if 3 bits are used to encode the coefficient, then 3 bits are added to the current amount of bits used for encoding, e.g., B curr,m =B curr,m +3.

[0103] Remaining band b=2,3,…,N bnd Regarding intra-frame prediction TIFF0007746431000018.tif9170 is the previously coded coherence value, i.e. Based on TIFF0007746431000019.tif13170.

[0104] In step 750, the inter-frame prediction TIFF0007746431000020.tif8170, is obtained based on the coherence vector elements previously reconstructed from one or more previous frames. If the background noise is stable or slowly varying, the coherence band value C b,m The frame-to-frame variation in σ becomes small. Therefore, inter-frame prediction using values ​​from the previous frame is often a good approximation resulting in a small prediction residual and a small residual coding bit rate. As an example, the last reconstructed value of band b can be used for the inter-frame prediction, i.e., TIFF0007746431000021.tif8170. An inter-frame linear predictor that considers two or more preceding frames is It can be formulated as TIFF0007746431000022.tif13170, where TIFF0007746431000023.tif8170 shows a column vector of predicted inter-frame coherence values ​​for all bands b in frame m, TIFF0007746431000024.tif8170 represents the reconstructed coherence values ​​of all bands b in frame mn, and g n is N interare the linear predictor coefficients over the previous frame. g n can be selected from a predefined set of predictors, in which case the predictor used needs to be represented by an index that can be communicated to the decoder.

[0105] In step 760, the weighted prediction TIFF0007746431000025.tif9170, is intra-frame prediction, TIFF0007746431000026.tif9170, Inter-frame prediction, TIFF0007746431000027.tif7170, and a prediction weighting factor α. In some embodiments, the weighted prediction is Given by TIFF0007746431000028.tif9170.

[0106] At step 770, the prediction residual is calculated and coded. In some embodiments, the prediction residual is calculated using a coherence vector and a weighted prediction, i.e., TIFF0007746431000029.tif9170. In some embodiments, a scalar quantizer quantizes the prediction residual with index I b,m In such an embodiment, the index is used to quantize I b,m =SQ(r b,m ), where SQ(x) is a scalar quantizer function with an appropriate range. An example of a scalar quantizer is shown in Table 1 below. Table 1 shows an example of reconstruction levels and quantizer indices of prediction residuals. TIFF0007746431000030.tif37170

[0107] In some embodiments, the index I b,mis coded with a variable length codeword scheme that consumes fewer bits for smaller values. Some examples of coding the prediction residual are Huffman coding, Golomb-Rice coding, and unary coding (unary coding is the same as Golomb-Rice coding with a divisor of 1). In the step of coding the prediction residual, the remaining bit allocation (B m -B curr,m ) must be taken into account. b,m The length of the codeword corresponding to code (I b,m ) fits within the remaining bit budget, i.e., L code (I b,m )≦B m -B curr,m , if index I b,m is the final index I * b,m The remaining bits are selected as index I b,m If the number of bits required to encode the reconstructed value is insufficient, a bit-rate truncation strategy is applied. In some embodiments, the bit-rate truncation strategy involves encoding the largest possible residual value, assuming that smaller residual values ​​consume fewer bits. Such a rate truncation strategy can be implemented by permuting the codebook as shown by table 800 in FIG. 8. FIG. 8 shows an example quantizer table 800 with unary codeword mapping for the scalar quantizer example shown in Table 1. In some embodiments, bit-rate truncation can be implemented by proceeding two steps up table 800 until codeword 0 is reached. That is, FIG. 8 shows a truncation scheme that moves upward from long codewords to shorter codewords. To maintain the correct sign of the reconstructed value, each truncation step proceeds two steps up table 800, as indicated by the dashed and solid arrows for negative and positive values, respectively. By proceeding two steps up table 800, the new truncated codebook index TIFF0007746431000031.tif8170 can be found. Continue until TIFF0007746431000032.tif9170 is filled or the top of table 800 is reached.

[0108] If the length of the codeword determined by the upward search does not exceed the bit budget, a final index is selected. TIFF0007746431000033.tif8170,I * b,m is output to the bitstream and the reconstructed residual is formed based on the final index, i.e., TIFF0007746431000034.tif8170.

[0109] If, after searching upward, the codeword length still exceeds the bit budget, TIFF0007746431000035.tif9170, which means the bit limit has been reached. m =B curr,m In such cases, the reconstructed residual is set to zero. TIFF0007746431000036.tif8170, no index is added to the bitstream. The decoder uses a synchronized bit counter, B curr,m , the decoder can detect this situation and TIFF0007746431000037.tif7170 can be used.

[0110] In an alternative embodiment, if the length of the codeword associated with the initial index exceeds the bit budget, the residual value is immediately set to zero, thereby refraining from the above upward search. This can be beneficial when computational complexity is critical.

[0111] In step 780, the reconstructed coherence values TIFF0007746431000038.tif9170 is formed based on the reconstructed prediction residuals and weighted predictions, i.e., TIFF0007746431000039.tif9170.

[0112] The bit counter is incremented accordingly in step 790. As previously mentioned, the bit counter is incremented throughout the encoding process 701.

[0113] In some embodiments, the frame-to-frame variation in the coherence vector is small. Therefore, inter-frame prediction using previous frame values ​​is often a good approximation, resulting in a small prediction residual and a small residual coding bit rate. In addition, the prediction weighting factor α serves the purpose of balancing bit rate versus frame loss resilience.

[0114] 9 is a flow diagram illustrating a decoding process 901 according to some embodiments. The decoding process 901, which corresponds to the encoding process 701, may be performed by a decoder according to the following steps:

[0115] In step 900, a bit counter, B, is set to record the bits consumed during the decoding process 901. curr,m , is initialized to zero, i.e., B curr,m = 0. For each frame m, the decoder calculates the last reconstructed coherence vector TIFF0007746431000040.tif8170 and bit allocation B m Get a copy of the.

[0116] In step 910, the selected predictor set p (q*) is decoded from the bitstream. A bit counter is incremented by the amount of bits needed to decode the selected predictor set. For example, if two bits are needed to decode the selected predictor set, then the bit counter, Bcurr,m , is increased by 2, i.e., B curr,m =B curr,m +2.

[0117] In step 920, predicted weighting factors α, which correspond to the weighting factors used in the encoder, are derived.

[0118] In step 930, band b=1, 2..N bnd For each of the following steps are performed:

[0119] In step 940, the internal predicted value, TIFF0007746431000041.tif9170 is obtained. The intra prediction of the first band is obtained similarly to step 740 of the encoding process 701. Accordingly, the intra prediction of the first frame may be set to zero. TIFF0007746431000042.tif10170, average value of the first band TIFF0007746431000043.tif9170 or coherence value may be decoded from the bitstream, and the intra-frame prediction of the first frame may be set to the reconstructed value TIFF0007746431000044.tif9170. When the coefficient is decoded, the bit counter, B curr,m , is increased by the amount of bits required for encoding. For example, if three bits are required for encoding a coefficient, then the bit counter, B curr,m , is increased by 3, i.e., B curr,m =B curr,m +3.

[0120] Remaining band b=2,3,..N bnd Regarding intra-frame prediction TIFF0007746431000045.tif9170 is based on previously decoded coherence values, i.e. TIFF0007746431000046.tif13170.

[0121] In step 950, the inter-frame prediction TIFF0007746431000047.tif8170 is obtained similarly to step 750 of the encoding process 701. As an example, the last reconstructed value of band b may be used for inter-frame prediction, i.e. TIFF0007746431000048.tif9170.

[0122] In step 960, the weighted prediction TIFF0007746431000049.tif10170, is intra-frame prediction, TIFF0007746431000050.tif9170, Inter-frame prediction, TIFF0007746431000051.tif8170, and a prediction weighting factor α. In some embodiments, the weighted prediction is Given by TIFF0007746431000052.tif10170.

[0123] In step 970, the reconstructed prediction residuals, TIFF0007746431000053.tif7170 is decoded. Bit counter, B curr、m If , is less than the bit limit, i.e., B curr、m m , the reconstructed prediction residual is derived from the available quantizer indices TIFF0007746431000054.tif8170. If the bit counter is equal to or exceeds the bit limit, the reconstructed prediction residual is set to zero, i.e., TIFF0007746431000055.tif7170.

[0124] In step 980, the coherence value TIFF0007746431000056.tif8170 is reconstructed based on the reconstructed prediction residuals and weighted prediction, i.e., ​TIFF0007746431000057.tif9170. In step 990, the bit counter is incremented.

[0125] In some embodiments, further enhancements to the CNG may be required at the encoder. In such embodiments, the local decoder may use the reconstructed coherence values TIFF0007746431000058.tif8170 will be used and will be run in the encoder.

[0126] FIG. 10 is a flow diagram illustrating a process 1000, according to some embodiments, performed by the encoder of the transmitting node 200a to encode a vector. The process 1000 may begin at step S1002, in which the encoder forms prediction weighting factors. The following steps S1004 through S1014 may be repeated for each vector element. In step S1004, the encoder forms a first prediction of the vector element. In some embodiments, the first prediction is an intraframe prediction based on a current vector in a sequence of vectors. In such embodiments, the intraframe prediction is formed by performing a process including selecting a predictor from a set of predictors, applying the selected predictor to a reconstructed element of the current vector, and encoding an index corresponding to the selected predictor. In step S1006, the encoder forms a second prediction of the vector element. In some embodiments, the second prediction is an interframe prediction based on one or more previous vectors in the sequence of reconstructed vectors.

[0127] In step S1008, the encoder combines the second prediction and the first prediction using the prediction weighting factor into a joint prediction.

[0128] In step S1010, the encoder forms a prediction residual using vector elements and joint prediction. In step S1012, the encoder encodes the prediction residual using a variable bit rate scheme. In some embodiments, the prediction residual is quantized to form a first residual quantizer index, where the first residual quantizer index is associated with a first codeword. In some embodiments, encoding the prediction residual using a variable bit rate scheme includes encoding the first residual quantizer index as a result of determining that the length of the first codeword does not exceed the amount of remaining bits. In some embodiments, encoding the prediction residual using a variable bit rate scheme includes obtaining a second residual quantizer index as a result of determining that the length of the first codeword exceeds the amount of remaining bits, where the second residual quantizer index is associated with a second codeword, where the length of the second codeword is shorter than the length of the first codeword. In such an embodiment, process 600 includes the further step of the encoder determining whether the length of the second codeword exceeds the determined amount of remaining bits.

[0129] In step S1014, the encoder reconstructs the vector elements based on the joint prediction and the prediction residual. In step S1016, the encoder transmits the encoded prediction residual. In some embodiments, the encoder also encodes the prediction weighting coefficients and transmits the encoded prediction weighting coefficients.

[0130] In some embodiments, process 1000 includes further steps in which the encoder receives a first signal on a first input channel, receives a second signal on a second input channel, determines spectral characteristics of the first signal and the second signal, determines spatial coherence based on the determined spectral characteristics of the first signal and the second signal, and determines a vector based on the spatial coherence.

[0131] FIG. 11 is a flow diagram illustrating a process 1100, according to some embodiments, performed by a decoder of receiving node 200b to decode a vector. Process 1100 may begin at step 1102, where the decoder obtains prediction weight coefficients. In some embodiments, obtaining prediction weight coefficients includes (i) deriving prediction weight coefficients or (ii) receiving and decoding prediction weight coefficients. The following steps S1104 through S1112 may be repeated for each element of the vector. In step S1104, the decoder forms a first prediction of the vector element. In some embodiments, the first prediction is an intraframe prediction based on a current vector in the sequence of vectors. In such embodiments, the intraframe prediction is formed by performing a process including receiving and decoding a predictor and applying the decoded predictor to a reconstructed element of the current vector. In step S1106, the decoder forms a second prediction of the vector element. In some embodiments, the second prediction is an interframe prediction based on one or more previous vectors in the sequence of vectors.

[0132] In step S1108, the decoder combines the second prediction and the first prediction using the prediction weighting factor into a joint prediction.

[0133] In step S1110, a decoder decodes the received coded prediction residual. In some embodiments, decoding the coded prediction residual includes determining an amount of remaining bits available for decoding and determining whether decoding the coded prediction residual will exceed the amount of remaining bits. In some embodiments, decoding the coded prediction residual includes setting the prediction residual to zero as a result of determining that decoding the coded prediction residual will exceed the amount of remaining bits. In some embodiments, decoding the coded prediction residual includes deriving the prediction residual based on a prediction index as a result of determining that decoding the coded prediction residual will not exceed the amount of remaining bits, where the prediction index is a quantization of the prediction residual.

[0134] In step S1112, the decoder reconstructs vector elements based on the joint prediction and the prediction residual. In some embodiments, the vector is one of a set of vectors. In some embodiments, process 1100 further includes a step in which the decoder generates signals for at least two output channels based on the reconstructed vectors.

[0135] 12 illustrates, in terms of several functional units, components of the transmitting node 200a according to one embodiment. The processing circuitry 210 is implemented using any combination of one or more suitable central processing units (CPUs), multiprocessors, microcontrollers, digital signal processors (DSPs), etc., capable of executing software instructions stored in a computer program product 1410 (such as in FIG. 14), for example, in the form of a storage medium 230. The processing circuitry 210 may further be provided as at least one application specific integrated circuit (ASIC), or field programmable gate array (FPGA).

[0136] Specifically, processing circuitry 210 is configured to cause transmitting node 200a to perform a set of operations, or steps, as described above. For example, storage medium 230 may store a set of operations, and processing circuitry 210 may be configured to retrieve the set of operations from storage medium 230 and cause transmitting node 200a to perform the set of operations. The set of operations may be provided as a set of executable instructions. Thus, processing circuitry 210 is thereby arranged to perform the methods disclosed herein.

[0137] In one embodiment, a transmitting node 200a for supporting generation of comfort noise for at least two audio channels at a receiving node comprises a processing circuit 210. The processing circuit is configured to cause the transmitting node to determine spectral characteristics of audio signals of at least two input audio channels and to determine spatial coherence between the audio signals of each input audio channel, where the spatial coherence is associated with a perceptual importance measure. The transmitting node is further caused to separate the spatial coherence into frequency bands, where a compressed representation of the spatial coherence is determined for each frequency band by weighting the spatial coherence within each frequency band according to the perceptual importance measure. The transmitting node is further caused to signal information regarding the spectral characteristics and information regarding the compressed representation of the spatial coherence for each frequency band to the receiving node to enable generation of comfort noise for the at least two audio channels at the receiving node.

[0138] The transmitting node 200a may further be caused to encode the spatial coherence vector by forming a first prediction of the vector, a second prediction of the vector, a prediction weighting factor, and a prediction residual using the vector and the joint prediction. The transmitting node may further be caused to encode the prediction residual in a variable bit rate manner and reconstruct the vector based on the joint prediction and the prediction residual. The transmitting node may further be caused to transmit the encoded prediction weighting factor and the encoded prediction residual to the receiving node 200b.

[0139] The storage medium 230 may also comprise persistent storage, which may be, for example, any one or combination of magnetic memory, optical memory, solid-state memory, or even remotely attached memory. The sending node 200a may further comprise a communication interface 220 configured at least for communication with the receiving node 200b. As such, the communication interface 220 may comprise one or more transmitters and receivers comprising analog and digital components. The processing circuit 210 controls the general operation of the sending node 200a, for example, by sending data and control signals to the communication interface 220 and the storage medium 230, by receiving data and reports from the communication interface 220, and by retrieving data and instructions from the storage medium 230. Other components of the sending node 200a, and related functions, are omitted so as not to obscure the concepts presented herein.

[0140] 13 shows a schematic diagram of components of a transmitting node 200a according to one embodiment with respect to several functional modules. The transmitting node 200a of FIG. 13 comprises several functional modules: a determining module 210a configured to perform steps S102 and S202, a determining module 210b configured to perform steps S104 and S204, a segmenting module 210c configured to perform step S106, and a signaling module 210d configured to perform steps S108 and S206. The transmitting node 200a of FIG. 13 may further comprise several optional functional modules (not shown in FIG. 8). The transmitting node may comprise, for example, a first forming unit for forming a first prediction of the vector, a second forming unit for forming a second prediction of the vector, a third forming unit and encoding unit for forming and encoding prediction weight coefficients, a combining unit for combining the second prediction and the first prediction using the prediction weight coefficients into a joint prediction, a fourth forming unit for forming a prediction residual using the vector and the joint prediction, and an encoding unit 1014 for encoding the prediction residual using a variable bit rate scheme. The signal module 210d may be further configured to transmit the coded prediction weight coefficients and the coded prediction residual.

[0141] Generally, each functional module 210a-210d may be implemented solely in hardware in one embodiment and implemented using software in another embodiment, i.e., the latter embodiment has computer program instructions stored in storage medium 230 that, when executed on a processing circuit, cause transmitting node 200a to perform the corresponding steps described above in connection with FIG. 12. It should also be noted that, although the modules correspond to portions of a computer program, they need not be separate modules therein, but the manner in which they are implemented in software depends on the programming language used. Preferably, one or more or all of functional modules 210a-210d may be implemented by processing circuit 210, possibly in conjunction with communication interface 220 and / or storage medium 230. Thus, processing circuit 210 may be configured to form storage medium 230 fetch instructions as provided by functional modules 210a-210d and execute these instructions, thereby performing any steps as disclosed herein.

[0142] The transmitting node 200a may be provided as a stand-alone device or as part of at least one additional device. For example, as in the example of FIG. 1, in some aspects the transmitting node 200a is part of a wireless transceiver device 200. Thus, in some aspects, a wireless transceiver device 200 is provided comprising the transmitting node 200a as disclosed herein. In some aspects, the wireless transceiver device 200 further comprises a receiving node 200b.

[0143] Alternatively, the functionality of the transmitting node 200a may be distributed among at least two devices or nodes. These at least two nodes or devices may be part of the same network portion or may span at least two such network portions. Thus, a first portion of the instructions executed by the transmitting node 200a may be executed on a first device, and a second portion of the instructions executed by the transmitting node 200a may be executed on a second device; the embodiments disclosed herein are not limited to any particular number of devices on which the instructions executed by the transmitting node 200a may be executed. Thus, methods according to embodiments disclosed herein are suitable for execution by a transmitting node 200a residing in a cloud computing environment. Thus, although a single processing circuit 210 is shown in FIG. 12, the processing circuit 210 may be distributed among multiple devices or nodes. The same applies to the functional modules 210a-210d of FIG. 13 and the computer program 1420 of FIG. 14 (see below).

[0144] The receiving node 200b includes a decoder for reconstructing the coherence and for creating a comfort noise signal having a stereo image similar to the original sound. The decoder may be further configured to form a first prediction of the vector and a second prediction of the vector and to obtain a prediction weighting factor. The decoder may be further configured to combine the second prediction and the first prediction using the prediction weighting factor into a joint prediction. The decoder may be further configured to reconstruct the vector based on the joint prediction and the received and decoded prediction residual.

[0145] 14 illustrates one example of a computer program product 1410 comprising a computer-readable storage medium 1430. The computer-readable storage medium 1430 may have stored thereon a computer program 1420 that can cause the processing circuitry 210 and entities and devices operatively coupled thereto, such as the communications interface 220 and the storage medium 230, to perform methods according to embodiments described herein. The computer program 1420 and / or the computer program product 1410 may thereby provide means for performing any steps as disclosed herein.

[0146] 14, computer program product 1410 is shown as an optical disc, e.g., a CD (compact disc) or DVD (digital versatile disc) or Blu-ray disc. Computer program product 1410 may also be embodied as memory, e.g., random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or electrically erasable programmable read-only memory (EEPROM), and more particularly as the non-volatile storage medium of the device in an external memory such as a Universal Serial Bus (USB) memory or flash memory, e.g., compact flash memory. Thus, although computer program 1420 is shown here schematically as a track on the illustrated optical disc, computer program 1420 may be stored in any manner suitable for computer program product 1410.

[0147] The proposed solution disclosed herein applies to stereo encoder and decoder architectures or multi-channel encoders and decoders, where channel coherence is considered in channel pairs.

[0148] FIG. 15 illustrates a parametric stereo encoding and decoding system 1500 according to some embodiments. The parametric stereo encoding and decoding system 1500 comprises a mono encoder 1503 including a CNG encoder 1504 and a mono decoder 1505 including a CNG decoder 1506. The encoder 1501 performs an analysis of input channel pairs 1507A-1507B, obtains a parametric representation of a stereo image via parametric analysis 1508, and reduces the channels to a single channel via downmixing 1509, thereby obtaining a downmixed signal. The downmixed signal is encoded using a mono encoding algorithm by the mono encoder 1503, and the parametric representation of the stereo image is encoded by a parameter encoder 1510. The encoded downmixed signal and the parametric representation of the stereo image are transmitted via a bitstream 1511. The decoder 1502 applies a mono decoding algorithm using the mono decoder 1505 to obtain a combined downmixed signal. The parameter decoder 1512 decodes the received parametric representation of the stereo image. The decoder 1502 converts the synthesized downmix signal into a synthesized channel pair using the decoded parametric representation of the stereo image. The parametric stereo encoding and decoding system 1500 further includes a coherence analysis 1513 within the parametric analysis 1508 and a coherence synthesis 1514 within the parameter synthesis 1515. The parametric analysis 1508 includes capabilities for analyzing the coherence of the input signals 1507A-1507B. When the mono encoder 1503 is configured to operate as the CNG encoder 1504, the parametric analysis 1508 can analyze the input signals 1507A-1507B. The mono encoder 1503 may further include a stereo encoder VAD according to some embodiments. The stereo encoder VAD can indicate to the CNG encoder 1504 that the signal contains background noise, thereby activating the CNG encoder 1504.In response, CNG analysis, including coherence analysis 1513, is initiated in parametric analysis 1508, and mono encoder 1503 initiates CNG encoder 1504. As a result, coded representations of coherence and mono CNG are compiled into bitstream 1511 for transmission and / or storage. Decoder 1502 identifies stereo CNG frames within bitstream 1511, decodes the mono CNG and coherence values, and synthesizes the target coherence. When decoding the CNG frames, decoder 1502 produces two CNG frames corresponding to two synthesized channels 1517A-1517B.

[0149] A set of exemplary embodiments follows to further illustrate the concepts presented herein.

[0150] 1. A method performed by a transmitting node for supporting comfort noise generation for at least two audio channels at a receiving node, comprising: determining spectral characteristics of an audio signal of at least two input audio channels; determining spatial coherence between audio signals of each input audio channel, the spatial coherence being related to a perceptual importance measure; dividing the spatial coherence into frequency bands, wherein a compressed representation of the spatial coherence is determined for each frequency band by weighting the spatial coherence values ​​within each frequency band according to a perceptual importance measure; and signaling to a receiving node information regarding spectral characteristics and information regarding a compressed representation of spatial coherence for each frequency band to enable generation of comfort noise for at least two audio channels at the receiving node.

[0151] 2. The method according to item 1, wherein the perceptual importance measure is based on the spectral characteristics of at least two input audio channels.

[0152] 3. The method according to item 2, wherein the perceptual importance measure is determined based on the power spectra of at least two input audio channels.

[0153] 4. The method according to item 2, wherein the perceptual importance measure is determined based on the power spectrum of a weighted sum of at least two input audio channels.

[0154] 5. The method of item 1, wherein the compressed representation of spatial coherence is one single value per frequency band.

[0155] 6. A method performed by a transmitting node for supporting comfort noise generation for at least two audio channels at a receiving node, comprising: determining spectral characteristics of audio signals of at least two input audio channels, the spectral characteristics being related to a perceptual importance measure; determining spatial coherence between audio signals of each input audio channel, the spatial coherence being divided into frequency bands and one single value of spatial coherence being determined for each frequency band by weighting the spatial coherence values ​​in each frequency band according to a perceptual importance measure of the corresponding value of the spectral characteristic; and signaling to a receiving node information regarding spectral characteristics and information regarding a single value of spatial coherence per frequency band to enable generation of comfort noise for at least two audio channels at the receiving node.

[0156] 7. The method according to item 1 or 6, wherein the perceptual importance measure of a given value of a spectral characteristic is defined by the power of the sum of the audio signals of at least two input audio channels.

[0157] 8. The method according to item 1 or 6, wherein the spatial coherence values ​​within each frequency band are weighted such that spatial coherence values ​​corresponding to values ​​of spectral features with higher energy have a greater influence on the single value of spatial coherence compared to spatial coherence values ​​corresponding to values ​​of spectral features with lower energy.

[0158] 9. For at least two audio channels, audio signals l(m,n), r(m,n) for frame index m and sample index n are filtered to obtain respective windowed signals l(m,n), ... win (m,n), r win 7. The method according to item 1 or 6, wherein the windowing is performed to form (m,n).

[0159] 10. The method of item 9, wherein spatial coherence C(m,k) for frame index m and sample index k is determined as follows: TIFF0007746431000059.tif13170 where L(m,k) is the windowed audio signal l win (m,n) and R(m,k) is the spectrum of the windowed audio signal r win is the spectrum of (m,n), and * denotes the complex conjugate.

[0160] 11. Energy spectrum of lr(m,n) = l(m,n) + r(m,n) |LR(m,k)| 2 Item 11. The method of item 10, wherein σ defines a perceptual importance measure in frame m and is used to weight the spatial coherence values.

[0161] 12. Each frequency band extends between a lower edge and an upper edge, and the one single value of spatial coherence for frame index m and frequency band b is C w Denoted by (m,b), it is determined as follows: TIFF0007746431000060.tif19170So, N bandItem 12. The method according to item 11, wherein limit(b) indicates the total number of frequency bands, and limit(b) indicates the lower frequency bin of frequency band b.

[0162] 13. The method according to item 12, wherein limit(b) is given as a function or a look-up table.

[0163] 14. The method according to item 1 or 6, wherein the spatial coherence is divided into frequency bands of unequal length.

[0164] 15. A transmitting node for supporting comfort noise generation for at least two audio channels at a receiving node, the transmitting node comprising: a processing circuit; determining spectral characteristics of an audio signal of at least two input audio channels; determining spatial coherence between audio signals of each input audio channel, the spatial coherence being related to a perceptual importance measure; dividing the spatial coherence into frequency bands, wherein a condensed representation of the spatial coherence is determined for each frequency band by weighting the spatial coherence values ​​within each frequency band according to a perceptual importance measure; signaling to a receiving node information about spectral characteristics and compressed representations of spatial coherence for each frequency band to enable generation of comfort noise for at least two audio channels at the receiving node; a transmitting node configured to cause the transmitting node to perform the following:

[0165] 16. A transmitting node according to item 15, further configured to perform the method according to any one of items 2 to 5.

[0166] 17. A method for supporting comfort noise generation for at least two audio channels at a receiving node, comprising: processing circuitry; determining spectral characteristics of audio signals of at least two input audio channels, the spectral characteristics being related to a perceptual importance measure; determining spatial coherence between audio signals of each input audio channel, the spatial coherence being divided into frequency bands and one single value of spatial coherence being determined for each frequency band by weighting the spatial coherence values ​​in each frequency band according to a perceptual importance measure of the corresponding value of the spectral characteristic; signaling to a receiving node information about spectral characteristics and a single value of spatial coherence per frequency band to enable generation of comfort noise for at least two audio channels at the receiving node; a transmitting node configured to cause the transmitting node to perform the following:

[0167] 18. A transmitting node according to item 17, further configured to perform the method according to any one of items 7 to 14.

[0168] 19. A wireless transceiver device comprising a transmitting node according to any one of items 15 to 18.

[0169] 20. The wireless transceiver device of item 19, further comprising a receiving node.

[0170] 21. A computer program for supporting comfort noise generation for at least two audio channels at a receiving node, the computer code comprising: computer code that, when executed in processing circuitry at a transmitting node, performs the following: determining spectral characteristics of an audio signal of at least two input audio channels; determining spatial coherence between audio signals of each input audio channel, the spatial coherence being related to a perceptual importance measure; dividing the spatial coherence into frequency bands, wherein a condensed representation of the spatial coherence is determined for each frequency band by weighting the spatial coherence values ​​within each frequency band according to a perceptual importance measure; signaling to a receiving node information about spectral characteristics and compressed representations of spatial coherence for each frequency band to enable generation of comfort noise for at least two audio channels at the receiving node; A computer program that causes a sending node to perform the above.

[0171] 22. A computer program for supporting comfort noise generation for at least two audio channels at a receiving node, comprising computer code, the computer code, when executed in processing circuitry at a transmitting node, determining spectral characteristics of audio signals of at least two input audio channels, the spectral characteristics being related to a perceptual importance measure; determining spatial coherence between audio signals of each input audio channel, the spatial coherence being divided into frequency bands and one single value of spatial coherence being determined for each frequency band by weighting the spatial coherence values ​​in each frequency band according to a perceptual importance measure of the corresponding value of the spectral characteristic; signaling to a receiving node information about spectral characteristics and a single value of spatial coherence per frequency band to enable generation of comfort noise for at least two audio channels at the receiving node; A computer program that causes a sending node to perform the above.

[0172] 23. A computer program product comprising a computer program according to at least one of items 21 and 22, and a computer-readable storage medium on which the computer program is stored.

[0173] In general, all terms used in the exemplary embodiments and the appended claims shall be interpreted according to their ordinary meaning in the art unless expressly stated otherwise. Unless expressly stated otherwise, all references to "an / one / the element, apparatus, component, means, module, step, etc." shall be interpreted as referring to at least one instance of the element, apparatus, component, means, module, step, etc. The steps of any method disclosed herein do not have to be performed in the exact order disclosed, unless explicitly stated otherwise.

[0174] The inventive concept has been described above primarily with reference to certain embodiments. However, those skilled in the art will readily recognize that embodiments other than the above-disclosed embodiments are equally possible within the scope of the inventive concept as defined by the accompanying list of enumerated embodiments.

Claims

1. 1. An audio signal processing method for encoding a vector, comprising the steps of: determining spatial coherence between audio signals of each audio channel, wherein at least one spatial coherence value C b,m for each frame m and frequency band b is determined to form a vector of predicted spatial coherence values; forming weighting factors based on the bit allocation B m available for encoding said vector of spatial coherence values ​​in each frame m; For each element of the vector, forming an intraframe prediction of the elements of said vector; forming an inter-frame prediction of the elements of said vector; combining the intra-frame prediction and the inter-frame prediction into a joint prediction using the weighting factors; forming a prediction residual using the elements of the vector and the joint prediction; encoding the prediction residuals in a variable bit rate manner; transmitting the encoded prediction residual; A method comprising:

2. The method of claim 1 , wherein the vector is one of a sequence of vectors.

3. The method of claim 1 or 2, further comprising reconstructing elements of the vector based on the joint prediction and a reconstructed prediction residual.

4. The method of claim 3 , wherein the intra-frame prediction is based on elements of the reconstructed vector.

5. The method of claim 2 , wherein the inter-frame prediction is based on one or more previously reconstructed vectors for the sequence of vectors.

6. selecting a predictor from a set of predictors; applying the selected predictor to elements of the reconstructed vector; encoding an index corresponding to the selected predictor; The method of claim 4 , wherein the intra-frame prediction is formed by performing a process including:

7. The method of claim 5 , wherein for the inter-frame prediction, values ​​from a previously reconstructed vector are used.

8. selecting a predictor from a set of predictors; applying the selected predictor to the one or more previously reconstructed vectors; encoding an index corresponding to the selected predictor; The method of claim 5 , wherein the inter-frame prediction is formed by performing a process including:

9. 1. An audio signal processing method for decoding vectors, comprising the steps of: determining spatial coherence between audio signals of each audio channel, wherein at least one spatial coherence value C b,m for each frame m and frequency band b is determined to form a vector of predicted spatial coherence values; obtaining weighting factors based on the bit allocation B m available for encoding said vector of spatial coherence values ​​in each frame m; For each element of the vector, forming an intraframe prediction of the elements of said vector; forming an inter-frame prediction of the elements of said vector; combining the intra-frame prediction and the inter-frame prediction into a joint prediction using the weighting factors; decoding the received coded prediction residual; and reconstructing elements of the vector based on the joint prediction and the decoded prediction residual.

10. The method of claim 9 , wherein the vector is one of a sequence of vectors.

11. The method according to claim 9 or 10, wherein the intra-frame prediction is based on elements of the reconstructed vector.

12. The method of claim 10 or 11, wherein the inter-frame prediction is based on one or more previously reconstructed vectors of the sequence of vectors.

13. receiving and decoding a predictor; applying the decoded predictor to elements of the reconstructed vector; The method of claim 11 , wherein the intra-frame prediction is formed by performing a process including:

14. The method of claim 12 , wherein for the inter-frame prediction, values ​​from a previously reconstructed vector are used.

15. receiving and decoding a predictor; applying the decoded predictor to the one or more previously reconstructed vectors; The method of claim 12 , wherein the inter-frame prediction is formed by performing a process including:

16. 16. The method of claim 9, wherein obtaining the weighting factors comprises one of: (i) deriving the weighting factors; and (ii) receiving and decoding the weighting factors.

17. 17. The method of claim 9, further comprising generating signals for at least two output channels based on the reconstructed vectors.

18. An encoder configured to perform the method of any one of claims 1 to 8.

19. A decoder configured to perform the method of any one of claims 9 to 17.

Citation Information

Patent Citations

  • Voice spectrum encoding system

    JP1989084300A

  • Perceptual noise replacement

    JP2005509926A

  • Comfort noise generation

    US20170047072A1

  • Ambisonic audio signal processing for bidirectional real-time communication

    US9865274B1