ADAPTIVE CHANNEL REDUCTION PROCESSING FOR CODING A MULTI-CHANNEL AUDIO SIGNAL

DE602016094146T2Active Publication Date: 2025-11-19ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602016094146
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-12-16
Filing Date
2016-12-13
Publication Date
2025-11-19
Estimated Expiration
2036-12-13

AI Technical Summary

Technical Problem

Existing parametric stereo encoding methods face challenges in preserving signal energy and phase alignment during stereo-to-mono conversion, particularly when channels are out of phase, leading to quality issues in the decoded mono signal.

Method used

A channel reduction processing method that includes extracting indicators of interchannel correlation (ICCr) and degree of phase opposition (ISD) to select appropriate downmix processing modes, such as passive, gain-compensated, or phase-aligned downmix, to ensure robust quality in the mono signal.

Benefits of technology

The method effectively maintains signal energy and phase alignment across various stereo signals, improving the quality of the decoded mono signal by adaptively handling phase relationships and channel correlations.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to the field of encoding / decoding digital signals.

[0002] The coding and decoding according to the invention is particularly suitable for the transmission and / or storage of digital signals such as audio frequency signals (speech, music or others).

[0003] More particularly, the present invention relates to parametric coding or processing of multichannel audio signals, for example stereophonic signals hereinafter referred to as stereo signals.

[0004] This type of coding is based on the extraction of spatial information parameters so that, upon decoding, these spatial characteristics can be reconstructed for the listener, in order to recreate the same spatial image as in the original signal.

[0005] Such a parametric encoding / decoding technique is described, for example, in the paper by J. Breebaart, S. van de Par, A. Kohlrausch, and E. Schuijers, entitled "Parametric Coding of Stereo Audio" in EURASIP Journal on Applied Signal Processing 2005:9, pp. 1305-1322. This example is repeated with reference to figures 1 And 2 describing respectively a parametric stereo encoder and decoder.

[0006] Document US 2015 / 030182 A1 discloses stereo encoding and modifies downmix modes between the summation of channels with and without gain.

[0007] Thus, the figure 1 describes a stereo encoder receiving two audio channels, a left channel (noted L for Left in English) and a right channel (noted R for Right in English).

[0008] Time signals L(n) and R(n), where n is the integer index of the samples, are processed by blocks 101, 102, 103 and 104 which perform a short-term Fourier analysis.

[0009] The transformed signals L [ k ] et R [k], Or k is the integer index of the frequency coefficients, are thus obtained.

[0010] Block 105 performs channel reduction or "downmix" processing to obtain, in the frequency domain, from the left and right signals, a monophonic signal, hereafter referred to as a mono signal.

[0011] An extraction of spatial information parameters is also carried out in block 105. The extracted parameters are as follows.

[0012] The ICLD parameters (for "InterChannel Level Difference" (in English), also called interchannel intensity differences, characterize the energy ratios per frequency sub-band between the left and right channels. These parameters allow for the positioning of sound sources in the horizontal stereo plane by "panning". They are defined in dB by the following formula: ICLD b = 10 . log 10 ∑ k = k b k b + 1 − 1 L k . L * k ∑ k = k b k b + 1 − 1 R k . R * k dB Or L [ k ] AndR [ k ] correspond to the (complex) spectral coefficients of the L and R channels, each frequency band with index b includes the frequency lines in the interval [ k b , k b +1 - 1] and the symbol * indicates the complex conjugate.

[0013] The ICPD parameters (for "InterChannel Phase Difference" (in English), also called phase differences, are defined according to the following relationship: ICPD b = ∠ ∑ k = k b k b + 1 − 1 L k . R * k where ∠ indicates the argument (the phase) of the complex operand.

[0014] An interchannel time delay called ICTD (for "InterChannel Time Difference" (in English) and whose known definition of the man of art is not recalled here.

[0015] Unlike ICLD, ICPD, and ICTD parameters, which are localization parameters, ICC parameters (for "InterChannel Coherence" (in English) represent the inter-channel correlation (or coherence) and are associated with the spatial width of the sound sources; their definition is not recalled here, but it is noted in the article by Breebart et al. that the ICC parameters are not necessary in sub-bands reduced to a single frequency coefficient - indeed the differences in amplitude and phase completely describe the spatialization in this "degenerate" case.

[0016] These ICLD, ICPD, and ICC parameters are extracted by analyzing the stereo signals, using block 105. If the ICTD or ITD parameters were also encoded, these could also be extracted by sub-band from the spectra. L [ k ] And R [ k However, the extraction of ITD parameters is generally simplified by assuming an identical inter-channel time offset for each sub-band, and in this case, a parameter can be extracted from the time channels.L ( n ) And R ( n ) through inter-correlations.

[0017] The mono signal M[k] is transformed in the time domain (blocks 106 to 108) after short-term Fourier synthesis (inverse FFT, windowing, and OverLap-Add or OLA), and mono coding (block 109) is then performed. In parallel, the stereo parameters are quantized and encoded in block 110.

[0018] In general, the spectrum of signals ( L [ k ] , R [ k ]) is divided according to a non-linear frequency scale of the ERB type (Equivalent Rectangudar Bandwidth) or Bark, with a number of sub-bands typically ranging from 20 to 34 for a signal sampled from 16 to 48 kHz according to the Bark scale. This scale defines the values ​​of k b And k b+1 for each sub-band b. The parameters (ICLD, ICPD, ICC, ITD) are coded by scalar quantization, possibly followed by entropy coding and / or differential coding. For example, in the previously cited article, the ICLD is coded by a non-uniform quantizer (ranging from -50 to +50 dB) with differential entropy coding. The non-uniform quantization step exploits the fact that the larger the ICLD value, the lower the auditory sensitivity to variations in this parameter.

[0019] For the coding of the mono signal (block 109), several quantization techniques with or without memory are possible, for example "Pulse Code Modulation" (PCM) coding, its version with adaptive prediction called "Adaptive Differential Pulse Code Modulation" (ADPM) or more advanced techniques such as perceptual transform coding or "Code Excited Linear Prediction" (CELP) coding or multi-mode coding.

[0020] This section focuses specifically on the 3GPP EVS recommendation (for "Enhanced Voice Services"), which uses multi-mode coding. The algorithmic details of the EVS codec are provided in the 3GPP TS specifications 26.441 to 26.451 and are therefore not repeated here. These specifications will subsequently be referred to as EVS.

[0021] The input signal of the EVS codec is sampled at a frequency of 8, 16, 32, or 48 kHz, and the codec can represent narrowband (NB), wideband (WB), super-wideband (SWB), or fullband (FB) telephone audio bands. EVS codec bitrates are divided into two modes: o "EVS Primary": o Fixed bitrates: 7.2, 8, 9.6, 13.2, 16.4, 24.4, 32, 48, 64, 96, 128 o Variable bitrate (VBR) mode with an average bitrate close to 5.9 kbit / s for active speech o Channel-aware mode at 13.2 in WB and SWB only o "EVS AMR-WB IO" whose bitrates are identical to the 3GPP AMR-WB codec (9 modes)

[0022] In addition, there is the discontinuous transmission mode (DTX) in which frames detected as inactive are replaced by SID frames (SID Primary or SID AMR-WB IO) which are transmitted intermittently, approximately once every 8 frames.

[0023] On decoder 200, in reference to the figure 2 , The mono signal is decoded (block 201), a decorrelator is used (block 202) to produce two versions M̂ ( n ) And M̂' ( n) of the decoded mono signal. This decorrelation, necessary only when the ICC parameter is used, allows for an increase in the spatial width of the mono source. M̂ ( n ) . These two signals M̂ ( n ) And M̂ '( n The parameters are converted to the frequency domain (blocks 203 to 206), and the decoded stereo parameters (block 207) are used by stereo synthesis (or shaping) (block 208) to reconstruct the left and right channels in the frequency domain. These channels are then reconstructed in the time domain (blocks 209 to 214).

[0024] Thus, as mentioned for the encoder, block 105 performs channel reduction or "downmixing" by combining the stereo channels (left, right) to obtain a mono signal, which is then encoded by a mono encoder. The spatial parameters (ICLD, ICPD, ICC, etc.) are extracted from the stereo channels and transmitted in addition to the binary stream from the mono encoder.

[0025] Several techniques have been developed for stereo-to-mono channel reduction or "downmixing." This downmixing can be performed in the time or frequency domain. Generally, two types of downmixing are distinguished: Passive "downmix" which corresponds to a direct matrixing of the stereo channels to combine them into a single signal - the coefficients of the downmix matrix are generally real and of predetermined (fixed) values; Active (adaptive) "downmix" which includes control of energy and / or phase in addition to the combination of the two stereo channels.

[0026] The simplest example of passive "downmixing" is given by the following time matrix: M n = 1 2 L n + R n = 1 / 2 0 0 1 / 2 L n R n

[0027] This type of "downmix," however, has the disadvantage of not preserving signal energy well after stereo-to-mono conversion when the L and R channels are not in phase: in the extreme case where L(n) = -R(n), the mono signal is zero, which is undesirable.

[0028] An active "downmix" mechanism that improves the situation is given by the following equation: M n = γ n L n + R n 2 Or γ(n) is a factor that compensates for any potential energy loss.

[0029] However, combining the signals L(n) And R(n) in the time domain does not allow fine control (with sufficient frequency resolution) of possible phase differences between L and R channels; when the L and R channels have comparable amplitudes and almost opposite phases, phenomena of "erasure" or "attenuation" (loss of "energy") on the mono signal can be observed by frequency sub-bands with respect to the stereo channels.

[0030] This is why it is often more advantageous in terms of quality to perform the "downmix" in the frequency domain, even if this involves calculating time / frequency transforms and induces additional delay and complexity compared to a time-domain "downmix".

[0031] We can thus transpose the previous active "downmix" with the spectra of the left and right channels, as follows: M k = γ k L k + R k 2 Or kcorresponds to the index of a frequency coefficient (Fourier coefficient, for example, representing a frequency sub-band). The compensation parameter can be set as follows: γ k = max 2 L k 2 + R k 2 L k + R k 2 / 2

[0032] This ensures that the overall energy of the downmix is ​​the sum of the energies of the left and right channels. The y[k] factor is saturated here at a 6dB gain.

[0033] The stereo-to-mono downmixing technique described in the previously cited document by Breebaart et al. is performed in the frequency domain. The mono signal M[k] is obtained by a linear combination of the L and R channels according to the equation: M k = w 1 L k + w 2 R k Or w 1 , w 2 are complex-valued gains. If w 1 = w 2 = 0.5, the mono signal is considered an average of the two channels L and R. The gains w 1 , w 2 are generally adapted according to the short-term signal, in particular to align the phases.

[0034] A particular case of this frequency downmixing technique is proposed in the document entitled "A stereo to mono downmixing scheme for MPEG-4 parametric stereo encoder" by Samsudin, E. Kurniawati, N. Boon Poh, F. Sattar, S. George, in Proc. ICASSP, 2006. In this document, the L and R channels are phase-aligned before performing the channel reduction processing.

[0035] More specifically, the phase of the L channel for each frequency sub-band is chosen as the reference phase, and the R channel is aligned according to the phase of the L channel for each sub-band using the following formula: R ′ k = e j . ICPD b R k Or j = − 1 , R ′ k is the R-channel aligned, k is the index of a coefficient in the b nth frequency sub-band, ICPD [ b ] is the inter-channel phase difference in the b i<th frequency sub-band given by to equation (1).

[0036] Note that when the sub-index band b is reduced to a frequency coefficient, we find: R ′ k = R k . e j ∠ L k

[0037] Finally, the mono signal obtained by the "downmix" of the document by Samsudin et al. cited previously is calculated by averaging the L channel and the aligned R' channel, according to the following equation: M k = L k + R ′ k 2

[0038] Phase alignment therefore conserves energy and avoids attenuation problems by eliminating the influence of phase. This "downmix" corresponds to the "downmix" described in the document by Breebart et al., where: M k = w 1 L k + w 2 R k with w 1 = 0 , 5 et w 2 = e j . ICPD b 2 in the case where the sub-band with index b contains only one frequency value with index k .

[0039] An ideal conversion from a stereo signal to a mono signal should avoid attenuation problems for all frequency components of the signal.

[0040] This "downmix" operation is important for parametric stereo encoding because the decoded stereo signal is just a spatial shaping of the decoded mono signal.

[0041] The frequency-domain downmixing technique described earlier effectively preserves the energy level of the stereo signal in the mono signal by aligning the right and left channels before processing. This phase alignment prevents situations where the channels are out of phase.

[0042] The method described in the Samsudin document referenced above, however, relies on a total dependence of the "downmix" processing on the channel (L or R) chosen to fix the reference phase.

[0043] In extreme cases, if the reference channel is zero (total silence) and the other channel is non-zero, the phase of the mono signal after downmixing becomes constant, and the resulting mono signal will generally be of poor quality. Similarly, if the reference channel is a random signal (ambient noise, etc.), the phase of the mono signal can become random or be poorly conditioned, again resulting in a mono signal that will generally be of poor quality. An alternative frequency downmixing technique was proposed in the paper entitled "Parametric stereo extension of ITU-T G.722 based on a new downmixing scheme" by TMN Hoang, S. Ragot, B. Kovësi, and P. Scalart, Proc. IEEE MMSP, 4-6 Oct. 2010. This paper proposes a downmixing technique that resolves some of the drawbacks of the downmix proposed by Samsudin et al. According to this paper, the mono signal M[k] is calculated from the stereo channels L[k] And R[k] by polar decomposition M [ k] = | M [ k ]| .e j.∠M [ k ]< , where the amplitude |M[ k ]| and the phase ∠M [ k ] for each sub-band are defined by: M k = L k + R k 2 ∠ M k = ∠ L k + ∠ R k

[0044] The amplitude of M[k] is the average of the amplitudes of the L and R channels. The phase of M[k] is given by the phase of the signal summing the two stereo channels (L+R).

[0045] The method of Hoang et al. preserves the energy of the mono signal like the method of Samsudin et al., and it avoids the problem of total dependence of one of the stereo channels (L or R) for the phase calculation ∠M [ k ] . However, it has a disadvantage when the L and R channels are almost out of phase in certain sub-bands (with L = -R as the extreme case). Under these conditions, the resulting mono signal will be of poor quality.

[0046] In the ITU-T G.722 Annex D codec and in the article "Parametric stereo coding scheme with a new downmix method and whole band inter-channel time / phase differences" by W. Wu, L. Miao, Y. Lang, and D. Virette, Proc. ICASSP, 2013, another method for managing phase opposition in stereo signals was described. This method relies in particular on estimating a full-band phase parameter. Experimentally, it can be verified that the quality of this method is unsatisfactory for stereo signals where the phase relationship between channels is complex or for stereo speech signals with AB-type recording (using two spaced omnidirectional microphones). Indeed, this method consists of calculating the phase of the downmix signal from the phases of the L and R signals, and this calculation can result in audio artifacts for certain signals because the phase defined by short-term FFT analysis is a parameter that is difficult to interpret and manipulate.

[0047] Furthermore, this method does not directly take into account phase changes that may appear in successive frames, which may eventually cause phase jumps.

[0048] There is therefore a need for a limited complexity encoding / decoding method that allows combining channels with "robust" quality, i.e. good quality regardless of the type of multichannel signal, while handling signals in opposite phase, signals whose phase is poorly conditioned (e.g. a null channel or a channel containing only noise), or signals whose channels have complex phase relationships that are best not "manipulated", to avoid the quality problems that these signals can create.

[0049] The invention improves upon the state of the art.

[0050] To this end, it proposes a channel reduction processing method applied to a stereophonic signal to obtain a monophonic signal, characterized in that it comprises, by frequency sub-band or by frame of the stereophonic signal, the following steps: extraction of at least one indicator characterizing an interchannel correlation (ICCr) or a degree of phase opposition (ISD) between the channels of the stereophonic signal; selection, from a set of channel reduction processing modes, of a channel reduction processing mode based on the value of at least one indicator, the channel reduction processing modes of the processing set being included in the following list: passive type channel reduction processing involving the calculation of a sum of signals without application of gain compensation; passive type channel reduction processing involving the calculation of a sum of signals with application of gain compensation; adaptive type channel reduction processing with phase alignment to a reference prior to a calculation of a sum of signals.

[0051] The invention relates to a channel reduction processing device applied to a stereophonic signal to obtain a monophonic signal, characterized in that it comprises: an extraction module capable of obtaining at least one indicator characterizing an interchannel correlation or a degree of phase opposition between the channels of the stereophonic signal, by frequency sub-band or by frame of the stereophonic signal; a selection module, capable of selecting, by frequency sub-band or by frame of the stereophonic signal, from a set of channel reduction processing modes, a channel reduction processing mode based on the value of at least one indicator characterizing the channels of the multichannel audio signal, the channel reduction processing modes of the processing set being included in the following list: passive type channel reduction processing comprising the calculation of a sum of signals without application of gain compensation; passive type channel reduction processing comprising the calculation of a sum of signals with application of gain compensation;Adaptive channel reduction processing with phase alignment to a reference prior to calculating a signal sum.

[0052] This device offers the same advantages as the process it implements.

[0053] Finally, the invention relates to a computer program comprising code instructions for implementing the steps of the channel reduction processing method according to the invention when these instructions are executed by a processor.

[0054] The invention also relates to a processor-readable storage medium on which is recorded a computer program comprising code instructions for executing the steps of the process as described.

[0055] Other features and advantages of the invention will become more apparent upon reading the following description, given solely by way of non-limiting example, and made with reference to the accompanying drawings, in which: there figure 1 illustrates an encoder implementing a known state-of-the-art parametric coding method described previously; the figure 2 illustrates a decoder implementing a parametric decoding method known from the state of the art and previously described; the figure 3 illustrates a stereo parametric encoder according to an embodiment of the invention; the figures 4a , 4b , 4c , 4d , 4e And 4f illustrate in flowchart form the steps of the channel reduction treatment according to different embodiments of the invention; the figure 5 illustrates an example of the evolution of an indicator characterizing the channels of a given multichannel signal used according to an embodiment of the invention, for a given signal; the figure 6 illustrates an example of possible weightings based on the value of an indicator characterizing the channels of a signal according to one embodiment of the invention; the figure 7 illustrates a stereo parametric decoder implementing decoding adapted to signals encoded according to the coding method of the invention; the figure 8 illustrates a device for processing a decoded audio signal in which channel reduction processing according to the invention is performed; and the figure 9 illustrates a material example of equipment incorporating an encoder capable of implementing the coding process, according to an embodiment of the invention.

[0056] With reference to the figure 3 ,A parametric stereo signal encoder according to an embodiment of the invention, delivering both a mono signal and spatial information parameters of the stereo signal, is now described.

[0057] This figure presents both the entities, hardware or software modules controlled by a processor of the coding device and the steps implemented by the coding process according to an embodiment of the invention.

[0058] This describes the case of a stereo signal. The invention also applies to the case of a multichannel signal with more than 2 channels.

[0059] This parametric stereo encoder, as illustrated, uses standardized EVS mono coding; it operates with stereo signals sampled at the sampling frequency. F s of 8, 16, 32 and 48 kHz, with 20 ms frames. Thereafter, without loss of generality, the description is primarily given for the case F s =16 kHz.

[0060] It should be noted that the choice of a frame length of 20 ms is in no way restrictive in the invention which applies equally in variants of the embodiment where the frame length is different, for example 5 or 10 ms, with a codec other than EVS.

[0061] Furthermore, the invention applies equally to other types of mono coding (e.g., IETF OPUS, ITU-T G.722) operating at identical or different sampling frequencies.

[0062] Each time channel ( L ( n ) and R(n)) sampled at 16 kHz is first pre-filtered by a high-pass filter (HPF for High Pass Filter (in English) typically eliminating components below 50 Hz (blocks 301 and 302). This pre-filtering is optional, but it can be used to avoid bias due to the DC component in the estimation of parameters such as ICTD or ICC.

[0063] The canals L'(n) and R'(n) from the pre-filtering blocks are analyzed in frequency by discrete Fourier transform with sinusoidal windowing with 50% overlap and a length of 40 ms, i.e., 640 samples (blocks 303 to 306). For each frame, the signal (L'(n), R'(n)) is therefore weighted by a symmetrical analysis window covering 2 frames of 20 ms, i.e. 40 ms (i.e. 640 samples for F s (=16 kHz). The 40 ms analysis window covers the current frame and the next frame. The next frame corresponds to a 20 ms "lookahead" signal segment, commonly referred to as a "future" signal. In embodiments of the invention, other windows may be used, for example, a low-delay asymmetric window called "ALDO" in the EVS codec. Furthermore, in embodiments, the analysis windowing may be made adaptive depending on the current frame, in order to use analysis with a long window on stationary segments and analysis with short windows on transient / non-stationary segments, possibly with transition windows between long and short windows.

[0064] For the current frame of 320 samples (20 ms to F s (=16 kHz), the spectra obtained, L [ k ] And R [ k ] ( k(=0...320), comprise 321 complex coefficients, with a resolution of 25 Hz per frequency coefficient. The index coefficient k=0 corresponds to the DC component (0 Hz); it is real. The index coefficient k =320 corresponds to the Nyquist frequency (8000 Hz for F s (=16 kHz), it is also real. The coefficients with index 0 < k <160 are complex and correspond to a sub-band with a width of 25 Hz centered on the frequency of k.

[0065] The ghosts L [ k ] And R [ k are combined in block 307, described later, to obtain a mono signal (downmix) M[k] in the frequency domain. This signal is converted into time by inverse FFT and windowing-overlap with the "lookahead" part of the previous frame (blocks 308 to 310).

[0066] The algorithmic delay of the EVS codec is 30.9375 ms at F s =8 kHz and 32 ms for other frequencies F s =16, 32 or 48 kHz. This delay includes the current 20 ms frame, so the additional delay relative to the frame length is 10.9375 ms. F s =8 kHz and 12 ms for the other frequencies (i.e., 192 samples at F s (=16 kHz), the mono signal is delayed (block 311) by T = 320 - 192 = 128 samples so that the accumulated delay between the mono signal decoded by EVS and the original stereo channels becomes a multiple of the frame length (320 samples). Consequently, to synchronize the stereo parameter extraction (block 314) and spatial synthesis from the mono signal performed at the decoder, the lookahead for calculating the mono signal (20 ms) and the mono encoding / decoding delay are added to the delay TTo align the mono synthesis (20 ms), this corresponds to an additional delay of 2 frames (40 ms) relative to the current frame. This 2-frame delay is specific to the implementation detailed here; in particular, it is related to the symmetric sinusoidal windows of 20 ms. This delay could be different. In an alternative embodiment, a delay of one frame could be achieved with an optimized window with less overlap between adjacent windows, using a 311 block that introduces no delay (T=0).

[0067] The shifted mono signal is then encoded (block 312) by the EVS mono encoder, for example, at a bit rate of 13.2, 16.4, or 24.4 kbit / s. In some variations, the encoding can be performed directly on the unshifted signal; in this case, the shifting can be done after decoding.

[0068] We consider, in a particular embodiment of the invention, illustrated here in the figure 3 that block 313 introduces a two-frame delay on the spectra L[k], R[k] And M[k] in order to obtain the spectra L buf [ k ] , R buf [ k ] And M buf [ k ] .

[0069] In a more advantageous way, in terms of the amount of data to be stored, we could shift the outputs of block 314 for parameter extraction or the outputs of quantization blocks 315, 316 and 317. We could also introduce this shift at the decoder upon receiving the stereo enhancement layers.

[0070] In parallel with mono coding, stereo spatial information coding is implemented in blocks 314 to 317.

[0071] The stereo parameters are extracted (block 314) and encoded (blocks 315 to 317) from the spectra L[k], R[k] et M[k] offset by two frames: L buf [ k ] , R buf [ k ] And M buf [ k ] .

[0072] The channel reduction processing block 307 or "downmix" is now described in more detail.

[0073] According to one embodiment of the invention, this device performs a "downmix" in the frequency domain to obtain a mono signal. M [ k ] .

[0074] This processing block 307 includes a module 307a for obtaining at least one indicator characterizing the channels of the multichannel signal, in this case the stereo signal. The indicator could, for example, be an interchannel correlation indicator or an indicator measuring the degree of phase opposition between channels. The method for obtaining these indicators will be described later.

[0075] Depending on the value of this indicator, the selection block 307b selects, from a set of "downmix" processing modes, a downmix processing mode which is applied in 307c to the input signals, here the stereo signal L [ k ] , R [k to give a mono signal M[k].

[0076] THE figures 4a à 4f illustrate different embodiments implemented by the processing block 307.

[0077] To present these figures and simplify their descriptions, several parameters are first defined: • Setting ICPD[k]

[0078] The parameter ICPD [ k ] is calculated in the current frame for each frequency line k according to the formula: ICPD k = ∠ L k . R * k

[0079] This parameter corresponds to the phase difference between the L and R channels. It is used here to define the parameter ICC r . • Setting ICCr[m]

[0080] A correlation parameter is calculated for the current frame as follows: ICCp = ∑ k = 1 N FFT 2 + 1 L k . R * k e j . ICPD k ∑ k = 1 L 2 + 1 L k . L * k ∑ k = 1 L 2 + 1 R k . R * k + ϵ where N FFT is the length of the FFT (here N FFT = 640 for F S (=16 kHz). In some variants, the complex modulus |.| may not be applied, but in that case the use of the parameter ICC p (or its derivatives) will have to take into account the signed value of this parameter.

[0081] It should be noted that the division in the calculation of the parameter ICC p can be avoided because the ICC p (smoothed according to equation (16) below) is then compared to a threshold; it is common to add a small, non-zero value ε to the denominator to avoid division by zero, but this precaution is actually unnecessary and ε can be set to 0 in practice if the numerator and denominator are calculated separately. In the embodiments of the invention, this division is not necessary because the parameter ICC p (or its possibly smoothed-out version) ICC r defined below) is compared to a threshold; the absence of division in the implementation is advantageous in terms of complexity. However, to simplify the description that follows, we retain the notation implying division.

[0082] This parameter can optionally be smoothed to mitigate temporal variations. If the current frame has index m, this smoothing can be calculated using a second-order MA (Average Adjusted) filter: ICCr m = 0 , 5 . ICCp m + 0 , 25 . ICCp m − 1 + 0 , 25 . ICCp m − 2

[0083] In practice, like the division in the definition of ICCr [ m ] does not have to be explicitly calculated, this MA filter will advantageously be applied separately to the values ​​of the numerator and the denominator.

[0084] Subsequently, the ICCr parameter will be used to designate ICCr[m] (without mentioning the current frame index); if smoothing is not applied, the parameter ICC r will correspond directly to ICC p . In variants other smoothing methods can be implemented, for example by using an AR (Autoregressive) filter, smoothing the signals.

[0085] The parameter ICCr allows us to quantify the level of correlation between L and R channels when the phase differences between these channels are ignored.

[0086] In some variants, the parameter ICC p can be defined by sub-band by simply changing the bounds of the sums, as follows: ICCp b = ∑ k = k b k b + 1 − 1 L k . R * k e j . ICPD k ∑ k = k b k b + 1 − 1 L k . L * k ∑ k = k b k b + 1 − 1 R k . R * k + ϵ where kb ... k b+1 - 1 represent the indices of the frequency lines in the sub-band with index b. Again, the parameter ICCp[b] can be smoothed out and in that case the invention will be implemented as follows: instead of having a single comparison to ICCr[m], there will be just as many comparisons to ICCp [ b that there are sub-bands of index b . • Setting SGN [ m ]

[0087] The dominant channel is also identified for use as a phase reference. For example, this dominant channel can be determined via a sign parameter. SGN calculated for the current frame as the sign of the difference in levels between the L and R channels: SGN d = sign ∑ k = 1 L 2 + 1 L k − ∑ k = 1 L 2 + 1 R k where the function sign ( . ) takes the value 1 or -1 if its operand is respectively ≥0 or <0.

[0088] It is important to note that the change of reference (L or R) for aligning the mono signal (from the downmix) to the phase of L or R only occurs under certain conditions. This prevents phase problems during the overlap-addition operation after inverse transformation, when the phase reference arbitrarily changes from L to R or vice versa.

[0089] In the preferred embodiment, switching is defined as being allowed only when the signal is weakly correlated and this phase is not used in the current frame because the downmix is ​​passive in this case (see below for details on the different downmixes used). Thus, the value of SGN d in the current frame will be ignored if this condition is not met; phase reference switching will only be allowed when the value of ICCr in the current frame is less than a predetermined threshold, for example ICC r <0,4.

[0090] We will therefore ask:

[0091] In some variations the value of 0.4 may be modified, however here it corresponds to the threshold th1 = 0.4 used later.

[0092] In some variations, the initial choice SGN [1] may be modified into SGN [1] = SGN d to ensure that the phase reference corresponds to the dominant signal in the first frame, even though by definition it only includes 20 ms of signal out of 40 ms used (for the frame size used here preferentially).

[0093] In some variants, the condition for allowing a phase reference switching can be defined by a frequency line and depend on the type of downmix used in the current frame (index m) and the type of downmix used in the previous frame (index m - 1); indeed, if the downmix for the frequency line with index k The downmix in frame m-1 was passive (with gain compensation), and if the downmix selected in frame m is a downmix with alignment to an adaptive phase reference, then phase reference switching will be permitted. In other words, phase reference switching is prohibited for the index line. kas long as the downmix explicitly uses the phase reference corresponding to the SGN parameter.

[0094] The sign parameter SGN [m] therefore only changes value when ICC r is below a threshold (in the preferred embodiment). This precaution avoids changing the phase reference in areas where the channels are highly correlated and potentially out of phase. In variants, another criterion may be used to define the phase reference switching conditions.

[0095] In variants of the invention, the binary decision associated with the calculation of SGN d This can be stabilized to avoid potentially rapid fluctuations. A tolerance, for example + / -3 dB, can be defined on the level values ​​of the L and R channels to implement hysteresis, preventing phase reference changes if the tolerance is not exceeded. Inter-frame smoothing can also be applied to the signal level.

[0096] In other variants, the parameter SGN d can be calculated using a different definition of channel level, for example: SGN d = sign ∑ k = 1 L 2 + 1 L k 2 − ∑ k = 1 L 2 + 1 R k 2 or, alternatively, from the ICLD parameters in the following form: SGN d = sign ∑ b = 1 B 20 ICPD k / 10 − B where B is the number of sub-bands, or, non-equivalently SGN d = sign ∑ b = 1 B ICPD k

[0097] In other variants, it will be possible to calculate the level of the different channels in the time domain.

[0098] In variants of the invention, the explicit calculation of SGN d will not be performed, and a parameter representing the level of each channel (L or R) will be calculated separately. When using SGN g We will perform a simple comparison between these respective levels. The implementation is in fact strictly equivalent, but it avoids explicitly calculating a sign. • Setting ISD [ k ]

[0099] A parameter ISD [ k ] defined for each line of the current frame and allowing the detection of a phase opposition is also calculated: ISD k = L k − R k L k + R k

[0100] When the L and R channels are in opposite phase, the value ISD becomes arbitrarily large.

[0101] It should be noted that the division in the calculation of the parameter ISD can be avoided because the ' ISD is then compared to a threshold; it is common to add a small non-zero value to the denominator to avoid division by zero, but this precaution is unnecessary here because in the embodiments of the invention this division is not implemented. Indeed, the comparison ISD [ k ] > th0 is equivalent to the comparison |L[ k ] - R[ k ]|>th0.|L[k] + R [ k ]| , which makes the downmix mode selection process attractive in terms of complexity.

[0102] In a first embodiment, the figure 4a illustrates the steps implemented for the channel reduction treatment of block 307.

[0103] At step E400, an indicator characterizing the channels of the multichannel audio signal is obtained. In the example shown here, this is the ICCr parameter as defined above, calculated from the ICPD parameter. The ICCr indicator corresponds to a correlation measure between the channels of the multichannel signal, in this particular case between the channels of the stereo signal.

[0104] As illustrated on this figure 4a The choice of downmix depends primarily on the indicator ICCr [ m ] calculated as explained previously from the L and R channels of the current frame and possible smoothing.

[0105] The choice between downmix processing modes is made according to the value of the indicator. ICCr [ m ] .

[0106] Several downmix processing modes are planned and are part of a set of channel reduction (downmix) processing modes.

[0107] The downmix signal is calculated line by line as follows, using three potential downmixes which are listed below: 1. Passive downmix (with gain compensation). This downmix M 1 [ k ] is defined as a summation sign with energy equalization in the form: M 1 k = L k + R k 2 . γ k where γ [k] is defined such that M 1 [ k ] which is equivalent to: M 1 k = L k + R k 2 ∠ M 1 k = ∠ L k + R k We define: γ k = L k + R k L k + R k

[0108] This downmix is ​​effective for stereo signals (and their frequency decompositions by lines or sub-bands) whose channels are not highly correlated and do not have a complex phase relationship. It is not used for problematic signals where the gain γ [ k ] could take arbitrary large values, no gain limitation is used here, however in variants an amplification limitation could be implemented.

[0109] In some variations, this equalization by gain γ [ k ] may be different. For example, it would be possible to use the value already mentioned: γ k = max 2 L k 2 + R k 2 L k + R k 2 / 2

[0110] The advantage of the γ gain [k] here it is important that it ensures the same level of amplitude for the downmix M 1 [ k ] than for the other downmixes used. It is therefore preferable to adjust the gain γ[ k ] to ensure a homogeneous level of amplitude or energy between the different downmixes. 2. Downmix with alignment to an adaptive phase reference

[0111] This downmix M 3 [ k is defined as follows: M 3 k = L k + R k 2 ∠ M 3 k = 1 + SGN 2 . ∠ L k + 1 − SGN 2 . ∠ R k where the value of SGN is to be understood as the value SGN [ m ] in the current frame, but to simplify the notation the frame index is not mentioned here.

[0112] As explained previously, the phase of this downmix can also be expressed equivalently as: ∠ M 3 k = ∠ L k si niveau L > niveau R ∠ R k si niveau R > niveau L

[0113] This downmix is ​​similar to the downmix proposed by the aforementioned Samsudin method, however here the reference phase is not given by the L channel and the phase is determined line by line and not at the level of a frequency band.

[0114] The phase is fixed here according to the dominant channel identified by the SGN parameter.

[0115] This downmix is ​​useful for highly correlated signals, for example, signals recorded with AB-type microphones or binaural recording. It can also happen that independent channels have a fairly strong correlation even if they are not the same signal recorded in the L and R channels; to avoid unwanted phase reference switching, it is best to only allow such switching when the signals do not present a risk of generating audio artifacts when this downmix is ​​used. This explains the constraint ICCr[m] <0.4 in the parameter calculation SGN [ m ] when the reference phase switching condition uses this criterion.

[0116] 3. Hybrid downmix between a passive downmix (with gain compensation) and a downmix with alignment to an adaptive phase reference, dependent on an indicator measuring the degree of phase opposition between the channels (ISD [ k ], (as defined above).

[0117] This downmix M 2 [ k is defined as follows:

[0118] This downmix is ​​applied here in cases where the signals are moderately correlated and potentially out of phase. The parameter ISD [ k ] is used here to detect a phase relationship close to phase opposition, and in this case it is preferable to select the downmix with alignment to an adaptive phase reference M 3 [k]; Otherwise, passive downmixing with gain compensation M 1 [k] is sufficient.

[0119] In variants the threshold th0=1.3 applied to ISD[k] may take on other values.

[0120] Note that the downmix M 2 [ k ] corresponds either to M 1 [ k ] either at M 3 [ k ] ,depending on the parameter value ISD[k]. It will be understood that in variants of the invention, it will therefore be possible not to explicitly define this downmix M 2 [ k but to combine the decisions on the downmix selection and the criterion on ISD [ k ] . Such an example is given to the figure 4c However, it is clear that this example applies of course to all the embodiments presented here.

[0121] Thus, according to the figure 4a , if at step E401 the indicator is less than a first threshold th1, then a first downmix processing mode M1 is implemented at step E402. Si ICCr m ≤ 0 , 4 (step E401 with th1=0.4) M k = M 1 k

[0122] If at step E403 the indicator is less than a second threshold th2, then a second downmix processing mode based on M1 and M2 is implemented at step E404. Si 0,4 < ICCr m ≤ 0 , 5 (Step E403 with th2=0.5) M k = f 1 M 1 k , M 2 k

[0123] If at step E405 the indicator is less than a third threshold th3, then a third downmix processing mode based on M2 and M3 is implemented at step E406. Si 0,5 < ICCr m ≤ 0 , 6 (Step E405 with th3 = 0.6) M k = f 2 M 2 k , M 3 k

[0124] Finally, if at step E405 the indicator is above the third threshold th3, then a fourth downmix processing mode M3 is implemented at step E407. Si ICCr m > 0 , 6 (Step E405, N) M k = M 3 k

[0125] In variants of the invention, the values ​​of the thresholds th1, th2, th3 may be set to other values; the values ​​given here typically correspond to a frame length of 20 ms.

[0126] The weighting functions of the combination functions f 1( ., .) And f 2( .,. ) are illustrated in the figure 6 These combination functions perform a "crossfade" between different downmixes to avoid threshold effects, i.e., overly abrupt transitions between the respective downmixes from one frame to the next for a given frequency spectrum. Any weighting functions with complementary values ​​between 0 and 1 are suitable within the defined range, but in this implementation, these functions are derived from the function: ρ = cos 2 π 2 . ICCr m − 0 , 5 0 , 1 pour 0,4 ≤ ICCr m ≤ 0 , 6 0 autrement with f 1 M 1 k , M 2 k = 1 − ρ . M 1 k + ρ . M 2 k And f 2 M 2 k , M 3 k = 1 − ρ . M 3 k + ρ . M 2 k

[0127] Note that the parameter ICCr [ m ] is defined here at the level of the current frame; in variants this parameter may be estimated by frequency band (for example according to the ERB or Bark scale).

[0128] In a second embodiment, the figure 4b illustrates the steps implemented for channel reduction treatment of block 307. This variant embodiment aims to simplify the decision on the downmix method to use and to reduce complexity by not implementing crossfading between two downmix methods.

[0129] Steps E400, E401, E402, E405 and E407 are identical to those described with reference to the figure 4a .

[0130] Thus, according to the figure 4b , if at step E401 the indicator is less than a first threshold th1, then a first downmix processing mode M1 is implemented at step E402. Si ICCr m ≤ 0 , 4 (step E401 with th1=0.4) M k = M 1 k

[0131] If at step E405 the indicator is below a threshold th3, then a second downmix processing mode M2 ​​is implemented at step E410. Si 0,4 < ICCr m ≤ 0 , 6 (Step E405 with th3 = 0.6) M k = M 2 k

[0132] Finally, if at step E405 the indicator is above the threshold th3, then a third downmix processing mode M3 is implemented at step E407. Si ICCr m > 0 , 6 (Step E405, N) M k = M 3 k

[0133] The downmix methods M1, M2 and M3 are, for example, those described previously.

[0134] Note that the M2 downmix is ​​a hybrid downmix between the M1 and M3 downmixes which incorporates another decision criterion on another ISD indicator as defined previously.

[0135] A strictly identical achievement in terms of the result of the figure 4b is shown at the figure 4c In this variant, the evaluation of selection parameters (block E450) and downmix selection decisions (block E451) are combined.

[0136] In a third embodiment, the figure 4d This illustrates the steps implemented for channel reduction treatment of block 307. This variant embodiment aims to simplify the decision on the downmixing method to use, this time by not using passive downmixing. M 1 [ k ] . In fact, this passive downmix is ​​already included in the hybrid downmix. M 2 [ k Furthermore, hybrid downmix can be considered a more robust variant than downmix M 1 [ k because it helps avoid phase opposition problems.

[0137] The downmix at the figure 4d is calculated as follows: If at step E403, the indicator is less than a threshold th2, then the downmix treatment M2 is implemented at step E410. Si ICCr m ≤ 0 , 5 (Step E403 with th2=0.5) M k = M 2 k

[0138] If at step E405 the indicator is less than a threshold th3, then a downmix processing mode based on M2 and M3 is implemented at step E406. Si 0,5 < ICCr m ≤ 0 , 6 (Step E405 with th3 = 0.6) M k = f 2 M 2 k , M 3 k

[0139] Finally, if at step E405 the indicator is greater than the threshold th3, then a downmix processing mode M3 is implemented at step E407. Si ICCr m > 0 , 6 (Step E405, N) M k = M 3 k

[0140] In a variant not shown here, it will be possible not to use a crossfade and thus eliminate decision E405 at the figure 4d .

[0141] It should be noted that the method of implementation of the figure 4d is strictly equivalent to that of the figure 4b by setting th1 to a value ≤0.

[0142] In a fourth embodiment, the figure 4e illustrates the steps implemented for channel reduction processing of block 307. In this embodiment, the indicator characterizing the channels of the multichannel digital audio signal is the ISD phase indicator representing a measure of the degree of phase opposition of the channels of the multichannel signal.

[0143] It is determined in step E420. For a stereo signal, this parameter is as defined in equation (18) for a spectral line calculation.

[0144] Thus, according to the figure 4e , if at step E421, the indicator ISD[k] is greater than a threshold th0, then a first downmix processing mode is implemented at step E422. Si ISD k > 1 , 3 (O of step E421 with th0=1.3) then the downmix treatment is defined as follows: ∠ M k = ∠ L k M k = L k + R k 2

[0145] If at step E421 the ISD[k] indicator is less than the threshold th0, then a second downmix processing mode is implemented at step E423. Si ISD k < 1 , 3 (N of step E421 with th0=1.3) then the downmix treatment M1[k] is applied. It is defined as follows: M k = L k + R k 2 ⋅ γ k

[0146] Finally, a variant of the determination of the downmix signal of the figure 4e is presented to the figure 4f . In this variant, the main criterion for selecting the downmix mode is defined as the parameter ISD as in the figure 4e However, this parameter is now defined by sub-band at step E430. ISD [ b ] Or b is the frequency subband index (typically ERB or Bark). In this variant, when the phase relationship between the L and R channels is close to phase opposition (threshold ISD[b] >1,3), at step E431, the selected downmix mode is this time similar to the method defined in Annex D of G.722 but in a more direct way, without using fullband IPD.

[0147] Thus, according to the figure 4f , if at step E431, the indicator ISD[b] is greater than a threshold th0, then a first downmix processing mode is implemented at step E432. Si ISD k > 1 , 3 (O of step E431 with th0=1.3) then the downmix processing is defined as follows (downmix with alignment to an adaptive phase reference, M3): pour k = k b … k b + 1 − 1 ∠ M k = ∠ L k . L k + ∠ R k . R k L k + R k M k = L k + R k 2

[0148] If at step E431, the indicator ISD[b] is below the threshold th0, then a second downmix processing mode is implemented at step E433. Si ISD b < 1 , 3 (N of step E431 with th0=1.3) then the downmix processing is defined as follows (passive downmix with gain compensation, M1): pour k = k b … k b + 1 − 1 M k = L k + R k 2 . γ k

[0149] In further variations, additional classification / decision criteria can be added to refine the downmix selection; however, at least one decision between at least two downmix modes will be maintained based on the value of at least one indicator characterizing the channels of the multichannel signal, such as the parameter ICC r or the parameter ISD (on the frame, by sub-band, or by line).

[0150] Examples of downmix selection illustrated in figures 4a to 4f are not exhaustive. Other combinations or applications of criteria may be considered.

[0151] For example, a crossfade could be applied in the embodiment where the criterion is the ISD indicator.

[0152] A downmix combining 3 types of downmix with adaptive weightings, of the type M[k] = p 1 . M 1 [ k ] + p2. M 2 [ k ] + p3. M 3 [k ] could also be chosen. The weightings p1, p2 and p3 would then be adjusted according to the selection criteria.

[0153] There figure 5 gives an example of parameter evolution ICC for a given signal with decision thresholds th3 and th1 set at 0.4 and 0.6 as described in the implementation example of the figure 4b Note that these predetermined values ​​are mainly valid for a 20ms frame and they can be modified if the frame length is different.

[0154] This figure shows the fluctuation of this indicator ICCr and the indicator SGN. The It is therefore wise to adapt the downmix processing as much as possible according to the evolution of this indicator. Indeed, a significant correlation of the signals for frames from 100 to 300, for example, can allow for adaptive downmixing with alignment to a phase reference. When the indicator ICCrIf the signal is located between the thresholds th1 and th3, this means that the signal channels are moderately correlated and potentially out of phase. In this case, the downmix to be applied depends on an indicator revealing a phase opposition between the channels. If the indicator reveals a phase opposition, then it is preferable to select the downmix with alignment to an adaptive phase reference defined above by M 3 [ k ] . Otherwise, the passive downmix with gain compensation defined above by M 1 [ k is sufficient.

[0155] The value of the parameter SGN which is also represented at the figure 5 is used to choose the correct phase reference in cases where the correlation indicator is below a threshold, for example 0.4. In the example of the figure 5 The phase reference therefore changes from L to R around frame 500.

[0156] We now return to the figure 3 To adapt the spatialization parameters to the mono signal as obtained by the "downmix" processing described above, a particular extraction of the parameters by block 314 is now described.

[0157] To adapt the spatialization parameters to the mono signal as obtained by the "downmix" processing described above, a specific extraction of the parameters by block 314 is now described with reference to the figure 3 .

[0158] For the extraction of ICLD parameters (block 314), the spectra L buff [ k ] and R buf [ k are divided into 35 frequency sub-bands. These sub-bands are defined by the following boundaries:

[0159] The table above delimits (in number of Fourier coefficients) the frequency sub-bands with index b = 0 to 34. For example, the first sub-band (b=0) goes from the coefficient kb = 0à k b+ 1 - 1 = 0; it is therefore reduced to a single coefficient representing 25 Hz. Similarly, the last sub-band (k=34) goes from the coefficient kb = 308 à k b+ 1 - 1 = 320, it comprises 12 coefficients (300 Hz). The frequency line with index k =321 which corresponds to the Nyquist frequency is not taken into account here.

[0160] For each frame, the ICLD of the sub-band b =0,...,34 is calculated according to the equation: ICLD b = 10 . log 10 σ L 2 b σ R 2 b Or σ L 2 b And σ R 2 b represent respectively the energy of the left channel ( L buf [k]) and the right canal ( R buf [ k ]): σ L 2 b = ∑ k = k b k b + 1 − 1 L k . L * k σ R 2 b = ∑ k = k b k b + 1 − 1 R k . R * k

[0161] According to a particular embodiment, the ICLD parameters are encoded by a differential non-uniform scalar quantization (block 315). This quantization will not be detailed here as it is beyond the scope of the invention.

[0162] Similarly, the ICPD and ICC parameters are coded by methods known to those skilled in the art, for example with uniform scalar quantization over the appropriate interval.

[0163] With reference to the figure 7 A decoder according to an embodiment of the invention is now described.

[0164] This decoder includes a 501 demultiplexer in which the encoded mono signal is extracted to be decoded into 502 by a mono EVS decoder in this example. The portion of the bitstream corresponding to the mono EVS encoder is decoded according to the encoder's bitrate. For simplicity, we assume there are no frame losses or bit errors in the bitstream; however, known frame loss correction techniques can certainly be implemented in the decoder.

[0165] The decoded mono signal corresponds to M̂ ( n) in the absence of channel errors. A short-term discrete Fourier transform analysis with the same windowing as at the encoder is performed on M̂ ( n (blocks 503 and 504) to obtain the spectrum M̂ [ k We assume here that a decorrelation in the frequency domain (block 520) is also applied.

[0166] The portion of the bitstream associated with the stereo extension is also demultiplexed. The ICLD, ICPD, and ICC parameters are decoded to obtain ICLD q< [ b ] , ICPD q< [ b ] And ICC q< [ b (blocks 505 to 507). Furthermore, the decoded mono signal can be decoupled, for example, in the frequency domain (block 520). The implementation details of block 508 are not presented here as they are beyond the scope of the invention, but conventional techniques known to those skilled in the art can be used.

[0167] The ghosts L̂ [ k ] And R̂ [ k ] are thus calculated and then converted into the time domain by inverse FFT, windowing, addition and overlapping (blocks 509 to 514) to obtain the synthesized channels L̂ ( n ) And R̂ ( n ).

[0168] The encoder presented in reference to the figure 3 and the decoder presented in reference to the figure 7The invention has been described in the specific application of stereo encoding and decoding. It is based on a decomposition of stereo channels using a discrete Fourier transform. The invention also applies to other complex representations, such as the Modulated Complex Lapped Transform (MCLT) decomposition combining a modified cosine discrete transform (MDCT) and a modified sine discrete transform (MDST), as well as to the case of Pseudo-Quadrature Mirror Filter (PQMF) banks. Therefore, the term "frequency coefficient" used in the detailed description can be extended to the notion of "sub-band" or "frequency band" without changing the nature of the invention.

[0169] Finally, the downmixing mechanism of the invention can be used not only for encoding but also for decoding in order to generate a mono signal at the output of a stereo decoder or receiver, thus ensuring compatibility with mono-only equipment. This might be the case, for example, when switching from headphone playback to speaker playback.

[0170] There figure 8 illustrates this implementation. A stereo signal, for example, is received and decoded. (L(n), R(n)). It is transformed by the respective blocks 601, 602 and 603, 604 to obtain the left and right spectra (L[k] And R[k]).

[0171] One of the methods as described with reference to figures 4a to 4f is then implemented in processing block 605, in the same way as for processing block 307 of the figure 3 .

[0172] This processing block 605 includes a module 605a for obtaining at least one indicator characterizing the channels of the received multichannel stereo signal, here the stereo signal. The indicator can, for example, be an interchannel correlation indicator or an indicator measuring the degree of phase opposition between channels.

[0173] Depending on the value of this indicator, the selection block 605b selects, from a set of downmix processing modes, a downmix processing mode which is applied in 605c to the input signals, in this case the stereo signal L [ k ] , R [ k to give a mono signal M[k].

[0174] The encoders and decoders as described with reference to figure 3 , 7 And 8They can be integrated into multimedia equipment such as set-top boxes, audio or video players. They can also be integrated into communication equipment such as mobile phones or communication gateways.

[0175] In some variations, the case of a 5.1 channel downmix to a stereo signal is considered. Instead of two input channels for the downmix, the case of a 5.1 surround signal is considered, defined as a set of six channels: L (Front Left), C (Center), R (Front Right), Ls (Left Surround or Rear Left), Rs (Right Surround or Rear Right), and LFE (Low Frequency effects or subwoofer). In this case, two variations of the stereo 5.1 downmix can be applied according to the invention: The C and LFE channels can be combined by passive downmixing, and the result can be combined separately with the L and R channels by applying the downmixing embodiments of 2 channels (stereo) to 1 channel (mono) to obtain channels L' and R', respectively. Then, channels L' and R' can also be combined with Ls and Rs, respectively, by applying the downmixing embodiments of 2 channels (stereo) to 1 channel (mono) to obtain channels L'' and R'', respectively, which constitute the downmix result. This implementation thus relies in a "hierarchical" manner (by successive steps) on an elementary 2-to-1 type downmix described previously in various forms. In a more general embodiment, the invention can be generalized to simultaneously combine 3 channels: on one side L, Ls, C+LFE, and on the other side R, Rs, C+LFE, where C+LFE is the result of a simple passive downmix to directly obtain two channels L' and R'.In this case, we can define several downmixes as in the stereo case: one downmix. M 1 [ k Passive of the 3 signals with gain compensation, a downmix M 3 [ k of the 3 signals with adaptive phase alignment to an adaptive reference (the dominant signal among the 3). In this case, the downmix is ​​obtained according to the generalization: M k = p 1 ICCr 12 , ICCr 13 , ICCr 23 . M 1 k + p 3 ICCr 12 , ICCr 13 , ICCr 23 . M 3 k where the weights p1 and p3 are functions of several variables, for example the correlation ICCrij between each pair of channels i and j (e.g., L, Ls, C+LFE) taken two at a time.

[0176] In other variants of the invention, the number of input and output channels of the downmix may be different from the stereo to mono or 5.1 to stereo cases illustrated here.

[0177] There figure 9 represents an example of the implementation of such equipment in which an encoder as described with reference to the figure 3or a treatment device as described with reference to the figure 8 According to the invention, this device is integrated. This device includes a processor PROC cooperating with a memory block BM comprising a storage and / or working memory MEM.

[0178] The memory block may advantageously include a computer program comprising code instructions for implementing the steps of the coding process within the meaning of the invention, or of the processing process when these instructions are executed by the PROC processor, and in particular the steps of extracting at least one indicator characterizing the channels of the multichannel digital audio signal and selecting, from a set of channel reduction processing modes, a channel reduction processing mode based on the value of at least one indicator characterizing the channels of the multichannel audio signal.

[0179] These instructions are executed for channel reduction processing during encoding of a multichannel signal or processing of a decoded multichannel signal.

[0180] The program may include the steps implemented to encode the information adapted to this processing.

[0181] The MEM memory can store the different downmix processing modes to be selected according to the process of the invention.

[0182] Typically, descriptions of figure 3 , 4a to 4f They reiterate the steps of an algorithm in such a computer program. The computer program can also be stored on a memory medium readable by a reader of the device or equipment, or downloadable into its memory space.

[0183] Such equipment or an encoder includes an input module capable of receiving a multichannel signal, for example, a stereo signal with R and L channels for right and left, either via a communication network or by reading content stored on a storage medium. This multimedia equipment may also include means for capturing such a stereo signal.

[0184] The device includes an output module capable of transmitting a mono signal M from the channel reduction processing selected according to the invention and in the case of a coding device, the coded spatial information parameters P c.

Claims

1. Method for achieving downmix processing (307), applied to a stereophonic signal to obtain a monophonic signal, characterized in that it comprises, per frequency sub-band or per frame of the stereophonic signal, the following steps: - extraction (307a) of at least one indicator characterizing an inter-channel correlation (ICCr) or a degree of phase opposition (ISD) between the channels of the stereophonic signal; - selection (307b), from a set of downmix processing modes, of a downmix processing mode as a function of the value of the at least one indicator, the downmix processing modes of the processing set being contained in the following list: - downmix processing of passive type, comprising the computation of a sum of signals without application of gain compensation; - downmix processing of passive type, comprising the computation of a sum of signals with application of gain compensation; - adaptive downmix processing with phase alignment on a reference prior to computation of a sum of signals.

2. Device for achieving downmix processing (307), applied to a stereophonic signal to obtain a monophonic signal, characterized in that it comprises: - an extraction module (307a) capable of obtaining at least one indicator characterizing an inter-channel correlation or a degree of phase opposition between the channels of the stereophonic signal, per frequency sub-band or per frame of the stereophonic signal; - a selection module (307b) capable of selecting, per frequency sub-band or per frame of the stereophonic signal, from a set of downmix processing modes, a downmix processing mode as a function of the value of the at least one indicator characterizing the channels of the multi-channel audio signal, the downmix processing modes of the processing set being contained in the following list: - downmix processing of passive type, comprising the computation of a sum of signals without application of gain compensation; - downmix processing of passive type, comprising the computation of a sum of signals with application of gain compensation; - adaptive downmix processing with phase alignment on a reference prior to computation of a sum of signals.

3. Method for coding a multi-channel digital audio signal comprising a downmix processing step according to Claim 1 and a step of coding the signal resulting from the downmix processing.

4. Method according to Claim 1, wherein downmix processing is applied to a decoded stereophonic signal.

5. Computer program comprising code instructions for implementing the steps of the method according to Claim 1 when these instructions are executed by a processor.

6. Storage medium able to be read by a processor and on which there is recorded a computer program comprising code instructions for executing the steps of the method according to Claim 1.