Transition between two downmixing modes with optimised level control

ZA202606444APending Publication Date: 2026-07-29ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ZA202606444
Authority / Receiving Office
ZA · ZA
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2026-06-18
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Current channel reduction processing methods, such as those used in the IVAS codec, introduce audible defects during transitions between Phase-only correlation (POC) and Phase compensation (PHA) modes due to level differences and lack of coherence in downmix levels.

Method used

A method that applies an adaptation gain to the signal resulting from the second channel reduction processing mode during transitions, combined with a crossfade between the two modes, to minimize level differences and ensure smooth transitions.

Benefits of technology

The proposed method effectively reduces audible defects by adapting the signal levels during transitions, maintaining coherence and minimizing amplitude fluctuations, thereby improving the quality of the monophonic output.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

NOT VISIBLE DUE TO STATUS OF PATENT
Need to check novelty before this filing date? Find Prior Art

Description

[0001]DESCRIPTION Title: Transition between two channel reduction processing modes with optimized level control. Technical Field The present invention relates to the general field of audio signal processing. The invention relates in particular to the channel reduction processing commonly called "downmix", of a multichannel audio signal. Particular interest is given to the processing of reducing a stereophonic signal to a monophonic signal and to the transition between two channel reduction processing modes. This type of processing finds applications generally in the field of audio technologies, and more precisely in the field of audio coding, whether at the encoding or decoding stage. Prior Art Channel reduction or downmix consists of deducing, from a combination of the C channels of a multichannel signal x, a signal y consisting of a smaller number D of channels. In practice, this consists of defining a function ^^(.) such that:^^1(^^) ^^1(^^) ^^ ^^ ^^ ^^ ^^and n. In this invention, we are interested in the particular case D=1 and C=2, we will then speak of stereo to mono downmix or simply mono downmix. In the case C=2, we speak of 2-channel content or signal which can be stereo content or binaural content, where channel 1 (^^1) is referenced as the left channel (Left), and channel 2 (^^2) as the right channel (for Right). Subsequently, the stereo case will be seen as a general 2-channel signal, including the binaural case to avoid repeating the two terms, even if technically a binaural signal has specific characteristics. In the case of the signal y obtained after reduction processing, we speak of a monophonic or mono signal, which will subsequently be noted m(n), in the time domain, or M(k), in the frequency domain. These stereo signals can come from a recording from a pair of stereo microphones or a binaural recording, or even from an artistic mix of audio tracks.There are several microphone pairs that can be used to create stereo content. Among the most popular are the XY pair, which consists of two cardioid microphones with a difference in orientation angle between 90° and 135°. The MS pair, for "Mid-Side," is another very common pair, consisting of a cardioid microphone and a second "figure-of-8" microphone oriented at 90°. By combining these two microphones (sum / difference), the left and right channels of stereo content can be created. These pairs are said to be coincident, meaning that the microphone capsules do not have any delay between them, the spatialization being perceived by the difference in sound intensity between the left and right channels. Another category of pairs, called phase stereo, consists of using microphones that are distant from each other, thus creating a phase difference between the channels for sources located closer to one of the microphones.The best known is the AB pair, which uses two omnidirectional microphones spaced a few centimeters to several meters apart. Another very common pair is the ORTF pair, which uses both a phase difference through a 17cm spacing of the microphones and an amplitude difference through cardioid directivities (the microphones have an angle around 90°). The binaural pair consists of placing omnidirectional microphones in the ears of a person or an artificial head: this type of device allows the native creation of binaural content that can be listened to with headphones. Stereo content also includes speech and audio signals from audio mixing or post-production (for example, channels stored on a CD or DVD, or broadcast on the Internet). There are also other methods for creating stereo content that are not reviewed here. The simplest method for creating a reduced signal by downmixing is known as passive downmixing.It consists of taking an average of the left and right channels of the stereo signal such as: ^^1(^^) + ^^2(^^). ^^ ^^ = ^^ (^^, ^^) + ^^ (^^, ^^)^^ ^ ^^ = 1 2where ^^^^(^^, ^^), ^^ = 1,2 is the short-term Fourier transform of ^^^^(^^) (or FFT in English for "Fast Fourier Transform"), k being the frequency index and t the frame index: ^^2^^^^^ ^^^^(^^, ^^) = ∑^^−1^ ^^(^^). ^̌^^^(^^, ^^)^^−^^ , for all ^^ ∈ {0,… , ^^ − 1} Où j the frame index, l is the time index in the current frame, N the size of the FFT, ^^(. ) an apodization window of size L, of sine or Hann type or other, adapted to the size of the frame, and ^^^^(^^, ^^) = ^^^^(^^, ^^). This definition extends to the case where N>L and ^^^^(. ) is an augmented version of ^^ ^^(^^) with 0s (0-padding). This very simple passive downmixing method works well in situations where the right and left channels are in phase. However, in situations where the microphones are not coincident and have phase differences, such mixing generates comb filtering, which manifests itself as a coloration of the original signal. This is due to the fact that depending on the microphone spacing, the position of the source relative to the microphone pair and the frequency, the left and right channels of the same source can be either in phase as shown in Figure 1a, or in phase opposition as shown in Figure 1b. Also, during downmixing, the left and right signals will either add together (constructive interference) or cancel each other out (destructive interference), thus creating level variations and coloration effects of the reduced signal ^^ ( ^^ ). The coloration comes from the effect of comb filtering: the effects of constructive and destructive interference appear in different frequency bands, thus modifying the balance and therefore the timbre of the signal from the downmix compared to the original. To correct this intensity level defect according to the frequencies, the downmix of the e-AAC+ codec implements a level correction ^^ ( ^^ ) which ensures that the energy level of the reduced signal, in each frequency band, remains comparable to the original one: ^^(^^) = ^^(^^).^^1(^^)+^^2(^^) , with 1(^^) 2 + ^^ (^^) ^^ ^^ 2 omitting the index of This approach allows to compensate to some extent the decrease in intensity of the downmix. However, this compensation leads to overamplifications when the signals are close to the phase opposition: therefore, in practice, the correction factor ^^(^^) is limited to a maximum value (for example an upper limit of 2). Stereo to mono IVAS downmix for EVS encoding The IVAS (Immersive Voice and Audio Services) codec is developed at 3GPP as an extension of the EVS (Enhanced Voice Services) codec to stereo and immersive audio; the floating-point source code of the IVAS codec is available in the 3GPP TS 26.258 V18.0.0 specification. It integrates a channel reduction process (stereo to mono) which allows to maintain interoperability with the EVS (Enhanced Voice Services) mono codec without additional delay. This encoding of a stereo signal in an EVS-compatible manner is partially described in section 5.3.2.2 of version V1.0.0 (2023-12) of 3GPP TS 26.253; this TS 26.253 specification is currently being written and will eventually provide the detailed algorithmic description of the IVAS codec. Depending on the characteristics of the stereo signal contained in the input frame (20 ms long), a particular processing mode is selected for this same frame from the 2 available downmix modes for EVS-compatible stereo coding: the phase-only correlation (POC) mode, hereinafter referred to as POC mode, and the phase compensation (PHA) mode, hereinafter referred to as PHA mode. • The first mode, POC, consists of a simple weighted sum of the 2 stereo channels, whose weighting coefficient is based on the phase correlation and the interchannel time difference (ITD) between these two signals.• The second mode, PHA, consists of mixing the 2 stereo channels after performing a phase compensation step by filtering on one or two of the input stereo channels. A detailed description of the two downmix modes (POC and PHA), the selection method as well as the mechanism for switching from one mode to the other, is partially available in section 5.3.2 of the 3GPP TS26.253 V1.0.0 document “Codec for Immersive Voice and Audio Services - Detailed Algorithmic Description” of December 2023. Figure 2 details a channel reduction processing selection compliant with the IVAS code. The principle here consists of calculating, from the input stereo signal (therefore composed of 2 channels ^^1(^^, ^^) and ^^2(^^, ^^)), the monophonic signals ^^^^^^^^(^^, ^^) and ^^^^^^^^(^^, ^^) according to 2 methods, a first in 201 called "POC" and a second in 202 called "PHA".The selection of the downmix method appropriate to the characteristics of the input signal is carried out in 204 on the basis of different parameters such as ParPOC, ParPHA and a detection of transients in the current frame (block 203) resulting in a binary indication noted "is_transient"; in the state-of-the-art method, the binary indication of binaural content (^^^^^^^^^^^^^^^^^^) is not used in 204. Finally, from the mode selected (POC or PHA) in 204, block 205 generates a monophonic signal ^^^^^^^^(^^, ^^); block 205 makes it possible in particular to operate a transition from one downmix to the other by cross-fade, when the mode selected in the current frame is different from that of the previous frame. Block 203 is implemented in the source code (3GPP TS 26.258 V18.0.0) as follows: By default, transient detection is initialized by setting the ^^^^^^^^^^^^^^^^^^^ flag to 0 ("False").The input frame is divided into 5 subframes of 4 ms. For the left and right channels, the subframe energy is calculated as follows: ^^ 5−1 ^^ 1. ,2 + ^^) > 75 (^^ = 1,… ,4),where ^^ ^^ depends on the sampling frequency (^^ ^^ = 80, 40, and 35 at 16, 32, and 48 kHz respectively), the transient indicator is set to 1: ^^^^^^^^^^^^^^^^^^^^^ ← 1 (“True”)The energy envelope of the entire frame ^̅^ ^ ^ ^ ^^^^^ for channel i, is calculated for each subframe with smoothing: ^̅^ ^^^^^^^^ ← ^^ ∗ ^̅^ ^^^^^^^^ + (1 − ^^) ∗ ^^^^(^^) where ^^ = 0.75. Blocks 201 and 202 are described in more detail below, with reference to Figures 3 and 4, respectively, before then detailing blocks 204 and 205. Phase-only correlation (POC) mode (block 201 and Figure 3) The detailed description of the Phase-Only Correlation mode can be found in section 5.3.2.2 of 3GPP TS26.253 (version 1.0.0). Only a simplified description is given here. This processing allows a mono output to be generated from a stereo signal via an adaptive weighted sum of the 2 input channels. The left-right weighting is determined based on an optimized calculation of the inter-channel time difference or ITD in the frequency domain.The output signal is produced with a greater weight given to the channel that is ahead of the other, when the estimated time difference is non-zero; furthermore, the weight is also adjusted according to the levels of the input channels and the intermediate output signal. Figure 3 provides a schematic description of the POC mode. The input 2-channel signal x(m), where m is the index of the interleaved samples, is deinterleaved (block 301) to find the two left and right channels.Then a frequency analysis with windowing and Fourier transform (blocks 302 and 303) is carried out to obtain the spectra of the 2 channels, and the cross-correlation based on the phase spectra is estimated (block 304) before determining the time difference between the channels (noted here ITD for “Interaural Time Difference” even if it is formally an ICTD between 2 channels for “Inter-Channel Time Difference”) by peak search (block 305); block 305 also provides the correlation level R* corresponding to the ITD. The mixing factor g is then determined as follows (block 206): 1 R. ∗ − ^^ > 0 where g represents the factor of is the interchannel time difference (mistakenly called ITD). Thus, the 2 channels are combined into a mono signal in block 307 according to this factor g (calculated in 306), with a smoothing of the value of the factor applied sample by sample at the start of each frame in block 307: mPOC(t,l) = ^(t,l) x1(t,l) + (1-^(t,l)) x2(t,l), where ^(t,l) is the factor derived from g after temporal smoothing with the gain of the previous frame. Then, the energies of the left x1(t,l) and right x2(t,l) channels and of the mono signal mPOC(t,l) are determined (blocks 308, 309, 310). The left and right channels are separately reinjected (added) in blocks 313 and 315 to the mono signal from block 307 according to an energy compensation determined in block 311 with respective scale factors defined in blocks 312 and 314.Phase compensation (PHA) mode (block 202 and figure 4) This second downmix mode of the IVAS codec consists of performing phase compensation in the time domain before summing in order to modify the input stereo signal. The left and right channels are adaptively filtered by filters whose impulse responses are short and causal, and which make it possible to reduce phase differences while operating without any additional delay. Unlike the POC mode, the description of the PHA mode is not detailed in the current version V1.0.0 of December 2023 of the 3GPP TS26.253 document. A brief description of the implementation present in the 3GPP TS26.258 V18.0.0 source code of the IVAS code is therefore provided below. There are 2 sub-modes within the PHA mode, called IPD and IPD2, described below, which are chosen adaptively to the frame using the ISD, IPD and ICCr parameters defined below.The calculation of the filters is performed in the frequency domain and is based on the discrete spectra of the 2 left channels ^^1. ( ^^ ) and right ^^2 ( ^^ ) previously calculated for the POC mode (blocks 302 and 303). Figure 4 provides a schematic description of the PHA mode. The ISD indicator for "Inter Spectral Difference", which is used here to detect when destructive interference due to phase opposition between the signals may occur, is calculated (block 401) as follows: |^^ (^^) − ^^ (^^)|^^^^^^^ ^^ = 1 2 respectively ^^^^^^(^^) >1.3 and ^^^^^^(^^) < √0.9 are determined in the frequency band 50-8000Hz, the ratio^^ ^ ℎ ^^^^^ is smoothed over time. The IPD indicator for “Inter-channel phase difference”, designating the inter-channel phase difference, is calculated (block 402) per sub-band as follows: ∑ ^^^^+1 −1 ^^ ^^^^^^ ^^^^^^ ∗ ^ 1(^^)^^2(^^) ^ ^ ^=^^ ^^ ^^ ^^ − The discrete spectra are divided into K-1 sub-bands. Each value of ^^ ^^^^^^ (^^) is smoothed with a subband-dependent forgetting factor. The IPD per subband is finally normalized to obtain a unitary complex value ^̅^ ^^^^^^ ( ^^ ) where ^^ is the sub-band index. The ICCr indicator for "Realigned inter-channel correlation", designating the realigned inter-channel correlation, is determined (block 403), before being smoothed, as follows: ^^−1 ^^ ^^+1 −1^^^^^^ ^^^^^^ ∗^^^^^^^^^^^(^^)The phase is preferable is calculated and is used for the selection of the PHA sub-mode (block 404) between IPD and IPD2. The coefficients of the filters for the current frame are determined (block 405) depending on the selected PHA sub-mode. The details on the determination of the filters are not presented here because they are not relevant to the invention. If the sub-mode of the current frame is IPD, only the filter for the right channel is calculated. The filter to be applied on the right channel is estimated for the ^^ sub-bands using ^̅^ ^^^^^^ (^^) and the ILD per sub-band b (Inter-channel Level Difference which translates as the difference in levels between the left and right channels). If the sub-mode of the current frame is IPD2, the filters for the left and right channels are calculated based on the value of ^^^^^^^^ and the PHA sub-mode of the previous frame - under certain conditions the POC mode can be forced. It can also be noted that ^^ ^ ^ ^^ ^ ^^^and the IPD (in sub-band) are used in the calculation of the filters to be applied for the IPD2 sub-mode. For the 2 IPD and IPD2 modes, the impulse response of the filter to be applied on the right channel is obtained by applying an inverse FFT of length 2K. For IPD2, the filter of the left channel is deduced from that of the right channel. Finally, the filter of each channel is truncated by windowing (specific window) and is normalized. As for the filtering and the downmix (block 406), the 2 input channels are filtered separately. If the filter is not defined for one of the channels, only a gain of 1⁄ √2 is applied. For each of the 2 channels (^^ = 1,2), the downmix in PHA mode is obtained as follows: the end of the previous frame for each of the channels is concatenated with the current frame. If the previous frame filter is set, samples on the cross-fade area are filtered with the previous frame filter (with a gain of 1⁄ √2 ).Then, the channel is filtered with the filter of the current frame (with a gain of 1⁄ √2 ). Finally, the signal from the downmix is ​​crossfaded with the contents of the memory from the previous frame. Selection between POC and PHA (block 204) We now return to Figure 2 which details a channel reduction processing selection compliant with the IVAS code. As described above, the POC downmix method is implemented at 201 and one of the PHA downmix methods (IPD or IPD2 depending on the selection made at block 404 described above) is implemented at 202. Here, the implementation consists of calculating the downmixes of each method for the current frame and making a transition at 205 from one downmix to the other by crossfading, which generates a monophonic signal m. out(n). The selection of POC and PHA modes depends on the following criteria (divided into 3 categories): • From POC processing: o ITD (block 305) o The E1 / E2 and E2 / E1 ratios which reflect the difference in level between the 2 channels (from the energies from blocks 308 and 309) • From PHA processing: o Information to force the POC mode linked to the value of ^^^^^^ ^^and the PHA sub-mode selection (blocks 403 and 404) • Other criteria: o Transient detection ^^^^^^^^^^^^^^^^^^^^The downmix mode (POC or PHA) for the current frame is determined as follows (204): • Firstly: oIf |^^^^^^| > ^^^^^^^^ (where ^^^^^^^^ = 55, 19, 29 at 16, 32, 48 kHz respectively) • The POC mode is selected from 2 consecutive frames fulfilling this condition. o Otherwise, the PHA mode is selected from 2 consecutive frames fulfilling this condition. •And finally, if (^^^^^^^^^^^^^^^^^^^^ = 1) or (^^1 > 1000 ∗ ^^2) or (^^2 > 1000 ∗ ^^1) or (POC PHA sub-modes), POC mode is forced (and the consecutive frame counter is set to 0). Mode transition (POC to PHA and PHA to POC) in 205 If the current mode is also the previous mode, the downmix signal is, for ^^ = 0,… , ^^ − 1, the output signal of the selected processing:^^ ^^^^^^^^(^^), if ^^^^^^^^(^^) = ^^^^^^^^^^^^(^^) = { ^^^^^^^^ ^^^^^^As illustrated in previous mode, the downmix signal is a crossfade for ^^ = 0,… , ^^^^^^^^ − 1:^^ ^^^^^^^^(^^)^^^^^^^^(^^) + (1 − ^^^^^^^^(^^))^^^^^^^^(^^) if ^^^^^^^^(^^) = ^^^^^^ ^^^^^^ ← in Figure 7a where we can see that during the transition time, "crossfade", the outgoing downmix signal DMX1 attenuates according to the gain ^^^^^^^^1 = (1 − ^^^^^^^^(^^)) while the incoming downmix signal DMX2 increases according to the gain ^^^^^^^^2 = (^^^^^^^^(^^)), concomitantly. Hereinafter, the output signal of a channel reduction process will be called downmix. The channel reduction process will be called indifferently process, method, mode or downmix processing, to differentiate it from the output signal. The downmix process integrated into IVAS described in Figure 2 is mainly based on processing choices (POC or PHA) in each 20 ms frame depending on the characteristics of the input stereo signal. This allows to take full advantage of the benefits of each of the methods (POC and PHA). However, the current state-of-the-art solution (described in section 5.3.2 of the TS 26.253 V1.0.0 specification and implemented in the 3GPP TS26.258 V18.0.0 code) introduces audible defects related to the transitions between POC and PHA modes and therefore to the transition from one method to another, particularly in terms of level.Indeed, the output levels of each of the 2 types of downmix depend on the characteristics of the input stereo signal, the type of downmix processing, and vary over time. In transitions (from POC to PHA and vice versa), this difference in the level of the downmix signal at the output of each mode and the lack of "coherence" in the downmix level can produce amplitude fluctuations or even artifacts linked to differences perceived as temporal discontinuities in level - the level is here to be taken in the sense of sound volume similar to an RMS (Root Mean Square) measurement, an amplitude envelope, etc. These defects are particularly audible for stereo input signals such as very harmonic instrument sounds, because they are temporally structured. Indeed, the ear is particularly sensitive to phase continuity (especially at low frequencies and for tonal signals) as well as to the continuity of the sound level.The slightest modification of this temporal structure (phase shift, difference or amplitude modulation, etc.) is quickly perceived by humans. These differences in output levels of the POC and PHA modes when there is a transition between these two modes can create particularly audible artifacts for certain critical signals. Statement of the invention The invention improves the state of the art.To this end, the invention relates to a method for processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal, the processing method comprising the following steps: - application of a determined adaptation gain, to the signal resulting from the channel reduction processing according to the second mode; - application of a crossfade, during a transition duration, between the signal resulting from the channel reduction processing according to the first mode and the signal resulting from the reduction processing according to the second mode modified by the application of the adaptation gain; - progressive variation of the adaptation gain applied to the signal resulting from the second processing mode during the transition duration and / or beyond the transition duration until converging towards a value of one.This adaptation gain makes it possible to adapt the amplitude of the signal at the output of the newly selected downmix, it makes it possible on the one hand to adjust the levels in the mixing phase between the two mono channels (during the crossfade) but also to eventually find the output level of the processing relating to the selected mode (method selected for the current frame). This progressive variation of the adaptation gain makes it possible to avoid the modulation effects which can appear when a level change is carried out too quickly. In one embodiment, the adaptation gain is determined from the respective energies of the signals resulting from the channel reduction processing according to the first and second modes. Thus, the levels of the signals resulting from the channel reduction processing of the two processing modes, during a transition are respectively adjusted. In a particular embodiment, the adaptation gain is limited to threshold values.This makes it possible to limit the energy estimation errors that may appear and that may generate an audible effect. In one embodiment, in the case where a second transition between the second mode and the first channel reduction processing mode occurs before the end of the duration of progressive variation of the adaptation gain determined for the previous transition, then each of the signals resulting from the channel reduction processing operations of the first and second modes respectively are modified by the application of their respective adaptation gain for the application of the crossfade. Thus, the signal levels are adapted as best as possible even in the case of successive channel reduction processing mode transitions. In a particular embodiment, the transition duration is defined as a function of a characteristic of the stereophonic signal.Such a parameter may be, for example, the detection of a transient that requires a short transition duration (of the order of a 20ms frame) to avoid the risk of altering, in particular, the attacks (music) or the plosives (speech) of the initial signal. In one embodiment, a first channel reduction processing mode is a processing mode by filtering and rephasing at least one of the channels and a second channel reduction processing mode is a processing without filtering using only a weighted sum of the two signals. Similarly, a first channel reduction processing mode is a processing mode without filtering using only a weighted sum of the two signals and a second channel reduction processing mode is a processing by filtering and rephasing at least one of the channels. This processing method is indeed well suited to this type of channel reduction processing mode transition.The invention also relates to a device for processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal, the processing device comprising a processing circuit for implementing the following steps: - application of a determined adaptation gain, to the signal resulting from the channel reduction processing according to the second mode; - application of a crossfade, during a transition duration, between the signal resulting from the channel reduction processing according to the first mode and the signal resulting from the reduction processing according to the second mode modified by the application of the adaptation gain,; - progressive variation of the adaptation gain applied to the signal resulting from the second processing mode during the transition duration and / or beyond the transition duration until converging towards a value of one.The processing device implements the method according to any one of the particular embodiments described above and has the same advantages described above. The invention also relates to a computer program comprising instructions for implementing the processing method according to the invention, according to any one of the particular embodiments described above, when said program is executed by a processor. Such instructions may be stored durably in a non-transitory memory medium of the channel reduction processing device implementing the channel reduction processing method according to the invention. This program may use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.The invention also relates to a recording medium or information medium readable by a computer, and comprising instructions of a computer program as mentioned above. The recording medium can be any entity or device capable of storing the program. For example, the medium can comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a mobile medium, a hard disk or an SSD. Furthermore, the recording medium can be a transmissible medium such as an electrical or optical signal, which can be conveyed via an electrical or optical cable, by radio or by other means, so that the computer program it contains is remotely executable. The program according to the invention can in particular be downloaded over a network, for example an Internet-type network.Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the aforementioned channel reduction processing method. According to an exemplary embodiment, the present technique is implemented by means of software and / or hardware components. With this in mind, the term "device" or "module" may correspond in this document to a software component, a hardware component or a set of hardware and software components. Brief description of the drawings Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which: [Fig. 1a] illustrates channels of a stereophonic signal, in phase as described previously; [Fig.1b] illustrates channels of a stereophonic signal, in phase opposition as previously described; [Fig. 2] illustrates the EVS-compatible stereo-to-mono downmix embodiment of the 3GPP IVAS codec, described previously; [Fig. 3] illustrates the first method of previously described existing channel reduction processing (POC); [Fig. 4] illustrates the second method of previously described existing channel reduction processing (PHA); [Fig. 5] illustrates an embodiment of a processing method and a device for processing an audio signal when transitioning between two downmix processing modes; [Fig. 6a] illustrates a flowchart of gain adaptation between two downmix processing modes according to an embodiment of the invention; [Fig. 6b] illustrates an embodiment of a gain adaptation module between two downmix processing modes according to an embodiment of the invention; [Fig.[Fig. 7a] illustrates the operation of the crossfade during the transition between two downmix processing modes according to the state of the art; [Fig. 7b] illustrates the operation of the gain adaptation between two downmix processing modes according to an embodiment of the invention; [Fig. 7c] illustrates the operation of the gain adaptation between two downmix processing modes according to an embodiment of the invention in the case of two simultaneous transitions and adaptations; [Fig. 7d] illustrates the operation of the gain adaptation between two downmix processing modes according to an embodiment of the invention in the case of several simultaneous transitions and adaptations, a new adaptation state extending the adaptation state of the same downmix being applied; [Fig. 8] illustrates an example of a structural embodiment of a processing device according to an embodiment of the invention; [Fig.9a] illustrates examples of transition durations for applying the crossfade between two downmix processing modes; [Fig. 9b] illustrates an example of implementing simultaneous crossfades in one embodiment of the invention; [Fig. 9c] illustrates a change in transition duration of simultaneous crossfades as a function of characteristics of the signal according to one embodiment of the invention. Description of the embodiments The invention will now be described below with reference to FIGS. 5 to 8. FIG. 5 details a method for processing an audio signal during a transition between two channel reduction processing modes according to one embodiment of the invention as well as a processing device implementing such a method. The downmix method of the POC mode is implemented in 201, and the downmix method of the PHA mode is implemented in 202. These methods have been described previously with reference to FIG. 2.The transient detection at 203 and the mode selection (POC or PHA) at 204 are here assumed to be as previously described for the state of the art. In variants of the invention, the implementations of blocks 201, 202, 203 and 204 may be modified without affecting the subject of the invention which relates to blocks 501 and 502. The downmixes ^^^^^^^^(^^, ^^) and ^^^^^^^^(^^, ^^) of each method are determined for the current frame of time index t, where ^^ is the time index of the samples in the current frame of index t with ^^ = 0,…, ^^ − 1 and ^^ is the frame length (20 ms) in number of samples which depends on the sampling frequency – ^^ = 320, 640 or 960 at 16, 32 or 48 kHz, respectively. The method according to one embodiment of the invention, as described below, makes it possible to carry out an optimized transition from one downmix to another, in order to generate the monophonic output signal ^^^^^^^^(^^, ^^) in the current frame of index t.The transition optimization is carried out as follows in block 501 detailed later: ▪ The level difference is minimized by applying a gain called "adaptation gain" to the selected (incoming) downmix signal, the gain being adapted to the outgoing downmix. ▪ This "adaptation gain" constraint is then relaxed to gradually return to the intrinsic amplitude of the selected downmix. The process, the principle of which is illustrated in figures 7b and 7c, is implemented in block 501 (also called SGC for Switch Gain Control). A beneficial effect on the transitions between two downmix modes has been observed empirically by also modifying the behavior of the crossfade (block 502), this improvement of block 205 - implemented in 502 - is detailed later.As shown in Figure 7a, and in accordance with the state of the art, without adaptation gain (in 501) according to the invention the transition between 2 downmix modes, called here DMX1 and DMX2, is ensured simply by a cross-fade (as in 205) between the signals ^^1(^^, ^^) and ^^2(^^, ^^) coming from the respective modes DMX1 and DMX2 to generate ^^^^^^^^(^^, ^^). Figure 7a represents, without loss of generality, the example of a transition from the DMX1 mode (previous frame) to the DMX2 mode (selected in the current frame); this figure easily adapts to the opposite case. The vertical dotted lines represent the boundaries of the time frames.Using the notations of Figure 7a, the fade is outgoing (with decreasing values ​​according to the gain^^^^^^^^1 = (1 − ^^^^^^^^(^^))) for downmix 1 (DMX1), i.e. the signal resulting from the channel reduction processing according to a first processing mode, and incoming (with increasing values ​​according to the gain ^^^^^^^^2 = (^^^^^^^^(^^))) for downmix 2 (DMX2), i.e. the signal resulting from the channel reduction processing according to a second processing mode. The first and second processing modes can be either a POC mode or a PHA mode, the transition being able to be carried out from one mode to the other regardless of the direction of transition.Since there is no level control in the transition phase in Figure 7a, a difference in the level between the signal from DMX1 mode and the signal from DMX2 mode is not actually compensated, but is gradually spread over the crossfade duration (which here is of duration equal to the frame length, but in the case of block 205 this duration is shorter); this approach illustrated in Figure 7a is not satisfactory when there are significant level differences between DMX1 and DMX2.As illustrated in Figure 7b, the general principle of the invention is, during a transition between two downmix modes, to apply a gain called "adaptation gain" to the incoming downmix (selected in the current frame and different from that of the previous frame) in order to compensate for the difference in level with the outgoing downmix (selected in the previous frame) in particular during the particularly sensitive crossfade duration, but also beyond this crossfade duration. An adaptation gain is determined (in block 501 in the case of the invention) for one of the downmixes (output signals of the DMX1 and DMX2 modes) as a function of the downmix processing mode selection conditions. This adaptation gain is applied to the signals before summation (in block 502 in the case of the invention).In Figure 7b, which is purely illustrative, the outgoing downmix (resulting from the processing noted DMX1) is only affected by the crossfade gain (Gain fading Dmx1) since the adaptation gain of this channel is constant and of value 1. The incoming downmix (resulting from the processing noted DMX2) is affected by the crossfade gain (Gain fading Dmx2) on the one hand but also by an adaptation gain noted GDMX2. The value of the GDMX2 gain depends directly on the level difference between the two downmixes (i.e. the signals resulting from the two channel reduction processing modes) in order to compensate for this difference, at the time of the crossfade (when its duration corresponds to the frame length) and beyond this crossfade (when the duration is shorter than the frame length).The level of each downmix (at the output of blocks 201 and 202) is specific to each downmix processing and depends for each of them on the characteristics of the input stereo signal and the processing. The adaptation gain GDMX1 or GDMX2 is determined as soon as there is a change of downmix in the current frame. It can be noted in this example of figure 7b that the gain ^^^^^^^^2 (^^, ^^) is greater than 1 (amplification). This indicates here a specific level of downmix 2 lower than that of downmix 1. A gain lower than 1 (attenuation) would indicate a specific level of downmix 2 higher than that of downmix 1. Without loss of generality, the adaptation gain varies progressively towards 1 with an update per sample according to a predetermined slope (with for example a value specified in dB or linearly); however in variants this gain could be updated in steps (by sample blocks, or even by frame).As the level difference depends on the characteristics of the input stereo signal and the downmix processing, there is no pre-established rule on the value of the adaptation gain when changing the processing type (DMX1 to DMX2 and vice versa). Furthermore, in order to avoid any instability and modulation effects due to chain gains following changes in downmix methods, it is essential to gradually converge towards the intrinsic level of the method selected in the current frame (to ultimately obtain ^^′^^^^^^(^^) = ^^^^^^^^(^^)) when this decision (selection) is kept over several consecutive frames. Thus, it is necessary that the adaptation gain gradually evolves towards 1, that is to say towards a value canceling its effect. When the adaptation gain (for the selected downmix, here for the incoming downmix) is different from 1, the downmix is ​​said here to be in "adaptation state".The duration during which a downmix is ​​in “adaptation state” corresponds to the duration of convergence of the adaptation gain towards the value of 1. In practice, it is not necessary to explicitly implement an indication of “adaptation state”; this concept of adaptation state is introduced here to facilitate the understanding of the invention. Figure 7c illustrates the particular case of a change of downmix during an ongoing adaptation state (here for downmix 2). Two changes of downmix mode are illustrated here, the second change taking place while the adaptation duration of the first change is not yet complete. There can therefore exist 2 simultaneous but independent adaptation states. In this example, the convergence of the adaptation gain GDMX2 is not affected by the change of downmix methods (from DMX2 to DMX1).As illustrated in Figure 7d, there may be a situation where the frequency of downmix changes implies a new calculation of the adaptation gain of DMX2 (using the previous example) with an adaptation state in progress for this same downmix, the adaptation gain not having yet converged towards the value 1 in the current frame (of index t) where a change of downmix mode occurs. In this case, the adaptation gain is indeed calculated for the downmix change that occurs in the current frame and a new adaptation state starts for the DMX2 downmix (still using the example) and therefore prolongs the adaptation state for this same DMX2. Moreover, the crossfade does not depend on the adaptation states nor on the value of the adaptation gains. The value of the GDMX1 gain, in the example of Figures 7c and 7d, corresponds to an attenuation; it depends directly on the level difference between the DMX1 downmix (^^. ^^^^^^1 ( ^^ )) and the adapted downmix 2 (^^′ ^^^^^^2 ( ^^ ) ), in order to minimize this difference when crossfading. It is important to consider the level of the DMX2 downmix after applying the GDMX2 adaptation gain because it is this signal which constitutes the monophonic output signal (^^ ^^^^^^ ( ^^ ) ) at the time of changing downmix methods (2 to 1). It can already be noted that the duration of an adaptation state is not predetermined. Indeed, it depends on the value of the adaptation gain generically noted G DMX (for G DMX1 or G DM2) when detecting a change in downmix method; this also depends on its convergence characteristics towards 1 (direction coefficient or slope in the case of a linear variation for example or slope in dB for an exponential variation). Block 501, also called SGC for Switch Gain Control, implementing the application of an adaptation gain determined according to an embodiment of the invention, is now described in more detail with reference to Figures 6a and 6b. More precisely, Figure 6a details according to the invention the sequencing of the steps useful for implementing the method for processing an audio signal during a transition between a first and a second channel reduction processing mode, in particular the step of determining and applying an adaptation gain during a transition between the POC and PHA downmixes.For each input frame, the energy of the monophonic signals coming from each of the POC and PHA downmixes is systematically calculated (block E110 for ^̅^1. ^^^^^^ (^^) and block E111 for ^̅^1 ^^^^^^ (^^)). The details of the calculations are given below in the description of blocks 601 and 602 in figure 6b. The sequence of processing for each of the POC and PHA downmixes, called here generically "^^ ^^^^^^ (^^) & ^̅^2 ^^^^^^ (^^) » where DMX takes the value POC or PHA depending on the case, is broken down as follows depending on the downmix selected: - The adaptation gain (^^ ^^^^^^ (^^) or ^^ ^^^^^^ (^^)) of each of the downmixes is applied (depending on the adaptation state of the downmix (as described in figures 7b to 7d). The details of the calculations are given during the description of blocks 607 and 608 of figure 6b. - And the energy of the resulting signal (^̅^2 ^^^^^^ (^^) or ^̅^2 ^^^^^^(^^)) is calculated. The details of the calculations are given in the description of blocks 603 and 604 of figure 6b. According to the invention, additional conditions are added as necessary conditions to start a new adaptation state in case of a downmix change (E100). These conditions are: - Non-detection of a transient signal in the current frame (it is preferable not to equalize the energies on the transients in order to keep the temporal profile of the original signal); - Sufficient energy level in the input stereo frame (see the description of block 605 – too low energies lead to estimation errors of the adaptation gain which risks being overestimated). In the preferred embodiment a binary value called "low_egy" is defined as follows: low_egy = (^̅^^^^^^^ 1(^^) < ^^) ^^^^ (^̅^^^^^^^ 1(^^) < ^^) where for example for a signal 5^^10 8 at 48kHz, ^^ = 3^^10 8 at 32kHz and ^^ = 2.5^^108 at 16kHz. Thus, in step E100, it is verified that the current frame triggers a change in downmix processing, in other words the transition between a first mode (selected in the previous frame) and a second channel reduction processing mode (selected in the current frame and different from the mode of the previous frame). It is also verified that no transient is detected in the current frame (the binary value is_transient indicating "False") and that the level of the two input stereo channels is not too low ("low_egy" value indicating "False"). If no new adaptation is required (N at E100), "^^ ^^^^^^ (^^) & ^̅^2 ^^^^^^(^^)" can be directly applied to each of the signals resulting from the two downmix modes (at E120 and E121) according to the respective adaptation states. In the case where a transition is proven, that is_transient = "False" and that "low_egy" = "False" (i.e. 0 at E100), step E101 is implemented. At this step, it is checked which processing mode is selected as the second mode, i.e. the one after the transition, also called incoming downmix. The following steps are then implemented sequentially, where DMX takes the value POC or PHA depending on the case: - "^^ ^^^^^^ (^^) & ^̅^2 ^^^^^^ (^^) » for the outgoing downmix; - The calculation of the new adaptation gain for the incoming downmix (before being able to apply it during « ^^ ^^^^^^ (^^) & ^̅^2 ^^^^^^ (^^) » ; - And finally, « ^^ ^^^^^^ (^^) & ^̅^2 ^^^^^^(^^) » for the incoming downmix. Thus, in the case where the second downmix mode is the POC mode (O to E101), step E140 is implemented. The adaptation gain ^^ is applied ^^^^^^ (^^) resulting from a previous transition to the signal from the outgoing PHA downmix and the energy calculation ^̅^2 ^^^^^^ (^^) is performed with this adapted signal. The calculation details will be described with reference to Figure 6b. At step E141, a new adaptation gain ^^ ^^^^^^ (^^) for the incoming POC downmix is ​​calculated. The details of the calculation of the new adaptation gain are given during the description of block 605 of Figure 6b. This is applied to the signal resulting from the incoming POC downmix at E142 and an energy calculation ^̅^2 ^^^^^^ (^^) with this adapted signal is performed. The same applies in the case where the second downmix mode is the PHA mode (N to E101). Step E130 is implemented where the adaptation gain ^^ is applied ^^^^^^(^^) resulting from a previous transition to the signal from the outgoing POC downmix and the energy calculation ^̅^2 ^^^^^^ (^^) is performed with this adapted signal. The calculation details will be described with reference to Figure 6b. At step E131, a new adaptation gain ^^ ^^^^^^ (^^) for the incoming PHA downmix is ​​calculated. The details of the calculation of the new adaptation gain are given during the description of block 605 of figure 6b. This is applied to the signal resulting from the incoming PHA downmix at E132 and an energy calculation ^̅^2 ^^^^^^with this adapted signal is carried out. In variants of the invention, the two conditions necessary to start an adaptation state (namely non-detection of transients and sufficient energy level) may be modified, for example to keep only one of these two conditions, adapting them (by modifying the transient detection or the low energy threshold) or by adding additional conditions. More precisely, Figure 6b details an adaptation module implementing in particular the step of determining and applying an adaptation gain during a transition between the POC and PHA downmixes according to an embodiment of the invention. Blocks 601, 602, 603 and 604 As previously explained, the principle of the invention is, in the event of a change of downmix, to apply an adaptation gain to the incoming monophonic signal. This adaptation gain is determined using an energy ratio (block 605).As the output energy of each downmix depends on the characteristics of the input signal and the processing applied, the first step consists of calculating the short-term energies of the 2 monophonic signals from the POC and PHA downmixes (energies (^̅^1. ^^^^^^ (^^) in 601 and ^̅^1 ^^^^^^ (^^) in 602). In addition, the energy of each of the monophonic channels at the input of block 502 of figure 5 (which notably ensures the crossfade between the 2 modes) depends on the adaptation gain applied to the signal. It is therefore advantageous to determine the output energies of the gain adaptation mechanism (^̅^2 ^^^^^^ (^^) in 603 and ^̅^2 ^^^^^^ (^^) in 604) in order to ensure continuity or level coherence between the modes. The energies of the current frame are determined as follows (with L the length of the current frame and DMX=POC, PHA): ^^−1 ^^ ^^^^^^ ^^−1 ^^ ^^^^^^ where ^^^^^^^^(^^, ^^) is the input (^^, ^^) its output. Then, to limit the effect induced by rapid variations in short-term energies, a smoothed version of each of these energies is used. In the preferred embodiment, these energies are smoothed by a low-pass filtering of type IIR of order 1 in each of the blocks 601, 602, 603 and 604 (with x=1,2 and DMX=POC, PHA) so that: ^̅^ ^^^^^^ ^^ (^^) ← ^^^̅^ ^^^^^^^^ (^^ − 1) + (1 − ^^)^^^ ^ ^ ^^^^^( ^^ ) The coefficient adjacent frames of length L=20ms, this coefficient is set to ^^ = 0.9. In variants, the forgetting coefficient could be different between energies ^̅^1 ^^^^^^ and ^̅^2 ^^^^^^ (with DMX = POC or PHA) or even between downmix. In variants, other short-term energy estimation and smoothing methods could be implemented. In particular, in one variant, it would be possible to use the energy calculation method of the POC method (blocks 208 and 209 of Figure 2), however this variant has the defect of depending on a particular windowing that is applied to the signal. In other variants, the energy calculation could be carried out more finely by sub-frames, with temporal smoothing, and the energy of the last sub-frame gives the energy of the frame. In the preferred embodiment, the energies are continuously estimated at the frame. In one variant, for a downmix considered (POC or PHA), the energy ^̅^2 ^^^^^^ could no longer be calculated (block 603 or 604), and consider when there is no adaptation phase in progress. Another variant would consist of checking that the following 2 conditions are met in order not to do the calculation (block 603 or 604): - When there is no adaptation phase in progress - And, the energy ^̅^2(^^) has already had time to converge towards ^̅^1(^^) (delay linked to the smoothing of the energies). In these 2 variants, the calculation of the energy ^̅^2(^^) must be resumed as soon as a decision to change the downmix method has been taken because an adaptation state restarts. These 2 variants make it possible to reduce the average complexity of block 501, however the maximum complexity remains the same. Blocks 605 According to the invention, in the event of a change of downmix, when ^^^^^^^^(^^) is different from ^^^^^^^^(^^ − 1), a new adaptation state for the incoming downmix (selected at the current frame) can be initialized except (see conditions of block E100 of figure 6a): - If a transient has been detected (203) - Or if the energies ^̅^1 ^^^^^^(^^) and ^̅^1 ^^^^^^ (^^) at the downmix output are considered insufficient, i.e. have simultaneously values ​​lower than a threshold. In the preferred embodiment, the threshold value is 5^^10 8 at 48kHz, from 3^^10 8 at 32kHz and 2.5^^10 8at 16kHz for a 16-bit signed integer input signal) but can be set experimentally. For other variants, different level criteria could be considered, with or without the criterion of the preferred embodiment, to not initialize an adaptation state when changing the downmix method. This could be for example if the output energy of the selected downmix at frame t is lower than a threshold, which may be different from the previous threshold (without link to the output energy of the outgoing downmix).… As explained previously, variants on the necessary conditions are possible. In case of starting (initialization) of a new adaptation state in the current frame of index t for the selected downmix, the adaptation gain is calculated as follows: For a selected downmix of POC type (i.e. ^^^^^^^^(^^) = ^^^^^^):^^^^^^ P our un downmix ^^^^^^^^(^^) = ^^^^^^)::^^^^^^ For example, according to Figure 6a, if a downmix change decision has been made and the mode selected in the current frame is PHA mode (and the mode in the previous frame was POC), the gain ^^ ^^^^^^ (^^) must be calculated in order to compensate for the difference in level between the signal ^^′^^^^^^(^^, ^^) and the signal ^^′^^^^^^(^^, ^^), particularly at the time of the crossfade (if this is of duration equal to the frame length, and beyond this duration otherwise). Initially it is preferable to apply the current adaptation gain ^^ ^^^^^^ (^^) (607) and determine the energy ^̅^2 ^^^^^^( ^^ ) (603). Then, the energy ^̅^1 ^^^^^^ (^^) of the signal from the PHA downmix being determined (601), the adaptation gain ^^ ^^^^^^ (^^) can be calculated using the previous equation. Once the gain ^^ ^^^^^^(^^) calculated and applied to the signal from the PHA downmix according to the modalities of block 608, the energy of the signal ^^′^^^^^^(^^, ^^) (therefore adapted in level) can be updated (604). In the embodiment, it is also provided to limit the value of the gains calculated in 605 in order to avoid an energy estimation error causing an audible effect and also to ensure a well-defined convergence time. In the preferred embodiment, the adaptation gain is limited to + / -3dB (i.e., an adaptation gain G included in the interval [-3dB; 3dB] or [0.7071; 1.4142] to the nearest rounding error).It may be noted that in other variants of the invention: - Different values ​​of minimum and maximum thresholds may be defined; - The attenuation (adaptation gain < 1) could be limited independently of the amplification (adaptation gain > 1), for example [-6dB; 3dB] or any other values ​​while respecting a first negative gain and a second positive gain in dB, therefore while respecting adaptation gain < 1 for the first and adaptation gain > 1 for the second in linear. - This saturation of the gain in a given interval may not be implemented. Blocks 607 and 608 In a preferred embodiment, if ^^. ^^^^^^ (^^) is in the interval [0.9885, 1.0116] (i.e. if an adaptation state is in progress for POC), an adaptation gain is applied to the frame at 607 for POC: ^^^^^^^^(^^, ^^) ← ^^ ^^^^^^(^^) × ^^^^^^^^(^^, ^^), ^^ = 0,… , ^^ − 1 Similarly if ^^ ^^^^^^(^^) is in the interval, [0.9885, 1.0116] (i.e. if an adaptation state is in progress for PHA), an adaptation gain is applied to the frame still in the preferred embodiment at 608 for PHA: ^^′^^^^^^(^^, ^^) ← ^^ ^^^^^^(^^) × ^^^^^^^^(^^, ^^) , ^^ = 0,… , ^^ − 1Otherwise (if the adaptation gain is not in the cited interval): ^^′^^^^^^(^^, ^^) ← ^^^^^^^^(^^, ^^) , ^^ = 0,… , ^^ − 1where DMX = POC or PHA The adaptation gain ^^ ^^^^^^ (^^) (with DMX=POC, PHA) tends to the value 1 as t increases. After applying it, blocks 607 and 608 update the adaptation gain ^^ ^^^^^^(^^) (where DMX=POC or PHA) in the preferred embodiment for a frame: ^^^^^^^^(^^) ← ^^^^^^^^(^^ − 1) / ^^^^ if ^^^^^^^^ > 1.^^^^^^^^(^^) ← ^^^^^^^^(^^ − 1) × ^^^^ if ^^^^^^^^ < 1.where ^^^^ is the frame adaptation gain update coefficient. The coefficient ^^^^ is for example fixed so as to observe an increase in the adaptation gain of 2dB / s for a gain < 1 and a decrease of -2dB / s for a gain > 1. Thus, for frames of 20ms, ^^^^ = 1.00461543. For other variants: - The interval [0.9885, 1.0116] could take different values ​​close to 1 (with always the first value < 1 and the second value > 1); - The update coefficient ^^^^ relating to a gain <1 could be different from that associated with a gain >1; - The adaptation gains ^^ ^^^^^^(where DMX = POC or PHA) could be updated so that they converge according to a different law (linear profile or any other function). - The adaptation gains ^^ ^^^^^^ could be updated so that they converge according to a different law (linear profile or any other function). - The adaptation gains could be applied and updated per block of samples or per sub-blocks of length less than the frame length. - The adaptation gains ^^ ^^^^^^could be applied and updated to the sample. After being applied to the downmix signal in the following way: ^^′^^^^^^(^^, ^^) ← ^^ ^^^^^^(^^, ^^) × ^^^^^^^^(^^, ^^)where DMX = POC or PHA The adaptation gain value could therefore be updated in the following way (with ^^^^ the adaptation gain update coefficient): -^^^^^^^^(^^, ^^) ← ^^^^^^^^(^^, ^^ − 1) / ^^^^ if ^^^^^^^^ > 1.- ^^^^^^^^(^^, ^^) ← ^^^^^^^^(^^, ^^ − 1) × ^^^^ if ^^^^^^^^ < 1.where DMX = POC or PHA With ^^^^ the coefficient for updating the adaptation gain to the sample. - If the adaptation gains are applied and updated to the sample, or by block of samples representing a duration less than the crossfade, their values ​​could be fixed during the crossfade period (value calculated according to the terms of block 605) then updated only at the end of this period.In the preferred embodiment, for a downmix the adaptation and crossfade gains are applied separately. A variant could combine the 2 gains into 1 single gain in order to apply it at once for example. Beyond the SGC processing module (block 501), there is an influence of the crossfade module (block 502) on the quality of transition between 2 downmixes, during a change of mode in the current frame. According to the invention, for signals which do not contain a transient, that is to say for is_transient = 0 (“False”) at the output of block 203, the crossfade duration in the current frame is longer than for the case where is_transient = 1 (“True”). This makes it possible to minimize the impact of the crossfade for transients, such as attacks (music) or plosives (speech).In a preferred embodiment (as illustrated in Figure 9a), the duration of the crossfade at 502 is 60ms for a frame without transient and 20ms for a frame with transient. Variants may propose other values ​​for these 2 crossfade durations or even create additional categories of signals which would require different durations. According to the invention, since the duration of the crossfading can extend beyond one frame, it is entirely possible for a change in downmix method to occur before the end of a crossfade. According to the invention (as illustrated in Figure 9b), in the case of a crossfade already in progress, to avoid any artifact, it is advantageous to keep the current fading gains for each of the downmixes but to modify the direction of the fade.For example, in the case of a linear crossfade, for the previous frame (t-2), a change from DMX2 to DMX1 was required, so the fade was in for DMX1 and out for DMX2. A downmix change, from DMX1 to DMX2, is detected for frame t. Thus the direction of the fade changes, it is out for DMX1 and in for DMX2. But, the crossfade (triggered in frame t-2) not having reached its end, the values ​​of the fade gains ^^^^ for DMX1 and (1 − ^^^^) for DMX2 are kept to start the new fade. The duration of the new crossfade is determined by the type of signal (transient or not) as determined by a classification or detection of the type of signal contained in the frame where a downmix change is triggered.For example (as illustrated in Figure 9c), if the initial crossfade duration in progress was 60ms, and a transient is detected in the frame that triggered the downmix change, the crossfade duration becomes 20ms. More precisely, as shown in Figure 9c, the crossfade of duration 60ms that starts at frame t-2 stops at the end of frame t-1 since frame t undergoes a downmix change. For this frame t, a fade is applied for a duration of 20ms. Thus, a crossfade of a predetermined duration may not be effective until the end if a new transition occurs before this duration. In variants, the crossfade duration in case of a transition in progress could be a function of the incoming-outgoing signal categories and their direction with associated rules.For example (for durations also given as examples), during a downmix change during an ongoing fade initially of 60ms, a transient would impose a crossfade duration of 20ms but the reverse would not be true, that is to say that during an ongoing fade initially of 20ms, a non-transient signal would not impose a crossfade duration of 60ms but would retain the initial 20ms. Depending on the preferred embodiment, and for considerations of memory size reduction and complexity, the crossfade duration for transients (for example 20ms) can be a multiple of the crossfade duration for other signals (for example 60ms) having the factor N (where N would be equal to 3 in the previous example).This makes it possible to use a single pre-calculated fade gain table (in the preferred embodiment, the table contains the pre-calculated fading gains corresponding to the fade-in over 20 ms), which can be used directly for transient signals, and to repeat the same gain from the table N times for N successive samples of the other types of signals. Other variants of the invention are possible, in particular in the case of downmix methods other than POC and PHA. According to the preferred embodiment, the invention applies to a downmix that is applied to stereo input signals. In variants, the downmix processing may be applied at the output of a stereo audio system (decoder or other) to ensure mono reproduction, for example.According to the preferred embodiment, the invention relates to downmix methods from stereo to mono, but variants could implement downmix methods (N channels to mono) other than stereo (2 channels to mono), even varying the downmix methods within the same system as long as the output signals are monophonic. The invention applies to a transition between two channel reduction processing modes with optimized level control. In variants, the number of processing modes could be greater than 2; For example, if the method comprises 3 downmix methods, such as POC, PHA and a third which would be, for example, a simple downmix of type (L+R) / 2, assuming that the decision block is adapted accordingly, the invention provides for defining an additional adaptation gain for the third downmix according to the methods already described. Illustrated in FIG. 8 is a processing device 800, within the meaning of the invention.The device 800 comprises a processing circuit typically including: - a memory MEM1 for storing instruction data of a computer program within the meaning of the invention; the computer program instructions stored in the memory MEM1, with a view to processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal; in particular, the processor being able to control the processing modules as described with reference to figures 5 to 7; and - a communication interface COM 1 for transmitting the reduced signals, resulting from the channel reduction processing, ^^^^^^^^(^^, ^^), to another processing module, for example a coding or decoding module of an audio signal coder or decoder. Of course, this figure 8 illustrates an example of a structural embodiment of a processing device within the meaning of the invention. Figures 5 to 7 commented above describe in detail functional embodiments of this device.Applications for this type of channel reduction processing and transition between two channel reduction processing modes are found, for example, in audio coding, for example when the remote party does not have the capacity to reproduce stereo sound on their terminal. In this case, it is not necessary to transport a stereo signal and this saves bandwidth. This type of process can be used during point-to-point conversations if one of the participants makes a stereo or binaural sound recording. This type of process can also be present during multi-party conferences: the conference bridge spatializes the scene by creating a stereo scene, but not all participants necessarily have stereo reproduction capabilities. It is then appropriate to downmix the scene in mono for these participants.The invention can also be applied to audio decoding: in this case, it is the user's renderer (restitution module) which will downmix the stereo content to adapt to the user's reduced capabilities (single speaker for example).

Claims

CLAIMS 1. Method for processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal, the processing method comprising the following steps: - application of a determined adaptation gain (501), to the signal resulting from the channel reduction processing according to the second mode; - application of a crossfade (502), during a transition duration, between the signal resulting from the channel reduction processing according to the first mode and the signal resulting from the reduction processing according to the second mode modified by the application of the adaptation gain; - progressive variation of the adaptation gain applied to the signal resulting from the second processing mode during the transition duration and / or beyond the transition duration until converging towards a value of one. 2.Method according to claim 1, in which the adaptation gain is determined from the respective energies of the signals resulting from the channel reduction processing according to the first and second modes.

3. Method according to claim 2, in which the adaptation gain is limited to threshold values.

4. Method according to one of the preceding claims, in which, in the case where a second transition between the second mode and the first channel reduction processing mode occurs before the end of the duration of progressive variation of the adaptation gain determined for the previous transition, then each of the signals resulting from the channel reduction processing respectively of the first and second modes are modified by the application of their respective adaptation gain for the application of the crossfade.

5. Method according to one of the preceding claims, in which the transition duration is defined as a function of a characteristic of the stereophonic signal.

6. Method according to one of the preceding claims, in which a first channel reduction processing mode is a processing mode by filtering and rephasing at least one of the channels and a second channel reduction processing mode is a processing without filtering using only a weighted sum of the two signals.

7. Method according to one of claims 1 to 5, in which a first channel reduction processing mode is a processing mode without filtering. using only a weighted sum of the two signals and a second channel reduction processing mode is a filtering and rephasing processing of at least one of the channels. 8.Device for processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal, the processing device comprising a processing circuit for implementing the following steps: - application of a determined adaptation gain, to the signal resulting from the channel reduction processing according to the second mode; - application of a crossfade, during a transition duration, between the signal resulting from the channel reduction processing according to the first mode and the signal resulting from the reduction processing according to the second mode modified by the application of the adaptation gain,; - progressive variation of the adaptation gain applied to the signal resulting from the second processing mode during the transition duration and / or beyond the transition duration until converging towards a value of one. 9.Storage medium, readable by a processor, storing a computer program comprising instructions for executing the processing method according to one of claims 1 to 7.