Transition between two channel reduction processing modes with optimized level control.
The method addresses audible defects in channel reduction processing by applying an adaptation gain and crossfading signals between POC and PHA modes, optimizing level transitions and reducing artifacts, thus improving the quality of the monophonic signal.
Patent Information
- Application Number
- FR2023015179
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current channel reduction processing methods, such as those used in the IVAS codec, experience audible defects due to level variations and artifacts during transitions between Phase-only Correlation (POC) and Phase Compensation (PHA) modes.
A method involving the application of an adaptation gain to the signal from the second channel reduction processing mode, followed by a crossfade between the signals from the first and second modes, with a progressive variation of the adaptation gain to converge towards unity, thereby minimizing level differences and avoiding modulation effects.
This approach optimizes the transition between channel reduction processing modes, reducing audible artifacts and ensuring smooth level adjustments, thereby enhancing the quality of the monophonic signal output.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Transition between two channel reduction processing modes with optimized level control. Technical field
[0001] The present invention relates to the general field of audio signal processing. The invention relates in particular to the channel reduction processing commonly called "downmixing" of a multichannel audio signal. Particular attention is paid to the processing of reducing a stereophonic signal to a monophonic signal and to the transition between two channel reduction processing modes.
[0002] This type of processing finds applications generally in the field of audio technologies, and more specifically in the field of audio coding, whether at the encoding or decoding stage. Prior art
[0003] Channel reduction or downmixing consists of deducing, from a combination of the C channels of a multichannel signal x, a signal y consisting of a smaller number D of channels. In practice, this consists of defining a function f (. ) such that:
[0004] m^n) / x^n), with D < C m(n) = ■mrM
[0005] and n representing a time index of the input and output signal.
[0006] In this invention, we are interested in the particular case D=1 and C=2, we will then speak of stereo to mono downmix or simply mono downmix. In the case C=2, we speak of 2-channel content or signal which can be stereo content or binaural content, where channel 1 (xi) is referenced as the left channel (Left), and channel 2 (x2) as the right channel (for Right). Subsequently, the stereo case will be seen as a general 2-channel signal, including the binaural case to avoid repeating the two terms, even if technically a binaural signal has specific characteristics. In the case of the signal y obtained after reduction processing, we speak of a monophonic or mono signal, which will subsequently be noted m(n), in the time domain, or M(k), in the frequency domain.
[0007] These stereo signals can come from a capture from a pair of stereo microphones or a binaural capture, or even from an artistic mix of audio tracks.
[0008] There are a number of microphone pairs that can be used to create stereo content. Among the most popular are the XY pair consisting of 2 cardioid microphones with a difference in orientation angle between 90° and 135°. The MS pair for "Mid-Side" is another very common pair composed of a cardioid microphone, and a second microphone called a "figure of 8" oriented at 90°. By combining these two microphones (sum / difference), we can create the left and right channels of stereo content. These pairs are said to be coincident, that is to say that the microphone capsules do not have any delay between them, the spatialization being perceived by the difference in sound intensity between the left and right channels.
[0009] Another category of so-called phase stereo pairs consists of using microphones that are distant from each other, thus creating a phase difference between the channels for sources located closer to one of the microphones. The best known is the AB pair which uses 2 omnidirectional microphones spaced a few centimeters to several meters apart. Another very widespread pair is the ORTF pair which uses both a phase difference through a 17cm spacing of the microphones and an amplitude difference through cardioid directivities (the microphones have an angle around 90°). The binaural pair consists of placing omnidirectional microphones in the ears of a person or an artificial head: this type of device makes it possible to natively create binaural content that can be listened to through headphones.
[0010] Stereo content also includes speech and audio signals resulting from audio mixing or post-production (e.g., channels stored on a CD or DVD, or streamed over the Internet).
[0011] There are also other methods of creating stereo content that are not reviewed here.
[0012] The simplest method for creating a downmixed signal is known as passive downmixing. It involves averaging the left and right channels of the stereo signal such that:
[0013] m(>0 = aa^
[0014] or in the frequency domain:
[0015] M(tk) =
[0016] where Xj(t, k), z = 1,2 is the short-term Fourier transform of X;(t) (or FFT in English for “Fast Fourier Transform”), k being the frequency index and t the frame index: X^t, k) = 12^00-^(^ for all {0, ..., N-1}
[0017] Where j is the imaginary number such that j — [, t is the frame index, l is the index temporal in the current frame, Via size of the FFT, w(.) an apodization window of size L, of type sine or Hann or other, adapted to the size of the frame, and X^t, l) — x{t, l)- This definition extends to the case where N>L and Xz(.) is a version augmented by x^t) with 0s (O-padding).
[0018] This very simple passive downmixing method works well in situations where the right and left channels are in phase. However, in situations where the microphones are not coincident and have phase differences, such mixing generates comb filtering, which manifests itself by a coloration of the original signal. This is due to the fact that depending on the spacing of the microphones, the position of the source relative to the microphone pair and the frequency, the left and right channels of the same source can be either in phase as shown in Figure 1a, or in phase opposition as shown in Figure 1b.Also, during downmixing, the left and right signals will either add up (constructive interference) or cancel each other out (destructive interference), thus creating level variations and coloration effects of the reduced signal m(n) - The coloration comes from the effect of comb filtering: the effects of constructive and destructive interference appear in different frequency bands, thus modifying the balance and therefore the timbre of the signal from the downmix compared to the original.
[0019] To correct this intensity level defect according to the frequencies, the downmix of the e-AAC+ codec implements a level correction p(k) which ensures that the energy level of the reduced signal, in each frequency band, remains comparable to the original one: M{k) =
[0021] with
[0022] =
[0023] omitting the frame index t to simplify the notations.
[0024] This approach allows to compensate to a certain extent the decrease in intensity of the downmix. However, this compensation leads to overamplifications when the signals are close to phase opposition: therefore, in practice, the correction factor y(k) is limited to a maximum value (for example an upper limit of 2). Stereo to mono IVAS downmix for EVS encoding
[0025] The IVAS (Immersive Voice and Audio Services) codec is developed at 3GPP as an extension of the EVS (Enhanced Voice Services) codec to stereo and immersive audio; the floating-point source code of the IVAS codec is available in the 3GPP TS 26.258 V18.0.0 specification. It incorporates a channel reduction process (stereo to mono) that allows interoperability with the EVS (Enhanced Voice Services) mono codec without additional delay. This coding of a stereo signal in an EVS-compatible manner is partially described in section 5.3.2.2 of version VI.0.0 (2023-12) of 3GPP TS 26.253; this TS 26.253 specification is currently being written and will eventually provide the detailed algorithmic description of the IVAS codec.
[0026] Depending on the characteristics of the stereo signal contained in the input frame (20 ms long), a particular processing mode is selected for this same frame from among the 2 available downmix modes for EVS-compatible stereo coding: the phase correlation mode, “Phase-only correlation” (POC) mode, hereinafter referred to as POC mode, and the phase compensation mode, “Phase compensation” (PHA) mode, hereinafter referred to as PHA mode. • The first mode, POC, consists of a simple weighted sum of the 2 stereo channels, whose weighting coefficient is based on the phase correlation and the interchannel time difference (ITD) between these two same signals. • The second mode, PHA, consists of mixing the 2 stereo channels after having carried out a phase compensation step by filtering on one or two of the input stereo channels.
[0027] A detailed description of the two downmix modes (POC and PHA), the selection method as well as the mechanism for switching from one mode to the other, is partially available in section 5.3.2 of the 3GPP TS26.253 Vl.0.0 document “Codée for Immersive Voice and Audio Services - Detailed Algorithmic Description” of December 2023.
[0028] Figure 2 details a selection of channel reduction processing in accordance with the IVAS code. The principle here consists of calculating, from the input stereo signal (therefore composed of 2 channels x^t, l) and X9( t, l )), the monophonic signals mpoc(.1- 0 ct 0 according to 2 methods, a first in 201 called "POC" and a second in 202 called “PHA”. The selection of the downmix method appropriate to the characteristics of the input signal is carried out in 204 on the basis of different parameters such as ParPOc, ParPHA and a detection of transients in the current frame (block 203) resulting in a binary indication noted “is_transient”; in the state-of-the-art method, the binary indication of binaural content (Fbînaurai) is not used in 204.
[0029] Finally, from the mode selected (POC or PHA) in 204, block 205 generates a monophonic signal mOUT{t, / ); block 205 makes it possible in particular to operate a transition from one downmix to the other by cross-fade, when the mode selected in the current frame is different from that of the previous frame.
[0030] Block 203 is implemented in the source code (3GPP TS 26.258 V18.0.0) as follows:
[0031] By default, transient detection is initialized by setting to 0 (“False”) the F transient indicator.
[0032] The input frame is divided into 5 subframes of 4 ms. For the 2 left and right channels, the energy of the subframe is calculated as follows:
[0033] E^m) m = Q,. ..,4;z= 1.2
[0034] If £.(,„) / 4), OR if E^m) / (E^m-1) + €) >75 (m — 1, ..., 4), where te depends on the sampling frequency (r£ 80, 40, and 35 at respectively 16, 32, and 48 kHz), the transient indicator is set to 1: F tranmeni 1 (“True”)
[0035] The energy envelope of the entire frame E"1' for channel i is calculated for each subframe with smoothing: Eu' <— a*E“UX + ( !-«)*£)(m) where « = 0.75.
[0036] Blocks 201 and 202 are described in more detail below, with reference to Figures 3 and 4, respectively, before then detailing blocks 204 and 205.
[0037] “Phase-only correlation (fPOC)” mode fblock 201 and [Fig.3] )
[0038] The detailed description of the "Phase-Only Correlation" mode is found in section 5.3.2.2 of the 3GPP TS26.253 document (version 1.0.0). Only a simplified description is given here.
[0039] This processing allows generating a mono output from a stereo signal via a weighted adaptive sum of the 2 input channels. The left-right weighting is determined based on an optimized calculation of the inter-channel time difference or ITD in the frequency domain. The output signal is produced with a greater weight given to the channel that is ahead of the other, when the estimated time difference is non-zero; furthermore, the weight is also adjusted based on the levels of the input channels and the intermediate output signal.
[0040] [Fig.3] provides a schematic description of the POC mode. The input 2-channel signal x(m), where m is the index of the interleaved samples, is deinterleaved (block 301) to find the two left and right channels. Then a frequency analysis with windowing and Fourier transform (blocks 302 and 303) is performed to obtain the spectra of the 2 channels, and the cross-correlation based on the phase spectra is estimated (block 304) before determining the time difference between the channels (denoted here ITD for "Interaural Time Difference" even if it is formally an ICTD between 2 channels for "Inter-Channel Time Difference") by peak search (block 305); block 305 also provides the correlation level R* corresponding to the ITD.
[0041] The mixing factor g is then determined as follows (block 206):
[0042] r<0 4 r=o
[0043] where g represents the mixing factor and T is the interchannel time difference (mistakenly called ITD). Thus, the 2 channels are combined into a mono signal in block 307 according to this factor g (calculated in 306), with a smoothing of the value of the factor applied sample by sample at each start of frame in block 307:
[0044] mP0C(t,l) = Y(t,l) X1(t,l) + (1-Y(t,l)) x2(t,l),
[0045] where Y(t,l) is the factor derived from g after temporal smoothing with the gain of the previous frame.
[0046] Then, the energies of the left Xi(t,l) and right x2(t,l) channels and of the mono signal mPOc(t,l) are determined (block 308, 309, 310). The left and right channels are separately reinjected (added) in blocks 313 and 315 to the mono signal from block 307 according to an energy compensation determined in block 311 with respective scale factors defined in blocks 312 and 314.
[0047] “Phase compensation (PHA)” mode (block 202 and [Fig.4])
[0048] This second downmix mode of the IVAS codec consists of performing a phase compensation in the time domain before summing in order to modify the input stereo signal. The left and right channels are adaptively filtered by filters whose impulse responses are short and causal, and which make it possible to reduce the phase differences while operating without any additional delay.
[0049] Unlike the POC mode, the description of the PHA mode is not detailed in the current version Vl.0.0 of December 2023 of the 3GPP TS26.253 document. It is therefore Below is provided a brief description of the implementation present in the 3GPP TS26.258 V18.0.0 source code of the IVAS code.
[0050] There are 2 sub-modes within the PHA mode, called IPD and IPD2, described below, which are chosen adaptively to the frame using the parameters ISD, IPD and ICCr defined below. The calculation of the filters is carried out in the frequency domain and is based on the discrete spectra of the 2 left channels X {( k ) and right channels X? ( k ) previously calculated for the POC mode (blocks 302 and 303).
[0051] [Fig.4] provides a schematic description of the PHA mode.
[0052] The ISD indicator for "Inter Spectral Difference", designating an interspectral difference, which is used here to detect when destructive interference due to phase opposition between the signals may occur, is calculated (block 401) as follows: ISD(k) = '• 7 \Xi(k)+X2(k) |
[0053] The ratios of the number of frequencies NjSD and N!1SD which respectively verify ISD(k) > 1.3 and ISD(k) < ^0.9 are determined in the frequency band 50-8000Hz, the ratio NljSD is smoothed over time.
[0054] The IPD indicator for “Inter-channel phase difference”, designating the inter-channel phase difference, is calculated (block 402) per sub-band in the following manner:
[0055] n =—rn--77--t
[0056] The discrete spectra are divided into Kl sub-bands. Each value of a LA is rPHA\ / smoothed with a sub-band dependent forgetting factor. The IPD per sub-band is finally normalized to obtain a unitary complex value ^pHA(b) where b is the sub-band index.
[0057] The ICCr indicator for “Realigned inter-channel correlation”, designating the realigned inter-channel correlation, is determined (block 403), before being smoothed, in the following manner: 100581 JCC .= I ~
[0059] The average percentage of frequencies noted for which the compensation of phase is preferred is calculated and is used for PHA submode selection (block 404) between IPD and IPD2.
[0060] The coefficients of the filters for the current frame are determined (block 405) depending on the selected PHA sub-mode. Details on the determination of the filters are not presented here because they are not relevant to the invention.
[0061] If the sub-mode of the current frame is IPD, only the filter for the right channel is calculated. The filter to be applied on the right channel is estimated for the K sub-bands using and the ILD per sub-band b (Inter-channel Level Difference which translates as the difference in levels between the left and right channels).
[0062] If the submode of the current frame is IPD2, the filters for the left and right channels are calculated based on the value of ICCr and the PHA submode of the frame. previous - under certain conditions POC mode can be forced. It can also be noted that Nl{SD and IPD (in sub-band) are used in the calculation of the filters to be applied for the IPD2 sub-mode.
[0063] For the 2 modes IPD and IPD2, the impulse response of the filter to be applied on the right channel is obtained by applying an inverse FFT of length 2K. For IPD2, the filter of the left channel is deduced from that of the right channel* Finally, the filter of each channel is truncated by windowing (specific window) and is normalized*
[0064] As for filtering and downmixing (block 406), the 2 input channels are filtered separately. If the filter is not defined for one of the channels, only a gain of j / ^2 is applied.
[0065] For each of the 2 channels (z = 1,2), the downmix in PHA mode is obtained as follows: the end of the previous frame for each of the channels is concatenated with the current frame. If the filter of the previous frame is defined, the samples on the cross-fade mixing area are filtered with the filter of the previous frame (with a gain of [ / ^2). Then, the channel is filtered with the filter of the current frame (with a gain of ] / Finally, the signal from the downmix is a cross-fade with the contents of the memory from the previous frame. Selection between POC and PHA (block 204)
[0066] We now return to [Fig. 2] which details a channel reduction processing selection compliant with the IVAS code. As described above, the POC downmix method is implemented at 201 and one of the PHA downmix methods (IPD or IPD2 depending on the selection made at block 404 described above) is implemented at 202.
[0067] Here, the implementation consists of calculating the downmixes of each method for the current frame and operating a transition in 205 from one downmix to the other by cross-fade, which generates a monophonic signal mout(n).
[0068] The selection of POC and PHA modes depends on the following criteria (divided into 3 categories): • From POC treatment: • ITD (block 305) • The Ei / E2 and E2 / Ei ratios which reflect the difference in level between the 2 channels (from the energies from blocks 308 and 309) • From PHA treatment: • Information to force POC mode linked to the value of / CCr and the PHA sub-mode selection (blocks 403 and 404) • Other criteria: • Transient detection F process
[0069] The downmix mode (POC or PHA) for the current frame is determined as follows (204): • To begin with : • \ITD\ > TjTD(where T / TD —55, 19, 29 to respectively 16, 32, 48 kHz) • POC mode is selected from 2 consecutive frames meeting this condition. Otherwise, PHA mode is selected from 2 consecutive frames meeting this condition. And finally, if (Ftransieil[ = 1) or (E, > 1000* E2) or (E2 > 1000* E^ or (POC required in the PHA sub-mode selection phase), the POC mode is forced (and the consecutive frame counter is set to 0).
[0070] Mode transition (POC to PHA and PHA to POC) in 205
[0071]
[0072] If the current mode is also the previous mode, the downmix signal is, for n = 0, ..., L - 1, the output signal of the selected processing: mpoc (n), modeÇt) = POC mPHA (n) ' if modetf) = PHA
[0073] As illustrated in Figure 7a, if the current mode is different from the previous mode, the downmix signal is a crossfade for n = 0, ..., Lfad " h
[0074] . x ^mode(t) = POC ^OUT\^) i \ ^d ^^ pha^ + (x^dmsn^) mpoc (n) if mode (t) = phase
[0075] where
[0076] and Lf(td is the length of the fade (i.e. 20 ms). These equations are in fact illustrated in figure 7a where we can see that during the transition duration, "crossfade", the outgoing downmix signal DMX1 attenuates according to the gain R — ( i-« (n\ ) while the incoming downmix signal DMX2 increases according to the gain R = ( R ( n ) L of concomitantly.
[0077] Hereinafter, the output signal of a channel reduction process will be called downmix. The channel reduction process will be called indifferently process, method, mode or downmix process, to differentiate it from the output signal.
[0078] The downmix method integrated into IVAS described in [Fig.2] is mainly based on processing choices (POC or PHA) in each 20 ms frame depending on the characteristics of the input stereo signal. This allows to take full advantage of the benefits of each of the methods (POC and PHA).
[0079] Nevertheless, the current state-of-the-art solution (described in section 5.3.2 of the TS 26.253 Vl.0.0 specification and implemented in the 3GPP TS26.258 V18.0.0 code introduces audible defects linked to the transitions between POC and PHA modes and therefore to the passage from one method to another, particularly in terms of level. Indeed, the output levels of each of the 2 types of downmix depend on the characteristics of the input stereo signal, the type of downmix processing, and vary over time. In the transitions (from POC to PHA and vice versa), this difference in the level of the downmix signal at the output of each mode and the lack of "coherence" in the downmix level can produce amplitude fluctuations or even artifacts linked to differences perceived as temporal discontinuities in level - the level is here to be taken in the sense of sound volume similar to an RMS measurement (Root Mean Square), an amplitude envelope, etc. These defects are particularly audible for stereo input signals such as highly harmonic instrument sounds, because they are temporally structured. Indeed, the ear is particularly sensitive to phase continuity (especially at low frequencies and for tonal signals) as well as to the continuity of the sound level. The slightest change in this temporal structure (phase shift, difference or amplitude modulation, etc.) is quickly perceived by humans.
[0080] These differences in output levels of the POC and PHA modes when there is a transition between these two modes can create particularly audible artifacts for certain critical signals. Statement of the invention
[0081] The invention improves the state of the art.
[0082] To this end, the invention relates to a method for processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal, the processing method comprising the following steps: - application of a determined adaptation gain to the signal resulting from the channel reduction processing according to the second mode; - application of a crossfade, for a transition duration, between the signal resulting from the channel reduction processing according to the first mode and the signal resulting from the reduction processing according to the second mode modified by the application of the adaptation gain; - progressive variation of the adaptation gain applied to the signal from the second processing mode during the transition duration and / or beyond the transition duration until converging towards a value of one.
[0083] This adaptation gain makes it possible to adapt the amplitude of the signal at the output of the newly selected downmix, it makes it possible on the one hand to adjust the levels in the mixing phase between the two mono channels (during the crossfade) but also to eventually find the output level of the processing relative to the selected mode (method selected for the current frame). This progressive variation of the adaptation gain makes it possible to avoid modulation effects which can appear when a level change is carried out too quickly.
[0084] In one embodiment, the adaptation gain is determined from the respective energies of the signals resulting from the channel reduction processing according to the first and second modes.
[0085] Thus, the levels of the signals resulting from the channel reduction processing of the two processing modes, during a transition, are respectively adjusted.
[0086] In a particular embodiment, the adaptation gain is limited to threshold values.
[0087] This makes it possible to limit energy estimation errors which may appear and which may generate an audible effect.
[0088] In one embodiment, in the case where a second transition between the second mode and the first channel reduction processing mode occurs before the end of the duration of progressive variation of the adaptation gain determined for the previous transition, then each of the signals resulting from the channel reduction processing operations of the first and second modes respectively are modified by the application of their respective adaptation gain for the application of the crossfade.
[0089] Thus, the signal levels are best adapted even in the event of successive channel reduction processing mode transitions.
[0090] In a particular embodiment, the transition duration is defined as a function of a characteristic of the stereophonic signal.
[0091] Such a parameter may be, for example, the detection of a transient which requires having a short transition duration (of the order of a 20ms frame) so as not to risk altering in particular the attacks (music) or the plosives (speech) of the initial signal.
[0092] In one embodiment, a first channel reduction processing mode is a processing mode by filtering and rephasing at least one of the channels and a second channel reduction processing mode is a processing without filtering using only a weighted sum of the two signals.
[0093] Similarly, a first channel reduction processing mode is a processing mode without filtering using only a weighted sum of the two signals and a second channel reduction processing mode is a processing by filtering and rephasing at least one of the channels.
[0094] This processing method is indeed well suited to this type of channel reduction processing mode transition.
[0095] The invention also relates to a device for processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal, the processing device comprising a processing circuit for implementing the following steps: - application of a determined adaptation gain to the signal resulting from the channel reduction processing according to the second mode; - application of a crossfade, for a transition duration, between the signal resulting from the channel reduction processing according to the first mode and the signal resulting from the reduction processing according to the second mode modified by the application of the gain adaptation, ; - progressive variation of the adaptation gain applied to the signal from the second processing mode during the transition duration and / or beyond the transition duration until converging towards a value of one.
[0096] The processing device implements the method according to any one of the particular embodiments described previously and has the same advantages described previously.
[0097] The invention also relates to a computer program comprising instructions for implementing the processing method according to the invention, according to any one of the particular embodiments described above, when said program is executed by a processor.
[0098] Such instructions may be stored permanently in a non-transitory memory medium of the channel reduction processing device implementing the channel reduction processing method according to the invention.
[0099] This program may use any programming language, and be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0100] The invention also relates to a recording medium or information medium readable by a computer, and comprising instructions of a computer program as mentioned above.
[0101] The recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a mobile medium, a hard disk or an SSD.
[0102] On the other hand, the recording medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means, so that the computer program it contains is remotely executable. The program according to the invention may in particular be downloaded over a network, for example an Internet-type network.
[0103] Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to perform or to be used in performing the aforementioned channel reduction processing method.
[0104] According to an exemplary embodiment, the present technique is implemented by means of software and / or hardware components. In this regard, the term “device” or “module” may correspond in this document to a software component, a hardware component or a set of hardware and software components. Brief description of the drawings
[0105] Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which:
[0106] [Fig. 1a] illustrates channels of a stereophonic signal, in phase as previously described;
[0107] [Fig.lb] illustrates channels of a stereophonic signal, in phase opposition as previously described;
[0108] [Fig.2] illustrates the EVS-compatible stereo-to-mono downmix embodiment of the 3GPP IVAS codec, described previously;
[0109] [Fig.3] illustrates the first existing channel reduction processing method described previously (POC);
[0110] [Fig.4] illustrates the second method of existing channel reduction processing described previously (PHA);
[0111] [Fig.5] illustrates an embodiment of a processing method and a device for processing an audio signal during the transition between two downmix processing modes;
[0112] [Fig.6a] illustrates a flowchart of gain adaptation between two downmix processing modes according to one embodiment of the invention;
[0113] [Fig.6b] illustrates an embodiment of a gain adaptation module between two downmix processing modes according to an embodiment of the invention;
[0114] [Fig.7a] illustrates the operation of the crossfade during the transition between two downmix processing modes according to the state of the art;
[0115] [Fig.7b] illustrates the operation of the gain adaptation between two downmix processing modes according to an embodiment of the invention;
[0116] [Fig.7c] illustrates the operation of the gain adaptation between two downmix processing modes according to an embodiment of the invention in the case of two simultaneous transitions and adaptations;
[0117] [Fig.7d] illustrates the operation of the gain adaptation between two downmix processing modes according to an embodiment of the invention in the case of several simultaneous transitions and adaptations, a new adaptation state extending the adaptation state of the same downmix being applied;
[0118] [Fig.8] illustrates an example of structural embodiment of a processing device according to an embodiment of the invention;
[0119] [Fig.9a] illustrates examples of transition durations for applying crossfading between two downmix processing modes;
[0120] [Fig.9b] illustrates an example of implementation of simultaneous crossfades in an embodiment of the invention;
[0121] [Fig.9c] illustrates a change in the duration of a crossfade transition if synchronized according to characteristics of the signal according to one embodiment of the invention. Description of the embodiments
[0122] The invention will now be described below with reference to Figures 5 to 8.
[0123] [Fig.5] details a method for processing an audio signal during a transition between two channel reduction processing modes according to an embodiment of the invention as well as a processing device implementing such a method.
[0124] The POC mode downmix method is implemented in 201, and the PHA mode downmix method is implemented in 202. These methods have been described previously with reference to [Fig.2]. The transient detection in 203 and the mode selection (POC or PHA) in 204 are here assumed to be as described previously for the state of the art. In variants of the invention, the implementations of blocks 201, 202, 203 and 204 may be modified without affecting the subject of the invention which relates to blocks 501 and 502.
[0125] The downmixes mPOC(t, l) and mPHA(t, l) of each method are determined for the current frame of time index t, where l is the time index of the samples in the current frame of index t with / = 0, ..., L - 1 and L is the frame length (20 ms) in number of samples which depends on the sampling frequency - L = 320, 640 or 960 at 16, 32 or 48 kHz, respectively. The method according to an embodiment of the invention, as described below, makes it possible to carry out an optimized transition from one downmix to another, in order to generate the monophonic output signal mOUT{t, l) in the current frame of index t. The optimization of the transition is carried out in the following manner in block 501 detailed later: • The level difference is minimized by applying a so-called “matching gain” to the selected (incoming) downmix signal, the gain being matched to the outgoing downmix. • This “adaptation gain” constraint is then released to gradually return to the intrinsic amplitude of the selected downmix.
[0126] The method, the principle of which is illustrated in figures 7b and 7c, is implemented in block 501 (also called SGC for Switch Gain Control). A beneficial effect on the transitions between two downmix modes has been observed empirically by also modifying the behavior of the crossfade (block 502), this improvement of block 205 - implemented in 502 - is detailed later.
[0127] As shown in Figure 7a, and in accordance with the state of the art, without adaptation gain (at 501) according to the invention the transition between 2 downmix modes, called here DMX1 and DMX2, is simply ensured by a cross-fade (as in 205) between the signals ( t, l) and m2( / , / ) from the respective modes DMX1 and DMX2 to generate mOUT( t, / ) • Figure 7a represents, without loss of generality, the example of a transition from DMX1 mode (previous frame) to DMX2 mode (selected in the current frame); this figure is easily adapted to the opposite case. The vertical dotted lines represent the boundaries of the time frames. Using the notations of Figure 7a, the fade is out (with decreasing values according to the gain « = f 1- / 7 ( n) j) for downmix 1 (DMX1), that is to say the signal resulting from the channel reduction processing according to a first processing mode, and entering (with increasing values according to the gain 2~ ( ) ) for downmix 2 (DMX2), i.e. the signal resulting from the channel reduction processing according to a second processing mode. The first and second processing modes can be either a POC mode or a PHA mode, the transition being able to take place from one mode to the other regardless of the transition direction. Since there is no level control in the transition phase in [Fig.7a], a difference in the level between the signal from DMX1 mode and the signal from DMX2 mode is not actually compensated, but is progressively distributed over the crossfade duration (which here is of duration equal to the frame length, but in the case of block 205 this duration is shorter); this approach illustrated in [Fig.7a] is not satisfactory when there are significant level differences between DMX1 and DMX2.
[0128] As illustrated in [Fig.7b], the general principle of the invention is, during a transition between two downmix modes, to apply a gain called “adaptation gain” to the incoming downmix (selected in the current frame and different from that of the previous frame) in order to compensate for the difference in level with the outgoing downmix (selected in the previous frame) in particular during the particularly sensitive crossfade duration, but also beyond this crossfade duration.
[0129] An adaptation gain is determined (in block 501 in the case of the invention) for one of the downmixes (output signals of the DMX1 and DMX2 modes) as a function of the downmix processing mode selection conditions. This adaptation gain is applied to the signals before summation (in block 502 in the case of the invention).
[0130] In [Fig.7b] which is purely illustrative, the outgoing downmix (resulting from the processing noted DMX1) is only affected by the crossfading gain (Gain fading Dmxl) since the adaptation gain of this channel is constant and of value 1. The incoming downmix (resulting from the processing noted DMX2) is affected by the crossfading gain (Gain fading Dmx2) on the one hand but also by an adaptation gain noted GDmx2- The value of the gain GDMX2 depends directly on the difference in level between the two downmixes (i.e. the signals resulting from the two processing modes of channel reduction) in order to compensate for this difference, at the time of the crossfade (when its duration corresponds to the frame length) and beyond this crossfade (when the duration is shorter than the frame length).
[0131] The level of each downmix (at the output of blocks 201 and 202) is specific to each downmix processing and depends for each of them on the characteristics of the input stereo signal and the processing. The adaptation gain GDMXi or GDMX2 is determined as soon as there is a change of downmix in the current frame.
[0132] It can be noted in this example of figure 7b that the gain GDmx^, l) is greater than 1 (amplification). This indicates here a natural level of downmix 2 lower than that of downmix 1. A gain lower than 1 (attenuation) would indicate a natural level of downmix 2 higher than that of downmix 1. Without loss of generality, the adaptation gain varies progressively towards 1 with an update per sample according to a predetermined slope (with for example a value specified in dB or linearly); however in variants this gain could be updated in stages (by sample blocks, or even by frame).
[0133] As the level difference depends on the characteristics of the input stereo signal and the downmix processing, there is no pre-established rule on the value of the adaptation gain when changing the processing type (DMX1 to DMX2 and vice versa).
[0134] Furthermore, in order to avoid any instability and modulation effect due to chain gains following changes in downmix methods, it is essential to gradually converge towards the intrinsic level of the method selected in the current frame (to ultimately obtain m ' DMX ( t ) = ( l) ) when this decision (selection) is maintained over several consecutive frames. Thus, it is necessary that the adaptation gain gradually evolves towards 1, that is to say towards a value canceling its effect.
[0135] When the adaptation gain (for the selected downmix, here for the incoming downmix) is different from 1, the downmix is said here to be in “adaptation state”. The duration during which a downmix is in “adaptation state” corresponds to the duration of convergence of the adaptation gain towards the value of 1. In practice, it is not necessary to explicitly implement an indication of “adaptation state”; this concept of adaptation state is introduced here to facilitate the understanding of the invention.
[0136] [Fig.7c] illustrates the particular case of a downmix change during an ongoing adaptation state (here for downmix 2). Two changes of downmix mode are illustrated here, the second change taking place while the adaptation duration of the first change is not yet complete. There can therefore be two simultaneous but independent adaptation states. In this example, the convergence of the gain GDMX2 adaptation is not affected by changing downmix methods (from DMX2 to DMX1).
[0137] As illustrated in [Fig.7d], there may be a situation where the frequency of downmix changes implies a new calculation of the adaptation gain of the DMX2 (resuming the previous example) with an adaptation state in progress for this same downmix, the adaptation gain not having yet converged towards the value 1 in the current frame (of index t) where a change of downmix mode occurs. In this case, the adaptation gain is indeed calculated for the downmix change which occurs in the current frame and a new adaptation state starts for the DMX2 downmix (still resuming the example) and therefore prolongs the adaptation state for this same DMX2.
[0138] Furthermore, the crossfade does not depend on the adaptation states or the value of the adaptation gains. The value of the gain GDMXi, in the example of figures 7c and 7d, corresponds to an attenuation; it depends directly on the difference in level between the downmix DMX1 t) ) and the adapted downmix 2 ) )' a^n minimize this difference when crossfading. It is important to consider the level of the DMX2 downmix after applying the adaptation gain G DMX2 because it is this signal which constitutes the monophonic output signal) when changing downmix methods (2 to 1).
[0139] It can already be noted that the duration of an adaptation state is not predetermined. Indeed, it depends on the value of the adaptation gain generically noted GDmx (for GDmxi or GDm2) when detecting a change in downmix method; it also depends on its convergence characteristics towards 1 (direction coefficient or slope in the case of a linear variation for example or slope in dB for an exponential variation).
[0140] Block 501, also called SGC for Switch Gain Control, implementing the application of an adaptation gain determined according to an embodiment of the invention, is now described in more detail with reference to figures 6a and 6b.
[0141] More precisely, [Fig.6a] details according to the invention the sequencing of the steps useful for implementing the method for processing an audio signal during a transition between a first and a second channel reduction processing mode, in particular the step of determining and applying an adaptation gain during a transition between the POC and PHA downmixes.
[0142] For each input frame, the energy of the monophonic signals coming from each of the POC and PHA downmixes is systematically calculated (block El 10 for and ^oc El 11 for The details of the calculations are given below during the description of blocks 601 and 602 of [Fig.6b].
[0143] The sequence of processing for each of the POC and PHA downmixes, here generically called “GDMX^& Ê^MX^j” where DMX takes the value POC or PHA depending on the case, is broken down as follows depending on the selected downmix: - The adaptation gain (GPOC(t^ or GPHA(lÿ of each of the downmixes is applied (depending on the adaptation state of the downmix (as described in figures 7b to 7d). The details of the calculations are given during the description of blocks 607 and 608 of [Fig.6b]. - And the energy of the resulting signal (ÊPOC[t^ or ËqHA^ is calculated. The details of the calculations are given during the description of blocks 603 and 604 of [Fig.6b].
[0144] According to the invention, additional conditions are added as necessary conditions for starting a new adaptation state in case of downmix change (E100). These conditions are: - Non-detection of a transient signal in the current frame (it is preferable not to equalize the energies on the transients in order to keep the temporal profile of the original signal); - Sufficient energy level in the input stereo frame (see the description of block 605 - too low energies lead to estimation errors of the adaptation gain which risks being overestimated). In the preferred embodiment a binary value called "low_egy" is defined as follows: 10W_egy= (Ei°C(t) < 5) and < s) where for example for a 16-bit signed integer input signal 5 = 5xlOs at 48kHz, <5 = 3x108 at 32kHz and <5 = 2.5xl08 at 16kHz.
[0145] Thus, in step E100, it is verified that the current frame triggers a change in downmix processing, in other words the transition between a first mode (selected in the previous frame) and a second channel reduction processing mode (selected in the current frame and different from the mode of the previous frame). It is also verified that no transient is detected in the current frame (the binary value is_transient indicating “False”) and that the level of the two input stereo channels is not too low (“low_egy” value indicating “False”).
[0146] If no new adaptation is required (N to E100), "GDMX(t^ E^^fy * Can be directly applied to each of the signals resulting from the two downmix modes (at E120 and E121) according to the respective adaptation states.
[0147] In the case where a transition is proven, that is_transient = "False" and that "low_egy" = "False" (i.e. 0 to E100), step E101 is implemented. At this step, it is checked which processing mode is selected as the second mode, i.e. the one after the transition also called incoming downmix.
[0148]
[0149]
[0150]
[0151]
[0152]
[0153]
[0154]
[0155]
[0156]
[0157]
[0158] The following steps are then implemented sequentially, where DMX takes the value POC or PHA depending on the case: _ « qDMX^ * For 'c outgoing downmix; - Calculating the new adaptation gain for the incoming downmix (before being able to apply it during " ' - And finally, "& E^1^^" For 'c downmix incoming. Thus, in the case where the second downmix mode is the POC mode (0 to E101), step E140 is implemented. The adaptation gain GPHA^ resulting from a previous transition is applied to the signal from the outgoing PHA downmix and the energy calculation ^PHA^ is carried out with this adapted signal. The calculation details will be described with reference to [Fig.6b]. At step E141, a new GPOC adaptation gain^ for the incoming POC downmix is calculated. The details of the calculation of the new adaptation gain are given during the description of block 605 of [Fig.6b]. This is applied to the signal resulting from the incoming POC downmix at E142 and an energy calculation with this adapted signal is carried out. The same applies in the case where the second downmix mode is the PHA mode (N to E101). Step E130 is implemented where the adaptation gain is applied resulting from a previous transition to the signal from the outgoing POC downmix and the energy calculation E^^t} is carried out with this adapted signal. The calculation details will be described with reference to [Fig.6b]. At step E131, a new adaptation gain GPHA[t^ for the incoming downmix PHA is calculated. The details of the calculation of the new adaptation gain are given during the description of block 605 of [Fig.6b]. This is applied to the resulting signal of the incoming PHA downmix at El32 and a £PHA energy calculation with this adapted signal is carried out. In variants of the invention, the two conditions necessary to start an adaptation state (namely non-detection of transients and sufficient energy level) may be modified, for example to keep only one of these two conditions, adapting them (by modifying the transient detection or the low energy threshold) or by adding additional conditions. More specifically, [Fig.6b] details an adaptation module implementing
[0159]
[0160]
[0161]
[0162]
[0163]
[0164]
[0165]
[0166]
[0167] in particular the step of determining and applying an adaptation gain during a transition between the POC and PHA downmixes according to an embodiment of the invention. Blocks 601, 602, 603 and 604 As previously explained, the principle of the invention is, in the event of a change in downmix, to apply an adaptation gain to the incoming monophonic signal. This adaptation gain is determined using an energy ratio (block 605). Since the output energy of each downmix depends on the characteristics of the input signal and the processing applied, the first step consists of calculating the short-term energies of the 2 monophonic signals from the POC and PHA downmixes (energies ( E^OC( t) in 601 EiHA(t) in 602). In addition, the energy of each of the monophonic channels at the input of block 502 of figure 5 (which notably ensures the crossfade between the 2 modes) depends on the adaptation gain applied to the signal. It is therefore advantageous to determine the output energies of the gain adaptation mechanism ( ^2^(0 in $03 and E^'î t) in 604) in order to ensure continuity or level coherence between the modes. The energies of the current frame are determined as follows (with L the length of the current frame and DMX=POC, PHA): A""A)=L«™DJ«(>. / )2 E^mx(0 = E / =( / n'Dmx( 0 2 where 0 is the input of block 501 in Figure 5 and ni' 0 is its output- Then, to limit the effect induced by rapid variations in short-term energies, a smoothed version of each of these energies is used. In the preferred embodiment, these energies are smoothed by a low-pass filtering of type IIR of order 1 in each of the blocks 601, 602, 603 and 604 (with x=1,2 and DMX=POC, PHA) so that: aEDMX + EDMX ( t ) The forgetting coefficient a is chosen ad-hoc. In the case of ad frames adjacent of length L=20ms, this coefficient is fixed at “— 0.9. In variants, the forgetting coefficient could be different between energies fJ)MX The and (with DMX = POC or PHA) or even between downmix. In variants, other short-term energy estimation and smoothing methods could be implemented. In particular, in one variant, it would be possible to use the energy calculation method of the POC method (blocks 208 and 209 of [Fig.2]), however this variant has the defect of depending on a particular windowing that is applied to the signal. In other variants, the energy calculation could be carried out more finely by sub-frames, with temporal smoothing, and the energy of the last subframe gives the energy of the frame.
[0168]
[0169] In the preferred embodiment, the energies are continuously estimated at the frame. In a variant, for a considered downmix (POC or PHA), the energy could no longer be calculated (block 603 or 604), and consider that ~DMX El _ when there is no adaptation phase in progress. Another variant would consist of verifying that the following 2 conditions are met in order not to do the calculation (block 603 or 604): - When there is no adaptation phase in progress - And, the energy Ë^t) has already had time to converge towards Ë^t) (delay linked to the smoothing of the energies).
[0170] In these 2 variants, the calculation of the energy Ë^t) must be resumed as soon as a decision to change the downmix method has been taken because an adaptation state restarts. These 2 variants make it possible to reduce the average complexity of the block 501, however the maximum complexity remains the same. Blocks 605
[0171] According to the invention, in the event of a change of downmix, when moddfy is different from modeit- 1), a new adaptation state for the incoming downmix (selected at the current frame) can be initialized except (see conditions of block E100 of [Fig.6a]): - If a transient has been detected (203) - Or if the energies and output of downmix are considered as insufficient, i.e. simultaneously have values below a threshold. In the preferred embodiment, the threshold value is 5xl08 at 48kHz, 3x108 at 32kHz and 2.5xl08 at 16kHz for a 16-bit signed integer input signal) but can be adjusted experimentally.
[0172] For other variants, different level criteria could be considered, with or without the criterion of the preferred embodiment, to not initialize an adaptation state during a change of downmix method. This could be for example if the output energy of the downmix selected at frame t is lower than a threshold, which may be different from the previous threshold (without link with the output energy of the outgoing downmix).... As explained previously, variants on the necessary conditions are possible.
[0173] In case of starting (initialization) of a new adaptation state in the current frame of index t for the selected downmix, the adaptation gain is calculated as follows:
[0174]
[0175]
[0176]
[0177] For a selected downmix of POC type (i.e. mode^t} = POO'- For a selected downmix of type PHA (i.e. mode(t) — PHA)'-'-
[0178] For example, according to Figure 6a, if a downmix change decision has been taken and the mode selected in the current frame is the PHA mode (and the mode of the previous frame was POC), the GPHA^ gain must be calculated in order to compensate for the difference in level between the signal m'pHA(t, l) and the signal m'POC(t, l), in particular at the time of the crossfade (if this is of duration equal to the frame length, and beyond this duration otherwise).
[0179] Firstly it is preferable to apply the current adaptation gain GPOC(t) (607) and determine the energy ) (603).
[0180] Then, the energy of the signal coming from the PHA downmix being determined
[0181] (601), the adaptation gain GPHA{t) can be calculated according to the previous equation.
[0182] Once the gain GFHA(t) has been calculated and applied to the signal from the PHA downmix according to the terms of block 608, the energy of the signal m'pjq^t, l) (therefore adapted in level) can be updated (604).
[0183] In the embodiment, it is also provided to limit the value of the gains calculated at 605 in order to prevent an energy estimation error from causing an audible effect and also to ensure a well-defined convergence time. In the preferred embodiment, the adaptation gain is limited to + / -3dB (i.e., an adaptation gain G included in the interval [-3dB; 3dB] or [0.7071; 1.4142] to the nearest rounding error).
[0184] It may be noted that in other variants of the invention: - Different minimum and maximum threshold values can be defined; - The attenuation (adaptation gain < 1) could be limited independently of the amplification (adaptation gain > 1), for example [-6dB; 3dB] or any other values while respecting a first negative gain and a second positive gain in dB, therefore while respecting adaptation gain < 1 for the first and adaptation gain > 1 for the second in linear. - This gain saturation in a given interval may not be implemented
[0185] artwork. Blocks 607 and 608 In a preferred embodiment, if GPOC(ij is in the interval [0.9885, 1.0116] (i.e. if an adaptation state is in progress for POC), an adaptation gain is applied to the frame at 607 for POC:
[0186] MPOC [T, / )*“Gpoc (t) xmpoc (t, = ql -1
[0187] Similarly if OPHA(tj is in the interval, [0.9885, 1.0116] (in other words if an adaptation state is in progress for PHA), an adaptation gain is applied to the frame still in the preferred embodiment at 608 for PHA:
[0188] M'PH^ t, GPHA (î) xmphj^ l) j = q
[0189] Otherwise (if the adaptation gain is not in the cited interval):
[0190] M'DM^t, 1) î), 1 = 0, ..., L- 1
[0191] where DMX = POC or PHA .
[0192] The adaptation gain GDMX[t] (with DMX=POC, PHA) tends towards the value 1 when t increases. After applying it, blocks 607 and 608 update the adaptation gain GDMX{t^ (where DMX=POC or PHA) in the preferred embodiment for a frame:
[0193] GDM \ t)^ GDMX (TY)} RTSïgdmx> 0
[0194]
[0195] where Rt is the frame adaptation gain update coefficient.
[0196] The coefficient Rt is for example fixed so as to observe an increase in the adaptation gain of 2dB / s for a gain < 1 and a decrease of -2dB / s for a gain > 1.
[0197] Thus, for 20ms frames, Rt = 1.00461543.
[0198] For other variants: - The interval [0.9885, 1.0116] could take different values close to 1 (with always the first value < 1 and the second value > 1); - The update coefficient Rt relating to a gain <1 could be different from that associated with a gain >1; - The adaptation gains qdmx (where DMX = POC or PHA) could be updated
[0199] so that they converge according to a different law (linear profile or any
[0200] other function). - The qdmx adaptation gains could be updated so that they converge according to a different law (linear profile or any other function). - Adaptation gains could be applied and updated per block of samples or per sub-blocks of length less than the frame length. - qdmx adaptation gains could be applied and updated to the sample. After being applied on the downmix signal in the manner next:
[0201] GDMX(t, / ) x ty
[0202] where DMX = POC or PHA
[0203] The value of the adaptation gain could therefore be updated in the following manner (with Re the adaptation gain update coefficient): GDMX(tJ)^ Grxwx ( l, l - 1) / Re S1 GDMX > 1-
[0204] where DMX = POC or PHA
[0205] With Re the sample adaptation gain update coefficient. - If the adaptation gains are applied and updated to the sample, or by block of samples representing a duration less than the crossfade, their values could be fixed during the crossfade period (value calculated according to the terms of block 605) then updated only at the end of this period.
[0206] In the preferred embodiment, for a downmix the adaptation and crossfade gains are applied separately. A variant could combine the 2 gains into 1 single gain in order to apply it at once for example.
[0207] Beyond the SGC processing module (block 501), there is an influence of the crossfade module (block 502) on the quality of transition between 2 downmixes, during a change of mode in the current frame.
[0208] According to the invention, for signals that do not contain a transient, i.e. for is_transient = 0 (“False”) at the output of block 203, the crossfade duration in the current frame is longer than for the case where is_transient = 1 (“True”). This makes it possible to minimize the impact of the crossfade for transients, such as attacks (music) or plosives (speech).
[0209] In a preferred embodiment (as illustrated in [Fig.9a]), the duration of the crossfade in 502 is 60ms for a frame without transient and 20ms for a frame with transient.
[0210] Variants may propose other values for these 2 crossfade durations or even create additional categories of signals which would require different durations.
[0211] According to the invention, since the duration of the cross-fading can extend beyond one frame, it is entirely possible for a change in downmix method to occur before the end of a cross-fade.
[0212] According to the invention (as illustrated in [Fig.9b]), in the case of a crossfade already in progress, to avoid any artifact, it is advantageous to keep the current fading gains for each of the downmixes but to modify the direction of the fade.
[0213] For example, in the case of a linear crossfade, for the previous frame (t-2), a change from DMX2 to DMX1 was required, so the fade was in for DMX1 and out for DMX2. A downmix change, from DMX1 to DMX2, is detected for frame t. Thus the direction of the fade changes, it is out for DMX1 and in for DMX2. But, the crossfade (triggered in frame t-2) not having reached its end, the values of the fade gains Gf for DMX1 and q _ qj for DMX2 are kept to start the new fade.
[0214] The duration of the new crossfade is determined by the type of signal (transient or not) as determined by a classification or detection of the type of signal contained in the frame where a downmix change is triggered. For example (as illustrated in [Fig.9c]), if the initial duration of the current crossfade was 60ms, and a transient is detected in the frame that triggered the downmix change, the duration of the crossfade becomes 20ms.
[0215] More precisely, as shown in [Fig.9c], the cross-fade of duration 60 ms which starts at frame t-2 stops at the end of frame t-1 since frame t undergoes a downmix change. For this frame t, a fade is applied for a duration of 20 ms. Thus, a cross-fade of a predetermined duration may not be effective until the end if a new transition occurs before this duration.
[0216] In variants, the duration of the crossfade in the event of a transition in progress may be a function of the categories of incoming-outgoing signals and their direction with associated rules. For example (for durations also given as examples), during a downmix change during a current fading initially of 60ms, a transient would impose a crossfade duration of 20ms but the reverse would not be true, i.e. during a current fading initially of 20ms, a non-transient signal would not impose a crossfade duration of 60ms but would retain the initial 20ms.
[0217] According to the preferred embodiment, and for memory size reduction and complexity considerations, the duration of the crossfade for the transients (for example 20ms) can be a multiple of the duration of the crossfade for the other signals (for example 60ms) having a factor of N (where N would be equal to 3 in the previous example). This makes it possible to use a single precalculated fade gain table (in the preferred embodiment, the table contains the precalculated fading gains corresponding to the incoming fade over 20ms), which can be directly used for the transient signals, and to repeat the same gain from the table N times for N successive samples of the other types of signals.
[0218] Other variants of the invention are possible, in particular in the case of downmix methods other than POC and PHA.
[0219] According to the preferred embodiment, the invention applies to a downmix which applies to stereo input signals. In variants, the downmix processing may be applied at the output of a stereo audio system (decoder or other) to ensure mono reproduction for example.
[0220] According to the preferred embodiment, the invention relates to downmix methods from stereo to mono, but variants could implement downmix methods (N channels to mono) other than stereo (2 channels to mono), even varying the downmix methods within the same system as long as the output signals are monophonic.
[0221] The invention applies to a transition between two channel reduction processing modes with optimized level control. In variants, the number of processing modes could be greater than 2; For example, if the method comprises 3 downmix methods, such as POC, PHA and a third which would be, for example, a simple downmix of type (L+R) / 2, assuming that the decision block is adapted accordingly, the invention provides for defining an additional adaptation gain for the third downmix according to the methods already described.
[0222] [Fig.8] illustrates a processing device 800, within the meaning of the invention.
[0223] The device 800 comprises a processing circuit typically including: - a memory MEM1 for storing instruction data of a computer program within the meaning of the invention; - an INT 1 interface for receiving a stereo audio signal; b^) - a processor PROC1 for receiving this signal and processing it by executing the computer program instructions stored in the memory MEM1, with a view to processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal; in particular, the processor being capable of controlling the processing modules as described with reference to figures 5 to 7; and - a COM 1 communication interface for transmitting the reduced signals, resulting from the channel reduction processing, mcwt( t, / ), to another processing module, for example a coding or decoding module of an audio signal coder or decoder.
[0224] Of course, this [Fig.8] illustrates an example of a structural embodiment of a treatment device within the meaning of the invention.
[0225] Figures 5 to 7 commented above describe in detail functional embodiments of this device.
[0226] Applications of this type of channel reduction and transition processing between Two channel reduction processing modes are found, for example, in audio coding, for example when the remote party does not have the capacity to reproduce stereo sound on their terminal. In this case, it is not necessary to transport a stereo signal and this saves bandwidth. This type of process can occur during point-to-point conversations if one of the participants makes a stereo or binaural sound recording. This type of process can also be present during multi-party conferences: the conference bridge spatializes the scene by creating a stereo scene, but not all participants necessarily have stereo reproduction capabilities. It is then appropriate to downmix the scene in mono for these participants.
[0227] The invention can also be applied for audio decoding: in this case, it is the user's renderer (restitution module) which will perform the downmix of the stereo content to adapt to the user's reduced capacities (single loudspeaker for example).
Claims
Claims
1. Method for processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal, the processing method comprising the following steps: - application of a determined adaptation gain (501), to the signal resulting from the channel reduction processing according to the second mode; - application of a crossfade (502), during a transition duration, between the signal resulting from the channel reduction processing according to the first mode and the signal resulting from the reduction processing according to the second mode modified by the application of the adaptation gain; - progressive variation of the adaptation gain applied to the signal resulting from the second processing mode during the transition duration and / or beyond the transition duration until converging towards a value of one.
2. Method according to claim 1, in which the adaptation gain is determined from the respective energies of the signals resulting from the channel reduction processing according to the first and second modes.
3. A method according to claim 2, wherein the adaptation gain is bounded to threshold values.
4. Method according to one of the preceding claims, in which, in the case where a second transition between the second mode and the first channel reduction processing mode occurs before the end of the duration of progressive variation of the adaptation gain determined for the previous transition, then each of the signals resulting from the channel reduction processing of the first and second modes respectively are modified by the application of their respective adaptation gain for the application of the crossfade.
5. Method according to one of the preceding claims, in which the transition duration is defined as a function of a characteristic of the stereophonic signal.
6. Method according to one of the preceding claims, in which a first channel reduction processing mode is a processing mode by filtering and rephasing at least one of the channels and a second channel reduction processing mode is a processing without filtering using only a weighted sum of the two signals.
7. Method according to one of claims 1 to 5, in which a first channel reduction processing mode is a processing mode without filtering using only a weighted sum of the two signals and a second channel reduction processing mode is a processing by filtering and rephasing at least one of the channels.
8. Device for processing an audio signal during a transition between a first and a second channel reduction processing mode of a stereophonic signal to obtain a monophonic signal, the processing device comprising a processing circuit for implementing the following steps: - application of a determined adaptation gain, to the signal resulting from the channel reduction processing according to the second mode; - application of a crossfade, during a transition duration, between the signal resulting from the channel reduction processing according to the first mode and the signal resulting from the reduction processing according to the second mode modified by the application of the adaptation gain,; - progressive variation of the adaptation gain applied to the signal resulting from the second processing mode during the transition duration and / or beyond the transition duration until converging towards a value of one.
9. Storage medium, readable by a processor, storing a computer program comprising instructions for executing the processing method according to one of claims 1 to 7.
Citation Information
Patent Citations
Adaptive channel-reduction processing for encoding a multi-channel audio signal
US20190156841A1