Process for reducing channels of a stereophonic audio signal by optimised rephasing

ZA202606441APending Publication Date: 2026-07-29ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ZA202606441
Authority / Receiving Office
ZA · ZA
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2026-06-18
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing downmixing methods for stereophonic signals to mono signals suffer from issues like comb filtering, coloration, and excessive reverberation due to phase differences between microphones, leading to degraded signal quality and intelligibility.

Method used

A method for channel reduction processing that involves filtering and applying a phase shift to one of the channels of a stereophonic signal, where the channel for the phase shift is selected based on an energy ratio between the channels, to avoid degradations in signal quality and timbre changes.

Benefits of technology

This approach effectively reduces audible discontinuities and timbre changes by stabilizing the selection of the channel for rephasing, thereby improving the quality of the downmixed signal.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

NOT VISIBLE DUE TO STATUS OF PATENT
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION Title: Channel reduction processing of a stereophonic audio signal by optimized rephasing Technical Field The present invention relates to the general field of audio signal processing. The invention relates in particular to the channel reduction processing commonly called "downmix", of a multichannel audio signal. Of particular interest is the processing of reducing a stereophonic signal to a monophonic signal. This type of processing finds applications generally in the field of audio technologies, and more precisely in the field of audio coding, whether at the encoding or decoding stage. Prior art Channel reduction or downmix consists of deducing, from a combination of the C channels of a multichannel signal x, a signal y consisting of a smaller number D of channels. In practice, this consists of defining a function ^^(. ) such that: ^^1 ( ^^) ^^1 ( ^^ ) representative In this invention, we are interested in the particular case D=1 and C=2, we will then speak of stereo to mono downmix or simply mono downmix. In the case C=2, we speak of 2-channel content or signal which can be stereo content or binaural content, where channel 1 (^^1) is referenced as the left channel (Left), and channel 2 (^^2) as the right channel (for Right). Subsequently, the stereo case will be seen as a general 2-channel signal, including the binaural case to avoid repeating the two terms, even if technically a binaural signal has specific characteristics. In the case of the signal y obtained after reduction processing, we speak of a monophonic or mono signal, which will subsequently be noted m(n), in the time domain, or M(k), in the frequency domain. These stereo signals can come from a recording from a pair of stereo microphones or a binaural recording, or even from an artistic mix of audio tracks.There are several microphone pairs that can be used to create stereo content. Among the most popular are the XY pair, which consists of two cardioid microphones with a difference in orientation angle between 90° and 135°. The MS pair, for "Mid-Side," is another very common pair, consisting of a cardioid microphone and a second "figure-of-8" microphone oriented at 90°. By combining these two microphones (sum / difference), the left and right channels of stereo content can be created. These pairs are said to be coincident, meaning that the microphone capsules do not have any delay between them, the spatialization being perceived by the difference in sound intensity between the left and right channels. Another category of pairs, called phase stereo, consists of using microphones that are distant from each other, thus creating a phase difference between the channels for sources located closer to one of the microphones.The best known is the AB pair, which uses two omnidirectional microphones spaced a few centimeters to several meters apart. Another very common pair is the ORTF pair, which uses both a phase difference through a 17cm spacing of the microphones and an amplitude difference through cardioid directivities (the microphones have an angle around 90°). The binaural pair consists of placing omnidirectional microphones in the ears of a person or an artificial head: this type of device allows the native creation of binaural content that can be listened to with headphones. Stereo content also includes speech and audio signals from audio mixing or post-production (for example, channels stored on a CD or DVD, or broadcast on the Internet). There are also other methods for creating stereo content that are not reviewed here.The stereo signals at the input of the downmix can, for example, come from a transmission or storage system, for example, the downmix can occur after stereo decoding to allow monophonic reproduction. The simplest method for creating a signal reduced by downmixing is called passive downmixing. It consists of averaging the left and right channels of the stereo signal such as: ^^ (^^) + ^^ (^^). ^^ ^^ 1 2 or in the frequency domain: ^^(^^, ^^) =^^1(^^,^^)+^^2(^^,^^) where ^^^^ ^^ = 1,2 is the short-term Fourier transform of ^^ ^^ Fast Fourier Transform »), k being the frequency index and t the frame index:^^ ^^−1 ^^ ^^ ^̌^ ^^ ^^− ^^2^^^^^^ ^^ all ^^ ∈ {0,… , ^^ −} Où j is the frame index, l is the time index in the current frame, N is the size of the FFT, ^^(. ) an apodization window of size L, of sine or Hann type or other, adapted to the frame size, and ^^^^(^^, ^^) = ^^^^(^^, ^^). This definition extends to the case where N>L and ^^^^(. ) is an augmented version of ^^ ^^(^^) with 0s (0-padding). This very simple passive downmixing method works well in situations where the right and left channels are in phase. However, in situations where the microphones are not coincident and have phase differences, such mixing generates comb filtering, which manifests itself as a coloration of the original signal. This is due to the fact that depending on the microphone spacing, the position of the source relative to the microphone pair and the frequency, the left and right channels of the same source can be either in phase as shown in Figure 1a, or in phase opposition as shown in Figure 1b. Also, during downmixing, the left and right signals will either add together (constructive interference) or cancel each other out (destructive interference), thus creating level variations and coloration effects of the reduced signal ^^ ( ^^ ). The coloration comes from the effect of comb filtering: the effects of constructive and destructive interference appear in different frequency bands, thus modifying the balance and therefore the timbre of the signal from the downmix compared to the original. To correct this intensity level defect according to the frequencies, the downmix of the e-AAC+ codec implements a level correction ^^ ( ^^ ) which ensures that the energy level of the reduced signal, in each frequency band, remains comparable to the original one:^^ ^^ = ^^ ^^ .^^1(^^)+^^2(^^) This approach allows to compensate to a certain extent the decrease in intensity of the downmix. However, this compensation leads to overamplifications when the signals are close to the phase opposition: also, in practice, the correction factor ^^(^^) is limited to a maximum value (for example an upper limit of 2). These methods, in addition to the coloration of the signal from the downmix, linked to the imperfectly corrected comb filtering, suffer from an excess of reverberation linked to the summation of signals when the input channels are in phase shift, and from a "click" type artifact when the channel reinjection is performed punctually in a single frame. Compensating for an identical delay for all frequencies is not realistic, this delay being linked to the different sources composing the content.In practice, this leads, in addition to additional coloration, to a potential loss of intelligibility of the sources, a defect that level compensation cannot correct. Other downmixing approaches seek to avoid the aforementioned defects. The principle of these methods is to re-phase the left and right signals before their summation. In the document entitled "A stereo to mono downmixing scheme for MPEG-4 parametric stereo encoder" by Samsudin, E. Kurniawati, N. Boon Poh, F. Sattar, S. George, in Proc. ICASSP, 2006, a method is proposed which consists of applying, in the frequency domain, a phase shift φ on one of the channels, generally the right, in order to re-phase it with the left channel which is not modified, such as: ^^ (^^) + ( )^^^^(^^) 1^^ ^^ . ^^. ^^ ^^ 2 The channel whose phase is then called the "reference channel". The ideal phase shift to be applied is given by what is called here the IPD for "Inter-channel Phase Difference" in English: ^^^^^^(^^) = ∠(^^1(^^). ^^2 ∗( ^^ ) ) where ^^ ∗is the conjugate of ^^ and ∠ indicates the phase of the complex operand. It is calculated in the frequency domain: this allows the spectral components of different sources to be re-phased, a source with its own IPD being able to be predominant in a frequency band k, while another source with a different IPD can be predominant in another frequency k'. In the case of the MPEG4 codec described in the Samsudin document cited above, the IPD applied is that estimated by averaging over a Bark band and not the IPD calculated for each frequency band, the latter being particularly noisy and variable from one frame to another. In addition, this method takes the left channel as a phase reference, and if the phase of this channel is poorly conditioned, the downmix has degraded quality. The re-phasing by the IPD in the method described above can sometimes significantly modify the timbre and cause degradations in the signal quality.These timbre modifications can be particularly audible when the most energetic signal is the right channel. Disclosure of the invention The invention improves the state of the art. To this end, the invention relates to a method for processing channel reduction of a stereophonic signal to obtain a monophonic signal comprising filtering of the stereophonic signal and a phase adjustment applied to one of the channels of the stereophonic signal, the method being such that the channel on which the phase adjustment is applied is selected according to a selection criterion depending on a value representative of an energy ratio between the channels of the stereophonic signal and the selection of the channel on which the phase adjustment is applied, from one frame to another, is carried out if the energy ratio between the channels of the stereophonic signal reaches a significant value.Selecting the channel to which the phase shift is applied helps avoid degradations in the quality of the stereo signal, particularly changes in timbre. In situations where the left and right channels are of comparable energy, resulting in an oscillation of the energy dominance from one channel to the other depending on the variations in the signals, taking into account a significant value of the energy ratio helps avoid rapid changes of the reference channel, and therefore frequent changes in timbre. This energy ratio between the channels makes it possible to determine the energy differences between the channels and, for example, to select the channel for which the energy is the lowest to apply the rephasing and therefore preserve the phase of the other channel whose energy is the highest. Thus, if the phase shift filter generates a change in timbre, it will be masked by the unmodified channel, with higher energy.In one embodiment, the energy calculation is performed per frame or per subframe of the stereophonic signal. Thus, the selection of the channel on which the rephasing is applied is performed per frame of the stereophonic signal, the calculation of the energy per subframe makes it possible to increase the responsiveness for the change from one channel to another. In an alternative embodiment, the energy value of a channel of the stereophonic signal per frame or per subframe is smoothed. The smoothing of the energy thus makes it possible to avoid audible discontinuities during excessively rapid switching from one channel to another. In an alternative embodiment, a change of channel on which the rephasing is applied, from one frame to another, is further conditioned on a stability value of the energy ratio over a number of frames.This makes it possible to stabilize over time the selection of the channel to be re-phased, and thus to avoid rapid changes in the reference channel which are sources of audible discontinuities. In another embodiment, a selection of the channel on which the re-phase is applied is carried out according to a value representative of the shape of the phase shift filter, in the case where the energy ratio between the channels of the stereophonic signal is between two thresholds. This other selection mode makes it possible to appropriately select a channel on which the re-phase is to be applied even if the left and right channels are of comparable energy. The spectral content of the channel which is not selected for re-phase, i.e. the channel which serves as a reference, is then better preserved since it does not undergo any phase modification and its spectral content is preserved.In a second embodiment, the selection criterion is based on a value representative of the shape of the phase shift filter. This other selection criterion makes it possible to detect the channel to be rephased on which there is a small modification of the timbre. In a particular embodiment, during a change of channel on which the rephasing is carried out, a transition on at least one frame is carried out by applying filtering and rephasing on the two channels of the stereophonic signal. This makes it possible to make a smoother transition during a change of channel to be rephased, avoiding significant phase jumps.In a variant, during a change of channel on which the rephasing is performed, a transition on at least one frame is performed by applying both filtering and rephasing on the channel selected for the previous frame, filtering and rephasing on the two channels of the stereophonic signal and a crossfade between the two filterings. This transition mode avoids significant phase jumps, thanks to the interpolation of the phase change performed by the crossfade. In a particular embodiment, the value representative of the shape of the phase shift filter is a flatness value of the impulse response of the preprocessed or truncated filter. This value makes it possible to detect the most marked peaks on the impulse response and thus indicate whether the channel on which the rephasing is applied is the one which presents the least modification of timbre.In an alternative embodiment, a change of channel on which the rephasing is applied, from one frame to the next, is further conditioned on a flatness value of the amplitude spectrum of at least one of the two channels. This alternative embodiment makes it possible to avoid a change of reference channel during periods of high harmonicity of the signals where the slightest change in timbre is audible and annoying. The invention relates to a channel reduction processing device comprising a processing circuit for implementing the steps of the channel reduction processing method as described above. The invention also relates to a computer program comprising instructions for implementing the channel reduction processing method according to the invention, according to any one of the particular embodiments described above, when said program is executed by a processor.Such instructions may be stored durably in a non-transitory memory medium of the channel reduction processing device implementing the channel reduction processing method according to the invention. This program may use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form. The invention also relates to a recording medium or information medium readable by a computer, and comprising instructions of a computer program as mentioned above. The recording medium may be any entity or device capable of storing the program.For example, the medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a mobile medium, a hard disk or an SSD. Furthermore, the recording medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means, so that the computer program contained therein is remotely executable. The program according to the invention may in particular be downloaded over a network, for example an Internet-type network. Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the aforementioned channel reduction processing method.According to an exemplary embodiment, the present technique is implemented by means of software and / or hardware components. In this regard, the term "device" or "module" may correspond in this document to a software component, a hardware component or a set of hardware and software components. Brief description of the drawings Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which: [Fig. 1a] illustrates channels of a stereophonic signal, in phase as described previously; [Fig. 1b] illustrates channels of a stereophonic signal, in phase opposition as described previously; [Fig. 2] illustrates in block diagram form a channel reduction processing chain in one embodiment of the invention; [Fig. 3a], [Fig.3b] and [Fig. 3c] illustrate exemplary embodiments of the selection of the channel on which the rephasing is applied for the channel reduction processing for a first embodiment; [Fig. 4a] and [Fig. 4b] illustrate exemplary embodiments of the selection of the channel on which the rephasing is applied for the channel reduction processing for a second embodiment; [Fig. 5] illustrates an embodiment of an adaptation phase of a channel reduction processing filter; [Fig. 6] illustrates an exemplary structural embodiment of a channel reduction processing device according to an embodiment of the invention. Description of the Embodiments The invention will now be described below with reference to FIGS. 2 to 6. For this purpose, the following symbols will be mainly used in the description which follows.INDICES k frequency index l time index of a signal (in the current frame) n time index of a signal t frame number CONSTANTS L number of samples in the current frame R filter delay N FFT length C number of input channels D number of output channels P filter length Q crossfade length SIGNALS xi(n) input channels m(n) downmix (in time) m(t,l) downmix in the current frame (in time) M(t,k) downmix in the current frame (in frequency) Figure 2 shows an example of a processing chain for a stereophonic audio signal, the processing including a channel reduction process called "downmix".At the input of this processing chain, a stereophonic signal x, composed of two channels (x1(n) and x2(n)) also called left channel and right channel (respectively), is initially divided in 201, into frames of L samples (x1(t,l) and x2(t,l)), t being the frame index and l the sample index. Block 202 applies windowing and an FFT (for "Fast Fourier Transform" in English), to obtain signals in the frequency domain (X1(t,k) and X2(t,k)), k being the frequency index). In 203, a selection of the channel (203b) on which a rephasing is applied for the channel reduction processing is carried out according to a selection criterion calculated in 203a. This selection criterion may be based on a value representative of an energy ratio between the channels of the stereophonic signal and / or a value representative of the shape of the phase shift filter implemented for the channel reduction processing.Several exemplary embodiments will be described with reference to Figures 3a to 3c. At the end of step 203, the filter to be applied to the stereophonic signal is determined according to the selection of the channel on which the rephasing is performed. This filter (defined by filter coefficients in the form of a finite length impulse response) thus determined is conditioned (or adapted) in 204 to make it causal or partly causal and to optimize it in order to reduce the processing complexity. This adaptation phase (of the filter impulse response) is described with reference to Figure 5. As described later, these adaptation steps can also be used to determine the shape of the phase shift filter which can, in an exemplary embodiment, be a criterion for selecting the channel on which the rephasing is performed.Once the filter(s) have been determined and adapted for the current frame, they are applied at 205 to this current frame to obtain a mono signal ^̂^^^(^^, ^^). Figures 3a to 3c describe examples of embodiments of block 203 of Figure 2. In this processing block, the channel (^^^^(^^), with a value of 1 for the left channel and 2 for the right channel) is selected on which the rephasing is applied for a channel reduction processing method comprising filtering of the stereophonic signal and a rephasing applied to one of the channels of the stereophonic signal. In a first approach, it is proposed in one embodiment to apply the rephasing to the weakest signal: thus, if the rephasing filter generates a timbre modification, this will be completely or partially masked by the unmodified channel.For this, in this embodiment, the selection criterion is a function of a value representative of an energy ratio between the channels of the stereophonic signal. This value is for example the inter-channel level difference called ILD for “Interchannel Level Difference”, which measures the energy ratio between the left channel and the right channel such as: ^. ^^^^ ‖ ‖2 ^ ^^1(^^) Where ‖^^ ^^ (^^)‖ 2 is the channel energy Figure 3a illustrates this selection step, in a first implementation. At step E301, the energy ratio between the left channel and the right channel ^^^^^^ ( ^^ )is calculated according to the above formula for the current frame t. In step E302, this ratio is compared to a threshold value S1 which in this embodiment example is equal to 1. At each frame, the channel with the lowest energy strength is selected, i.e. depending on whether the ILD value is less than a threshold value of 1, to apply a phase adjustment in each frequency or frequency band based on the calculation of the IPD described previously, i.e. according to the expression: ^^^^^^(^^) = ∠(^^1(^^). ^^2 ∗( ^^ ) ) where ^^ ∗ is the conjugate of ^^ and ∠ indicates the phase of the complex operand. To limit the effect induced by rapid variations in the IPD, we can use a smoothed version of the IPD over time to calculate the filters ^^^^(^^, ^^). One way to smooth the IPD is to apply IIR type low-pass filtering such that: ̅^̅^ ̅^̅^^ ̅^(^^, ^^) = ^^(^^)̅^̅^ ̅^̅^^ ̅^(^^ − 1, ^^) + (1 − ^^(^^))^^^^^^(^^, ^^)The coefficient interesting to choose a high forgetting coefficient in the low frequencies, typically below 5kHz, because this allows to maintain a phase coherence in the low part of the spectrum from one frame to another, the spectrum of the speech signals being stationary, in particular in its low part and the particular ear sensitive to the phase in this part of the spectrum. Conversely, it is interesting to choose a low forgetting coefficient in the high frequencies: in the high frequencies, typically above 5kHz, the phase continuity of the speech signals is less marked and the ear is less sensitive to phase discontinuities in this part of the spectrum. We can choose a forgetting coefficient of the following form: ^^^^^^^^, ^^ < ^^^^^^^^The values ​​^^ ^^^^^^ , In the case of adjacent frames of length L=20ms, we can for example choose the following set of parameters:^^^^^^^^ = 0.94, ^^^^^^^^ = 0.86, ^^^^^^^^ = 0, ^^ℎ^^^^ℎ corresponds to 5^^^^^^. In practice, the calculation of the IPD requires an arctangent function which can be expensive. In a variant, we will calculate the complex exponential of the IPD, ^^^^^^^^^^(^^,^^), via the normalized inter-spectrum: ^ ^^^^^^^ ^^1(^^, ^^). ^^∗ 2(^^, ^^) ^^ Where |^^| is the modulus of the As with its direct version, we can also smooth the complex exponential of the IPD: ^^ ^ ̅ ^^ ̅ ̅ ^^ ̅ ̅ ^^ ̅ ^ (^^,^^) = ^^(^^)^^^̅^^̅ ̅^^̅ ̅^^̅^ (^^−1,^^) + (1 − ^^(^^))^^^^^^^^^^(^^,^^) In the remainder of the description, to simplify the notations and equations, the notation ^^^^^^^^^^(^^,^^) will define the exponential of the phase difference, obtained by direct calculation of the IPD or by the normalized inter-spectrum, in its smoothed or unsmoothed version. In this embodiment illustrated in Figure 3a, in the case where ^^^^^^(^^) < 1, the right channel is selected to be put back in phase with the left channel called "reference", i.e. channel ^^^^(^^) = 2 in E303. Otherwise, (N in E302), the left channel is selected to be put back in phase with the right channel called "reference", i.e. ^^^^(^^) = 1 in E304. In an alternative embodiment, another threshold value may be chosen. The filtering for channel reduction processing is then determined according to the following equations: −^^^^^^^^(^^,^^)^ ^^^^^ ^^ ^^ ^ ^ 1 The level between the initial stereophonic content and the monophonic downmix. Different normalization values ​​can be taken depending on the implementations: in practice, the values ​​2 or √2 are often used (one being adapted to an energy average and the other to an amplitude average). The calculation of the energy of each channel on the current frame t can be done by a simple integration on the samples of the current frame: ^^−1 In practice, in coders are about ten milliseconds in size. With such short frame durations, the ILD calculated in this way can lead to rapid switching from one channel to another, resulting in audible discontinuities. To avoid this, the energies of each frame can be smoothed with a recursive low-pass filter of 1 er order, as follows: ^^−1 Where ^^ is a factor of 1 to avoid rapid switching from one channel to another. With frames around 20ms, such smoothing can however lead to poor responsiveness during a major channel change. To limit this lack of responsiveness, we can divide the frame into S sub-frames, and apply the smoothing to the energies of the sub-frames: ^^∗ ^^^^ ^ ^ −1 Avec la L‘énergie The comparison of the energy ratio of the stereo signal channels is carried out according to the calculation per frame (^^1 ( ^^ ) / ^^2 ( ^^ )or vice versa) or according to the calculation on the last subframe of the current frame (^^2(^^, ^^) / ^^2(^^, ^^) or vice versa). This choice of predominant channel based on the ILD thus calculated makes it possible to avoid audible timbre changes in most cases where the signals are unbalanced in energy but nevertheless present a certain correlation, as in the case, for example, of binaural signals. However, despite the proposed smoothing of the energies, the choice of the predominant channel can lead to rapid and repeated switching when the left and right channels are of comparable energies, which can result in audible discontinuities such as clicking. To avoid this, we can limit the choice of the channel to be rephased to periods where the level difference is significant: −^^^^^^^ ( )^ ^^^^^ ^^ 1 ^ ^ ^^ ^^ ^ ^^,^^ ^ ^ ^^ 1 ^^^^^^(^^) <1: ^^^^(^^) = 1 one in decibels, typically 3dB, √2 in linear. This exemplary embodiment is illustrated in Figure 3b, where in step E301, the energy ratio between the left channel and the right channel ^^^^^^(^^) is calculated. This ratio is compared to a first threshold S2 (1 ^ ^ ^^^^^^ ) in E305. In the case where the ratio is lower than this threshold then the left channel is selected for rephasing, i.e. ^^^^(^^) = 1 in E306, the rephasing being carried out by applying to the left channel the filter =^^−^^^^^^^^(^^,^^) the right channel phase, called "reference" in this case, remaining is potentially modified by the normalization coefficient ^^). Otherwise, the ratio is compared to a second threshold S3 (^^ ^^^^^^) in E307. In the case where the ratio is higher than this threshold then the right channel is selected for rephasing, i.e. ^^^^(^^) = 2 in E308, the rephasing being carried out by applying to the right channel the filter ^^2(^^, ^^) =^^^^^^^^^^(^^,^^)^^ , the left channel phase, called "reference" in this case, remaining unchanged (only the module is potentially modified by the normalization coefficient ^^). Otherwise, that is to say when the ratio is between the two thresholds, the channel which was selected for rephasing in the previous frame (^^^^(^^ − 1)), is again selected in E309, to which the rephasing filter is applied with the new IPD, i.e. ^^ (^^, ^^) =^^−^^^^^^^^(^^,^^) for the left channel ( )^^^^^^^^^^(^^,^^)1 ^^uche, or ^^2 ^^, ^^ =^^ for the right channel.The application of this decision freeze zone (combined with the temporal smoothing of parameters such as IPD, energies, etc.) makes it possible in practice to stabilize the choice of the channel to rephase and thus avoid audible discontinuities. To further stabilize the choice of channel, the switching from one channel to rephase to another can be made conditional when the ILD is stable over time. In practice, it can be decided to choose a channel as the reference channel if it remains the most energetic for a sufficient number of consecutive frames, typically of the order of several hundred milliseconds. Figure 3c illustrates this scenario. Step E301 remains unchanged compared to Figures 3a and 3b, the comparison with the first threshold S2 is carried out in E305 as in Figure 3b.In the case where the ratio is lower than this threshold then a step of incrementing a frame counter C1 is implemented in E310 and a verification step in E311 is carried out by comparing the frame counter C1 with a threshold T of a desired number of frames (corresponding for example to a duration of 100ms). If this counter C1 exceeds this threshold T, then the selection of the channel is authorized and the rephasing is carried out on the left channel (^^^^(^^) = 1), in E312. If the ratio is higher than the threshold S2 in E305 (N to E305), then the frame counter C1 is reset to 0 in E317. This counter C1 makes it possible to authorize the selection of a channel to be rephased as soon as this channel has been selected for a time T.In the case where the ratio is greater than the threshold S3 at E307, then a step of incrementing a frame counter C2 is implemented at E313 and a verification step at E314 is carried out by comparing this frame counter C2 with a threshold T of a desired number of frames (corresponding for example to a duration of 100 ms). If this counter C2 exceeds this threshold T, then the channel switch is authorized and the rephasing is carried out on the right channel (^^^^(^^) = 2), at E315. If the ratio is less than the threshold S3 at E307 (N at E307), then the. of frame C2 is reset to 0 in E318. Thus, this counter C2 makes it possible to authorize switching to a channel to be rephased as soon as this channel has been selected for a time T. In all other cases, (N in step E311, in step E314 and in step E307), the channel which was selected in the previous frame (^^^^(^^ − 1)) for rephasing, is again selected for the current frame in E316, and rephased with the new value ^^^^^^(^^) calculated for the current frame. This rule makes it possible to avoid switching very quickly from one channel to another, for example in the case of sources distributed on either side of the sound scene. In a second embodiment, the channel selection criterion is different. It is based on the shape of the phase shift filter implemented for the channel reduction processing.Indeed, in situations where the energy of the left and right channels remain very close to each other, no decision is made except to keep the channel of the previous frame. In the latter case, the decision to freeze the channel switching can sometimes lead to non-optimal situations. Indeed, when the left and right signals are of comparable energies, one can no longer rely on energy masking to mask the spectral modifications that can cause timbre modifications to appear. To avoid this, it is proposed in one embodiment to base the choice of the channel to be rephased on a minimal modification of the timbre, by choosing the channel whose filter best preserves the spectral content of said channel. In practice, it is known that filters whose impulse responses resembling a Dirac are the most neutral in terms of timbre modification, because they maintain phase coherence between the different frequencies.This is the case, for example, of a binaural signal where a source is positioned slightly to the left: it appears that the rephasing filter based on the IPD of the right channel ^^2(^^, ^^) is better conditioned than that of the left channel ^^1(^^, ^^), it has a more linear phase, which guarantees a low modification of the timbre, the frequencies not being out of phase with each other. In this second embodiment, the channel to be rephased is selected as the one whose impulse response of the rephasing filter has the most marked peak. To do this, we propose to use the "spectral flatness" function which measures the degree of flatness of a curve, as follows: ^^ ^. ^ Where P is the size causal and finite length) When SF(h) takes values ​​close to 0, this means that the filter h has one or more marked peaks, synonymous with a small modification of the timbre. Conversely, a value close to 1 indicates a spread filter and potentially a source of audible modification of the timbre. Thus, we propose to select the channel in the following manner: ̿ ̿ ^^−^^^^^^^^(^^,^^)1 where in detail with reference to Figure 5 later: ℎ̃ ^^ , ℎ̌ ^^ or ℎ̆ ^^ (the filters ℎ̌ ^^ and ℎ̆ ^^give the same degree of flatness). Figure 4a illustrates this embodiment. In step E401, the shape of the rephasing filter is determined. In the example described here, the degree of flatness of the rephasing filter is calculated for each of the left and right channels, i.e. the criteria ^^^^(̿ℎ̿1 ̿) and ^^^^(̿ℎ̿2 ̿). In E402, a comparison is made between the two calculated flatness values. In the case where the degree of flatness of the rephasing filter on the left channel is greater than the degree of flatness of the rephasing filter on the right channel, then the channel selected for rephasing is the channel ^^^^(^^) = 2 in E403 and the rephasing is carried out on the right channel ( ^^^^^^^^^^(^^, )2(^^, ^^) =^^ ^^^^ ). Otherwise, the channel selected for rephasing is channel ^^^^(^^) =1 in E404 and rephasing is performed on the left channel ( ^^1(^^, ^^) =^^−^^^^^^^^(^^,^^)^^ ). In the same way as for the selection by the ILD, a time delay by counter can also be set up to avoid switching too quickly from one predominant channel to the other. In another embodiment, the criterion based on the shape of the phase shift filter can be taken into account in the case where, for the embodiment described in figure 3b, the value of the energy ratio between the two channels of the stereophonic signal is between two thresholds: In this case, the the shape of the phasing filter is implemented as in Figure 4a. Figure 4b illustrates this embodiment where steps E301 to E308 correspond to the steps described with reference to Figure 3b and steps E401 to E404 correspond to the steps described with reference to Figure 4a but applied to signals for which the value of the ILD is between thresholds S2 and S3. Counters can be added (as in the case illustrated in Figure 3c) to each of the decisions to avoid switching too quickly from one channel to another. When switching from one channel to another (where the choice of the rephased channel is changed), a discontinuity may sometimes be audible.Also, in one embodiment, to avoid switching directly between a mode with left channel phasing to a mode with right channel phasing, a transition step is performed over a duration of one or more frames using another channel reduction processing mode, for example a mode where the rephasing is performed on both channels, for example with a rephasing with an angle of IPD / 2 or IPD / 4. This downmix with rephasing on both channels can be seen, in the frequency domain, as a filtering according to the following equation, with an angle of IPD / 4: ^^(^^, ^^) = ^^1(^^, ^^). ^^1(^^, ^^) + ^^2(^^, ^^). ^^2(^^, ^^)^−. ^^^^^^ ( ^^,^^ ) ^^^^^^ ( ^^,^^ ) ^ ^^ 4 ^^ ^^ 4 ^^(^^, ^^) = ^^1(^^, ^^). ^^1(^^, ^^) + ^^2(^^, ^^). ^^2(^^, ^^) ^^^^^^(^^,^ ) ( ) −^^ ^ ^^^^^^ ^^,^^ ^^ 2 ^^ ^^ 2 Thus, during a channel change on which the phase shift is applied, a transition on at least one frame is performed by applying filtering and phase shifting on the two channels of the stereophonic signal. In an alternative embodiment, during a channel change on which the phase shift is applied, a transition on at least one frame is performed by applying both filtering and phase shifting on the channel selected for the previous frame, filtering and phase shifting on the two channels of the stereophonic signal and a crossfade between the two filtering operations. This transition mode avoids significant phase jumps, thanks to the interpolation of the phase change performed by the crossfade. If the use of an intermediate mode makes it possible to globally mask the discontinuities, the channel change proposed in the described embodiments can nevertheless become audible.This discontinuity is particularly evident if the change occurs over a period when one of the signals ^^. ^^ ( ^^ ) is highly harmonic: in this case, the slightest change in coloration becomes particularly audible. This is the case for voiced periods for a speech signal, or generally musical signals. Also, to avoid these audible changes in coloration, in one embodiment, it is proposed to restrict the transitions from one channel to be rephased to the other at times when the signal(s) do not present strong harmonicity. To detect this harmonicity, a criterion based on the degree of flatness ("spectral flatness") described above is used, applied to the amplitude spectrum of the left or right channel, or a combination of the degrees of flatness calculated on each channel: ^^√∏^^ (^^,^^) this criterion on the power spectrum of at least one of the channels. A low value is synonymous with a signal having a low number of frequency lines, and therefore a high harmonicity. Also, in the case of a low value of ^^^^(^^^^(^^, ^^)), ^^ = 1.2, this embodiment proposes not to allow the transition from one channel to be rephased to the other. The decision can be taken on one of the values ​​of ^^^^(^^1(^^, ^^)), ^^ = 1.2, or on a combination of ^^^^(^^^^(^^, ^^)). An optimal decision consists of basing the decision on the ^ m^=1in,2(^^^^(^^^^(^^, ^^)) , that no channel is highly harmonic. To limit the calculations, we can base ourselves on the degree of flatness ^^^^(^^^^(^^, ^^)) of an arbitrarily chosen channel, which avoids the calculation of the degree of flatness of the other channel. In practice, the spectral flatness taking its values ​​in the interval [0; 1], we can compare the chosen decision criterion, i.e. one of the spectral flatnesses ^^^^(^^^^(^^, ^^)), the^ m ^=1 in ,2(^^^^(^^^^(^^, ^^))), or any other combination of ^^^^(^^^^(^^, ^^)), at a threshold ^^^^^^^^ ∈ [0; 1]. A threshold value of the order of 0.05 seems a good compromise. Figure 5 now details block 204 of Figure 2. The filters ^^^^(^^, ^^), as defined above according to the channel on which the phase shift is placed are reconditioned or even adapted to avoid circular convolution. In a first step, in 501, we calculate the impulse responses ℎ^^(^^, ^^) of the filters ^^^^(^^, ^^), such that:ℎ^^(^^, ^^) = ^^^^^^−1{^^^^(^^, ^^)}, ^^ = 1,2 where the filters ℎ^^(^^, ^^) are In the preferred embodiment, we have L = 20 ms, i.e. K = 320 samples at 16 kHz, 640 samples at 32 kHz and 960 samples at 48 kHz, ^^^^ = 2, i.e. B = 160 samples at 16 kHz, 320 samples at 32 kHz and 480 samples at 48 kHz. The impulse responses thus defined are classically non-causal, they correspond to filters with a finite impulse response. A classic technique to make them causal consists of performing a rotation of ^^ / 2 samples to center the response at ^^ / 2. The disadvantage is that the filter thus reconstructed then has a latency of ^^ / 2. In the preferred embodiment, a truncation of the impulse response will be carried out: =ℎ^^(^^, ^^ − ^^), ^^ ≤ ^^ < ^^^^ (^^, ^^ − ^^ + ^^) 0 ≤ ^^ < ^^This operation of (circular) can also be implemented by applying an appropriate phase term to ^^^^(^^, ^^) before inverse FFT, as known to those skilled in the art. In this embodiment, the resulting filter has a finite impulse response with a delay of R samples. This delay allows the truncated filter to be properly conditioned, either in terms of gain by minimizing the influence of the truncation on the unity gain of the phase-shifting filter, or in phase by minimizing the deviation from the IPD. This delay R can be adapted according to the implementation constraints: in practice, a delay of around one millisecond is sufficient to properly condition the truncated filter. In an extreme variant, the truncation is asymmetric to avoid this latency. We choose to truncate the filter asymmetrically, at 502, by keeping only a portion of the first half of the impulse response ℎ^^(^^, ^^) in the following way: ℎ̃^^(^^, ^^) = ℎ^^(^^, ^^), 0 ≤ ^^ < ^^ ≤^^ ⁄This filter presents (in absolute value) on its first sample, which guarantees processing without additional delay. In terms of frequency response, it has a phase roughly equal to the optimal rephasing filter ℎ^^(^^, ^^). The fact that the anticausal part of the impulse response is suppressed can possibly be compensated for by additionally applying a factor of 2 to the values ​​of ℎ̃^^(^^, ^^) for ^^ > 0. In this variant, the resulting filter has a finite impulse response with virtually no delay. The choice of P depends on various criteria: a value of P close to N⁄2 guarantees near-optimal rephasing, while a smaller value limits the computing power required for filtering. A smaller value of P also smooths the phase of the filter.This has an advantage with filters defined by the channel reduction processing as defined above where the phase varies very rapidly from one frequency to another: these variations can result in a very high group delay perceptible to the human ear. A low value of P makes it possible to remove all or part of these artifacts. The filter thus truncated ℎ̃^^(^^, ^^), in particular when P <N⁄2, voire P≪N⁄2, présente unequeue de réponse impulsionnelle qui ne tend pas vers 0. Cela crée des discontinuités audibles de type cliquetis à la transition entre 2 trames. Aussi, pour éviter ces artefactsaudibles, on pondère ℎ̃^^(^^, ^^) avec une fenêtre ^^(^^), en 503, tel que :ℎ̌(^^, ^^) = ^^(^^). ℎ̃^^(^^, ^^)où les coefficients de ^^(^^) tendent vers 0 lorsque l tend vers 0 ou P et vaut 1 en l=R. Cela permet de réduire l’énergie des échantillons de la fin du filtre et d’éviter des discontinuités lors de la transition entre des filtres de trames consécutives.In the preferred embodiment, where the truncation preserves an anti-causal part, we apply for example ½ triangular windows on the causal and anti-causal parts ^^ or any other window of. In the variant where the truncation is asymmetric, this window can be a half-triangular window, or even a half-Hann window: ^^(^^) = 1 − ^^⁄ ^^ , ^^^^ cos(^^^^⁄ 2^^ )Finally, in the case where we only keep P samples from the causal part of the filter, we can also use a combination of a rectangular window of size ^^ ^^ , concatenated with a half Hann window of size P − ^^^^ 1 ^^ = 0 1The A value in the interval ^^ ∈ [1,2]. Typically, a value ^^ = 1.8 is a good compromise. Due to its truncation and windowing, the filters ℎ̃^^(^^, ^^) have a reduced energy compared to that of the optimal filter ℎ^^(^^, ^^): the latter, as a pure phase-shifting filter, has a norm ½ by construction. Also, we can proceed to a normalization of the energy of the filter thus truncated and windowed, in 504, such that: ℎ̌^^(^^, ^^) ^^ ^^ Where ‖^^(. )‖ is the 2-norm of a Alternatively, another method can be used in 204 and the method in Figure 5 can be replaced. For example, the impulse response can be determined directly by least squares minimization and matrix inversion using the (complex) Levinson algorithm as described in section 2 of the article by Mathias C. Lang, Design of nonlinear phase FIR digital filters using quadratic problems, Proc. ICASSP, 1997. An implementation of this method can be found, for example, in the thesis by Mathias C. Lang, Algorithms for the Constrained Design of Digital Filters with Arbitrary Magnitude and Phase Responses, June 1999 (see routine levin.m in Appendix B and lslevin.m in section 2.1.4). In this case, the inverse FFT, truncation, and weighting are not necessary; This alternative method directly estimates a finite impulse response of length P (without a minimum phase guarantee). However, a normalization in the form ℎ can be applied. ̌(^ )^^ ^^ ^, ^^ Illustrated in Figure 6 is a channel reduction processing device 600, within the meaning of the invention. The device 600 comprises a processing circuit typically including: - a memory MEM1 for storing instruction data of a computer program within the meaning of the invention; the computer program instructions stored in the memory MEM1, with a view to carrying out channel reduction processing; in particular, the processor being able to control the processing modules as described with reference to figures 2 to 5; and - a communication interface COM 1 for transmitting the reduced signals, resulting from the channel reduction processing, ^̂^(^^, ^^), to another processing module, for example a coding or decoding module of an audio signal coder or decoder. Of course, this figure 6 illustrates an example of a structural embodiment of a channel reduction processing device within the meaning of the invention. Figures 2 to 5 commented on above describe in detail functional embodiments of this device. The applications of this type of channel reduction processing are found for example in audio coding, for example when the remote does not have the capacity to reproduce stereo sound on its terminal.In this case, it is not necessary to transport a stereo signal and this saves bandwidth. This type of process can be used during point-to-point conversations if one of the participants makes a stereo or binaural sound recording. This type of process can also be present during multi-party conferences: the conference bridge spatializes the scene by creating a stereo scene, but not all participants necessarily have stereo reproduction capabilities. It is then appropriate to downmix the scene in mono for these participants. The invention can also be applied to audio decoding: in this case, it is the user's renderer (reproduction module) that will downmix the stereo content to adapt to the user's reduced capabilities (single speaker for example).

Claims

CLAIMS 1. Method for processing channel reduction of a stereophonic signal ( (x1(n), x2(n)) pour obtenir un signal monophonique (^̂^^^(^^, ^^)) comprenant un filtrage du signal stéréophonique ( ^^1(^^, ^^), ^^2(^^, ^^) en et ℎ1(^^), ℎ2(^^) en temporel) et une appliquée sur un des canaux of the stereophonic signal, in which the channel on which the rephasing is applied is selected (203b) according to a selection criterion (203a) depending on a value representative of an energy ratio between the channels of the stereophonic signal and the selection of the channel on which the rephasing is applied, from one frame to another, is carried out if the energy ratio between the c anaux du signal stéréophonique atteint une valeur significative (> ^^^^^^^^ ou < 1 ⁄ ^^^^^^^^ ).

2. Method according to claim 1, wherein the channel selected for rephasing is the one for which the energy is the lowest.

3. Method according to one of claims 1 to 2, wherein the calculation of the energy is carried out per frame or per subframe of the stereophonic signal.

4. Method according to claim 3, wherein the energy value of a channel of the stereophonic signal per frame or per subframe is smoothed.

5. Method according to one of claims 1 to 4, wherein a change of channel on which the rephasing is applied, from one frame to another, is further conditioned on a stability value of the energy ratio over a number of frames. 6.Method according to one of claims 1 to 5, in which a selection of the channel on which the rephasing is applied is carried out according to a value representative of the shape of the phase shift filter, in the case where the energy ratio between the channels of the stereophonic signal is between two thresholds.

7. Method according to one of the preceding claims, in which, during a change of channel on which the rephasing is carried out, a transition on at least one frame is carried out by applying filtering and rephasing on the two channels of the stereophonic signal.

8. Method according to one of claims 1 to 6, in which, during a change of channel on which the rephasing is carried out, a transition on at least one frame is carried out by applying both filtering and rephasing on the channel selected for the previous frame, filtering and. rephasing on the two channels of the stereophonic signal and a crossfade between the two filterings.

9. Method according to claim 6, in which the value representative of the shape of the phase shift filter is a flatness value of the impulse response of the preprocessed or truncated filter.

10. Method according to one of the preceding claims, in which a change of channel on which the rephasing is applied, from one frame to another, is further conditioned on a flatness value of the amplitude spectrum of at least one of the two channels.

11. Channel reduction processing device comprising a processing circuit for implementing the steps of the channel reduction processing method according to one of claims 1 to 10.

12. Storage medium, readable by a processor, storing a computer program comprising instructions for executing the method according to one of claims 1 to 10.