Channel reduction processing of a stereophonic audio signal by optimized rephasing

The method addresses the challenges of channel reduction by selectively rephasing channels based on energy ratios, enhancing signal quality and reducing timbre changes in stereophonic audio downmixing.

FR3157767A1Inactive Publication Date: 2025-06-27ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
FR2023015175
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing channel reduction methods, such as downmixing of stereophonic audio to monophonic, suffer from issues like comb filtering, level variations, and timbre changes due to phase differences between audio channels, leading to degraded signal quality and potential loss of intelligibility.

Method used

A method for channel reduction processing that involves filtering and rephasing of stereophonic signals, where the channel for rephasing is selected based on an energy ratio criterion to minimize timbre changes and preserve signal quality.

Benefits of technology

The proposed method effectively reduces timbre changes and maintains signal quality by selectively rephasing channels based on energy ratios, thereby improving the accuracy of channel reduction processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Channel reduction processing of a stereophonic audio signal by optimized rephasing The invention relates to a method for channel reduction processing of a stereophonic signal (x1(n), x2(n)) to obtain a monophonic signal () comprising filtering of the stereophonic signal () and rephasing applied to one of the channels of the stereophonic signal, the method being such that a step of selecting (203b) the channel on which the rephasing is applied is carried out according to a selection criterion (203a). The invention also relates to a channel reduction processing device implementing the method. Abstract figure: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Channel reduction processing of a signal stereophonic audio by optimized phase resetting Technical field

[0001] The present invention relates to the general field of audio signal processing. The invention relates in particular to the channel reduction processing commonly called "downmixing" of a multichannel audio signal. Of particular interest is the processing of reducing a stereophonic signal to a monophonic signal.

[0002] This type of processing finds applications generally in the field of audio technologies, and more specifically in the field of audio coding, whether at the encoding or decoding stage. Prior art

[0003] Channel reduction or downmixing consists of deducing, from a combination C channels of a multichannel signal x, a signal y consisting of a smaller number D of channels. In practice, this consists of defining a function f (.) such that:

[0004] ' Tl ) \ , with D < C and n representing m(n) .xc(n) ! a time index of the input and output signal.

[0005] In this invention, we are interested in the particular case D=1 and C=2, we will then speak of stereo to mono downmix or simply mono downmix. In the case C=2, we speak of 2-channel content or signal which can be stereo content or binaural content, where channel 1 (Jt) is referenced as the left channel (Left), and channel 2 (x2) as the right channel (for Right). Subsequently, the stereo case will be seen as a general 2-channel signal, including the binaural case to avoid repeating the two terms, even if technically a binaural signal has specific characteristics. In the case of the signal y obtained after reduction processing, we speak of a monophonic or mono signal, which will subsequently be noted m(n), in the time domain, or M(k), in the frequency domain.

[0006] These stereo signals can come from a capture from a pair of stereo microphones or a binaural capture, or even from an artistic mix of audio tracks.

[0007] There are a number of microphone pairs that can be used to create stereo content. Among the most popular are the XY pair, which consists of two cardioid microphones with a difference in orientation angle between 90° and 135°. Another very common pair is the MS pair, which stands for "Mid-Side", which consists of a cardioid microphone, and a second microphone called a "figure of 8" oriented at 90°. By combining these two microphones (sum / difference), we can create the left and right channels of stereo content. These pairs are said to be coincident, that is to say that the microphone capsules do not have any delay between them, the spatialization being perceived by the difference in sound intensity between the left and right channels.

[0008] Another category of so-called phase stereo pairs consists of using microphones that are distant from each other, thus creating a phase difference between the channels for sources located closer to one of the microphones. The best known is the AB pair which uses 2 omnidirectional microphones spaced a few centimeters to several meters apart. Another very widespread pair is the ORTF pair which uses both a phase difference through a 17cm spacing of the microphones and an amplitude difference through cardioid directivities (the microphones have an angle around 90°). The binaural pair consists of placing omnidirectional microphones in the ears of a person or an artificial head: this type of device makes it possible to natively create binaural content that can be listened to through headphones.

[0009] Stereo content also includes speech and audio signals resulting from audio mixing or post-production (e.g., channels stored on a CD or DVD, or streamed over the Internet).

[0010] There are also other methods of creating stereo content that are not reviewed here. The stereo signals at the input of the downmix can, for example, come from a transmission or storage system; for example, the downmix can occur after stereo decoding to allow monophonic reproduction.

[0011] The simplest method for creating a downmixed signal is known as passive downmixing. It involves averaging the left and right channels of the stereo signal such that:

[0012]

[0013] or in the frequency domain:

[0014] tJc)+x2(tJc) where Xy{t, k), / =1,2 is the Fourier transform at short-term of x^t) (or FFT in English for “Fast Fourier Transform”), k being the frequency index and t the frame index: k) = 12^00-^(^ 0^"^' For all / this {0, ..., N-1}

[0015] Where j is the imaginary number such that j — [, t is the frame index, l is the index temporal in the current frame, Via size of the FFT, w(.) an apodization window of size L, of type sine or Hann or other, adapted to the size of the frame, and Xf(t, l) — x{t, l)- This definition extends to the case where N>L and Xz(.) is a version augmented by x(f) with 0 (O-padding).

[0016] This very simple passive downmixing method works well in situations where the right and left channels are in phase. However, in situations where the microphones are not coincident and have phase differences, such mixing generates comb filtering, which manifests itself by a coloration of the original signal. This is due to the fact that depending on the spacing of the microphones, the position of the source relative to the microphone pair and the frequency, the left and right channels of the same source can be either in phase as shown in Figure 1a, or in phase opposition as shown in Figure 1b.Also, during downmixing, the left and right signals will either add up (constructive interference) or cancel each other out (destructive interference), thus creating level variations and coloration effects of the reduced signal m(t)- The coloration comes from the effect of comb filtering: the effects of constructive and destructive interference appear in different frequency bands, thus modifying the balance and therefore the timbre of the signal from the downmix compared to the original.

[0017] To correct this intensity level defect according to the frequencies, the downmix of the e-AAC+ codec implements a level correction p(k) which ensures that the energy level of the reduced signal, in each frequency band, remains comparable to the original one: [00i8] ™ Y0.5|Xt(À)+X2(Æ)|2

[0020] omitting the frame index t to simplify the notations.

[0021] This approach allows to compensate to a certain extent the decrease in intensity of the downmix. However, this compensation leads to overamplifications when the signals are close to phase opposition: also, in practice, the correction factor y(k) is limited to a maximum value (for example an upper limit of 2).

[0022] These methods, in addition to the coloration of the signal from the downmix, linked to the imperfectly corrected comb filtering, suffer from an excess of reverberation linked to the summation of signals when the input channels are out of phase, and from a "click" type artifact when the channel reinjection is carried out punctually in a single frame. Compensating for an identical delay for all frequencies is not realistic, this delay being linked to the different sources making up the content. In practice, this leads, in addition to additional coloration, to a potential loss of intelligibility of the sources, a defect that level compensation cannot correct.

[0023] Other downmix approaches seek to avoid the aforementioned defects. The principle of these methods is to re-phase the left and right signals before their summation. In the paper "A stereo to mono downmixing scheme for MPEG-4 parametric stereo encoder" by Samsudin, E. Kurniawati, N. Boon Poh, F. Sattar, S. George, in Proc. ICASSP, 2006, a method is proposed which consists of applying, in the frequency domain, a phase shift q> on one of the channels, generally the right, in order to put it back in phase with the left channel which is not modified, such that:

[0024] M{k) =

[0025] The channel whose phase remains unchanged, here the left channel, is then called the “reference channel”. The ideal phase shift to be applied is given by what is called here the IPD for “Inter-channel Phase Difference” in English:

[0026] IPD(k)= ^X^X^k))

[0027] where y* is the conjugate of X and L indicates the phase of the complex operand.

[0028] It is calculated in the frequency domain: this allows the spectral components of different sources to be re-phased, a source with its own IPD being able to be predominant in a frequency band k. while another source with a different IPD can be predominant in another frequency k'.

[0029] In the case of the MPEG4 codec described in the Samsudin document cited above, the IPD applied is that estimated by averaging over a Bark band and not the IPD calculated for each frequency band, the latter proving to be particularly noisy and variable from one frame to another. In addition, this method takes the left channel as phase reference, and if the phase of this channel is poorly conditioned, the downmix has degraded quality.

[0030] The rephasing by the IPD in the method described above can sometimes significantly change the timbre and cause degradations in the signal quality. These timbre changes can be particularly audible when the most energetic signal is the right channel. Statement of the invention

[0031] The invention improves the state of the art.

[0032] To this end, the invention relates to a method for processing channel reduction of a stereophonic signal to obtain a monophonic signal comprising filtering of the stereophonic signal and rephasing applied to one of the channels of the stereophonic signal, the method being such that a step of selecting the channel to which the rephasing is applied is carried out according to a selection criterion.

[0033] Selecting the channel on which the phase shift is applied makes it possible to avoid degradation of the quality of the stereo signal, in particular changes in timbre.

[0034] In a first embodiment, the selection criterion is a function of a value representative of an energy ratio between the channels of the stereophonic signal.

[0035] This energy ratio between the channels makes it possible to determine the energy differences between the channels and, for example, to select the channel for which the energy is the lowest to apply the phase shift and therefore preserve the phase of the other channel whose energy is the highest. Thus, if the phase shift filter generates a timbre modification, it will be masked by the unmodified channel, with higher energy.

[0036] In one embodiment, the energy calculation is performed per frame or per subframe of the stereophonic signal.

[0037] Thus, the selection of the channel on which the rephasing is applied is carried out per frame of the stereophonic signal, the calculation of the energy per sub-frame makes it possible to increase the reactivity for the change from one channel to another.

[0038] In an alternative embodiment, the energy value of a channel of the stereophonic signal per frame or per subframe is smoothed.

[0039] Energy smoothing thus makes it possible to avoid audible discontinuities during excessively rapid switching from one channel to another.

[0040] In a particular embodiment, the selection of the channel on which the rephasing is applied, from one frame to another, is only effective if the energy ratio between the channels of the stereophonic signal exceeds a significant difference threshold.

[0041] Thus, in situations where the left and right channels are of comparable energy, which results in an oscillation of the energy dominance from one channel to the other depending on the variations in the signals, this makes it possible to avoid rapid changes of reference channel, and therefore frequent changes in timbre.

[0042] In an alternative embodiment, a change of channel on which the rephasing is applied, from one frame to another, is further conditioned on a stability value of the energy ratio over a number of frames.

[0043] This makes it possible to stabilize over time the selection of the channel to be re-phased, and thus to avoid rapid changes in the reference channel which are sources of audible discontinuities.

[0044] In another embodiment, a selection of the channel on which the rephasing is applied is carried out according to a value representative of the shape of the phase shift filter, in the case where the energy ratio between the channels of the stereophonic signal is between two thresholds.

[0045] This other selection mode makes it possible to appropriately select a channel on which the rephasing is to be applied even if the left and right channels are of comparable energy. The spectral content of the channel which is not selected for rephasing, i.e. the channel which serves as a reference, is then better preserved since it does not undergo any phase modification and its spectral content is preserved.

[0046] In a second embodiment, the selection criterion is based on a value representative of the shape of the phase shift filter.

[0047] This other selection criterion makes it possible to detect the channel to be rephased on which there is a slight modification of the timbre.

[0048] In a particular embodiment, during a change of channel on which the rephasing is performed, a transition on at least one frame is performed by applying filtering and rephasing on the two channels of the stereophonic signal.

[0049] This allows for a smoother transition when changing the channel to be rephased, avoiding significant phase jumps.

[0050] In a variant, during a channel change on which the rephasing is performed, a transition on at least one frame is performed by applying both filtering and rephasing on the channel selected for the previous frame, filtering and rephasing on the two channels of the stereophonic signal and a crossfade between the two filterings.

[0051] This transition mode avoids significant phase jumps, thanks to the interpolation of the phase change carried out by the crossfade.

[0052] In a particular embodiment, the value representative of the shape of the phase shift filter is a flatness value of the impulse response of the preprocessed or truncated filter.

[0053] This value makes it possible to detect the most marked peaks on the impulse response and thus indicate whether the channel on which the phase adjustment is applied is the one which presents the least modification of timbre.

[0054] In an alternative embodiment, a change of channel on which the rephasing is applied, from one frame to another, is further conditioned on a flatness value of the amplitude spectrum of at least one of the two channels.

[0055] This embodiment variant makes it possible to avoid a change of reference channel during periods of high signal harmonicity where the slightest change in timbre is audible and annoying.

[0056] The invention relates to a channel reduction processing device comprising a processing circuit for implementing the steps of the channel reduction processing method as described previously.

[0057] The invention also relates to a computer program comprising instructions for implementing the channel reduction processing method according to the invention, according to any one of the particular embodiments described above, when said program is executed by a processor.

[0058] Such instructions may be stored permanently in a non-transitory memory medium of the channel reduction processing device implementing the channel reduction processing method according to the invention.

[0059] This program can use any programming language, and be under the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0060] The invention also relates to a recording medium or information medium readable by a computer, and comprising instructions of a computer program as mentioned above.

[0061] The recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a mobile medium, a hard disk or an SSD.

[0062] On the other hand, the recording medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means, so that the computer program it contains is remotely executable. The program according to the invention may in particular be downloaded over a network, for example an Internet-type network.

[0063] Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the aforementioned channel reduction processing method.

[0064] According to an exemplary embodiment, the present technique is implemented by means of software and / or hardware components. In this regard, the term "device" or "module" may correspond in this document to a software component, a hardware component or a set of hardware and software components. Brief description of the drawings

[0065] Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which:

[0066] [Fig. 1a] illustrates channels of a stereophonic signal, in phase as previously described;

[0067] [Fig.lb] illustrates channels of a stereophonic signal, in phase opposition as previously described;

[0068] [Fig.2] illustrates in block diagram form a channel reduction processing chain in one embodiment of the invention;

[0069] [Fig.3a], [Fig.3b] and [Fig.3c] illustrate exemplary embodiments of the selection of the channel on which the rephasing is applied for the channel reduction processing for a first embodiment;

[0070] [Fig.4a] and [Fig.4b] illustrate examples of implementing channel selection on wherein rephasing is applied for channel reduction processing for a second embodiment;

[0071] [Fig.5] illustrates an embodiment of an adaptation phase of a channel reduction processing filter;

[0072] [Fig.6] illustrates an exemplary structural embodiment of a channel reduction processing device according to an embodiment of the invention. Description of the embodiments

[0073] The invention will now be described below with reference to Figures 2 to 6. For this purpose, the following symbols will be mainly used in the description which follows. INDICATIONS

[0074] k frequency index

[0075] 1 time index of a signal (in the current frame)

[0076] n time index of a signal

[0077] t frame number CONSTANTS

[0078] The number of samples in the current frame

[0079] R filter delay

[0080] N length of FFT

[0081] C number of input channels

[0082] D number of output channels

[0083] P filter length

[0084] Q crossfade length SIGNALS

[0085] Xi(n) input channels

[0086] m(n) downmix (in time)

[0087] m(t,l) downmix in the current frame (in time)

[0088] M(t,k) downmix in the current frame (in frequency)

[0089] [Fig.2] shows an example of a processing chain for a stereophonic audio signal, the processing comprising a channel reduction processing called “downmix”.

[0090] At the input of this processing chain, a stereophonic signal x, composed of two channels (xi(n) and x2(n)) also called left channel and right channel (respectively), is initially cut in 201, into frames of L samples (xi(t,Z) and x2(t,Z)), t being the index of the frame and 1 the index of the sample. Block 202 applies windowing and a FFT type transform (for “Fast Fourier Transform” in English), to obtain signals in the frequency domain

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097] (X^k) and X2(t,k)). k being the index of the frequency). In 203, a selection of the channel (203b) on which a rephasing is applied for the channel reduction processing is carried out according to a selection criterion calculated in 203a. This selection criterion can be based on a value representative of an energy ratio between the channels of the stereophonic signal and / or a value representative of the shape of the phase shift filter implemented for the channel reduction processing. Several exemplary embodiments will be described with reference to FIGS. 3a to 3c. At the end of step 203, the filter to be applied to the stereophonic signal is determined according to the selection of the channel on which the rephasing is carried out.This filter (defined by filter coefficients in the form of a finite length impulse response) thus determined is conditioned (or adapted) in 204 to make it causal or partly causal and to optimize it in order to reduce the processing complexity. This adaptation phase (of the filter impulse response) is described with reference to [Fig.5]. As described later, these adaptation steps can also be used to determine the shape of the phase shift filter which can, in an exemplary embodiment, be a criterion for selecting the channel on which the rephasing is performed. Once the filter(s) have been determined and adapted for the current frame, they are applied at 205 to this current frame to obtain a mono signal mk (t, l ). Figures 3a to 3c describe examples of embodiments of block 203 of Figure 2. In this processing block, the channel (CH ( t ), with a value of 1 for the left channel and 2 for the right channel) on which the rephasing is applied is selected for a channel reduction processing method comprising filtering of the stereophonic signal and rephasing applied to one of the channels of the stereophonic signal. In a first approach, it is proposed in one embodiment to apply the rephasing to the weakest signal: thus, if the rephasing filter generates a timbre modification, this will be completely or partially masked by the unmodified channel. For this purpose, in this embodiment, the selection criterion is a function of a value representative of an energy ratio between the channels of the stereophonic signal. This value is for example the inter-channel level difference called ILD for “Interchannel Level Difference”, which measures the energy ratio between the left channel and the right channel such that: ZLD(f) ■ ■ iwoir

[0098] Where || x.( fj || 2 is the energy of channel i on frame t.

[0099] Figure 3a illustrates this selection step, in a first implementation. In step E301, the energy ratio between the left channel and the right channel ILD(t) is calculated according to the formula above for the current frame t.

[0100] In step E302, this ratio is compared to a threshold value SI which in this exemplary embodiment is equal to 1. The channel with the lowest energy content is selected for each frame, i.e. depending on whether the value of the ILD is less than a threshold value of 1, to apply a phase adjustment to it in each frequency or frequency band based on the calculation of the IPD described previously, i.e. according to the expression:

[0106] a(k) =

[0101] where X* is the conjugate of A and Z indicates the phase of the complex operand.

[0102] To limit the effect induced by rapid variations in the IPD, a smoothed version of the IPD over time can be used to calculate the filters t, k) ■ One way to smooth the IPD is to apply a low-pass filter of type IIR of order 1, such that:

[0103] IPD(kk) = a(k)IPD(ï -1, k) + (la(k))IPD(kk)

[0104] The forgetting coefficient a(k) is chosen ad-hoc. It may be interesting to choose a high forgetting coefficient in the low frequencies, typically below 5kHz, because this allows phase coherence to be maintained in the low part of the spectrum from one frame to another, the spectrum of speech signals being stationary, particularly in its low part and the particular ear sensitive to the phase in this part of the spectrum. Conversely, it is interesting to choose a low forgetting coefficient in the high frequencies: in the high frequencies, typically above 5kHz, the phase continuity of speech signals is less marked and the ear is less sensitive to phase discontinuities in this part of the spectrum.

[0105] We can choose a forgetting coefficient of the following form: k 4 k > kfoig h amin + 'ki»wk^ khigh

[0107] The values ​​amin, kj^ and klow are determined experimentally. In the case of adjacent frames of length L=20ms, we can for example choose the following set of parameters:

[0108] 0^^ = 0.94, amin = 0.86, k!m. = 0, k^j,^ corresponds to 5kHz,

[0109] In practice, the calculation of the IPD requires an arctangent function which can be expensive. In a variant, we will calculate the complex exponential of the IPD, , via the normalized inter-spectrum:

[0110] [YES]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123] Where | x| is the modulus of the complex x. As with its direct version, we can also smooth the complex exponential of the IPD: eEpD(tk) = a(k)eEPD(rU) + (!-«(£) In the following description, to simplify the notations and equations, the notation gEPiEt-k) will define the exponential of the phase difference, obtained by direct calculation of the IPD or by the normalized inter-spectrum, in its smoothed or unsmoothed version. In this embodiment illustrated in Figure 3a, in the case where ILD(t) < 1, the right channel is selected to be put back in phase with the so-called “reference” left channel, i.e. channel CH(t) — 2 in E303. Otherwise, (N in E302), the left channel is selected to be put back in phase with the so-called “reference” right channel, i.e. CH(t) = 1 in E304. In an alternative embodiment, another threshold value may be chosen. The filtering for channel reduction processing is then determined according to the following equations: ILD(t) < l(E304): H2(t,k)=± ILD(t) > 1 (E303): H2(t,k)=^^- The coefficient a is a normalization coefficient that allows for level consistency between the initial stereophonic content and the monophonic downmix. Different normalization values ​​can be taken depending on the implementation: in practice, the values ​​2 or are often used (one being adapted to an energy average and the other to an amplitude average). The calculation of the energy of each channel on the current frame t can be done by a simple integration on the samples of the current frame: i-Xt) In practice, in encoders for example, frames are about ten milliseconds long. With such short frame durations, the calculated ILD can lead to rapid switching from one channel to another, resulting in audible discontinuities. To avoid this, the energies of each frame can be smoothed with a first-order recursive low-pass filter, as follows: = aE^t-1) + (1-n) Where a is a forgetting factor. We will take values ​​close to 1 to avoid rapid switching from one channel to another.

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134] With frames around 20ms, such smoothing can however lead to poor responsiveness during a major channel change. To limit this lack of responsiveness, we can divide the frame into S sub-frames, and apply the smoothing to the energies of the sub-frames: = aEj(t,sl) + Vse {l, With the notation q) 1, S)2 The energy of the frame t can be taken as equal to E:(t, S) The comparison of the energy ratio of the channels of the stereophonic signal is carried out according to the calculation per frame (Æj(^) / E2(t) or vice versa) or according to the calculation on the last subframe of the current frame (E2(t, s) / E2(t, S1) or vice versa). This choice of predominant channel based on the calculated ILD makes it possible to avoid audible timbre changes in most cases where the signals are energy unbalanced but nevertheless present a certain correlation, as in the case, for example, of binaural signals. However, despite the proposed energy smoothing, the choice of the dominant channel can lead to rapid and repeated switching when the left and right channels are of comparable energies, which can result in audible discontinuities such as clicking. To avoid this, the choice of the channel to be rephased can be limited to periods where the level difference is significant: [lLD < 2LD(t) < alLD: we keep the previous channel Let f ILD(t) <cd-              ="l" ild(t)>alLD- CH(t)=2 \iILD <ILD(t)<aILD : CH(t) = CH(t-1) where aiLD is a threshold value, in practice a few decibels, typically 3dB, or Jï in linear terms. This exemplary embodiment is illustrated in Figure 3b, where in step E301, the energy ratio between the left channel and the right channel ILD(t) is calculated. This ratio is compared to a first threshold S2 Q) in E305. In the case where the ratio is lower than this threshold then the left channel is selected for rephasing, i.e. CH(t) = 1 in E306, the rephasing being carried out by applying to the left channel the filter .jipp(tk) , |a phase Right channel, called “reference” in this case, remaining unchanged (only the module is potentially modified by the normalization coefficient °). Otherwise, the ratio is compared to a second threshold S3 (aiLD) in E307. In the case where the ratio is higher than this threshold then the right channel is selected for rephasing, i.e. CH( t ) = 2 in E308, the rephasing being carried out by applying to the right channel the filter h Çt k)= , the left channel phase, called “reference” in this case, remaining unchanged (only the module is potentially modified by the normalization coefficient °).

[0135] In the opposite case, that is to say when the ratio is between the two thresholds, the channel which was selected for rephasing in the previous frame (CH(l- 1 ) ), is again selected at E309, to which the rephasing filter is applied with the new IPD, i.e. H k) = for the left channel, or tk) = for the right channel.

[0136] The application of this decision freeze zone (combined with the temporal smoothing of parameters such as the IPD, energies, etc.) makes it possible in practice to stabilize the choice of the channel to be rephased and thus avoid audible discontinuities. To further stabilize the choice of channel, the switching from one channel to be rephased to another can be made conditional when the ILD is stable over time. In practice, it can be decided to choose a channel as the reference channel if it remains the most energetic for a sufficient number of consecutive frames, typically of the order of several hundred milliseconds.

[0137] Figure 3c illustrates this scenario. Step E301 remains unchanged compared to Figures 3a and 3b, the comparison with the first threshold S2 is carried out at E305 as in Figure 3b. In the case where the ratio is lower than this threshold then a step of incrementing a frame counter C1 is implemented at E310 and a verification step at E311 is carried out by comparing the frame counter C1 with a threshold T of a desired number of frames (corresponding for example to a duration of 100 ms). If this counter C1 exceeds this threshold T, then the channel selection is authorized and the rephasing is carried out on the left channel (CH(t) - 1), at E312. If the ratio is higher than the threshold S2 at E305 (N to E305), then the frame counter C1 is reset to 0 at E317. This Cl counter allows the selection of a channel to be rephased as soon as this channel has been selected for a time T.

[0138] In the case where the ratio is greater than the threshold S3 in E307, then a step of incrementing a frame counter C2 is implemented in E313 and a verification step in E314 is carried out by comparing this frame counter C2 with a threshold T of a desired number of frames (corresponding for example to a duration of 100ms). If this counter C2 exceeds this threshold T, then the channel switchover is

[0139]

[0140]

[0141]

[0142]

[0143]

[0144] authorized and the rephasing is carried out on the right channel = 2), at E315. If the ratio is lower than the threshold S3 at E307 (N to E307), then the frame counter C2 is reset to 0 at E318. Thus, this counter C2 allows switching to a channel to be rephased to be authorized as soon as this channel has been selected for a time T. In all other cases, (N in step E311, in step E314 and in step E307), the channel which was selected in the previous frame (CH{ f-1)) for rephasing, is again selected for the current frame in E316, and rephasing with the new IPD(t) value calculated for the current frame. This rule helps avoid switching very quickly from one channel to another, for example in the case of sources distributed on either side of the sound stage. In a second embodiment, the channel selection criterion is different. It is based on the shape of the phase shift filter implemented for channel reduction processing. Indeed, in situations where the energy of the left and right channels remain very close to each other, no decision is made except to keep the channel of the previous frame. In the latter case, the decision to freeze the channel switching can sometimes lead to non-optimal situations. Indeed, when the left and right signals are of comparable energies, one can no longer rely on energy masking to mask the spectral modifications that can cause timbre modifications to appear. To avoid this, it is proposed in one embodiment to base the choice of the channel to be rephased on a minimal modification of the timbre, by choosing the channel whose filter best preserves the spectral content of said channel.In practice, it is known that filters with impulse responses resembling a Dirac are the most neutral in terms of timbre modification, because they maintain phase coherence between the different frequencies. This is the case, for example, of a binaural signal where a source is positioned slightly to the left: it appears that the rephasing filter based on the IPD of the right channel H2(f, k) is better conditioned than that of the left channel k ), it has a more linear phase, which guarantees a small modification of the timbre, the frequencies not being out of phase with each other. In this second embodiment, the channel to be rephased is selected as the one whose impulse response of the rephasing filter has the most marked peak. To do this, we propose to use the "spectral flatness" function which measures the degree of flatness of a curve, as follows: C / î \ geometric mean ( h ) _ 5 / tlwpÏT ' ' arithmetic mean^ h ) Where P is the size of the filter h (defined as a causal impulse response of finite length)

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155] When SF(h) takes values ​​close to 0, this means that the h filter has one or more marked peaks, synonymous with a slight modification of the timbre. Conversely, a value close to 1 indicates a spread filter and potentially a source of audible modification of the timbre. So, we propose to select the channel in the following way: SF^IPD^t)) : H2(t,k)=i SF(Tt[lPD(t))>SF(^IPD^ Hl(t,k)=i. where 7i. is a version of the filters calculated by block 204 of Figure 2 and described in more detail with reference to Figure 5 later: hj or hj (filters hj and hj give the same degree of flatness). Figure 4a illustrates this embodiment. In step E401, the shape of the rephasing filter is determined. In the example described here, the degree of flatness of the rephasing filter is calculated for each of the left and right channels, i.e. the criteria SF^) ^SF(Ti2)- In E402, a comparison is made between the two calculated flatness values. In the case where the degree of flatness of the rephasing filter on the left channel is greater than the degree of flatness of the rephasing filter on the right channel, then the channel selected for rephasing is the channel CH( t) = 2 in E403 and the rephasing is performed on the right channel ( k) = ). Otherwise, the channel selected for rephasing is channel CH( t ) = 1 in E404 and rephasing is performed on the left channel ( In the same way as for the selection by the ILD, we can also set up a timer by counter to avoid switching too quickly from one predominant channel to another. In another embodiment, the criterion based on the shape of the phase shift filter can be taken into account in the case where, for the embodiment described in Figure 3b, the value of the energy ratio between the two channels of the stereophonic signal is between two thresholds ■ ~ ILD <ILD(t) < arLD . In this case, the selection criterion based on the shape of the phasing filter is implemented as in [Fig.4a]. [Fig.4b] illustrates this embodiment where steps E301 to E308 correspond to the steps described with reference to [Fig.3b] and steps E401 to E404 correspond to the steps described with reference to [Fig.4a] but applied to signals for which the value of the ILD is between thresholds S2 and S3. Counters can be added (as in the case illustrated in [Fig.3c]) to each of the decisions to avoid too rapid switching from one channel to another.

[0156] When switching from one channel to another (where the choice of the rephased channel is changed), a discontinuity may sometimes be audible. Also, in one embodiment, to avoid switching directly between a mode with left channel phasing to a mode with right channel phasing, a transition step is performed over a duration of one or more frames using another channel reduction processing mode, for example a mode where the rephasing is performed on both channels, for example with a rephasing with an IPD / 2 or IPD / 4 angle.

[0157] This downmix with rephasing on the two channels can be seen, in the frequency domain, as a filtering according to the following equation, with an angle of IPD / 4:

[0158] M(t,k) = Hl(t,k).Xi(t,k)+H2(t,k)X2(t,k)

[0159]

[0160] Or again with an angle of IPD / 2:

[0161] M(t,k) =Hi(t,k)Xl{t,k)+H2(t,k)X2(t,k) [0^2]

[0163] Thus, during a change of channel on which the rephasing is applied, a transition on at least one frame is carried out by applying filtering and rephasing on the two channels of the stereophonic signal.

[0164] In an alternative embodiment, during a change of channel on which the rephasing is applied, a transition on at least one frame is carried out by applying both filtering and rephasing on the channel selected for the previous frame, filtering and rephasing on the two channels of the stereophonic signal and a crossfade between the two filterings.

[0165] This transition mode avoids significant phase jumps, thanks to the interpolation of the phase change achieved by the crossfade.

[0166] If the use of an intermediate mode makes it possible to globally mask the discontinuities, the channel change proposed in the described embodiments can nevertheless become audible. This discontinuity is particularly evident if the change occurs over a period where one of the signals t) is highly harmonic: in this case, the slightest change in coloration becomes particularly audible. This is the case for voiced periods for a speech signal, or generally musical signals. Also, to avoid these audible changes in coloration, in one embodiment, it is proposed to restrict the transitions from one channel to be rephased to the other at times when the signal(s) do not present a strong harmonic. monicity. To detect this harmonicity, a criterion based on the degree of flatness ("spectral flatness") described above is used, applied to the amplitude spectrum of the left or right channel, or a combination of the degrees of flatness calculated on each channel:

[0167]

[0168] This criterion can also be calculated on the power spectrum of at least one of the channels.

[0169] A low value is synonymous with a signal having a low number of frequency lines, and therefore a high harmonicity. Also, in the case of a low value of the SF (X^kk)), t = 1,2, this embodiment proposes not to authorize the transition from one channel to be rephased to the other. The decision can be taken on one of the values ​​of SF(X}(t, k)), i= 1,2, or on a combination of the SFÇX^t, k) ) • An optimal decision consists of basing the decision on the min(SF(Xz( t, k) ) ), which ensures that no channel is strongly harmonic. To limit the complexity of the calculations, we can base ourselves on the degree of flatness SF(Xj( t. k) ) of a chosen channel arbitrarily, which avoids calculating the degree of flatness of the other channel. In practice, since spectral flatness takes its values ​​in the interval [ (>, 1], we can compare the chosen decision criterion, i.e. one of the spectral flatnesses SF^X^t, k) ) , the k) ) ), or any other combination of SF^X^t, k) ), at a threshold asfx e [0,1 ]. A threshold value of the order of 0.05 seems a good compromise.

[0170] Figure 5 now details block 204 of Figure 2. The filters H^t, k), such that defined above according to the channel on which the phase shift is implemented, are reconditioned or even adapted to avoid circular convolution.

[0171] In a first step, at 501, we calculate the impulse responses hj(t, l) of the filters Hj(j, k), such that:

[0172] =FFr1{Hi(t,k)}, 1,2

[0173] where the filters t, l) are of size Lh = B.

[0174] In the preferred embodiment, we have L = 20 ms, or K = 320 samples at 16 kHz, 640 samples at 32 kHz and 960 samples at 48 kHz, BK = 2, or B = 160 samples at 16 kHz, 320 samples at 32 kHz and 480 samples at 48 kHz.

[0175] The impulse responses thus defined are classically non-causal, they correspond to filters with finite impulse response. A classic technique to make them causal consists of performing a rotation of B / 2 samples to center the response in B / 2. The disadvantage is that the filter thus reconstructed has then a latency of B / 2.

[0176] In the preferred embodiment, a truncation of the impulse response will be carried out:

[0177] hi(t, / ) = M< KP h^pB-M + l) 0< KM

[0178] This (circular) time shift operation can also be implemented by applying an appropriate phase term to t, k) before inverse FFT, as known to those skilled in the art. In this embodiment, the obtained filter has a finite impulse response with a delay of R samples. This delay makes it possible to correctly condition the truncated filter, whether in terms of gain by minimizing the influence of the truncation on the unity gain of the phase shifter filter, or in phase by minimizing the deviation from the IPD. This delay R can be adapted according to the implementation constraints: in practice, a delay around one millisecond is sufficient to correctly condition the truncated filter.

[0179] In an extreme variant, the truncation is asymmetrical to avoid this latency. We choose to truncate the filter asymmetrically, at 502, by retaining only a portion of the first half of the impulse response t, l) in the following manner:

[0180] KP

[0181] This filter has the advantage of generally having a maximum (in absolute value) on its first sample, which guarantees processing without additional delay. In terms of frequency response, it has a phase roughly equal to the optimal rephasing filter h^t, l). It may be possible to compensate for the fact that the anticausal part of the impulse response is suppressed by additionally applying a factor of 2 to the values ​​of t, l) for / > 0. In this variant, the resulting filter has a finite impulse response with virtually no delay. ​

[0182] The choice of P depends on different criteria: a value of P close to N / 2 guarantees a near-optimal phase re-phasing, while a smaller value limits the computing power required for filtering. A reduced value of P also smooths the phase of the filter. This has an advantage with filters defined by the channel reduction processing as defined above where the phase varies very quickly from one frequency to another: these variations can result in a very high group delay perceptible to the human ear. A low value of P allows all or part of these artifacts to be eliminated.

[0183] The filter thus truncated t, l), in particular when P <n 2, voire p«:n présente une queue de réponse impulsionnelle qui ne tend pas vers 0. cela crée des discontinuités audibles type cliquetis à la transition entre 2 trames. aussi, pour éviter ces artefacts audible, we weight t, Q with a window w(l), in 503, such that:

[0184] =w(l

[0185] where the coefficients of w(l) tend towards 0 when 1 tends towards 0 or P and are equal to 1 at l=R. This makes it possible to reduce the energy of the samples at the end of the filter and to avoid discontinuities during the transition between filters of consecutive frames.

[0186] In the preferred embodiment, where the truncation preserves an anticausal part, for example, U2 triangular windows are applied to the causal and anticausal parts.

[0187] |,0 <kr R ^1 <p

[0188] or any other Hann type window for example*

[0189] In the variant where the truncation is asymmetrical, this window can be a triangular half-window, or even a Hann half-window:

[0190] = ouc0s(br / 2P)

[0191] Finally, in the case where only P samples of the causal part of the filter are kept, we can also use a combination of a rectangular window of size LR, concatenated with a half-Hann window of size P - LR

[0192] [ 1 ZZ = O n=l-LR 10.5*( 1 + / 3 11 = ^ ....P-1

[0193] The coefficient P makes it possible to compensate for the energy loss of the anti-causal part. A value in the interval fj e [1,2]. Typically, a value / 7=1.8 is a good compromise.

[0194] Due to its truncation and windowing, the filters t, l) have a reduced energy compared to that of the optimal filter h^t, / ): the latter, as a pure phase-shifting filter, has a norm ^2 by construction. Also, we can proceed to a normalization of the energy of the filter thus truncated and windowed, in 504, such that:

[0195] .....

[0196] Where || x(.) || is the 2-norm of a vector.

[0197] In variants, another method may be used in 204 and the method of [Fig.5] may be replaced. For example, the impulse response may be determined directly by least squares minimization and matrix inversion using the Levinson algorithm (complex) as described in section 2 of the article by Mathias C. Lang, Design of nonlinear ph / ase FIR digital filters using quadratic problems, Proc. ICASSP, 1997. For example, an implementation of this method in Mathias C. Lang's thesis, Algorithms for the Constrained Design of Digital Filters with Arbitrary Magnitude and Phase Responses, June 1999 (see routine levin.m in Appendix B and Islevin.m in section 2.1.4). In this case the inverse FFT, truncation and weighting are not necessary; this alternative method directly estimates a finite impulse response of length P (without a minimum phase guarantee). However, a normalization in the form can be applied 1144011

[0198] [Fig.6] illustrates a channel reduction processing device 600, within the meaning of the invention.

[0199] The device 600 comprises a processing circuit typically including: - a memory MEM1 for storing instruction data of a computer program within the meaning of the invention; - an INT 1 interface for receiving a stereo audio signal (x / m); X2(n) - a processor PROC1 for receiving this signal and processing it by executing the computer program instructions stored in the memory MEM1, with a view to carrying out channel reduction processing; in particular, the processor being capable of controlling the processing modules as described with reference to figures 2 to 5; and - a COM 1 communication interface for transmitting the reduced signals, resulting from the channel reduction processing, m( / ), to another processing module, for example a coding or decoding module of an audio signal coder or decoder.

[0200] Of course, this [Fig.6] illustrates an example of a structural embodiment of a channel reduction processing device within the meaning of the invention.

[0201] Figures 2 to 5 commented above describe in detail functional embodiments of this device.

[0202] Applications of this type of channel reduction processing are found for example in audio coding, for example when the remote party does not have the capacity to reproduce stereo sound on its terminal. In this case, it is not necessary to transport a stereo signal and this saves bandwidth. This type of process can be used during point-to-point conversations if one of the participants makes a stereo or binaural sound recording. This type of process can also be present during multi-party conferences: the conference bridge spatializes the scene by creating a stereo scene, but not all participants necessarily have stereo reproduction capabilities. It is then appropriate to perform a downmix of the scene in mono for these participants.

[0203] The invention can also be applied for audio decoding: in this case, it is the user's renderer (restitution module) which will perform the downmix of the stereo content to adapt to the user's reduced capacities (single loudspeaker for example). < / kr> < / n>

Claims

Claims

1. Method for processing channel reduction of a stereophonic signal ((xi(n), x2(n)) to obtain a monophonic signal l) ) comprising filtering of the stereophonic signal (H^k), H2(t, k) in frequency and h^n), h2(n) in time) and a rephasing applied to one of the channels of the stereophonic signal, the method being characterized in that a step of selection (203b) of the channel on which the rephasing is applied, is carried out according to a selection criterion (203a).

2. Method according to claim 1, in which the selection criterion is a function of a value representative of an energy ratio between the channels of the stereophonic signal.

3. The method of claim 2, wherein the channel selected for rephasing is the one for which the energy is lowest.

4. Method according to one of claims 2 to 3, in which the calculation of the energy is carried out per frame or per sub-frame of the stereophonic signal.

5. A method according to claim 4, wherein the energy value of a channel of the stereophonic signal per frame or per subframe is smoothed.

6. Method according to one of claims 2 to 5, in which a change of channel on which the rephasing is applied, from one frame to another, is carried out if the energy ratio between the channels of the stereophonic signal reaches a significant value ( > aiLD or < l / aHD )•

7. The method of claim 6, wherein a change of channel on which the rephasing is applied, from one frame to another, is further conditioned on a stability value of the energy ratio over a number of frames.

8. Method according to one of claims 5 to 6, in which a selection of the channel on which the rephasing is applied is carried out according to a value representative of the shape of the phase shift filter, in the case where the energy ratio between the channels of the stereophonic signal is between two thresholds.

9. The method of claim 1, wherein the selection criterion is based on a value representative of the shape of the phase shift filter.

10. Method according to one of the preceding claims, in which, during a change of channel on which the rephasing is carried out, a transition on at least one frame is carried out by applying a filtering and rephasing on both channels of the stereo signal.

11. Method according to one of claims 1 to 9, in which, during a change of channel on which the rephasing is carried out, a transition on at least one frame is carried out by applying both filtering and rephasing on the channel selected for the previous frame, filtering and rephasing on the two channels of the stereophonic signal and a crossfade between the two filterings.

12. Method according to one of claims 9 to 11, in which the value representative of the shape of the phase shift filter is a flatness value of the impulse response of the preprocessed or truncated filter.

13. Method according to one of the preceding claims, in which a change of channel on which the rephasing is applied, from one frame to another, is further conditioned on a flatness value of the amplitude spectrum of at least one of the two channels.

14. A channel reduction processing device comprising a processing circuit for implementing the steps of the channel reduction processing method according to one of claims 1 to 13.

15. Storage medium, readable by a processor, storing a computer program comprising instructions for executing the method according to one of claims 1 to 13.

Citation Information

Patent Citations

  • Adaptive channel-reduction processing for encoding a multi-channel audio signal

    US20190156841A1

  • Apparatus and method for downmixing or upmixing a multichannel signal using phase compensation

    WO2018086948A1