Optimized processing to reduce the number of channels in a stereo audio signal.

BR112025021284A2Pending Publication Date: 2026-08-25
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025021284
Authority / Receiving Office
BR · BR
Patent Type
Applications
Publication Date
2026-08-25

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

1 / 31 Optimized processing to reduce channels of a stereo audio signal. Field of Technique

[001] The present invention relates to the general field of audio signal processing. The invention relates, in particular, to the reduction mixing of a multichannel audio signal. The reduction mixing of a stereophonic signal to a monophonic signal is of particular interest.

[002] This type of processing is generally applicable in the field of audio technologies and, more specifically, in the field of audio-to-code conversion, whether in the encoding or decoding stage. Previous Technique

[003] Channel reduction or downmixing consists of deducing, based on a combination of the channels C of a multichannel signal x, a signal y that consists of a smaller number D of channels. In practice, this consists of defining a function y(.) such that: ni(n) = with d < cen represent a time index of the input and output signal.

[004] In this invention, the particular case D=1 and C=2 is of interest; reference will be made herein to stereo-to-mono reduction mixing or simply mono reduction mixing. In the case C=2, reference is made to 2-channel content or signal, which may be stereo content or binaural content, wherein channel 1 (%x) is referred to as the left channel and channel 2 (Çv2) as the right channel. Below, the stereo case will be viewed as a general 2-channel signal, which includes the binaural case, in order to avoid repeating the two terms, even though technically a binaural signal has a different character. Petition 870250089800, dated 02 / 10 / 2025, page 9 / 61 2 / 31 specific characteristics. In the case of the signal y obtained after the reduction processing, reference is made to a monophonic or mono signal, which will be noted m(n) in the time domain, or M(k) in the frequency domain, below.

[005] These stereo signals can result from capture by a pair of stereo microphones or binaural capture, or even from artistic mixing of audio tracks.

[006] There are a number of microphone pairs that make it possible to create stereo content. Among the most popular, one can mention the XY pair, composed of 2 cardioid microphones that exhibit a difference in orientation angle between 90° and 135°. The MS (mid-side) pair is another widely used pair composed of a cardioid microphone and a second microphone, called a Figure 8 microphone, oriented at 90°. By combining these 2 microphones (sum / difference), the left and right stereo content channels can be created. These pairs are called coincident pairs, meaning that the microphone capsules do not exhibit a delay between them, with spatialization perceived by the difference in sound intensity between the left and right channels.

[007] Another category of pairs, called phase-shifted stereo pairs, involves the use of microphones that are far apart, thus creating a phase difference between the channels for sources located closer to one of the microphones. The best known is the AB pair, which uses two omnidirectional microphones separated by a few centimeters to several meters. Another widely used pair is the ORTF pair, which uses a phase difference through a 17 cm spacing of the microphones and an amplitude difference through cardioid directivities (the microphones exhibit an angle of about 90°). The binaural pair consists of placing omnidirectional microphones in the ears of a person or an artificial head. Petition 870250089800, dated 02 / 10 / 2025, p. 10 / 61 3 / 31 ciai: This type of device makes it possible to create binaural content natively, which can be heard through headphones.

[008] Stereo content also includes speech and audio signals resulting from audio mixing or audio post-production (e.g., channels stored on a CD or DVD or transmitted over the Internet).

[009] There are also other methods for creating stereo content, which are not reviewed here.

[0010] The simplest method for creating a reduced signal by reduction mixing is known as passive reduction mixing. It consists of obtaining an average of the left and right channels of the stereo signal such that: .. XiíÜ+xaít) m(t) =--------or, in the frequency domain: where k),i = 1,2 is the fast Fourier transform (FFT), with k being the frequency index and t being the frame index: Ll Σ / 2π-ΒίΓwCO.XjCt,í)e Artpara qualquer k E (0,. ..,JV - 1} ϊ=0 onde j é o número imaginarios como que j = y—1 , t é o rfc, l é o rfc(.) é o rfc(.) a senó or Hann ou o rfc(.) janela de apodização de tamanho L, conformed a tamanho de quadro, e xf(tj) = This definition extends to the case where N>L and íf(.) é um versão augmentada de xf(t) com Os (0-filling).

[0011] This passive reduction mixing method, which is very simple, works well in situations where the right and left channels are in phase. However, in situations where the microphones are not in phase and exhibit phase differences, such mixing Petition 870250089800, dated 02 / 10 / 2025, p. 11 / 61 4 / 31 generates combined filtering, which is manifested by the coloration of the original signal. This is due to the fact that, depending on the microphone spacing, the position of the source relative to the microphone pair, and the frequency, the left and right channels of the same source may be in phase, as shown in Figure 1a, or out of phase as shown in Figure 1b. Thus, in the downmix, the left and right signals will add (constructive interference) or cancel each other out (destructive interference), thus creating variations in the level and coloration effects of the downmixed signal m(t). The coloration is due to the effect of combined filtering: The effects of constructive and destructive interference manifest themselves in different frequency bands, thus modifying the balance and therefore the tone of the resulting downmixed signal relative to the original.

[0012] In order to correct this intensity level defect according to frequencies, the e-AAC+ codec's reduction mixing implements y(k) level correction which ensures that the energy level of the reduced signal, in each frequency band, remains comparable to the original: Aí(Jc) = comr' JO'SIXiU) + x2{*)|2

[0013] This approach makes it possible to compensate, to some extent, for the drop in intensity of the reduction mix. However, this compensation leads to overamplification when the signals are close to phase opposition: Also, in practice, the correction factor yú') is limited to a maximum value (for example, an upper limit of 2).

[0014] Figure 7 describes a method for reducing stereo channels to mono that is integrated into the public candidate IVAS codec which is available at: Petition 870250089800, dated 02 / 10 / 2025, page 12 / 61 5 / 31 https: / / forge.3gpp.org / rep / ivas-codec-pc / ivas-codec / - / tree / main

[0015] The signal x(m) with 2 input channels, where m is the index of interleaved samples, is deinterleaved (block 701) to find the left and right channels. Then, a frequency analysis with windowing and Fourier transform is performed (blocks 702 and 703) in order to obtain the spectra of the 2 channels, and the intercorrelation based on the phase spectra is estimated (block 704) before determining the time difference between the channels (here marked as ITD although it is formally an ICTD between 2 channels) looking for a peak (block 705); block 705 also provides the correlation level R* corresponding to the ITD.

[0016] The mixing factor g is then determined as follows (block 706): <1 R* - — — τ > 0 2 2 R*S=2+2 r<0 1 t = 0 2 where g represents the mixing factor and τ is the time difference between channels (incorrectly called ITD). Thus, the 2 channels are combined into a mono signal in block 706 according to this factor g, with smoothing of the factor value applied sample by sample at the beginning of each frame after block 707: m(n) = v(n) xi(n) + (1-γ(η)) x2(n), where γ(η) is the factor derived from g after temporal smoothing with the gain from the previous frame.

[0017] Next, the energies of the left channel xi(n) and right channel x2(n) and the mono signal m(n) are determined (blocks 708, 709, 710). The left and right channels are reinjected separately (added) in blocks 713 and 715 into the resulting mono signal from block 706, depending on the energy compensation determined in Petition 870250089800, dated 02 / 10 / 2025, page 13 / 61 6 / 31 block 711 with respective scale factors defined in blocks 712 and 713.

[0018] These methods, in addition to the signal coloration resulting from reduction mixing, linked to imperfectly corrected combined filtering, suffer from excessive reverberation linked to the summing of signals when the input channels are out of phase and from a click artifact when the channels are pointwise reinjected into a single frame. Compensating for an identical delay for all frequencies is unrealistic; this delay is linked to the various sources that make up the content. In practice, this leads, in addition to further coloration, to a potential loss of intelligibility of the sources, whose defect level compensation cannot correct.

[0019] Other downsampling approaches seek to avoid the aforementioned defects. The principle of these methods is to rephase the left and right signals before they are summed. In the document entitled A stereo to mono downsampling scheme for MPEG-4 parametric stereo encoder by Samsudin, E. Kurniawati, N. Boon Poh, F. Sattar, S. George, in Proc. ICASSP, 2006, a method is proposed that consists of applying, in the frequency domain, a phase shift φ to one of the channels, usually the right one, in order to rephase it with the left channel, so that: fT, XiOc) Mík) = -------1-------

[0020] The ideal phase shift is determined by what is called here IPD (phase difference between channels): IPDÍk) = 4^0 / ). where X* is the conjugate of x and indicates the phase of the complex operand.

[0021] It is calculated in the frequency domain: this makes it possible to reframe the spectral components of multiple sources, a source with its own IPD being able to be predominant in a frequency band. Petition 870250089800, dated 02 / 10 / 2025, page 14 / 61 7 / 31 frequency k, while another source with a different IPD may be predominant at another frequency k.

[0022] In the case of the MPEG4 codec described in the Samsudin document cited above, the IPD applied is that estimated by averaging over a Bark band and not the IPD calculated for each frequency band, the latter being particularly noisy and variable from one frame to the next. Furthermore, this method takes the left channel as a phase reference and, if the phase of this channel is poorly conditioned, the downscaling mix will have degraded quality.

[0023] In published patent application no. W02017103418, a reduction mixing method is proposed that combines the method proposed by Samsudin and passive reduction mixing. Notably, it proposes an ISD (interspectral distance) indicator that allows the selection, frequency band by frequency band, of the most suitable method for rephasing the signals: ()ΐχ,ω+χ,ωι

[0024] This indicator makes it possible to measure whether the signals are in phase (ISD < 1), that is, with a phase difference or IPD in the interval , or rather, in opposite phase (ISD > 1), that is, with an IPD in the interval . Lz 2 J

[0025] When they are in opposite phase, a high ISD value is synonymous with a comparable level between the left and right channels, and rephasing is important to avoid coloration and dropout effects. On the other hand, an ISD value close to 1 is representative of one channel being predominant over the other; in the latter case, rephasing is unimportant, since the reduction mix is ​​almost equal to the predominant signal. It is, however, important to avoid modifying the phase of the right channel if the latter is predominant. Thus, in this patent application, it is proposed Petition 870250089800, dated 02 / 10 / 2025, page 15 / 61 8 / 31 to the following reduction mix selection mechanism: Xi(fc) + MW =-----------Juí 4 ír p = ------------------------If lSD{k) > 1.3 If ISD{k) < 1,3

[0026] Passive reduction mixing, proposed in patent application no. W02017103418, for low ISD situations, sometimes introduces increased reverberation.

[0027] Furthermore, the method proposed in this patent application requires the application of a frequency-domain processing operation, whereby the reduced signal is reconstructed based on the short-time inverse Fourier transform (STFT) of the signal M(t,k). This type of filtering generates a circular convolution that creates audible artifacts: Echo, pre-echo, and static. In order to mask these artifacts, an implementation with frame overlap associated with appropriate windowing (analysis and / or synthesis window) can be used, which guarantees the reconstruction of the filtered signal. Generally, this overlap is not compatible with the operation of audio encoders that work with adjacent frames, i.e., without overlap; furthermore, this type of reconstruction leads to a delay, usually half a frame, which is the conventional overlap, which is not acceptable for a code encoder that has strong latency constraints.

[0028] Another implementation of overlay-addition (OLA) is also possible: this method requires a zero-filling step of the signals and filters to avoid circular convolution. However, tests show that with these filters, whose phase varies very rapidly from one frame to another, an OLA approach does not make it possible to completely mask artifacts in the transition of certain frames; the artifacts remain audible.

[0029] There are also, in the methods of the previous technique, in which the mi Petition 870250089800, dated 02 / 10 / 2025, page 16 / 61 9 / 31 Reduction mixing is achieved by switching, for example, in the frequency domain in the published patent application WO2017103418, different reduction mixing methods. In this case, it is important to ensure that the switching occurs continuously, i.e., without discontinuities or differences in levels between methods to avoid artifacts. Description of the Invention

[0030] The invention aims to improve the prior art

[0031] For this purpose, the invention targets a method for processing a channel reduction of a stereophonic signal in order to obtain a monophonic signal, which comprises the following: - to apply, to a current frame of said signal, two channel reduction processing operations, one of the processing operations using rephasing filtering in at least one of the channels of the stereophonic signal and the other processing operation not using rephasing filtering, - Select one of the aforementioned processing operations to be applied to the current frame, with said selection being implemented depending on: - the presence or absence of at least one transient in the stereophonic signal, or - a level of quality of stereophonic signal filtering used when implementing the processing operation with the use of stereophonic signal filtering.

[0032] The invention advantageously provides two decision criteria that make it possible to choose between two different processing operations, one of which uses filtering of the stereophonic signal and the other does not. Petition 870250089800, dated 02 / 10 / 2025, page 17 / 61 10 / 31

[0033] In one embodiment, if the presence of at least one transient is detected, or not detected, in the stereo signal, the channel reduction processing operation, which does not use stereo signal filtering or which uses stereo signal filtering, is selected, respectively.

[0034] This option allows: - To obtain better rendering of one or more signal transitions, when one or more transitions are detected, by selecting the channel reduction processing operation that does not use filtering and, as a result, preserves the nature of the stereophonic signal. - When no transition is detected, favor the channel reduction processing operation that uses filtering, which is more suitable for stereo signal components that are not in phase.

[0035] In one embodiment, said transient is detected before either of the two channel reduction processing operations is applied to the stereo signal.

[0036] This mode makes it possible, even before a channel reduction processing operation is applied, to deduce a priori whether at least one transition is present or not present in the stereophonic signal, without a particular additional processing operation being applied to that signal.

[0037] In one embodiment, the detection of said at least one transient comprises the following steps: - to decompose the current stereophonic signal frame into subblocks; - Calculate the energy of each of the sub-blocks obtained; - compare the energies obtained with an energy threshold (th61); and Petition 870250089800, dated 02 / 10 / 2025, page 18 / 61 11 / 31 If the energy is less than or greater than the stated energy threshold, respectively, the presence of a transient will either not be detected or will be detected, respectively.

[0038] In one embodiment, said filter quality level is compared with a filter quality threshold and, if said filter quality level is below, or above, said filter quality threshold, the channel reduction processing operation that does not use stereo signal filtering or that uses stereo signal filtering will be selected.

[0039] According to one embodiment, the said level of filtration quality is measured as follows: - Calculate the energy of the stereo signal, referred to as the input energy, before filtering is applied. - Calculate the energy of the monophonic signal, referred to as the output energy, after filtering has been applied. - Calculate the ratio between the input energy and the output energy.

[0040] In one mode, a representative indicator of the presence or absence of a transient is generated.

[0041] In one embodiment, two levels of said processed current frame are respectively obtained after two processing operations, one using filtering of the stereophonic signal and the other not, are applied to said current frame of said signal, - adjust the said level of the current frame to which the channel reduction processing operation, which uses said filtering, was applied, wherein said adjusted level is similar to or equal to said level of the current frame to which the channel reduction processing operation, which does not use said filtering, was applied; - compare the current processed frame that has the said adjusted level and the current processed frame that has the said level obtained after Petition 870250089800, dated 02 / 10 / 2025, page 19 / 61 12 / 31 the channel reduction processing operation, which does not use stereo signal filtering, to be applied to said current frame of said signal, - select one of the aforementioned processing operations to be applied to the current table, based on the aforementioned comparison.

[0042] The invention advantageously enables, when the channel reduction processing operation in a current stereophonic signal frame is altered, maintaining the stereophonic signal level as stable as possible, without leading to discontinuities in the perceived signal level.

[0043] The invention targets a channel reduction processing device comprising a processing circuit to implement the steps of the channel reduction processing method as described above.

[0044] The invention also relates to a computer program comprising instructions for implementing the channel reduction processing method that is in accordance with the invention, according to any of the particular embodiments described above, when said program is executed by a processor.

[0045] Such instructions can be durably stored in a non-transient memory medium of the channel reduction processing device that implements the channel reduction processing method according to the invention.

[0046] This program can use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form or in any desired form.

[0047] The invention also targets a computer-readable storage medium or information medium comprising instructions of a computer program as described above. Petition 870250089800, dated 02 / 10 / 2025, page 20 / 61 13 / 31

[0048] The storage medium can be any entity or device capable of storing the program. For example, the medium can comprise storage media, such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or indeed a magnetic storage medium, for example, a mobile medium, a hard disk or an SSD.

[0049] Furthermore, the storage medium may be a transmissible medium, such as an electrical or optical signal, which may be routed through an electrical or optical cable, by radio or by other means, so that the computer program it contains may be executed remotely. The program according to the invention may, in particular, be transferred by download from a network, for example, an Internet network.

[0050] Alternatively, the storage medium may be an integrated circuit in which the program is embedded, the circuit being adapted to execute or to be used in the execution of the aforementioned channel reduction processing method.

[0051] According to an example of an embodiment, the present technique is implemented by means of software components and / or hardware components. With this in mind, the term device or module may correspond, in this document, equally to a software component, a hardware component, or a set of hardware and software components. Brief description of the drawings

[0052] Other features and advantages of the invention will become more clearly apparent upon reading the following description of particular embodiments, which are given by way of illustrative and non-limiting examples only, and the accompanying drawings, among which: [Figure 1A] illustrates channels of a stereophonic signal, in phase as described above; Petition 870250089800, dated 02 / 10 / 2025, page 21 / 61 14 / 31 [Figure 1B] illustrates channels of a stereophonic signal, in opposite phase as described above; [Figure 2] illustrates, in the form of a block diagram, a channel reduction processing sequence in one embodiment of the invention; [Figure 3] illustrates one embodiment of a filter calculation for reduction mixing, as well as a selection of a channel reduction processing operation; [Figure 4] illustrates one embodiment of an adaptation phase of a channel reduction processing filter; [Figure 5] illustrates one mode of applying reduction mixing during a transition between two stereo signal frames; [Figure 6a] illustrates another embodiment of a filter calculation for reduction mixing, as well as a selection of a channel reduction processing operation; [Figure 6b] illustrates another embodiment of a filter calculation for reduction mixing, as well as a selection of a channel reduction processing operation; [Figure 7] illustrates an existing channel reduction processing method described above; [Figure 8] illustrates an example of a structural embodiment of a channel reduction processing device according to an embodiment of the invention. Description of the modalities

[0053] The invention will now be described below with reference to Figures 2 to 8. For this purpose, the following symbols will be used primarily in the following description. k-indices frequency index Petition 870250089800, dated 02 / 10 / 2025, page 22 / 61 15 / 31 l time index of a signal (in the current frame) n time index of a signal t frame number CONSTANTS L is the number of samples in the current table. N is the number of input channels. M is the number of output channels. P filter length Q crossfading length SIGNS Xi(n) input channels m(n) downmixing (in the time domain) m(t,l) downmixing in the current frame (in the time domain) M(t,k) reduction mixing in the current frame (in the frequency domain)

[0054] Figure 2 shows an example of a sequence for processing a stereophonic audio signal, where the processing operation comprises a channel reduction (downmix) processing operation.

[0055] At the input of this processing sequence, a stereophonic signal x, composed of two channels (x1(n) and x2(n)), also called the left channel and right channel (respectively), is, in a first stage, divided, in 201, into sample frames L (xi(t,l) and x2(t,l), t being the frame index and the sample index. Block 202 applies windowing and an FFT (abbreviation for Fast Fourier Transform), in order to obtain signals in the frequency domain (Xi(t,k) and X2(t,k), k being the frequency index). In 203, a channel reduction filter to be applied to a signal frame is then selected and determined. This step will be described with reference to Fig. Petition 870250089800, dated 02 / 10 / 2025, page 23 / 61 16 / 31 ra 3. The filter thus determined is conditioned (or adapted) in 204 in order to make it causal and optimize it in order to reduce processing complexity. In one example of a modality, windowing and FFT are similar to steps 702 and 703 of Figure 7.

[0056] This adaptation phase is described with reference to Figure 4.

[0057] Once the filter has been determined and adapted to the current frame, it is applied, in 205, to this current frame. In order to avoid audible artifacts that are due to a change of filters between two frames, a crossfading is performed taking into consideration the filters of the previous frame.

[0058] This step will be described with reference to Figure 5.

[0059] Figure 3 describes a detailed embodiment of block 203 of Figure 2. In this processing block, a first or second channel reduction processing operation to be applied to a current frame of the stereophonic signal is selected and determined.

[0060] A first method of processing T1 channel reduction is defined here by shifting the left and right channels of the stereo signal simultaneously by an angle defined by half the phase difference (IPD) determined between the two channels of the stereo signal. Thus, there is a reduction mixing signal, as follows: , .IPDfck) .IPDfck) k) = —--------------------

[0061] This method makes it possible to avoid the excess reverberation degradation generated by conventional reduction mixing (averaging the two signals) or by combined filtering in the case of a time shift between the two channels.

[0062] This reduction mix defined in this way can be Petition 870250089800, dated 02 / 10 / 2025, page 24 / 61 17 / 31 viewed in the frequency domain as filtering according to the following equation: Λί(ίΛ) = Ηχίί, + Η2(ίΛ).ΑΤ2(ί^) with .IPD _ .IPJÇtÀ) and =e-^~ and w2(t,fc) =

[0063] A second T2 reduction mixing method is, for example, defined as that proposed in the Samsudin document cited above.

[0064] In this second method, then, = 1 and .-jlPStftô H2CtA) = —-—

[0065] The stereo signal channels, in the frequency domain (Xi(t,k) and X2(t,k)), are used, on the one hand, to calculate an ISD phase indicator in 301, which is representative of a measure of the degree of phase opposition between the stereo signal channels and, on the other hand, the phase difference between the channels (IPD) in 303.

[0066] The phase indicator is, for example, defined by the ISD (interspectral distance) indicator, as defined above.

[0067] In the mode described here, the ISD is determined by the stereo signal frame so that the decision to select a first channel reduction processing (first mix) or a second one can be made frame by frame.

[0068] This frame-by-frame decision makes it possible to eliminate substantial phase shifts between frequencies and other audible artifacts that can appear when the decision varies from one frequency to another, as in the patent application cited above.

[0069] To apply a frame-by-frame decision, a decision criterion according to the ISD phase indicator is defined according to the following equation in 302: Petition 870250089800, dated 02 / 10 / 2025, page 25 / 61 18 / 31 _ , mure-,,πο total number of frequency lines where l£Dlt,k) > 1,3 iSDÍt) = , total number of frequency lines

[0070] This criterion measures the percentage of frequency lines with index k for which the criterion favors downmixing with rephasing, i.e., when !SD(t1k)> 1.3. The decision is to apply the first downmixing as defined above, with rephasing, when fâbit) exceeds a certain threshold ÍSD^, which is defined, for example, at 70%.

[0071] There is then, in 305, the following decision to select a reduction mix between T1 and T2: , .IPD(.tfrJ .... _. . . and 'J2 . . eJ2 íSDt.t) < !SDth: k) =---------, K21.t, k) =--------jSD(t) > > H^k} = -t / ί2·Χ / ) =------

[0072] In 303, the phase difference IPD is obtained between the two channels of the stereo signal by !PD(t,k) = Ãx^t,k^.X^k)) where z indicates the phase of the complex operand.

[0073] However, the decision of the framework may sometimes be insufficient, as the IPD, calculated according to the formula above, varies very rapidly over time. In addition, the filters H^k) can vary very rapidly from one frame to the next, generating discontinuities that are sometimes audible during the transition from one frame to the next. To limit this effect, a time-smoothed version of the IPD can be used to calculate the H^k filters. One way to smooth the IPD is to apply first-order IIR low-pass filtering, at 304, such that: JPDÍt, k) = aíkyiPDÍt, k) + (1- aÇk^IPDÍt, k)

[0074] The a(k) forgetting coefficient is chosen ad hoc. It may be advantageous to choose a high forgetting coefficient at lower frequencies, typically below 5 kHz, as this makes it possible to maintain phase coherence in the lower part of the spectrum. Petition 870250089800, dated 02 / 10 / 2025, p. 26 / 61 19 / 31 One frame for the next, given that the spectrum of speech signals is stationary, notably in its lower part, and the specific ear is sensitive to phase in this part of the spectrum. Conversely, it is advantageous to choose a low forgetting coefficient at higher frequencies: at higher frequencies, typically above 5 kHz, the phase continuity of speech signals is less marked, and the ear is less sensitive to phase discontinuities in this part of the spectrum.

[0075] A forgetting coefficient of the following form can be chosen: j-ydl·· k ** kaít0Λ-ϋαίχο , , , ^ηιίη τ' τ __tf ι Κ. C. lía [high - low

[0076] The values ​​kaJtae kbaÍxasão are determined experimentally. In the case of adjacent frames of length L = 20 ms, the following set of parameters can, for example, be chosen: =0.94, íTmin=6.86, ^π)-τα= 0, , corresponds to 5 kHz.

[0077] In some variants, very low frequencies and very high frequencies will be forced to remain intact. Thus, in some variants, H^k) = H2{.t,k) θ forced for the first rows k = o or k = o, 1 as well as for the rows of index k > kmÁX, where k corresponds to 8 kHz.

[0078] Figure 4 currently details block 204 of Figure 2. The H^tk filters, such as those defined in the T1 and T2 reduction mixing methods defined above, are reconditioned or, in fact, adapted in order to avoid circular convolution.

[0079] In a first step, in 401, the impulse responses h^t, í) of the filters Hf(t, k) are calculated such that: Petition 870250089800, dated 02 / 10 / 2025, p. 27 / 61 20 / 31 (t, I) = = 1.2 where the h^l) filters are of size Λ / , which corresponds to the length of the FFT. Here, N=L.

[0080] In the preferred mode, L = 20 ms, or 320 samples at 16 kHz, 640 samples at 32 kHz and 960 samples at 48 kHz.

[0081] Impulse responses defined in this way are conventionally non-causal; they correspond to finite impulse response filters. A conventional technique to make them causal consists of rotating L / 2 samples to center the response at L / 2. The disadvantage is that the filter reconstructed in this way then exhibits a latency of L / 2. To avoid this latency, the choice is made to truncate the filter asymmetrically at 402, retaining only a portion of the first half of the impulse response ^-(t, l) as follows: < l < P ^ / 2

[0082] This filter exhibits the advantage of generally having a maximum (in absolute value) in its first sample, 0, which guarantees 0 processing without additional delay. In terms of frequency response, it exhibits a phase that is almost equal to the ideal rephasing filter.

[0083] The choice of P depends on several criteria: a value of P that is close to N / 2 guarantees almost ideal rephasing, while a smaller value makes it possible to limit the computational power required for filtering. A low value of P also results in smoothing the filter phase. This exhibits an advantage with filters defined by the T2 processing operation, as defined above, where the phase varies very rapidly from one frequency to the next: these variations can result in a very high group delay that is perceptible to the human ear. A low value of P makes it possible to remove all or some of these artifacts. Petition 870250089800, dated 02 / 10 / 2025, page 28 / 61 21 / 31

[0084] The filter, therefore, is truncated (h^l), notably when P <N / 2, ou mesmo P«N / 2, exibe uma cauda de resposta de impulso que não tende em direção a 0. Isso cria descontinuidades de clique audíveis na transição entre 2 quadros. Assim, a fim de evitar esses artefatos audíveis, h&í) é ponderado com uma janela w(Z), em 403, de modo que: where the coefficients of w(Z) tend to 0 when I tends to P. This makes it possible to reduce the energy of the samples at the end of the filter and avoid discontinuities during the transition between filters of consecutive frames.

[0085] This window could be a triangular half-window, or indeed a Hann half-window: iz / p wQ) = cos{> / 2P)

[0086] Because of their truncation and windowing, the fyCtJ filters have reduced energy compared to the ideal filter; the latter, as a pure phase-shift filter, has a 1 / 2 norm by construction. Thus, the energy of the truncated and windowed filter can be normalized, in 404, so that: where ||%(.)|| is the 2-norm of a vector

[0087] According to the invention, in 405, a level of filtration quality that has been implemented is determined. For this purpose: - an energy E401 of the filter before truncation 402 is calculated; - An E403 energy of the temporal filter after 403 windowing is calculated; Petition 870250089800, dated 02 / 10 / 2025, page 29 / 61 22 / 31 - A ratio R4 between E401 and E403 is calculated.

[0088] A filtering quality can also be determined by comparing an E401 energy and an E404 energy of the temporal filter after normalization.

[0089] Through non-exhaustive alternatives, the level of filtration quality can be determined: - directly in the Hi(t,k) filter and / or - a frequency version of this same filter before normalization (H^k) or after normalization (H^k). - etc.

[0090] In some variants, it is possible to calculate: Pl (k^l) 1=0

[0091] This criterion can be normalized by: O2 I

[0092] Another quality variant may be concerned with measuring the amount of energy lost during truncation in order to make the filter causal. In this variant, E401 can be calculated as the anticausal energy of the original filter: AT-1 E401(O = h-fitA)

[0093] The E403 (or E404) energy can be calculated as the energy of the synthesized filter, before normalization (or after normalization): pi £403(0 = I=1 Pl £404(0 = I=1

[0094] In order to avoid the polarization associated with the size of the ja Petition 870250089800, dated 02 / 10 / 2025, p. 30 / 61 23 / 31 in it, the energy of the anticausal part can also be calculated over the same number of samples, with or without windowing by the w(Z) window: Jf-l£401(O= 2 i=N / aP Jf-1 EmCO = X wQv - í + Plkfí&l) I=N —P

[0095] In some variants, the frequency response Hfy.k) of the filter obtained after truncation (at the output of 403) can be calculated and compared with the target response H^tk) in order to define the quality level, for example: k

[0096] This criterion can be normalized by: k

[0097] Another variant of the quality criterion may consist of comparing the phase coherence of the filters, since the objective is to rephase the signals. Thus, the coherence between the 2 filters can be calculated:

[0098] A high coherence (i.e., close to 1) indicates that the synthesized filter is close to the ideal filter and therefore a high quality factor, a low value (close to 0) indicates, conversely, a low quality criterion.

[0099] Figure 5 currently depicts block 205 of Figure 2, where an OLS (overlay-save) implementation of the filtering is Petition 870250089800, dated 02 / 10 / 2025, page 31 / 61 24 / 31 performed in the time or frequency domain, in detail. The choice of domain will be made depending on the complexity, which depends on the size of the filter relative to the frame to be processed.

[00100] An implementation of OLS in the time domain is presented here, with adjacent frames of size L.

[00101] In a first stage, the current frame is concatenated, in 501, with the P-1 samples from the previous frame (saving principle): ) = IX-(tL — P + 2), ...xf((t + 1)L — 1)]

[00102] Each channel is filtered by its phase-shift filter in 503: Pl yf(t,2) = l + P-pl.O £1 <L p=0

[00103] The mono reduction mix is ​​then created by summing the two re-phased channels in 505: ?n(t, 0 = ^(tj) + £ í < £

[00104] In practice, the filters can be very different between two successive frames, for example, when a new source appears. In this case, the transition between l) and lrl can cause differences in the form of audible artifacts. To avoid these artifacts, a crossfading step can be applied. In order to perform this crossfading, the filters h^t — lrl) are applied to the current frame over a number of samples Q: Pl hi(t - 1 + Pp),G < Z < Q p=0

[00105] This is accomplished by block 502 in Figure 5.

[00106] The number of samples required for cross fading will have to be determined experimentally, depending on the size of the filters, the sampling frequency, etc.

[00107] Next, the mono reduction mix ?nr_1(tj) is created Petition 870250089800, dated 02 / 10 / 2025, p. 32 / 61 25 / 31 of 504, based on the filters in the previous table: = yit-i&Í) + ?2Λ-ιΟ,Ο,0 <L<Q

[00108] The resulting signal from the reduction mixing fn(tj) is created in 506, based on crossfading according to the following equation: q=+ fl -β0 < Z < Q ( inCt, l)rQ < I < L where β{1) is a decreasing monotonic function, with / ?(o) = 1 and - i) =£ with ε being a value close to 0 or equal to zero. Normally, a half Hann window can be chosen: β{ΐ) = CGS í < Q

[00109] It should be noted that, in an integration with many other reduction mixing methods, the aforementioned reduction mixing methods T1 and T2 cannot always be represented by a filtering operation. This is the case for method T3 in Figure 7 or any other method whose gain depends on the instantaneous signal level.

[00110] According to the invention, in the case where a selection of the reduction mixing methods T1 and T2 or the reduction mixing method T3 is implemented, as will be described in the remainder of the description, a 5040 and 5050 scaling step of each of the signals and Tn(t,i), respectively, is implemented. For this purpose, the signal (n) obtained at the output of the reduction mixing T3 of Figure 7 which has a level N3, the level N1, or N2, of each of the signals ™Γ-ι(<τ0 θ mfaí) to which the reduction mixing T1, or T2, was applied is adjusted to the level N3, so as to be equal to or similar to this level N3. Instead of being applied to each of the signals and ττι(ίτΖ), such scaling can also be performed on the signal ϊη(ί,Ζ).

[00111] Figure 6a currently details a channel reduction processing method that is in accordance with the invention, in Petition 870250089800, dated 02 / 10 / 2025, page 33 / 61 26 / 31 that one of the reduction mixing methods T1 or T2, chosen at the end of the selection method in Figure 3, is implemented in 610, and another reduction mixing method is implemented in 611, whose gain depends on the instantaneous signal level, such as, for example, the T3 method in Figure 7.

[00112] In this case, one solution is to calculate the reduction mixes of each method for the current time index frame / and perform a transition, in 612, from one reduction mix to the next by crossfading, which generates a monophonic signal d(n). More specifically, in the case of a transition from a reduction mix T1 (or T2) to a reduction mix T3, the reduction mix can be performed between the signal fn(tj) obtained at the output of T1 (or T2) and the signal obtained at the output of T3. The reduction mix can also be performed between the output of T1 (or T2) and the signal d(n). In this last case, the energy compensation of the T3 method (adapt_gain in Figure 7) would be applied to all methods continuously.The decision to change from a T1 or T2 reduction mix to a T3 reduction mix may be motivated by criteria other than ISD, for example Par3 parameters already calculated by the T3 method, such as, for example, ITD, choosing to apply the T3 reduction mix by default when the ITD is very large (e.g., ITD > 600 microseconds in absolute value), or any other criterion or combination of criteria.

[00113] According to the invention, the reduction mix T1 (or T2) or T3 is selected depending on: - the presence or absence of at least one transition in the stereophonic signal, or - of a quality level of stereo signal filtering used when T1 (or T2) reduction mixing is implemented.

[00114] In the case of a 613 detection of the presence or absence of Petition 870250089800, dated 02 / 10 / 2025, page 34 / 61 27 / 31 at least one transition in the stereophonic signal, according to a first embodiment, transient detection is applied separately to each of the input signals xi(n) and X2(n), or to the current frame xi(f, / ) and X2(f, / ). Since transient detection is a conventional problem in audio-to-code conversion, the module already implemented in the EVS codec and described in the 3GPP TS 26.445 standard clause 5.1.8 is taken as an example. For this purpose, each of the input signals xi(n) and X2(n), or the current frame xi(f, / ) and X2(f, / ) is divided into sub-blocks. The energy of each sub-block is calculated, then smoothing is optionally implemented. The energy obtained is then compared with an energy threshold thsi. If this energy obtained is below thsi, the presence of a transition is considered not to be detected. If this obtained energy is above thsi, the presence of a transition is considered to be detected.For this purpose, a parameter Par 1 (if T1 was chosen at the end of the selection in Figure 3) or Par2 (if T2 was chosen at the end of the selection in Figure 3) is defined: - to a first value, for example 0, in order to indicate that the presence of a transition was not detected, - to a second value, for example 1, in order to indicate that the presence of a transition has been detected.

[00115] This parameter Par1 / Par2 is transmitted, as a decision criterion, to the cross-fading block 612.

[00116] In the case where the T1 (or T2) reduction mix or the T3 reduction mix is ​​selected based on a quality level of the stereo signal filtering used when the T1 (or T2) reduction mix is ​​implemented, with reference to Figure 4, the following steps are implemented: - calculate the energy Eβι of the signal resulting from filtering XiítJ) by Petition 870250089800, dated 02 / 10 / 2025, pp. 35 / 61 28 / 31 - calculate the energy Eδ2 of the signal - Calculate a ratio between Εβι and Εδ², - compare the said ratio with an SE energy threshold.

[00117] If the ratio between Δι and Δβ2 is below SE, the presence of a transition is considered detected. If the ratio between Δι and Δθ2 is above SE, the presence of a transition is considered detected. For this purpose, a parameter Par 1 (if T1 was chosen at the end of the selection in Figure 3) or Par2 (if T2 was chosen at the end of the selection in Figure 3) is defined: - to a first value, for example 0, in order to indicate that the presence of a transition was not detected, - to a second value, for example 1, in order to indicate that the presence of a transition has been detected.

[00118] This parameter Par1 / Par2 is transmitted, as a decision criterion, to the cross-fading block 612.

[00119] Thus, according to the invention, when a transient is detected in the current frame in one or both channels, the T3 reduction mix is ​​applied to the current frame instead of the T1 (or T2) reduction mix. Specifically, the T1 and T2 reduction mixes use impulse response filtering which sometimes has a tendency to spread the signal envelope; this problem is more pronounced in transient sounds, such as castanets, scissors clicking in a binaural recording simulating a haircut in a hair salon, etc.

[00120] In order to limit potential artifacts due to crossfading (block 612) between reduction mixing methods, the invention may provide for the application, in 614, of signal scaling at the output of T1 (or T2). According to the invention, such scaling 614 implements the following steps: - Determine the energy E' of the reduction mixing signal. Petition 870250089800, dated 02 / 10 / 2025, pp. 36 / 61 29 / 31 at the output of T1 (or T2), preferably with smoothing, performing the same steps as steps 708, 709 and 710 of Figure 7, - Determine the scaling factor: g = -—V 2Sf - Apply the g' factor to the aforementioned reduction mixing signal, returning to the operating principle (with smoothing at the beginning of the frame) of blocks 712 and 714 in Figure 7.

[00121] Unlike the reduction mix in Figure 7, a direct compensation of the reduction mix level at the output of T1 or T2 is applied here, without reinjecting an input signal.

[00122] Figure 6b illustrates an alternative to the channel reduction processing method described with reference to Figure 6a.

[00123] The mode illustrated in Figure 6b is distinguished from Figure 6a only by the fact that a transition is detected according to a second mode. According to this second mode, such detection is implemented during the T1 (or T2) reduction mixing step 610 instead of being applied separately to each of the input signals xi(n) and X2(n), or to the current frame xi(f, / ) and X2(f, / ). For this purpose, with reference to Figure 5, the energy of each 2 ms sub-block is compared between the input and output signals of blocks 502 and 503. If the energy has dropped a threshold, for example 3 dB, it is considered that a transient has been sufficiently attenuated and, in this case, the T3 reduction mixing is applied to the current frame. If the energy has not dropped a threshold, for example 3 dB, it is considered that a transient has not been attenuated and, in this case, the T1 (or T2) reduction mix is ​​applied to the current frame.

[00124] Figure 8 illustrates an 800 channel reduction processing device, within the meaning of the invention.

[00125] The 800 device comprises a processing circuit that typically includes: Petition 870250089800, dated 02 / 10 / 2025, pp. 37 / 61 30 / 31 - a MEM1 memory for storing instruction data of a computer program within the meaning of the invention; - an INT1 interface to receive a stereo audio signal: íxiOO. Ιχ2{ή)' - a PROC1 processor to receive this signal and process it by executing the instructions of the computer program stored in MEM1 memory, in order to perform the reduction mixing; in particular, the processor is capable of controlling the processing modules as described with reference to Figures 2 to 5; and - a COM 1 communication interface to transmit the reduced signals, resulting from the reduction mixing, fn(tj), to another processing module, for example, a code conversion module or a decoding module of an audio signal decoder or converter.

[00126] Naturally, this Figure 8 illustrates an example of a structural embodiment of a channel reduction processing device within the meaning of the invention.

[00127] Figures 2 to 7, discussed above, describe functional modes of this device in detail.

[00128] Applications of this type of channel reduction processing include, for example, audio-to-code conversion, such as when the remote control lacks the capabilities to render stereo sound on its terminal. In this case, it is not necessary to transport a stereo signal, thus saving bandwidth. This type of method can work during point-to-point conversations if one of the participants makes a stereo or binaural sound recording. This type of method can also be present during a multi-party call: the conference bridge spatializes the scene. Petition 870250089800, dated 02 / 10 / 2025, pp. 38 / 61 31 / 31 creating a stereo scene, but not all participants necessarily have stereo rendering capabilities. The scene must then be reduced to mono for those participants.

[00129] The invention can also be applied to audio decoding: In this case, it is the user's renderer (rendering module) that will mix the stereo content by downscaling to adapt to the user's reduced capabilities (a mere loudspeaker, for example). Petition 870250089800, dated 02 / 10 / 2025, pp. 39 / 61

Claims

1 / 3 CLAIMS 1. A method for processing a stereophonic signal to obtain a monophonic signal, characterized in that it comprises the following: - applying (610, 611) to a current frame of said signal, two channel reduction processing operations (T1 or T2, T3), one of the channel reduction processing operations using rephasing filtering in at least one of the channels of the stereophonic signal and the other reduction processing operation not using rephasing filtering, - selecting (612) one of said processing operations to be applied to the current frame, said selection being implemented depending on: - the presence or absence of at least one transient in the stereophonic signal, or - a quality level of the stereophonic signal filtering used when implementing the processing operation using stereophonic signal filtering.

2. Channel reduction processing method, according to claim 1, characterized in that, if the presence of said at least one transient is detected (613), or not detected (613), in the stereophonic signal, the channel reduction processing, which does not use stereophonic signal filtering or which uses stereophonic signal filtering, is selected (612), respectively.

3. A channel reduction processing method according to claim 1, characterized in that said at least one transient is detected before either of the two channel reduction processing operations is applied to said stereophonic signal.

4. Channel reduction processing method, according to any of the preceding claims, characterized by the fact that the detection of said at least one transient comprises the following steps: - decomposing the current frame of the stereophonic signal into sub-blocks; - calculating the energy of each of the sub-blocks obtained; - comparing the energies obtained with an energy threshold (th61); and - if the energy is less than or greater than said energy threshold, respectively, the presence of a transient will not be detected or will be detected, respectively.

5. A channel reduction processing method, according to claim 1, characterized in that said filtering quality level is compared with a filtering quality threshold and, if said filtering quality level is below, or above, said filtering quality threshold, the channel reduction processing operation, which does not use stereo signal filtering or which uses stereo signal filtering, will be selected.

6. Channel reduction processing method, according to claim 5, characterized in that said filtering quality level is measured as follows: - calculate the energy of the stereophonic signal, referred to as the input energy, before filtering is applied, - calculate the energy of the monophonic signal, referred to as the output energy, after filtering is applied, - calculate a ratio between the input energy and the output energy.

7. Channel reduction processing method, according to claim 1, characterized in that an indicator (Par1 or Par2) is generated representing the presence or absence of a transient. Petition 870250089800, dated 02 / 10 / 2025, page 41 / 61 3 / 3 8. Channel reduction processing method, according to any one of claims 1 to 6, characterized in that two levels (E', E1+E2) of said processed current frame are respectively obtained after the two channel reduction processing operations (T1 or T2, T3), one using filtering of the stereophonic signal and the other not, are applied to said current frame of said signal, - adjusting (614) said level (E') of the current frame to which the channel reduction processing operation, which uses said filtering, was applied, said adjusted level (E') being similar to or equal to said level (E1+E2) of the current frame to which the channel reduction processing operation, which does not use said filtering, was applied;- Compare the current processed frame that has the said adjusted level and the current processed frame that has the said level obtained after the channel reduction processing operation, which does not use stereophonic signal filtering, is applied to the said current frame of the said signal; - Select one of the said processing operations to be applied to the current frame, based on the said comparison.

9. Channel reduction processing device, characterized in that it comprises a processing circuit for implementing the steps of the channel reduction processing method, as defined in any one of claims 1 to 7.

10. A storage medium that can be read by a processor, characterized in that it stores a computer program comprising instructions for executing the method, as defined in any one of claims 1 to 7. Petition 870250089800, dated 10 / 02 / 2025, pp. 42 / 61