Optimized processing to reduce the number of channels in a stereo audio signal.

BR112025020986A2Pending Publication Date: 2026-08-25
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025020986
Authority / Receiving Office
BR · BR
Patent Type
Applications
Publication Date
2026-08-25

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

1 / 36 Optimized processing to reduce channels of a stereo audio signal. DESCRIPTIVE REPORT Technical Field

[0001] The present invention relates to the general field of audio signal processing. The invention relates, in particular, to the reduction mixing of a multichannel audio signal. The reduction mixing of a stereophonic signal to a monophonic signal is of particular interest.

[0002] This type of processing is generally applicable in the field of audio technologies and, more specifically, in the field of audio encoding, whether in the encoding or decoding stage. Prior Technique

[0003] Reduction mixing consists of deducing, based on a combination of the channels C of a multichannel signal x, a signal y that consists of a smaller number D of channels. In practice, this consists of defining a function f(.) such that: m(n) = m1(n)' .mD(n). f (x(n)) = f / Γ%ι(η)1\ \Lxc(n)J / with D < C en representing a time index of the input and output signal.

[0004] In this invention, the particular case D=1 and C=2 is of interest; Reference will then be made to stereo-to-mono reduction mixing or simply mono reduction mixing. In the case C=2, reference is made to 2-channel content or signal, which can be stereo content or binaural content, where channel 1 (x1) is referred to as the left channel and channel 2 (x2) as the right channel. Below, the stereo case will be seen as a general 2-channel signal, which includes the binaural case, in order to avoid repeating the two terms, even though technically a binaural signal has specific characteristics. In the case of the signal y obtained after reduction mixing, reference is made to Petition 870250088391, dated 09 / 30 / 2025, p. 9 / 63 2 / 36 a monophonic or mono signal, which will be denoted m(n) in the time domain, or M(k) in the frequency domain, below.

[0005] These stereo signals can result from capture by a pair of stereo microphones or binaural capture, or even from artistic mixing of audio tracks.

[0006] There are a number of microphone pairs that make it possible to create stereo content. Among the most popular, we can mention the XY pair, composed of 2 cardioid microphones that exhibit a difference in orientation angle between 90° and 135°. The MS (mid-side) pair is another widely used pair composed of a cardioid microphone and a second microphone, called a Figure 8 microphone, oriented at 90°. By combining these 2 microphones (sum / difference), the left and right stereo content channels can be created. These pairs are called coincident pairs, meaning that the microphone capsules do not exhibit a delay between them, with spatialization perceived by the difference in sound intensity between the left and right channels.

[0007] Another category of pairs, called phased stereo pairs, consists of using microphones that are far apart from each other, thus creating a phase difference between the channels for sources located closer to one of the microphones. The best known is the AB pair, which uses 2 omnidirectional microphones separated by a few centimeters to several meters. Another widely used pair is the ORTF pair, which uses a phase difference through a 17 cm spacing of the microphones and an amplitude difference through cardioid directivities (the microphones exhibit an angle of about 90°). The binaural pair consists of placing omnidirectional microphones in the ears of a person or an artificial head: this type of device makes it possible to create native binaural content, which can be heard through headphones. Petition 870250088391, dated 09 / 30 / 2025, p. 10 / 63 3 / 36

[0008] Stereo content also includes speech and audio signals resulting from audio mixing or audio post-production (e.g., channels stored on a CD or DVD or transmitted over the Internet).

[0009] There are also other procedures for creating stereo content, which are not reviewed here.

[0010] The simplest procedure for creating a reduced signal by reduction mixing is known as passive reduction mixing. It consists of obtaining an average of the left and right channels of the stereo signal such that: X1(t)+%2(t) m(t) =----------or, in the frequency domain:

[0011] M(t,k) =Xi(t'k)+X2(t'k)where Xi(t,k'),i = 1,2 is the fast Fourier transform xi(t) (FFT), where k is the frequency index and t is the frame index: j2nlk Xi(t, k) = YlÍ-ow(^)-xi(t> 0e^-, for any k E {0,..., N — 1} where j is the imaginary number such that j = V—1, t is the frame index, l is the time index in the current frame, N is the size of the FFT, w(.) is a sine or Hann or other apodization window of size L, adapted to the frame size, and xt(t, l) = xi(t, l). This definition extends to the case where N>L and jci(.) is an augmented version of xi(t) with 0s (0-filling).

[0012] This passive reduction mixing procedure, which is very simple, works well in situations where the right and left channels are in phase. However, in situations where the microphones are not coincident and exhibit phase differences, such mixing generates combined filtering, which is manifested by the coloration of the original signal. This is due to the fact that, depending on the spacing of the microphones, the position of the source relative to the microphone pair, and the Petition 870250088391, dated 09 / 30 / 2025, page 11 / 63 4 / 36 frequency, the left and right channels of the same source can be in phase, as shown in Figure 1a, or out of phase as shown in Figure 1b. Thus, when performing a reduction mix, the left and right signals will add (constructive interference) or cancel each other out (destructive interference), creating variations in the level and coloration effects of the reduced signal m(t). The coloration results from the combined filtering effect. The effects of constructive and destructive interference manifest themselves in different frequency bands, thus modifying the balance and therefore the tone of the resulting signal from the reduction mix compared to the original.

[0013] In order to correct this intensity level defect according to frequencies, the e-AAC+ codec's reduction mixing implements y(k) level correction which ensures that the energy level of the reduced signal, in each frequency band, remains comparable to the original: M(k) = r(k).X1(k) + X2(k), with + y^ ' a| 0.5lXi(k)+X2(k)l2

[0014] This approach makes it possible to compensate, to some extent, for the drop in intensity of the reduction mix. However, this compensation leads to overamplification when the signals are close to phase opposition. Also, in practice, the correction factor y(k) is limited to a maximum value (for example, an upper limit of 2).

[0015] Figure 7 describes a procedure for reducing stereo channels to mono, which is integrated into the public candidate IVAS codec that is available at: https: / / forge.3gpp.org / rep / ivas-codec-pc / ivas-codec / - / tree / main

[0016] The signal x(m) with 2 input channels, where m is the index of the interleaved samples, is deinterleaved (block 701) to find the left and right channels. Then, an analysis is performed. Petition 870250088391, dated 09 / 30 / 2025, page 12 / 63 5 / 36 frequency with windowing and Fourier transform (blocks 702 and 703) in order to obtain the spectra of the 2 channels, and the intercorrelation based on the phase spectra is estimated (block 704) before determining the time difference between the channels (here marked as ITD although it is formally an ICTD between 2 channels) looking for a peak (block 705); block 705 also provides the correlation level R* corresponding to the ITD.

[0017] The mixing factor g is then determined as follows (block 706): R* 1R* where g represents the mixing factor and τ is the time difference between channels (incorrectly called ITD). Thus, the 2 channels are combined into a mono signal in block 706 according to this factor g, with smoothing of the factor value applied sample by sample at the beginning of each frame after block 707: m(n) = γ(η) X1(n) + (1-γ(η)) xs(n), where γ(η) is the factor derived from g after temporal smoothing with the gain from the previous frame.

[0018] Next, the energies of the left channel X1(n) and right channel X2(n) and the mono signal m(n) are determined (blocks 708, 709, 710). The left and right channels are separately reinjected (added) in blocks 713 and 715 into the resulting mono signal from block 706, depending on the energy compensation determined in block 711 with respective scaling factors defined in blocks 712 and 713.

[0019] These procedures, in addition to the coloration of the resulting signal from the reduction mix, linked to imperfectly corrected combined filtering, suffer from excessive reverberation related to the summing of signals when the input channels are out of phase and from an artifact of Petition 870250088391, dated 09 / 30 / 2025, page 13 / 63 6 / 36 click when channels are reinjected point by point into a single frame. Compensating for an identical delay for all frequencies is unrealistic; this delay is linked to the various sources that make up the content. In practice, this leads, in addition to further coloration, to a potential loss of intelligibility of the sources, which defect-level compensation cannot correct.

[0020] Other downmixing approaches seek to avoid the aforementioned defects. The principle of these procedures is to rephase the left and right signals before they are summed. In the document entitled 'A stereo to mono downmixing scheme for MPEG-4 parametric stereo encoder' by Samsudin, E. Kurniawati, N. Boon Poh, F. Sattar, S. George, in Proc. ICASSP, 2006, a procedure is proposed that consists of applying, in the frequency domain, a phase shift φ to one of the channels, usually the right one, in order to rephase it with the left channel, so that: X1(k')+X2(k).e-^P^ M(fc) =------------------

[0021] The ideal phase shift is determined by what is referred to here as IPD (phase difference between channels): iPD(k) = z(X1(k).%2(k)) where X£ is the conjugate of X and z indicates the phase of the complex operand.

[0022] It is calculated in the frequency domain: this makes it possible to rephase the spectral components of several sources, a source with its own IPD being able to be preponderant in a frequency band k, while the other source with a different IPD may be preponderant at another frequency k.

[0023] In the case of the MPEG4 codec described in Samsudin's document cited above, the IPD applied is that estimated by averaging over a Bark band and not the IPD calculated for each frequency band, the latter being particularly noisy and variable from one frame to the next. Furthermore, this procedure takes the left channel. Petition 870250088391, dated 09 / 30 / 2025, page 14 / 63 7 / 36 as a phase reference, and if the phase of this channel is poorly conditioned, the reduction mix will have degraded quality.

[0024] In published patent application no. WO2017103418, a reduction mixing procedure is proposed that combines the procedure proposed by Samsudin and passive reduction mixing. Notably, it proposes an ISD (interspectral distance) indicator that allows the selection, frequency band by frequency band, of the most suitable procedure for rephasing the signals: |Xi(k)-x2(fc)I ISD(fc) = -——---ijιχ^+^ωι

[0025] (ISD < This indicator makes it possible to measure whether the signals are in phase 1), that is, with a phase difference or IPD in the interval or better in the opposite phase (ISD > 1), that is, with an IPD in the interval π 3π .2,2. When they are in opposite phase, a high value ISD is synonymous with a comparable level between the left and right channels, and rephasing is important to avoid coloration and dropout effects. On the other hand, an ISD value close to 1 is representative of one channel being predominant over the other; in this latter case, rephasing is unimportant, since the reduction mix is ​​almost equal to the predominant signal. However, it is important to avoid modifying the phase of the right channel if it is predominant. Thus, in this patent application, the following reduction mix selection mechanism is proposed: X,U) + X2U).e;H>ÍM(fc) = —-------, if ISD(fc) > 1.3 Xi(fc)+X2(fc) M(fc) =------------if ISD (fc) < 1.3

[0026] The passive reduction mixing, proposed in the aforementioned patent application, for low ISD situations, sometimes introduces increased reverberation.

[0027] Furthermore, it was observed that in certain situations where Petition 870250088391, dated 09 / 30 / 2025, p. 15 / 63 8 / 36 The signals are somewhat out of phase, that is, when ISD(k) < 1.3, and in particular when the channels are highly unbalanced (source mainly on the left or right), a potentially unpleasant loss of timbre occurs and therefore it may be important in these cases to keep the signal intact, i.e., not to apply a phase shift.

[0028] Furthermore, the procedure proposed in the aforementioned patent application requires the application of a frequency domain processing operation, whereby the reduced signal is reconstructed based on the short-term inverse Fourier transform (STFT) of the signal M(t,k). This type of filtering generates a circular convolution that creates audible artifacts: echo, pre-echo, and static. In order to mask these artifacts, an implementation with frame overlap associated with appropriate windowing (analysis and / or synthesis window) can be used, which guarantees the reconstruction of the filtered signal. Generally, this overlap is not compatible with the operation of audio encoders that work with adjacent frames, i.e., without overlap; furthermore, this type of reconstruction leads to a delay, usually half a frame, which is the conventional overlap, which is not acceptable for a code encoder that has strong latency constraints.

[0029] Another implementation of overlay and addition (OLA) is also possible: this procedure requires a zero-filling step of the signals and filters to avoid circular convolution. However, tests show that with these filters, whose phase varies very rapidly from one frame to another, an OLA approach does not make it possible to completely mask artifacts in the transition of certain frames; the artifacts remain audible.

[0030] There are also prior art procedures in which reduction mixing is achieved by switching, for example, in the frequency domain in published patent application no. WO2017103418, Petition 870250088391, dated 09 / 30 / 2025, page 16 / 63 9 / 36 of different reduction mixing procedures. In this case, it is important to ensure that switching occurs continuously, i.e., without discontinuities or differences in levels between procedures to avoid artifacts. Description of the Invention

[0031] The invention aims to improve the previous technique.

[0032] For this purpose, the invention relates to a method of reducing a stereophonic signal in order to obtain a monophonic signal comprising, for a current frame: - a first stage of selecting a reduction mixing operation from two procedures that use filtering of the stereo signal and rephasing of the signals, with the selection being made according to a phase indicator, representative of a measure of a degree of phase opposition between the frequency components of the stereo signal channels, - a second selection stage between one of the two procedures, selected at the end of the first selection stage, and a reduction mixing procedure without rephasing, with the selection being made according to the value of an energy ratio between the channels of the stereo signal.

[0033] These selection steps thus make it possible to determine the type of reduction mix that is most suitable for the characteristics of the stereophonic signal.

[0034] In one mode, the power ratio between the channels of the stereo signal is calculated by frequency band, and the second selection stage is performed by frequency band.

[0035] Thus, only the relevant frequency bands, whose energy ratio exceeds a threshold, apply reduction mixing without rephasing to achieve finer matching.

[0036] In a particular embodiment, the mixing operation of Petition 870250088391, dated 09 / 30 / 2025, page 17 / 63 10 / 36 reduction without rephasing is selected for relevant frequency bands only if the number of relevant frequency bands exceeds a threshold.

[0037] This makes it possible to avoid excessive fluctuations from one frame to another.

[0038] In one embodiment, a third selection stage is performed between one of the two procedures selected at the end of the second selection stage and a third reduction mixing procedure that does not use stereo signal filtering, with said selection being made depending on: - the presence or absence of at least one transition in the stereophonic signal, or - a level of quality of stereophonic signal filtering used when implementing a processing operation that uses stereophonic signal filtering, or - an indication of the presence of a binaural signal.

[0039] This option allows: - When one or more transitions are detected, obtain a better rendering of the signal transitions by selecting the reduction mixing operation that does not use filtering and, as a result, preserves the nature of the stereophonic signal. - When no transition is detected, favor the reduction mixing operation that uses filtering, which is more suitable for the components of the stereo signal that are not in phase.

[0040] In one embodiment, the detection of at least one transition is implemented before either of the two reduction mixing operations is applied to the stereo signal.

[0041] This method makes it possible, even before a reduction mixing operation is applied, to deduce a priori whether at least one transition is present or not present in the stereophonic signal, Petition 870250088391, dated 09 / 30 / 2025, page 18 / 63 11 / 36 without any particular additional processing operation being applied to that signal.

[0042] In one embodiment, the detection of said at least one transition is implemented as follows: - Calculate the energy of the stereo signal, referred to as the input energy, before filtering is applied. - Calculate the energy of the monophonic signal, referred to as the output energy, after filtering has been applied. - Calculate the ratio between input energy and output energy. - compare said ratio with an energy threshold, and - If the said ratio is below, or above, the said energy threshold, the reduction mixing operation, which does not use stereo signal filtering or which uses stereo signal filtering, will be selected, respectively.

[0043] In one embodiment, said filter quality level is compared with a filter quality threshold and, if said filter quality level is below, or above, said filter quality threshold, the reduction mixing operation that does not use stereo signal filtering or that uses stereo signal filtering will be selected, respectively.

[0044] In one embodiment, an indicator is generated that is representative of the result of said detection, of comparing said ratio with an energy threshold, or of comparing the level of filtering quality with a filtering quality threshold.

[0045] In one embodiment, the method comprises a preliminary phase of adapting reduction mixing filters before application to the stereo signal through the following steps: - Obtain impulse responses from the filters corresponding to the reduction mixing operation; Petition 870250088391, dated 09 / 30 / 2025, page 19 / 63 12 / 36 - truncate part of the impulse responses; - to weigh the remaining portion by applying a weighting window; - Normalize the impulse responses resulting from windowing to obtain the appropriate filters that should be applied to a stereo signal frame.

[0046] In a particular embodiment, the truncation step consists of preserving a causal part and a non-causal part of the impulse responses.

[0047] The invention targets a reduction mixing device comprising a processing circuit to implement the steps of the reduction mixing method as described above.

[0048] The invention also relates to a computer program comprising instructions for implementing the reduction mixing method that is in accordance with the invention, according to any of the particular embodiments described above, when said program is executed by a processor.

[0049] Such instructions can be durably stored in a non-transient memory medium of the reduction mixing device that implements the reduction mixing method according to the invention.

[0050] This program can use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form or in any desired form.

[0051] The invention also targets a computer-readable storage medium or an information medium comprising instructions of a computer program as described above.

[0052] The storage medium may be any entity or Petition 870250088391, dated 09 / 30 / 2025, page 20 / 63 13 / 36 device capable of storing the program. For example, the medium may comprise storage media such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or indeed a magnetic storage medium, for example, a mobile medium, a hard disk or an SSD.

[0053] Furthermore, the storage medium may be a transmissible medium, such as an electrical or optical signal, which may be routed through an electrical or optical cable, by radio or by other means, so that the computer program it contains may be executed remotely. The program according to the invention may, in particular, be transferred by download from a network, for example, an Internet network.

[0054] Alternatively, the storage medium may be an integrated circuit in which the program is embedded, the circuit being adapted to execute, or to be used in the execution of, the aforementioned reduction mixing method.

[0055] According to an example of an embodiment, the present technique is implemented by means of software components and / or hardware components. With this in mind, the term device or module may correspond, in this document, equally to a software component, a hardware component, or a set of hardware and software components. Brief description of the drawings

[0056] Other features and advantages of the invention will become more clearly apparent by reading the following description of particular embodiments, which are given by way of illustrative and non-limiting examples only, and the accompanying drawings, among which: [Figure 1A] illustrates channels of a stereophonic signal, in phase as described above; [Figure 1B] illustrates channels of a stereophonic signal, in Petition 870250088391, dated 09 / 30 / 2025, page 21 / 63 14 / 36 phase opposition as described above; [Figure 2] illustrates, in the form of a block diagram, a reduction mixing sequence in one embodiment of the invention; [Figure 3] illustrates one embodiment of a filter calculation for reduction mixing, as well as a selection of a reduction mixing operation; [Figure 4] illustrates one embodiment of an adaptation phase of a reduction mixing filter; [Figure 5] illustrates one mode of applying reduction mixing during a transition between two stereo signal frames; [Figure 6a] illustrates another embodiment of a filter calculation for downmixing, as well as a selection of a downmixing operation; [Figure 6b] illustrates another embodiment of a filter calculation for downmixing, as well as a selection of a downmixing operation; [Figure 7] illustrates an existing reduction mixing procedure described above; [Figure 8] illustrates an example of a structural embodiment of a reduction mixing device according to an embodiment of the invention. Description of the modalities

[0057] The invention will now be described below with reference to Figures 2 to 8. For this purpose, the following symbols will be used primarily in the following description. INDICES k frequency index l time index of a signal (in the current frame) n time index of a signal Petition 870250088391, dated 09 / 30 / 2025, page 22 / 63 15 / 36 t number of frames CONSTANTS L is the number of samples in the current table. R filter delay N length of FFT C number of input channels D number of output channels P filter length Q crossfading length SIGNALS xi(n) input channels m(n) downmixing (in the time domain) m(t,l) downmixing in the current frame (in the time domain) M(t,k) downmixing in the current frame (in the frequency domain)

[0058] Figure 2 shows an example of a sequence for processing a stereophonic audio signal, where the processing operation comprises a downmixing operation.

[0059] At the input of this processing sequence, a stereophonic signal x, composed of two channels (x1(n) and x2(n)), also called the left channel and right channel (respectively), is, in a first stage, divided, in 201, into sample frames L (xi(t,l) and x2(t,l), t being the frame index and the sample index. Block 202 applies windowing and an FFT (for the Fast Fourier Transform) in order to obtain signals in the frequency domain (Xi(t,k) and X2(t,k), k being the frequency index). In 203, a reduction mixing filter to be applied to a signal frame is then selected and determined. This step will be described with reference to Figure 3. The filter thus determined is conditioned (or adapted) in 204 in order to make it causal or partially causal, and to optimize it in order to reduce processing complexity. In an example of a modality, the window Petition 870250088391, dated 09 / 30 / 2025, page 23 / 63 16 / 36 mento and FFT are similar to steps 702 and 703 in Figure 7.

[0060] This adaptation phase is described with reference to Figure 4.

[0061] Once one or more filters have been determined and adapted for the current frame, they are applied, in 205, to this current frame. In order to avoid audible artifacts that are due to a change of filters between two frames, a crossfading is performed leÍá1(t- 1, taking into account the filters of the previous frame {1\ ' when al ^2 Çt 1, L) the previous frame's reduction mix is ​​of the same type (T1 / T2 defined below).

[0062] This step will be described with reference to Figure 5.

[0063] Figure 3 describes a detailed embodiment of the block 203 of Figure 2. In this processing block, in a first selection step, a first or second reduction mixing operation to be applied to a current frame of the stereophonic signal is selected and determined.

[0064] A first T1 reduction mixing procedure is defined here by shifting the left and right channels of the stereo signal simultaneously by an angle defined by a quarter of the phase difference (IPD) determined between the two channels of the stereo signal. Thus, there is a reduction mixing signal, as follows: .IPD(t,k) _ ,IPD(t,k) M(t fc) =Xi(t'k^.e 1 4+X2(t,k\e1 4

[0065] This procedure makes it possible to avoid the degradation of the excess reverberation type generated by conventional reduction mixing (averaging the two signals) or by combined filtering in the case of a time shift between the two channels.

[0066] This reduction mixing defined in this way can be seen, in the frequency domain, as filtering according to the following equation: Petition 870250088391, dated 09 / 30 / 2025, page 24 / 63 17 / 36 M(t,k) = H1(t,k').X1(t,k') + H2(t,k).X2(t,k) with H1(t,k) = / PD(t,fc)e—— and H2(t,k) = .IPD (t,k) e-]---4--2

[0067] In another embodiment, the first T1 reduction mixing procedure is defined by shifting the left and right channels of the stereo signal simultaneously by an angle defined by half the phase difference (IPD) determined between the two channels of the stereo signal.

[0068] So: M(t, k) ,IPD(t,k) _ ,IPD(t,k) 71(^).6+ 2 +X2(t,k').e12 or equivalently: M(t,k) = H1(t,k').X1(t,k') + H2(t,k').X2(t,k') with H1(t,k') = ^PD(tk) + > 2 . . —-— and H2(t,k = .IPD (t,k) e-}---2--

[0069] A second T2 reduction mixing procedure is, for example, defined as that proposed in the cited Samsudin paper. In this second procedure, then, H1(t, k)=1 ee-jIPD(t,k) . 2 As described below, in one embodiment, the corresponding impulse responses, hi(t,l'), are in fact shifted in the above.

[0070] H2(t,k) =

[0071] time to maintain causal and anticausal parts. This shift can, in variants, be integrated directly into the definition of Hi(t,k).

[0072] Furthermore, in variants, it is equivalently possible to remove the factor1Λ from the definition of Hi(t,k) and integrate this factor1Λ during the mixing of the channels processed in 504 and 505 (see Figure 5 described below).

[0073] In practice, the IPD or IPD / 4 or IPD / 2 phase difference does not need to be determined. In an efficient implementation, the filters Petition 870250088391, dated 09 / 30 / 2025, page 25 / 63 18 / 36 11,(.,k) can be determined by normalizing the (complex) frequency lines - with or without smoothing - by their modulus (in the complex number sense).

[0074] The stereo signal channels, in the frequency domain (X1(t,k) and X2(t,k)), are used, on the one hand, to calculate an ISD phase indicator in 301, which is representative of a measure of the degree of phase opposition between the stereo signal channels and, on the other hand, the phase difference between the channels (IPD) in 303.

[0075] The phase indicator is, for example, defined by the indicator ISD (interspectral distance), as defined above.

[0076] In the mode described here, the ISD is determined by the stereophonic signal frame so that the decision to select a first reduction mix or a second one can be made frame by frame.

[0077] This frame-by-frame decision makes it possible to eliminate substantial phase shifts between frequencies and other audible artifacts that can appear when the decision varies from one frequency to another, as in the patent application cited above.

[0078] To apply a frame-by-frame decision, a decision criterion according to the ISD phase indicator is defined according to the following equation in 302: _ total number of frequency lines where ISD(t, k) > 1.3 total number of frequency lines

[0079] This criterion measures the percentage of frequency lines with index k for which the criterion favors the reduction mix with rephasing, i.e., when ISD(t,k) > 1.3. The decision is to apply the first reduction mix as defined above, with rephasing, when ISD(t) exceeds a certain threshold ISDth, which is defined, for example, at 70%.

[0080] In another modality, the ISD indicator will potentially Petition 870250088391, dated 09 / 30 / 2025, page 26 / 63 19 / 36 determined in a predefined frequency band:IIsB(t) =1+ιΣ',^15B(ί^) 'max 'min+ 1 ml“ where, for example, kmin= 2 and kmax is the frequency line number corresponding to a maximum frequency, for example, 4 kHz.

[0081] In this case, this criterion is an average of the ISD value per frequency line. In this variant, the threshold does not correspond to a percentage, but to an ISD value and, for example, will possibly be defined as ísBth= 1.5.

[0082] There is, then, in 305, the following decision to select a reduction mix between T1 and T2: <|,IPD(t,k) IPD(t,k) isB(t)<ísBth·· H1(t,k') =e* , H2(t,k) =e4 2 ísB(t) > ísBth : H1(t,k) =1, H2(k,f) =e ]^'^

[0083] In 303, the phase difference IPD is obtained between the two channels of the stereo signal by IPB(t,k) = z(x1(t,k).X2,(t,k)') where z indicates the phase of the complex operand.

[0084] However, the decision of the framework may sometimes be insufficient, as the IPD, calculated according to the formula above, varies very rapidly over time. In addition, the filters

[0085] Hi(t,k) can vary very rapidly from one frame to the next, which generates discontinuities that are sometimes audible during the transition from one frame to the next.

[0086] To limit this effect, according to a first approach, the IPD will potentially be calculated in frequency sub-bands instead of in each band independently, while averaging each channel over wider sub-bands. A number B of sub-bands of width equal to BK = K / B will potentially be chosen, where B is an integer greater than or equal to 1. To calculate the IPD in each sub-band b, initially the spectra X^t,^ in Petition 870250088391, dated 09 / 30 / 2025, page 27 / 63 20 / 36 each sub-band are calculated by averaging: X(it,b) = T,'=^-1X<:t,i + b»Bk)

[0087] The average spectra of the sub-bands corresponding to zero frequency (b=0) and Fe / 2 frequency (b = B / 2) are processed separately by direct deduction from the spectra in these frequency bands: X'(t,0) =X'(t,0) )l'(t,B / 2) = X'(t,K / 2)

[0088] The B number of sub-bands must be determined experimentally. You need to maintain sufficient frequency resolution to be able to apply proper rephasing to the various sources: in practice, a frequency resolution of about one hundredth of a hertz seems like a good bandwidth. For example, for a stereo signal sampled at 32 kHz, with an FTT size of 640, this resolution corresponds to a B number of 320 sub-bands with BK = 2 frequency bands in each sub-band.

[0089] In variants, the sub-bands can be divided to have a non-uniform width, for example, according to the Bark scale.

[0090] Next, the IPD is calculated in the same way as above, that is: IPD(b) = z(J?i(b).Jf2*(b))

[0091] Below, the index k will indicate a frequency band or a frequency sub-band, where the case BK=1 is equivalent to performing processing in each frequency band separately.

[0092] Again, to limit the effect of rapid variations in the IPD, a time-smoothed version of the IPD will potentially be used to calculate the H'(t,k) filters. One way to smooth the IPD is to apply first-order IIR low-pass filtering at 304, such that: Petition 870250088391, dated 09 / 30 / 2025, page 28 / 63 21 / 36 TPÕ(t,k) = a(k')TPD(t,k') + (1- a(k'))lPD(t,k)

[0093] The forgetting coefficient a(k) is chosen ad hoc. It can be advantageous to choose a high forgetting coefficient at lower frequencies, typically below 5 kHz, as this allows for maintaining phase coherence in the lower part of the spectrum from one frame to the next. This is because the spectrum of speech signals is stationary, notably in its lower part, and the ear is particularly sensitive to phase in this part of the spectrum. Conversely, it is advantageous to choose a low forgetting coefficient at higher frequencies: at higher frequencies, typically above 5 kHz, the phase continuity of speech signals is less pronounced, and the ear is less sensitive to phase discontinuities in this part of the spectrum.

[0094] A forgetting coefficient of the following form can be chosen: a(k) =amax,amín, . kk-lowamín + ,kaltokbottomk <kbaixok>kaltokbaixo <k<kalto

[0095] The values ​​a, amín, kaito and kbaiXo are determined experimentally. In the case of adjacent frames of length L = 20 ms, the following set of parameters can, for example, be chosen: amax — 0.94, amin — 0.86, klow — 0, khigh, corresponds to 5KHz.

[0096] In some variants, very low frequencies and very high frequencies will be forced to remain intact. Thus, in some variants, H1(t, k) — H2(t, k) — 1 is forced for the first k — 0 or k — 0.1 lines as well as for the lines with index k > kmax, where k corresponds to 8 kHz.

[0097] Reduction mixing with rephasing (T1 or T2) can sometimes substantially alter the timbre and cause a decrease in signal quality. These timbre changes are particularly Petition 870250088391, dated 09 / 30 / 2025, p. 29 / 63 22 / 36 audible when a signal is predominant in any channel, that is, when the ILD (Interchannel Level Difference), which measures the energy ratio between the left and right channels, is close to 0 or much greater than 1. This is, for example, the case of a predominant signal in the right channel to which the T1 rephasing mode based on the IPD is applied. To limit these timbre changes, it is advantageous, when the energies of the left and right channels are very different, to avoid rephasing by applying conventional reduction mixing.

[0098] In an initial implementation, an independent decision will potentially be made in each frequency band, allowing multiple sources (one of which may be predominant in one channel and not require rephasing, while the others are distributed between the left and right channels and require rephasing) to be processed independently. In this case, a passive reduction mix will only be enforced in frequency bands that exhibit a high-energy disparity: H1(t, k) = H2(t,k) = —,Vk,se ILD(t,k) > aILDou ILD(t,k) < 1 / aILDem que ILD(t,k) =IX1(t,k)le aILDé um threshold de ILD de \X2(t,k)\2experimental. In practice, a value of around 100 (20 dB) will possibly be chosen.

[0099] Thus, the ILD is calculated in 306 in order to implement a second selection step in 307, and thus determine the reduction mix to be applied with or without rephasing T1, T'1, T2 or T'2 according to the frequency bands.

[0100] The decision to apply passive reduction mixing independently to each sub-band can sometimes cause phase fluctuations from one frame to another: this is particularly audible when few frequency bands are involved. To avoid these Petition 870250088391, dated 09 / 30 / 2025, page 30 / 63 23 / 36 fluctuations, the application of passive reduction mixing will potentially be limited to the relevant bands only if a sufficient number of relevant bands are observed, indicative of the actual presence of a left or right source. In practice, it is possible to calculate the percentage of relevant bands ILDprc(t) and limit the application of passive reduction mixing to the relevant bands defined above in the case where this percentage is greater than a predefined threshold aILDprc, or: ILDprc(t) is the number of bands where ILD(t, k) > aILD or ILD(t, k) < 1 / aILDtotal. The frequency bands are numbered, and aILDprc is a threshold determined experimentally. In practice, a threshold of 30% is a good compromise.

[0101] Figure 4 currently details block 204 of Figure 2. The Hi(t,k) filters, such as, for example, those defined in the T1 and T2 reduction mixing procedures defined above, are reconditioned or, in fact, adapted in order to avoid circular convolution.

[0102] In a first step, in 401, the impulse responses hi(t, l) of the filters Hi(t, k) are calculated such that: hi(t,l') = FFT-1{Hi(t,k')},i = 1,2 where the filters hi(t, l) are of size Lh= B.

[0103] In the preferred mode, L = 20 ms, that is, K = 320 samples at 16 kHz, 640 samples at 32 kHz and 960 samples at 48 kHz, and BK = 2, that is, B = 160 samples at 16 kHz, 320 samples at 32 kHz and 480 samples at 48 kHz.

[0104] Impulse responses defined in this way are conventionally non-causal; they correspond to finite impulse response filters. A conventional technique to make them causal consists of a B / 2 sample rotation to center the response at B / 2. The disadvantage is that the filter reconstructed in this way then exhibits Petition 870250088391, dated 09 / 30 / 2025, page 31 / 63 24 / 36 a latency of B / 2.

[0105] In preferred mode, the impulse response will be truncated: h((t,l) = { hi(t,l -M), M < l <P hi(t,B - M + l) 0 < l <M

[0106] This time-shift (circular) operation can also be implemented by applying an appropriate phase term Hi(t,k) before the inverse FFT, as known to those skilled in the art. In this embodiment, the resulting filter has a finite impulse response with a sample delay R. This delay allows the truncated filter to be correctly conditioned, either in terms of gain by minimizing the influence of truncation on the unity gain of the phase-shift filter, or in phase by minimizing the IPD deviation. This delay R can be adapted depending on the implementation constraints: in practice, a delay of about one millisecond is sufficient to correctly condition the truncated filter.

[0107] In an extreme variant, the truncation is asymmetric to avoid this latency. The choice is made to truncate the filter asymmetrically, at 402, retaining only a portion of the first half of the impulse response hi(t, l) as follows: iii(t,l) = hi(t,l),0< l <P<B / 2

[0108] This filter exhibits the advantage of generally having a maximum (in absolute value) in its first sample, which guarantees processing without additional delay. In terms of frequency response, it exhibits a phase that is almost equal to the ideal rephasing filter hi(t, l). The fact that the anticausal part of the impulse response is removed can optionally be compensated by applying an additional factor of 2 to the values ​​of hi(t, l) for l > 0. In this variant, the resulting filter has a finite impulse response, practically without delay.

[0109] The choice of P depends on several criteria: a value of P Petition 870250088391, dated 09 / 30 / 2025, p. 32 / 63 25 / 36, which is close to N / 2, ensures near-ideal rephasing, while a lower value allows limiting the computational power required for filtering. A low P value also results in smoothing the filter phase. This exhibits an advantage with filters defined by the T2 processing operation, as defined above, where the phase varies very rapidly from one frequency to the next: these variations can result in a very high group delay that is perceptible to the human ear. A low P value makes it possible to remove all or some of these artifacts.

[0110] The filter, therefore, truncated fy(t, l), notably when P <N / 2, ou mesmo P«N / 2, exibe uma cauda de resposta de impulso que não tende em direção a 0. Isso cria descontinuidades de cliques audíveis na transição entre 2 quadros. Assim, a fim de evitar esses artefatos I' X. 'II I Zí\ Λ ΛΛ I I audíveis, hi(t, l) é ponderado com uma janela w(l), em 403, de modo que: h(t,l) = w(l).'hi(t,l) where the coefficients of w(l) tend to 0 when l tends to 0 or P and are equal to 1 when l=R. This makes it possible to reduce the energy of the samples at the end of the filter and avoid discontinuities during the transition between filters of consecutive frames.

[0111] In the preferred embodiment, in which truncation preserves an anticausal part, the ½ triangular windows are, for example, applied to the causal and anticausal parts w(l) = < i.o <l<R— —,R < l < P RP or any other type of window, a Hann window, for example.

[0112] In the variant where the truncation is asymmetric, this window can be a triangular half-window or, in fact, a Hann half-window: w(l) = 1 — l / P, or cos(ln / 2P) Petition 870250088391, dated 09 / 30 / 2025, p. 33 / 63 26 / 36

[0113] Because of its truncation and windowing, the filters 7^ X . aΛI ** ' I ÍM1 I 1 7 X . 7 X hi(t,l) have reduced energy compared to the ideal filter hi(t, l): Finally, as a pure phase-shift filter, it has a 1 / 2 norm by construction. Thus, the energy of the truncated and windowed filter can be normalized to 404, so that: hf(t, l) = hj(t,l) 2||7 / í(t,Z)|| where ||x(. )|| is the norm 2 of a vector.

[0114] In variations, another procedure can be used in 204 instead of the procedure in Figure 4. For example, the impulse response will possibly be determined directly by least squares minimization and matrix inversion using the (complex) Levinson algorithm as described in section 2 of the paper by Mathias C. Lang, Design of nonlinear phase FIR digital filters using quadratic problems, Proc. ICASSP, 1997. The interested reader can, for example, find an occurrence of this procedure in the thesis by Mathias C. Lang, Algorithms for the Constrained Design of Digital Filters with Arbitrary Magnitude and Phase Responses, June 1999 (see the levin.m routine in Appendix B and lslevin.m in Section 2.1.4). In this case, an inverse FFT, truncation, and weighting are not necessary; this alternative procedure estimates a finite impulse response of length P directly (without guarantee of a minimum phase).However, a normalization with the following format will potentially be applied:. hi(t,l) = hj(t,l) 2||Kí(t,i)||

[0115] According to the invention, in 405, a quality level of the implemented filtering is determined. In the preferred embodiment with truncation with a causal part, the level of | / íi(t, l)| will potentially be determined for l = R. When this value is less than a predetermined threshold (e.g., 0.25), this indicates a Petition 870250088391, dated 09 / 30 / 2025, page 34 / 63 27 / 36 indicates a low quality criterion, and conversely, when this value is greater than a predetermined threshold (for example, 0.25), this indicates a good quality criterion.

[0116] In some variants, it is possible to determine for this purpose: - an E401 energy of the filter before 402 truncation is calculated; - an E403 energy of the temporal filter after windowing 403; is calculated; - A ratio R4 between E401 and E403 is calculated.

[0117] A filtering quality can also be determined by comparing an E401 energy and an E404 energy of the temporal filter after normalization.

[0118] Through non-exhaustive alternatives, the level of filtration quality can be determined: - directly in the filter H i (t, k) and / or - a frequency version of this same filter before normalization / / ,( / , / <) or after normalization Híi(t,k). - etc.

[0119] In other variants, it is possible to calculate: Σ / 'Umm> -ímh

[0120] This criterion can be normalized by: Σι hi(t,l)2

[0121] Another quality variant may be related to measuring the amount of energy lost during truncation in order to make the filter causal. In this variant, E401 can be calculated as the 'anticausal energy' of the original filter: ^401(t)= Σφ=N / 2-i(t,l)

[0122] The E403 (or E404) energy can be calculated as the energy of the synthesized filter, before normalization (or after normalization): Petition 870250088391, dated 09 / 30 / 2025, p. 35 / 63 28 / 36 ^403(ί)=ΣΓ=-11ξ2(ί, / ) M) = ΣΓ=-A2(M)

[0123] In order to avoid the polarization associated with the window size, the energy of the anticausal part can also be calculated over the same number of samples, with or without windowing by the w(l) window: Ε40ΐ(ΐ)=Σ^ / 2-Ρ^2(ΐ,1') W) = Σι-—μ(Ν - l + ^)hl(t, l)

[0124] In some variants, the frequency response Hj(t,k) of the filter obtained after truncation (at the output of 403) can be calculated and compared with the target response / / , (t, fc) in order to define the quality level, for example: Σ|Hí(t,fc)- / 7í(t,fc)|2

[0125] This criterion can be normalized by: ΣM&νΐ2

[0126] Another variant of the quality criterion may consist of comparing the phase coherence of the filters, since the objective is to rephase the signals. Thus, the coherence CH^Ct) between the 2 filters can be calculated: CHiJh(t)=^k Hi(t,k~)Hi(t,k~) |Hí(t,k) / íí*(t,k)|

[0127] A high coherence (that is, close to 1) indicates that the synthesized filter is close to the ideal filter and therefore a high quality factor, a low value (close to 0) indicates, conversely, a low quality criterion.

[0128] Figure 5 currently describes block 205 of Figure 2, in which an OLS (overlay-save) implementation of filtering is performed in the time or frequency domain, in detail. The choice of domain will be made depending on the complexity, which depends on the size of the filter relative to the frame to be processed.

[0129] An implementation of OLS in the time domain is Petition 870250088391, dated 09 / 30 / 2025, pp. 36 / 63 29 / 36 shown here, with adjacent L-sized frames.

[0130] In a first stage, the current frame is concatenated, in 501, with the P-1 samples from the previous table (saving principle): Xt(k,.) = [xi(tL - P + 2), ...,xt(tL - 1),%í(tL), ^%í((t + 1)L - 1)]

[0131] Each channel is filtered by its phase-shift filter in 503: yi (t, ΐ) = Σ^ο1iii(t / p).Xi(t,l + Pp),0 <l<L

[0132] The mono reduction mix is ​​then created by summing the two re-phased channels in 505: m(t,l) = yi(t,l) + y2(t,l),0 <l < L

[0133] In practice, the filters can be very different between two successive frames, for example, when a new source appears. In this case, the transition between mt-i(t, l) and (t,l) can cause differences in the form of audible artifacts. To avoid these artifacts, a crossfading step can be applied. In order to perform this crossfading, the filters hi(t - 1,l) are applied to the current frame over a number of samples Q: yi,ti(t,D = Σ, 1ii(t - 1,p\Xi(t,l + pp\o <l<Q

[0134] Isso é realizado pelo bloco 502 da Figura 5.

[0135] The number of samples required for cross fading will have to be determined experimentally, depending on the size of the filters, the sampling frequency, etc.

[0136] Next, the mono reduction mix mt-i(t, l) is created in 504, based on the filters from the previous table: mt_i(t, l) = yi:ti(t, l) + y2,ti(t, l),0 <l<Q

[0137] The resulting signal from the m(t,l) reduction mix is ​​created in 506, based on a crossfading according to the following equation: rnít D = [β^·™--^ + (1-P(D).m(t,l), 0 <l<Q,m(t, l), Q < l < L Petition 870250088391, dated 09 / 30 / 2025, p. 37 / 63 30 / 36 where β(ϊ) is a monotonically decreasing function, with β(0) = 1 and β(-1) = ε, with ε being a value close to or equal to zero. Typically, a Hann half-window can be chosen: «l) = cos(£),0Sl<«

[0138] It should be noted that, in an integration with other reduction mixing procedures, the aforementioned reduction mixing procedures T1 and T2 cannot always be represented by a filtering operation. This is the case for procedure T3 in Figure 7 or any other procedure whose gain depends on the instantaneous signal level.

[0139] In the preferred mode, procedure T3 in Figure 7 is modified to add a time offset (delay) in order to synchronize the output reduction mix with the reduction mix of procedures T1 / T2. In addition, energy compensation is applied to procedure T3 in the same way as to procedures T1 / T2 to ensure level consistency and allow for a better transition between procedures.

[0140] According to the invention, in the case where a selection of the reduction mixing procedures T1 and T2 or the reduction mixing T3 is implemented, as will be described in the remainder of the description, a 5040 and 5050 scaling step of each of the signals mt-1(t,l) in(t, l), respectively, is implemented. For this purpose, the signal m'(n) obtained at the output of the reduction mixing T3 of Figure 7, which has a level N3, the level N1, or N2, of each of the signals mt-1(t, l) in(t, l) to which the reduction mixing T1, or T2, was applied, is adjusted to the level N3, so as to be equal to or similar to this level N3. Instead of being applied to each of the signals mt-1(t, l) in(t, l), such scaling can also be performed on the signal m(t, l).

[0141] Figure 6a currently details a reduction mixing method that is in accordance with the invention, in which one of the processes Petition 870250088391, dated 09 / 30 / 2025, p. 38 / 63 31 / 36 reduction mixing procedures T1 or T2 (or T'1 / T'2 in the second selection stage) chosen at the end of the selection method in Figure 3, is implemented in 610, and another reduction mixing procedure is implemented in 611, whose gain depends on the instantaneous signal level, such as, for example, procedure T3 in Figure 7.

[0142] In this case, one solution is to calculate the reduction mixes of each procedure for the current time index frame and perform a transition, in 612, from one reduction mix to the next by crossfading, which generates a monophonic signal m(n). More specifically, in the case of a transition from a reduction mix T1 (or T2) to a reduction mix T3, the reduction mix can be performed between the signal m(t, 0) obtained at the output of T1 (or T2) and the signal mT3(t, 0) obtained at the output of T3. Crossfading can also be performed between the output of T1 (or T2) and the signal d(n) as described in Figure 7. In this last case, the energy compensation of procedure T3 (adapt_gain in Figure 7) would be applied to all procedures continuously.The decision to change from a T1 or T2 reduction mix to a T3 reduction mix (which is a third selection step) can be motivated by a criterion such as ISD, defined above, or by other criteria, for example Par3 parameters already calculated by the T3 procedure, such as ITD, choosing to apply the T3 reduction mix by default when the ITD is very large (e.g., ITD > 600 microseconds in absolute value). The ILD(t) criterion can also be used when one channel is much higher in level than the other (e.g., absence of signal to the left or right): reduction mixing without rephasing will be preferable to avoid any loss of timbre. Other criteria or combinations of criteria could also be used.

[0143] In particular, the third step of selecting the T1 (or T2) or T3 reduction mix is ​​implemented depending on: Petition 870250088391, dated 09 / 30 / 2025, page 39 / 63 32 / 36 - the presence or absence of at least one transition in the stereophonic signal, or - a level of quality of stereo signal filtering used when T1 (or T2) reduction mixing is implemented, or - an indication of the presence of a binaural signal. This last indication is external information Fbinaural (for example, defined as a binary value: 1=binaural, 0=other) that explicitly indicates that the input signal is binaural.

[0144] In the case of a 613 detection of the presence or absence of at least one transition in the stereophonic signal, according to a first embodiment, the detection of a transient is applied separately to each of the input signals x1(n) and x2(n), or to the current frame xi(t,l) and x2(t,l). Since transient detection is a conventional problem in audio-to-code conversion, the module already implemented in the EVS codec and described in the 3GPP TS 26.445 standard clause 5.1.8 is taken as an example. For this purpose, each of the input signals x1(n) and x2(n), or the current frame xi(t,l) and x2(t,l) is divided into sub-blocks. The energy of each sub-block is calculated, then smoothing is optionally implemented. The energy obtained is then compared with an energy threshold th61. If the energy obtained is below th61, the presence of a transition is considered undetectable.If this obtained energy is above th61, the presence of a transition is considered to be detected. For this purpose, a parameter Par 1 (if T1 was chosen at the end of the selection in Figure 3) or Par2 (if T2 was chosen at the end of the selection in Figure 3) is defined: - to a first value, for example 0, in order to indicate that the presence of a transition was not detected, - to a second value, for example 1, in order to indicate that the presence of a transition has been detected.

[0145] This Par1 / Par2 parameter is transmitted as a criterion of Petition 870250088391, dated 09 / 30 / 2025, p. 40 / 63 33 / 36 decision, to the cross-fading block 612.

[0146] In the case where the T1 (or T2) reduction mix or the T3 reduction mix is ​​selected based on a quality level of the stereo signal filtering used when the T1 (or T2) reduction mix is ​​implemented, with reference to Figure 4, the following steps are implemented: - calculate the energy E61 of the signal h^t, / ), - Calculate the energy E62 of the signal hf(t, / ), - Calculate a ratio between E61 and E62, - Compare this ratio with an energy threshold SE.

[0147] If the ratio between E61 and E62 is below SE, the presence of a transition is considered undetected. If the ratio between E61 and E62 is above SE, the presence of a transition is considered detected. For this purpose, a parameter Par 1 (if T1 was chosen at the end of the selection in Figure 3) or Par2 (if T2 was chosen at the end of the selection in Figure 3) is defined: - to a first value, for example 0, in order to indicate that the presence of a transition was not detected, - to a second value, for example 1, in order to indicate that the presence of a transition has been detected.

[0148] This Par1 / Par2 parameter is transmitted, as a decision criterion, to the cross-fading block 612.

[0149] Thus, according to the invention, when a transient is detected in the current frame in one or both channels, the T3 reduction mix is ​​applied to the current frame instead of the T1 (or T2) reduction mix. Specifically, the T1 and T2 reduction mixes use impulse response filtering which sometimes has a tendency to spread the signal envelope; this problem is more pronounced in transient sounds, such as castanets, scissors clicking in a binaural recording simulating a haircut in a barbershop. Petition 870250088391, dated 09 / 30 / 2025, page 41 / 63 34 / 36 reiro, etc.

[0150] When the external information Fbinaural explicitly indicates that the signal with 2 input channels is binaural, i.e., Fbinaural = 1, it is possible, according to the invention, to choose a particular reduction mix, for example, the reduction mix T3.

[0151] In order to limit potential artifacts due to crossfading (block 612) between reduction mixing procedures, the invention may provide for the application, in 614, of signal scaling at the output of T1 (or T2). According to the invention, such scaling 614 implements the following steps: - Determine the energy E' of the reduction mixing signal at the output of T1 (or T2), preferably with smoothing, by performing the same steps as steps 708, 709 and 710 of Figure 7, - Determine the scaling factor: g' , / E1+E2 \ 2 Ef - Apply the g' factor to the aforementioned reduction mixing signal, returning to the operating principle (with smoothing at the beginning of the frame) of blocks 712 and 714 in Figure 7.

[0152] Unlike the reduction mix in Figure 7, a direct compensation of the reduction mix level at the output of T1 or T2 is applied here, without reinjecting an input signal.

[0153] Note that when the T1 or T2 reduction mix introduces a delay, the power compensation must take this delay into account in order to align the power of the reduction mix with the input channels. This can be achieved by shifting the input channels to synchronize them with the reduction mix. Alternatively, the power used for power compensation can also be calculated by sub-block. For example, with a delay of R = 1 ms, it is possible to divide the 20 ms frame into 20 sub-blocks and take the power of the last sub-block of the previous frame and the Petition 870250088391, dated 09 / 30 / 2025, pp. 42 / 63 35 / 36 first 19 subblocks of the current frame for input channels.

[0154] Figure 6b illustrates an alternative to the reduction mixing method described with reference to Figure 6A.

[0155] The modality illustrated in Figure 6B is distinguished from Figure 6A only because a transition is detected according to a second method. According to this second method, such detection is implemented during the T1 (or T2) reduction mixing step 610 instead of being applied separately to each of the input signals X1(n) and X2(n), or to the current frame xi(t,l) and X2(t,l). For this purpose, with reference to Figure 5, the energy of each 2 ms sub-block is compared between the input signal and the output signal of blocks 502 and 503. If the energy drops below a threshold, for example by 3 dB, a transient is considered to have been sufficiently attenuated and, in this case, the T3 reduction mixing is applied to the current frame. If the energy does not drop below a threshold, for example by 3 dB, a transient is considered not to have been attenuated and, in this case, the T1 (or T2) reduction mixing is applied to the current frame.

[0156] Once again, when the external information Fbinaural explicitly indicates that the signal with 2 input channels is binaural, that is, Fbinaural = 1, it is possible, according to the invention, to choose a particular reduction mix, for example, the reduction mix T3.

[0157] Figure 8 illustrates a reduction mixing device. 800, within the meaning of the invention.

[0158] The 800 device comprises a processing circuit that typically includes: - a MEM1 memory for storing instruction data of a computer program within the meaning of the invention; - an INT1 interface to receive a stereo audio signal f%i(n). l%2(n)' - a PROC1 processor to receive this signal and Petition 870250088391, dated 09 / 30 / 2025, pp. 43 / 63 36 / 36 process it by executing the instructions of the computer program that the MEM1 memory stores, in order to perform the reduction mixing; in particular, the processor is capable of controlling the processing modules as described with reference to Figures 2 to 5; and - a COM 1 communication interface to transmit the reduced signals, resulting from the reduction mixing, m(t, / ), to another processing module, for example, a code conversion module or a decoding module of an audio signal decoder or converter.

[0159] Naturally, this Figure 8 illustrates an example of a structural embodiment of a reduction mixing device within the meaning of the invention.

[0160] Figures 2 to 7, discussed above, describe functional modes of this device in detail.

[0161] Applications of this type of reduction mixing include, for example, audio-to-code conversion, such as when the remote control lacks the capabilities to render stereo sound on its terminal. In this case, it is not necessary to transport a stereo signal, thus saving bandwidth. This type of method can work during point-to-point conversations if one of the participants makes a stereo or binaural sound recording. This type of method can also be present during a multi-party call: the conference bridge spatializes the scene, creating a stereo scene, but not all participants necessarily have stereo rendering capabilities. The scene must then be subjected to a mono reduction mix for these participants.

[0162] The invention can also be applied to audio decoding: in this case, it is the user's renderer (rendering module) that will perform the reduction mixing of the stereo content to adapt to the user's reduced capabilities (a mere loudspeaker, for example). Petition 870250088391, dated 09 / 30 / 2025, pp. 44 / 63< / kbaixok>

Claims

1 / 3 CLAIMS 1. A method for reducing a stereophonic signal to obtain a monophonic signal, characterized in that it comprises, for a current frame: - a first stage of selecting a reduction mixing operation from two procedures that use filtering of the stereophonic signal and rephasing of the signals, the selection being made according to a phase indicator, representative of a measure of a degree of phase opposition between the frequency components of the stereo signal channels, - a second stage of selection between one of the two procedures, selected at the end of the first stage of selection, and a reduction mixing procedure without rephasing, the selection being made according to the value of an energy ratio between the stereo signal channels.

2. Method, according to claim 1, characterized in that the power ratio between the channels of the stereo signal is calculated by frequency band, and the second selection step is performed by frequency band.

3. A method according to claim 2, characterized in that the reduction-without-rephasing mixing operation is selected for relevant frequency bands only in the case where the number of relevant frequency bands exceeds a threshold.

4. A method, according to any of the preceding claims, characterized in that a third selection step is performed between one of the two procedures selected at the end of the second selection step and a third reduction mixing procedure that does not use stereophonic signal filtering, wherein said selection is implemented depending on: - the presence or absence of at least one transition in the stereophonic signal, or - a quality level of the stereophonic signal filtering used when implementing a processing operation that uses stereophonic signal filtering, or - an indication of the presence of a binaural signal.

5. Reduction mixing method, according to claim 4, characterized in that, if the presence of said at least one transition is detected (613), or not detected (613), in the stereophonic signal, the reduction mixing operation, which does not use filtering of the stereophonic signal or which uses filtering of the stereophonic signal, is selected (612), respectively.

6. A reduction mixing method according to claim 4, characterized in that said at least one transition is detected before either of the two reduction mixing operations is applied to said stereophonic signal.

7. A reduction mixing method, according to claim 1, characterized in that said at least one transition is detected as follows: - calculating the energy of the stereophonic signal, referred to as the input energy, before filtering is applied, - calculating the energy of the monophonic signal, referred to as the output energy, after filtering is applied, - calculating a ratio between the input energy and the output energy, - comparing said ratio with an energy threshold, and - if said ratio is below, or above, said energy threshold, the reduction mixing operation, which does not use filtering of the stereophonic signal or which uses filtering of the stereophonic signal, will be selected, respectively.

8. Reduction mixing method, according to Petition 870250088391, dated 09 / 30 / 2025, page 46 / 63 3 / 3 claim 4, characterized in that said filtering quality level is compared with a filtering quality threshold and, if said filtering quality level is below, or above, said filtering quality threshold, the reduction mixing operation, which does not use stereo signal filtering or which uses stereo signal filtering, will be selected, respectively.

9. A method, according to any of the preceding claims, characterized in that it comprises a preliminary phase of adapting reduction mixing filters before application to the stereo signal through the following steps: - obtaining impulse responses from the filters corresponding to the reduction mixing operation; - truncating part of the impulse responses; - weighting the remaining part by applying a weighting window; - normalizing the impulse responses resulting from the windowing to obtain the adapted filters that should be applied to a stereo signal frame.

10. Method, according to claim 9, characterized in that the truncation step consists of preserving a causal part and a non-causal part of the impulse responses.

11. A reduction mixing device, characterized in that it comprises a processing circuit for implementing the steps of the reduction mixing method, as defined in any one of claims 1 to 10.

12. A storage medium that can be read by a processor, characterized in that it stores a computer program comprising instructions for executing the method, as defined in any one of claims 1 to 10. Petition 870250088391, dated 09 / 30 / 2025, pp. 47 / 63