Optimized processing to reduce the number of channels in a stereo audio signal.

The method optimizes stereo to mono downmixing by selecting phase-corrected and non-phase-corrected procedures based on inter-channel phase and energy ratios, addressing issues of comb filtering and artifacts in existing techniques, ensuring high-quality mono conversion.

JP2026516657APending Publication Date: 2026-05-26オランジュ

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
オランジュ
Filing Date
2024-04-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing downmixing techniques for stereo to mono conversion suffer from issues such as comb filtering, excessive reverberation, and audio artifacts like echoes and statics, particularly when microphones are not in phase, leading to signal coloration and loss of intelligibility.

Method used

A method for downmixing stereo signals to mono that selects between phase-corrected and non-phase-corrected procedures based on inter-channel phase and energy ratios, using filtering and phase correction indicators to optimize the downmix process for each frequency band, and applies adaptive filtering to minimize artifacts.

Benefits of technology

This approach reduces frame-to-frame fluctuations and audible artifacts, maintaining signal quality and intelligibility by tailoring the downmix process to the specific characteristics of the stereo signal, avoiding excessive reverberation and coloration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516657000001_ABST
    Figure 2026516657000001_ABST
Patent Text Reader

Abstract

The present invention relates to a processing method for reducing the number of channels in a stereo signal in order to obtain a monaural signal, comprising: a first step (305) of selecting a channel reduction processing operation for the current frame from among two methods using filtering of the stereo signal and phase correction of the signal, wherein the selection is made according to a phase indicator (302) representing a measure of the degree of out-of-phase relationship between the frequency components of the channels of the stereo signal; and a second step (307) of selecting between one of the two methods selected at the end of the first selection step and a channel reduction processing method without phase correction, wherein the selection is made according to a value of the energy ratio (306) between the channels of the stereo signal. The present invention also relates to a processing device for carrying out the method and reducing the number of channels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the general field of audio signal processing. In particular, the present invention relates to the downmixing of multi-channel audio signals. The downmixing of stereo signals to monaural signals is particularly targeted.

[0002] This type of processing is generally applicable in the field of audio coding, regardless of whether it is a coding step or a decoding step, more specifically in the field of audio coding.

Background Art

[0003] Downmixing involves inferring a signal y consisting of a smaller number of D channels based on the combination of C channels of a multi-channel signal x. In practice, this involves defining a function f(.)

Number

[0004] In the present invention, a specific case of D = 1 and C = 2 is targeted, and reference is made to downmixing from stereo to mono or simply mono downmixing. In the case of C = 2, reference is made to a two-channel content or signal that can be stereo content or binaural content. Channel 1 (x1) is called the left channel (left), and channel 2 (x2) is called the right channel (right). In the following, the stereo case, including the binaural case, is considered as a general two-channel signal in order to avoid repeating these two terms even when the binaural signal has specific characteristics technically. In the case of the signal y obtained after downmixing, a monaural signal or mono signal is referred to and is denoted as m(n) in the time domain or M(k) in the frequency domain below.

[0005] These stereo signals can be generated from a pair of stereo microphones, binaural capture, or further from artistic mixing of audio tracks.

[0006] There are specific pairs of microphones that allow for the creation of stereo content. Among the most common are the XY pair, which consists of two cardioid microphones with an azimuth angle difference of 90° to 135°. The MS (mid-side) pair is another very common pair, consisting of a cardioid microphone and a second microphone, called a "figure-eight" microphone, pointed at 90°. By combining (sum / difference) these two microphones, left and right stereo content channels can be created. These pairs are called co-incident pairs, meaning that the microphone capsules do not exhibit delay between them, and spatialization is perceived by the difference in sound intensity between the left and right channels.

[0007] Another category of pairs, called phase stereo pairs, involves the use of microphones where the microphones are far apart from each other, resulting in a phase difference between channels for sources closer to one microphone. The most well-known is the AB pair, which utilizes two omnidirectional microphones spaced a few centimeters to several meters apart. Another very popular pair is the ORTF pair, which utilizes both phase difference through a 17cm distance between microphones and amplitude difference through a cardioid directivity (the microphones exhibit an angle of approximately 90°). Binaural pairs involve placing omnidirectional microphones over the ears of a person or artificial head. This type of device allows for the natural delivery of binaural content, which can be heard through a headset.

[0008] Stereo content also includes speech and audio signals resulting from audio mixing or audio post-production (e.g., channels stored on CDs or DVDs, or broadcasts over the Internet).

[0009] Other procedures for creating stereo content exist that are not discussed in this specification.

[0010] The simplest procedure for creating a signal reduced by downmixing is called passive downmixing. Passive downmixing is,

number

number

number

number

number

number

[0011] This passive downmixing procedure is very simple and works well when the right and left channels are in phase. However, when the microphones are not in conjunction but exhibit a phase difference, such mixing results in comb filtering, which manifests as coloration of the original signal. This is because, depending on the spacing of the microphones, the position of the source relative to the microphone pair, and the frequency, the left and right channels of the same source can be either in phase as shown in Figure 1a or out of phase as shown in Figure 1b. Therefore, when downmixing, the left and right signals are either added together (constructive interference) or canceled out (cancellative interference), resulting in variations in the level of the reduction signal m(t) and a coloration effect. The coloration is due to the effect of comb filtering. The effects of constructive or cancelative interference appear in different frequency bands, thus altering the balance and resulting in a change in the tone of the signal due to the downmix compared to the original.

[0012] To compensate for this flaw in intensity levels across frequencies, the e-AAC+ codec's downmixing implements a level correction γ(k) to ensure that the reduced signal energy levels remain equivalent to the original across each frequency band.

number

number

[0013] This technique allows for some degree of compensation for the reduction in intensity caused by downmixing. However, this compensation can lead to over-amplification when the signal is close to out of phase. In practice, the correction coefficient γ(k) is also limited to a maximum value (e.g., upper limit of 2).

[0014] Figure 7 illustrates the procedure for reducing stereo channels to mono, which is integrated into the public candidate IVAS codec available at address:https: / / forge.3gpp.org / rep / ivas-codec-pc / ivas-codec / - / tree / main.

[0015] A signal x(m) with two input channels, where m is the index of the interleaved sample, is deinterleaved (block 701) to find both the left and right channels. Next, frequency analysis using windowing and Fourier transform (blocks 702 and 703) is performed to obtain the spectra of the two channels, and the cross-correlation based on the phase spectrum is estimated (block 704). Then, the time difference between the channels is identified by searching for peaks (although formally it is ICTD between the two channels, it is written as ITD herein) (block 705), and block 705 is the correlation level R corresponding to ITD. * We also offer this.

[0016] Next, the mixing coefficient g is determined as follows (block 706).

number

[0017] Next, the energies of the left channel x1(n), the right channel x2(n), and the mono signal m(n) are determined (blocks 708, 709, and 710). In blocks 713 and 715, the left and right channels are separately reinjected (added) to the mono signal generated from block 706, according to the energy compensation determined in block 711, using the respective scale factors defined in blocks 712 and 713.

[0018] These procedures have drawbacks, in addition to signal coloration resulting from downmixing associated with imperfectly corrected comb filtering, such as excessive reverberation associated with signal aggregation when input channels are out of phase and "click" artifacts when channels are reinjected into a single frame in time. Compensating for delays that are identical at all frequencies is impractical, as this delay is related to the various sources that make up the content. In practice, this leads to a potential loss of intelligibility of sources whose defect levels cannot be compensated for, in addition to the added coloration.

[0019] Other downmixing techniques attempt to avoid the above-mentioned drawbacks. The principle of these procedures is to phase-correct the left and right signals before summing them. In a document titled (Non-Patent Literature 1),

number

[0020] The ideal phase shift is given by what is referred to herein as IPD (Inter-channel Phase Difference).

number

[0021] This is calculated in the frequency domain. This makes it possible to phase-correct the spectral components of various sources, so that a source with its own IPD may be dominant in frequency band k, while another source with a different IPD may be dominant at a different frequency k'.

[0022] In the case of the MPEG4 codec described in Samsudin's document cited above, the applied IPD is estimated by an average calculation across the bark bandwidth, not by an IPD calculated for each frequency band, the latter of which has been found to be particularly noisy and variable from frame to frame. In addition, this procedure takes the left channel as the phase reference, and if the phase adjustment of this channel is poor, the quality of the downmix will be degraded.

[0023] A published patent application (Patent Document 1) proposes a downmixing procedure that combines the procedure proposed by Samsudin with passive downmixing. In particular, it proposes an ISD (Interspectral Distance) indicator that allows for the selection of the procedure best suited for signal phase correction for each frequency band.

number

[0024] This indicator shows that the signals are in phase (ISD<1), i.e., interval

number

number

number

[0025] The passive downmix proposed in the above patent application may introduce increased reverberation in the case of low ISD.

[0026] Furthermore, in certain situations where the signals are somewhat in phase, i.e., when ISD(k) < 1.3, and especially when the channels are highly unbalanced (the source is mainly on the left or right), it has been observed that a potentially undesirable loss of timbre occurs, and therefore, in such cases, it may be important to leave the signal as is, i.e., not apply a phase shift.

[0027] Furthermore, the procedure proposed in the above patent application requires the application of processing operations in the frequency domain, and the reduced signal is reconstructed based on the short-time inverse Fourier transform (STFT) of the signal M(t,k). This type of filtering results in cyclic convolution that produces audio artifacts: echoes, pre-echoes, and statics. To mask these artifacts, implementations using frame superposition associated with appropriate windowing (analysis and synthesis windows) to ensure the reconstruction of the filtered signal may be used. Generally, this superposition is incompatible with the operation of audio coders that work with adjacent frames, i.e., without superposition. In addition, this type of reconstruction leads to a delay of generally half a frame, which is unacceptable for coders with strict latency constraints.

[0028] Another implementation of superimposed addition (OLA) is possible. This procedure requires a step of zero-padding the signal and filter to avoid cyclic convolution. However, experiments have shown that when using such filters whose phase changes very rapidly from frame to frame, the OLA method cannot completely mask artifacts in transitions of a particular frame, and the artifacts remain audible.

[0029] A published patent application (Patent Document 1) describes prior art procedures in which downmixing is achieved, for example, by switching between different downmixing procedures in the frequency domain. In this case, it is important to ensure that the switching is "seamless," that is, without any discontinuities or differences in levels between procedures, in order to avoid artifacts. [Prior art documents] [Patent Documents]

[0030] [Patent Document 1] International Publication No. 2017103418 [Non-patent literature]

[0031] [Non-Patent Document 1] “A stereo to mono downmixing scheme for MPEG-4 parametric stereo encoder” by Samsudin,E.Kurniawati,N.Boon Poh,F.Sattar,S.George,in Proc.ICASSP,2006 [Overview of the Initiative] [Problems that the invention aims to solve]

[0032] The present invention aims to improve upon the prior art. [Means for solving the problem]

[0033] For this purpose, the present invention provides a method for downmixing a stereo signal to obtain a monaural signal, with respect to the current frame, - A first step of selecting a downmix operation from two procedures using filtering of a stereo signal and phase correction of a signal, wherein the selection is made according to a phase indicator representing a measure of the degree of out-of-phase relationship between the frequency components of the channels of the stereo signal, - A second step of selecting between one of two procedures selected at the end of the first selection step and a downmix procedure without phase correction, wherein the selection is made according to the value of the inter-channel energy ratio of the stereo signal. Regarding methods including

[0034] Therefore, these selection steps make it possible to identify the type of downmix best suited to the characteristics of the stereo signal.

[0035] In one embodiment, the energy ratio between channels of a stereo signal is calculated for each frequency band, and the second selection step is performed for each frequency band.

[0036] Therefore, only the relevant frequencies where the energy ratio exceeds the threshold are downmixed without phase correction to achieve a finer fit.

[0037] In one particular embodiment, a downmix operation without phase compensation is selected only for relevant frequency bands if the number of relevant frequency bands exceeds a threshold.

[0038] This makes it possible to avoid excessive fluctuations from frame to frame.

[0039] In one embodiment, a third step is performed in which a selection is made between one of two procedures selected at the end of the second selection step and a third downmix procedure that does not use filtering of the stereo signal, and the selection is made between the two procedures selected at the end of the second selection step and a third downmix procedure that does not use filtering of the stereo signal. - The presence or absence of at least one transition in the stereo signal, or - The quality level of stereo signal filtering used when performing processing operations that utilize stereo signal filtering, or - Indication of the presence of binaural signals It will be implemented accordingly.

[0040] In this embodiment, - If one or more transitions are detected, a better rendering of one or more transitions in the signal is obtained by selecting a downmix operation that preserves the properties of the stereo signal, without using filtering. - If no transition is detected, support a downmix operation that uses filtering better suited to the non-in-phase stereo signal components. This will become possible.

[0041] In one embodiment, the detection of at least one transition is performed before one or the other of the two downmix operations is applied to the stereo signal.

[0042] This embodiment makes it possible to predict in advance whether or not at least one transition is present in the stereo signal, even before the downmix operation is applied, without applying any additional processing operations to the signal.

[0043] In one embodiment, the detection of the at least one transition is - Before filtering is applied, calculate the energy of the stereo signal, called the input energy. - After filtering is applied, calculate the energy of the monaural signal, called the output energy. - Calculating the ratio of input energy to output energy. - Comparing the ratio with the energy threshold, and - If the ratio falls below or exceeds the energy threshold, a downmix operation without filtering of the stereo signal or a downmix operation with filtering of the stereo signal is selected, respectively. It will be implemented as described.

[0044] In one embodiment, the filtering quality level is compared with a filtering quality threshold, and if the filtering quality level is below or above the filtering quality threshold, a downmix operation without filtering of the stereo signal or a downmix operation with filtering of the stereo signal is selected, respectively.

[0045] In one embodiment, the detection results generate an indicator that represents the result of comparing the ratio with an energy threshold or the result of comparing the filtering quality level with a filtering quality threshold.

[0046] In one embodiment, this method applies a downmix filter before applying it to the stereo signal. - A step to obtain the impulse response of the filter corresponding to the downmix operation, - A step to truncate a portion of the impulse response. - The step of weighting the remaining portion by applying a weighted window, - A step to normalize the impulse response resulting from windowing to obtain a fitted filter to be applied to the stereo signal frame. This includes a preliminary step of adapting via [a specific method / platform].

[0047] In one particular embodiment, the truncation step includes preserving the causal and non-causal portions of the impulse response.

[0048] The present invention relates to a downmix device that includes a processing circuit for carrying out the steps of the downmix method described above.

[0049] The present invention also relates to a computer program which, when executed by a processor, includes instructions for carrying out a downmix method according to the present invention in accordance with any one of the particular embodiments described above.

[0050] Such instructions can be permanently stored in a non-temporary storage medium of a downmix device that implements the downmix method according to the present invention.

[0051] This program can use any programming language and may be in the form of source code, object code, or intermediate code between source code and object code, such as a partially compiled form or any other desired form.

[0052] The present invention also covers computer-readable storage media or information media containing the instructions of the computer program described above.

[0053] A storage medium can be any entity or device capable of storing a program. For example, the medium may include ROM, such as a CD-ROM or micro-electronic circuit ROM, or in practice, magnetic storage means, such as a mobile medium, hard disk, or SSD.

[0054] Furthermore, the storage medium may be a transmittable medium such as an electrical signal or optical signal that can be routed wirelessly or by other means via an electrical cable or optical cable, so that the computer program contained in the storage medium can be executed remotely. In particular, the program according to the present invention may be downloaded from a network, such as the Internet network.

[0055] Alternatively, the storage medium could be an integrated circuit into which the program is embedded, and the circuit is adapted to perform or be used in performing the downmixing method described above.

[0056] In one embodiment, the technique is implemented by software components and / or hardware components. With this in mind, the terms “device” or “module” may, as used herein, correspond equally to a software component, a hardware component, or a set of hardware and software components.

[0057] Other features and advantages of the present invention are merely illustrative and will become clearer by reading the specific embodiments and accompanying drawings given as non-limiting examples. [Brief explanation of the drawing]

[0058] [Figure 1a] The channels of the in-phase stereo signals mentioned above are shown. [Figure 1b] The channels of the inverted-phase stereo signal described above are shown. [Figure 2] The downmix sequence in one embodiment of the present invention is shown in the form of a block diagram. [Figure 3] This document describes one embodiment of the selection of downmix filter calculation and downmix operation. [Figure 4] This shows one embodiment of the process of adapting a downmix filter. [Figure 5]This illustrates one embodiment of the transition of downmixing during the transition between two frames of a stereo signal. [Figure 6a] Another embodiment of the selection of downmix filter calculation and downmix operation is shown. [Figure 6b] Another embodiment of the selection of downmix filter calculation and downmix operation is shown. [Figure 7] The existing downmixing procedure described above is shown below. [Figure 8] An example of a structural embodiment of a downmix device according to one embodiment of the present invention is shown. [Modes for carrying out the invention]

[0059] The present invention will now be described below with reference to Figures 2 to 8. For this purpose, the following symbols will be mainly used in the following description. index k frequency index l The time index of the signal (in the current frame) n Signal time index t Frame number constant L Number of samples in the current frame R filter delay N FFT length Number of C input channels D Output Channel Count P filter length Q Crossfade length signal x i (n) Input channel m(n) Downmixing (in the time domain) m(t,l) Downmixing in the current frame (in the time domain) M(t,k) Downmixing in the current frame (in the frequency domain)

[0060] Figure 2 shows an example sequence for processing a stereo audio signal, where the processing operation includes a downmix operation.

[0061] At the input to this processing sequence, a stereo signal x, consisting of two channels (x1(n) and x2(n)), also called the left channel and the right channel (respectively), is divided in the first step at 201 into frames of L samples (x1(t,l) and x2(t,l)), where t is the frame index and l is the sample index. Block 202 applies windowing and FFT (Fast Fourier Transform) to obtain the signal in the frequency domain (x1(t,k) and x2(t,k), where k is the frequency index). At 203, a downmix filter to be applied to the signal frames is then selected and determined. This step is illustrated with reference to Figure 3. The thus determined filter is adjusted (or fitted) at 204 to be causal or partially causal and optimized to reduce the complexity of the processing. In one embodiment, the windowing and FFT are similar to steps 702 and 703 in Figure 7.

[0062] This adaptation stage will be explained with reference to Figure 4.

[0063] Once one or more filters are determined and fitted to the current frame, in 205 they are applied to this current frame. To avoid audible artifacts caused by filter changes between two frames, if the downmix of the previous frame is of the same type (T1 / T2 as defined below), the filter of the previous frame is applied.

number

[0064] This step will be explained with reference to Figure 5.

[0065] Figure 3 illustrates a detailed embodiment of block 203 in Figure 2. In this processing block, a first selection step selects and determines a first or second downmix operation to be applied to the current frame of the stereo signal.

[0066] The first downmixing procedure T1 is defined herein as simultaneously shifting the left and right channels of a stereo signal by an angle defined by 1 / 4 of the phase difference (IPD) identified between the two channels of the stereo signal. Thus, the following downmixed signals exist:

number

[0067] This procedure avoids the excessive reverberation caused by conventional downmixing (averaging two signals) or the comb filtering-type degradation that occurs when time-shifting between two channels.

[0068] This downmixing, as defined, is expressed in the frequency domain by the following equation: M(t,k)=H1(t,k).X1(t,k)+H2(t,k).X2(t,k) This can be considered a form of filtering,

number

[0069] In another embodiment, the first downmixing procedure T1 is defined by simultaneously shifting the left and right channels of the stereo signal by half the phase difference (IPD) identified between the two channels of the stereo signal.

[0070] In that case,

number

Number

[0071] The second down - mixing procedure T2 is defined as, for example, that proposed in the Samsudin reference cited above.

[0072] In this second procedure,

Number

[0073] As will be described below, in one embodiment, the corresponding impulse response h i (t, l) is actually time - shifted to retain the causal and non - causal parts. This shift can, in a variant form, be directly incorporated into the definition of H i (t, k).

[0074] Furthermore, in a variant form, the coefficient 1 / 2 is removed from the definition of H i (t, k) and this coefficient 1 / 2 can equally well be incorporated during the mixing of the channels processed at 504 and 505 (see Figure 5 described below).

[0075] In fact, there is no need to determine the phase difference IPD or IPD / 4 or IPD / 2. In an efficient implementation form, the filter H i (t, k) can be determined by normalizing by the absolute value (in the sense of complex numbers) the (complex) frequency line - with or without smoothing.

[0076] The channels of the stereo signal in the frequency domain (x1(t,k) and x2(t,k)) are used, on the one hand, in 301 to calculate an ISD phase indicator that represents a measure of the degree of phase inversion between the channels of the stereo signal, and on the other hand, in 303 to calculate the (IPD) phase difference between channels.

[0077] The phase indicator is defined, for example, by the ISD (interspectral distance) indicator defined above.

[0078] In the embodiments described herein, the ISD is determined for each stereo signal frame so that a decision to select either a first downmix or a second downmix can be made on a frame-by-frame basis.

[0079] This per-frame determination eliminates significant phase shifts between frequencies, as well as other audible artifacts that may arise when the determination changes from frequency to frequency, as in the patent application cited above.

[0080] To apply the decision frame by frame, the decision criterion using the ISD phase indicator is defined in 302 according to the following formula:

number

[0081] This criterion measures the proportion of frequency lines at index k that prioritizes a phase-corrected downmix, i.e., ISD(t,k) > 1.3.

number

number

[0082] In another embodiment, the ISD indicator is potentially determined within a predefined frequency band.

number

[0083] In this case, the criterion is the average of the ISD values ​​for each frequency line. In this variant, the threshold corresponds to the ISD value, not a percentage, and for example, in some cases,

number

[0084] Next, in 305, the following downmix selection between T1 and T2 is determined.

number

[0085] In the 303, the phase difference IPD is,

number

[0086] However, the IPD calculated according to the above formula changes rapidly over time, which may result in insufficient frame determination. Filter H i (t,k) can also change very rapidly from frame to frame, which can result in audible discontinuities during transitions between frames.

[0087] To limit this effect, according to the first method, the IPD is potentially calculated in frequency subbands rather than independently in each band, while the average of each channel is calculated over a wider subband. Potentially, BK B subbands with widths equal to =K / B are selected, where B is an integer greater than or equal to 1. To calculate the IPD in each subband b, first, the spectrum in each subband is calculated.

number

number

[0088] The average spectra of the subbands corresponding to frequency zero (b=0) and frequency Fe / 2 (b=B / 2) are processed separately by direct inference from the spectra in these frequency bands.

number

number

[0089] The number of subbands B should be determined experimentally. Sufficient frequency resolution must be maintained so that appropriate phase correction can be applied to various sources. In practice, a frequency resolution of approximately 1 / 100 Hz seems to be a good subbandwidth. For example, for a stereo signal sampled at 32 kHz with an FFT size of 640, this resolution is B in each subband. K = It has two frequency bands, and B corresponds to 320 subbands.

[0090] In the modified form, the subband can be divided, for example, to have non-uniform widths according to the bark scale.

[0091] Next, the IPD is calculated in the same manner as above, i.e.,

number

[0092] In the following, index k represents either a frequency band or a frequency subband, B K The case where =1 is equivalent to performing processing separately for each frequency band.

[0093] Here too, to limit the effects of fast fluctuations in IPD, a temporally smoothed version of IPD is potentially filtered H i Used in the calculation of (t,k). One way to smooth the IPD is in 304,

number

[0094] The forgetting coefficient α(k) is selected ad hoc. Typically, it may be advantageous to select a high forgetting coefficient at low frequencies below 5 kHz because this allows for maintaining phase coherence at the bottom of the spectrum frame by frame, the spectrum of the speech signal does not change particularly at its bottom, and certain ears are sensitive to phase in this part of the spectrum. In contrast, it may be advantageous to select a low forgetting coefficient at high frequencies. Typically, at high frequencies above 5 kHz, the phase continuity of the speech signal is not very pronounced, and the ear is not very sensitive to phase discontinuities in this part of the spectrum.

[0095] The following forms of forgetting coefficients may be selected.

number

[0096] Value α max , α min , k high and k low This is determined experimentally. For adjacent frames of length L=20ms, for example, the following set of parameters may be selected. α max =0.94, α min =0.86, k low =0, k highThis corresponds to 5kHz.

[0097] In some variants, very low and very high frequencies are forced to remain unchanged. Therefore, in some variants, the first line k=0 or k=0,1 and index k>k max On this line,

number

[0098] Phase-corrected downmixing (T1 or T2) can substantially alter the timbre and degrade signal quality. These timbre changes are particularly audible when the signal is dominant in one channel, i.e., when the ILD (Inter-Channel Level Difference), which measures the energy ratio between the left and right channels, is close to 0 or much higher than 1. This is the case, for example, with a dominant signal in the right channel when IPD-based phase correction mode T1 is applied. To limit these timbre changes, it is advantageous to avoid phase correction by applying conventional downmixing when the energies of the left and right channels are significantly different.

[0099] In the first implementation, independent decisions are potentially made in each frequency band, allowing for the independent processing of various sources (one of which may be dominant in one channel and not require phase correction, while other sources are dispersed between the left and right channels and require phase correction). In this case, passive downmixing is enforced only in frequency bands exhibiting high energy imbalance.

number

number

[0100] Therefore, the second selection step is performed in 307, and the ILD is calculated in 306 to determine the downmixing T1, T'1, T2, or T'2 to be applied with or without phase compensation according to the frequency band.

[0101] The decision to apply passive downmixing independently to each subband can sometimes result in frame-by-frame phase shifts. This is particularly audible when a small number of frequency bands are involved. To avoid these shifts, the application of passive downmixing is limited to the relevant bands only when a sufficient number of relevant bands are observed to indicate the actual presence of a source to the left or right. In particular, the relevant band ILD prc The ratio of (t) is calculated and applied to the relevant band defined above for passive downmixing, where this ratio is a predefined threshold α. ILDprc It is possible to impose restrictions when it exceeds that, that is,

number

[0102] Here, Figure 4 details block 204 of Figure 2. For example, filter H defined in the downmixing procedures T1 and T2 defined above. i (t,k) is readjusted or actually adapted to avoid cyclic convolution.

[0103] In the first step, 401, filter H i Impulse response h (t,k) i (t,l) is h i (t,l)=FFT -1 {H i (t,k)},i=1,2 It is calculated so that, where the filter h i The size of (t,l) is L h = B.

[0104] In a preferred embodiment, L=20ms, i.e., K=320 samples at 16kHz, 640 samples at 32kHz, and 960 samples at 48kHz, B K = 2, which means that at 16kHz, B = 160 samples, at 32kHz, 320 samples, and at 48kHz, 480 samples.

[0105] The impulse response thus defined is conventionally non-causal and corresponds to a finite impulse response filter. One conventional technique to make the impulse response causal involves rotating the B / 2 samples to center the response at B / 2. The drawback is that the thus reconstructed filter exhibits the latency of B / 2.

[0106] In a preferred embodiment, the impulse response is truncated.

number

[0107] This (cyclic) time shift operation involves adding an appropriate phase term H before the inverse FFT, as is known to those skilled in the art. i This can also be implemented by applying it to (t,k). In this embodiment, the resulting filter has a finite impulse response with a delay of R samples. This delay allows the truncated filter to be properly tuned either with respect to gain by minimizing the effect of truncation on the unit gain of the phase-shift filter, or with respect to phase by minimizing the deviation from IPD. This delay R can be adapted according to implementation constraints. In practice, a delay of about 1 millisecond is sufficient to properly tune the truncated filter.

[0108] In an extreme deformation mode, the truncation is asymmetric to avoid this waiting time. In 402, [Number] As in, only a portion of the first half of the impulse response h i (t, l) is selected to truncate the filter asymmetrically so as to retain only.

[0109] This filter generally shows the advantage of having a maximum (in absolute value) at its first sample, guaranteeing processing without additional delay. With respect to the frequency response, it shows a phase approximately equal to that of the optimal phase correction filter h i (t, l). The removal of the non-causal part of the impulse response can optionally, when l > 0, in addition, apply the coefficient 2 to [Number] the value of to be compensated. In this deformation mode, the resulting filter has a finite impulse response with little delay.

[0110] The choice of P depends on various criteria. A value of P close to N / 2 guarantees a phase correction closest to optimal, while a low value allows the computational power required for filtering to be limited. A low value of P also leads to a smoothing of the filter's phase. This shows an advantage with the filter defined by the processing operation T2 defined above, where the phase changes rapidly for each frequency. These variations can lead to a very high group delay audible to the human ear. A low value of P allows all or some of these artifacts to be removed.

[0111] In particular, when P < N / 2 or even P << N / 2, the thus-truncated filter [Number] This exhibits an impulse response tail that does not tend towards 0. This results in an audible click discontinuity in the transition between two frames. Therefore, to avoid these audible artifacts,

number

number

[0112] In a preferred embodiment where the truncate retains the non-causal portion, a triangular 1 / 2 window is applied, for example, to the causal portion and the non-causal portion.

number

[0113] In the deformed form where the truncate is asymmetrical, this window is a triangular half-window or, in fact, a Hanning half-window. w(l) = 1 - l / P or cos(lπ / 2P)

[0114] For that truncation and windowing process, filter

number

number

[0115] In the variant form, in 204, an alternative procedure may be used instead of the procedure in Figure 4. For example, the impulse response may be directly determined by least-squares minimization and matrix inversion using the (complex) Levinson algorithm, as described in Section 2 of the article, *Design of nonlinear phase FIR digital filters using quadratic problems*, Proc. ICASSP, 1997, by Mathias C. Lang. Interested readers may find an example of this procedure in the paper, for example, *Algorithms for the Constrained Design of Digital Filters with Arbitrary Magnitude and Phase Responses*, June 1999, by Mathias C. Lang (see routine levin.m in Appendix B and lslevin.m in Section 2.1.4). In this case, inverse FFT, truncation, and weighting are not required. This alternative procedure estimates a finite impulse response of length P (without guarantee of minimum phase). However, normalization in the following form may be potentially applicable.

number

[0116] According to the present invention, in 405, the quality level of the filtering that has been implemented is identified. In a preferred embodiment using causal part truncation,

number

[0117] In some variant forms, it is possible to identify the following for this reason: - Energy E before Truncate 402 401 The calculation is performed, - Energy E of the time filter after windowing 403 403 The calculation is performed, - E 401 and E 403 The ratio R4 is calculated.

[0118] Energy E of the normalized time filter 401 and energy E 404 By comparing these, the filtering quality can also be determined.

[0119] Due to non-extensive alternative forms, the quality level of filtering is, - Filter H i (t,k) directly and / or - Before normalization

number

number

[0120] In other transformation forms,

number

[0121] This standard is

number

[0122] Another variation of quality may involve measuring the amount of energy lost during truncation in order to make the filter causal. In this variation, E 401 This can be calculated as the "non-causal" energy of the original filter.

number

[0123] Energy E 403 (or E 404 ) can be calculated as the energy of the composite filter before (or after) normalization.

number

number

[0124] To avoid bias associated with the window size, the non-causal portion of the energy can also be calculated over the same number of samples with or without windowing by window w(l).

number

number

[0125] In some variations, the frequency response of the filter obtained after truncation (at the output of the 403)

number

number

[0126] This criterion

Number

[0127] Another variation of the quality criterion may involve comparing the phase coherence of the filters, given that the objective is signal phase correction. Thus, the coherence

Number

Number

[0128] High (i.e., close to 1) coherence indicates that the composite filter is close to an ideal filter, and thus shows a high quality factor, while low (close to 0) values, in contrast, indicate a low quality criterion.

[0129] Here, Figure 5 details block 205 of Figure 2 where the OLS (Overlap-Save) implementation of filtering is performed in the time or frequency domain. The choice of domain depends on the complexity that depends on the size of the filter relative to the size of the frame being processed.

[0130] The OLS implementation in the time domain is presented here, and the size of the adjacent frames is L.

[0131] In the first step, at 501, the current frame is concatenated with P - 1 samples of the previous frame (the save principle).

Number

[0132] At 503, each channel is filtered by its phase shift filter.

Number

[0133] Next, a mono downmix is created by adding two channels phase-corrected at 505. m(t, l) = y1(t, l) + y2(t, l), 0 ≤ l < L

[0134] In practice, the filter can be very different between two consecutive frames, for example, when a new source appears. In this case, the transition between m t-1 (t, l) and m(t, l) can cause a difference in the form of audible artifacts. To avoid these artifacts, a crossfade step can be applied. To perform this crossfade, the filter

Number

Number

[0135] This is executed by block 502 in FIG. 5.

[0136] The number of samples required for the crossfade needs to be determined experimentally according to the filter size, sampling frequency, etc.

[0137] Next, at 504, a mono downmix m t-1 (t, l) is created based on the filter of the previous frame. m t-1 (t, l) = y 1,t-1 (t, l) + y 2,t-1 (t, l), 0 ≤ l < Q

[0138] At 506, the downmix

Number

number

number

[0139] In integration with other downmixing procedures, it should be noted that the downmixing procedures T1 and T2 described above cannot always be represented by a filtering operation. This is the case with procedure T3 in Figure 7 or any other procedure where the gain depends on the instantaneous level of the signal.

[0140] In a preferred embodiment, step T3 in Figure 7 is modified to add a time shift (delay) to synchronize the output downmixing with the downmixing of steps T1 / T2. In addition, energy compensation is applied to step T3 in the same way as in steps T1 / T2 to ensure level consistency and enable better transitions between steps.

[0141] According to the present invention, when a selection of downmixing procedures T1 and T2 or downmix T3 is performed, as described in the remainder of the description, the signal m t-1 Steps 5040 and 5050 are performed to scale (t,l) and m(t,l), respectively. For this purpose, the signal m'(n) obtained at the output of downmix T3 in Figure 7, having level N3, and the signal m to which downmix T1 or T2 is applied are performed. t-1 The respective levels N1 or N2 of (t,l) and m(t,l) are adjusted to level N3 so that they are equal to or similar to this level N3. t-1 Instead of applying such scaling to (t,l) and m(t,l) respectively, such scaling is applied to the signal

Number

[0142] Here, FIG. 6a details a downmixing method according to the present invention in which one of the downmixing procedures T1 or T2 (or T'1 / T'2 in the second selection step) selected at the end of the selection method of FIG. 3 is performed at 610, the other downmixing procedure is performed at 611, and the gain thereof depends on the instantaneous level of a signal such as procedure T3 in FIG. 7, for example.

[0143] In this case, one solution is to calculate the downmixing of each procedure in the current frame of the time index l, perform the transition between the downmixings by cross-fading at 612, thereby generating the monaural signal m(n). More specifically, in the case of the transition from downmixing T1 (or T2) to downmixing T3, the downmixing is the signal obtained at the output of T1 (or T2)

Number

[0144] In particular, the third step of selecting downmix T1 (or T2) or T3 is, - The presence or absence of at least one transition in the stereo signal, or - The quality level of filtering of the stereo signal used when downmix T1 (or T2) is performed, or - Indication of the presence of binaural signals This is carried out accordingly. The latter instruction is external information F that explicitly indicates that the input signal is binaural. binaural (For example, binary values ​​are defined as 1 = binaural, 0 = other).

[0145] When detecting the presence or absence of at least one transition in a stereo signal (613), according to the first embodiment, the transition detection is applied separately to each of the input signals x1(n) and x2(n), or to the current frames x1(t,l) and x2(t,l). Since transition detection is a conventional problem in audio coding, it is already implemented in the EVS codec, and an example of the module described in Section 5.1.8 of the 3GPP® standard TS26.445 is taken. For this purpose, each of the input signals x1(n) and x2(n) or the current frames x1(t,l) and x2(t,l) is divided into subblocks. The energy of each subblock is calculated, and then smoothing is optionally performed. The obtained energy is then subjected to an energy threshold th 61 This is compared to the energy obtained. 61 If it is less than th, the presence of the transition is considered not to be detected. The obtained energy is th 61 If it exceeds this value, the existence of a transition is considered detected. For this reason, the parameter Par1 (if T1 is selected last in the selection in Figure 3) or Par2 (if T2 is selected last in the selection in Figure 3) is set as follows: - To indicate that no transition was detected, the first value is set to, for example, 0. - A second value, for example 1, is set to indicate that the existence of a transition has been detected.

[0146] Such parameters Par1 / Par2 are sent to the crossfade block 612 as decision criteria.

[0147] If downmix T1 (or T2) or downmix T3 is selected based on the quality level of filtering of the stereo signal used when downmix T1 (or T2) is performed, refer to Figure 4 and perform the following steps. - signal h i Energy E of (t,l) 61 Steps to calculate - Signal

number

[0148] E 61 and E 62 If the ratio to SE is less than SE, the existence of the transition is considered not detected. 61 and E 62 If the ratio exceeds SE, the presence of a transition is considered detected. For this reason, the parameter Par1 (if T1 is selected last in the selection in Figure 3) or Par2 (if T2 is selected last in the selection in Figure 3) is, - To indicate that no transition was detected, the first value is set to, for example, 0. - A second value, for example 1, is set to indicate that the existence of a transition has been detected.

[0149] Such parameters Par1 / Par2 are sent to the crossfade block 612 as decision criteria.

[0150] Therefore, according to the present invention, if a transition is detected in the current frame of one or both channels, downmix T3 is applied to the current frame rather than downmix T1 (or T2). Specifically, downmixes T1 and T2 use filtering by an impulse response, which can tend to spread the signal envelope. This problem is more pronounced with transition sounds such as castanets or the clicking of scissors in a binaural recording simulating a haircut in a salon.

[0151] External information F indicates that a signal with two input channels is binaural. binaural This explicitly shows, namely F binaural If = 1, the present invention makes it possible to select a specific downmix, for example, downmix T3.

[0152] To limit potential artifacts resulting from crossfades between downmixing procedures (block 612), the present invention may provide signal scaling at the output of T1 (or T2) in 614. According to the present invention, such scaling 614 performs the following steps: - A step of determining the energy E' of the downmix signal at the output of T1 (or T2), preferably with smoothing, by performing the same steps as in steps 708, 709, and 710 in Figure 7. - Scale factor:

number

[0153] Unlike the downmix in Figure 7, direct compensation of the downmix level at the output of T1 or T2 is applied here without reinjection of the input signal.

[0154] Note that if the downmix T1 or T2 introduces a delay, the energy compensation must take this delay into account so that the downmix energy aligns with the input channel energy. This can be achieved by shifting the input channel to synchronize with the downmix. In a variant, the energy used for energy compensation can also be calculated equally for each subblock. For example, with a delay R=1ms, a 20ms frame can be divided into 20 subblocks, and the energy of the last subblock of the previous frame and the first 19 subblocks of the current frame can be taken.

[0155] Figure 6b shows an alternative method to the downmix method described with reference to Figure 6A.

[0156] The embodiment shown in Figure 6B is distinguished from the embodiment in Figure 6A solely by the fact that the transition is detected by the second embodiment. According to this second embodiment, such detection is performed during step 610 of downmix T1 (or T2), instead of being applied separately to each of the input signals x1(n) and x2(n) or to the current frames x1(t,l) and x2(t,l). For this purpose, referring to Figure 5, the energy of each 2ms subblock is compared between the input and output signals of blocks 502 and 503. If the energy falls below a threshold, e.g., 3dB, the transition is considered to have been significantly attenuated, in which case downmix T3 is applied to the current frame. If the energy does not fall below a threshold, e.g., 3dB, the transition is considered not to have been attenuated, in which case downmix T1 (or T2) is applied to the current frame.

[0157] Here too, external information F indicates that a signal with two input channels is binaural. binaural This explicitly shows, namely F binaural If = 1, the present invention makes it possible to select a specific downmix, for example, downmix T3.

[0158] Figure 8 shows a downmix device 800 within the scope of the present invention.

[0159] Device 800 typically includes processing circuitry that includes the following: - Memory MEM1 for storing instruction data of a computer program within the scope of the present invention, - Stereo audio signal

number

number

[0160] Naturally, Figure 8 shows an example of a structural embodiment of a downmix device within the scope of the present invention.

[0161] Figures 2 to 7 above illustrate in detail the functional embodiments of this device.

[0162] This type of downmixing is applied, for example, in audio coding, when, for instance, a remote device does not have the capability to render stereo sound at its terminal. In this case, there is no need to transport a stereo signal, thereby saving bandwidth. This type of method can work during a two-point conversation if one of the participants is performing stereo or binaural sound recording. This type of method can also be present during a multi-party simultaneous call. A conference bridge spatializes the scene by bringing in a stereo scene, but not all participants necessarily have the capability to render stereo. In that case, the scene should be downmixed to mono for these participants.

[0163] The present invention can also be applied to audio decoding. In this case, it is the user's renderer (rendering module) that downmixes the stereo content to suit the user's limited capabilities (e.g., a simple loudspeaker).

Claims

1. A method for downmixing a stereo signal to obtain a mono signal, for the current frame, - A first step of selecting a downmix operation from two procedures using filtering the stereo signal and phase correction of the signal, wherein the selection is made according to a phase indicator representing a measure of the degree of out-of-phase relationship between the frequency components of the channels of the stereo signal. - A second step of selecting between one of the two procedures selected at the end of the first selection step and a downmix procedure without phase correction, wherein the selection is made according to the value of the inter-channel energy ratio of the stereo signal. A method that includes this.

2. The method according to claim 1, wherein the energy ratio between the channels of the stereo signal is calculated for each frequency band, and the second selection step is performed for each frequency band.

3. The method according to claim 2, wherein the downmix operation without phase compensation is selected only for relevant frequency bands if the number of relevant frequency bands exceeds a threshold.

4. A third step is performed in which a selection is made between one of the two procedures selected at the end of the second selection step and a third downmix procedure that does not use filtering of the stereo signal, and the selection is, - The presence or absence of at least one transition in the stereo signal, or - The quality level of the filtering of the stereo signal used when performing a processing operation that uses the filtering of the stereo signal, or - Indication of the presence of binaural signals The method according to any one of claims 1 to 3, which is carried out accordingly.

5. The downmixing method according to claim 4, wherein if the presence of at least one transition is detected in the stereo signal (613) or not detected (613), the downmixing operation without filtering of the stereo signal or the downmixing operation with filtering of the stereo signal is selected, respectively (612).

6. The downmixing method according to claim 4, wherein the at least one transition is detected before any of the two downmixing operations is applied to the stereo signal.

7. At least one transition is, - Before the filtering is applied, calculate the energy of the stereo signal, called the input energy. - After the filtering is applied, calculate the energy of the monaural signal, called the output energy. - Calculate the ratio of the input energy to the output energy. - Comparing the aforementioned ratio with the energy threshold, and - If the ratio falls below or exceeds the energy threshold, either a downmix operation without filtering the stereo signal or a downmix operation with filtering the stereo signal is selected, respectively. The downmixing method according to claim 1, which is detected as described above.

8. The downmixing method according to claim 4, wherein the filtering quality level is compared with a filtering quality threshold, and if the filtering quality level is below or above the filtering quality threshold, the downmixing operation without filtering the stereo signal or the downmixing operation with filtering the stereo signal is selected, respectively.

9. The downmix filter is applied to the stereo signal before it is applied to the stereo signal. - A step of obtaining the impulse response of the filter corresponding to the downmix operation, - A step of truncating a portion of the impulse response, - The step of weighting the remaining portion by applying a weighted window, - Steps to normalize the impulse response resulting from the windowing process to obtain a fitted filter to be applied to the stereo signal frame. The method according to any one of claims 1 to 8, comprising a preliminary step of adapting via

10. The method according to claim 9, wherein the truncating step includes preserving the causal and non-causal portions of the impulse response.

11. A downmix device including a processing circuit for carrying out the steps of the downmix method according to any one of claims 1 to 10.

12. A processor-readable storage medium for storing a computer program which includes instructions for performing the method according to any one of claims 1 to 10.