System and method for generating a multi-channel audio signal from a stereo audio signal, and corresponding system and method for spatialization
By dividing the sound panorama into angular sectors and attenuating frequency components based on amplitude ratios, the method generates multi-channel audio signals that accurately reproduce the original sound panorama, addressing compatibility and spatialization issues in existing technologies.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-03-27
AI Technical Summary
Existing methods for generating multi-channel audio signals from stereo audio signals fail to accurately reproduce the original sound panorama due to assumptions about non-overlapping sound sources in the time-frequency domain, leading to inadequate spatialization and compatibility issues with stereo signals.
A method that divides the sound panorama into angular sectors, calculates amplitude and end-gain ratios, and attenuates frequency components outside defined ranges to generate multi-channel signals that accurately represent the original sound panorama, compatible with stereo signals.
Generates high-quality multi-channel audio signals that reproduce the original sound panorama with good spatial continuity and compatibility with stereo audio sources, suitable for wide-area sound reproduction.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: System and method for generating a multi-channel audio signal from a stereo audio signal, and corresponding spatialization system and method. FIELD OF THE INVENTION
[0001] The present invention relates generally to the generation of a multi-channel audio signal from a stereo audio signal comprising a left and a right channel signal. The invention also relates to the distribution of a generated multi-channel audio signal to a plurality of loudspeakers. EARLIER ART
[0002] US patent 7567845B proposes a method for generating a multi-channel audio signal from a stereo audio signal. A similarity measure is used to identify the panning coefficients corresponding to the different individual instruments in the stereo mix. The proposed method specifically involves identifying energy peaks in the stereo signal (Figure 10 of US patent 7567845B) and their positions corresponding to panning indices, in order to identify the main sound sources corresponding to these peaks within the stereo signal. Signals corresponding to the identified peaks in the signal are then extracted.
[0003] Such a method implies the assumption that the signals from the sound sources (instruments or voices) do not overlap in the time-frequency domain. However, this assumption is generally not valid.
[0004] The multi-channel signal formed by the extractions of the main sound sources therefore does not allow for a sound reproduction sufficiently close to the original sound panorama, which has been encoded in the stereo sound signal, and with good continuity of panorama reproduction.
[0005] It is also known from the prior art, for example from documents EP2530956A8 and US5594800, to generate a multi-channel audio signal, for example by matrix methods, so as to obtain a center signal, a left signal and a right signal, from a stereo audio signal.
[0006] However, the methods used in these documents merely add or subtract the left and right signals, which does not allow for the generation of a multi-channel signal suitable for reproducing a sound scene corresponding to the soundscape encoded in the stereo signal. In particular, the user does not perceive, when playing back such a multi-channel signal obtained using these methods, the precise location of the different sound sources as they appear in the original soundscape.
[0007] Sound spatialization aims to recreate a sound scene using a speaker system.
[0008] Known sound spatialization methods involve manipulating sound objects and require possessing all the audio tracks of a sound scene, such as a piece of music, independently before mixing. These sound spatialization methods do not allow the use of an audio signal already mixed on two or more channels without a dedicated demixing solution.
[0009] Known multichannel spatialization methods, such as Dolby, 5.1, and DTS, require a dedicated encoder from the mixing stage onward, resulting in incompatibility with stereo signals, meaning that little content is compatible with these methods. Furthermore, these methods require a standardized speaker configuration, i.e., one whose positions are defined by the standard used, and do not allow for uniform reproduction over a wide area.
[0010] Furthermore, upmix solutions that allow the creation of several channels from a reduced number of channels in order to generate spatializable signals, creating envelopment, exhibit low spatial extraction quality, as well as the introduction of artifacts in the generation of signals, which does not allow the recreation of a satisfactory soundscape.
[0011] The present invention aims to overcome at least partially all or part of the problems set out above. Summary of the invention
[0012] To this end, the invention relates to a method for generating a multichannel audio signal, denoted S(t), said audio signal S(t) being generated from a supplied stereo audio signal, the supplied stereo audio signal comprising: - a left-hand signal, denoted l(t), which was obtained by mixing monophonic signals according to a left-hand panning law, denoted panL, and - a right-hand signal, denoted r(t), which was obtained by mixing monophonic signals according to a right-hand panning law, denoted panR, in which the process includes the following steps: 1) dividing a soundscape (PS) into a given number N of angular sectors, denoted Pi, preferably contiguous, 2) for a given angular sector Pi = [pu ; pri] of the soundscape (PS), and for a given frequency band Bj, - calculation of the value of the amplitude ratio, called left-right amplitude ratio, denoted lllBj(t)ll / llrBj(t)ll, between the left signal in said frequency band Bj, denoted lBj(t), and the right signal in said frequency band Bj, denoted rBj(t), - for each end pu; pri of said angular sector Pi, calculation of the ratio, called the end gain ratio, panL(pii) / panR(pu) and panL(pri) / panR(pri), between the gain provided by the left panning law panL for said end of said angular sector Pi, and the gain provided by the right panning law panR for said end of said angular sector Pi; - comparison of the left-right amplitude ratio lllBj(t)ll / llrBj(t)Il, with the calculated end-gain ratios panL(pii) / panR(pii) and panL(pri) / panR(pri); - generation, in said frequency band Bj, of an audio signal, denoted SPijBj(t), from said left signal lBj(t) in said given frequency band and said right signal rBj(t) in said given frequency band, and according to the result of the comparison of the left right amplitude ratio lllBj(t)ll / llrBj(t)Il, with the end gain ratios panL(fold) / panR(fold) and panL(pri) / panR(pri); the generation step comprising the attenuation of the left signal lBj(t) in said frequency band and of the right signal rBj(t) in said frequency band when the left-right amplitude ratio lllBj(t)ll / llrBj(t)Il is outside the range [panL(pri) / panR(pri) ; panL(pii) / panR(pii)] defined by said end-gain ratios: the attenuation value, denoted Gmask Bj(t), applied being a function of the distance between the left-right amplitude ratio lllBj(t)ll / llrBj(t)Il and the range defined by said end-gain ratios; 3) repetition of step 2) for several other frequency ranges Bj; summation of the SPijBj(t) signals which were generated for said angular sector Pi and said frequency bands Bj, in order to obtain a new audio signal Si(t) corresponding to one of the N components of the soundscape to be realized, repetition of steps 2), 3) and 4) for the other angular sectors Pi in order to obtain said N new audio signals Si(t).
[0013] Such a method makes it possible to generate a multi-channel audio signal from a stereo signal, the N channels of which correspond respectively to the sound content present in the N angular sectors according to which the sound panorama, encoded in the stereo signal, has been divided. In other words, each generated channel corresponds to the sound content present in the angular sector of the panorama encoded in the stereo file from which the channel was generated. The generation of such a multi-channel signal provides a high-quality signal for good reproduction of the original sound panorama, which allows the signal to be distributed across a set of loudspeakers, offering a good spatial listening experience with the reproduction of a stable soundstage that can be as close as possible to the original sound panorama encoded in the stereo signal.
[0014] In prior art solutions that use similarity functions, the aim is to identify in the signal the angular positions where there is a peak in sound energy, i.e., where the main sound sources are located. The sound from a main sound source present at the identified position is then extracted from the signal. Focusing primarily on the main sound sources does not allow for a qualitative reproduction of the original soundscape.
[0015] Conversely, in the solution according to the invention, angular sectors are first defined by dividing the soundscape, independently of their sound composition. Then, for each angular sector, a signal is generated, retaining the frequency components of the left and right signals for which the corresponding left-to-right amplitude ratio is contained within the range defined by the end-gain ratios corresponding to that angular sector, and attenuating the frequency components of the left and right signals whose left-to-right amplitude ratio lies outside said range. In particular, the further the amplitude ratio of a frequency component of the left and right signals is from said range, the greater the attenuation applied.
[0016] By summing the signals generated for a given angular sector and for several frequency bands, from the frequency components of the left and right signals, part of which has been preserved and another part has been attenuated according to the result of comparing their gain ratio with the range defined by the end gain ratios corresponding to the angular sector, we thus obtain a component (a channel) of the multi-channel signal which corresponds more precisely and completely, compared to the known methods of the prior art, to the sound (sound content) located in said angular sector of the sound panorama which has been encoded in the stereo signal.
[0017] The N signals thus generated (extracted) for the N angular sectors of the sound panorama which has been encoded in the stereo signal, thus form the N components of the multi-channel signal.
[0018] The multi-channel signal can then be used to reproduce (restore) a stable sound scene close to the soundscape that was encoded in the stereo signal, with a speaker system, at least three speakers, arranged around a listening area.
[0019] The generation process may also include one or more of the following features taken in any technically permissible combination.
[0020] According to one embodiment, the frequency band decomposition of the left signal l(t) and the right signal r(t) is carried out with a short-term Fourier transform (stft).
[0021] According to one embodiment, to smooth out abrupt variations resulting from attenuation, the attenuation value, denoted Gmasque_Bj(t), is smoothed over time with: - a first-order filter, and / or - a first time constant, named attack, for increasing variations, and / or - a second time constant, named release, for decreasing variations.
[0022] According to one embodiment, the attenuation Gmasque_Bj (0 is set with a selectivity factor which is defined as a function of the angular sector Pi.
[0023] According to one embodiment, the selectivity factor, denoted selectivity;, associated with the given angular sector P,, is calculated by the formula: selectivity;= max (selectivityL,; ; selectivityR,i) with selectivityL,i = Grej / (threshLi-RdBL>i) selectivityR,i = Grej / (threshRi-RdBR>i)
[0024] with
[0025] threshLi = 20.1ogi0(panL(pii) / panR(pii)) corresponding to a left threshold value in dB associated with the left end pu of the angular sector Pi,
[0026] threshRi = 20.1ogio(panL(pri) / panR(pri)) corresponding to a right-hand threshold value in dB associated with the right end pri of the angular sector Pi,
[0027] RdBL,i = 20.1ogio(panL(pi-Wi) / panR(pi-Wi)) RdBR,i = 20.1og10(panL(pi+wi) / panR(pi+wi)) with Wi= pri- pu and pi = i.Wi-1 + w / 2
[0028] Grej: a minimum rejection value in dB.
[0029] According to one embodiment, the angular sectors Pi are contiguous. It can also be envisaged that sectors partially overlap (i.e., cover one another).
[0030] The invention also relates to a data set comprising: - a multichannel audio signal S(t) generated according to any one of the preceding embodiments, and - for each of the N new signals of said multichannel audio signal S(t), characteristics of the angular sector, preferably the central direction and width of said angular sector, of the original sound panorama encoded in the stereo file from which said new signal was generated.
[0031] Otherwise, said set comprises: a multi-channel signal generated as proposed above, and associated characteristic (representative) data of the angular sectors of the sound panorama encoded in the stereo signal, from which the N new signals (channels) were generated to generate the multi-channel signal.
[0032] Said angular sector characteristics are for example stored in a computer file or accessible from a database.
[0033] The invention also relates to a computer program product, characterized in that it comprises program code instructions for implementing the steps of a process according to any one of the embodiments proposed above, when said program is executed by a computer system, such as a computer.
[0034] The invention thus relates to a method for broadcasting N audio signals Si(t) from a multichannel audio signal S(t), preferably generated according to any one of the embodiments proposed above, onto a plurality of loudspeakers, the method (also called the spatialization method) comprising the steps of:
[0035] - definition of a given number NHp of HPk loudspeakers, and, for each of the loudspeakers speakers, definition of the position, orientation and directivity of said speaker HPk;
[0036] - definition of the positions of N virtual sound sources SV; around the loudspeakers HPk, each virtual sound source SV; corresponding to the ith audio channel Si(t) generated from the multichannel audio signal S(t);
[0037] - determination for each of the HPk loudspeakers and each of the SVi sources, of a gain gk>iet of a delay ôk>i as a function of the distance dk>ientre the virtual source SV; and said loudspeaker HPk;
[0038] - supplying said generated audio signals S;(t), to each speaker HPk in applying to said loudspeaker HPk, for each audio signal Si(t), the gain gkji and the delay ôkji are determined for said loudspeaker. According to one embodiment,
[0039] The process may also include one or more of the following features taken in any technically permissible combination.
[0040] According to one embodiment, the determination for each of the speakers HPk and each of the sources SVi, of the gain gk>i and the delay ôkji to be applied, is also a function of a parameter, called the Rolloff parameter, denoted R; said determination of gain and delay is also carried out, for at least one listening position (Pp), depending also on:
[0041] - said Rolloff parameter R;
[0042] - of the distance between said loudspeaker HPk and the listening position Pp; and
[0043] - of the angle ak>p between the axis of the loudspeaker HPk and the listening position Pp.
[0044] According to one embodiment, the determination for each of the loudspeakers HPk and each of the sources SVi, of the gain gk>i and the delay ôkji to be applied, is carried out as a function of a set of directivity data of the loudspeaker HPk, called the directivity diagram, which provides for the angle ak>p, an attenuation value wakp, the attenuation value wakp being used, with also the distance d'k>pet as a function of said Rolloff parameter R, to determine an energy vector rE ip, as a function of which are calculated an angular error 0ip and an angular overlap ipi>p.
[0045] According to one embodiment, the calculation of a gain gk>i and a delay ôk>i being carried out for a set NP of given listening positions Pp, - a cost function Q(R) is defined for all positions as a function of the angular errors 0i>p and the corresponding angular overlaps rpip, - and we identify the value, denoted Ropt, of the Rolloff parameter R for which the cost function Q(R) is minimal,
[0046] The gain gkji and the delay ôkji to be applied are calculated by the formulas:
[0047] gk i=m / (dk>i )a
[0048] with a= R / (20.1ogi0(2)) And 1 m = ■ ôki— dk>i / c with c = 340 ms 1
[0049] According to one embodiment, the cost function Q(R) is calculated by the formula: Q(R) = a fi3+ p.oS+ Y.^ip+
[0050] with a, [3, y, and ô: weightings;
[0051] q0 = average of the angular errors I0ipl for all sources i at all listening positions p; ^0 = F standard deviation of the angular en-eur for all sources i and all listening positions p; p,1a average of successive overlaps for all sources i and all listening positions p; successive angular overlaps for all sources i and all listening positions p; NP being the number of listening positions p and Ns the number of sources i.
[0052] According to one embodiment, the N virtual sound sources SViSont are positioned around the speakers HPk on an arc of a circle.
[0053] According to one embodiment, the loudspeakers can be arranged in a regular or irregular manner.
[0054] The sound scene can be reproduced over a wide hearing area.
[0055] This process allows for spatialized sound reproduction in locations equipped with multipoint electro-acoustic reproduction systems, such as a loudspeaker network.
[0056] The method of generating a multi-channel signal from a stereo signal, as well as the method of broadcasting the corresponding multi-channel signal, is compatible with existing stereo audio sources (signals, files) which can be, for example, audio content streamed over the Internet, or live stereo mixing.
[0057] The invention also relates to a computer program product characterized in that it comprises program code instructions for implementing the steps of a broadcasting method according to any one of the broadcasting embodiments described above, when said program is executed by a computer system, such as a computer. The computer system includes a memory in which the stereo audio signal is stored or downloaded.
[0058] The invention also relates to a system for generating a multichannel audio signal, denoted S(t), from a stereo audio signal, the system comprising a stereo audio signal acquisition module; the supplied stereo audio signal comprising a left signal, denoted l(t), which was obtained by mixing monophonic signals according to a left panning law, denoted panL, and a right signal, denoted r(t), which was obtained by mixing monophonic signals according to a right panning law, denoted panR, in which the system comprises: a cutting module configured to cut the sound panorama (PS) into a given number N of angular sectors, denoted P;, preferably contiguous; a gain ratio calculation and comparison module, configured for: a given angular sector P; = [pu; pri] of the sound panorama (PS), and a given frequency band, denoted Bj, - calculate the value of the amplitude ratio, called the left-right amplitude ratio, denoted lllBj(t) ll / llrBj(t)ll, between the left signal lBj(t) in said frequency band Bj and the right signal rBj(t) in said frequency band Bj,
[0059] - for each end (pu; pri) of said angular sector P, calculate the ratio, called the end gain ratio panL(pii) / panR(pu) and panL(pri) / panR(pri), between the gain provided by the left panning law panL for said end of said angular sector R, and the gain provided by the right panning law panR for said end of said angular sector R;
[0060] - compare the left-right amplitude ratio lllBj(t)ll / llrBj(t)Il, with the ratios of calculated end gain panL(pri) / panR(pri) and panL(pu) / panR(pii); and a signal generation module configured for: - generate, in said given frequency band B j, an audio signal SK,Bj(t) from said left signal lBj(t) in said given frequency band Bj and said right signal rBj(t) in said given frequency band Bj, and according to the result of the comparison of the left right amplitude ratio lllBj(t)ll / llrBj(t)ll, with the end gain ratios panL(fold) / panR(fold) and panL(pri) / panR(pri);
[0061] the generation comprising the attenuation of the left signal lBj(t) in said frequency band Bj and of the right signal rBj(t) in said frequency band B when the left-right amplitude ratio lllBj(t)ll / llrBj(t)Il is outside the range [panL(pri) / panR(pri) ; panL(pii) / panR(pii)] defined by said end-gain ratios, the attenuation value, denoted Gmask Bj (t), applied being a function of the distance between the left-right amplitude ratio lllBj(t)ll / llrBj(t)Il and the range defined by said end-gain ratios; the system being configured for: - (i) for a given angular sector P;, execute the modulus of several times calculation and comparison of gain ratio for several frequency ranges Bj; - (ü) sum the SPijBj(t) signals that were generated for said sector angular P; and said frequency bands Bj, in order to obtain a new audio signal S;(t) corresponding to one of the N components of the soundscape to be created, - (iii) repeat steps (i) and (ii) for the other angular sectors P; so as to obtain said N new audio signals Si(t).
[0062] The invention also relates to a system for broadcasting N audio signals Si(t) of a multichannel audio signal S(t), preferably generated according to any one of the embodiments proposed above, onto a plurality of loudspeakers, the system comprising: - a user interface allowing the position and orientation of each of said HPk speakers to be defined for a number NHp of HPk speakers; each HPk speaker being associated with directivity data; - a module for defining the positions of N virtual sound sources SV; around the speakers HPk, each virtual sound source SVi corresponding to the i-th audio channel Si(t) generated from the multichannel audio signal S(t); - a determination module for each of the HPk speakers and each of the SVi sources, of a gain gkji and a delay ôkji as a function of the distance dkji between the virtual source SV and said HPk speaker; - a transmission module configured to transmit said generated audio signals Si(t), to each speaker HPk by applying to said speaker HPk for each audio signal Si(t), the gain gkji and the delay ôkji determined for said speaker.
[0063] According to a particular aspect, the diffusion system is configured to implement any one of the characteristics of the diffusion process presented above.
[0064] According to a particular aspect, the audio signal generation system is configured to implement any one of the generation process characteristics presented above.
[0065] Advantageously, the diffusion system which allows the multi-channel audio signal to be broadcast to the loudspeakers includes in particular a digital-to-analog conversion system for converting the multi-channel audio signal, and an amplification system connected to the loudspeaker system.
[0066] It can be foreseen that the computer program or each computer program is downloadable from a communication network or stored on a computer-readable medium. Brief description of the drawings
[0067] Other features and advantages of the invention will become apparent from the following description, which is purely illustrative and not limiting and should be read in conjunction with the accompanying drawings, on which:
[0068] - [Fig. 1] [Fig. 1] is a schematic view of an example of a sound panorama which was encoded into a stereo audio signal, the sound panorama comprising different sound sources, such as maracas, guitar, saxophone, percussion, located at different angular positions (corresponding to panning indices (or panning index) p / defined in a range between -1 and 1) of the sound panorama;
[0069] - [Fig.1A] [Fig.1A] is a circular arc graph on which are mentioned panning indices pt from -1 to 1, corresponding to angular positions in the sound panorama schematized in [Fig.l] which has been encoded in the stereo audio signal, w / defining the width of an angular sector P t centered on a panning index p h with ph and p ri the end values of this sector;
[0070] - [Fig.2] [Fig.2] is a graph representing a panning law, also called panoramic law or distribution law, with constant power including a law of left panning law with which the left signal of a stereo signal was encoded and a right panning law with which the right signal of a stereo signal was encoded, with on the x-axis the panning index p; corresponding to an angular position in the sound panorama that was encoded in the stereo file, and on the y-axis the corresponding dB gains giL and giR for the left panning law and the right panning law;
[0071] - [Fig.3] [Fig.3] is a graph illustrating different attenuation masks noted Gmasque_slct (1,2,5,10,20) applicable to a frequency component of a stereo signal as a function of the distance from the search area for different values of the selectivity factor (in dB), with the panning index on the x-axis;
[0072] - [Fig.4] [Fig.4] is a synoptic diagram of an extraction program (method) spectral according to one embodiment;
[0073] - [Fig. 5] [Fig. 5] is a synoptic diagram of the signal extraction program (method) atmosphere;
[0074] - [Fig.6] [Fig.6] is a diagram showing the positions symbolized by triangles, of a set of loudspeakers at the front of the stage in a room, such as a theatre hall;
[0075] - [Fig.6A] [Fig.6A] is a diagram showing the positions of the loudspeakers of the [Fig.6], behind which virtual sound sources have been distributed, corresponding to N signals of a multi-channel signal, at angular positions relative to a listening area, for which it is desired that the listener perceive said virtual sound sources, the positions being defined on an arc of a circle, according to the corresponding angular sectors of the original sound panorama on the basis of which the N signals were generated;
[0076] - [Fig.7] [Fig.7] is a diagram illustrating another speaker configuration distributed around a listening area, and a distribution of virtual sound sources on an arc around and behind the speakers, as well as a positioning of two ambient sources at the ends of the arc;
[0077] - [Fig.8] [Fig.8] is a diagram showing a plurality of loudspeakers and the distance of a virtual source corresponding to a new signal generated from the stereo file, relative to these speakers;
[0078] - [Fig.8A] [Fig.8A] is a diagram reproducing that of [Fig.8] to which are added two listening positions, the diagram showing the distance between a speaker and a listening position, and the listening angle relative to the axis of that speaker;
[0079] - [Fig.8B] [Fig.8B] is a diagram reproducing that of [Fig.8A] illustrating a energy vector localization model for the same signal duplicated on multiple speakers;
[0080] - [Fig.9] [Fig.9] is a diagram showing three signals generated from the stereo signal, their application to a set of three speakers with gain and corresponding delays, a gain and delay pair being defined for each signal and each speaker to which said signal is applied. DETAILED DESCRIPTION
[0081] Embodiments are described below with reference to the accompanying drawings. Similar numbers refer to similar features in all drawings. However, the invention can be implemented in many different forms and should not be construed as being limited to the embodiments shown here. The scope of the invention is defined by the accompanying claims.
[0082] A reference throughout the specification to "an embodiment" means that a particular feature, structure, or characteristic described in relation to an embodiment is included in at least one embodiment of the present invention. Thus, the appearance of the phrase "in an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0083] A sound panorama, or sound scene, can be defined by: - a center, and - a distribution of several sound sources between two extremities around said center, each source being located at a given position or within a given angular range, also called the spatialization angle or spatialization angular range.
[0084] A stereo mix (stereo signal) consists of a left signal denoted l(t) and a right signal denoted r(t), which when played back on a suitable speaker system allows a sound image to be recreated.
[0085] To construct a stereo signal of a sound scene (or sound panorama) comprising several sound sources, each sound source (instrument or voice) is placed in the stereo mix by applying a gain law, also called panning law or panning law, to each of the left and right channels of the sound source.
[0086] Different gain laws can be used. Preferably, the gain law used is a constant power gain law, as illustrated for example in [Fig. 2]. Different variants of the gain law can be used, such as a linear gain law or a -4.5 dB gain law.
[0087] According to a preferred embodiment, the audio signals emitted by the sound sources that make up the original soundscape have been encoded in the stereo signal according to a constant power gain law.
[0088] In the remainder of the description, the supplied stereo audio signal is thus considered to comprise: - a left-hand signal l(t) which was obtained by mixing monophonic signals according to a left-hand panning law, denoted panL, and - a right-hand signal r(t) which was obtained by mixing monophonic signals according to a right-hand panning law, denoted panR (see [Fig.2]).
[0089] From said stereo signal, we seek to generate several audio signals (channels) corresponding to different angular sectors of said original sound panorama which has been encoded in the stereo signal, so that each generated channel makes it possible to reproduce the sound content present in the corresponding angular sector of the original panorama.
[0090] The channels thus generated, as detailed below, make it possible to generate a new global audio signal, called a multi-channel signal, composed of the signals generated for the different angular sectors of the sound panorama which has been encoded in the stereo signal.
[0091] This overall audio signal, which corresponds to the soundscape that we wish to reconstruct or reproduce, can then be broadcast on a speaker system arranged and configured as detailed later.
[0092] I. Generation of a multi-channel audio signal from a stereo audio signal
[0093] A number N of angular sectors, called angular extraction sectors, are defined according to which the encoded sound panorama is cut.
[0094] The sound panorama encoded in the stereo signal extends over an angular range, for example 180°, which translates into a panning index range, for example between -1 and 1, or between 0 and 1.
[0095] Each angular extraction sector can thus be defined by a central panning index p; of the angular sector (between -1 and 1 or between 0 and 1), and a width parameter w; corresponding to the width of said angular sector, as illustrated for example in [Fig.1A].
[0096] The sound panorama PS (see [Fig. 1]) is thus divided into a given number N of angular sectors P, preferably contiguous. The angular sector P is defined by the range [pu; pri] (see [Fig. 1A]) with pli: the left end of the angular sector, and pri: the right end of the angular sector.
[0097] We thus obtain N angular sectors P;, with i ranging from 1 to N. An embodiment for the calculation of N is subsequently proposed as a function of the number of loudspeakers on which we wish to spatialize the global multi-channel signal that we are going to generate.
[0098] The left signal l(t) and right signal r(t) of the stereo signal are subjected to a frequency decomposition (spectral decomposition) using a Fourier transform. The Fourier transform used is preferably a short-time Fourier transform (stft or STFT hereafter) with possible overlap.
[0099] The number of points in the Fourier transform can depend on the application. For real-time applications requiring very low latency (for example, broadcasting a spatialized concert on a given number of speakers from a stereo mixing console), a Fourier transform of 128 points provides good results for latency below 3 ms at a sampling frequency of 48,000 Hz. For applications not subject to high latency constraints (playing stereo recordings, etc.), increasing the number of points improves definition and therefore audio quality. A 1024-point transform is therefore a viable option.
[0100] Frequency decomposition allows the left and right signals to be decomposed into several components for different given frequency bands.
[0101] Thus, for a given frequency band Bj, we obtain the left signal lBj (t) in said frequency band Bj and the right signal rBj(t) in said frequency band Bj. j = 1 to D with D the number of frequency bands according to which the left and right signals are decomposed.
[0102] The generation (extraction) of a new signal for a given frequency band and a given angular extraction sector is described below in relation to the block diagram of [Fig.4].
[0103] For a given frequency band Bj, the value of the amplitude ratio, called left-right amplitude ratio, denoted lllBj(t)ll / llrBj(t)Il, is calculated between the left signal lBj(t) in said frequency band Bj and the right signal rBj(t) in said frequency band Bj.
[0104] For each end pu; pri of said angular sector R, the ratio, called the end gain ratio: panL(pii) / panR(pu) and panL(pri) / panR(pri), is calculated between the gain provided by the left panning law PanL for said end pu; pri of said angular sector R, and the gain provided by the right panning law PanR for said end pu; pri of said angular sector P;
[0105] The left-right amplitude ratio lllBj(t)ll / llrBj(t)ll is then compared with the calculated end-gain ratios panL(pii) / panR(pu) and panL(pri) / panR(pri).
[0106] An audio signal SPijBj(t), associated with said angular sector Pi and with the frequency band Bj, is generated from said left signal lBj (t) and said right signal rBj (t) in said given frequency band Bj, according to the result of the comparison above.
[0107] The generation of the audio signal SPijBj(t) includes a step of attenuating the left signal lBj(t) in said frequency band Bj and the right signal rBj(t) in said frequency band Bj when the left right amplitude ratio lllBj(t)ll / llrBj(t)Il, noted PBj(t), in [Fig.4], is outside the range [panL(pri) / panR(pri) ; panL(pii) / panR(pii)] defined by said end gain ratios.
[0108] The attenuation value denoted Gmasque_Bj(t) applied to the left signal lBj(t) and the right signal rBj(t) in the given frequency band Bj to generate the new signal SPijBj is a function of the distance between the left-right amplitude ratio lllBj(t)ll / llrBj(t)Il and the range defined by the said end-gain ratios. In other words, we check whether the left-right amplitude ratio is within the range (between the two end-gain ratio values) or whether the left-right amplitude ratio is outside this range and by how much.
[0109] In particular, the further the left-right amplitude ratio lllBj(t)ll / llrBj(t)ll is from the interval defined between the end-gain ratios panL(pii) / panR(pu) and panL(pri ) / panR(pri), the more the left signal lBj (t) and the right signal rBj (t) in said given frequency band Bj are attenuated for the generation of the corresponding audio signal S^ / t).
[0110] Such treatment makes it possible to attenuate sound sources located outside the angular sector P; (extraction sector).
[0111] The extraction process can be performed for each signal frame. The attenuation to be applied from one frame to another for the same frequency or frequency band can vary abruptly, resulting in an undesirable pumping effect. Time smoothing can therefore be applied to the attenuation value (mask) from one frame to the next.
[0112] According to a preferred embodiment, to smooth out abrupt variations resulting from attenuation, the attenuation value is smoothed over time with:
[0113] - a first-order filter,
[0114] - a first time constant, called attack, for increasing variations, And
[0115] - a second time constant, called relaxation, for the variations decreasing.
[0116] According to a preferred embodiment, the smoothing is characterized by two time constants, namely the attack and release times, to allow for different smoothing depending on the direction of change of the attenuation. A short attack time (for example, on the order of 1 ms) is preferred to preserve transients. A release time on the order of 40 ms helps to avoid overly noticeable pumping effects.
[0117] The attenuation Gmasque_Bj (0 can be adjusted with a selectivity factor which is defined as a function of the angular sector R.
[0118] The selectivity factor, denoted selectivity, associated with the given angular sector R, can be calculated by the formula:
[0119] selectivity;= max (selectivityL,; ; selectivityR,;) with selectivityL,i = Grej / (threshLi-RdBL>i) selectivityR,i = Grej / (threshRi-RdBR>i)
[0120] With threshLi = 20.1ogi0(panL(pii) / panR(pii)) corresponding to a left threshold in dB of the search area (angular sector) Pi, threshRi = 20.1ogi0(panL(pri) / panR(pri)) corresponding to a right threshold in dB of the search area (angular sector) P;. RdBLj = 20.1og10(panL(pi-wi) / panR(pi-wi)) RdBRj = 20.1og10(panL(pi+wi) / panR(pi+wi)) with Wi= p, pu and pi = i.Wi-1 + w / 2 Grej (in dB): a minimum rejection value
[0121] In other words, depending on the distance of the amplitude ratio from the thresholds defined by the end-gain ratios, a frequency attenuation mask can be defined to attenuate frequencies located far from the thresholds and to preserve those located in the search area (angular extraction sector). The selectivity factor increases the applied attenuation as a function of distance. It can be compared to the slope of a bandpass filter.
[0122] The extraction mask (attenuation value) can be calculated as follows:
[0123] The following variables are defined: - L(f,t)=STFT(l(t)), where STFT is the short-term Fourier transform function, and f is the frequency. - R(f,t)=STFT(r(t)), - P(f,t)=20.1ogio(HL(f,t) / R(f,t)Il), the amplitude difference in dB between the L and R signals in the frequency domain, - threshLi=20.1ogio(panL(pii) / panR(pu)), the left threshold in dB of the search area, - threshRi=20.1ogio(panL(pri) / panR(pri)), the right threshold in dB of the search area, - At: Interval between two frames of the short-time Fourier transform (STFT) ' Tattack, Trelectse of the two attack and release time constants - selectivityi: the selectivity factor of the filter for the angular sector centered on the panning index pi.
[0124] The frequency mask (attenuation value) does not modify the phase of the signals l(t) and r(t) and is therefore real. It is a gain in dB to be applied at each instant to each frequency or frequency band.
[0125] The gain is defined as:
[0126] GdB^f,t)=min(threshu-P^f,t),0)+min (P(f,t)-threshRi,O)
[0127] The gain in dB GdB(f,t) can then be: - amplified by the selectivity factor to obtain a gain G(f,t) - converted into linear amplitude, - preferably, smoothed to obtain a gain Gsmooth(f,t). The smoothing is, for example, a first-order exponential smoothing. Note that other smoothing methods can also be used (ramp, second order, etc.). - possibly, the gain Gsmooth(f,t) can be corrected by a compensation gain Gcomp (pi) depending on the panning law used and the index pi.
[0128] with
[0129] G(f,t)=WdB^t) selectivity, G smoothÇf ,t)— aG(f ,t)+(1 ” Cf) .Gsmooth(.f ,t~ 1 ) with / ________.èl______ 1((;( / ; 0 g(f, ti)] > o u j ________________ Gmask (yGsmooth^f it'j-Gcomp^Pt)
[0130] We thus obtain for a frequency band Bj: GmaSqUe_Bj(t) Gsmooth _ Bj(t)*Gcomp(Pi)
[0131] The mask can thus be applied in the frequency domain to the left and / or right signal of the stereo signal or to a mix of the two signals depending on the value of p; in order to perform the extraction. As illustrated in [Fig. 4], before applying the mask, a first weighting can be applied to the left signal lBj(t) and a second weighting can be applied to the right signal rBj(t) depending on the panning index p;.
[0132] Examples of masks, defined for different selectivity values, are illustrated in [Fig.3] with the panning index value pi on the abscissa and the gain value in dB on the ordinate.
[0133] Conversely, when the left-right amplitude ratio lllBj(t)ll / llrBj(t)ll is within the defined range between the end-gain ratios panL(pii) / panR(pii) and panL(pri) / panR(pri), the audio signal SFl Bj,(t) can be generated from said left signal ZBj(t) and said right signal rBj(t) in said given frequency band Bj without attenuation. The left signal lBj(t) and the right signal rBj(t) in said given frequency band Bj can thus be summed without attenuation.
[0134] A gain can be applied to the filtered signal in order to compensate for an energy loss related to the panning laws used and to ensure a balanced spatial amplitude.
[0135] The resulting signal can finally be synthesized in the time domain using an inverse short-term Fourier transform (istft) to obtain the signal SPijBj(t).
[0136] The preceding steps are repeated for several other frequency ranges Bj, to generate other audio signals SPijBj(t), preferably so as to cover the frequency range [20 Hz - 20000 Hz].
[0137] The SPijBj(t) signals which have been generated for said angular sector Pi and said frequency bands Bj, are summed in order to obtain a new audio signal Si(t) corresponding to one of the N components (the ith) of the sound panorama to be realized.
[0138] Such an operation of generating the signal Si(t) associated with a given angular extraction sector, P, makes it possible to extract from the stereo audio signal, a signal of quality corresponding to the audio content which is in said given angular sector of the sound panorama encoded in the stereo mix.
[0139] The preceding steps are repeated for the other angular sectors Pt so as to obtain the N new audio signals Si(t) corresponding to the N channels of the multi-channel signal that can be played back on a loudspeaker system as detailed below. In other words, the extraction steps described above, applied to the stereo audio signal, make it possible to generate N signals corresponding to as many tracks, each track corresponding to a given angular sector of the soundscape. The N signals Si(t) with i = 1 to N form the multi-channel audio signal S(t).
[0140] As explained later, for each loudspeaker, a gain and delay pair can be defined for each signal S;(t), and the signals S;(t) can be summed to be applied to said loudspeaker with the corresponding gain and delay pairs.
[0141] The extraction steps allow the stereo signal panorama to be divided into angular sectors and a signal to be generated for each angular sector, which then contains the sources mixed in that spatial area. The sound sources located in Sources within a given angular sector are thus preserved in the generated signal based on the extracted signal associated with that angular sector, while sources outside said angular sector are attenuated. In particular, the further the source is from said sector, the more its amplitude in the generated signal is modified in the direction of attenuation.
[0142] The extractions described above thus correspond to the application of a spatial bandpass filter to a given angular sector. Adjusting the attenuation of sound sources located outside the extraction sector corresponds to adjusting the filter's selectivity. By combining several of these filters, each associated with an angular sector, the stereo image can therefore be decomposed into several sectors.
[0143] The number of sectors used, the setting of the out-of-band attenuation, the possible overlap of the different sectors makes it possible to define a filter bank whose signals from it allow the continuous recomposition of the panorama.
[0144] According to one embodiment, and as illustrated for example in [Fig. 5], the method also includes a step of extracting the left and right ambience signals. Ambient signals are defined as the out-of-phase components of the left and right signals located in the left or right channel of the original mix, excluding a central area. As in the example illustrated in [Fig. 5], the method can be similar to that of [Fig. 4] by working on the difference or ratio between the left and right signals and applying a mask, denoted Gmask_left_Bj(t), to the difference or ratio of the left and right signals to obtain the left ambience signal, and a mask, denoted Gmask_right_Bj(t), to the right signal, to the difference or ratio of the left and right signals to obtain the right ambience signal.
[0145] In other words, according to the same principle as the spatial bandpass filter described above, a right-pass filter and a left-pass filter can be defined from the compared amplitudes of the right and left signals and the distance to a threshold.
[0146] According to one embodiment, the extraction method and system uses Nsources parameterized low-pass filters with a width Wi = 2 / Nsources and a panning index p = i.Wi - 1 + W / 2, and preferably a left-pass and right-pass filter for extracting the two ambient signals. As explained above, the selectivity parameter of each filter can be adjusted to achieve a rejection of a minimum value Grej (in dB) at a distance of w / 2 from the extraction zone.
[0147] The audio processing method described can perhaps be carried out using an electronic and computer processing system.
[0148] The processing system is presented for example in the form of a processor and a data memory in which computer instructions executable by said processor are stored, or in the form of a microcontroller.
[0149] In other words, the functions and steps described can be implemented as a computer program or via hardware components (e.g., programmable gate arrays). In particular, the functions and steps described above can be implemented by instruction sets or computer modules implemented in a processor or controller, or by dedicated electronic components, or by components such as field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs). It is also possible to combine computer and electronic components.
[0150] The processing system is thus an electronic and / or computer unit. When it is specified that the system is configured to perform a given operation, this means that the system includes computer instructions and the corresponding means of execution that enable the said operation to be carried out and / or that the unit includes corresponding electronic components.
[0151] IL Multi-channel signal diffusion on a plurality of loudspeakers
[0152] A method for broadcasting N audio signals S;(t) of a signal is described below. Multichannel audio S(t) on a plurality of speakers. S(t) is formed by the N signals Si(t) with i = 1 to N
[0153] According to one embodiment, the multichannel audio signal S(t) was generated using the generation (extraction) process described above. Each of the N audio signals is associated with an angular sector corresponding to the angular sector P of extraction.
[0154] Speakers
[0155] A number NHp of HPk loudspeakers is defined, with k = 1 to NHp.
[0156] According to one embodiment, the number of sources N, also denoted Nsources, is defined depending on the number of speakers NHp by the relation: Nsources = min (Nsmax,NHp+l-mod(NHp,2))
[0157] Nsmax being a given maximum number of sound sources, for example equal to 13.
[0158] Thus, in the previously described extraction process for generating a The new multi-channel signal S(t), the number N of audio signals S;(t) extracted from the stereo file, can be determined from the number of available or desired speakers. The number N does not include any additional ambient signals.
[0159] Advantageously, the number N, which also corresponds to the number of sources Nsource s that we wish to broadcast, is odd in order to keep a signal corresponding to the center of the stereo image, and not to exceed a maximum number Nsmax of sound sources, at the risk of creating artifacts by trying to extract sectors of the panorama that are too close together. Preferably N is less than or equal to 13.
[0160] Each of the loudspeakers is characterized by: - its position (in the Earth's frame of reference and / or in relation to other loudspeakers or in relation to landmarks in a room), - the orientation and directivity of said HPk speaker.
[0161] In the example of [Fig.6], the positions of a plurality of loudspeakers for an arrangement in a theatre at the front of the stage have been illustrated by triangles.
[0162] Positioning of N virtual sound sources SV;
[0163] The positions of Nsource virtual sound sources SV are defined around the NHplou-louvers HPk. Each virtual sound source SVi corresponds to the ith audio channel Si(t) of the multichannel audio signal S(t) that we wish to make a listener perceive at a given angular position relative to a listening area when the multichannel audio signal S(t) is played by the loudspeakers, in particular when this ith audio channel Si(t) is supplied to the loudspeakers.
[0164] According to one embodiment, the positions of the virtual sound sources are defined, preferably in a regular manner, on an arc of a circle surrounding all the loudspeakers and whose distance to each loudspeaker is as constant as possible.
[0165] In the graph of [Fig.6A], the position of the virtual sound sources SV has been illustrated; which are distributed around the speakers HPk, positioned, for example as in [Fig.6], on an arc of a circle spaced apart from the speakers and behind the speakers with respect to the listening area symbolized by the star in the center of the graph.
[0166] The placement in an arc of the virtual sound sources SVi allows, thanks to the delays of the spatialization algorithm, to keep as much in phase as possible all the signals of the multi-channel signal when they are played by the loudspeakers and thus to maintain a coherent panorama for the central area of the listening area.
[0167] Advantageously, the center of the arc of the circle is as equidistant as possible from all the loudspeakers.
[0168] The center of the circular arc can be defined by an optimization method (e.g., gradient descent, Broyden-Fletcher-Goldfarb-Shanno (BFGS), quadratic optimization) by minimizing the standard deviation of the distances from each loudspeaker to the center of the arc. The centroid of the loudspeaker positions can be given as the initial position. If all the points are collinear (speakers perfectly aligned), the initial position is given on the listening area side.
[0169] Once the center is determined, the radius of the circular arc is defined as the maximum distance from the center to each speaker plus an offset, for example equivalent to 10% of the distance between the most distant speakers.
[0170] Advantageously, and as illustrated for example in [Fig. 7], the virtual sound sources SV are placed on the circle at preferably regular angles:
[0171] - the angle Q between the center, the first and the last speaker is calculated;
[0172] - the arc of a circle going from the 1st speaker to the last is then divided into a number of sectors depending on the number of sources. Advantageously, especially when both ambient signals are present in the multi-channel signal to be played on the speakers, the number of sectors is equal to Nsources+ 2 sectors and each sector has an angle 0, with 0 = Q / (Nsources+ 2).
[0173] The two ambient signals can be spatialized at the ends of the arc of the circle.
[0174] The use of DBAP (distance based amplitude panning) and / or WFS (wave field synthesis) algorithms allows for a good rendering of the re-spatialization of the channels of the multi-channel signal over a wide listening area.
[0175] As mentioned above, each channel of the multi-channel signal is reproduced at a given virtual position in space. In other words, when this channel is supplied to all the loudspeakers, the aim is for a listener to perceive the sound source corresponding to this channel as being located at a given position.
[0176] For this purpose, each signal (channel) of the multi-channel signal is played on all N hp loudspeakers of the playback system with a gain (amplitude panning) and a delay calculated as a function of the distance between the virtual position of the sound source (virtual sound source) and each of the loudspeakers.
[0177] In other words, we determine for each of the HPk loudspeakers and each of the SV sources:
[0178] - a gain gk>i and
[0179] - a delay ôk>i
[0180] depending on the distance dkji between the virtual source SV; and said speaker HPk.
[0181] As illustrated in the example in [Fig. 9] for three signals Sb S2, S3, the signals audio generated S;(t) can thus be provided to each speaker HPk by applying to said speaker HPk for each audio signal Si(t), the gain gkji and the delay ôkji determined for said speaker.
[0182] Calculation of profit and time
[0183] According to a preferred embodiment, for each virtual sound source SVi (a channel of the multi-channel signal), the gains gk>i to be applied to the speakers HPk are calculated as follows based on the distance dk>i between the virtual sound source SV; and the speaker HPk:
[0184] gk i=m / (dk 4)a
[0185] with a= R / (20.1ogi0(2))
[0186] and i m= , 1 J
[0187] R (Rolloff) is a global calibration parameter of the spatialization system. Preferably, R is a parameter calculated several times until an optimization value Ropt is reached as proposed below.
[0188] The delays are calculated as follows:
[0189] ôk> ;= dkji / c with c = 340 ms 1
[0190] Optimization of the Rolloff parameter
[0191] As explained above, the determination for each of the speakers HPket each of the sources SVi, of the gain gkji and of the delay ôkji to be applied, is thus a function of a parameter, called the Rolloff parameter, denoted R.
[0192] To tune the spatialization algorithm, the perceived spatialization of each source at each listening position is simulated. This is done using an energy vector model, preferably an extended one. Those skilled in the art may refer to the extended energy vector model described in the publication by Jacquet, Damien; Dominguez, Johana; Petiot, Jean-François; 2024; entitled "Perceptual validation of the extended energy vector model for audio objects localization in large audience areas"; available at: https: / / aes2.org / publications / elibrary-page / ?id=22359. When the same audio signal is duplicated and reaches a listener via different acoustic paths, if the replicas arrive with a delay relative to the first signal that is less than a threshold (around 30 ms), they are perceived as a single merged signal from a single source.The apparent direction of this merged source can be modeled by summing the energies from each source as a vector.
[0193] For a given listening position Pp, the determination of the gain gk i and the delay ôk i to be applied to the signal Si for the speaker HPk (in order to obtain the virtual sound source SVi at the desired position) is also carried out as a function of the distance d k>p between said speaker HPk and the listening position Pp; and of the angle akp between the axis of the speaker HPk and the listening position Pp ([Fig.8A]).
[0194] Advantageously, the determination of the gain gk>i and the delay ôk>i to be applied is carried out also as a function of a set of directivity data of the loudspeaker HPk, called the directivity diagram, provides, for the angle akp, an attenuation value wakjP. The attenuation value wak, p is used (along with the distance d k>p and as a function of the Rolloff parameter R) to determine an energy vector rE ip, from which an angular error 0i>p and an angular overlap are calculated. <p-
[0195] Figure 8B illustrates the energy vector localization model for the same signal duplicated across multiple loudspeakers with different loudspeaker gains. Figure 8B shows the energy vectors from each loudspeaker at position P1 and the energy vectors from each loudspeaker at position P2. For each listening position p, the set of energy vectors provides the apparent direction DAip of the source i for the corresponding position p.
[0196] The perceived energy from each loudspeaker depends on: - the spatialization gain of the source on the speaker; - the level decrease due to acoustic propagation between the loudspeaker and the listener; - attenuation due to the directivity of the loudspeaker and the angle between the axis of the loudspeaker and the listener.
[0197] In order to take into account the effects of precedence and temporal masking, the model can be extended so as to weight the energies with their time of arrival relative to the first auditory event.
[0198] We therefore obtain the following model ÿÇ_,i,p, denoted r?, : / f —A With: _Mdir,k' ^ac,k' Sk) ^k—1 ^d.ir,k' ®ac,k- = resulting vector indicating the apparent direction DAipde of the source SVi (at listening point p) = unit vector indicating the direction of the loudspeaker k at the listening point p
[0199] gk = signal spatialization gain on loudspeaker k W&c.k = attenuation due to acoustic propagation from the loudspeaker k at the listening point p djQc.k^re / ^dre / ' / d'k.p with sref the sound level of loudspeaker k measured along the axis of loudspeaker k at the distance dref. Since intensity calculations are relative, we can take Sref = dref = 1 = attenuation due to the directivity of the loudspeaker k and the angle "dir" between the axis of the loudspeaker k and the listener at the listening point p. The attenuation can result from a measured or approximated using a generic directivity function. To simulate the directivity of a loudspeaker with an opening angle "dir", one can, for example, use a cardioid function whose attenuation at an angle "dir" / 2 is equivalent to -6dB:
[0200] 6 edir i A + cos(dir)\ 20. log10 ---- J / 1 -F cos akxedir ^dir,kM = (------J \ Z /
[0201] By simplification ak>pest denoted ak in the equation above.
[0202] j 'c weighting factor based on the delay of the signal coming of the loudspeaker k at the listening position p relative to the first auditory event.
[0203] &
[0204] with ôack= acoustic propagation delay of the loudspeaker k at the listening position p
[0205] ôac>k= d'k>p / c, d'k>p being the distance of the loudspeaker k to the listening point p.
[0206] and ôspat> k= spatialization delay on loudspeaker k.
[0207] r = slope of the echo threshold, r depends on the nature of the signal. An average value of -0.25 dB / ms can be used.
[0208] The perceived width of the SVi source at listening point p can be modeled by the following equation:
[0209] 5 180° W = —.----. 2 c os 1U ey 8 TT
[0210] The energy vector model, in particular extended, presented above makes it possible to simulate the perception of the reconstructed panorama at each seat in an audience.
[0211] Simulation of the perception of the reconstructed panorama at each seat in an audience
[0212] According to one embodiment, the hearing area is sampled by defining the position of virtual seats for a given seat density per m2.
[0213] For each seat, we simulate with given spatialization parameters the perceived directions and perceived widths of the sources at each virtual seat.
[0214] A correct reproduction of the panorama can be defined as follows:
[0215] - the angle of perception error at the listening position between the direction of the source virtual and the perceived direction from this same source must be minimal for all seats;
[0216] - the widths perceived at the listening position of each source do not create neither recovery nor disjunction.
[0217] This can be translated mathematically as follows:
[0218] For the listening position Pp and the source SV;, the perceived source (also called the apparent source) is denoted S'i, such that S'i = Pp + pr- * E ip
[0219] corresponding to the apparent direction given by the extended energy vector model for the source SVi at the listening position Pp.
[0220] The angular error is defined as the vertex angle Pp formed by the points SVi, Pp, S'i and is denoted 0,p
[0221] The successive apparent angular overlap of two apparent sources, denoted S'i and S'i+i, is defined as the following angle:
[0222] Wp ;= nï- - w / 2- wi+1 / 2 AWt
[0223] With w; and wi+iles the apparent widths of the sources S'i and S'i+i
[0224] If W>0, the sources are considered disjoint, if W<0 the sources overlap. For correct reconstruction, 104 must tend towards 0.
[0225] The calculation of a gain gk>i and a delay ôk>i is performed for a given number NP of listening positions Pp with p = 1 to NP.
[0226] A cost function Q(R) is defined for the set of positions as a function of the angular errors 0i>p and the corresponding angular overlaps rpi,P.
[0227] We identify the value, denoted Ropt, of the Rolloff parameter R for which the cost function Q(R) is minimal.
[0228] The cost function Q(R) is calculated using the formula:
[0229] Q(R) = a .^0+ ^.rrB+ y.
[0230] with a, [3, y, and ô: weightings;
[0231] q0 = average of the angular errors I0ipl for all sources i at all listening positions p;
[0232] 6 = Vpe of the angular error for all sources i and all listening positions p;
[0233] pj, ,1a average of successive recoveries for all the sources i and the set of listening positions p;
[0234] <7ip= yP successive angular overlaps for all sources i and all listening positions p;
[0235] NP being the number of listening positions p and Ns the number of sources i.
[0236] The value of R that minimizes the function Q(R) corresponds to the optimized value Ropt, and the gain gkji and the delay ôkji to be applied are thus calculated by the formulas:
[0237] gki= m / dk>ia
[0238] with a = Ropt / 20.1ogi0(2)
[0239] m = 1 / yj dk r i A 2a
[0240] ôk>i= dk,; / c
[0241] with c = 340 m / s
[0242] The invention is not limited to the embodiments illustrated in the drawings.
[0243] Furthermore, the term "including" does not exclude other elements or steps. In addition, features or steps that have been described with reference to one of the embodiments set out above may also be used in combination with other features or steps from other embodiments set out above.
[0244] The spatial extraction algorithm can be viewed as a filter and therefore used upstream of any processing where it is necessary to extract from the original signal the audio components located within a specific panning area. The algorithm will thus function as the first stage of multiband processing. Signal recomposition at the end of the processing chain must include a respatialization method, such as the one proposed above or another method (binaural, Dolby, etc.).
[0245] The spatialization method presented above is also applicable to a non-regular spacing of virtual sources.
[0246] The panning and selectivity indices of each bandpass filter can be individually adjusted to recreate a continuous panorama using the same simulation method as the perceived or audible panorama. In this case, the extraction and reproduction widths may not be identical for all sources.
[0247] The rendering width can then be reproduced as a function of a source width parameter in the spatialization algorithm (method).
[0248] The diffusion (spatiation) process described may be implemented using an electronic and computer processing system.
[0249] The processing system is presented for example in the form of a processor and a data memory in which computer instructions executable by said processor are stored, or in the form of a microcontroller.
[0250] In other words, the functions and steps described can be implemented as a computer program or via hardware components (e.g., programmable gate arrays). In particular, the functions and steps described above can be implemented by instruction sets or computer modules implemented in a processor or controller, or by dedicated electronic components, or by components such as field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs). It is also possible to combine computer and electronic components.
[0251] The processing system is thus an electronic and / or computer unit. When it is specified that the system is configured to perform a given operation, this means that the system includes computer instructions and the corresponding means of execution that enable the said operation to be carried out and / or that the unit includes corresponding electronic components.
[0252] The invention is not limited to the embodiments illustrated in the drawings. Accordingly, it should be understood that, where the features mentioned in the appended claims are followed by reference numerals, these numerals are included solely for the purpose of improving the intelligibility of the claims and are in no way limiting the scope of the claims.
[0253] Furthermore, the term "including" does not exclude other elements or steps. In addition, features or steps that have been described with reference to one of the embodiments set out above may also be used in combination with other features or steps from other embodiments set out above.
Claims
1. Demands Method for generating a multichannel audio signal, denoted S(t), said audio signal S(t) being generated from a supplied stereo audio signal, the supplied stereo audio signal comprising: - a left-hand signal, denoted l(t), which was obtained by mixing monophonic signals according to a left-hand panning law, denoted panL, and - a right-hand signal, denoted r(t), which was obtained by mixing monophonic signals according to a right-hand panning law, denoted panR, in which the process includes the following steps: 1) cutting a sound panorama (PS) into a given number N of angular sectors, denoted Pb, preferably contiguous, 2) for a given angular sector Pt = [pu; pri] of the soundscape (PS), and for a given frequency band Bj, - calculation of the value of the amplitude ratio, called left-right amplitude ratio, denoted lllBj(t)ll / llrBj(t)ll, between the left signal in said frequency band Bj, denoted lBj(t), and the right signal in said frequency band Bj, denoted Bj(t), - for each end pu; p„ of said angular sector P t, calculation of the ratio, called end gain ratio, panL(pu) / panR(pü) and panL(pri) / panR(pri), between the gain provided by the left panning law panL for said end of said angular sector Pi, and the gain provided by the right panning law panR for said end of said angular sector P;; - comparison of the left-right amplitude ratio lllBj(t)ll / llrBj(t)ll, with the calculated end-gain ratios panL(pii) / panR(pii) and panL(pri) / panR(pri); - generation, in said frequency band Bj, of an audio signal, denoted SK,Bj(t), from said left signal lBj(t) in said given frequency band and said right signal rBj(t) in said given frequency band, and according to the result of the comparison of the left right amplitude ratio lllBj(t)ll / llrBj(t)Il, with the end gain ratios panL(pii) / panR(pii) and panL(pri) / panR(pri); the generation step comprising the attenuation of the left signal lBj(t) in said frequency band and of the right signal rBj(t) in said frequency band when the left amplitude ratio right lllBj(t)ll / llrBj(t)It is outside the range [panL(pri) / panR(pri) ; panL(pii) / panR(pii)] defined by said end gain ratios: the attenuation value, denoted Gmasque_Bj (t), applied being a function of the distance between the left right amplitude ratio lllBj(t)ll / llrBj(t)ll and the range defined by said end gain ratios; 3) repetition of step 2) for several other frequency ranges Bj; summation of the SPijBj(t) signals which were generated for said angular sector P; and said frequency bands Bj, in order to obtain a new audio signal Si(t) corresponding to one of the N components of the sound panorama to be realized, repetition of steps 2), 3) and 4) for the other angular sectors P; so as to obtain said N new audio signals Si(t).
2. Method according to claim 1, wherein the frequency band decomposition of the left signal l(t) and the right signal r(t) is carried out with a short-term Fourier transform (stft).
3. A method according to any one of the preceding claims, wherein, in order to smooth out abrupt variations resulting from attenuation, the attenuation value, denoted Gmask Bj(t), is smoothed over time with: - a first-order filter, and / or - a first time constant, named attack, for increasing variations, and / or - a second time constant, named release, for decreasing variations.
4. A method according to any one of the preceding claims, wherein the attenuation Gmask Bj(t) is tuned with a selectivity factor that is defined as a function of the angular sector Pi.
5. A method according to claim 4, wherein the selectivity factor, denoted selectivityi, associated with the given angular sector Pt, is calculated by the formula: selectivityi = max(selectivityL,i ; selectivityR,i) with selectivityL,i = Grej / (threshLi-RdBLji) selectivityR,i = Grej / (threshRi-RdBR>i) with threshLi = 20.1ogi0(panL(pii) / panR(pii)) corresponding to a left threshold value in dB associated with the left end pu of the angular sector Pi, threshRi = 20.1ogi0(panL(pri) / panR(pri)) corresponding to a right threshold value in dB associated with the right end pri of the angular sector Pi, RdBLj = 20.1og10(panL(pi-wi) / panR(pi-wi)) RdBR,i = 20.1og10(panL(pi+wi) / panR(pi+wi)) with Wi= pri- Pu and pi = i.Wi-1 + w / 2 Grej: a minimum rejection value in dB.
6. Data set comprising: - a multichannel audio signal S(t) generated according to any one of claims 1 to 5, and - for each of the N new signals of said multichannel audio signal S(t), characteristics of the angular sector, preferably the center direction and width of said angular sector, of the original soundscape encoded in the stereo file from which said new signal was generated.
7. Computer program product, characterized in that it comprises program code instructions for implementing the steps of a process according to any one of claims 1 to 5 when said program is executed by a computer system, such as a computer.
8. Method of broadcasting N audio signals Si(t) of a multichannel audio signal S(t) generated according to any one of claims 1 to 5, on a plurality of loudspeakers, the method comprising the steps of: - defining a given number NHp of loudspeakers HPk, and, for each of the loudspeakers, defining the position, orientation and directivity of said loudspeaker HPk; - defining the positions of N virtual sound sources SV; around the loudspeakers HPk, each virtual sound source SV iCorresponding to the ith audio channel Si(t) generated from the multichannel audio signal S(t); - determination for each of the speakers HPk and each of the sources SVi, of a gain gkji and a delay ôkji as a function of the distance dk i between the virtual source SV; and said speaker HPk; - provision of said generated audio signals Si(t), to each speaker HPk by applying to said speaker HPk for each audio signal Si (t), the gain gk>i and the delay ôk>i determined for said speaker.
9. A method according to claim 8, wherein the determination, for each loudspeaker HPk and each source SVi, of the gain gk>i and the delay ôk>i to be applied, is also a function of a parameter, called the Rolloff parameter, denoted R; said determination of gain and delay is also carried out, for at least one listening position (Pp), as a function of: - said Rolloff parameter R; - the distance dkp between said loudspeaker HPk and the listening position Pp; and - the angle ak>p between the axis of loudspeaker HPk and the listening position Pp
10. 'p- Method according to claim 9, wherein the determination for each of the loudspeakers HPket each of the sources SVi, of the gain gk>i and the delay ôk>i to be applied, is carried out as a function also of a set of directivity data of the loudspeaker HPk, called directivity diagram, which provides for the angle ak>p, an attenuation value wakp, the attenuation value wakp being used, with also the distance d'k>pet as a function of said Rolloff parameter R, to determine an energy vector rE i>p, as a function of which are calculated an angular error 0i>p and an angular overlap ipi>p.
11. A method according to claim 10, wherein, after calculating a gain gkji and a delay ôkji for a set NP of given listening positions Pp, - a cost function Q(R) is defined for the set of positions as a function of the angular errors 0i>p and the corresponding angular overlaps ipijP, - and the value, denoted Ropt, of the Rolloff parameter R is identified for which the cost function Q(R) is minimal, the gain gkji and the delay ôkji to be applied being calculated by the formulas: gkji=m / (dk>i)a with a = R / (20.1ogi0(2)) and im = । ---- J ôk>i = dk>i / c with c = 340 ms 1
12. A method according to claim 11, wherein the cost function Q(R) is calculated by the formula: Q(R) = a.$10+ with a, [3, y, and ô: weights; q0 = average of the angular errors I0i>pl for the set of sources i at the set of listening positions p; ae= ^pws))' recar'type of the angular error for the set of sources i and the set of listening positions p; p |, ,1a average of the successive overlaps for the set of sources i and the set of listening positions p; V(Z(|V\P|-^) 2 / (Np+Ns+Ns-1 ))' the standard deviation of the successive angular overlaps for the set of sources i and the set of listening positions p; NP being the number of listening positions p and Ns the number of sources i.
13. A method according to any one of claims 8 to 12, wherein the N virtual sound sources SVi are positioned around the speakers HPk on an arc of a circle.
14. Computer program product characterized in that it comprises program code instructions for implementing the steps of a process according to any one of claims 8 to 13 when said program is executed by a computer system, such as a computer.
15. A system for generating a multichannel audio signal, denoted S(t), from a stereo audio signal, the system comprising a stereo audio signal acquisition module; the supplied stereo audio signal comprising a left-hand signal, denoted l(t), which was obtained by mixing of monophonic signals according to a left panning law, denoted panL, and a right signal, denoted r(t), which was obtained by mixing monophonic signals according to a right panning law, denoted panR, in which the system comprises: a cutting module configured to cut the sound panorama (PS) into a given number N of angular sectors, denoted P;, preferably contiguous; a module for calculating and comparing gain ratios, configured for: for a given angular sector P; = [pu; pri] of the sound panorama (PS), and for a given frequency band, denoted Bj, - calculate the value of the amplitude ratio, called the left-right amplitude ratio, denoted lllBj(t)ll / llrBj(t)ll, between the left signal lBj(t) in said frequency band Bj and the right signal rBj(t) in said frequency band Bj, - for each end (pn; pri) of said angular sector P, calculate the ratio, called the end gain ratio panL(pii) / panR(pu) and panL(pri) / panR(pri), between the gain provided by the left panning law panL for said end of said angular sector Pi, and the gain provided by the right panning law panR for said end of said angular sector P;; - compare the left-right amplitude ratio lllBj(t)ll / llrBj(t)ll, with the calculated end-gain ratios panL(pri) / panR(pri) and panL(pii) / panR(pii); and a signal generation module configured for: - generate, in said given frequency band B j, an audio signal SPijBj(t) from said left signal lBj(t) in said given frequency band Bj and said right signal rBj(t) in said given frequency band Bj, and according to the result of the comparison of the left right amplitude ratio lllBj(t)ll / llrBj(t)ll, with the end gain ratios panL(fold) / panR(fold) and panL(pri) / panR(pri); the generation including the attenuation of the left signal lBj(t) in said frequency band Bj and of the right signal rBj(t) in said frequency band B7 when the left right amplitude ratio lllBj (t) ll / llrBj(t)ll is outside the range [panL(pri) / panR(pri) ; panL(pu) / panR(pü)] defined by said end gain ratios, The attenuation value, denoted Gmask_Bj(t), applied is a function of the distance between the left-right amplitude ratio lllBj(t)ll / llrBj(t)Il and the range defined by said end-gain ratios; the system being configured for: - (i) for a given angular sector P;, execute several times the calculation module and gain ratio comparison for several frequency ranges Bj; - (ii) sum the signals SpijBj(t) that were generated for said angular sector P; and said frequency bands Bj, in order to obtain a new audio signal Si(t) corresponding to one of the N components of the soundscape to be created, - (iii) repeat steps (i) and (ii) for the other angular sectors P; so as to obtain said N new audio signals Si(t).
16. System for broadcasting N audio signals Si(t) of a multichannel audio signal S(t) generated according to any one of claims 1 to 5, on a plurality of loudspeakers, the system comprising: - a user interface allowing the position and orientation of each of said loudspeakers HPk to be defined for a number Nhp of loudspeakers HPk; each loudspeaker HPk being associated with directivity data; - a module for defining the positions of N virtual sound sources SV; around the speakers HPk, each virtual sound source SVi corresponding to the i-th audio channel Si(t) generated from the multichannel audio signal S(t); - a determination module for each of the HPk speakers and each of the SVi sources, of a gain gkji and a delay ôkji as a function of the distance dk>i between the virtual source SV; and said HPk speaker; - a transmission module configured to transmit said generated audio signals Si(t), to each speaker HPk by applying to said speaker HPk for each audio signal Si(t), the gain gk>i and the delay ôkji determined for said speaker.
Citation Information
Patent Citations
Method for generating a surround audio signal from a mono / stereo audio signal
EP2530956A1
Sound reproduction system having a matrix converter
US5594800A
Spatial disassembly processor
US20100086136A1
Segment-wise adjustment of spatial audio signal to different playback loudspeaker setup
US20150248891A1
Ambience generation for stereo signals
US7567845B1