Crosstalk cancellation using a virtual sound field
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SCUOLA UNIVRIA PROFESSIONALE DELLA SVIZZERA ITAL (SUPSI)
- Filing Date
- 2025-01-31
- Publication Date
- 2026-08-06
Smart Images

Figure EP2025052582_06082026_PF_FP_ABST
Abstract
Description
[0001] " METHOD FOR PERFORMING AND OPTIMIZING CROSSTALK CANCELLATION BY USING A VIRTUAL SOUND FIELD"
[0002] D E S C R I P T I O N
[0003] The present invention relates generally to a method for performing and optimizing crosstalk cancellation by using a virtual sound field.
[0004] Definitions
[0005] In the context of the present description, the expression "crosstalk cancellation" (in acronym, XTC) denotes, as it is known, a technique for performing a full three-dimensional (3D) audio reproduction from two (or more) speakers.
[0006] Further, the expression "sound source" denotes an object that has both an associated audio signal (i. e., an audio signal or a sound that the "sound source" is supposed to emit) and a given location in a 3D space. In case of moving sound sources, the given location in the 3D space may vary over time.
[0007] It is also known that the expression "binaural signals" denotes audio signals that - when emitted by the sound source - are received separately at a left and a right ear of a listener. In other words, the binaural signals reflect how the listener' s ears perceive the 3D audio signal emitted by the sound source.In practice, the "binaural signals" depend not only- on which audio signal is emitted by the sound source, but also on the sound source' s location and the listener' s position and orientation.
[0008] It is also known that the word "binauralizer" denotes an algorithm (or a device configured to carry out such algorithm) which is adapted to operate signal processing (i. e., a sort of signal filtering) in order to enable a computation of the "binaural signals". Moreover, the term "binauralization" and the verb "to binauralize" denote the action of processing one or more audio signals by using at least one binauralizer.
[0009] In practice, a "binauralization" is a process of computing which audio signals should be received at the listener' s ears based on the sound source' s location, as well as on the listener' s position and / or orientation. Accordingly, a "binauralization" process computes the binaural signals starting from:
[0010] - an audio signal emitted by the sound source, - the sound source' s location,
[0011] - the listener' s position, and
[0012] - the listener' s orientation.
[0013] Furthermore, the expression "physical speakers" denotes real speakers, i. e., tangible devices which areconfigured to emit sound within a listening environment such as the 3D space. Typically, the "physical speakers" are disposed so that they surround a zone where the listener is allowed to be present and to move.
[0014] In practice, the "physical speakers" are the sole physical devices which can effectively emit sound within the listening environment. When the "physical speakers" emit sound, a real sound field is obtained.
[0015] By contrast, the expression "virtual speakers" denotes virtual objects that exist only within a virtual sound field (i. e., within a virtual environment such as a computer program).
[0016] Background of the Invention
[0017] As it is known, any XTC process has the goal to enable the listener to perceive a given sound source at a correct location in the 3D space, whatever the listener' s position and orientation. In this framework, a first underlying principle of XTC resides in that any human listener can perceive 3D audio from two ears, namely, a right ear and a left ear. For example, when a plurality of physical speakers emits (i. e. reproduces) sound signals for a single listener, sound emitted from any physical speaker can reach both the left and the right ears.As it is also known, a second underlying principle of XTC resides in that, by controlling signals at the listener' s ears, full 3D audio can be rendered. In this framework, XTC exploits projecting binaural signals to the listener' s ears and cancelling one (or more) "crosstalk" effect (s) between the physical speakers and the listener' s ears in order to project a binaural signal to the listener' s ears. For example, when the plurality of physical speakers includes a left speaker and a right speaker, XTC utilizes additional signals for "killing" any crosstalk contribution from the left speaker to the right ear and / or from the right speaker to the left ear. As a result, the left ear will hear only a first signal of the binaural signals, and the right ear will hear only a second signal of the binaural signals. In Figures 1 and 2, exemplary crosstalk contributions are denoted by dashed lines.
[0018] In the context of the present description, those "killing" signals are called "cancellation signals", whereas the original left and right speaker signals (i. e., the signals that need to be reproduced at the listener' s ears) are called "direct signals".
[0019] Typically, the known standard XTC processes utilize a setup including:- two or more direct signals physical speakers, which are configured to emit respective binaural signals, and
[0020] - two or more cancellation signals physical speakers, which are configured to emit respective cancellation signals (i. e., signals adapted to "kill" one or more cross-talk effects between the physical speakers and the listener' s ears).
[0021] The direct signals physical speakers and the cancellation signals physical speakers can be colocalized provided that some constraints are fulfilled. Properly, "cancellation" and "direct" are more designation of roles of physical objects than physical objects themselves. Accordingly, one and a same real speaker may play both the role of a direct signals physical speaker and the role of a cancellation signals physical speaker.
[0022] Usually, both the direct signals physical speakers and the cancellation signals physical speakers are fixed speakers (i. e., speakers which are not movable).
[0023] Therefore, when a data processing apparatus or an equivalent apparatus comprises means for carrying out an XTC process, the XTC process provides a way to control which audio signal is received at the left ear of thelistener, and which audio signal is received at the right ear of the listener. If such audio signals mimic binaural signals which are received from a source located at any point in space, the listener will have the impression to hear that source originating from its ef fective location.
[0024] From another perspective, when a system comprising the physical speakers is analyzed in the frequency domain (i. e., in the so-called "z-domain" ), the signals E ( z ) which are received at the listener' s ears can be expressed as the product of a forward path matrix H ( z ) and an input signals vector S ( z ), based on following relation R.0:
[0025] E(z) = H(z)·S(z) (R.0).
[0026] In particular, S ( z ) can be expressed as a function of L ( z ) and R ( z ), wherein:
[0027] - L ( z ) represents the signal intended for the left ear, and
[0028] - R ( z ) represents the signal intended for the right ear.
[0029] When a XTC filter C ( z ) ( i. e., a filter adapted to perform XTC) is disposed at the input of the system, relation R.0 becomes following relation R.1:
[0030] E(z) = H(z)·C(z)·S(z) (R.1).As relation R. l suggests, a perfect cancellation can be achieved if:
[0031] E(z) = k·z-d·S(z) (R.2),
[0032] - k is a gain factor, and
[0033] - z-dis a pure delay.
[0034] In other words, XTC is achieved when the filter C (z) approximates the inverse of the forward path matrix H (z), up to the gain factor k, combined with a delay for causality reasons.
[0035] However, inversion of the forward path matrix H (z) is a non-trivial problem, thereby representing a key issue for achieving proper XTC performance.
[0036] Additionally, if the listener moves, the geometry of the system will change. In fact, since the direct signals physical speakers and the cancellation signals physical speakers are fixed speakers, any listener' s movement will change a relative positioning between the same listener and the fixed speakers. As a result, XTC performance will degrade, and the cancellation signals will have to be modified for maintaining the proper XTC performance.
[0037] It is also known that XTC is now appearing in products from both big corporations and start-ups suchas, for example, Apple™, Samsung™ / Harman™, Silentium™, and Theoretica™.
[0038] However, most current XTC solutions suffer from limitations in terms of supported user positions and / or supported movements. For instance, although Apple™ has introduced 3D audio rendering for Dolby Atmos™ content on iPhones™, iPads™ and selected Mac™ models, supported user movements are rather limited in such a 3D audio rendering.
[0039] Therefore, new XTC approaches are needed which could overcome (or at least mitigate) one or more of the problems and limitations discussed above.
[0040] Summary of the Invention
[0041] The aim of the present invention is to improve the background art in one or more of the aspects indicated above.
[0042] Within the scope of this aim, an object of the present invention is to maximize XTC performance.
[0043] In particular, an object of the present invention is to minimize crosstalk effects between speakers and the listener' s ears.
[0044] Moreover, an object of the present invention is to provide optimal rendering of binaural signals at the listener' s ears.Another object of the present invention is to provide the listener with an enhanced 3D audio immersion experience.
[0045] A further object of the present invention is to provide a panning process which is capable of mapping signals from a virtual sound field to the real sound field while maintaining the benefits provided by executing an XTC process in the virtual sound field.
[0046] In particular, the present invention aims at providing a panning process that is capable of minimizing coloration at a target ear of the listener and, at the same time, maximizing cancellation of signals at a nontarget ear of the listener.
[0047] Another object of the present invention is to maximize quality of resulting 3D audio perception at the listener' s ears irrespective of listener' s position and orientation.
[0048] A still further object of the present invention is to provide a method for performing XTC which is highly versatile, reliable, and relatively easy to implement and at competitive costs when compared to prior art solutions.
[0049] This aim, as well as these and other objects that will become better apparent hereinafter, are achieved bya method according to claim 1. Preferred embodiments are defined by the dependent claims.
[0050] This aim and these objects are also achieved by the subject-matter of any of the appended claims.
[0051] Brief Description of the Drawings
[0052] The foregoing, as well as further characteristics and advantages of the present invention, will become better apparent to those skilled in the art from the following description of a preferred, but not exclusive, embodiment of the method according to the invention, illustrated by way of nonlimiting example in the accompanying drawings, wherein:
[0053] Figure 1 illustrates the principle of XTC for a static user;
[0054] Figure 2 illustrates the principle of XTC for a moving user;
[0055] Figure 3 illustrates a signal chain in a preferred embodiment of the method according to the invention;
[0056] Figure 4A illustrates a general principle of an approach to XTC, which can be applied by the method according to the invention when using two binauralization processes (front and back) and - for each of these two binauralization processes - two virtual speakers for direct signals (i. e., two "direct signalsvirtual speakers") and two virtual speakers for cancellation signals (i. e., two "cancellation signals virtual speakers");
[0057] Figure 4B illustrates a possible arrangement of physical speakers within a real sound field in which a listener is present;
[0058] Figure 4C illustrates a sound source emitting sound within a listening environment including the listener;
[0059] Figure 5 provides an overview of an exemplary binauralization process for use in the method according to the invention;
[0060] Figures 6A and 6B illustrate a preferred localization of cancellation signals virtual speakers in a situation considering two binauralization zones (front and back) and two cancellation signals virtual speakers for each binauralization zone;
[0061] Figures 7A, 7B and 7C illustrate preferred positions of virtual speakers for an XTC process using two direct signals virtual speakers and two cancellation signals virtual speakers; and
[0062] Figure 8 illustrates a preferred setup for implementing a panning process for the method according to the invention.
[0063] Detailed Description of the InventionThe following detailed description and appended figures describe and illustrate an exemplary embodiment of the invention. The description and figures serve to enable one skilled in the art to make and use the invention, and are not intended to limit the present invention, and its applications or uses. It should also be understood that throughout the figures, corresponding reference signs, numerals or symbols indicate like or corresponding parts and features.
[0064] With reference to the cited figures, each method according to the present invention comprises the following steps:
[0065] - inputting at least one input signal to an XTC module 20;
[0066] - executing an XTC process on the at least one input signal, so as to obtain at least one processed signal;
[0067] - panning the at least one processed signal, so as to obtain at least one panned signal; and - outputting the at least one panned signal to two or more physical speakers (which are configured to emit sound for at least one listener LI). In each embodiment, the XTC process is executed in a virtual speakers' domain, by using two or more virtualspeakers which are included in the XTC module 20.
[0068] In Figure 3, the virtual speakers' domain can be represented by any of reference numerals 22 and 24. As will become more apparent from the discussion further below, performing the XTC process in the virtual speakers' domain entails the determination of optimized positioning and number of virtual speakers. As a result, XTC performance can be maximized.
[0069] In particular, each virtual speaker (which may be utilized both for direct signals and for cancellation signals) is arranged according to an optimized placement with respect to position and orientation of at least one listener LI (and, optionally, with further respect to other system constrains), and the XTC process is carried out by each virtual speaker (or, in optimal embodiments, by the combination of all virtual speakers).
[0070] In some preferred embodiments, the XTC process is carried out by the combination of all virtual speakers.
[0071] In Figure 3, the number of virtual speakers, as well as signals and positions for each virtual speaker are defined within the XTC module 20.
[0072] The number of virtual speakers within the XTC module 20 may appropriately change depending on usage conditions, so that the arrangement shown in Fig. 3 shallnot be considered limiting for all the embodiments. In practice, the XTC module 20 comprises at least two virtual speakers, so the number of virtual speakers may vary between two and M (wherein M is an integer number greater than two).
[0073] In some preferred embodiments, the optimized arrangement of each virtual speaker can take account not only of actual position and orientation of the listener LI, but also of physical constraints of the system such as, for example, one or more walls of the 3D space or listening environment (so as to take account of, e. g., compensation of sound reflection of any wall of the listening environment).
[0074] In particular, the XTC process (which occurs in the virtual speakers' domain) can include a real-time optimization of positions and / or number of the virtual speakers. Optionally, the real-time optimization can further involve an optimization of orientations of the two or more virtual speakers.
[0075] In practice, the position and / or the number of the virtual speakers can be optimized dynamically in realtime, thereby leading to optimized XTC conditions which are always optimal irrespective of any user' s position and orientation.In simpler embodiments, the XTC is carried out by a static XTC filter. In other words, in the simpler embodiments, the optimization takes account of a static user such as, for example, the listener LI shown in Fig. 1. Hence, in these simpler embodiments, there is no change in the system over time (because the user is static), so that the optimization is performed once for all.
[0076] In more sophisticated embodiments, the XTC is carried out by a dynamic XTC filter, i. e., by a filter having dynamic adaptation. In other words, in the more sophisticated embodiments, the real-time optimization can be adaptive, so as to take account of a moving user such as, for example, the listener LI shown in Fig. 2.
[0077] In practice, the dynamic filter is configured to operate as a first filter Cl (z) when the listener LI is in a first user position Pl, and to operate as a second filter C2 (z) when the listener LI is in a second user position P2 (cf. Fig. 2). In other words, the dynamic filter is configured so that, when the listener LI is in the first user position Pl, the signals received at the listener' s ears El (z) can be expressed by means of following relation R.3:
[0078] El(z) = Hl(z)·Cl(z)·S(z) (R.3).Conversely, when the listener LI is in the second user position P2, the signals received at the listener' s ears E2 (z) can be expressed by means of following relation R.4:
[0079] E2(z) = H2(z)·C2(z)·S(z) (R.4).
[0080] In practice, the method according to the present invention executes the XTC process in the virtual speakers' domain, where optimized XTC conditions are achievable. In fact, since virtual speakers can be placed anywhere in space, they can track the listener' s position and orientation, thereby guaranteeing optimal XTC performance even in case of any listener' s movement. Subsequently, in the panning step, the method maps the virtual situation back to the physical speakers which are present in the system (for example, to the physical speakers which are disposed within the 3D space / listening environment).
[0081] Further aspects concerning preferred embodiments of the present invention are discussed in subsections 1 to 4 below.
[0082] 1 – Binauralization
[0083] The present subsection 1 relates to binauralization. Accordingly, the aspects discussed in subsection 1 mainly relate to a binauralization modulethat, in Figure 3, is represented by a block denoted by reference numeral 10.
[0084] The method according to the present invention can utilize any kind of binauralizat ion or similar / compatible downmixing technique.
[0085] In practice, the presented approach to XTC renders the method agnostic to (i. e., independent from) any specific binauralization process used. In other words, the method according to the present invention can work with any kind of binauralization (or downmixing) process.
[0086] In practice, some embodiments may utilize a single binauralization, whereas other embodiments may utilize multiple binauralizations (for example, multiple binauralizations for different spatial zones around the listener). Accordingly, in certain embodiments, the audio signal may be processed by a single binauralizer, whereas in other embodiments the audio signal may processed by multiple binauralizers.
[0087] Nevertheless, in certain circumstances (e. g., when the input signal is a stereo signal), the use of one or more binauralizers is not necessary. Hence, in some embodiments, a step of binauralization is entirely optional.If a single binauralizer is used, the single binauralizer must be able to cover all spatial regions where a sound source may be located. For example, if the single binauralizer is front only, the single binauralizer will only support any sound source located in front of the listener LI.
[0088] By contrast, in other embodiments entailing a multiple-binaural approach, the audio signal may be processed by multiple binauralizers (for example, by two or more binauralizers). In other words, the binauralization can include considering multiple spatial regions.
[0089] In those embodiments entailing the multiple-binaural approach, each binauralizer is preferably optimized for a different zone around the listener. Moreover, multiple binauralizers may be utilized simultaneously.
[0090] In particular, the multiple-binaural approach entails:
[0091] - selecting two or more specific bianuralizers among the multiple binauralizers, and
[0092] - performing the binauralization based on the selected (specific) bianuralizers.
[0093] In practice the selection of the two or morespecific bianuralizers is based on a sound source location.
[0094] In particular, each specific bianuralizer is selected so as to cover a zone where a specific sound source is placed. For example, the selected bianuralizers can include:
[0095] - a first binauralizer, configured to binauralize one or more audio signals of any sound source located in front of the listener; and
[0096] - a second binauralizer, configured to binauralize one or more audio signals of any sound source located behind the listener.
[0097] When the multiple-binaural approach (with separated binauralization for different zones around the listener) is used, real advantages in terms of resulting 3D audio perception are achieved, and a 3D audio immersion effect for the listener is maximized.
[0098] In practice, when a multiple-binauralization approach is used, sound sources in different zones around the listener can be binauralized separately.
[0099] In particular, performing the binauralization can include:
[0100] - a first binauralization process, for any sound source located in front of the listener; and- a second binauralization process, for any sound source located at the back of the listener.
[0101] In practice, when there is a plurality of sound sources, different binauralization processes may be applied depending on the location of the various sound sources relative to the listener LI.
[0102] In particular, the different binauralization processes can utilize different binauralizers, respectively. For example, a front binauralizer can be used for any sound source located in the front of the listener LI. By contrast, a back binauralizer can be used for any sound source located behind the listener LI.
[0103] In preferred embodiments, the audio signal of each sound source is input to the at least one binauralizer to compute the corresponding 3D audio signals to be received at the listener' s ears (cf., e. g., Fig. 3: "3D audio").
[0104] Optionally, the binauralization can be considered with respect to the actual position and orientation of a listener' s head (i. e., with respect to the actual position and orientation of the head of the listener LI).
[0105] In practice, each sound source is located at arespective location with respect to a global orthogonal coordinates system. Nevertheless, for performing the binauralization, the location of each sound source shall be expressed with respect to a local orthogonal coordinates system attached to the listener.
[0106] In some preferred and illustrated embodiments, the global orthogonal coordinates system is denoted by axes x, y and z ( i. e., by an x-axis, an y-axis and a z-axis ). By contrast, the local orthogonal coordinates system is denoted by axes x', y', z ' ( i. e., by an x' -axis, an y' -axis and a z ' -axis ).
[0107] In particular, the local orthogonal coordinates system (x', y', z ' ) is attached to the head of the listener LI ( cf., for example, any of Figs. 4A to 4C). Hence, the local orthogonal coordinates system (x', y', z ' ) can represent ( i. e., be indicative of ) a reference system centered at the listener' head.
[0108] Preferably, the local orthogonal coordinates system (x', y', z ' ) has its origin ( 0', 0', 0' ) at the center of the listener's head.
[0109] In practice, when the listener LI is disposed at a position having coordinates (xl, yl, zl ), and at an orientation defined by a pitch angle (pil ), a yaw angle (yal ) and a roll angle ( rol ) which are defined withrespect to the global orthogonal (or orthonormal ) coordinates system, the position / orientation of each sound source shall be expressed in the local orthogonal coordinates system (x', y', z ' ) attached to the listener LI.
[0110] In the preferred and illustrated embodiments, the x' -axis points in front of the listener' s head, the z' -axis points towards the top of the listener' s head, and plane x' y' ( i. e. the plane defined by the x' -axis and the y' -axis ) passes through the listener' s ears.
[0111] Additionally, when:
[0112] - the position (xl, yl, zl ) of the listener LI is considered as the result of a translation T from the origin ( 0, 0, 0 ) of the global orthogonal coordinates system (x, y, z ), and
[0113] - the orientation (pil, yal, rol ) of the listener LI is considered as the result of a rotation R from the ( 0, 0, 0 ) angling,
[0114] the coordinates S = (xs, ys, zs ) of a given sound source SS in the global orthogonal coordinates system (x, y, z ) can be expressed in the local orthogonal coordinates system (x', y', z ' ), as:
[0115] S' = R-1• T-1• S (R. 5 ), and the binauralization is applied considering the soundsource S' expressed in the local orthogonal coordinates system (x', y', z' ).
[0116] In Fig. 4C, reference sign A denotes the audio signal associated to the sound source SS having the coordinates (xs, ys, zs). Moreover, the reference signs R and L denote binaural signals that are received at the listener' s right and left ears, respectively, when the sound source emits the audio signal A.
[0117] In the system, there is also a number N of physical speakers Si (with i = 1,..., N) which are located at coordinates (xl, yl, zl),..., (xN, yN, zN), respectively. In other words, each physical speaker Si is located at respective coordinates (xi, yi, zi) with respect to the global orthogonal coordinates system (x, y, z) (cf., e. g., Figs. 4A-4C).
[0118] In Fig. 4A, four physical speakers S1-S4 are shown. Thus, in Fig. 4A, the number N is equal to four. Nevertheless, the number N of physical speakers may be appropriately varied, e. g. based on physical constraints of the 3D space or listening environment in which the physical speakers may be located). For example, when the 3D space or listening environment is in the shape of a parallelepiped, a setup with N=8 entails the presence of eight physical speakers S1-S8, and each physical speakermay be disposed at a respective corner of the parallelepiped (cf., e. g., Fig. 4B).
[0119] In practice, in a given system, the number N of physical speakers is fixed, and does not change over time. By contrast, as discussed in greater detail below, the number of virtual speakers may dynamically vary.
[0120] The main steps of the binauralization process for a given sound source S are illustrated in Fig. 5.
[0121] In particular, processing the audio signal by using the at least one binauralizer includes a complete generation of binaural signals.
[0122] In some preferred embodiments, the complete generation of binaural signals includes one or more (and, more preferably, all) of following major elements El to E3:
[0123] El) Simulation of source Direction of Arrival (DOA): this is typically implemented using Head-Related Transfer Functions (HRTFs) corresponding to the effective DOA of the source for the listener' s left and right ears. For any given DOA, a pair of HRTFs corresponding to the user' s left and right ears are required. These HRTFs can be typically provided as impulse responses with 256 to 4096 taps at 48kHz. As such, they can be implemented as simple FIR filters.E2) Source distance simulation (i. e., simulation of a source distance in terms of gain / delay, possibly completed by low-pass characteristics): HRTFs are generally considered for sources located at a given distance (for instance, 3 meters) of the center of the listener's head. To take the effective distance of the sound source into consideration, following processing E2.1, E2.2 can be applied.
[0124] - E2.1) Modeling an amplitude variation with source distance: a point source radiates with a 1 / d pattern, where d is the distance to the source. By comparing the desired source distance with the nominal distance of the HRTFs, a corresponding attenuation / gain factor can be defined as
[0125] g = r / d (R. 6),
[0126] wherein d is the source distance and r is the nominal distance used in the definition of the HRTFs.
[0127] - E2.2) Applying a delay correction to mimic propagation time from source to said listener LI: considering that a source at distance r (nominal HRTF distance) would arrive at time t at the listener' s ear, the sound from a source in the same direction at distance d would arrive at listener' s ear at a time t' computed as:t' = t + (d-r) / c (R.7),
[0128] where c is the speed of sound.
[0129] E3) Simulation of ambient characteristics: early reflections and diffuse reverberation are key elements for the natural rendering of an audio source.
[0130] The simulation of the ambient characteristics can be carried out in two different ways W1 or W2:
[0131] Wl) Convolution of signals with effective impulse responses corresponding to a given listening environment. Such impulse responses can be acquired in a physical environment having the desired characteristics.
[0132] W2) Simulation of early reflections and diffuse reverberation. Early reflections can be simulated using a multi-tap delay line, with different weights / delays corresponding to the different reflections with the addition of optional low-pass filters to simulate decaying frequency content in the reflections. Optional HRTF filters may be added on each tap to simulate the DOA of early reflections. Diffuse reverberation is typically implemented using Feedback Delay Networks (FDN) where all-pass processed, gain-reduced signals are feedback with different delays.
[0133] In some preferred embodiments, the binauralizationprocess includes the simulation of ambient characteristics. In other preferred embodiments, by contrast, this kind of simulation is not considered by the binauralization process.
[0134] If the simulation of ambient characteristics is included in the binauralization process, one must take care that any interaction between the ambient characteristics embedded in the binauralized process does not interfere with local reproduction and listening ambient characteristics in a way to negatively affect audio rendering. A key element to achieve this is to equalize local room' s early reflections, which could be achieved by cancellation signals dedicated to the cancelling of early reflections. In these latter circumstances, cancellation signals would be integrated directly within an XTC algorithm underlying the XTC process.
[0135] By contrast, if the simulation of ambient characteristics is not included in the binauralization process, it is assumed that the local ambient characteristics (i. e., the ambient characteristics of the 3D space / listening environment) will provide the desired ambient characteristics.
[0136] In some preferred embodiments, the sound sourcesinclude directional sources and ambient sources, and the binauralization is only applied to the directional sources.
[0137]
[0138] The present subsection 2 relates to execution of the XTC process in the virtual speakers' domain. Accordingly, the aspects discussed in subsection 2 mainly relate to the XTC module 20. Nevertheless, a reader will clearly understand that some of the aspects discussed in subsection 2 may further relate to the module concerning binauralization.
[0139] In preferred embodiments, the positions of the virtual speakers are optimized with respect to the position (xl, yl, zl) and the orientation (pil, yal, rol) of the listener LI (and, more precisely, an upper body part including the listener' s head, in particular the listener' s ears).
[0140] In particular, as will be discussed in greater detail below, the two or more virtual speakers can be located at respective optimized positions with respect to the local orthogonal coordinates system (x', y', z' ). For example, the optimized position of a first virtual speaker with respect to the local orthogonal coordinates system (x', y', z' ) can be denoted as (x'vsl, y'vslz ' vsl ), whereas the optimized position of a second virtual speaker with respect to the local orthogonal coordinates system (x', y', z ' ) can be denoted as (x' vs2, y' vs2, z ' vs2 ). In more general words, the optimized position of a j-th virtual speaker with respect to the local orthogonal coordinates system (x', y', z ' ) can be denoted as (x' vs j, y' vs j, z ' vs j ), wherein j is a positive integer comprised between 1 and M ( i. e., j = 1, M).
[0141] Optionally, the positions of the virtual speakers can be further optimized with respect to the physical constraints of the listening environment.
[0142] In practice, computation of the signals to be emitted by the virtual speakers can be optimized based on both the position and the orientation of the listener LI (and, optionally, based on the physical constraints of the listening environment, thereby including the position of each physical speaker and of the walls of the listening environment ). Hence, the listener LI can always be considered according to optimized position and orientation for proper XTC operation for one or more binauralization zones which are taken into consideration.
[0143] In Figs. 6A-6B, a system 50 includes two binauralization zones ( in particular a frontbinauralization zone and a "back" binauralization zone), each with two cancellation signals virtual speakers (i. e., two virtual speakers adapted to reproduce cancellation signals). Nevertheless, in some embodiments, the number and arrangement of the cancellation signals virtual speakers may differ from those of the system 50 shown on Figs. 6A-6B.
[0144] Further, although Figs. 6A-6B illustrate an arrangement in which cancellation signals virtual speakers are shown, the same or similar arrangement may be utilized for positioning direct signals virtual speakers (i. e. virtual speakers adapted to reproduce the direct signals).
[0145] In practice, the XTC process enables locating both the direct signals virtual speakers and the cancellation signals virtual speakers according to respective optimized coordinates, and computation of such optimized coordinates takes account of the XTC (and, optionally, of the physical constraints of the 3D space / listening environment).
[0146] In particular, when two direct signals virtual speakers dsl, ds2 and two cancellation signals virtual speakers csl, cs2 are considered (cf., for example, Figs.
[0147] 7A-7B), the following coordinates (expressed withrespect to the local orthogonal coordinates system (x', y', z ' ) ) can be obtained:
[0148] - (x' dsl, y' dsl, z ' dsl ) for a first direct signals virtual speaker dsl;
[0149] - (x' ds2, y' ds2, z ' ds2 ) for a second direct signals virtual speaker ds2;
[0150] - (x' csl, y' csl, z ' csl ) for a first cancellation signals virtual speaker csl; and
[0151] - (x' cs2, y' cs2, z ' cs2 ) for a second cancellation signals virtual speaker cs2.
[0152] Additionally or alternatively, the positioning of the direct signals virtual speakers dsl, ds2 and of the cancellation signals virtual speakers csl, cs2 can be expressed in polar or spherical coordinates ( i. e., by determining, for each speaker, a distance D and at least one angle A at which the same speaker is arranged with respect to predefined reference directions ). For example, in Figs. 6A and 7A, when each virtual speaker lies within the plane x' y':
[0153] - " Des" denotes the distance of each cancellation signals virtual speaker from a first reference direction corresponding to the y' -axis; and - " Acs" denotes an angle at which a specific cancellation signals virtual speaker is arrangedwith respect to a second reference direction corresponding to the x' -axis.
[0154] Moreover, when at least one virtual speaker lies outside the plane x' y', determining an elevation angle (not shown) is also necessary to express the positioning of this latter at least one virtual speaker with respect to the same plane x' y'.
[0155] In Figs. 6A and 7A, the x' -axis is also a symmetry axis.
[0156] Additionally, in Fig. 7A:
[0157] - " Dds" denotes a distance of any direct signals virtual speaker with respect to the first reference direction ( i. e., with respect to the y ' -axis ); and
[0158] - " Ads" denotes an angle at which a specific direct signals virtual speaker is arranged with respect to the second reference direction ( i. e., with respect to the x' -axis ).
[0159] In particular, " Ads" applies for any direct signals virtual speaker arranged symmetrically with respect to the x' -axis ( of., e. g., Fig. 7A: " Direct signals speakers" dsl and ds2 ).
[0160] In practice, by using virtual speakers, it can be ensured that the "direct signals virtual speakers" ( i. e.those virtual speakers which are configured to emit the direct signals) are optimally placed for best rendering of direct binaural signals at the listener' s ears (if necessary, taking account of any physical realities of the system). Moreover, the positioning of the "cancellation signals virtual speakers" (i. e., those virtual speakers which are configured to emit the cancellation signals) can be optimized for both a most efficient cancellation of signals at a Non-Target Ear of the listener LI and a least coloration on a Target Ear of the listener LI (if necessary, also by taking account of any physical realities of the system such as, for example, the physical constraints discussed above), wherein:
[0161] - the Non-Target Ear (in acronym, NTE) is the ear that is meant to receive a crosstalk contribution to be cancelled; and
[0162] - the Target Ear (in acronym, TE) is the ear that is meant to be reached by a desired direct signal.
[0163] Therefore, preferred embodiments of the method according to the present invention can provide an optimized placement for both the direct signals virtual speakers and the cancellation signals virtual speakers(and, optionally, also considering system constraints such as physical speakers locations).
[0164] As the direct signals virtual speakers are supposed to virtually emit binaural signals, in each embodiment there are at least two direct signals virtual speakers.
[0165] Where necessary, the number of cancellation signals virtual speakers and / or the number of direct signals virtual speakers may be appropriately varied based on an expected position / orientation of the listener LI with respect to the physical speakers (for example, by taking account of possible changes of position / orientation by the listener LI between the first user position Pl and the second user position P2). Optionally, the positioning of both the listener LI and each physical speaker with respect to the global orthogonal coordinates system (x, y, z) can be taken into account.
[0166] The possibility of varying the number of direct signals virtual speakers and / or cancellation signals virtual speakers may provide advantages when panning them to the physical speakers present in the system.
[0167] Therefore, unlike than existing solutions which provide XTC from physical speakers with fixed positions and count, the method according to the present invention provides optimized conditions for XTC in the virtualspeakers' domain. In fact, the existing solutions using fixed positions speakers for XTC result in sub-optimal XTC performance depending on the position / orientation of the speakers and the user (i. e, of the listener).
[0168] In some preferred embodiments, effective positions of any physical speaker in the listening environment may further be used, to guide selection of placement and number of virtual speakers depending on XTC conditions.
[0169] In any case, the approach to XTC proposed by the present invention is quite flexible with respect to physical speakers' layout.
[0170] Since human ears are arranged in a (substantially) symmetric manner, an overall symmetry of virtual speakers' placement with respect to a user' s line of sight is preferred. In fact, XTC generally works best if there is overall symmetry of speakers' placement with respect to the user' s line of sight (as the ears are symmetrically located on the head of the listener). In practice, in a preferred setup having two direct signals virtual speakers dsl, ds2 and two cancellation signals virtual speakers csl, cs2 (cf., e. g., Figs. 7A-7C):
[0171] - the direct signals virtual speakers dsl and ds2 may be located symmetrically Dds meters in frontof the listener LI;
[0172] - the direct signals virtual speakers dsl and ds2 may be angled Ads degrees with respect to a line connecting a center of the listener' s head to a middle point of a line connecting one of the same direct signals virtual speakers dsl, ds2 to the other direct signals virtual speaker ds2, dsl; - the cancellation signals virtual speakers csl, cs2 may be located symmetrically Des meters in front of the listener LI;
[0173] - the cancellation signals virtual speakers csl, cs2 may be angled Acs degrees with respect to the line connecting the center of the listener' s head to a middle point of a line connecting one of the same cancellation signals virtual speakers csl, cs2 to the other cancellation signals virtual speaker cs2, csl.
[0174] In preferred embodiments, if the multiple-binaural approach is used which separates sound sources into different zones (or regions) to be binauralized individually, separate XTC signals and virtual speakers (for direct signals and cancellation signals) can be used for the different zones. In other words, in the multiple-binaural approach, the XTC process can includemultiple XTC instances for the multiple spatial regions, respectively (i. e., specific XTC instances for respective different zones).
[0175] In any case, the separate XTC signals may however share (part of) the virtual speakers.
[0176] Where necessary, an effective number of virtual speakers for virtual direct and virtual cancellation signals may be varied. This allows for optimized rendering, also considering final locations of the physical speakers and other physical constraints of the 3D space / listening environment (for instance, by considering that, when a certain virtual speaker is placed close - but not superposed - to a given physical speaker but far from other physical speakers, performance of any panning system will degrade). Hence, in some cases, it may be of advantage to use more than two direct signals virtual speakers or more than two cancellation signals virtual speakers (for any of the possibly multiple binauralization processes).
[0177] In preferred embodiments in which the at least one input signal relates to sound emitted by at least one sound source located at its respective coordinates (xs, ys, zs) with respect to the global orthogonal coordinates system (x, y, z), and the two or more virtualspeakers include at least two direct signals virtual speakers and at least two cancellation signals virtual speakers (and wherein, optionally, the at least two cancellation signals virtual speakers are colocalized with the at least two direct signals virtual speakers ), the XTC process can further comprise:
[0178] - expressing the location of each sound source with respect to the local orthogonal coordinates system (x', y', z ' ) ( i. e., considering the system from the user' s location and orientation point of view); and
[0179] - computing optimized positions and / or optimized orientations for each direct signals virtual speaker and for each cancellation signals virtual speaker, based on:
[0180] ■ the location of each sound source, expressed with respect to the local orthogonal coordinates system (x', y', z ' ), and ■ actual positioning and actual orientation of the local orthogonal coordinates system (x', y', z ' ) with respect to the global orthogonal coordinates system (x, y, z ).
[0181] In this framework, the XTC process may further comprise:- computing direct signals for the direct signals virtual speakers, based on the optimized positions (and, optionally, on the optimized orientations) that are computed for each direct signals virtual speaker; and
[0182] - computing cancellation signals for the cancellation signals virtual speakers, based on the optimized positions (and, optionally, on the optimized orientations) that are computed for each cancellation signals virtual speaker.
[0183] In practice, since the actual positioning and orientation of the local orthogonal coordinates system (x', y', z' ) with respect to the global orthogonal coordinates system (x, y, z) can represent the actual positioning and orientation of the head of the listener LI with respect to the global orthogonal coordinates system (x, y, z ):
[0184] - expressing the location of each sound source with respect to the local orthogonal coordinates system (x', y', z' ) means expressing the location of each sound source with respect to the listener' s head, and
[0185] - the actual positioning and orientation of the local orthogonal coordinates system (x', y', z' )with respect to the global orthogonal coordinates system (x, y, z) corresponds to the actual positioning and orientation of the listener' s head with respect to the global orthogonal coordinates system (x, y, z).
[0186] Expressing each sound source location with respect to the local orthogonal coordinates system (x', y', z' ) can be achieved by combining a translation T and a rotation R of the 3D space, in particular by using the above-outlined relation R.5.
[0187] Computing the optimized positions for each direct signals virtual speaker and / or for each cancellation signals virtual speaker can also take account of system constraints such as, for example, the physical constraints discussed above.
[0188] In some preferred embodiments, the binauralization process (i. e., the processing of the audio signal by using the at least one binauralizer ) and the XTC process may comprise performing a binauralization of the sound sources, based on the location of each sound source with respect to the local orthogonal coordinates system (x', y', z' ). Moreover, the XTC process comprises:
[0189] - computing corresponding virtual direct signals considering the at least two direct signalsvirtual speakers, and
[0190] - computing corresponding virtual cancellation signals considering the at least two cancellation signals virtual speakers.
[0191] In practice, performing the binauralization of the sound sources based on the location of each sound source with respect to the local orthogonal coordinates system (x', y', z' ) gives (as a result) those direct signals that the direct signals virtual speakers will have to reproduce. In this framework, the best results can be achieved when multiple spatial regions are considered, and each spatial region of these multiple spatial regions has both its own direct signals virtual speaker (s) and its own cancellation signals virtual speaker (s). In other words, the best results can be achieved when each spatial region has:
[0192] - two or more own direct signals virtual speakers, and
[0193] - two or more own cancellation signals virtual speakers.
[0194] For example, when a first spatial region has first direct signals virtual speakers and first cancellation signals virtual speakers, a second spatial region can have second direct signals virtual speakers and secondcancellation signals virtual speakers.
[0195] In general, the second direct signals virtual speakers are not colocalized with the first direct signals virtual speakers. Nevertheless, alternative embodiments are conceivable in which the first direct signals virtual speakers and the second direct signals virtual speakers are in the same positions.
[0196] In a similar manner, the second cancellation signals virtual speakers generally are not colocalized with the first cancellation signals virtual speakers. Nevertheless, alternative embodiments are conceivable in which the first cancellation signals virtual speakers and the second cancellation signals virtual speakers are in the same positions.
[0197] Additionally, the computation of the virtual cancellation signals gives (as a result) those cancellation signals which the cancellation signals virtual speakers will have to emit for achieving XTC.
[0198] In some preferred embodiments, virtual direct signals and virtual cancellation signals may be emitted by different virtual speakers.
[0199] In particular, the arrangement shown in Fig. 4A is based on two binauralization processes (namely, one binauralization process for sound sources located in thefront of the listener LI and the other binauralization processes for sound sources located at the back of the listener LI).
[0200] The preferred embodiment according to Fig. 4A illustrates the case where two virtual speakers are used for direct signals and two virtual speakers are used for cancellation signals for each binauralization process. Nevertheless, alternative embodiments are also conceivable which utilize more binauralization processes and / or more virtual speakers per binauralization process.
[0201] Where necessary, an effective number of direct signals virtual speakers and cancellation signals virtual speakers may appropriately be varied depending on many conditions (such as, for example, user position and orientation, proximity of physical speakers, etc. ). Hence, in some embodiments, the effective number of direct signals virtual speakers and cancellation signals virtual speakers may be larger than two.
[0202] 3
[0203]
[0204] Virtual-to-physical mapping (panning)
[0205] The present subsection 3 relates to panning, i. e., to mapping from virtual speakers to physical speakers (in short, "virtual-to-physical mapping"). Accordingly, the aspects discussed in subsection 3 mainly relate toa panning module that, in Figure 3, is represented by a block denoted by reference numeral 30. Nevertheless, a reader will clearly understand that some of the aspects discussed in subsection 3 may further relate to the XTC module 20.
[0206] As it is known, panning is a process which map the virtual sound field to the real sound field (i.e., a process which maps the virtual speakers to the physical speakers). More precisely, the panning process defines which signals have to be emitted by the physical speakers to mimic what would be perceived by the user (e. g., the listener LI) from the virtual speakers. As a result, the signals of the virtual speakers are "panned", i. e., mapped from the virtual speakers to the physical speakers.
[0207] In practice, the panning includes - for any given virtual speaker - a computation of those signals which the physical speakers should emit in order to mimic the binaural signals that would be received at the listener' s ears if that / those given virtual speaker (s) would emit sound.
[0208] In general, the method according to the present invention is quite flexible as regards how to implement panning. In other words, the method according to thepresent invention can be used with any kind of "virtual-to-physical mapping" technology.
[0209] In simpler variants, virtual-to-physical mapping is carried out by using existing panning technologies such as the so-called " Vector Base Amplitude Panning" (in acronym, " VBAP")1or the approach proposed in the so-called " ITU-R advanced audio renderer" (hereinafter referred to as "advanced ITU-R" approach)2.
[0210] In the simpler variants, using the VBAP technology or the "advanced ITU-R" approach ensures a good point source localization. In other words, in the simpler variants, using classical panning processes can provide an easy solution for the virtual to physical mapping.
[0211] In more advanced variants, panning is carried out by using more sophisticated panning processes such as the "extended VBAP" process discussed below in subsubsection 3.1.
[0212] On the one hand, the simpler variants allow for a
[0213] ¹cf. Ville Pulkki, Virtual Sound Source Positioning Using Vector Base Amplitude Panning, J.AudioEng.Soc., Vol.45, No.6, 1997, June
[0214] 2cf. Recommendation ITU-R BS. 2127-0, Audio Definition Model renderer for advanced sound systems, ITU-R (
[0215]
[0216] 06 / 2019), https: / / www.itu.int / dms_pubrec / itu-r / rec / bs / R-REC-BS.2127-0-201906-S!!PDF-E.pdflow complexity implementation of the panning. On the other hand, usage of the "extended VBAP" process proposed herein ensures a mapping from virtual to physical speakers which provides both a close approximation of a signal at the TE of the listener LI (i. e., a panning process in which the signal from the physical speakers closely approximates the signal from a considered virtual speaker) and an analytical description of a sound field at the NTE of the listener.
[0217] Based on the analytical description, additional cancellation signals may be emitted to kill signals at NTE, thereby leading to an XTC with minimum coloration at the TE and maximum cancellation at the NTE.
[0218] 3.1 - " Extended VBAP"
[0219] The present subsection 3.1 presents an "extended VBAP" process which corresponds to an enhanced variation on both the VBAP and the "advanced ITU-R" approach. The extended VBAP process can be particularly suited to XTC on virtual speakers. Main goals G1 and G2 of the "extended VBAP" process are discussed below.
[0220] Gl) Given a specific (j-th) virtual speaker VSj, the listener LI with given position and orientation, and a set of physical speakers Si (with i = 1,..., N) disposed such that the specific virtual speaker VSj is locatedinside the solid angle defined by the set of physical speakers Si and by (a central head position of) the listener LI, determine what gains giand delays ki(with i = 1, …, N) must be applied to the set of i = 1, …, N physical speakers Si such that a real signal received at a TE of the listener LI is equal to a virtual signal that would be received from the specific virtual speaker VSj at the same TE of the listener LI. As a result, if the specific virtual speaker VSj were (a digital twin of) a physical speaker, the TE of the listener LI would hear the real signal corresponding to the virtual signal of the virtual speaker VSj.
[0221] G2) Given the same conditions of goal G1 above, provide an analytical description of the sound field at the NTE of the listener LI (i. e., a description of what is happening at the NTE the listener LI).
[0222] For achieving the two goals Gl, G2, the following assumptions Al to A4 are necessary.
[0223] Al) For each specific physical speaker Si (with i = 1,..., N):
[0224] - let θi and φi be the azimuth and elevation, respectively, of the same specific physical speaker Si with respect to head's position and orientation of the listener L1 (i.e. withrespect to the local orthogonal coordinates system (x', y', z' ) ); and
[0225] - let di represent the distance of the specific physical speaker Si from the center of the head of the listener LI.
[0226] A2) For the specific virtual speaker VSj:
[0227] - let θvsj and φvsj be the azimuth and elevation, respectively, of the same specific virtual speaker VSj with respect to the listener's head position and orientation; and
[0228] - let dvsj represent the distance of the specific virtual speaker VSj from the center of the head of the listener LI.
[0229] A3) Let' s define, within the virtual speakers' domain, a set of "additional speakers" S' i (with i = 1,..., N), wherein each "additional speakers" S' i is oriented in the same direction of each respective physical speaker Si but at a respective distance d' i (measured from the center of the head of the listener LI) which is equal to the distance dvsj.
[0230] In practice, each "additional speaker" S' i can be considered to represent a virtual clone of a respective physical speaker Si, and each virtual clone is located at the same distance dvsj (measured from the center ofthe head of the listener LI) in the direction of its respective physical speaker Si.
[0231] A4) For each i = 1,..., N, let' s also define a respective ratio Ri = d' i / di (wherein, in line with the terminology adopted for the above-outlined assumption Al, "di" denotes the distance of the specific physical speaker Si from the center of the head of the listener LI). An exemplary scheme for visualizing assumptions Al to A4 with N=2 is illustrated in Fig. 8. Hence, in Fig. 8:
[0232] - SI and S2 are a first and a second physical speaker, respectively;
[0233] - the first physical speaker S1 is located at a distance d1 from the center (0',0',0') of the head of the listener L1;
[0234] - the second physical speaker S2 is located at a distance d2 from the head's center (0',0',0'); - θ1 and θ2 are the azimuth of the first and second physical speakers S1 and S2, respectively, with respect to head's position and orientation of the listener L1;
[0235] - S'1 is a first additional speaker, located in the same direction of the first physical speaker S1 but at a distance dvsj from the head's center(0',0',0'); and
[0236] - S'2 is a second additional speaker, located in the same direction of the second physical speaker S2 but at a distance dvsj from the head's center (0',0',0').
[0237] In practice, in Fig. 8, the three speakers VS j, S ' l and S ' 2 are all arranged at the same distance dvsj from the head' s center ( 0', 0', 0' ).
[0238] Although Fig. 8 shows a bidimensional (2D) arrangement (i.e., a situation wherein the virtual speaker VSj is in the same plane (x',y') as the physical speakers S1, S2), any extension to a 3D arrangement should be considered trivial.
[0239] Additionally, despite the arrangement shown by Fig. 8, the panning preferably occurs in 3D.
[0240] In particular, the panning involves three physical speakers (e. g., SI, S2 and S3 ), such that the direction of the virtual speaker VS j from the center ( 0', 0', 0' ) of the head of the listener LI is located inside a solid angle formed by the positioning of these three physical speakers with respect to the center ( 0', 0', 0' ) of the head.
[0241] For achieving the two goals G1 and G2, the following assumptions A5 to A7 about sound propagation, soundsource locations, gains and delays are also necessary. A5) All sound sources are supposed to be sufficiently far from the head of the listener L1, so that sound propagation from a sound source located at distance d, at azimuth θ and elevation
[0242]
[0243] with respect to the center of the head of the listener L1 to a specific ear E of the same listener L1 can be described as:
[0244] 9dZ~kdg9,<p, EZ~ke'<l,'E
[0245] wherein:
[0246] - gdz-krepresents sound propagation in free-field as a gain (i.e., a volume) and a delay which are dependent on the distance d, and
[0247] >99,< P, EZ~KE,< L>'Erepresents a sound propagation contribution of the head (in short: "head contribution") for azimuth 0, elevation
[0248]
[0249] and specific ear E, also modelled as a gain and a delay.
[0250] In other words, gdz-kexpresses a contribution that air exerts on the propagation of any sound emitted by the sound source from the same sound source to the listener's head.
[0251] By contrast, ge,<p, Ez~ke,cl>’Eexpresses the HRTF, i. e., how any feature (e. g., the orientation) of the listener' s head influences the propagation of any sound emitted bythe sound source before the same sound reaches the specific ear E. For example, when a first sound reaches the ears of the listener L1 from behind and a second sound reaches the same ears from ahead, the head provides a filtering effect on the propagation of the first sound which is different from a filtering effect provided by the head on the propagation of the second sound (due to the anatomical conformation of the ears' auricles).
[0252] In practice, both the contribution of the sound propagation in free-field and the sound propagation contribution of the head can be modeled as respective filters which affect gain and delay.
[0253] A5) Head contributions of the form gθ,φ,Ez-kare available for any azimuth θ, elevation
[0254]
[0255] and ear E.
[0256] A6) The constraint set by the above-mentioned "sufficiently far" sound sources is expressed as:
[0257]
[0258] + ^0,<p, E »0•
[0259] A7) Let g± and ki be a gain and a delay (both to be defined), respectively, which are applied to the signal of a respective specific physical speaker Si of the set of physical speakers Si (with i = 1, …, N).
[0260] Given the assumptions Al to A6 above, goal G1 can be formulated using following relation R.8:9 ad JVSjZ7kdVSj 9 n0.VSj,<p,VSj, „TE„Z7keVSj’<l>VSj’TE
[0261] N (R.8) = ^giZ^ddiZ^1ge^uTEZ^8^^
[0262]
[0263] i = l In practice, relation R.8 basically states that the signal at the listener's TE from the specific virtual speaker VSj is equal to the sum of each contribution by each physical speaker Si (with i = 1, …, N).
[0264] Considering coherent time of arrival at the listener's TE for all signals, following condition C.1 shall be respected:
[0265] k
[0266]
[0267] dvsj+ ^eVSj,<pVS]-, TE= kt+ kdi+kei,<pi, TE i = l,..., N (C. l). Accordingly, each delay ki can be computed based on following equation EQ. l:
[0268] ki= (kd+ kθ,φ,TE) − (kd+ kθ,φ,TE) i = 1,..., N (EQ.1). In practice, in EQ. l:
[0269] - kidenotes each single, i-th, delay among a plurality of delays k1,…,kN,
[0270] - kddenotes a contribution given to the i-th delay by a respective i-th physical speaker Si of the set of physical speakers S1,…,SN when the same i-th physical speaker Si is considered as emitting sound in free-field at a respective distance di with respect to the head's center (0',0',0'),>kθ,φ,TEdenotes a contribution given to the i-th delay by the head contribution at the TE of the listener L1 corresponding to the direction of the i-th physical speaker Si, the i-th physical speaker Si being located at respective azimuth θi and elevation φi with respect to the head's center (0',0',0'),
[0271] - kddenotes a contribution given to the i-th delay by the specific virtual speaker VSj (which is included within the virtual speakers' domain) when the same specific virtual speaker VSj is considered as emitting sound in free-field at its respective distance dVSjfrom the head's center (0',0',0'), and
[0272] >kθ,φ,TEdenotes a contribution given to the i-th delay by the head of the listener L1 at the TE (i.e., the "head contribution" at the TE of the listener L1) when considering that the specific virtual speaker VSj is located at respective azimuth θVSjand elevation φVSjwith respect to the head's center (0',0',0'). On a gains side, it shall be considered that, at the listener' s TE, sound power from the specific virtual speaker VSj and sound power from each physical speakerSi are equal. This can be expressed by following relation R. 9:
[0273] N {9dVSj 9eVSj,<t>VSj, TEY =S^\9i9dl9el,<t>l, TE Y (R- 9) •
[0274]
[0275] i = l In practice, relation R.9 expresses a constraint on the normalization of the gains gi(with i = 1, …, N).
[0276] Nevertheless, relation R. 9 does not allow a proper definition of each gain gi. To do so, the concept of "first computing gain factors" g' i (with i = 1,..., N) shall be introduced. Specifically, each "first computing gain factor" g' i expresses a gain contribution provided by a respective "additional speaker" S' i which is arranged in the same direction as a respective specific physical speaker Si but at distance d from the center of the listener' s head (i = 1,..., N). Typical values for N are:
[0277] - N=3, in which case a 3D VBAP can be applied to (three) additional speakers S' l, S' 2, S' 3 as presented by Ville Pulkki in his AES article on VBAP3;
[0278] - N = 4, which corresponds to a 4 speakers-based
[0279] 3cf. Ville Pulkki, Virtual Sound Source Positioning Using Vector Base Amplitude Panning, J. AudioEng. Soc., Vol. 45, No. 6, 1997, Junepanning for point sources which can be applied to four additional speakers S' l, S' 2, S' 3 and S' as presented in the "advanced ITU-R" approach, in particular section 6. 1.2. 3 of the related document.
[0280] In practice, the "first computing gain factors" provide a number N of gain values g' i corresponding to panning gains for those N additional speakers S' i which are located at distance d from the center of the listener's head. At this stage, two more normalization steps are needed, namely:
[0281] - a first normalization step, in which each gain is adapted to take account of the physical speakers distances di; and
[0282] - a second normalization step, taking account of the constraint given by relation R. 9.
[0283] As regards the first normalization step, it shall be recalled that, for any spherical radiating point source, sound pressure decreases with distance. Accordingly, new intermediary gain coefficients g"i can be defined as:
[0284] g''i = Rig’i t = i. N (R. IO) wherein Riis the above-mentioned ratio between the distances di and d.The second normalization step is obtained by combining relations R. 9 and R.10. Accordingly, the (final) gains g± can be computed as:
[0285] _nn (9dVSJ9OVSJ,VSJ, TE )2<yi o i v’N z n \? ^ A-,..., IN
[0286]
[0287] A|Li=i(g i9dig0i,(pi, TE)z(EQ. 2). In practice, in EQ. 2:
[0288] - gidenotes each single, i-th, gain among a plurality of gains g1,…,gN.
[0289] - gddenotes a contribution given to the i-th gain by the specific virtual speaker VSj when the same specific virtual speaker VSj is considered as emitting sound in free-field at its specific distance dVSjfrom the head's center (0',0',0'), - gθ,φ,TEdenotes a contribution given to the i-th gain by the head of the listener L1 at the TE (i.e., the "head contribution" at the TE of the listener L1) when considering that the specific virtual speaker VSj is located at its respective azimuth θVSjand elevation φVSjwith respect to the head's center (0',0',0'),
[0290] - gddenotes a contribution given to the i-th gain by the i-th physical speaker Si when the same i-th physical speaker Si is considered as emitting sound in free-field at its respectivedistance di with respect to the head's center (0',0',0'),
[0291] - gθ,φ,TEdenotes a contribution given to the i-th gain by the head of the listener L1 at the TE (i.e., the "head contribution" at the TE of the listener L1) when considering that the i-th physical speaker Si is located at respective azimuth θi and elevation φi with respect to the head's center (0',0',0'), and
[0292] - g''idenotes, for each i-th gain, a respective intermediary gain coefficient which takes account of:
[0293] ■ the distance di at which the i-th physical speaker Si is located with respect the head' s center ( 0 ’, 0 ’, 0 ’ ),
[0294] ■ the specific distance dvsj, measured from the specific virtual speaker VSj to the head's center (0',0',0'), and
[0295] ■ a (respective) computing gain factor g' i, which expresses a gain contribution provided by an "additional speaker" S' i arranged (within the virtual speakers' domain) at the specific distance dvsj in a same direction as the i-th physical speaker Si with respect tothe head' s center ( 0', 0', 0' ).
[0296] Therefore, in those preferred embodiments of the present invention which utilize the "extended VBAP" disclosed herein, goal G1 can be achieved by computing each delay kiwith EQ.1, and by computing each gain giwith EQ.2.
[0297] By contrast, goal G2 can be formulated as expressing the sound field at the listener's NTE, given the gains giand the delays kifrom goal G1 above. To do so, it shall be first recalled that sound propagations from a generic sound source to the TE and to the NTE, respectively, are given by following expressions EX. l and EX. 2, respectively:
[0298] (EX.1)
[0299] g
[0300]
[0301] gdz-kgθ,φ,NTEz-k(EX.2). Thanks to expressions EX. l and EX. 2 above, a transfer function H between the signal at the TE and the signal at the NTE can be expressed as:
[0302] n, — ^e’‘t>’NTEp-tkerfjwE-ke^TE) / p i i \ TE-»NTE,8,<p —z(K. ll ).
[0303]
[0304] ge,<p, TE In view of R.8 and R.11, a sound field NTEθ,φat the NTE can be thus expressed as:N MTP — n 7~kin 7~kdi n7-^ffi,<pi, TE 99bcpbNTE -(ikd <pNTE-kd <pTE) Nl b0,<p - / 9izddjZl9el,^l, TEz 1 1-z 1 1 1 l90b<j>bTE N
[0305] _ \ -7_fed; n -7~kS i,0 i, NT E ~.9iz9dtz 190b<t>bNTEz 1 1
[0306]
[0307] i = l In practice, thanks to the above-outlined expression of the NTEθ,φ, a single pulse at the TE will result in multiple pulses at the NTE (provided that differences in sound propagation delays between the TE and the NTE are not constant across the various speakers Si). This behavior is in departure of the behavior of the specific virtual speaker VSj, for which a single pulse at the TE would also result in a single pulse at the NTE.
[0308] Therefore, the "extended VBAP" process proposed herein allows for a close approximation between virtual source and rendered signals at the TE with an analytical description of an error signal at the NTE.
[0309] 3.2 – Refinements regarding panning and XTC
[0310] For achieving full XTC, one may require that the NTE be silent and that the TE receives the desired signal obtained by one of the panning methods (for example, the "extended VBAP") described above. Hence, in some refinements of the method, silencing the NTE could be advantageously achieved by applying an additional XTC(for instance, from the physical speakers) to cancel out any residual signal at the NTE. In other words, some preferred embodiments may further include an additional XTC round (i. e., an additional round of the XTC process), to be performed after having panned the audio signal.
[0311] As a result of the additional round of XTC, minimal coloration at the TE and high cancellation at the NTE can be achieved. Hence, the additional round of XTC enables a "killing" (i. e., a silencing) of one or more audio signals at the NTE.
[0312] 4 – Implementation of the method
[0313] With reference to Figure 3, each method according to the present invention can be implemented by any system 100 which comprises the block 10 (i. e., the binauralization module), the block 20 (i. e, the XTC module), and the block 30 (i. e., the panning module).
[0314] In practice, the system 100 is configured to perform XTC in a virtual sound field.
[0315] In particular, the method according to the present invention can be implemented by:
[0316] - any data processing apparatus comprising means configured for carrying out the same method; - any computer program product comprising instructions which, when the program is executedby a computer, cause this latter computer to carry out the same method; and / or
[0317] - any computer-readable storage medium comprising instructions which, when executed by a computer, cause this latter computer to carry out the same method.
[0318] Each binauralizer (if any) can be implemented in a software form and / or in a hardware form. For example, each binauralizer can be either implemented as an integrated functionality of the virtual sound field or applied upstream of the virtual sound field (e. g., as / within an external device which does not correspond to the device which is configured to generate the virtual sound field).
[0319] Any XTC process disclosed herein can be implemented in a software form and / or in a hardware form (e. g., either by a same hardware / device which is also configured to perform the binauralization, or by a separate device).
[0320] Similarly, the panning can be implemented in a software form and / or in a hardware form (e. g., either by the same hardware / device configured to perform any XTC process and / or any binauralization process, or by a separate device).
[0321] In practice, in simpler variants, all theprocessing regarding the method according to the present invention can be handled by a dedicated software running on a computer (or on another device) capable of executing software, wherein the computations carried out by the dedicated software mimic what would happen in the virtual sound field. For example, all the processing regarding the claimed method may be handled by a software application installed on a general-purpose processor (e. g., of the DSP type). Moreover, the processing regarding the claimed method may be in form of hardware such as, for example, ASIC, FPGA, etc..
[0322] Nevertheless, more sophisticated variants are also conceivable in which the method is implemented in a distributed manner. For example, each functionality (e. g., binauralization, XTC, panning, etc. ) may be implemented in a respective device or software / plug-in application, and all the respective devices and / or software / plug-in applications can be operatively connected among each other.
[0323] In the light of the above, it has been ascertained that each embodiment achieves the intended aim in an effective manner, by optimizing XTC performance in a virtual sound field.
[0324] The invention so devised is susceptible of numerousmodifications and variations, all of which are within the scope of the appended claims; all the details may furthermore be replaced with other technically equivalent elements.
[0325] In practice, the dimensions (e. g., positions and distances) described herein may be appropriately selected according to the requirements and the state of the art.
[0326] Although the present description is focused on embodiments which take account of a single listener (i. e. the listener LI), complex embodiments are also conceivable which are adapted for supporting more than one listener.
[0327] Where technical features mentioned in any claim are followed by references signs, the reference signs have been included for the sole purpose of increasing the intelligibility of the claims and accordingly, neither the reference signs nor their absence have any limiting effect on the technical features as described above or on the scope of any claim elements.
[0328] Except where otherwise specified, in the cited figures, elements having a same or equivalent structure and / or a same or equivalent function are designated either by a same reference numeral / sign or, in case ofreference signs comprising letters and numbers, by reference signs sharing the same letters.
[0329] One skilled in the art will realize the invention may be embodied in other specific forms without departing from the invention or essential characteristics thereof. For example, although specific aspects of the present invention are described in the context of a system including a single sound source, these specific aspects are equally applicable to cases of multiple sound sources, provided each sound source is considered individually. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting of the invention described herein.
[0330] Except where otherwise specified, the various embodiments described above may be combined to provide additional and / or alternative embodiments.
[0331] Further, expressions such as "in some embodiments", "in some preferred embodiments" or the like which are mentioned in this description mean that the specific features, structures or characteristics described in conjunction with some embodiments can also be included in at least one further embodiment. Thus, these latter expressions do not necessarily all refer to a same embodiment.Scope of the invention is thus indicated by the appended claims, rather than the foregoing description, and all changes that come within the meaning and range of equivalence of the claims are therefore intended to be embraced therein.
Claims
C L A I M S1. Method for performing and optimizing crosstalk cancellation, XTC, by using a virtual sound field, said method comprising:inputting at least one input signal to an XTC module (20);executing an XTC process on said at least one input signal, so as to obtain at least one processed signal;panning said at least one processed signal, so as to obtain at least one panned signal; andoutputting said at least one panned signal to two or more physical speakers (S1-S8);characterized in that said XTC process is executed in a virtual speakers' domain (22, 24), by using two or more virtual speakers (csl, cs2, cs3, cs4, dsl, ds2, ds3, ds4, VSj ) which are included in said XTC module (20).
2. The method according to claim 1, wherein said XTC process includes a real-time optimization of positions, orientations, and / or number of said two or more virtual speakers (csl-cs4, dsl-ds4, VSj ).
3. The method according to any of the preceding claims, wherein executing said XTC process comprises computing virtual XTC signals to be output by said twoor more virtual speakers (csl-cs4, dsl-ds4, VSj ).
4. The method according to any of the preceding claims, wherein said two or more virtual speakers (csl-cs4, dsl-ds4, VSj ) comprise cancellation signals virtual speakers (csl-cs4, VSj ) and direct signals virtual speakers (VSj, dsl-ds4), and wherein executing said XTC process comprises defining number, positions, and optionally orientations, of said cancellation signals virtual speakers (csl-cs4, VSj ) and of said direct signals virtual speakers (VSj, dsl-ds4) in said virtual speaker domain.
5. The method according to claims 3 and 4, wherein said virtual XTC signals are computed after having determined said number, positions, and optionally orientations, of said cancellation signals virtual speakers (csl-cs4, VSj ) and of said direct signals virtual speakers (VSj, dsl-ds4) in said virtual speaker domain.
6. The method according to claims 4 or 5, wherein, when said at least one input signal relates to sound emitted by at least one sound source (SS) located at a respective distance (Dss) with respect to a global orthogonal coordinates system (x, y, z), executing said XTC process comprises:expressing the location of each sound source with respect to a local orthogonal coordinates system (x', y', z' );computing optimized positions and / or optimized orientations for each direct signals virtual speaker (dsl-ds4) and for each cancellation signals virtual speaker (csl-cs4), based on:- the location of each sound source, expressed with respect to said local orthogonal coordinates system (x', y', z' ), and- actual positioning and orientation of said local orthogonal coordinates system (x', y', z' ) with respect to said global orthogonal coordinates system (x, y, z);computing direct signals for said direct signals virtual speakers (dsl-ds4), based on said optimized positions and, optionally, on said optimized orientations; andcomputing cancellation signals for said cancellation signals virtual speakers (csl-cs4), based on said optimized positions and, optionally, on said optimized orientations.
7. The method according to any of the preceding claims, wherein, before being input to said at least oneXTC module (20), said at least one input signal is subjected to binauralization (10).
8. The method according to claim 7, wherein said binauralization (10) comprises one or more of:a simulation of a direction of arrival of one or more sounds;a source distance simulation; anda simulation of ambient characteristics.
9. The method according to claim 8, wherein said source distance simulation includes:modeling an amplitude variation with source distance; andapplying a delay correction to mimic propagation time of at least one audio signal from said two or more physical speakers (S1-S8) to said at least one listener (LI).
10. The method according to claims 8 or 9, wherein said simulation of said ambient characteristics is carried out as one of:convolution of signals with effective impulse responses corresponding to a given listening environment; andsimulation of early reflections and diffuse reverberation.
11. The method according to any of claims 7 to 10, wherein said binauralization includes considering multiple spatial regions, each spatial region having at least two own direct signals virtual speakers (dsl, ds2; ds3, ds4) and at least two own cancellation signals virtual speakers (csl, cs2; cs3, cs4).
12. The method according to claim 11, wherein said XTC process includes multiple XTC instances for said multiple spatial regions, respectively.
13. The method according to any of claims 7 to 12 when dependent on claim 6, wherein said binauralization includes:a first binauralization process, for sound sources located in front of said listener (LI); anda second binauralization process, for sound sources located at behind said listener (LI).
14. The method according to any of claims 7 to 13 when dependent on claim 6, wherein, when said at least one sound source (SS) includes one or more directional sources and one or more ambient sources, said binauralization is performed only with respect to any sound emitted by said one or more directional sources.
15. The method according to any of the preceding claims, further including an additional XTC round, saidadditional XTC round being performed after having panned said at least one processed signal.
16. The method according to any of the preceding claims, wherein said panning of said at least one processed signal is based on an analytical description of a sound field at a Non-Target Ear, NTE, of said listener (LI).
17. The method according to claim 16, wherein said panning of said at least one processed signal is based on a Vector Base Amplitude Panning, VBAP, process; and wherein, given:- a specific virtual speaker (VSj ) among said two or more virtual speakers (csl-cs4, dsl-ds4, VSj ), and- a set of physical speakers (Sl-SN) among said two or more physical speakers (S1-S8), said set of physical speakers (Sl-SN) being disposed such that said specific virtual speaker (VSj ) is located inside a solid angle defined by said listener (LI) and said set of physical speakers (Sl-SN),said VBAP process includes:determining gains (gi-gN) and delays (ki-kN) to apply to said set of physical speakers (Sl-SN) such that areal signal received at a Target Ear, TE, of said listener (LI) equals a virtual signal to be received from said specific virtual speaker (VSj ) at said TE of said listener (LI); andproviding said analytical description of said sound field at said NTE of said listener (LI).
18. The method according to claim 17, wherein said delays (ki-kn) are computed based on following equation EQ. l:ki= (kd+ kθ,φ,TE) − (kd+ kθ,φ,TE) i = 1,..., N (EQ.1) wherein:- k denotes each single, i-th, delay among said delays (kl-kN),- kddenotes a contribution given to said i-th delay by a respective i-th physical speaker (Si) of said set of physical speakers (S1-SN) when said i-th physical speaker (Si) is considered as emitting sound in free-field at a respective distance (di) with respect to a head's center (0',0',0') of said listener (L1),- kθ,φ,TEdenotes a contribution given to said i-th delay by said i-th physical speaker (Si) when considering sound propagation from said i-th physical speaker (Si) to said TE of said listener(LI), said i-th physical speaker (Si) being located at respective azimuth (0) and elevation (4>) with respect to said head' s center (O', 0', 0' ),- kdvsjdenotes a contribution given to said i-th delay by said specific virtual speaker (VSj ) when said specific virtual speaker (VSj ) is considered as emitting sound in free-field at a specific distance (dvsj) from said head' s center ( 0 ', 0 ', 0 ' ), and- kθ,φ,TEdenotes a contribution given to said i-th delay by the head of said listener (L1) at said TE when considering that said specific virtual speaker (VSj) is located at respective azimuth (θVSj) and elevation (φVSj) with respect to said head's center (0',0',0'); and wherein said gains (g1-gN) are computed based on following equation EQ.2:gi= g''i√((gdgθ,φ,TE)² / ΣNi=1(g''igdgθ,φ,TE)²) i = 1,...,N (EQ.2)wherein:- gidenotes each single, i-th, gain among said gains (g1-gN),- gddenotes a contribution given to said i-thgain by said specific virtual speaker (VSj ) when said specific virtual speaker (VSj ) is considered as emitting sound in free-field at said specific distance (dvsj) from said head' s center (O', 0', 0' ),- gθ,φ,TEdenotes a contribution given to said i-th gain by said head of said listener (L1) at said TE when considering that said specific virtual speaker (VSj) is located at its respective azimuth (θVSj) and elevation (φVSj) with respect to said head's center (0',0',0'),- gddenotes a contribution given to said i-th gain by said i-th physical speaker (Si) when said i-th physical speaker (Si) is considered as emitting sound in free-field at its respective distance (di) with respect to said head's center (0',0',0'),>ddi.ipi. TE denotes a contribution given to said i-th gain by said head of said listener (LI) at said TE when considering that said i-th physical speaker (Si) is located at respective azimuth (0i) and elevation (4>i) with respect to the head's center (0',0',0'), and- g''idenotes, for each i-th gain, a respectiveintermediary gain coefficient which takes account of:■ the distance (di) at which said i-th physical speaker (Si) is located with respect to said head' s center (O', O', O' ),■ said specific distance (dVSj), measured from said specific virtual speaker (VSVSj) to said head's center (0',0',0'), and■ a computing gain factor (g'i), which expresses a gain contribution provided by an additional speaker (S'i), said additional speaker (S'i) being arranged within said virtual speakers' domain (22, 24) at said specific distance (dVSj) in a same direction as said i-th physical speaker (Si) with respect to said head's center (0',0',0').
19. A data processing apparatus comprising means configured for carrying out the method of any of claims 1 to 18.
20. A computer program product comprising instructions which, when the program is executed by a computer, cause said computer to carry out the method of any of claims 1 to 18.
21. A computer-readable storage medium comprisinginstructions which, when executed by a computer, cause said computer to carry out the method of any of claims 1 to 18.