Sound reproduction systems
The system addresses the challenge of generating accurate virtual sound images for multiple listeners by employing multiple arrays of sound sources with distinct frequency ranges and optimized signal processing, achieving reduced crosstalk and efficient frequency coverage with fewer sources.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-03-12
AI Technical Summary
Existing sound reproduction systems struggle to accurately generate virtual sound images for multiple listeners by effectively reproducing sound from multiple discrete sources, particularly in terms of minimizing crosstalk and ensuring optimal frequency coverage with a reduced number of sound sources.
The system employs multiple arrays of sound sources, with each array emitting a distinct frequency range, and utilizes a signal processor to generate drive signals that minimize the condition number of the transfer function matrix between sound sources and listener positions, incorporating bandpass filters and inverse filters to achieve precise sound reproduction.
This approach allows for accurate binaural sound reproduction for multiple listeners with reduced crosstalk and optimized frequency coverage, using fewer sound sources than traditional systems, while maintaining a low condition number for improved accuracy.
Smart Images

Figure GB2025051959_12032026_PF_FP_ABST
Abstract
Description
[0001] SOUND REPRODUCTION SYSTEMS Technical Field The present invention relates generally to sound reproduction systems. Background The problem of accurately reproducing sound at the ears of a listener in order to generate a virtual sound image from multiple discrete sources has been studied by numerous authors [1- 12]. A number of previous authors have also addressed the problem of reproducing binaural sound for multiple listeners [13-19]. We have devised an improved method of sound reproduction which is applicable for two or more listeners. Summary According to a first aspect of the invention there is provided a sound reproduction apparatus,which includes a plurality of spaced-apart sound sources, and the sound sources providing afirst array arranged to emit at least a first acoustic frequency range and a second array arranged to emit at least a second acoustic frequency range; and wherein the first array is arranged to emit a sound frequency range which is higher than the frequency range emitted by the second array, and further wherein a lateral extent of the first array is less than a lateral extent of the second array. An array may be considered as a set of the sound sources which is arranged to emit a respective frequency band, and may include that the same sound source may be part of the first array and the second array (and more generally one or more of the sound sources may each serve for emitting multiple frequency bands) The first aspect of the invention may omit one or more of the above features, and one, some or all of any omitted features may be replaced by additional features. The first aspect of the invention may include features additional to those as stated.The invention may be considered as comprising multiple arrays of sound sources with eacharray being designed to deliver sound over a dominant frequency range, with lower frequencies being delivered by more widely spaced sources and higher frequency ranges being delivered by successively less widely spaced sources, the highest frequencies being delivered by the most closely spaced sources. The apparatus may be arranged to generate binaural sound reproduction for at least twolisteners. The apparatus may be arranged to generate binaural sound for each of threelisteners. Each array may emit a respective sound frequency range, which is different to the sound frequency range emitted by another array. There may be substantially no overlap between the frequency range emitted by one array as compared to the frequency range emitted by another array. The frequency range emitted by one array as compared to the frequency range emitted by another array may be separate and distinct. Where there are three or more arrays, each array may emit a frequency range which is substantially wholly distinct from the frequency ranges emitted by each of the other arrays. At least one of the sound sources may be arranged to emit multiple frequency bands (and therefore one or more sound sources need not necessarily exclusively emit a single frequencyband). The sound sources may be arranged in a single group configuration to emit the firstfrequency band and the second frequency band, and the configuration may comprise a single group / set arranged in horizontal alignment (for example as a single row). In such a configuration, there may advantageously be fewer sound sources as compared to a configuration in which each sound source exclusively emits the first frequency or the secondfrequency, and are configured in a layout of two vertically spaced rows / alignments. It willbe appreciated that the sound sources may be arranged in two distinct groups (such as two vertically spaced rows), however some of the sound sources in one row may output two frequency bands (one for each of the arrays).The arrays may be operatively combined to reduce or minimize the total number of soundsources that would otherwise be required, used with all sound sources contributing to at leastone frequency range and a plurality of sound sources contributing to two or more frequencybands.The arrays may be configured to reduce or minimize both condition number over a range offrequencies and the total number of sound sources used.By ‘lateral extent’ of an array, we include an extent of an end-most sound source of an array to an (opposite) endmost sound source of the array, which may be understood as the distal extent or (overall) width of the array. Lateral extent may be along a horizontal plane of the sound reproduction apparatus. Broadly, the lateral extent of each of multiple arrays may decrease as the frequency band level(s) output by an array increases. For example, an array arranged to output a lower frequency band may be larger / wider than for an array which outputs a higher frequency band. The feature of the first array having an inferior lateral extent to that of the second array may alternatively or additionally be described as the extent of the first array being more centrally located than the second array. This may alternatively be stated as the first array being more densely / closely grouped to the midpoint as compared to the second array. The sound sources of any further arrays, each emitting a respective frequency range, may each be progressivelymore widely spaced in relation to emitting a lower frequency range than the immediatelyprecedingly ordered array. The progressive change in lateral extent of each of the arrays may be from the highest frequency emitting array to the lowest frequency emitting array. Sound sources closer the midpoint of an array may be described as inner sound sources, and those further away from the midpoint may be described as outer sound sources. An array may have inner sound sources and outer sound sources. The first array and the second array may have midpoints which are substantially (vertically) aligned. The midpoint of an array may be the geometrical location of substantially midway between distal-most sound sources of the array. The first array and the second array may be substantially symmetrical about a vertical midline. By ‘spaced-apart’ we include that the sound sources of an array are spaced on a plane, or in a horizontal direction. The first array and / or the second array of sound sources may be substantially co-planar in relation to a respective plane, and may occupy a respective horizontal plane. The sound sources of at least one of the first array and the second array may be arranged in a respective row. More generally, for multiple arrays, the sound sources of each may be arranged in a respective row. Even more generally, the sound sources may be arranged in one or more geometrical groups, wherein one group of sound sources may be positionally / geometrically distinct from another group of sound sources. One or more of the sound sources of different arrays may form part of at least two different groups. At least one / some of the sound sources of the first array may be vertically spaced from the second array. At least one / some of the sound sources of the first array may be arranged at ahigher vertical position than at least one / some of the sound sources of the second array.Each array may be arranged as a respective group which is vertically spaced from the respective groups of other arrays. It will be appreciated that the arrays do not need to be arranged in a vertical hierarchy in relation to the respective emitted frequency ranges, and may be arranged in any vertical order (including the possibility of one or more additional arrays). Spacing of the sound sources of an array may be defined as the (lateral) distance(s) between adjacent sound sources of an array, for example the distance between the midpoint / mid-line of one sound source, and the midpoint / midline of a neighbouring sound source. For an array of sound sources, the spacings may be substantially uniform. Alternatively, or in addition, some of the spacings may be substantially the same whereas other spacings of an array may have one or more different spacings. All of the spacings between sound sources of an array may be different. The spacings between sound sources of an array may increase (in size) in relation to distance from a midpoint or (vertical) mid-line of an array. For example, the further a sound source is away from the midpoint, the greater the spacing to a neighbouring sound source of the array may be.At least some or all of the spacings between the sound sources of the first array may beinferior to the spacings between the sound sources of the second array. The spacings between sources may be up to 1m, may be approximately one or more of any value in the range 0.01 to 0.9m, and all such values are included in this disclosure. One or both of the arrays may include five or more sound sources. Further arrays may include five or more sound sources. The first array and the second array may have the same number of sound sources arranged to output the respective frequency band. Alternatively, in some implementations the first array and the second array may have a different number of sound sources. More broadly, multiple arrays may all comprise the same number of sound sources which provide the acoustic output at the respective frequency range, or may each comprise different respective number of sound sources or one or more arrays may comprise the same number of sound sources, and one more other arrays may each comprise different number of sound sources. The apparatus may include a third array of sound sources. The third array may have a wider lateral extent than both the first array and the second array. The third array may be arranged to emit a sound frequency range which is lower than the sound sources of the second array. More generally, there may be more than two arrays. Four or five arrays may be provided.The arrays may be configured to use the same number of sound sources as the intendednumber of ears of the listeners, optionally with an additional loudspeaker used substantiallycentral of the arrays to broaden the frequency range over which a condition number is low.The sound reproduction apparatus may present a condition number of 10bB or less. All sound sources may be provided, or packaged, in a single physical housing, such as a cabinet. Alternatively, two or more sound sources of one or more of the arrays may be provided in physically separate housings. An array of sound sources which generates the lowest frequency range emitted by the apparatus may include one or more sound sources which are provided in physically separate housings. Said physically separate housings may be arranged to be located laterally outwardly of a (principal) housing which contains a majority of the sound sources of the at least one array. Each, or at least some, of the sound sources may comprise an electro-mechanical transducer which is configured, on receipt of a drive signal, to emit sound. Such an electro-mechanical transducer may comprise a loudspeaker. The apparatus may comprise a signal processor which is arranged to generate drive signals to the arrays, based on an input (digital) signal, or file, representative of a sound recording or sound data which is to be reproduced. The signal processor may be configured toimplement a filter, such as an inverse filter. The signal processor may be arranged to applyone or more inverse matrices to incoming signals, and to subsequently apply bandpass filters. The signal processor may be arranged to apply bandpass filters to incoming signals that are to be reproduced at the ears of the listeners in order that each frequency range is radiated by the appropriate loudspeaker array. A bandpass filter may have a frequency response function. There may be a bandpass filter for each array. A bandpass filter may comprise at least one high-pass filter and / or may comprise at least one low-pass filter.A matrix of bandpass filters may be used to enable the sources to deliver the (required)frequency bands. The bandpass filters may be configured to minimize the condition numberof the product of the matrix of bandpass filters and a matrix of transfer functions relating thesound source outputs to the sound pressures at the ears of a number of listeners.Inverse filters may be configured to ensure (accurate) inversion of the product of the matrixof bandpass filters and the matrix of transfer functions relating the loudspeaker outputs to the sound pressures at the ears of a number of listeners.A regularization factor may be included in the configuration of inverse filters to ensureaccurate inversion of the product of the matrix of bandpass filters and the matrix of transfer functions relating the loudspeaker outputs to the sound pressures at the ears of a number of listeners.A group of four loudspeakers may be configured to deliver a frequency range that includesthe frequency at which a matrix of transfer functions relating the sound source outputs to thesound pressures at the ears of two listeners has a condition number of substantially unity (equivalent to substantially zero on a decibel scale).Multiple arrays, each consisting of four sound sources, may be combinable so as to ensurethat a matrix of transfer functions relating the sound source outputs to the sound pressures atthe ears of two listeners have a condition number of substantially unity at a number of frequencies. Various implementations are now summarized which comprise four sound sources, spaced apart from left to right and referenced source 1, source 2, source 3 and source 4 respectively, for use by two listeners, spaced apart and in front of the sound sources, and referenced listener 1 and listener 2 respectively.A frequency at which a matrix of transfer functions relating the outputs of four sound sourcesto the sound pressures at the ears of two listeners has a condition number of substantiallyunity may be defined when the magnitudes of the path length differences between each of thefour sources and the ears of two listeners sum to substantially one acoustic wavelength at that frequency.A frequency at which the matrix of transfer functions relating the outputs of fourloudspeakers to the sound pressures at the ears of two listeners has a condition number of substantially unity is defined when the magnitudes of the four path length differences between each of the four sources and the ears of two listeners sum to substantially (2n+1) multiples of one acoustic wavelength at that frequency, where n is an integer.A frequency at which a matrix of transfer functions relating the outputs of four sound sourcesto the sound pressures at the ears of two listeners has a condition number of substantially unity is defined when the magnitudes of the path length differences between source 1 and listener 1 and source 2 and listener 2 add to substantially one half wavelength, and the path length differences between source 1 and listener 2 and source 2 and listener 1 add to substantially one half wavelength.A frequency at which the matrix of transfer functions relating the outputs of four soundsources to the sound pressures at the ears of two listeners has a condition number of unitymay be defined when the magnitudes of the path length differences between source 1 andlistener 1 and source 2 and listener 2 sum to substantially (2n+1) multiples of one halfwavelength, and the path length differences between source 1 and listener 2 and source 2 andlistener 1 sum to substantially (2n+1) multiples of one half wavelength, where n is an integer.An advantageous arrangement of sources which implement the arrays may be such that whenthe path length differences between source 2 and listener 1 is set to be very small comparedto the wavelength, and preferably close to or substantially zero. By symmetry, under theseconditions, the path length difference between source 3 and listener 2 may also be set to bevery small compared to the wavelength, and preferably close to or substantially zero. Undersuch conditions source 2 and source 3 may be placed directly in front of the listeners and canparticipate in one or some of the arrays, or possibly all of the arrays, and therefore may bearranged to transmit substantially all frequencies in the frequency bands delivered by thearrays. A range of path length differences of which the magnitudes sum to substantially onewavelength may includes 0.1^ to 3 λ, where λ is the acoustic wavelength, and or include0.25^ to 1.75λ .There may be provided a substantially centrally located sound source speaker which isadditional to four arrays of sound sources to enable the condition number of the matrix oftransfer functions relating the loudspeaker outputs to the sound pressures at the ears of twolisteners to be reduced or minimized over a wider frequency range (as compared to theabsence of the central speaker). Where a sound source is arranged to output multiple frequency bands, it may be arranged to receive a drive signal which combines each or some of the multiple frequency bands. The combination of the different elements may be achieved by linear superposition. A sound source for two or more different arrays may conveniently be provided in a respective electro-mechanical transducer sub-assembly, which is a distinct / unitary physical entity, in respect of a main / principal cabinet / housing (which may contain the majority of sound sources of the arrays).The signal processor may be arranged to implement the multiple signal vectors that arecalculated for input into each array. Alternatively, or in addition, the signal processor may be arranged to implement a single vector of signals that is input to each array. The first aspect of the invention may include any filter disclosed in the Detailed Description and / or as shown in the Figures. Some or all of the sound sources may be arranged substantially linearly (that is the respective sound sources are arranged substantially in a respective line or linear path, as viewed in plan)and / or may be arranged in a part-circular fashion (that is they trace a respective substantiallyarcuate path, as viewed in plan). The sound sources may so arranged into one or more groups, and the or each group may include sound sources exclusively part of one array, or which include one or more sound sources which serve to output the frequency bands of two or more arrays. There may be multiple groups which relate to one array only, and / or which relate to multiple arrays. All groups may each relate exclusively to a respective array, or all groups may each relate to multiple arrays. Sources and listeners may be geometrically arranged in order to minimize a condition number of the matrix of transfer functions relating the complex strengths of the sources (loudspeakers) to the complex acoustic signals at the ears of the listeners.The frequency band of each array may be configured to ensure a sufficiently small conditionnumber over the range of frequencies predominantly delivered by the array, where thecondition number is that of a matrix of transfer functions relating the source / loudspeakeroutputs to the sound pressures at the ears of a number of listeners. For a given arrangement of two listeners, an arrangement of four loudspeakers may be determined which ensures that the condition number of the matrix of transfer functions is below a certain value over a given frequency range. The frequency range over which the condition number is low may be determined by the spacing between the sources. Low frequency ranges may be provided by widely spaced loudspeakers and increasingly high frequency ranges may require a correspondingly reduced spacing between the loudspeakers. In the case of a symmetrical arrangement of two listeners and four sources / loudspeakers, the frequency of the lowest condition number in a given frequency range may be such that the path length difference between the left ear of the left listener and the left-most and right- most sources of the loudspeaker array is of the order of an acoustic wavelength. The sound fields generated by well-conditioned arrangements of loudspeakers and listeners advantageously show well-defined beam patterns that ensure the required crosstalk cancellation. The use of a plurality number of different arrays of sources loudspeakers may advantageously allow coverage of a wide frequency range, with each array covering a particular range of frequencies. One or more loudspeaker unitary sub-assemblies may be arranged to emit frequencies from two or more frequency ranges (and thereby contribute to the acoustic output of two or more arrays). It may be desirable to restrict the range of frequencies over which cross-talk cancellation is produced by using only mid and high frequency ranges and so if required the size of the arraymay be reduced. In such cases, low frequency sound may be delivered by a separate singlelow frequency loudspeaker unit or units. The apparatus may be configured such that the arrays with more closely spaced loudspeakers can be housed within a single cabinet and the more widely spaced loudspeakers can be housed in separate independent cabinets.The inclusion of a central loudspeaker placed in addition to an arrangement of sound sourcesmay provide that the condition number is lower over a wider range of frequencies than by using only four loudspeakers. A central loudspeaker sub-assembly may form part of one or more of the arrays and is dedicated to transmitting the different frequency ranges. A regularization parameter may be used to (further) improve the conditioning of a given arrangement of either four or five sound sources.An arrangement of four sound sources and two listeners may ensure optimal conditioning ofthe matrix at a given frequency such that the condition number is unity. Such an optimally conditioned arrangement may be used as a reference point for system design.An arrangement of six sound sources may be used for three listeners. A central sound sourcemay be included, to therefore result in an arrangement of seven sound sources.The invention may be viewed as an improved sound reproduction system comprising multiple arrays for multi-listener virtual sound imaging, which includes multiple (discrete) sources and reproduction at multiple points. According to a further aspect of the invention, there is provided a sound reproductionapparatus which comprises multiple sound sources, and further comprises a signal processorto generate drive signals to the sound sources from a sound recording so as to generate virtualsound images to each of two or more listeners, and the apparatus further arranged todetermine listener position information of each of the listeners using sound emitted (directlyor indirectly) by each of the listeners or from respective listener positions, and to thenpreferably tailor characteristics of signal processing applied to the audio material to bereproduced, which characteristics take account of the determined listener positioninformation. This may be termed listener localization. The sound reproduction apparatus of the first aspect of the invention may be configured toimplement listener localization.The sound(s) or signal(s) emitted by a listener may be termed listener localization sounds / signals. A way of a listener directly emitting a listener localization sound includes uttering or vocalizing one or more sounds or words. A way of a listener indirectly emitting a listener localization sound / signal includes the use of an emitter device, which for example is activated at or near to the head of a listener.The sound sources may include an electro-mechanical transducers, such as a loudspeakers.Listener position information may include an angular position of a listener relative to thesound sources / loudspeakers. Listener position information may include azimuthal position and / or elevational position relative to the sound sources / loudspeakers / the sound reproduction apparatus.Listener position information may include a range or a distance of a listener from the soundsources / loudspeakers / the sound reproduction apparatus. A listener sound (or listenerlocalization signal) emanating from at least one listener position may be used to determinerange / distance, said sound / signal may have a known signal strength / output level at the pointof generation (which can be used to determine distance / range by the sound reproductionapparatus). A range / distance may be calculated for each listener.Listener localization sounds (such as a wake word, for example) may be detected and thenprocessed by, for example, a neural network. The neural network may be trained to recognizefeatures in a spectrogram formed from undertaking a short time Fourier transform (STFT) ofa signal from a single microphone. The neural network may comprise a trained convolutionalneural network. Initiation of the listener localization process may be by way of a listener uttering apredetermined word or phrase, such as a wake word. Following initiation, the first of thelisteners may then be instructed / prompted to utter a phrase such as “listener position one” or“changing listener position” which will be detected by the microphones. The microphone signals will then be processed to yield the angular position of the first listener. Once thelocation of the first listener is established, the process may be repeated for the otherlistener(s). Each of the listeners may be prompted (e.g. by an audio and / or visual output of the sound reproduction apparatus) to utter a word or phrase, which is then processed so as to determine listener location. The sound reproduction apparatus may be arranged to output audio instructions to lead listeners through the listener localization process. The audio instructions may be stored in a memory of the sound reproduction apparatus. The sound reproduction apparatus may be arranged to output an audio request for a response from a listener. The microphones may be spaced on or within the sound reproduction apparatus. The microphones may be located in spaced apart locations of a frontmost portion of the soundreproduction apparatus. Three, four, five, or more, microphones may be provided (disposedin the sound reproduction apparatus).Signals from the microphones may be processed by first undertaking a fast Fourier transform(FFT) and then evaluating the cross-power spectrum between microphone channels. A further aspect of the invention is a listener location process, including one or more the steps disclosed herein, for use with the sound reproduction system of the first aspect of the invention. The listener position information may be used to determine characteristics of one or more filters used in the generation of loudspeaker drive signals. Filter characteristics may be selected from a repository of predetermined filters. For example,once the positions of the listeners are established, a matrix of inverse filters appropriate tothe given combination of listener positions may be retrieved from a stored dictionary offilters. The invention may include one or more features disclosed in the description and / or the drawings, either singularly or in combination. Where a feature or features is part of an embodiment which includes other features, the present disclosure includes that any of the above aspects of the invention many include any such a feature or features such that they are not inextricably linked to any other features of an embodiment. The invention may include one or more features of multiple different embodiments. Brief Description of the Drawings There are now described various embodiments of the invention, given by way of example only in which: Figure 1 is a schematic representation of four sources of sound in relation to twolisteners, Figure 2 is schematic representation of the geometrical arrangement of five sources ofsound and two listeners; Figures 3 and 4 are representations in graph form of condition number of the matrix ^for the source and listener arrangement of Figure 1; Figure 5 is a representation in graph form of the condition number of the matrix B forthe source and listener arrangement of Figure 2 for different values of the source spacing; Figure 6 is a representation in graph form of the condition number of the matrix B for the source and listener arrangement of Figure 2 for different values of the source spacing showing only the frequency range in the region of the minimum value of the condition number; Figure 7 shows various representations of the sound field at the frequency of theminimum condition number for each array at different source spacings; Figure 8 shows various plots of the value of the minimum condition number fordifferent array spacings in the four-source case; Figure 9 shows plots of the frequency at which the minimum condition number occursfor different array spacings in the four-source case; Figure 10 shows graphical representations of the ratio of the longest path length difference to the acoustic wavelength for the four- source arrangement of Figure 1 as afunction of the position of the centre of the listeners;Figure 11 shows an exemplary arrangement of sources and listeners;Figure 12 is a representation in graphical form of the variation of the condition numberof the arrangement of Figure 11 as a function of the distance of the listeners from the source array;Figure 13 is a representation in graph form of the path length difference between eachsource and the left listener as a function of the distance of the listeners from the source array shown in Figure 11;Figure 14 is a representation in graph form of the condition number of the matrix B forthe source and listener arrangement of Figure 1;Figure 15 is a representation in graph form of the condition number of the matrix B forthe source and listener arrangement of Figure 2;Figure 16 is a schematic of a front elevation of an embodiment of a sound reproductionapparatus which comprises multiple arrays of sources of sound output;Figure 17 is a schematic of a front elevation of a further embodiment of a soundreproduction apparatus which comprises multiple arrays of sources of sound output;Figure 18 is a schematic of a front elevation of a yet a further embodiment of a soundreproduction apparatus which comprises multiple arrays of sources of sound output.Figure 19 shows the arrangement of six sources and three listeners;Figure 20 shows the arrangement of seven sources and three listeners;Figure 21 shows the condition number of the matrix B for the six source and threelistener arrangement shown in Figure 19 using five arrays having the spacings ls between the sources with the listeners spaced ll = 0.5m apart and at a distance lls = 2m from the source arrays;Figure 22 shows the condition number of the matrix B for the six source and threelistener arrangement shown in Figure 19 using five arrays having the spacings ls between the sources with the listeners spaced ll= 0.5m apart and at a distance lls= 2m from the source arrays. This shows only the frequency range in the region of the minimum value of the condition number;Figure 23 shows the condition number of the matrix B for the seven source and threelistener arrangement shown in Figure 2 using five arrays having the spacings ls between the sources with the listeners spaced ll= 0.5m apart and at a distance lls= 2m from the source arrays; Figure 24 shows the condition number of the matrix B for the six source and three listener arrangement shown in Figure 2 using five arrays having the spacings ls between the sources with the listeners spaced ll= 0.5m apart and at a distance lls= 2m from the source arrays. This shows only the frequency range in the region of the minimum value of the condition number;Figure 25 shows the arrangement of seven sources and three listeners with a non-uniform array geometry such that the innermost five sources are spaced at adistance ls / 2 whilst the two outermost sources are spaced at a distance ls from the fiveinnermost sources;Figure 26 shows the condition number of the matrix ^ for the seven source and threelistener arrangement shown in Figure 25 using five arrays having the spacings definedin Figure 25 with the listeners spaced ll= 0.5m apart and at a distance lls= 2m from the source arrays;Figure 27 shows the condition number of the matrix ^ for the seven source and threelistener arrangement shown in Figure 25 using five arrays having the spacings definedin Figure 25 with the listeners spaced ll = 0.5m apart and at a distance lls = 2m from the source arrays. This shows only the frequency range in the region of the minimum value of the condition number;Figure 28 shows five arrays of seven sources each having non-uniform spacingillustrated in Figure 25 can be combined to give a single array of fifteen sources;Figure 29 shows the sound field at the frequency of the minimum condition number foreach array where the frequency of the minimum is shown in brackets (a) ^^= 0.8 m(1208 Hz), (b) ^^= 0.4 m (2137 Hz), (c) ^^= 0.2 m (4113 Hz), (d) ^^ = 0.1 m (8148 Hz),(e) ^^= 0.05 m (16258 Hz);Figure 30 shows the design of the magnitude response of five bandpass filters in orderto approximately align to the regions of minimum condition number for each of the five arrays each consisting of seven non-uniformly spaced sources as illustrated inFigure 25;Figure 31 shows Signal flow diagram illustrating the signal processing system to beused in the implementation of multiple loudspeaker arrays for multiple listeners;Figure 32 shows fifth order bandpass filters used in the design of the loudspeakersystem for two listeners illustrated in Figure 18; Figure 33 shows a comparison of the condition numbers of the matrices ^^^^and ^^^^F for the loudspeaker system for two listeners illustrated in Figure 18 when using thefifth order bandpass filters illustrated in Figure 14 to comprise the matrix F; Figure 34 shows Second order bandpass filters used in the design of the loudspeaker system for two listeners illustrated in Figure 18;Figure 35 shows a comparison of the condition numbers of the matrices ^^^^and ^^^^Ffor the loudspeaker system for two listeners illustrated in Figure 18 when using the second order bandpass filters illustrated in Figure 16 to comprise the matrix F;Figure 36 is a schematic representation of symmetric geometry of the source andlistener arrangement showing (a) the source-receiver path lengths and (b) the source and listener separation distances;Figure 37 is a plot of variation of condition number of the matrix B as a function offrequency. The geometrical parameters in this case are given by R=3m, l=0.8m,s=0.55m, d=0.3m;Figures 38 illustrate the variation in the condition number of B when both the innersource spacing s and outer source spacing h are varied but the listener spacing l is held constant at (a) 0.4m, (b) 0.6m, (c) 0.8m, (d) 1.0m (e) 1.175m, wherein the black dots on the figures denote the ten smallest condition numbers found in each of the simulations;Figures 39 illustrate the frequency at which the condition number is unity as a functionof the source spacing s with the optimal choice of h, wherein results are shown for listener spacings l of (a) 0.4m (b) 0.8m (c) 1.175m; Figure 40 is a schematic illustration of an arrangement of sources designed tosynthesize five sub-arrays, each covering a different frequency range, wherein most of the sources can participate in two arrays, thus reducing the number of sources required (from twenty to twelve);Figure 41 is a plot of condition number of the matrix B as a function of frequency forthe arrangement of five arrays illustrated in Figure 40, wherein two listeners are spacedapart by l=0.8m at a distance of R=3m with assumed head width a=0.2m; Figure 42 is a plot of the condition number of the matrix B as a function of frequencyfor the arrangement of five arrays illustrated in Figure 40, wherein two listeners arespaced apart by l=0.8m at a distance of R=2m with assumed head width a=0.2m; Figures 43 are plots of the condition number of the array shown in Fig. 42 with (a) anadditional centre source compared with (b) the array without the additional centre source; andFigure 44 are representations of sound fields produced by the arrangement of sourcesshown in Figure 40 at three of the optimally conditioned frequencies shown inFigure 41.Detailed Description There are now described various embodiments of a sound reproduction system, including analysis of the system design for a virtual sound imaging system for at least two listeners.The strength of a distributed array of acoustic sources is defined by a vector ^ of order ^ andthe pressure is defined at a number of points in the sound field by the vector ^ of order ^ .In general ^ = ^^ where ^ is an ^ × ^ matrix. However, it is helpful to partition this matrix^ into two other matrices ^ and ^ and also partition the vector ^ of reproduced signals sothatThe vector ^^ is of order ^ and defines the reproduced signals at a number of pairs of pointsin the sound field at which crosstalk cancellation is sought. Thus ^^ = ^^ where ^ definesthe ^ × ^ transmission path matrix relating the strength of the ^ sources to these reproducedsignals. The vector ^^ is of order ^ and defines the reproduced signals sampled at theremaining points in the sound field. Thus ^ = ^ − ^ and the reproduced field at theseremaining points can be written as ^^ = ^^, where ^ is the ^ × ^ transmission path matrixbetween the sources and these points.Now note that the desired pressure at the ^ points in the sound field at which crosstalkcancellation is required can be written as the vector ^^. This vector can in turn be written as^^ = ^^ , where the matrix ^ defines the reproduced signals required in terms of the desiredsignals. As a simple example, suppose that crosstalk cancellation is desired at two pairs of points in the sound field such that ^^^ 1 0^ ^ ^^^^^^ ^ = ^ 0 11 0 ^ ^ ^^^^ (2) ^^^ 0 1The matrix ^ has elements of either zero or unity, is of order ^ × 2, and may be extended byadding further pairs of rows if cross talk cancellation is required at further pairs of points. It is assumed that the inputs to the sources are determined by operating on the two desiredsignals defined by the vector ^ via an ^ × 2 matrix ^ of inverse filters. The specific taskaddressed here is to find the source strength vector ^ that generates crosstalk cancellation ata number of pairs of positions in the sound field. Crosstalk cancellation at multiple pairs of points It is first assumed that cross-talk cancellation is required at the specific sub-set of all thepoints in the sound field where the number ^ of points in the sound field is smaller than thenumber of sources ^ available to reproduce the field. As noted previously
[0019] , one can thenseek to ensure that we make ^^ = ^^ = ^^ whilst minimising the “effort” ‖^‖^^ made by the acoustic sources. The problem is thus ^^^‖^‖^^ subject to ^^ = ^^ (3)where ‖^‖^ denotes the 2-norm of the vector ^. Note that the square of the 2-norm of acomplex vector is equal to the sum of the squared magnitudes of the elements of the vector such that‖^‖^^ = ^^^ where the superscript H denotes the Hermitian transpose. Thesolution
[0019] to this minimum norm problem is given by the optimal vector of sourcestrengths defined by ^^^^ = ^^[^^^]^^^^ (4)Thus one possible solution to the problem can be found that requires only specification ofthe points at which cross-talk cancellation is required in the sound field. Note that it is alsopossible to include a regularization factor ^ into this solution such that^ ^ ^ ^^^^ = ^ [^^ + ^^] ^^^ (5)where ^ is the identity matrix. Also note that if the number of sources used to reproduce thefield is equal to the number of points at which cross talk cancellation is required, then the solution given by equation (4) reduces to ^^^^ = ^^^^^ (6)An aspect of a solution of the kind given above is the conditioning of the matrix to beinverted. For example, first write ^^ = ^^ , for the case of a “square” system of equationswhere ^^ = ^^ and the number of sources is equal to the number of points in the sound fieldat which the acoustic pressure is prescribed. The singular value decomposition (SVD)of thematrix ^ can be written as where ^ and ^ are respectively the matrices of left and right singular vectors given by^ = [^^, ^^, … … .. ^^] and ^ = [^^, ^^, … … .. ^^] (8)These are orthogonal matrices having the properties ^^^ = ^ and ^^^ = ^ whilst ^ is thediagonal matrix of singular values given by where the singular values are ordered from the maximum to the minimum ^^. The solution for the source strength in terms of the SVD can then be written as
[0020] ^^^^^^ = ^ ^^ = [^^^^] As pointed out by Golub and Van Loan
[0020] relatively small changes in ^ or ^^ can inducerelatively large changes in the solution for ^^^^ if the singular value ^^ is very small. Goluband Van Loan
[0020] go on to analyze the case where there are some errors in the estimation ofthe matrix ^ or the vector ^^ by writing(^ + ^^)^(^) = ^^ + ^^ (11)where ^ is a small parameter. It then follows
[0020] that the relative error in the solution for ^can be written as where ^(^) is the condition number of the matrix ^ given by The condition number of the matrix ^ is thus central to the accuracy with which the desiredacoustic pressures can be produced in the sound field. A small value of the condition number will give a solution for the source strengths that are less sensitive to errors (for example in the estimation of the transmission path matrix ^) whereas a large value of the condition number will amplify any errors in this estimation and result in correspondingly large errors in the source strengths used for reproduction. Note that in the above relationships the 2-normof a matrix ^ is given by the square root of the largest eigenvalue of ^^^ (which is in turnequal to the largest singular value of ^). Loudspeaker listener geometry and condition number First consider the geometry illustrated in Figure 1. This shows an arrangement of four sources (i.e. M = 4) which are modelled here as simple monopole sources of sound spaced apart uniformly by a distance ^^. Also shown are the positions of the ears of two listeners(i.e. P = 4) spaced apart by a distance of 0.2 m and where the scattering of the sound by thelistener’s head is neglected. The sound fields generated by the sources are modelled by assuming that the complex acoustic pressure produced by the m’th source at the p’th ear is given by ^^^^^^^^^^^ = ^^(14) 4^^^^where ^^ is the density, ^ = ^⁄ ^^ is the wavenumber, ^^ is the sound speed, ^ is theangular frequency and ^^^is the distance from the m’th source to the p’th ear. If it is nowassumed that the pressures desired at the listener’s ears are defined by ^^ = ^^ then thematrix ^ can in general be written as^^^^ ^^^^é^^^^^^⋯ ù ^=^^ê ^^^^^^ú ê⋮ ⋱ ⋮ú (15) 4^ê ê ^^^^^^^^^^^^^^ú ⋯ ú ë ^^^^^^û The condition number of this matrix is shown in Figure 3 for five different spacings ^^of the four sources in the geometry of Figure 1. It is evident that at low frequencies, for all sourcearrangements the matrix ^ is very badly conditioned. However, it can also be seen that thecondition number falls to a minimum value at a frequency that depends on the spacing between the sources. The closer together are the sources, the higher the frequency at which this minimum occurs. Figure 4 shows the same plots, but with only those frequency ranges surrounding the first minimum of each source arrangement. Now consider the geometry of Figure 2. This arrangement has five sources instead of four, but still aims to reproduce the pressures at the four ears of just two listeners. The conditionnumbers of the matrix ^ for five different source spacings ^^ are also shown in Figure 5. Thisshows a similar behavior to the four-source case, with very poor conditioning at low frequencies but with a condition number that drops to a minimum value, this minimum value being higher in frequency for more closely spaced sources. Figure 6 shows the condition number for each source spacing but only over the range of frequencies surrounding the first minimum value. It is notable that the minimum condition number in the five-source case is lower than that in the four-source case. It is also notable that the frequency range over which the condition number is low is also considerably extended by the inclusion of the additional centre source.Finally, it is worth noting that the matrix ^ for symmetrical arrangements of equal numbersof sources and listeners will be a complex square symmetric matrix. Such matrices can be decomposed using the Autonne-Takagi decomposition (see, for example,
[0021] ) and have well- known properties
[0022] . The sound field at the minimum condition number Figure 7 shows the form of the sound field produced in the four-source case for a number of different frequencies that correspond to the minimum condition number for each source spacing. It is notable that the sound field shows a very similar form at each of these frequencies and that crosstalk cancellation appears to be very effectively delivered at these frequencies due to the structure of the sound field that exhibits strong beams of positive and negative interference. Sensitivity of the frequency and minimum value of the condition number It is interesting to note the sensitivity of both the minimum value of the condition number and the frequency at which the minimum occurs to the position of the listeners relative to the arrays having a range of source spacings. Figure 8 shows the value of the minimum of thecondition number as the listeners are moved around relative to the array for the four-sourcecase. Note that the listeners remain spaced apart by = 0.5 m horizontally and remainparallel with the x-axis in Figure 8. The plots in Figure 8 show the variation in the minimum value of condition number as a function of the coordinate position defining the centre of the two listeners. It is evident that the minimum condition number value is relatively insensitive to movements of the listeners, but increases as the listeners become closer to the array. Similarly, Figure 9 shows the variation in the frequency of the minimum condition number as a function of the centre of the two listeners. This shows that the frequency is relatively insensitive to lateral movement of the listeners, but that the frequency increases as the listeners move closer to the source array. The significance of source to receiver path length differenceFirst note that the matrix ^ can be expressed in terms of the path length differences betweenthe sources and the listeners ears. Suppose for example that the longest path from the right most source to the left most receiver in Figure 1 is given by ^^^then the matrix can be written in the form The important parameters determining the condition numbers are likely to be the terms givenby the exponent terms in the matrix defined where ∆^^ is the pathlength difference between each source and each ear. Note that ^∆^^= 2^∆^^ / ^ where ^ isthe acoustic wavelength. It is interesting to compute the ratio ∆^^ / ^ for the four-sourcearrangement in Figure 1 as a function of the position of the listeners. This is shown in Figure10 which shows the ratio ∆^^ / ^ for the longest of the path length differences at the frequencyof the minimum condition number for four different source spacings. It is noteworthy that this ratio remains very close to unity for all listener positions where the coordinates of thecentre of the two listeners remains central to the source array, for all the four arrays shown,irrespective of the values of source separation. Geometry for optimal conditioning It is also worth noting that there is a particular symmetric arrangement of sources and listeners that leads to an optimally conditioned matrix at a certain frequency. That is, there is an arrangement that leads to a unit value of condition number. An example of such an arrangement is illustrated in Figure 11. This shows as an example a four-source array havinga spacing ^^ = 0.4 m and a listener spacing of = 0.4 m. The condition number of thisarrangement as a function of frequency is shown in Figure 12 for different values of the distance between the source array and the listeners. As the distance between sources and listeners increases, the frequency of the minimum condition number also increases. This arrangement ensures that for the left listener there is zero path length difference between source 2 (counting sources from left to right) and the ears of the left listener. It also turns out that for this arrangement, the difference in path length between both source 1 and source 3and the ears of the left listener is ^ / 4 whilst the path length difference between source 4 andthe ears of the left listener is ^ / 2. A plot showing the exact variation of these path length differences as a function of the distance is shown in Figure 13. It will be appreciated that the symmetry of the source and listener arrangement ensures that exactly similar path length differences exist between the sources and the ears of the right listener. Design of systems for wider listener separations The systems described above assume listener separations of 0.4m in the optimally conditioned case or 0.5m in the first example presented. It is also worthwhile to consider acase where listeners are more widely spaced such that = 0.8m for example. In this case, thesame dips in condition number in the four-source case are illustrated in Figure 14 although the spacings in the source arrays must be chosen to again ensure that the condition number falls below a certain value. These array spacings are shown in the figure and in general thesehave to be chosen to be slightly smaller than those in the case where where = 0.5m, withthe exception of the lowest frequency range. Figure 15 shows that the inclusion of the centre speaker again helps to ensure that the condition number is lower over a wider frequency range than in the four-source case. Further embodiments including multiple arrays of sound sources Reference is now made to Figures 16, 17 and 18. These each show embodiments of a sound reproduction apparatus each comprising respective multiple arrays of sound sources. In those figures, the black circular features represent sound sources, each configured to emit (along with each of the other sound sources of an array / arrays) at least one particular frequency range of sound output. Each sound source may be embodied by a loudspeaker comprising an electro-mechanical transducer. Turning initially to Figure 16, this shows an arrangement of loudspeakers housed in a single cabinet 20 i.e. it is a single physical unit. The arrangement of loudspeakers is represented as front elevation. Three arrays of speakers are included, which are arranged in three aligned respective groups. As viewed in Figure 16, the uppermost array and the (vertically) intermediate array comprise five speakers. The lowermost array comprises nine speakers. Thespeakers of each array are spaced apart. Spacing s between speakers is measured fromgeometric mid-point of one speaker to the geometrical mid-point of a neighbouring speaker. The speakers of the uppermost array, as shown in Figure 16, are each arranged to output only the highest band of frequencies, relative to the frequency bands output by each of the other arrays. The speakers of the intermediate array are arranged to emit sound only in a frequency range which is lower than that of the band emitted by the uppermost array. The speakers of the lowermost array are arranged to emit sound only in a frequency band which is lower than that of the intermediate array (and therefore also lower than the uppermost frequency band). The speakers of the uppermost array are uniformly spaced apart, that is the distance from one speaker of the array to its neighbour or neighbours is the same for all of the speakers. Similarly, the speakers of the intermediate array are also evenly spaced apart, but by a larger spacing than the uppermost array. Some of the speakers of the lowermost array are uniformly spaced, whereas the outermost four speakers are more widely spaced as compared to the speakers located centrally in the array. The spacing between the centremost five speakers is the same (or substantially so). The spacing of the two speakers which are next further laterally outward relative to their immediately inward neighbours is greater than the spacing of the five central speakers. And the spacing of the outermost speakers of the lowermost array is greater than all of the spacings in the lowermost array are greater than the spacings of the intermediate array and the uppermost array.Each of the three speaker arrays has a lateral extent or overall width. As shown by w this isthe distance between the outermost speakers. The uppermost array has the smallest lateral extent, the intermediate array has a larger extent, and the lowermost array has the greatest lateral extent. Each of the three speaker arrays is symmetrical about a common midline, as shown by the broken line ML in Figure 16. The device 15 comprises a signal processor which is arranged to generate drive signals to the arrays, based on a digital signal, or file, representative of a sound recording or sound data which is to be reproduced. The signal processor implements an inverse filter. The sound field which is generated by the apparatus creates high quality binaural sound reproduction for each of two listeners. The device 15 comprises a data port, through which sound data is received, from a device such as a computer, a smartphone, a tablet computer, or a personal electronic device (PED), or a solid-state memory device (on which is stored sound data). The data port may comprise a physical connector (for coupling to a counterpart connector) by wired connection and / or the port may comprise a wireless connector which is arranged to receive data over an air interface. The device 15 may comprise a network connector which is configured to connect to a local network and / or the internet, and a user may control the device to input sound data to be reproduced from either or both. The device 15 may comprise an internal memory which may be configured to store sound data which is accessible to the signal processor to be reproduced by the arrays of loudspeakers. Turning to Figure 17, this shows a different implementation of the arrangement of Figure 16 in which each of the outermost speakers of the lowermost array is provide in a respective and physically distinct housing 15b which is apart from the (principal) housing 15a which contains the majority of speakers. It will appreciated that other possibilities in relation to modularization of the speaker set are possible. The sound reproduction device 250 shown in Figure 18, is arranged to provide five sound source arrays, each outputting a respective frequency range (of five frequency bands in total). In overview, this is achieved by causing some of the speakers to emit multiple sound frequency bands. Like the embodiments in Figures 16 and 17, the device 250 presents three rows of five speakers, but in this embodiment 250 additional frequency bands are output by the same number of speakers. In Figure 18, shows a schematic layout of five arrays, each array having a respective group of five speakers, each group aligned in a row. This layout represents a way in which five arrays to cover five frequency bands could be implemented. Figure 18 shows how the required sound outputs can be achieved using the layout of the embodiment in Figure 16. This is achieved by some of the speakers providing output for multiple frequency bands. An array 201 emits the highest frequency band, an array 205 emitting the lowest frequency band, and the arrays 202 to 204 emit distinct frequency bands intermediate of the uppermost band and the lowermost band. In this implementation, some of the output relating to subsets of the arrays 203, 204 and 205 are provided by a single speaker in the unit 20. Specifically, the speakers 212, 213, 241, 215 and 216, which are speakers of the lowest row of the unit 20, each provide a combined / shared sound output for some of the vertically aligned speakers of the arrays 203 to 205, as shown in Figure 18. It is noteworthy that each said speakers of the unit 20 conforms to the required geometrical positioning. As will evident, the use of singular speakers to provide combined output for multiple speakers results in a device with significantly smaller packaging as compared to if each of the arrays was implemented by a respective row of speakers. The input signals to such shared output speakers comprises two signals, one for each of the respective frequency bands, which are combined used linear supposition. The spacings between the speakers of the different arrays may be as follows, the firstarray 201 may have a speaker spacing of 0.05m, the second array may have a speaker spacingof 0.1m, the third array may have a speaker spacing of 0.2m, the third array may have a speaker spacing of 0.4m and the fifth array may have a speaker spacing of 0.8m. Extension to three listeners The same principles as described above in relation to two listeners can be applied to thedesign of systems that will generate virtual sound images for three listeners. Figure 19 showsa typical arrangement of loudspeaker sources and listeners. In this case the number of loudspeakers required for any given frequency band will be at least six to ensure crosstalk cancellation at the ears of all three listeners. The arrangement for three listeners may include a centre channel such that seven loudspeakers may be used to ensure crosstalk cancellationfor all three listeners. Such an arrangement is illustrated in Figure 20. The condition numbersof the matrix ^ for the geometrical arrangement of loudspeakers and listeners shown inFigures 19 and 20 are also shown in the plots shown in Figures 21, 22, 23 and24. Thecondition number dependence determines the frequency ranges to be delivered by each of theloudspeaker arrays. It can be seen from the plots in Figures 21, 22, 23 and 24 that using theadditional centre loudspeaker (so that seven speakers are used in the array) provides anenhancement to the arrangement shown in Figure 19 of six loudspeakers by reducing thecondition number over a wide frequency range.As previously described above in relation to the embodiments above in relation to two listeners, it is also possible to use some loudspeakers to assist in the delivery of more than asingle frequency range in the case of three listeners. There is considerable advantage in usinga slightly non-uniform spacing in each of the component arrays. This is illustrated inFigure 25 which shows that by spacing the outermost speakers in a given array at a separationdistance which is double that separating the five innermost speakers, then the total numberof loudspeakers required in combining all the arrays is reduced. The condition numbers ofthe matrix ^ for this series of arrays is shown in Figures 26and 27. Using independentloudspeakers for each of the five arrays requires thirty-five loudspeakers in total, with eacharray realized by a respective row of seven loudspeakers (as shown in the upper region ofFigure 28). However, the use of non-uniform loudspeaker spacing enables the number ofloudspeakers required to be reduced to fifteen arranged in a single row, with some speakersthereof outputting frequency ranges of two or more arrays. This ‘condensed’ arrangementimproves considerably the condition number at very high frequencies and minimizes theoverall length of the array. The arrangement of fifteen non-uniformly spaced loudspeakersserving as five arrays of seven speakers each, is illustrated in the lower portion of Figure 28.As is also the case for the two-listener arrangements described earlier, the sound fieldsgenerated by the above described three-listener arrangement provide advantageouscharacteristics shown in the plots presented in Figure 29 (for an array consisting of sevenloudspeakers). These figures show, at a number of representative frequencies, the form of thesound field generated with the distinct interference patterns that ensure crosstalk cancellation at the ears of all three listeners. Design of bandpass filtersIn implementations of the systems designed to operate using the principles described above,it will be necessary to apply filters to the incoming signals that are to be reproduced at theears of the listeners in order that each frequency range is radiated by the appropriateloudspeaker array. For example, this can be accomplished by using the arrangement of filtersillustrated in Figure 30 which shows how the bandwidth of the filters can be approximatelyaligned to cover the range of frequencies within which each array results in a low conditionnumber of the matrix B which relates the loudspeaker input signals to the listener ear signals.This shows the case of the five non-uniformly spaced arrays of seven sources each, but thesame principles will apply to the case of the arrays designed for just two listeners.The signal flow in a real system implementation is illustrated in Figure 31. Firstly, thedesired signals given by the matrix product ^^ = ^^ are processed via a matrix of inversefilters ^ whose output signals can be defined as the vector of signals ^ such that ^ = ^^^.These signals are then filtered by a number of bandpass filters each having a respectivefrequency response function ^^ where the index ^ denotes the array. The roll-off of thebandpass filters is designed to ensure that the magnitude of the overall frequency responsefunction of the system is as uniform (flat) as possible. Thus, at any given frequency, thesignals at the listeners ears will be dominated by the contribution from the array designed todeliver that frequency, but there will also be contributions from the arrays designed topredominantly radiate other frequencies. This occurs as a result of the overlap of thefrequency response functions of the filters illustrated in Figure 30.In order to illustrate a possible the approach,, first assume that there are only two arrays (ofsay five sources each) and that the input signals to the loudspeakers in each array are denoted as the vectors ^^and ^^respectively. These signals are generated by passing the vectors ^^and ^^ via bandpass filters having frequency response functions ^^ and ^^ respectively. (Notethat these signal vectors are distinct from the vectors defining the singular value decomposition described above). Writing the individual inputs to the speakers in each arrayas ^^^ where the index ^ varies from one to two, whilst the index ^ varies from one to five, thespeaker inputs can be written as Now note that it is also possible that (for example ) the speaker input ^^^consists of a combination of inputs from say ^^^^^and ^^^^^, whilst the speaker input ^^^consists of a combination of say ^^^^^and ^^^^^and the speaker input ^^^consists of a combination ofsay ^^^^^ and ^^^^^. The number of loudspeakers is then reduced, and the relationshipbetween the composite vectors can be written as the matrix operation given by The above relationship demonstrates how the use of a given loudspeaker can participate in more than one array. In general, one can define a composite vector of loudspeaker inputsignals ^ and a composite vector of filter input signals ^ such that where, as illustrated above, the matrix ^ will have a form determined by the extent to whicheach of the bandpass filter input signals contributes to the loudspeaker input signals.Whilst it is possible to calculate the signal vectors input to each array as described above, itis also possible to calculate a single vector of signals that is input to each array. This has theadvantage of reducing the output dimension of the inverse filter matrix and thus reduces theprocessing necessary in the implementation of the system. To illustrate this point, nowassume that only a single vector ^ is input to two arrays. The relationship that defines theloudspeaker input signals can then be written as Again, there can be a modification to the matrix of bandpass filters in order to reduce thenumber of loudspeakers necessary. For example, it is also possible that (say) the speakerinput ^^^ consists of a combination of inputs from say ^^^^ and ^^^^, whilst the speakerinput ^^^ consists of a combination of say ^^^^ and ^^^^ and the speaker input ^^^ consists ofa combination of say ^^^^ and ^^^^. The number of loudspeakers is then reduced, and therelationship between the composite vectors can be written as the matrix operation given byThe above relationship again demonstrates how the use of a given loudspeaker can participatein more than one array. In general, one can define a composite vector of loudspeaker inputsignals ^ such that where, as illustrated above, the matrix ^ will have a form determined by the extent to whicheach of the bandpass filter input signals contributes to the loudspeaker input signals.An example of the configuration of the matrix ^ is given here for the case of the array ofloudspeakers depicted in Figure 10. The individual fifteen loudspeakers can be identified by counting the loudspeakers shown in the box in Figure 10 from left to right and designatingtheir indices from 1 to 15. The coordinate positions of the loudspeakers in this case is givenin the table below. Loudspeaker1 2 3 4 5 6 7index Coordinate (m) -1.6 -0.8 -0.4 -0.2 -0.1 -0.05 -0.025Loudspeaker8 9 10 11 12 13 14 15index Coordinate (m) 0 0.025 0.05 0.1 0.2 0.4 0.8 1.6The table below shows which loudspeakers are used to transmit the signals from the five component arrays illustrated above the box in Figure 10. Source array Loudspeaker index1 Low frequency 1 2 3 8 13 14 152 Low-mid frequency 2 3 4 8 12 13 143 Mid frequency 3 4 5 8 11 12 134 High-mid frequency 4 5 6 8 10 11 125 High frequency 5 6 7 8 9 10 11In this case the matrix F has seven input signals and outputs the values to fifteen loudspeakersand the matrix relationship between the single input signal vector ^ and the full loudspeakerinput signal vector ^ takes the following form. Now note that the signals reproduced at the listeners ears is given by the relationship ^^ =^^^^^ where the matrix ^^^^ now relates all the loudspeaker input signals (i.e. the signals inputto all the arrays) to the signals generated at the listeners ears. Thus, for example, if we denotethe relationship between the ^ ‘th array and the signals at the ears of the listeners as ^^^ =^^^^ then it follows that since the total signal ^^ will be given by the matrix product then the matrix ^^^^is given by the composite matrix ^^^^ = [^^ ^^ … ^^]In general, this matrix will include the effect of the listeners head, which will often bemodelled by a population average frequency response function, possibly represented by astandard mannequin or dummy head. It therefore follows that the signals at the listeners earsare given by ^^ = ^^^^^ = ^^^^^^Note that the analysis that follows is applicable to the case where the vector ^ is either acomposite vector of signals input to each array, or a single vector of bandpass filter inputsignals as illustrated above. In signal processing terms, one can now seek to determine thebandpass filter input signals ^ that provide a solution that minimizes ‖^‖^^ subject to theconstraint of ^^ = ^^ = ^^ . The solution in this case is therefore given by Since the bandpass filter input signals are given by ^ = ^^^ where ^ is the matrix of inversefilters it also follows that ^^ ^^^^ = (^^^^^)^[(^^^^^^)(^^^^^)+ ^^]where ^^^^is the optimal matrix of inverse filters. This is therefore computed from aknowledge or measurement of the matrix ^^^^ and the chosen structure of the matrix ^ for thearray design being employed. It is also possible to make use of the regularization factor ^ asis well known in the field of inverse filter design, in order to ensure that the inverse filterresponses are well-contained in the time domain. Finally note that in designing the inversefilter matrix ^^^^ , the causality of the filters in this matrix can be ensured by using theconstraint that the signals reproduced at the ears of the listeners are equal to a delayed versionof the desired signals ^^. That is, one ensures that, in the frequency domain, ^^ = ^^e^^^∆where ∆ represents a suitably chosen modelling delay.An example of bandpass filter design can be presented for the case of the array ofloudspeakers depicted in Figure 18 intended to provide virtual imaging for two listeners. Theindividual thirteen loudspeakers can be identified by counting the loudspeakers shown in thebox in Figure 18 from left to right and designating their indices from 1 to 13. The coordinatesof the loudspeakers along the horizontal axis are shown in the table below (where the centreloudspeaker is placed at the origin). Loudspeaker1 2 3 4 5 6 7 8 9 10 11 12 13indexCoordinate (m) -1.6 -- - --0.05 0 0.05 0.1 0.2 0.4 0.8 1.60.8 0.4 0.2 0.1The contributions from each of the loudspeakers to each of the component arrays of sourcesare given in the table below Source Array Loudspeaker index1 Low frequency 1 2 7 12 132 Low-mid frequency 2 3 7 11 123 Mid frequency 3 4 7 10 114 High-mid frequency 4 5 7 9 105 High frequency 5 6 7 8 9The relationship between the input signal vector ^ and the loudspeaker input signal vector ^therefore takes the following form, with the values of ^^ , ^^ etc. again denoting the frequencyresponse of the bandpass filters where ^^^^ = ^^ + ^^ + ^^ + ^^ + ^^. This therefore enables the definition of the matrix ^ for thegiven loudspeaker configuration.Now note how the choice of the bandpass filters influences the condition number of the matrix^^^^^ that is subject to the least squares inversion process defined above. It is to be noted thatthe condition number of this matrix is influenced by the design of the bandpass filters. Thus,for example, if an array of fifth order bandpass filters is used, with the frequency responseshown in Figure 14, then the condition number of the matrix ^^^^^ that results is plotted inFigure 15 as a function of frequency and is compared with the condition number of theunfiltered plant matrix ^^^^ . In general, the condition number of the matrix is advantageouslymade lower over much of the frequency range.. However, also note that several undesirablepeaks occur in this figure, indicating that at the frequencies in the region of these peaks willresult in poor quality reproduction at the listeners ears. Changing the filter types to be secondorder, however, enables these undesirable peaks to be suppressed. Figures 16 and 17respectively show the modification the filter frequency response functions and the modified condition number of the matrix ^^^^^. The characteristics of the bandpass filters can beoptimized to ensure that the conditioning of the matrix inversion can be minimized to ensuregood reproduction of the desired signals. Acoustic localization of listener positionThe design the filter matrix ^^^^ for the loudspeaker arrays described above whether foreither two or three listeners, depends on the position of the listeners relative to theloudspeaker array. As is evident from the form of the sound field radiated by the arrays, theangular position of the listener relative to the array is of principle significance, assuming thatthe listeners are placed within a certain distance from the arrays. One method of detectingthe angular position of a listener (in both azimuth and elevation) is to have the listener tospeak from their desired listening position such that one or more microphones incorporatedinto the sound reproduction apparatus can detect the sound of the words spoken by thelistener. The respective angular position of each listener uses the microphones which are builtinto a front face of the housing of the sound reproduction apparatus, which also contains theloudspeakers. The positions of the listeners are then deduced from the spoken voices of thelisteners, the sound of which is detected by the microphones. A user may utter a “wake word”to alert the sound reproduction apparatus to the presence of a user. A user may therefore usea phrase such as “hello sound bar” or “hey loudspeakers” detected by a single microphone tothen trigger the acquisition of the signals by all the microphones. The wake words may bedetected by, for example, first training a convolutional neural network to recognize thefeatures in spectrogram formed from undertaking a short time Fourier transform (STFT) of asingle microphone signal. The training of such a network can be accomplished by firstpresenting a number of such utterances from a wide range of spoken voices from a largepopulation of potential users. The weights in the network are updated during training throughuse of the backpropagation algorithm. Considerable research has been undertaken on theefficient detection of such wake words and the field generally of keyword spotting and thedetails of many methods are disclosed in, for example, references [23-31].Following initiation of the process by the wake word, the first of the listeners may then beinstructed to utter a phrase such as “listener position one” or “changing listener position”which will be detected by the microphones. The microphone signals will then be processedto yield the angular position of the first listener. Once the location of the first listener isestablished, the process is repeated for either one or two further listeners. A very convenientmethod for determining the angular position of the listener is to process the signals from themicrophones that will typically consist of a relatively small number of microphones, such asthree, four or five. The signals from the microphones are then processed by first undertakinga fast Fourier transform (FFT) and then evaluating the cross-power spectrum betweenmicrophone channels. Thus, if there are two microphones in the array producing outputsignals whose discrete Fourier transforms are given by X^(^) and X^(^) where ^ representsthe discrete frequency index, the normalized cross spectral density function is given by This complex valued function can be evaluated for each pair of microphones in the array togive a complex vector, which in the case of a three-microphone array, will have the form^^(^) = [Φ^^(^) Φ^^(^) Φ^^(^) Φ^^(^) Φ^^(^) Φ^^(^)]This vector consists of the normalized cross-power spectra of all the combinations of themicrophone signals available. It is also possible to undertake some time averaging of thevalues of Φ^^(^) from a number of FFT blocks and thus form the expected value of thiscomplex quantity. The complex vector ^(^) can be presented to the input of a complex valuedneural network which may either be a complex multilayer perceptron or a complex valuedrecurrent neural network. Such a network is first trained by using a number of human speechsignals generated either by a loudspeaker placed at defined listener positions relative to theloudspeaker arrays or by a number of human speakers at such defined positions. During thetraining process, the vector defined above that is deduced from the measured microphonesignals is repeatedly input to the neural network for a range of listener positions and for eachposition the weights in the network are updated by using the complex backpropagation algorithm. The loss function used in the backpropagation algorithm can be based on the meansquared error between the actual source location and the location estimated by the network.An alternative is for the network to have a number of outputs, each of which corresponds toa defined location of the source, such that the network is configured to solve a classificationproblem, the classes involved being the various locations of the speech sound source. In thiscase the backpropagation algorithm can be used to minimize, for example, a loss functionbased on the cross entropy of the desired output vector and the estimated output vector for agiven position of the speech sound source. Details of such a complex valued network andassociated loss functions are given in reference
[0032] .Alternatively, rather than make use of complex valued neural networks, the real andimaginary parts of the vector ^(^) can be concatenated into a single vector and presented tothe input of a real valued multilayer perceptron or a real valued recurrent neural networkwhich are then trained using the same process. Once such networks are trained, they can thenbe employed within an additional signal processing apparatus of the loudspeaker array inorder to detect listener location. Once the positions of the listeners are established, using any of the above described processes, the matrix of inverse filters appropriate to the given combination of listenerpositions is retrieved from a stored dictionary of filters and downloaded in order to processthe desired signals as described above.Another acoustical method of localizing the positions of the listeners is to employ a remote-control device of the type that typically works with coded infrared signals often used intelevision control. Such a remote control may be provided with a high frequency sound sourcewhose output is set at a fixed / predetermined sound level. Such a sound source could be usedas the source whose position is detected, rather than using the voices of the listenersthemselves. Such a device needs to be located at or near to the head position of each listenerin turn and a suitable acoustic signal emitted for detection via the microphone array. Thesignal emitted may be ultrasonic and the microphone array may be designed to mostefficiently detect the signals emitted by the sound source. Having a source with a knownfixed output level will then also enable the detection of the range of the listeners from theloudspeaker array in addition to their angular location in azimuth and elevation. The signalprocessing method described above can also be used in conjunction with the appropriatelytrained neural network to yield the coordinates of the listeners in three dimensions.Further embodiments of the four-source two-listener arrangement Further embodiments are now described which build on and / or are variants of the disclosure above. Consideration of the multiple listener crosstalk cancellation problem is again discussed, with a specific focus on the four-source two-listener problem coupled with further observations regarding the optimal conditioning of this problem. We have discovered that a number of frequencies can be identified where the problem is optimally conditioned such that the matrix inversion required can be accomplished with minimum error. We have also discovered that the sound field radiated at these frequencies is very similar to that associated with the optimal source distribution. Furthermore, a design is disclosed below that enables good quality crosstalk cancellation to be produced simultaneously for two listeners. First, the arrangement of four sources and two listeners to be analysed in detail below is illustrated in Figure 36. As a preliminary to the study of this two-listener problem, it is again helpful to describe briefly the origin of the OSD for binaural reproduction for a single listener. An initial analysis is thus presented of the system consisting of Source 1, Source 2 and Listener 1 depicted in Figure 36. The spacings of the sources and listeners are linear as indicated and the sources are assumed to be point monopoles. In analysing this problem, the same notation will be used as that adopted above. A harmonic time dependence is assumed throughout and the desired signals for reproduction at the listener’s ears, the source signals, and the reproduced signals are defined by the complex vectors given respectively by ^= [^ ^ ^^, ^^] , ^ = [^^, ^^] , ^^ = [^^, ^^]^The transmission path matrix ^ defines the relationship ^^ = ^^ between the source inputsignals and the reproduced signals whilst the inverse filter matrix ^ defines the relationshipbetween the source input signals ^ and the desired signals ^ such that ^ = ^^. It thereforefollows that ^ = ^^^ and that the reproduced signals can in principle be made equal to thedesired signals provided ^ = ^^^. Note that the inverse filter matrix can be made causalprovided there is a vector of target signals ^ at the listener’s ears that are sufficiently delayedversions of the desired signals ^. The matrix ^ is given by where the terms ^^^ denote the radial distance from the source, ^^ is the density and ^ = ^⁄ ^^is the wavenumber, where ^ is the angular frequency and ^^ the sound speed. In the firstinstance the diffraction of the listener’s heads will be ignored. As has been amplydemonstrated in previous studies of virtual sound imaging, a key property of the matrix ^ isits condition number defined by where ^^^^(^) and ^^^^(^) are respectively the maximum and minimum singular values ofthe matrix ^ and ‖^‖^ denotes the 2-norm of ^. The inversion of this matrix is necessary toachieve crosstalk cancellation, and the singular values of the matrix are central to the effectiveness of the system in reproducing the desired signals. It is to be noted that thesingular values of ^ can be found either by taking the positive square roots the eigenvaluesof the matrix ^^^, or from the positive square roots of the eigenvalues of ^^^, where thesuperscript H denotes the Hermitian transpose. As demonstrated below, working with ^^^enables important properties of the matrix to be expressed in terms of the path lengthdifferences between each source and the ears of the listeners. The eigenvalues ^ can thus bedetermined by finding the roots ofdet^^^^ − ^^^^ = 0,where ^^denotes the identity matrix. Note that the matrix ^^^is Hermitian (equal to its Hermitian transpose) and normal (it commutes with its Hermitian transpose). To simplifymatters, it is assumed that the listener is in the far field of the sources at a distance ^ suchthat the denominator terms in ^ can be approximated ≈ ^^^ ≈ ^^^ ≈ ^^^ ≈ ^. As will bedemonstrated below, this still enables the essential physics of the problem to be captured by the analysis. Using this assumption shows that where Δ^ = ^^^ − ^^^ and Δ^ = ^^^ − ^^^ respectively define the path length differences in thedistances from the sources 1 and 2 to the ears of the listener. The path length differences can be positive or negative. The eigenvalues of a 2 x 2 matrix having the form are given by A condition number of unity for the matrix ^ can only be achieved when the two eigenvalues of the matrix ^^^are equal. This implies that the term under the square root sign in the above equation must be equal to zero. Examination of the terms in the matrix ^^^shows that(^ + ^)^ − 4(^^) = 0 and that since ^ is the complex conjugate of ^, then ^^ = |^|^ which inturn must be equal to zero. Now note that|^|^can only be equal to zero provided that boththe real and imaginary parts of ^ are equal to zero. This then implies that the followingconditions must be satisfied: cos^Δ^ + cos^Δ^ = 0sin^Δ^ + sin^Δ^ = 0These conditions can also be written as These identities imply that both the real and imaginary parts of ^ will be zero if the cosineterm in both above equations is zero. The cosine function is zero when the angle is equalto (2^ + 1)(^ ^) where ^ is a positive or negative integer. Thus, unity condition number willfollow provided that ^Δ^ − ^Δ^ = (2^ + 1)^Since ^ = 2^ / ^ where ^ is the acoustic wavelength, the above equation shows that thecondition for optimality becomes Δ^ − Δ^ = (2^ + 1)^ 2If the sources are arranged relative to the listener as depicted in Figure 36 then Δ = − ^^will be negative whilst Δ^ = ^^^ − ^^^ will be positive. Thus, under these circumstances onecan write the above condition as The left side of this equation must be positive and thus the right side of this equation mustalso be positive, so ^ must always be chosen to be negative and non-zero. Thus for ^ =−1, −2, −3 … When the sources are arranged symmetrically with respect to the listener, such that Δ^ = ^^^ −^^^ and that Δ^ = ^^^ − ^^^ such that Δ^ = −Δ^ then the following simple condition results:|Δ^| = |Δ^| = (2^ + 1)^ 4 This is entirely consistent with the well-known result that optimal conditioning is achievedat a given frequency when the path length differences at that frequency are equal to (2^ + 1)multiples of one quarter of the acoustic wavelength. However, for asymmetric arrangements of two sources and a single listener, it seems that the sum of the magnitudes of the path lengthdifferences must add to odd multiples of one half of the acoustic wavelength.The arrangement of four sources and two listeners is illustrated in Figure 36. And it is assumed from the outset that this arrangement will be symmetric about the vertical axis. The spacings of the sources and listeners are linear as indicated and the sources are again assumedto be point monopoles. In analysing this problem, the same notation as above will be used.The desired signals for reproduction at the listener’s ears, the source signals, and the reproduced signals are defined by the complex vectors of order four given respectively by^, ^ and ^^. As in the two-source single-listener case described above, the four-by-fourtransmission path matrix ^ defines the relationship ^^ = ^^ between the source input signalsand the reproduced signals whilst the inverse filter matrix ^ defines the relationship betweenthe source input signals ^ and the desired signals ^ such that ^ = ^^. It therefore againfollows that ^^ = ^^^ and that the reproduced signals can in principle be made equal to thedesired signals provided ^ = ^^^. Note again that the inverse filter matrix can be made causalprovided there is a vector of target signals ^^at the listener’s ears that are sufficiently delayed versions of the desired signals ^. Finally, also note that the desired signals forreproduction can be written in terms of a pair of signals, defined by the vector ^^^ = [^^ where ^^and ^^are the desired signals at the left and right ears of both listeners such that^ = ^^^^. Thus, for example, the vector of desired signals can be written as Note that specification of the elements of the matrix ^ enables different levels of signal tobe replicated for the two listeners, which may be important in some applications. It is also possible to include frequency dependent functions in this matrix, if for example, one listener was troubled by reduced hearing in a particular frequency range. Again, the diffraction of the listener’s heads will be ignored, and if each source acts as a free field monopole, the signal^ produced at the ^ ‘th location due to the ^ ‘th source having volume acceleration ^ is^ ^again given by ^ ^^^ ^ =^e^^^^^^, 4^^^^where ^^^is the radial distance from the source. It will also help the analysis below by usingthe notation ^ = ^ (^ ⁄ 4^) ^ where ^ = ^^^^^^ ^ ^ ^^ ^^ ^^⁄ ^^^ . It then follows that the matrixrelating the pressures at the listener’s ears to the strengths of the four sources can be written as Now it follows from the symmetry of the arrangement illustrated in Fig.1 that the following equalities hold: ^= ^ , ^ = ^ , ^ = ^ , ^ = ^ ,^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^^ = ^ , ^ = ^ , ^ = ^ , ^ = ^ .^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^These relationships, together with the interchange of the ordering of both ^ , ^ and ^ , ^ ,^ ^ ^ ^enable the matrix above to be written in the form Thus, the relationship between the vectors of sound pressures ^ and source strengths ^ canbe written in the form ^^ = ^^ where the matrix ^ can be written in the form where the block matrices ^ and ^ are respectively given by As in the two-source single listener case described above, it is helpful to determine thecircumstances under which the matrix ^ has unit condition number. The eigenvalues can thusbe determined by finding the roots ofdet^^^^ − ^^^^ = 0,where ^^is the four-by-four identity matrix. Furthermore, it is worth noting that the blockmatrix structure of ^ shows that where the matrices ^ and ^ are respectively given by^ = ^^^ + ^^^,^ = ^^^ + ^^^.It is important it what follows to note the form of the matrices ^ and ^ when the listeners areeffectively placed in the far field of the sources such that all the terms 1 / ^^^can beapproximated by 1 / ^ as described above. In this case the matrices comprising ^ and ^ canbe written as First note the exponents in these expressions that are associated with ^^^ and with ^^^.These distances respectively define the path length difference from Source 1 to the ears of listener 1, from Source 2 to the ears of listener 1, from Source 1 to the ears of listener 2, and from Source 2 to the ears of listener 2. Thus, these are given by Δ^ = (^^^ − ^^^), Δ^ = (^^^ − ^^^),Δ^ = (^^^ − ^^^), Δ^ = (^^^ − ^^^).There are a further eight path length differences given by Γ^ = (^^^ − ^^^), Γ^ = (^^^ − ^^^), Γ^ = (^^^ − ^^^), Γ^ = (^^^ − ^^^),Γ^ = (^^^ − ^^^), Γ^ = (^^^ − ^^^), Γ^ = (^^^ − ^^^), Γ^ = (^^^ − ^^^),all of which are associated with ^^^and the negative values of these path length differencesare associated with ^^^. Now turn to the matter of finding the eigenvalues of ^^^. First notethat using the block matrix structure of ^^^ shows that the eigenvalues are given by the rootsof where ^^is the two-by-two identity matrix. It can also be shown that if the block matrices are square and of the same order
[0018] the determinant of a block matrix having the form of ^^^is given by det ^^ ^^ ^^ = det(^ + ^) det(^ − ^),and the eigenvalues of ^^^can therefore be found from the roots ofdet[(^ + ^) − ^^^)] det[(^ − ^) − ^^^] = 0.The eigenvalues can be found by using the same process as described above in the two-source single-listener case. Consider the first term in the above product. Setting this to zero amountsto finding the eigenvalues of (^ + ^). First the matrices involved can be written as^ = ^^^^^^^^ ^^^^^^ ^ , ^ = ^ ^^^^^^^^^^^^ and the matrix sum in the form In this case therefore the eigenvalues are given by (^^ − ^^)]^At an optimally conditioned frequency, these two eigenvalues must be equal. As noted above, this implies that In terms of the elements of the matrices ^ and ^ this therefore implies[(^^^ + ^^^) + (^^^ + ^ ^^^)] − 4[(^^^ + ^^^)(^^^ + ^^^) − (^^^ + ^^^)(^^^ + ^^^)] = 01 ^^,^= 2[(^^^ + ^^^) + (^^^ + ^^^)]Similarly, one can write expressions for the eigenvalues of the matrix difference (^ − ^). Inthis case the condition for equal eigenvalues can be written as [(^^^ − ^^^) + (^^^ − ^^^)]^ − 4[(^^^ − ^^^)(^^^ − ^^^) − (^^^ − ^^^)(^^^ − ^^^)] = 0 ^^^) + (^^^ − ^^^)]First note that optimal conditioning requires that ^^,^ = ^^,^. This therefore requires that ^^^ +^^^ = 0 . Now note that using this relationship, together with the observation that ^^^ = ^^^ =4, can be used in adding the above two conditions for optimality to show, after some algebra, that ^^^^^^ + ^^^^^^ = 0Similarly, subtraction of the two conditions for optimality shows that ^^^^^^ + ^^^^^^ = 0Furthermore, adding these conditions can then be used to show that (^^^ + ^^^)(^^^ + ^^^) = 0Also note that from the matrices ^ and ^, it follows that ^ = ^∗^^ ^^ and that ^^^ = ^^∗^ . This shows that under optimal conditions it must follow that ^^^ + ^^^ = 0Also note that the condition ^^^^^^ + ^^^^^^ = 0 can then be written as^∗^^^^^ − ^^^(−^∗^^ ) = 0which shows that 2|^^^|^ = 0 the condition for optimal conditioning is therefore given simplyby |^ |^^^ = 0Now note that the term ^^^ is given by adding the relevant terms in the matrices ^^^and ^^^defined above. This shows that and therefore, the condition for optimality depends on the four path length differences from the four sources to the ears of the two listeners. Again, the condition for the modulus squared of ^^^to be equal to zero implies that both the real and imaginary parts of ^^^must be zero. Thereforecos^Δ^ + cos^Δ^ + cos^Δ^ + cos^Δ^ = 0sin^Δ^ + sin^Δ^ + sin^Δ^ + sin^Δ^ = 0It thus appears that there are very similar conditions for optimality as in the case of the two- source single-listener case discussed above. One approach, for example would be to add two pairs of terms in each of the above expressions and seek conditions of optimality exactly as suggested above. It is helpful to undertake some further numerical investigations. Simulations can be used to evaluate both the effect of listener spacing and the effect of source spacing in the determination of condition numbers. A plot of the variation of condition number with frequency is shown in Figure 37 for a particular geometrical arrangement of sources and listeners. This is clearly a complicated function of frequency, but in the simulations that follow, the numerical simulations are intended to find the frequency andvalue of the condition number at the first minimum as a function of frequency. The elementsof the matrix ^ were again assumed to be approximated by using 1 / ^^^ = 1 / ^. The value of^ for these simulations is assumed to be 3m. It is also assumed that the value of the headdiameter (2^ in Fig.1) was 0.2m. For these simulations, the listener spacing is fixed and asearch was again conducted for the first minimum in the condition number of the matrix ^ asit varies as a function of frequency. For each chosen listener spacing, the minimum conditionnumber is evaluated over a range of values of both the spacing ^ between the inner sourcesand the spacing ℎ between the inner and outer sources, as illustrated in Figure 36. The resultsof this simulation are shown in Figure 38 which shows the variation in condition number forlistener spacings of (a) 0.4m, (b) 0.6m (c) 0.8m (d) 1.0m and (e) 1.175m. There are distinctpatterns in the values of ^ and ℎ that produce minimum condition numbers. However, it isonly for listener separations ^ of 0.4 m, 0.8m and 1.175m that condition numbers are veryclose to unity, although for other listener separations of ^ of 0.6m and 1.0m, there is a widevalley of low condition numbers (in the region of 1.5 - 2.0). The dependence on frequency ofthe unity condition numbers is illustrated in Figure 44 for source spacings ^ of 0.4m, 0.8mand 1.175m respectively. The ratio of optimal ratio of ℎ / ^ also depends on listener spacing,as exemplified by the slope of the lines tracing the minimum condition number inFigures. 38(a), 38(c) and 38(e). Examination of the geometry of these arrangements leadingto unity condition number shows that, in almost all cases, |Δ^| + |Δ^| ≈ ^ / 2 and that |Δ^| +|Δ^| ≈ ^ / 2 although in this case, generally |Δ^| ≪ |Δ^| and is often close to zero. However,note that the sum of the magnitudes of the path length differences consistently add to one wavelength ^. The results disclosed firstly suggest that the criteria for optimality given above (in terms of sums of four cosine functions and four sine functions) are indeed satisfied by thosegeometries that give rise to values of condition number that are very close to unity. This hasbeen verified by evaluating each of the sine and cosine terms for a given source and listener geometry and shown that their sum is equal to zero. There is also the suggestion from the simulations that the sums of cosine and sine terms might be written as and that the conditions for optimality are given by both Thus, using similar arguments to those presented above in the two-source single-listener case, these suggest optimality criteria given by which appear to be consistent with the results of the numerical simulations undertaken to date. The above findings may advantageously assist in the design of practical source arrays for given listener spacing. Here, by way of illustration, the focus will be on designing for aspecific listener spacing of ^ = 0.8m which is typical of the spacing between two seatedlisteners in a domestic situation (watching TV from a couch for example). In any practical application, it is helpful to minimize the number of sources used. As demonstrated above, the frequencies associated with low condition numbers vary with source spacing. As a general principle, low frequencies require widely spaced sources, whilst high frequencies require narrowly spaced sources. One method of designing such an arrangement is to define theseparation distance ^ that will ensure good conditioning at the high end of the frequencyrange. For example, a practical choice of an inner source spacing is ^ = 0.2m. For the listenerpositions chosen it follows (see Fig 38(c)) that the spacing of the outer sources should begiven by ℎ = 0.33^ = 0.066m. This specifies the high frequency array. The outer sources canthen be used to form the inner sources of a lower frequency array with ^ = 0.2 + 2(0.066) =0.33m. The process is repeated sequentially, with increasing source spacings being used for progressively lower frequencies. An example of the geometry of an array designed in this manner shown in Figure 40, and the corresponding variation of condition number withfrequency for each sub-array is shown in Figure 41. In Figure 40 multiple spaced apart soundsources (designated by circles are spaced apart in a single row in a sound reproduction unit 20), and the sound sources in operation synthesise five arrays, 301, 302, 303, 304 and 305.The array 305 emits the highest frequency band, and array 301 emits the lowest frequencyband. The arrays 302 to 304 emit distinct frequency bands intermediate of the uppermostband and the lowermost band. In this implementation, audio outputs relating to differentarrays are provided by a respective single speaker in the unit 20, as shown by the broken line mapping in Figure 40. It is noteworthy that each said speakers of the unit 20 conforms to the required geometrical positioning, and in particular in relation the speakers which provide output for multiple arrays. The variation in condition number for the above array when the listeners are moved to adistance ^ = 2m from the array is shown in Figure 42. Generally, this shows that the matrixremains well-conditioned. However, there is a distinct shift in the minima associated with the condition numbers, with the high frequency array exhibiting a fall in operational bandwidth, although this is compensated by a fall in the lowest frequency of low condition number. Notethat in these computations, the terms (1 / ^^^) in the matrix ^ are again approximated by(1 / ^). Numerical simulations demonstrate that there is a negligible difference in the results from making this assumption. The introduction of an additional source in the centre of the array can prove to be beneficial. In this case the optimization problem is aimed at minimizing the sum of squared source strengths with the constraint of ensuring the desired crosstalk at the ears of the two listeners.The condition number of the matrix ^ in this case is shown as a function of frequency inFigure 43 (a). Figure 43(b) also shows for comparison the same condition number plot for the array as that shown in Figure 42, but with the plots truncated at a value of condition number of four. This demonstrates that the frequency range of lower condition numbers can be made broader by the addition of the centre source. The form of the sound field generated at three of the optimally conditioned frequencies is illustrated in Figure 44. The sound field has the highly desirable property of well-defined zones of constructive and destructive interference in the region of the listeners. The scattering effect of the head may need to be considered, but we understand that there is little influence on the basic physical properties of the sound field. Despite being at different frequencies, these three sound fields all have very similar radiation patterns, all of which appear to be able to accommodate an appropriately spaced number of listeners. These radiation patterns are also very similar to those produced by the optimal source distribution. This source distribution is known to have a frequency independent radiation pattern when the component monopole sources are optimally spaced. REFERENCES [1] B. B. Bauer, Stereophonic earphones and binaural loudspeakers, Journal of the Audio Engineering Society 9 (2) (1961) 148–151. [2] B. S. Atal, M. R. Schroeder, United States Patent 3,236,949a, Apparent sound source translator (1966). [3] P. Damaske, Head-related two-channel stereophony with loudspeaker reproduction, Journal of the Acoustical Society of America 50 (4B) (1971) 1109–1115. [4] Y. Ando, S. Shidara, Z. Maekawa, K. Kido, Some basic studies on the acoustic design of room by computer, Journal of the Acoustical Society of Japan 29 (1973). [5] S. Sakamoto, T. Gotoh, T. Kogure, M. Shimbo, A. Clegg, Controlling sound-image localisation in stereophonic reproduction, Journal of the Audio Engineering Society 29 (11) (1981) 794–799. [6] H. Hamada, Construction of orthosterophonic system for the purposes of quasi-insitu recording and reproduction, Journal of the Acoustical Society of America 39 (5) (1983) 337– 348. [7] J. L. Bauck, D. H. Cooper, Generalised transaural stereo and applications, Journal of the Audio Engineering Society 44 (9) (1996) 683–705. [8] O. Kirkeby, P. A. Nelson, H. Hamada, F. Orduna-Bustamante, Fast deconvolution of multichannel systems using regularisation, IEEE Transactions on Speech and Audio Processing 6 (2) (1998) 189–194. [9] E. C. Hamdan, Theoretical advances in multichannel crosstalk cancellation systems, Ph.D. thesis, University of Southampton (2020).
[0010] T. Takeuchi, P. A. Nelson, Optimal source distribution for binaural synthesis over loudspeakers, Journal of the Acoustical Society of America 112 (6) (2002) 2786–97. doi:10.1121 / 1.1513363. URL https: / / www.ncbi.nlm.nih.gov / pubmed / 12508999
[0011] T. Takeuchi, M. Teschl, P. A. Nelson, Objective and subjective evaluation of the optimal source distribution for virtual acoustic imaging, Journal of the Audio Engineering Society 55 (11) (2007) 981–987.
[0012] P. Mannerheim, P. A. Nelson, Virtual sound imaging using visually adaptiveloudspeakers, Acta Acustica united with Acustica 94 (2008) 1024 – 1039.
[0013] D. G. Morgan, T. Takeuchi, K. R. Holland, Off-axis cross-talk cancellation evaluation of 2-channel and 3-channel Opsodis soundbars, in: Proceedings of Reproduced Sound 2016, Southampton UK, 2016.
[0014] L. A. T. Haines, T. Takeuchi, K. R. Holland, Investigating multiple off-axis listening positions of an Opsodis.
[0015] M. Yairi, T. Takeuchi, K. R. Holland, D. G. Morgan, L. Haines, Binaural reproduction capability for multiple off-axis listeners based on the 3-channel optimal source distribution principle, in: Proceedings of the 23rd International Congress on Acoustics, Aachen, Germany, 2019.
[0016] Y. Kim, O. Deille, P. A. Nelson, Crosstalk cancellation in virtual acoustic imaging systems for multiple listeners, Journal of Sound and Vibration 297 (1-2) (2006) 251–266. doi:10.1016 / j.jsv.2006.03.042.
[0017] B. Masiero, X. Qiu, Two listeners crosstalk cancellation system modelled by four point sources and two rigid spheres, Acta Acustica united with Acustica 95 (2) (2009) 379–385.
[0018] Philip Nelson, Takashi Takeuchi, Sound reproduction. US Patent 11,337,001 B2 (2022)
[0019] P.A. Nelson, T. Takeuchi, P. Couturier, X. Zhou, Sound field control for multiple listener virtual imaging, Journal of Sound and Vibration 539 (2022) 117259
[0020] G. E. Golub, C. F. V. Loan, Matrix Computations, Third Edition, The John Hopkins University Press, 1996.
[0021] R.A. Horn and C.R. Johnson, Matrix Analysis, Second Edition, Cambridge University Press (2013) Ch.4 pp 263, 278.
[0022] A. Bunse-Gerstner and W.B. Cragg, Singular value decompositions of complex symmetric matrices, Journal of Computational and Applied Mathematics 21 (1988) 21-54
[0023] Arnav Kundu, Mohammad Samragh Razlighi, Minsik Cho, Priyanka Padmanabhan,Devang Naik. “HeiMDaL: Highly efficient method for the detection and localization of wakewords.” Proceedings of ICASSP 2023. Rhodes Island, Greece, 2023, pp. 1-5.
[0024] S. Panchapagesan, M. Sun, A. Khare, S. Matsoukas, A. Mandal, B. Hoffmeister, and S. Vitaladevuni, “Multi-task learning and weighted cross-entropy for DNN-based keyword spotting,” in INTERSPEECH, 2016.
[0025] M. Sun, D. Snyder, Y. Gao, V. Nagaraja, M. Rodehorst, S. Panchapagesan, N. Strom, S. Matsoukas, and S. Vitaladevuni, “Compressed time delay neural network for small-footprint keyword spotting,” 082017, pp. 3607–3611.
[0026] M. Sun, A. Schwarz, M. Wu, N. Strom, S. Matsoukas, and S. Vitaladevuni, “An empirical study of cross-lingual transfer learning techniques for small-footprint keyword spotting,” 201716th IEEE International Conference on Machine Learning and Applications (ICMLA), pp. 255–260, 2017.
[0027] K. Kumatani, S. Panchapagesan, M.Wu, M. Kim, N. Strom, G. Tiwari, and A. Mandal, “Direct modeling of raw audio with DNNs for wake word detection,” 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp. 252–257, 2017.
[0028] J. Guo, K. Kumatani, M. Sun, M. Wu, A. Raju, N. Strom, and A. Mandal, “Time-delayedbottleneck highway networks using a DFT feature for keyword spotting,” 04 2018.
[0029] M.Wu, S. Panchapagesan, M. Sun, J. Gu, R. Thomas, S. N. Prasad Vitaladevuni, B. Hoffmeister, and A. Mandal, “Monophone-based background modeling for two-stage on- device wake word detection,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5494–5498.
[0030] G. Chen, C. Parada, and G. Heigold, “Small-footprint keyword spotting using deep neural networks,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014, pp. 4087–4091.
[0031] T. N. Sainath and C. Parada, “Convolutional neural networks for small-footprint keyword spotting.” in INTERSPEECH, 2015, pp. 1478–1482.
[0032] Vlad S. Paul and Philip A. Nelson, 2024, “Efficient design of complex-valued neural networks with application to the classification of transient acoustic signals” The Journal of The Acoustical Society of America, 156(2), 1099-1110
Claims
CLAIMS 1. A sound reproduction apparatus comprising a plurality of spaced-apart sound sources, and the sound sources providing multiple arrays comprising a first array arranged to emit at least a first acoustic frequency range and a second array arranged to emit at least a second acoustic frequency range; and wherein the first array is arranged to emit a sound frequency range which is higher than the frequency range emitted by the second array, and further wherein a lateral extent of the first array is less than a lateral extent of the second array, and yet further wherein the lateral extent of each of the multiple arraysdecreases as the frequency band output by an array increases.2.The apparatus as claimed in claim 1 which is arranged to generate binaural sound reproduction for at least two listeners.
3. The apparatus as claimed in claim 1 or claim 2 in which at least one of or some of the sound sources is arranged to emit multiple frequency bands. 4.The apparatus as claimed claim 1 or claim 2 in which the sound sources are arranged in a single group configuration to emit the first frequency band and the second frequency band, and the configuration comprises a single group / set arranged in horizontal alignment.
5. The apparatus as claimed in any of claims 1, 2 or 3, in which the sound sources are grouped in a layout of at two or three vertically spaced rows / alignments.
6. The apparatus as claimed in claim 5 in which some of the sound sources in one or more rows output two frequency bands.
7. The apparatus as claimed in claim 5 or 6 in which at least some of the sound sources in different groups / rows are substantially vertically aligned.
8. The apparatus as claimed in any of claims 5 to 7 in which a first row has a first spacing between sound sources and a second row has a second spacing between sound sources, and the first spacing is inferior to the second spacing.
9. The apparatus as claimed in claim 8 which comprises a third row of sound sources which has a third spacing, and the third spacing is greater than the second spacing.
10. The apparatus as claimed in any preceding claim in which the first array and the second array have midpoints which are substantially aligned, wherein the midpoint of an array is the geometrical location of substantially midway between distal-most sound sources of the array.
11. The apparatus as claimed in in any preceding claim in which the first array and the second array have the same number of sound sources arranged to output the respective frequency band.
12. The apparatus as claimed in any preceding claim in which the cumulative effective number of sound sources of the arrays is greater than the number of physical sound sources.
13. The apparatus as claimed in any preceding claim in which the sound sources are configured to synthesize a greater number of sound sources of multiple arrays.
14. The apparatus as claimed in any preceding claim which comprises a signal processor which is arranged to generate drive signals to the arrays, based on an input signal representative of a sound recording or sound data which is to be reproduced.
15. The apparatus as claimed in claim 14 in which where a sound source is arranged to output multiple frequency bands, the sound source is arranged to receive a drive signal which combines each or some of the multiple frequency bands.
16. A sound reproduction apparatus which comprises multiple sound sources, and further comprises a signal processor to generate drive signals to the sound sources from a sound recording so as to generate virtual sound images to each of two or more listeners, and the apparatus further arranged to determine listener position information of each of the listeners using sound emitted directly or indirectly by each of the listeners or from respective listener positions, and to then preferably tailor characteristics of signal processing applied to the audio material to be reproduced, which characteristics take account of the determined listener position information.
17. The apparatus as claimed in claim 16 which is arranged to receive a listener localization sound which includes uttering or vocalizing one or more sounds or words.
18. The apparatus as claimed in claim 16 or claim 17 which is arranged to receive a listener localization sound using an emitter device.
19. The apparatus as claimed in claim 18 in which the listener localization sound is emitted at or near to the head of a listener.
20. The apparatus as claimed in any of claims 16 to 19 which comprises one or more microphones for receiving listener localization sounds.
21. The apparatus as claimed in any of claims 16 to 20 in which the listener position information may be used to determine characteristics of one or more filters used in the generation of loudspeaker drive signals.
Citation Information
Patent Citations
Sound reproduction
US11337001B2
Apparent sound source translator
US3236949A
Sound reproduction system
US20190090060A1
Sound reproduction
US20210152938A1
Audio device and method for producing a sound field
US20240098418A1