System and method for controlling the sound stage rendered by loudspeakers
The SSC processor addresses sound stage inaccuracies in reflective environments by using matched filters and frequency-dependent time-windowing, ensuring stable sound localization and improved sound quality in vehicle cabins with asymmetric loudspeaker setups.
Patent Information
- Application Number
- JP2025551899
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-08
- Filing Date
- 2024-03-05
- Publication Date
- 2026-03-06
AI Technical Summary
In acoustically reflective environments, such as vehicle cabins, existing digital filters for controlling the sound stage rendered by distributed loudspeakers face challenges due to reflection sensitivity, listener position changes, and asymmetric loudspeaker placement, leading to inaccurate and aesthetically undesirable sound localization and tonal characteristics.
A sound stage control (SSC) processor is designed using time-matched and level-matched filters, combined with frequency-dependent time-windowing and regularized pseudoinverse transformation, to stabilize the sound stage despite listener movement and loudspeaker asymmetry, enabling effective crosstalk cancellation and real-time head-tracking.
The SSC processor enhances sound stage depth, symmetry, and height control, maintaining stable sound localization and reducing computational complexity, even in reflective environments with unevenly positioned loudspeakers.
Smart Images

Figure 2026507881000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 489,124, filed March 8, 2023, entitled "System and Method for Controlling the Soundstage Rendered by Loudspeakers Distributed Arbitrarily in a Reflective Environment," the disclosure of which is incorporated herein by reference in its entirety.
[0002] The present disclosure relates to a system and method for designing a sound stage control (SSC) processor to control the sound stage rendered by an array of distributed loudspeakers in an acoustically reflective environment. [Background technology]
[0003] Without proper processing, audio rendered from multiple loudspeakers in a highly reflective environment, such as the interior of a vehicle cabin, can result in an inaccurate and often asymmetric "sound stage," resulting in sound quality with undesirable tonal and temporal characteristics. As used herein, "sound stage" generally refers to the spatial region perceived by one or more listeners into which an audio system renders sound sources. An inaccurate sound stage can result in a listener being unable to accurately localize the rendered sound source, or, even if the sound source is clearly localizable, its perceived location can be aesthetically undesirable. An asymmetric sound stage can generally refer to a sound stage in which the left and right ends of the sound stage are perceived asymmetric to the listener. Undesirable tonal characteristics can be quantified as deviations (in amplitude and / or phase) from a target or desired frequency response as measured at any listening position.
[0004] Digital filters can be used to process audio signals and control the sound stage rendered by loudspeakers. The basic principle is to use one or more digital filters to process the audio signal sent to each loudspeaker so that the acoustic response measured at each listener's ear matches the desired target response. Such filters are called sound stage control (SSC) filters. A major challenge is that filter performance degrades in the presence of reflections (the highly reflective environment of a vehicle cabin is a typical example). While it is possible in principle to design filters to compensate for reflections, the sound stage and tonal characteristics of sounds reproduced with such filters are highly sensitive to changes in listening position. Even slight deviations from the exact listening position can, in effect, change the arrival times of reflections at the intended listener's ears, resulting in unacceptable acoustic response at the ears and / or loss of sound stage control (SSC) performance. While this can be addressed by using cameras and tracking software to track the listener's position and dynamically (i.e., in real time) updating the filters accordingly, the large variations in reflection arrival times and levels necessitate either 1) switching between many filters to cover small ranges of listener movement, or 2) frequent real-time filter updates. This not only increases computational and implementation cost and complexity, but also makes it more difficult to effectively implement spatial audio techniques such as crosstalk cancellation and real-time filter interpolation (for rendering three-dimensional audio with only two channels) because the filter responses are smeared in time, as discussed below.
[0005] Another challenge arises when loudspeakers are asymmetrically positioned relative to each listener and have limited operating bandwidths. This is common, for example, in automobile cabins, where practical considerations such as mounting space limitations and safety dictate that individual loudspeaker sizes be limited. Therefore, the distance from the tweeter to the listener's ear may be significantly different from the distance from the midrange and low-frequency drivers. This results in impulse responses (IRs) between individual loudspeakers and the listener's ears that are less consistent in time. In other words, exciting a single channel (e.g., the left channel of a stereo input) may result in IRs that are less clear in time than those obtained from exciting a conventional stereo loudspeaker setup, for example, when all drivers (e.g., woofers and tweeters) for each speaker are closely spaced (e.g., housed in the same cabinet). Such smeared IRs not only make equalization (a necessary step for sound stage control) more difficult, but also result in sound stage control (SSC) filters themselves being temporally smeared, making advanced solutions such as crosstalk cancellation (for rendering spatial audio with only two channels) and real-time filter interpolation for head-tracked acoustic response compensation for increased performance robustness to head movement more difficult to implement.
[0006] A key challenge in designing an SSC filter is selecting the time window necessary to include or exclude reflections. On the one hand, longer windows tend to include more reflections and therefore increase the energy of the delay time (i.e., the impulse response moves from zero more significantly in the delay time direction), resulting in: 1) increased latency; 2) greater CPU and RAM requirements; 3) less robustness to head movement; and 4) the requirement for more aggressive equalization to ensure acceptable sound quality outside the sweet spot, leading to a loss of dynamic range. Longer windows also tend to reduce the perceived distance to the soundstage when using anechoic target responses (e.g., HRTFs). Therefore, controlling the soundstage distance with longer windows requires room information (whether measured or modeled), which increases complexity and the potential for other errors (e.g., increased processing requirements when using real-time room models, or a lack of individualization when using measured room responses). On the other hand, shorter windows tend to maintain the perceived distance to the soundstage because they introduce fewer reflections. However, short windows limit the spectral control necessary to meet the target spectral response, especially at low frequencies. Summary of the Invention
[0007] The embodiments disclosed herein each have multiple aspects, no single one of which is solely responsible for the desirable attributes of the disclosure. Without limiting the scope of the disclosure, its more prominent features will now be briefly described. After considering this description, and particularly after reading the section entitled "Detailed Description of the Invention," one will appreciate how the features of the embodiments described herein offer advantages over existing digital filter designs.
[0008] Embodiments of the present invention relate to systems and methods for designing a sound stage control (SSC) processor that provides effective SSC in reflective environments while largely avoiding many of the drawbacks discussed above. Such SSC also facilitates 1) existing crosstalk cancellation (which enables rendering spatial / 3D audio through two-channel signals, improving the perceived depth of the sound stage), and 2) effective implementation of dynamic filtering using head-tracking techniques. Embodiments also include an SSC processor and a vehicle equipped with such a processor.
[0009] Some embodiments include methods for designing a digital signal processor to control the spatial characteristics (depth, azimuthal extent, height, and symmetry) of a sound stage rendered by loudspeakers in an acoustically reflective environment, even when the loudspeakers are unevenly positioned in space and operate in different spectral bands, an example of which is the array of speakers in a modern car cabin. Depth, as used herein, refers to the perceived distance between the listener's head and the sound stage (perceived depth can also be enhanced by spatial audio rendering techniques such as crosstalk cancellation). Embodiments can also be used to individually control multiple sound stages for multiple listeners in the same space (e.g., passengers in a car). Embodiments can also take into account the movement of one or more listeners to ensure that the sound stage remains stable even in the presence of such movement.
[0010] In some embodiments, the method comprises the following five general steps: I) Design a set of time-matched and level-matching filters for an array of arbitrarily distributed loudspeakers, which constitutes the first stage of signal processing (the matching stage). II) Measure the response of the system at the listener's ear with the filter applied. III) Preprocessing the measured (or simulated) response of a matched stage filter by time-windowing the reflections in a frequency-dependent manner, thereby exploiting the reflection effect and providing effective control of the perceived sound stage depth. IV) Using the preprocessed response of the matched filter with the spatial target response (either head-related transfer function (HRTF) or measured acoustic response), the second stage (SSC) filter is designed by regularized least squares pseudoinverse transformation. V) By cascading the two stages of filters (after optional SSC filter equalization), the final SSC processor is obtained, which allows effective control of sound stage depth, azimuthal spread, symmetry, and height.
[0011] In some embodiments, the method includes the following two additional steps. I) Design a final bank of SSC processors for a separate set of listener head positions (including combinations of positions if there are multiple listeners). II) Using cameras and tracking software to track the movement of one or more listeners' heads, and dynamically (i.e., in real time) updating the active final SSC processors by interpolating between two or more processors in a row based on the head-tracking data.
[0012] One element of the aforementioned method is the frequency-dependent time windowing of reflections in step III (which can be equivalently achieved using complex smoothing), which allows soundstage distance control by correcting for low-frequency components of the direct sound, early reflections, and subsequent reflections (which contribute to room modes that are easier to control) while exploiting the higher-frequency components of subsequent reflections, which are more difficult to equalize and more sensitive to listener position. This windowing step also enables the simultaneous activation of anechoic spatial targets, such as HRTFs, and soundstage distance control. Without such windowing, distance control can only be achieved using spatial targets, which also include room responses, and thus suffer from many of the drawbacks of the prior art mentioned above. Furthermore, SSC filters designed based on the matching stage in step I and the windowed responses in step III are more compact in time, allowing their interpolation to be performed more accurately and efficiently in real time, a key step in dynamically updating filters based on head position via head tracking.
[0013] Any of the features of an embodiment may be applicable to all embodiments identified herein. Furthermore, any of the features of an embodiment may be independently combined in any way, partially or wholly, with other embodiments described herein, for example, one, two, or three or more embodiments may be combined in whole or in part. Furthermore, any of the features of an embodiment may be optional with respect to other embodiments.
[0014] One aspect is a method for programming a sound stage control (SSC) processor to control a sound stage rendered by an array of arbitrarily distributed loudspeakers in an acoustically reflective environment, comprising the steps of: (a) designing matched stage filters to obtain time-matched and level-matched responses from different loudspeakers as measured at specific control points in the environment; (b) measuring the responses to the same control points from the different loudspeakers with the matched stage filters applied; (c) preprocessing the responses by time-windowing the impulse response in a frequency-dependent manner to correct one or more of the direct sound, early reflections, and low-frequency components of the subsequent reflections while excluding high-frequency components of the subsequent reflections from the correction; (d) inverting the preprocessed IR resulting from step (c) using a regularized pseudo-inverse transform and cascading with a spatial target response to obtain a filter for the SSC stage; and (e) programming the SSC processor by cascading the matched stage and SSC stage. Some embodiments include a sound stage control processor programmed according to any of the above methods. [Brief explanation of the drawings]
[0015] [Figure 1A] The application of the SSC processor is demonstrated to provide an improved sound stage for two listeners in an acoustically reflective environment using an array of loudspeakers typical of a car cabin.
[0016] [Figure 1B] FIG. 1 is an example flow diagram of a process for designing an SSC digital filter.
[0017] [Figure 2] 1 is a flowchart illustrating an example of a general signal flow including a matched stage filter and an SSC stage filter configured in an embodiment of a sound stage control filter.
[0018] [Figure 3] 1 is a flowchart illustrating an example of a method for designing a matched stage filter.
[0019] [Figure 4] 1 is a flowchart illustrating an example of a method for designing an SSC stage filter.
[0020] [Figure 5] 5 is a flow chart illustrating an example of substeps involved in performing step 32 of FIG. 4.
[0021] [Figure 6] 5 is a flow chart illustrating an example of substeps involved in performing step 34 of FIG. 4.
[0022] [Figure 7] 5 is a flow chart illustrating an example of substeps involved in performing step 36 of FIG. 4. DETAILED DESCRIPTION OF THE INVENTION
[0023] The following detailed description of certain embodiments presents various descriptions of specific embodiments. However, the innovations described herein can be embodied in many different ways, for example, as defined and encompassed by the claims. This description refers to the drawings, where like reference numbers and / or terminology may indicate identical or functionally similar elements. It will be understood that the elements depicted in the drawings are not necessarily drawn to scale. Furthermore, it will be understood that certain embodiments can include more elements and / or a subset of the elements depicted in the drawings. Furthermore, some embodiments can incorporate any suitable combination of features from two or more drawings.
[0024] Embodiments of the present invention relate to a sound stage control (SSC) processor and associated method for designing digital filters implemented in such processors for processing audio, particularly in vehicle embodiments. For example, the method includes five steps, detailed below, resulting in the design of two stages of digital filters that, when combined, result in an SSC processor that provides effective spatial control of a sound stage for multiple listeners in an acoustically reflective environment using an arbitrary distributed array of loudspeakers 110A-110I (see FIG. 1A). FIG. 1A illustrates an application in which listeners 102A and 102B perceive a sound stage consisting of left (L), center (C), and right (R) images from an array of N loudspeakers 110A-110I in a vehicle cabin. Listener 1 102A perceives sound stage L1, C1, and R1 (70), and listener 2 102B perceives sound stage L2, C2, and R2 (71). As will be described below, the process of designing a filter may use fewer or more individual steps, or may combine certain steps together, without departing from the spirit of the invention.
[0025] 1B is a flow diagram of a process for designing a digital filter. The process for designing a digital filter disclosed herein can include five steps, as shown in FIG. 1B. These steps are shown as examples; fewer or more individual steps can be used to design a digital filter, or one or more steps can be combined based on a specific application. In this disclosure, Roman numerals are used to indicate the five general steps of the method, and Arabic numerals are used to indicate the detailed steps, some of which are optional or may be implemented differently in different embodiments of the method described below.
[0026] Step I (162): Design Matched Stage Filters: The matched stage filters are linear-phase bandpass filters (one for each loudspeaker 110A-110I) with predetermined passbands, with passband gain and group delay set so that all loudspeakers 110A-110I located to the left and right of all listeners 102A, 102B result in an acoustic response with the same overall delay and average level when measured at the left and right ears of the left-most and right-most listeners 102A / 102B, respectively. For all other loudspeakers not located to the left and right of listeners 102A, 102B, the gain and delay are set so that the acoustic response, when further averaged, results in the same overall delay and average level when measured at the ears of all listeners closest to each loudspeaker 110A-110I. In both cases, each listener is assumed to be located in a predetermined standard listening position. For example, in a vehicle, the listeners are assumed to be located in the front seats.
[0027] Step II (164): Measuring the response with matched stage filters applied: This step II (164) involves measuring the transfer functions between the different loudspeakers and the same control points (usually the listener's ear positions) used in step I (162), with matched stage filters applied.
[0028] Step III (166): Frequency-Dependent Time-Windowing of Reflections: In this step III (166), the impulse response obtained from step II (164) is processed by time-windowing the reflections in a frequency-dependent manner, or equivalently by using methods such as Complex Fractional Octave Smoothing (CFOS). This part of the method can correct early components (direct sound, early reflections, etc.) but progressively low-pass filter subsequent components, excluding high frequencies (which are more difficult to equalize and more sensitive to listener position) from the correction.
[0029] Step IV (168): Use of Regularized Pseudoinverse Transform with Spatial Target Response Design Filter: The SSC stage filter is then calculated using the impulse response obtained from Step III via a regularized least-squares method. The spatial target response is specified as either a head-related transfer function (HRTF) or a measured acoustic response, and may optionally apply cross-path constraints. The SSC stage filter is then post-processed by applying either a minimum-phase or linear-phase equalization filter to ensure that the response corresponding to the mono input (or stereo input with perfectly correlated left and right channels, or the center channel of a multi-channel input) matches the target spectral response.
[0030] Step V (170) Cascading Filters: The overall SSC signal processor is then implemented as a cascade of the matched stage filters and SSC stage filters described above.
[0031] Step I (162) is used to make the final SSC stage filter more compact in time. Step III (166) effectively controls the perceived depth of the sound stage while maintaining temporal compactness. The spatial targets introduced in Step IV (168) allow control of the perceived width, symmetry, and height of the sound stage, while the regularization optimization in that step ensures maximized SSC performance with minimal coloration, making the SSC processor largely immune to the drawbacks listed above for the prior art.
[0032] In some embodiments, the above five steps (162-170) are repeated for various positions of the listener's head, creating the bank of cascaded filters described in step V (170). Real-time linear interpolation is then used to smoothly change the filters based on the listener's head movements tracked by any head-tracking technology, although any other form of interpolation may be used. Instead of interpolating between banks of filters, a single set of cascaded filters may be updated in real time without departing from the spirit of the present invention.
[0033] One embodiment of the method of the present invention is illustrated in the SSC processor shown in Figure 2. In step 10, the acoustic responses between each of the N loudspeakers and Q control points (e.g., the entrances of each listener's ear canal) are measured. These responses are represented in the frequency domain as a QxN transfer function (TF) matrix B. The measurements are performed using an exponential sine sweep, as described, for example, in the article "A new approach to impulse response measurements at high sampling rates" by Joseph G. Tylka, Rahulram Sridhar, Braxton B. Boren, and Edgar Y. Choueiri et al., presented at the 137th Audio Engineering Society Conference in 2014, and in U.S. Patent No. 9,959,883. Matrix B is then used to calculate a matched stage filter, represented by an NxN filter matrix A, in step 12. Finally, the overall response of A cascaded with B, which may be determined either by measurement or simulation, along with a predetermined target response, is used to calculate an SSC stage filter, represented by an N×P filter matrix S, in step 14. The overall SSC processor is then implemented as a cascade of filter matrices S and A, and is therefore an N×P filter matrix, where P is the number of input channels (e.g., P=2 for a stereo input).
[0034] To calculate the filter matrix A, the first step involves processing the transfer function (TF) of B, as shown in step 20 of Figure 3. More specifically, the corresponding IR (typically calculated from B via an inverse fast Fourier transform) is windowed in time (typically using a raised cosine window) and then grouped according to the position of each loudspeaker relative to all listeners. The grouping of loudspeakers is performed so that all speakers to the left and right of all listeners belong to their own group (one on the left and one on the right), while all other speakers are grouped according to their proximity to individual control points, which is application-specific (i.e., depending on the number and relative positions of the loudspeakers and control points).
[0035] Next, in step 22, delays and gains of time-aligned and level-matching responses are calculated, both within and between groups, with the goal of normalizing the average level and onset of each loudspeaker IR as measured at appropriately selected control points. For example, step 22 may include calculating a time-aligned delay 22A and a level-matching gain 22B. This is done so that the IRs between the individual loudspeakers and the listener's ears are, on average, more time-coherent, an important step for achieving soundstage control filters with more compact time-domain responses. Such responses are desirable, for example, because they enable better-performing crosstalk cancellation filters (for rendering spatial audio using only two channels).
[0036] The delay is calculated based on the onset of the IR, estimated, for example, as the first time the absolute amplitude of the IR exceeds 20% of its absolute peak amplitude (the so-called "IR thresholding method"), although any other technique for extracting or estimating the delay may be used. For example, the absolute amplitude of the IR may exceed 10%, 15%, 20%, 25%, 30% or more of its absolute peak amplitude and still be within the scope of the present invention. The gain is calculated as the linear average of the response amplitude within a predetermined passband of the loudspeaker, although any other equivalent metric may be used.
[0037] From the calculated passband gains and group delays of the individual loudspeakers, level-matching gains and time-matching delays are calculated. Typically, these gains and delays are applied so that the acoustic responses for all loudspeakers to the left and right of all listeners have the same (or nearly the same) overall delay and average level when measured at the left and right ears of the left-most and right-most listeners, respectively. For all other loudspeakers, gains and delays are applied so that the acoustic responses, when further averaged with the right and left ears of the left-most and right-most listeners, have the same (or nearly the same) overall delay and average level. Finally, as shown in step 24, the calculated gains and delays are applied to a zero-phase bandpass filter (apply delay 24B and apply gain 24C) (24A) to generate the desired matched stage filter matrix A (24D). In both cases, each listener is considered to be in a predetermined standard listening position. However, alternative groupings of loudspeakers, listener placements, and gain and delay applications may be possible.
[0038] To calculate the filter matrix S, the first step involves measuring or simulating the response with the matched stage filter applied. This is shown as step 30 in FIG. 4, where measuring A involves performing acoustic measurements as described above, where the signal chain includes a filter for A, while simulating A involves convolving the filter for A with the response at B (in the time domain). These responses, along with responses corresponding to selected spectral and spatial targets, are preprocessed as described in step 32 in FIG. 4. For example, step 32 can include preprocessing of a spectral target IR32A, a simulated or measured IR32B, and a spatial target IR32C. The preprocessing in this step consists of several substeps, as shown in FIG. 5.
[0039] A typical example of the first two of these substeps involves applying a bandpass filter with a passband between 30 Hz and 18 kHz (see step 40, e.g., filtering the spectral target IR40A, filtering the measured or simulated IR40B, and filtering the spatial target IR40C), followed by applying an initial Tukey window of length approximately 341 ms (see step 42). This filtering and windowing is used to perform a subsequent complex smoothing operation, shown as step 44. Filtering allows for a more accurate estimation of IR onset and removal of any overall delay in the IR using the thresholding technique described above (operations performed as part of step 22 shown in FIG. 3), which is a step used prior to complex fractional-octave smoothing according to a paper by Panagiotis D. Hatziantoniou and John N. Mourjopoulos et al. entitled "Generalized fractional-octave smoothing of audio and acoustic responses," published in Volume 48, Number 4, of the Journal of Audio Engineering Society, April 2000. Windowing is performed to reduce the computation time of the smoothing.
[0040] The complex smoothing implemented in step 44 is an improved version of that described in the paper by Hatziantoniou and Mourjopoulos in that it preserves log-frequency symmetry, which is achieved using the algorithm described in the paper "A Generalized Method for Fractional-Octave Smoothing of Transfer Functions that Preserves Log-Frequency Symmetry" by Joseph G. Tylka, Braxton B. Boren, and Edgar Y. Choueiri, et al., published in Volume 65, Issue 3 of the Journal of Audio Engineering, March 2017.
[0041] As described in Hatziantoniou and Mourjopoulos's paper, complex smoothing (performed in the frequency domain) has the time-domain effect of gradually windowing and blocking subsequent reflections, emphasizing the high-frequency components of these reflections more strongly. The greater the amount of smoothing, the more pronounced the windowing effect. As an alternative or supplement to smoothing, an additional design window may optionally be applied to the measured or simulated matching stage response. This is step 46 in Figure 5, depicted by a dotted block, indicating that it is an optional step.
[0042] One aspect of the present invention is the filter designer's ability to control the degree of complex smoothing / optional windowing applied, which directly affects the perceived depth of the sound stage. This is because increasing the smoothing or applying optional design windows will remove some of the reflections in the simulated or measured matched stage response, and therefore leave them uncorrected in the design of the SSC stage filter, as described below. Therefore, instead of traditional approaches that attempt to completely correct for reflections (i.e., implementing room response equalization), the method of the present invention exploits some of these reflections in a frequency-dependent manner to contribute to controlling the depth of the sound stage. An advantage of CFOS is that reflections are partially corrected in a frequency-dependent manner, allowing for sound stage depth control by correcting the low-frequency components of the direct sound, early reflections, and subsequent reflections (which contribute to room modes that are easier to control) while exploiting the high-frequency components of subsequent reflections, which are more difficult to equalize and are more sensitive to listener position. It is important to note that CFOS is equivalent to using multiple time windows, each operating in a different frequency band, and this step of the present invention can be implemented using either the method or another equivalent method.
[0043] Step 48 of Figure 5 is another optional step that involves applying either a minimum-phase or linear-phase equalization filter to the spatial targets IR. The equalization filter is designed to equalize either the amplitude response of an individual spatial target or an average response calculated from one or more individual responses to the corresponding amplitude response of a spectral target. This step is optional because equalization to a spectral target can alternatively be performed exclusively as part of step 36 shown in Figure 4.
[0044] Returning to FIG. 4, once all IRs are preprocessed in step 32, intermediate SSC stage filters are designed in step 34. This is shown in FIG. 6, where the processed simulated / measured IRs 602 are divided into groups 606 and further filtered 608 according to predetermined available passbands. The same filters are applied to the spatial target IRs (e.g., processed spatial target IRs 604), and then intermediate SSC stage filter matrices are designed for each group using a regularized least-squares minimization method. Once the SSC stage filter matrices for each group are calculated, these filters are combined into a single SSC stage filter by applying an appropriately selected Linkwitz-Riley crossover filter.
[0045] The final step in the design of the SSC stage filters is step 36, which post-processes the intermediate filters as shown in FIG. 4. FIG. 7 shows a flowchart of the post-processing substeps applied to these filters. In step 60, from the processed spectral target IR 604, an average response M is estimated (using the process shown in FIG. 5), which represents the response perceived by all listeners to a mono input (or a stereo input with perfectly correlated left and right channels, or the center channel of a multi-channel input). Something similar occurs in step 62, but this time starting with a simulation or acoustic measurement of the overall response with the matched stage and intermediate SSC stage filters applied in cascade. For example, matched stage filter matrix A 702 and intermediate SSC stage filter matrix 704 are applied in cascade and simulated in step 706. Both of these average responses are then used to calculate the equalization filter in step 64. The equalization filter first applies frequency-dependent regularization to the average response TIFF2026507881000002.tif76 and cascading the resulting inverse filter with the mean response M. Next, a minimum-phase or linear-phase version of the equalization filter is applied to the intermediate SSC stage filter matrix 704 (step 708). These steps allow the response of the SSC stage filter to be corrected with minimal or no effect on the perceived spatial characteristics of the sound stage. In step 66, a level matching gain is applied to ensure that switching the SSC stage filter on and off does not cause a noticeable change in the perceived overall loudness of the reproduced audio. The SSC stage filter matrix 710 is generated in step 710. This level matching gain is calculated by estimating the perceived loudness with and without the SSC stage filter applied using a loudness-based binaural model described in the paper "Analysis of Virtual Sound Source Attributes Using a Binaural Auditory Model" by Ville Pulkki, Matti Karjalainen, and Jyri Huopaniemi et al., published in Volume 47, Issue 4 of the Journal of Audio Engineering, and essentially calculating the ratio of these quantities.
[0046] Finally, the matching stage and SSC stage filters are cascaded to produce the desired SSC processor, which may be incorporated into a vehicle, for example, to provide an improved performance soundstage. Furthermore, this integration may be in the form of a single SSC processor that is updated either statically or dynamically (whether by interpolation from a bank of SSC processors or not) based on tracking data representing the listener's instantaneous head position.
[0047] In the foregoing specification, the present disclosure has been described with reference to specific embodiments. It will, however, be apparent that various modifications and changes can be made thereto without departing from the broader spirit and scope of the present disclosure. The specification and drawings are therefore to be regarded in an illustrative rather than a restrictive sense.
[0048] Indeed, while the present disclosure is directed to specific embodiments and examples, it will be understood by those skilled in the art that the present invention extends beyond the specifically disclosed embodiments to other alternative embodiments and / or uses of the present invention and equivalents thereof. Additionally, while several variations of the embodiments have been shown and described in detail, other modifications within the scope of the present disclosure will be readily apparent to those skilled in the art based on this disclosure. It is also contemplated that various combinations or subcombinations of specific features and aspects of the embodiments may be made and still fall within the scope of the present disclosure. It should be understood that various features and aspects of the disclosed embodiments can be combined with or substituted for one another to form varying forms of the embodiments disclosed herein. Any methods disclosed herein need not be performed in the order described. Therefore, it is not intended that the scope of the present disclosure should be limited by the specific embodiments described above.
[0049] It will be understood that the systems and methods of the present disclosure each have multiple innovative aspects, no single one of which is solely responsible for or necessary for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure.
[0050] Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a particular combination and initially claimed as such, one or more features may be deleted from a claimed combination, and a claimed combination may refer to a subcombination or a variation of a subcombination. No single feature or group of features is necessary or essential to each embodiment.
[0051] It will also be understood that conditional language used herein, particularly "can," "could," "might," "may," "e.g.," and the like, is generally intended to convey that certain embodiments include certain features, elements, and / or steps, but not others, unless otherwise specified or understood within the context of use. Thus, such conditional language is generally not intended to suggest that features, elements, and / or steps are somehow essential in one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps should be included in or performed in any particular embodiment, with or without author input or prompting. Terms such as "comprising," "including," and "having" are synonymous and are used inclusively and without limitation, and do not exclude additional elements, features, acts, operations, etc. Additionally, the term "or" is used in an inclusive sense (and not an exclusive sense); for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Additionally, the articles "a," "an," and "the," as used in this application and the appended claims, should be interpreted to mean "one or more" or "at least one," unless otherwise specified. Similarly, while operations are depicted in the figures in a particular order, it should be recognized that such operations need not be performed in the particular order or sequential order depicted, nor need all depicted operations be performed, to achieve desirable results. Furthermore, the figures may schematically illustrate one or more exemplary processes in flowchart form. However, other operations not depicted may be incorporated into the schematically illustrated exemplary methods and processes. For example, one or more additional operations may be performed before, after, concurrently with, or between any of the depicted operations.Additionally, operations may be rearranged or reordered in other embodiments. Multitasking and parallel processing may be advantageous in certain situations. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. Additionally, other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve desirable results.
[0052] Furthermore, the methods and apparatuses described herein are susceptible to various modifications and alternative forms, specific examples of which are shown in the drawings and described in detail herein. However, it should be understood that the disclosure is not limited to the particular forms or methods disclosed; rather, the disclosure encompasses all modifications, equivalents, and alternatives falling within the spirit and scope of the various implementations described and the appended claims. Furthermore, any particular feature, aspect, method, attribute, property, quality, attribute, element, etc., disclosed herein relating to an implementation or embodiment can be used in all other implementations or embodiments described herein. Any method disclosed herein need not be performed in the order described. The methods disclosed herein may include specific actions performed by a practitioner; however, the methods may also include, explicitly or implicitly, any third-party instructions regarding those actions. Ranges disclosed herein also encompass all overlaps, subranges, and combinations thereof. Expressions such as "up to," "at least," "greater than," "less than," "between," and the like, include the recited numbers. Numbers preceded by terms such as "about" or "approximately" are inclusive of the stated number and should be interpreted in accordance with the context (e.g., as precisely as reasonably possible under the circumstances, e.g., ±5%, ±10%, ±15%, etc.). Phrases preceded by terms such as "substantially" are inclusive of the stated phrase and should be interpreted in accordance with the context (e.g., as closely as reasonably possible under the circumstances). For example, "substantially constant" includes "constant." Unless otherwise specified, all measurements are at standard conditions, including temperature and pressure.
[0053] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single members. By way of example, "at least one of: A, B, or C" is intended to encompass A, B, C, A and B, A and C, B and C, and A, B, and C. Transposed language such as "at least one of X, Y, and Z" is understood in its commonly used context to convey that an item, term, etc., can be at least one of X, Y, or Z, unless otherwise specified. Thus, such transposed language is generally not intended to suggest that a particular embodiment requires that at least one of X, at least one of Y, and at least one of Z, respectively, be present. Headings provided herein, if any, are for convenience only and do not necessarily affect the scope or meaning of the apparatus and methods disclosed herein.
[0054] Thus, the scope of the claims is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the present disclosure, principles, and novel features disclosed herein.
[0055] Many variations and modifications may be made to the above-described embodiments, and the elements thereof should be understood as being among other acceptable examples. All such modifications and variations are intended to be within the scope of the present disclosure. The above description details specific embodiments. However, no matter how detailed the foregoing appears in text, it will be understood that the systems and methods can be implemented in many ways. Also, as noted above, the use of a particular term in describing a particular feature or aspect of the systems and methods should not be construed as meaning that the term is redefined herein to include only any specific characteristics of the feature or aspect of the systems and methods to which the term pertains.
Claims
1. 1. A method for programming a sound stage control (SSC) processor to control a sound stage rendered by an array of arbitrarily distributed loudspeakers in an acoustically reflective environment, comprising: (a) designing matched stage filters to obtain time-matched and level-matched responses from different loudspeakers; (b) measuring the response to the same control point from the different loudspeakers with the matched stage filter applied; (c) pre-processing the impulse response by time-windowing the response in a frequency-dependent manner to correct one or more of the direct sound, early reflections, and low frequency components of subsequent reflections while excluding high frequency components of subsequent reflections from said correction; (d) inverting the preprocessed IR obtained from step (c) using a regularized pseudo-inverse transform and a cascade of spatial target responses to obtain the filter for the SSC stage; (e) programming an SSC processor by cascading the matching stage and the SSC stage; A method comprising:
2. The method of claim 1 , wherein the matched stage filter in step (a) is obtained using a measured transfer function.
3. The method of claim 1 , wherein the matched stage filter in step (a) is obtained using a calculated or numerically simulated transfer function.
4. The method of claim 1 , wherein the matched stage filter in step (a) is based on measurements of sound pressure level and delay.
5. The method of claim 1 , wherein the matched stage filter in step (a) is modified by acoustic measurements.
6. The method of claim 1 , wherein the response measured in step (b) is obtained from a calculation or a numerical simulation.
7. The method of claim 1 , wherein the preprocessing in step (c) includes complex fractional octave smoothing.
8. The method of claim 1 , wherein the preprocessing in step (c) comprises applying different time windows applied to various frequency bands.
9. The method of claim 1 , wherein the spatial target responses used in step (d) are measured or calculated using head-related transfer functions (HRTFs) of a human or dummy head.
10. The method of claim 1 , wherein the spatial target response used in step (d) is a measured or calculated acoustic response.
11. The method of claim 1 , wherein the regularized pseudo-inverse transform in step (d) is a constant.
12. The method of claim 1 , wherein the regularized pseudo-inverse transform in step (d) is frequency dependent.
13. 2. The method of claim 1, wherein the SSC processor in step (e) is programmed through a single-stage filter obtained by convolving the matched stage filter with an SSC stage filter.
14. 2. The method of claim 1, wherein the SSC processor in step (e) is programmed by cascading a signal through the matching stage and the SSC stage.
15. The method of claim 1 , wherein the SSC processor in step (e) is a CPU, GPU, DSP, or FPGA, or part of a network of such devices.
16. The method of claim 1 , wherein the SSC processor in step (e) is implemented as software running on a computational cloud.
17. 10. The method of claim 1, wherein the SSC processor in step (e) is implemented as one or more analog circuits or a combination of digital and analog circuits.
18. The method of claim 1 , wherein the method is repeated for different head positions, synthesizing one or more banks of filters based on the different positions of each listener's head.
19. The method of claim 18 , wherein the one or more filters are synthesized by interpolation between filters in the sequence.
20. 20. The method of claim 18, wherein the different positions of each listener's head are determined using a camera and head tracking software.
21. 20. The method of claim 18, wherein the different positions of each listener's head are determined acoustically.
22. A sound stage control processor programmed according to the method of claim 1.
23. 20. A sound stage control processor programmed according to the method of claim 18.