Methods and apparatus for rendering spatial audio over loudspeakers using monoaural perceptual CUE filters
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-02
- Publication Date
- 2026-08-13
Smart Images

Figure US2026013455_13082026_PF_FP_ABST
Abstract
Description
D25008W001 RENDERING SPATIAL AUDIO OVER LOUDSPEAKERS USING MONAURAL PERCEPTUAL CUE FILTERSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from United States Provisional Patent Application No. 63 / 753,517, filed on 4 February 2025, European Patent Application No. 25187392.3, filed on 3 July 2025, and United States Provisional Patent Application No.63 / 867,873, filed on 21 August 2025, each of which is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] The disclosure pertains to systems and methods for rendering audio for playback by one or more loudspeakers.BACKGROUND
[0003] Audio devices, including but not limited to smart audio devices, have been widely deployed and are becoming common features of many homes, stores, offices and other environments. Although existing systems and methods for controlling audio devices provide benefits, improved systems and methods would be desirable.NOTATION AND NOMENCLATURE
[0004] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0005] Throughout this disclosure, including in the claims, “speaker” and “loudspeaker” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers.D25008W001
[0006] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).
[0007] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.SUMMARY
[0008] At least some aspects of the present disclosure may be implemented via methods, such as audio processing methods. In some instances, the methods may be implemented, at least in part, by a control system such as those disclosed herein. Some disclosed methods involve receiving, by a control system and via an interface system, audio data. The audio data may include one or more audio signals and associated spatial data. The spatial data may indicate an intended perceived spatial position corresponding to the one or more audio signals. Some disclosed methods involve rendering, by the control system, the audio data for reproduction via one or more loudspeakers of an environment, to produce rendered audio signals. Rendering each of the one or more audio signals included in the audio data may involve applying a monaural perceptual cue filter to each audio signal of the one or more audio signals. The monaural perceptual cue filter may be configured to insert one or more spectral perceptual cues corresponding to the intended perceived spatial position of a corresponding audio signal. The monaural perceptual cue filter may be configured to at least partially remove one or more spectral perceptual cues corresponding to at least one spatial position of the one or more loudspeakers for which the rendered audio signals are produced.
[0009] According to some examples, the method may involve obtaining each spatial position of the one or more loudspeakers. In some examples, the method may involve providing, by the control system and via the interface system, the rendered audio signals to the one or more loudspeakers.
[0010] In some examples, the monaural perceptual cue filter may be configured as a ratio of an audio signal perceptual cue corresponding to an intended perceived spatial positionD25008W001 of the one or more audio signals and a composite loudspeaker perceptual cue of at least one of the one or more loudspeakers associated with rendering the audio data. According to some examples, the audio signal perceptual cue may be based, at least in part, on a monaural head-related transfer function (HRTF) corresponding to the intended perceived spatial position of the one or more audio signals and the composite loudspeaker perceptual cue may be based, at least in part, on a weighted combination of monaural HRTFs corresponding to a spatial position of each of the one or more loudspeakers. In some examples, the monaural HRTF corresponding to the intended perceived spatial position of the one or more audio signals may be computed as a sum of squared magnitudes of a left ear HRTF and a right ear HRTF for the intended perceived spatial position. According to some examples, the method may involve computing, by the control system, a weighting for the weighted combination of the monaural HRTFs as a function of rendering filters for the one or more audio signals.
[0011] According to some examples, the monaural HRTF corresponding to the intended perceived spatial position of the one or more audio signals and the composite loudspeaker perceptual cue may each additionally include a diffuse sound component. In some such examples, a variable scaling of the diffuse sound component may control a strength of the monaural perceptual cue filter.
[0012] In some examples, rendering the audio data may be based at least in part on Center of Mass Amplitude Panning. Alternatively, or additionally, rendering the audio data may be based at least in part on Flexible Virtualization. According to some examples, rendering the audio data may be based on Flexible Virtualization for frequencies below a frequency threshold. In some examples, rendering the audio data may be based on Center of Mass Amplitude Panning for frequencies above the frequency threshold. According to some examples, the frequency threshold may be in a range from 1 kHz to 5 kHz, inclusive.
[0013] According to some examples, the one or more audio signals may be, or may include, one or more audio object signals. In some such examples, the audio data may include audio object metadata indicating the intended perceived spatial position of a corresponding audio object signal. In some examples, the intended perceived spatial position may be indicated by a channel of a channel-based audio format.
[0014] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include one or more memory devices such as those described herein, including but not limited to one or more random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented in oneD25008W001 or more non-transitory media having software stored thereon which, when executed by one or more devices, cause the one or more devices to perform one or more of the disclosed methods.
[0015] At least some aspects of the present disclosure may be implemented via apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus may include an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof. According to some examples, the control system may be configured to execute instructions stored on a non-transitory computer-readable storage medium. The instructions may cause the apparatus to perform one or more of the disclosed methods.
[0016] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A shows an example of a cone of confusion.
[0018] Figure 1B is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0019] Figure 2 shows loudspeaker positions according to a first example.
[0020] Figure 3 includes graphs that represent monaural Head Related Transfer Functions (HRTFs) corresponding to the three loudspeaker positions shown in Figure 2.
[0021] Figure 4 includes graphs that depict monaural perceptual cues corresponding to the loudspeaker and desired perceived audio object positions shown in Figure 2.
[0022] Figure 5 shows graphs that represent monaural perceptual cue filters computed according to the three components depicted in Figure 4.
[0023] Figure 6 shows the loudspeaker and desired perceived audio object positions for a second example.
[0024] Figure 7 shows graphs of loudspeaker monaural HRTFs for the loudspeaker positions shown in Figure 6.D25008W001
[0025] Figure 8 includes graphs that depict monaural perceptual cues corresponding to the loudspeaker and desired perceived audio object positions shown in Figure 6.
[0026] Figure 9 shows graphs that represent the monaural perceptual cue filters computed according to the three components depicted in Figure 8.
[0027] Figure 10 is a flow diagram that outlines one example of a method that may be performed by an apparatus or system such as those disclosed herein.DETAILED DESCRIPTION OF EMBODIMENTS
[0028] Amplitude panning is a common method for rendering spatial audio to an array of loudspeakers. Amplitude panning involves varying the amplitudes of panning gains for a given audio signal as a function of the desired spatial position of the audio signal such that the perceived position of the sound energy aligns with the target signal position. When the desired audio signal position lies between two speakers, this can be accomplished with pairwise panning, meaning only the two speakers on either side of the desired audio signal position participate in its rendering. When the desired signal position aligns with the physical location of a loudspeaker, only that loudspeaker participates in its rendering. Amplitude panning is an apt method for rendering to a myriad of speaker layouts.
[0029] Playback of spatial audio in a consumer environment has typically been tied to a prescribed number of loudspeakers placed in prescribed positions, such as positions corresponding to Dolby 5.1 or 7.1 surround sound. In these cases, content is authored specifically for the associated loudspeakers and encoded as discrete channels, one for each loudspeaker (e.g., Dolby Digital™, Dolby Digital Plus™, etc.) More recently, immersive, object-based spatial audio formats have been introduced (such as Dolby Atmos™) which break this association between the content and specific loudspeaker locations. Instead, the content may be described as a collection of individual audio objects, each with possibly time varying metadata describing the desired perceived location of these audio objects in three-dimensional space and, in some examples, other properties of the audio object. At playback time, the audio content is transformed into loudspeaker feeds by a renderer which, in some examples, is configured to adapt to the number and location of loudspeakers in the playback system.
[0030] Techniques for rendering spatial audio over loudspeakers typically do not accurately create perceived spatial impressions distant from the physical loudspeaker locations (for example, making a sound seem as if the sound is coming from behind a listener using loudspeakers in front of the listener). Trans-aural audio, which attempts to independently control the audio signals arriving at the left and right ears of a listener from the loudspeakers, is a common approach to this problem. This technique generally involvesD25008W001 synthesizing a binaural signal corresponding to the desired perceived spatial location and then feeding this binaural signal through a cross-talk canceller which inverts the inherent cross talk between the two ears from each loudspeaker.
[0031] The cross-talk cancellation process typically becomes unstable at higher frequencies, and practical systems often revert to simple amplitude panning above some cutoff frequency (e.g., a frequency in a range from 1kHz–5kHz). In the lower-frequency region of stable operation, a trans-aural system can effectively impart important interaural level (ILD) and interaural time differences (ITD) of the binaural signal critical to eliciting the desired spatial impression. When switching to a previously-disclosed rendering strategy at higher frequencies, however, features of the binaural signal in this high-frequency region are lost. High-frequency features tend to be spectral cues imparted by the pinna of the ear rather than by interaural differences imparted by the head. These high-frequency features are important for differentiating the location of sounds lying on a “cone of confusion”, a term for a contour in space along which ILD and ITD are roughly the same. An example of a cone of confusion is shown in Figure 1 A and is described below.
[0032] As will be disclosed in more detail below, in some disclosed methods discarded high-frequency binaural spectral cues may be approximated as a monaural filtering process and applied in combination with, or independently from, a trans-aural system to improve the perceived spatial impressions of existing spatial audio rendering systems.
[0033] Figure 1 A shows an example of a cone of confusion. Cones of confusion are three-dimensional (3D) surfaces around a listener’s head on which sound sources will impart equivalent interaural cues. Cones of confusion extend to the left from a listener’s left ear and to the right from a listener’s right ear. In the example shown in Figure 1 A, the origin of the coordinate system 105 is at the center of the listener’s head 130. Here, the listener’s head 130 is facing in the positive direction of the y axis, the positive direction of the x axis extends from the left of the listener’s head 130 and the positive direction of the z axis extends through the top of the listener’s head 130.
[0034] The open end of the cone of confusion 110 forms a circle 115. Any sound source lying on this circle 115 imparts approximately the same ILD and ITD to the listener. To disambiguate locations of sound sources along this circle, a listener relies on the unique spectral cues imparted by the pinna at higher frequencies. Absent these spectral cues, a listener may struggle to accurately localize sounds lying on circle 115.
[0035] Figure 1B is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 1B are merely providedD25008W001 by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 150 may be, or may include, a smart audio device that is configured for performing at least some of the methods disclosed herein. In other implementations, the apparatus 150 may be, or may include, another device that is configured for performing at least some of the methods disclosed herein, such as a laptop computer, a cellular telephone, a tablet device, a “smart home hub” that is configured to control multiple audio devices, etc. In some such implementations the apparatus 150 may be, or may include, a server.
[0036] In this example, the apparatus 150 includes an interface system 155 and a control system 160. The interface system 155 may, in some implementations, be configured for receiving audio data. The audio data may include audio signals that are to be reproduced by at least some speakers of an environment. The audio data may include one or more audio signals and associated spatial data. The spatial data corresponding to an audio signal may indicate the intended perceived spatial position of that audio signal. In some examples, the spatial data may be, or may include, audio object metadata. In some examples, the spatial data may correspond with a channel of a channel-based audio format. According to some examples, the intended perceived spatial position may be derived from the audio format, such as with higher-order Ambisonics (HO A) or other spherical harmonic based sound field representations.
[0037] The interface system 155 may be configured for providing rendered audio signals to at least some loudspeakers of the set of loudspeakers of the environment. The interface system 155 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
[0038] The interface system 155 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 155 may include one or more wireless interfaces. The interface system 155 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 155 may include one or more interfaces between the control system 160 and a memory system, such as the optional memory system 165 shown in Figure 1B. However, the control system 160 may include a memory system in some instances.
[0039] The control system 160 may, for example, include a general purpose single- or multi-chip processor, a digital signal processor (DSP), an application specific integratedD25008W001 circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0040] In some implementations, the control system 160 may reside in more than one device. For example, a portion of the control system 160 may reside in a device within a playback environment and another portion of the control system 160 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 160 may reside in a device within a playback environment and another portion of the control system 160 may reside in one or more other devices of the playback environment. For example, control system functionality may be distributed across multiple smart audio devices of a playback environment, or may be shared by an orchestrating device (such as what may be referred to herein as a smart home hub) and one or more other devices of the playback environment. The interface system 155 also may, in some such examples, reside in more than one device.
[0041] In some implementations, the control system 160 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 160 may be configured for receiving (e.g., via the interface system 155) audio data including one or more audio signals and associated spatial data. The spatial data may indicate a desired perceived spatial position corresponding to an audio signal.
[0042] According to some examples, the control system 160 may be configured for rendering, by the control system, the audio data for reproduction via one or more loudspeakers of an environment (such as a playback environment), to produce rendered audio signals. In some examples, rendering each of the one or more audio signals included in the audio data may involve applying a monaural perceptual cue filter to each audio signal of the one or more audio signals. The monaural perceptual cue filter may be configured to insert one or more spectral perceptual cues corresponding to the intended perceived spatial position of the one or more audio signals, and at least partially remove one or more spectral perceptual cues corresponding to each spatial position of the one or more loudspeakers for which the rendered audio signals are produced. In some examples, the control system 160 may be configured for providing, via the interface system 155, the rendered audio signals to the one or more loudspeakers.
[0043] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in theD25008W001 optional memory system 165 shown in Figure 1B and / or in the control system 160.Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, include instructions for controlling at least one device to process audio data. The software may, for example, be executable by one or more components of a control system such as the control system 160 of Figure 1B.
[0044] In some examples, the apparatus 150 may include the optional microphone system 170 shown in Figure 1B. The optional microphone system 170 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc.
[0045] According to some implementations, the apparatus 150 may include the optional loudspeaker system 175 shown in Figure 1B. The optional loudspeaker system 175 may include one or more loudspeakers. Loudspeakers may sometimes be referred to herein as “speakers.” In some examples, at least some loudspeakers of the optional loudspeaker system 175 may be arbitrarily located. For example, at least some speakers of the optional loudspeaker system 175 may be placed in locations that do not correspond to any standard prescribed speaker layout, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc. In some such examples, at least some loudspeakers of the optional loudspeaker system 175 may be placed in locations that are convenient to the space (e.g., in locations where there is space to accommodate the loudspeakers), but not in any standard prescribed loudspeaker layout.
[0046] In some implementations, the apparatus 150 may include the optional sensor system 180 shown in Figure 1B. The optional sensor system 180 may include one or more cameras, touch sensors, gesture sensors, motion detectors, etc. According to some implementations, the optional sensor system 180 may include one or more cameras. In some implementations, the cameras may be free-standing cameras. In some examples, one or more cameras of the optional sensor system 180 may reside in a smart audio device, which may be a single purpose audio device or a virtual assistant. In some such examples, one or more cameras of the optional sensor system 180 may reside in a TV, a mobile phone or a smart speaker.
[0047] In some implementations, the apparatus 150 may include the optional display system 185 shown in Figure 1B. The optional display system 185 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some instances, the optional display system 185 may include one or more organic light-emitting diode (OLED)D25008W001 displays. In some examples wherein the apparatus 150 includes the display system 185, the sensor system 180 may include a touch sensor system and / or a gesture sensor system proximate one or more displays of the display system 185. According to some such implementations, the control system 160 may be configured for controlling the display system 185 to present a graphical user interface (GUI), such as one of the GUIs disclosed herein.
[0048] According to some examples the apparatus 150 may be, or may include, a smart audio device. In some such implementations the apparatus 150 may be, or may include, a wakeword detector. For example, the apparatus 150 may be, or may include, a virtual assistant.
[0049] Spatial Audio Rendering System
[0050] A spatial audio rendering system may receive one or more audio object signals Oj, each signal i=1...N having an associated desired perceived spatial position d̄ᵢ. The rendering system may be configured to distribute each audio object signal Oj into a set of loudspeaker signals sⱼ, each signal j=1...M intended for playback over a loudspeaker located at a spatial position Sj with respect to an assumed or actual listening position. The distribution of audio object signals into loudspeaker signals may be represented generically as a timevarying filtering and summation of the audio object signals, e.g., as follows:
[0051] (1)
[0052] In Equation 1,t) represents a complex value for filtering audio object signal Oj(, t) into loudspeaker signal sy(, t) at frequency f and time t. Note that, in general, all loudspeaker signals may contain elements of all audio object signals, but an audio object signal may be excluded from a loudspeaker signal ift) = 0.
[0053] Configuration of the rendering filterst) may be represented generically as an algorithmic operator 3(v) that depends at least on the desired audio object position dj(t), here represented as possibly varying over time, and the set of loudspeaker positions {s7}, j=l... M
[0054] Hᵢⱼ(f,t) = ℑ⟨d̄ᵢ(t), {s̄ⱼ}⟩ (2)
[0055] At a high level, the goal of the algorithm is to compute rendering filters t) such that a listener perceives each audio object signal substantially at its intended position when the loudspeaker signals computed according to Equation 1 are auditioned over a set of loudspeakers at their specified positions. Numerous algorithms and techniques exist to achieve this goal. Some such methods implement a Center of Mass Amplitude Panner (CMAP), which computes broadband gains such that the center of mass of the loudspeakerD25008W001 positions weighted by their associated gains is equal to the desired audio object position. In this case, the rendering filters are real values constant across frequency. Vector Based Amplitude Panning (VBAP) is another broadband amplitude panning algorithm with similar perceptual properties. VBAP is described in V. Pulkki, “Virtual sound source positioning using vector base amplitude panning,” Journal of the Audio Engineering Society, vol. 45, no.6, pp. 456–466, 1997, which is hereby incorporated by reference. C. Q. Robinson, S. Mehta, and N. Tsingos, “Scalable Format and Tools to Extend the Possibilities of Cinema Audio,” SMPTE Motion Imaging Journal, vol. 121, no. 8, pp. 63–69, Nov. 2012, which is hereby incorporated by reference, outlines another broadband amplitude panner for spatial audio, but with constraints imposed on the allowed positions of the loudspeakers. Another rendering method is called Flexible Virtualization (FV), which is a generalized trans-aural system with constraints added to promote solutions that are sparse in the number of loudspeakers utilized for each object signal. With FV, the rendering filters are fully complex and variant across frequency. Additional details regarding CMAP, FV and another panning method that combines aspects of both, are provided below.
[0056] Given the stability issues at higher frequencies cited earlier with trans-aural systems, in some implementations, FV is used up to a threshold frequency, generally in a range from 1kHz–5kHz, inclusive, and above this threshold frequency CMAP is employed. While the disclosed methods are particularly well suited to this rendering strategy, the disclosed methods are equally applicable to all those listed above, as well as any others described by Equations 1 and 2.
[0057] As mentioned above, binaural control using loudspeakers is challenging at higher frequencies, and many rendering systems revert to alternate strategies in this region. Some disclosed implementations of the present disclosure involve imparting the resulting absent binaural cues in a manner that is compatible with a selected rendering system.
[0058] For the sake of simplicity moving forward, Equation 1 will be represented with the time and frequency variables removed from the rendering filters, the audio object signals, and the loudspeaker signals, with the understanding that all subsequent developments may operate in a time and frequency varying manner:
[0059] (3)
[0060] In general, the perceived spatial position of audio object signal i is dictated largely by the relative values of the rendering filter Hij across loudspeakers j. The rendering algorithm is generally carefully designed to properly set these relative values, and therefore any strategy for introducing the absent binaural perceptual cues should not disrupt this balance.D25008W001
[0061] In accordance with this requirement, some aspects of the present disclosure involve approximating the absent binaural perceptual cues as an additional monaural perceptual cue filter Mtthat is applied equally across all loudspeakers for audio object signal i, e.g., as follows:
[0062] sj=liHijMioi(4)
[0063] Because the filter is the same for each loudspeaker in such examples, the relative balance of the new composite filtering HijMi across loudspeakers is the same as the original rendering filteringand the perceived location of the audio object signal is largely undisturbed.
[0064] The FV trans-aural renderer relies on a set of Head Related Transfer Functions (HRTFs), indexed by perceived spatial position, to both synthesize the desired binaural signal at the listener’s ears and to model the acoustic transmission from the loudspeakers to the listener’s ears. The HRTF at the left and right ears for perceived position p may be represented as HRTFLp) and HRTFR(p), respectively. The set of HRTFs may be obtained in a variety of manners, for example, measurements from a single individual, an average over a set of individuals, a parametric model of a human head, etc. The University of California’s Center for Image Processing and Integrated Computing (CIPIC) HRTF Database, Release 1.1, October 21, 2001, is an example of a publicly available database of HRTF sets, including individuals, measurements from a dummy head, and an average over all measurements.
[0065] The monaural perceptual cue filterof the present disclosure may be derived from such a set of HRTFs. For this derivation, it is useful to define a monaural response at perceived position p as the sum of the squared magnitudes of the response at the left and right ears:
[0066] HRTFM{p) = \HRTFL{p)\2+ \HRTFR{p)\2(5)
[0067] This monaural response discards any interaural difference cues contained in the HRTF, but it maintains an approximation of higher-frequency spectral cues as an average between the left and right ears.
[0068] The main purpose of the monaural perceptual cue filteris to impart to the listener perceptual spectral cues associated with the desired perceived audio object position dj. However, the audio object signal is reproduced by the set of loudspeakers at positions {s weighted by the rendering filters Hij. Sounds corresponding to these signals traveling from the loudspeakers to the listener will impart their own spectral perceptual cues according to the acoustic transmission from each loudspeaker to the listener, and therefore the filtershould be configured to simultaneously remove the spectral perceptual cues caused by the loudspeakers. This process of imparting the spectral perceptual cues of the audio objectD25008W001 position and removing the spectral perceptual cues of the loudspeakers may be expressed as a ratio between a monaural representation of the audio object spectral perceptual cues Pg. and a composite monaural representation of the loudspeaker spectral perceptual cues dependent on the loudspeaker locations and weighted by the rendering filters yI Pd-
[0069] Mt= - - ‘ - (6)J (filN
[0070] The spectral perceptual cues of the audio object may be defined as the monaural HRTF response at the object position:
[0071] Pg. = HRTFM{di) (7)
[0072] The composite spectral perceptual cues of the loudspeakers may be defined as a weighted sum of the monaural HRTF responses at the loudspeaker positions, with the weighting given by the rendering filters Htj:
[0073] (8)= HRTFM{SJ}
[0074] With this definition of the loudspeaker composite spectral perceptual cues, the spectral perceptual cues from loudspeakers contributing more to therendering of the audio object will contribute more to the composite response.
[0075] Constructing the monaural perceptual cue filters according to Equations 6-8 yields reasonable results in general. However, combinations of deep notches and strong peaks in either Pg or P may result in extreme variations in Mtacross frequency that areperceptually unpleasant or unnatural sounding. To eliminate these artifacts, the ratio in Equation 6 may be augmented with regularization to smooth out any extreme variations across frequency.
[0076] A useful form of regularization is inspired by a simple model of acoustic transmission in a room, where the sound arriving at a listening position from a source is given by a direct sound component (sound travelling directly from the source to the listener) added with a diffuse sound component (sound arriving from all directions resulting from reflections in the room). In Equation 6, the numerator and denominator can be viewed as the direct sound from the object and the loudspeakers, respectively. To regularize, the same diffuse sound component Pdif may be added to both:aPoi+^-^Pdif
[0077] M,- =J c. UH.. \JD25008W001
[0078] In Equation 9, the parameter a controls a direct-to-diffuse sound ratio and may vary between 0 and 1, inclusive. Note that when a = 1, Equation 9 simplifies to the unregularized definition ofin Equation 8. When a =0, = 1, and therefore the filter has no effect. Viewed this way, a maybe interpreted as a “strength” control for the monaural perceptual cue filtersand may be referred to herein as a filter strength parameter. In practice, the inventor has found that values of a between 0.5 and 1 generally yield acceptable results, depending on the HRTF set employed.
[0079] Keeping in line with the definition of the audio object and loudspeaker direct sound components in Equations 7 and 8 as a function of monaural HRTF responses from the object and loudspeaker positions, the diffuse sound component may be defined as the average of the monaural HRTF responses across all possible positions p:
[0080] Pdif= f. HRTFM(p) (10)
[0081] Examples of Monaural Perceptual Cue Filters
[0082] Two examples are now presented of the monaural perceptual cue filters disclosed in the present invention. For these examples, both audio object and loudspeaker positions are constrained to lie on a unit circle around a listener at zero degrees elevation. Hence, position is specified purely as an azimuth angle, with 0 degrees directly in front of the listener, negative azimuth to the left of the listener, and positive azimuth to the right of the listener. Many HRTF sets are indexed by positions specified in spherical coordinates with the listener positioned at the center of the sphere, and these examples operate under this assumption. The disclosed methods can, however, operate with positions specified in any coordinate system, such as standard 3 -dimensional cartesian space, and are not limited to arrangements of loudspeakers that are equidistant from a listening location.
[0083] Figure 2 shows loudspeaker positions according to a first example. Figure 2 shows three loudspeakers positioned at azimuth angles -15, 30, and 180 degrees, with Loudspeakers 1 and 2 in front of the listener 210 and Loudspeaker 3 behind the listener 210. The desired perceived object position 215 for this example is specified as -90 degrees azimuth, directly to the left side of the listener 210. In these examples, the azimuth angles are measured with reference to axis 201, which passes through the center of the listener 210’s head, with zero degrees corresponding to the listener 210’s viewing direction. According to this example, the listener 210 is located at the center of a unit circle 205, upon which Loudspeakers 1, 2 and 3, as well as the desired perceived audio object position 215, are located.
[0084] Figure 3 includes graphs that represent the monaural HRTFs from Equation 5 corresponding to the three loudspeaker positions shown in Figure 2. The HRTF set used forD25008W001 this example is one that has been derived by averaging across measurements from a group of individuals. Noted is a significant variation at the higher frequencies, indicative of the variations imparted by the pinna of the ear to sound arriving from different positions. In addition, these variations are reasonably smooth across frequency, a consequence of averaging over numerous individuals.
[0085] Figure 4 includes graphs that depict monaural perceptual cues corresponding to the loudspeaker and desired perceived audio object positions shown in Figure 2. In this example, the graphs of Figure 4 correspond to the three monaural perceptual cues from Equation 9 used to compute the monaural perceptual cue filter Mt, which are the audio object cue Pg., the composite loudspeaker cueand the diffuse cue Pdif. As shown inEquation 8, the composite loudspeaker cue is computed as a weighted sum of themonaural loudspeaker HRTFs depicted in Figure 3, where the weighting is derived from the rendering filters Htj corresponding to the desired perceived audio object position of -90 degrees azimuth. In this example, a simple broadband amplitude panner, CMAP, is employed to derive the rendering filters H, and the resulting weightings for the three loudspeakers are 0.55 for Loudspeaker 1, 0 for Loudspeaker 2, and 0.45 for Loudspeaker 3. Note that because the desired perceived audio object position lies between Loudspeaker 1 and Loudspeaker 3, only the weightings for those two are non-zero. Loudspeaker 2 is not activated in this example and therefore does not contribute to the composite loudspeaker monaural perceptual cue.
[0086] Figure 5 shows graphs that represent the monaural perceptual cue filterscomputed according to Equation 9 from the three components depicted in Figure 4. In this example, the monaural perceptual cue filterswere computed for three different values of the filter strength parameter a 1.0, 0.75, and 0.5. With a strength of 1.0, the monaural perceptual cue filtershows the largest degree of variation, and as the filter strength parameter a decreases towards 0, the monaural perceptual cue filter approaches OdB, corresponding to= 1. Because the HRTF set from this example is inherently smooth due to averaging across individuals, the filter Mtis reasonably smooth across frequency, even at full strength. As such, a relatively high filter strength parameter value, such as a = 1.0, may be an appropriate strength for such an HRTF set.
[0087] Figure 6 shows the loudspeaker and desired perceived audio object positions for a second example. In this case, Loudspeakers 1, 2 and 3 are positioned at 0, -30, and 30 degrees azimuth, respectively. These positions correspond to the positions of the center, left, and right loudspeakers of a standard surround sound layout. For this second example, theD25008W001 desired perceived object position 615 is specified as 180 degrees, directly behind the listener 210.
[0088] Figure 7 shows graphs of loudspeaker monaural HRTFs for the loudspeaker positions shown in Figure 6. According to this example, the HRTF set employed is one measured from the author of this disclosure. Because the HRTF set does not involve averaging over numerous listeners, one immediately notices more rapid variations across frequency in the loudspeaker monaural HRTFs depicted in Figure 7. Employing the same broadband amplitude panner from the first example to derive the rendering filters H, the corresponding weights for the composite loudspeaker cue are 1 for Loudspeaker 1, 0 for Loudspeaker 2, and 0 for Loudspeaker 3.
[0089] Figure 8 includes graphs that depict monaural perceptual cues corresponding to the loudspeaker and desired perceived audio object positions shown in Figure 6. In this case, only the Loudspeaker 1 at 0 degrees azimuth is activated, and hence the composite loudspeaker perceptual cue depicted in Figure 8 is equal to the monaural HRTF of the first loudspeaker depicted in Figure 7.
[0090] Figure 9 shows graphs that represent the monaural perceptual cue filterscomputed according to Equation 9 from the three components depicted in Figure 8. In this example, the monaural perceptual cue filters Mtwere computed for three different values of the filter strength parameter a 1.0, 0.75, and 0.5. At full strength (a = 1.0), one notes some extreme and peaky variations across frequency, particularly above 8kHz, but these variations are smoothed out as the filter strength parameter a decreases. As such, with this individualized HRTF set, a lower filter strength parameter than that for the first example, e.g., a filter strength parameter of 0.75, may be appropriate.
[0091] Figure 10 is a flow diagram that outlines one example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of method 1000, like other methods described herein, are not necessarily performed in the order indicated. In some implementation, one or more of the blocks of method 1000 may be performed concurrently. Moreover, some implementations of method 1000 may include more or fewer blocks than shown and / or described. The blocks of method 1000 may be performed by a control system — such as the control system 160 that is shown in Figure 1B and described above — of one or more devices, which may be (or may include) an instance of the apparatus 150 of Figure 1B.
[0092] According to this example, block 1005 involves receiving, by a control system and via an interface system, audio data. In this example, the audio data includes one or more audio signals and associated spatial data. Here, the spatial data indicates an intendedD25008W001 perceived spatial position corresponding to at least one of the one or more audio signals. In some examples the spatial data may be, or may include, spatial metadata of an object-based audio format such as Dolby Atmos™. In some examples, the intended perceived spatial position may be represented as dj(t) or simply as Oj, as disclosed elsewhere herein. In some instances, the spatial data may be, or may correspond with, channels of a channel-based audio format such as a Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4 or Dolby 9.1 format.Accordingly, the intended perceived spatial position may correspond with a channel of a channel-based audio format, may correspond with metadata, or may correspond with both the channel and the metadata. In some examples, the intended perceived spatial position may be derived from the audio format, such as with higher-order Ambisonics (HO A) or other spherical harmonic based sound field representations.
[0093] In this example, block 1010 involves rendering, by the control system, the audio data for reproduction via one or more loudspeakers of an environment, to produce rendered audio signals. The environment may, for example, be a playback environment such as one or more rooms of a home, a retail location, a restaurant, an office or other business environment, etc. According to this example, rendering each of the one or more audio signals included in the audio data comprises applying a monaural perceptual cue filter to each audio signal of the one or more audio signals. The monaural perceptual cue filter may be an instance of the monaural perceptual cue filterthat is described in this disclosure. In this example, the monaural perceptual cue filter is configured to insert one or more spectral perceptual cues corresponding to the intended perceived spatial position of a corresponding audio signal, and at least partially remove one or more spectral perceptual cues corresponding to at least one spatial position of the one or more loudspeakers for which the rendered audio signals are produced.
[0094] According to some examples (as noted in optional block 1015, shown with a dashed outline), method 1000 may involve providing, by the control system and via the interface system, the rendered audio signals to the one or more loudspeakers.
[0095] In some examples, the monaural perceptual cue filter may be configured as a ratio of an audio signal perceptual cue corresponding to an intended perceived spatial position of the one or more audio signals and a composite loudspeaker perceptual cue of at least one of the one or more loudspeakers associated with rendering the audio data. For example, the process of imparting the spectral perceptual cues of the audio position and removing the spectral perceptual cues of the loudspeakers may be expressed as a ratio between the monaural representation of the audio object spectral perceptual cues Pg and a composite monaural representation of the loudspeaker spectral perceptual cues dependent onD25008W001 the loudspeaker locations and weighted by the rendering filters e.g.,asdescribed above with reference to Equations 6-10. Accordingly, in some examples the audio signal perceptual cue may be based, at least in part, on a monaural head-related transfer function (HRTF) corresponding to the intended perceived spatial position of the one or more audio signals and the composite loudspeaker perceptual cue may be based, at least in part, on a weighted combination of monaural HRTFs corresponding to a spatial position of each of the one or more loudspeakers.
[0096] According to some examples, the monaural HRTF corresponding to the intended perceived spatial position of the one or more audio signals may be computed as a sum of squared magnitudes of a left ear HRTF and a right ear HRTF for the intended perceived spatial position, e.g., as described herein with reference to Equation 5.
[0097] In some examples, method 1000 may involve computing, by the control system, a weighting for the weighted combination of the monaural HRTFs as a function of rendering filters for the one or more audio signals, e.g., as described herein with reference to Equation 8.
[0098] According to some examples, the monaural HRTF corresponding to the intended perceived spatial position of the one or more audio signals and the composite loudspeaker perceptual cue each additionally include a diffuse sound component. In some such examples, a variable scaling of the diffuse component controls a strength of the monaural perceptual cue filter. Some examples are described above with reference to Equations 9 and 10.
[0099] In some examples, rendering the audio data may be based, at least in part, on Center of Mass Amplitude Panning (CMAP). Alternatively, or additionally, rendering the audio data may be based, at least in part, on Flexible Virtualization (FV). In some examples, rendering the audio data may be based on FV for frequencies below a frequency threshold. In some such examples, rendering the audio data may be based on CMAP for frequencies above the frequency threshold. The frequency threshold may be in a range from 1 kHz to 5 kHz, inclusive.
[0100] According to some examples, method 1000 may involve obtaining, by the control system, loudspeaker location data indicating locations of the one or more loudspeakers. According to some such examples, the location data may indicate locations of the one or more loudspeakers relative to a desired listening position or area, an assumed listening position or area, or an actual listening position or area. According to some examples, method 1000 may involve obtaining the loudspeaker location data and / or location data corresponding to the listening position or area from a data structure stored in a memoryD25008W001 of, or accessible by, the control system. In other examples, method 1000 may involve determining the loudspeaker location data and / or location data corresponding to the listening position or area.
[0101] The loudspeaker location data and / or location data corresponding to the desired listening position or area may be obtained through numerous mechanisms known in the art. According to some examples, loudspeaker location data and / or location data corresponding to the desired listening position or area may be specified according to a standard loudspeaker layout, such as a Dolby 5.1 loudspeaker layout. In some such examples, a user could, for example, provide input to a device — such as an audio / video receiver (AVR) — indicating that that they have a set of loudspeakers in a Dolby 5.1 layout. The AVR may then assume a "canonical" Dolby 5.1 layout and apply one or more disclosed methods according to the corresponding loudspeaker layout.
[0102] In some applications, such as an automobile cabin, loudspeaker location data and location data corresponding to the desired listening position or area are fixed and can be physically measured, e.g. with a tape measure, or obtained from layout information such as computer assisted drafting CAD data.
[0103] In other examples, method 1000 may involve a more adaptable approach that can automatically detect these loudspeaker and / or user locations and orientations through a one-time setup procedure or even dynamically across time. In Hess, Wolfgang, Head-Tracking Techniques for Virtual Acoustic Applications, (AES 133rd Convention, October 2012), which is hereby incorporated by reference, numerous commercially available techniques for tracking both the position and orientation of a listener’s head in the context of spatial audio reproduction systems are presented. One particular example discussed is the Microsoft Kinect. With its depth sensing and standard cameras along with a publicly available software (Windows Software Development Kit (SDK)), the positions and orientations of the heads of several listeners in a space can be simultaneously tracked using a combination of skeletal tracking and facial recognition. Although the Kinect for Windows has been discontinued, the Azure Kinect developer kit (DK), which implements the next generation of Microsoft’s depth sensor, is currently available.
[0104] In U. S. Patent No. 10,779,084, entitled “Automatic Discovery and Localization of Speaker Locations in Surround Sound Systems,” which is hereby incorporated by reference, a system is described which can automatically locate the positions of loudspeakers and microphones in a listening environment by acoustically measuring the time-of-arrival (TOA) between each speaker and microphone. A listening position or area may be detected by placing and locating a microphone at a desired listening position (aD25008W001 microphone in a mobile phone held by the listener, for example), and an associated listening orientation may be defined by placing another microphone at a point in the viewing direction of the listener, e.g. at the TV. Alternatively, the listening orientation may be defined by locating a loudspeaker in the viewing direction, e.g. the loudspeakers on the TV.
[0105] In Shi, Guangi el al, Spatial Calibration of Surround Sound Systems including Listener Position Estimation, (AES 137thConvention, October 2014), which is hereby incorporated by reference, a system is described in which a single linear microphone array associated with a component of the reproduction system whose location is predictable, such as a soundbar a front center speaker, measures the time-difference-of-arrival (TDOA) for both satellite loudspeakers and a listener to locate the positions of both the loudspeakers and listener. In this case, the listening orientation is inherently defined as the line connecting the detected listening position and the component of the reproduction system that includes the linear microphone array, such as a sound bar that is co-located with a television (placed directly above or below the television). Because the sound bar’s location is predictably placed directly above or below the video screen, the geometry of the measured distance and incident angle can be translated to an absolute position relative to any point in front of that reference sound bar location using simple trigonometric principles. The distance between a loudspeaker and a microphone of the linear microphone array can be estimated by playing a test signal and measuring the time of flight (TOF) between the emitting loudspeaker and the receiving microphone. The time delay of the direct component of a measured impulse response can be used for this purpose. The impulse response between the loudspeaker and a microphone array element can be obtained by playing a test signal through the loudspeaker under analysis. For example, either a maximum length sequence (MLS) or a chirp signal (also known as logarithmic sine sweep) can be used as the test signal. The room impulse response can be obtained by calculating the circular cross-correlation between the captured signal and the MLS input. Fig. 2 of this reference shows an echoic impulse response obtained using a MLS input. This impulse response is said to be similar to a measurement taken in a typical office or living room. The delay of the direct component is used to estimate the distance between the loudspeaker and the microphone array element. For loudspeaker distance estimation, any loopback latency of the audio device used to playback the test signal should be computed and removed from the measured TOF estimate.
[0106] International Publication Number WO 2022 / 118072, entitled “Pervasive Acoustic Mapping,” which is hereby incorporated by reference, discloses additional methods for estimating loudspeaker location data and location data corresponding to the desired listening position or area. This disclosure describes multiple techniques that may be used inD25008W001 various combinations in order to provide automated acoustic mapping. The acoustic mapping may be pervasive and ongoing. Such acoustic mapping may sometimes be referred to as “continuous,” in the sense that the acoustic mapping may be continued after an initial set-up process and may be responsive to changing conditions in the audio environment, such as changing noise sources and / or levels, loudspeaker relocation, the deployment of additional loudspeakers, the relocation and / or re-orientation of one or more listeners, etc. Some disclosed methods involve generating calibration signals that are injected (e.g., mixed) into the audio content being rendered by audio devices in an audio environment. In some such examples, the calibration signals may be, or may include, acoustic direct sequence spread spectrum (DSSS) signals. In other examples, the calibration signals may be, or may include, other types of acoustic calibration signals, such as swept sinusoidal acoustic signals, white noise, “colored noise,” such as pink noise (a spectrum of frequencies that decreases in intensity at a rate of three decibels per octave), acoustic signals corresponding to music, etc.
[0107] United States Patent Application Publication No. 2023 / 0040846 Al, entitled “Audio Device Auto-Location,” which is hereby incorporated by reference, discloses additional methods for estimating loudspeaker location data and location data corresponding to the desired listening position or area. Some disclosed methods for estimating an audio device location in an environment involve obtaining direction of arrival (DOA) data for each audio device of a plurality of audio devices in the environment and determining interior angles for each of a plurality of triangles based on the DOA data. Each triangle has vertices that correspond with audio device locations. The method involves determining a side length for each side of each of the triangles, performing a forward alignment process of aligning each of the plurality of triangles produce a forward alignment matrix and performing a reverse alignment process of aligning each of the plurality of triangles in a reverse sequence to produce a reverse alignment matrix. A final estimate of each audio device location is based, at least in part, on values of the forward alignment matrix and values of the reverse alignment matrix. Some such methods may yield a result that is correct up to an unknown scale and rotation. In many applications, absolute scale is unnecessary, and rotations can be resolved by placing additional constraints on the solution. For example, some multi-speaker environments may include television (TV) speakers and a couch positioned for TV viewing. After locating the speakers in the environment, some methods may involve finding a vector pointing to the TV and locating the speech of a user sitting on the couch by triangulation. Some such methods may then involve having the TV emit a sound from its speakers and / or prompting the user to walk up to the TV and locating the user’s speech by triangulation. Some implementations may involve rendering an audio object that pans around theD25008W001 environment. A user may provide user input (e.g., saying “Stop”) indicating when the audio object is in one or more predetermined positions within the environment, such as the front of the environment, at a TV location of the environment, etc. According to some such examples, after locating the speakers within an environment and determining their orientation, the user may be located by finding the intersection of directions of arrival of sounds emitted by multiple speakers. Some implementations involve determining an estimated distance between at least two audio devices and scaling the distances between other audio devices in the environment according to the estimated distance.
[0108] As can be seen, there exist numerous mechanisms through which loudspeaker location data and location data corresponding to the desired listening position or area may be obtained, and all such methods (as well as relevant future methods that may be developed) are meant to be applicable to the implementations of the present disclosure. Accordingly, the specific details disclosed herein should merely be regarded as examples.
[0109] Further Details Regarding CMAP and FV
[0110] As noted above, existing flexible rendering techniques include CMAP and FV. From a high level, both of these techniques render a set of one or more audio signals, each with an associated desired perceived spatial position, for playback over a set of two or more speakers, where the relative activation of speakers of the set is a function of a model of perceived spatial position of said audio signals played back over the speakers and a proximity of the desired perceived spatial position of the audio signals to the positions of the speakers. The model ensures that the audio signal is heard by the listener near its intended spatial position, and the proximity term controls which speakers are used to achieve this spatial impression. In particular, the proximity term favors the activation of speakers that are near the desired perceived spatial position of the audio signal. For both CMAP and FV, this functional relationship is conveniently derived from a cost function written as the sum of two terms, one for the spatial aspect and one for proximity:C(g)—Cspatial(£> {■${}) + CproximityCS’ O, {■$[}) (H)
[0111] Here, the set {s denotes the positions of a set of AT loudspeakers, o denotes the desired perceived spatial position of the audio signal, and g denotes an AT dimensional vector of speaker activations. For CMAP, each activation in the vector represents a gain per speaker, while for FV each activation represents a filter (in this second case g can equivalently be considered a vector of complex values at a particular frequency and a different g is computed across a plurality of frequencies to form the filter). The optimal vector of activations is found by minimizing the cost function across activations:gopt = mm C(g, 6, {sj) (12a)D25008W001
[0112] With certain definitions of the cost function, it is difficult to control the absolute level of the optimal activations resulting from the above minimization, though the relative level between the components of goptis appropriate. To deal with this problem, a subsequent normalization of goptmay be performed so that the absolute level of the activations is controlled. For example, normalization of the vector to have unit length may be desirable, which is in line with a commonly used constant power panning rules:(12b>
[0113] The exact behavior of the flexible rendering algorithm is dictated by the particular construction of the two terms of the cost function, Cspatiaiand Cproximity. For CMAP, Cspatiaiis derived from a model that places the perceived spatial position of an audio signal playing from a set of loudspeakers at the center of mass of those loudspeakers’ positions weighted by their associated activating gainsg, (elements of the vector g):(13)
[0114] Equation 13 may then be manipulated into a spatial cost representing the squared error between the desired audio position and that produced by the activated loudspeakers:2 2 csp«ttaI(g. a. K}) = = 11^131 (5 - ^)11 (14)
[0115] With FV, the spatial term of the cost function is defined differently. There the goal is to produce a binaural response b corresponding to the audio object position 6 at the left and right ears of the listener. Conceptually, b is a 2x1 vector of filters (one filter for each ear) but is more conveniently treated as a 2x1 vector of complex values at a particular frequency. Proceeding with this representation at a particular frequency, the desired binaural response may be retrieved from a set of HRTFs indexed by object position:b = HRTF{p} (15)
[0116] At the same time, the 2x1 binaural response e produced at the listener’s ears by the loudspeakers is modelled as a 2xAf acoustic transmission matrix H multiplied with the Afxl vector g of complex speaker activation values:e = Hg (16)
[0117] The acoustic transmission matrix H is modelled based on the set of loudspeaker positions {s with respect to the listener position. Finally, the spatial component of the cost function is defined as the squared error between the desired binaural response (Equation 5) and that produced by the loudspeakers (Equation 16):Cspatiai^ o, &}) = (b - Hg)*(b - Hg) (17)D25008W001
[0118] Conveniently, the spatial term of the cost function for CMAP and FV defined in Equations 4 and 7 can both be rearranged into a matrix quadratic as a function of speaker activations g:Cspatial(.& o, {sj) = g*Ag + Bg + C (18)
[0119] In Equation 18, A represents an M M square matrix, B represents a l x, W vector, and C represents a scalar. The matrix A is of rank 2, and therefore when M > 2 there exist an infinite number of speaker activations g for which the spatial error term equals zero. Introducing the second term of the cost function, Cproximity, removes this indeterminacy and results in a particular solution with perceptually beneficial properties in comparison to the other possible solutions. For both CMAP and FV, Cproximity is constructed such that activation of speakers whose position Sj is distant from the desired audio signal position o is penalized more than activation of speakers whose position is close to the desired position. This construction yields an optimal set of speaker activations that is sparse, where only speakers in close proximity to the desired audio signal’s position are significantly activated, and practically results in a spatial reproduction of the audio signal that is perceptually more robust to listener movement around the set of speakers.
[0120] To this end, the second term of the cost function, Cproximity, may be defined as a distance-weighted sum of the absolute values squared of speaker activations. This is represented compactly in matrix form as:Cproximity(&> {^j})—g Dg (19a)
[0121] In Equation 19A, D represents a diagonal matrix of distance penalties between the desired audio position and each speaker:d 0D =: di = distance(o, Sj) (19b).0
[0122] The distance penalty function can take on many forms, but the following is a useful parameterization:distance (o, Sj) = a d%S1'h (19c)
[0123] In Equation 19c, ||o — sf|| represents the Euclidean distance between the desired audio position and speaker position and a and ft represent tunable parameters. The parameter a indicates the global strength of the penalty; d0corresponds to the spatial extent of the distance penalty (loudspeakers at a distance around d0or futher away will be penalized), and ft accounts for the abruptness of the onset of the penalty at distance d0. Combining the two terms of the cost function defined in Equations 18 and 19a yields the overall cost function:D25008W001 C(g) = g*Ag + Bg + C + g*Dg = g*(A + D)g + Bg + C (20)
[0124] Setting the derivative of this cost function with respect to g equal to zero and solving for g yields the optimal speaker activation solution:gopt= |(A + D)-1B (21)
[0125] In general, the optimal solution in Equation 11 may yield speaker activations that are negative in value. For the CMAP construction of the flexible renderer, such negative activations may not be desirable, and thus Equation (11) may be minimized subject to all activations remaining positive.
[0126] Some disclosed implementations include a system or device configured (e.g., programmed) to perform any embodiment of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing any embodiment of the disclosed methods or steps thereof. For example, the disclosed system can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including an embodiment of the disclosed method or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform an embodiment of the disclosed method (or steps thereof) in response to data asserted thereto.
[0127] Some embodiments of the disclosed system are implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of an embodiment of the disclosed method. Alternatively, embodiments of the disclosed system (or elements thereof) are implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including an embodiment of the disclosed method. Alternatively, elements of some embodiments of the disclosed system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform an embodiment of the disclosed method, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform an embodiment of the disclosed method would typically be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.
[0128] Various aspects of the present disclosure may be appreciated by way of the following Enumerated Example Embodiments (EEEs):D25008W001
[0129] EEE 1. A method comprising: receiving, by a control system and via an interface system, audio data, wherein the audio data comprises one or more audio signals and associated spatial data, wherein the associated spatial data indicates an intended perceived spatial position corresponding to the one or more audio signals; and rendering, by the control system, the audio data for reproduction via one or more loudspeakers of an environment, to produce rendered audio signals, wherein rendering each of the one or more audio signals included in the audio data comprises applying a monaural perceptual cue filter to each audio signal of the one or more audio signals, wherein the monaural perceptual cue filter is configured to insert one or more spectral perceptual cues corresponding to the intended perceived spatial position of a corresponding audio signal, and at least partially remove one or more spectral perceptual cues corresponding to at least one spatial position of the one or more loudspeakers for which the rendered audio signals are produced.
[0130] EEE 2. The method of EEE 1, wherein the monaural perceptual cue filter is configured as a ratio of an audio signal perceptual cue corresponding to an intended perceived spatial position of the one or more audio signals and a composite loudspeaker perceptual cue of at least one of the one or more loudspeakers associated with rendering the audio data.
[0131] EEE 3. The method of EEE 2, wherein the audio signal perceptual cue is based, at least in part, on a monaural head-related transfer function (HRTF) corresponding to the intended perceived spatial position of the one or more audio signals and the composite loudspeaker perceptual cue is based, at least in part, on a weighted combination of monaural HRTFs corresponding to a spatial position of each of the one or more loudspeakers.
[0132] EEE 4. The method of EEE 3, wherein the monaural HRTF corresponding to the intended perceived spatial position of the one or more audio signals is computed as a sum of squared magnitudes of a left ear HRTF and a right ear HRTF for the intended perceived spatial position.
[0133] EEE 5. The method of EEE 3 or EEE 4, further comprising computing, by the control system, a weighting for the weighted combination of the monaural HRTFs as a function of rendering filters for the one or more audio signals.
[0134] EEE 6. The method of any one of EEEs 3-5, wherein the monaural HRTF corresponding to the intended perceived spatial position of the one or more audio signals and the composite loudspeaker perceptual cue each additionally include a diffuse sound component.
[0135] EEE 7. The method of EEE 6, wherein a variable scaling of the diffuse sound component controls a strength of the monaural perceptual cue filter.D25008W001
[0136] EEE 8. The method of any one of EEEs 1-7, wherein rendering the audio data is based at least in part on Center of Mass Amplitude Panning.
[0137] EEE 9. The method of any one of EEEs 1-7, wherein rendering the audio data is based at least in part on Flexible Virtualization.
[0138] EEE 10. The method of any one of EEEs 1-9, wherein rendering the audio data is based on Flexible Virtualization for frequencies below a frequency threshold.
[0139] EEE 11. The method of EEE 10, wherein rendering the audio data is based on Center of Mass Amplitude Panning for frequencies above the frequency threshold.
[0140] EEE 12. The method of EEE 10 or EEE 11, wherein the frequency threshold is in a range from 1 kHz to 5 kHz, inclusive.
[0141] EEE 13. The method of any one of EEEs 1-12, further comprising obtaining each spatial position of the one or more loudspeakers.
[0142] EEE 14. The method of any one of EEEs 1-13, further comprising providing, by the control system and via the interface system, the rendered audio signals to the one or more loudspeakers.
[0143] EEE 15. The method of any one of EEEs 1-1, wherein the one or more audio signals comprise one or more audio object signals and wherein the associated spatial data includes audio object metadata indicating the intended perceived spatial position of a corresponding audio object signal.
[0144] EEE 16. The method of any one of EEEs 1-15, wherein the intended perceived spatial position is indicated by a channel of a channel-based audio format.
[0145] EEE 17. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any of EEEs 1-16.
[0146] EEE 18. A computer program including instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any of EEEs 1-16.
[0147] EEE 19. A computing apparatus, comprising: at least one processor; a display; and memory storing instructions, which when executed by the at least one processor, cause the computing apparatus to perform the method of any of EEEs 1-16.
[0148] Another aspect of the present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) any disclosed method or steps thereof.
[0149] While specific embodiments and applications have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodimentsD25008W001 and applications described herein are possible without departing from the scope described and claimed herein. It should be understood that while certain forms have been shown and described, the scope of the present disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.
Claims
D25008W001CLAIMSWhat Is Claimed Is:
1. A method comprising:receiving, by a control system and via an interface system, audio data, wherein the audio data comprises one or more audio signals and associated spatial data, wherein the associated spatial data indicates an intended perceived spatial position corresponding to the one or more audio signals; andrendering, by the control system, the audio data for reproduction via a plurality of loudspeakers of an environment, to produce rendered audio signals, wherein rendering each of the one or more audio signals included in the audio data comprises applying a monaural perceptual cue filter to each audio signal of the one or more audio signals, wherein the monaural perceptual cue filter is configured to insert one or more spectral perceptual cues corresponding to the intended perceived spatial position of a corresponding audio signal, and remove spectral perceptual cues caused by the loudspeakers for which the rendered audio signal is produced.
2. The method of claim 1, wherein the monaural perceptual cue filter is configured as a ratio of an audio signal perceptual cue corresponding to an intended perceived spatial position of the one or more audio signals and a composite loudspeaker perceptual cue of the plurality of loudspeakers associated with rendering the audio data.
3. The method of claim 2, wherein the audio signal perceptual cue is based, at least in part, on a monaural head-related transfer function (HRTF) corresponding to the intended perceived spatial position of the one or more audio signals and the composite loudspeaker perceptual cue is based, at least in part, on a weighted combination of monaural HRTFs corresponding to a spatial position of each of the plurality of loudspeakers.
4. The method of claim 3, wherein the monaural HRTF corresponding to the intended perceived spatial position of the one or more audio signals is computed as a sum of squared magnitudes of a left ear HRTF and a right ear HRTF for the intended perceived spatial position.D25008W001 5. The method of claim 3 or claim 4, further comprising computing, by the control system, a weighting for the weighted combination of the monaural HRTFs as a function of rendering filters for the one or more audio signals.
6. The method of any one of claims 3-5, wherein the monaural HRTF corresponding to the intended perceived spatial position of the one or more audio signals and the composite loudspeaker perceptual cue each additionally include a diffuse sound component.
7. The method of claim 6, wherein a variable scaling of the diffuse sound component controls a strength of the monaural perceptual cue filter.
8. The method of any one of claims 1-7, wherein rendering the audio data is based at least in part on Center of Mass Amplitude Panning.
9. The method of any one of claims 1-7, wherein rendering the audio data is based at least in part on Flexible Virtualization.
10. The method of any one of claims 1-9, wherein rendering the audio data is based on Flexible Virtualization for frequencies below a frequency threshold.
11. The method of claim 10, wherein rendering the audio data is based on Center of Mass Amplitude Panning for frequencies above the frequency threshold.
12. The method of claim 10 or claim 11, wherein the frequency threshold is in a range from 1 kHz to 5 kHz, inclusive.
13. The method of any one of claims 1-12, further comprising obtaining each spatial position of the plurality of loudspeakers.
14. The method of any one of claims 1-13, further comprising providing, by the control system and via the interface system, the rendered audio signals to the plurality of loudspeakers.
15. The method of any one of claims 1-14, wherein the one or more audio signals comprise one or more audio object signals and wherein the associated spatial data includesaudio object metadata indicating the intended perceived spatial position of a corresponding audio object signal.
16. The method of any one of claims 1–15, wherein the intended perceived spatial position is indicated by a channel of a channel-based audio format.
17. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any of claims 1–16.
18. A computer program including instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any of claims 1–16.
19. A computing apparatus, comprising:at least one processor;a display; andmemory storing instructions, which when executed by the at least one processor, cause the computing apparatus to perform the method of any of claims 1–16.