A system and method for rendering personal sound zones with head tracking using multiple loudspeakers
The SANN-based method dynamically adjusts PSZ filters in real-time to listener movements, enhancing audio quality and reducing interference in dynamic environments.
Patent Information
- Application Number
- PCT/US2025/052376
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-25
- Filing Date
- 2025-10-24
- Publication Date
- 2026-04-30
AI Technical Summary
Existing personal sound zone (PSZ) systems fail to dynamically adapt to listener movement in dynamic environments, leading to insufficient audio quality and interference from other audio programs.
A deep learning-based method using a Spatially Adaptive Neural Network (SANN) generates and updates audio filters in real-time based on listener head positions, minimizing interference by dynamically adjusting PSZs.
The method provides efficient, flexible, and robust audio filtering that adapts to listener movements, maintaining high audio quality and minimizing interference across various environments.
Smart Images

Figure US2025052376_30042026_PF_FP_ABST
Abstract
Description
[0001] Princeton - 104476
[0002] A SYSTEM AND METHOD FOR RENDERING PERSONAL SOUND ZONES WITH HEAD TRACKING USING MULTIPLE LOUDSPEAKERS
[0003] CROSS-REFERENCE TO RELATED APPLICATIONS
[0004] The present application claims priority to US 63 / 711,934, filed October 25, 2024, the contents of which are incorporated by reference herein in its entirety.
[0005] TECHNICAL FIELD
[0006] This application relates to a system and method for rendering personal sound zones (PSZs) with head tracking using multiple loudspeakers.
[0007] BACKGROUND PSZs refer to distinct spatial regions where listeners can experience personalized audio programs with minimal interference from other programs. These zones are categorized into two types: acoustically bright zone (BZ), where the audio program is delivered to the intended listener, and acoustically dark zone (DZ), where the program is attenuated. In BZ, the audio program may be personalized in attributes such as content, volume, and spatialization to suit the listener’s preferences. Conventionally, there are known ways to rendering personal sound zones (PSZs), but there is an ongoing demand for improved systems and methods for producing PSZs.
[0008] SUMMARY
[0009] In some cases, digital audio filters are generated for rendering PSZs by solving optimization problems that consider the acoustic transfer functions (ATFs) between the loudspeakers and the control points in both BZ and DZ. The goal is to minimize the sound energy in DZ while preserving the sound quality in BZ, for example. These filters can then be applied to the audio programs played back through the loudspeakers. Some filter design methods include pressure matching (PM) and acoustic contrast control (ACC), each of which approaches the optimization with different objectives.
[0010] Some PSZ filter design methods, including PM and ACC, assume that the zones are static and do not account for listener movement. However, in dynamic environments like home entertainment, where listeners may move freely across a larger area than in confined spaces like car cabins, a static PSZ configuration can be insufficient. Thus, it can be beneficial to use Princeton - 104476
[0011] systems and methods that can dynamically update PSZs by adjusting the audio filters based on the listener's real-time head position.
[0012] Some approaches to updating audio filters with head tracking can involve either precomputing the filters for a discrete set of head positions or computing the filters in real-time as the listeners move. In the former approach, numerous filter sets may be used to accommodate many possible head positions and PSZ combinations. In some cases, the filters are interpolated, while in other cases, such combinations are constrained. The latter approach, on the other hand, can be computationally demanding as it may use the inversion of ATF matrices for every new head position. This complexity can result in compromises, such as simplified filter designs or adaptive solutions that are less computationally demanding.
[0013] The present disclosure aims to provide a novel deep learning-based method for rendering PSZs with head tracking. In some deep learning-based methods for designing PSZ filters, see for instance the article “Digital filters design for personal sound zones: A neural approach” by G. Pepe et al., published at 2022 International Joint Conference on Neural Networks (IJCNN), the listener movement aspect is not considered in the neural network architecture and the training process. Instead, the present disclosure relies on a novel architecture that generates filters that adapt to listeners’ movements. Furthermore, a novel network training process is adopted, which ensures filter robustness in unknown rendering environments and allows for customization in a specific environment.
[0014] This application discloses a system and method for rendering personal sound zones (PSZs) with head tracking using multiple loudspeakers. To deliver head-tracked, personalized audio to listeners while minimizing interference from other audio programs, the method relies on a Spatially Adaptive Neural Network (SANN) to dynamically generate and update audio filters in real-time based on the listeners' head positions.
[0015] We categorize the rendered PSZs into two types: acoustically bright zone (BZ) and acoustically dark zone (DZ). In BZ, an audio program is delivered to the intended listener, while in DZ, the program is attenuated. The audio program rendered in BZ may be personalized in attributes such as content, volume, and spatialization (e.g., spatial enhancement with crosstalk cancellation technique) to suit the listener’s preferences. The system can be configured to render a single BZ for one listener or multiple PSZs that adapt simultaneously to the movements of multiple listeners.
[0016] An example method for rendering PSZs with head tracking comprises the following six general steps: Princeton - 104476
[0017] Step I: System Configuration: Specify i) the loudspeaker layout, ii) the size and shape of the PSZ and rendering area, iii) the range of head movement to be tracked, and / or iv) the range of room acoustics parameters (e.g., room size, absorption coefficients, reverberation time) in which the system is designed to operate.
[0018] Step II: SANN Architecture Design: Specify the architecture of SANN based on the system parameters configured in Step I, including i) the size of input and output layers, ii) the layers for processing and encoding the SANN input, and / or iii) the number, size, and connectivity pattern of intermediate layers.
[0019] Step III: Loss Function Definition: Define a unified loss function that incorporates relevant PSZ design objectives, e.g., maximizing acoustic isolation between PSZs, minimizing unwanted filter artifacts (e.g., pre- and post-ringing), and rendering audio programs with specific spatialization properties in BZ.
[0020] Step IV: Training Data Preparation: Create numerically simulated acoustic transfer functions (ATFs) corresponding to the loudspeakers used in the system and control points in the rendering area. In some implementations, if a specific rendering environment is considered, measuring corresponding ATFs using a single microphone, a microphone array, or a mannequin head with in-ear microphones.
[0021] Step V: SANN Training: Train the specified SANN model using simulated ATFs and / or measured ATFs (if the latter are available) to generate audio filters with the defined loss function, through backpropagation and optimization algorithms.
[0022] Step VI: Real-time PSZ Rendering: Given a stream of head position data from a head tracking device, generate PSZ filters in real-time using the trained SANN and convolving them with the audio programs to be played through the loudspeakers. In the case of band-limited audio programs, cascade the generated PSZ filters with a band-pass filter before convolution.
[0023] The method offers significant advantages in terms of storage and computational efficiency. By integrating filter generation and interpolation into a single step using the SANN, it eliminates both the need to pre-compute and store numerous filter sets for different head positions and the computational overhead of evaluating closed-form filter solutions for every new head position. Furthermore, the method allows for flexible customization to specific rendering environments by incorporating measured ATFs into the training process, while maintaining robust performance in a variety of environments. Princeton - 104476
[0024] BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Fig. 1 is an illustration of an embodiment of the system and environment for rendering PSZs.
[0026] Fig. 2 is a flowchart depicting an embodiment of the SANN architecture design (Step II).
[0027] Fig. 3 is a flowchart depicting an embodiment of the loss function calculation process (Step III).
[0028] Fig. 4 is a flowchart depicting an embodiment of the training data preparation process (Step IV).
[0029] Fig. 5 is a flowchart depicting an embodiment of the SANN training process (Step V). Fig. 6 is a flowchart depicting an embodiment of the real-time PSZ rendering process (Step VI).
[0030] Fig. 7 is an illustration of an embodiment of a SANN architecture and a forward pass of a training process. Dotted arrows indicate the paths along which the gradients are back-propagated.
[0031] Fig. 8 is an illustration of a PSZ system layout, showing possible placement of BZ and DZ. The black triangles represent the transducers. The dashed rectangle indicates the rendering area within which the BZ and DZ can be positioned.
[0032] DETAILED DESCRIPTION
[0033] The system and method for rendering personal sound zones (PSZs) with head tracking using multiple loudspeakers may involve six general steps, outlined below. These steps result in the real-time generation of PSZ filters using a spatially adaptive neural network (SANN) based on listeners’ head positions. The process of SANN training and filter generation could involve fewer individual steps, or combine certain steps, without departing from the core principles of this disclosure.
[0034] We use Roman numerals to denote the six general steps of the example method, and Arabic numerals to denote detailed steps, some of which are optional or may vary in different embodiments of the method described below. Various steps described herein can be combined, omitted, or divided into sub-steps.
[0035] Step I: System Configuration: In this step, the PSZ rendering system can be configured by specifying the following set of inputs (or a subset of them): i) the loudspeaker layout, ii) the size and shape of the PSZ and rendering area, iii) the range of head movement to be tracked, and / or iv) the range of room acoustics parameters in which the system is designed Princeton - 104476
[0036] to operate. The loudspeaker layout includes both the number and the arrangement of loudspeakers. The PSZ rendering area may be defined as a 3D volume or a 2D plane, depending on the application. The head tracking may be restricted to a specific region, such as a room or a defined listening space. The room acoustics parameters may include room size, absorption coefficients, reverberation time, and the position of the rendering system in the room. In some implementations, the system can provide a user interface for a user to provide the configuration parameters. In some cases, the system can determine some configuration parameters automatically. For example, room acoustics parameters can be determined based on sound recorded by a microphone of the system or of an external device. The system can output one or more test sounds and can record sounds received in response to the test sounds, which can be used to determine acoustical parameters and / or the size and / or shape of the space. In some cases, a user interface can permit a user to specify a sub-area in the space to be used as the PSZ rendering area.
[0037] Fig. 1 depicts an example of the PSZ system configuration defined in a 2D plane, including loudspeakers, the PSZ rendering area, two rendered PSZs (one BZ and one DZ), and control points within the rendering area where the training data is collected. The size of PSZ can be determined by balancing the isolation performance (smaller zone) with the robustness against head misalignment (larger zone). The system can include a user interface to enable a user to provide input to specify a PSZ size. In some cases, the system can monitor movement of listener and determine the appropriate PSZ size. For example, a listener that shifts position more can result in the system prescribing a larger PSZ size. The spacing between control points in the rendering area can be determined by the desired rendering spatial resolution and the maximum rendering audio frequency to avoid spatial aliasing. The system can include a user interface to enable a user to provide input to specify the control point spacing. For more detailed discussion on the spatial sampling of ATFs, see the article "Spatial sampling of binaural room transfer functions for head-tracked personal sound zones" by Y. Qiao et al., published in the Journal of the Audio Engineering Society, volume 72, no. 7 / 8, pages 455-466, July / August 2024.
[0038] Step II: SANN Architecture Design: In this step, given the system parameters configured in Step I, the SANN architecture can be designed by specifying i) the size of input and output layers, ii) the layers for processing and encoding the SANN input, and / or iii) the number, size, and connectivity pattern of intermediate layers. The input layer can be set to receive head position data (e.g.. in 3D coordinates) from one or multiple listeners. When the coordinates only include head translations in three orthogonal directions (X / Y / Z), the input has Princeton - 104476
[0039] a size of 3N, where N is the number of listeners to be tracked. In this case, the coordinates are considered as the center of rendered PSZs. The output layer can be set to generate concatenated PSZ filters for all loudspeakers (e.g., PSZ filter coefficients or parameters). If time-domain impulse responses are considered, the output layer has a size of L times the length of the taps of a single filter, where L is the number of loudspeakers; if frequency-domain filter responses are considered, it has a size of 2L times the length of complex-valued filter coefficients, as the real and imaginary parts of the coefficient are concatenated in the output. If band-limited audio programs are considered, the output layer may only include frequency-domain coefficients for the desired frequency band. After the input layer, a series of non-leamable layers can be cascaded for processing and encoding the input coordinates. As an example, in an embodiment of the SANN architecture design depicted in Fig. 2, a normalization layer (label 10) can be set to normalize the head coordinates by the extents of the PSZ rendering area specified in Step I to a range of (-1, 1), for example. Then, a positional encoding layer (label 12) can be cascaded to convert the coordinates into a higher-dimensional representation. One example implementation of the positional encoding layer may use Fourier features, as discussed in the article "Fourier features let networks leam high frequency functions in low dimensional domains" by M. Tancik et al., published in Advances in Neural Information Processing Systems, 2020. After the non-leamable layers, a series of intermediate layers (label 14 in Fig.
[0040] 2) can be cascaded to leam the nonlinear mapping from the encoded positional information to the corresponding PSZ filters. One example implementation of the intermediate layers to use a series of multi-layer perceptrons (MLPs, also referred to as fully-connected layers) with Rectified Linear Unit (ReLU) activation functions. The size of each intermediate layer and the number of layers can be set to balance computational complexity with filter generation accuracy.
[0041] Step III: Loss Function Definition: In this step, a unified loss function for SANN training can be defined by incorporating relevant PSZ design objectives, e.g., maximizing acoustic isolation between PSZs, minimizing unwanted filter artifacts (e.g., pre- and postringing), and rendering audio programs with specific spatialization properties (e.g., mono, stereo, crosstalk cancelled binaural, etc.) in BZ. The loss function can include any combination of three terms: i) a BZ loss term that corresponds to the rendering accuracy in BZ, ii) a DZ loss term that corresponds to the energy in DZ, and / or iii) a filter loss term that corresponds to the desired filter properties. These loss terms may be formulated in either the time or frequency domain, and may be weighted based on the importance of each objective. Princeton - 104476
[0042] Fig. 3 depicts an embodiment of the loss function calculation process. In Step 16, the BZ loss term can be calculated as the LI -norm difference between the target and estimated pressure at the control points within BZ. The target pressure can correspond to different spatialization settings, such as mono, stereo, or binaural rendering with crosstalk cancellation. More details on setting the target pressure can be found in Section 4 of the article "Isolation performance metrics for personal sound zone reproduction systems" by Y. Qiao et al., published in JAS A Express Letters, volume 2, No. 10, 2022. In Steps 18 and 20, sound pressure at control points within both BZ and DZ can be computed by convolving the corresponding acoustic transfer functions (ATFs) with the PSZ filters from the SANN output and summing over all loudspeakers, respectively. In Step 22, the DZ loss term can be calculated as the LI norm of the pressure within DZ. In Step 24, the filter loss term can be computed to constrain the filter properties, such as impulse response compactness. This can be achieved by applying a windowing function to the generated filter impulse responses and computing the energy of the windowed responses. An example of computing such a loss term is given in the article "Digital filters design for personal sound zones: A neural approach" by G. Pepe et al, published at 2022 International Joint Conference on Neural Networks (IJCNN). Finally, in Step 26, a unified loss can be computed as a weighted sum of the BZ, DZ, and filter loss terms, with weights adjusted based on the importance of each term in the optimization process.
[0043] Step IV: Training Data Preparation: This step involves creating ATFs for training the SANN model, as depicted in Fig. 4 (in some embodiments, the dashed steps in Fig. 4 are optional). In Step 28, the ATFs corresponding to the loudspeakers and control points in the PSZ rendering area are numerically simulated under different room acoustics conditions, given the parameters specified in Step I. The simulated ATFs can ensure that SANN is trained to generate PSZ filters that are robust to variations in various rendering environments. The room acoustics parameters can include room size, absorption coefficients, and / or reverberation time, and can be randomly sampled from the previously specified range. The ATF simulation may only consider the room-related effects (e.g., reflections), or may incorporate head and / or torso scattering effects using head-related transfer functions (HRTFs). An example simulation method is the image source method, see the article "Py roomacoustics: A Python package for audio room simulations and array processing algorithms" by R. Scheibler et al., published at 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), for an example implementation. If a specific rendering environment is considered where the physical system same as the one specified in Step I is deployed, additional ATFs can be measured (Step 30) using microphones to improve system performance. The microphone used Princeton - 104476
[0044] for the measurement can be a single omnidirectional microphone, a microphone array, or in-ear microphones inserted in a mannequin or a human head, depending on whether the HRTFs are considered in the simulation. If both simulated and measured ATFs are available, the measured ATFs can replace those simulated at control points where measurements were taken. After combination of ATFs (Step 32), the resulting ATF dataset is used to train the SANN model. The combination allows for both robustness outside the measurement region and customization within the region. Some implementations could use only measured ATFs and could omit the simulated ATFs.
[0045] Step V: SANN Training: In this step, the SANN model can be trained using the ATF dataset from Step IV and / or the loss function defined in Step III. Fig. 5 depicts an embodiment of the training process. Prior to training, the SANN model weights can be randomly initialized. During training, each mini-batch of the training examples can include head positions of one or multiple listeners sampled within the rendering area and their corresponding ATFs. In Step 34, given the listener head coordinates as input, the ATFs corresponding to control points within the PSZs centered at the head positions are selected from the ATF dataset. These ATFs may be either simulated (e.g., under random room acoustics conditions) or measured in the existing environment, depending on the availability of measured ATFs at the control points. These selected ATFs can be further augmented in Step 36 by adding noise or amplitude / phase perturbations (e.g., from random loudspeaker movements). In Step 38, the input is passed through the SANN to generate PSZ filters for each loudspeaker. These filters, combined with the ATFs, can be used to compute the loss function in Step 40. Finally, in Step 42, the SANN model weights can be updated, such as by using backpropagation and optimization algorithms, such as the mini-batch gradient descent algorithm and the ADAM optimization method. The latter is described in the article "Adam: A method for stochastic optimization" by D. P. Kingma and J. Ba, published at the International Conference on Learning Representations, 2015.
[0046] Step VI: Real-time PSZ Rendering: In this step, given a stream of head position data from a head tracking device (e.g., camera, infrared sensor, or wearable), PSZ filters can be generated in real-time by the trained SANN, for example as depicted in Fig. 6. First, in Step 44, head position data from a head tracking device can be received and processed (e.g., smoothing and / or outlier removal) to obtain the head coordinates of one or multiple listeners. In Step 46, the head position data is fed into the trained SANN, which generates PSZ filters, for example in real-time. In Step 48, optional bandpass filtering process may be cascaded to the generated PSZ filters, such as if only band-limited audio programs (and filter coefficients) are considered. In Step 50, the audio programs to be played through the loudspeakers are Princeton - 104476
[0047] convolved with the generated PSZ filters. To enable real-time rendering, partitioned convolution algorithms may be adopted, as detailed in the thesis "‘Partitioned convolution algorithms for real-time aurahzation” by Frank Wefers, published by Logos Verlag Berlin, 2015. Finally, in Step 52, the filtered audio programs are played through the loudspeakers to render the PSZs. To render different audio programs for multiple listeners, multiple instances of the trained SANN models can be used, each rendering a BZ for the intended listener. The PSZ filters can be updated at a rate no less than the head tracking frequency to adapt to listener movement. The rendering process may be implemented on a dedicated hardware, such as a digital signal processor (DSP), field-programmable gate array (FPGA), or a high-performance computer with an audio interface, or any other ty pe of DSP processor.
[0048] Examples provided herein of training a machine learning model are described in connection with a Spatially Adaptive Neural Network (SANN). However, the system may additionally or alternatively train any suitable machine learning approach or model, including but not limited to: an artificial neural network (NN), neural-network-based regression models, linear regression models, logistic regression models, decision trees, random forests, support vector machines (SVMs), Naive or anon-Naive Bayes network (also referred to as a “Bayesian classifier”), L-nearest neighbors (KNN) models, / t-means models, clustering models, or any combination thereof. For brevity,, aspects of training will not be described with respect to each possible machine learning model that may be used. In practice, however, many or all of the aspects of the disclosure may apply to other machine learning models, including but not limited to those listed herein.
[0049] Generally described, NNs — including deep neural networks, transformer neural networks, and the like — have multiple layers of nodes, also referred to as “neurons.” Illustratively, a NN may include an input layer, an output layer, and any number of intermediate, internal, or “hidden” layers between the input and output layers. The individual layers may' include any number of separate nodes. Nodes of adjacent layers may be logically connected to each other, and each logical connection between the various nodes of adjacent layers may be associated with a respective weight. Conceptually, a node may be thought of as a computational unit that computes an output value as a function of a plurality of different input values. Nodes may be considered to be “connected” when the input values to the function associated with a current node include the output of functions associated with nodes in a previous layer, multiplied by a set of weights associated with the individual “connections” between the current node and the nodes in the previous layer. When a NN is used to process input data in the form of an input vector or a matrix of input vectors (e.g., an operational input Princeton - 104476
[0050] vector, or a batch of training data input vectors), the NN may perform a “forward pass'’ to generate an output vector or a matrix of output vectors, respectively. The input vectors may each include n separate data elements or “dimensions,” corresponding to the n nodes of the NN input layer (where n is a positive integer). Each data element may be a value, such as a floatingpoint number or integer. A forward pass ty pically includes multiplying the matrix of input vectors by a matrix representing the set of weights associated with connections between the nodes of the input layer and nodes of the next layer, and applying an activation function to the results. The process is then repeated for each subsequent NN layer. Some NNs have hundreds of thousands or millions of nodes, and millions or billions of weights for connections between the nodes of all of the adjacent layers. Some NNs, such as transformer neural networks, have a multi-head attention mechanism allowing the effect of some portions of input (e.g., key tokens, contextual tokens, other contextual data, etc.) to be amplified or diminished during processing pass.
[0051] Training data to train the model may include labeled data, such as manually or automatically labeled samples of previously received, manually generated, or automatically generated data items. The labels may include correct or desired output to be provided by the model being trained, such as the correct classification, metric value, etc. Training data can include input vectors representing the model input, and corresponding reference data output vectors representing class labels (in the case of classification models) or numeric output values (in the case of regression models) associated with each input vector. Training a machine learning model on training data can result in the machine learning model being able to identify patterns in the training data and make predictions concerning which class label or numeric output value should apply to an input vector.
[0052] The connections between individual nodes of adjacent layers can each be associated with a trainable parameter, such as a weight and / or bias term, that is applied to the value passed from the prior layer node to the activation function of the subsequent layer node. For example, the weights associated with the connections from an input layer to a first internal layer to which it is connected may be arranged in a weight matrix W with a size m x n, where m denotes the number of nodes in an internal layer and n denotes the dimensionality of the input layer. The individual rows in the weight matrix W may correspond to the individual nodes in the input layer, and the individual columns in the weight matrix W may correspond to the individual nodes in the internal layer. The weight w associated with a connection from any node in the input layer to any node in the internal layer may be located at the corresponding intersection location in the weight matrix W. Princeton - 104476
[0053] Illustratively, a training data input vector may be provided to a computer processor that stores or otherwise has access to the weight matrix W. The processor then multiplies the combined training data input vector by the weight matrix W to produce an intermediary vector. The processor may adjust individual values in the intermediary vector using an offset or bias that is associated with the internal layer (e.g., by adding or subtracting a value separate from the weight that is applied). In addition, the processor may apply an activation function to the individual values in the intermediary vector (e.g., by using the individual values as input to a sigmoid function or a rectified linear unit (“ReLU”) function).
[0054] There may be a set of “hidden” layers, including multiple internal hidden layers in some implementations, and each internal layer may or may not have the same number of nodes as each other internal layer. The weights associated with the connections from one internal layer (also referred to as the “preceding internal layer”) to the next internal layer (also referred to as the “subsequent internal layer”) may be arranged in a weight matrix similar to the weight matrix W, with a number of rows equal to the number of nodes in the subsequent internal layer and a number of columns equal to the number of nodes in the preceding internal layer. The weight matrix may be used to produce another intermediary vector using the process described above with respect to the input layer and first internal layer. The process of multiplying intermediary vectors by weight matrices and applying activation functions to the individual values in the resulting intermediary vectors may be performed for each internal layer subsequent to the initial internal layer.
[0055] The output layer of the NN can make output determinations from the hidden layers. Weights associated with the connections from the last internal layer to the output layer may be arranged in a weight matrix similar to the weight matrix W, with a number of rows equal to the number of nodes in the output layer and a number of columns equal to the number of nodes in the last internal layer. The weight matrix may be used to produce an output vector using the process described above with respect to the input layer and first internal layer.
[0056] The output vector may include data representing the classification determinations or regression determinations made by the NN for the training data input vector. Some NNs are configured to make u classification determinations corresponding to u different classifications (where u is a number corresponding to the number of nodes in the output layer, and may be less than, equal to, or greater than the number of nodes n in the input layer). The data in each of the u different dimensions of the output vector may be a confidence score indicating the probability that the combined training data input vector is properly classified in a corresponding classification. Some NNs are configured to generate values based on regression Princeton - 104476
[0057] determinations. The output value(s) is / are based on a mapping function modeled by the NN. Thus, an output value from a NN-based regression model is the value that corresponds to the combined training data input vector or the ordinary training data input vectors.
[0058] During training, a model manager may determine the difference between the output vectors and corresponding reference data output vectors, and then adjust parameters of the NN (e.g., weights and / or bias terms of the NN) such that the NN will subsequently produce output vectors that are closer to the corresponding reference data output vectors. The modification of parameter values may be performed through a process referred to as “back propagation.” Back propagation includes determining the difference between the expected model output (e.g., the reference data output vectors) and the obtained model output, and then determining how to modify the values of some or all parameters of the model to reduce the difference between the expected model output and the obtained model output. In some embodiments, a computing system may compute the difference using a loss function, such as a cross-entropy loss function, a L2 Euclidean loss function, a logistic loss function, a hinge loss function, a square loss function, or a combination thereof. The computing system can compute a derivative, or “gradient,” that corresponds to the direction in which each parameter of the machine learning model is to be adjusted in order to improve the model output (e.g., to produce output that is closer to the correct or preferred output for a given combined training data input vector, as represented by the reference data output vector). The computing system can update one or more parameters of the machine learning model based on the gradient. For example, the computing system can update some or all parameters of the machine learning model using a gradient descent method. The adjustments may be propagated back through the NN layer-by-layer.
[0059] The process of generating output vectors from combined training data input vectors, determining the differences, and adjusting the parameters of the NN may be repeated until one or more termination criteria are met. For example, the termination criteria can be based on the accuracy of the NN as determined using the loss function, a number of iterations performed, a duration of time, or the like.
[0060] Additional Information
[0061] The methods and tasks described herein may be performed and automated by a computer system. The computer system may, in some cases, include multiple distinct computers or computing devices (e.g., physical servers, workstations, storage arrays, cloud computing resources, etc.) that communicate and interoperate over a network to perform the described functions. Each such computing device typically includes a processor (or multiple Princeton - 104476
[0062] processors) that executes program instructions or modules stored in a memory or other non-transitory computer-readable storage medium or device (e.g., solid state storage devices, disk drives, etc.). The various functions disclosed herein may be embodied in such program instructions, or may be implemented in application-specific circuitry (e.g., ASICs or FPGAs) of the computer system. Where the computer system includes multiple computing devices, these devices may, but need not, be co-located. The results of the disclosed methods and tasks may be persistently stored by transforming physical storage devices, such as solid-state memory chips or magnetic disks, into a different state. In some embodiments, the computer system may be a cloud-based computing system whose processing resources are shared by multiple distinct business entities or other users.
[0063] Depending on the embodiment, certain acts, events, or functions of any of the processes or algorithms described herein can be performed in a different sequence, can be added, merged, or left out altogether (e.g., not all described operations or events are necessary for the practice of the algorithm). Moreover, in certain embodiments, operations or events can be performed concurrently, e.g.. through multi-threaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially.
[0064] The various illustrative logical blocks, modules, routines, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, or combinations of electronic hardware and computer software. To clearly illustrate this interchangeability, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware, or as software that runs on hardware, depends upon the particular application and design constraints imposed on the overall system. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure.
[0065] In some embodiments, the methods, techniques, microprocessors, and / or controllers described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be hard-wired to perform the techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) that are persistently programmed to perform the techniques, or may include one or more general purpose hardware processors programmed to perform the techniques pursuant to program instructions in firmware, memory, other storage, or a combination thereof. The instructions can reside in RAM memory, flash memory, ROM memory, EPROM memory,, EEPROM memory, registers, hard disk, a Princeton - 104476
[0066] removable disk, a CD-ROM, or any other form of a non-transitory computer-readable storage medium. Such special-purpose computing devices may also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to accomplish the techniques. The specialpurpose computing devices may be desktop computer systems, server computer systems, portable computer systems, handheld devices, networking devices or any other device or combination of devices that incorporate hard-wired and / or program logic to implement the techniques.
[0067] The microprocessors or controllers described herein can be coordinated by operating system software, such as iOS, Android, Chrome OS, Windows XP, Windows Vista, Windows 7, Windows 8, Windows 10, Windows Server, Windows CE, Unix, Linux, SunOS, Solaris, iOS, Blackberry OS, VxWorks. or other compatible operating systems. A universal media server (UMS) can be used in some instances. In other embodiments, the computing device may be controlled by a proprietary operating system. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file system, networking, I / O services, and provide a user interface functionality, such as a graphical user interface ("GUT’), among other things.
[0068] The microprocessors and / or controllers described herein may implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic which causes microprocessors and / or controllers to be a special-purpose machine. According to one embodiment, parts of the techniques disclosed herein are performed a controller in response to executing one or more sequences instructions contained in a memory. Such instructions may be read into the memory' from another storage medium, such as storage device. Execution of the sequences of instructions contained in the memory causes the processor or controller to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.
[0069] Moreover, the various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processor device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardw are components, or any combination thereof designed to perform the functions described herein. A processor device can be a microprocessor, but in the alternative, the processor device can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor device can include electrical circuitry Princeton - 104476
[0070] configured to process computer-executable instructions. In another embodiment, a processor device includes an FPGA or other programmable device that performs logic operations without processing computer-executable instructions. A processor device can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor device may also include primarily analog components. For example, some or all of the techniques described herein may be implemented in analog circuitry or mixed analog and digital circuitry..
[0071] Unless the context clearly requires otherwise, throughout the description and the claims, the words "‘comprise,” “comprising.” “include,” “including,” and the like are to be construed in an inclusive sense, as opposed to an exclusive or exhaustive sense; that is to say, in the sense of “including, but not limited to.” The words “coupled” or connected,” as generally used herein, refer to two or more elements that can be either directly connected, or connected by way of one or more intermediate elements. Additionally, the words “herein.” “above,” “below,” and words of similar import, when used in this application, shall refer to this application as a whole and not to any particular portions of this application. Where the context permits, words in the Detailed Description using the singular or plural number can also include the plural or singular number, respectively. The words “or” in reference to a list of two or more items, is intended to cover all of the following interpretations of the word: any of the items in the list, all of the items in the list, and any combination of the items in the list. All numerical values provided herein are intended to include similar values within a range of measurement error.
[0072] Although this disclosure contains certain embodiments and examples, it will be understood by those skilled in the art that the scope extends beyond the specifically disclosed embodiments to other alternative embodiments and / or uses and obvious modifications and equivalents thereof. In addition, while several variations of the embodiments have been shown and described in detail, other modifications will be readily apparent to those of skill in the art based upon this disclosure. It is also contemplated that various combinations or subcombinations of the specific features and aspects of the embodiments may be made and still fall within the scope of this disclosure. It should be understood that various features and aspects of the disclosed embodiments can be combined with, or substituted for, one another in order to form varying modes of the embodiments. Any methods disclosed herein need not be performed Princeton - 104476
[0073] in the order recited. Thus, it is intended that the scope should not be limited by the particular embodiments described above.
[0074] Conditional language, such as, among others, “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or steps. Thus, such conditional language is not generally intended to imply that features, elements and / or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements and / or steps are included or are to be performed in any particular embodiment. Any headings used herein are for the convenience of the reader only and are not meant to limit the scope.
[0075] Further, while the devices, systems, and methods described herein may be susceptible to various modifications and alternative forms, specific examples thereof have been show n in the drawings and are herein described in detail. It should be understood, however, that the disclosure is not to be limited to the particular forms or methods disclosed, but, to the contrary, this disclosure covers all modifications, equivalents, and alternatives falling within the spirit and scope of the various implementations described. Further, the disclosure herein of any particular feature, aspect, method, property,, characteristic, quality,, attribute, element, or the like in connection with an implementation or embodiment can be used in all other implementations or embodiments set forth herein. Any methods disclosed herein need not be performed in the order recited. The methods disclosed herein may include certain actions taken by a practitioner; how ever, the methods can also include any third-party instruction of those actions, either expressly or by implication.
[0076] Example
[0077] System Configuration and Datasets.
[0078] The PSZ system for experimental evaluation comprises a linear array of eight loudspeakers and a rectangular rendering area with dimensions [xmin, xmax] x [ymin, ymax] = [— 1, l]m x [0.5, 2]m. See FIGS. 7 and 8. The loudspeaker array layout is based on an inhouse system. The PSZs are defined as horizontal circular areas with a radius of 0.1 m, approximating the upper limit of human head size.
[0079] Both simulated and measured ATF datasets were utilized in the evaluation. The simulated ATF datasets are generated within the rendering area with a spatial sampling resolution of 5 cm in both x and dimensions. They are computed by treating the loudspeakers as point sources under both anechoic and reverberant conditions (the latter case is simulated Princeton - 104476
[0080] with the image source method). The measured ATF datasets are obtained with the in-house PSZ system in a listening room, with RT60« 0.24s calculated in the range 1300-6300 Hz. The ATFs are measured with a linear array of 16 Earthworks M30 microphones, covering the region of [-0.84, 0.84] m X [0.9, 1.3] m.
[0081] Model Specifications and Training
[0082] The model outputs were set to the real and imaginary parts of the frequency-domain filter coefficients, which correspond to an 8192-length impulse response at 48 kHz sampling frequency, for discrete frequencies between 100 and 1500 Hz for all eight loudspeakers. The frequency upper limit is empirically chosen as pilot experiments show that DZ is difficult to render in some areas above this frequency, potentially due to the loudspeaker spacing. The highest Fourier encoding order, K. is set to 3. To compute the BZ loss term Z}, the target magnitude |pT B|was set as the average of the ATF magnitude responses from one edge loudspeaker and one center loudspeaker, depending on the position of the BZ (e.g., choosing the first and the fourth loudspeakers from the left if xi < 0). To compute the loss term Jfi, a 4th-order 8192-length Butterworth bandpass filter with cutoff frequencies at 100 and 1500 Hz was used. In the implementation, the filter coefficients outside the frequency range (100-1500 Hz) were augmented with the coefficients of the bandpass filters, after normalization by the amplitude at the cutoff frequencies to avoid discontinuity in the spectrum. The weighting function w[n] was taken to be an "inverted" Hamming window with a length of 8192, varying in a cosine fashion from 1 to around 0.1 and back to 1 across the 8192 samples. Because the filters generated in the frequency domain are non-causal by nature, a circular shifting of 4096 samples was applied to w[n] to match the filter response in the loss computation.
[0083] All the models evaluated below have an input size of 4 (2D coordinates for BZ and DZ), an output size of 3824 (8 loudspeakers x 2 real / imaginary parts x 239 frequency bins), and three fully connected layers with 512 neurons for each layer, totaling 2.5 M parameters. A total of 10,000 random combinations of the BZ and DZ positions within the rendering area are used for training each model, with 8.000 samples for training and 2,000 samples for validation. The models are trained with the Adam optimizer with a learning rate of 10'3and a batch size of 32 for a maximum of 400 epochs. The models are implemented with PyTorch and trained on an NVIDIA Al 00 GPU.
[0084] Evaluation Metrics
[0085] Three performance metrics were adopted for evaluation, namely the inter-zone isolation (IZI), inter-program isolation (IPI), and normalized mean squared error (NMSE). Assuming Princeton - 104476
[0086] two PSZs Zi, Z2 with their corresponding ATF sub-matrices as Hi and H2, the filters that render BZ in Zi (or Z2) and DZ in Z2 (or Zi) were denoted as g (or g^). Here, as the filters render mono audio programs in BZ (i.e., a single vector pr), IZI is defined in as:
[0087]
[0088] where the subscript 1 (or 2) of IZI refers to the case of rendering BZ in Zi (or Z2) and DZ in the other. In this particular case, IZI was shown to be equivalent to the commonly used Acoustic Contrast (AC) metric. Correspondingly, IPI for two different BZ / DZ assignments is expressed as
[0089]
[0090] Compared to IZI, which indicates the difference between BZ and DZ, IPI indicates the interference level a listener would perceive when both target and interfering programs are rendered at the same time. The NMSE between the rendered and target amplitude responses in BZ is defined as
[0091]
[0092] All the metrics will be plotted with a logarithmic scale (i.e., taking 10 log10(’)) for better visualization.
[0093] Results
[0094] The effects of loss function hyperparameters on the model’s performance are first examine by evaluating it with simulated ATFs. Next, the models are evaluated with different robustness enhancement methods using measured ATFs. Furthermore, the results of the model customization are shown with measured ATFs that are partially available only in certain areas. Finally, the best-performing model is compared with the traditional methods (PM and AM) for all three metrics. All results below, except for the filter response plots, are either processed by a log-weighted average function in the frequency domain, defined as
[0095] logMean(
[0096]
[0097] if shown as spatial maps or applied with a 1 / 6-octave smoothing if shown as functions of frequency, after taking the logarithm.
[0098] Effects of Loss Function Hyper parameters
[0099] The effects of the gain limit hyperparameter gmaxin the loss term 23 is first investigated on the model’s performance. Princeton - 104476
[0100] Four models were trained with gmaxvaried and the other hyperparameters fixed. These models are trained with ATFs simulated under anechoic conditions and tested with ATFs simulated in a shoebox room with dimensions of 4 m x 7.5 m x 2.5 m and RTeo = 0.24 s; the purpose is to evaluate the performance under unknown reverberant conditions similar to those of the actual system. To evaluate the spatial dependency of the performance, one zone (Zi) is fixed at (-0.5, 1.0) m and the position of the other zone (Z2) is varied within the rendering area.
[0101] First, it was observed from the IZI plots that, although the goal was set of minimizing the energy in DZ regardless of its position for training, it is practically infeasible to achieve high IZI values in the regions that are close to, in front of, or behind the static BZ, due to the limited filter gains and the room reflections. Comparing the four IZI plots, it was observed that IZI gradually decreases as gmaxincreases, especially in the regions near x = 0 and behind the static zone. This suggests that the filters generated w ith a more relaxed gain limit are less robust under reverberant conditions. It was also noted that a region for low IZI values near the static zone in all four IZI plots extends further to the left as it is further away from the static zone. This is because the size of the loudspeaker array limits the capability of rendering DZ further away from the center axis in the far field. The IPI plots have similar trends as the IZI plots, except for minor magnitude differences due to the different definitions. The NMSE plots also show similar spatial patterns as the IZI plots, except when the two zones overlap. In this case, the NMSE values are relatively low because the DZ loss term is abandoned in the training. Moreover, NMSE values are higher in front of the static zone than behind it, suggesting that it is more challenging to render DZ in front of BZ than behind. This is expected as once the controlled wavefronts are formed in BZ, they are less likely to be canceled out from behind.
[0102] Next, the effects of the weighting parameter y associated with Z4, which reflects the filter compactness, were investigate on the performance. Hyperparameters were chosen that correspond to the best model from the previous experiment and y was changed from 0 to 0.5. Looking at the magnitude responses (between 100 and 1500 Hz) and impulse responses (after a circular shift of 4097 samples) of the filters generated by the two models when the centers of BZ and DZ are set at (-0.5. 1.0) m and (0.5, 1.0) m, respectively, it can be seen that adding the loss term ℒ4not only leads to more compact impulse responses but also better reflects the bandpass characteristics in the frequency responses. Further comparisons of the two models in terms of IZI, IPI, and NMSE show similar levels in all three metrics, except for decreased NMSE around 100 Hz in the model with y = 0.5. which is likely due to the resulting shorter Princeton - 104476
[0103] impulse responses. This suggests that adding the loss term ℒ4does not significantly affect the performance but helps to reduce the filter artifacts in the time domain.
[0104] Performance with Simulated Training, Data
[0105] The real-world performance of the models trained with different strategies intended for ensuring robustness was tested. Measured ATFs are used to determine which strategy yields the best performance given minimal knowledge of the actual system. Four models are evaluated:
[0106] (a) Model trained with simulated anechoic ATFs (a = 0.5, / ? = 0.5, gmax= 1 / 8, y = 0.5);
[0107] (b) Model trained with simulated anechoic ATFs after perturbations (a = 0.5, / ? = 0, y = 0.5);
[0108] (c) Model trained with simulated reverberant ATFs ( a = 0.5, / ? = 0, y = 0.5); and (d) Model trained with simulated reverberant ATFs after perturbations ( a — 0.5, fl — 0,y = 0.5).
[0109] More specifically, the ATF perturbations were model as a combination of loudspeaker displacement error d~‘l (— 0.3, 0.3) m, loudspeaker frequency response mismatch 6A~‘l (0.79, 1.26) (corresponding to ±2 dB), and background noise en~JV'(0, cr2) with cr corresponding to 40 dB of signal-to-noise ratio. Tl and JV denote the uniform and normal distributions, respectively. The simulated reverberant ATFs are from 51 different shoebox room configurations with randomized room geometries (7£(4, 7)m X 1 / (4, 7)m X 1 / (2, 4) m) and placement of the loudspeaker array, but the same RTeo as the actual room. 50 room configurations are used for training and one for validation. Note that the gain limit loss term Z3is not included in the training of model (b)-(d) by setting / ? = 0 for evaluating the model’s performance without explicit regularization.
[0110] The IZI, IPI, and NMSE results of the four models evaluated with the measured ATFs were generated, with Zi fixed at (0.5, 1.0) m. Comparing the results of models (a) and (b), it was observed that both perform similarly when the two zones are far apart (corresponding to x< 0), (b) performs clearly worse in all three metrics when Z2 moves near (especially behind) Zi. This indicates that (b) is less robust against the actual system uncertainties without explicit regularization and that ATF perturbations are not sufficient to ensure robustness. However, comparing (a) with (c) and (d), which are trained with reverberant ATFs without explicit regularization, it was observed that the models trained with reverberant ATFs lead to higher IZI and IPI when Z2 is behind and to the right of Zi. (c) and (d) also yield significantly lower Princeton - 104476
[0111] NMSE values than (a) in the region where the two zones overlap, but at a cost of slightly higher NMSE values where x ≤ 0. This suggests that using reverberant ATFs for training may replace the explicit regularization term and achieve higher robustness against the actual system uncertainties. Comparing (c) with (d), it was observed that the ATF perturbations do not significantly affect the model's performance, except for the improvement of IPI near the right boundary of the rendering area and NMSE in the overlapping region. This suggests that room reflections may be the dominant source of uncertainties in the actual system.
[0112] Performance with Mixed Simulated / Measured Training Data
[0113] In contrast to the previous evaluation where measured data is unavailable for model training, three additional models were trained with a mixture of simulated ATFs and the measured ATFs from the previous evaluation. The amount of measured ATFs in the training data is varied for each model by spatially sampling the measured ATFs with different spacings. As the measured ATFs are not evenly distributed in the rendering area, one can sample the data by specifying circular regions with the same size as the PSZs and with spacings of A = {0.2, 0.4, 0.6} m between the centers of the circles and selecting the measured ATFs that fall within the circles. The rest of the rendering area, once the measured ATFs are selected, is still filled with simulated ATFs from multiple room configurations. All three models use the same settings as model (d), except that the weighting of the DZ loss term (ℒ2) corresponding to the measured ATFs is increased by a factor of 2 to prioritize the isolation performance in the measurement region.
[0114] Spatial maps were created of IZI, IPI, and NMSE for model (d) and the three models trained with different sets of measured ATFs, with Zi fixed at (0.7, 1.0) m. Comparing the results of the four models, a gradual improvement in IZI and IPI was observed as more measured ATFs are included in the training data. For the cases of A = 0.6 m and A = 0.4 m. the increase in IZI mostly appears in the regions where the measured ATFs are selected because the models are designed to prioritize the minimization of DZ energy in these regions; the increase in IPI is more evenly distributed across the rendering area because the emphasized DZ loss is associated with the region near the fixed Zi. therefore affecting the overall performance. This indicates that the model can still benefit from the limited measured data when the PSZ is in the vicinity of the measurement region. When measured ATFs are available for the entire rendering area (A = 0.2 m), the model achieves the highest IZI and IPI values (over 15 dB in most regions). For NMSE, it was observed that the models with partially available measured ATFs yield higher NMSE in the measurement regions but lower NMSE in the rest of the rendering area. This is due to prioritizing the DZ loss term (and therefore sacrificing the BZ Princeton - 104476
[0115] loss term) when Z2 moves into the measurement region. However, the compromised NMSE performance is still comparable to that of the baseline model (trained with simulated ATFs only). In practice, the trade-off between the isolation performance and the rendering quality can be adjusted by varying the hyperparameter a based on the system requirements.
[0116] Comparison with Traditional Methods
[0117] Lastly, the best-performing SANN model was compared with the traditional methods (Pressure Matching (PM) and Amplitude Matching (AM)) for all three metrics. It was assumed that measured data is unavailable for training and simulated ATFs were used for filter generation. As the robustness enhancement methods do not directly apply to the traditional methods, anechoic ATFs were use for filter generation with PM and AM and explicit regularization was relied upon to ensure robustness.
[0118] A typical cost function for PM may be defined as
[0119] 풥PM = ‖pT- p‖2+ β‖g‖2= ‖pT- Hg‖2+ λ‖g‖2, (5) where pTis the target sound pressure at the control points and A = A(w) is the regularization parameter.
[0120] A typical cost function for AM may be given by
[0121] 풥AM = ‖|pT| - |Hg|‖2+ λ‖g‖2, (6) The regularization parameter A in the cost functions of PM (Eq. 5) and AM (Eq. 6) are set to σmax(H(w)) × 0.05 for each frequency w. σmaxdenotes the largest singular value of the matrix. While not shown here, it was found that the filter performance with PM and AM is insensitive to the choice of A as long as it is within a reasonable range. The model evaluated is the same as model (d). The AM method is implemented with a majorization-minimization algorithm.
[0122] Plots were created that, for simplicity, only showed the results for the case where Zi is fixed at (0.6, 1.0) m and Z2 is at (-0.6, 1.0) m, showing, e.g., filter magnitude responses for the case of Zi being BZ, and the IZI, IPI, and NMSE results for the SANN model (model (d)) and the PM and AM methods.
[0123] It was observed from the filter magnitude response plots that the filters generated by the SANN model have a lower average magnitude at low frequencies than those generated by PM or AM; this is likely the outcome of optimizing the robustness against room reflections during model training. Moreover, the filters generated by AM have a significantly larger variance in magnitude than the other two approaches as phase is not constrained in the optimization. For the performance metrics, it was seen that the SANN model has equal or better Princeton - 104476
[0124] (especially below 200 Hz and around 1000 Hz) IZI and IPI performance than PM and AM. It also yields lower NMSE than PM and AM above 300 Hz but higher NMSE below 300 Hz, which is likely due to the filter compactness enforced by the loss term ℒ4. This suggests that by effectively utilizing the limited knowledge of the environment (e.g., reverberation level and rough room geometry), the SANN model can yield more robust filters than the traditional methods. However, when measured data is available and only stationary’ PSZ rendering is required, the traditional methods may yield comparable or better performance than the SANN model as robustness is not a maj or concern in that case.
[0125] It is also worth comparing the actual implementation cost of the model to that of the traditional methods. The SANN model, once trained, requires only a forward pass of the neural network to generate the filters. Under the current experiment setup, the evaluated model has a total of 2.5 M parameters (which takes about 10 MB of memory) and takes about 0.65 ms on an Apple M1 Pro processor (using CPU only) for a single filter generation. In contrast, with the traditional methods, it would require either ~5891 MB to store the precomputed filters for all possible combinations of BZ and DZ positions (assuming a spatial grid with 5 cm resolution in the region of [-1, 1] m x [0.5, 2] m), or significantly more computation resources to generate the filters on-the-fly as the ATF matrices are constantly changing with the positions of the zones. For example, a single filter generation based on the closed-form PM solution takes about 8.4 ms on the same CPU. The AM method would be even more computationally expensive as it requires iterative optimization for each filter generation.
[0126] As can be seen, the disclosed deep learning-based approach utilizes a spatially adaptive neural network (SANN) model for rendering head-tracked PSZs. Depending on whether measured ATFs are available, the SANN model can be trained either with simulated ATFs for robustness in unknown environments or with a mixture of simulated and measured ATFs for better isolation performance (i.e., IZI and IPI) in the measurement region.
[0127] In evaluating the model's performance, it was found that although the listeners can move freely within the rendering area, the isolation is fundamentally limited by the loudspeaker array size and the relative positions of the PSZs. It was also found that the model is able to reduce the filter's artifacts (e.g., pre-ringing) by adding the time-domain loss term without significantly affecting the isolation performance, and that training with randomized room configurations yields better filter robustness compared to limiting the filter gains.
[0128] In comparison with traditional methods, it was found that the SANN model yields similar or better isolation performance than the PM and AM methods with explicit regularization when no measured ATFs are available, and has fewer artifacts in the filter Princeton - 104476
[0129] responses. The model is also more efficient in terms of computation and storage costs, making the real-time implementation of head-tracked PSZ rendering more feasible.
[0130] It should be noted that the model's architecture adopted in the example and its associated loss function are only an example of the possible design choices, and can be easily modified to meet other system requirements. For example, the model outputs can be set to be time-domain filter impulse responses, or to frequency-domain filter coefficients of different lengths; in addition, thanks to the flexibility of the loss calculation, other loss terms that are formulated in both the frequency and time domains can also be added to the model to further improve the performance. Moreover, the proposed approach applies also to a wide range of other audio rendering tasks that require head tracking or position-dependent processing, such as crosstalk cancellation, loudspeaker equalization, and sound field synthesis.
[0131] To simplify the analysis, an assumption was made that the actual listeners are absent in both filter generation and evaluation. Although the acoustic scattering effects due to the listener's head are not dominant in the frequency range of interest (100-1500 Hz), the model's performance may be significantly affected by torso-related effects and the interlistener scattering effect at higher frequencies in realistic scenarios. Therefore, the model may advantageously incorporate listener-specific information, such as the head-related transfer functions (HRTFs) and / or the anthropometric data of the listener, to achieve better isolation performance at higher frequencies.
[0132] The ranges disclosed herein also encompass any and all overlap, sub-ranges, and combinations thereof. Language such as “up to,” “at least,” “greater than,” “less than,” “between,” and the like includes the number recited. Numbers preceded by a term such as “about” or “approximately” include the recited numbers and should be interpreted based on the circumstances (e.g., as accurate as reasonably possible under the circumstances, for example ±5%, ±10%, ±15%, etc.). For example, “about 3.5 mm” includes "3.5 mm.” Phrases preceded by a term such as “substantially” include the recited phrase and should be interpreted based on the circumstances (e.g., as much as reasonably possible under the circumstances). For example, “substantially constant” includes “constant.” Unless stated otherwise, all measurements are at standard conditions including ambient temperature and pressure.
Claims
Princeton - 104476What is claimed is:
1. A method for rendering a personal sound zone (PSZ) with head tracking using multiple loudspeakers, comprising:i) configuring a PSZ rendering system by determining a loudspeaker layout, a size and shape of the PSZ and a rendering area of the PSZ, a range of head movement to be tracked, and a range of room acoustics parameters in which the PSZ rendering system is designed to operate;ii) providing a spatially adaptive neural network (SANN) that takes in listener head coordinates and outputs PSZ filters;iii) defining a unified loss function that incorporates one or more PSZ design objectives;iv) generating acoustic transfer functions (ATFs) between the multiple loudspeakers and control points in the rendering area for training the SANN;v) training the SANN using the ATFs to generate the PSZ filters with the unified loss function; andvi) generating the PSZ filters in real-time based on head position data, and convolving the PSZ filters with audio programs to be played through the multiple loudspeakers.
2. The method of claim 1, wherein the SANN comprises:a) an input layer that receives head coordinates from one or multiple listeners; b) a normalization layer that scales the head coordinates to provide normalized head coordinates;c) a positional encoding layer that converts the normalized head coordinates into a higher-dimensional representation to provide encoded features;d) a series of intermediate layers for processing the encoded features; and e) an output layer that generates the PSZ filters for each of the multiple loudspeakers based at least on information from the series of intermediate layers.
3. The method of claim 1, wherein the SANN comprises:an input layer that receives head coordinates from one or more listeners;a series of intermediate layers for processing at least information based on the head coordinates from the one or more listeners to provide intermediate information; andPrinceton - 104476an output layer that generates the PSZ filters for each of the multiple loudspeakers based at least on the intermediate information.
4. The method of claim 3, wherein the input layer is operable to scales the head coordinates to provide normalized head coordinates and to converts the normalized head coordinates into a higher-dimensional representation.
5. The method of claim 1, wherein the SANN comprises:an input layer that receives head coordinates from one or more listeners;a positional encoding layer that converts the head coordinates into a higherdimensional representation to provide encoded features;a series of intermediate layers for processing the encoded features; andan output layer that generates the PSZ filters for each of the multiple loudspeakers based at least on information from the series of intermediate layers.
6. The method of claim 1, wherein the PSZ rendering system is configured to render a single bright zone (BZ) for one listener or both at least one BZ and at least one dark zone (DZ) that adapt simultaneously to movements of multiple listeners.
7. The method of claim 1, wherein generating ATFs comprises:i) providing simulated ATFs by numerically simulating ATFs between the multiple loudspeakers and control points in the rendering area under various room acoustics conditions;ii) providing measured ATFs by measuring ATFs using a microphone array or a mannequin head with microphones in specific rendering environments; andiii) combining the simulated ATFs and the measured ATFs by replacing simulated ATFs with measured ATFs in regions where measurements were taken.
8. The method of claim 1, wherein generating ATFs only includes simulating ATFs.
9. The method of claim 1, wherein generating ATFs only includes measuring ATFs.
10. The method of claim 1, wherein the unified loss function comprises:Princeton - 104476a) a bright zone (BZ) loss term that minimizes a difference between a target pressure and an estimated pressure at control points within the BZ:b) a dark zone (DZ) loss term that minimizes a pressure within the DZ; and c) a filter loss term that constrains properties of the generated PSZ filters.
11. The method of claim 10, wherein the unified loss function contains multiple BZ and DZ loss terms for multiple defined PSZs.
12. The method of claim 1, wherein the unified loss function comprises a bright zone (BZ) loss term that minimizes a difference between a target pressure and an estimated pressure at control points within the BZ.
13. The method of claim 1, wherein training the SANN comprises:i) selecting ATFs corresponding to control points within PSZs centered around input listener head positions:ii) augmenting the selected ATFs with noise or distortions to improve SANN robustness;iii) generating PSZ filters with the input listener head positions;iv) computing the unified loss function based on the generated PSZ filters and the selected ATFs: andv) updating model weights of the SANN using backpropagation and optimization algorithms.
14. The method of claim 1, wherein training the SANN comprises:selecting ATFs corresponding to control points within PSZs centered around input listener head positions;generating PSZ filters with the input listener head positions;computing the unified loss function based on the generated PSZ filters and the selected ATFs: andupdating model weights of the SANN using backpropagation and optimization algorithms.
15. The method of claim 1, wherein generating the PSZ filters and convolving the PSZ filters with audio programs comprises:Princeton - 104476i) receiving the head position data from a head tracking device;ii) generating PSZ filters in real-time using the trained SANN based on the received head position data;iii) convolving audio programs with the generated PSZ filters to produce filtered audio programs; andiv) playing the filtered audio programs through the multiple loudspeakers to render PSZs.
16. The method of Claim 15, further comprising bandpass filtering the generated PSZ filters.
17. The method of Claim 1, wherein generating the PSZ filters and convolving the PSZ filters with audio programs comprises:receiving a stream of head position data;generating updated PSZ filters in real-time using the trained SANN based on the stream of head position data; andconvolving the audio programs with the updated PSZ filters to render dynamic PSZs.
18. The method of claim 1, wherein generating the PSZ filters and convolving the PSZ filters with audio programs comprises:accessing pre-defined head position data;generating PSZ filters in real-time using the trained SANN based on the pre-defined head position data;convolving audio programs with the generated PSZ filters to produce filtered audio programs; andplaying the filtered audio programs through the loudspeakers to render PSZs.
19. A system for rendering personal sound zones (PSZs) with head tracking using multiple loudspeakers, comprising:a) multiple loudspeakers arranged in a configuration to render PSZs;b) a head tracking device that provides head position data of one or multiple listeners; andc) a processor that receives the head position data and generates signals for the multiple loudspeakers by:Princeton - 104476i) configuring a PSZ rendering system by determining a loudspeaker layout, a size and shape of a PSZ and a rendering area of the PSZ, a range of head movement to be tracked, and a range of room acoustics parameters in which the system is designed to operate;ii) providing a spatially adaptive neural network (SANN) that takes in listener head coordinates and outputs PSZ filter coefficients;iii) defining a unified loss function that incorporates one or more PSZ design objectives;iv) generating acoustic transfer functions (ATFs) between the multiple loudspeakers and control points in the rendering area for training the SANN;v) training the SANN using the ATFs to generate the PSZ filters with the unified loss function; andvi) generating the PSZ filters in real-time given a stream of head position data from a head tracking device, and convolving the PSZ filters with audio programs to be played through the multiple loudspeakers.
20. The system of claim 19, wherein the rendered audio program is personalized in content and volume to suit a preference of a listener.
21. The system of claim 19, wherein the rendered audio program comprises spatial enhancement with crosstalk cancellation technique.
22. A computer-implemented method comprising:under control of a computing system comprising one or more computing devices configured to execute specific instructions, obtaining training data that includes location data for a plurality of control points and a plurality of acoustic transfer functions (ATFs) configured to use a set of loudspeakers to attenuate sound energy at dark zones (DZs) that include one or more of the control points and to preserve sound energy at bright zones (BZs) that include one or more of the control points; andtraining a machine learning model using the obtained training data, wherein the machine learning model is trained to generate model output data for personal sound zone (PSZ) filter parameters.Princeton - 10447623. The computer-implemented method of claim 22, wherein the machine learning model comprises a spatially adaptive neural network (SANN).
24. The computer-implemented method of claim 22, wherein obtaining the training data comprises numerically simulating the ATFs.
25. The computer-implemented method of claim 22, wherein obtaining the training data comprises measuring the ATFs using a microphone.
26. The computer-implemented method of claim 22, further comprising generating PSZ filter parameters based on listener position data using the trained machine learning model.
27. The computer-implemented method of claim 26, further comprising applying the PSZ filter parameters to audio content to render a PSZ.
28. The computer-implemented method of claim 27, further comprising:receiving updated listener position data;generating updated PSZ filter parameters based on the updated listener position data using the trained machine learning model; andapplying the updated PSZ filter parameters to the audio content so that the rendered PSZ moves with a position of a listener.