APPARATUS AND METHOD FOR PERCEPTUALLY BASED CLUSTERING OF OBJECT-BASED AUDIO SCENE - Patent application

The apparatus and method for perceptual audio scene clustering address the limitations of existing algorithms by grouping audio objects based on perceptual models, achieving efficient reduction of audio objects while maintaining high quality and reducing computational demands.

JP2025533617APending Publication Date: 2025-10-07FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025518553
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-27
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

State-of-the-art object-based audio clustering algorithms do not consider the perceptual properties corresponding to a listener, neglecting the position dependency of spatial localization accuracy in human hearing, which affects the efficiency and quality of immersive audio experiences.

Method used

An apparatus and method for perceptual audio scene clustering that groups audio objects into clusters based on perceptual models, using techniques such as Gaussian Mixture Models (GMM), hierarchical clustering, and Just Noticeable Difference (JND) to reduce the number of audio objects while maintaining high perceptual quality, incorporating optimizations for temporal stability and signal processing.

Benefits of technology

The proposed method effectively reduces the number of audio objects while preserving the perceptual quality of immersive audio experiences, enhancing transmission efficiency and reducing computational demands by leveraging human auditory masking effects and localization accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533617000001_ABST
    Figure 2025533617000001_ABST
Patent Text Reader

Abstract

According to one embodiment, an apparatus (100) is provided. The apparatus (100) comprises an input interface (110) for receiving information relating to three or more audio objects. The apparatus (100) further comprises a cluster generator (120) for generating two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that, for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that, for each of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster. The cluster generator (120) is configured to generate the two or more audio object clusters according to a perceptually based model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus and method for perceptual clustering of object-based audio scenes. [Background technology]

[0002] Modern audio playback systems enable immersive three-dimensional (3D) sound experiences. One common format for 3D sound reproduction is channel-based audio, where individual channels associated with defined speaker positions are generated by multi-microphone recording or studio-based production. Another common format for 3D sound reproduction is object-based audio, which utilizes so-called audio objects, which are placed in the listening room by the producer and converted by the rendering system into speaker or headphone signals for playback. Object-based audio offers a high degree of flexibility in terms of sound scene design and playback. It should be noted that channel-based audio can be considered a special case of object-based audio, where sound sources (=objects) are placed at fixed positions corresponding to defined speaker positions.

[0003] To improve the efficiency of transmission and storage of object-based immersive audio scenes, as well as to reduce the computational demands of real-time rendering, it is beneficial, and even necessary, to reduce or limit the number of audio objects. This is achieved by identifying groups or clusters of adjacent audio objects and combining them into a smaller number of sound sources. This process is called object clustering or object consolidation.

[0004] It has been shown in the literature that the localization accuracy of human hearing is limited and dependent on sound source position (e.g., horizontal localization is more accurate than vertical localization), and that auditory masking effects can be observed between spatially distributed sound sources. By exploiting these limitations in localization accuracy and auditory masking effects for object clustering, a significant reduction in the number of audio objects can be achieved while maintaining high perceptual quality. To reduce the number of audio objects while maintaining high perceptual quality, methods and algorithms have been developed to perform object-based audio clustering based on the perceptual characteristics of the audio scene corresponding to the listener.

[0005] In the state of the art, auditory masking and localization models are known. Furthermore, the state of the art presents directional loudness maps (DLMs). Examples include: C. Avendano, “Frequency-domain source identification and manipulation in stereo mixes for enhancement,suppression and re-panning applications,” 2003 IEEE Workshop on Applications of Signal Processing to Audio, and P. Delgado, J. Herre, "Objective Assessment of Spatial Audio Quality using Directional Loudness Maps," in Proc. 2019 IEEE ICASSP,.

[0006] Furthermore, the state of the art presents object clustering algorithms, e.g. J. Herder. "Optimization of Sound Spatialization Resource Management through Clustering", The Journal of Three Dimensional Images, 1999. Nicolas Tsingos, Emmanuel Gallo, George Drettakis, "Perceptual Audio Rendering of Complex Virtual Environments", SIGGRAPH, 2004. Breebaart, Jeroen, Cengarle, Giulio, Lu, Lie, Mateos, Toni, Purnhagen, Heiko, and Tsingos, Nicolas, "Spatial Coding of Complex Object-Based Program Material," JAES Volume 67 Issue 7 / 8, pages 486-497, July 2019.

[0007] Furthermore, the state of the art presents the GMM expectation-maximization algorithm (EM algorithm). Summary of the Invention [Problem to be solved by the invention]

[0008] State-of-the-art algorithms for object-based audio clustering consider the spatial properties of audio objects relative to each other, but they do not consider the perceptual properties corresponding to a listener and therefore do not consider the position dependency of spatial localization accuracy in human hearing. [Means for solving the problem]

[0009] The object of the present invention is to provide an improved concept for object-based audio scene clustering, which is solved by an apparatus according to claim 1, a decoder according to claim 20, a method according to claim 21, a method according to claim 22 and a computer program according to claim 23.

[0010] According to one embodiment, an apparatus is provided, comprising an input interface for receiving information relating to three or more audio objects. The apparatus further comprises a cluster generator, the cluster generator generating the two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that, for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that, for each of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster. The cluster generator is configured to generate the two or more audio object clusters according to a perceptually based model.

[0011] Further provided is a decoder comprising: a decoding unit for decoding the encoded information to obtain information relating to two or more audio object clusters, wherein the two or more audio object clusters are generated by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that, for each of the two or more audio object clusters, at least one of three or more audio objects is associated with said audio object cluster and, for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with said audio object cluster, wherein the two or more audio object clusters are generated in accordance with a perceptually based model; and a signal generator for generating two or more audio output signals in accordance with the information relating to the two or more audio object clusters.

[0012] There is also provided a method according to an embodiment, the method comprising: receiving information about three or more audio objects; and generating two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with said audio object cluster, and such that for each of the at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with said audio object cluster, wherein the generation of the two or more audio object clusters is performed according to a perceptually based model.

[0013] Further, there is provided a method according to another embodiment, the method comprising: - decoding the encoded information to obtain information about two or more audio object clusters, wherein the two or more audio object clusters are generated by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that, for each of the two or more audio object clusters, at least one of three or more audio objects is associated with said audio object cluster, and such that, for each of the two or more audio object clusters, at least two of the three or more audio objects are associated with said audio object cluster, wherein the two or more audio object clusters are generated according to a perceptually based model; and generating two or more audio output signals in response to information about two or more audio object clusters;

[0014] Furthermore, computer programs are provided, each computer program being configured to perform one of the methods described above when run on a computer or signal processor. According to one embodiment, a perceptually based clustering algorithm groups audio objects in an audio scene into clusters and combines the original objects into fewer output objects, e.g., by combining their signals and by selecting common centroid positions as output object positions, e.g., based on perceptual model criteria. Based on the target use case, the goal can be to achieve a given (maximum) number of output clusters or to reduce the number of objects in the scene without introducing perceptible differences beyond a given tolerance. This can be achieved using various embodiments presented below.

[0015] Some embodiments relate to clustering of audio objects. According to one embodiment, Gaussian Mixture Model (GMM) based clustering is provided. This generative clustering method can, for example, compute a 3D directional loudness map (3D-DLM) for the entire acoustic scene to represent the overall spatial characteristics of the scene. A GMM is fitted to approximate the original DLM with a given number of components representing the corresponding number of clusters. Thus, the algorithm aims to reproduce the overall spatial characteristics of the acoustic scene rather than considering individual object characteristics. This method is particularly useful when a dense acoustic scene consisting of a large number of objects needs to be represented by only a small number of cluster locations, e.g., for low-complexity / low-bitrate applications.

[0016] In one embodiment, hierarchical clustering is provided. This "agglomerative" clustering approach iteratively combines objects based on a perceptual distance metric, for example, until a target number of clusters is reached and / or a given limit on the distance metric is reached (e.g., all imperceptible differences are eliminated). This approach is computationally efficient and provides flexibility for configuration to constant quality or constant rate applications. Furthermore, it scales well enough that, for example, if the number of participating audio objects falls below the maximum number of allowed clusters, their presence becomes imperceptible.

[0017] According to one embodiment, clustering based on JND (Just Noticeable Difference) is provided. This can be considered a simplified special case of hierarchical clustering methods. That is, if objects are so close that their positions cannot be distinguished, they can be combined to reduce redundancy, for example, without perceptible difference in the overall acoustic scene. Therefore, a JND-based clustering method determines groups of all objects that are within JND of each other in terms of a perceptual distance metric and combines them into clusters. This method requires low computational complexity and produces a variable number of output clusters with perceptual quality that makes the intervening clusters (almost) imperceptible.

[0018] Further embodiments provide improved performance. Additionally, several optimizations regarding temporal stability and resulting cluster output positions are developed. For example, according to one embodiment, temporal stabilization is provided. Because clustering algorithms typically operate on a frame-by-frame basis, several measures can be taken to improve the temporal stability of the cluster algorithm's results. That is, object membership to clusters can be stabilized, for example, by a penalty factor for reassigning objects to clusters in a perceptual distance metric. In the case of DLM-based approaches, the DLM can be temporally smoothed to improve temporal stability, for example. Permutations of the cluster index order can be identified and optimized to improve the stabilization of the output signal and position metadata, for example.

[0019] And / or, for example, in one embodiment, centroid position optimization is provided. Clustering algorithms typically result in cluster centroid positions and object cluster membership. However, the output cluster positions can be further optimized using perceptual criteria, for example, taking into account the target playback scenario.

[0020] According to some embodiments, a signal mixing and processing concept is provided. Based on the results of the proposed clustering algorithm, the signals of the input audio objects can be, for example, mixed and combined to obtain an output cluster signal. The signal processing in this mixing stage can also be perceptually optimized by several aspects, such as, for example, crossfading to avoid signal discontinuities, and / or processing correlations between signals, and / or taking into account distance-based gain differences, and / or equalization to compensate for changes in spectral localization cues. In the following, embodiments of the invention will be explained in more detail with reference to the drawings. [Brief explanation of the drawings]

[0021] [Figure 1] 1 illustrates an apparatus according to one embodiment. [Figure 2] 1 illustrates a decoder according to one embodiment. [Figure 3] 1 illustrates a system according to one embodiment. [Figure 4] We show a one-dimensional example where the directional loudness map produced by 10 sound sources is approximated by a Gaussian mixture model with only two components. [Figure 5] 1 illustrates three different distance model levels for JND-based clustering according to an embodiment. [Figure 6a] 1 illustrates a small example of a level 2 JND-based clustering algorithm according to one embodiment. [Figure 6b] 1 illustrates a small example of a level 2 JND-based clustering algorithm according to one embodiment. [Figure 6c] 1 illustrates a small example of a level 2 JND-based clustering algorithm according to one embodiment. [Figure 6d] 1 illustrates a small example of a level 2 JND-based clustering algorithm according to one embodiment. [Figure 6e] 1 illustrates a small example of a level 2 JND-based clustering algorithm according to one embodiment. [Figure 6f] 1 illustrates a small example of a level 2 JND-based clustering algorithm according to one embodiment. [Figure 6g] 1 illustrates a small example of a level 2 JND-based clustering algorithm according to one embodiment. [Figure 7] 10 illustrates cluster index permutation according to one embodiment due to slight changes in the scene. [Figure 8] 1 illustrates cluster assignment permutation and optimization according to one embodiment. [Figure 9] The barycentric projection on the unit sphere in the horizontal plane and the barycentric projection on the perceptual coordinate system in the horizontal plane are shown. [Figure 10]10 illustrates a projection of the centroid onto a cone of confusion in a side view according to one embodiment. [Figure 11] 10 illustrates a height-preserving centroid projection onto a cone of confusion at a lateral side, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0022] FIG. 1 illustrates an apparatus 100 according to one embodiment. The device 100 comprises an input interface 110 for receiving information about three or more audio objects. The apparatus 100 further comprises a cluster generator 120 that generates two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with said audio object cluster, and such that for each of the at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with said audio object cluster. The cluster generator 120 is configured to generate the two or more audio object clusters according to a perceptually based model.

[0023] According to an embodiment, the cluster generator 120 may be configured to generate the two or more audio object clusters according to a perceptually based model, for example by generating the two or more audio object clusters according to at least one of a perceptual distance metric, a directional loudness map, a perceptual coordinate system, and a spatial masking model. In one embodiment, the cluster generator 120 may be configured to generate two or more audio object clusters according to a perceptual distance metric, for example, by determining, for a pair of two audio objects of the three or more audio objects, whether the two audio objects have a perceptual distance according to a perceptual distance metric that is less than or equal to a threshold, and if the perceptual distance is less than or equal to the threshold, by associating the two audio objects with the same one of the two or more audio object clusters.

[0024] According to one embodiment, the cluster generator 120 may be configured to generate two or more audio object clusters according to a perceptual distance metric, for example by iteratively associating two perceptually closest audio objects of the three or more audio objects according to the perceptual distance metric, until a predetermined target number of audio object clusters is reached or a predetermined maximum perceptual distance according to the perceptual distance metric is exceeded. In one embodiment, the cluster generator 120 may be configured to generate two or more audio object clusters, for example, in response to a three-dimensional directional loudness map.

[0025] According to an embodiment, the cluster generator 120 may be configured to generate two or more audio object clusters, for example by using a Gaussian mixture model. Further, the cluster generator 120 may be configured to determine two or more audio object clusters, for example by determining components of the Gaussian mixture model such that a three-dimensional directional loudness map is approximated. In one embodiment, the cluster generator 120 may be configured to generate two or more audio object clusters, for example, by using a Gaussian mixture model. Further, the cluster generator 120 may be configured to determine two or more audio object clusters, for example, by using an expectation-maximization algorithm to fit weighted data points onto an arbitrary grid of a Gaussian mixture model.

[0026] According to one embodiment, the cluster generator 120 may be configured to perform, for example, a perceptual optimization of the centroid positions resulting from the clustering. In one embodiment, the cluster generator 120 may be configured to perform cluster assignment and centroid location optimization, for example, depending on the spectral matching of two or more audio object clusters.

[0027] According to an embodiment, the cluster generator 120 may be configured to generate two or more audio object clusters as a first plurality of audio object clusters, for example by creating an association of each of the three or more audio objects with at least one of the two or more audio object clusters. Further, the cluster generator 120 may be configured to generate a second plurality of two or more audio object clusters, for example, whereby at least one audio object of the three or more audio objects is associated with an audio object cluster of the second plurality of audio object clusters that is different compared to an audio object cluster of the first plurality of audio object clusters with which said at least one audio object was associated.

[0028] In one embodiment, the cluster generator 120 may be configured to generate the second plurality of two or more audio object clusters, for example, depending on temporal smoothing and / or depending on one or more penalty factors in the perceptual distance metric. According to one embodiment, the cluster generator 120 may be configured to generate the second plurality of two or more audio object clusters, for example by optimizing the cluster assignment permutation according to the energy distributions of the three or more audio objects. In one embodiment, the cluster generator 120 may be configured to generate the second plurality of two or more audio object clusters, for example by stabilizing the resulting cluster centroid positions by hysteresis.

[0029] According to one embodiment, the cluster generator 120 may be configured to generate the second plurality of two or more audio object clusters, for example by performing a perceptual optimization of the centroid positions resulting from the clustering to generate the first plurality of two or more audio object clusters. In one embodiment, the cluster generator 120 may be configured to generate the second plurality of two or more audio object clusters, for example by performing cluster assignment and centroid location optimization according to spectral matching of the first plurality of audio object clusters.

[0030] According to one embodiment, the cluster generator 120 may be configured to perform signal processing, for example, for each audio object cluster to which at least two of the three or more audio objects are associated, by combining the audio object signals of each audio object associated with said audio object cluster. In one embodiment, the cluster generator 120 may be configured to perform, for example, at least one of the following: Crossfading to prevent signal discontinuities in the reassignment of object membership to clusters; Considering signal correlation to achieve energy conservation, distance-based gain adjustment, Equalization to compensate for perceptual differences due to spectral cues.

[0031] According to one embodiment, the cluster generator 120 may be configured to generate two or more audio object clusters, for example depending on the actual or assumed location of the listener. In one embodiment, the cluster generator 120 may be configured to determine one or more characteristics of each audio object cluster of two or more audio object clusters depending on one or more characteristics of three or more audio objects associated with said audio object cluster, for example, said one or more characteristics being: an audio signal associated with said audio object cluster; a location associated with said audio object cluster; Contains at least one of the following: According to an embodiment, the apparatus 100 may further comprise an encoding unit for generating encoded information, for example encoding information about two or more audio object clusters.

[0032] FIG. 2 shows a decoder 200 according to one embodiment. The decoder 200 comprises a decoding unit 210 for decoding the encoded information to obtain information regarding two or more audio object clusters, the two or more audio object clusters being generated by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that for each of the two or more audio object clusters, at least one of three or more audio objects is associated with said audio object cluster, and such that for each of the two or more audio object clusters, at least two of the three or more audio objects are associated with said audio object cluster, the two or more audio object clusters being generated according to a perceptually based model. Furthermore, the decoder 200 comprises a signal generator 220 for generating two or more audio output signals in response to information about the two or more audio object clusters.

[0033] FIG. 3 illustrates a system according to one embodiment. The system comprises the apparatus 100 of Figure 1. The apparatus 100 of Figure 1 further comprises an encoding unit for generating encoded information that encodes information about the two or more audio object clusters. Furthermore, the system comprises a decoding unit 210 for decoding the encoded information to obtain information about two or more audio object clusters. Furthermore, the system comprises a signal generator 220 for generating two or more audio output signals in response to information about the two or more audio object clusters.

[0034] Before describing the preferred embodiment in more detail, some consideration of the background art on which the embodiment of the present invention is based is provided. Perceptual models will now be considered and an overview of the perceptual models underlying clustering algorithms and methods according to embodiments will be presented. The presented psychoacoustic model can, for example, include the following main components that correspond to different aspects of human perception: a 3D directional loudness map, a perceptual coordinate system, a spatial masking model, and a perceptual distance metric.

[0035] We first describe the 3D directional loudness map (3D-DLM). The underlying concept of the directional loudness map (DLM) is to find a representation of how much loudness is perceived to come from a given direction. This concept has already been presented as a one-dimensional approach to represent binaural localization in binaural DLM (Delgado et al., 2019). Here, this concept is extended to three-dimensional (3D) localization by creating a 3D-DLM on a surface surrounding the listener to uniquely represent perceived loudness as a function of the angle of incidence relative to the listener. Note that while binaural DLMs are derived by analyzing signals at the ears, 3D-DLMs are synthesized for object-based audio by utilizing a priori known sound source positions and signal characteristics.

[0036] Here, a perceptual coordinate system (PCS) is presented. Human sound source localization accuracy varies with spatial direction. To represent this in a computationally efficient manner, a perceptual coordinate system (PCS) is introduced. To obtain this PCS, spatial locations are distorted to accommodate the non-uniform characteristics of localization accuracy. Thus, distances within the PCS correspond not to physical distances but to "perceived distances" between locations, e.g., the number of just noticeable differences (JNDs). This principle is similar to the use of psychoacoustic frequency scales in perceptual audio coding, such as the Bark scale or the ERB scale.

[0037] Next, we describe spatial masking models (SMMs). Monophonic time-frequency auditory masking models are a fundamental element of perceptual audio coding and are often refined by binaural (un)masking models to improve stereo coding. Spatial masking models extend this concept for immersive audio to incorporate and exploit masking effects between arbitrary sound source positions in three dimensions. Regarding perceptual distance metrics, it should be noted that the above components can be combined, for example, to obtain a perceptually based distance metric between spatially distributed sound sources. These can be used in a variety of applications, e.g., as a cost function in object clustering algorithms, to control bit distribution in perceptual audio coders, and to obtain objective quality assessments. These metrics address questions such as "How perceptible is it if a sound source's position changes?", "How perceptible is the difference between two different representations of an audio scene?", and "How important is a given sound source in the overall audio scene (and how noticeable would removing it be)?"

[0038] In the following, the developed clustering concepts and algorithms are presented. In applications using object-based audio, it is desirable to reduce the number of objects required to represent an audio scene while maintaining high perceptual quality in order to improve transmission, storage efficiency, and computational complexity in rendering applications. Thus, for example, perceptually based clustering of audio objects can be used. In other words, based on a proposed perceptual model, audio objects with similar perceptual properties can be grouped and combined into, for example, fewer audio objects. Depending on the use case, there is a wide range of desired object characteristics and how much the number of objects in a scene should be reduced. In the field of audio coding, there are well-known methodologies that aim for constant quality with a variable bit rate (VBR) or for providing variable quality with a constant bit rate (CBR). Correspondingly, object clustering can be configured to aim for constant quality resulting in a variable number of clusters (= output objects), or for a constant number of simultaneous objects with variable quality.

[0039] The most conservative approaches aim only at removing redundancy and irrelevance in the scene representation. This means, for example, that only objects that can be combined without introducing audible changes to the scene can be consolidated to reduce the number of objects without affecting perceived quality ("transparent" clustering). This approach can also be extended to further reduce the total number of objects, for example, by clustering objects within a selected threshold of a perceptual distance metric, i.e., a maximum distance (e.g., a multiple of the JND distance). These approaches can, for example, result in a variable number of clusters and therefore a variable number of output objects.

[0040] On the other hand, in many applications, the maximum number of objects may be determined by external factors, such as the maximum transport channels of an audio codec profile or the number of signals that can be processed by a real-time renderer. Depending on the use case, this may result in stringent requirements for the reduction rate, such as a movie scene authored with up to 128 objects may be reduced to a channel bed plus 4-8 objects (e.g., to be transmitted as 7.1 + 4 channels + 4 objects, with a maximum of 16 transport channels per MPEG-H LC Level 3). In these use cases, the clustering algorithm may result in a given constant or maximum number of clusters. Maximum number clustering can be obtained directly from the maximum distance-based approach, for example, by increasing the allowable distance until the resulting number of clusters is below a limit. However, this can lead to ambiguity and possibly a lower-than-target number of output clusters, resulting in an unnecessary decrease in quality.

[0041] According to one embodiment, an iterative hierarchical clustering algorithm is presented in which the number of objects in a scene is reduced by iterative pairwise grouping using perceptual distance as optimization criterion. Furthermore, for very stringent reduction rates, it may be beneficial to recreate the entire acoustic scene in a "generative" way, for example by approximating the spatial distribution of loudness rather than individual sound sources.

[0042] In the following, clustering based on Gaussian mixture models (GMM) is considered. Clustering based on mixture models can be considered, for example, as a generative method. For example, a given DLM is approximated by a given number of components of a GMM. In other words, the method assumes a given / predetermined (maximum) number of available sound sources and aims to reproduce the overall loudness distribution of a given / predetermined scene, rather than considering individual source positions. Therefore, it can be considered a scene-based method (and should not be confused with Ambisonics, which is often called "scene-based audio"). This approach is particularly useful when a large number of objects need to be represented by only a few cluster positions (e.g., for low bitrate applications), for example, when many input objects are typically assigned to one cluster. Conversely, it is computationally inefficient to represent a large number of positions by a similarly large number of distributions.

[0043] Figure 4 shows a simplified one-dimensional example where a DLM generated by 10 sources is approximated by a GMM with only two components. Such GMM-based methods yield not only centroid locations and memberships, but also the probability that a point belongs to a given cluster. This can be advantageous for identifying cases where cluster membership is ambiguous (such as the sound source at approximately location 45 in the illustrated example). This information can be used to exploit temporal stabilization with hysteresis for fluctuations in membership assignments, and can even be used in the context of audio object clustering to enable soft clustering methods where an object can be mixed into two output clusters.

[0044] The expectation-maximization (EM) algorithm is a well-known technique for fitting a GMM to the distribution density of a given set of data points. The underlying model assumption can be, for example, that the input data points are located by a random process with a probability distribution density that is a mixture of Gaussian distributions in a given coordinate system. In other words, the GMM aims to approximate the probability that a data point is located at a given location. The EM algorithm is an iterative method for fitting such a probability distribution to a given set of data points. In principle, this method is similar to the well-known k-means clustering algorithm, which iteratively assigns points to their nearest centroid locations and then updates the centroid locations based on the updated cluster membership. Simply put, the EM algorithm is a "soft" version of that method: instead of a "hard" membership assignment of points to clusters, the parameters of the Gaussian distribution (centroid location and standard deviation) are updated based on the point's probability of belonging to each individual Gaussian component. Thus, at each update step, a point can affect the centroid locations of more than one component. Conversely, the results of the EM algorithm yield not only centroid locations and membership, but also the "spread" (standard deviation) of the individual components and, thereby, the probability that the point belongs to a given cluster.

[0045] The EM algorithm involves two named steps, expectation and maximization, which are repeated iteratively until a convergence criterion is reached. Roughly speaking (omitting the underlying statistics), the iterative steps are the expectation step and the maximization step. In the expectation step, given the distribution parameters, eg, centroid location and eg, standard deviation, the membership probabilities, eg, the probability that each point belongs to each of the individual Gaussian components, are calculated.

[0046] In the maximization step, given the membership probabilities, the distribution parameters, eg, centroid and distribution width from the mean and variance, are updated and calculated and weighted by the respective membership probabilities. As a termination criterion for the iterations, the log-likelihood of the distribution can be used, for example, as a measure of "goodness of fit." Also, the number of iterations may be limited, for example, typically to control the maximum computation time.

[0047] Existing DLM-based object clustering methods have limitations. Fitting a GMM to data points is, in principle, a common task for which algorithms and toolboxes are available (e.g., provided by Matlab toolboxes). However, a typical application is fitting a model to a random distribution of unweighted points with varying densities. Conversely, DLMs represent a regular grid of points with varying weights. This discrepancy precludes the direct use of available algorithms and toolboxes for GMM fitting. To make existing toolboxes usable, this discrepancy can be addressed through data preprocessing, for example, by emulating varying distribution densities by repeating points based on DLM values. However, this results in substantial data bloat due to the repeated points and is therefore inefficient in terms of memory requirements and computational complexity. Furthermore, the chosen sampling grid of the DLM can have a detrimental effect on the results of feeding the preprocessed data into existing GMM fitting algorithms: if the sampling grid, and therefore the relative point density, is not uniformly distributed, the centroids of the resulting GMM will be biased towards regions of higher sampling point density, e.g., concentrated at the poles due to uniform sampling in the azimuth / elevation region.

[0048] As an aside, note that the statistical analogy for using the EM algorithm on grid-based data is the analysis of histogram data rather than the underlying point distribution, as required for DLM fitting. Interestingly, however, there is little literature on using the EM algorithm on grid-based / histogram data. Because histograms are first generated from the underlying data, binning data into histograms reduces accuracy and appears to be done only for computational efficiency or data acquisition reasons (e.g., Chiang et al., "Where are passengers? Grid-based Gaussian Mixture Models for Taxi Bookings," 2015) and is not supported by available toolboxes. Also, histogram-based methods assume a uniformly sampled grid, which is not necessarily the case for DLMs sampled on a sphere. Furthermore, fitting a model to represent the probability of a random distribution results in a distribution where the sum (or integral) over all positions is always normalized to 1, i.e. equal to 1. However, in DLM, the overall sum is determined by the sum of the loudness of the individual sound sources, which is not normalized to a constant value. Therefore, an improved EM algorithm was developed that was modified to fit a GMM to a set of weighted points in an arbitrary grid of locations.

[0049] In the case of PCS-based DLM, distance is actually modeled to fit the Euclidean distance between two given points rather than the angular distance (e.g., to account for anterior / posterior confusion). Thus, the underlying distribution model is a 3D Gaussian distribution rather than a surface distribution (such as a spherical distribution).

[0050] Below, an improved EM algorithm according to one embodiment for weighted data points is described. As a specific embodiment, detailed exemplary operation of the developed algorithm is given in the following pseudo-code representation. The algorithm parameters may be, for example, one, more, or all of the following parameters: The input parameters include, for example: Pre-generated loudness map (sampled grid point positions p i and the corresponding loudness value DLM(p i ))、 The target number of clusters is k. Output parameters include, for example: Cluster centroid position c_l Membership probability for each input position to each component: clusterProb(i,l) "Hard" membership assignment of locations to clusters (to provide interface compatibility with other clustering methods that generate centroids and memberships) mem(i) Distribution parameters for determining the Gaussian components of the model DLM (scaling the weights to represent different loudness of different components), e.g., the center of gravity c_l, the spread parameter sigma_l (=standard deviation of the Gaussian distribution), and the weight parameter a_l The resulting GMM approximation of the DLM distribution is DLM_GMM(p i ) Error metric: Input DLM(p i ) and the approximated DLM_GMM(p i ) distribution and the sum of squared errors (SSE)

[0051] The following describes algorithm initialization (initial value setting) according to one embodiment. Note that, in general, the membership probabilities of individual points to clusters and the corresponding contribution weights are not available at initialization (as they are the result of probability estimation), so initialization is performed using "hard" membership and geometric distances. The weights and width distribution parameters of the Gaussian components are then determined and refined in subsequent iterative steps. There are several options for initializing the centroid positions c_j (c_1...c_k). For example, the initialization of the centroid positions can be performed as follows: For the first processed frame, the k loudest input objects can be selected, or initialization can be performed with random positions, or a (computationally fast) k-means clustering algorithm can be performed with the random initialization, and the result can be used as a better guess of the initial centroid positions, for example, to increase the convergence speed of the EM algorithm (e.g., coarse clustering with k-means, subsequent EM algorithm for improvement). For subsequent frames, initialization can be performed using previous centroid positions, for example, for improved temporal stability, or re-initialization can be performed with one of the above methods, for example, based on scene change detection. Optionally, multiple instances of the EM algorithm with different initialization methods can be run (e.g., previous positions and the current loudest object), and the result with the lower error metric can be selected.

[0052] The initialization of the membership mem(i) may be performed, for example, by assigning all points closest to the centroid based on the Euclidean distance d_i(j) = d(p_i,c_j) = |p_i-c_j|, or may be given previously, for example if the initialization is done by k-means. The initialization of the distribution width parameter sigma can be calculated, for example, as a first option, based on the standard deviation of the distribution of the initial centroids, i.e., sigma(j,dim) = std({c_1(dim), ... c_k(dim)}), which is the same for all components, or as a second option, based on the standard deviation of the positions of the initialized cluster members, sigma(j,dim) = std(p(mem==j)). Note that in the case of multidimensional data, the Gaussian distribution is assumed to be separable in each dimension, i.e., the distribution width controlled by the standard deviation parameter sigma(j,dim) is determined independently for each dimension dim cluster index j (i.e., the three degrees of freedom in the case of 3D-DLM can be reduced to 2D, e.g., for use cases with sound sources only in the horizontal plane).

[0053] Regularization of sigma may be performed, for example, by restricting it to a value between regmin, regmax (e.g., [1, 5]) to prevent an overly narrow or overly wide distribution that would hinder the convergence of the algorithm, e.g., for stability reasons. (For example, if a cluster has only one member at initialization, the distribution width is effectively 0, preventing other members from agglomerating into the cluster.) In addition to algorithm stability, this is also motivated by psychoacoustic considerations, since the distribution width representing the membership probability, or conversely, the "uncertainty," should not be narrower than the localization accuracy of the underlying perceptual model.

[0054] A weight a_j can be assigned to each cluster to represent, for example, the difference in distribution weights. To initialize the weights a_j, we can first compute the joint probability density function (PDF) across all dimensions for each data point, jointPdf(i), as the product of the individual PDFs, e.g., the PDF of the Gaussian-normal distribution, given by normpdf(x, mu, sigma), using the corresponding distribution parameters c_i, sigma initialized above. TIFF2025533617000002.tif1680 The cluster weights a_j can then be calculated, for example, from the ratio of the sum of the jointPdf weighted by the values ​​of the data points to the unweighted sum of the jointPdf, e.g. The file is TIFF2025533617000003.tif1451. The sum of all distributions at a data point location, sumPdf(i), can be calculated, for example, as the sum of the weighted distributions of all Gaussian components to obtain an approximation of the overall DLM(p_i). TIFF2025533617000004.tif1354

[0055] The iterative steps according to one embodiment are described below. In the expectation step, the probability that each data point belongs to a given cluster, clusterProb(i,j), can be calculated, for example, as the ratio of the individual cluster's contribution to the overall PDF. clusterProb(i,j)=jointPdf(i,j)*alpha(j) / sumPdf(i) In other words, this is similar to calculating the ratio between the DLM of an individual component and the overall DLM. In the maximization step, the centroid position c_l can be updated (for each dimension separately) as, for example, the weighted average position of all points weighted by their probability of belonging to a given cluster, TIFF2025533617000005.tif1346, To improve numerical stability (and avoid division by 0), a small offset can be added, for example, and the positions are further weighted by the data point values, e.g. TIFF2025533617000006.tif1570. Optionally, to represent data originally sampled on a sphere or ellipse, the centroid positions are projected onto the spherical surface, e.g., assuming a distribution on a unit sphere, by normalizing the position vector to 1.

[0056] Similarly, the distribution width sigma_l can be updated based on, for example, the mean weighted variance, e.g., TIFF2025533617000007.tif15109, and jointPdf, cluster weights a_j, and sumPdf may be updated as described above for initialization, for example. The expectation and maximization steps may continue, eg, iteratively, until, eg, a termination criterion is met. The termination criterion may be, for example, reaching a maximum number of iterations, for example 50. Such a termination criterion guarantees an upper bound on the overall computation time.

[0057] Alternatively, the termination criterion may be, for example, a criterion based on the sum of squared errors (SSE) between DLM and sumPdf (instead of the log-likelihood commonly used in EM algorithms for unweighted data). For example, the termination criterion may be, for example, that the overall SSE is sufficiently small, i.e., that the fitted model is sufficiently good. Alternatively, the termination criterion may be, for example, that the SSE is no longer decreasing (i.e., the SSE difference between two consecutive iterations is below a given threshold, e.g., 0.1*std(DLM)), e.g., that the algorithm has converged and more iterations will not bring further improvement.

[0058] With respect to output data collection, after completion, the algorithm may, for example, collect model parameters and generate additional output values, for example, distribution parameters (e.g., centroid location c_l, diffusion parameter sigma_l, weight parameter a_l), and, for example, membership probabilities for each location to each component; further, "hard" membership assignments mem(i) may be determined based on the highest membership probability for each point, for example, to provide interface compatibility with other clustering methods that also generate centroids and memberships.

[0059] Improved versions of the EM algorithm result in "soft" clustering by providing the centroid location and "hard" cluster membership for a given input location, as well as membership probabilities, which are common output parameters of clustering algorithms. Furthermore, they provide parameters for a weighted GMM model that approximates the input distribution (DLM). Key improvements over state-of-the-art EM algorithms are the incorporation of weighted input points with variable overall weights, consideration of inputs at uniform or non-uniform grid locations, and the fitting of locations on a sphere.

[0060] In the following, hierarchical clustering is considered. Generative clustering methods, such as those based on GMM, can be very efficient for fitting a small number of clusters to a large number of input objects. However, generative methods do not scale well for larger numbers of clusters (and therefore target quality) because the computational complexity increases with the target number of clusters. On the one hand, the number of calculations for mutual probability estimation increases. On the other hand, the increased degrees of freedom may require more iterations to converge to a stable solution. For example, if the target number of clusters is already close to the original number of input objects, a large number of iterations may be required to converge to a solution in which most objects ultimately remain unchanged.

[0061] According to one embodiment, an iterative hierarchical clustering algorithm is introduced. Briefly, iteratively selects two "closest" objects (preferably based on psychoacoustic metrics) and merges them until a target number of clusters is reached and / or a minimum distance threshold between the closest objects is exceeded. Thus, in each iteration, the number of output objects is reduced by one, resulting in N objects reduced to k clusters within (Nk) iterations, thus resulting in deterministic computational complexity. The general concept of hierarchical clustering is well known in the literature, and the developed algorithm according to an embodiment may apply known concepts, for example, in the context of object-based audio clustering, but according to one embodiment includes concepts and refinements that may, for example, use psychoacoustic metric(s) as cost function.

[0062] The distance metric for hierarchical clustering may be given, for example, by the connectivity within the cluster, e.g., the distance considered as a cost function of the membership within the cluster. Common connectivity models are, for example, "full connectivity", which is the maximum distance between any two objects within a cluster, or "centroid connectivity", which is, for example, given by the distance between their respective centroids. In the presented algorithm according to one embodiment, for example, a greedy iterative approach can be chosen where the pairwise distance is minimized and then the centroids are updated, which corresponds to the centroid linkage model.

[0063] The following describes a hierarchical clustering algorithm according to one embodiment. The input parameters and preprocessing include, for example: Input object position p i , input object energy (optionally perceptually weighted, e.g., by pre-filtering in the time domain to apply A-weighting frequency weighting); previous centroid location and object membership in subsequent frames, For example, a maximum number of clusters k, or a target condition such as, for example, an upper limit on the distance metric threshold (either or both may be specified). The output parameters include, for example: Cluster centroid position c_l, Cluster membership mem(i).

[0064] The following describes algorithm initialization according to one embodiment. For example, a masking model between input objects can be calculated. A cost function / distance metric, e.g., an inter-object distance matrix, can be calculated. For example, a baseline model can be determined, e.g., based on the Euclidean distance between object positions in a world coordinate system. Or, for example, a perceptually improved model can be determined, e.g., based on the Euclidean distance between object positions in a PCS. Or, for example, a complete model can be determined, e.g., taking into account the masking effect from the entire scene, and for example, a pairwise perceptual distance D_perc can be calculated.

[0065] The following describes iteration according to one embodiment. Note that the iteration may be performed "in place," e.g., two objects may be merged at the index position of one of the objects, and the other may be marked as invalid. This results in an updated centroid, which may be considered by the next iteration step, e.g., like any other object. In other words, during the iteration, each object may be considered, e.g., a centroid, and vice versa, and therefore these terms are used synonymously herein. The iteration may include, e.g., the following: For example, the minimum distance in the distance matrix can be selected.

[0066] Two corresponding objects may be merged, for example. Objects may be merged to one index of the two objects based on one or more of the following criteria: for example, to a smaller object index position (regression), for example, to an object / cluster with more energy, for example, to a cluster that already has more members. The centroid position may be updated, for example, as the average position of the two merged objects weighted by object energy, or alternatively, as the geometric midpoint position, or based on a weighted average of all member positions.

[0067] The parameters and distance metrics may be updated, for example. Note that the updated centroid is treated like any object in the next iteration. All row and column elements in the distance matrix of the "removed" object may be invalidated, for example, marked to be excluded from further search iterations. The energy of the combined object may be calculated, for example, as the sum of the merged object energies. The masking threshold at the new centroid position may be updated, for example, in a high-complexity model by recalculating the masking at the updated position, or in a low-complexity model by estimating the masking threshold at the centroid position as the maximum, sum, or weighted average of the thresholds of the merged objects. For example, the PE (perceptual entropy) of the combined object may be calculated from the updated energy and masking threshold. The rows and columns of the distance matrix for updating the distance to the combined object calculated in the initialization step of the input object may be recalculated, for example.

[0068] The iterations may continue, for example, until a termination condition is met. The termination criterion may be, for example, whether a target number of clusters has been reached, or, alternatively, the termination criterion may be, for example, whether the minimum distance is above a given threshold, for example, 1 JND. How the termination criteria are combined may depend on the target use case, for example, to achieve various goals, e.g., a constant quality, a constant number of output clusters, or as a compromise, a maximum number of clusters with nearly constant quality (assumed to be rarely the case).

[0069] Thus, the termination criteria can be combined in various AND / OR conditions to achieve one of the following options: The first basic case is the "constant rate" case: the iterations can continue, for example, until a target number of clusters is reached. This will always result in k clusters (unless the input number of objects is already N<=k), but with varying quality depending on the number and distribution of input objects. The second basic case is the "constant quality" case. The iterations can continue, for example, until the smallest distance in the distance matrix exceeds a given threshold. This results in (nearly) constant quality and can be used, for example, to remove only differences that are already below or close to the JND, or below a tolerance appropriate for a given use case. However, the number of output clusters can vary and, in the worst case, can be equal to the number of input objects.

[0070] The first combined AND case is the "remove irrelevant at a constant maximum rate" case (low target number of clusters, low distance threshold). The iterations can continue indefinitely, for example, until the target number of clusters is reached. If the minimum distance is below a given threshold (e.g., some JND), the iterations continue to remove irrelevant from the scene. The second combined AND case is the "constant quality with rate upper limit" case (high target number of clusters, high distance threshold). However, in terms of the same (Boolean) definition of the termination criteria as the first combined AND case, the main parameters are primarily the distance threshold to achieve constant quality, and the target number of clusters is set relatively high to provide an upper limit on the number of output clusters (e.g., to avoid exceeding the transport channel or renderer input capabilities).

[0071] The combined OR case is the "constant rate with quality disruption limit" case. This case is mentioned mainly for completeness, as its possible use cases are limited. The iterations can continue, for example, until any one of the termination criteria is met, i.e., the number of clusters or the distance metric indicates termination. This results in a variable rate with variable quality output. A possible use case is one in which the number of clusters (i.e., the rate) is intended to be nearly constant, but excessively large disruptions in quality should be avoided, and therefore more output clusters are temporarily allowed (e.g., in the case of file-based storage, where the average rate is more important than the peak rate). In the following, clustering based on JND (Just Noticeable Difference) is considered.

[0072] In contrast to "constant rate" clustering methods that use a given maximum number of clusters, JND-based clustering methods only aim to remove irrelevant and redundancy from a scene in order to reduce computational complexity and / or transmission bitrate (similar to the VBR mode of perceptual audio coders) while maintaining a perceptually imperceptible result or at least a constant quality. This can be achieved, for example, by only clustering objects together if the position change does not exceed a given threshold, eg, some JND.

[0073] Since the localization accuracy of human hearing can be revealed, this technique can be used to remove meaningless separation between objects that are already close to each other, and therefore can even be performed based solely on location metadata, without the need for actual signal measurements.

[0074] JND-based clustering can be performed, for example, with varying levels of stringency. In level 1 centroid distance, the distance between the cluster centroid and the clustered object must not exceed a threshold. In level 2 object distance, the distance between all pairs of objects in the cluster operation must not exceed a threshold. At level 3 total distance, the combined variation of all objects in the auditory scene must not exceed a threshold (e.g., to achieve a perceptually imperceptible quality). Note that levels 1 and 2 roughly correspond to "centroid-linkage" and "full-linkage" in hierarchical clustering methods, while level 3 corresponds to global scene analysis tasks (e.g., measuring distance sums or global DLM divergence).

[0075] FIG. 5 illustrates three different distance model levels (labeled L1-L3) for JND-based clustering according to an embodiment. In the given example, for level 1, all objects that can be within JND distance of the resulting centroid can be combined, for example. At level 2, objects may have to be closer, for example, within JND distance of each other, to be combined. At level 3, even if all objects are within JND distance, for example, only two of three objects can be combined, because otherwise the sum of their distances would exceed the JND.

[0076] Level 1 (centroid distance) may be implemented as a variant of the above-mentioned hierarchical clustering algorithm, for example, by not setting a target number of clusters in the termination criteria and considering only the smallest component of the distance matrix, min(D_perc), as being below a given threshold, or alternatively, by considering the perceptual spatial distance D_PCS, independent of masking and energy characteristics, e.g., below 1 JND. The latter allows clustering in applications where only location metadata is known to the algorithm, but not signal energy.

[0077] Level 3 (sum distance) may be implemented, for example, by a hierarchical clustering algorithm, where the sum of the distances may be used as the termination criterion instead of, for example, the minimum distance, or the divergence of the DLM over the entire scene is used as the termination criterion. Note, however, that the iterative calculation of the DLM divergence results in high computational complexity and is therefore more suitable for encoding and conversion tasks than for real-time applications.

[0078] Level 2 (object distance) offers a favorable compromise between the strictness of Levels 1 and 3. Because it relies only on initial object positions, it can be implemented with low computational complexity and is therefore the recommended mode of operation for most applications. Because only pairwise distance metrics between objects are considered, it can be performed, for example, based on only one initial calculation of the distance matrix, without iteratively updating centroid positions and distances. To improve the computational complexity of an object clustering system, such object distance-based JND clustering can be performed as a preprocessing step to reduce the number of initial clusters with low computational effort while maintaining its imperceptible quality, for example, before applying an iterative (hierarchical or GMM-based) clustering algorithm to achieve a target number of clusters. Note that, because different groupings are possible (e.g., A+B and B+C may be combined, but A+C should not), there is generally no unique solution to such clustering. Optimizing such a "fully connected" clustering problem to minimize the number of clusters is known in the literature as the "exact covering problem," which has been shown to be NP-complete. However, for object clustering applications, the distance metric offers an alternative optimization criterion, on the basis of which a greedy algorithm with low computational complexity can be obtained. An algorithm according to one embodiment can be implemented, for example, as follows:

[0079] For example, an initial distance matrix can be calculated. Depending on the use case, this can be based on D_PCS, for example, to consider only spatial relationships, or on D_perc, for example, to additionally consider masking properties. The advantage of using D_PCS is that the JND clustering step is independent of signal energy, i.e., it can be performed with very low computational complexity. The advantage of using D_perc is that perceptual properties are modeled more accurately. Furthermore, since silent or inaudible objects are assigned a PE of 0 (or close to 0), this implicitly serves as a culling stage to consolidate irrelevant objects.

[0080] All elements in the distance matrix (outside the main diagonal) below a selected threshold may be marked as pairs that may be combined, for example, in a Boolean combination matrix. The threshold can be selected depending on the use case. For clustering based on D_PCS distance, a threshold of 1 [JND] can be selected to integrate only objects that are within the localization accuracy of human hearing. For clustering based on D_perc, masking characteristics are additionally incorporated into the distance metric via PE. Assuming the signal is exactly at the masking threshold, the resulting PE is log2(1 + 1 / 1) = 1 [bit]. Therefore, a D_perc threshold of 1 [bit * JND] can be selected as a simple approximation. All elements for which the connection matrix is ​​true can be considered, for example, as candidate pairs. Cluster creation can begin, for example, by initializing a cluster of two objects by selecting from the candidate pairs the one with the smallest entry in the distance matrix.

[0081] Iteratively, objects can be aggregated into clusters, for example, by: Select the corresponding true entries in the connectivity matrix to create a candidate object list (candidate list) of objects that can be added to the cluster, e.g., objects that can be combined with all objects already in the cluster (but not necessarily all can be combined with each other). Select the candidate object with the smallest absolute distance or the smallest sum of distances to all objects in the cluster.

[0082] The selected object is added to the current list and the candidate list is updated based on the connection matrix of the new object, e.g., objects that may not be connected to the just-added object are removed from the candidate list. Repeat until there are no more items left in the candidate list. After the iterations are finished, the connectivity matrix of all objects in the just-created cluster may be set to, for example, false so that they are no longer assigned to another cluster. The search may be repeated for additional clusters, for example, starting from the initial cluster creation, until no true components remain in the connection matrix.

[0083] 6a-6g show a small scale example of a Level 2 JND based clustering algorithm according to one embodiment. Figure 6a shows the initial distance matrix calculated based on D_PCS. Figure 6b shows a distance matrix, where all components outside the main diagonal in the distance matrix that are below a selected threshold may be marked as pairs that may be connected, for example, in a Boolean connection matrix. In Figure 6b, the threshold selected for marking components is 1 or less. Figure 6c shows the connection matrix. FIG. 6d shows the selection of the candidate pair with the smallest distance matrix components to initialize a cluster of two objects. Figure 6e illustrates finding candidates in the join matrix that can join with both objects in the cluster and adding them to the cluster (adding the first object in the illustrated example) until the list of candidate objects is empty. Figure 6e shows that for objects already assigned to a cluster, each row / column is analyzed to determine which other candidate objects can join with the objects in the cluster. For example, object 2 can join with (1,3,5). Object 3 can join with (1,2). Therefore, (1,3,5) AND (1,2) = (1). Therefore, add object 1 from the candidate list to the cluster, the candidate list is empty, and proceed to the next cluster. Figure 6f shows the connection matrix where the element in row / column (1,2,3) is invalidated when the cluster is completed. Figure 6g shows the connectivity matrix from which the next cluster is selected. If the candidate list is empty, the algorithm is complete.

[0084] In the following, improvements according to specific embodiments are discussed. First, temporal stabilization according to one embodiment will be described. The presented clustering algorithm may be performed, for example, on a frame-by-frame basis. In addition to the perceptual distance in each frame, the temporal stability of the scene in successive frames is also important for perceptual quality. For example, if an originally stationary object position becomes unstable and starts to move around, or if an audible "jump" is introduced instead of an originally smooth transition, this also affects perceptual quality.

[0085] This leads to a trade-off in terms of optimization objectives between minimizing the instantaneous distance metric and temporal stability. For example, consider a sound source with an originally fixed position that is located, say, near the "border" between two clusters. Without temporal stabilization, slight changes across the scene can cause the object's membership assignment to switch between different clusters, thus frequently jumping between centroid positions. Such destabilization may be perceived as more annoying than a larger but stable object's shifting position.

[0086] For example, for offline ("file-to-file") applications, such as encoding or converting pre-created scenes (e.g., object-based audio mixes for movies), some look-ahead or even multi-pass encoding techniques can be taken to optimize temporal stability. However, for real-time functionality (e.g., interactive virtual reality (VR) applications), temporal stabilization may need to operate with little or no look-ahead, for example, to avoid introducing additional delays into the system.

[0087] The temporal stabilization concept according to some embodiments presented below does not require look-ahead, as it relies on applying smoothing or hysteresis to past frames. First, consider the concept of utilizing temporal penalties in hierarchical clustering according to one embodiment. To avoid switching object membership assignments for objects where the optimal assignment is ambiguous, in one embodiment, an additional penalty is introduced for objects to change cluster membership. Thus, for example, a temporal penalty can be applied to the perceptual distance D_perc between objects that previously belonged to different clusters.

[0088] There are several options for implementing the time penalty. For example, a fixed offset may be added to D_perc (for example, 30 [JND*bit]). Alternatively, for example, a multiplication factor may be applied to D_perc (eg, 2). Alternatively, for example, the (lateral) distance of an object relative to a previous centroid of another cluster can be used, e.g., taking into account the actual resulting centroid position and not just the distance between the objects (e.g., to take into account that two objects that may be close to each other may be on exactly opposite sides of the boundary between two clusters). Or, for example, one could use the (weighted) distance between previous cluster centroids (e.g., assuming the worst case that reassigning object membership will move the object's position from one centroid to the other, where this has a small impact on the object's centroid position).

[0089] Next, DLM smoothing and centroid initialization in GMM-based clustering according to one embodiment will be described. GMM-based clustering methods can take into account the sluggishness of spatial hearing, for example, by temporally smoothing the DLM. Thus, the smoothed DLM is calculated as a weighted average of the DLM of the current frame and the previous DLM (either using the previous frame's DLM for short FIR-type smoothing, or using the previous smoothed DLM for IIR-type smoothing with a longer decay).

[0090] In addition to smoothing the DLM, the EM algorithm for GMM fitting may be initialized with the centroid position of the previous frame, for example. To prevent temporal smearing, for example, for scene changes (e.g., movie cuts), a threshold on the overall difference (e.g., SAD, sum of absolute differences) of the DLM between two subsequent frames can be set to trigger re-initialization of the centroid position.

[0091] Next, cluster replacement optimization according to an embodiment will be described. In addition to source locations, the temporal stability of the combined output signal is also important, especially when the signal is transmitted by a perceptual audio codec. Even if the cluster centroid locations and object assignments remain mostly stable within a scene, slight changes in cluster membership can result in permutations of the cluster index order (because the cluster index order depends on the lowest member object index in hierarchical clustering or can be the result of random position initialization in GMM-based clustering methods).

[0092] Such a permutation is shown in Figure 7, where only the central object moves slightly and is reassigned to a cluster from left to right, but the cluster indexes are swapped. In particular, Figure 7 shows the permutation of cluster indices according to one embodiment due to a slight change in the scene (the circles pointed to by the arrows in Figure 7 are the cluster centroid locations, and the outer circle from which the arrows originate in Figure 7 are the input objects).

[0093] Typically, object signals are mixed into a continuous waveform, resulting in one signal (e.g., a transport channel) for each cluster. If object signals are assigned to different output signals in subsequent frames due to substitution, discontinuities may be introduced into the output signal. Repeated crossfades between signals may be necessary, which may introduce transients into the otherwise continuous signal (which are not actually perceived as transients in the overall audio scene). These "false" transients may hinder the performance of perceptual audio codecs and should therefore be prevented. In addition to affecting the output signal, the substitution / swapping of cluster indices may also result in unnecessarily large and frequent changes in the corresponding centroid positions, which may cause artifacts in the renderer (e.g., when positions are interpolated between frames) and, for example, reduce the efficiency of time-shifted coding of cluster positions. Therefore, measures may be taken to stabilize the assignment of cluster indices against the effects of substitution in successive frames.

[0094] Because the assignment of multiple objects to clusters and centroid locations can, for example, change over time, especially when larger scene changes occur, the replacement assignment can be ambiguous and requires an appropriate optimization strategy. However, the optimization goal of the replacement strategy is use case dependent.

[0095] According to one embodiment, for example, a baseline approach may be used to count and minimize the number of objects that are reallocated between clusters. Alternatively, to stabilize the location metadata, according to another embodiment, the sum of absolute or squared distances between the previous and current cluster centroids may be minimized, for example. However, one explicit goal is also to stabilize the resulting output signal waveform. Thus, according to one embodiment, for example, signal characteristics can also be taken into consideration. As an illustrative example, for example, consider a scene with two very loud objects and several more nearly silent objects. Here, for example, it may be preferable to keep the allocation of loud objects stable (rather than minimizing the number of object reassignments). Simply put, the optimization goal in this case is to maintain the magnitude of the signal energy allocated to the previous location.

[0096] According to one embodiment, a permutation optimization is performed with the goal of stabilizing the energy distribution from objects to clusters. First, the algorithm calculates a matrix of how much total energy of a given object is reallocated between individual clusters for cluster assignments in two consecutive frames. Based on this energy permutation matrix, a greedy algorithm is used to minimize the amount of energy reallocated between clusters.

[0097] Figure 8 illustrates cluster assignment permutation and optimization according to one embodiment. In particular, Figure 8 illustrates an example of cluster permutation optimization according to one embodiment, assuming 10 objects are assigned to three clusters. The direction of the arrows indicates the assignment of objects to clusters (e.g., cluster indexes). The cluster membership of an object in the previous frame, corresponding to its previous cluster assignment, is shown in Figure 8a. The weight of the arrow indicates the expected energy of the object in the current frame (the energy is also shown numerically in the square on the left).

[0098] Figure 8b) shows, for example, the cluster assignments for the current frame as they result from a clustering algorithm, with the cluster index order determined by the lowest member object index. Note that, as in the previous frame, the three loudest objects are still assigned separately to three separate clusters. However, because the grouping of the objects has changed, the assigned order has changed, resulting in a reassignment of the output signal.

[0099] Thus, according to one embodiment, permutation optimization is performed based on the energy permutation matrix shown in Figure 8c. The highlighted cells indicate the optimized permutation assignment (e.g., row 1, column 2 indicates that most of the energy previously in cluster 1 is now in cluster 2). The resulting permutation-optimized cluster assignment is shown in Figure 8 d). Thus, in this (purposely chosen) illustrative example, the assignment of the three loudest objects remains stable with respect to the previous frame.

[0100] In particular, an algorithm according to one embodiment may be implemented, for example, as follows. Assuming that a fixed number k of clusters is obtained from the clustering algorithm, a squared energy permutation matrix M_Eperm of size k×k can be initialized, for example, with the value 0. M_Eperm=zeros(k,k) For each object index i, the current energy E(i) can be added to the matrix elements corresponding to, for example, the row of the current cluster membership index and the column of the previous cluster membership index, mem_new(i), mem_prev(i). M_Eperm(mem_new(i), mem_prev(i))+=E(i) This can result in, for example, a matrix representing how much energy is reallocated to different indices. If no reallocation occurs, this reduces to a diagonal matrix. If the object grouping remains the same but a permutation of the cluster index order occurs, this results in a sparse matrix with only k non-zero entries. However, in the general case where objects from different groups are combined, this is not a sparse matrix (especially if many objects are combined into a small number of clusters, i.e., N>>k).

[0101] The permutations may be optimized, for example, by a greedy search in a permutation matrix, which may include, for example: Initialize a permutation vector of length k with the value 0. Find the maximum element in the matrix and get the indices rowMax, colMax. Define a permutation vector for each position permutation(colMax)=rowMax Set the component, row rowMax and column colMax to 0 (to indicate that the corresponding input index has already been assigned and the output index has already been obtained). M_Eperm(rowMax,:)=0 M_Eperm(:,colMax)=0 Repeat until all k permutations have been assigned.

[0102] The permutation can be done, for example, by directly reassigning the centroid indices: c_perm(j)=c(permutation(j)) and, for example, by selecting and replacing the corresponding membership indexes. If mem(i)=permutation(j), then mem_perm(i)=j It can be applied to the assignment of centroids and membership indices. In applications where the energy of the objects is not known to the algorithm, the algorithm may be used to minimize the number of objects that are reassigned, for example, by assuming that all object energies are equal to 1. This effectively uses the energy permutation matrix M_Eperm to count the objects.

[0103] Cluster centroid position optimization according to the embodiment will be described below. Clustering algorithms yield the membership (or membership probabilities) of individual objects, as well as cluster centroids. Clustering of 3D object positions can yield clusters that include objects in front and behind, especially in the case of clustering based on perceptual metrics that exploit the limited spatial resolution of human hearing for ascending along the cones of confusion and front-to-back confusion. Assuming that the centroids are calculated as a weighted average of the original positions on a convex hull around the listener, e.g., a unit sphere or PCS ellipsoid, the resulting average position may lie inside the sphere / ellipsoid. However, in most applications, it is desirable for the output cluster positions to also lie on the sphere. This is especially essential in loudspeaker playback scenarios, where the sphere corresponds to the convex hull of the loudspeakers; otherwise, this would require internal panning, which is not supported by many renderers (e.g., the VBAP implementation in MPEG-H). Therefore, the resulting cluster positions need to be shifted from the internal centroid positions onto the spherical surface.

[0104] The technique is to project positions onto the unit sphere by normalizing their coordinate vectors to length 1 (they are warped from / to PCS coordinates before and after normalization), as shown in Figure 9. In particular, Figure 9a) shows the centroid projection onto the unit sphere in the horizontal plane ("top view"). Figure 9b) shows the centroid projection onto the perceptual coordinate system (PCS) in the horizontal plane. However, this results in perceptually incorrect output locations because locations that were initially on the same cone of confusion (CoC) are projected outward. Thus, combining source locations that are perceptually different only in spectral cues alters the left / right characteristic and, therefore, binaural cues.

[0105] Thus, according to one embodiment, for example, a perceptually optimized arrangement of cluster output positions can be utilized, where the left-right coordinate of the centroid position is preserved and the cluster positions are optimized along the corresponding cone of confusion. The optimization along the CoC may also depend on the intended playback scenario, e.g., a different strategy may be chosen for e.g., binaural rendering than for loudspeaker rendering. Therefore, in the following, several options for center of gravity placement are presented.

[0106] The following describes normalization of the lateral centroid position according to one embodiment. The baseline projection method can project positions outward by normalizing the position vectors within a side surface to match the radius of the corresponding circle along the unit sphere, as shown in Figure 10.

[0107] 10 shows a projection of the center of gravity at the side ("side view") onto the cone of confusion, according to one embodiment. Note how the foreground and background objects are projected upwards as a result. The radius of the circle representing the CoC within the side is calculated, and the centroid coordinate vector is normalized within the side to match the radius of the CoC while maintaining the original left-right coordinates. If PCS coordinates are used, the centroid location is first converted back to unit coordinates. (This mode can be advantageous in playback scenarios with sparse immersive speaker setups, where intermediate positions are recreated by amplitude panning, for example. In this case, the energy of the object is redistributed between front and back by exploiting the properties of the target rendering.) If the coordinate axes are positioned as c = "front / back" (+1 = front), y = "left / right" (+1 = left), and z = "up / down" (+1 = up), this is calculated as follows:

[0108] Azimuth=sin -1 (y_center of gravity) radius_coc=cos(azimuth) Radius_center of gravity=sqrt(x_center of gravity 2 +z_center of gravity 2 ) x_projection=x_center of gravity*radius_coc / radius_center of gravity z_projection=z_center of gravity*radius_coc / radius_center of gravity In the following, a height preservation mode according to one embodiment is presented.

[0109] Psychoacoustic experiments have shown that "height" spectral cues for vertical localization are distinct from "front / back" spectral cues. In other words, perceptually "up" is not halfway between "front" and "back." As a result, baseline normalization of center of gravity locations within the CoC's lateral plane is not an ideal placement of cluster locations for many applications, e.g., binaural rendering (HRTFs with "height" spectral cues can be used to reproduce front and back objects at ear height). Therefore, a projection mode that preserves height cues is introduced. In order to preserve perceptual cues for height perception and front-back confusion resolution, both magnitudes can be considered separately, for example.

[0110] FIG. 11 shows a height-preserving centroid projection onto the CoC in a side view (= “side view”) according to one embodiment. As shown in Figure 11, the height component may be preserved, for example, from the centroid position, and the position may be projected onto a cone of confusion, for example, parallel to the horizontal plane. However, this means that the decision whether to project forward or backward is difficult. If the centroid is close to the transition between forward and backward (e.g., y_centroid is close to 0), the projected position may jump between forward and backward, for example, if the energy of the forward and backward objects changes slightly over time. To stabilize the resulting position, for example, hysteresis can be used in the sign of the front / back coordinates to prevent the cluster position from switching.

[0111] Note that this mode is particularly well suited for binaural rendering applications: it prioritizes preserving height cues over resolving front / back confusion. In loudspeaker rendering applications, front / back confusion can be easily resolved by binaural cues introduced by slight head movements, whereas in binaural rendering, for example, only spectral cues may be available to resolve front / back confusion.

[0112] The following describes a spectral matching ("EQ matching") mode according to one embodiment. The basic idea behind the fact-based spectral matching mode is that position along the CoC corresponds to variations in spectral cues. Therefore, the perception of a position change depends on the frequency range affected and the actual amount of spectral content the signal has in each frequency range. This means that a position change will be easier to perceive for objects that have more energy in the affected frequency range than other objects, and vice versa. Therefore, the spectral matching approach according to one embodiment optimizes the position to minimize the spectral difference of the sum of the signals at the ears. Another interpretation is to consider the variation of object position between CoCs as a multi-equalizer (EQ) curve, where the task is to match the overall spectral envelope, and therefore this mode is also called "equalizer (EQ) matching".

[0113] The EQ matching mode may require higher computational complexity than, for example, the centroid projection mode, since it considers the positions and signal characteristics of all member objects of a cluster, rather than just the centroid positions. For the configuration and calibration of this mode, for example, appropriate frequency bands can be selected, and an average elevation gain curve for each band can be calculated, for example, based on an analysis of an HRTF (Head Related Transfer Function) database (e.g., similar to the calibration of a PCS). During operation, signal energies can be calculated, for example, for each band and object, and an optimized position is selected by numerically minimizing the difference of the weighted energy sums, or by minimizing a ratio, for example, the sum of the logarithmic differences.

[0114] To improve computational complexity, for example, first-order component analysis can be used to derive a limited number of "eigenspectras" for positions along the CoC. This can be interpreted as a predefined equalizer curve for the entire spectrum, with its intensity adjusted based on position, rather than determining individual components for each position and frequency band. These can then be correlated with the spectral envelopes of the individual signals to generate lower-dimensional representations that can be minimized with lower computational complexity.

[0115] The following describes mixing and processing of output signals according to some embodiments. After cluster membership and centroid locations are determined, the object signals are combined to generate one output signal for each output cluster. One approach, for example, may be to sum the signals of all members in a cluster. However, further precautions and improvements must be considered to avoid audible artifacts and optimize perceptual quality. Because cluster assignments are determined on a frame-by-frame basis, membership may change from one frame to the next. Crossfades may be applied when membership changes, for example, to prevent audible clicks due to signal discontinuities.

[0116] There may be correlation between the signals of the objects within a cluster, which may, for example, result in positive or negative interference in the downmix signal. To achieve an energy-preserving downmix, for example, the signal correlation may be taken into account. Clustering algorithms such as GMM-based clustering yield not only memberships but also membership probabilities: objects with ambiguous membership can be mixed into two or more clusters, for example, to achieve a "soft" clustering approach.

[0117] A crossfade according to one embodiment will now be described. According to one embodiment, when object membership changes between subsequent frames, the downmix signal may be cross-faded to prevent, for example, hard signal cuts that may cause audible clicks due to signal discontinuities. To avoid the need for additional look-ahead for cluster allocation in the next frame, the cross-fade may be performed, for example, at the start of the current frame. To avoid unnecessary crossfades, the cluster membership of each object can be stored and compared, for example, with the previous membership and current membership, and a crossfade applied if and only if the membership has changed.

[0118] For the crossfade, a complementary window function may be applied, for example, to fade in the object signals in the newly assigned cluster signal and fade out from the previously assigned output signal. The crossfade may, for example, be selected to be energy-conserving, so for example, a sinusoidal shaped window may be used. In one embodiment, the crossfade duration may, for example, be long enough to prevent audible clicks, but may also be as short as possible to prevent audible delays at the sound source positions. Thus, in certain embodiments, for example, a crossfade length of 128 samples (2.7 ms at a sampling rate of approximately 48 kHz) may be used.

[0119] In the following, correlation-aware downmixing according to some embodiments is described. The basic assumption for object-based audio clustering is that audio objects represent individual, uncorrelated sound sources, which are typically rendered as individual point sources by object-based audio renderers (e.g., VBAP, vector-based amplitude panning). However, this assumption may be violated, for example, when two or more object signals are correlated. This may result in positive or negative interference when calculating a downmix signal for correlated object signals within a cluster. Therefore, additional precautions may be taken, for example, when calculating a downmix in a scene expected to contain correlated objects. Note that strong correlation between sound sources may also result in the perception of phantom sound sources. However, this also affects the placement of the resulting cluster positions and is therefore not discussed within the scope of signal downmixing.

[0120] Generally, there will be a randomly small amount of correlation between originally independently created / recorded audio signals (unless the signals are obviously generated to be orthogonal, e.g., as independent random noise), but this is usually not significant. However, a greater correlation between the signals may be introduced depending on, for example, the production methodology used to create the object-based sound scene. For example, in some cases, objects are created from signals originating from two or more channels of a stereo or multi-microphone recording within a sound scene. Another way of looking at this is that an object-based audio scene can include an "unmarked channel bed," e.g., a recording or production originally created for loudspeaker playback that is reused and placed at object positions roughly corresponding to the intended loudspeaker positions. This is typically known at the time of production, but depending on the metadata transmission format, may not be known to the clustering algorithm. Similarly, to a lesser extent, correlation can occur when objects are selected from multiple spot microphones within a single physical scene, for example, for different performers or instruments on stage. This is not typically considered a channel-based recording, but crosstalk between individual microphone signals can still occur.

[0121] Furthermore, signal correlation can occur even for separately recorded or synthesized signals due to content relationships, for example when multiple instruments follow the same melody line. In some of these different cases, correlations between signals can be predicted during manufacturing and marked with appropriate metadata. However, when correlations are introduced more casually, no additional metadata is available. Therefore, object clustering algorithms cannot rely solely on external information and must be able to properly detect and process correlations when downmixing object signals, even without available metadata.

[0122] If there is correlation between object signals combined within one cluster signal, for example, the signal amplitudes are summed rather than the signal energies, which may lead to boosts or decreases in signal energy and thus differences in perceived loudness. According to some embodiments, for example, correlation-aware downmixing can be applied to preserve the perception of loudness of the original scene. However, it must be recognized that the perceived effect of correlation between object signals also depends on the playback scenario and renderer algorithm used.

[0123] According to one embodiment, for example, energy summation can be performed. In an idealized playback-independent scenario, objects represent physical sound sources at different spatial positions. Here, actual sound waves are physically superimposed in the playback environment and the ear. Because a typical listening environment is not anechoic (e.g., a BS1116 room), the correlation between signals reaching the ear is reduced, especially at higher frequencies, due to different propagation paths (i.e., room reverberation and HRTFs). As a simplified model, energy summation may be assumed for this case, for example. In an applied playback scenario, this may be the case, for example, of binaural headphone playback, where different BRIRs (binaural room impulse responses) may be applied to, for example, different sound source positions. In the case of loudspeaker playback, this may be assumed, for example, when the distance between objects is large enough relative to the loudspeaker placement, resulting in the objects being played by separate loudspeakers.

[0124] In one embodiment, amplitude summation can be performed. In the case of amplitude panning-based rendering (e.g., VBAP) with a relatively sparse speaker arrangement (e.g., a typical home cinema arrangement), for example, different sound source positions can be panned and reproduced between the same pair of speakers. In this case, signal amplitudes can be summed, for example, in the rendering algorithm, resulting in correlation-dependent behavior of the energy sum. Renderer-independent object clustering algorithms assume an idealized situation of independent sound sources and therefore energy summation. However, the goal of an object clustering algorithm is often to get as close as possible to a reference rendering for a given rendering in a given target reproduction scenario. This means that the reference behavior is intentional or not, and the goal is to reproduce the energy or amplitude summation characteristics of the target rendering and reproduction.

[0125] Two downmix modes can be chosen based on the target use case. According to the first downmix mode, for example, direct signal addition may be performed. If the object signals are assumed to be uncorrelated and / or the target playback scenario is loudspeaker playback with amplitude panning, the object signals are simply added to the cluster output signals. This mode also avoids the additional computational complexity of correlation analysis and is therefore preferred for real-time applications. According to a second downmix mode, for example, correlation-aware signal summation may be performed: the aim is to sum in an energy-preserving manner, and if correlation between signals is expected, energy-preserving weighting is applied. To achieve scene-wide energy conservation, one approach is to calculate the energy of all objects before mixing, calculate the resulting energy of the downmix signal, and apply a gain correction factor to the downmix signal. However, a pitfall of such a simple approach is that not all objects in a cluster are necessarily correlated in the same way. Therefore, such a global energy gain correction also reduces the energy of uncorrelated signals, thus still resulting in an over-representation of correlated signals in the final mix.

[0126] Therefore, according to one embodiment, an advanced downmix algorithm based on, for example, signal correlation can be used, for which, for example, a cross-correlation matrix between all objects in a cluster can be calculated. Based on this, for example, a downmix gain correction factor for each individual object can be calculated. Thus, for example, the overall energy relationship between correlated and uncorrelated objects can be preserved.

[0127] In particular, in a particular embodiment, for example, downmix coefficients may be calculated, which may include, for example: The cross-correlation matrix C between all member objects of a cluster can be calculated, for example, as a dot product from the signal samples. Furthermore, a normalized correlation matrix C_norm can be calculated, for example, to include the respective Pearson correlation coefficients (thus, the main diagonal of C corresponds to the signal energy, while the main diagonal of C_norm is all equal to 1). For the purpose of energy-preserving downmixing, for example, only medium to high correlations may be important. Even the inclusion of low, negligible correlations due to random influences may hinder the stability of the downmix algorithm and therefore the perceived quality. Therefore, low correlations can be removed, for example, by applying a threshold and setting all components in C to 0 if the absolute value of C_norm is less than 0.5.

[0128] Optionally, the correlation may be limited, for example, to only positive correlations, so that only the energy increase due to correlation is compensated (e.g., to avoid clipping of the signal before downmixing in applications where there is not enough headroom), but no boost is applied in case of signal cancellation. For each object, the energy weighting factor w_En can be calculated, for example, as the ratio between the sum over the corresponding row of the correlation matrix and the signal energy. In other words, this coefficient approximates how much each object's energy is boosted due to its correlation with other signals. If all signals have correlations below the threshold, there will be only non-zero elements on the main diagonal, and all coefficients will be 1.

[0129] Each weighting factor w_A for scaling the signal amplitude can be calculated, for example, as the square root of the reciprocal of the energy weight. TIFF2025533617000009.tif1630 The coefficient w_A is applied as a scalar multiplication to the signal before summation in the time domain. In an exemplary implementation, the weighting factor w_A may be limited to a maximum value of, e.g., 2, to prevent excessively large boost factors in case of strong signal cancellation (alternatively, |w_En| may be limited to a minimum value of, e.g., 0.25, to prevent division by 0 as well). Note that when signal cancellation occurs, a large weighting factor will result in boosting of the remaining background noise rather than reconstruction of the canceled signal components.

[0130] An improvement to prevent signal cancellation is to detect strong negative correlations by an appropriate threshold (e.g., C(i,j)<-0.8) and either set the weighting factor of one of the negatively correlated signals to 0 (e.g., to consider only one of the otherwise canceling signals) or apply a negative weight. Note that even for negative correlations in scenarios of reproduction of individual point sources in non-anechoic environments, it can be assumed that the signals do not completely cancel at the listener's position due to decorrelation from room reverberation, etc. Stronger signal cancellation may occur in renderings with sparse loudspeakers. As a further improvement, the correlation analysis and summation can be applied in the frequency domain, for example, using an STFT (Short-Time Fourier Transform) filter bank with appropriate band grouping.

[0131] The following describes distance-based gain considerations according to one embodiment. Depending on the target use case, the rendering algorithm may also take into account the distance of the reproduced sound source. A basic implementation is to apply a distance-based gain to account for the radial distance between the listener and the sound source. If the target renderer is known to apply a distance-dependent gain, this can be compensated for when downmixing clusters, for example, to prevent perceptible loudness differences in the reproduced scene. If the actual distance gain function of the renderer is known to the clustering algorithm, a simple solution is to calculate the gain at the original source positions and the merged cluster positions, and correct for the resulting gain difference before downmixing.

[0132] A generalized, computationally efficient approach for PCS-based clustering can utilize the radial distance component from the PCS, which can be previously modeled, for example, from distance-dependent gain differences. Thus, the difference in the radial distance component between the object position and the cluster position can be calculated directly and applied, for example, as a gain difference, for example in dB.

[0133] Further embodiments are described below. According to a first embodiment, for example, object-based audio scene clustering based on a listener-related perceptual model can be performed. In a second embodiment, the clustering algorithm of the first embodiment may be based on, for example, a perceptual distance metric / perceptual distortion metric (PDM). According to a first variant of the second embodiment, for example, the identification and joining of clusters of objects within a given maximal PDM linkage can be performed for example on all pairs just below a noticeable difference. According to a second variant of the second embodiment, clustering, e.g. by iterative agglomeration of nearest objects in the PDM, may be performed, e.g., until a target number of clusters is achieved, or, e.g., until a given maximum value of the distortion metric is exceeded.

[0134] In a third embodiment, the clustering algorithm of the first embodiment can be based on, for example, 3D-DLM similarity. According to a first variant of the third embodiment, a 3D-DLM reconstruction of the original scene can be performed, for example by fitting a Gaussian Mixture Model (GMM). According to a second variant of the third embodiment, for example, an improved Expectation Maximization (EM) algorithm for GMM fitting of weighted data points on an arbitrary grid can be utilized.

[0135] In the fourth embodiment, for example, one or more improvements for temporal stability may be made to the object-based clustering of the first to third embodiments. According to a first variant of the fourth embodiment, for example, temporal smoothing and a penalty factor in the perceptual distance metric can be implemented. According to a second variant of the fourth embodiment, for example, optimization of cluster assignment permutation based on energy distribution can be performed. According to a third variant of the fourth embodiment, the resulting cluster centroid positions can be stabilized, for example, by hysteresis.

[0136] In the fifth embodiment, for example, a perceptual optimization of the centroid positions resulting from the clustering of one of the first to third embodiments can be performed. According to the sixth embodiment, for example, it is possible to perform cluster assignment and optimization of the centroid position based on spectrum matching ("EQ matching of HRTFs") for the clustering of the first embodiment. In the seventh embodiment, for example, signal processing can be performed on the union of audio objects resulting from the clustering of the first embodiment.

[0137] According to a first variant of the seventh embodiment, for example, a cross-fade can be performed to prevent signal discontinuities in the reassignment of object membership to clusters. According to a second variant of the seventh embodiment, for example, consideration of signal correlation can be made in order to achieve energy conservation. According to the third modification of the seventh embodiment, for example, it is possible to adjust the gain based on the distance. According to a fourth variant of the seventh embodiment, for example, equalization can be performed to compensate for perceptual differences due to spectral cues.

[0138] While some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with a block or device corresponding to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0139] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partly in hardware, or at least partly in software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, on which electronically readable control signals are stored, which cooperates (or is capable of cooperating) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable.

[0140] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0141] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code can be stored on, for example, a machine-readable carrier. Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an embodiment of the inventive methods is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0142] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or sequence of signals can for example be adapted to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0143] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver. In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0144] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the immediately following claims and not by the specific details presented by way of illustration and description of the embodiments herein.

Claims

1. An apparatus (100) comprising: an input interface (110) for receiving information about three or more audio objects; a cluster generator (120) for generating the two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that, for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that, for each of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster; The apparatus (100), wherein the cluster generator (120) is configured to generate the two or more audio object clusters according to a perceptually based model.

2. the cluster generator (120) is configured to generate the two or more audio object clusters according to a perceptually based model by generating the two or more audio object clusters according to at least one of a perceptual distance metric, a directional loudness map, a perceptual coordinate system, and a spatial masking model. The apparatus (100) of claim 1.

3. the cluster generator (120) is configured to generate, for a pair of two audio objects among the three or more audio objects, the two or more audio object clusters according to the perceptual distance metric by determining whether the two audio objects have a perceptual distance according to the perceptual distance metric that is less than or equal to a threshold, and if the perceptual distance is less than or equal to the threshold, by associating the two audio objects with the same one of the two or more audio object clusters. The apparatus (100) of claim 2.

4. the cluster generator (120) is configured to generate the two or more audio object clusters according to the perceptual distance metric by iteratively associating two perceptually closest audio objects among the three or more audio objects according to the perceptual distance metric until a predetermined target number of audio object clusters is reached or until a predetermined maximum perceptual distance according to the perceptual distance metric is exceeded. The apparatus (100) of claim 2.

5. the cluster generator (120) is configured to generate the two or more audio object clusters in response to a three-dimensional directional loudness map. An apparatus (100) according to any one of claims 1 to 4.

6. the cluster generator (120) is configured to generate the two or more audio object clusters by using a Gaussian mixture model; the cluster generator (120) is configured to determine two or more audio object clusters by determining components of the Gaussian mixture model such that the three-dimensional directional loudness map is approximated.

6. The apparatus (100) of claim 5.

7. the cluster generator (120) is configured to generate the two or more audio object clusters by using a Gaussian mixture model; the cluster generator (120) is configured to determine two or more audio object clusters by using an expectation-maximization algorithm to fit weighted data points onto an arbitrary grid of the Gaussian mixture model; 6. The apparatus (100) of claim 5.

8. the cluster generator (120) is configured to perform a perceptual optimization of the centroid positions resulting from the clustering, and / or the cluster generator (120) is configured to optimize cluster assignment and centroid locations according to spectral matching of the two or more audio object clusters; An apparatus (100) according to any one of claims 1 to 7.

9. the cluster generator (120) is configured to generate the two or more audio object clusters as a first plurality of audio object clusters by creating an association between each of the three or more audio objects and at least one of the two or more audio object clusters; the cluster generator (120) is configured to generate the second plurality of two or more audio object clusters such that at least one audio object of the three or more audio objects is associated with a different audio object cluster of a second plurality of audio object clusters compared to the audio object cluster of the first plurality of audio object clusters with which the at least one audio object was associated. An apparatus (100) according to any one of claims 1 to 8.

10. the cluster generator (120) is configured to generate the second plurality of two or more audio object clusters in response to temporal smoothing and / or in response to one or more penalty factors in the perceptual distance metric.

10. The apparatus (100) of claim 9.

11. the cluster generator (120) is configured to generate the second plurality of two or more audio object clusters by optimizing cluster assignment permutations according to energy distributions of the three or more audio objects. An apparatus (100) according to claim 9 or claim 10.

12. the cluster generator (120) is configured to generate the second plurality of two or more audio object clusters by stabilizing the resulting cluster centroid positions with hysteresis. An apparatus (100) according to any one of claims 9 to 11.

13. the cluster generator (120) is configured to generate the second plurality of two or more audio object clusters by perceptually optimizing centroid positions resulting from the clustering to generate the first plurality of two or more audio object clusters; and / or the cluster generator (120) is configured to generate the second plurality of two or more audio object clusters by optimizing cluster assignment and centroid locations according to spectral matching of the first plurality of audio object clusters. An apparatus (100) according to any one of claims 9 to 12.

14. the cluster generator (120) is configured to, for each audio object cluster to which at least two of the three or more audio objects are associated, perform signal processing by combining the audio object signals of each audio object associated with the audio object cluster; An apparatus (100) according to any one of claims 1 to 13.

15. The cluster generator (120) Crossfading to prevent signal discontinuities in the reassignment of object membership to clusters; Considering signal correlation to achieve energy conservation, distance-based gain adjustment, Equalization to compensate for perceptual differences due to spectral cues; The apparatus (100) of claim 14, configured to perform at least one of the following:

16. the cluster generator (120) is configured to generate the two or more audio object clusters depending on an actual or assumed location of a listener; An apparatus (100) according to any one of claims 1 to 15.

17. The cluster generator (120) is configured to determine one or more characteristics of each audio object cluster of the two or more audio object clusters in response to one or more characteristics of the three or more audio objects associated with the audio object cluster, the one or more characteristics comprising: an audio signal associated with said audio object cluster; a location associated with said audio object cluster; 17. The apparatus (100) of any one of claims 1 to 16, comprising at least one of:

18. the apparatus (100) further comprises an encoding unit for generating encoded information encoding information about the two or more audio object clusters; 18. Apparatus (100) according to any one of claims 1 to 17.

19. 1. A system comprising: An apparatus (100) according to claim 18, a decoding unit (210) for decoding the encoded information to obtain information about the two or more audio object clusters; a signal generator (220) for generating two or more audio output signals in response to the information about the two or more audio object clusters.

20. A decoder (200), a decoding unit (210) for decoding the encoded information to obtain information regarding two or more audio object clusters, the two or more audio object clusters being generated by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that, for each of the two or more audio object clusters, at least one of three or more audio objects is associated with the audio object cluster and, for each of the at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster, the two or more audio object clusters being generated according to a perceptually based model; and a signal generator (220) for generating two or more audio output signals in response to said information about said two or more audio object clusters; A decoder (200) comprising:

21. 1. A method comprising: receiving information about three or more audio objects; and generating the two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that, for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that, for each of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster; A method, wherein generating the two or more audio object clusters is performed according to a perceptually based model.

22. 1. A method comprising: decoding the encoded information to obtain information regarding two or more audio object clusters, the two or more audio object clusters being generated by associating each of the three or more audio objects with at least one of the two or more audio object clusters such that, for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster and, for each of the at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster, the two or more audio object clusters being generated according to a perceptually based model; and generating two or more audio output signals in response to said information about said two or more audio object clusters; A method comprising:

23. 23. A computer program for performing the method of claim 21 or claim 22 when the computer program is run on a computer or signal processor.

Citation Information

Patent Citations

  • Device for generating vector quantization code book

    JP1993304478A

  • Sound reproducing device

    JP2000013900A

  • Object Clustering for Rendering Object-Based Audio Content Based on Perceptual Criteria

    JP2016509249A

  • Spatial error metric for audio content

    JP2017508175A

  • Directional Loudness Map Based Audio Processing

    JP2022505964A