Apparatus and method for perception-based clustering of object-based audio scenes
Through the audio object clustering algorithm based on the perception model, audio object clustering is generated and optimized, the problem of not considering auditory characteristics in the prior art is solved, and the effect of maintaining high perceptual quality when reducing the number of audio objects is achieved, which is suitable for audio encoding and storage.
Patent Information
- Application Number
- CN202380080935.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-29
- Filing Date
- 2023-09-27
- Publication Date
- 2025-07-04
AI Technical Summary
Existing object-based audio clustering algorithms fail to effectively consider the spatial positioning accuracy and perceptual characteristics of human hearing, resulting in affecting perceptual quality when reducing the number of audio objects.
Audio object clustering is adopted using perception-based models, audio objects are grouped into clusters through generators, and a smaller number of audio object clustering is generated based on perceived distance metrics, directional loudness maps, perceived coordinate systems and spatial masking models, optimizing centroid position and signal processing to maintain high perceived quality.
While reducing the number of audio objects, the perceived quality of audio scenes is maintained or improved, and the computing needs of transmission and rendering are reduced. It is suitable for different application scenarios such as audio encoding and storage.
Smart Images

Figure CN120266499A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus and method for perception-based clustering of object-based audio scenes. Background Art
[0002] Modern audio reproduction systems are capable of providing immersive three-dimensional (3D) sound experiences.
[0003] A common format for 3D sound reproduction is channel-based audio, where individual channels associated with defined speaker positions are generated through multi-microphone recording or studio-based production. Another common format for 3D sound reproduction is object-based audio, which utilizes so-called audio objects that are placed by a producer in a listening room and converted into speaker or headphone signals for playback through a rendering system. Object-based audio offers a high degree of flexibility when it comes to the design and reproduction of sound scenes. Note that channel-based audio can be considered a special case of object-based audio, where the sound sources (= objects) are located at fixed positions corresponding to the defined speaker positions.
[0004] To improve the efficiency of transmission and storage of object-based immersive sound scenes, and to reduce the computational requirements for real-time rendering, it is beneficial or even necessary to reduce or limit the number of audio objects. This is achieved by identifying groups or clusters of adjacent audio objects and combining them into a smaller number of sound sources. This process is called object clustering or object consolidation.
[0005] The literature shows that the localization accuracy of human hearing is limited and depends on the position of the sound source (e.g., horizontal localization is more accurate than vertical localization), and an auditory masking effect can be observed between spatially distributed sound sources. By exploiting the limitations of the localization accuracy of human hearing and the auditory masking effect for object clustering, the number of audio objects can be significantly reduced while maintaining high perceptual quality.
[0006] To reduce the number of audio objects while maintaining high perceptual quality, methods and algorithms have been developed to perform clustering of object-based audio with respect to the perceptual attributes of the audio scene relative to the listener.
[0007] In the prior art, auditory masking and localization models are known.
[0008] Moreover, a directional loudness map (DLM) has been proposed in the prior art, examples being:
[0009] C. Avendano, “Frequency-domain source identification and manipulation in stereo mixes for enhancement, suppression and re-panning applications,” 2003 IEEE Workshop on Applications of Signal Processing to Audio, and
[0010] P. Delgado, J. Herre, “Objective Assessment of Spatial Audio Quality using Directional Loudness Maps”, Proc. 2019 IEEE ICASSP.
[0011] Moreover, object clustering algorithms have been proposed in the prior art, examples being:
[0012] J. Herder. "Optimization of Sound Spatialization Resource Management through Clustering", The Journal of Three Dimensional Images, 1999,
[0013] Nicolas Tsingos, Emmanuel Gallo, George Drettakis: "Perceptual Audio Rendering of Complex Virtual Environments", SIGGRAPH, 2004,
[0014] Breebaart, Jeroen; Cengarle, Giulio; Lu, Lie; Mateos, Toni; Purnhagen, Heiko; Tsingos, Nicolas: “Spatial Coding of Complex Object-Based Program Material”; JAES Volume 67 Issue 7 / 8 pp. 486 - 497; July 2019.
[0015] Moreover, in the prior art, the GMM Expectation Maximization algorithm (EM algorithm) has been proposed.
[0016] Existing algorithms for clustering object - based audio consider the spatial attributes of audio objects relative to each other. However, they do not consider the perceptual characteristics relative to the listener, and thus do not consider the position - dependence of the spatial localization accuracy of human hearing. Summary of the Invention
[0017] The object of the present invention is to provide an improved concept for clustering object - based audio scenes. The object of the present invention is achieved by the device according to claim 1, the decoder according to claim 20, the method according to claim 21, the method according to claim 22, and the computer program according to claim 23.
[0018] According to an embodiment, a device is provided. The device includes an input interface for receiving information about three or more audio objects. Moreover, the device includes a clustering generator for generating two or more audio object clusters by associating each of the three or more audio objects with at least one of two or more audio object clusters, such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster. The clustering generator is used to generate two or more audio object clusters according to a perception - based model.
[0019] Moreover, a decoder is provided. The decoder includes a decoding unit for decoding encoded information to obtain information about two or more audio object clusters, where the two or more audio object clusters are generated by associating each of the three or more audio objects with at least one of two or more audio object clusters, such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster, where the two or more audio object clusters are generated according to a perception - based model. Moreover, the decoder includes a signal generator for generating two or more audio output signals according to the information about the two or more audio object clusters.
[0020] Moreover, a method according to an embodiment is provided. The method includes:
[0021] - receiving information about three or more audio objects; and
[0022] - Generating two or more clusters of audio objects by associating each of three or more audio objects with at least one of two or more clusters of audio objects such that for each of the two or more clusters of audio objects, at least one of the three or more audio objects is associated with the cluster of audio objects and such that for each of at least one of the two or more clusters of audio objects, at least two of the three or more audio objects are associated with the cluster of audio objects. The two or more clusters of audio objects are generated according to a perception-based model.
[0023] Moreover, a method according to another embodiment is provided. The method includes:
[0024] - Decoding encoded information to obtain information about two or more clusters of audio objects, wherein the two or more clusters of audio objects are generated by associating each of three or more audio objects with at least one of two or more clusters of audio objects such that for each of the two or more clusters of audio objects, at least one of the three or more audio objects is associated with the cluster of audio objects and such that for each of at least one of the two or more clusters of audio objects, at least two of the three or more audio objects are associated with the cluster of audio objects, wherein the two or more clusters of audio objects are generated according to a perception-based model; and
[0025] - Generating two or more audio output signals based on the information about the two or more clusters of audio objects.
[0026] Moreover, a computer program is provided, each of which is configured to implement one of the above methods when executed on a computer or a signal processor.
[0027] According to an embodiment, a perception-based clustering algorithm groups audio objects in an audio scene into clusters and combines the original objects into fewer output objects by combining their signals and by selecting a common centroid location as the output object location based on perception model criteria. Based on the target use case, the goal can be to achieve a given (maximum) number of output clusters or to reduce the number of objects in the scene without introducing a perceivable difference beyond a given limit. This can be achieved using different embodiments presented below.
[0028] Some embodiments relate to clustering of audio objects.
[0029] According to an embodiment, clustering based on a Gaussian mixture model (GMM) is provided.
[0030] In this clustering generation method, a three-dimensional directional loudness map (3D-DLM) can be calculated for the entire sound scene to represent the overall spatial attributes of the scene. A GMM is fitted to approximate the original DLM with a given number of components to represent the corresponding number of clusters. Thus, the algorithm aims to reconstruct the overall spatial attributes of the sound scene rather than considering the attributes of individual objects. This method is particularly beneficial for dense sound scenes consisting of a large number of objects that need to be represented by only a few cluster positions, for example, for low-complexity / low-bitrate applications.
[0031] In an embodiment, hierarchical clustering is provided. In this "agglomerative" clustering method, objects are iteratively combined, for example, based on a perceptual distance metric, until a target number of clusters is reached and / or a given limit of the distance metric is reached (e.g., all imperceptible differences are eliminated). This method is computationally efficient and provides flexibility for configuration for constant quality or constant rate applications. Moreover, this method can scale well to transparency, for example, when the number of active audio objects is below the maximum allowed number of clusters.
[0032] According to an embodiment, clustering based on the just noticeable difference (JND) is provided. This can be considered a simplified special case of the hierarchical clustering method: when objects are so close that their positions cannot be distinguished, they can be combined together to reduce redundancy without creating a perceptible difference in the entire sound scene. Thus, the JND-based clustering method determines groups of objects that are within the JND of each other for the perceptual distance metric and combines the groups of objects into clusters. This method requires a lower computational complexity and produces a variable number of output clusters with (near) transparent perceptual quality.
[0033] Enhancements are provided in a further embodiment.
[0034] In addition, some optimization schemes regarding temporal stability and the resulting cluster output positions have been developed:
[0035] For example, according to an embodiment, temporal stability is provided. Since clustering algorithms typically run frame by frame, several measures can be taken to improve the temporal stability of the results of the clustering algorithm: for example, the membership of objects to clusters can be stabilized by a penalty factor for reassigning objects to clusters in the perceptual distance metric. For DLM-based methods, the DLM can be temporally smoothed to improve temporal stability. For example, the permutation of the cluster index order can be identified and optimized to enhance the stability of the output signal and position metadata.
[0036] And / or, for example, according to an embodiment, centroid position optimization is provided. Clustering algorithms typically generate cluster centroid positions and object cluster memberships. However, in the case of considering the target reproduction scene, the output cluster positions can be further optimized using perceptual criteria.
[0037] According to some embodiments, a signal mixing and processing concept is provided. Based on the results of the above clustering algorithm, the signals of the input audio objects can be mixed and combined to obtain an output clustering signal. The signal processing in this mixing stage can also be perceptually optimized in several aspects, for example, cross-fading to avoid signal discontinuities, and / or processing the correlation between signals, and / or considering distance-based gain differences, and / or equalization to compensate for changes in spectral localization cues. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings, in which:
[0039] Figure 1 An apparatus according to an embodiment is shown.
[0040] Figure 2 A decoder according to an embodiment is shown.
[0041] Figure 3 A system according to an embodiment is shown.
[0042] Figure 4 A one-dimensional example is shown, in which the direction loudness map generated by ten sound sources is approximately represented by a Gaussian mixture model containing only two components.
[0043] Figure 5 Three different distance model levels of JND-based clustering according to an embodiment are shown.
[0044] Figures 6a - 6g A small-scale example of a two-level JND-based clustering algorithm according to an embodiment is shown.
[0045] Figure 7 The clustering index arrangement according to an embodiment resulting from a slight change in the scene is shown.
[0046] Figure 8 The clustering assignment arrangement and optimization according to an embodiment are shown.
[0047] FIG. 9 shows the centroid projection in the unit sphere in the horizontal plane and the centroid projection in the perceptual coordinate system in the horizontal plane.
[0048] Figure 10 The centroid to cone of the confusion projection in the transverse plane according to an embodiment is shown.
[0049] Figure 11 The centroid projection maintaining the height of the cone of confusion in the transverse plane according to an embodiment is shown. DETAILED DESCRIPTION
[0050] Figure 1 An apparatus 100 according to an embodiment is shown.
[0051] The apparatus 100 includes an input interface 110 for receiving information about three or more audio objects.
[0052] Moreover, the apparatus 100 includes a clustering generator 120 for generating two or more audio object clusters by associating each of the three or more audio objects with at least one of two or more audio object clusters, such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster. The clustering generator 120 is used to generate two or more audio object clusters according to a perception-based model.
[0053] According to an embodiment, the clustering generator 120 can be used to generate two or more audio object clusters according to a perception-based model by generating two or more audio object clusters according to at least one of a perceptual distance metric, a directional loudness map, a perceptual coordinate system, and a spatial masking model.
[0054] In an embodiment, the clustering generator 120 can be used to generate two or more audio object clusters according to a perceptual distance metric by, for a pair of two audio objects among the three or more audio objects, determining whether the perceptual distance between the two audio objects is less than or equal to a threshold according to the perceptual distance metric, and by, if the perceptual distance is less than or equal to the threshold, associating the two audio objects to the same audio object cluster among the two or more audio object clusters.
[0055] According to an embodiment, the clustering generator 120 can be used to generate two or more audio object clusters according to a perceptual distance metric by iteratively associating the two perceptually closest audio objects among the three or more audio objects until a preset target number of audio object clusters is reached or until a preset maximum perceptual distance according to the perceptual distance metric is exceeded.
[0056] In an embodiment, the clustering generator 120 can be used to generate two or more audio object clusters according to a three-dimensional directional loudness map.
[0057] According to an embodiment, the clustering generator 120 can be used to generate two or more audio object clusters by adopting a Gaussian mixture model. Moreover, the clustering generator 120 can be used to determine two or more audio object clusters by determining the components of the Gaussian mixture model to approximate the three-dimensional directional loudness map.
[0058] In an embodiment, the clustering generator 120 can be used to generate two or more audio object clusters by employing a Gaussian mixture model. Moreover, the clustering generator 120 can be used to determine two or more audio object clusters by fitting weighted data points on an arbitrary grid of the Gaussian mixture model by employing an expectation maximization algorithm.
[0059] According to an embodiment, the clustering generator 120 can be used to perform perceptual optimization on the centroid positions of the generated clusters.
[0060] In an embodiment, the clustering generator 120 can be used to optimize the cluster assignment and centroid positions according to the spectral matching of two or more audio object clusters.
[0061] According to an embodiment, the clustering generator 120 can be used to generate two or more audio object clusters as a first plurality of audio object clusters by associating each of three or more audio objects with at least one of two or more audio object clusters. Moreover, the clustering generator 120 can be used to generate a second plurality of two or more audio object clusters such that at least one audio object among the three or more audio objects is associated with a different audio object cluster in the second plurality of audio object clusters than the audio object cluster in the first plurality of audio object clusters associated with the at least one audio object.
[0062] In an embodiment, the clustering generator 120 can be used to generate a second plurality of two or more audio object clusters according to temporal smoothing and / or according to one or more penalty factors in a perceptual distance metric.
[0063] According to an embodiment, the clustering generator 120 can be used to generate a second plurality of two or more audio object clusters by optimizing the cluster assignment arrangement according to the energy distribution of three or more audio objects.
[0064] In an embodiment, the clustering generator 120 can be used to generate a second plurality of two or more audio object clusters by stabilizing the generated cluster centroid positions via a hysteresis effect.
[0065] According to an embodiment, the clustering generator 120 can be used to generate a second plurality of two or more audio object clusters by performing perceptual optimization on the centroid positions resulting from the clusters that generate the first plurality of two or more audio object clusters.
[0066] In an embodiment, the clustering generator 120 can be used to generate a second plurality of two or more audio object clusters by optimizing the cluster assignment and centroid positions according to the spectral matching of the first plurality of audio object clusters.
[0067] According to an embodiment, the clustering generator 120 can be used to cluster each audio object associated with at least two of three or more audio objects, and perform signal processing by combining the audio object signals of each audio object associated with the audio object clustering.
[0068] In an embodiment, the clustering generator 120 can be used to perform at least one of the following:
[0069] Cross-fading to prevent signal discontinuity when object-to-cluster membership is reallocated;
[0070] Consider signal correlation to achieve energy conservation;
[0071] Adjust the distance-based gain; or
[0072] Equalization to compensate for perceptual differences caused by spectral cues.
[0073] According to an embodiment, the clustering generator 120 can be used to generate two or more audio object clusters according to the actual position or assumed position of the listener.
[0074] In an embodiment, the clustering generator 120 can be used to, for each audio object cluster among two or more audio object clusters, determine one or more attributes of the audio object cluster according to one or more attributes of the audio objects associated with the audio object cluster among three or more audio objects, where the one or more attributes include at least one of the following:
[0075] The audio signal associated with the audio object cluster; or
[0076] The position associated with the audio object cluster.
[0077] According to an embodiment, the apparatus 100 may further include an encoding unit for generating encoded information encoding information about two or more audio object clusters.
[0078] Figure 2 Show a decoder 200 according to an embodiment.
[0079] The decoder 200 includes a decoding unit 210 configured to decode encoded information to obtain information about two or more audio object clusters, wherein the two or more audio object clusters are generated by associating each of three or more audio objects with at least one of the two or more audio object clusters such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster and such that for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster, and wherein the two or more audio object clusters are generated according to a perception-based model.
[0080] Moreover, the decoder 200 includes a signal generator 220 configured to generate two or more audio output signals based on the information about the two or more audio object clusters.
[0081] Figure 3 Fig. shows a system according to an embodiment.
[0082] The system includes Figure 1 the apparatus 100 in Figure 1 The apparatus 100 further includes an encoding unit configured to generate encoded information encoding the information about the two or more audio object clusters.
[0083] Moreover, the system includes a decoding unit 210 configured to decode the encoded information to obtain the information about the two or more audio object clusters.
[0084] In addition, the system includes a signal generator 220 configured to generate two or more audio output signals based on the information about the two or more audio object clusters.
[0085] Before describing the preferred embodiments in more detail, some background considerations on which the embodiments of the present invention are based are described.
[0086] Now, consider the perception model and provide an overview of the perception model that forms the basis of the clustering algorithm and method according to an embodiment.
[0087] The presented psychoacoustic model may for example include the following core components corresponding to different aspects of human perception, namely, a 3D directional loudness map, a perception coordinate system, a spatial masking model, and a perception distance metric.
[0088] First, the 3D-Directional Loudness Map (3D-DLM) is described. The basic idea of the Directional Loudness Map (DLM) is to find a representation of "how loud is perceived from a given direction". This concept has been proposed as a one-dimensional method for representing binaural localization in the binaural DLM (Delgado et al., 2019). This concept is now extended to three-dimensional (3D) localization by creating a 3D-DLM on the surface around the listener to uniquely represent the perceived loudness according to the angle of incidence relative to the listener. It should be noted that the binaural DLM is obtained by analyzing the signals at the ears, while the 3D-DLM is utilized for object-based audio synthesis using the a priori known source positions and signal characteristics.
[0089] Now, a Perceptual Coordinate System (PCS) is proposed. The human source localization accuracy varies with different spatial directions. To represent this in a computationally efficient way, the Perceptual Coordinate System (PCS) is introduced. To obtain the PCS, the spatial positions are distorted to correspond to the non-uniform characteristics of the localization accuracy. Thus, the distances in the PCS correspond to the "perceptual distances" between corresponding positions, such as the number of Just Noticeable Differences (JNDs), rather than physical distances. This principle is similar to the use of psychoacoustic frequency scales, such as the Bark scale or the ERB scale, in perceptual audio coding.
[0090] Now, the Spatial Masking Model (SMM) is described. The monaural time-frequency auditory masking model is a basic element of perceptual audio coding and is typically enhanced by binaural (non-)masking models to improve stereo coding. The Spatial Masking Model extends this concept to immersive audio to combine and utilize the masking effects between arbitrary source positions in 3D.
[0091] Regarding the perceptual distance metric, it should be noted that the above components can be combined to obtain a perception-based distance metric between spatially distributed sound sources. These can be used in various applications, such as as a cost function in object clustering algorithms, to control the bit distribution in perceptual audio encoders, and for obtaining objective quality measurements. These metrics address some problems, such as, "How easily is it noticed if the position of a sound source changes?", "How distinct is the difference between two different sound scene presentations?", "How important is a given sound source in the whole sound scene? (How obvious would it be if the given sound source were removed?)".
[0092] The developed clustering concepts and algorithms are introduced below.
[0093] In applications using object-based audio, it is desirable to reduce the number of objects required to represent a sound scene while maintaining high perceptual quality, in order to improve the efficiency of transmission, storage, and the computational complexity of the rendering application. Accordingly, perception-based audio object clustering can be employed. In other words, based on the proposed perception model, audio objects with similar perceptual attributes can, for example, be grouped and combined into fewer audio objects.
[0094] Depending on the use case, the range of desired target attributes is wide, as is the range of how much the number of objects in the scene is reduced. In the field of audio coding, there are well-known paradigms that aim to achieve constant quality at variable bit rate (VBR), or variable quality at constant bit rate (CBR). Accordingly, object clustering can, for example, be configured to target constant quality, which will result in a variable number of clusters (= output objects), or a constant number of concurrent objects with variable quality.
[0095] The most conservative approach aims to remove only the redundancy and irrelevance in the scene representation. This means that, for example, only objects that can be combined without introducing an audible change to the scene can be merged, thus reducing the number of objects without affecting the perceptual quality ("transparent" clustering). For example, this approach can also be extended to further reduce the object count by clustering objects within a selected threshold of a perceptual distance metric, i.e., a maximum distance (e.g., a multiple of the JND distance). For example, these methods may result in a variable number of clusters and output objects.
[0096] On the other hand, in many applications, for example, the maximum number of objects may be determined by external factors such as the maximum number of transmission channels in an audio codec profile or the number of signals that can be processed by a real-time renderer. Depending on the use case, this may lead to stringent requirements for the reduction factor. For example, a movie scene written with up to 128 objects can be reduced to a channel bed plus 4 to 8 objects (e.g., for transmission via MPEG-H LC level 3 with a maximum of 16 transmissions, e.g., 7.1 + 4 channels + 4 objects). For these use cases, the clustering algorithm may, for example, produce a given constant or maximum number of clusters.
[0097] For example, the maximum number of clusters can be directly obtained from a maximum-distance-based method by increasing the allowed distance until the number of resulting clusters is below the limit. However, this may lead to ambiguity and may result in the number of output clusters being lower than the target number, which will lead to an unnecessary reduction in quality.
[0098] According to an embodiment, an iterative hierarchical clustering algorithm is proposed, in which the perceptual distance is used as the optimization criterion, and the number of objects in the scene is reduced by iterative pairwise grouping. Moreover, for very severe reduction factors, for example, instead of individual sound sources, by approximating the spatial distribution of the loudness, it may be beneficial to "generate" the entire sound scene by a "generation" method.
[0099] The following considers clustering based on a Gaussian mixture model (GMM).
[0100] For example, clustering based on a mixture model can be considered a generation method. For example, a given DLM is approximated by a given number of components in a GMM. In other words, this method assumes a given / predefined (maximum) number of available sound sources, aiming to reconstruct the overall loudness distribution of a given / predefined scene, rather than looking at the individual sound source positions. Therefore, it can be considered a scene-based method (not to be confused with high-fidelity stereophonic sound reproduction (Ambisonics), which is usually called "scene-based audio").
[0101] This method is particularly advantageous when a large number of objects need to be represented by only a few clustering positions (e.g., for low-bitrate applications), for example, when many input objects are usually assigned to one cluster. On the contrary, recreating a large number of positions through a similar large distribution is not computationally efficient.
[0102] Figure 4 A simplified one-dimensional example is shown, in which the DLM generated by ten sound sources is approximated by a GMM with only two components.
[0103] This GMM-based method not only produces the centroid positions and memberships, but also produces the probability that a certain point belongs to a given cluster. This is advantageous for identifying cases where the cluster membership is ambiguous (e.g., the sound source at position 45 in the shown example). This information can be used to achieve temporal stability by lagging the fluctuations of the membership assignment, and can even be used to implement a soft clustering method, in which, in the context of audio object clustering, an object can be mixed into two output clusters.
[0104] The expectation maximization (EM) algorithm is a well-known method for fitting a GMM to the distribution density of a set of given data points. For example, a potential model assumption may be that the input data points are placed by a random process, and its probability distribution density is a mixture of Gaussian distributions within a given coordinate system. In other words, the GMM aims to approximate the probability that the data points are located at given positions.
[0105] The Expectation-Maximization (EM) algorithm is an iterative method for fitting such a probability distribution to a given set of data points. In principle, this method is similar to the well-known k-means clustering algorithm, which iteratively assigns points to the nearest centroid locations and then updates the centroid locations based on the updated cluster memberships. Simply put, the EM algorithm is the "soft" version of this method, where instead of a "hard" membership of points to clusters, the parameters of the Gaussian distribution (centroid locations and standard deviations) are updated based on the probabilities of points belonging to each individual Gaussian component. Thus, in each update step, a point can influence the centroid location for a component. Vice versa, the result of the EM algorithm will not only yield centroid locations and memberships, but also the "spread width" (standard deviation) of the individual components, thus yielding the probability of a point belonging to a given cluster.
[0106] The EM algorithm consists of two named steps: Expectation and Maximization, which are iteratively repeated until a convergence criterion is met. As a high-level explanation (omitting the underlying statistics), the steps that are iteratively repeated are the Expectation step and the Maximization step.
[0107] In the Expectation step, the distribution parameters are assumed to be given, e.g., centroid locations and standard deviations, and the membership probabilities, i.e., the probabilities of each point belonging to each individual Gaussian component, are calculated.
[0108] In the Maximization step, the membership probabilities are assumed to be given, and the distribution parameters are updated, e.g., the centroid and the distribution width are calculated from the mean and variance and weighted by their respective membership probabilities.
[0109] As an exit criterion for the iteration, the log-likelihood of the distribution can be used as a "goodness-of-fit" measure. Additionally, the iteration count may typically be limited to control the maximum computation time.
[0110] Existing DLM-based object clustering has limitations. In principle, fitting a GMM to data points is a common task and can be done using algorithms and toolboxes (e.g., provided by Matlab toolboxes). However, typical applications are fitting the model to a random distribution of unweighted points with different densities. In contrast, DLM represents a regular grid of points with different weights. This difference hinders the direct use of available algorithms and toolboxes for GMM fitting. To be able to utilize existing toolboxes, this mismatch can be addressed through data preprocessing, e.g., by repeating points based on DLM values to simulate different distribution densities. However, due to point repetition, there will be a significant expansion of the data, and thus it is not efficient in terms of memory requirements and computational complexity. Moreover, the chosen sampling grid of the DLM may hinder the result of feeding the preprocessed data into existing GMM fitting algorithms: if the sampling grid and the distribution of relative point densities are non-uniform, the centroids of the resulting GMM will be biased towards regions with higher sampling point densities, e.g., concentrated at the poles where the azimuth / elevation domain is uniformly sampled.
[0111] As a side note, it should be noted that the grid-based data required for DLM fitting uses an analogy of the statistical method of the EM algorithm for the analysis of histogram data, rather than the analysis of the underlying point distribution. However, interestingly, there is not much literature on using the EM algorithm to process grid / histogram data. Since the histogram is first generated from the underlying data, binning the data to generate a histogram reduces accuracy and is done only for reasons of computational efficiency or data acquisition (e.g., CHIANG et al.: "Where are the passengers? A Grid-Based Gaussian Mixture Model for taxi bookings", 2015), and there does not seem to be any available toolbox support. Moreover, the histogram-based method assumes a uniformly sampled grid, which is not necessarily given for the DLM sampled on the sphere.
[0112] Moreover, fitting a model to represent the probability of a random distribution results in a distribution where the sum (or integral) over all positions is always normalized to unity, e.g., equal to 1. However, in the DLM, the sum is determined by the sum of the loudness of individual sound sources and is not normalized to a constant value.
[0113] Therefore, an enhanced EM algorithm has been developed that is modified to make the GMM suitable for a set of weighted points in a grid at arbitrary positions.
[0114] For the PCS-based DLM, the distance is actually modeled to fit the Euclidean distance between two given points, rather than the angular distance (e.g., considering front / back confusion). Therefore, the underlying distribution model is a three-dimensional Gaussian distribution, rather than a surface distribution (e.g., spherical distribution).
[0115] The enhanced EM algorithm for weighting data points according to an embodiment is described below.
[0116] As a specific embodiment, a detailed example operation of the developed algorithm is shown in the following pseudocode representation:
[0117] The algorithm parameters can include, for example, one or more or all of the following parameters:
[0118] The input parameters can include, for example:
[0119] Pre-generated loudness map (sampled grid point positions p i and corresponding loudness values DLM(p i )); and
[0120] The target number of clusters k.
[0121] Output parameters may include, for example:
[0122] The position of the cluster centroid c_l;
[0123] The membership probability clusterProb(i, l) of each input position for each component;
[0124] The "hard" membership assignment of positions to clusters (providing interface compatibility with other clustering methods that produce centroids and memberships);
[0125] The distribution parameters that determine the Gaussian components of the model DLM: for example, the centroid position c_l; the distribution parameter sigma_l (= the standard deviation of the Gaussian distribution); the weight parameter a_l (scaling the weights to represent the different loudness of different components);
[0126] The GMM approximation DLM_GMM(pi) of the generated DLM distribution; and
[0127] The error metric: the sum of squared errors (SSE) between the input DLM(p i ) and the approximate DLM_GMM(p i ) distributions.
[0128] The algorithm initialization according to an embodiment is described below.
[0129] As a general note, it should be noted that since the membership probabilities and corresponding contribution weights of each point to the cluster cannot be obtained during initialization (because they are the results of probability estimation), "hard" membership and geometric distances are used for initialization, and then the weight and width distribution parameters of the Gaussian components are determined and refined in subsequent iterative steps.
[0130] For the initialization of the centroid positions c_j (c_1...c_k), there are multiple options. For example, the initialization of the centroid positions can be carried out as follows: for the first processing frame, for example, k of the loudest input objects can be selected, the initialization can be performed at random positions, the k-means clustering algorithm with (faster calculation speed) random initialization can be performed, and for example, the result can be used as a better guess for the initial centroid positions to improve the convergence speed of the EM algorithm (for example, rough clustering by k-means and then refinement by the EM algorithm). In subsequent frames, the previous centroid positions can be used for initialization to improve temporal stability, and re-initialization can be performed using one of the above methods, for example, based on scene change detection. Optionally, multiple instances of the EM algorithm using different initialization methods (for example, the previous positions and the current loudest objects) can be run, and the result with a lower error metric can be selected.
[0131] The initialization of the membership relationship mem(i) can be carried out, for example, by assigning all points closest to the centroid based on the Euclidean distance d_i(j) = d(p_i, c_j) = |p_i - c_j|, or, if the initialization is done by k-means, the initialization of the membership relationship may already be provided.
[0132] The distribution width parameter sigma can be initialized, for example, as the standard deviation. As a first option, based on the distribution of the initial centroids, i.e., the same for all components: sigma(j, dim) = std({c_1(dim), … c_k(dim)}); or as a second option, based on the standard deviation of the positions of the initialized cluster members: sigma(j, dim) = std(p(mem == j)). It should be noted that for multi-dimensional data, it is assumed that the Gaussian distribution is separable in each dimension, i.e., the distribution width controlled by the standard deviation parameter sigma(j, dim) is determined independently for each dimension dim and cluster index j (e.g., in the case of 3D-DLM, the 3 degrees of freedom can be reduced to 2D, e.g., for use cases with only sound sources in the horizontal plane).
[0133] For example, regularization of Sigma can be performed. For example, it can be restricted to values between regmin and regmax (e.g., [1, 5]) to maintain stability and prevent too narrow or too wide width distributions, which would hinder the convergence of the algorithm (e.g., if a cluster has only one member during initialization, the distribution width will effectively be zero, hindering other members from belonging to that cluster). Besides the stability of the algorithm, this is also for psychoacoustic considerations, because the distribution width representing the membership probability, i.e., the "uncertainty" vice versa, should not be less than the localization accuracy of the underlying perceptual model.
[0134] For example, a weight a_j can be assigned to each cluster to represent the difference in distribution weights.
[0135] To initialize the weight a_j, first, for example, using the corresponding distribution parameters c_i and sigma of the above initialization, the joint probability density function (PDF) of all dimensions of each data point jointPdf(i) can be calculated as the product of individual PDFs given by the PDF of the Gaussian normal distribution normpdf(x, mu, sigma):
[0136]
[0137] For example, the cluster weight a_j can be calculated from the ratio of the sum of jointPdfs weighted by the values of the data points to the sum of unweighted jointPdfs, e.g.:
[0138]
[0139] For example, the sum sumPdf(i) of all distributions at the data point location can be calculated as the sum of the weighted distributions of all Gaussian components to obtain an approximation of the overall DLM(p_i):
[0140]
[0141] The iterative steps according to an embodiment are described below.
[0142] In the expectation step, the probability clusterProb(i,j) that each data point belongs to a given cluster can be calculated as the contribution ratio of a single cluster to the overall PDF:
[0143] clusterProb(i,j) = jointPdf(i,j) * alpha(j) / sumPdf(i)
[0144] In other words, this is similar to calculating the ratio between the DLM of a single component and the overall DLM.
[0145] In the maximization step, the centroid position c_l can be updated, for example, as a weighted average position, weighted by the probability that all data points belong to a given cluster (independently for each dimension).
[0146]
[0147] To improve numerical stability (and avoid division by zero), a small offset can be added, for example, and the position is additionally weighted by the data point value, for example:
[0148]
[0149] Optionally, to represent data initially sampled on a sphere or ellipsoid, the centroid position is projected onto the sphere, for example, by normalizing the position vector, assuming a distribution on the unit sphere.
[0150] Similarly, the distribution width sigma_l can also be updated, for example, based on the average weighted variance, for example:
[0151]
[0152] jointPdf, the cluster weight a_j, and sumPdf can be updated as described above for initialization.
[0153] The expectation step and the maximization step can be continued iteratively until an exit criterion is met.
[0154] For example, the exit criterion can be reaching a maximum number of iterations (e.g., 50). Such an exit criterion ensures an upper bound on the overall computation time.
[0155] Alternatively, the exit criterion can be, for example, a criterion based on the sum of squared errors (SSE) between the DLM and sumPdf (instead of the log-likelihood typically used for unweighted data in the EM algorithm). For example, the exit criterion can be that the overall SSE is small enough, i.e., the fitted model is good enough. Or, the exit criterion can be, for example, that the SSE no longer decreases (i.e., the difference in SSE between two consecutive iterations is below a given threshold, e.g., 0.1*std(DLM)), meaning that the algorithm has converged and further iterations do not bring further improvement.
[0156] For output data collection, after termination, the algorithm can, for example, collect the model parameters and generate additional output values, such as distribution parameters (e.g., centroid position c_l; diffusion parameter sigma_l; and weight parameter a_l), and the membership probabilities of each position to each component. Additionally, the "hard" membership assignment mem(i) can be determined, for example, based on the highest membership probability for each point in order to provide interface compatibility with other clustering methods that also produce centroids and memberships.
[0157] The enhanced version of the EM algorithm produces the centroid positions and "hard" cluster memberships for the given input positions (which are common output parameters of clustering algorithms), as well as "soft" clustering by providing membership probabilities. Moreover, it provides the parameters of a weighted GMM model that approximates the input distribution (DLM). Compared with existing EM algorithms, the main enhancements are combining weighted input points with a variable total weight, considering inputs at uniform or non-uniform grid positions, and adapting to positions on a sphere.
[0158] Hierarchical clustering will be considered below.
[0159] Generating clustering methods, such as GMM-based methods, can be very effective in fitting a small number of clusters to a large number of input objects. However, since the computational complexity increases with the number of target clusters, this generation method does not scale well to higher numbers of clusters (and target qualities). On the one hand, the number of computations for mutual probability estimation increases; on the other hand, due to the increase in degrees of freedom, more iterations may be required to converge to a stable solution. For example, if the number of target clusters is already close to the original number of input objects, a large number of iterations may be required to converge to a solution where most objects remain unchanged in the end.
[0160] According to an embodiment, an iterative hierarchical clustering algorithm is introduced. Briefly, this algorithm iteratively selects two "closest" objects (preferably based on psychoacoustic metrics) and combines them until a target number of clusters is reached and / or until a minimum distance threshold between the closest objects is exceeded. Thus, in each iteration, the number of output objects is reduced by 1, and thus N objects will be reduced to k clusters in (N - k) iterations, providing a deterministic computational complexity.
[0161] The general concept of hierarchical clustering is well known in the literature. The algorithm developed according to an embodiment includes concepts and enhancements that can apply known concepts in the context of object-based audio clustering. However, according to an embodiment, (one or more) psychoacoustic metrics can be used, for example, as a cost function.
[0162] The distance metric for hierarchical clustering can be given, for example, by the linkage within the cluster, for example, whose distance is considered as a cost function of the members within the cluster. Common linkage models are "complete linkage", for example, the maximum distance between any two objects in the cluster, or "centroid linkage", for example, given by the distance between the respective centroids.
[0163] In the algorithm proposed according to an embodiment, a greedy iterative method can be selected, where the pairwise distance is minimized and then the centroid is updated. This corresponds to the centroid linkage model.
[0164] The hierarchical clustering algorithm according to an embodiment is described below.
[0165] The input parameters and preprocessing can include, for example:
[0166] Input object position p i ;
[0167] Input object energy (optionally, perceptually weighted, for example, by pre-filtering in the time domain to apply A-weighting); previous centroid position and object membership in subsequent frames; and
[0168] Target conditions (one or two can be specified), for example, the maximum number of clusters k, or, for example, an upper limit of the distance metric threshold.
[0169] The output parameters can include, for example:
[0170] Cluster centroid position c_l; and
[0171] Cluster membership mem(i).
[0172] The algorithm initialization according to an embodiment is described below.
[0173] For example, a masking model between input objects can be calculated. For example, a cost function / distance metric can be calculated, such as an inter-object distance matrix. For example, a baseline model can be determined, such as the Euclidean distance between object positions in world coordinates. Or for example, a perception-enhanced model can be determined, such as the Euclidean distance between object positions in PCS. Or for example, a complete model can be determined, such as, considering the masking effect of the entire scene, the pairwise perception distance D_perc can be calculated.
[0174] The iteration according to an embodiment is described below.
[0175] It should be noted that the iterative process can be carried out "in-place", for example, combining two objects into the index position of one of the objects and marking the other object as invalid. In this way, an updated centroid is formed, which can be regarded as any other object in the next iteration step. In other words, during the iteration, each object can be considered as a centroid and vice versa, so these two terms are synonymous here. The iteration can for example include:
[0176] For example, the minimum distance in the distance matrix can be selected.
[0177] For example, the corresponding two objects can be merged. For example, objects can be combined into the index of one of the two objects based on one or more of the following criteria: for example, combined into the smaller object index position (fallback); for example, combined into the object / cluster with more energy; for example, combined into the cluster with more existing members. The centroid position can for example be updated to the average position of the two merged objects, weighted by object energy, or alternatively, updated to the geometric mid-position, or updated based on the weighted average of all member positions.
[0178] For example, the parameters and distance metrics can be updated. It should be noted that the updated centroid will be processed like any object in the next iteration. All row and column entries of the "removed" object in the distance matrix may be invalid, for example, marked as excluded from further search iterations. The energy of the combined object can for example be calculated as the sum of the energies of the combined objects. For example, in a high-complexity model, the masking threshold of the new centroid position can be updated by recalculating the masking of the updated position, or, for example, in a low-complexity model, the masking threshold at the centroid position can be estimated as the maximum value, sum, or weighted average of the thresholds of the combined objects to update the masking threshold of the new centroid position. For example, the PE (perceptual entropy) of the combined object can be calculated from the updated energy and masking threshold. For example, the rows and columns of the distance matrix calculated in the initialization step of the input objects can be recalculated to update the distances to the combined object.
[0179] For example, the iteration can continue until an exit condition is met.
[0180] For example, the exit criterion can be whether the target number of clusters is reached. Alternatively, the exit criterion can be whether the minimum distance is higher than a given threshold, e.g., 1 JND.
[0181] For example, how the exit criteria are combined can depend on the target use case to achieve different purposes, e.g., constant quality, a constant number of output clusters, or, as a compromise, substantially constant quality and the maximum number of clusters (assuming this is rarely reached).
[0182] Therefore, the exit criteria can be combined in different AND / OR conditions to achieve one of the following options:
[0183] The first basic case is the "constant rate" case. For example, the iteration can continue until the target number of clusters is reached. This always results in k clusters (unless the number of input objects is already N <= k), but the quality will vary depending on the number and distribution of the input objects.
[0184] The second basic case is the "constant quality" case. For example, the iteration can continue until the minimum distance in the distance matrix exceeds a given threshold. This results in (approximate) constant quality and can be used, for example, to only remove differences that are already below or close to the JND, or below the appropriate tolerance for a given use case. However, the number of output clusters varies and can be equal to the number of input objects in the worst case.
[0185] The first combined AND case is the "constant maximum rate with irrelevance reduction" case (low target cluster number, low distance threshold). For example, the iteration can continue until the target number of clusters is reached. If the minimum distance is below a given threshold (e.g., one JND), the iteration continues to remove irrelevance from the scenario.
[0186] The second combined AND case is the "constant quality with rate ceiling" case (high target cluster number, high distance threshold). According to the (boolean) definition of the exit criteria, it is the same as the first combined AND case; however, the main parameter is the distance threshold that mainly achieves constant quality, while the target number of clusters is set relatively high to provide an upper limit on the number of output clusters (e.g., so as not to exceed the transmission channel or renderer input capacity).
[0187] The combined OR case is the "constant rate with quality barrier limit" case. This case is mentioned mainly for completeness as its possible use cases are limited. For example, the iteration can continue until one of the exit criteria is met, i.e., the number of clusters or the distance metric indicates exit. This results in a variable rate and variable quality output. Possible use cases are applications where the number of clusters (i.e., the rate) is usually constant, but large quality barriers are to be avoided, so temporarily allowing more output clusters (e.g., for file-based storage where the average rate is more important than the peak rate).
[0188] Consider clustering based on JND (Just Noticeable Difference) below.
[0189] Compared with the clustering method of "constant rate" with a given maximum number of clusters, the JND-based clustering method aims to remove only irrelevance and redundancy from the scene in order to reduce the computational complexity and / or the transmission bit rate while maintaining a perceptually transparent result or at least a constant quality (similar to the VBR mode in perceptual audio encoders).
[0190] For example, this can be achieved by clustering only the objects whose position changes do not exceed a given threshold (e.g., one JND).
[0191] This method can be used to eliminate the irrelevant separation between objects, where the distance between the objects has approached the limit distance that can be distinguished by the localization accuracy of human hearing. Therefore, it can even be performed based only on the position metadata without measuring the actual signal.
[0192] For example, JND-based clustering can be performed at different levels of strictness:
[0193] For level 1 centroid distance, the distance between the clustering centroid and the clustering objects must not exceed the threshold.
[0194] For level 2 pairwise distance between objects, the pairwise distance between all objects in the clustering operation must not exceed the threshold.
[0195] For level 3 sum of distances, the combined change of all objects in the auditory scene must not exceed the threshold (e.g., in order to achieve perceptually transparent quality).
[0196] It should be noted that level 1 and level 2 approximately correspond to "centroid linkage" and "complete linkage" in hierarchical clustering methods, while level 3 corresponds to the overall scene analysis task (e.g., measuring the sum of distances or the overall DLM divergence).
[0197] Figure 5 Show three different distance model levels of JND-based clustering according to an embodiment (titled L1 to L3)
[0198] In the given example, for level 1, all objects that may be within the JND distance of the resulting centroid can be combined. At level 2, objects may have to be close to each other within the JND distance to be combined. At level 3, even if all objects are within the JND distance, for example, only two out of three objects can be combined because otherwise the sum of the distances would exceed the JND.
[0199] For example, level 1 (centroid distance) can be implemented as a variant of the hierarchical clustering algorithm above by not setting the target number of clusters in the exit criterion and only considering that the minimum entry in the distance matrix min(D_perc) is below a given threshold, or only considering that the perceptual space distance D_PCS is below 1 JND, regardless of the masking and energy attributes. The latter supports clustering in applications where the algorithm only knows the position metadata and not the signal energy.
[0200] For example, level 3 (sum of distances) can be implemented by a hierarchical clustering algorithm where, for example, the sum of the distances can be used as the exit criterion instead of the minimum distance, or the divergence of the DLM of the entire scene can be used as the exit criterion. However, it should be noted that the repeated calculation of the DLM divergence results in a relatively high computational complexity, so it is more suitable for encoding and conversion tasks rather than real-time applications.
[0201] Between the strictness of levels 1 and 3, level 2 (distance between objects) presents a favorable compromise. Since level 2 only depends on the initial object positions, it can be implemented, for example, with low computational complexity. Therefore, in most applications, level 2 is the recommended operating mode. Since only the pairwise distance metric between objects is considered, it can be performed based on only one initial calculation of the distance matrix without iteratively updating the centroid positions and distances. To improve the computational complexity of the target clustering system, this JND clustering based on object distances can be performed as a preprocessing step before applying an iterative (hierarchical or GMM-based) clustering algorithm to achieve the target number of clusters, reducing the initial number of clusters with a relatively low computational cost while maintaining the transparency quality. It should be noted that generally, there is no unique solution for such clustering because different groupings are possible (for example, A + B and B + C can be combined, but A + C cannot). Optimizing such a "complete-link" clustering problem to minimize the number of clusters is known as the "exact cover problem" in the literature, which has been proven to be NP-complete. However, in the application of object clustering, the distance metric presents an alternative optimization criterion, based on which a greedy algorithm with a relatively low computational complexity is derived. For example, the algorithm according to an embodiment can be implemented as follows:
[0202] For example, an initial distance matrix can be calculated. Based on the use case, this can, for example, be based on D_PCS to only consider spatial relationships, or, for example, be based on D_perc to additionally consider masking attributes. The advantage of using D_PCS is that the JND clustering step is independent of the signal energy, i.e., it can be performed with very low computational complexity. The advantage of using D_perc is that it models the perceptual attributes more accurately. Additionally, since silent or inaudible objects are assigned zero (or near-zero) PEs, this implicitly acts as a screening stage for integrating irrelevant objects.
[0203] All entries in the distance matrix below a selected threshold (outside the main diagonal) can, for example, be marked in pairs and can potentially be combined in a boolean combination matrix. The threshold can be selected according to the use case. For clustering based on the D_PCS distance, the threshold can be selected as 1 [JND] to only merge objects within the localization accuracy of human hearing. For clustering based on D_perc, the masking attribute is additionally incorporated into the distance metric through the PE. Assuming the signal is exactly at the masking threshold, the resulting PE is log2(1 + 1 / 1) = 1 [bit]. Thus, similarly, the threshold of 1 [bit*JND] for D_perc can be selected as a simple approximation.
[0204] For example, all elements for which the combination matrix is true can be considered candidate pairs.
[0205] For example, clustering creation can be started by initializing the clustering of two objects by selecting the candidate pair with the smallest entry in the distance matrix from the candidate pairs.
[0206] For example, iterative objects can be integrated into the cluster in the following way:
[0207] Select the corresponding true entries in the combination matrix to create a list of candidate objects (candidate list) that can be added to the cluster, i.e., objects that can be combined with all objects already existing in the cluster (although not necessarily all combined with each other).
[0208] Select the candidate object with the smallest absolute distance or the smallest sum of distances to all objects in the cluster.
[0209] Add the selected object to the current list and update the candidate list based on the combination matrix of the new object, i.e., remove objects from the candidate list that may not be combinable with the most recently added object.
[0210] Iterate until there are no more entries in the candidate list.
[0211] After the iteration ends, the combination matrix of all objects in the most recently created cluster can, for example, be set to false because they may no longer be assigned to another cluster.
[0212] Iterative search of other clusters can be started from the start of clustering creation until there are no true entries in the combination matrix.
[0213] Figures 6a to 6g Shows a small-scale example of a two-level JND-based clustering algorithm according to an embodiment.
[0214] Figure 6a Shows the initial distance matrix calculated based on D_PCS.
[0215] Figure 6b Shows the distance matrix, where all entries outside the main diagonal in the distance matrix that are below the selected threshold are marked as candidate pairs that can potentially be combined in the boolean combination matrix. In Figure 6b the selected threshold for marking entries ≤ 1.
[0216] Figure 6c Shows the combination matrix.
[0217] Figure 6d Shows selecting the candidate pair with the smallest entry in the distance matrix from the candidate pairs to initialize the clustering of two objects.
[0218] Figure 6e Shows finding candidate objects in the combination matrix that can be combined with the two objects in the cluster and adding them to the cluster until the candidate object list becomes empty (adding the first object in the shown example). Figure 6e Shows that, for the objects already assigned to the cluster, the corresponding rows / columns are analyzed to determine which other candidate objects can be combined with the objects in the cluster. For example, object 2 can be combined with (1, 3, 5); object 3 can be combined with (1, 2). Therefore, (1, 3, 5) AND (1, 2) = (1). Therefore, adding object 1 to the cluster => the candidate list is empty, and continue to the next cluster.
[0219] Figure 6f Shows the combination matrix, where when the clustering is completed, the entries in rows / columns (1, 2, 3) are invalid.
[0220] Figure 6g Shows the combination matrix, where the next cluster is selected. When the candidate list is empty, the algorithm is completed.
[0221] Enhancements according to a particular embodiment will be considered below.
[0222] First, the time stability according to an embodiment is described.
[0223] For example, the proposed clustering algorithm can be executed frame by frame. In addition to the perceived distance in each frame, the temporal stability of the scene in consecutive frames is also crucial for the perception quality. For example, if the position of an originally stationary object becomes unstable and starts to move around, or an audible "jump" will be introduced into an originally smooth movement, this will also have an impact on the perception quality.
[0224] This results in a trade-off in the optimization objective between minimizing the instantaneous distance metric and temporal stability. For example, a sound source with an initially fixed position can be considered, for example, which is located near the "boundary" between two clusters. Without temporal stability, a small change in the entire scene may cause the assignment of the object's membership to switch between different clusters, resulting in frequent jumps between the centroid positions. This instability may be considered more annoying than a larger but stable movement of the object's position.
[0225] For offline ("file-to-file") applications, such as the encoding or conversion of pre-produced scenes (e.g., audio mixing based on movie objects), some look-ahead or even multi-channel encoding methods can be adopted to optimize temporal stability. However, for real-time capabilities (e.g., for interactive virtual reality (VR) applications), temporal stability may, for example, need to operate with little or no look-ahead to avoid introducing additional latency to the system.
[0226] The temporal stability concept according to some embodiments presented below does not require look-ahead because it relies on smoothing with respect to past frames or applying a hysteresis effect.
[0227] First, consider the concept of using temporal penalties in hierarchical clustering according to an embodiment.
[0228] To avoid the switching of the object's membership assignment in cases where the optimal assignment is ambiguous, in an embodiment, an additional penalty for an object changing its cluster membership is introduced. Thus, the temporal penalty can, for example, be applied to the perceived distance D_perc between objects that previously belonged to different clusters.
[0229] There are multiple options for implementing the temporal penalty:
[0230] For example, a constant offset (e.g., 30 [JND * bits]) can be added to D_perc.
[0231] Alternatively, for example, a multiplicative factor can be applied to D_perc (e.g., 2).
[0232] Alternatively, for example, the (lateral) distance of the object to the previous centroid of another cluster can be used, for example, considering not only the distance between objects but also the actually obtained centroid positions (e.g., considering that two objects close to each other may be just on opposite sides at the boundary between two clusters).
[0233] Alternatively, for example, the (weighted) distance between the previous cluster centroids can be used (e.g., assuming the worst-case scenario where the influence of the object on the centroid position is small, reassigning the membership of the object will cause the object position to move from one centroid to another).
[0234] Now, the DLM smoothing and centroid initialization in GMM-based clustering according to an embodiment are described.
[0235] For a GMM-based clustering method, the hysteresis of spatial hearing can be considered by temporally smoothing the DLM. Thus, the smoothed DLM is calculated as the weighted average of the DLM of the current frame and the previous DLM (using the DLM of the previous frame for short FIR type smoothing, or using the previously smoothed DLM for IIR type smoothing with a longer decay).
[0236] In addition to smoothing the DLM, the EM algorithm for GMM fitting can be initialized, for example, with the centroid positions of the previous frame. To prevent temporal smearing of, for example, scene changes (e.g., cuts in a movie), a threshold for the overall difference in the DLM between two subsequent frames (e.g., SAD, sum of absolute values) can be set to trigger re-initialization of the centroid positions.
[0237] Now, the clustering arrangement optimization according to an embodiment is described.
[0238] In addition to the sound source position, the temporal stability of the combined output signal is also important, especially when the signal is transmitted through a perceptual audio codec. Even if the cluster centroid positions and object assignments remain basically stable in the scene, minor changes in the cluster membership can lead to a rearrangement of the cluster index order (since the cluster index order depends on the lowest member object index in hierarchical clustering, or may be the result of random position initialization in a GMM-based clustering method).
[0239] Such a rearrangement is as Figure 7 shown, where only the middle object moves slightly and is reassigned from the left cluster to the right cluster, but this causes the cluster indices to be swapped. In particular, Figure 7 shows the cluster index rearrangement due to minor changes in the scene according to an embodiment (where Figure 7 the circles pointed by the arrows in Figure 7 are the cluster centroid positions, and the outer circle from which the arrows originate in
[0240] Typically, object signals can be mixed into a continuous waveform, generating one signal (e.g., a transmission channel) for each cluster. When object signals are assigned to different output signals in subsequent frames due to permutations, discontinuities may be introduced, for example, in the output signals. For example, repeated cross-fades between signals may be required, but this may introduce transients in an otherwise continuous signal (which are not actually perceived as transients across the audio scene). These "false" transients may impede the performance of a perceptual audio codec and should therefore be prevented. In addition to affecting the output signals, permutation / exchange of cluster indices may also lead to unnecessarily large and frequent changes in the corresponding centroid positions, which may result in artifacts in the renderer (e.g., when interpolating positions between frames), and may, for example, reduce the efficiency of time-difference encoding of cluster positions. Therefore, measures can be taken, for example, to stabilize the assignment of cluster indices to counteract the permutation effects in consecutive frames.
[0241] Since the assignment of multiple objects to clusters and the centroid positions can vary over time, for example, especially when significant changes occur in the scene, the permutation assignment can be ambiguous and requires an appropriate optimization strategy. However, the optimization goal of the permutation strategy depends on the use case.
[0242] According to an embodiment, for example, a baseline method can be employed to count and minimize the number of objects that are reassigned between clusters.
[0243] Alternatively, to stabilize the position metadata, according to another embodiment, for example, the sum of the absolute or squared distances between the previous and current cluster centroids can be minimized.
[0244] However, an explicit goal is also to stabilize the generated output signal waveform. Therefore, according to an embodiment, signal characteristics can also be considered, for example. By way of illustration, consider, for example, a scenario with two very loud objects and several other almost silent objects. Here, for example, it is preferably to keep the assignment of the loud objects stable (rather than minimizing the number of object reassignments). Put simply, the optimization goal in this case is to keep as much signal energy as possible assigned to the previous positions.
[0245] According to an embodiment, permutation optimization is performed with the aim of stabilizing the energy distribution from objects to clusters. First, the algorithm calculates a matrix that represents how much energy of the objects is redistributed in total between the respective clusters for a given object-to-cluster assignment in two consecutive frames. Based on this energy permutation matrix, a greedy algorithm is used to minimize the energy redistributed between clusters.
[0246] Figure 8 Shows the permutation and optimization of cluster assignments according to an embodiment. Specifically, Figure 8Shows an example of clustering arrangement optimization according to an embodiment for a hypothetical scenario of assigning ten objects to three clusters. The direction of the arrows shows the assignment of objects to clusters (e.g., to cluster indices).
[0247] The cluster membership of the object in the previous frame, corresponding to the previous cluster assignment, as Figure 8 shown in a). The weight of the arrow indicates the assumed energy of the object in the current frame (the energy is also represented numerically in the square on the left).
[0248] Figure 8 b) shows the cluster assignment in the current frame, e.g., generated by a clustering algorithm, where the cluster index order is determined by the lowest member object index. It should be noted that, similar to the previous frame, the three loudest objects are still assigned to three separate clusters respectively. However, since the grouping of the objects has changed, the assignment order has also changed, which will result in a reallocation of the output signal.
[0249] Therefore, according to an embodiment, permutation optimization is performed based on Figure 8 the energy permutation matrix shown in c). The highlighted cells represent the optimized permutation assignment (e.g., the cell in the first row and second column indicates that most of the energy previously found in cluster 1 is now found in cluster 2).
[0250] The resulting permutation-optimized cluster assignment is as Figure 8 shown in d). Thus, in this (intentionally chosen) illustrative example, the assignment of the three loudest objects remains stable compared to the previous frame.
[0251] Specifically, for example, the algorithm according to an embodiment can be implemented as follows:
[0252] Assuming that the number of k clusters generated by the clustering algorithm is constant, for example, a square energy permutation matrix M_Eperm of size k x k with values of 0 can be initialized:
[0253] M_Eperm = zeros(k,k)
[0254] For each object index i, the current energy E(i) can be added, for example, to the matrix entry corresponding to the column where the previous cluster membership index mem_new(i), mem_prev(i) is located and the row where the current cluster membership index is located:
[0255] M_Eperm(mem_new(i),mem_prev(i)) += E(i)
[0256] This can, for example, result in a matrix representing how much energy is redistributed to different indices. If no redistribution occurs, it is reduced to a diagonal matrix. If the grouping of objects remains the same, but there is a permutation of the clustering index order, this will result in a sparse matrix with only k non-zero entries. However, in the general case, when objects from different groups are combined, it is no longer a sparse matrix (especially when many objects are combined into a few clusters, i.e., N >> k).
[0257] For example, the permutation can be optimized by a greedy search in the matrix, which, for example, includes:
[0258] Initialize a permutation vector of length k with values of 0.
[0259] Find the maximum entry in the matrix, obtaining the indices rowMax and colMax.
[0260] Set the permutation vector at the corresponding positions:
[0261] permutation(colMax) = rowMax
[0262] Set the entry in row rowMax and column colMax to 0 (to indicate that the corresponding input index has been assigned and the output index has been obtained):
[0263] M_Eperm(rowMax, :) = 0
[0264] M_Eperm(:, colMax) = 0
[0265] Iterate until all k permutations are assigned.
[0266] For example, the permutation can be applied to the assignment of centroid and membership indices by directly reassigning the centroid indices:
[0267] c_perm(j) = c(permutation(j))
[0268] And by selecting and replacing the corresponding membership indices, for example:
[0269] If (mem(i) == permutation(j)) then mem_perm(i) = j
[0270] In applications where the algorithm does not know the energy of the objects, the algorithm can be adopted, for example, by assuming that the energy of all objects is equal to 1, to minimize the number of objects being redistributed. Thus, the energy permutation matrix M_Eperm is effectively used to count the objects.
[0271] The optimization of the clustering centroid positions according to an embodiment is described below.
[0272] The clustering algorithm produces the membership (or membership probability) of each object and the cluster centroids. Clustering of three-dimensional object positions may result in clusters that include front and back objects, especially when clustering based on perceptual metrics that exploit the limited spatial resolution of human hearing along the cone of confusion and the front-back cone of confusion.
[0273] Assuming that the centroid is calculated as the weighted average of positions initially on the convex hull (e.g., unit sphere or PCS ellipsoid) around the listener, the resulting average position can be inside the sphere / ellipsoid. However, in most applications, the position of the output clusters is also desired to be on the sphere. This is particularly important for speaker playback scenarios, where the sphere corresponds to the convex hull of the speakers, and otherwise internal translations would be required, which are not supported by many renderers (e.g., the VBAP implementation in MPEG-H). Therefore, the resulting cluster positions need to be shifted from the internal centroid positions to the sphere surface.
[0274] As shown in Figure 9, one method is to project the position onto the unit sphere by normalizing its coordinate vector to a length of 1 (and warping from / to PCS coordinates before and after normalization). In particular, Figure 9a ) shows the projection of the centroid in the unit sphere on the horizontal plane ("top view"). Figure 9b ) shows the projection of the centroid in the perceptual coordinate system (PCS) on the horizontal plane.
[0275] However, this will result in perceptually incorrect output positions because positions initially on the same cone of confusion (CoC) are projected outwards. Therefore, when combining sound source positions that are only different in spectral cues perceptually, the left / right attributes as well as the binaural cues will change.
[0276] Therefore, according to an embodiment, a perceptually optimized placement of the cluster output positions can be utilized, where the left / right coordinates of the centroid position are retained and the cluster positions are optimized along the corresponding cones of confusion.
[0277] For example, the optimization along the CoC also depends on the expected playback scenario. For example, different strategies can be selected for binaural rendering compared to speaker rendering. Therefore, multiple options for centroid placement are presented below.
[0278] The normalization of the centroid position in the lateral plane according to an embodiment is described below.
[0279] As Figure 10 shown, the baseline projection method is to project the position outwards by normalizing the position vector in the lateral plane to match the radius of the corresponding circle along the unit sphere.
[0280] Figure 10Shows the centroid of the cone of the confused projection in the lateral plane according to an embodiment (side view). Note how the front and back objects result in an upward projection.
[0281] Calculate the radius of the circle representing the CoC in the lateral plane and normalize the centroid position coordinate vector within the lateral plane to match the radius of the CoC while maintaining the original left / right coordinates.
[0282] When using PCS coordinates, the centroid position is first converted back to unit coordinates.
[0283] (This mode can be advantageous for playback scenarios on a sparse immersive speaker setup, where the middle position will be reproduced, for example, by amplitude translation. In this case, the energy of the object will be redistributed to the front and back by exploiting the properties of the target rendering.)
[0284] Assume axis alignment: c = "front / back" (+1 = front), y = "left / right" (+1 = left), z = "up / down" (+1 = up), and the calculation is as follows:
[0285] Azimuth = sin -1 (y_centroid)
[0286] Radius_coc = cos(azimuth)
[0287] Radius_centroid = sqrt(x_centroid 2 + z_centroid 2 )
[0288] x_projection = x_centroid * radius_coc / radius_centroid
[0289] z_projection = z_centroid * radius_coc / radius_centroid
[0290] The following presents a height-preserving mode according to an embodiment.
[0291] Psychoacoustic experiments have shown that for vertical positioning, the spectral cues for "height" are different from those for "front / back". Or, in other words, the perceived "above" is not the middle of "front" and "back". Therefore, the baseline normalization of the centroid position within the lateral plane of the CoC is not an ideal placement for clustering positions for many applications, such as for binaural rendering (where HRTFs with spectral cues for "height" may be used to reproduce front and back objects at ear level).
[0292] Therefore, a projection mode that preserves height cues is introduced. To preserve the perceptual cues of height perception and resolve front / back confusion, these two dimensions can be considered separately.
[0293] Figure 11Shows the height - maintaining centroid projection (side view) to the CoC in the lateral plane according to an embodiment.
[0294] As Figure 11 shown, the height component can be retained from the centroid position, and this position can be projected onto the cone of confusion, for example, parallel to the horizontal plane. However, this means making a difficult choice between forward or backward projection. When the centroid is close to the transition between front and back (e.g., y_centroid is close to zero), the projection position may jump between front and back, and the energy of the front and back objects changes slightly over time. To stabilize the resulting position, for example, hysteresis can be applied to the sign of the front / back coordinates to prevent the clustering position from switching.
[0295] Note that this mode is particularly suitable for binaural rendering applications. It gives priority to retaining height cues rather than resolving front - back ambiguity. For speaker rendering applications, since slight head movements introduce binaural cues, front - back ambiguity can be easily resolved, but for binaural rendering, for example, only spectral cues are available to resolve front - back ambiguity.
[0296] The spectral matching (EQ matching) mode according to an embodiment is described below.
[0297] The basic idea of the spectral matching mode is based on the fact that the position along the CoC corresponds to the variation of spectral cues. Therefore, the perception of position change depends on the affected frequency region and the actual amount of the spectral content of the signal in the respective frequency region. This means that an object with more energy than other objects in the affected frequency region will be more likely to perceive a position change, and vice versa.
[0298] Thus, the spectral matching method according to an embodiment optimizes the position to minimize the spectral difference between the sums of the signals at the two ears. Another interpretation is to consider the change of the object position in the CoC as multiple equalizer (EQ) curves, and the task is to match the entire spectral envelope, so this mode is also called "equalizer (EQ) matching".
[0299] Since the EQ matching mode takes into account the positions and signal attributes of all member objects in the cluster, not just the centroid position, it may require a higher computational complexity than the centroid projection mode.
[0300] For the setup and calibration of this mode, for example, appropriate frequency bands can be selected, and the average elevation gain curve for each frequency band can be calculated, for example, based on the analysis of the head - related transfer function (HRTF) database (e.g., comparable to the calibration of the PCS). During operation, for example, the signal energy for each band and object can be calculated, and the optimized position can be selected by numerically minimizing the difference of the sum of the weighted energies or by minimizing the ratio (e.g., the sum of the logarithmic differences).
[0301] To improve computational complexity, for example, principal component analysis can be utilized to derive a limited number of "eigen spectra" along the position of the CoC. This can be interpreted as a preset equalizer curve for the entire spectrum that adjusts the intensity based on position, rather than determining separate factors for each position and frequency band. For example, these can be associated with the spectral envelopes of individual signals in order to generate a minimized lower-dimensional representation with lower computational complexity.
[0302] The mixing and processing of output signals according to some embodiments are described below.
[0303] After determining the cluster membership and centroid positions, the target signals are combined in order to generate an output signal for each output cluster. The method can be to sum the signals of all members within a cluster. However, to avoid audible artifacts and optimize perceptual quality, further precautions and improvements need to be considered:
[0304] Since the cluster assignment is determined frame by frame, the membership can change from one frame to the next. For example, when the membership changes, cross-fading can be applied to prevent audible clicks due to signal discontinuities.
[0305] There may be correlations between the signals of the objects within a cluster, which may, for example, result in positive or negative interference in the downmixed signal. To achieve energy-preserving downmixing, signal correlations can be considered, for example.
[0306] Cluster algorithms such as GMM-based clustering not only produce memberships but also membership probabilities. For example, an object with fuzzy membership can be mixed into more than one cluster to implement a "soft" clustering method.
[0307] The cross-fading according to an embodiment is described below.
[0308] When the membership of an object changes between subsequent frames, according to an embodiment, the downmixed signal can be cross-faded, for example, to prevent a hard signal cut that results in an audible click due to signal discontinuity.
[0309] To not require additional look-ahead for the cluster assignment in the next frame, for example, the cross-fading can be performed at the start of the current frame.
[0310] To avoid unnecessary cross-fading, the cluster memberships of the previous and current memberships of each object can be saved and compared. The cross-fading is applied if and only if the membership changes.
[0311] For cross-fading, for example, complementary window functions can be applied to fade in the target signal in the newly assigned cluster signal and fade it out from the previously assigned output signal. For example, the cross-fading can be chosen to preserve energy, so a sine-shaped window can be used, for example. In an embodiment, the cross-fading duration can be, for example, long enough to prevent audible clicks, but can also be as short as possible, for example, to prevent an audible delay in the source position.
[0312] Thus, in a particular embodiment, for example, a cross-fading length of 128 samples (about 2.7 milliseconds at a 48 kHz sampling rate) can be employed.
[0313] The relevance-aware downmix according to some embodiments is described below.
[0314] The basic assumption of clustering for object-based audio is that audio objects represent separate and uncorrelated sound sources, which are typically rendered as separate point sources by an object-based audio renderer (e.g., Vector Base Amplitude Panning (VBAP)). However, there are cases where this assumption is violated, for example, when two or more object signals are correlated. For example, when calculating the downmix signal of correlated object signals within a cluster, this can result in positive or negative interference. Thus, for example, additional precautions can be taken when calculating the downmix in a scenario that is expected to contain correlated objects. It should be noted that strong correlations between sound sources can also lead to the perception of phantom sound sources. However, this also involves the placement of the resulting cluster positions and is thus not discussed within the scope of signal downmixing.
[0315] In general, a small amount of correlation may randomly occur between audio signals that are initially created / recorded independently (when the signals are not explicitly created as orthogonal (e.g., independent random noise)), although this is usually not significant.
[0316] However, for example, according to the production paradigm used to create an object-based sound scene, more substantial correlations between signals can be introduced.
[0317] For example, in some cases, objects are created from signals from two or more channels of a stereo or multi-microphone recording in a sound scene. Another perspective is that an object-based audio scene may include an "unlabeled channel bed", for example, recordings or works that were initially produced for speaker playback and have been reused and placed at object positions that roughly correspond to the intended speaker positions. This is usually known at the time of production, but the clustering algorithm used may be unknown, depending on the metadata transmission format. Similarly, but to a lesser extent, correlations can also occur when objects are captured from multiple spot microphones within a physical scene, such as different actors or instruments on a stage. This is not typically considered a channel-based recording, but crosstalk can still occur between the individual microphone signals.
[0318] Moreover, due to content correlation, for example, when multiple musical instruments follow the same melody line, signal correlation may occur even for individually recorded or synthesized signals.
[0319] In some different situations, the correlation between signals can be predicted during the production stage and marked with appropriate metadata. However, when the correlation is introduced more accidentally, no additional metadata can be obtained. Therefore, the object clustering algorithm cannot rely solely on external information and needs to be able to detect and handle the correlation when downmixing object signals in the absence of available metadata.
[0320] When there is correlation between object signals combined within a cluster signal, for example, the amplitudes of the signals rather than the energies of the signals can be summed, which can lead to an increase or loss of signal energy and thus a difference in perceived loudness. According to some embodiments, in order to maintain the loudness perception of the original scene, for example, correlation-aware downmixing can be applied.
[0321] However, it should be recognized that the perceived effect of the correlation between object signals also depends on the playback scenario and rendering algorithm used.
[0322] According to an embodiment, for example, energy summation can be performed. In an ideal playback-agnostic scenario, the objects represent physical sound sources at different spatial positions. Here, the actual sound waves are physically superimposed at the reproduction environment and the ears. Since a typical listening environment is not anechoic (e.g., a BS1116 room), especially for higher frequencies, due to different propagation paths (i.e., room reverberation and HRTF), the correlation between the signals arriving at the ears is reduced. As a simplified model, for this case, for example, energy summation can be assumed. In the applied playback scenario, this may be the case for example in binaural headphone reproduction, where different binaural room impulse responses (BRIRs) can be applied for different sound source positions. For speaker playback, for example, for the case where the distance between the objects is large enough relative to the speaker placement so that the objects are reproduced by different speakers, this can be assumed.
[0323] In an embodiment, for example, amplitude summation can be performed. For amplitude-based panning rendering (e.g., VBAP) on a relatively sparse speaker setup (e.g., a typical home theater setup), different sound source positions can be panned and reproduced, for example, between the same speaker pair. In this case, for example, the signal amplitudes can be added in the rendering algorithm, resulting in a correlation-dependent behavior of the energy sum.
[0324] The renderer-independent object clustering algorithm will assume the ideal case of independent sound sources and thus perform energy summation. However, the purpose of the object clustering algorithm is usually to be as close as possible to the reference rendering on a given rendering in a given target playback scenario. This means that the aim is to replicate the energy or amplitude summation characteristics of the target rendering and playback, regardless of whether the behavior of the reference is deliberately designed or not.
[0325] Based on the target use case, two downmixing modes can be selected:
[0326] According to the first downmixing mode, for example, direct signal summation can be performed. If it is assumed that the object signals are uncorrelated and / or if the target playback scenario is played back with amplitude-shifted speakers, then the object signals are directly summed into the clustered output signal. This mode also avoids the additional computational complexity of correlation analysis and is therefore more preferable for real-time applications.
[0327] According to the second downmixing mode, for example, correlation-aware signal summation can be performed. If the aim is energy-preserving summation and signal correlation is expected, then energy-preserving weighting is applied.
[0328] To achieve the preservation of the overall scene energy, the method is to calculate the energy of all objects before mixing, calculate the final energy of the downmixed signal, and apply a gain correction factor to the downmixed signal. However, the drawback of this simple method is that not all objects in the cluster have to be associated in the same way. Therefore, this global energy gain correction also reduces the energy of uncorrelated signals, resulting in overrepresentation of correlated signals in the final mix.
[0329] Therefore, according to an embodiment, for example, an advanced downmixing algorithm based on signal correlation can be employed. For example, the cross-correlation matrix between all objects within the cluster can be calculated. Based on this, for example, the downmix gain correction factor for each individual object can be calculated. Therefore, for example, the overall energy relationship between correlated and uncorrelated objects can be preserved.
[0330] Specifically, in a particular embodiment, for example, the downmixing coefficients can be calculated, where the calculation can, for example, include:
[0331] For example, the cross-correlation matrix C between all member objects of the cluster can be calculated as the dot product of signal samples. In addition, for example, the normalized correlation matrix C_norm can be calculated therefrom, including the corresponding Pearson correlation coefficients. (Therefore, the main diagonal of C corresponds to the signal energy, while the main diagonal of C_norm is all equal to 1).
[0332] For the purpose of energy-conserving downmixing, for example, only medium to high correlations may be of interest. Low and negligible correlations due to random effects may even impair the stability of the downmixing algorithm and thus affect the perceived quality. Therefore, for example, a threshold may be applied by setting all entries in C to zero where the absolute value of C_norm is below 0.5.
[0333] Optionally, for example, the correlation may be restricted to only positive correlations, so that only the energy increase due to correlation is compensated, but no enhancement is applied in case of signal cancellation (e.g., in applications with insufficient headroom in order to avoid clipping of the signal prior to downmixing).
[0334] For each object, for example, the energy weight factor w_En may be calculated as the ratio of the sum of the corresponding row in the correlation matrix to the signal energy.
[0335]
[0336] In other words, this factor approximates how much the energy of each object is boosted due to correlation with other signals. If the correlations of all signals are below the threshold, only non-zero entries exist on the main diagonal and all factors are 1.
[0337] For example, the respective weighting factor w_A for scaling the signal amplitude may be calculated as the square root of the inverse energy weight:
[0338]
[0339] The factor w_A is applied as a scalar multiplier to the signal before adding in the time domain.
[0340] In a typical implementation, for example, the weighting factor w_A may be restricted, e.g., with a maximum value of 2, to prevent an overly large boost factor in case of strong signal cancellation (or correspondingly, e.g., |w_En| may be restricted to avoid division by zero, e.g., with a minimum value set to 0.25). It should be noted that in case of signal cancellation, an overly large weighting factor would rather lead to an enhancement of the residual background noise than to the reconstruction of the cancelled signal components.
[0341] To prevent enhancement of signal cancellation: strong negative correlations are detected by an appropriate threshold (such as C(i,j) < -0.8) and the weighting factor of one of the negatively correlated signals is set to zero (e.g., only considering one of the cancelled signals), or a negative weight is applied. It should be noted that for negative correlations in the playback scenario of a single point source in a non-anechoic environment, due to decorrelation caused by room reverberation etc., the signals will not completely cancel at the listener's position. Strong signal cancellation may occur in sparse loudspeaker rendering.
[0342] As a further enhancement, the correlation analysis and addition can be applied, for example, in the frequency domain, e.g., using a short-time Fourier transform (STFT) filter bank with appropriate frequency band grouping.
[0343] The consideration of distance-based gain according to an embodiment is described below.
[0344] Depending on the target use case, the rendering algorithm can also consider the distance of the reproduced sound sources. A basic implementation is to apply distance-based gain to compensate for the radial distance between the listener and the sound source. If it is known that the target renderer will apply distance-dependent gain, for example, this can be compensated for during downmix clustering to avoid perceivable loudness differences in the playback scenario.
[0345] If the clustering algorithm knows the actual distance gain function of the renderer, a straightforward solution is to calculate the gain at the original sound source positions and the integrated cluster positions and compensate for the resulting gain differences before downmixing.
[0346] As a generalized, computationally efficient clustering method based on PCS, for example, the radial distance component from PCS can be utilized, which may have been modeled, for example, after the distance-dependent gain difference. Thus, the difference in the radial distance component between the object and the cluster position can be calculated directly, and this difference can be used, for example, as the gain difference, e.g., in dB.
[0347] Further embodiments are described below.
[0348] According to a first embodiment, for example, an object-based audio scene can be clustered based on a perception-based model relative to the listener.
[0349] In a second embodiment, the clustering algorithm of the first embodiment can be based on, for example, a perceptual distance metric / perceptual distortion metric (PDM).
[0350] According to a first variant of the second embodiment, for example, the clustering of objects in a given maximum PDM link can be identified and combined, e.g., the identification and combination of all pairs below the just noticeable difference.
[0351] According to a second variant of the second embodiment, for example, clustering can be performed by iterative aggregation of the closest objects in the PDM until the target number of clusters is reached, or until a given maximum of the distortion metric is exceeded.
[0352] In a third embodiment, the clustering algorithm of the first embodiment can be based on, for example, 3D-DLM similarity.
[0353] According to a first variant of the third embodiment, for example, the reconstruction of the 3D-DLM of the original scene can be performed by fitting a Gaussian mixture model (GMM).
[0354] According to the second variant of the third embodiment, for example, an enhanced Expectation Maximization (EM) algorithm for GMM fitting of weighted data points on an arbitrary grid can be employed.
[0355] In the fourth embodiment, for example, one or more enhancements for time stability of object-based clustering for the first to third embodiments can be performed.
[0356] According to the first variant of the fourth embodiment, for example, temporal smoothing and penalty factors in the perceptual distance metric can be implemented.
[0357] According to the second variant of the fourth embodiment, for example, optimization of the clustering assignment arrangement based on energy distribution can be performed.
[0358] According to the third variant of the fourth embodiment, for example, the resulting cluster centroid positions can be stabilized by hysteresis.
[0359] In the fifth embodiment, for example, perceptual optimization of the centroid positions resulting from clustering of one of the first to third embodiments can be performed.
[0360] According to the sixth embodiment, for example, optimization of the clustering assignment and centroid positions based on spectral matching (EQ matching of HRTF) for the clustering of the first embodiment can be performed.
[0361] In the seventh embodiment, for example, signal processing of the combination of audio objects resulting from clustering of the first embodiment can be performed.
[0362] According to the first variant of the seventh embodiment, for example, cross-fading can be performed to prevent signal discontinuities during object-to-cluster membership reallocation.
[0363] According to the second variant of the seventh embodiment, for example, signal correlation can be considered to achieve energy preservation.
[0364] According to the third variant of the seventh embodiment, for example, adjustment of the distance-based gain can be performed.
[0365] According to the fourth variant of the seventh embodiment, for example, equalization can be performed to compensate for perceptual differences due to spectral cues.
[0366] Although certain aspects are described in the context of an apparatus, it is clear that these aspects also represent a description of a corresponding method, where a module or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or use) a hardware device, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0367] According to the requirements of certain implementations, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware, or at least partially in software. The implementation may be performed using a digital storage medium, for example, a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory, which has electronically readable control signals stored thereon, and these electronically readable control signals cooperate (or are capable of cooperating) with a programmable computer system to perform the corresponding method. Therefore, the digital storage medium may be computer-readable.
[0368] Some embodiments according to the present invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0369] Generally, embodiments of the present invention may be implemented as a computer program product having program code that, when the computer program product runs on a computer, can be used to perform one of the above methods. The program code may be stored on a machine-readable carrier.
[0370] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the methods described herein.
[0371] In other words, therefore, embodiments of the method of the present invention are computer programs having program code for performing one of the methods described herein when the computer program is executed on a computer.
[0372] Therefore, another embodiment of the method of the present invention is a data carrier (or a digital storage medium, or a computer-readable medium) including a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is generally tangible and / or non-transitory.
[0373] Therefore, another embodiment of the method of the present invention is a data stream or a signal sequence representing a computer program for performing one of the methods described herein. For example, the data stream or signal sequence may be configured to be transmitted through a data communication connection, for example, transmitted through the Internet.
[0374] Another embodiment includes a processing device, e.g., a computer or a programmable logic device, configured or adapted to execute one of the methods described herein.
[0375] Another embodiment includes a computer having installed thereon a computer program for executing one of the methods described herein.
[0376] Another embodiment according to the present invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for executing one of the methods described herein to a receiver. For example, the receiver may be a computer, a mobile device, a storage device, etc. For example, the apparatus or system may include a file server for transmitting the computer program to the receiver.
[0377] In some embodiments, a programmable logic device (e.g., a field programmable array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable array may cooperate with a microprocessor to execute one of the methods described herein. Generally, the methods are preferably executed by any hardware device.
[0378] The apparatus described herein may be implemented using a hardware device, or a computer, or a combination of a hardware device and a computer.
[0379] The methods described herein may be executed using a hardware device, or a computer, or a combination of a hardware device and a computer.
[0380] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the above arrangements and details will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the following claims, rather than by the specific details shown by the description and explanation of the above embodiments.
Claims
1. An apparatus (100), comprising: an input interface (110) for receiving information about three or more audio objects; and a clustering generator (120) for generating two or more audio object clusters by associating each of the three or more audio objects with at least one of two or more audio object clusters, such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster, wherein the clustering generator (120) is configured to generate the two or more audio object clusters according to a perception-based model.
2. The apparatus (100) according to claim 1, Among them, wherein the clustering generator (120) is configured to generate the two or more audio object clusters according to the perception-based model by generating the two or more audio object clusters according to at least one of a perceptual distance metric, a directional loudness map, a perceptual coordinate system, and a spatial masking model.
3. The apparatus (100) according to claim 2, Among them, wherein the clustering generator (120) is configured to generate the two or more audio object clusters according to the perceptual distance metric by, for a pair of two audio objects among the three or more audio objects, determining whether the perceptual distance between the two audio objects is less than or equal to a threshold according to the perceptual distance metric, and by, if the perceptual distance is less than or equal to the threshold, associating the two audio objects with the same audio object cluster among the two or more audio object clusters.
4. The apparatus (100) according to claim 2, Among them, wherein the clustering generator (120) is configured to generate the two or more audio object clusters according to the perceptual distance metric by iteratively associating the two perceptually closest audio objects among the three or more audio objects until a preset target number of audio object clusters is reached or until a preset maximum perceptual distance according to the perceptual distance metric is exceeded.
5. The apparatus (100) according to any one of the above claims, Among them, wherein the clustering generator (120) is configured to generate the two or more audio object clusters according to a three-dimensional directional loudness map.
6. The apparatus (100) according to claim 5, Among them, wherein the clustering generator (120) is configured to generate the two or more audio object clusters by employing a Gaussian mixture model; and wherein the clustering generator (120) is configured to determine the two or more audio object clusters by determining the components of the Gaussian mixture model so as to approximate the three-dimensional directional loudness map.
7. The apparatus (100) according to claim 5, Among them, wherein the clustering generator (120) is configured to generate the two or more audio object clusters by employing a Gaussian mixture model; and Wherein, the clustering generator (120) is configured to fit weighted data points on any grid of the Gaussian mixture model by using the expectation maximization algorithm to determine the two or more audio object clusters.
8. The apparatus (100) according to any one of the above claims, Among them, the clustering generator (120) is configured to perform perceptual optimization on the centroid positions of the generated clusters; and / or, wherein, the clustering generator (120) is configured to optimize the cluster assignment and the centroid positions according to the spectral matching of the two or more audio object clusters.
9. The apparatus (100) according to any one of the above claims, Among them, the clustering generator (120) is configured to generate the two or more audio object clusters as a first plurality of audio object clusters by creating an association between each of the three or more audio objects and at least one of the two or more audio object clusters; and wherein, the clustering generator (120) is configured to generate a second plurality of two or more audio object clusters such that at least one audio object among the three or more audio objects is associated with a different audio object cluster in the second plurality of audio object clusters compared to the audio object cluster in the first plurality of audio object clusters associated with the at least one audio object.
10. The apparatus (100) according to claim 9, Among them, the clustering generator (120) is configured to generate the second plurality of two or more audio object clusters according to temporal smoothing and / or according to one or more penalty factors in the perceptual distance metric.
11. The apparatus (100) according to claim 9 or 10, Among them, the clustering generator (120) is configured to generate the second plurality of two or more audio object clusters by optimizing the cluster assignment arrangement according to the energy distribution of the three or more audio objects.
12. The apparatus (100) according to any one of claims 9 to 11, Among them, the clustering generator (120) is configured to generate the second plurality of two or more audio object clusters by stabilizing the generated cluster centroid positions via a hysteresis effect.
13. The apparatus (100) according to any one of claims 9 to 12, Among them, the clustering generator (120) is configured to generate the second plurality of two or more audio object clusters by performing perceptual optimization on the centroid positions of the clusters generating the first plurality of two or more audio object clusters; and / or, wherein, the clustering generator (120) is configured to generate the second plurality of two or more audio object clusters by optimizing the cluster assignment and the centroid positions according to the spectral matching of the first plurality of audio object clusters.
14. The apparatus (100) according to any one of the above claims, Among them, The clustering generator (120) is configured to perform signal processing for each audio object cluster associated with at least two of the three or more audio objects by combining the audio object signals of each audio object associated with the audio object cluster.
15. The apparatus (100) according to claim 14, Among them, The clustering generator (120) is configured to perform at least one of the following: Cross-fading to prevent signal discontinuities when object-to-cluster membership is re-assigned; Considering signal correlation to achieve energy conservation; Adjusting distance-based gain; or Equalizing to compensate for perceptual differences caused by spectral cues.
16. The apparatus (100) according to any of the above claims, Among them, The clustering generator (120) is configured to generate the two or more audio object clusters according to the actual or assumed position of the listener.
17. The apparatus (100) according to any of the above claims, Among them, The clustering generator (120) is configured to, for each audio object cluster of the two or more audio object clusters, determine one or more attributes of the audio object cluster according to one or more attributes of the audio objects among the three or more audio objects associated with the audio object cluster, where the one or more attributes include at least one of the following: The audio signal associated with the audio object cluster; or The position associated with the audio object cluster.
18. The apparatus (100) according to any of the above claims, Among them, The apparatus (100) further includes an encoding unit for generating encoded information encoding information about the two or more audio object clusters.
19. A system, comprising: The apparatus (100) according to claim 18; A decoding unit (210) for decoding the encoded information to obtain information about two or more audio object clusters; And A signal generator (220) for generating two or more audio output signals according to the information about the two or more audio object clusters.
20. A decoder (200), comprising: A decoding unit (210) for decoding the encoded information to obtain information about two or more audio object clusters, where the two or more audio object clusters are generated by associating each of the three or more audio objects with at least one of the two or more audio object clusters, such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster, where the two or more audio object clusters are generated according to a perception-based model; and A signal generator (220) for generating two or more audio output signals according to the information about the two or more audio object clusters.
21. A method, comprising: Receiving information about three or more audio objects; and generating the two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters, such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster, wherein the two or more audio object clusters are generated according to a perception-based model.
22. A method, comprising: decoding encoded information to obtain information about two or more audio object clusters, wherein the two or more audio object clusters are generated by associating each of three or more audio objects with at least one of the two or more audio object clusters, such that for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and such that for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster, wherein the two or more audio object clusters are generated according to a perception-based model; and generating two or more audio output signals based on the information about the two or more audio object clusters.
23. A computer program for implementing the method of claim 21 or 22 when executed on a computer or a signal processor.