Apparatus and method for using a perceptually based distance metric for spatial audio

A perceptual distance metric using a perceptual coordinate system and spatial masking model optimizes object-based audio clustering, addressing the need for efficient perceptual impact estimation in real-time spatial audio applications.

JP2025533618APending Publication Date: 2025-10-07FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025518554
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-28
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

Existing audio technologies lack a computationally efficient method to estimate the perceptual impact of spatial audio changes in real-time applications like virtual reality, failing to optimize object-based audio clustering for immersive sound experiences.

Method used

A perceptual distance metric is introduced, utilizing a perceptual coordinate system, 3D directional loudness map, and spatial masking model to predict and optimize perceptual differences in spatial audio, enabling efficient object clustering and output signal generation.

Benefits of technology

This approach allows for computationally efficient prediction of perceptual differences in spatial audio, optimizing object-based audio clustering and maintaining high perceptual quality in real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533618000001_ABST
    Figure 2025533618000001_ABST
Patent Text Reader

Abstract

According to one embodiment, an apparatus (100) is provided. The apparatus comprises an input interface (110) for receiving a plurality of audio objects of an audio sound scene. The apparatus (100) further comprises a processor (120). Each of the plurality of audio objects represents a sound source that is different from any other sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same sound source at different locations. The processor (120) is configured to obtain information about a perceptual difference between two audio objects of the plurality of audio objects in response to a distance metric, the distance metric representing a perceptual difference in spatial characteristics of the audio sound scene. And / or the processor (120) is configured to process the plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects in response to the distance metric.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] The present invention relates to an apparatus and method for using a perceptually based distance (distortion) metric for spatial audio.

[0002] Modern audio playback systems enable an immersive three-dimensional (3D) sound experience.

[0003] One common format for 3D sound reproduction is channel-based audio, where individual channels associated with defined speaker positions are created by multi-microphone recording or studio-based production. Another common format for 3D sound reproduction is object-based audio, which utilizes so-called audio objects, which are placed in the listening room by the producer and converted by the rendering system into speaker or headphone signals for playback. Object-based audio allows for a high degree of flexibility in terms of sound scene design and playback. It should be noted that channel-based audio can be considered a special case of object-based audio, where sound sources (=objects) are positioned at fixed positions corresponding to defined speaker positions.

[0004] To increase the efficiency of transmission and storage of object-based immersive sound scenes, as well as to reduce the computational requirements of real-time rendering, it is beneficial or even necessary to reduce or limit the number of audio objects. This is achieved by identifying groups or clusters of adjacent audio objects and combining them into a smaller number of sound sources. This process is called object clustering or object consolidation.

[0005] It has been shown in the literature that the localization accuracy of human hearing is limited and dependent on sound source position (e.g., horizontal localization is more accurate than vertical localization), and that auditory masking effects can be observed between spatially distributed sound sources. By exploiting these limitations in localization accuracy in human hearing and auditory masking effects for object clustering, a significant reduction in the number of audio objects can be achieved while maintaining high perceptual quality.

[0006] Auditory masking and localization models are known in the art.

[0007] Directional loudness maps (DLMs) are presented in "Frequency-domain source identification and manipulation in stereo mixes for enhancement, suppression, and re-panning applications" by C. Avendano, in the 2003 IEEE Workshop on Applications on Signal Processing to Audio, and "Objective Assessment of Spatial Audio Quality using Directional Loudness Maps" by P. Delgado and J. Herre, in Proc. 2019 IEEE ICASSP.

[0008] Object clustering algorithms are presented in "Optimization of Sound Spatialization Resource Management through Clustering" by J. Herder, The Journal of Three Dimensional Images, 1999; "Perceptual Audio Rendering of Complex Virtual Environments" by Nicolas Tsingos, Emmanuel Gallo, and George Drettakis, SIGGRAPH, 2004; and "Spatial Coding of Complex Object-Based Program Material" by Jeroen Breebaart, Giulio Gengarle, Lie Lu, Toni Mateos, Heiko Purnhagen, and Nicolas Tsingos, JAES, Volume 67, Issue 7 / 8, pp. 486-497, July 2019.

[0009] The state-of-the-art includes psychoacoustic models for localization cues, masking, and saliency, but they do not provide a way to estimate the perceptual impact of changes to the spatial properties of individual sound sources in a scene relative to the listener's position in a computationally efficient representation suitable for real-time applications such as audio for virtual reality (VR). Summary of the Invention [Problem to be solved by the invention]

[0010] It is an object of the present invention to provide an improved concept for distance metrics for spatial audio. [Means for solving the problem]

[0011] The object of the present invention is solved by an apparatus according to claim 1, a decoder according to claim 20, a method according to claim 21, a method according to claim 22 and a computer program according to claim 23.

[0012] According to one embodiment, an apparatus is provided. The apparatus comprises an input interface for receiving a plurality of audio objects of an audio sound scene. The apparatus further comprises a processor. Each of the plurality of audio objects represents a (real or virtual) sound source that is different from any other (real or virtual) sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same (real or virtual) sound source at different locations. The processor is configured to obtain information regarding a perceptual difference between two audio objects of the plurality of audio objects in response to a distance metric, the distance metric representing a perceptual difference in spatial characteristics of the audio sound scene. And / or the processor is configured to process the plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects in response to the distance metric.

[0013] Further, a decoder according to an embodiment is provided, comprising a decoding unit and a signal generator, wherein each of a plurality of audio objects of an audio sound scene represents a (real or virtual) sound source different from any other (real or virtual) sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same (real or virtual) sound source at different locations, the decoding unit is configured to decode the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects, the plurality of audio object clusters or the plurality of processed audio objects being dependent on the plurality of audio objects of the audio sound scene and on a distance metric representing a perceptual difference in spatial characteristics of the audio sound scene, and the signal generator is configured to generate two or more audio output signals in response to the plurality of audio object clusters or in response to the plurality of processed audio objects. and / or the decoding unit is configured to decode the encoded information to obtain a plurality of audio objects of the audio sound scene and to obtain information regarding a perceptual difference between two of the plurality of audio objects, the perceptual difference being dependent on a distance metric; and the signal generator is configured to generate two or more audio output signals in response to the plurality of audio objects and in response to the perceptual difference between said two audio objects. Further, a method according to one embodiment is provided, the method comprising: receiving information about a plurality of audio objects in an audio sound scene; and obtaining information regarding a perceptual difference between two of the plurality of audio objects as a function of the distance metric.

[0014] Each of the plurality of audio objects represents a (real or virtual) sound source that is different from any other (real or virtual) sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same (real or virtual) sound source at different locations. The distance metric represents a perceptual difference in spatial characteristics of the audio sound scene and / or processing the plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects according to the distance metric.

[0015] Further, according to another embodiment, there is provided a method, wherein each of a plurality of audio objects of an audio sound scene represents a (real or virtual) sound source that is different from any other (real or virtual) sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same (real or virtual) sound source at different locations, the method comprising: decoding the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects, the plurality of audio object clusters or the plurality of processed audio objects being dependent on a plurality of audio objects of the audio sound scene and dependent on a distance metric representing a perceptual difference in spatial characteristics of the audio sound scene; and generating two or more audio output signals in response to the plurality of audio object clusters or in response to the plurality of processed audio objects; and / or obtaining a plurality of audio objects of an audio sound scene, decoding the encoded information to obtain information regarding a perceptual difference between two of the plurality of audio objects, the perceptual difference being dependent on a distance metric; and generating two or more audio output signals in response to the plurality of audio objects and in response to the perceptual difference between said two audio objects. Further provided are computer programs, each configured to implement one of the above methods when run on a computer or signal processor.

[0016] To predict the perceptual effects of localization changes in a sound scene, some embodiments provide a perceptual model that represents perceptual differences in a computationally efficient manner. This model can be utilized to optimize the perceptual quality of object-based audio clustering algorithms, as well as for objective measures that quantify the perceptual differences between different representations of a sound scene.

[0017] The perceptual distance metric according to some embodiments answers questions such as: How perceptible is it if the position of a sound source changes? How perceptible is the difference between two different sound scene representations? How important is a given sound source in the overall sound scene (and how noticeable would removing it be?). A psychoacoustic model according to some embodiments may include, for example, one or more of the following components that correspond to different aspects of human perception: a perceptual coordinate system, a 3D directional loudness map, a spatial masking model, and a perceptual distance metric.

[0018] According to some embodiments, a perceptual coordinate system (PCS) is provided. Sound source localization accuracy in humans varies with spatial direction. To represent this in a computationally efficient manner, a perceptual coordinate system (PCS) is introduced. To obtain this PCS, spatial locations are distorted to accommodate the heterogeneous characteristics of localization accuracy. This allows distances in the PCS to correspond to "perceived distances" between locations, e.g., just noticeable differences (JNDs), rather than physical distances. This principle is similar to the use of psychoacoustic frequency scales, e.g., the Bark scale or the ERB scale (equivalent rectangular bandwidth scale), in perceptual audio coding.

[0019] According to some embodiments, a 3D directional loudness map (3D-DLM) is provided. The underlying concept of a directional loudness map (DLM) is to find a representation of the "perceived loudness coming from a given direction." This concept has already been presented as a one-dimensional approach to represent binaural localization in binaural DLM (Delgado et al., 2019). This concept is now extended to three-dimensional (3D) localization by creating a 3D-DLM on a surface surrounding the listener to uniquely represent the perceived loudness as a function of the angle of incidence relative to the listener. Note that while binaural DLMs have been obtained by analyzing signals at the ears, 3D-DLMs are synthesized for object-based audio by utilizing a priori known sound source locations and signal characteristics.

[0020] In some embodiments, a spatial masking model (SMM) is provided. The monophonic time-frequency auditory masking model is a fundamental element of perceptual audio coding and is often augmented by a binaural (un)masking model to improve stereo coding. The spatial masking model extends this concept of immersive audio to incorporate and exploit masking effects between arbitrary sound source positions in 3D. According to some embodiments, a perceptual distance metric is provided. Note that the above components can be combined, for example, to obtain a perceptual-based distance metric between spatially distributed audio sources. These can be utilized in various applications, for example, to control bit distribution in perceptual audio coders, to obtain objective quality measures, as a cost function in object clustering algorithms, etc.

[0021] In the following, embodiments of the invention will be explained in more detail with reference to the drawings. [Brief explanation of the drawings]

[0022] [Figure 1] 1 illustrates an apparatus according to one embodiment. [Figure 2] 1 illustrates a decoder according to one embodiment. [Figure 3] 1 illustrates a system according to one embodiment. [Figure 4] 1 illustrates a two-dimensional example of perceptual coordinate system coordinates being distorted according to one embodiment. [Figure 5] 10 illustrates perceptual coordinates obtained via multidimensional scaling of modeled differences in the CIPIC HRTF database according to one embodiment. [Figure 6] 1 illustrates a polynomial model-based perceptual coordinate system according to one embodiment. [Figure 7] 1 illustrates an ellipsoid model-based perceptual coordinate system according to one embodiment. [Figure 8] 1 illustrates an example of synthesis of a one-dimensional directional volume map based on known object positions and volumes, according to one embodiment. [Figure 9] 1 illustrates an example of a 3D directional volume map synthesized from known sound source positions according to an embodiment. [Figure 10] 1A and 1B illustrate different sampling methods of a unit spherical grid according to an embodiment, where (a) shows azimuth / elevation sampling and (b) shows an ICO sphere. [Figure 11] 1 illustrates a masking model calculation in perceptual coordinates according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0023] FIG. 1 illustrates an apparatus 100 according to one embodiment. According to one embodiment, an apparatus 100 is provided. The device comprises an input interface 110 for receiving a plurality of audio objects of an audio sound scene. Furthermore, the device 100 comprises a processor 120 . Each of the plurality of audio objects represents a real or virtual sound source that is different from any other real or virtual sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same real or virtual sound source at different positions. For example, the same real or virtual sound source may be considered at different positions because different points in time are considered. Alternatively, the same real or virtual sound source may be considered at different positions because, for example, a position before position quantization can be compared with a position after position quantization.

[0024] The processor 120 is configured to obtain information regarding a perceptual difference between two of the plurality of audio objects as a function of a distance metric, the distance metric representing a perceptual difference in spatial properties of the audio sound scene. And / or the processor 120 is configured to process the plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects depending on the distance metric. According to one embodiment, the audio sound scene may for example be a three-dimensional audio sound scene.

[0025] In one embodiment, the processor 120 may be configured to obtain information about a perceptual difference between two audio objects, for example, according to a perceptual coordinate system, and / or the processor 120 may be configured to process a plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects, for example, according to a perceptual coordinate system, where the distance in the perceptual coordinate system represents a perceptible localization difference. According to an embodiment, the processor 120 may be configured to obtain information about a perceptual difference between two objects, for example, according to an inverse mapping function, and / or the processor 120 may be configured to process a plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects, for example, according to an inverse mapping function. Further, the processor 120 may be configured to convert coordinates of the physical coordinate system into coordinates of the perceptual coordinate system, for example, using an inverse mapping function. In one embodiment, the inverse mapping function may be determined by, for example, head-related transfer function data.

[0026] According to one embodiment, the processor 120 may be configured to obtain information about the perceptual difference between two audio objects, for example, according to a spatial masking model for spatially distributed sound sources, and / or the processor 120 may be configured to process multiple audio objects to obtain multiple audio object clusters or multiple processed audio objects, for example, according to a spatial masking model. The spatial masking model may depend, for example, on a masking threshold. The processor 120 may be configured to determine the masking threshold, for example, according to an attenuation function and one or more distances in a perceptual coordinate system. In one embodiment, the processor 120 may be configured to determine the masking threshold as a function of, for example, a Gaussian-shaped decay function as the decay function and as a function of an offset of the minimum masking. According to one embodiment, the processor 120 may be configured to, for example, identify one or more inaudible audio objects of the plurality of audio objects.

[0027] In an embodiment, processor 120 may be configured to obtain information about a perceptual difference between two audio objects, for example, depending on a perceptual distortion metric, and / or processor 120 may be configured to process multiple audio objects to obtain multiple audio object clusters or multiple processed audio objects, for example, depending on a perceptual distortion metric. Further, processor 120 may be configured to determine the perceptual distortion metric, for example, depending on a distance in a perceptual coordinate system and depending on a spatial masking model. According to one embodiment, the processor 120 may be configured to determine the perceptual distortion metric as a function of, for example, the perceptual entropy of one or more of the plurality of audio objects. In one embodiment, the processor 120 may be configured to determine the perceptual distortion metric, for example, in response to a first distance between a first one of two audio objects of the plurality of audio objects and the center of gravity of the two audio objects, and in response to a second distance between a second one of the two audio objects and the center of gravity of the two audio objects.

[0028] According to an embodiment, the processor 120 may be configured to obtain information about the perceptual difference between two audio objects, for example, according to a three-dimensional directional loudness map, and / or the processor 120 may be configured to process a plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects, for example, according to a directional loudness map, which may for example depend on direction-dependent loudness perception. In one embodiment, the processor 120 may be configured to synthesize a directional loudness map onto a uniformly sampled grid on a surface around the listener, for example, depending on the positions and energies of multiple audio objects.

[0029] According to one embodiment, the directional loudness map may depend, for example, on a grid that depends on a perceptual coordinate system and one or more attenuation curves. In one embodiment, the processor 120 may be configured to determine, for example, the sum of the differences between the three-dimensional directional volume map and another three-dimensional directional volume map as the distance metric for the audio sound scene and the other audio sound scene.

[0030] According to one embodiment, the distance metric may rely, for example, on a 3D directional loudness map and a spatial masking model. In one embodiment, the processor 120 may be configured to, for example, process a plurality of audio objects to obtain a plurality of audio object clusters. Further, the processor 120 may be configured to obtain the plurality of audio object clusters, for example, by associating each of three or more audio objects of the plurality of audio objects with at least one of two or more audio object clusters, whereby, for each of the two or more audio object clusters, at least one of the three or more audio objects may be associated with, for example, the aforementioned audio object cluster, and for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects may be associated with, for example, the aforementioned audio object cluster. Further, the processor 120 may be configured to obtain the plurality of audio object clusters, for example, according to a distance metric representing a perceptual difference in spatial characteristics of an audio sound scene.

[0031] According to an embodiment, the apparatus 100 may further comprise, for example, an encoding unit, which may be configured to generate encoded information encoding, for example, a plurality of audio object clusters or a plurality of processed audio objects, and / or which may be configured to generate, for example, encoded information encoding a plurality of audio objects of an audio sound scene and information regarding a perceptual difference between two of the plurality of audio objects.

[0032] 2 shows a decoder 200 according to one embodiment. The decoder 200 comprises a decoding unit 210 and a signal generator 220. Each of the plurality of audio objects of the audio sound scene represents a real or virtual sound source that is different from any other real or virtual sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same real or virtual sound source at different locations.

[0033] The decoding unit 210 is configured to decode the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects, wherein the plurality of audio object clusters or the plurality of processed audio objects depend on the plurality of audio objects of the audio sound scene and on a distance metric representing a perceptual difference in spatial characteristics of the audio sound scene, and the signal generator 220 is configured to generate two or more audio output signals in response to the plurality of audio object clusters or in response to the plurality of processed audio objects. and / or the decoding unit 210 is configured to decode the encoded information to obtain a plurality of audio objects of the audio sound scene and to obtain information regarding a perceptual difference between two of the plurality of audio objects, the perceptual difference being dependent on a distance metric; and the signal generator 220 is configured to generate two or more audio output signals according to the information of the plurality of audio objects and according to the perceptual difference between said two audio objects.

[0034] 3 illustrates a system according to one embodiment, which comprises the device 100 of FIG. The apparatus 100 of Fig. 1 further comprises an encoding unit configured to generate encoded information encoding the plurality of audio object clusters or the plurality of processed audio objects and / or configured to generate encoded information encoding the plurality of audio objects of the audio sound scene and information regarding a perceptual difference between two of the plurality of audio objects.

[0035] Furthermore, the system comprises a decoding unit 210 and a signal generator 220 . The decoding unit 210 is configured to decode the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects, and the signal generator (220) is configured to generate two or more audio output signals in response to the plurality of audio object clusters or in response to the plurality of processed audio objects.

[0036] and / or the decoding unit 210 is configured to decode the encoded information to obtain a plurality of audio objects of the audio sound scene and to obtain information regarding a perceptual difference between two of the plurality of audio objects, and the signal generator 220 is configured to generate two or more audio output signals according to the plurality of audio objects and according to the perceptual difference between said two audio objects.

[0037] Specific embodiments are described in detail below. According to some embodiments, a perceptual distance model is provided. The task of the developed perceptual distance model is to obtain a distance metric that represents perceptual differences in spatial properties of 3D audio sound scenes in a computationally efficient way. This can be achieved, for example, by transforming geometric coordinates into a coordinate system that takes into account the direction-dependent localization accuracy of human hearing. Furthermore, the distance model can incorporate perceptual properties of the entire scene that contribute to localization uncertainty as well as masking effects.

[0038] According to some embodiments, a Perceptual Coordinate System (PCS) is provided. The localization accuracy of human spatial hearing is known to be heterogeneous. For example, localization accuracy has been shown to be higher in front of the listener than to the side, higher for horizontal localization than for vertical localization, and higher in front of the listener than behind the listener. This property can be exploited to optimize perceptual quality for, for example, quantization schemes or object clustering algorithms. To model non-uniform characteristics for spatial audio processing, a perceptual coordinate system (PCS) according to one embodiment is provided. The PCS may utilize, for example, a distorted coordinate system in which distances within the coordinate system (e.g., Euclidean distance) correspond to "perceptible differences" between sound source positions rather than their physical distances. In other words, instead of considering localization accuracy according to absolute positions, the non-uniform characteristics of perception may be represented, for example, by distorting the coordinate system itself. This is similar to using a psychoacoustic frequency scale (e.g., the Bark scale or, for example, the ERB scale) to represent the non-uniformity of frequency resolution in human hearing.

[0039] Figure 4 shows a two-dimensional example of perceptual coordinate system coordinates distorted according to one embodiment. In particular, Figure 4 shows two-dimensional perceptual coordinates distorted for sound source positions (dots) assumed to be perceptually equally spaced apart in a horizontal plane. More specifically, Figure 4 shows sound source positions spaced perceptually equally spaced (e.g., an exemplary JND) within a unit circle in the median plane. In the case of the geometric coordinates in Figure 4a), the distance depends on the absolute orientation of the sound source. In the perceptual coordinates in Figure 4b), the positions are distorted so that the Euclidean distance between the sound sources remains constant.

[0040] A perceptual coordinate system according to one embodiment can, for example, approximate the perceived difference between arbitrary sound source positions and enable deriving updated positions with low computational complexity, for example for fast spatial audio processing algorithms, e.g., real-time clustering of object-based audio. The mapping from geometric coordinates to perceptual coordinates is designed to be unique and invertible, e.g., a bijective mapping function. For example, all calculations and updates of sound source positions may be performed, e.g., in the perceptual domain, and the final result may be transformed back, e.g., to the physical space domain.

[0041] According to one embodiment, a method is provided for deriving a PCS based on an analysis of HRTF data, e.g., using models of binaural and spectral localization cues and a multidimensional scaling (MDS) approach to pairwise differences. This can, for example, generate a mapping of a grid of positions provided by the analyzed HRTF database, which can be used, for example, for table lookup and interpolation. For a closed-form representation, a mapping function can, for example, be curve-fitted to the analysis grid data, and a simplified mapping model can, for example, be derived.

[0042] In the case of a generalized PCS model, the analysis can be calculated and averaged, for example, using HRTF data from many subjects. Furthermore, it should be noted that the presented analysis method can be calculated, for example, specifically for a binaural renderer that uses a known HRTF data set in the target application, e.g., general-purpose or personalized HRTF data. Existing models can estimate localization cues and perceived differences between sound source positions. However, for spatial audio processing algorithms (such as object clustering), they require iterative calculation of the localization model, which is computationally inefficient and disadvantageous for real-time applications. By considering and representing the perceptual model in the analysis and construction steps of PCS, the computationally expensive parts of the model can be calculated in an offline preprocessing step, resulting in a computationally efficient model suitable for real-time processing. Furthermore, PCS allows for the direct manipulation of sound source locations in the perceptual domain (e.g., optimization of cluster centroid locations).

[0043] Additionally, the PCS can be modeled based on, for example, HRTF data analysis, thereby providing a tailored, perceptually optimized model for a target application with a given HRTF data set. The "resolution" of the human auditory system varies with changes in azimuth and elevation, and depends on the absolute position of the sound source. The baseline model considers only the angles of incidence, eg, azimuth and elevation, relative to the listener, while assuming a constant distance of the sound source (see extension below for distance model).

[0044] The position along the interaural axis ("left / right") is determined by binaural cues (ICC, ILD, ITD, IPD), resulting in the so-called cone of confusion (CoC), along which the binaural cues are approximately constant. Note that if the radius is assumed to be constant, the cone reduces to a "circle of confusion" along a sphere with a given radius. Along the CoC, the spectral coloration introduced by the pinna, head, and shoulders can be used as a primary cue, for example, for elevation localization and front-to-back confusion resolution. Note that spectral filtering is not necessarily the same for both ears at a given elevation angle, thus introducing potential additional binaural cues.

[0045] To represent this separation of cues, one could use, for example, a "binaural spherical polar coordinate system," where azimuth represents "left / right" position along the horizontal plane between ±90° and elevation represents "up / down" position along the CoC ranging from 0°...360°, e.g., a polar coordinate system with the rotation axes aligned with the ear positions, e.g., the poles are placed to the left and right of the listener rather than in vertical polar coordinates, and the poles are above / below the listener as in geographic coordinates. The just noticeable difference (JND) is significantly smaller for azimuth differences (about 1°) than for elevation differences (about 4° for noise, and up to 10-15° for spectrally sparse content). Furthermore, localization accuracy also depends on the absolute position, being more accurate, for example, in front of the listener than above.

[0046] Thus, neither the Euclidean distance between Cartesian coordinate positions (e.g., on a unit sphere) nor the angular distance in polar coordinates corresponds to perceived distance. Although position can be represented by a 2D coordinate system (e.g., spanned by azimuth and elevation angles) that parameterizes a 2D surface (e.g., the unit sphere), the "wraparound" property of a closed sphere (i.e., 360° = 0°) cannot be represented when calculating distance in a 2D coordinate system, and therefore a generalized PCS requires (at least) three dimensions.

[0047] The following describes the concepts for generating a PCS. A primary target application of a PCS is to consistently express the JND of localization accuracy for a given location, e.g., to determine whether two locations are close enough to each other that they can be combined into one without any perceptual change. Thus, a selected design goal for a PCS may be, for example, the property that a Euclidean distance of 1 from a given location always corresponds to a JND for the respective direction. The JND for elevation angles along the cone of confusion can be predicted from the JND to distinguish spectral differences between HRTFs (see ICASSP19), and the JND for azimuth angles in the horizontal plane can be estimated from the JND for ILDs, which have been extensively investigated by experiments in the literature. Based on the position-dependent JNDs, the PCS can be constructed as an absolute coordinate system that is scaled by accumulating JNDs between positions. In other words, the Euclidean distance between any two positions can correspond to the accumulated number of JNDs between them.

[0048] Note that this concept is loosely based on the Weber-Fechner law. Although the Weber-Fechner law shows a logarithmic relationship, the positional distance is measured in the linear domain. However, the perceptual cues considered, such as ILD or spectral difference, are already measured in the logarithmic domain. For example, assuming a JND of 1 dB, a PCS distance of 10 JNDs corresponds to 10 dB. Based on this concept, according to one embodiment, the perceptual distance (PD) = "number of just-in-time differences" between two given locations can be calculated, for example, from HRTF measurements at each location. Using an available set of HRTF databases, the complete set of pairwise distances between given HRTF measurement locations can be calculated, for example, and averaged across multiple subjects. This results in a matrix of pairwise perceptual distances between a given grid of geometric input locations.

[0049] To derive absolute coordinates from a given set of pairwise differences, machine learning techniques using, for example, multidimensional scaling (MDS) can be used, whereby one or more coordinates in selected dimensions, for example, three dimensions, that approximate a given distance can be calculated, for example. According to one embodiment, the MDS technique may provide, for example, a set of PCS locations for the spatial locations of the corresponding HRTF measurements.

[0050] FIG. 5 shows perceptual coordinates obtained via multidimensional scaling of modeled differences in the CIPIC HRTF database according to one embodiment. In applications where only grid locations are of interest, the resulting locations can be used, for example, as a lookup table. Calculating the distance between any two locations can employ, for example, interpolation in the lookup table. To obtain a continuous closed-form solution where the PCS coordinates are invertible to the geometric coordinates, according to one embodiment, a lower-dimensional model can be fitted to the MDS results, for example.

[0051] The following describes pre-processing, particularly coordinate alignment, according to one embodiment. In such pre-processing, in one embodiment, the MDS coordinates are inherently misaligned with the geometric characteristics of the input location (eg, left-right, front-back). Because MDS is based on relative distance, the resulting PCS positions can be, for example, mirrored, translated and rotated without affecting their fit to the underlying relative distance measurements. However, for intuitive understanding of the coordinate system, it may be preferable for PCS coordinates to be aligned as closely as possible with actual spatial locations, e.g., a clear correspondence of what is "left," "right," "front," and "up." MDS can result in coordinates sorted by their contribution to the variance of the input dataset, similar to the energy compaction properties in, for example, primary component analysis (PCA). Since binaural cues have a substantial effect on perceptible differences and have an approximately monotonic relationship with azimuthal position, typically the first coordinate may be mirrored with respect to spatial coordinates, but may correspond, for example, to a "left / right" axis.

[0052] However, because spectral cues have no inherent relationship to elevation position and are subject to wraparound, MDS results can exhibit arbitrary rotations, e.g., coordinates may correspond to axes pointing from "low back" to "top front," and in some cases may exhibit some deformation between coordinates, see, e.g., the "D-shape" of the mid-plane coordinates in the diagram in Figure 5. Thus, prior to curve fitting, coordinates from the MDS results can be aligned to correspond to desired properties of geometric coordinates on the unit sphere by, for example, reflection (e.g., to align left-right flips), translation (e.g., to align anterior / posterior or superior / inferior hemispheres), and rotation (e.g., to align points in a horizontal plane).

[0053] The following describes a curve fitting technique, specifically a polynomial nonlinear regression, according to one embodiment. To obtain a continuous mapping function from spatial coordinates to perceptual coordinates, for example, curve fitting techniques can be used. According to one embodiment, multidimensional nonlinear regression can be used, for example, to fit polynomial approximations or spline representations to the MDS results. However, since the available positions in the HRTF database are typically sparsely sampled, the parameterization may be chosen appropriately to, for example, avoid overfitting. Furthermore, most HRTF databases do not include measurements of the region below the listener, so care must be taken to ensure that this extrapolated region is well behaved. Otherwise, for example, the lower back region may result in large overshoots in spline or polynomial fitting.

[0054] To maintain the underlying model assumptions of binaural and spectral cues, for example, a separate fitting approach can be applied. For example, one aspect corresponds to binaural cues, which are clearly separated between left and right and have no "wraparound." Therefore, this is fitted to be represented by a single coordinate. For example, another aspect corresponds to monaural spectral cues along a cone of confusion that essentially includes cyclic wraparound. Therefore, the anterior / posterior axis and the superior / inferior axis may be fitted together to represent a cross section along the cone of confusion.

[0055] To avoid overfitting, a linear model is chosen for the first coordinate U (left / right) and a second order polynomial for the second and third coordinates V and W. u p (x)=27.8x v p (y)=8.15y 4 -1.75y 3 -3.46y 2 +4.61y-0.60 w p (z)=-6.94z 4 +4.03z 3 +3.11z 2 +3.92z-1.13 Figure 1 shows MDS results (points) and curve fitting (surface) of the CIPIC HRTF database.

[0056] FIG. 6 shows a polynomial model-based perceptual coordinate system according to one embodiment, whose surface represents a distorted unit sphere. In the following, an efficient model fitting technique, specifically linear fitting of an ellipsoid, according to one embodiment is described. Especially for real-time applications, such as real-time object clustering, a coordinate system that is computationally simple and efficiently invertible is required. The MDS results and polynomial fitting may resemble an ellipsoid, except for, for example, a "dip" in the anterior / posterior confusion, and a "skirt" in the lower dorsal position close to the main body. As a simplified model approximation, for example, an ellipsoid can be used. This can be constructed efficiently, for example, by scaling the Cartesian coordinates of the unit sphere by an appropriate factor, which can also be easily inverted by inverse scaling.

[0057] Here, the mapping function is, for example, U=c u *X V=c v *Y W=c w *Z With appropriate weights as above, this can be reduced to, for example, a scalar scaling of the individual coordinates.

[0058] The scaling factors can be derived from the MDS results, for example, by linear fitting of the respective mapping functions, which can be reduced to scalar weightings of the coordinates of the unit sphere. However, the scaling coefficients of the selected ellipsoidal model may be fitted directly to approximate the underlying distance matrix, for example, without computing the MDS. This reduces the computation time and minimizes the approximation error, since otherwise two fitting operations would be performed (distance → MDS → ellipsoid).

[0059] FIG. 7 shows an ellipsoidal model-based perceptual coordinate system according to one embodiment, whose surface represents a distorted unit sphere. The following describes input data selection for parameter fitting according to one embodiment. Note that in general, for ellipsoidal models, trade-offs must be considered when selecting the range of input locations. MDS results may show a "tail" at lower locations that emphasizes the distance between low front and low rear, for example. Because these locations are separated by the listener's torso, the shadow of the torso can, for example, provide additional spectral cues between these locations, thus making the distinction easier than the front / rear at higher locations. However, this cannot be represented by an ellipsoid, so the anterior / posterior coefficient is a compromise between the lower and upper hemisphere, with the anterior / posterior confusion being more pronounced in the horizontal plane and at higher positions.

[0060] This can be taken into account if the target application scenario (= playback system) is known. For example, in the case of an immersive loudspeaker setup, loudspeaker positions are mainly located in the upper hemisphere, so lower hemisphere positions may, for example, be omitted (or weighted lower) in the parameter fitting. Conversely, in VR applications, playback of sound sources below the listener is more common, so lower hemisphere positions need to be incorporated into the model fitting. The resulting distortion coefficients may depend, for example, on the database, the frequency range analyzed, and / or the input considered. Parameter fitting of the CIPIC HRTF database yields, for example, c_u=28.1, c_v=5.81, and c_w=8.56. A set of coefficients averaged across multiple HRTF databases is, for example, c_u=25 (left / right), c_v=6 (front / back), and cw=5 (top / bottom).

[0061] For binaural rendering applications where the playback HRTFs are known, the PCS can be modeled directly on the HRTFs in use, instead of using generic approximations from a database, for example. The PCS model can be updated for real-time applications where the HRTFs can be personalized, for example, when a new set of HRTFs is loaded. Therefore, as mentioned above, high computational efficiency of the model fitting itself is also desirable.

[0062] For more advanced modeling, the PCS may be constructed frequency-dependently to reflect, for example, larger HRTF differences for higher frequencies (see Blauert's directional bands). This is particularly relevant for coordinates representing spectral cues (V / W). Psychoacoustic experiments in the literature have shown that the left / right localization of physical sound sources is less dependent on frequency. While ILD differences are smaller at lower frequencies, ILD / IPD cues become more relevant. Therefore, non-frequency-dependent scaling of the left / right axis can be used in combination with frequency-dependent scaling, for example, along the cone of confusion.

[0063] The transformation from geometric coordinates to PCS coordinates may be applied, for example, to transform spatially distributed sound source positions into a domain that represents the perceptual properties of sound source localization in human hearing. In the PCS domain, the perceptibility of sound source position differences can be expressed, for example, by the Euclidean distance between PCS coordinates, which allows for computationally efficient estimation of perceptual differences in sound source localization. Furthermore, the PCS domain may be calibrated, for example, to represent 1 JND as a PCS distance of 1. This allows for estimating the limits of localization accuracy for any given position. This can be applied, for example, to control the resolution of the quantization scheme.

[0064] To convert a sound source position given in geometric coordinates (X, Y, Z) into perceptual coordinates (U, V, W), a mapping function can be applied, which may be expressed in general terms, for example, as follows: U=f U (X,Y,Z) V=f V (X,Y,Z) W=f W (X,Y,Z) To convert coordinates from the perceptual domain back to geometric coordinates, an inverse mapping function can for example be applied, which may for example be in general notation as follows: X=f -1 X (U,V,W) Y=f -1 Y (U,V,W) Z=f -1 Z (U, V, W) The inverse mapping function allows operations, such as manipulating sound source positions and calculating tolerances, to be performed directly in the perceptual domain. This enables, for example, computationally efficient perceptual-based algorithms for processing spatial audio to operate entirely directly in the perceptual domain without requiring repeated calculation of perceptual models. The resulting spatial positions in the perceptual domain can then be converted back to geometric coordinates, for example, via the inverse mapping function.

[0065] The appropriate mapping function is derived as above. For a computationally efficient implementation, for example, a separable ellipsoid approximation approach may be preferred, and the mapping function may be simplified, for example, to: U=c u *X V=c v *Y W=c w *Z Therefore, the inverse mapping function may be simplified, for example, to: Y=U / c u Y=V / c v Z=W / c w Note that the ellipsoidal mapping function is valid for positions on the unit sphere and the corresponding ellipsoidal surface. If a spatial operation results in a position outside the surface, the position can be mapped back to the defined surface, for example, by projecting onto the unit sphere in geometric coordinates, or by choosing the closest point on the ellipsoidal surface in the PCS domain.

[0066] In the following, a 3D directional volume map (3D-DLM) according to some embodiments is described. The purpose of DLM is to represent the "amount of sound coming from a given direction." In other words, it represents the composite volume perceived from the superposition of all sound sources in a scene, taking into account the localization accuracy of human hearing. In the context of object-based audio, the sound source locations and corresponding signal characteristics are known. Based on that, DLM can be calculated, for example, as the cumulative contribution of all active sound sources weighted by a distance-based decay function, e.g., a Gaussian or linear decay function.

[0067] 8 shows an example of synthesis of a one-dimensional directional loudness map (1D-DLM) based on known object positions and loudness, according to one embodiment. Note that this example shows that the accumulation of, for example, four closely spaced sound sources on the right side results in a higher synthesized loudness than individually loud sound sources around a central position. According to one embodiment, DLM synthesis can be extended to localization in 3D space into a 3D-DLM, for example, by using a sampling grid (e.g., a unit sphere) on a surface surrounding the listener and calculating the cumulative contribution of all sound sources for each grid point, resulting in a 3D-DLM as shown in the example calculation in Figure 9.

[0068] 9 shows an example of a 3D directional loudness map synthesized from known sound source positions (marked with crosses) according to an embodiment, where (a) shows the 3D-DLM on the unit sphere and (b) shows the 3D-DLM in perceptual coordinates. Known binaural one-dimensional DLMs represent perceived loudness based on binaural cues, i.e., "left / right" spatial images. However, according to some embodiments, immersive audio applications can also take into account spatial properties in 3D space, such as top / bottom and front / back relationships, which can be made possible, for example, by utilizing a 3D DLM. Furthermore, known DLMs require a scene analysis step, in which a binaural downmix of the entire sound scene is computed and processed by binaural cue analysis to extract a binaural 1D-DLM. In the context of object-based audio, signal characteristics such as source positions and signal energy are known a priori. According to one embodiment, a 3D-DLM may be synthesized directly from this information without requiring the computational complexity of computing, for example, a binaural downmix and scene analysis step.

[0069] Below, a baseline concept for the generation of a 3D-DLM according to one embodiment is provided. The 3D-DLM can, for example, be calculated on a grid on a surface around the listener, with each point corresponding to, for example, a unique spherical coordinate angle, e.g., a uniformly sampled unit sphere. Further details and different embodiments regarding sampling and surface geometry are described below. The energy of each sound source can be calculated (e.g., as described below) and can be spread around its position with a given decay curve. Following the convention of one-dimensional DLM, the decay curve is modeled according to a Gaussian distribution. For lower computational complexity, a linear decay curve in the logarithmic domain can be used instead. Attenuation may be determined by the Euclidean distance between positions in 3D space, as opposed to angular distance or distance along the surface of a sphere / ellipsoid, to account for perceptual effects such as before / after confusion. For example, the energy contribution of each sound source weighted by the magnitude of the attenuation function can be calculated for each sound source and each grid point and accumulated for each grid point to calculate a directed energy map (DEM).

[0070] This approach assumes uncorrelated sources; if correlation between sources is expected, phantom source extraction is performed in a preprocessing step (see, e.g., below). To account for the increased blurring of phantom sources, the decay curves can be adjusted, for example, to represent a wider spread. From the energy sum at each grid location, the respective loudness can be calculated as, for example, Energy^0.25=sqrt(sqrt(Energy)), an approximation of the exponent 0.23 given by the Zwicker loudness model. Note that in a real-world playback environment, assuming uncorrelated sound sources, the physical energy of the sound sources is superimposed on the ear, rather than a perceptual measure of loudness, so the summation may be done, for example, in the energy domain, and not, for example, in the loudness domain.

[0071] The standard deviation of the spread attenuation curve, e.g., a Gaussian distribution, can be determined, e.g., psychoacoustically, which corresponds to the JND of the localization accuracy. To achieve low computational complexity, e.g., for real-time applications, a baseline model of the 3D-DLM can be obtained, e.g., using time-domain energy calculations, e.g., full-band energy, e.g., for each frame. To incorporate the frequency dependence of human loudness perception, the signal is pre-filtered, e.g., using A-weighting or K-weighting. Otherwise, high energy in the low-frequency region, for example, would be over-represented. The perceptual weighting can be implemented computationally efficiently, e.g., in the form of a relatively low-order IIR filter, e.g., a 7th-order filter for A-weighting.

[0072] Extensions and further embodiments are contemplated here. To reduce computational complexity, for example, if the tails of the Gaussian distribution are below a given threshold, the attenuation curve may be truncated and a simpler spread function, e.g., linear attenuation, can be used, and attenuation curve weights can be buffered and / or pre-computed for fixed source positions corresponding to speaker positions in a defined configuration, e.g., 5.1, 7.1+4, 22.2. For advanced perceptual models where higher spectral resolution is required, a frequency-dependent DLM can be calculated. The DLM calculation may then be performed for each spectral band, e.g., at ERB resolution. As an extension, for a frequency-dependent DLM, the diffusion coefficient may be frequency-dependent, e.g., to account for the different localization accuracy of human hearing in different frequency regions.

[0073] As an extension, correlations between sound sources that result in phantom sound sources are taken into account, for example when the sound sources correspond to two or more channels in a stereo or multi-channel channel-based production. According to an embodiment, the direct and diffuse signal parts may, for example, be extracted. For this purpose, for example, the cross-correlation between the individual channels can be calculated. For correlations above a given threshold, eg 0.7, phantom sources may eg be inserted and direct and diffuse partial decomposition may eg be performed. The positions of the phantom sound sources can be calculated, for example, based on the energy ratio between the original sound source positions, for example by a weighted average of the positions, or by an inverse panning law, for example sine law panning. To account for the reduced localization accuracy of phantom sound sources, the spreading coefficient of the spatial attenuation function may be widened by, for example, a factor appropriate for the phantom sound source, which may be fixed (e.g., 2 JND) or may be scaled based on the amount of correlation (i.e., use a narrower spread for higher correlation since phantom sound sources are better localizable).

[0074] To take into account the remaining uncorrelated parts of the signal, e.g. the diffuse part, the overall signal energy can be distributed between the additionally inserted phantom sources and the original source positions, e.g. based on correlation coefficients. To take into account the diffuse properties of the remaining (uncorrelated) signal parts, the diffusion coefficient of the original sound source position can also be adjusted, for example by an appropriate factor, which can be, for example, fixed, e.g., 2 JND, or can be, for example, scaled based on the amount of correlation, e.g., resulting in a wider spread for higher correlations, as opposed to the spread of phantom sound sources, e.g., since the remaining parts correspond to a diffuse field rather than a sound source at the original position. To take into account the "slowness" of human hearing in terms of temporal localization accuracy, for example, a temporal diffusion factor can be used that weights and adds the DLM of the previous frame to the current frame. The temporal diffusion factor can be determined, for example, by the temporal characteristics of human hearing and therefore needs to be adapted to the frame length and sample rate.

[0075] Next, the sampling grid of the DLM according to one embodiment will be described. 10 illustrates different sampling methods for a unit spherical grid, according to an embodiment, where (a) shows azimuth / elevation sampling and (b) shows an ICO sphere. See, e.g., https: / / en.wikipedia.org / wiki / Geodesic-_polyhedron. See also, https: / / medium.com / @qinzitan / mesh-deformation-study-with-a-sphere-ceee37d47e32. For numerical computation, the DLM may be sampled, for example, on a grid surrounding the listener. The sampling resolution of the grid is a trade-off between spatial accuracy and computational complexity and therefore needs to be optimized by observing geometric and perceptual properties.

[0076] According to one embodiment, the generation of the grid for computing the DLM is done by uniformly sampling along the azimuth and elevation angles along the unit sphere (e.g., 360 x 180 = 64800 points). However, spherical coordinates are much denser towards the poles, thus resulting in non-uniform oversampling and generating an unnecessary large number of points. This results in significant overhead in computational complexity. Furthermore, subsequent algorithms (e.g., Gaussian mixture models) may be hindered by the non-uniform sampling, for example, due to the increased density of values ​​at the poles.

[0077] A method of uniformly sampling a sphere (e.g., for computer graphics) may be, for example, a "geodesic / polyhedron," "geosphere," or "ICOsphere," derived by subdividing a regular icosahedron. To maintain a resolution of approximately 1°, for example, a 5-division ICO sphere may be used, resulting in a grid of 10242 points (a uniform grid of approximately 16% of the azimuth / elevation angles), which significantly reduces computational and memory requirements while maintaining comparable perceptual quality. For many applications, a lower order may be sufficient, for example using only three partitions corresponding to 642 points.

[0078] The following describes a spatial masking model (SMM) according to some embodiments. FIG. 11 illustrates masking model calculation in perceptual coordinates according to one embodiment. The masking effect occurring in human hearing between loud and soft sounds is an important aspect of psychoacoustic models for audio coding. Existing models typically estimate masking thresholds for mono or stereo coding. However, in immersive audio applications, the masking effect between arbitrary sound source positions is of interest. Subjective listening test experiments typically cover only a limited selection of position pairs for which the masking effect is measured. To estimate the masking effect between arbitrary sound source positions for immersive audio, a generalized spatial masking model (SMM) is provided in one embodiment. Findings in subjective experiments suggest that masking differences may be related to, for example, available localization cues, and that the differences may be related to localization accuracy. PCS and 3D-DLM have been introduced as models of localization accuracy and the spread of loudness perception. Based on this, a spatial masking model for arbitrary sound source positions has been derived, in which the distance between sound sources can be calculated in the PCS domain to estimate the difference in localization cues, and spatial attenuation curves are applied to model the unmasking effect. This is shown in Figure 11 for locations in the median plane and for a masker at -30° azimuth. It can be seen that due to the smaller distance in the PCS representation, stronger masking is incorporated for anterior-posterior symmetric locations, while there is substantially less masking for laterality, with interaural cues contributing more to unmasking.

[0079] Masking models aimed at perceptual audio coding may, for example, need to be time- and frequency-dependent in order to control the spectral shaping of the introduced quantization noise. Conversely, object clustering affects the spatial location of sound sources. Modifying sound source locations as a whole may, for example, be an essentially "full-band" operation. It should be recognized that masking between individual sound sources may still be frequency dependent, for example. However, changing the spatial location of a sound source changes localization cues rather than introducing additional noise. In other words, a masking model for localization changes may have different requirements than a masking model for additional signals, such as quantization noise.

[0080] For real-time applications, for example, a computationally efficient model may be required, and therefore a simplified full-band masking model based on time-varying signal energy may be applied, for example, in the context of object clustering. To take into account the frequency-dependent sensitivity of human hearing, frequency weighting can be applied, e.g., A-weighting, which can be achieved by time-domain filtering with a relatively short filter, e.g., an IIR filter of order 7. It should be noted that operations capable of removing signal components like the culling of inaudible sources in the context of object-based audio are used, preferably operations that utilize frequency-dependent masking models, as this is more similar to the use cases of adding signal components (quantization noise) or removing them (quantizing to zero) in perceptual audio coding.

[0081] Next, a masking model overview according to some embodiments is provided. The SMM can assume, for example, a maximum masking threshold at the location of the masker, e.g., intra-source masking. The masking threshold can be reduced for spatially separated sources, e.g., weighted by an attenuation function depending on the spatial distance. The attenuation function may be, for example, a linear attenuation in the logarithmic domain ("dB per distance"), or a Gaussian-shaped attenuation curve, which allows, for example, the DLM calculations to be reused or shared to reduce computational complexity.

[0082] In addition to distance-dependent masking, for example, a position-independent offset may be added to the masking threshold, which depends on the sum of the energy of all sound sources in the scene weighted by a maximum unmasking factor (e.g., -15 dB). This is done to reflect the fact that there is always some residual masking between sound sources. (Psychoacoustic experiments have shown that the maximum level of binaural / spatial unmasking is approximately 15 dB BMLD with headphones.) In other words, masking between spatially separated sound sources cannot be zero, for example, because the amount of spatial unmasking is limited (maximum BMLD has been found in the literature to be approximately 15 dB in headphone experiments). However, spatial masking experiments show that there is still a fairly steep initial decay for unmasking of spatially separated sound sources, so the decay curve needs to reflect that as well. Therefore, especially when using a Gaussian model for the decay curve, the curve should not be chosen to be very wide in order to accommodate maximum unmasking at maximum distance, but rather should be chosen so that it is steep enough locally around the sound source, but later only drops off to a given minimum value rather than zero. Similar to localization accuracy, there may be differences in spatial unmasking between, for example, horizontal and vertical separation. To reflect this, the distance of the SMM attenuation curves can be calculated, for example, by PCS rather than geometric distance. This allows interaural (left / right) differences to unmask more than elevation differences, while preserving significant masking between anterior / posterior symmetric sound sources.

[0083] Detailed calculations according to a specific embodiment will now be described. the local energy diffusion map M of the sound source represented by the object with index k local (k) is, for example, the Euclidean distance D PCSA weighted object energy for all object indices i weighted by a Gaussian-shaped decay function depending on (k,i) and the (adjustable) diffusion coefficient s This can be calculated from the sum of TIFF2025533618000002.tif68. Note that in contrast to the parameterization of the normal distribution density function, the decay function in the masking model is not normalized; for example, the diffusion coefficient scales only the width of the distribution, not the height (and therefore the overall sum of the source contributions). In other words, a higher diffusion coefficient implies "more masking power," similar to the diffusion function in frequency-domain masking. (This should not be confused with affecting the overall loudness of the scene, especially given the context of DLM calculations.) Optionally, according to certain embodiments, the diffusion coefficient is, for example, 2s for all sound sources. 2 = 5 (which will result in a corresponding standard deviation of the normal distribution, e.g. TIFF2025533618000004.tif623, resulting in a diffusion width of 1-2 JND), or alternatively for wider diffusions, e.g., s=6 (e.g., 2s 2 =72).

[0084] Furthermore, optionally, according to another particular embodiment, as a further improvement of the model accuracy, if suitable detectors are available, for example in a given implementation, the diffusion coefficient may depend on the signal characteristics and masking capabilities (noise-like, tonal, transient, ...) of individual objects. In addition to local masking, the minimum residual masking between sound sources (and vice versa, corresponding to the maximum binaural unmasking) is calculated, e.g., by the energy diffusion map M min can be incorporated as a global minimum of

[0085] According to one embodiment, the minimum masking may be, for example, direction-independent. In other words, it may reflect, for example, the overall sound energy of the scene that limits the resolution of the ear. This can be estimated from the sum of signal energies weighted by the worst-case BLMD value found in the literature of 15 dB [Blauert]. Alternatively, this may be calculated as the sum of local energy masking maps at the source location, e.g., the source's energy and the local contributions of neighboring sources. This models the increased masking ability of groups of sources closer to each other. Furthermore, if the diffusion coefficient can be modeled as signal-dependent, this also models sources with wider diffusion coefficients as having a greater impact on overall (minimal) masking. TIFF2025533618000006.tif1746 Combined masking threshold T k can be calculated as follows, for example, using 20 dB as an upper estimate of the masking threshold (from Hellman72 for tonal masking noise at 60 dB SPL): Note that computing the joint masking as the sum of the local and global maskings has the advantage of preserving the smoothness of the Gaussian falloff and saturating at the offset. Alternatively, this can be done by, for example, computing the joint masking as (M min M, which allows us to cut off the evaluation of the Gaussian function for longer distances (using a calculation based only on the energy of min ,M local It may be implemented as a maximum operation between , thus saving computational complexity.

[0086] The following describes a perceptual distance metric according to some embodiments. An underlying question for perceptual distance metrics in the context of audio object clustering can be, for example, "How perceptible is it when multiple objects are combined into one?" This leads to a more detailed question: "When combining two candidate objects into one, how far will each object be moved, and how audible is the difference introduced by this position change in the context of the whole scene?" The PCS provides a model of the perceptibility of spatial location changes of sound sources, and the SMM provides a model of the audibility of sound sources given the masking effects of the overall sound scene. According to one embodiment, these models can be combined, for example, to derive a measure of the perceptual distance between two sound sources (e.g., objects in this context). Thus, the perceptual distance between two objects can be calculated, for example, based on the inter-object distance in the PCS (to take into account localization differences) and weighted by an estimate of the perceptual relevance of the objects (with respect to the masking effects of the overall sound scene).

[0087] Key concerns for such distance metrics are robustness and numerical stability. Because real-world implementations operate only with limited numerical precision, metrics can be made robust to numerical inaccuracies and borderline cases, such as values ​​close to or equal to zero. For example, if the number of active sound sources varies over time, some audio scene representations may always contain metadata and audio for a maximum number of active objects (similar to a fixed number of tracks in a DAW). This results in "inactive" objects whose signal's PCM data contains only digital zeros or (potentially worse) noise due to numerical inaccuracies (LSB noise). A preferred approach may be to detect and remove these inactive objects in a preprocessing culling step, for example, before the actual clustering. However, this may not be feasible in all applications. Thus, according to one embodiment, the distance metric can be designed to be robust to small / zero energy, for example, by adding an appropriate offset value as needed (e.g., without requiring explicit detection of such cases).

[0088] Next, a definition of a perceptual distance model according to one embodiment is provided. In the field of perceptual audio coding, perceptual entropy (PE) [JJ88] is a well-known measure for assessing the "amount of audible signal content relative to the masking threshold." Here, a simplified, computationally efficient estimate of PE for each object can be calculated using the full-band energy derived, for example, by SMM (which can apply frequency weighting before energy calculation to account for the frequency dependence of human hearing) and the masking threshold.

[0089] Note that, as mentioned above, object position is frequency independent, so frequency-dependent calculations can improve the accuracy of the masking model but cannot add to the degrees of freedom of the clustering algorithm. The PE of an object with index k can be calculated, for example, as follows: TIFF2025533618000008.tif1244 Distance metric D between two object indices k and l Prec (k,l) is, for example, the distance D in PCS as follows: PCS (k,l) can be used to calculate TIFF2025533618000009.tif1540 TIFF2025533618000010.tif839 TIFF2025533618000011.tif14119 Model parameters are, for example, thr offs =33[dB] and d offs = 0.1 [bit].

[0090] Next, a detailed derivation of the model equation according to one embodiment will be described. To avoid numerical instabilities for small energies, for example, an offset can be added to the object energy. Since the energy range can span several orders of magnitude depending on the PCM data scaling, the offset may, for example, be scaled to the overall energy sum (or the maximum energy). For example, for applications with pre-normalized scaling, a constant value may be used. As the offset, a worst-case estimated masking threshold of -33 dB (e.g., assuming an average BMLD of 27 dB + 6 dB for the tonal masking noise) may be selected, for example, by adding a constant offset ε (e.g., ε = FLT_MIN = 1e-37) depending on the calculation accuracy. TIFF2025533618000012.tif1541 TIFF2025533618000013.tif627When combining two objects, for example, the new centroid c k,l Here, it can be assumed that the position is selected as the average position weighted by the energy of the objects. As a result, the centroid position depends on the ratio between the energies of the objects. In other words, the position change of the first object can be larger if, for example, the second object has more energy, and vice versa. Therefore, the candidate centroid c k,l The perceived location distance D of the first candidate object with index k to PCS (k,c k,l ) is, for example, the energy E' of the second object relative to the sum of the energies of both objects. l From the ratio, for example, it can be estimated as follows: TIFF2025533618000014.tif1254 To take into account the perceptual relevance of an object in the context of masking from the entire sound scene, the estimated position distance can be weighted, for example, by the object's PE. The units of the distance metric can be thought of as, for example, "number of bits times JND." With this metric, for example, given two pairs of candidate objects with the same distance, combining the objects with a lower PE can be assigned, for example, a lower penalty.

[0091] To avoid instability of objects with negligible PE or energy, an offset that depends only on the inter-object distance can be added. The offset coefficient can be selected as, for example, 0.1 [bit] (which is smaller than the masking threshold (approximately 0.3 dB = 10 log 10 2 -0.1 ) corresponds to the PE of the signal slightly above. TIFF2025533618000016.tif6108 Expanding and simplifying the above equation gives us: The perceptual distance with radius according to one embodiment is described as TIFF2025533618000017.tif12114 extension. The PCS described above may, for example, only consider the angle of incidence of the sound source relative to the listener in order to model differences in spectral and binaural cues. However, in various applications (e.g., binaural rendering in VR), the distance between the listener and the sound source is also of interest. Therefore, according to one embodiment, additional coordinates may be introduced into the modeled PCS to reflect, for example, the JND of radius change.

[0092] While absolute distance determinations have been shown to be less accurate, relative changes in distance can be more easily detected, for example, based on three main cues: level changes, direct-to-reverberant ratio, and the Doppler effect. Regarding level changes, the intensity of a sound source may for example decrease with longer distances (in free field conditions, SPL decreases as 1 / r^2, in closed environments the level decrease is usually lower due to reverberation). Regarding the direct-to-reverberation ratio (DRR), in a reverberant environment, a distant sound source can, for example, have more reverberation. The Doppler effect causes the pitch of a sound source to change due to the Doppler effect when the relative distance between the listener and the sound source changes at a certain speed.

[0093] The cues from level changes and DRR changes are related. In a reverberant environment, level changes are reduced, but additional cues due to DRR changes may occur. Therefore, for example, an environment-independent radial distance model can be used based on distance-dependent levels. The psychoacoustic literature reports a JND of 1 dB for the detection of level changes. Therefore, the radius-dependent gain can be calculated, for example, as a ratio to a reference radius and converted to the logarithmic domain. Therefore, a 1 dB difference in relative gain directly corresponds to a 1 JND of perceptible distance change in this model.

[0094] The radial distance coordinates can be calculated, for example, as follows. d_r = 20 * log10(r / 0.2 + FLT_MIN) (assuming a reference radius of 0.2 m close to the head, for example) The Doppler effect can cause a pitch shift, for example, when the distance between the sound source and the listener changes over time. For a given frequency f, sound source velocity v_S, listener velocity v_L, and speed of sound c, the resulting frequency can be, for example, f’ = f * (c + v_L) / (c + v_S) and the signs of v_S and v_L depend on whether they are moving towards each other or away from each other.

[0095] Note that the equation depends on both relative and absolute velocities. However, for v << c, it can be simplified by considering only the relative velocity. The human ear is quite sensitive to relative changes in frequency and can detect a change of about 5 percent (5% of a semitone). The relative pitch change can be derived, for example, from the Doppler effect equation. deltaPitch=12 *log2((c+v_L) / (c+v_S))[semitone] =1200*log2((c+v_L) / (c+v_S))[cents] Solving the equation for Doppler pitch shift for a 5 cent JND gives a JND of approximately 1 m / s for both listener and source motion (slow speed). Thus, the velocity component of the PCS can directly model the relative velocity between, for example, the listener and the sound source, with 1 m / s equaling 1 JND.

[0096] Further embodiments are provided below. According to a first embodiment, a distance metric is provided that represents perceptual differences in spatial properties of a 3D audio sound scene. According to a second embodiment, a perceptual coordinate system (PCS) is provided in which geometric distances, for example Euclidean or angular distances, represent perceptible localization differences according to the first embodiment. According to a first variant of the second embodiment, a parametric inverse mapping function is provided for transforming geometric (physical) coordinates into the perceptual coordinate system of the second embodiment. According to a particular variant of the second embodiment, there is provided a method for deriving the mapping parameters of the first variant of the second embodiment based on analysis of HRTF data.

[0097] According to a third embodiment, a masking model for spatially distributed sound sources is provided that uses the perceptual distance-based spatial attenuation curve of the second embodiment. In a first variation of the third embodiment, a masking model of the third embodiment is provided that uses a Gaussian attenuation curve with an offset for minimum masking. A second variant of the third embodiment provides for the calculation of the masking effect of the entire sound scene as a sum of monaural masking thresholds weighted by the position-dependent masking model of the third embodiment. In a third variant of the third embodiment, an estimation of the contribution of a sound source to the sound scene information based on the perceptual entropy (PE) and the sound source energy calculated from the masking model of the third embodiment is provided. In a fourth variant of the third embodiment, identification of inaudible sound sources for culling irrelevant audio objects is provided.

[0098] According to a fourth embodiment, a perceptual distortion metric (PDM) for changes in spatial characteristics of a 3D audio sound scene is provided based on the perceptual distance of the second embodiment and the spatial masking model of the third embodiment. According to a first variant of the fourth embodiment, a distortion metric for the position change of a single source is provided as a weighted combination of the PCS distance from the masking model and the PE. According to a second variant of the fourth embodiment, a distortion metric for integrating two or more sound sources is calculated based on the estimated centroid positions and a weighted sum of the individual distortion metrics.

[0099] According to a fifth embodiment, a 3D directional loudness map (3D-DLM) is provided for representing direction-dependent loudness perception. According to a first variant of the fifth embodiment, the synthesis of 3D-DLMs of known sound source positions and energies on a uniformly sampled grid on the surface around the listener is provided. According to a second variant of the fifth embodiment, a 3D-DLM is provided that is based on the grid and attenuation curve in PCS coordinates of the second embodiment. According to a third variant of the fifth embodiment, the sum of the differences between two 3D-DLMs as distortion metric of the first embodiment for two sound scene representations is provided. According to a fourth variant of the fifth embodiment, a combination of the 3D-DLM as a PE-based difference metric between two sound scene representations and the masking model of the third embodiment is provided.

[0100] While some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0101] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partially in hardware, or at least partially in software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable. Some embodiments according to the invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to cause one of the methods described herein to be performed. Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code can be stored on, for example, a machine-readable carrier.

[0102] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an embodiment of the inventive methods is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer readable medium) comprising a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can for example be adapted to be transferred via a data communication connection, for example via the Internet.

[0103] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. Further embodiments according to the invention comprise an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0104] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus. The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0105] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the pending claims and not by the specific details presented by the description and illustration of the embodiments herein.

Claims

1. An apparatus (100) comprising: an input interface (110) for receiving a plurality of audio objects of an audio sound scene; a processor (120); each of the plurality of audio objects represents a sound source that is different from any other sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same sound source at different locations; the processor (120) is configured to obtain information about a perceptual difference between two audio objects of the plurality of audio objects as a function of a distance metric, the distance metric representing a perceptual difference in a spatial characteristic of the audio sound scene; and / or The apparatus (100), wherein the processor (120) is configured to process the plurality of audio objects according to the distance metric to obtain a plurality of audio object clusters or a plurality of processed audio objects.

2. The apparatus (100) of claim 1, wherein the audio sound scene is a three-dimensional audio sound scene.

3. the processor (120) is configured to obtain the information on a perceptual difference between two audio objects according to a perceptual coordinate system, and / or the processor (120) is configured to process the audio objects according to the perceptual coordinate system to obtain the audio object clusters or the processed audio objects, The device (100) of claim 1 or 2, wherein the distance in the perceptual coordinate system represents a perceptible localization difference.

4. the processor (120) is configured to obtain the information on a perceptual difference between two audio objects in response to an inverse mapping function, and / or the processor (120) is configured to process the audio objects in response to the inverse mapping function to obtain the audio object clusters or the processed audio objects, The apparatus (100) of claim 3, wherein the processor (120) is configured to use the inverse mapping function to transform coordinates in a physical coordinate system into coordinates in the perceptual coordinate system.

5. The apparatus (100) of claim 4, wherein the inverse mapping function depends on head-related transfer function data.

6. the processor (120) is configured to obtain the information on a perceptual difference between two audio objects according to a spatial masking model of spatially distributed sound sources, and / or the processor (120) is configured to process the audio objects according to the spatial masking model to obtain the audio object clusters or the processed audio objects, the spatial masking model is dependent on a masking threshold; The apparatus (100) of any one of claims 3 to 5, wherein the processor (120) is configured to determine the masking threshold as a function of an attenuation function and as a function of one or more distances in the perceptual coordinate system.

7. The apparatus (100) of claim 6, wherein the processor (120) is configured to determine the masking threshold in response to a Gaussian-shaped decay function as the decay function and in response to a minimum masking offset.

8. The apparatus (100) of claim 6 or 7, wherein the processor (120) is configured to identify one or more inaudible audio objects of the plurality of audio objects.

9. the processor (120) is configured to obtain the information on a perceptual difference between two audio objects in response to a perceptual distortion metric, and / or the processor (120) is configured to process the audio objects to obtain the audio object clusters or the processed audio objects in response to the perceptual distortion metric, The apparatus (100) of any one of claims 6 to 8, wherein the processor (120) is configured to determine the perceptual distortion metric as a function of distance in the perceptual coordinate system and as a function of the spatial masking model.

10. The apparatus (100) of claim 9, wherein the processor (120) is configured to determine the perceptual distortion metric as a function of a perceptual entropy of one or more of the plurality of audio objects.

11. 11. The apparatus of claim 10, wherein the processor is configured to determine the perceptual distortion metric in response to a first distance between a first one of two audio objects of the plurality of audio objects and a center of gravity of the two audio objects, and in response to a second distance between a second one of the two audio objects and a center of gravity of the two audio objects.

12. the processor (120) is configured to obtain the information on the perceptual difference between two audio objects in response to a three-dimensional directional loudness map; and / or the processor (120) is configured to process the plurality of audio objects to obtain the plurality of audio object clusters or the plurality of processed audio objects according to the directional loudness map; The device (100) according to any one of claims 3 to 11, wherein the three-dimensional directional loudness map relies on direction-dependent loudness perception.

13. 13. The apparatus of claim 12, wherein the processor is configured to synthesize the directional loudness map onto a uniformly sampled grid on a surface around a listener according to positions and energies of the plurality of audio objects.

14. 14. The apparatus (100) of claim 12 or 13, wherein the directional loudness map depends on a grid that depends on the perceptual coordinate system and on one or more attenuation curves.

15. The apparatus (100) of any one of claims 12 to 14, wherein the processor (120) is configured to determine the sum of differences between the three-dimensional directional volume map and another three-dimensional directional volume map as the distance metric between the audio sound scene and another audio sound scene.

16. An apparatus (100) according to any one of claims 12 to 15, further dependent on claim 6, wherein the distance metric depends on the three-dimensional directional loudness map and the spatial masking model.

17. the processor (120) is configured to process the plurality of audio objects to obtain the plurality of audio object clusters; the processor (120) is configured to obtain the plurality of audio object clusters by associating each of three or more audio objects of the plurality of audio objects with at least one of the two or more audio object clusters, whereby for each of the two or more audio object clusters, at least one of the three or more audio objects is associated with the audio object cluster, and for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with the audio object cluster; 17. The apparatus (100) of claim 1, wherein the processor (120) is configured to obtain the plurality of audio object clusters as a function of the distance metric representing the perceptual difference in the spatial characteristics of the audio sound scene.

18. the device (100) further comprises an encoding unit; the encoding unit is configured to generate encoded information encoding the plurality of audio object clusters or the plurality of processed audio objects; and / or The apparatus (100) of any one of claims 1 to 17, wherein the encoding unit is configured to generate encoded information encoding the plurality of audio objects of the audio sound scene and information regarding a perceptual difference between two audio objects of the plurality of audio objects.

19. An apparatus (100) according to claim 18, a decoding unit (210); a signal generator (220); the decoding unit (210) is configured to decode the encoded information to obtain the plurality of audio object clusters or the plurality of processed audio objects, and the signal generator (220) is configured to generate two or more audio output signals in response to the plurality of audio object clusters or in response to the plurality of processed audio objects; and / or the decoding unit (210) is configured to decode the encoded information to obtain a plurality of audio objects of the audio sound scene and to obtain information regarding a perceptual difference between two of the plurality of audio objects, and the signal generator (220) is configured to generate the two or more audio output signals in response to the plurality of audio objects and in response to the perceptual difference between the two audio objects.

20. a decoding unit (210); a signal generator (220); each of a plurality of audio objects of an audio sound scene represents a sound source that is different from any other sound source represented by any other audio object of said plurality of audio objects, or at least two of said plurality of audio objects represent the same sound source at different locations; the decoding unit (210) is configured to decode the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects, the plurality of audio object clusters or the plurality of processed audio objects being dependent on the plurality of audio objects of the audio sound scene and on a distance metric representing a perceptual difference in spatial properties of the audio sound scene, and the signal generator (220) is configured to generate two or more audio output signals in response to the plurality of audio object clusters or in response to the plurality of processed audio objects; and / or a decoder (200) configured to: decode the encoded information to obtain the plurality of audio objects of the audio sound scene; and to obtain information regarding a perceptual difference between two of the plurality of audio objects, the perceptual difference depending on a distance metric; and a signal generator (220) configured to generate the two or more audio output signals according to the plurality of audio objects and according to the perceptual difference between the two audio objects.

21. receiving a plurality of audio objects of an audio sound scene; obtaining information about a perceptual difference between two audio objects of the plurality of audio objects as a function of a distance metric; each of the plurality of audio objects represents a sound source that is different from any other sound source represented by any other audio object of the plurality of audio objects; or at least two of the plurality of audio objects represent the same sound source at different locations; The method, wherein the distance metric represents a perceptual difference in spatial characteristics of the audio sound scene, and / or the method processes the plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects depending on the distance metric.

22. each of the plurality of audio objects represents a sound source that is different from any other sound source represented by any other audio object of the plurality of audio objects, or at least two of the plurality of audio objects represent the same sound source at different locations; decoding the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects, the plurality of audio object clusters or the plurality of processed audio objects being dependent on the plurality of audio objects of the audio sound scene and on a distance metric representing a perceptual difference in spatial characteristics of the audio sound scene; and generating two or more audio output signals in response to the plurality of audio object clusters or in response to the plurality of processed audio objects; and / or 1. A method comprising: obtaining the plurality of audio objects of the audio sound scene; decoding the encoded information to obtain information regarding a perceptual difference between two of the plurality of audio objects, the perceptual difference depending on a distance metric; and generating the two or more audio output signals in response to the plurality of audio objects and in response to the perceptual difference between the two audio objects.

23. A computer program which, when run on a computer or signal processor, implements the method of claim 21 or 22.

Citation Information

Patent Citations

  • Stereo sound signal processing method and apparatus

    JP2003516555A

  • Object Clustering for Rendering Object-Based Audio Content Based on Perceptual Criteria

    JP2016509249A

  • Receiver, content transfer system, and program

    JP2021136465A

  • Audio object clustering based on renderer-aware perceptual difference

    US20190182612A1

  • Directional loudness map based audio processing

    WO2020084170A1