Apparatus and method employing perception-based spatial audio distance metrics
By calculating the perceptual difference between sound sources through the perceptual coordinate system and the spatial masking model, the clustering of audio objects is optimized, and the problem of low computing efficiency in the existing technology is solved, and efficient rendering and high perceptual quality in virtual reality audio applications are achieved.
Patent Information
- Application Number
- CN202380081285.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-29
- Filing Date
- 2023-09-28
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art cannot efficiently calculate the perceived impact of human hearing on the change of sound source position, especially in virtual reality audio applications, real-time rendering computing demands are high, and effective spatial audio distance measurement methods are lacking.
The perception-based spatial audio distance measurement device and method are adopted to calculate the perceived difference between sound sources by perceiving coordinate systems, direction loudness maps and spatial masking models, optimize audio object clustering, reduce the number of audio objects, and improve computing efficiency.
Improves computing efficiency and perceptual quality in virtual reality audio applications, reduces the computing requirements for real-time rendering, while maintaining high perceptual quality.
Smart Images

Figure CN120283419A_ABST
Abstract
Description
[0001] Specification
[0002] The present invention relates to an apparatus and method for employing a distance (distortion) metric for perceptual-based spatial audio.
[0003] Modern audio reproduction systems are capable of providing immersive three-dimensional (3D) sound experiences.
[0004] A common format for 3D sound reproduction is channel-based audio, where individual channels associated with defined loudspeaker positions are generated by multi-microphone recording or studio-based production. Another common format for 3D sound reproduction is object-based audio, which utilizes so-called audio objects that are placed by a producer in a listening room and converted into loudspeaker or headphone signals for playback by a rendering system. Object-based audio offers high flexibility when it comes to the design and reproduction of sound scenes. Note that channel-based audio can be considered a special case of object-based audio, where the sound sources (= objects) are located at fixed positions corresponding to the defined loudspeaker positions.
[0005] To improve the transmission and storage efficiency of object-based immersive sound scenes and to reduce the computational requirements for real-time rendering, it is beneficial, and even necessary, to reduce or limit the number of audio objects. This is achieved by identifying groups or clusters of adjacent audio objects and combining them into a smaller number of sound sources. This process is referred to as object clustering or object consolidation.
[0006] The literature shows that the localization accuracy of human hearing is limited and depends on the position of the sound source (e.g., horizontal localization is more accurate than vertical localization), and an auditory masking effect can be observed between spatially distributed sound sources. By exploiting the limitations of the localization accuracy of human hearing and the auditory masking effect for object clustering, the number of audio objects can be significantly reduced while maintaining high perceptual quality.
[0007] In the prior art, both auditory masking and localization models are known.
[0008] The Directional Loudness Map (DLM) has been introduced in the following articles: C. Avendano, "Frequency-domain source identification and stereo mixing processing for enhancement, suppression, and relocalization applications", Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio, 2003; and P. Delgado, J. Herre, "Objective assessment of spatial sound quality using the directional loudness map", Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, 2019.
[0009] Object clustering algorithms have been introduced in the following articles: J. Herder, "Optimization of Clustering-Based Sound Spatialization Resources Management, 3D Image Journal, 1999; Nicolas Tsingos, Emmanuel Gallo, George Drettakis, "Perceptual Audio Rendering for Complex Virtual Environments," International Conference on Computer Graphics and Interactive Techniques, 2004; Breebaart, Jeroen; Cengarle, Giulio; Lu, Lie; Mateos, Toni; Purnhagen, Heiko; Tsingos, Nicolas, "Spatial Encoding of Complex Program Materials Based on Objects," Journal of Applied Acoustics, Volume 67, Issue 7 / 8, pages 486 - 497, July 2019.
[0010] The prior art includes psychoacoustic models for localization cues, masking, and saliency. However, it does not provide a method for estimating the perceptual impact of changes in the spatial attributes of individual sound sources in a scene relative to the listener's position, which is computationally efficient and suitable for real-time applications such as virtual reality (VR) audio.
[0011] The object of the present invention is to provide an improved concept for distance metrics in spatial audio. The object of the present invention is achieved by the device of claim 1, the decoder of claim 20, the method of claim 21, the method of claim 22, and the computer program of claim 23.
[0012] According to an embodiment, a device is provided. The device includes an input interface for receiving a plurality of audio objects of an audio sound scene. Moreover, the device further includes a processor. Each of the plurality of audio objects represents a (real or virtual) sound source that is different from any other (real or virtual) sound source represented by any other audio object among the plurality of audio objects; or at least two of the plurality of audio objects represent the same (real or virtual) sound source at different positions. The processor is configured to obtain information about the perceptual difference between two audio objects among the plurality of audio objects according to a distance metric, where the distance metric represents the perceptual difference in the spatial attributes of the audio sound scene; and / or the processor is configured to process the plurality of audio objects according to the distance metric to obtain a plurality of audio object clusters or a plurality of processed audio objects.
[0013] Moreover, according to an embodiment, a decoder is also provided. The decoder includes a decoding unit and a signal generator. Each audio object among a plurality of audio objects of an audio sound scene represents a (real or virtual) sound source different from any other (real or virtual) sound source represented by any other audio object among the plurality of audio objects; or, at least two audio objects among the plurality of audio objects represent the same (real or virtual) sound source at different positions. The decoding unit is configured to decode encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects; wherein the plurality of audio object clusters or the plurality of processed audio objects depend on the plurality of audio objects of the audio sound scene and a distance metric representing a perceived difference in the spatial attributes of the audio sound scene; and wherein the signal generator is configured to generate two or more audio output signals based on the plurality of audio object clusters or based on the plurality of processed audio objects; and / or, the decoding unit is configured to decode encoded information to obtain the plurality of audio objects of the audio sound scene and information about the perceived difference between two audio objects among the plurality of audio objects, wherein the perceived difference depends on the distance metric; and wherein the signal generator is configured to generate two or more audio output signals based on the plurality of audio objects and the perceived difference between the two audio objects.
[0014] Moreover, according to an embodiment, a method is provided, including:
[0015] - receiving a plurality of audio objects of an audio sound scene; and
[0016] - obtaining information about the perceived difference between two audio objects among the plurality of audio objects according to a distance metric.
[0017] Each audio object among the plurality of audio objects represents a (real or virtual) sound source different from any other (real or virtual) sound source represented by any other audio object among the plurality of audio objects; or, at least two audio objects among the plurality of audio objects represent the same (real or virtual) sound source at different positions. The distance metric represents a perceived difference in the spatial attributes of the audio sound scene; and / or, the plurality of audio objects are processed according to the distance metric to obtain a plurality of audio object clusters or a plurality of processed audio objects.
[0018] Moreover, according to another embodiment, a method is provided. Each audio object among the plurality of audio objects represents a (real or virtual) sound source different from any other (real or virtual) sound source represented by any other audio object among the plurality of audio objects; or, at least two audio objects among the plurality of audio objects represent the same (real or virtual) sound source at different positions. The method includes:
[0019] - Decode the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects; wherein the plurality of audio object clusters or the plurality of processed audio objects depend on a plurality of audio objects of an audio sound scene and on a distance metric representing a perceptual difference in the spatial attributes of the audio sound scene; and generate two or more audio output signals based on the plurality of audio object clusters or the plurality of processed audio objects; and / or,
[0020] - Decode the encoded information to obtain a plurality of audio objects of an audio sound scene and information on the perceptual difference between two audio objects among the plurality of audio objects, wherein the perceptual difference depends on a distance metric; and generate two or more audio output signals based on the plurality of audio objects and based on the perceptual difference between the two audio objects.
[0021] Moreover, a computer program is provided, each program of which is configured to implement any of the above methods when executed on a computer or a signal processor.
[0022] To predict the perceivable impact of localization changes in a sound scene, according to some embodiments, a perceptual model is provided that represents perceptual differences in a computationally efficient manner. The model can be used to optimize the perceptual quality of object-based audio clustering algorithms and as an objective measurement means to quantify the perceivable differences between different representations of a sound scene.
[0023] According to some embodiments, the perceptual distance metric answers the following questions: If the position of a sound source changes, how perceivable is it? How perceivable is the difference between two different representations of a sound scene? How important is a given sound source in the overall sound scene? (How noticeable would it be if it were removed?)
[0024] For example, according to some embodiments, a psychoacoustic model can include one or more of the following components corresponding to different aspects of human perception, namely a perceptual coordinate system, a 3D direction loudness map, a spatial masking model, and a perceptual distance metric.
[0025] According to some embodiments, a Perceptual Coordinate System (PCS) is provided. Humans' localization accuracy of sound sources varies with different spatial directions. To represent this situation in a computationally efficient way, the Perceptual Coordinate System (PCS) is introduced. To obtain the PCS, the spatial positions are distorted to conform to the non-uniform characteristics of the localization accuracy. Therefore, the distance in the PCS corresponds to the "perceptual distance" between positions, e.g., the number of Just Noticeable Differences (JNDs), rather than the physical distance. This principle is similar to using psychoacoustic frequency scales in perceptual audio coding, e.g., the Bark-Scale or the Equivalent Rectangular Bandwidth Scale (ERB-Scale).
[0026] According to some embodiments, a three-dimensional Direction Loudness Map (3D-DLM) is provided. The core idea of the Direction Loudness Map (DLM) is to find a representation of "how loud the sound is perceived to be coming from a given direction". This concept has been proposed as a one-dimensional method to represent binaural localization in a binaural DLM (Delgado et al., 2019). This concept is now extended to three-dimensional (3D) localization by creating a 3D-DLM on the surface around the listener, which uniquely represents the perceived loudness according to the angle of incidence relative to the listener. It should be noted that the binaural DLM is obtained by analyzing the signals at the ears, while the 3D-DLM is utilized for object-based audio synthesis using the priori known sound source positions and signal attributes.
[0027] In some embodiments, a Spatial Masking Model (SMM) is provided. The monaural time-frequency auditory masking model is a basic element of perceptual audio coding and is usually improved for stereo coding by a binaural (non-)masking model. The Spatial Masking Model extends this concept to immersive audio to incorporate and utilize the masking effect between any sound source positions in 3D.
[0028] According to some embodiments, a perceptual distance metric is provided. It should be noted that the above components can be combined to obtain a perception-based distance metric between spatially distributed sound sources. These components can be used in various applications, e.g., as a cost function in an object clustering algorithm, controlling the bit distribution in a perceptual audio encoder, and obtaining objective quality measurements.
[0029] Embodiments of the present invention will be described in more detail below with reference to the schematic diagrams.
[0030] Figure 1 An apparatus according to an embodiment is shown.
[0031] Figure 2 A decoder according to an embodiment is shown.
[0032] Figure 3 Shows a system according to an embodiment.
[0033] Figure 4 Shows a two - dimensional example for sensing coordinate distortion of a coordinate system according to an embodiment.
[0034] Figure 5 Shows the perceived coordinates obtained by multidimensional scaling through modeling differences in the CIPIC HRTF database according to an embodiment.
[0035] FIG. 6 shows a perceived coordinate system based on a polynomial model according to an embodiment.
[0036] Figure 7 Shows a perceived coordinate system based on an ellipsoid model according to an embodiment.
[0037] Figure 8 Shows an example of synthesizing a one - dimensional direction loudness map based on known object positions and loudness according to an embodiment.
[0038] FIG. 9 shows an example of a 3D direction loudness map synthesized according to known sound source positions according to an embodiment.
[0039] FIG. 10 shows different sampling methods for a unit sphere grid according to an embodiment, where (a) depicts azimuth / elevation sampling and where (b) depicts an icosahedron.
[0040] Figure 11 Shows the calculation of a masking model in perceived coordinates according to an embodiment.
[0041] Figure 1 Shows apparatus 100 according to an embodiment.
[0042] Provides apparatus 100 according to an embodiment.
[0043] The apparatus includes an input interface 110 for receiving a plurality of audio objects of an audio sound scene.
[0044] Moreover, apparatus 100 further includes a processor 120.
[0045] Each audio object among the plurality of audio objects represents a real sound source or a virtual sound source that is different from any other real sound source or virtual sound source represented by any other audio object among the plurality of audio objects; or, at least two audio objects among the plurality of audio objects represent the same real sound source or virtual sound source at different positions. For example, the same real sound source or virtual sound source can be considered at different positions because different time points are considered. Or, the same real sound source or virtual sound source can be considered at different positions because the position before position quantization can be compared with the position after position quantization.
[0046] The processor 120 is configured to obtain information about the perceived difference between two audio objects among a plurality of audio objects according to a distance metric, where the distance metric represents the perceived difference in the spatial attributes of the audio sound scene.
[0047] And / or, the processor 120 is configured to process a plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects according to a distance metric.
[0048] According to an embodiment, the audio sound scene may be, for example, a three-dimensional audio sound scene.
[0049] In an embodiment, the processor 120 may be configured to obtain information about the perceived difference between two audio objects according to a perceived coordinate system; and / or, wherein the processor (120) may be configured to process a plurality of audio objects according to the perceived coordinate system to obtain a plurality of audio object clusters or a plurality of processed audio objects. The distance in the perceived coordinate system represents the perceivable localization difference.
[0050] According to an embodiment, the processor 120 may be configured to obtain information about the perceived difference between two audio objects according to a reversible mapping function; and / or, wherein the processor 120 may be configured to process a plurality of audio objects according to the reversible mapping function to obtain a plurality of audio object clusters or a plurality of processed audio objects. In addition, the processor 120 may also be configured to use the reversible mapping function to convert the coordinates of the physical coordinate system into the coordinates of the perceived coordinate system.
[0051] In an embodiment, the reversible mapping function may depend, for example, on head-related transfer function data.
[0052] According to an embodiment, the processor 120 may be configured to obtain information about the perceived difference between two audio objects according to a spatial masking model of spatially distributed sound sources; and / or, wherein the processor 120 may be configured to process a plurality of audio objects according to the spatial masking model to obtain a plurality of audio object clusters or a plurality of processed audio objects. The spatial masking model may depend, for example, on a masking threshold. The processor 120 may be configured to determine the masking threshold according to an attenuation function and according to one or more distances in the perceived coordinate system.
[0053] In an embodiment, the processor 120 may be configured to determine the masking threshold according to a Gaussian-type attenuation function as an attenuation function and according to an offset of minimum masking.
[0054] According to an embodiment, the processor 120 may be configured to identify one or more inaudible audio objects among a plurality of audio objects.
[0055] In an embodiment, the processor 120 may be configured to obtain information about the perceptual difference between two audio objects based on a perceptual distortion metric; and / or, the processor 120 may be configured to process multiple audio objects based on the perceptual distortion metric to obtain multiple audio object clusters or multiple processed audio objects. Moreover, the processor 120 may be configured to determine the perceptual distortion metric based on the distance in a perceptual coordinate system and based on a spatial masking model.
[0056] According to an embodiment, the processor 120 may be configured to determine the perceptual distortion metric based on the perceptual entropy of one or more audio objects among multiple audio objects.
[0057] In an embodiment, the processor 120 may be configured to determine the perceptual distortion metric based on a first distance between a first audio object of two audio objects among multiple audio objects and the centroid of the two audio objects, and based on a second distance between a second audio object of the two audio objects and the centroid of the two audio objects.
[0058] According to an embodiment, the processor 120 may be configured to obtain information about the perceptual difference between two audio objects based on a three-dimensional direction loudness map; and / or, the processor 120 may be configured to process multiple audio objects based on the direction loudness map to obtain multiple audio object clusters or multiple processed audio objects. The three-dimensional direction loudness map may, for example, depend on the loudness perception related to direction.
[0059] In an embodiment, the processor 120 may be configured to synthesize a direction loudness map on a uniformly sampled grid on the surface around the listener based on the positions and energies of multiple audio objects.
[0060] According to an embodiment, the direction loudness map may depend on the grid and one or more attenuation curves, and the one or more attenuation curves depend on the perceptual coordinate system.
[0061] In an embodiment, the processor 120 may be configured to use the sum of the differences between a three-dimensional direction loudness map and another three-dimensional direction loudness map as a distance metric between an audio sound scene and another audio sound scene.
[0062] According to an embodiment, the distance metric may depend on the three-dimensional direction loudness map and the spatial masking model.
[0063] In an embodiment, the processor 120 may be configured to process a plurality of audio objects to obtain a plurality of audio object clusters. Further, the processor 120 may be configured to obtain a plurality of audio object clusters by associating each of three or more audio objects among the plurality of audio objects with at least one audio object cluster among two or more audio object clusters, such that for each audio object cluster among the two or more audio object clusters, at least one audio object among the three or more audio objects is associated with the audio object cluster, and such that for each audio object cluster among at least one audio object cluster of the two or more audio object clusters, at least two audio objects among the three or more audio objects are associated with the audio object cluster. Additionally, the processor 120 may also be configured to obtain a plurality of audio object clusters based on a distance metric representing a perceived difference in spatial attributes of an audio sound scene.
[0064] According to an embodiment, the apparatus 100 may also include, for example, an encoding unit. The encoding unit may be configured to generate encoding information, for example, that encodes the plurality of audio object clusters or the plurality of processed audio objects; and / or, the encoding unit may be configured to generate encoding information that encodes the plurality of audio objects of the audio sound scene and information about the perceived difference between two audio objects among the plurality of audio objects.
[0065] Figure 2 A decoder 200 is shown according to an embodiment. The decoder 200 includes a decoding unit 210 and a signal generator 220.
[0066] Each audio object among the plurality of audio objects of the audio sound scene represents a real sound source or a virtual sound source that is different from any other real sound source or virtual sound source represented by any other audio object among the plurality of audio objects; or, at least two audio objects among the plurality of audio objects represent the same real sound source or virtual sound source at different positions.
[0067] The decoding unit 210 is configured to decode the encoding information to obtain the plurality of audio object clusters or the plurality of processed audio objects; wherein the plurality of audio object clusters or the plurality of processed audio objects depend on the plurality of audio objects of the audio sound scene and on a distance metric representing a perceived difference in spatial attributes of the audio sound scene; and the signal generator 220 is configured to generate two or more audio output signals based on the plurality of audio object clusters or based on the plurality of processed audio objects.
[0068] And / or, the decoding unit 210 is configured to decode the encoded information to obtain a plurality of audio objects of the audio sound scene and obtain information about the perceived difference between two audio objects among the plurality of audio objects, wherein the perceived difference depends on a distance metric; and, the signal generator 220 is configured to generate two or more audio output signals according to the plurality of audio objects and the perceived difference between the two audio objects.
[0069] Figure 3 illustrates a system according to an embodiment. The system includes Figure 1 the device 100 in
[0070] Figure 1 The device 100 of also includes an encoding unit. The encoding unit is configured to generate encoded information that encodes a plurality of audio object clusters or a plurality of processed audio objects; and / or, the encoding unit is configured to generate encoded information that encodes a plurality of audio objects of the audio sound scene and information about the perceived difference between two audio objects among the plurality of audio objects.
[0071] Moreover, the system further includes a decoding unit 210 and a signal generator 220.
[0072] The decoding unit 210 is configured to decode the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects; and the signal generator 220 is configured to generate two or more audio output signals according to the plurality of audio object clusters or the plurality of processed audio objects.
[0073] And / or, the decoding unit 210 is configured to decode the encoded information to obtain a plurality of audio objects of the audio sound scene and obtain information about the perceived difference between two audio objects among the plurality of audio objects; and the signal generator 220 is configured to generate two or more audio output signals according to the plurality of audio objects and the perceived difference between the two audio objects.
[0074] Specific embodiments are described in detail below.
[0075] According to some embodiments, a perceived distance model is provided.
[0076] The task of the developed perceived distance model is to obtain a distance metric that represents the perceived difference in the spatial attributes of a 3D audio sound scene in a computationally efficient manner. For example, this can be achieved by transforming geometric coordinates in a coordinate system that takes into account the directions related to the localization accuracy of human hearing. In addition, the distance model can include the perceived attributes of the entire scene, which cause localization uncertainty and masking effects.
[0077] According to some embodiments, a perceived coordinate system (PCS) is provided.
[0078] The localization accuracy of human spatial hearing is non-uniform. For example, studies have shown that the localization accuracy in front of the listener is higher than that on the side, the horizontal localization accuracy is higher than the vertical localization accuracy, and the localization accuracy in front of the listener is higher than that behind. This property can be used to optimize the perceptual quality, such as quantization schemes or object clustering algorithms.
[0079] To model the non-uniform property of spatial audio processing, a Perceptual Coordinate System (PCS) is provided according to an embodiment. For example, the PCS can utilize a distorted coordinate system, where the distance in the coordinate system (e.g., Euclidean distance) is modeled as the "perceptual difference" between the corresponding sound source positions, rather than the physical distance between them. In other words, the non-uniform characteristics of perception can be represented by distorting the coordinate system itself, rather than considering the localization accuracy based on absolute positioning. This is similar to using psychoacoustic frequency scales (such as the Bark-Scale or the Equivalent Rectangular Bandwidth Scale (ERB-Scale)) to represent the non-uniformity of frequency resolution in human hearing.
[0080] Figure 4 A two-dimensional example of coordinate distortion for a perceptual coordinate system according to an embodiment is shown. In particular, Figure 4 A two-dimensional perceptual coordinate distortion of sound source positions (points) that are separated by a hypothetically equal perceptual distance on a horizontal plane is shown. More specifically, Figure 4 It is shown that the sound source positions are separated by a perceptually equal distance (e.g., an exemplary JND) in a unit circle in the middle plane. For Figure 4 the geometric coordinates in a), the distance depends on the absolute azimuth angle of the sound source. For Figure 4 the perceptual coordinates in b), the positions have been distorted so that the Euclidean distance between the sound sources is constant.
[0081] For example, the perceptual coordinate system according to an embodiment can approximately perceive the differences between arbitrary sound source positions and obtain updated positions with low computational complexity, such as for fast spatial audio processing algorithms, such as real-time clustering for object-based audio.
[0082] The mapping from geometric coordinates to perceptual coordinates is designed to be unique and invertible, such as a bijective mapping function. For example, all calculations and updates of sound source positions can be performed in the perceptual domain, and the final result can be transformed back to the physical space domain.
[0083] According to an embodiment, a method for deriving PCS based on HRTF data analysis is provided, using a binaural sum and spectral localization cue model and a multidimensional scaling (MDS) method to process pairwise differences. This can generate a mapping of a grid of positions provided by analyzing an HRTF database, which can be used for table look-up and interpolation. For a closed-form representation, the mapping function can be curve-fitted to the analyzed grid data, and a simplified mapping model can be derived therefrom.
[0084] For a general PCS model, the analysis can be computed and averaged using HRTF data from many subjects. Additionally, it should be noted that the analysis methods presented can be computed for a known HRTF data set in a target application, such as a binaural renderer using general or personalized HRTF data.
[0085] Existing models can estimate the localization cues and perceptual differences between sound source positions. However, for spatial audio processing algorithms (such as object clustering), these algorithms need to repeatedly compute the localization model, resulting in low computational efficiency, and this is a major drawback for real-time applications.
[0086] By considering and representing the perceptual model in the analysis and construction steps of PCS, the computationally expensive parts of the model can be computed in an offline preprocessing step, thus obtaining a computationally efficient model suitable for real-time processing. Additionally, using PCS can also directly operate on the sound source positions in the perceptual domain (such as optimizing the cluster centroid positions).
[0087] Furthermore, since PCS can be modeled based on HRTF data analysis, a customized perceptually optimized model can be provided for a target application with a given HRTF data set.
[0088] The "resolution" of the human auditory system varies with changes in azimuth and elevation angles and depends on the absolute position of the sound source.
[0089] The reference model only considers the angle of incidence relative to the listener, such as azimuth and elevation angles, while assuming a constant distance of the sound source (for an extended scheme of the distance model, see below).
[0090] Positions along the interaural axis ("left / right") are determined by binaural cues (interaural coherence (ICC), interaural level difference (ILD), interaural time difference (ITD), interaural phase difference (IPD)), resulting in a so-called cone of confusion (CoC), along which these binaural cues are approximately constant. It should be noted that when the radius is assumed to be constant, the cone simplifies to a "circle of confusion" of a sphere along a given radius.
[0091] Along the CoC, the spectral coloring introduced by the pinna, head, and shoulders can be used as the main cue for elevation localization and elimination of the front / back ambiguity effect. Note that at a given elevation, the spectral filtering of the two ears may not be the same, so additional binaural cues can be introduced.
[0092] To represent this cue separation, a “binaural spherical coordinate system” can be adopted, where the azimuth angle describes the “left / right” position between ±90° along the horizontal plane; the elevation angle describes the “elevation” position within the range of 0° to 360° along the CoC. This coordinate system is a polar coordinate system whose axis of rotation is aligned with the tail position (i.e., the pole is located at the left and right positions of the listener), rather than the vertical polar coordinate system (such as the form where the pole is located above / below the listener in the geographical coordinate system).
[0093] The just noticeable difference (JND) of azimuth angle difference (about 1°) is significantly smaller than that of elevation angle difference (about 4° for noise and up to 10 - 15° for spectrally sparse content). In addition, the localization accuracy also depends on the absolute position. For example, the localization accuracy in front of the listener is higher than that above.
[0094] Therefore, neither the Euclidean distance between Cartesian coordinate positions (e.g., on the unit sphere) nor the angular distance in polar coordinates conforms to the perceived distance.
[0095] Although the position can be represented by a 2D coordinate system (e.g., spanned by azimuth and elevation angles) and the 2D surface (e.g., the unit sphere) can be parameterized, when calculating the distance in the 2D coordinate system, the “wrapping” property of the closed sphere (i.e., 360° = 0°) cannot be reflected. Therefore, a general PCS requires (at least) three dimensions.
[0096] The concept of generating PCS is described below.
[0097] The main target application of PCS is to consistently characterize the JND of localization accuracy at a given position. For example, to determine whether two positions are close enough to be combined without perceiving a change. Therefore, the selected design goal of PCS can be, for example, that the property of having a Euclidean distance of 1 from a given position should always correspond to the JND in the corresponding direction.
[0098] The JND of the elevation angle along the confusion cone can be predicted from the JND to distinguish the spectral differences between HRTFs (see ICASSP19). The JND of the azimuth angle in the horizontal plane can be estimated from the JND of ILD and has been extensively studied experimentally in the literature.
[0099] Based on the JND related to the position, PCS can be constructed as an absolute coordinate system, which is scaled by accumulating the JND between positions. In other words, the Euclidean distance between two arbitrary positions can correspond to the number of accumulated JNDs between the two positions.
[0100] Note that this concept is generally based on the Weber-Fechner law. Although the Weber-Fechner law represents a logarithmic relationship, the position distance is measured in a linear domain. However, the perceptual cues considered (such as ILD and spectral differences, etc.) have been measured in a logarithmic domain. For example, assuming the JND is 1 dB, the PCS distance of 10 JNDs corresponds to 10 dB.
[0101] Based on this principle, according to an embodiment, the perceptual distance (PD) between two given positions = 'number of JNDs', which can be calculated, for example, by measuring the HRTFs of the respective positions. Using the available set of HRTF databases, a complete set of pairwise distances between the measured positions of the given HRTFs can be calculated and averaged over multiple subjects.
[0102] A matrix of pairwise perceptual distances between the grids of the given geometric input positions is thus generated.
[0103] To derive the absolute coordinates from a given set of pairwise differences, a machine learning method using multidimensional scaling (MDS) can be employed. Thus, the coordinates of the selected dimensions (e.g., three dimensions) approximating the given distances can be calculated.
[0104] According to an embodiment, the MDS method can provide a set of PCS positions for the spatial positions of the corresponding HRTF measurements.
[0105] Figure 5 The perceptual coordinates obtained by multidimensional scaling through the modeled differences in the CIPIC HRTF database according to an embodiment are shown.
[0106] In applications where only the grid positions are of interest, the obtained positions can be used as a lookup table. For example, for calculating the distance between any positions, interpolation in the lookup table can be employed.
[0107] According to an embodiment, in order to obtain a continuous, closed-form solution where the PCS coordinates can be inverted to geometric coordinates, a low-dimensional model can be fitted to, for example, the MDS results.
[0108] The preprocessing according to an embodiment is described below, especially the coordinate alignment.
[0109] In such preprocessing, in an embodiment, the MDS coordinates do not inherently match the geometric attributes (such as left-right, front-back) of the input positions.
[0110] Since MDS is based on relative distances, the obtained PCS positions can be mirrored, translated, and rotated without affecting the fit based on the relative distance measurements.
[0111] However, for an intuitive understanding of the coordinate system, for example, the PCS coordinates are preferably aligned as much as possible with the actual spatial positions, for example, clearly corresponding to "left", "right", "front", "top".
[0112] For example, MDS can sort the coordinates according to the variance contribution of the input data set, similar to the energy compression property of principal component analysis (PCA).
[0113] Given that binaural cues have a significant impact on the perceivable differences and have a largely monotonic relationship with the azimuthal position, the first coordinate usually corresponds to the "left / right" axis, although it can be mirrored with respect to the spatial coordinates.
[0114] However, the spectral cues do not have a unique relationship with the elevation position and wraparound occurs. Therefore, the MDS results may be arbitrarily rotated. For example, the coordinates may correspond to an axis from "below the back" to "above the front", and deformation may occur between the coordinates. For example, see Figure 5 the "D shape" of the midplane coordinates in
[0115] Therefore, before curve fitting, the coordinates from the MDS results can be aligned through reflection (such as aligning left / right inversion), translation (such as aligning the front / back hemispheres or the upper / lower hemispheres), and rotation (such as aligning points in the horizontal plane) to correspond to the desired properties of the unit sphere geometric coordinates.
[0116] The following describes a curve fitting method according to an embodiment, especially polynomial nonlinear regression.
[0117] To obtain a continuous mapping function from spatial coordinates to perceptual coordinates, a curve fitting method can be used, for example.
[0118] According to an embodiment, multidimensional nonlinear regression can be used, for example, to fit a polynomial approximation or a spline representation to the MDS results.
[0119] However, since the available positions in the HRTF database are usually sparsely sampled, for example, an appropriate parameterization method needs to be selected to avoid overfitting.
[0120] Moreover, most HRTF databases do not include measurement data in the area below the listener. Therefore, special attention needs to be paid to ensuring good performance in this inferred area. Otherwise, in spline fitting or polynomial fitting, the area below the back may cause a large overshoot.
[0121] To retain the basic model assumptions of binaural cues and spectral cues, a separate fitting method can be adopted. For example, one aspect corresponds to binaural cues, which are clearly separated on the left / right sides and have no "wrapping property", so they can be represented by a single coordinate. For example, on the other hand, it corresponds to monaural spectral cues along the confusion cone, which essentially includes circular wrapping, so the front / back axis and the up / down axis can be jointly fitted to represent the cross-section along the confusion cone.
[0122] To avoid overfitting, a linear model is selected for the first coordinate U (left / right), and second-order polynomials are selected for the second coordinate V and the third coordinate W.
[0123] u p (x) = 27.8x
[0124] ν p (y) = 8.15y 4 -1.75y 3 -3.46y 2 +4.61y - 0.60
[0125] w p (z) = -6.94z 4 +4.03z 3 +3.11z 2 +3.92z - 1.13
[0126] The MDS results (dots) and curve fitting (surface) of the CIPIC HRTF database are shown above.
[0127] Figure 6 shows a perceptual coordinate system based on a polynomial model according to an embodiment, where the surface represents the distorted unit sphere.
[0128] An efficient model fitting method according to an embodiment is described below, especially the linear fitting of an ellipsoid.
[0129] Especially for real-time applications, such as real-time object clustering, a computationally concise and efficient reversible coordinate system needs to be adopted.
[0130] For example, the MDS results and polynomial fitting can be similar to an ellipsoid, except for the "depression" of front / back confusion and the "tail" at the lower rear position near the body.
[0131] As an approximation of a simplified model, an ellipsoid can be adopted, for example.
[0132] For example, this coordinate can be efficiently constructed by scaling the Cartesian coordinates of the unit sphere by an appropriate factor, and can also be easily inverted by inverse scaling.
[0133] Here, the mapping function can be simplified to a scalar scaling of a single coordinate with appropriate weights. For example,
[0134] ·U = c u *X
[0135] ·V = c v *Y
[0136] ·W = c w *Z
[0137] The scale factors can be derived from the MDS results by linear fitting of the respective mapping functions, which can be reduced to scalar weighting of the unit sphere coordinates.
[0138] However, the scale factors of the selected ellipsoid model can be directly fitted to approximate the underlying distance matrix without computing MDS.
[0139] This reduces the computation time and minimizes the approximation error, which would otherwise perform two fitting operations (distance -> MDS -> ellipsoid).
[0140] Figure 7 A coordinate system based on the perceptual ellipsoid model according to an embodiment is shown, where the surface represents a distorted unit sphere.
[0141] The selection of input data for parameter fitting according to an embodiment is described below.
[0142] It should be noted that, generally speaking, for the ellipsoid model, a trade-off needs to be considered when selecting the input position range: for example, the MDS results may exhibit a "tail" effect at lower positions, highlighting the distance between the lower front and lower rear positions. Since these positions are separated by the listener's torso, the torso shadow may provide additional spectral cues between these positions, making them more distinguishable than the front / back at elevated positions.
[0143] However, the ellipsoid model cannot characterize this phenomenon. Therefore, the front / back factor is a compromise between the lower and upper hemispheres due to more significant front / back confusion in the horizontal plane and at elevated positions.
[0144] This can be taken into account when the target application scenario ( = playback system) is known. For example, for an immersive speaker setup, the speaker positions are mainly in the upper hemisphere, so the positions in the lower hemisphere can be omitted (or given lower weights) in the parameter fitting. Conversely, in VR applications, sound source reproduction below the listener is more common, so the positions in the lower hemisphere need to be included in the model fitting.
[0145] For example, the resulting distortion factor may depend on the database, the frequency range analyzed, and / or the input considered. For example, the results of the parametric fitting of the CIPIC HRTF database are: c_u = 28.1, c_v = 5.81, c_w = 8.56. For example, a set of average factors for multiple HRTF databases are: c_u = 25 (left / right), c_v = 6 (front / back), c_w = 5 (up / down).
[0146] For binaural rendering applications that reproduce known HRTFs, the PCS can be directly modeled according to the HRTF in use, rather than being modeled as a general approximation based on a database. For example, in applications where HRTFs can be personalized, the PCS model can be updated when a new HRTF dataset is loaded. Therefore, as described above, the model fitting itself also requires high computational efficiency.
[0147] For more advanced modeling, the PCS can be constructed according to frequency. For example, to reflect the greater HRTF differences caused by high-frequency elevation angles, see Blauert's direction bands. This is particularly relevant for the coordinates representing the spectral cues (V / W). Psychoacoustic experiments in the literature show that the left / right localization of physical sound sources has little relation to frequency. Although the ILD differences are smaller at low frequencies, the ILD / IPD cues become more relevant. Therefore, the frequency-independent left / right axis scaling can be combined with the frequency-dependent cone of confusion scaling.
[0148] For example, the conversion from geometric coordinates to PCS coordinates can be used to transform the positions of spatially distributed sound sources in the domain representing the perceptual attributes of sound source localization in human hearing.
[0149] In the PCS domain, the perceptibility of the difference in sound source positions can be represented by the Euclidean distance between PCS coordinates. This provides a computationally efficient estimation method for the perceptual differences in sound source localization.
[0150] Moreover, the PCS domain can be calibrated: for example, representing 1 JND as a PCS distance of 1. From this, the limits of the localization accuracy at any given position can be estimated. For example, this can be used to control the resolution of the quantization scheme.
[0151] To convert the sound source position given by geometric coordinates (X, Y, Z) to perceptual coordinates (U, V, W), a mapping function can be applied, for example, which is represented in general notation as:
[0152] ·U = f U (X, Y, Z)
[0153] ·V = f V (X, Y, Z)
[0154] ·W = f W (X, Y, Z)
[0155] To transform coordinates from the perceptual domain back to the geometric domain, an inverse mapping function, for example, can be applied, which is represented by the general symbol as:
[0156] · X = f -1 X (U, V, W)
[0157] · Y = f -1 Y (U, V, W)
[0158] · Z = f -1 Z (U, V, W)
[0159] The invertible mapping function allows operations to be directly performed within the perceptual domain, such as operating on the sound source position and tolerance calculation. In this way, efficient computational algorithms based on perception can directly and fully process spatial audio within the perceptual domain. For example, there is no need to repeatedly calculate the perception model. The resulting spatial positions in the perceptual domain can be transformed back to geometric coordinates through the inverse mapping function, for example.
[0160] As described above, a suitable mapping function can be obtained.
[0161] To achieve computational efficiency, a separable ellipsoidal approximation method can be preferably adopted, where the mapping function can be simplified to:
[0162] · U = c u * X
[0163] · V = c v * Y
[0164] · W = c w * Z
[0165] Therefore, the inverse mapping function can be simplified to:
[0166] · Y = U / c u
[0167] · Y = V / c v
[0168] · Z = W / c w
[0169] It should be noted that the ellipsoidal mapping function is valid for positions on the unit sphere and the corresponding ellipsoidal surface. When spatial operations result in positions outside this surface, the positions can be mapped back to the defined surface. For example, by projecting the geometric coordinates onto the unit sphere or by selecting the nearest point on the ellipsoidal surface in the PCS domain.
[0170] A three-dimensional direction loudness map (3D-DLM) according to some embodiments is described below.
[0171] The purpose of DLM is to represent "how much sound comes from a specific direction". In other words, considering the localization accuracy of human hearing, it represents the combined perceived loudness after the superposition of all sound sources in the scene. In an object-based audio environment, the sound source positions and corresponding signal attributes are known. Based on this, DLM can be calculated as the cumulative contribution of all effective sound sources and weighted by a distance-based attenuation function (such as a Gaussian function or a linear attenuation function).
[0172] Figure 8 An example of synthesizing a one-dimensional direction loudness map (1D-DLM) based on known object positions and loudness according to an embodiment is shown. It should be noted that this example illustrates, for example, that the combined loudness generated by the cumulative effect of four closely spaced sound sources on the right is higher than that of a single louder sound source near the center position.
[0173] According to an embodiment, the synthesis of DLM can be extended to localization in three-dimensional space to form a three-dimensional direction loudness map (3D-DLM) by using a sampling grid on the surface (such as a unit sphere) around the listener and calculating the cumulative contribution of all sound sources at each grid point. As shown in the example calculation in FIG. 9, a 3D-DLM is generated therefrom.
[0174] FIG. 9 shows an example of a 3D direction loudness map synthesized from known sound source positions (marked as x) according to an embodiment, where (a) shows the 3D-DLM on the unit sphere and (b) shows the 3D-DLM in the perceptual coordinate system.
[0175] The known binaural one-dimensional DLM represents the perceived loudness based on binaural cues, that is, the "left / right" spatial image.
[0176] However, according to some embodiments, for immersive audio applications, the spatial attributes of the three-dimensional space, such as the elevation angle and the front / back relationship, can also be considered. For example, this can be achieved by using 3D-DLM.
[0177] In addition, the known DLM also requires a scene analysis step, in which the binaural downmix of the entire sound scene is calculated and processed through binaural cue analysis to extract the binaural 1D-DLM. In object-based audio, the sound source positions and signal attributes (such as signal energy) are known in advance.
[0178] According to an embodiment, the 3D-DLM can be directly synthesized based on this information without the computational complexity of calculating the binaural downmix and performing the scene analysis step.
[0179] The following provides a benchmark concept for generating 3D-DLM according to an embodiment.
[0180] For example, the 3D-DLM can be computed on a mesh of the surfaces around the listener, where each point can correspond to a unique spherical coordinate angle, e.g., a uniformly sampled unit sphere. Different embodiments regarding sampling and surface shapes will be described in detail below.
[0181] The energy of each sound source can be computed (as described below) and spread around its position with a given attenuation curve. According to the convention of the 1D DLM, the attenuation curve is modeled as a Gaussian distribution. To reduce the computational complexity, a linear attenuation curve in the log domain can also be used.
[0182] For example, the attenuation can be determined by the Euclidean distance between positions in 3D space, rather than the angular distance or the distance along the sphere / ellipsoid, in order to account for perceptual effects such as front / back ambiguity.
[0183] For example, the energy contribution of each sound source can be computed for each sound source and each mesh point (weighted, e.g., according to the magnitude of the attenuation function), and the energy contributions of each mesh point can be accumulated to compute the Directional Energy Map (DEM).
[0184] This method assumes that the sound sources are uncorrelated. If the sound sources are expected to be correlated, the phantom sound source extraction needs to be performed in a preprocessing step, see below. To account for the increased localization ambiguity due to the phantom sound sources, the attenuation curve can be adjusted to represent a wider spread.
[0185] Based on the sum of the energies at each mesh position, the respective loudness can be computed, e.g., using energy^0.25 = sqrt(sqrt(energy)), which is an approximation with an exponent of 0.23 given in the Zwicker loudness model.
[0186] It should be noted that the summation can be performed in the energy domain, rather than in the loudness domain, because in a real-world playback environment, assuming uncorrelated sound sources, the physical energies of the sound sources will be superimposed on the ear, rather than the perceptual measurement of loudness.
[0187] The spread of the attenuation curve, e.g., the standard deviation of the Gaussian distribution, can be determined based on psychoacoustics, e.g., corresponding to the JND of the localization accuracy.
[0188] To achieve low computational complexity, e.g., for real-time applications, the reference model of the 3D-DLM can be obtained by computing the energy in the time domain, e.g., frame-by-frame, e.g., using the full-band energy. To incorporate the frequency dependence of the human ear's loudness perception, the signal can be pre-filtered, e.g., using A-weighting or K-weighting. Otherwise, e.g., the high energy in the low-frequency region will be overrepresented. The perceptual weighting can be implemented in a computationally efficient way, e.g., using an IIR filter with a relatively low order, e.g., a 7th-order filter for A-weighting.
[0189] Now consider the extended schemes and more embodiments.
[0190] To reduce the computational complexity, the attenuation curve can be truncated. For example, when the tail of the Gaussian distribution is below a given threshold, a more simplified spreading function, such as linear attenuation, can be employed, and the attenuation curve weights for fixed sound source positions corresponding to the speaker positions in a defined configuration (such as 5.1, 7.1+4, 22.2) can be cached and / or pre-computed.
[0191] For advanced application perception models that require a high spectral resolution, a frequency-dependent DLM can be computed: for example, the DLM computation can be performed per spectral band (such as ERB resolution). As an extension, for a frequency-dependent DLM, the spreading factor can also be frequency-dependent to account for the different localization accuracies of human hearing in different frequency regions.
[0192] For example, in stereo / multi-channel production, when a sound source corresponds to two or more channels, the correlation between the sound sources can be considered to generate virtual sound sources. According to an embodiment, for example, the direct signal and the diffuse signal parts can be extracted.
[0193] To this end, the cross-correlation between the channels can be computed.
[0194] When the cross-correlation exceeds a given threshold (such as 0.7), a virtual source can be inserted and the direct / diffuse part decomposition can be performed.
[0195] For example, the position of the virtual source can be calculated based on the energy ratio between the original sound source positions: for example, by weighted averaging of the positions, or by an inverse translation rule, such as the sine rule.
[0196] To account for the reduced localization accuracy of the virtual source, the spreading factor of the spatial attenuation function can be broadened for the virtual source by an appropriate factor. This factor can be a fixed value (such as 2JND), or scaled based on a correlation-based quantity (for example, the higher the correlation, the narrower the spread due to the higher localization accuracy of the virtual source).
[0197] To account for the remaining uncorrelated part in the signal, such as the diffuse part, the overall signal energy can be distributed between the additionally inserted virtual sound source and the original sound source positions according to the correlation coefficient.
[0198] To account for the diffuse nature of the remaining (uncorrelated) signal part, the spreading factor for the original sound source positions can be adjusted, for example, by an appropriate factor. This factor can be a fixed value, such as 2JND, or scaled based on a correlation-based quantity, for example, inversely proportional to the spread of the virtual source. For example, the greater the spread, the higher the correlation, since the remaining part corresponds to a diffuse field rather than a sound source at the original position.
[0199] As an extension considering the "sluggishness" of human hearing in terms of time localization accuracy, a time expansion factor can be used to weight the DLM of the previous frame and add it to the current frame. The time expansion factor can be determined by the time attributes of human hearing and thus needs to be adapted to the frame length and sampling rate.
[0200] The sampling grid of the DLM will now be described according to an embodiment.
[0201] FIG. 10 shows different sampling methods of a unit sphere grid according to an embodiment, where (a) describes azimuth / elevation sampling, and where (b) describes an icosahedron, e.g., https: / / en.wikipedia.org / wiki / Geodesic-_polyhedron, and https: / / medium.com / @qinzitan / mesh-deformation-study-with-a-sphere-ceee37d47e32.
[0202] For numerical calculations, the DLM can be sampled on a grid around the listener. The sampling resolution of the grid is a trade-off between spatial accuracy and computational complexity and thus needs to be optimized according to geometric and perceptual attributes.
[0203] According to an embodiment, the grid for calculating the DLM is generated by uniformly sampling the azimuth and elevation along the unit sphere, e.g., 360×180 = 64,800 points.
[0204] However, the density of the spherical coordinate system increases significantly near the poles, resulting in non-uniform oversampling and causing an unnecessary large number of points. This will lead to a significant increase in computational complexity. In addition, subsequent algorithms (such as Gaussian mixture models) may be affected by factors such as non-uniform sampling and the increased density of pole values.
[0205] For example, methods for uniformly sampling the sphere (such as for computer graphics) can use "geodesic polyhedra", "geographic spheres", or "icosahedra", which are obtained by subdividing the icosahedron.
[0206] For example, to maintain a resolution of about 1°, 5 subdivided icosahedra can be used to generate a grid with 10242 points (about 16% of the azimuth / elevation uniform grid). This can significantly reduce the computational and storage requirements while maintaining a comparable perceptual quality.
[0207] In many applications, even a lower subdivision is sufficient, e.g., just using 3 subdivisions (corresponding to 642 points).
[0208] The spatial masking model (SMM) according to some embodiments will be described below.
[0209] Figure 11 Shows the masking model calculation performed in perceptual coordinates according to an embodiment.
[0210] The masking effect produced by human hearing between loud and soft sounds is an important aspect of audio coding psychoacoustic models. Existing models typically estimate the masking threshold for mono or stereo coding. However, for immersive audio applications, the masking effect between any sound source positions is of interest. Subjective listening experiments can usually only measure the masking effect for a limited number of position pairs. To estimate the masking effect between any sound source positions in immersive audio, a general spatial masking model (SMM) according to an embodiment is provided. Subjective experiment results show that: masking differences may be related to the available localization cues and the differences, and thus related to the localization accuracy. PCS and 3D-DLM are introduced as models for extending localization accuracy and loudness perception. Based on this, a spatial masking model applicable to any sound source position is obtained, where the distance between sound sources can be calculated in the PCS domain to estimate the difference in localization cues, and a spatial attenuation curve is applied to model the unmasking effect. Figure 11 Shows the position of the masker with an azimuth angle of -30° on the median plane. It can be seen that due to the smaller distance in the PCS representation, the strong masking effect at symmetric positions in the front and back is stronger, while the masking effect of the left-right difference is significantly weaker. In this case, the interaural cues play a greater role in unmasking.
[0211] For example, the masking model for perceptual audio coding may need to be related to time and frequency in order to control the spectral shaping of the introduced quantization noise. Conversely, object clustering affects the spatial position of the sound source. For example, changing the overall position of the sound source is essentially a "full-band" operation.
[0212] It should be recognized that the masking between individual sound sources still depends on factors such as frequency. However, changing the spatial position of the sound source changes the localization cues rather than introducing additional noise. In other words, the masking model for localization changes may have different requirements from the masking model for additional signals (such as quantization noise).
[0213] For real-time applications, a computationally efficient model may be required. Therefore, a simplified full-band masking model based on the time-varying signal energy can be used for aspects such as object clustering.
[0214] Considering that human auditory sensitivity is frequency-dependent, frequency weighting such as A-weighting can be adopted, which can be achieved by performing time-domain filtering using a relatively short filter (such as a 7th-order IIR filter).
[0215] Note that operations that can remove signal components, such as removing inaudible sound sources in object-based audio, preferably use a frequency-dependent masking model, as this is more similar to the use cases of adding signal components (quantization noise) or removing signal components (quantizing to zero) in perceptual audio coding.
[0216] Now, an overview of the masking model is provided according to some embodiments.
[0217] For example, the SMM can assume a maximum masking threshold at the position of the masker, such as in-masker masking. The masking threshold can be weighted by an attenuation function according to the spatial distance to reduce for spatially separated sound sources.
[0218] The attenuation function can be a linear attenuation in the logarithmic domain ('decibels per distance') or a Gaussian-type attenuation curve, so that the calculations of the DLM can be reused or shared to save computational complexity.
[0219] In addition to distance-dependent masking, a position-independent offset can be added to the masking threshold, which depends on the total energy of all sound sources in the scene and is weighted by a maximum dereverberation factor (e.g., -15 dB). This is done to reflect that there is always a certain amount of residual masking between sound sources (psychoacoustic experiments have found that the maximum level of binaural / spatial dereverberation on headphones is about 15 dB BMLD).
[0220] In other words, the masking effect between spatially separated sound sources may never drop to zero because the amount of spatial dereverberation is limited (the maximum BMLD found in the literature is about 15 dB in headphone experiments). However, spatial masking experiments show that the initial slope of the dereverberation of spatially separated sound sources is still quite steep, so the attenuation curve needs to reflect this characteristic. Therefore, especially when using a Gaussian model to plot the attenuation curve, the curve should not be chosen to be very wide to accommodate the maximum dereverberation at the maximum distance, but should be locally steep enough around the sound source and then only drop to a given minimum value rather than zero.
[0221] Similar to the positioning accuracy, there may also be differences in spatial dereverberation between horizontal separation and vertical separation. To reflect this characteristic, the distance of the attenuation curve in the SMM can be calculated using PCS instead of Euclidean distance. Thus, the interaural (left / right) difference causes more dereverberation than the elevation difference, and a large amount of masking between front / back symmetric sound sources is also retained.
[0222] Now, the detailed calculations according to specific embodiments are described.
[0223] The local energy spread map M local (k) of the sound source is represented by an object with index k. For example, it can be based on the A-weighted object energy of all object indices i The sum is calculated, and the weighted Gaussian decay function depends on the Euclidean distance D in the PCS PCS (k,i) and an (adjustable) expansion factor s, for example, as follows:
[0224]
[0225] It should be noted that, unlike the parameterization of the normal distribution density function, the decay function in the masking model is not normalized. For example, the expansion factor only scales the width of the distribution rather than the height (and thus the sum of the source contributions). In other words, the larger the expansion factor, the stronger the "masking ability", which is similar to the expansion function in frequency-domain masking (especially in the scenario of calculating the DLM, it should be noted that it should not be confused with the overall loudness affecting the scenario).
[0226] Optionally, according to a specific embodiment, the expansion factor 2s is selected for all sound sources 2 = 5 (considering the standard deviation of the corresponding normal distribution obtained whose expansion width is 1 to 2 JND), or s = 6 can be selected for a wider expansion (for example, 2s 2 = 72).
[0227] Furthermore, optionally, according to another specific embodiment, in order to further improve the accuracy of the model, for example, when a suitable detector can be used in a given embodiment, the expansion factor can depend on the signal characteristics and masking ability (such as noise-like, tonal, transient, etc.) of a single object.
[0228] In addition to local masking, the minimum residual masking between sound sources (and vice versa, corresponding to the maximum binaural unmasking) can be taken as the global minimum M of the energy expansion map min incorporated.
[0229] According to an embodiment, the minimum masking can be direction-independent. In other words, it can reflect the overall sound energy in the scene that limits the human ear's resolution ability. The sum of the signal energies can be weighted and estimated based on the worst-case BLMD value of 15 dB [Blauert] found in the literature.
[0230]
[0231] Alternatively, it can also be calculated as the sum of the local energy masking maps at the sound source positions, such as the energy of the sound source plus the local contributions of adjacent sound sources. This simulates that a group of sound sources closer together has a stronger masking ability. Moreover, when the expansion factor is modeled as signal-dependent, this also simulates that sound sources with a wider expansion factor have a greater impact on the overall (minimum) masking.
[0232]
[0233] For example, the combined masking threshold T k A 20 dB upper limit estimate of the masking threshold can be used (from Hellman72 for the case of tonal masking noise at 60 dB SPL), and the calculation formula is as follows:
[0234]
[0235] It should be noted that calculating the combined masking as the sum of the local masking and the global masking is beneficial for maintaining the smoothness of the Gaussian decay and achieving saturation at the offset. Optionally, it can also be implemented as the maximum operation between M min and M local In this way, the evaluation of the Gaussian function can be cut off at a large distance (using the energy-based calculation of M min ), thus saving computational complexity.
[0236] The following describes the perceptual distance metric according to some embodiments.
[0237] In the context of audio object clustering, the core problem of the perceptual distance metric may be "When we merge multiple objects, how perceptible is it?", which leads to more detailed questions "If we merge two candidate objects, how far will each object move? In the context of the entire scene, how large is the sound difference brought about by this position change?"
[0238] PCS provides a model for the perceptibility of the change in the spatial position of the sound source, and SMM provides a model for the audibility of the sound source under the masking effect of the entire sound scene. According to the embodiments, these models can be combined to obtain a measurement of the perceptual distance between two sound sources (such as objects in the context). Therefore, the perceptual distance between two objects can be calculated based on the distance between the objects in PCS (considering the localization difference) and weighted by the estimated perceptual correlation of the objects (relative to the masking effect in the overall sound scene).
[0239] An important issue for this distance metric is robustness and numerical stability. Since real-world implementations only operate under limited numerical precision calculations, this metric can be designed to be robust against numerical inaccuracies and edge cases (such as values close to or equal to zero). For example, when the number of effective sound sources changes over time, some audio scene representations may always include metadata and audio for the maximum number of effective objects (similar to a fixed number of tracks in a DAW). This will result in "non-effective" objects, where the PCM data of the signal only includes digital zeros or noise (LSB noise) generated by numerical inaccuracies (which may be worse). The preferred method may be to detect and remove these non-effective objects in a preprocessing screening step before actual clustering. However, not all applications support this operation.
[0240] Thus, according to an embodiment, the distance metric can be designed to be robust to small / zero energy by adding appropriate offset values where necessary (e.g., without explicitly detecting such a situation).
[0241] Now, the definition of the perceptual distance model is described according to an embodiment.
[0242] In the field of perceptual audio coding, perceptual entropy (PE) [JJ88] is a well-known measure for evaluating "how much of the audible signal content is related to the masking threshold". Here, a simplified and computationally efficient estimate of the PE for each object can be made, for example, using the full-band energy and the masking threshold obtained from the SMM (the SMM can apply frequency weighting before energy calculation to account for the frequency dependence of human hearing).
[0243] It should be noted that, as described above, the object position is independent of frequency. Therefore, frequency-related calculations can improve the accuracy of the masking model but do not increase the degree of freedom of the clustering algorithm.
[0244] For example, the PE of the k-th object can be calculated by the following formula:
[0245]
[0246] For example, the distance metric D Perc (k, l) between two object indices k and l can be calculated using the distance D PCS (k, l) in the PCS as follows:
[0247]
[0248]
[0249] The model parameters can be optionally set as thr offs = 33 [dB] and d offs = 0.1 [bit].
[0250] Now, the derivation process of the model formula is described according to an embodiment.
[0251] To avoid numerical instability for small energies, an offset can be added to the object energy.
[0252] The offset can be scaled according to the sum (or maximum energy) of the total energy, because the energy range can span multiple orders of magnitude according to the PCM data scaling. For example, for the application of pre-normalization scaling, a constant value can be used. As the offset, a worst-case estimated masking threshold of -33 dB can be selected (e.g., assuming a tonal masking noise of 27 dB + an average BMLD of 6 dB), and for example, a constant offset ε depending on the calculation precision can be added (such as ε = FLT_MIN = 1e-37).
[0253]
[0254] When two objects are merged, a new centroid c can be determined. k,l . Here, it is assumed that the centroid position is the average position weighted by the energy of the object. Therefore, the centroid position depends on the ratio of the object energies. In other words, when the energy of the second object is greater, the position change of the first object may be greater, and vice versa. Therefore, the perceived position distance D k,l from the first candidate object of index k to the candidate centroid c PCS (k, c k,l ) can be estimated by the ratio of the energy E l ′ of the second object to the sum of the energies of the two objects, for example, as follows:
[0255]
[0256] To account for the perceived relevance of the objects in the masking background of the entire sound scene, the estimated position distance can be weighted according to the PE of the objects, etc.:
[0257] D′ Perc (k, l) = PE(k)D PCS (k, c k,l ) + PE(l)D PCS (l, c k,l )
[0258] For example, the unit of the distance metric can be considered as "bit × JND". For example, in this metric, assuming that the distances of two pairs of candidate objects are the same, combining the objects with lower PE will incur a smaller cost.
[0259] To avoid the instability of objects with negligible PE or energy, an offset related only to the distance between the objects can be added. The offset factor can be set to 0.1 [bit] (corresponding to the PE of a signal just above the masking threshold (about 0.3 dB = 10 log 10 2 -0.1 ).
[0260] D Perc (k, l) = PE(k)D PCS (k, c) + PE(l)D PCS (l, c) + 0.1D PCS (k, l)
[0261] Expanding and simplifying the above equation gives:
[0262]
[0263] As an extension, a radius-perceived distance according to an embodiment is described.
[0264] For example, PCS can simulate the differences in spectrum and binaural cues by only considering the angle of incidence of the sound source relative to the listener.
[0265] However, in various applications (such as VR binaural rendering), the distance between the listener and the sound source is also important.
[0266] Therefore, according to an embodiment, additional coordinates can be introduced in PCS, which are modeled to reflect the JND in radius changes.
[0267] Although the absolute distance has been proven to be less accurate, the relative change in distance is easier to detect, for example, based on three main cues: sound level change, direct-to-reverberation ratio, and Doppler effect.
[0268] Regarding the sound level change, the intensity of the sound source may decay with the increase in distance (under free-field conditions, SPL decreases as 1 / r^2; while in a closed environment, due to the presence of reverberation, the decrease in sound level is usually smaller).
[0269] Regarding the direct-to-reverberation ratio (DRR), in a reverberant environment, a distant sound source may have more reverberation.
[0270] Regarding the Doppler effect, when the relative distance between the listener and the sound source changes at a given speed, the pitch of the sound source changes due to the Doppler effect.
[0271] The cues from the sound level change and the DRR change are related. In a reverberant environment, the sound level change weakens, but the DRR change may bring additional cues.
[0272] Therefore, an environment-independent radial distance model can be adopted based on the sound level related to the distance. According to the psychoacoustic literature, the JND for detecting the sound level change is 1 dB. Therefore, the gain related to the radius can be calculated as the ratio to the reference radius and transformed into the logarithmic domain. Therefore, in this model, a 1 dB relative gain difference directly corresponds to a 1 JND perceivable distance change.
[0273] For example, the radial distance coordinate can be calculated according to the following formula:
[0274] d_r = 20 * log10(r / 0.2 + FLT_MIN)
[0275] (assuming the reference radius is 0.2 meters, for example, a position close to the head)
[0276] For example, when the distance between a sound source and a listener changes over time, the Doppler effect may cause a pitch shift. For a given frequency f, source velocity v_S, listener velocity v_L, and speed of sound c, the resulting frequency can be:
[0277] f’ = f * (c + v_L) / (c + v_S)
[0278] The signs of v_S and v_L depend on whether they are moving towards or away from each other.
[0279] Note that this formula depends on two absolute velocities, not just the relative velocity. However, for the case of v << c, considering only the relative velocity can simplify the formula.
[0280] The human ear is quite sensitive to relative changes in frequency and can detect changes of about 5 cents (5% of a semitone).
[0281] For example, the relative pitch change can be derived from the Doppler effect formula:
[0282] deltaPitch = 12 * log2((c + v_L) / (c + v_S)) [semitone]
[0283] = 1200 * log2((c + v_L) / (c + v_S)) [cent]
[0284] Solving the Doppler pitch shift formula for the JND of 5 cents gives a JND of approximately 1 m / s for the movement of the listener and the sound source (low speed).
[0285] Therefore, the velocity component of the PCS can be directly modeled according to the relative velocity between the listener and the sound source, and 1 m / s is equivalent to 1 JND.
[0286] More embodiments are provided below.
[0287] According to a first embodiment, a distance metric representing the perceptual difference in the spatial attributes of a 3D audio sound scene is provided.
[0288] According to a second embodiment, a Perceptual Coordinate System (PCS) is provided, where a geometric distance (such as Euclidean distance or angular distance) represents the perceivable localization difference according to the first embodiment.
[0289] According to a first variant of the second embodiment, a parameterized reversible mapping function is provided for transforming geometric (physical) coordinates in the Perceptual Coordinate System of the second embodiment.
[0290] According to a specific variant of the second embodiment, a method for obtaining the mapping parameters of the first variant of the second embodiment based on the analysis of HRTF data is provided.
[0291] According to the third embodiment, a spatial distribution sound source masking model using a spatial decay curve based on the perceived distance of the second embodiment is provided.
[0292] In a first variant of the third embodiment, a masking model of the third embodiment is provided, which uses a Gaussian decay curve with a minimum masking offset.
[0293] In a second variant of the third embodiment, a method for calculating the masking effect of an entire sound scene by summing the mono masking thresholds weighted by the position-dependent masking model of the third embodiment is provided.
[0294] In a third variant of the third embodiment, a perceptual entropy (PE) calculated based on the masking model of the third embodiment and the sound source energy is provided to estimate the contribution of the sound source to the sound scene information.
[0295] In a fourth variant of the third embodiment, a method for identifying inaudible sound sources for screening out irrelevant audio objects is provided.
[0296] According to the fourth embodiment, a perceptual distortion measure (PDM) of changes in the spatial attributes of a 3D audio sound scene based on the perceived distance of the second embodiment and the spatial masking model of the third embodiment is provided.
[0297] According to a first variant of the fourth embodiment, a distortion measure of the position change of a single sound source is provided as a weighted combination of the PCS distance from the masking model and the PE.
[0298] According to a second variant of the fourth embodiment, a distortion measure for the combination of two or more sound sources is provided based on the estimated centroid position and the weighted sum of individual distortion measures.
[0299] According to the fifth embodiment, a three-dimensional direction loudness map (3D-DLM) representing the direction-related loudness perception is provided.
[0300] According to a first variant of the fifth embodiment, the 3D-DLM is synthesized for known sound source positions and energies on a grid uniformly sampled around the listener.
[0301] According to a second variant of the fifth embodiment, a 3D-DLM based on the grid and decay curve in the PCS coordinates of the second embodiment is provided.
[0302] According to a third variant of the fifth embodiment, the sum of the differences between two 3D-DLMs is provided as a distortion measure of the two sound scene representations of the first embodiment.
[0303] According to a fourth variant of the fifth embodiment, a combination of the 3D-DLM and the masking model in the third embodiment is provided as a PE-based difference measure between two sound field representations.
[0304] Although some of the technical solutions are described in the context of an apparatus, it is evident that these solutions also constitute a description of the corresponding method, where the modules or components correspond to method steps or features of method steps. Similarly, the technical solutions described by method steps also represent a description of the corresponding modules or items or features of the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware apparatuses, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0305] According to the requirements of certain implementations, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware, or at least partially in software. This implementation may be performed using a digital storage medium, such as a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory, which has electronically readable control signals stored thereon, and these electronically readable control signals can cooperate (or be capable of cooperating) with a programmable computer system to perform their respective methods. Therefore, the digital storage medium may be computer-readable.
[0306] Some embodiments according to the present invention include a data carrier having electronically readable control signals, which can cooperate with a programmable computer system to perform one of the above methods.
[0307] Generally, embodiments of the present invention may be implemented as a computer program product having program code, and when the computer program product runs on a computer, the program code can be used to perform one of the above methods. The program code may be stored on a machine-readable carrier.
[0308] Other embodiments include a computer program stored on a machine-readable carrier for performing one of the above methods.
[0309] In other words, therefore, the method embodiments of the present invention are computer programs, and the computer programs have program code for performing one of the above methods when the computer programs run on a computer.
[0310] Therefore, another method embodiment of the present invention is a data carrier (or a digital storage medium, or a computer-readable medium) on which a computer program for performing one of the above methods is recorded. The data carrier, the digital storage medium, or the recording medium is generally tangible and / or non-transitory.
[0311] Therefore, another embodiment of the method of the present invention is a data stream or a signal sequence representing a computer program for performing one of the above methods. The data stream or the signal sequence may be configured to be transmitted through a data communication connection, such as through the Internet.
[0312] Another embodiment includes a processing device, such as a computer or a programmable logic device, configured to or adapted to perform one of the above methods.
[0313] Another embodiment includes a computer on which a computer program for performing one of the above methods is installed.
[0314] Another embodiment of the present invention includes an apparatus or system for transmitting (e.g., electronically or optically) a computer program for performing one of the above methods to a receiver. The receiver can be a computer, a mobile device, a storage device, etc. The apparatus or system may include a file server for transmitting the computer program to the receiver.
[0315] In some embodiments, a programmable logic device (e.g., a field programmable array) can be used to perform some or all of the functions of the above methods. In some embodiments, a field programmable array can cooperate with a microprocessor to perform one of the above methods. Generally, the above methods are preferably performed by any hardware device.
[0316] The above apparatus can be implemented using a hardware device, or a computer, or a combination of a hardware device and a computer.
[0317] The above methods can be performed using a hardware device, or a computer, or a combination of a hardware device and a computer.
[0318] The above embodiments are only illustrative of the principles of the present invention. It should be understood that modifications and variations of the above arrangements and details are obvious to those skilled in the art. Therefore, it is intended to be limited only by the scope of the following claims, rather than by the specific details shown by the description and explanation of the above embodiments.
Claims
1. A device (100), comprising: An input interface (110) for receiving a plurality of audio objects of an audio sound scene; And A processor (120), Wherein, each audio object among the plurality of audio objects represents a sound source different from any other sound source represented by any other audio object among the plurality of audio objects; or, wherein at least two audio objects among the plurality of audio objects represent the same sound source at different positions; Wherein, the processor (120) is configured to obtain information on the perceived difference between two audio objects among the plurality of audio objects according to a distance metric, wherein the distance metric represents the perceived difference in the spatial attributes of the audio sound scene; and / or, Wherein, the processor (120) is configured to process the plurality of audio objects according to the distance metric to obtain a plurality of audio object clusters or a plurality of processed audio objects.
2. The device (100) according to claim 1, Among them, The audio sound scene is a three-dimensional audio sound scene.
3. The device (100) according to claim 1 or 2, Among them, The processor (120) is configured to obtain the information on the perceived difference between two audio objects according to a perceived coordinate system; and / or, wherein the processor (120) is configured to process the plurality of audio objects according to the perceived coordinate system to obtain the plurality of audio object clusters or the plurality of processed audio objects, Wherein the distance in the perceived coordinate system represents a perceivable positioning difference.
4. The device (100) according to claim 3, Among them, The processor (120) is configured to obtain information on the perceived difference between two audio objects according to a reversible mapping function; and / or, wherein the processor (120) is configured to process the plurality of audio objects according to the reversible mapping function to obtain the plurality of audio object clusters or the plurality of processed audio objects; and Wherein, the processor (120) is configured to use the reversible mapping function to convert the coordinates of the physical coordinate system into the coordinates of the perceived coordinate system.
5. The device (100) according to claim 4, Among them, The reversible mapping function depends on head-related transfer function data.
6. The device (100) according to any one of claims 3 to 5, Among them, The processor (120) is configured to obtain information on the perceived difference between two audio objects according to a spatial masking model of spatially distributed sound sources; and / or, wherein the processor (120) is configured to process the plurality of audio objects according to the spatial masking model to obtain the plurality of audio object clusters or the plurality of processed audio objects, Wherein the spatial masking model depends on a masking threshold; Wherein, the processor (120) is configured to determine the masking threshold according to an attenuation function and according to one or more distances in the perceived coordinate system.
7. The device (100) according to claim 6, Among them, The processor (120) is configured to determine the masking threshold according to a Gaussian-type attenuation function as the attenuation function and according to an offset of minimum masking.
8. The apparatus (100) according to claim 6 or 7, Among them, wherein the processor (120) is configured to identify one or more inaudible audio objects among the plurality of audio objects.
9. The apparatus (100) according to any one of claims 6 to 8, Among them, wherein the processor (120) is configured to obtain information on a perceptual difference between two audio objects according to a perceptual distortion metric; and / or, wherein the processor (120) is configured to process the plurality of audio objects according to the perceptual distortion metric to obtain the plurality of audio object clusters or the plurality of processed audio objects; and wherein the processor (120) is configured to determine the perceptual distortion metric according to the distance in the perceptual coordinate system and according to the spatial masking model.
10. The apparatus (100) according to claim 9, Among them, wherein the processor (120) is configured to determine the perceptual distortion metric according to the perceptual entropy of one or more audio objects among the plurality of audio objects.
11. The apparatus (100) according to claim 10, Among them, wherein the processor (120) is configured to determine the perceptual distortion metric according to a first distance between a first audio object among two audio objects in the plurality of audio objects and the centroid of the two audio objects, and according to a second distance between a second audio object among the two audio objects and the centroid of the two audio objects.
12. The apparatus (100) according to any one of claims 3 to 11, Among them, wherein the processor (120) is configured to obtain information on a perceptual difference between two audio objects according to a three-dimensional directional loudness map; and / or, wherein the processor (120) is configured to process the plurality of audio objects according to the directional loudness map to obtain the plurality of audio object clusters or the plurality of processed audio objects, wherein the three-dimensional directional loudness map depends on direction-dependent loudness perception.
13. The apparatus (100) according to claim 12, Among them, wherein the processor (120) is configured to synthesize the directional loudness map on a uniformly sampled grid on the surface around the listener according to the positions and energies of the plurality of audio objects.
14. The apparatus (100) according to claim 12 or 13, Among them, wherein the directional loudness map depends on the grid and one or more attenuation curves, and the one or more attenuation curves depend on the perceptual coordinate system.
15. The apparatus (100) according to any one of claims 12 to 14, Among them, wherein the processor (120) is configured to use the sum of the differences between the three-dimensional directional loudness map and another three-dimensional directional loudness map as a distance metric between the audio sound scene and another audio sound scene.
16. The apparatus (100) according to any one of claims 12 to 15 and further according to claim 6, Among them, wherein the distance metric depends on the three-dimensional directional loudness map and the spatial masking model.
17. The apparatus (100) according to any one of the above claims, Among them, wherein the processor (120) is configured to process the plurality of audio objects to obtain the plurality of audio object clusters; Wherein, the processor (120) is configured to obtain the plurality of audio object clusters by associating each of three or more audio objects among the plurality of audio objects with at least one audio object cluster among two or more audio object clusters, such that for each audio object cluster among the two or more audio object clusters, at least one audio object among the three or more audio objects is associated with the audio object cluster, and such that for each audio object cluster among at least one audio object cluster among the two or more audio object clusters, at least two audio objects among the three or more audio objects are associated with the audio object cluster; Wherein, the processor (120) is configured to obtain the plurality of audio object clusters according to a distance metric representing the perceptual difference among the spatial attributes of the audio sound scene.
18. The apparatus (100) according to one of the above claims, Among them, The apparatus (100) further includes an encoding unit, Wherein, the encoding unit is configured to generate encoding information that encodes the plurality of audio object clusters or the plurality of processed audio objects; and / or, Wherein, the encoding unit is configured to generate encoding information that encodes the plurality of audio objects of the audio sound scene and information about the perceptual difference between two audio objects among the plurality of audio objects.
19. A system, comprising: The apparatus (100) according to claim 18; A decoding unit (210); And A signal generator (220), Wherein, the decoding unit (210) is configured to decode the encoding information to obtain the plurality of audio object clusters or the plurality of processed audio objects; and wherein the signal generator (220) is configured to generate two or more audio output signals according to the plurality of audio object clusters or the plurality of processed audio objects; and / or, Wherein, the decoding unit (210) is configured to decode the encoding information to obtain the plurality of audio objects of the audio sound scene and obtain information about the perceptual difference between two audio objects among the plurality of audio objects; and wherein the signal generator (220) is configured to generate the two or more audio output signals according to the plurality of audio objects and the perceptual difference between the two audio objects.
20. A decoder (200), comprising: A decoding unit (210); And A signal generator (220), Wherein, each of the plurality of audio objects of the audio sound scene represents a sound source of any other sound source different from any other representation of the plurality of audio objects; or, at least two of the plurality of audio objects represent the same sound source at different positions; Wherein, the decoding unit (210) is configured to decode the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects; wherein, the plurality of audio object clusters or the plurality of processed audio objects depend on the plurality of audio objects of the audio sound scene and on a distance metric representing a perceived difference in the spatial attributes of the audio sound scene; and wherein the signal generator (220) is configured to generate two or more audio output signals based on the plurality of audio object clusters or based on the plurality of processed audio objects; and / or, Wherein, the decoding unit (210) is configured to decode the encoded information to obtain the plurality of audio objects of the audio sound scene and to obtain information about the perceived difference between two audio objects among the plurality of audio objects, wherein the perceived difference depends on a distance metric; and wherein the signal generator (220) is configured to generate the two or more audio output signals based on the plurality of audio objects and based on the perceived difference between the two audio objects.
21. A method, comprising: Receiving a plurality of audio objects of an audio sound scene; And Obtaining information about the perceived difference between two audio objects among the plurality of audio objects according to a distance metric, Wherein, each audio object among the plurality of audio objects represents a sound source of any other sound source different from any other audio object among the plurality of audio objects; or, wherein at least two audio objects among the plurality of audio objects represent the same sound source at different positions; and Wherein, the distance metric represents a perceived difference in the spatial attributes of the audio sound scene; and / or, processing the plurality of audio objects according to the distance metric to obtain a plurality of audio object clusters or a plurality of processed audio objects.
22. A method, wherein, Each audio object among the plurality of audio objects represents a sound source of any other sound source different from any other audio object among the plurality of audio objects; Or, at least two audio objects among the plurality of audio objects represent the same sound source at different positions; the method comprises: Decoding the encoded information to obtain a plurality of audio object clusters or a plurality of processed audio objects; wherein, the plurality of audio object clusters or the plurality of processed audio objects depend on the plurality of audio objects of the audio sound scene and on a distance metric representing a perceived difference in the spatial attributes of the audio sound scene; and generating two or more audio output signals based on the plurality of audio object clusters or based on the plurality of processed audio objects; and / or, Decoding the encoded information to obtain the plurality of audio objects of the audio sound scene and to obtain information about the perceived difference between two audio objects among the plurality of audio objects, wherein the perceived difference depends on the distance metric; and generating the two or more audio output signals based on the plurality of audio objects and based on the perceived difference between the two audio objects.
23. A computer program for implementing the method according to claim 21 or 22 when executed on a computer or a signal processor.