Metadata preserving audio object clustering
The proposed audio object clustering method classifies and allocates audio objects based on metadata to preserve artistic intent, addressing the loss of metadata in conventional methods and improving rendering accuracy in systems with fewer channels.
Patent Information
- Application Number
- JP2025178178
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-01-06
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-27
AI Technical Summary
Conventional audio object clustering methods fail to preserve metadata such as zone mask, snapping, and rendering mode, leading to a loss of artistic intent and suboptimal rendering results, especially when downmixing to systems with fewer channels.
A method and system for audio object clustering that classifies audio objects into categories based on metadata to be preserved, assigns a predetermined number of clusters to these categories, and allocates audio objects within those clusters, ensuring that metadata like zone mask, snapping, and rendering mode are maintained during the clustering process.
Preserves metadata during audio object clustering, resulting in improved rendering accuracy and artistic intent preservation, particularly in systems with limited channels, enhancing the audio experience.
Smart Images

Figure 2026012835000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority from Chinese Patent Application No. 201410765578.6, filed December 11, 2014, and U.S. Provisional Patent Application No. 62 / 100,183, filed January 6, 2015, the contents of each of which are incorporated herein by reference in their entirety.
[0002] technology FIELD The exemplary embodiments disclosed herein relate generally to audio content processing, and more particularly to methods and systems for audio object clustering that allow metadata to be preserved. [Background technology]
[0003] The advent of object-based audio has significantly increased the amount of audio data and the complexity of rendering this data within high-end playback systems. For example, a movie soundtrack may contain many different sound elements, dialogue, noises, and sound effects, corresponding to the images on the screen. These soundtracks also emanate from different locations on the screen and combine with background music and ambient effects to create an overall auditory experience. Accurate reproduction requires that sounds be reproduced in a manner that corresponds as closely as possible to what is shown on the screen in terms of source location, intensity, motion, and depth. Object-based audio represents a significant improvement over traditional channel-based audio systems, which route audio content in the form of speaker feeds to individual speakers within a listening environment and are therefore relatively limited in terms of the spatial reproduction of specific audio objects.
[0004] The introduction of digital cinema and the development of three-dimensional ("3D") content have created new standards for sound, such as the incorporation of multi-channel audio, allowing greater creativity for content creators and a more immersive, realistic auditory experience for audiences. Expanding beyond traditional speaker feeds and channel-based audio as a means of delivering spatial audio is crucial. Furthermore, there has been significant interest in model-based audio descriptions that allow listeners to select a desired playback configuration by using audio rendered specifically for the chosen configuration. Spatial presentation of sound utilizes audio objects. An audio object is an audio signal with associated parametric source descriptions, such as apparent source position (e.g., 3D coordinates), apparent source width, and other parameters. As a further advancement, next-generation spatial audio (also known as "adaptive audio") formats are being developed by including the blending of audio objects with traditional channel-based speaker feeds (audio beds) along with positional metadata for the audio objects.
[0005] As used herein, the term "audio object" refers to a discrete audio element that exists in the sound field for a defined duration, while the term "audio bed" or "bed" refers to an audio channel that is intended to be reproduced at predefined, fixed speaker positions.
[0006] Some soundtracks may have several (e.g., 7, 9, or 11) bed channels containing audio. Furthermore, based on the capabilities of the authoring system, there may be dozens or even hundreds of individual audio objects that are combined during rendering to create a spatially diverse and immersive audio experience. In other distribution and transmission systems, there may be a large enough available bandwidth to transmit all audio beds and objects with little or no audio compression. However, in some cases, such as Blu-ray Disc, broadcast (cable, satellite, terrestrial), mobile (3G and 4G), and over-the-top (OTT or Internet) distribution, there may be significant limitations in the available bandwidth to digitally transmit all of the bed and object information generated at the time of authoring. Audio coding methods (lossless or lossy) may be applied to the audio to reduce the required bandwidth, but audio coding may not be sufficient to reduce the bandwidth required to transmit audio over very limited networks, such as mobile 3G and 4G networks.
[0007] Several conventional methods have been developed to reduce the number of input objects to a smaller set of output objects through clustering. Generally, some clustering process requires metadata such as size, zone mask, and snapping to be pre-rendered into the internal channel layout. The clustering of audio objects is based solely on the spatial position of the audio objects, and the output objects contain only positional metadata. This type of output object may not work well for some playback systems, as the loss of metadata may violate the expected artistic intent.
[0008] The subject matter discussed in the Background section should not be assumed to be prior art merely because of its disclosure in the Background section. Similarly, it should not be assumed that the problems mentioned in or related to the subject matter of the Background section have been previously recognized in the prior art. The subject matter in the Background section is merely representative of various approaches, which may themselves be exemplary embodiments. Summary of the Invention [Problem to be solved by the invention]
[0009] To address the above and other potential problems, the illustrative embodiments propose a method and system for audio object clustering in which metadata is preserved. [Means for solving the problem]
[0010] In one aspect, exemplary embodiments provide a method for audio object clustering in which metadata is preserved. The method includes classifying a plurality of audio objects into categories based on information to be preserved in metadata associated with the plurality of audio objects. The method further includes assigning a predetermined number of clusters to the categories and allocating audio objects in each of the categories to at least one of the clusters based on the assignment. Embodiments in this regard also include corresponding computer program products.
[0011] In another aspect, exemplary embodiments provide a system for audio object clustering in which metadata is stored, the system including an audio object classification unit configured to classify a plurality of audio objects into categories based on information to be stored in metadata associated with the plurality of audio objects, a cluster assignment unit configured to assign a predetermined number of clusters to the categories, and an audio object allocation unit configured to assign audio objects in each of the categories to at least one of the clusters based on the assignment.
[0012] From the above description, it can be understood that, according to the exemplary embodiments disclosed herein, input audio objects are classified into corresponding categories depending on the information to be stored in the metadata, such that different metadata to be stored or unique combinations of metadata to be stored are associated with different categories. After clustering, audio objects in a certain category are less likely to be mixed with audio objects associated with different metadata. In this regard, the metadata of the audio objects can be preserved after clustering. Other advantages achieved by the exemplary embodiments will be apparent from the following description. [Brief explanation of the drawings]
[0013] The above and other objects, features and advantages of the embodiments will become more apparent through the following detailed description taken in conjunction with the accompanying drawings, in which several exemplary embodiments are shown by way of illustration and not limitation. [Figure 1] 1 is a flowchart of a method for audio object clustering with metadata preservation according to an example embodiment. [Figure 2]FIG. 1 is a schematic diagram for an audio object clustering process according to an example embodiment. [Figure 3] 1 is a block diagram of a system for audio object clustering with metadata preservation according to an example embodiment. [Figure 4] 1 is a block diagram of an exemplary computer system suitable for implementing the embodiments. Throughout the drawings, like or corresponding reference numerals refer to like or corresponding parts. DETAILED DESCRIPTION OF THE INVENTION
[0014] The principles of the exemplary embodiments will now be described with reference to various exemplary embodiments illustrated in the drawings. It should be understood that the illustrations of these embodiments are merely to enable those skilled in the art to better understand and further implement the exemplary embodiments, and are not intended to limit the scope in any way.
[0015] As mentioned above, due to limitations in encoding / decoding rates and transmission bandwidth, clustering can reduce the number of audio objects used to generate adaptive audio content. In addition to metadata describing its spatial location, audio objects typically have other metadata describing their attributes, such as size, zone mask, snapping, and content type. Each of these attributes describes the artistic intent of how the audio object should be processed when rendered. However, in some conventional methods, only the positional metadata remains after the audio objects are clustered. Other metadata is sometimes pre-rendered into internal channel layouts, such as 7.1.2 or 7.1.4 systems, but this does not work well for all systems. The artistic intent of the audio object may be broken when rendered, especially when the audio object is downmixed to, for example, a 5.1 or 7.1 system.
[0016] Take the "Zone Mask" metadata as an example. It has multiple modes, each of which defines an area where an audio object should not be rendered. One mode of the Zone Mask is "No Sides," which specifies that the side speakers should be masked when rendering that audio object. If an audio object at spatial location z=1 is rendered to a 5.1 system using the "No Sides" metadata using traditional clustering methods, the side speakers may be activated in the 5.1 rendering because the sound from the ceiling speakers may be convoluted to the sides. To address this issue, the "Zone Mask" metadata should be preserved during the clustering process so that it can be correctly handled by the audio renderer.
[0017] In another example, dialogue objects may be expected to be separated from other objects after clustering. This can have many benefits for subsequent audio object processing. For example, in subsequent audio processing such as dialogue enhancement, the separated dialogue object clusters can be easily enhanced by simply applying gain(s). Otherwise, it can be very difficult to separate dialogue objects when they are mixed with other objects in a cluster. In dialogue replacement applications, the dialogues in each language may be completely separated from each other. For such purposes, in the clustering process, dialogue objects should be preserved and assigned to separate, distinct clusters.
[0018] Additionally, audio objects may be associated with metadata describing their rendering mode, for example, to render as left-sum / right-sum (Lt / Rt) when processed in a headphone renderer or as binaural with head-related transfer functions (HRTFs). These rendering modes are also expected to be preserved after clustering to yield the best rendering results.
[0019] Therefore, in order to achieve a better audio experience, it is desirable to allow metadata to be preserved in audio object clustering.The exemplary embodiments disclosed herein propose a method and system for object clustering in which metadata is preserved.
[0020] Reference is first made to Figure 1, which depicts a flowchart of a method 100 for audio object clustering with metadata preservation according to an exemplary embodiment.
[0021] In S101, a plurality of audio objects are classified into categories based on information to be stored in metadata associated with the plurality of audio objects. These audio objects are provided as input, and there may be tens, hundreds, or even thousands of input audio objects.
[0022] As used herein, information to be stored in the metadata associated with each audio object may indicate the processing intent when the audio object is rendered. The information may describe how the audio object should be processed when rendered. In some embodiments, the information may include one or more of the audio object's size information, zone mask information, snap information, content type, or rendering mode. The size information may be used to indicate the spatial area or volume occupied by the audio object. The zone mask information indicates the mode of the zone mask, which defines the area in which the audio object should not be rendered. For example, the zone mask information may indicate modes such as "no sides," "surround only," or "front only." The snap information indicates whether the audio object should be panned directly to the nearest speaker.
[0023] It should be noted that while some examples of information to be stored in the metadata are described, other information contained in the metadata (non-limiting examples include spatial location, spatial width, etc.) may also be considered in audio object classification, depending on the preferences of the user or audio engineer. In some embodiments, all information in the metadata associated with an audio object may be considered.
[0024] The number of categories may depend on the information in the metadata of those audio objects and may be one or more. In some embodiments, audio objects with no information to be preserved may be classified into one category, and audio objects with different information to be preserved may be classified into a different category. That is, depending on the different information to be preserved, corresponding audio objects are classified into different categories. Alternatively, a category may represent a unique combination of different information to be preserved in the metadata. All other audio objects without interesting information may, in some cases, be included in one category or multiple categories. The scope of the example embodiments is not limited in this respect.
[0025] Categories may be assigned manually, automatically, or a combination thereof. For example, a user or audio engineer may label audio objects associated with different types of metadata with different flags, and the labeled audio objects may then be classified into different categories according to the flags. As another example, information to be stored in the metadata may be automatically identified. A user or audio engineer may pre-configure their preferences or expectations, such as separating dialogue objects, separating different dialogue languages, and / or separating different modes of zone masks. Depending on the pre-configuration, audio objects may be classified into different categories.
[0026] Suppose there are O audio objects. In the classification process, the information to be stored in the metadata of the audio objects may be derived from (1) manual labels of the metadata given by user input, such as zone masks or snaps or content type or language labels, and / or (2) automatic identification / labeling of the metadata, such as, but not limited to, content type identification information. The number N of possible categories may be determined according to the derived information, with each category consisting of a unique combination of information to be stored. After classification, each audio object will have an associated category identification n o It can have.
[0027] Referring to Figure 2, a schematic diagram of audio object clustering is shown. As shown in Figure 2, multiple input audio objects are classified into five categories, categories 0 to 4, based on the information to be preserved in the metadata. An example of the categories may be given as follows: Category 0: all audio objects without any information to be saved; Category 1: Music objects, no zone mask; Category 2: Sound effect objects, with zone mask "Surround only"; Category 3: English dialogue objects; Category 4: Spanish dialogue objects with "Ahead Only" zone mask.
[0028] An input audio object may contain one or more frames. A frame is the processing unit for audio content, and its duration may vary and depend on the configuration of the audio processing system. Because the audio objects to be classified may change over time for different frames, and their metadata may also change, the value of the number of categories may also change over time. Categories representing different types of information to be preserved may be predefined by the user or by default, and in that case, input audio objects in one or more frames may be classified into the predetermined categories based on that information. In subsequent processing, categories with classified audio objects may be considered, and those without audio objects may be ignored. For example, in Figure 2, if there are no audio objects with no information to be preserved, the corresponding category 0 may be omitted. The number of audio objects classified into each category may change over time.
[0029] In S102, a predetermined number of clusters are assigned to the categories. The predetermined number may be greater than 1 and may depend on the transmission bandwidth and the encoding / decoding rate of the audio processing system. There may be a trade-off between the transmission bandwidth (and / or the encoding rate and / or the decoding rate) and the error criterion of the output audio objects. For example, the predetermined number may be 11 or 16. Other values, such as 5, 7, or 20, may also be determined, and the scope of the exemplary embodiment is not limited in this respect.
[0030] In some embodiments, the predetermined number may remain constant within the same processing system, while in other embodiments, the predetermined number may vary for different audio files to be processed.
[0031] In exemplary embodiments disclosed herein, audio objects are first classified into categories according to metadata in S101, such that each category may represent a different piece of information or a unique combination of different pieces of information to be stored. The audio objects in these categories may then be clustered in a subsequent process. There may be various approaches to assigning / allocating a predetermined overall number of clusters to categories. In some exemplary embodiments, the overall number of clusters is predetermined and fixed, so that the number of clusters to be assigned to each category can be determined before clustering the audio objects. Some exemplary embodiments will now be discussed.
[0032] In an exemplary embodiment, cluster assignment may depend on the importance of the plurality of audio objects. Specifically, the predetermined number of audio objects from the plurality of audio objects may be first determined based on the importance of each audio object relative to other audio objects, and then a distribution of the predetermined number of audio objects among the categories may be determined. The predetermined number of clusters may be assigned to the categories correspondingly according to the distribution.
[0033] The importance of each audio object may be related to one or more of the type of content, the local loudness level, or the energy level of that audio object. An audio object with high importance may represent that the audio object is perceptually prominent among the input audio objects, for example, due to its local loudness level or energy level. In some use cases, one or more types of content may be considered important, and the corresponding audio object may then be given high importance. For example, dialogue objects may be assigned higher importance. It should be noted that there are many other ways to determine or define the importance of each audio object. For example, the importance level of some audio objects may be specified by a user. The scope of the illustrative embodiments is not limited in this respect.
[0034] Assume that the predetermined overall number of clusters is M. In the first step, up to M most important audio objects among the input audio objects are selected. Since all input audio objects are classified into corresponding categories in S101, in the second step, the distribution of the M most important audio objects among the categories may be determined. Based on how many of the M audio objects are distributed in a category, the same number of clusters may be assigned to that category.
[0035] Referring to Figure 2, for example, 11 of the most important audio objects (shown as circles 201) are determined from a plurality of input audio objects (shown as a collection of circles 201 and 202). After classifying all the input audio objects into five categories, categories 0 to 4, it can be seen from Figure 2 that four of the most important audio objects are classified into category 0, three of the most important audio objects are classified into category 1, one of the most important audio objects is classified into category 2, two of the most important audio objects are classified into category 3, and one of the most important audio objects is classified into category 4. As a result, as shown in Figure 2, categories 0 to 4 are assigned 4, 3, 1, 2, and 1 cluster, respectively.
[0036] It should be noted that the above example of importance criteria according to this exemplary embodiment of the exemplary embodiments may not be very strict. That is, it is not necessary that the most important audio objects are selected. In some embodiments, an importance threshold may be configured. The predetermined number of audio objects may be randomly selected from among audio objects whose importance is higher than a threshold.
[0037] In addition to importance criteria, cluster assignment may also be performed based on reducing the overall spatial distortion for the category, i.e., the predetermined number of clusters may be assigned to the category based on reducing or even minimizing the overall spatial distortion for the category.
[0038] In an exemplary embodiment, the overall spatial distortion for a category may comprise a weighted sum of the individual spatial distortions of those categories. The weight of a corresponding category may represent the importance of that category or the importance of the information associated with that category to be preserved. For example, a category with higher importance may have a larger weight. In another embodiment, the overall spatial distortion for the categories may comprise the largest spatial distortion among the individual spatial distortions of those categories. It should be considered that it is not necessary to select only the largest, and in some embodiments, other spatial distortions among the categories, such as the second largest spatial distortion, the third largest spatial distortion, etc., may also be considered as the overall spatial distortion.
[0039] The spatial distortion for each category may be represented by the distortion level of the audio objects contained in that category, and the distortion level of each audio object may be measured by the difference between its original spatial location and its clustered location. Generally, the clustered location of an audio object depends on the spatial location of the cluster(s) to which it is assigned. In this sense, the spatial distortion for each category is related to the original spatial location of each audio object in that category and the spatial location of the cluster(s). The original spatial location of the audio object may be included in the audio object's metadata and may consist of, for example, three Cartesian coordinates (or similarly, for example, polar coordinates, cylindrical and spherical coordinates, homogeneous coordinates, line number coordinates, etc.). In one embodiment, to calculate the spatial distortion for each category, the reconstructed spatial location of each audio object in the category may be determined based on the spatial location of the cluster(s). The spatial distortion for each category may then be calculated based on the distance between the original spatial location of each audio object in the category and the reconstructed spatial location of the audio object. The reconstructed spatial location of an audio object is the spatial location of the audio object represented by one or more corresponding spatial clusters. One exemplary approach for determining the reconstructed spatial location is described below.
[0040] To obtain the overall spatial distortion, the spatial distortion for different numbers of clusters may first be calculated for each category. There are many approaches to determining the spatial distortion for a category of audio objects. One approach is given below as an example. It should be noted that other existing ways of measuring the spatial distortion of an audio object (and therefore of a category) may also be applied.
[0041] For category n, spatial location
number
number
number
number
[0042] Once the spatial distortion for each category is obtained, in one embodiment, the overall spatial distortion for the categories may be determined as a weighted sum of the individual spatial distortions of the categories, as described above. For example, the overall spatial distortion may be determined as follows:
number
[0043] In another embodiment, the overall spatial distortion for the categories may be determined as the maximum spatial distortion among the individual spatial distortions of those categories. For example, the overall spatial distortion may be determined as follows:
number
number
[0044] The input audio objects are generally within one frame of the audio signal. Due to the typical dynamic nature of audio signals, and considering that the number of audio objects varies in each category, the number of clusters assigned to each category may typically vary over time. Since the varying number of clusters for each category may cause some instability issues, a modified spatial distortion that takes into account cluster number consistency is utilized in the cost metric. As a result, the cost metric may be defined as a function of time. Specifically, the spatial distortion for each category is further based on the difference between the number of clusters assigned to that category in the current frame and the number of clusters assigned to that category in the previous frame. In this regard, the overall spatial distortion in equation (2) may be modified as follows:
number
number
[0045] If the number of clusters assigned to a category changes in the current frame compared to the previous spatial distortion, the modified spatial distortion may be increased to prevent the change in the number of clusters. In one embodiment, f(D n (M n ),M n ,M n ') may be determined as follows:
number
[0046] Since reducing the number of clusters in a category is more likely to introduce spatial instability than increasing the number of clusters, in another embodiment, f(D n (M n ),M n ,M n ') may be determined as follows:
number
[0047] In the above description, with regard to cluster assignment based on reducing the overall spatial distortion, determining the optimal number of clusters for each category may involve a large amount of computational effort. To efficiently determine the number of clusters for each category, in one embodiment, an iterative process is proposed. That is, the optimal number of clusters for each category is estimated by maximizing the cost reduction in each iteration of the cluster assignment process. Thereby, the overall spatial distortion for the category may be iteratively reduced or even minimized.
[0048] By iterating from 1 to a predetermined number of clusters M, at each iteration, one or more clusters are assigned to the category that needs them most. The overall spatial distortion at the (m-1)th and mth iterations is denoted as Cost(m-1) and Cost(m). At the mth iteration, one or more new clusters may be assigned to category n* that can most reduce the overall spatial distortion. Thus, n* may be determined by expanding or maximizing the reduction in overall spatial distortion. This may be expressed as follows:
number
[0049] For an overall spatial distortion obtained by a weighted sum of all spatial distortions of the categories, the iterative process may be based on the difference between the spatial distortion for a category in the current iteration and the previous iteration. In each iteration, at least one cluster may be assigned to a category such that, when the category is assigned with the at least one cluster, its spatial distortion in the current iteration is significantly lower (according to a first predetermined level) than its spatial distortion in the previous iteration. In one embodiment, the at least one cluster may be assigned to the category that has the most reduced spatial distortion when assigned with the at least one cluster. For example, in this embodiment, n* may be determined as follows:
number
[0050] With the overall spatial distortion determined as the maximum spatial distortion among all categories, the iterative process may be based on the amount of spatial distortion for a category in a previous iteration. In each iteration, at least one cluster may be assigned to a category that had spatial distortion higher than a second predetermined level in the previous iteration. In one embodiment, the at least one cluster may be assigned to the category with the maximum spatial distortion in the previous iteration. For example, in this embodiment, n* may be determined as follows:
number
[0051] Note that the decisions given in equations (9) and (10) may be used jointly in one iterative process. For example, in one iteration, equation (9) may be used to assign new cluster(s) in that iteration. In another iteration, equation (10) may be used to assign other new cluster(s).
[0052] Two ways of cluster assignment are described above: one based on the importance of audio objects, and the other based on reducing overall spatial distortion. Additionally or alternatively, user input can also be used to guide cluster assignment. This can greatly improve the flexibility of the clustering process, as users may have different requirements for different content for different use cases. In some embodiments, cluster assignment may be further based on one or more of: a first threshold for the number of clusters to be assigned to each category, a second threshold for spatial distortion for each category, or the importance of each category relative to other categories.
[0053] The first threshold may be predefined for the number of clusters to be assigned to each category. The first threshold may be a predetermined minimum or maximum number of clusters for each category. For example, a user may specify that a category should have a certain minimum number of clusters. In this case, at least the specified number of clusters should be assigned to that category during the assignment process. If a maximum threshold is set, at most the specified number of clusters can be assigned to that category. The second threshold may be set to ensure that spatial distortion for a category is reduced to a reasonable level. The importance of each category may also be specified by the user or may be determined based on the importance of the audio objects classified in that category.
[0054] In some cases, the spatial distortion for a category may be high after cluster assignments are made, which may introduce audible artifacts. To address this issue, in some embodiments, at least one audio object in a category may be reclassified to another category based on the spatial distortion for that category. In one exemplary embodiment, if the spatial distortion for one of the categories is higher than a predetermined threshold, some audio objects within that category may be reclassified to another category until the spatial distortion is less than (or equal to) the threshold. In some examples, audio objects may be reclassified to a category containing audio objects with no information to be preserved in metadata, such as category 0 in FIG. 2. In some embodiments, where cluster assignment is based on minimizing overall spatial distortion in a successive iterative process, object reassignment also involves minimizing the overall spatial distortion for that category in each iteration until the spatial distortion criterion is met for that category. n (I C n (1), C n (2),…,C n (M n )}) may be an iterative process in which audio objects are reclassified.
[0055] Due to the typically dynamic nature of audio signals, the importance or spatial location (and therefore spatial distortion) of audio objects change over time. As a result, cluster assignments may be time-varying, in which case the number of clusters assigned to each category may change over time. In this case, the category identity associated with cluster m may change over time. In particular, cluster m may represent one language (e.g., Spanish) during the first frame, while the category identity, and therefore the language, may change (e.g., English) for the second frame. This is in contrast to legacy channel-based systems, in which the language is statically bound to the channel rather than changing dynamically.
[0056] Cluster assignment in S102 was described above.
[0057] Returning to reference to FIG. 1, at S103, audio objects in each of the categories are allocated to at least one of the clusters based on the allocation.
[0058] In the following description, two approaches are provided for clustering audio objects after they are classified into categories in S101 and each category is assigned a cluster in S102.
[0059] In one approach, audio objects in each category may be assigned to at least one of the clusters assigned to one or more of those categories based on reducing the distortion cost associated with those categories. That is, due to limitations on the number of clusters assigned to each category, some leakage across clusters and categories is allowed to reduce the distortion cost and avoid artifacts for complex audio content. This approach may be referred to as fuzzy category clustering. In this fuzzy category clustering approach, an audio object may be softly partitioned into different clusters in different categories using the benefits and corresponding costs. During the clustering process, the distortion cost is expected to be minimal with respect to the overall spatial distortion and the penalties or inconsistencies of assigning objects within a category to clusters of different categories. Therefore, there is a trade-off between the cluster budget and the complexity of the audio content. The fuzzy category clustering approach may be suitable for audio objects with metadata such as zone masks and snaps, for which strict separation from other metadata is not required. The fuzzy category clustering approach may be described as follows:
[0060] In a fuzzy category clustering approach, the number of clusters assigned to each category may be determined in S102 based on the importance of the audio objects or based on minimizing overall spatial distortion. For importance-based cluster assignment, there may be some categories to which no clusters are assigned. In these cases, a fuzzy category clustering approach may be applied when clustering the audio objects, since the objects may be softly clustered into cluster(s) of other categories. There may not be a necessary correlation between the approach applied in the cluster assignment stage and the approach applied in the audio object clustering stage.
[0061] In the fuzzy categorical clustering approach, the distortion cost is calculated by: (1) the original spatial location of each audio object;
number
number
number
number
[0062] In some embodiments, the gain g o,m teeth,
number
number
number
[0063]
number
number
[0064] In some embodiments, the cost function is:
number
[0065] The cost function may typically be minimized under some additional criterion. When allocating audio signals, one criterion may be to maintain the total amplitude or energy of the input audio objects. For example,
number
[0066] In the following, we may discuss the cost function E. By minimizing the cost function, the gain g o,m can be determined.
[0067] The cost function, as mentioned above, is
number
number
[0068] The cost function is
number
number
number
number
number
number
number
number
number
number
number
[0069] Based on these four terms in the cost function, the payoff g o,m The gain g can be determined. o,m An example of the calculation for is given below. It should be noted that other methods of calculation are possible.
[0070] The gain g of the oth audio object for each of the M clusters o,m may be written as a vector:
number
number
number
number
number
[0071] Audio Object o and n m The second term E represents the discrepancy between C may be reformulated as follows:
number
[0072] The third term E represents the sum of the gains of the audio objects and a deviation of +1. N may be reformulated as follows:
number
[0073] The fourth term, E, represents the distance between the original spatial position of the audio object and the reconstructed spatial position. P may be reformulated as follows:
number
number
number
number
number
[0074] The oth audio object is divided into M clusters and the gain vector is determined as
number
[0075] The reconstructed spatial positions of the audio objects may be obtained by equation (17) when the gain vectors are determined. In this regard, the process of determining the gains may also be applied in cluster assignment based on minimizing the overall spatial distortion described above to determine the reconstructed spatial positions, and thus the spatial positions of each category.
[0076] It should be noted that a quadratic polynomial is used as an example to determine the minimum in the cost function, and in other exemplary embodiments many other exponent values may be used, such as 1, 1.5, 3, etc.
[0077] A fuzzy category clustering approach for audio object clustering was described above. In another approach, audio objects in each category may be assigned to at least one of the clusters assigned to that category based on reducing the spatial distortion cost associated with that category. That is, cross-category leakage is not allowed. Audio object clustering is performed within each category, and audio objects cannot be grouped into clusters assigned to other categories. This approach may be referred to as a hard category clustering approach. In some embodiments using this approach, an audio object may be assigned to two or more of the clusters assigned to the category corresponding to that audio object. In a further embodiment, cross-cluster leakage is not allowed, and an audio object may be assigned to only one of the clusters assigned to the corresponding category.
[0078] The hard category clustering approach may be suitable for some specific applications, such as dialogue replacement or dialogue enhancement, which require audio objects (dialogue objects) to be separated from one another.
[0079] In a hard category clustering approach, audio objects in a given category cannot be clustered into one or more clusters in other categories, so it is expected that each category will have been assigned at least one cluster in the previous cluster assignment. To this end, in some embodiments, cluster assignment by minimizing the overall spatial distortion described above may be more preferable. In other embodiments, importance-based cluster assignment may also be used when hard category clustering is applied. Additional criteria may be used in cluster assignment to ensure that each category is assigned at least one cluster, as discussed above. For example, a minimum threshold of clusters or a minimum threshold of spatial distortion for each category may be utilized.
[0080] Because categories represent the same type of metadata, within a category, audio objects may be clustered into only one cluster or into multiple clusters in one or more exemplary embodiments. For example, as shown in FIG. 2, audio objects in category 1 may be clustered into one or more of clusters 4, 5, or 6. In scenarios where audio objects are clustered into multiple clusters within a category, the corresponding gain may also be determined to reduce or even minimize the distortion cost associated with that category (this may be similar to that described with respect to the fuzzy category clustering approach). The difference is that the determination is performed within a single category. In some embodiments, each input audio object may be allowed to be clustered into only one cluster assigned to that category.
[0081] Two approaches to audio clustering are discussed above. It should be noted that these two approaches can be used separately or in combination. For example, after audio object classification in S101 and cluster assignment in S102, for some of the categories, a fuzzy category clustering approach may be applied to cluster the audio objects within them; for the remaining categories, a hard category clustering approach may be applied. That is, some leakage across categories may be acceptable within some categories, and leakage across categories is not acceptable for other categories.
[0082] After the input audio objects are assigned to clusters, for each cluster, the audio objects may be combined to obtain a clustered audio object, and the metadata of the audio objects in each cluster may be combined to obtain the metadata of the clustered audio object. The clustered audio object may be a weighted sum of all audio objects in a cluster using corresponding gains. The metadata of the clustered audio object may be corresponding metadata representing its category in some examples, or may be metadata of any audio object or the most important audio object within the cluster or its category in other examples.
[0083] Since all input audio objects are classified into corresponding categories based on the information to be preserved in the metadata before audio object clustering, different metadata to be preserved or unique combinations of metadata to be preserved are associated with different categories. After clustering, audio objects in a certain category are less likely to be mixed with audio objects associated with different metadata. In this regard, the metadata of the audio objects can be preserved after clustering. Furthermore, spatial distortion or distortion cost is taken into account during the cluster assignment and audio object allocation processes.
[0084] 3 illustrates a block diagram of a system 300 for audio object clustering with metadata storage according to an example embodiment. As illustrated in FIG. 3, the system 300 includes an audio object classification unit 301 configured to classify a plurality of audio objects into categories based on information to be stored in metadata associated with the plurality of audio objects. The system 300 further includes a cluster assignment unit 302 configured to assign a predetermined number of clusters to the categories, and an audio object allocation unit 303 configured to assign audio objects in each of the categories to at least one of the clusters based on the assignment.
[0085] In some embodiments, the information may include one or more of audio object size information, zone mask information, snap information, content type, or rendering mode.
[0086] In some embodiments, the audio object classification unit 301 may be further configured to classify audio objects with no information to be preserved into one category; and classify audio objects with different information to be preserved into a different category.
[0087] In some embodiments, the cluster allocation unit 302 may further comprise: an importance-based determination unit configured to determine the number of audio objects from the plurality of audio objects based on the importance of each audio object relative to the other audio objects; and a distribution determination unit configured to determine a distribution of the number of audio objects among the categories. In these embodiments, the cluster allocation unit 302 may be further configured to allocate the number of clusters to the categories based on the distribution.
[0088] In some embodiments, the cluster assignment unit 302 may be further configured to assign the predetermined number of clusters to a category based on reducing an overall spatial distortion for the category.
[0089] In some embodiments, the overall spatial distortion for the category may comprise a maximum spatial distortion among the individual spatial distortions of the category or a weighted sum of the individual spatial distortions of the category. The spatial distortion for each category may be associated with the original spatial location of each audio object in that category and with the spatial location of at least one of the clusters.
[0090] In some embodiments, the reconstructed spatial position of each audio object may be determined based on the spatial positions of the at least one cluster, and the spatial distortion for each category may be determined based on the distance between the original spatial position of each audio object in that category and the reconstructed spatial position of that audio object.
[0091] In some embodiments, the multiple audio objects may be within one frame of the audio signal, and the spatial distortion for each category may be further based on the difference between the number of clusters assigned to that category in the current frame and the number of clusters assigned to that category in the previous frame.
[0092] In some embodiments, the cluster assignment unit 302 may be further configured to iteratively reduce the overall spatial distortion for the categories based on at least one of: the amount of spatial distortion for a category in a previous iteration or the difference between the spatial distortion for a category in a current iteration and a previous iteration.
[0093] In some embodiments, the cluster assignment unit 302 may be further configured to assign the predetermined number of clusters to the categories based on one or more of: a first threshold for the number of clusters assigned to each category, a second threshold for spatial distortion for each category, or an importance of each category relative to other categories.
[0094] In some embodiments, the system 300 may further comprise an audio object reclassification unit configured to reclassify at least one audio object in one category to another category based on the spatial distortion for that category.
[0095] In some embodiments, the audio object allocation unit 303 may further allocate the audio objects in each category to at least one of the clusters assigned to that category based on reducing the distortion cost associated with that category.
[0096] In some embodiments, the audio object allocation unit 303 may further allocate the audio objects in each category to at least one of the clusters assigned to one or more of said categories based on reducing the distortion cost associated with those categories.
[0097] In some embodiments, the distortion cost may be associated with one or more of the original spatial location of each audio object, the spatial location of the at least one cluster, the identity of a category into which each audio object is classified, or the identity of each category to which the at least one cluster is assigned.
[0098] In some embodiments, the distortion cost may be determined based on one or more of: the distance between the original spatial location of each audio object and the spatial location of the at least one cluster; the distance between the original spatial location of each audio object and a reconstructed spatial location of that audio object determined based on the spatial location of the at least one cluster; or the mismatch between the identity of the category into which each audio object is classified and the identity of each category to which the at least one cluster is assigned.
[0099] In some embodiments, the system 300 may further comprise an audio object combination unit configured to combine the audio objects in each cluster to obtain clustered audio objects, and a metadata combination unit configured to combine metadata of the audio objects in each cluster to obtain metadata of the clustered audio objects.
[0100] For clarity, some additional components of system 300 are not depicted in FIG. 3 . However, it should be understood that all of the features described above with reference to FIG. 1 are applicable to system 300. Furthermore, the components of system 300 may be hardware modules, software unit modules, or the like. For example, in some embodiments, system 300 may be implemented partially or entirely in software and / or firmware implemented as a computer program product embodied in a computer-readable medium. Alternatively or additionally, system 300 may be implemented partially or entirely based on hardware, such as an integrated circuit (IC), an application-specific integrated circuit (ASIC), a system-on-chip (SOC), a field-programmable gate array (FPGA), or the like.
[0101] 4 illustrates a block diagram of an exemplary computer system 400 suitable for implementing embodiments. As shown, the computer system 400 includes a central processing unit (CPU) 401 that can execute various processes based on programs stored in a read-only memory (ROM) 402 or loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 also stores data needed by the CPU 401 to execute various processes. The CPU 401, the ROM 402, and the RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0102] The following components are connected to the I / O interface 405: an input unit 406 including a keyboard, a mouse, etc.; an output unit 407 including a display such as a cathode ray tube (CRT) or a liquid crystal display (LCD) and a speaker, etc.; a storage unit 408 including a hard disk, etc.; and a communication unit 409 including a network interface card such as a LAN card or a modem, etc. The communication unit 409 executes communication processes over a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory is mounted on the drive 410 as needed, and a computer program read from the removable medium 411 is installed in the storage unit 408 as needed.
[0103] Specifically, according to exemplary embodiments disclosed herein, the process described above with reference to Figure 1 may be implemented as a computer software program. For example, embodiments of the exemplary embodiments include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing method 100. In such embodiments, the computer program may be downloaded from a network via communication unit 409, mounted, and / or installed from removable media 411.
[0104] In general, various exemplary embodiments may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. Although various aspects of the exemplary embodiments are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, the blocks, devices, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing device, or some combination thereof.
[0105] Additionally, the various blocks illustrated in the flowcharts may be viewed as method steps and / or as actions resulting from operations of computer program code and / or as multiple coupled logic circuit elements configured to perform the associated function(s). For example, embodiments include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the method described above.
[0106] In the context of this disclosure, a machine-readable medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0107] Computer program code for implementing the methods of the exemplary embodiments may be written in any combination of one or more programming languages. Such computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus. When executed by the processor of the computer or other programmable data processing apparatus, the program code implements the functions / acts specified in the flowcharts and / or block diagrams. The program code may execute entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer, partially on a remote computer, or entirely on a remote computer or server. The program code may also be distributed on specially programmed devices, sometimes generally referred to herein as "modules." The software component portions of a module may be written in any computer language and may be part of a monolithic code base or developed in more discrete code portions, as is typical of object-oriented computer languages. Furthermore, modules may be distributed across multiple computer platforms, servers, terminals, mobile devices, etc. A given module may be implemented such that the described functions are performed by separate processors and / or computing hardware platforms.
[0108] As used herein, the term "circuitry" refers to all of the following: (a) hardware-only circuit implementations (e.g., implemented solely with analog and / or digital circuitry) and (b) combinations of circuitry and software (and / or firmware), such as (as appropriate): (i) a combination of processor(s) or (ii) processor(s) / software (including digital signal processors), software, and memory(s) that together cause a device such as a cell phone or server to perform various functions, and (c) circuitry that requires software or firmware for operation, even if the software or firmware is not physically present, such as microprocessor(s) or portions of microprocessor(s). Additionally, those skilled in the art will be familiar with the concepts of communication media, which typically embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and include any delivery media.
[0109] Furthermore, although acts are depicted in a particular order, this should not be understood as requiring such acts to be performed in the particular order shown, or sequentially, or that all of the shown acts be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Similarly, while certain specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of what may be claimed. Rather, they should be construed as descriptions of features that may be specific to particular exemplary embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination.
[0110] Various modifications and adaptations to the above exemplary embodiments may become apparent to those skilled in the art in light of the above description, when read in conjunction with the accompanying drawings. Any and all modifications fall within the scope of the non-limiting exemplary embodiments. Furthermore, other exemplary embodiments described herein will come to mind to those skilled in the art having the benefit of the teachings presented in the above description and drawings.
[0111] Thus, the exemplary embodiments disclosed herein may be embodied in any of the forms described herein. For example, the following enumerated example embodiments (EEE) describe some structure, features, and functionality of some aspects of the exemplary embodiments disclosed herein. [EEE1] A method for preserving object metadata in audio object clustering, comprising: assigning audio objects to categories, each category representing one or a unique combination of metadata that requires preservation; and generating several clusters for each category through a clustering process, subject to an overall (maximum) number of available clusters and an overall error criterion, the method further comprising: fuzzy object category separation or hard object category separation. [EEE2] The fuzzy object category separation involves the steps of: determining the output cluster centroids, for example by selecting the most important objects; and (1) determining the location metadata of each object.
number
number
[0112] It will be understood that the embodiments of the exemplary embodiments disclosed herein are not limited to the particular embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0113] Several aspects will be described. [Aspect 1] 1. A method for audio object clustering in which metadata is preserved, comprising: classifying a plurality of audio objects into categories based on information to be stored in metadata associated with the plurality of audio objects; assigning a predetermined number of clusters to said categories; and allocating audio objects in each of the categories to at least one of the clusters based on the assignments. method. [Aspect 2] 2. The method of aspect 1, wherein the information includes one or more of audio object size information, zone mask information, snap information, content type, or rendering mode. Aspect 3 Classifying a plurality of audio objects into categories based on information to be stored in metadata associated with the plurality of audio objects, comprising: Classifying audio objects with no information to be preserved into one category; This involves classifying audio objects into different categories with different information to be preserved. The method of embodiment 1. Aspect 4 The step of assigning a predetermined number of clusters to the categories comprises: determining the predetermined number of audio objects from the plurality of audio objects based on the importance of each audio object relative to the other audio objects; determining a distribution of said predetermined number of audio objects among said categories; assigning the predetermined number of clusters to the categories based on the distribution. The method of embodiment 1. Aspect 5 The step of assigning a predetermined number of clusters to the categories comprises: assigning the predetermined number of clusters to the categories based on reducing overall spatial distortion for the categories. The method of embodiment 1. Aspect 6 the overall spatial distortion for the category comprises a maximum spatial distortion among the individual spatial distortions of the category or a weighted sum of the individual spatial distortions of the category; the spatial distortion for each category is related to the original spatial location of each audio object in that category and the spatial location of at least one of said clusters; The method of embodiment 5. Aspect 7 A method as described in aspect 6, wherein the reconstructed spatial position of each audio object is determined based on the spatial position of the at least one cluster, and the spatial distortion for each category is determined based on the distance between the original spatial position of each audio object in that category and the reconstructed spatial position of that audio object. Aspect 8 7. The method of claim 6, wherein the plurality of audio objects are within one frame of the audio signal, and the spatial location for each category is further based on the difference between the number of clusters assigned to that category in the current frame and the previous frame. Aspect 9 Allocating the predetermined number of clusters to the categories based on reducing overall spatial distortion for the categories comprises: The amount of spatial distortion for a category in the previous iteration, or The difference between the spatial distortion for a category in the current iteration and the previous iteration and iteratively reducing the overall spatial distortion for the category based on at least one of: The method of embodiment 5. Aspect 10 The step of assigning a predetermined number of clusters to said categories further comprises: a first threshold for the number of clusters to be assigned to each category; A second threshold for spatial distortion for each category or The importance of each category relative to the others 10. The method of any one of embodiments 4 to 9, based on one or more of: Aspect 11 further comprising reclassifying at least one audio object in one category into another category based on the spatial distortion for that category. The method of embodiment 1. Aspect 12 allocating audio objects in each of the categories to at least one of the clusters based on the assignments, comprising: allocating audio objects in each category to at least one of the clusters assigned to that category based on reducing a distortion cost associated with that category. The method of embodiment 1. Aspect 13 allocating audio objects in each of the categories to at least one of the clusters based on the assignments, comprising: allocating audio objects in each category to at least one of the clusters assigned to one or more of the categories based on reducing a distortion cost associated with the category; The method of embodiment 1. Aspect 14 A method as described in aspect 12 or 13, wherein the distortion cost is related to one or more of the original spatial location of each audio object, the spatial location of the at least one cluster, the identification of a category into which each audio object is classified, or the identification of each category to which the at least one cluster is assigned. Aspect 15 wherein the distortion cost is: the distance between the original spatial location of each audio object and the spatial location of said at least one cluster; a distance between the original spatial location of each audio object and a reconstructed spatial location of that audio object determined based on the spatial locations of the at least one cluster; or a mismatch between the identity of the category into which each audio object is classified and the identity of each category to which said at least one cluster is assigned; 20. The method of embodiment 14, wherein the determination is based on one or more of: Aspect 16 combining the audio objects in each cluster to obtain clustered audio objects; and combining metadata of the audio objects in each cluster to obtain metadata of the clustered audio objects. The method of embodiment 1. Aspect 17 1. A system for audio object clustering in which metadata is preserved, comprising: an audio object classification unit configured to classify a plurality of audio objects into categories based on information to be stored in metadata associated with the plurality of audio objects; a cluster allocation unit configured to allocate a predetermined number of clusters to said categories; an audio object allocation unit configured to allocate audio objects in each of the categories to at least one of the clusters based on the allocation; system. Aspect 18 20. The system of claim 17, wherein the information includes one or more of audio object size information, zone mask information, snap information, content type, or rendering mode. Aspect 19 The system of aspect 17, wherein the audio object classification unit is further configured to classify audio objects without information to be preserved into one category and to classify audio objects with different information to be preserved into a different category. Aspect 20 The cluster allocation unit: an importance-based determination unit configured to determine the predetermined number of audio objects from the plurality of audio objects based on an importance of each audio object relative to other audio objects; a distribution determination unit configured to determine a distribution of the predetermined number of audio objects among the categories, the cluster allocation unit is further configured to allocate the predetermined number of clusters to the categories based on the distribution. 20. The system of embodiment 17. Aspect 21 20. The system of claim 17, wherein the cluster assignment unit is further configured to assign the predetermined number of clusters to the categories based on reducing an overall spatial distortion for the categories. Aspect 22 the overall spatial distortion for the category comprises a maximum spatial distortion among the individual spatial distortions of the category or a weighted sum of the individual spatial distortions of the category; the spatial distortion for each category is related to the original spatial location of each audio object in that category and the spatial location of at least one of said clusters; 22. The system of embodiment 21. Aspect 23 23. The system of claim 22, wherein the reconstructed spatial position of each audio object is determined based on the spatial position of the at least one cluster, and the spatial distortion for each category is determined based on the distance between the original spatial position of each audio object in that category and the reconstructed spatial position of that audio object. Aspect 24 23. The system of claim 22, wherein the plurality of audio objects are within one frame of the audio signal, and the spatial location for each category is further based on the difference between the number of clusters assigned to that category in the current frame and the previous frame. Aspect 25 the cluster allocation unit further comprising: The amount of spatial distortion for a category in the previous iteration, or The difference between the spatial distortion for a category in the current iteration and the previous iteration and iteratively reducing the overall spatial distortion for the category based on at least one of 22. The system of embodiment 21. Aspect 26 The cluster assignment unit further assigns the predetermined number of clusters to the categories: a first threshold for the number of clusters to be assigned to each category; A second threshold for spatial distortion for each category or The importance of each category relative to the others 26. The system of any one of aspects 20 to 25, configured to perform one or more of the following: Aspect 27 an audio object reclassification unit configured to reclassify at least one audio object in one category to another category based on a spatial distortion for that category; 20. The system of embodiment 17. Aspect 28 20. The system of claim 17, wherein the audio object allocation unit is further configured to allocate audio objects in each category to at least one of the clusters assigned to that category based on reducing a distortion cost associated with that category. Aspect 29 20. The system of claim 17, wherein the audio object allocation unit is further configured to allocate audio objects in each category to at least one of the clusters assigned to one or more of the categories based on reducing a distortion cost associated with the category. Aspect 30 A system as described in aspect 28 or 29, wherein the distortion cost is related to one or more of the original spatial location of each audio object, the spatial location of the at least one cluster, the identification of a category into which each audio object is classified, or the identification of each category to which the at least one cluster is assigned. Aspect 31 wherein the distortion cost is: the distance between the original spatial location of each audio object and the spatial location of said at least one cluster; a distance between the original spatial location of each audio object and a reconstructed spatial location of that audio object determined based on the spatial locations of the at least one cluster; or a mismatch between the identity of the category into which each audio object is classified and the identity of each category to which said at least one cluster is assigned; The system of embodiment 30, wherein the determination is based on one or more of: Aspect 32 an audio object combination unit configured to combine the audio objects in each cluster to obtain clustered audio objects; a metadata combination unit configured to combine metadata of the audio objects in each cluster to obtain metadata of the clustered audio objects. 20. The system of embodiment 17. Aspect 33 17. A computer program product having embodied on a machine-readable medium a computer program comprising program code for performing the method of any one of aspects 1 to 16.
Claims
1. 1. A method for decoding an encoded audio signal, the method comprising: receiving the encoded audio signal and determining an audio object from the encoded audio signal; receiving zone metadata indicating areas in which the audio object should not be rendered; classifying the audio objects into categories based on manual assignments; and rendering the audio object based on the category and the zone metadata. method.
2. A computer program product for causing a computer to carry out the method of claim 1.
3. The method of claim 1 , wherein the manual assignment includes metadata including pre-configuration setting information identifying a category assigned to the audio object.
4. 1. An apparatus for decoding an encoded audio signal, comprising: a first receiver that receives the encoded audio signal and determines an audio object from the encoded audio signal; a second receiver for receiving zone metadata indicating areas in which the audio object should not be rendered; a classification unit for classifying the audio objects into categories based on manual assignment; a renderer for rendering the audio object based on the category and the zone metadata. Device.
5. 5. The apparatus of claim 4, wherein the manual assignment includes metadata including pre-configuration setting information identifying a category assigned to the audio object.