Metadata-preserved audio object clustering
The method classifies audio objects by metadata and assigns clusters to preserve zone mask, snap, and rendering mode, addressing metadata loss in conventional clustering and ensuring accurate rendering.
Patent Information
- Application Number
- JP2025044207
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-01-06
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2035-12-10
AI Technical Summary
Conventional audio object clustering methods fail to preserve essential metadata such as zone mask, snap, and rendering mode, leading to violations of artistic intent and suboptimal rendering in downmixed audio systems.
A method and system for audio object clustering that classifies audio objects into categories based on metadata, assigns a predetermined number of clusters to these categories, and allocates objects within each category while preserving metadata like zone mask, snap, and rendering mode, using fuzzy or hard category clustering approaches to minimize spatial distortion.
Preserves critical metadata during clustering, ensuring accurate rendering and maintaining artistic intent, even in downmixed systems, thereby enhancing the audio experience.
Smart Images

Figure 2025108436000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority of Chinese Patent Application No. 201410765578.6, filed on December 11, 2014, and U.S. Provisional Patent Application No. 62 / 100,183, filed on January 6, 2015. The content of each application is hereby incorporated by reference in its entirety.
[0002] Technique The exemplary embodiments disclosed herein relate generally to audio content processing, and more particularly to methods and systems for audio object clustering that allow metadata to be stored.
Background Art
[0003] The advent of object - based audio has significantly increased the amount of audio data and the complexity of rendering this data within high - end playback systems. For example, a movie soundtrack may include many different sound elements, dialogues, noises, and sound effects corresponding to the images on the screen. These soundtracks may also originate from different positions on the screen and combine with background music and ambient effects to create an overall auditory experience. Accurate playback requires that the sound be reproduced in a manner that corresponds as closely as possible to what is shown on the screen with respect to the position, intensity, movement, and depth of the sound source. Object - based audio represents a significant improvement over traditional channel - based audio systems. Channel - based audio systems send audio content to individual speakers in the listening environment in the form of speaker feeds, and thus are relatively limited with respect to the spatial reproduction of specific audio objects.
[0004] The introduction of digital cinemas and the development of three-dimensional ("3D") content have created new standards for sound. For example, the incorporation of multi-channel audio that allows for greater creativity for content creators, and a more immersive, realistic auditory experience for the audience. It is crucial to extend beyond traditional speaker feeds and channel-based audio as a means of delivering spatial audio. Furthermore, there has been significant interest in model-based audio descriptions that allow listeners to select a desired playback configuration by using audio rendered specifically for a chosen configuration. The spatial presentation of sound utilizes audio objects. An audio object is an audio signal with an associated parametric source description, such as an apparent source position (e.g., 3D coordinates), an apparent source width, and other parameters. As a further advancement, a next-generation spatial audio (also referred to as "adaptive audio") format has been developed by including a mix of audio objects and traditional channel-based speaker feeds (audio beds) along with position metadata for the audio objects.
[0005] In the usage in this paper, the term "audio object" refers to an individual audio element that exists over a defined duration in a sound field. The term "audio bed" or "bed" refers to an audio channel that is intended to be played back at a predefined, fixed speaker position.
[0006] In some sound tracks, there may be several (e.g., 7, 9, or 11) bed channels that include audio. Further, based on the capabilities of an authoring system, there may be dozens or even hundreds of individual audio objects that are combined during rendering to create a spatially diverse and immersive audio experience. In other delivery and transmission systems, there may be a sufficiently large available bandwidth to transmit all audio beds and objects with little or no audio compression. However, in some cases such as Blu-ray Disc, broadcast (cable, satellite, terrestrial), mobile phone (3G and 4G), and over-the-top (OTT or Internet) delivery, there may be significant limitations to the available bandwidth for digitally transmitting all of the bed and object information generated at the time of authoring. An audio encoding method (reversible or irreversible) may be applied to the audio to reduce the required bandwidth, but audio encoding may not be sufficient to reduce the required bandwidth for transmitting audio over highly restricted networks such as mobile 3G and 4G networks.
[0007] Several conventional methods have been developed for reducing the number of input objects by clustering into a smaller set of output objects. Generally, in some clustering processes, metadata such as size, zone mask, and snap should be pre-rendered to the internal channel layout. Clustering of audio objects is based only on the spatial position of the audio objects, and the output objects contain only position metadata. This type of output object may not function well for some playback systems because the loss of metadata may break the artistic intent where loss is expected.
[0008] The subject matter discussed in the background section should not be assumed to be prior art merely for the purpose of disclosure in the background section. Similarly, the problems mentioned in the background section or related to the subject matter of the background section should not be assumed to have been previously recognized in the prior art. The subject matter in the background section merely represents various approaches, and those approaches themselves may also be exemplary embodiments.
Summary of the Invention
Problems to be Solved by the Invention
[0009] To address the above and other potential problems, exemplary embodiments propose a method and system for audio object clustering in which metadata is stored.
Means for Solving the Problems
[0010] In one aspect, an exemplary embodiment provides a method for audio object clustering in which metadata is stored. The method includes classifying a plurality of audio objects into several categories based on information to be stored in the metadata associated with the plurality of audio objects. The method further includes assigning a predetermined number of clusters to the categories and allocating the audio objects in each of the categories to at least one of the clusters based on the assignment. Embodiments in this regard further include a corresponding computer program product.
[0011] In another aspect, the exemplary embodiment provides a system for audio object clustering in which metadata is stored. The system includes an audio object classification unit configured to classify a plurality of audio objects into several categories based on information to be stored in the metadata associated with the plurality of audio objects. The system further includes a cluster assignment unit configured to assign a predetermined number of clusters to the categories, and an audio object allocation unit configured to allocate the audio objects in each of the categories to at least one of the clusters based on the assignment.
[0012] Through the above description, according to the exemplary embodiments disclosed in this document, it will be understood that the input audio object is classified into corresponding categories depending on the information to be stored in the metadata, whereby different metadata to be stored or unique combinations of metadata to be stored are associated with different categories. After clustering, for the audio objects within a certain category, the possibility that it is mixed with audio objects associated with different metadata is reduced. In this regard, the metadata of the audio object can be stored after clustering. Other advantages achieved by the exemplary embodiments will become apparent from the following description.
Brief Description of the Drawings
[0013] Through the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the embodiments will become more understandable. In the drawings, several exemplary embodiments are shown in an illustrative and non-limiting manner.
Figure 1
Figure 2
Figure 3
Figure 4
Mode for Carrying Out the Invention
[0014] The principles of the exemplary embodiments will now be described with reference to the various exemplary embodiments shown in the drawings. It should be understood that the depiction of these embodiments is solely for the purpose of enabling those skilled in the art to better understand and further implement the exemplary embodiments, and is not intended to limit the scope in any way.
[0015] As described above, due to limitations in the encoding / decoding rate and transmission bandwidth, the number of audio objects used to generate adaptive audio content may be reduced by clustering. In addition to the metadata describing its spatial position, audio objects typically have other metadata describing their attributes such as size, zone mask, snap, and content type. Each of these attributes describes the artistic intent regarding how the audio object should be processed when rendered. However, in some conventional methods, after the audio objects are clustered, only the position metadata remains. Other metadata may be pre-rendered into an internal channel layout such as a 7.1.2 or 7.1.4 system, but this does not work well for all systems. In particular, when an audio object is downmixed to, for example, a 5.1 or 7.1 system, the artistic intent of the audio object may be violated when rendered.
[0016] Taking the metadata "zone mask" as an example. This has multiple modes, and each mode defines an area where the audio object should not be rendered. One mode of the zone mask is "no side", which describes that the side speakers should be masked when rendering the audio object. When an audio object at spatial position z = 1 is rendered to a 5.1 system using the metadata "no side" by utilizing traditional clustering methods, the side speakers may be activated in the 5.1 rendering. This is because the sound from the ceiling speakers may be convolved laterally. To address this problem, the metadata "zone mask" should be preserved in the clustering process and be correctly processable in the audio renderer.
[0017] In another example, it may be desirable for the dialog object to be separated from other objects after clustering. This can have many benefits for subsequent audio object processing. For example, in subsequent audio processing such as dialog enhancement, the separated dialog object cluster can be easily enhanced by simply applying the gain(s). Otherwise, it can be very difficult to separate the dialog object if it is mixed with other objects in the cluster. In the application of dialog replacement, the dialogs in each language may be completely separated from each other. For such purposes, in the clustering process, the dialog objects should be preserved and assigned to separate individual clusters.
[0018] Furthermore, the audio object may be associated with metadata that describes rendering it in its rendering mode, such as left total / right total (Lt / Rt) or binaurally using a head-related transfer function (HRTF) when processing in a headphone renderer. These rendering modes are also expected to be preserved after clustering to yield the best rendering results.
[0019] Therefore, to achieve a better audio experience, it is desirable to have the metadata preserved in the audio object clustering. The exemplary embodiments disclosed herein propose methods and systems for object clustering where metadata is preserved.
[0020] First, refer to FIG. 1. FIG. 1 depicts a flowchart of a method 100 for audio object clustering where metadata is preserved, based on an exemplary embodiment.
[0021] In S101, a plurality of audio objects are classified into several categories based on the information to be stored in the metadata associated with the plurality of audio objects. These audio objects are provided as input, and there may be dozens, hundreds, or sometimes thousands of input audio objects.
[0022] In the usage in this article, the information to be stored in the metadata associated with each audio object may indicate the processing intention when the audio object is rendered. That information may describe how the audio object should be processed when it is rendered. In some embodiments, that information may include one or more of the size information, zone mask information, snap information, content type, or rendering mode of the audio object. The size information may be used to indicate the spatial area or volume occupied by the audio object. The zone mask information indicates the mode of the zone mask that defines the area where the audio object should not be rendered. For example, the zone mask information may indicate modes such as "no side", "surround only", "front only", etc. The snap information indicates whether the audio object should be panned directly to the nearest speaker.
[0023] Some examples of the information to be stored in the metadata are described, and it should be noted that other information included in the metadata (non-limiting examples include spatial position, spatial width, etc.) may also be considered in the audio object classification according to the preferences of the user or audio engineer. In some embodiments, all the information in the metadata associated with the audio object may be considered.
[0024] The number of categories may depend on the information in the metadata of those audio objects and can be one or more. In some embodiments, audio objects with no information to be stored may be classified into one category, and audio objects with different information to be stored may be classified into different categories. That is, depending on the different information to be stored, the corresponding audio objects are classified into different categories. Alternatively, a category may represent a unique combination of different information to be stored in the metadata. All other audio objects with no information of interest may, in some cases, be included in one category or multiple categories. The scope of the exemplary embodiments is not limited in this regard.
[0025] Categories may be given by manual assignment, automatic assignment, or a combination thereof. For example, a user or an audio engineer may label audio objects associated with different types of metadata with different flags, and then those labeled audio objects may be classified into different categories according to those flags. As another example, the information to be stored in the metadata may be automatically identified. A user or an audio engineer may pre-configure their preferences or expectations, such as separating dialog objects, separating different dialog languages, and / or separating different modes of zone masks. Depending on the pre-configuration, audio objects may be classified into different categories.
[0026] Assume there are zero audio objects. In the classification process, the information to be stored in the metadata of the audio object may be derived from (1) manual labels of the metadata given by user input, such as zone masks or snaps or content type or language labels, and / or (2) automatic identification / labeling of the metadata, such as but not limited to content type identification information. The number N of possible categories may be determined according to the derived information, and each category consists of a unique combination of the information to be stored. After classification, each audio object may have an associated category identification information n o Thereof.
[0027] Referring to FIG. 2, a schematic diagram of audio object clustering is shown. As shown in FIG. 2, based on the information to be stored in the metadata, a plurality of input audio objects are classified into five categories, categories 0 to 4. An example of a category may be given as follows: · Category 0: All audio objects with no information to be stored; · Category 1: Music objects, without zone masks; · Category 2: Sound effect objects, with zone mask "surround only"; · Category 3: English dialogue objects; · Category 4: Spanish dialogue objects, with zone mask "front only".
[0028] The input audio object may include one or more frames. A frame is a processing unit for audio content, and the duration of a frame may vary and may depend on the configuration of the audio processing system. Since the audio objects to be classified may vary temporally for different frames and their metadata may also vary, the value of the number of categories may also vary over time. The categories representing different types of information to be stored may be predefined by the user or by default, in which case the input audio objects in one or more frames may be classified into predefined categories based on that information. In subsequent processing, the categories of the classified audio objects may be considered, and those without audio objects may be ignored. For example, in FIG. 2, category 0 corresponding to audio objects with no information to be stored may be omitted when there are no such audio objects. The number of audio objects classified into each category may be considered to vary over time.
[0029] In S102, a predetermined number of clusters are assigned to the categories. The predetermined number may be greater than 1 and may depend on the transmission bandwidth and the encoding / decoding rate of the audio processing system. There may be a trade-off between the transmission bandwidth (and / or encoding rate and / or decoding rate) and the error criterion of the output audio object. For example, the predetermined number may be 11 or 16. Other values such as 5, 7 or 20 may be determined, and the scope of the exemplary embodiments is not limited in this regard.
[0030] In some embodiments, the predetermined number may not vary within the same processing system. In some other embodiments, the predetermined number may vary for different audio files to be processed.
[0031] In the exemplary embodiments disclosed in this document, audio objects are first classified into categories according to metadata in S101, such that each category can represent different information to be stored or a unique combination of different information to be stored. Then, the audio objects in those categories may be clustered in subsequent processing. There can be various approaches to assign / allocate a predetermined total number of clusters to the categories. In some exemplary embodiments, since the total number of clusters is predetermined and fixed, it is possible to determine the number of clusters to be assigned to each category before clustering the audio objects. Some exemplary embodiments will be discussed hereinafter.
[0032] In an exemplary embodiment, the cluster assignment may depend on the importance of the plurality of audio objects. Specifically, a predetermined number of audio objects from the plurality of audio objects may first be determined based on the importance of each audio object relative to other audio objects, and then the distribution of the predetermined number of audio objects among the categories may be determined. The predetermined number of clusters is correspondingly assigned to the categories according to the distribution.
[0033] The importance of each audio object may be related to one or more of the type of the content of the audio object, the partial loudness level or the energy level. An audio object of high importance may indicate that the audio object is perceptually prominent among the input audio objects, for example, due to its partial loudness level or energy level. In some use cases, one or more types of content may be considered important, and then a high importance may be assigned to the corresponding audio object. For example, a higher importance may be assigned to a dialog object. It should be noted that there are many other ways to determine or define the importance of each audio object. For example, the importance levels of some audio objects may be specified by the user. The scope of the exemplary embodiments is not limited in this regard.
[0034] Assume that the predetermined total number of clusters is M. In the first step, up to the M most important audio objects among the input audio objects are selected. Since all input audio objects are classified into corresponding categories in S101, in the second step, the distribution of those M important audio objects among the categories may be determined. Based on the number of those M audio objects assigned to a certain category, the same number of clusters may be assigned to that category.
[0035] Referring to FIG. 2, for example, 11 of the most important audio objects (shown as circles 201) are determined from a plurality of input audio objects (shown as a set of circles 201 and 202). After classifying all the input audio objects into five categories, categories 0 to 4, it can be seen from FIG. 2 that four of the most important audio objects are classified into category 0, three of the most important audio objects are classified into category 1, one of the most important audio objects is classified into category 2, two of the most important audio objects are classified into category 3, and one of the most important audio objects is classified into category 4. As a result, as shown in FIG. 2, 4, 3, 1, 2, and 1 clusters are assigned to categories 0 to 4, respectively.
[0036] It should be noted that the above example of the importance criteria based on this exemplary embodiment of the various exemplary embodiments may not be so strict. That is, it is not essential that the most important audio object be selected. In some embodiments, an importance threshold may be configured. From among the audio objects whose importance is higher than the threshold, the predetermined number of audio objects may be randomly selected.
[0037] In addition to the importance criteria, the cluster assignment may be performed based on reducing the overall spatial distortion for the category. That is, the predetermined number of clusters may be assigned to the category based on reducing or even minimizing the overall spatial distortion for the category.
[0038] In one exemplary embodiment, the overall spatial distortion for categories may include a weighted sum of the individual spatial distortions of those categories. The weight of the corresponding category may represent the importance of that category or the importance of the information to be preserved associated with that category. For example, a category with a higher importance may have a larger weight. In another embodiment, the overall spatial distortion for those categories may include the maximum spatial distortion among the individual spatial distortions of those categories. It should be considered that it is not essential that only the maximum is selected, and in some embodiments, other spatial distortions among the categories, such as the second largest spatial distortion, the third largest spatial distortion, etc., may be considered as the overall spatial distortion.
[0039] The spatial distortion for each category may be represented by the distortion level of the audio objects included in that category, and the distortion level of each audio object may be measured by the difference between its original spatial position and its position after clustering. Generally, the clustered position of an audio object depends on the spatial position of the cluster (s) to which it is assigned. In this sense, the spatial distortion for each category is related to the original spatial position of each audio object within that category and the spatial position of the cluster (s). The original spatial position of an audio object may be included in the metadata of the audio object and may consist of, for example, three Cartesian coordinates (or similarly, for example, polar coordinates or cylindrical and spherical coordinates, homogeneous coordinates, line number coordinates, etc.). In one embodiment, in order to calculate the spatial distortion for each category, the reconstructed spatial position of each audio object within the category may be determined based on the spatial position of the cluster (s). Then, the spatial distortion for each category may be calculated based on the distance between the original spatial position of each audio object within the category and the reconstructed spatial position of the audio object. The reconstructed spatial position of an audio object is the spatial position of the audio object represented by one or more corresponding spatial clusters. One exemplary approach for determining the reconstructed spatial position is described below.
[0040] To obtain the overall spatial distortion, the spatial distortion for different numbers of clusters may first be calculated for each category. There are many approaches for determining the spatial distortion for a category of audio objects. One approach is given as an example below. It should be noted that other existing ways of measuring the spatial distortion of an audio object (and thus of a category) may be applied.
[0041] For category n, the spatial position [Number] M having n cluster centroids, and let them be represented as {C n (1), C n (2), …, C n (M n )}. dis(o n (i), {C n (1), C n (2), …, C n (M n )}) may represent the spatial distortion when clustering the audio object o n (i) into the M n cluster centroids (in this case, assume that an audio object within a certain category is assigned only to the cluster associated with that category). The spatial distortion for category n may be expressed as follows: [Number] Here, O n represents the number of audio objects in category n, and o n (i) represents the i-th audio object in category n. In some embodiments, C n (m) may be the spatial position of the audio object having the m-th greatest importance within the category, and the spatial position of C n (m) may be the spatial position of that audio object. The spatial distortion dis(o n (i), {C n (1), C n (2), …, C n (M n )}) is the spatial position n of each audio object o [Number] and the reconstructed spatial position of that audio object when clustered into the M n cluster centroids
Number
[0042] When the spatial distortion for each category is obtained, in one embodiment, the overall spatial distortion for those categories may be determined as the weighted sum of the individual spatial distortions of those categories as described above. For example, the overall spatial distortion may be determined as follows:
Number
[0043] In another embodiment, the overall spatial distortion for the category may be determined as the maximum spatial distortion among the individual spatial distortions of those categories. For example, the overall spatial distortion may be determined as follows:
Number
Number
[0044] The input audio object is generally within one frame of the audio signal. Due to the typical dynamic nature of the audio signal and considering that the number of audio objects varies in each category, the number of clusters assigned to each category can typically change over time. The varying number of clusters for each category can cause some instability issues, so a modified spatial distortion that takes into account cluster number consistency is utilized in the cost metric. As a result, the cost metric may be defined as a function of time. Specifically, the spatial distortion for each category is further based on the difference between the number of clusters assigned to that category in the current frame and the number of clusters assigned to that category in the previous frame. In this regard, the overall spatial distortion in Equation (2) may be modified as follows: [Number] The overall spatial distortion in Equation (3) may be modified as follows: [Number] In Equations (4) and (5), M n represents the number of clusters of category n in the current frame, M n ' represents the number of clusters of category n in the previous frame, and f(D n (M n ), M n , M n ) represents the modified overall spatial distortion.
[0045] When the number of clusters assigned to a category changes in the current frame compared to the previous spatial distortion, the modified spatial distortion may be increased to prevent the change in the number of clusters. In one embodiment, f(D n (M n ), M n , M n ) may be determined as follows:
Mathematics
[0046] Since reducing the number of clusters in a category is more likely to introduce spatial instability than increasing the number of clusters, in another embodiment, f(D n (M n ),M n ,M n ) may be determined as follows:
Mathematics
[0047] In the above description, with respect to cluster assignment based on reducing the overall spatial distortion, a large amount of computational effort may be involved in determining the optimal number of clusters for each category. To efficiently determine the number of clusters for each category, in one embodiment, a sequential iterative process is proposed. That is, the optimal number of clusters for each category is estimated by maximizing the cost reduction in each iteration step of the cluster assignment process. Thereby, the overall spatial distortion for the category may be reduced sequentially iteratively and even minimized.
[0048] By sequentially iterating from 1 to a predetermined number of clusters M, in each iteration step, one or more clusters are assigned to the category that most requires it. Let Cost(m - 1) and Cost(m) be denoted as the overall spatial distortion in the (m - 1)-th and m-th iteration steps. In the m-th iteration step, one or more new clusters may be assigned to the category n* that can most reduce the overall spatial distortion. Therefore, n* may be determined by expanding or maximizing the reduction of the overall spatial distortion. This may be expressed as follows: [Number] The sequential iteration process may be based on at least one of the difference between the spatial distortion for a category in the current iteration step and the previous iteration step or the amount of spatial distortion for a category in the previous iteration step.
[0049] Regarding the overall spatial distortion obtained by the weighted sum of the spatial distortions of all categories, the sequential iteration process may be based on the difference between the spatial distortion for a category in the current iteration step and the previous iteration step. In each iteration step, at least one cluster is assigned to a category, and the category may be such that its spatial distortion in the current iteration step is sufficiently lower (according to a certain first predetermined level) than its spatial distortion in the previous iteration step when the at least one cluster is assigned together with it. In an embodiment, the at least one cluster may be assigned to the category with the most reduced spatial distortion when assigned together with the at least one cluster. For example, in this embodiment, n* may be determined as follows: [Number] Here, M n*,m-1 and D n* (M n*,m-1) represents the number of clusters and the spatial distortion for category n* after the (m - 1)-th iteration step. M n*,m-1 +1 represents the number of clusters of category n* in the m-th iteration step when one new cluster is assigned / added to category n* in this iteration step, D n* (M n*,m-1 +1) represents the spatial distortion for category n* in the m-th iteration step. Note that in each iteration step, two or more new clusters may be assigned, and category n* may be determined similarly.
[0050] Regarding the overall spatial distortion determined as the maximum spatial distortion among all categories, the sequential iteration process may be based on the amount of spatial distortion of a certain category in the previous iteration step. In each iteration step, at least one cluster may be assigned to a category having a spatial distortion higher than a second predetermined level in the previous iteration step. In one embodiment, the at least one cluster may be assigned to the category having the maximum spatial distortion in the previous iteration step. For example, in this embodiment, n* may be determined as follows: [Number] Since the category having the maximum spatial distortion in the previous iteration step has its spatial distortion reduced in the current iteration step (if one or more clusters are assigned in the current iteration step), the overall spatial distortion determined by the maximum spatial distortion among all categories may also be reduced in the current iteration step.
[0051] Note that the decisions given in equations (9) and (10) may be used jointly in a single sequential iteration process. For example, in a certain iteration step, equation (9) may be used to assign new cluster(s) in this iteration step. In another iteration step, equation (10) may be used to assign other new cluster(s).
[0052] Two ways of cluster assignment have been described above. One is based on the importance of the audio object, and the other is based on reducing the overall spatial distortion. Additionally or alternatively, user input may also be used to guide the cluster assignment. Since users may have different requirements for different contents for different use cases, this can significantly improve the flexibility of the clustering process. In some embodiments, the cluster assignment may further be based on: a first threshold for the number of clusters to be assigned to each category, a second threshold for the spatial distortion for each category, or one or more of the importance of each category compared to other categories.
[0053] The first threshold may be predefined for the number of clusters to be assigned to each category. The first threshold may be a predetermined minimum or maximum number of clusters for each category. For example, the user may specify that a certain category should have a certain minimum number of clusters. In this case, during the assignment process, at least the specified number of clusters should be assigned to that category. If a maximum threshold is set, at most the specified number of clusters can be assigned to that category. The second threshold may be set to ensure that the spatial distortion for a certain category is reduced to a reasonable level. The importance of each category may also be specified by the user, or determined based on the importance of the audio objects classified in that category.
[0054] In some cases, the spatial distortion for a certain category may be high after the cluster assignment, which may introduce audible artifacts. To address this issue, in some embodiments, at least one audio object in a certain category may be reclassified into another category based on the spatial distortion for that category. In an exemplary embodiment, if the spatial distortion of one of the categories is higher than a predetermined threshold, some audio objects within that category may be reclassified into another category until the spatial distortion becomes less than (or equal to) the threshold. In some examples, the audio object may be reclassified into a category that includes audio objects without information to be stored in the metadata, such as category 0 in FIG. 2. In some embodiments where the cluster assignment is based on minimizing the overall spatial distortion in an iterative process, the object reallocation may also be performed in each iteration step until the criterion for the spatial distortion for the category is satisfied, with the maximum spatial distortion dis(o n (i),{C n (1),C n (2),…,C n (M n )}) having an audio object may be a sequential iterative process of reclassification.
[0055] Due to the typical dynamic nature of audio signals, the importance or spatial position of an audio object (and thus the spatial distortion as well) changes over time. As a result, the cluster assignment may change over time, in which case the number of clusters assigned to each category may vary over time. In this case, the category identification information associated with cluster m may change over time. In particular, cluster m may represent a certain language (e.g., Spanish) during the first frame, while for the second frame, the category identification information, and thus the language, may be changed (e.g., English). This is in contrast to legacy channel - based systems where the language is statically associated with a channel rather than changing dynamically.
[0056] The cluster assignment in S102 was described above.
[0057] Returning to the reference of FIG. 1, in S103, for each of the categories, the audio objects are assigned to at least one of the clusters based on the assignment.
[0058] In the following description, after the audio objects are classified into categories in S101 and clusters are assigned to each category in S102, two approaches for clustering the audio objects are provided.
[0059] In one approach, audio objects in each category may be assigned to at least one of the clusters assigned to one or more of those categories based on reducing the distortion cost associated with those categories. That is, due to the limitation in the number of clusters assigned to each category, some leakage across clusters and categories is allowed to reduce the distortion cost and avoid artifacts for complex audio content. This approach may be referred to as fuzzy category clustering. In this fuzzy category clustering approach, an audio object may be softly split using the gains and corresponding costs to different clusters in different categories. During the clustering process, it is expected that the distortion cost is minimal with respect to the overall spatial distortion as well as the disadvantage or mismatch of assigning objects within a category to clusters of different categories. Thus, there is a trade-off between the cluster budget and the complexity of the audio content. The fuzzy category clustering approach may be suitable for audio objects with metadata such as zone masks and snaps, because a strict separation from other metadata is not required for them. The fuzzy category clustering approach may be described as follows.
[0060] In the fuzzy category clustering approach, the number of clusters assigned to each category may be determined in S102 based on the importance of the audio object or based on minimizing the overall spatial distortion. For cluster assignment based on importance, there may be some categories to which no cluster is assigned. In these cases, the fuzzy category clustering approach may be applied when clustering the audio objects. This is because the objects may be softly clustered into the cluster(s) of other categories. There may be no required correlation between the approach applied at the stage of cluster assignment and the approach applied at the stage of audio object clustering.
[0061] In the fuzzy category clustering approach, the distortion cost is: (1) the original spatial position of each audio object
Number
Number
Number
Number
[0062] In some embodiments, the gain g o,m is
Number
Number
Number
[0063]
Number
Number
[0064] In some embodiments, the cost function may be expressed as a cumulative contribution using
Number
[0065] The cost function may typically be minimized under some additional criterion. When allocating the audio signal, one criterion may be to maintain the summed amplitude or energy of the input audio object. For example,
Number
[0066] Hereinafter, the cost function E may be discussed. By minimizing the cost function, the gain g o,m can be determined.
[0067] The cost function, as described above,
Number
[0068] The cost function [Number] may also be associated with the discrepancy between, which is the second term E in the cost function C and may be regarded as. E C may represent the cost of clustering audio objects across clusters within different categories and may be determined as follows: [Number] where n m != n0 may be determined as follows: [Number] As described above, when minimizing the cost function, one criterion is to maintain the total amplitude or energy of the input audio objects. Thus, the cost function may also be related to the gain or loss of energy; that is, the sum of the gains for specific audio objects and the deviation from +1. The deviation is the third term E in the cost function N and may be regarded as and may be determined as follows: [Number] Furthermore, the cost function is the original spatial position of each audio object [Number] and the reconstructed spatial position of that audio object [Number] It may also be based on the distance to. This reconfigured spatial position is the spatial position of the cluster in which the audio object is clustered with gain g o,m with [Number] It may be determined according to. For example, it may be determined as follows: [Number] [Number] The distance between may be regarded as the fourth term E in the cost function P and may be expressed as follows: [Number] According to the first, second, third and fourth terms, the cost function may be expressed as the weighted sum of these terms and may be expressed as follows: [Number] Here, the weights w D , w C , w N , w P may represent the importance of different terms in the cost function.
[0069] Based on these four terms in the cost function, the gain g o,m can be determined. An example of the calculation for the gain g o,m is given below. Note that other methods of calculation are also possible.
[0070] The gain g o,m for the o-th audio object for M clusters may be written as a vector: [Number] The spatial positions of the M clusters may be written as a matrix:
Number
Number
Number
Number
[0071] The second term E representing the mismatch between n o and n m of the audio object C may be reformulated as follows:
Number
[0072] The third term E representing the deviation of the sum of the gains of the audio object from +1 N may be reformulated as follows:
Number
[0073] The fourth term E representing the distance between the original spatial position and the reconstructed spatial position of the audio object P may be reformulated as follows:
Number
Number
Number
Number
Number
[0074] The o-th audio object may be clustered into the M clusters with the determined gain vector
Number
[0075] The reconstructed spatial position of the audio object may be obtained by Equation (17) when the gain vector is determined. In this regard, the process of determining the gain may also be applied in cluster assignment based on minimizing the overall spatial distortion described above to determine the reconstructed spatial position, and thus the spatial positions of each category.
[0076] It should be noted that a quadratic polynomial is used as an example to determine the minimum in the cost function. In other exemplary embodiments, many other exponential values, such as 1, 1.5, 3, etc. may also be used.
[0077] A fuzzy category clustering approach for audio object clustering has been described above. In another approach, the audio objects in each category may be assigned to at least one of the clusters assigned to that category based on reducing the spatial distortion cost associated with that category. That is, leakage across categories is not allowed. Audio object clustering is performed within each category, and an audio object cannot be grouped into a cluster assigned to another category. This approach may be referred to as a hard category clustering approach. In some embodiments where this approach is applied, an audio object may be assigned to two or more of the clusters assigned to the category corresponding to that audio object. In a further embodiment, leakage across clusters is not allowed, and an audio object may be assigned to only one of the clusters assigned to the corresponding category.
[0078] The hard category clustering approach may be suitable for some individual applications that require an audio object (dialog object) to be separated from others, such as dialog replacement or dialog enhancement.
[0079] In the hard category clustering approach, since an audio object within a certain category cannot be clustered into one or more clusters of other categories, it is expected that at least one cluster is assigned to each category in the previous cluster assignment. For this purpose, in some embodiments, it may be more suitable to perform cluster assignment by minimizing the overall spatial distortion described above. In other embodiments, when hard category clustering is applied, a cluster assignment based on importance may also be used. Some additional conditions may be used in the cluster assignment to ensure that at least one cluster is assigned to each category as discussed above. For example, a minimum threshold for the cluster or a minimum threshold for the spatial distortion for each category may be utilized.
[0080] Since categories represent the same type of metadata, within a category, an audio object can be clustered into one or more clusters in one or more exemplary embodiments. For example, as shown in FIG. 2, the audio objects of category 1 may be clustered into one or more of clusters 4, 5, or 6. In a scenario where an audio object is clustered into multiple clusters within a single category, the corresponding gain may also be determined to reduce or even minimize the distortion cost associated with that category (this may be similar to what was described for the fuzzy category clustering approach). The difference lies in that the determination is made within a single category. In some embodiments, each input audio object may be allowed to be clustered into only one cluster assigned to its category.
[0081] Two approaches to audio clustering were discussed above. It should be noted that these two approaches can be used separately or in combination. For example, after audio object classification in S101 and cluster assignment in S102, for some categories, a fuzzy category clustering approach may be applied to cluster the audio objects within them; for the remaining categories, a hard category clustering approach may be applied. That is, some leakage across categories may be acceptable within some categories, while leakage across categories is not acceptable for other categories.
[0082] After the input audio object is assigned to a cluster, for each cluster, an audio object obtained by combining and clustering the audio objects may be obtained, and the metadata of the audio objects in each cluster may be combined to obtain the metadata of the clustered audio object. The clustered audio object may be the weighted sum of all the audio objects in the cluster using the corresponding gains. The metadata of the clustered audio object may, in some examples, be the corresponding metadata representing the category, or in other examples, may be the metadata of any audio object or the most important audio object within the cluster or its category.
[0083] Since all input audio objects are classified into corresponding categories depending on the information to be stored in the metadata before audio object clustering, different metadata to be stored or unique combinations of stored metadata are associated with different categories. After clustering, for audio objects within a certain category, the possibility of being mixed with audio objects associated with different metadata is reduced. In this regard, the metadata of audio objects can be stored after clustering. Further, during the cluster assignment and audio object allocation process, spatial distortion or distortion cost is considered.
[0084] Figure 3 depicts a block diagram of a system 300 for audio object clustering in which metadata is stored based on an exemplary embodiment. As depicted in Figure 3, the system 300 has an audio object classification unit 301 configured to classify a plurality of audio objects into several categories based on the information to be stored in the metadata associated with the plurality of audio objects. The system 300 further has a cluster assignment unit 302 configured to assign a predetermined number of clusters to those categories, and an audio object allocation unit 303 configured to allocate the audio objects in each of those categories to at least one of the clusters based on the assignment.
[0085] In some embodiments, the information may include one or more of size information of the audio object, zone mask information, snap information, content type, or rendering mode.
[0086] In some embodiments, the audio object classification unit 301 may further be configured to classify audio objects without information to be stored into one category, and classify audio objects with different information to be stored into different categories.
[0087] In some embodiments, the cluster assignment unit 302 may further include: an importance-based determination unit configured to determine a predetermined number of audio objects from the plurality of audio objects based on the importance of each audio object relative to other audio objects; and an assignment determination unit configured to determine the distribution among the categories of the predetermined number of audio objects. In these embodiments, the cluster assignment unit 302 may further be configured to assign the predetermined number of clusters to the categories based on the distribution.
[0088] In some embodiments, the cluster assignment unit 302 may further be configured to assign the predetermined number of clusters to the categories based on reducing the overall spatial distortion for the categories.
[0089] In some embodiments, the overall spatial distortion for the categories may include the maximum spatial distortion among the individual spatial distortions of the categories, or the weighted sum of the individual spatial distortions of the categories. The spatial distortion for each category may be associated with the original spatial position of each audio object in that category and the spatial position of at least one of the clusters.
[0090] In some embodiments, the reconstructed spatial positions of each audio object may be determined based on the spatial positions of the at least one cluster, and the spatial distortion for each category may be determined based on the distance between the original spatial position of each audio object in that category and the reconstructed spatial position of that audio object.
[0091] In some embodiments, the plurality of audio objects may be within one frame of the audio signal, and the spatial distortion for each category may further be based on the difference between the number of clusters assigned to that category in the current frame and the number of clusters assigned to that category in the previous frame.
[0092] In some embodiments, the cluster assignment unit 302 may further be configured to sequentially and iteratively reduce the overall spatial distortion for the category based on at least one of the amount of spatial distortion for a category in the previous iteration step or the difference between the spatial distortion for a category in the current iteration step and the previous iteration step.
[0093] In some embodiments, the cluster assignment unit 302 may further be configured to assign the predetermined number of clusters to the category based on one or more of a first threshold for the number of clusters assigned to each category, a second threshold for the spatial distortion for each category, or the importance of each category relative to other categories.
[0094] In some embodiments, the system 300 may further include an audio object reclassification unit configured to reclassify at least one audio object in a category into another category based on the spatial distortion for that category.
[0095] In some embodiments, the audio object allocation unit 303 may further allocate the audio objects in each category to at least one of the clusters assigned to that category based on reducing the distortion cost associated with that category.
[0096] In some embodiments, the audio object allocation unit 303 may further allocate the audio objects in each category to at least one of the clusters assigned to one or more of the categories based on reducing the distortion cost associated with those categories.
[0097] In some embodiments, the distortion cost may be associated with one or more of the original spatial position of each audio object, the spatial position of the at least one cluster, the identification information of the category to which each audio object is classified, or the identification information of each category to which the at least one cluster is assigned.
[0098] In some embodiments, the distortion cost may be determined based on one or more of the distance between the original spatial position of each audio object and the spatial position of the at least one cluster, the distance between the original spatial position of each audio object and the reconstructed spatial position of that audio object determined based on the spatial position of the at least one cluster, or the mismatch between the identification information of the category to which each audio object is classified and the identification information of each category to which the at least one cluster is assigned.
[0099] In some embodiments, system 300 may further include an audio object combination unit configured to combine audio objects within each cluster to obtain a clustered audio object, and a metadata combination unit configured to combine the metadata of the audio objects in each cluster to obtain the metadata of the clustered audio object.
[0100] For clarity, some additional components of system 300 are not depicted in FIG. 3. However, it should be understood that all of the features described above with reference to FIG. 1 are applicable to system 300. Further, the components of system 300 may be hardware modules or software unit modules, etc. For example, in some embodiments, system 300 may be implemented, in part or in whole, as software and / or firmware implemented as a computer program product embodied in a computer-readable medium. Alternatively or additionally, system 300 may be implemented, in part or in whole, based on hardware, such as an integrated circuit (IC), an application-specific integrated circuit (ASIC), a system-on-chip (SOC), a field-programmable gate array (FPGA), etc.
[0101] FIG. 4 depicts a block diagram of an exemplary computer system 400 suitable for implementing an embodiment. As shown in the figure, computer system 400 includes a central processing unit (CPU) 401 that can execute various processes based on a program stored in a read-only memory (ROM) 402 or a program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, data required when the CPU 401 executes various processes, etc., is also stored as needed. The CPU 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0102] The following components are connected to the I / O interface 405: an input unit 406 including a keyboard, a mouse, etc.; an output unit 407 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker; a storage unit 408 including a hard disk, etc.; and a communication unit 409 including a network interface card such as a LAN card, a modem, etc. The communication unit 409 executes a communication process via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 410 as needed, and a computer program read therefrom is installed in the storage unit 408 as needed.
[0103] Specifically, according to the exemplary embodiments disclosed in this document, the process described above with reference to FIG. 1 may be implemented as a computer software program. For example, an embodiment of the exemplary embodiment includes a computer program product including a computer program embodied tangibly on a machine-readable medium and including program code for executing method 100. In such an embodiment, the computer program may be downloaded and mounted from a network via the communication unit 409, and / or installed from the removable medium 411.
[0104] In general, various illustrative embodiments may be implemented in hardware or special purpose circuitry, software, logic, or any combination thereof. Some aspects may be implemented in hardware, and other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. Although various aspects of the illustrative embodiments are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, the blocks, devices, systems, techniques, or methods described herein may be implemented, by way of non-limiting example, in hardware, software, firmware, special purpose circuitry or logic, general purpose hardware or controllers or other computing devices, or any combination thereof.
[0105] Furthermore, the various blocks shown in the flowcharts may be viewed as method steps, and / or acts resulting from operations of computer program code, and / or multiple coupled logic circuit elements configured to perform the associated functions(s). For example, an embodiment includes a computer program product that includes a computer program embodied tangibly on a machine-readable medium and including program code configured to perform the above method.
[0106] In the context of the present disclosure, a machine-readable medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium may include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0107] The computer program code for carrying out the method of the illustrative embodiments may be written in any combination of one or more programming languages. These computer program codes may be provided to the processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus. Thereby, when the program code is executed by the processor of the computer or other programmable data processing apparatus, it causes the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on the remote computer or server. The program code may be distributed over specially programmed devices, which may generally be referred to herein as "modules". The software component part of the module may be written in any computer language, may be part of a monolithic code base, or may be developed in more discrete code parts, as is typical in object-oriented computer languages. Further, the modules may be distributed across multiple computer platforms, servers, terminals, mobile devices, etc. A given module may be implemented such that the described functions are executed by separate processors and / or computing hardware platforms.
[0108] As used in this application, the term "circuit" refers to all of the following: (a) a hardware-only circuit implementation (e.g., implemented only with analog and / or digital circuits) and (b) a combination of a circuit and software (and / or firmware), e.g., as appropriate: (i) a combination of one or more processors or (ii) a part of one or more processors / software (including a digital signal processor), software, and one or more memories that together perform various functions in a device such as a mobile phone or a server and (c) a circuit that requires software or firmware to operate, such as one or more microprocessors or a part of one or more microprocessors, even if the software or firmware does not physically exist. Further, it is well known to those skilled in the art that a communication medium typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transfer mechanism and includes any delivery medium.
[0109] Further, although operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all of the shown operations be performed, to achieve desirable results. In some circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of any of the claimed subject matter. Rather, these should be construed as descriptions of features that may be specific to a particular exemplary embodiment. Certain features described in the context of separate embodiments herein can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination.
[0110] Various modifications and adaptations to the above exemplary embodiments may become apparent to those skilled in the art upon reading the above description in connection with the accompanying drawings. Any and all modifications fall within the scope of the exemplary embodiments without limitation. Further, other exemplary embodiments within the teachings presented in the above description and drawings will occur to those skilled in the art.
[0111] Therefore, the exemplary embodiments disclosed herein may be embodied in any of the forms described herein. For example, the following numbered examples (EEE: enumerated example embodiment) describe some of the structures, features, and functions of some aspects of the exemplary embodiments disclosed herein. 〔EEE1〕 A method of storing object metadata in audio object clustering, comprising: assigning audio objects to categories, each category representing one or a unique combination of metadata requiring storage; and generating several clusters for each category through a clustering process based on the overall (maximum) number of available clusters and the overall error criterion, the method further comprising: fuzzy object category separation or hard object category separation. 〔EEE2〕 The fuzzy object category separation comprises: determining the centroids of the output clusters, for example, by selecting the most important objects; (1) the position metadata of each object
Number
Number
[0112] It will be understood that the embodiments of the exemplary embodiments disclosed herein are not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are used herein, they are not for purposes of limitation and are used only in a general, descriptive sense.
[0113] Some aspects are described below. 〔Aspect 1〕 A method for audio object clustering in which metadata is stored, comprising: classifying a plurality of audio objects into several categories based on information to be stored in the metadata associated with the plurality of audio objects; assigning a predetermined number of clusters to the categories; allocating the audio objects in each of the categories to at least one of the clusters based on the assignment; A method. 〔Aspect 2〕 The method according to aspect 1, wherein the information includes one or more of size information of the audio object, zone mask information, snap information, content type, or rendering mode. 〔Aspect 3〕 The step of classifying a plurality of audio objects into several categories based on information to be stored in the metadata associated with the plurality of audio objects is: classifying audio objects without information to be stored into one category; including classifying audio objects with different information to be stored into different categories. The method according to aspect 1. 〔Aspect 4〕 The step of assigning a predetermined number of clusters to the categories is: determining the predetermined number of audio objects from the plurality of audio objects based on the importance of each audio object relative to other audio objects; determining the distribution of the predetermined number of audio objects among the categories; including assigning the predetermined number of clusters to the categories based on the distribution. The method according to aspect 1. 〔Aspect 5〕 The step of assigning a predetermined number of clusters to the categories is: Assigning the predetermined number of clusters to the category based on reducing an overall spatial distortion for the category. The method according to aspect 1. [Aspect 6] The overall spatial distortion for the category includes a maximum spatial distortion among individual spatial distortions of the category or a weighted sum of individual spatial distortions of the category. The spatial distortion for each category is related to an original spatial position of each audio object within that category and a spatial position of at least one of the clusters. The method according to aspect 5. [Aspect 7] The method according to aspect 6, wherein a reconstructed spatial position of each audio object is determined based on the spatial position of the at least one cluster, and the spatial distortion for each category is determined based on a distance between an original spatial position of each audio object within that category and the reconstructed spatial position of that audio object. [Aspect 8] The method according to aspect 6, wherein the plurality of audio objects are within one frame of an audio signal, and the spatial position for each category is further based on a difference between the number of clusters assigned to that category in the current frame and the previous frame. [Aspect 9] Assigning the predetermined number of clusters to the category based on reducing the overall spatial distortion for the category is: an amount of spatial distortion for a category in a previous iteration step, or a difference between spatial distortions for a category in the current iteration step and the previous iteration step including sequentially and iteratively reducing the overall spatial distortion for the category based on at least one of them. The method according to aspect 5. [Aspect 10] The step of assigning the predetermined number of clusters to the category further: a first threshold for the number of clusters to be assigned to each category, a second threshold for the spatial distortion for each category or the importance of each category relative to other categories The method according to any one of Aspects 4 to 9, based on one or more of the above. 〔Aspect 11〕 The method according to Aspect 1, further comprising reclassifying at least one audio object within a category into another category based on the spatial distortion for that category. The method according to Aspect 1. 〔Aspect 12〕 The step of allocating the audio objects in each of the categories to at least one of the clusters based on the allocation is: Allocating the audio objects in each category to at least one of the clusters assigned to that category based on reducing the distortion cost associated with that category. The method according to Aspect 1. 〔Aspect 13〕 The step of allocating the audio objects in each of the categories to at least one of the clusters based on the allocation is: Allocating the audio objects in each category to at least one of the clusters assigned to one or more of the categories based on reducing the distortion cost associated with that category. The method according to Aspect 1. 〔Aspect 14〕 The distortion cost is related to one or more of the original spatial position of each audio object, the spatial position of the at least one cluster, the identification information of the category to which each audio object is classified, or the identification information of each category to which the at least one cluster is assigned, according to the method described in Aspect 12 or 13. 〔Aspect 15〕 The distortion cost is: The distance between the original spatial position of each audio object and the spatial position of the at least one cluster; The distance between the original spatial position of each audio object and the reconstructed spatial position of that audio object determined based on the spatial position of the at least one cluster; or The discrepancy between the identification information of the category to which each audio object is classified and the identification information of each category to which the at least one cluster is assigned, The method according to aspect 14, determined based on one or more of the above. [Aspect 16] Combining the audio objects within each cluster to obtain clustered audio objects; Further comprising combining the metadata of the audio objects within each cluster to obtain the metadata of the clustered audio objects, The method according to aspect 1. [Aspect 17] A system for audio object clustering where metadata is stored, comprising: An audio object classification unit configured to classify a plurality of audio objects into several categories based on information to be stored in the metadata associated with the plurality of audio objects; A cluster assignment unit configured to assign a predetermined number of clusters to the categories; An audio object allocation unit configured to allocate the audio objects in each of the categories to at least one of the clusters based on the assignment, The system. [Aspect 18] The system according to aspect 17, wherein the information includes one or more of size information of the audio object, zone mask information, snapshot information, content type, or rendering mode. [Aspect 19] The audio object classification unit further classifies audio objects with no information to be stored into one category, and classifies audio objects with different information to be stored into different categories, in the system according to Aspect 17. [Aspect 20] The cluster assignment unit includes: An importance-based determination unit configured to determine the predetermined number of audio objects from the plurality of audio objects based on the importance of each audio object with respect to other audio objects; An assignment determination unit configured to determine the assignment of the predetermined number of audio objects among the categories; and The cluster assignment unit is further configured to assign the predetermined number of clusters to the categories based on the assignment. The system according to Aspect 17. [Aspect 21] The system according to Aspect 17, wherein the cluster assignment unit is further configured to assign the predetermined number of clusters to the categories based on reducing the overall spatial distortion for the categories. [Aspect 22] The overall spatial distortion for the categories includes the maximum spatial distortion among the individual spatial distortions of the categories or the weighted sum of the individual spatial distortions of the categories, The spatial distortion for each category is related to the original spatial position of each audio object in that category and the spatial position of at least one of the clusters. The system according to Aspect 21. [Aspect 23] The reconstructed spatial position of each audio object is determined based on the spatial position of the at least one cluster, and the spatial distortion for each category is determined based on the distance between the original spatial position of each audio object within that category and the reconstructed spatial position of that audio object, the system according to aspect 22. [Aspect 24] The plurality of audio objects are within one frame of an audio signal, and the spatial position for each category is further based on the difference between the number of clusters assigned to that category in the current frame and the previous frame, the system according to aspect 22. [Aspect 25] The cluster assignment unit is further configured to: The amount of spatial distortion for a certain category in the previous iteration step, or The difference between the spatial distortion for a certain category in the current iteration step and the previous iteration step Based on at least one of which, it is configured to sequentially and iteratively reduce the overall spatial distortion for that category. The system according to aspect 21. [Aspect 26] The cluster assignment unit is further configured to assign the predetermined number of clusters to the category: A first threshold value for the number of clusters to be assigned to each category, A second threshold value for the spatial distortion for each category or The importance of each category relative to other categories Based on one or more of which, it is configured to perform the assignment, the system according to any one of aspects 20 to 25. [Aspect 27] The system further includes an audio object reclassification unit configured to reclassify at least one audio object within a certain category into another category based on the spatial distortion for that category. The system according to aspect 17. [Aspect 28] The system according to aspect 17, wherein the audio object allocation unit is further configured to allocate audio objects in each category to at least one of the clusters assigned to that category based on reducing the distortion cost associated with that category. [Aspect 29] The system according to aspect 17, wherein the audio object allocation unit is further configured to allocate audio objects in each category to at least one of the clusters assigned to one or more of the categories based on reducing the distortion cost associated with that category. [Aspect 30] The system according to aspect 28 or 29, wherein the distortion cost is related to one or more of the original spatial position of each audio object, the spatial position of the at least one cluster, the identification information of the category to which each audio object is classified, or the identification information of each category to which the at least one cluster is assigned. [Aspect 31] The distortion cost is: The distance between the original spatial position of each audio object and the spatial position of the at least one cluster; The distance between the original spatial position of each audio object and the reconstructed spatial position of that audio object determined based on the spatial position of the at least one cluster; or The mismatch between the identification information of the category to which each audio object is classified and the identification information of each category to which the at least one cluster is assigned, The system according to aspect 30, determined based on one or more of the above. [Aspect 32] An audio object combination unit configured to combine audio objects within each cluster to obtain a clustered audio object; A metadata combination unit configured to combine the metadata of the audio objects within each cluster to obtain the metadata of the clustered audio objects, and further having The system according to aspect 17. [Aspect 33] A computer program product embodied on a machine-readable medium and including computer program code for executing the method according to any one of aspects 1 to 16.
Claims
Claim 1 A method for decoding an encoded audio signal, comprising: receiving the encoded audio signal and determining at least one audio object from the encoded audio signal; receiving zone metadata indicating an area where the audio object should not be rendered; classifying the at least one audio object into at least one category based on rendering mode metadata associated with the at least one audio object; for each category, determining at least one cluster; rendering the at least one audio object based on the rendering mode metadata for the at least one cluster, wherein the rendering is further based on the zone metadata. A method.
Citation Information
Patent Citations
Scalable downmix design with feedback for object-based surround codec
US20140023196A1
Object clustering for rendering object-based audio content based on perceptual criteria
WO2014099285A1
Efficient coding of audio scenes comprising audio objects
WO2014187990A1
Transmission device, transmission method, reception device, and reception method
WO2016039287A1