Headphone rendering metadata-preserving spatial coding with speaker optimization

By integrating a speaker optimization process into object-based audio systems, the challenges of speaker dropout during solo playback are addressed, resulting in enhanced audio quality and reduced dropout issues.

WO2025128413A1PCT designated stage expired Publication Date: 2025-06-19DOLBY LABORATORIES LICENSING CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058803
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-25
Filing Date
2024-12-06
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing object-based audio systems face challenges in transmitting and rendering audio content with hundreds of individual objects due to bandwidth limitations and computational resource constraints, leading to issues like speaker dropout during solo playback.

Method used

The implementation of a speaker optimization process within the object-based audio system, which includes a cluster selection module, a speaker dropout monitoring module, and an object-to-speaker gain calculation module, to adjust the clustering of audio objects and mitigate speaker dropout issues.

Benefits of technology

This approach effectively reduces the impact of speaker dropout by optimizing the clustering of audio objects based on speaker dropout metrics, ensuring improved single-channel quality and maintaining overall audio quality across various playback scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058803_19062025_PF_FP_ABST
    Figure US2024058803_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods of clustering audio objects. Example systems and devices are described that include a cluster selection module and a speaker dropout monitoring module. Systems and devices may also include a speaker optimization process that is selectively enabled when the speaker dropout monitoring modules determines that one or more speakers in the object-based audio system has a dropout issue. Example methods for clustering audio objects are described that may include the steps of receiving an input audio block that includes a plurality of audio objects and calculating object-to-speaker gains for the plurality of audio objects. The methods may include identifying a speaker as experiencing a dropout issue based on the object-to-speaker gains, and clustering the audio objects based on whether the speaker is experiencing dropout issues.
Need to check novelty before this filing date? Find Prior Art

Description

HEADPHONE RENDERING METADATA-PRESERVING SPATIAL CODING WITH SPEAKER OPTIMIZATION CROSS-REFERENCE TO RELATED APPLICAITONS

[0001] This application claims the benefit priority of International Application No. PCT / CN2023 / 137864 filed December 11, 2023, U.S. Provisional Application No.63 / 613,203 filed December 21, 2023, International Application No. PCT / CN2024 / 114322 filed August 23, 2024, and U.S. Provisional Application No.63 / 698,765 filed September 25, 2024, each of which is hereby incorporated by reference in their entireties. TECHNICAL FIELD

[0002] This application relates generally to audio processing systems, methods, and devices that include object clustering. BACKGROUND

[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.

[0004] An object-based audio system implements an object-based audio format that includes both beds and objects. Audio beds refer to audio channels that are meant to be reproduced in predefined, fixed locations while audio objects refer to individual audio elements that may exist for a defined duration in time but also have spatial information of each object, such as position, size, and the like. During transmission, beds and objects may be sent together and then used by a spatial reproduction system to recreate the artistic intent. These reproduction systems often include a variable number of speakers or headphones.

[0005] It is with respect to these and other considerations that the disclosure made herein is presented. BRIEF SUMMARY OF THE DISCLOSURE

[0006] Techniques are described for clustering audio objects in object-based audio systems. Various embodiments described herein provide systems, methods, and devices with a speaker optimization process that is employed for clustering audio objects in object-based audio systems.

[0007] In some examples, systems and devices are described that include a cluster selection module and a speaker dropout monitoring module. Systems and devices may also include a 1  speaker optimization process that is selectively enabled when the speaker dropout monitoring modules determines that one or more speakers in the object-based audio system has a dropout issue.

[0008] In some additional examples, a speaker optimization process may include the steps of receiving a plurality of audio objects, evaluating objects in the audio system to determine if a speaker optimization process should be enabled, and optimizing the clustering of audio objects for speaker dropout mitigation when the speaker optimization process is enabled. When a dropout issue is identified by the speaker optimization process, the selection of audio objects in clustering may be adjusted to minimize the impact of the speaker dropout issue by way of object- to-speaker gain calculations.

[0009] In yet further examples, a method for clustering audio objects may include the steps of receiving an input audio block that includes a plurality of audio objects and calculating object-to- speaker gains for the plurality of audio objects. The method may include identifying a speaker as experiencing a dropout issue based on the object-to-speaker gains, and clustering the audio objects based on whether the speaker is experiencing dropout issues.

[0010] According to some examples, speaker optimization may be enabled based on the number of remaining clusters to be allocated and the number of speakers having dropout issues. For example, the speaker optimization determination module may enable a speaker optimization process in the cluster selection module based on the number of remaining clusters to be allocated and based on the speakers experiencing dropout issues. Dropout issues may be identified by a speaker dropout monitoring module based on prior speaker loudness and posterior speaker loudness, both of which are calculated based on the object-to-speaker gains. The posterior speaker loudness may be estimated based on object-to-cluster gains and / or cluster-to-speaker gains, or estimated based on a spatial distance metric between object pairs.

[0011] In further examples, the speaker optimization is implemented by adding weights to object importance based on object-to-speaker gains. Accordingly, when speaker dropout is identified and speaker optimization is enabled, the clustering of audio objects may be adjusted (e.g., altered). The clustering of audio objects may also include using a centroid determination process that selects the most perceptually important objects, such as referring to speaker loudness.

[0012] In some examples, when the speaker optimization process is enabled, clustering of audio objects is based on a number of remaining clusters to be allocated and the number of speakers having dropout issues. A dropout speaker-concentration metric may be applied for cluster 2  selection when the speaker optimization process is enabled. When the speaker optimization process is not enabled, clustering of the audio objects is based on a partial loudness of the plurality of audio objects in a critical band. The partial loudness may be based on an excitation of each audio object. Information and / or concepts relating to the critical bands (for example, perceptual frequency banding such as Bark or Equivalent Rectangular Bandwidth [ERB]) may be utilized to calculate the excitation of each audio object.

[0013] In yet other examples, a preset speaker layout is implemented for cluster selection, speaker optimization determination, and / or speaker dropout monitoring. Speaker layouts may include a number of speakers, a location of speakers, and combinations thereof.

[0014] In additional examples, generating clusters includes applying a cost function. The overall cost of the cost function may be determined for the cluster positions based on the object-to- cluster gains. In some instances, the overall cost includes penalty terms, such as using an extended hybrid distance metric. The overall cost may also be a linear combination of a sub-cost of each penalty term, and may combine at least one positional distance metric describing differences in object position, a metric representing similarity of dissimilarity in HRM, and a loudness, level, or importance metric of the audio objects. In some instances, the audio objects are rendered to the cluster positions by minimizing the overall cost. In other instances, an iterative greedy approach is applied to select the audio objects for clustering based on a maximum partial loudness, overall loudness, energy, level, salience, or importance of the audio objects.

[0015] In this manner, various aspects of the present disclosure provide for processing of audio signals, and effect improvements in at least the technical fields of audio encoding, audio decoding, virtual reality, and the like.

[0016] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.

[0017] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associated drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims. 3  DESCRIPTION OF THE DRAWINGS

[0018] These and other more detailed and specific features of various embodiments are more fully disclosed in the following description, reference being had to the accompanying drawings, in which:

[0019] FIG.1 illustrates a block diagram of an example audio coding system in which various aspects of the present disclosure can be practiced.

[0020] FIG.2 illustrates a block diagram of an example spatial coding method that can be implementing in the audio system of FIG.1 in accordance with various aspects of the present disclosure.

[0021] FIG.3 illustrates a block diagram of various example methods for rendering clusters of audio objects, which may be performed by the spatial coding component of FIG.1, in accordance with various aspects of the present disclosure.

[0022] FIG.4 illustrates a block diagram of an example spatial coding component, such as the spatial coding component of FIG.1.

[0023] FIG.5A illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.

[0024] FIG.5B illustrates a schematic block diagram of an example CPU implemented in the device architecture of FIG.5A that may be used to implement various aspects of the present disclosure. DETAILED DESCRIPTION

[0025] In the following description, numerous details are set forth, such as audio device configurations, timings, operations, and the like, in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to one skilled in the art that these specific details are merely examples and not intended to limit the scope of this application.

[0026] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least 4  the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0027] Various Acronyms which may appear throughout this disclosure and in the associated claims and / or drawings are listed below. Other commonly used acronyms and terms of art may be excluded from this list in the interest of brevity. Thus, a short list of acronyms is provided below as an easy reference for the reader. HRM – Headphone Rendering Mode HRMpSC – Headphone Rendering Mode-preserving Spatial Coding ETSI – European Telecommunications Standard Institute ERB – Equivalent Rectangular Bandwidth

[0028] For entertainment content (including both linear and interactive content), transmitting the original object-based audio signal, which may contain hundreds of individual objects, can be challenging because endpoints that support object-based audio typically have limitations with respect to the maximum number of audio objects that can be supported, e.g., due to the limited amounts of computational resources and / or memory. Additionally, the bitrate required to transmit the audio content to an endpoint over a bandwidth-limited transmission channel may be limited due to the limited amounts of computational resources and / or memory. In such cases, application of efficient object-based audio-scene management is desirable.

[0029] In some cases, “spatial coding” may be used to reduce the complexity of an audio scene. Some spatial coding methods employ clustering techniques that aim to reduce the number of input objects and / or beds to a smaller set of output objects (hereafter referred to as clusters) with minimal impact on the audio quality.

[0030] Herein, the terms “clustering,” “grouping,” and “combining” may be used interchangeably to describe combinations of objects and / or beds (or channels) configured to reduce the amount of data in a unit of adaptive audio content for transmission and rendering in an audio playback system. The terms “compression” and “reduction” may be used to refer to an 5  act of performing audio scene simplification, e.g., via such clustering. The terms “clustering,” “grouping,” and “combining” throughout this description are not limited to a strictly unique assignment of an audio object or bed to a single cluster only. In some cases, an audio object or bed may be distributed over more than one output bed or cluster using weights or gain vectors that determine the relative contribution of an object or bed signal to the output cluster or output bed signal.

[0031] FIG.1 illustrates a block diagram of an example audio coding system 100 in which various aspects of the present disclosure can be practiced. The audio system 100 includes a spatial coding component 110. The audio system 100 also includes an audio encoder 120 and an audio decoder 140 coupled via a bandwidth-limited communication channel 130. The spatial coding component 110 receives input signals 102 and processes the input signals 102 to generate input signals 112. For example, the spatial coding component 110 applies spatial coding techniques to the input signals 102 to simplify the audio scene represented by the input signals 102, e.g., via clustering.

[0032] The spatial coding component 110 provides the input signals 112 to the encoder 120. The input signals 104 bypass the spatial coding component 110 and are coupled directly to the encoder 120. The encoder 120 receives input signals 104, 112 and processes the received input signals to generate an encoded bitstream 132, which is transmitted over the communication channel 130 to the decoder 140. The decoder 140 receives the encoded bitstream 132 via the communication channel 130 and decodes the received bitstream to generate output audio signals 142. The output audio signals 142 are applied to an audio rendering component 150 that operates to render and / or playback the audio content represented by the audio signals 142.

[0033] In various examples, the audio rendering component 150 may include any professional or consumer-grade audio system, such as a home theater (e.g., including an A / V receiver, a soundbar, a Blu-ray player, etc.), one or more E-media devices (e.g., a computer, a tablet, a mobile phone equipped with headphones and / or speakers, etc.), a TV set, and a sound reproduction system. In some examples, the audio rendering component 150 provides an audio environment for playback of audio or audio / visual content using a plurality of speakers and suitable playback devices. In some examples, the audio rendering component 150 may represent any environment in which a listener is experiencing playback of the audio content, such as a cinema, a concert hall, an outdoor theater, a home theater or room, a listening booth, a car, a game console, a headset device, a public address (PA) system, or other audio playback environment. In other examples, the audio rendering component is an audio process that renders 6  the output audio signals 142 and provides the rendered audio to a separate playback device, such as one or more speakers, a set of headphones, or another processing device.

[0034] The audio coding system 100 may implement object-based audio formats that include both beds and objects. Audio beds refer to audio channels that are meant to be reproduced in predefined, fixed locations. Audio objects refer to individual audio elements that may exist for a defined duration in time but also have spatial information of each object, such as position, size, and the like. During transmission, beds and objects may be sent separately and then used by a spatial reproduction system to recreate the artistic intent. The reproduction system may include a variable number of speakers or headphones. Object-based audio as referred to herein may include interactive content that is typically generated at runtime as the user interacts with the corresponding virtual scene. Object-based audio as referred to herein may also include linear or static content, such as music content, that has at least a set position in space.

[0035] Due to the bandwidth limitation of distribution and transmission systems, transmitting the original object-based audio signal which may contain hundreds of individual objects may be challenging and processing intensive. To reduce the complexity of the audio scene, a series of clustering techniques may be implemented and may be referred to as “Spatial Coding”. These Spatial Coding techniques aim to reduce the number of input objects and beds into a small set of output objects (hereafter referred to as clusters) via clustering techniques with minimal impact (e.g., non-perceivable by a user) on audio quality. In general, object clustering methods described herein include two steps: a cluster centroid determination step to determine the cluster position and associated metadata, and a cluster generation step to calculate the object to cluster gains and generate the determined clusters.

[0036] In addition to the metadata describing spatial position, other types of metadata may also indicate the rendering requirements, such as a snap and zone mask in speaker rendering scenarios. In headphone rendering scenarios, the object may be associated with metadata describing the headphone rendering modes (HRMs). HRMs as described herein are a specific example of Headphone Rendering Metadata, which may be understood as a general class of metadata for controlling headphone rendering. The HRMs are usually created by artists and content creators during the content creation phase. The HRMs indicate whether the virtualization techniques should be applied or not (e.g., a “bypass” mode) for binaural rendering. Further, the desired room effects for virtualization may also be indicated. For example, an object may carry the HRM with either “near”, “far”, or “middle”, which indicate three types of distance 7  scale from object to the head center. The HRMs including “bypass”, “near”, “far”, and “middle” should be preserved through clustering to preserve the artist’s intention.

[0037] Embodiments described herein may implement Headphone Rendering Mode-preserving Spatial Coding (HRMpSC), which supports the delivery of audio content with HRM. The HRMpSC algorithm generates a number of clusters to achieve good perceptual quality for both headphone and speaker playback. The terms “speaker” and “channel”, as referred to herein, may be used interchangeably.

[0038] In previous evaluations of speaker playback, efforts were focused on overall quality, e.g., playback of all active speakers / channels (hereafter referred to as “full playback”). However, issues are identified when soloing one or more specific channel(s) / speaker(s), including the issue of dropout where the channel that was active becomes nearly silent after performing HRMpSC. Although dropout issues may be hardly perceived in speaker full playback, it becomes conspicuous for quality assurance where soloing channel is one approach.

[0039] Accordingly, optimization of the HRMpSC in terms of good single channel quality for soloing speakers is desired without causing significant quality degradation for full playback. Embodiments described herein provide a “Speaker Optimized HRMpSC” technique to use a fixed number of clusters to achieve good quality for headphone playback, speaker full playback, and speaker solo playback simultaneously.

[0040] FIG.2 illustrates a block diagram of an example spatial coding method 200 that can be implemented in the spatial coding component 110 of the audio system 100 according to some examples. The method 200 provides a spatial coding of input audio objects into clusters of objects in a manner based on a centroid determinization process. The method 200 includes a cluster centroid determination module 202 and an object clustering module 204.

[0041] The inputs of the cluster centroid determination module 202 correspond to a first path 212 and a second path 214. The output of the cluster centroid determination module 202 corresponds to a third path 216. The cluster centroid determination module 202 is configured to receive input audio objects from the first path 212 and a first extended hybrid distance d2 from the second path 214. An output of the cluster centroid determination module 202 is connected to the object clustering module 204 via the third path 216. The cluster centroid determination module 202 evaluates the input audio objects (e.g., processes, analyzes) using the cluster position and associated metadata to generate centroid information. The cluster centroid determination module 202 outputs the centroid information to the object clustering module 204 via the third 8  path 216. The inputs of the object clustering module 204 correspond to the third path 216 and a fourth path 218. The output of the object clustering module 204 corresponds to a fifth path 220. The object clustering module 204 receives the centroid information from the cluster centroid determination module 202 via the third path 216 and receives a second extended hybrid distance d2’ via the fourth path 218. The object clustering module 204 calculates object to cluster gains and generates the clusters based on the object to cluster gains. The clusters are output via the fifth path 220 from the spatial coding, for example, as the input signals 112 that are received by the encoder 120 from the spatial coding component 110. The method 200 repeats as the centroid positions and HRM are determined one by one until a target cluster count is reached. The first extended hybrid distance d2 and the second extended hybrid distance d2’ may be, for example, penalty terms used by the cluster centroid determination modules 202 and the object clustering module 204, respectively.

[0042] FIG.3 illustrates a block diagram of various example methods 300 for rendering clusters of audio objects, which may be performed by the spatial coding component 110 of FIG.1. The methods 300 may be performed by a processor, which may be configured to perform methods 300 via machine-executable instructions. The methods 300 may be broken into various blocks or partitions, such as blocks 302, 304, 306, and 308. The various process blocks illustrated in FIG. 3 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 302.

[0043] At block 302, “Receiving Audio Objects”, an example method 300 may include receiving a plurality of audio objects. Each audio object of the plurality of audio objects may be associated with respective object metadata that indicates respective spatial position information and an HRM. The HRM may be a value indicating, for example, “bypass”, “near”, “far”, or “middle”. Processing may continue from block 302 to block 304.

[0044] At block 304, “Determining Cluster Positions”, an example method 300 may include determining, by computation or estimation, a plurality of cluster positions. For example, a plurality of cluster positions is determined by applying an extended hybrid distance metric to a spatial coding algorithm to calculate a partial loudness for each of the audio objects. In some embodiments, the extended hybrid distance metric integrates an HRM distance into a hybrid distance. In some embodiments, the hybrid distance combines Euclidean and angular distance. In some embodiments, a computation of the HRM distance is adaptive to different audio scenes in 9  terms of spatial complexity. In some embodiments, the HRM distance functions as a scaling factor for calculating a distance between pairs of the audio objects when determining the cluster positions. In some embodiments, the HRM distance functions as a scaling factor for calculating a distance between each of the audio objects and each of the clusters when rendering the audio objects to the cluster positions. In some embodiments, the extended hybrid distance metric is applied to the spatial coding algorithm to ensure positional correctness and preserve the HRM. In some embodiments, the cluster positions are determined according to a target cluster count. In some embodiments, the target cluster count is set according to an available bandwidth or an expected bitrate. In some embodiments, each of the cluster positions is determined by an iterative greedy approach. In some embodiments, the iterative greedy approach includes selecting the audio object with a maximum partial loudness, overall loudness, energy, level, salience, or importance. Processing may continue from block 304 to block 306.

[0045] At block 306, “Rendering The Audio Objects To The Cluster Positions To Form Clusters”, an example method 300 may include rendering the plurality of audio objects to the plurality of cluster positions to form a plurality of clusters. For example, the audio objects are rendered to the cluster positions to form a plurality of clusters by applying the extended hybrid distance metric to the spatial coding algorithm to calculate object-to-cluster gains. In some embodiments, an overall cost when calculating the object-to-cluster gains includes a plurality of penalty terms. In some embodiments, at least one of the penalty terms uses the extended hybrid distance metric. In some embodiments, the overall cost is defined as a linear combination of a sub-cost of each of the penalty terms. In some embodiments, the overall cost combines at least one positional distance metric describing differences in object position; a metric representing similarity or dissimilarity in HRM; and a loudness, level, or importance metric of the audio objects. In some embodiments, the audio objects are rendered to the cluster positions by minimizing the overall cost. In some embodiments, a first set of parameters is used when applying the extended hybrid distance metric to determine the cluster positions. In some embodiments, a second set of parameters is used when applying the extended hybrid distance metric to render the audio objects to the cluster positions. In some embodiments, each of the clusters includes cluster audio data and associated cluster metadata. In some embodiments, the cluster audio data is determined by applying the object-to-cluster gains to audio data of each of the audio objects rendered to the respective cluster. In some embodiments, the cluster metadata includes the cluster position of the associated cluster and a cluster HRM. In some embodiments, at least one of the object metadata associated with each of the audio objects rendered to a cluster 10  is preserved to the respective associated cluster metadata.  Processing may continue from block 306 to block 308.

[0046] At block 308, “Transmitting The Clusters To A Reproduction System”, an example method 300 may include transmitting the plurality of clusters to a spatial reproduction system. For example, the spatial coding component 110 provides the plurality of clusters to the encoder 120 as a portion of the input signals 112.

[0047] Further details regarding the spatial coding method 200 and the example methods 300 can be found in PCT Patent Application Publication No. WO2023039096, “Systems And Methods For Headphone Rendering Mode-Preserving Spatial Coding,” incorporated herein by reference in its entirety.

[0048] To improve single channel quality, embodiments described herein implement a novel HRMpSC technique. FIG.4 illustrates a block diagram of an example spatial coding component, such as the spatial coding component 110 of FIG.1. The spatial coding component 110 may perform operations such as that performed by the cluster centroid determination module 202. Additionally, the spatial coding component 110 may implement HRMpSC techniques described herein. The spatial coding component 110 selects objects to be in the cluster in an iterative manner up to a preset cluster count. In an initial iteration, the “most important” object is selected for the cluster, as described below in more detail. In the subsequent iteration, the next “most important” object is selected from the remaining objects for the cluster. This continues until the preset cluster count is satisfied.

[0049] The example spatial coding component 110 includes an object-to-speaker gain calculation module 402, a speaker dropout monitoring module 404, a speaker optimization determination module 406, and a cluster selection module 408. The input of the object-to- speaker gain calculation module 402 corresponds to an input path 410. The output of the object- to-speaker gain calculation module 402 corresponds to an object-to-speaker gains signal path 412. The object-to-speaker gain calculation module 402 is configured to receive a plurality of audio objects via input signal path 410. The object-to-speaker gain calculation module 402 may first determine (for example, by calculation or estimation) a target (or reference) speaker layout as the optimization target. The target speaker layout may be a preset layout for a particular application that includes, for example, a number of speakers and / or a location of the speakers. For example, the target speaker layout may be 7.1.4. As speaker layouts used in practical playback is unpredictable, the object-to-speaker gain calculation module 402, the target speaker layout provides an assumed (high standard) layout by which the method 200 aims to optimize. If 11  single channel quality is well preserved for the target layout (e.g., 7.1.4), then playback on layouts with fewer speakers (such as 5.1.2 or 5.1) do not experience dropout effects.

[0050] For given speaker layout and object metadata, the object-to-speaker gains may be obtained by, for example, the methods proposed for AC-4 in section 5.2 of the technical standard TS 103448 from the European Telecommunications Standard Institute (ETSI). The object-to- speaker gains may be derived based on either amplitude preserving or energy preserving. Herein, the notation Gs,j denotes the object-to-speaker gains (rather than taking the square, i.e., G2s,j), where s is the speaker index and j is the object index. The object-to-speaker gain calculation module 402 is further configured to provide an object-to-speaker gains signal that indicates the object-to-speaker gains to the speaker dropout monitoring module 404 and the cluster selection module 408 via object-to-speaker gains signal path 412.

[0051] The input of the speaker dropout monitoring module 404 corresponds to the object-to- speaker gains signal path 412. The input of the speaker dropout monitoring module 404 also corresponds to the input signal path 410. The input of the monitoring module 404 further corresponds to the output signal path 418 from the cluster selection module 408. The output of the speaker dropout monitoring module 404 corresponds to an identified dropout speakers signal path 414. The speaker dropout monitoring module 404 is configured to receive the object-to- speaker gains signal from the object-to-speaker gain calculation module 402. The speaker dropout monitoring module 404 is further configured to receive the plurality of audio objects via the input signal path 410. The speaker dropout monitoring module 404 is yet further configured to receive clusters from the cluster selection module 408, as will be described below in more detail. The speaker dropout monitoring module 404 identifies dropout speakers based on prior speaker loudness and posterior speaker loudness. For example, to monitor the single channel quality of each speaker, the speaker dropout monitoring module 404 estimates the loudness of sound assigned to each speaker via direct rendering of the objects to speakers without clustering. The speaker dropout monitoring module 404 estimates the loudness prior to cluster selection. The object total loudness is calculated according to Equation 1: ^^^^^^ఈ ^^^ ൌ ∑ ^ ൬^^^^^^^^ [Equation 1] ^where ^^^^^ of object j at an initial iteration, ^^^^^^ ^^^^ to denote the excitation of object ^^ in critical band ^^, and ^^ is a model parameter. Since the excitation is updated in each iteration, the notation of superscript with square brackets ^^^^ denotes the iteration index. The 12  initial ^^^^^^ ^^^^ represents the original excitation, which may be used for the 1stiteration, i.e., ^^^^^^ ^^^^ ൌ ^^^^^^ ^^^^. dropout monitoring module 404 may calculate the prior speaker loudness bytaking linear combination of the object total loudness with object-to-speaker gain as the coefficients, as provided by Equation 2: Γൌ ∑^^^ ^^ ^^^,^ ^^^[Equation 2] where Γ^isloudness for a speaker s.

[0053] In each iteration, the posterior speaker loudness may also be calculated by the speaker dropout monitoring module 404 to monitor how much loudness is accumulated for each speaker. The posterior speaker loudness may be updated once the cluster has been determined for the given iteration. For example, suppose ^^∗^^^was determined as the cluster in the i-th iteration. The posterior speaker loudness isaccording to Equation 3: ^^^^^ ^^ି^^^ ൌ ^^^ ^ ∑ ^^^^ ^^^,^ ^^ ∗^ ^1 െ ^^ா൫^^, ^^^^^൯^ [Equation 3] where ^^ ^^,j and ^^∗^^^, which is defined by^^ ^^^, ^^^ ൌ 1 െ ^ ^ఈ ா൫ ^^ா ^^, ^^ ൯[Equation 4] where ^^ாis similar to ^^ுexcept for that it only takes the Euclidean (or spatial) distance into account. One example implementation of ^^ாis provided by Equation 5: ^^ ^^ ^గௗಶா ^, ^^ ൌ cosଶ ^ఛಶ^ [Equation 5]where ^^ cut-off threshold, ^^ாis the Euclidean distance between objects i and and⌊⋅⌋represents the function to restrict values to ^0,1^.

[0054] According to the above Equations 1-5, for a given object j, the term ^1 െ ^^ ∗ா൫^^, ^^^^^൯^inEquation 3 determines (e.g., defines) the percentage of object loudness thaton speakers in the current iteration. For example, for the selected object ^^ ൌ ^^∗^^^ (i.e., the cluster),^^ா^^∗^^^, ^^∗^^^൯ ൌ 0, and therefore ^1 െ ^^ ∗^^^ா൫^^, ^^ ൯^ ൌ 1.example(e.g., the selected cluster) will unreservedly contribute its loudness (e.g., contribute 100% of its loudness) to the speakers. 13

[0055] As another example, if there is a nearby object ^^′ which is very close to the cluster, the object ^^′ may also contribute a portion of its loudness to speakers, where the percentage depends on the Euclidean distance to the cluster. Nearby objects may be (at least) partially rendered to the selected object in the real clustering step performed by the object clustering module 204, and hence the final cluster may contain a certain percentage of loudness contributed by its nearby objects. Once the cluster has been determined in this iteration, the selected object ^^∗^^^as well as the nearby objects are considered to have positive impact on the relevant

[0056] The speaker dropout monitoring module 404 may update the object loudness ^^^^^^ at the end of each iteration according to Equation 6:^^^^ା^^ ∗^^^^ൌ ^^ா൫^^, ^^^^^൯ ^^^[Equation 6]

[0057] The^^^^ା^^^ is then implemented by the speaker dropout monitoring module 404 in the next Additionally, objects nearby to the cluster areaccounted for, as otherwise the loudness ^^^^^^ is partially removed for nearby objects prior to the next iteration, and the removed not accumulated into the calculation of the posteriorspeaker loudness.

[0058] With both the prior speaker loudness Γ and the posterior s^^^^ peaker loudness ^^^calculated, the speaker dropout monitoring module 404 calculates a speaker dropout metric according to Equation 7: ^ ^ ^^^ ൌఊ^ ^^^ೞ^ೞ[Equation 7]

[0059] The speaker dropout monitoring module 404 may identify dropout speakers by identifying all speakers s in which ^^^^^^ is less than a threshold. For example, the speaker dropout monitoring module 404 calculatesspeaker dropout metric ^^^^^^ for each speaker, and compares the speaker dropout metric ^^^^^to a pre^^^^ determined threshold. Speakers s for which ^^^is less than the predeterminedare identified as dropout speakers. The speakermonitoring module 404 may be configured to generate an identified dropout speakers signal that indicates which speakers s are dropout speakers. The speaker dropout monitoring module 404 is configured to provide an identified dropout speakers signal to the speaker optimization determination module 406 via the identified dropout speakers signal path 414. 14

[0060] The input of the speaker optimization determination module 406 corresponds to the identified dropout speakers signal path 414. The input of the speaker optimization determination module 406 also corresponds to the output signal path 418 from the cluster selection module 408. The output of the speaker optimization determination module 406 corresponds to a speaker optimization enablement signal path 416. The speaker optimization determination module 406 is configured to receive the identified dropout speakers signal from the speaker dropout monitoring module 404. The speaker optimization determination module 406 is also configured to receive the clusters from the cluster selection module 408.

[0061] The speaker optimization determination module 406 determines (via calculation or estimation) whether speaker optimization should be enabled in the current iteration. In some instances, the speaker optimization determination module 406 may enable speaker optimization in the current iteration based on the number of cluster budget remaining (denoted by ^^^^^) and based on the number of speakers experiencing dropout (denoted by ^^^^^). The number of speakers experiencing dropout may be determined (via calculation or estimation) by the speaker optimization determination module 406 by summing the number of speakers indicated in the identified dropout speakers signal from the speaker dropout monitoring module 404.

[0062] For example, when ^^^^^> ^^^^^, there are opportunities to optimize dropout channels in later (e.g., subsequent) iterations, and speaker optimization may or may not be enabled in the current iteration. When ^^^^^= ^^^^^, there may not be adequate clusters available to optimize the dropout speakers in later iterations, and speaker optimization may be enabled for the current iteration.

[0063] While first performing speaker optimization and then performing cluster selection as previously understood by one skilled in the art, regardless of whether ^^^^^= ^^^^^(or ^^^^^≤ ^^^^^), may ensure a good single channel quality, the overall quality of themay beparticularly when the selected clusters are instable over multiple frames. This degradation of audio quality impacts headphone playback where speaker dropout is not a concern. In contrast, when performing speaker optimization only when ^^^^^≤ ^^^^^, the most perceptually important objects may be selected in the former iterations.good overall quality is achieved with high priority while optimizing the cluster selection strategy in the last few iterations to alleviate speaker dropout issues.

[0064] When ^^^^^≤ ^^^^^, the speaker optimization determination module 406 may also be configured towhich speaker is optimized in the given iteration when multiple 15  speakers are experiencing dropout. In one instance, the speaker optimization determination module 406 identifies which speaker is optimized based on a preset order. As one example (such as for 7.1.4), the preset speaker order may be L, R, C, Lss, Rss, Lrs, Rrs, Ltf, Rtf, Ltr, Rtr.

[0065] The speaker optimization determination module 406 is configured to generate a speaker optimization enablement signal indicating whether speaker optimization is enabled. The speaker optimization enablement signal may also indicate which speaker should be optimized. The speaker optimization determination module 406 is configured to provide the speaker optimization enablement signal to the cluster selection module 408 via the speaker optimization enablement signal path 416.

[0066] The input of the cluster selection module 408 corresponds to the input signal path 410. The input of the cluster selection module 408 also corresponds to the speaker optimization enablement signal path 416. The cluster selection module 408 is configured to receive the plurality of audio objects via the input signal path 410. The cluster selection module 408 is also configured to receive the speaker optimization enablement signal from the speaker optimization determination module 406. The output of the cluster selection module 408 corresponds to an output signal path 418. The cluster selection module 408 generates the clusters of audio objects.

[0067] For example, in each iteration, the cluster selection module 408 calculates the perceptual importance for each audio object included in the plurality of audio objects. The most perceptually important object is selected by the cluster selection module 408 as the cluster. Clustering is then performed based on whether speaker optimization is enabled for the given iteration, as indicated by the speaker optimization enablement signal from the speaker optimization determination module 406.

[0068] When speaker optimization is not enabled for the iteration, the cluster selection module 408 refers to partial loudness of the audio objects for generating the clusters. For example, the specific loudness ^^^ᇱ^^^^ of object j can be calculated according to Equation 8: ^^ఈ ఈ ^ᇱ ^^^ ^^^^ ൌ ^^^ ^ ∑^ ^^^^^^^^^^^ െ ^^^ ^ ∑^ ^^ ^^^^ ^^^^ ൫1 െ ^^ு^^^, ^^^൯ ^[Equation 8]whereof masking (whichdependent on the distance of the two objects ^^ and ^^). For the HRMpSC, ^^ுis the function intowhich the HRM distance is incorporated. For example, the case ^^ு^^^, ^^^ → 1 means the twoobjects are close in position and similar in HRM. In contrast, large spatial distance and / or HRMsimilarity results in ^^ு^^^, ^^^ → 0. In some instances, HRMpSC combines the factors by means of16  ^^ு^^^, ^^^ ൌ ^^^^^ு, where ^^^ and ^^ு are the inverse of hybrid distance 1 and HRM distance,respectively.

[0069] The cluster selection module 408 calculates the partial loudness of the object i by taking the sum of the specific loudness ^^^ᇱ^^^^ across auditory filters b, as provided by Equation 9: ^^^^^^ ൌ ∑^ ^^ᇱ^^^^ ^^^^ [Equation 9]

[0070] The408 selects the cluster by selecting the object with maximum partial loudness, as provided by Equation 10: ^^∗^^^ ൌ argmax ^^^ ^^^^^[Equation 10]^

[0071] When the cluster isthe cluster selection module 408 updates the excitation of all the audio objects, as provided by Equation 11: ^^^^ା^^^^^^ ൌ ^ ^^^ ∗^^^^^^ ^1 െ ^^ு൫^^, ^^൯^ [Equation 11] The updatedin the next iteration.

[0072] Alternatively,speaker optimization is enabled such that cluster selection module 408 optimizes speaker ^^∗^^^in the ^^-th iteration, a dropout speaker-concentrated metric is applied by the cluster selection module 408 for cluster selection. The dropout speaker-concentrated metric ^^^is provided by Equation 12: ^^ ^^^^^ ൌ ^^^∗,^^^^ ^^^^^^ ^ [Equation 12] where ^^^∗,^gain of object ^^ to speaker ^^∗, and where ^^^is the Sone-to- Phon conversion function. In this manner, the cluster selection module 408 provides speaker optimization by adding a weight value or gain to the importance of objects. The weight value may be based on the object-to-speaker gains.

[0073] The cluster selection module 408 selects the cluster by selecting the object which maximizes the dropout-speaker concentrated metric, as provided by Equation 13: ^^∗^^^ ൌ argmax^^^^^^^ ^ [Equation 13]

[0074] When the cluster is determined, the cluster selection module 408 updates the excitation of all the audio objects, as described with respect to Equation 11. The cluster selection module 408 outputs the clusters via the output signal path 418.

[0075] FIG.5A illustrates a schematic block diagram of an example device architecture 500 (e.g., an apparatus 500) that may be used to implement various aspects of the present disclosure. Architecture 500 includes but is not limited to servers and client devices, systems, and methods as described in reference to FIGS.1-4. As shown, the architecture 500 includes central processing unit (CPU) 501 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 502 or a program loaded from, for example, storage unit 508 to random access memory (RAM) 503. The CPU 501 may be, for example, an electronic processor 501. In RAM 503, the data required when CPU 501 performs the various processes is also stored, as required. CPU 501, ROM 502, and RAM 503 are connected to one another via bus 504. Input / output interface 505 is also connected to bus 504.

[0076] The following components are connected to I / O interface 505: input unit 506, that may include a keyboard, a mouse, or the like; output unit 507 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 508 including a hard disk, or another suitable storage device; and communication unit 509 including a network interface card such as a network card (e.g., wired or wireless).

[0077] In some implementations, input unit 506 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0078] In some implementations, output unit 507 include systems with various number of speakers. Output unit 507 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0079] In some embodiments, communication unit 509 is configured to communicate with other devices (e.g., via a network). Drive 510 is also connected to I / O interface 505, as required. Removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 510, so that a computer program read therefrom is installed into storage unit 508, as required. A person skilled in the art would understand that although apparatus 500 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these 18  components and all these modifications or alteration all fall within the scope of the present disclosure.

[0080] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 509, and / or installed from the removable medium 511, as shown in FIG.5A.

[0081] FIG.5B illustrates a schematic block diagram of an example CPU 501 implemented in the device architecture 500 of FIG.5A that may be used to implement various aspects of the present disclosure. The CPU 501 includes an electronic processor 520 and a memory 521. The electronic processor 520 is electrically and / or communicatively connected to the memory 521 for bidirectional communication. The memory 521 stores spatial coding software 522. In some examples, memory 521 may be located internal to the electronic processor 520, such as for an internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, memory 521 may be located external to the electronic processor 520, such as in a ROM 502, a RAM 503, flash memory or a removable medium 511, or another non-transitory computer readable medium that is contemplated for device architecture 500. In some instances, the electronic processor 520 may implement the spatial coding software 522 stored in the memory 521 to perform, among other things, any of the methods 300 of FIG.3 or the example operations of the spatial coding component 110 of FIG.4.

[0082] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 501 in combination with other components of FIG.5A), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, 19  firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0083] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0084] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine- readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0085] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0086] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims. Various aspects and implementations of 20  the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.

[0087] EEE1. A method for clustering audio objects, the method comprising: a cluster selection module with an optional speaker optimization process; a speaker optimization determination module to determine if the speaker optimization process should be enabled in the cluster selection module; and a speaker dropout monitoring module to determine if a speaker has a dropout issue.

[0088] EEE2. The method according to EEE1, wherein the speaker optimization determination module makes decisions based on the number of remaining clusters to be allocated and the number of speakers having dropout issues.

[0089] EEE3. The method according to any one of EEE1 to EEE2, wherein the speaker dropout monitoring module examines if a speaker has a dropout issue based on a prior speaker loudness and a posterior speaker loudness.

[0090] EEE4. The method according to EEE3, wherein the prior speaker loudness is estimated based on object-to-speaker gains.

[0091] EEE5. The method according to any one of EEE3 to EEE4, wherein the posterior speaker loudness is estimated based on object-to-cluster gains and cluster-to-speaker gains.

[0092] EEE6. The method according to any one of EEE3 to EEE4, wherein the posterior speaker loudness is estimated based on a spatial distance metric between object pairs.

[0093] EEE7. The method according to EEE6, wherein the posterior speaker loudness is estimated based on a spatial distance metric between an object and a cluster, wherein the cluster corresponds to one or more objects combined in a set of objects.

[0094] EEE8. The method according to any one of EEE1 to EEE7, wherein a preset speaker layout can be used for the cluster selection module, speaker optimization determination module, and speaker dropout monitoring module.

[0095] EEE9. The method according to any one of EEE1 to EEE8, wherein the cluster selection module provides speaker optimization by adding weights to the object importance.

[0096] EEE10. The method according to EEE9, wherein the weights are based on object-to- speaker gains. 21

[0097] EEE11. A method for mitigating speaker dropouts in an object-based audio system, the method comprising: receiving a plurality of audio objects, wherein each of the audio objects is associated with a corresponding object metadata that indicates respective spatial position information for the corresponding audio object; evaluating the plurality of audio objects to determine if a speaker optimization process should be enabled; and optimizing the clustering of audio objects for speaker dropout mitigation when the speaker optimization process is enabled.

[0098] EEE12. The method according to EEE11, further comprising: iteratively determining cluster positions for the plurality of audio objects up to a maximum cluster count.

[0099] EEE13. The method according to any one of EEE11 to EEE12, wherein evaluating the plurality of audio objects further comprises: identifying a speaker dropout associated with one or more of the plurality of audio objects, and adjusting the clustering of the audio objects when the speaker dropout is identified.

[0100] EEE14. The method according to any one of EEE11 to EEE13, wherein the clustering of audio objects further comprises using a centroid determination process that selects the most perceptually important audio objects based on speaker loudness.

[0101] EEE15. The method according to any one of EEE11 to EEE14, wherein the clustering of audio objects comprises making clustering decisions based on the number of remaining clusters to be allocated and the number of speakers having dropout issues.

[0102] EEE16. The method according to EEE15, wherein clustering decisions for a speaker experiencing a dropout is determined based on a prior speaker loudness and a posterior speaker loudness.

[0103] EEE17. The method according to EEE16, wherein the prior speaker loudness is estimated based on object-to-speaker gains.

[0104] EEE18. The method according to any one of EEE16 to EEE17, wherein the posterior speaker loudness is estimated based on object-to-cluster gains and cluster-to-speaker gains.

[0105] EEE19. The method according to any one of EEE16 to EEE17, wherein the posterior speaker loudness is estimated based on a spatial distance metric between object pairs.

[0106] EEE20. The method according to EEE19, wherein the posterior speaker loudness is estimated based on a spatial distance metric between an object and a cluster, wherein the cluster corresponds to one or more objects combined in a set of objects. 22

[0107] EEE21. The method according to any one of EEE11 to EEE20, wherein a preset speaker layout is employed in the cluster selection and speaker optimization process.

[0108] EEE22. The method according to EEE21, wherein the preset speaker layout includes one or more of a numbers of speakers and a location of one or more speakers.

[0109] EEE23. The method according to any one of EEE11 to EEE22, wherein a speaker optimization process is employed that identifies each object of importance and adds a corresponding weighting factor to each of the identified objects of importance.

[0110] EEE24. The method according to EEE23, wherein the weighting factors may be based on object-to-speaker gains.

[0111] EEE25. The method according to any one of EEE11 to EEE24, further comprising clustering audio objects based on a distance between pairs of audio objects when determining the cluster positions.

[0112] EEE26. The method according to EEE25, further comprising applying a cost function in determining the cluster positions, wherein an overall cost of the cost function is determined for cluster positions by calculating object-to-cluster gains.

[0113] EEE27. The method according to EEE26, wherein the overall cost includes a plurality of penalty terms, and wherein at least one of the penalty terms uses the extended hybrid distance metric.

[0114] EEE28. The method according to EEE27, wherein the overall cost is defined as a linear combination of a sub-cost of each of the penalty terms, and wherein the overall cost combines at least one positional distance metric describing differences in object position; a metric representing similarity of dissimilarity in headphone rendering modes (HRM); and a loudness, level, or importance metric of the audio objects.

[0115] EEE29. The method according to any one of EEE27 to EEE28, wherein the audio objects are rendered to the cluster positions by minimizing the overall cost.

[0116] EEE30. The method according to any one of EEE27 to EEE29, further comprising applying an iterative greedy approach to selecting the audio objects with a maximum partial loudness, overall loudness, energy, level, salience, or importance.

[0117] EEE31. A method for clustering audio objects, the method comprising: receiving an input audio block, the audio block including a plurality of audio objects; calculating object-to-speaker 23  gains for the plurality of audio objects; identifying whether a speaker is experiencing a dropout issue; and performing clustering of the audio objects based on whether the speaker is experiencing the dropout issue.

[0118] EEE32. The method according to EEE31, wherein the dropout issue of the speaker is identified based on a prior speaker loudness and a posterior speaker loudness.

[0119] EEE33. The method according to EEE32, wherein the prior speaker loudness is estimated based on object-to-speaker gains, and wherein the posterior speaker loudness is estimated based on object-to-cluster gains and cluster-to-speaker gains.

[0120] EEE34. The method according to any one of EEE32 to EEE33, wherein the prior speakerloudness is calculated according to: Γ^ ൌ ∑ ^ ^^^,^ ^^^^^ ^ , where s denotes the speaker index, j denotes the object index, ^^ denotes gains and ^^^^^^ denotes an initialtotal loudness.

[0121] EEE35. The method according to any one of EEE32 to EEE34, wherein the posterior speaker loudness is calculated according to: ^^^^^ ^^ି^^^ ൌ ^^^ ^ ∑ ^^^ ∗^ ^^^,^ ^^^^^1 െ ^^ா൫^^, ^^^^൯^ , where s denotes the speaker index, j denotes the object index,^ି^^denotes posterior speaker loudness from a previous iteration, ^^^,^denotes object-to-speaker gains and ^^^^^ denotes object loudnes ∗^^^^s for an audio object j, and ^^ா൫^^, ^^൯ denotes the proximity of the object ^^ and ^^∗^^^.

[0122] EEE36. The method according to any one of EEE32 to EEE35, wherein identifying whether the speaker is experiencing the dropout issue includes: dividing the posterior speaker loudness by the prior speaker loudness to obtain a speaker dropout metric; and comparing the speaker dropout metric to a threshold.

[0123] EEE37. The method according to any one of EEE31 to EEE36, wherein performing clustering of the audio objects includes determining whether to enable a speaker optimization process, wherein, when the speaker optimization process is enabled, clustering of the audio objects is based on the number of remaining clusters to be allocated and the number of speakers having dropout issues.

[0124] EEE38. The method according to EEE37, wherein, when the speaker optimization process is not enabled, clustering of the audio objects is based on an excitation of each of the plurality audio objects in a critical band. 24

[0125] EEE39. The method according any one of EEE37 to EEE38, wherein a loudness of each audio object is calculated according to: ^^ᇱ^^^ఈ ^^^^^ ൌ ^^^ ^ ∑^^^^^^^^^^^ ^െ ^^^ ^ ∑ ^^^^ ^^^ ^^^^ ൫1 െఈ^^ு^^^, ^^^൯ ^, where ^^^^^^ ^^^^ denotes the^^^^^^^denotes an original excitation of the audio object, ^^ and ^^ denote parameters, and ^^ு^^^, ^^^denotes an amount of masking.

[0126] EEE40. The method according to any one of EEE37 to EEE39, wherein a partial loudness of each audio object j is calculated according to: ^^^^^^ ൌ ∑ᇱ^^^^^^^^^^^, and whereinclusters are determined by selection an audio object loudness.

[0127] EEE41. The method according to any one of EEE37 to EEE40, wherein, when the speaker optimization process is enabled, a dropout speaker-concentrated metric is applied for cluster selection.

[0128] EEE42. The method according to EEE41, wherein the dropout speaker-concentrated metric is calculated according to: ^^^^^^ ൌ ^^^∗,^^^^ ^^^^^^^ ^, where ^^^∗,^denotes a rendering gain of audio object j to speaker s*,loudness of audio object j in an i, and where ^^ is the Sone-to-

[0129] EEE43. An apparatus comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the apparatus to carry out the method according to any one of EEE1 to EEE44.

[0130] EEE44. A non-transitory computer-readable storage medium storing a program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one EEE1 to EEE42.

[0131] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims. 25

[0132] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0133] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0134] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter. 26

Claims

CLAIMS What is claimed is:

1. A method (400) for clustering audio objects, the method comprising: a cluster selection module (408) with an optional speaker optimization process; a speaker optimization determination module (406) to determine if the speaker optimization process should be enabled in the cluster selection module (408); and a speaker dropout monitoring module (404) to determine if a speaker has a dropout issue.

2. The method (400) of claim 1, wherein the speaker optimization determination module (406) makes decisions based on the number of remaining clusters to be allocated and the number of speakers having dropout issues.

3. The method (400) of claim 1 or 2, wherein the speaker dropout monitoring module (404) examines if a speaker has a dropout issue based on a prior speaker loudness and a posterior speaker loudness.

4. The method (400) of claim 3, wherein the prior speaker loudness is estimated based on object-to-speaker gains.

5. The method (400) of claim 3 or 4, wherein the posterior speaker loudness is estimated based on object-to-cluster gains and cluster-to-speaker gains.

6. The method (400) of claim 3 or 4, wherein the posterior speaker loudness is estimated based on a spatial distance metric between object pairs.

7. The method (400) of claim 6, wherein the posterior speaker loudness is estimated based on a spatial distance metric between an object and a cluster, wherein the cluster corresponds to one or more objects combined in a set of objects.

8. The method (400) of any of claims 1 to 7, wherein a preset speaker layout can be used for the cluster selection module (408), speaker optimization determination module (406), and speaker dropout monitoring module (404). 27  9. The method (400) of any of claims 1 to 8, wherein the cluster selection module (408) provides speaker optimization by adding weights to the object importance.

10. The method (400) of claim 9, wherein the weights are based on object-to-speaker gains.

11. A method for mitigating speaker dropouts in an object-based audio system, the method comprising: receiving a plurality of audio objects, wherein each of the audio objects is associated with a corresponding object metadata that indicates respective spatial position information for the corresponding audio object; evaluating the plurality of audio objects to determine if a speaker optimization process should be enabled; and optimizing the clustering of audio objects for speaker dropout mitigation when the speaker optimization process is enabled.

12. The method of claim 11, further comprising: iteratively determining cluster positions for the plurality of audio objects up to a maximum cluster count.

13. The method of any of claims 11 to 12, wherein evaluating the plurality of audio objects further comprises: identifying a speaker dropout associated with one or more of the plurality of audio objects, and adjusting the clustering of the audio objects when the speaker dropout is identified.

14. The method of any of claims 11 to 13, wherein the clustering of audio objects further comprises using a centroid determination process that selects the most perceptually important audio objects based on speaker loudness.

15. The method of any of claims 11 to 14, wherein the clustering of audio objects comprises making clustering decisions based on the number of remaining clusters to be allocated and the number of speakers having dropout issues.

16. The method of claim 15, wherein clustering decisions for a speaker experiencing a dropout is determined based on a prior speaker loudness and a posterior speaker loudness. 28  17. The method of claim 16, wherein the prior speaker loudness is estimated based on object- to-speaker gains.

18. The method of claim 16 or 17, wherein the posterior speaker loudness is estimated based on object-to-cluster gains and cluster-to-speaker gains.

19. The method of claims 16 or 17, wherein the posterior speaker loudness is estimated based on a spatial distance metric between object pairs.

20. The method of claim 19, wherein the posterior speaker loudness is estimated based on a spatial distance metric between an object and a cluster, wherein the cluster corresponds to one or more objects combined in a set of objects.

21. The method of any of claims 11 to 20, wherein a preset speaker layout is employed in the cluster selection and speaker optimization process.

22. The method of claim 21, wherein the preset speaker layout includes one or more of a number of speakers and a location of one or more speakers.

23. The method of any of claims 11 to 22, wherein a speaker optimization process is employed that identifies each object of importance and adds a corresponding weighting factor to each of the identified objects of importance.

24. The method of claim 23, wherein the weighting factors may be based on object-to- speaker gains.

25. The method of any of claims 11 to 24, further comprising clustering audio objects based on a distance between pairs of audio objects when determining the cluster positions.

26. The method of claim 25, further comprising applying a cost function in determining the cluster positions, wherein an overall cost of the cost function is determined for cluster positions by calculating object-to-cluster gains. 29  27. The method of claim 26, wherein the overall cost includes a plurality of penalty terms, and wherein at least one of the penalty terms uses an extended hybrid distance metric.

28. The method of claim 27, wherein the overall cost is defined as a linear combination of a sub-cost of each of the penalty terms, and wherein the overall cost combines at least one positional distance metric describing differences in object position; a metric representing similarity of dissimilarity in headphone rendering modes (HRM); and a loudness, level, or importance metric of the audio objects.

29. The method of any of claims 27 to 28, wherein the audio objects are rendered to the cluster positions by minimizing the overall cost.

30. The method of any of claims 27 to 29, further comprising applying an iterative greedy approach to selecting the audio objects with a maximum partial loudness, overall loudness, energy, level, salience, or importance.

31. A method for clustering audio objects, the method comprising: receiving an input audio block, the audio block including a plurality of audio objects; calculating object-to-speaker gains for the plurality of audio objects; identifying whether a speaker is experiencing a dropout issue; and performing clustering of the audio objects based on whether the speaker is experiencing the dropout issue.

32. The method of claim 31, wherein the dropout issue of the speaker is identified based on a prior speaker loudness and a posterior speaker loudness.

33. The method of claim 32, wherein the prior speaker loudness is estimated based on object- to-speaker gains, and wherein the posterior speaker loudness is estimated based on object-to- cluster gains and cluster-to-speaker gains.

34. The method of any of claims 32 to 33, wherein the prior speaker loudness is calculated according to: Γ ^^^ ^ ^^^ where s denotes the speaker index, j denotes the object index, ^^^,^denotes object-to- speaker gains and ^^^^^^ denotes an initial object total loudness.

35. Theany of claims 32 to 34, wherein the posterior speaker loudness is calculated according to: ^^^^^^ ൌ ^^^^ି^^^ ^^^^^,^^^^^^^ ^1 െ ^^ ∗ா൫^^, ^^^^^൯^ ^ where s denotes i denotes the iterationndex, ^^^i^ି^^^ denotes ^^^,^denotes object- to- and ^^^^^^ denotes object loudness for an audio object j, and ^^ா൫^^, ^^∗^^^൯ denotes the proximity of the ^^ and ^^∗^^^.

36. The method of any of claims 32 to 35, wherein identifying whether the speaker is experiencing the dropout issue includes: dividing the posterior speaker loudness by the prior speaker loudness to obtain a speaker dropout metric; and comparing the speaker dropout metric to a threshold.

37. The method of any of claims 31 to 36, wherein performing clustering of the audio objects includes determining whether to enable a speaker optimization process, wherein, when the speaker optimization process is enabled, clustering of the audio objects is based on the number of remaining clusters to be allocated and the number of speakers having dropout issues.

38. The method of claim 37, wherein, when the speaker optimization process is not enabled, clustering of the audio objects is based on an excitation of each of the plurality audio objects in a critical band.

39. The method of any of claims 37 to 38, wherein a loudness of each audio object is calculated according to: ఈ ఈ ^^ᇱ^^^^^^^ ^ ^^^^^^^^^^ ^^ ^ ^^^^^^^^^^ where ^^^^^^ ^^^^ denotes the excitation of audio object j in critical band b, ^^^^^^ ^^^^ denotesan original excitation of the audio object, ^^ and ^^ denote parameters, and ^^ு^^^, ^^^ denotes anamount of masking.

40. The method of any of claims 37 to 39, wherein a partial loudness of each audio object j is calculated according to: ^^^^^^ ൌ ∑^ ^^ᇱ^^^^ ^^^^ , and wherein clusters are an audio object having the greatest partial loudness.

41. The method of any of claims 37 to 40, wherein, when the speaker optimization process is enabled, a dropout speaker-concentrated metric is applied for cluster selection.

42. The method of claim 41, wherein the dropout speaker-concentrated metric is calculated according to: ^^^^^^^^^ൌ ^^^∗,^^^^ ^^^^ ^where ^^ ∗ denotes a j to speaker s*, whe^^^^,^re ^^^denotes a partial loudness of audio object j in an iteration i, and where ^^ is the Sone-to-function.

43. An apparatus comprising a processor and a memory coupled to the processor, wherein the processor is adapted to cause the apparatus to carry out the method according to any one of the preceding claims.

44. A non-transitory computer-readable storage medium storing a program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of preceding claims. 32

Citation Information

Patent Citations

  • Object Clustering for Rendering Object-Based Audio Content Based on Perceptual Criteria

    US20150332680A1

  • Audio object clustering based on renderer-aware perceptual difference

    US20190182612A1

  • Panning of audio objects to arbitrary speaker layouts

    WO2015017037A1

  • Systems and methods for headphone rendering mode-preserving spatial coding

    WO2023039096A1