Low-latency gain interpolation for audio object clustering
The low-latency audio object clustering system addresses resource limitations in interactive entertainment systems by dynamically applying gain interpolation, reducing complexity and latency in audio object transmission while maintaining quality.
Patent Information
- Application Number
- PCT/US2025/020744
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-03-20
- Publication Date
- 2025-09-25
AI Technical Summary
Interactive entertainment systems face challenges in transmitting object-based audio signals due to limited computational resources and memory, necessitating efficient audio scene management and low-latency audio object clustering to reduce complexity while maintaining audio quality.
A low-latency audio object clustering system that dynamically applies gain interpolation based on energy changes in audio objects, using an audio object clustering unit to calculate object-to-cluster gains, a gain buffering unit to store these gains, and a gain interpolation unit to adjust gains based on onset/offset detection, allowing quick adaptation to loudness and positional changes.
The system effectively reduces audio complexity and latency, ensuring high-quality audio rendering by quickly adapting to energy and positional changes in audio objects, thus enhancing user experience in interactive entertainment systems.
Smart Images

Figure US2025020744_25092025_PF_FP_ABST
Abstract
Description
LOW-LATENCY GAIN INTERPOLATION FOR AUDIO OBJECT CLUSTERINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from International Patent Application No. PCT / CN2024 / 083367, filed on 22 March 2024, International Patent Application No.PCT / CN2024 / 100208, filed on 19 June 2024, US Provisional Application Ser. No 63 / 669,883 filed on 11 July 2024, and European Application No. 24191986.9 filed on 31 July 2024, each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] This application relates generally to object-based audio signal processing and, more specifically, dynamically applying gain interpolation to object-based audio signals.BACKGROUND
[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.
[0004] In interactive entertainment systems, object-based audio signals are transmitted between system endpoints, where the object-based audio signals may contain hundreds of individual audio objects. Consequently, transmitting these audio signals may be challenging. For example, some system endpoints may have limited resources (e g., limited processing power, limited memory, etc.) that result in strict requirements as to the maximum number of audio objects supported at the endpoint. The present disclosure appreciates these and other system limitations, and therefore, an efficient object-based audio scene management approach is desired.
[0005] To reduce the complexity of a given audio scene, clustering techniques may be employed to reduce the number of input audio objects and beds into a small set of output objects (hereafter called clusters) , and with minimum impact on audio quality. In general, an object clustering method may include two steps: 1) cluster determination, and 2) clustering generation. Cluster determination is used to determine the cluster position and associated metadata, while cluster generation is used to calculate the object to cluster gains and generate the clusters.
[0006] In an example clustering method, a first step may include determining the clusters by selecting the most perceptually important objects. Both the object loudness and the spatial position of the object may be considered to identify the measure of perceptual importance. In a second step, the clusters may be generated by calculating the object- to-cluster gains and applying the calculated gains to the input audio objects. In a further example, the gain calculations may include a method that minimizes a cost function, where position correctness, distance and amplitude preservation may be jointly considered.
[0007] In gaming and the other interactive use-cases, latency is a primary concern for a good user experience, which may lead to strict requirements for all running algorithms. The presently disclosure thus appreciates that audio object clustering methods are expected to be optimized in terms of 1) limited access to advanced information including the audio signal and related metadata, and 2) quickly adapted to temporal changes in terms of audio loudness / energy and positional and non-positional metadata.
[0008] It is with respect to these and other considerations that the disclosure made herein is presented.BRIEF SUMMARY OF THE DISCLOSURE
[0009] Techniques are described for processing audio signals. Examples found herein provide for systems, devices, and methods to process audio objects and, more specifically, to dynamically apply gain interpolation to object-based audio signals. For example, embodiments provide a dynamic gain interpolation based on energy changes in audio objects between audio blocks.
[0010] One example described herein provides a low-latency audio object clustering system. The system includes an audio object clustering unit (or module) configured to calculate object-to-cluster gains of an object, a gain buffering unit configured to store the object-to-cluster gains provided by the audio object clustering unit, an onset / offset detection unit configured to set a flag based on at least one of a loudness and a position of the object, and a gain interpolation unit configured to calculate interpolated gain for an object based on the flag and the object-to-cluster gains.
[0011] In some aspects, the audio object clustering unit selects audio objects, and in particular the most important audio objects, as cluster centroids. The object-to-cluster gains for each object in a selected cluster may be calculated by the audio object clustering unit. Object-to-cluster gains arereferred to by the gain interpolation unit when the gain interpolation unit is calculating interpolated gain for objects. Calculated object-to-cluster gains may also be stored in the gain buffering unit such that the gain interpolation unit is provided both current object-to-cluster gains and previous object-to-cluster gains.
[0012] In some aspects, the onset / offset detection unit determines whether a gain interpolation method implemented by the gain interpolation module for calculating gain interpolation should be changed or adjusted. For example, the onset / offset detection unit compares the energy of received audio frames or blocks of audio to a threshold. When the energy is above the threshold, the onset / offset detection unit sets a flag to HIGH. When the energy is below the threshold, the onset / offset detection unit sets the flag to LOW. The onset / offset detection unit may also set the flag based on changes in position of audio objects.
[0013] In some aspects, the gain interpolation unit calculates the interpolated gain for audio objects based on the flag set from the onset / offset detection unit. For example, an interpolation ramp length may be selected by the gain interpolation unit based on the value of the flag. The interpolation ramp may be shorter or longer based on how quickly the energy of the frame increases, and is set based on a difference in the initial gain and the target gain for the current frame.
[0014] Another example described herein provides a method for clustering audio objects. The method includes receiving an input audio block included in a frame, the audio block including a plurality of audio objects, calculating current object-to-cluster gains for the plurality of audio objects, and setting a flag based on an energy of the frame. The method includes calculating, for each audio object of the plurality of audio objects, an interpolated gain based on the flag, the current object-to-cluster gains, and previous object-to-cluster gains stored in a buffer, and processing the input audio block using the interpolated gain.
[0015] In some aspects, previous object-to-cluster gains are associated with a previous audio block in the current frame or in the previous frame. Current object-to-cluster gains may be stored in the buffer to be used as previous object-to-cluster gains for a future iteration. Object to cluster gains for the plurality of audio objects may be represented as N-dimensional vectors.
[0016] Various aspects of the present disclosure provide for processing of stereo audio signals, and effect improvements in at least the technical fields of audio processing, audio encoding, audio decoding, virtual reality, object-based audio, and the like.
[0017] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.
[0018] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associate drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims.DESCRIPTION OF THE DRAWINGS
[0019] These and other more detailed and specific features of various embodiments are more fully disclosed in the following description, reference being had to the accompanying drawings, in which:
[0020] FIG. 1 illustrates a block diagram of an example audio coding system in which various aspects of the present disclosure can be practiced.
[0021] FIG. 2 illustrates a block diagram of an example spatial coding method that can be implemented in the audio system of FIG. 1 according to some aspects of the present disclosure.
[0022] FIG. 3 illustrates a block diagram of a framework for audio object clustering according to some aspects of the present disclosure.
[0023] FIGS. 4A-4B illustrate example gain interpolation functions according to some aspects of the present disclosure.
[0024] FIGS. 5A-5B illustrate example cross-fade functions according to some aspects of the present disclosure.
[0025] FIG. 6 illustrates a block diagram of another framework for audio object clustering according to some aspects of the present disclosure.
[0026] FIG. 7 illustrates a block diagram of various example methods for clustering audio objects, which may be performed by the audio coding system of FIG. 1, in accordance with various aspects of the present disclosure.
[0027] FIG. 8A illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.
[0028] FIG. 8B illustrates a schematic block diagram of an example CPU implemented in the device architecture of FIG. 8A that may be used to implement various aspects of the present disclosure.DETAILED DESCRIPTION
[0029] In the following description, numerous details are set forth, such as audio device configurations, timings, operations, and the like, in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to one skilled in the art that these specific details are merely examples and not intended to limit the scope of this application.
[0030] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0031] For interactive entertainment content, transmitting the original object-based audio signal, which may contain hundreds of individual objects, can be challenging because endpoints thatsupport object-based audio typically have limitations with respect to the maximum number of audio objects that can be supported, e.g., due to the limited amounts of computational resources and / or memory. In such cases, application of efficient object-based audio-scene management is desirable.
[0032] In some cases, “spatial coding” may be used to reduce the complexity of an audio scene. Some spatial coding methods employ clustering techniques that aim to reduce the number of input objects and / or beds to a smaller set of output objects (hereafter referred to as clusters) with minimal impact on the audio quality.
[0033] Herein, the terms “clustering,” “grouping,” and “combining” may be used interchangeably to describe combinations of objects and / or beds (or channels) configured to reduce the amount of data in a unit of adaptive audio content for transmission and rendering in an audio playback system. The terms “compression” and “reduction” may be used to refer to an act of performing audio scene simplification, e.g., via such clustering. The terms “clustering,” “grouping,” and “combining” throughout this description are not limited to a strictly unique assignment of an audio object or bed to a single cluster only. In some cases, an audio object or bed may be distributed over more than one output bed or cluster using weights or gain vectors that determine the relative contribution of an object or bed signal to the output cluster or output bed signal.
[0034] FIG. 1 illustrates a block diagram of an example audio coding system 100 in which various aspects of the present disclosure can be practiced. The audio coding system 100 includes a spatial coding component 110. The audio coding system 100 also includes an audio encoder 120 and an audio decoder 140 coupled via a bandwidth-limited communication channel 130. The spatial coding component 110 receives input signals 102 and processes the input signals 102 to generate input signals 112. For example, the spatial coding component 110 applies spatial coding techniques to the input signals 102 to simplify the audio scene represented by the input signals 102, e.g., via clustering.
[0035] The spatial coding component 110 provides the input signals 112 to the encoder 120. The input signals 104 bypass the spatial coding component 110 and are coupled directly to the encoder 120. The encoder 120 receives input signals 104, 112 and processes the received input signals to generate an encoded bitstream 132, which is transmitted over the communication channel 130 to the decoder 140. The decoder 140 receives the encoded bitstream 132 via the communication channel 130 and decodes the received bitstream to generate output audio signals 142. The output audiosignals 142 are applied to an audio rendering component 150 that operates to render and / or playback the audio content represented by the output audio signals 142.
[0036] In various examples, the audio rendering component 1 0 may include any professional or consumer-grade audio system, such as a home theater (e.g., including an A / V receiver, a soundbar, a Blu-ray player, etc.), one or more E-media devices (e.g., a computer, a tablet, a mobile phone equipped with headphones and / or speakers, etc.), a TV set, and a sound reproduction system. In some examples, the audio rendering component 150 provides an audio environment for playback of audio or audio / visual content using a plurality of speakers and suitable playback devices. In some examples, the audio rendering component 150 may represent any environment in which a listener is experiencing playback of the audio content, such as a cinema, a concert hall, an outdoor theater, a home theater or room, a listening booth, a car, a game console, a headset device, a public address (PA) system, or other audio playback environment. In other examples, the audio rendering component is an audio process that renders the output audio signals 142 and provides the rendered audio to a separate playback device, such as one or more speakers, a set of headphones, or another processing device.
[0037] The audio coding system 100 may implement object-based audio formats that include both beds and objects. Audio beds refer to audio channels that are meant to be reproduced in predefined, fixed locations. Audio objects refer to individual audio elements that may exist for a defined duration in time but also have spatial information of each object, such as position, size, and the like. During transmission, beds and objects may be sent separately and then used by a spatial reproduction system to recreate the artistic intent. The reproduction system may include a variable number of speakers or headphones. Object-based audio as referred to herein may include interactive content that is typically generated at runtime as the user interacts with the corresponding virtual scene, and may also include music content.
[0038] FIG. 2 illustrates a block diagram of an example spatial coding method 200 that can be implemented in the audio coding system 100 according to some aspects of the present disclosure. The method 200 may include a seed selection module 210 and a cluster generation module 220. The input of the seed selection module 210 corresponds to a first path 202. The seed selection module 210 is configured to receive audio objects via the first path 202. The output of the seed selection module 210 corresponds to a second path 212. The seed selection module 210 may be configured toperform operations for cluster-seed selection, where input audio objects received via the first path 202 may be evaluated to identify a set of seed objects for output clusters. Once the set of seed objects is identified by the seed selection module 210, the cluster positions may be determined (e.g., calculated or estimated), and the associated metadata may be generated. The set of seed objects, the cluster positions, and the associated metadata are transmitted by the seed selection module 210 to the cluster generation module 220 via the second path 212.
[0039] The input of the cluster generation module 220 corresponds to the second path 212. The output of the cluster generation module 220 corresponds to a third path 222. The cluster generation module 220 is configured to receive the set of seed objects, the cluster positions, and the associated metadata from the seed selection module 210 via the second path 212. The cluster generation module 220 may be configured to perform cluster generation, where object to cluster gains are calculated and the corresponding clusters are generated from the set. The output clusters are provided by the cluster generation module 220 via the third path 222.
[0040] In some examples, operations of the seed selection module 210 may include evaluating the input audio objects and selecting perceptually most-important objects as cluster seeds. Both the object loudness and spatial position may be considered as factors in the process of determining the object’s importance. Operations of the cluster generation module 220 may include calculating the object-to-cluster gains and applying the gains to the input audio objects to generate the corresponding output clusters. In some examples, the object-to-cluster gains may be calculated in the cluster generation module 220 based on minimizing a cost function, which may consider one or more factors such as the position correctness, distance preservation, amplitude preservation, or any combinations thereof.
[0041] Examples, aspects, and instances described herein provide a low latency audio objects clustering framework. With this described framework, audio object clustering can be implemented by using the current (and / or the previous) audio blocks and the associated metadata without requiring access to any future audio blocks and metadata. The term audio block may refer to fixed samples of an audio signal (for example, 256, 512, 1024, 1536, or 2048 samples per packet).
[0042] One of the challenges identified by the present disclosure is to quickly adapt the object loudness and / or position change of an audio object. For example, if there is any onset / offset occurring in the current block, then reaching the target gain earlier would be better than reaching thetarget at the block end. To achieve these and other goals, the onset / offset detection module 304 may be employed to determine (e.g., calculate or estimate) if there exists an abrupt object energy and / or metadata change for the current audio object, and then set the transition flag accordingly. The transition flag may be used to steer the way of gain interpolation.
[0043] FIG. 3 illustrates a block diagram of a framework 300 for audio object clustering according to some aspects of the present disclosure. The framework 300 includes a gain interpolation module 302, an onset / offset detection module 304, an audio object clustering module 306, a gain buffer module 308, and a summation module 310.
[0044] The proposed audio object clustering framework 300 may be operated in a frame-by-frame basis. In each frame, the audio blocks and the associated metadata may be generated at run-time. In an example operating frame, there may be M audio objects generated at runtime. A fixed number of clusters, denoted by N clusters, may be generated through the proposed approach.
[0045] The input of the audio object clustering module 306 corresponds to a first path 322. The output of the audio object clustering module 306 corresponds to a second path 324. The audio object clustering module 306 is coupled to the gain buffer module 308 and the gain interpolation module 302 via the second path 324. The audio object clustering module 306 may receive input audio blocks and associated metadata via the first path 322. The audio object clustering module 306 is configured to process the input audio blocks and the associated metadata to generate object-to- cluster gains. In some instances, both the current and historical audio blocks and associated metadata are used for analysis, where no advanced audio block and metadata may be available.
[0046] The audio object clustering performed by the audio object clustering module 306 may comprise two steps. As a first step, the audio object clustering module 306 may select the most important N objects as cluster centroids. Secondly, the audio object clustering module 306 may calculate the object-to-cluster gains for each object to the N clusters. Therefore, the object-to-cluster gains for each object can be represented by an N-dimensional vector. The audio object clustering module 306 is configured to provide the calculated object-to-cluster gains for each object to the gain interpolation module 302 and the gain buffer module 308 via the second path 324.
[0047] The input of the gain buffer module 308 corresponds to the second path 324. The gain buffer module 308 is coupled to the audio object clustering module 306 via the second path 324.The output of the gain buffer module 308 corresponds to a third path 326. The gain buffer module 308 is coupled to the gain interpolation module 302 via the third path 326. The gain buffer module 308 is configured to receive the calculated object-to-cluster gains for each object from the audio object clustering module 306 via the second path 324. The gain buffer module 308 may be configured to store the calculated object-to-cluster gains in the operating frame, where the saved object-to-cluster gains may be retrieved and implemented in a later frame (for example, in the next frame). The gain buffer module 308 is configured to provide the stored object-to-cluster gains to the gain interpolation module 302 via the third path 326.
[0048] The input of the onset / offset detection module 304 corresponds to the first path 322. The output of the onset / offset detection module 304 corresponds to a fourth path 328. The onset / offset detection module 304 is coupled to the gain interpolation module 302 via the fourth path 328. The onset / offset detection module 304 may receive input audio blocks and associated metadata via the first path 322. The onset / offset detection module 304 is configured to process the input audio blocks and associated metadata to detect and indicate if the abrupt change in terms of loudness and / or position occurred in the current frame for one or more of the objects in the current frame. The onset / offset detection module 304 may indicate the abrupt change by setting an onset / offset flag (e.g., a transition flag). The onset flag may be set to indicate an abrupt increase in loudness and the offset flag may be set to indicate an abrupt decrease in loudness. In some instances, a single flag is implemented to indicate an abrupt change in loudness. The change in the current frame may be a temporal change, such as changes in loudness and / or position between frames.
[0049] As one example that takes changes in loudness (e.g., volume) into account, the onset / offset flag may be set based on the following rules: 1) the onset flag is set to TRUE, if and only if the current frame energy is C times larger than the previous frame energy (otherwise, the onset flag is set to FALSE); or 2) the offset flag is set to TRUE, if and only if the current frame energy is no larger than 1 / C times previous frame energy (otherwise, the offset flag is set to FALSE). In the example, C is a preset constant (for example, C = 1000). Values of C may be obtained via other methods, including statistical calculation, experimental data collection, lookup tables, and the like. The onset / offset detection module 304 provides the onset / offset flag to the gain interpolation module 302 via the fourth path 328. C times the previous frame energy is one example of an onset threshold and 1 / C times the previous frame energy is one example of an offset threshold.
[0050] As another example that takes both loudness and position change into account, the onset / offset flag may be set based on the following rules: 1) the onset flag is set to TRUE, if and only if the current frame energy is C times larger than the previous frame energy and the distance from the current frame position to the last frame position is greater than D (otherwise, the onset flag is set to FALSE); or 2) the offset flag is set to TRUE, if and only if the current frame energy is no larger than 1 / C times previous frame energy and the distance from the current frame position to the last frame position is greater than D (otherwise, the offset flag is set to FALSE). In the example, C and D are both preset constants (for example, C = 1000 and D = 0.1).
[0051] The inputs of the gain interpolation module 302 correspond to the first path 322, the second path 324, the third path 326, and the fourth path 328. The gain interpolation module 302 may receive input audio blocks and associated metadata via the first path 322. The gain interpolation module 302 may receive calculated object-to-cluster gains from the audio object clustering module 306 via the second path 324. The gain interpolation module 302 may receive previous object-to- cluster gains (for example, gains associated with a previous frame) from the gain buffer module 308 via the third path 326. The gain interpolation module 302 may receive the onset / offset flag from the onset / offset detection module 304 via the fourth path 328. The output of the gain interpolation module 302 corresponds to a fifth path 330. The gain interpolation module 302 may be configured to process the input audio blocks based on the associated metadata, the calculated object-to-cluster gains, the previous object-to-cluster gains, and the onset / offset flag to generate processed audio blocks. In some instances, processing the input audio blocks includes determining (e.g., calculating or estimating) the interpolated gain for each audio object.
[0052] For each audio object to a specific cluster, the interpolated gain can be represented by a L- dim vector, where M is the number of samples per audio block. To determine (e g., calculate or estimate) the interpolated gain, the gain interpolation module 302 may be configured to obtain (for example, receive from the audio object clustering module 306 and the gain buffer module 308) the current and saved object-to-cluster gains as the target gain and initial gain, respectively, and perform interpolation between the initial gain and target gain values. The gain interpolation module 302 may select an interpolation method for determining the interpolated gain for each audio object based on the onset / offset flag, as will be described in further detail. The gain interpolation module 302 is configured to provide the processed audio blocks to the summation module 310 via the fifth path 330.
[0053] The input of the summation module 310 corresponds to the fifth path 330. The summation module 310 is coupled to the gain interpolation module 302 via the fifth path 330. The output of the summation module 310 corresponds to a sixth path 332. The summation module 310 is configured to receive processed audio blocks from the gain interpolation module 302 via the fifth path 330 and responsively generate output clusters based on a weighted summation, as will be described in further detail below. The summation module 310 provides the output clusters via the sixth path 332.
[0054] In some instances, processed audio block may be generated by the gain interpolation module 302 by taking an element-wise product to the interpolated gain and the current audio block. The cluster signal is then constructed by the summation module 310 by taking summation of all processed audio blocks for the respective cluster. The audio block of cluster], denoted by y7-, can be obtained by Equation 1 :where xtthe audio block of object z, is object i to cluster j's gain vector, and ® denotes element product of two vectors.
[0055] The framework 300 (and more specifically the gain interpolation module 302) is configured to quickly adapt to abrupt changes in the energy of received input audio objects. In some instances, the gain interpolation module 302 selects a gain interpolation function (e.g., a gain interpolation method) based on the received onset / offset flag.
[0056] FIGS. 4A-4B illustrate example gain interpolation functions according to some aspects of the present disclosure. FIG. 4A illustrates a first gain interpolation function 400 for a first audio block 402 included in a frame N. In the first gain interpolation function 400, an initial gain 404 (e.g., a previous obj ect-to-cluster gain) is shown at the beginning of the first audio block 402 and a target gain 406 (e.g., a current obj ect-to-cluster gain) is shown at the end of the first audio block 402. The interpolated gain 408 represents the change in gain over the duration of the first audio block 402 to increase from the initial gain 404 to the target gain 406.
[0057] FIG. 4B illustrates a second gain interpolation function 410 for a second audio block 412 included in the frame N. In the second gain interpolation function 410, an initial gain 414 is shown at the beginning of the second audio block 412 and a target gain 416 is shown (e.g., achieved)partway through the second audio block 412. The interpolation gain 418 represents the change in gain to increase from the initial gain 414 to the target gain 406.
[0058] In one example, when both the onset flag and the offset flag are identified as FALSE (e.g., LOW), the gain interpolation module 302 may select the first gain interpolation function 400, where the length to reach the target gain 406 (hereafter referred to as the ramp length) is determined (e.g., set by the gain interpolation module 302) to be equal to the block length of the first audio block 402. In another example, when either the onset flag or the offset flag is identified as TRUE, the gain interpolation module 302 may select the second gain interpolation function 410, where the ramp length is determined to be shorter than the block length of the second audio block 412. The ramp length in the second gain interpolation function 410 may correspond to a preset fixed value. The ramp length may be smaller than the attack time for compressing the input audio signal.
[0059] In another example, the gain interpolation module 302 selects the interpolation method based on the gain change from a previous gain value to a current gain value. For example, the gain interpolation module 302 compares the absolute value of the difference between a current object-to- cluster gain and a previous object-to-cluster gain to a predefined threshold. When the absolute value of the difference between a current object-to-cluster gain and a previous object-to-cluster gain is greater than or equal to a predefined threshold (e.g., a value in a range of about 0.5 to about 0.95, such as 0.8), the gain interpolation module 302 selects the second gain interpolation function 410. Otherwise, when the difference is less than the predefined threshold, the gain interpolation module 302 selects the first gain interpolation function 400.
[0060] In some instances, the interpolation method is selected using a combination of the aforementioned conditions. For example, when either of the above two conditions is true, i.e., either the onset flag or the offset flag is TRUE or the calculated gain difference is greater than or equal to the predefined threshold, then the gain interpolation module 302 selects the second gain interpolation function 410. Otherwise, the gain interpolation module 302 selects the first gain interpolation function 400.
[0061] In another implementation, the gain interpolation module 302 performs gain interpolation by taking a linear combination of initial gain and target gain with sum-to-one cross-fade functions as coefficients, as provided in Equation 2:9 interpolated fdown9init fup 9 target (2) where fdown and fupare cross-fade functions. FIGS. 5A-5B illustrate example cross-fade functions according to some aspects of the present disclosure. In FIGS. 5A-5B, the y-axis illustrates example gain coefficient values and the x-axis illustrates samples indexes.
[0062] In some instances, the ramp length may be dynamic or adaptive according to the onset time. In one example, when the onset is determined by the gain interpolation module 302 to occur in audio blocks at or near the end of the frame, the ramp length for the interpolated gain is longer as shown in the first gain interpolation function 400 of FIG. 4A. In another example, when the onset is determined by the gain interpolation module 302 to occur in audio blocks at or near the beginning of the frame, the ramp length is selected by the gain interpolation module 302 to be shorter as shown by the second gain interpolation function 410 of FIG. 4B.
[0063] In yet another implementation, a non-linear interpolation method may be implemented by the gain interpolation module 302 between the initial gain and the target gain. For example, the gain interpolation module 302 may calculate the square between the initial gain and the target gain, or apply a sine function to produce a non-linear gain, or apply another non-linear function. In some implementations, a combination of linear and non-linear interpolation methods may be applied.
[0064] While examples described herein primarily refer to changes in energy between adjacent audio blocks (for example, a current audio block and a previous audio block), examples described herein may also relate to changes in energy between frames (for example, a current frame and a previous frame). Additionally, detected changes in energy may be between audio blocks or frames that are not adjacent.
[0065] FIG. 6 illustrates a block diagram of another example framework 600 for audio object clustering according to some aspects of the present disclosure. The framework 600 includes the gain interpolation module 302, the audio object clustering module 306, the gain buffer module 308, and the summation module 310. Rather than including the onset / offset detection module 304 as in the framework 300, the framework 600 includes an audio change detection module 602. The input of the audio object clustering module 306 corresponds to a first path 612. The audio object clustering module 306 is configured to receive input audio blocks and associated metadata via the first path 612. The output of the audio object clustering module 306 corresponds to a second path614. The audio object clustering module 306 is configured to provide calculated object-to-cluster gains to the gain buffer module 308, the audio change detection module 602, and the gain interpolation module 302 via the second path 614. Operation of the audio object clustering module 306 is substantially similar to that as described with respect to the framework 300.
[0066] The input of the gain buffer module 308 corresponds to the second path 614. The gain buffer module 308 is configured to receive object-to-cluster gains from the audio object clustering module 306 via the second path 614. The output of the gain buffer module 308 corresponds to a third path 616. The gain buffer module 308 is configured to provide previous gains (e.g., stored object-to-cluster gains) to the audio change detection module 602 and the gain interpolation module 302 via the third path 616. Operation of the gain buffer module 308 is substantially similar to that as described with respect to the framework 300.
[0067] The inputs of the audio change detection module 602 correspond to the second path 614 and the third path 616. The audio change detection module 602 is configured to receive object-to-cluster gains from the audio object clustering module 306 via the second path 614 and receive the previous gains from the gain buffer module 308 via the third path 616. The output of the audio change detection module 602 corresponds to a fourth path 618. The audio change detection module 602 is configured to detect changes in the received object-to-cluster gains, and based on the detected changes, set an onset flag and / or an offset flag. The audio change detection module 602 is configured to provide the onset / offset flag to the gain interpolation module 302 via the fourth path 618.
[0068] The inputs of the gain interpolation module 302 correspond to the first path 612, the second path 614, the third path 616, and the fourth path 618. The gain interpolation module 302 is configured to receive input audio blocks and associated metadata via the first path 612. The gain interpolation module 302 is configured to receive object-to-cluster gains from the audio object clustering module 306 via the second path 614. The gain interpolation module 302 is configured to receive previous gains from the gain buffer module 308 via the third path 616. The gain interpolation module 302 is configured to receive the onset / offset flag from the audio change detection module 602 via the fourth path 618. The output of the gain interpolation module 302 corresponds to a fifth path 620. The gain interpolation module 302 is configured to provide processed audio blocks to the summation module 310 via the fifth path 620. Operation of the gaininterpolation module 302 is substantially similar to that as described with respect to the framework 300.
[0069] The input of the summation module 310 corresponds to the fifth path 620. The summation module 310 is configured to receive processed audio blocks from the gain interpolation module 302 via the fifth path 620. The output of the summation module 310 corresponds to a sixth path 622. The summation module 310 is configured to provide output clusters via the sixth path 622. Operation of the summation module 310 is substantially similar to that as described with respect to the framework 300.
[0070] The illustrated blocks and modules in the spatial coding method 200, the framework 300, and the framework 600 are merely examples for clustering audio objects. In other examples, the spatial coding method 200, the framework 300 and / or the framework 600 may include additional blocks, may omit blocks, may combine the functions of blocks, or may divide portions of the blocks into additional blocks.
[0071] FIG. 7 illustrates a block diagram of various example methods 700 for clustering audio objects, which may be performed by the framework 300 of FIG. 3. While primarily described with respect to the framework 300, the methods 700 may also be performed by the framework 600 of FIG. 6. The methods 700 may be performed by a processor, which may be configured to perform methods 700 via machine-executable instructions. The methods 700 may be broken into various blocks or partitions, such as blocks 705, 710, 715, 720, and 725. The various process blocks illustrated in FIG. 7 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 705.
[0072] At block 705, “Receiving An Input Audio Block Included In A Frame”, an example method 700 may include receiving an input audio block included in a frame. The audio block includes a plurality of audio objects. For example and with reference to FIG. 3, an input audio block and associated metadata is received by the framework 300 via the first path 322. Processing may proceed from block 705 to block 710.
[0073] At block 710, “Calculating Current Object-To-Cluster Gains”, an example method 700 may include calculating current obj ect-to-cluster gains for the plurality of audio objects, as previously described with respect to the audio object clustering module 306. Processing may proceed from block 710 to block 715.
[0074] At block 715, “Setting A Flag Based On An Energy Of The Frame,” an example method 700 may include setting a flag based on an energy of the frame, as previously described with respect to the onset / offset detection module 304. Processing may proceed from block 715 to block 720.
[0075] At block 720, “Calculating, For Each Audio Object, An Interpolated Gain,” an example method 700 may include calculating, for each audio object of the plurality of audio objects, an interpolated gain. The interpolated gain may be based on the flag, the current obj ect-to-cluster gains, and previous obj ect-to-cluster gains stored in a buffer, as previous described with respect to the gain buffer module 308 and the gain interpolation module 302. Processing may proceed from block 720 to block 725.
[0076] At block 725, “Processing The Input Audio Block Using The Interpolated Gain,” an example method 700 may include processing the input audio block using the interpolated gain, as previously described with respect to the gain interpolation module 302 and the summation module 310.
[0077] FIG. 8A illustrates a schematic block diagram of an example device architecture 800 (e.g., an apparatus 800) that may be used to implement various aspects of the present disclosure.Architecture 800 includes but is not limited to servers and client devices, systems, and methods as described in reference to FIGS. 1-6. As shown, the architecture 800 includes central processing unit (CPU) 801 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 802 or a program loaded from, for example, storage unit 808 to random access memory (RAM) 803. The CPU 801 may be, for example, an electronic processor 801. In RAM 803, the data required when CPU 801 performs the various processes is also stored, as required. CPU 801, ROM 802, and RAM 803 are connected to one another via bus 804. Input / output interface 805 is also connected to bus 804.
[0078] The following components are connected to VO interface 805: input unit 806, that may include a keyboard, a mouse, or the like; output unit 807 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 808 including a hard disk, or anothersuitable storage device; and communication unit 809 including a network interface card such as a network card (e.g., wired or wireless).
[0079] In some implementations, input unit 806 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0080] In some implementations, output unit 807 include systems with various number of speakers. Output unit 807 (depending on the capabilities of the hose device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0081] In some embodiments, communication unit 809 is configured to communicate with other devices (e.g., via a network). Drive 810 is also connected to I / O interface 805, as required. Removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 810, so that a computer program read therefrom is installed into storage unit 808, as required. A person skilled in the art would understand that although apparatus 800 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.
[0082] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 809, and / or installed from the removable medium 811, as shown in FIG. 8 A.
[0083] In some implementations, which gain interpolation function (e.g. which out of the first or second gain interpolation function as shown in FIGS. 4A-4B) that is to be applied by the gain interpolation module is determined based on the onset / offset flag. A first value of the onset / offset flag may indicate that a gain interpolation function which rapidly transitions from the initial gain to the target gain is to be applied (e.g. the second gain interpolation function 410 of FIG. 4B) and a second value of the onset / offset flag may indicate that a gain interpolation function which slowlytransitions from the initial gain to the target gain is to be applied (e.g. the first gain interpolation function 400 of FIG. 4A).
[0084] The rate at which a gain interpolation function transitions from the initial gain to a target gain, i.e. an object- to-cluster gain of a preceding audio block to an object-to-cluster gain of a subsequent audio block, is governed by the interpolation ramp length of the gain interpolation function. Accordingly, the onset / offset flag may indicate a gain interpolation function, or at least an interpolation ramp length thereof, that is to be applied by the gain interpolation module. A shorter interpolation ramp length is associated with a more rapid transition and a longer interpolation ramp length is associated with a less rapid transition. The interpolation ramp length may be defined using a time and / or a number of samples.
[0085] The onset / offset flag may be set based on detection of a temporal change for an audio block or associated metadata for an audio object. A temporal change may refer to a temporal change in at least one of loudness of the object (e.g., volume, energy or amplitude), position of the object and the object-to-cluster gain of the object. In the framework of FIG. 3, the onset / offset detection module 304 sets the onset / offset flag based on a temporal change identifiable from the input audio blocks and associated metadata. For example, the onset / offset detection module 304 may set the onset / offset flag for an audio object based on the position (as indicated in the metadata) and / or the loudness changing from a preceding audio block to subsequent audio block (e.g. a current audio block). The audio blocks may be of the same frame or of different frames.
[0086] Additionally or alternatively, the onset / offset detection module 304 may set the onset / offset flag based on a temporal change detected for the object-to-cluster gain of an audio object, analogous to operation of the audio change detection module 602 of the example framework 600 of FIG. 6.
[0087] A temporal change for any one of loudness, position and object-to-cluster gains may be determined by calculating the absolute value of a difference between the loudness, the position or object-to-cluster gain of a current audio block and a preceding audio block. In response to the absolute value of the difference being above a predetermined threshold it is determined that a temporal change has occurred and the onset / offset flat is set to a value indicating a shorter interpolation ramp length. In response to the absolute value of the difference being below the predetermined threshold it is determined that no temporal change (of sufficient magnitude) has occurred and the onset / offset flat is set to a value indicating a longer interpolation ramp length. Thatis, the flag indicates a shorter interpolation ramp length when the difference in loudness, object-to- cluster gain or position of the object increases, and vice versa.
[0088] Comparing the absolute value of the difference of the loudness, position or object-to-cluster gains from different audio blocks may be equivalent to comparing a loudness, position or object-to- cluster gain of a current audio block to a threshold, wherein the threshold is based on the loudness, position or object-to-cluster gain of the preceding audio block.
[0089] The setting of the flag may in implementations be performed individually for each audio object and frame or audio block. Alternatively, a single flag may be set for all audio objects of the same frame wherein the single flag is based on a change in total loudness or energy of all audio objects from a preceding frame to a current frame.
[0090] The onset / offset flag may be set to one of two values. For example, the onset / offset flag may be a Boolean meaning that the onset / offset flag is either TRUE or FALSE. A Boolean flag may be used to indicate when an onset and / or offset condition has been met. Wherein each value of the onset / offset flag indicates a respective interpolation ramp length.
[0091] The onset / offset flag may in some implementations be set to one out of more than two values. For example the onset / offset flag may be set to one of three values, one value indicating onset, another value indicating offset and yet another value indicating no offset or onset. Wherein each value of the onset / offset flag indicates a respective interpolation ramp length.
[0092] The onset / offset flag may in some implementations indicate the interpolation ramp length directly. For example, as described above it is envisaged that the onset / offset detection module determines the interpolation ramp length based on when in the current frame or current audio block an onset / offset condition is met and uses the flag to indicate the determined interpolation ramp length to the gain interpolation module for application. Hereby, the onset / offset flag may be set to a value selected from a range of multiple values, wherein each value is associated with a respective interpolation ramp length.
[0093] FIG. 8B illustrates a schematic block diagram of an example CPU 801 implemented in the device architecture 800 of FIG. 8 A that may be used to implement various aspects of the present disclosure. The CPU 801 includes an electronic processor 820 and a memory 821. The electronic processor 820 is electrically and / or communicatively connected to the memory 821 for bidirectionalcommunication. The memory 821 stores encoding software 822 and / or decoding software 823. In some examples, memory 821 may be located internal to the electronic processor 820, such as for an internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, memory 821 may be located external to the electronic processor 820, such as in a ROM 802, a RAM 803, flash memory or a removable medium 811, or another non-transitory computer readable medium that is contemplated for device architecture 800. In some instances, the electronic processor 820 may implement the spatial coding software 822 stored in the memory 821 to perform, among other things, any of the methods 700 of FIG. 7.
[0094] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units and modules discussed above can be executed by control circuitry (e.g., CPU 801 in combination with other components of FIG. 8 A), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as nonlimiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0095] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0096] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a randomaccess memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD- ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0097] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.
[0098] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims. Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.
[0099] EEE1. A low-latency audio object clustering system, comprising: an audio object clustering unit configured to calculate object-to-cluster gains of an object; a gain buffering unit configured to store the object-to-cluster gains provided by the audio object clustering unit; an onset / offset detection unit configured to set a flag based on at least one of a loudness and a position of the object; and a gain interpolation unit configured to calculate interpolated gain for an object based on the flag and the object-to-cluster gains.
[0100] EEE2. The system according to EEE1, wherein the gain interpolation unit is configured to calculate interpolated gain for an object with a gain interpolation function based on the flag and the object-to-cluster gains, wherein the flag indicates an interpolation ramp length for the gain interpolation function.
[0101] EEE3. The system according to EEE1 or EEE2, wherein the audio object clustering unit is configured to select the most important audio objects as cluster centroids.
[0102] EEE4. The system according to any one of EEE1 to EEE3, wherein the audio object clustering unit is configured to calculate the object-to-cluster gains for each object for a selected cluster.
[0103] EEE5. The system according to any one of EEE1 to EEE4, wherein the gain interpolation unit is configured to calculate the interpolated gain using a current object-to-cluster gains calculated by the audio object clustering unit and using a saved object-to-cluster gains stored by the gain buffering unit.
[0104] EEE6. The system according to any one of EEE1 to EEE5, wherein the onset / offset detection unit determines whether temporal change occurs in a current audio block.
[0105] EEE7. The system according to EEE6, wherein whether the temporal change occurs is determined based on at least one of a received audio block and a metadata associated with the received audio block.
[0106] EEE8. The system according to EEE7, wherein the received audio block and / or the received metadata are associated with one of a prior received frame or a current frame.
[0107] EEE9. The system according to EEE7, wherein the gain interpolation unit is configured to select an interpolated gain function based on whether the temporal change occurs.
[0108] EEE10. The system according to EEE9, wherein the interpolated gain function corresponds to one or more of a linear function, a non-linear function, or combinations thereof.
[0109] EEE11. The system of any of EEE 1 to EEE10, wherein the onset / offset detection unit (304) is configured to set the flag to a first value when a temporal change occurs in the current audio block; and set the flag to a second value when no temporal change occurs in the current audio block.
[0110] EEE12. The system of any of EEE 1 to EEE10, wherein the onset / offset detection unit is configured to set the flag to a first value when an energy of a current audio block is above an onset threshold or when an energy of the audio block is below an offset threshold and set the flag to a second value when the energy of the audio block is between the onset and offset threshold.
[0111] EEE13. The system of EEE12, wherein the offset threshold and / or onset threshold is based on the energy of a previous audio block or wherein the gain interpolation unit is configured to calculate the interpolated gain for the object based on flags set for the energy of a previous audio block.
[0112] EEE14. The system according to any one of EEE1 to EEE10, wherein the onset / offset detection unit is configured to: set the flag to a first value when an energy of a current frame is above a threshold; and set the flag to a second value when the energy of the current frame is below the threshold.
[0113] EEE15. The system according to EEE14, wherein the threshold is based on the energy of a previous frame.
[0114] EEE16. The system according to EEE14, wherein the gain interpolation unit is configured to calculate the interpolated gain for the object based on flags set for the energy of a previous frame.
[0115] EEE17. The system of any of EEE11 to EEE16, wherein the gain interpolation unit (302) is configured to calculate the interpolated gain for the current audio block using an interpolation gain function having an interpolation ramp length, wherein the gain interpolation unit (302) is configured to calculate the interpolated gain using a first interpolated gain function having a first interpolation ramp length in response to the flag being set to the first value, and wherein the gain interpolation unit (302) is configured to calculate the interpolated gain using a second interpolated gain function having a second interpolation ramp length in response to the flag being set to the first value, and wherein the first interpolation ramp length is shorter than the second interpolation ramp length.
[0116] EEE18. The system according to EEE17, wherein the second interpolation ramp length corresponds to the length of the current audio block.
[0117] EEE19. The system according to any one of EEE1 to EEE18, wherein the onset detection unit is configured to select an interpolation ramp length based on a detected onset condition for a current frame.
[0118] EEE20. The system according to EEE19, wherein the interpolation ramp length is determined from a difference in an initial gain and a target gain for the current frame.
[0119] EEE21. The system according to EEE20, wherein the interpolation ramp length is determined to be shorter when the onset occurs towards a beginning of the current frame, and the interpolation ramp length is determined to be longer when the onset occurs towards an end of the current frame.
[0120] EEE22. The system of any of EEE1 to EEE21, wherein the object- to-cluster gains comprises a current obj ect-to-cluster gain associated with a current audio block of a frame and a previous obj ect-to-cluster gain associated with a previous audio block included in the frame or wherein the previous obj ect-to-cluster gain is associated with a previous audio block included in a previous frame.
[0121] EEE23. A method for clustering audio objects, the method comprising: receiving an input audio block included in a frame, the audio block including a plurality of audio objects; calculating current obj ect-to-cluster gains for the plurality of audio objects; setting a flag based on an energy of the frame; calculating, for each audio object of the plurality of audio objects, an interpolated gain based on the flag, the current obj ect-to-cluster gains, and previous obj ect-to-cluster gains stored in a buffer; and processing the input audio block using the interpolated gain.
[0122] EEE24. The method according to EEE23, wherein the previous obj ect-to-cluster gains are associated with a previous audio block included in the frame.
[0123] EEE25. The method according to EEE23, wherein the previous obj ect-to-cluster gains are associated with a previous audio block included in a previous frame.
[0124] EEE26. The method according to any one of EEE23 to EEE25, further comprising: storing the current obj ect-to-cluster gains in the buffer.
[0125] EEE27. The method according to any one of EEE23 to EEE26, wherein setting the flag includes: comparing an energy of the current frame to an energy of the previous frame; and setting the flag based on the comparison.
[0126] EEE28. The method according to any one of EEE23 to EEE27, wherein calculating the interpolated gain includes selecting a ramp function based on the flag, wherein the previous object- to-cluster gains is set as an initial gain value of the audio block, and wherein the current object-to- cluster gains is set as a target gain value of the audio block.
[0127] EEE29. The method according to any one of EEE23 to EEE28, further comprising: selecting a set of audio objects from the plurality of audio objects as cluster centroids, wherein calculating the current object-to-cluster gains for the plurality of audio objects is performed for each audio object of the set of audio objects.
[0128] EEE30. The method according to EEE29, wherein the object-to-cluster gains for the plurality of audio objects are represented as N-dimensional vectors.
[0129] EEE31. The method according to any one of EEE23 to EEE30, wherein calculating the interpolated gain includes: calculating a difference between the current object-to-cluster gains and the previous object-to-cluster gains; comparing the difference to a gain threshold; and selecting a ramp function based on the comparison.
[0130] EEE32. The method according to EEE31, wherein selecting the ramp function based on the comparison includes selecting a first ramp function when the comparison is greater than the gain threshold or when the flag is set to indicate that the energy of the frame is above an energy threshold, and selecting a second ramp function when the comparison is less than the gain threshold and the flag is set to indicate that the energy of the frame is less than the energy threshold.
[0131] EEE33. The method according to any one of EEE23 to EEE32, further comprising selecting an interpolation ramp length based on a detected onset condition for the current frame.
[0132] EEE34. The method according to EEE33, wherein the interpolation ramp length is determined from a difference between an initial gain and a target gain for the current frame.
[0133] EEE35. The method according to EEE34, wherein the interpolation ramp length is determined to be shorter when the onset occurs towards a beginning of the current frame, and theinterpolation ramp length is determined to be longer when the onset occurs towards an end of the current frame.
[0134] EEE36. A method for clustering audio objects, the method comprising: receiving an audio block including an audio object; calculating object-to-cluster gains for the audio object; setting a flag based on one or more of a loudness of the audio object, a position of the audio object or an object-to-cluster gain of the audio object; calculating, for the audio object, an interpolated gain based on the flag and the object-to-cluster gains, with a gain interpolation function, wherein the flag indicates an interpolation ramp length for the gain interpolation function; and processing the input audio block using the interpolated gain.
[0135] EEE37. The method of EEE36, wherein the object-to-cluster gains comprises a current object-to-cluster gain associated with a current audio block of a frame and a previous object-to- cluster gain associated with a previous audio block included in the frame or wherein the previous object-to-cluster gain is associated with a previous audio block included in a previous frame.
[0136] EEE38. An apparatus comprising: an electronic processor configured to perform operations including the method according to any one of EEE23 to EEE37.
[0137] EEE39. A non-transitory computer-readable storage medium recording a program of instructions that is executable by a device to perform the method according to any one of EEE23 to EEE37.
[0138] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.
[0139] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference tothe above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0140] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.
[0141] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.
Claims
CLAIMSWhat is claimed is:
1. A low-latency audio object clustering system, comprising: an audio object clustering unit (306) configured to calculate object-to-cluster gains of an object; a gain buffering unit (308) configured to store the object-to-cluster gains provided by the audio object clustering unit; an onset / offset detection unit (304) configured to set a flag based on at least one of a loudness, an object-to-cluster gain and a position of the object; and a gain interpolation unit (302) configured to calculate interpolated gain for an object with a gain interpolation function based on the flag and the object-to-cluster gains, wherein the flag indicates an interpolation ramp length for the gain interpolation function.
2. The system of claim 1, wherein the onset / offset detection unit (304) is configured to set the flag for a current audio block based on a difference in loudness, object-to-cluster gain or position of the object between the current audio block of a current frame and a preceding audio block of the current frame or a preceding frame, and wherein the flag indicates a shorter interpolation ramp length with an increasing difference in loudness, object-to-cluster gain or position of the object.
3. The system of claim 1 or claim 2, wherein the audio object clustering unit (306) is configured to select the most important audio objects as cluster centroids.
4. The system of any of claims 1-3, wherein the audio object clustering unit (306) is configured to calculate the object-to-cluster gains for each object for a selected cluster.
5. The system of any of claims 1-4, wherein the gain interpolation unit (302) is configured to calculate the interpolated gain using a current object-to-cluster gains calculated by the audio object clustering unit (306) and using a saved object-to-cluster gains stored by the gain buffering unit (308).
6. The system of any of claims 1-4, wherein the onset / offset detection unit (304) determines whether temporal change occurs for at least one of the loudness, the object-to-cluster gain and theposition of the object in a current audio block, and optionally wherein whether the temporal change occurs is determined based on at least one of a received audio block and a metadata associated with the received audio block.
7. The system of claim 6, wherein the received audio block and / or the received metadata are associated with one of a prior received frame or a current frame.
8. The system of claim 6, wherein the gain interpolation unit (302) is configured to select an interpolated gain function based on whether the temporal change occurs.
9. The system of any of claims 1 - 8, wherein the onset / offset detection unit (304) is configured to: set the flag to a first value when a temporal change occurs in the current audio block; and set the flag to a second value when no temporal change occurs in the current audio block.
10. The system of any of claims 1- 8, wherein the onset / offset detection unit (304) is configured to: set the flag to a first value when an energy of a current audio block is above an onset threshold or when an energy of the audio block is below an offset threshold; and set the flag to a second value when the energy of the audio block is between the onset and offset threshold.
11. The system of claim 10, wherein the offset threshold and / or onset threshold is based on the energy of a previous audio block or wherein the gain interpolation unit (302) is configured to calculate the interpolated gain for the object based on flags set for the energy of a previous audio block.
12. The system of any of claims 9 - 11, wherein the gain interpolation unit (302) is configured to calculate the interpolated gain for the current audio block using an interpolation gain function having an interpolation ramp length, wherein the gain interpolation unit (302) is configured to calculate the interpolated gain using a first interpolated gain function having a first interpolation ramp length in response to the flag being set to the first value, andwherein the gain interpolation unit (302) is configured to calculate the interpolated gain using a second interpolated gain function having a second interpolation ramp length in response to the flag being set to the first value, and wherein the first interpolation ramp length is shorter than the second interpolation ramp length.
13. The system according to claim 12, wherein the second interpolation ramp length corresponds to the length of the current audio block.
14. The system of any of claims 1 - 7, wherein the onset / offset detection unit (304) is configured to determine an interpolation ramp length based on a detected onset condition for a current frame wherein the flag indicates the determined interpolation ramp length, and optionally wherein the interpolation ramp length is determined from a difference in an initial gain and a target gain for the current frame.
15. The system of claim 14, wherein the determined interpolation ramp length is shorter when the onset occurs towards a beginning of the current frame, and the selected interpolation ramp length is determined to be longer when the onset occurs towards an end of the current frame.
16. The system of any of claims 1 - 15, wherein the object-to-cluster gains comprises a current object- to-cluster gain associated with a current audio block of a frame and a previous object-to-cluster gain associated with a previous audio block included in the frame or wherein the previous object-to-cluster gain is associated with a previous audio block included in a previous frame.
17. A method (700) for clustering audio objects, the method (700) comprising: receiving (705) an audio block including an audio object; calculating (710) object-to-cluster gains for the audio object; setting (715) a flag based on one or more of a loudness of the audio object, a position of the audio object or an object-to-cluster gain of the audio object; calculating (720), for the audio object, an interpolated gain based on the flag and the object- to-cluster gains, with a gain interpolation function, wherein the flag indicates an interpolation ramp length for the gain interpolation function; and processing (725) the input audio block using the interpolated gain.
18. The method (700) of claim 17, wherein the object- to-cluster gains comprises a current object- to-cluster gain associated with a current audio block of a frame and a previous object- to-cluster gain associated with a previous audio block included in the frame or wherein the previous object-to-cluster gain is associated with a previous audio block included in a previous frame.
19. The method (700) of any of claims 18, wherein calculating (720) the interpolated gain includes selecting a ramp function based on the flag, wherein the previous object-to-cluster gain is set as an initial gain value of the audio block, and wherein the current object-to-cluster gain is set as a target gain value of the audio block.
20. The method (700) of any of claims 17-19, wherein calculating (720) the interpolated gain includes: calculating a difference between a current object-to-cluster gain and a previous object-to- cluster gain; comparing the difference to a gain threshold; and selecting a ramp function based on the comparison, and optionally wherein selecting the ramp function based on the comparison includes selecting a first ramp function when the comparison is greater than the gain threshold or when the flag is set to indicate that the energy of the frame is above an energy threshold, and selecting a second ramp function when the comparison is less than the gain threshold and the flag is set to indicate that the energy of the frame is less than the energy threshold.
Citation Information
Patent Citations
Spatial error metrics of audio content
US20160337776A1
Efficient coding of audio scenes comprising audio objects
US20170180905A1
Adaptive loudness normalization for audio object clustering
US20220159395A1
Improved main-associated audio experience with efficient ducking gain application
US20230247382A1