Improving primary-associated audio experience with efficient ducking gain applications

By implementing built-in ramping operations and subframe gain smoothing techniques in the audio renderer, the artifact problem caused by excessive frame-level gain changes is solved, improving the quality and consistency of audio presentation.

CN115668364BActive Publication Date: 2026-01-23DOLBY INTERNATIONAL AB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180038468.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-26
Filing Date
2021-05-20
Publication Date
2026-01-23
Estimated Expiration
2041-05-20

AI Technical Summary

Technical Problem

During audio processing, excessive changes in ducking gain between frames can cause audible artifacts, such as a "zipper" effect, which affects the quality of audio presentation.

Method used

By implementing a built-in ramp operation in the audio renderer, gain variations are smoothed using a finer time scale than the audio frame. This includes calculating subframe gain to reduce abrupt changes in frame-level gain and combining the ramp length generated by the built-in or decoder for gain smoothing.

Benefits of technology

It effectively reduces or prevents audible artifacts, improves the smoothness and quality of audio presentation, and meets the expectations of content creators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0003964929230000011
    Figure HDA0003964929230000011
  • Figure HDA0003964929230000021
    Figure HDA0003964929230000021
  • Figure HDA0003964929230000031
    Figure HDA0003964929230000031
Patent Text Reader

Abstract

An audio bitstream is decoded into an audio object and audio metadata for the audio object. The audio object comprises a particular audio object. The audio metadata specifies a frame-level gain comprising a first gain and a second gain for a first audio frame and a second audio frame, respectively. It is determined, based on the first gain and the second gain, whether sub-frame gains are to be generated for the particular audio object. If so, a ramp length of a ramp used to generate the sub-frame gains for the particular audio object is determined. The sub-frame gains are generated for the particular audio object using the ramp having the ramp length. An audio speaker renders a soundfield represented by the audio object having the sub-frame gains.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to the following prior applications: U.S. Provisional Application 63 / 029,920 (reference number: D20015USP1), filed May 26, 2020, and European Application 20176543.5 (reference number: D20015EP), filed May 26, 2020, which are incorporated herein by reference. Technical Field

[0003] This invention relates generally to processing audio signals, and more specifically to improving the master-associated audio experience by utilizing efficient ducking gain applications. Background Technology

[0004] Multiple audio processors are distributed across an end-to-end audio processing chain to deliver audio content to end-user devices. Different audio processors may perform different, similar, and / or even repetitive media processing operations. Some of these operations can easily introduce audible artifacts. For example, an audio bitstream generated by an upstream encoding device can be decoded to provide a presentation of audio content consisting of a "master audio" and "associated audio." To control the balance between the master audio and associated audio in the decoded presentation, the audio bitstream may carry audio metadata specifying "dodge gain" at the audio frame level. In audio rendering operations, without adequate smoothing of the gain values, large changes in the dodge gain between frames can lead to audible degradation in the decoded presentation, such as "zipper" artifacts.

[0005] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise specified, no method described in this section should be assumed to be prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise specified, it should not be assumed that the problems identified with respect to one or more methods have already been identified in any prior art based on this section. Attached Figure Description

[0006] The invention is illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings:

[0007] Figure 1 The illustration shows an example audio encoding device;

[0008] Figures 2A to 2C The illustration shows an example downstream audio processor;

[0009] Figures 3A to 3D The illustration shows an example subframe gain smoothing operation;

[0010] Figure 4 An example process flow is illustrated; and

[0011] Figure 5 An example hardware platform is illustrated, on which a computer or computing device as described herein can be implemented. DETAILED DESCRIPTION

[0012] Example embodiments are described herein that relate to improving primary-associated audio experiences with efficient ducking gain applications. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present application. It will be apparent, however, that the present application can be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail, in order to avoid unnecessarily

[0013] Example embodiments are described herein according to the following outline:

[0014] 1. OVERALL SUMMARY

[0015] 2. UPSTREAM AUDIO PROCESSOR

[0016] 3. DOWNSTREAM AUDIO PROCESSOR

[0017] 4. SUBFRAME GAIN GENERATION

[0018] 5. EXAMPLE PROCESS FLOW

[0019] 6. IMPLEMENTATION MECHANISMS - HARDWARE OVERVIEW

[0020] 7. EQUIVALENTS, EXTENSIONS, ALTERNATIVES, AND OTHERS

[0021] 1. OVERALL SUMMARY

[0022] This Summary introduces aspects of some embodiments of the present application. As such, this Summary is not an extensive or exhaustive overview of aspects of embodiments. Furthermore, this Summary is not intended to be construed as identifying any specifically important aspects of embodiments or to be used to limit the scope of embodiments. This Summary is merely intended to introduce some concepts of example embodiments in a simplified and abbreviated format, and is not intended to be construed as identifying any specifically important aspects of embodiments, or to be used to limit the scope of example embodiments. Note that, although discussion is made herein of individual embodiments, any combination of the embodiments and / or partial embodiments discussed herein can be combined to form further embodiments.

[0023] An audio bitstream as described herein can be encoded with audio signals containing object essence of audio objects and audio metadata (or object audio metadata) of audio objects, including but not limited to side information for reconstructing audio objects. The audio bitstream can be encoded according to a media coding syntax, such as an AC-4 coding syntax, an MPEG-H coding syntax, etc.

[0024] An audio object in the audio bitstream can be a static-only audio object, a dynamic-only audio object, or a combination of static and dynamic audio objects. An example static audio object can include, but is not necessarily limited to, any one of the following: a bed object, a channel content, an audio bed, an audio object whose respective spatial position is fixed by assignment to an audio speaker in an audio channel configuration, etc. An example dynamic audio object can include, but is not necessarily limited to, any one of the following: an audio object with time-varying spatial information, an audio object with time-varying motion information, an audio object whose position is not fixed by assignment to an audio speaker in an audio channel configuration, etc.

[0025] Spatial information of a static audio object (such as a spatial position of a static audio object) can be inferred from a (audio) channel ID of the static audio object. Spatial information of a dynamic audio object (such as a time-varying or time-constant spatial position of a dynamic audio object) can be indicated or specified in audio metadata of the dynamic audio object or a specific portion thereof.

[0026] One or more audio programs can be represented or included in the audio bitstream. Each audio program in the audio bitstream can include a corresponding subset or combination of audio objects among all audio objects represented in the audio bitstream.

[0027] The audio bitstream can be transmitted / delivered, directly or indirectly, to a recipient decoding device and decoded by the recipient decoding device. The decoding device can operate with an audio Tenderer, such as an object audio Tenderer, to drive audio speakers (or output channels) in an audio rendering environment to reproduce a soundfield (or sound scene) that depicts sound sources represented by audio objects of the audio bitstream.

[0028] In some operational scenarios, the audio metadata of the audio bitstream can include an audio metadata parameter (encoded or embedded into the audio bitstream by an upstream encoding device according to a media coding syntax) for indicating time-varying frame-level gain values of one or more audio objects in the audio bitstream.

[0029] For example, an audio object in the audio bitstream can be specified in the audio metadata to undergo a transient change in gain value from the previous audio frame to the subsequent audio frame in the audio bitstream. The audio object can be part of a "main audio" program that is to be simultaneously mixed with the "associated audio" program via the time-varying gain value during a ducking operation. In some embodiments, the "main audio" program or content includes separate "music and effects" content / programming and separate "dialogue" content / programming, each distinct from the "associated audio" program or content. In some embodiments, the "main audio" program or content includes "music and effects" content / programming (e.g., excluding "dialogue" content / programming, etc.), and the "associated audio" program includes "dialogue" content / programming (e.g., excluding "music and effects" content / programming, etc.).

[0030] Upstream encoding devices can generate time-varying ducking (attenuation) gains for some or all audio objects in the "main audio" to successively reduce the loudness level of the "main audio". Correspondingly, upstream encoding devices can generate time-varying ducking (enhancement) gains for some or all audio objects in the "related audio" to successively increase the loudness level of the "related audio".

[0031] Instantaneous changes in gain, indicated at the frame level, can be performed by the audio decoding device at the receiver of the audio bitstream. In some methods, without sufficient smoothing by the receiver's audio decoding device, relatively large changes in gain can easily introduce audible artifacts, such as a "zipper" effect, into the decoded presentation.

[0032] Conversely, techniques as described herein can be used to provide smooth operation that prevents or reduces these audible artifacts. Under these techniques, an audio renderer in the receiver audio decoding device (with a built-in ability to handle dynamic changes in audio objects related to motion) can be adapted to utilize this built-in ability to smooth instantaneous changes in gain specified for an audio object on a timescale much finer than the timescale of an audio frame. For example, the audio renderer can be adapted to implement a built-in ramp to smooth gain changes in an audio object, where multiple additional subframe gains are calculated on the built-in ramp. A ramp length can be input to the audio renderer for use with the built-in ramp. The ramp length represents a time interval within which subframe gains, in addition to or replacing the frame-level gains sent by the encoder, can be calculated or generated using one or more gain smoothing / interpolation algorithms. Instead of applying the same frame-level gain to all subframe units in a frame, the subframe gains described herein can include smoothly differentiated values ​​of different QFM slots and / or different PCM samples within the same audio frame. As used herein, "encoder-sent" operation parameters, such as frame-level gain, can refer to operation parameters or gains encoded by an upstream device (including, but not limited to, an audio encoder) into the audio bitstream or in its audio metadata. In one example, such "encoder-sent" operation parameters or gains can be generated and encoded into the audio bitstream by the upstream device without receiving the parameter / gain or its specific value. In another example, such "encoder-sent" operation parameters or gains can be received, converted, translated, and / or encoded into the audio bitstream by the upstream device from input parameters / gains (or input values ​​for them). Input parameters / gains (or input values ​​for them) can be received or specified in user input or input content received by the upstream device.

[0033] The audio object receiving the audio bitstream, along with time-varying gains such as ducking gain, can be a static audio object (or bed object) that is part of the channel content. The audio metadata received from the bitstream may not specify the ramp length for the static audio object. The audio decoding device can modify the received audio metadata to add specifications for the ramp length of the built-in ramp. The frame-level ducking gain in the received audio metadata can be used to set or obtain a target gain. The ramp length and target gain enable the audio renderer to perform gain smoothing operations for static audio using the built-in ramp.

[0034] An audio object that receives time-varying gains, such as ducking gain, along with the audio bitstream, can be a dynamic audio object that is part of the object's audio. Similar to a static audio object, the frame-level ducking gain received in the audio bitstream can be used to set or obtain a target gain.

[0035] In some operational scenarios, the encoder sends a ramp length along with the audio bitstream to a dynamic audio object. The ramp length and target gain sent by the encoder can be used by the audio renderer to perform gain smoothing on the dynamic audio object using built-in ramps. Using the ramp length sent by the encoder may or may not effectively prevent audible artifacts. It should be noted that in various embodiments, the ramp length may or may not be generated directly or entirely by the encoder for the audio object. In some operational scenarios involving film content, the encoder may not generate the ramp length directly or entirely for the audio object. The ramp length may be received by the encoder as part of the input (including but not limited to the audio content itself including audio samples and metadata), and the encoder then encodes, transforms, or translates the input including the ramp length of the audio object into an output bitstream according to an applicable bitstream syntax. In some operational scenarios involving broadcast content, the encoder may generate the ramp length directly or entirely for the audio object, encoding the ramp length of the audio object along with the audio samples and metadata obtained from the input into the output bitstream according to an applicable bitstream syntax.

[0036] In some operational scenarios, regardless of whether the ramp length sent by the encoder is received, the audio decoding device still modifies the audio metadata to add a specification of the ramp length generated by the decoder with a built-in ramp. Using the ramp length generated by the decoder can effectively prevent audible artifacts, but there is a risk of altering some aspects of the audio rendering of dynamic audio objects, because intermediate frame level gains can be received in the audio bitstream within the time interval corresponding to the ramp length generated by the decoder and can be ignored in the audio rendering of dynamic audio objects.

[0037] In some operational scenarios, regardless of whether the ramp length sent by the encoder is received, the audio decoding device still modifies the audio metadata to add a specification of the ramp length generated by the decoder with a built-in ramp. Using the ramp length generated by the decoder can effectively prevent audible artifacts. Additionally, optionally or alternatively, the audio renderer can implement a smoothing / interpolation algorithm that merges or enforces intermediate frame-level gain received with the audio bitstream within a time interval corresponding to the ramp length generated by the decoder. This effectively prevents audible artifacts while maintaining the audio rendering of dynamic audio objects as intended by content creators.

[0038] Some or all of the techniques described can be widely applied to various media systems that implement various audio processing techniques (including, but not limited to, those related to AC-4, DD+JOC, MPEG-H, etc.).

[0039] In some embodiments, the apparatus as described herein forms part of a media processing system, which includes, but is not limited to: audiovisual equipment, flat-screen TVs, handheld devices, game consoles, televisions, home theater systems, soundbars, tablet computers, mobile devices, laptops, netbooks, cellular wireless phones, e-book readers, point-of-sale terminals, desktop computers, computer workstations, media streaming devices, computer kiosks, various other types of terminals and media processors, etc.

[0040] Various modifications to the preferred embodiments, general principles, and features described herein will be readily apparent to those skilled in the art. Therefore, this disclosure is not intended to be limited to the illustrated embodiments, but rather to be consistent with the maximum scope of the principles and features described herein.

[0041] 2. Upstream audio processor

[0042] Figure 1 The illustration depicts an example upstream audio processor, such as an audio encoding device (or audio encoder) 150. The audio encoding device (150) may include a source audio content interface 152, an audio metadata generator 154, an audio bitstream encoder 158, etc. The audio encoding device 150 may be part of a broadcast system, an Internet-based media streaming server, an over-the-air (OTA) network operator system, a film production system, a local media content server, a media transcoding system, etc. Some or all of the components of the audio encoding device (150) may be implemented in hardware, software, a combination of hardware and software, etc.

[0043] The audio encoding device uses a source audio content interface (152) to obtain or receive source audio content from one or more content sources and / or systems, the source audio content including one or more source audio signals 160 representing the object nature of one or more source audio objects, source object spatial information 162 of one or more audio objects, etc.

[0044] The received source audio content can be used by an audio encoding device (150) or a bitstream encoder (158) therein to generate an audio bitstream 102 encoded with one or more of the following: a single audio program, multiple audio programs, commercials, movies, coexisting main audio programs and associated audio programs, continuous audio programs, audio portions of media programs (e.g., video programs, audiovisual programs, pure audio programs, etc.).

[0045] The object essence of the source audio object in one or more source audio signals (160) of the received source audio content may include positionless PCM-coded audio sample data. The spatial information (162) of the source object in the received source audio content may be received by the audio encoding device (150) alone (e.g., in auxiliary source data input, etc.) or jointly with the object essence of the source audio object in one or more source audio signals (160). Example source audio signals carrying the object essence of the audio object (and possibly spatial information of the audio object) as described herein may include, but are not limited to, some or all of the following: source channel content signals, source audio bed channel signals, source object audio signals, audio feeds, audio tracks, dialogue signals, ambient sound signals, etc.

[0046] Source audio objects may include one or more of the following: static audio objects (which may be referred to as “bed objects” or “channel content”), dynamic audio objects, etc. A static audio object or bed object may refer to a non-moving object mapped to a specific speaker or channel location in an audio channel configuration (e.g., output, input, intermediate, etc.). A static audio object, as described herein, may represent or correspond to some or all of the audio beds to be encoded into the audio bitstream (102). A dynamic audio object, as described herein, may move freely within some or all of a 2D or 3D sound field to be depicted by rendering audio data in the audio bitstream (102).

[0047] Source object spatial information (162) includes some or all of the following: the location and extent of the source audio object, importance, spatial exclusion, divergence, etc.

[0048] The audio metadata generator (154) generates audio metadata to be included or embedded in the audio bitstream (102) from the received source audio content (such as the source audio signal (160) and source object spatial information (162)). The audio metadata includes object audio metadata, auxiliary information, etc. Some or all of the audio metadata can be carried in the audio metadata container, field, parameters, etc., and are separated from the audio sample data encoded in the audio bitstream (102) according to bitstream encoding syntax such as AC-4 and MPEG-H.

[0049] The audio metadata transmitted to the receiving audio playback system may include an audio metadata portion that guides the receiving playback system's object audio renderer (which performs some or all of the audio rendering phases) to render audio data (corresponding to the audio metadata) within the specific playback (or audio rendering) environment in which the receiving playback system operates. Different audio metadata portions reflecting variations in different audio scenes may be sent to the receiving playback system for rendering audio scenes or their subdivisions.

[0050] The Object Audio Metadata (OAMD) in the audio bitstream (102) can specify or be used to obtain audio operation parameters for the receiving device of the audio bitstream (102) to render the audio object. Auxiliary information in the audio bitstream (102) can specify or be used to obtain audio operation parameters for the receiving device of the audio bitstream (102) to reconstruct the audio object from the audio signal, which is encoded by the audio encoding device (150) and decoded by the receiving device from the audio bitstream (102).

[0051] Example audio operation parameters represented in the audio metadata of the audio bitstream (102) (e.g., sent by the encoder, generated by the upstream device, etc.) may include, but are not limited to, the following: object gain, ducking gain, dialogue normalization gain, dynamic range control gain, peak limit gain, frame level / resolution gain, position, media description data, renderer metadata, translation coefficients, submix gain, downmix coefficients, upmix coefficients, reconstruction matrix coefficients, timing control data, etc., some or all of which may change dynamically as a function of time.

[0052] In some operating scenarios, each audio operating parameter (e.g., gain, timing control data, etc.) among some or all of the audio operating parameters represented in the audio bitstream (102) can be wideband or broadband, applicable to all frequencies, samples, or subbands in the audio frame.

[0053] The audio objects represented or encoded in the audio bitstream (102) (such as audio objects generated by the audio encoding device (150)) may or may not be the same as the source audio objects represented in the source audio content received by the audio encoding device (150). In some operational scenarios, spatial analysis is performed on the source audio objects to combine or cluster one or more source audio objects into (encoded) audio objects represented in the audio bitstream (102) having spatial information of the encoded audio objects. The spatial information of the encoded audio objects formed by combining or clustering one or more source audio objects can be obtained from the source spatial information of one or more source audio objects in the source object spatial information (162).

[0054] The audio signal representing an audio object (which may be the same as, or derived from, or clustered from, a source audio object) can be encoded in the audio bitstream (102) based on a reference audio channel configuration (e.g., 2.0, 3.0, 4.0, 4.1, 4.1, 5.1, 6.1, 7.1, 7.2, 10.2, 10-60 speaker configuration, 60+ speaker configuration, etc.). For example, the audio object can be shifted to one or more reference audio channels (or speakers) in the reference audio channel configuration. Submixes (or downmixes) for the reference audio channels (or speakers) in the reference audio channel configuration can be generated by shifting from some or all of the contributions from some or all of the audio objects in the audio object. The submixes can be used to generate corresponding audio signals for the reference channels (or speakers) in the reference audio channel configuration. The reconstruction operation parameters can be obtained at least in part from translation coefficients, spatial information of audio objects, etc., for encoder-side translation and submixing / downmixing operations, and are passed in audio metadata (e.g., auxiliary information, etc.) so that the receiving device of the audio bitstream (102) can reconstruct the audio objects represented in the audio bitstream (102).

[0055] The audio bitstream (102) may be transmitted directly or indirectly or otherwise delivered to the receiving device in a series of transmission frames. Each transmission frame may include one or more audio frames that carry a series of PCM samples or encoded audio data, such as a QMF matrix, within the same (frame) time interval (e.g., 20 milliseconds, 10 milliseconds, short or long frame intervals, etc.) for all audio channels (or speakers) in a reference audio channel configuration. The audio bitstream (102) may include a continuous sequence of audio frames that includes PCM samples or encoded audio data covering the continuous (frame) time interval sequence. The continuous (frame) time interval sequence may constitute the duration of a media program (e.g., replay, playback, live broadcast, video streaming, etc.), the audio content of which is at least partially encoded in or provided in the audio bitstream (102).

[0056] The time interval represented by an audio frame as described herein may include multiple subframe time intervals represented by multiple corresponding QMF (time) slots. Each subframe time interval of the audio frame may correspond to a corresponding QMF slot in the multiple corresponding QMF slots. The QMF slots as described herein may be represented by matrix columns in the QMF matrix of the audio frame and include spectral elements of multiple frequencies or subbands that together constitute a frequency band or broadband (e.g., covering some or all of the entire frequency band audible to the human auditory system).

[0057] The audio encoding device (150) can perform a number of (encoder-side) audio processing operations that alter the gain of one or more audio objects (among all audio objects) represented in the audio bitstream (102). These gains can be applied directly or indirectly by the receiving device of the audio bitstream (102) to one or more audio objects in the audio rendering operation—for example, to change the loudness level or dynamics of one or more audio objects.

[0058] Example (encoder side) audio processing operations may include, but are not limited to, ducking, dialogue enhancement, user-controlled gain shifting (e.g., based on user input provided by the content creator or producer), downmixing, dynamic range control, peak limiting, crossfading, continuous or concurrent program mixing, gain smoothing, fade-in / fade-out, program switching, or other gain shifting operations.

[0059] By way of example, and not limitation, the audio bitstream (102) may cover a (gain transition) time period during which a first audio program of type "main audio" (referred to as the "main audio" program) and a second audio program of type "related audio" (referred to as the "related audio" program) are encoded or included in the audio bitstream (102) for use by a receiving device that simultaneously renders the audio bitstream (102). The "main audio" program may include a first subset of audio objects encoded or represented in the audio bitstream (102) or one or more of its first audio substreams. The "related audio" program may include a second subset of audio objects, different from the first subset of audio objects, encoded or represented in the audio bitstream (102) or one or more of its second audio substreams. The first subset of audio objects may mutually exclude or alternatively partially overlap with the second subset of audio objects.

[0060] An audio encoding device (150) or a frame-level gain generator (156) therein (which may be, but is not limited to, part of an audio metadata generator (154)) may perform ducking operations to (e.g., dynamically, over a time period, etc.) change or control the dynamic balance of loudness between the "main audio" program and the "associated audio" program during the (gain shift) time period. For example, these ducking operations may be performed to reduce the loudness level of some or all audio objects in a first subset of audio objects carried in one or more first substreams of the "main audio" program, while simultaneously increasing the loudness level of some or all audio objects in a second subset of audio objects in one or more second substreams of the "associated audio" program.

[0061] To control the balance between the "main audio" program and the "related audio" program in the decoded presentation, the audio metadata included in the audio bitstream (102) can provide or specify ducking gains for a first subset of audio objects in the "main audio" program and a second subset of audio objects in the "related audio" program, according to the bitstream encoding syntax. Content creators or producers can use ducking gains to scale or "duck" the "main audio" program content and simultaneously scale or "enhance" the "related audio" program content to make the "related audio" program content easier to understand than otherwise.

[0062] Dodge gains can be transmitted in the audio bitstream (102) at the frame level or on a per-frame basis (e.g., two gains for the main audio and associated audio of each frame, a gain for each frame where the gain changes from the previous value to the next different value, etc.). As used herein, “at the frame level” (or “at the frame resolution”) may mean providing or specifying a single instance / value of an operating parameter for a single audio frame or for multiple audio frames—e.g., a single instance / value of an operating parameter per frame. Specifying gains at the frame level can reduce bitrate usage associated with encoding, transmitting, receiving, and / or decoding the audio bitstream (102) (e.g., relative to specifying a higher resolution gain).

[0063] The audio encoding device (150) can avoid or reduce large variations in gain between frames (e.g., for one or more audio objects, etc.) to improve the user's listening experience. The audio encoding device (150) can cap the gain variation to no more than the maximum permissible gain variation between two consecutive audio frames. For example, a -12dB gain variation can be distributed over six consecutive audio frames, for example, by the frame-level gain generator (156) of the audio encoding device (150), with a step size of -2dB for each audio frame, which is below the maximum permissible gain variation.

[0064] 3. Downstream audio processor

[0065] Figure 2A An example downstream audio processor, such as audio decoding device 100, is illustrated. This audio decoding device includes an audio bitstream decoder 104, a subframe gain calculator 106, an audio renderer 108 (e.g., integrated, distributed, etc.), etc. Some or all of the components in the audio decoding device (100) may be implemented in hardware, software, a combination of hardware and software, etc.

[0066] The bitstream decoder (104) receives the audio bitstream (102) and performs demultiplexing and decoding operations on the audio bitstream (102) to extract the audio signal and audio metadata that have been encoded in the audio bitstream (102) by the audio encoding device (150).

[0067] The audio metadata extracted from the audio bitstream (102) may include, but is not limited to, the following: object gain, ducking gain, dialogue normalization gain, dynamic range control gain, peak limit gain, frame level / resolution gain, position, media description data, renderer metadata, translation coefficients, submix gain, downmix coefficients, upmix coefficients, reconstruction matrix coefficients, timing control data, etc. Some or all of this audio metadata may change dynamically as a function of one or more times.

[0068] The extracted audio signal and some or all of the extracted audio metadata (including but not limited to auxiliary information) can be used to reconstruct the audio object represented in the audio bitstream (102). In some operational scenarios, the extracted audio signal can be represented in a reference audio channel configuration. A reconstruction matrix that varies or remains constant over time can be created based on the auxiliary information and can be applied to the extracted audio signal in the reference audio channel configuration to generate or obtain an audio object. The reconstructed audio object can include one or more of the following: static audio objects (e.g., audio bed objects, channel content, etc.), dynamic audio objects (e.g., having a spatial location that varies or remains constant over time), etc. Object characteristics such as position and degree, importance, spatial exclusion, dispersion, etc., can be specified as part of the audio metadata received through the audio bitstream (102) or the object audio metadata (OAMD) therein.

[0069] The audio decoding device (100) can perform a number of (decoder-side) audio processing operations related to decoding and rendering audio objects in an output audio channel configuration (e.g., 2.0, 3.0, 4.0, 4.1, 4.1, 5.1, 6.1, 7.1, 7.2, 10.2, 10-60 speaker configuration, 60+ speaker configuration, etc.). Example (decoder-side) audio processing operations may include, but are not limited to, ducking, dialogue enhancement, user-controlled gain shifting (e.g., based on user input provided by a content creator or producer), downmixing, or other gain shifting operations.

[0070] Some or all of these decoder-side operations may involve applying differential gain (or differential gain value) to the decoder-side audio object at a finer instantaneous resolution than the frame-level instantaneous resolution. Examples of instantaneous resolutions finer than the frame-level instantaneous resolution may include, but are not limited to, those associated with one or more of the following: subframe level, per QMF slot, per PCM sample, etc. These decoder-side operations applied at a fairly fine instantaneous resolution may be referred to as gain smoothing operations.

[0071] For example, the audio bitstream (102) may cover gain variation / conversion durations (e.g., time intervals, intervals, sub-intervals, etc.), wherein a “main audio” program and a “related audio” program are encoded or included in the audio bitstream (102) for simultaneous rendering by a receiving device of the audio bitstream (102) with a gain varying over time. As previously mentioned, the “main audio” program and the “related audio” program may be respectively included in a first subset and a second subset of audio objects encoded or represented in the audio bitstream (102) or its audio substreams.

[0072] Upstream audio encoding devices (e.g., Figure 1 The audio bitstream (102) can perform ducking operations to (e.g., dynamically, during the gain change / conversion duration, etc.) change or control the dynamic balance of (loudness) between the "main audio" program and the "related audio" program during the (gain conversion) time period. Therefore, the time-varying gain (e.g., ducking, etc.) can be specified in the audio metadata of the audio bitstream (102). These gains can be provided in the audio bitstream (102) at the frame level or on a per-frame basis.

[0073] The frame-level gain sent by the encoder, transmitted in the bitstream—in this example related to the ducking operation, but generally can be extended to the time-varying gain associated with any gain change / conversion operation performed by the upstream encoding device—can be decoded from the audio bitstream (102) by the audio decoding device (100).

[0074] As the content creator intended, in the decoded presentation (or audio rendering) of the audio content in the audio bitstream (102), a ducking gain can be applied to the “main audio” program or content represented in the audio bitstream (102), while a corresponding (e.g., enhancement, etc.) gain can be applied to the accompanying “related audio” program or content represented in the audio bitstream (102).

[0075] Additionally, optionally, or alternatively, in some operational scenarios, the audio decoding device (100) may receive user input 118 from one or more user controls (or user interface components) provided with the audio decoding device (100) and interacting with the listener. The user input (118) may specify or be used to obtain user adjustments to be applied to a time-varying frame-level gain received in the audio bitstream (102), such as the ducking gain in this example. Through one or more user controls, the listener may cause the primary / associative balance to be changed, for example, making the "primary audio" more audible than the "associative audio," or vice versa, or another balance between the "primary audio" and the "associative audio." The listener may also choose to listen to the "primary audio" or the "associative audio" individually or as a whole; in this case, during the duration during which both the "primary audio" and the "associative audio" programs are represented in the audio bitstream (102), it may only be necessary to decode and render one of the "primary audio" and the "associative audio" programs in the decoded presentation of the audio bitstream (102).

[0076] For illustrative purposes only, if the audio object decoded or generated from the audio bitstream (102) includes a specific audio object, the time-varying gain at the frame level for that specific audio object is specified in or obtained from the audio metadata in the audio bitstream (102), and this gain may be further adjusted or modified, at least in part, based on user input (118).

[0077] A specific audio object can refer to any audio object whose gain over time is specified in the audio metadata of the audio bitstream (102). In some operational scenarios, a first subset of audio objects decoded or generated from the audio bitstream (102) represents the "main audio" program, while a second subset of audio objects decoded or generated from the audio bitstream (102) represents the "related audio" program. A specific audio object can belong to either the first subset or the second subset.

[0078] The time-varying gain at the frame level for a specific audio object may include a first gain (value) and a second gain (value) for the first and second audio frames in the audio frame sequence carried in the audio bitstream (102), respectively.

[0079] The first audio frame may correspond to a first time point in the sequence of time points (e.g., frame indices, etc.) in the decoded presentation (e.g., represented by the first frame index logic, etc.), and includes a first audio signal portion for obtaining a first object-essential portion (e.g., PCM samples, transform coefficients, positionless audio data portions, etc.) of a particular audio object. Similarly, the second audio frame may correspond to a second time point in the sequence of time points (e.g., frame indices, etc.) in the decoded presentation (e.g., represented by the second frame index logic, after or following the first time point, etc.), and includes a second audio signal portion for obtaining a second object-essential portion (e.g., PCM samples, transform coefficients, positionless audio data portions, etc.) of a particular audio object.

[0080] In the example, the first audio frame and the second audio frame can be two consecutive audio frames in the audio frame sequence encoded in the audio bitstream (102). In another example, the first audio frame and the second audio frame can be two non-consecutive audio frames in the audio frame sequence encoded in the audio bitstream (102); the first audio frame and the second audio frame can be separated by one or more intermediate audio frames in the audio frame sequence.

[0081] The first and second gains can be associated with one of the following: ducking, dialogue enhancement, user-controlled gain shift, downmixing, or other gain shift, as well as any combination thereof.

[0082] An audio decoding device (100) or a subframe gain calculator (106) therein can determine whether to perform a subframe gain smoothing operation on a first gain and a second gain. This determination can be performed at least in part based on a minimum gain difference threshold, which can be zero or a non-zero value. In response to determining that the difference between the first gain and the second gain (e.g., absolute value, amplitude, etc.) exceeds the minimum gain difference threshold (e.g., absolute value, amplitude, etc.), the subframe gain calculator (106) applies a subframe gain smoothing operation to audio frames between the first audio frame and the second audio frame (e.g., including the first audio frame and the second audio frame, excluding the first audio frame and the second audio frame, etc.).

[0083] In some operating scenarios, the minimum gain difference threshold can be non-zero; therefore, when the difference between the first gain and the second gain is relatively small compared to the non-zero minimum threshold, the gain smoothing operation or corresponding calculation may not be invoked, because small differences are unlikely to cause audible artifacts.

[0084] Alternatively, this determination may be performed at least in part based on a minimum gain change rate threshold. In response to determining that the rate of change (e.g., absolute value, amplitude, etc.) between the first and second gains exceeds the minimum gain change rate threshold (e.g., absolute value, amplitude, etc.), a subframe gain calculator (106) applies a subframe gain smoothing operation to audio frames between the first and second audio frames (e.g., including the first and second audio frames, excluding the first and second audio frames, etc.). The rate of change between the first and second gains can be calculated as the difference between the first and second gains divided by the time difference between the first and second gains. In some operational scenarios, the time difference can be logically represented or calculated based on the difference between the first frame index of the first audio frame and the second frame index of the second audio frame.

[0085] In some operating scenarios, the minimum gain rate of change threshold can be non-zero; therefore, when the rate of change between the first gain and the second gain is relatively small compared to the minimum gain rate of change threshold, the gain smoothing operation or corresponding calculation may not be invoked, because a small rate of change is unlikely to cause audible artifacts.

[0086] In some operational scenarios, determining whether to perform subframe gain smoothing can be symmetrical. For example, the same minimum gain difference threshold or the same minimum gain change rate threshold can be used to determine whether the change or rate of change in the gain value is positive (e.g., enhancement or improvement) or negative (e.g., dodge or reduction). In the determination, the absolute value of the difference can be compared with a threshold of the absolute value.

[0087] The human auditory system can respond by increasing and decreasing loudness levels using varying integration times. In some operational scenarios, determining whether to perform subframe gain smoothing can be asymmetric. For example, different minimum gain difference thresholds or different minimum gain rate of change thresholds (e.g., converted to absolute or amplitude values) can be used to make a determination based on whether the change or rate of change in the gain value is positive (e.g., enhancement or boost) or negative (e.g., ducking or reduction). The change or rate of change in the gain value can be converted to absolute or amplitude values ​​and then compared to a specified threshold from among the different minimum gain difference thresholds or different minimum gain rate of change thresholds.

[0088] Alternatively, or alternatively, one or more other determining factors may be used to determine whether to perform gain smoothing operations such as interpolation. Examples of determining factors may include, but are not limited to, any of the following: aspects and / or characteristics of the audio content, aspects and / or characteristics of the audio object, system resource availability of the audio decoding device and / or audio encoding device or its processing components, and the system resource usage of the audio decoding device and / or audio encoding device or its processing components.

[0089] In response to determining that a gain smoothing operation should be performed on a specific audio object related to a first gain specified for a first audio frame and a second gain specified for a second audio frame, the subframe gain calculator determines the ramp length (e.g., decoder-side insertion, timing data, etc.) for smoothing or interpolating the gain ramp to be applied to the specific audio object between the first gain specified for the first audio frame and the second gain specified for the second audio frame. The example gain smoothing / interpolation algorithms described herein may include, but are not limited to, combinations of one or more of the following: piecewise constant interpolation, linear interpolation, polynomial interpolation, spline interpolation, etc. Additionally, optionally, or alternatively, gain smoothing / interpolation algorithm operations may be applied individually to a single audio channel, a single audio object, a single time period / interval, etc. In some operational scenarios, the smoothing / interpolation algorithms described herein may implement a smoothing / interpolation function modified or modulated using a psychoacoustic function, which may be a nonlinear function describing or representing a perceptual model of the human auditory system. Smoothing / interpolation algorithms, or the timing control implemented therein, can be specifically designed to provide smoothed loudness levels with little or no perceptible audio artifacts such as “zipper” effects.

[0090] Audio metadata in an audio bitstream (102) provided by an upstream encoding device may be unaffected by the ramp length specification. In some operational scenarios, audio metadata may specify the ramp length sent by a separate encoder for a particular audio object; this separate encoder-sent ramp length may differ from the ramp length determined by a subframe gain calculator (106) (e.g., decoder generation, etc.). In the example, the particular audio object is a dynamic audio object in a film media program (e.g., a non-bed object, non-channel content, with spatial information that changes over time, etc.). In another example, the particular audio object is a static audio object in a broadcast media program. By comparison, in some operational scenarios, audio metadata may not specify the ramp length sent by any separate encoder for a particular audio object. In the example, the particular audio object is a static audio object in a broadcast media program that is not a broadcast media program or for which the encoder has not yet specified the ramp length (e.g., a bed object, channel content, with a fixed position corresponding to a channel ID in the audio channel configuration, etc.). In another example, the particular audio object is a dynamic audio object in a non-film media program for which the encoder has not yet specified the ramp length.

[0091] To implement the gain smoothing operation as described herein, the subframe gain calculator (106) can calculate or generate the subframe gain based on a first gain, a second gain, and a ramp length. Example subframe gains may include, but are not limited to, any of the following: wideband gain, broadband gain, narrowband gain, frequency-specific gain, binary-specific gain, time-domain gain, transform-domain gain, frequency-domain gain, gain applicable to encoded audio data in a QFM matrix, gain applicable to PCM sample data, etc. The subframe gain may differ from the frame-level gain obtained from the audio bitstream (102). For example, a subframe gain generated or calculated for a time interval covering a ramp with a ramp length may be a superset of any frame-level gain specified for the same time interval in the audio stream (102). Subframe gains may include one or more interpolated gains at the subframe level, per QFM slot, per PCM sample, etc. Within an audio frame between the first and second frames (including the first and second frames), two different subframe units (e.g., two different QFM slot bases, two different PCM samples, etc.) can be assigned to two different subframe gains (or different subframe gain values).

[0092] In some operational scenarios, the subframe gain calculator (106) interpolates a first gain specified for a first audio frame to a second gain specified for a second audio frame to generate a subframe gain for a specific audio object within a time interval represented by the ramp length. The contributions from different subframe units (such as QMF slots or PCM samples between the first and second audio frames) to a specific audio object can be allocated using different (or differentiated) subframe gains in the calculated subframe gain.

[0093] The subframe gain calculator (106) can generate or obtain subframe gains for some or all of the audio objects represented in the audio bitstream (102) based at least in part on the frame-level gains specified for audio frames that contain audio data contributing to the audio objects. These subframe gains (e.g., those including specific audio objects) of some or all of the audio objects represented in the audio bitstream (102) can be provided by the subframe gain calculator (106) to the audio renderer (108).

[0094] In response to the subframe gain of the received audio object, the audio renderer (108) performs a gain smoothing operation to apply differential subframe gain to the audio object at a finer instantaneous resolution than the frame-level instantaneous resolution (e.g., at the subframe level, on a per-QMF slot basis, on a per-PCM sample basis, etc.). Additionally, optionally, or alternatively, the audio renderer (108) causes a set of audio speakers operated by the audio decoding device (100) to render the sound field represented by the audio object (where the subframe gain is applied to the audio object) in a particular playback environment (or a particular output audio channel configuration).

[0095] In some methods, the decoder can apply variations in gain values, such as those associated with skipping the "main audio" program while simultaneously enhancing the "related audio" program at the frame level. Frame-level gains, as specified in the audio bitstream, can be applied on a per-frame basis. Therefore, each subframe unit in an audio frame (such as a QMF slot or PCM sample) can implement the same wideband or broadband (e.g., perceptual, non-perceptual, etc.) gain as specified for the audio frame, without gain smoothing or interpolation. Without subframe gain smoothing, this results in a "zipper" artifact, where discontinuous changes in loudness levels can be perceived by the listener (as an audible artifact).

[0096] Conversely, with the techniques described herein, gain smoothing can be implemented or performed, at least in part, based on subframe gains calculated at a finer instantaneous resolution than the frame level. Therefore, audible artifacts such as "zipper" artifacts can be eliminated or significantly reduced.

[0097] In some approaches, the upstream device of the audio renderer can perform or apply interpolation operations such as linear interpolation with linear gain to the QMF slots or PCM samples in the audio frame. However, this would be computationally expensive, complex, and / or repetitive, especially when the audio frame may contain numerous contributions from many audio signals, many audio objects, etc., to the audio data portion.

[0098] Conversely, in the techniques described herein, gain smoothing operations (including, but not limited to, interpolating to generate smoothly varying subframe gains over a time period or interval of a ramp) can be partially performed by an audio renderer (e.g., an object audio renderer, etc.) that may have been delegated to process the audio data of audio objects at a finer instantaneous scale than the frame level, for example, based on one or more built-in ramps (which may have been implemented by the audio renderer to handle the decoded presentation of audio objects or the motion of any audio object from one spatial location to another in audio rendering). These techniques can be implemented to generate or compute the subframe gain for each of some or all of the audio objects to be provided to the audio renderer using frame-level gains received from upstream devices. In response to the time-varying frame-level gains, subframe gain smoothing operations based on these subframe gains can be implemented as part of or incorporated into subframe audio rendering operations performed by the audio renderer.

[0099] Alternatively, the audio sample data (such as PCM audio data) representing the audio data of an audio object does not necessarily need to be decoded before applying subframe gains to the audio sample data as described herein. The audio metadata or OAMD to be input to or used by the audio renderer can be modified or generated. In other words, in some operational scenarios, these subframe gains can be generated without decoding the encoded audio data carried in the audio bitstream into audio sample data. The audio renderer can then decode the encoded audio data into audio sample data and apply the subframe gains to the audio data portion of the subframe unit within the audio sample data as part of rendering the audio object using audio speakers configured with (actual) output audio channels.

[0100] Therefore, with the techniques described herein, there is no or almost no additional computational cost. Furthermore, upstream devices (e.g., before the audio renderer) do not need to perform these subframe audio processing operations in response to time-varying frame-level gain. Thus, with the techniques described herein, repetitive and complex computations or operations at the subframe level can be avoided or significantly reduced.

[0101] 4. Subframe gain generation

[0102] In some operational scenarios, such as the audio streams described in this article (e.g., Figure 1 or Figure 2A 102, etc.) includes a set of audio objects and audio metadata of the audio objects. To generate a decoded rendering or audio rendering of the audio objects decoded from the audio stream (102), an audio renderer such as an object audio renderer (e.g., Figure 2A (e.g., 108, etc.) can be used with audio decoding devices (e.g.,Figure 2A 100, etc.) or with the use of audio decoding devices (e.g., Figure 2B Equipment that operates (e.g., 100-1, etc.) Figure 2C Integrate 100-2, etc.

[0103] The audio decoding device (100, 100-1) can set the object's audio metadata as input to the audio renderer (108) to instruct the integrated audio renderer (108) to perform audio processing operations to render the audio object. The object's audio metadata can be generated at least in part from the audio metadata received from the audio bitstream (102).

[0104] Audio objects, such as dynamic audio objects, can move within an audio rendering environment (e.g., a home, movie theater, amusement park, music bar, opera house, concert hall, bar, multiple homes, auditorium, etc.). An audio decoding device (100) can generate timing data to be input into the audio renderer (108) as part of the object's audio metadata. The timing data generated by the decoder can specify a ramp length for a built-in ramp implemented by the audio renderer (108) to handle transformations of the audio object caused by its movement, such as spatial and / or transient changes (e.g., in object gain, translation coefficients, submix / downmix coefficients, etc.).

[0105] Built-in ramps can operate at subframe instantaneous scales (e.g., down to the sample level in some operational scenarios) and smoothly transition audio objects from one place to another within an audio rendering environment. Once the audio renderer (108) or its algorithm has determined the target gain that reflects or represents the final destination of the ramp used to smooth the gain of the audio objects, the built-in ramps in the audio renderer (108) can be applied to compute or interpolate the gain on subframe units such as QMF slots, PCM samples, etc.

[0106] Compared to any ramp outside the audio renderer (108), this built-in ramp offers the significant advantage of being active in the signal path for rendering the (actual) audio of all audio objects to the (actual) output audio channel configuration that utilizes the audio renderer (108). Therefore, audible artifacts such as the "zipper" effect can be prevented or reduced quite effectively and easily by the built-in ramp implemented in the audio decoding device.

[0107] By comparison, in other methods, any ramping or interpolation process implemented at the frame level, such as at an upstream device like the audio encoding device (150), may not be based on information about the actual audio channel configuration, but may be based on an assumed reference audio channel configuration that differs from the actual audio channel configuration (or audio rendering capabilities). Therefore, audible artifacts such as the "zipper" effect may not be effectively prevented or reduced by such ramping or interpolation processes in the upstream device.

[0108] Subframe gain smoothing operations as described in this article (e.g., using built-in ramps, etc.) can be applied to various input audio content to be rendered, having different combinations of audio objects and / or audio object types. Example input audio content may include, but is not limited to, any of the following: channel content, object content, combinations of channel content and object content, etc.

[0109] For channel content represented by one or more static audio objects (or bed objects), the object audio metadata input to the audio renderer (108) may include audio metadata parameters (e.g., encoder transmission, bitstream transmission, etc.) specifying the static audio object and its associated channel ID. The spatial location of the static audio object may be given by or inferred from the channel ID specified for the static audio object.

[0110] The audio decoding device (100) can generate or regenerate audio metadata parameters and use the decoder-generated audio metadata parameters (or parameter values) to control audio rendering operations of channel content or static audio objects therein via (e.g., integrated, separate, etc.) the audio renderer (108). For example, for some or all of the static audio objects in the channel content, the audio decoding device (100) can set or generate timing control data, such as one or more ramp lengths to be used by the built-in ramps implemented in the audio renderer (108). Thus, for static audio objects corresponding to channels in the output audio channel configuration, the audio decoding device (100) can provide frame-level gain (e.g., ducking gain received in the audio bitstream (102)) and one or more decoder-generated ramp lengths in the object audio metadata input to the audio renderer (108), as well as spatial information of these static audio objects corresponding to the channel ID, to perform gain smoothing using ramps with one or more decoder-generated ramp lengths.

[0111] For example, for ducking operations associated with the “main audio” program and the “associated audio” program represented in the audio bitstream (102), the audio decoding device (100) (e.g., a subframe gain calculator (106), an audio renderer (108), a combination of processing elements in the audio decoding device (100), etc.) can calculate or generate a first set of gains and calculate or generate a second set of gains. The first set of gains includes subframe gains generated by a first decoder to be applied to a first subset of audio objects constituting the “main audio” program, and the second set of gains includes subframe gains generated by a second decoder to be applied simultaneously to a second subset of audio objects constituting the “associated audio” program. The first set of gains and the second set of gains can reflect the attenuation of the amount of ducking transmitted in the overall rendering of the “main audio” content and the corresponding enhancement of the amount of enhancement transmitted in the overall rendering of the “associated audio” content.

[0112] Figure 3A The illustration shows example gain smoothing operations for audio objects (e.g., static audio objects) that are part of the channel content. These operations can be performed at least partially by the audio renderer (108). For illustrative purposes, Figures 3A to 3D The horizontal axis represents time 200. Figures 3A to 3D The vertical axis represents the gain of 204.

[0113] The frame-level gain of a static audio object can be compared with that of the audio bitstream (e.g., Figure 1 or Figure 2A The frame-level gains are specified in the audio metadata received together with (e.g., 102). These frame-level gains may include a first frame-level gain 206-1 for a first audio frame and a second frame-level gain 206-2 for a second distinct audio frame. The first and second audio frames may be part of an audio frame sequence in the audio bitstream (102). This audio frame sequence may cover the playback duration. In one example, the first and second audio frames may be two consecutive audio frames in the audio frame sequence. In another example, the first and second audio frames may be two discontinuous audio frames separated by one or more intermediate audio frames in the audio frame sequence. The first audio frame may include a first audio data portion of a first frame time interval starting at a first playback time point 202-1, while the second audio frame may include a second audio data portion of a second frame time interval starting at a second playback time point 202-2.

[0114] The audio metadata received in the audio bitstream (102) may be unaffected by the specification, or may not carry timing control data, such as ramp length, for applying gain smoothing to the first frame level gain and the second frame level gain (206-1 and 206-2).

[0115] An audio decoding device (100), including an audio renderer (108) and / or operating using an audio renderer, can determine (e.g., based on a threshold, based on the inequality of the first and second gains, based on additional determining factors, etc.) whether subframe gain smoothing should be performed for the first and second gains. In response to determining that subframe gain smoothing should be performed for the first and second gains, the audio decoding device (100) generates timing control data for applying subframe gain smoothing to the first and second frame-level gains (206-1 and 206-2), such as the ramp length of ramp 216. Additionally, optionally, or alternatively, the audio decoding device (100) can set a final or target gain 212 at the end of ramp (216). The final or target gain (212) can, but is not limited to, be the same as the second frame-level gain (206-2).

[0116] The ramp length (216) can be specified as a (gain change / conversion) time interval in the object audio metadata input to the audio renderer (108) during which subframe gain smoothing is to be performed. The ramp length or time interval (216) can be input to or used by the audio renderer (108) to determine the final or target time point 208 representing the end of the ramp (216). The final or target time point (208) of the ramp (216) may or may not be the same as the second time point (202-2). The final or target time point (208) of the ramp (216) may or may not be aligned with the frame boundary separating two adjacent audio frames. For example, the final or target time point (208) of the ramp (216) may be aligned with a subframe unit such as a QFM slot or a PCM sample.

[0117] In response to receiving object audio metadata, the audio renderer (108) performs a gain smoothing operation to calculate or obtain the individual subframe gain on the ramp (216). For example, these individual subframe gains may include different gains (or different gain values) for different subframe units (e.g., subframe units corresponding to subframe time point 210 in the ramp (216), such as subframe gain 214.

[0118] For object content represented by one or more dynamic audio objects (e.g., non-channel content, non-bed objects, etc.), the object audio metadata input to the audio renderer (108) may include audio metadata parameters specifying one or more (e.g., encoder-sent, bitstream-transmitted, etc.) ramp lengths and time-varying frame-level gain. This is achieved by upstream audio processing devices (e.g., Figure 1Some or all of the ramp lengths specified (e.g., 150) may be important for rendering dynamic audio objects or for timing aspects of such rendering. It should be noted that in some operational scenarios, encoders supporting movie applications may not specify one or more ramp lengths for the object content. Additionally, optionally, or alternatively, in some operational scenarios, encoders supporting broadcast applications may (freely) specify one or more ramp lengths for the channel content.

[0119] In some operating scenarios, such as audio bitstreams (e.g., Figure 1 or Figure 2A The ramp length sent by the encoder, specified by the gain over time in the audio metadata (e.g., 102), can be determined by an audio renderer as described herein (e.g., Figure 2A (e.g., 108) to be used and implemented.

[0120] Figure 3B The illustration shows example gain smoothing operations for audio objects, such as dynamic audio objects within the object content. These operations can be performed at least in part by the audio renderer (108).

[0121] The frame-level gain of an audio object (e.g., static or dynamic) can be specified in the audio metadata received along with the audio bitstream (102). These frame-level gains may include a third frame-level gain 206-3 for a third audio frame and a fourth frame-level gain 206-4 for a fourth distinct audio frame. The third and fourth audio frames may be part of an audio frame sequence in the audio bitstream (102). This audio frame sequence may cover the playback duration. In one example, the third and fourth audio frames may be two consecutive audio frames in this audio frame sequence. In another example, the third and fourth audio frames may be two discontinuous audio frames separated by one or more intermediate audio frames in this audio frame sequence. The third audio frame may include a third audio data portion of a third frame time interval starting at a third playback time point 202-3, while the fourth audio frame may include a fourth audio data portion of a fourth frame time interval starting at a fourth playback time point 202-4.

[0122] The audio metadata received in the audio bitstream (102) may specify, or may carry timing control data for applying gain smoothing to the third frame level gain and the fourth frame level gain (206-3 and 206-4), such as the ramp length of ramp 216-1.

[0123] The ramp length of the ramp (216-1) (e.g., sent by the encoder, transmitted by the bitstream, etc.) can be specified as a (gain change / conversion) time interval in the object audio metadata input to the audio renderer (108) within which a subframe gain smoothing operation is to be performed. The ramp length or time interval of the ramp (216-1) can be input to or used by the audio renderer (108) to determine the final or target time point 208-1 representing the end of the ramp (216-1). Alternatively, the audio decoding device (100) can set the final or target gain 212-1 at the end of the ramp (216-1).

[0124] In response to receiving object audio metadata of a specified ramp length from a specified encoder, the audio renderer (108) performs a gain smoothing operation, for example, using the built-in ramp function, to calculate or obtain the individual subframe gain on the ramp (216-1). These individual subframe gains may include different gains (or different gain values) for different subframe units in the ramp (216-1).

[0125] In some operational scenarios, in audio bitstreams (e.g., Figure 1 or Figure 2A The ramp length sent by the encoder is specified in the audio metadata (e.g., 102, etc.) for the time-varying gain. This is not specified in the audio bitstream (e.g., Figure 1 or Figure 2A The ramp length specified by the decoder for the time-varying gain in the audio metadata (e.g., 102, etc.) can be generated as described herein by modifying the received audio metadata and by the audio renderer (e.g., Figure 2A (e.g., 108) to be used or implemented.

[0126] Figure 3C The illustration shows example gain smoothing operations for audio objects (e.g., dynamic audio objects) that are part of the object's content. These operations can be performed at least in part by the audio renderer (108).

[0127] For illustrative purposes only, Figure 3C Here, the dynamic audio object in the audio metadata received along with the audio bitstream (102) is specified as... Figure 3B The same frame-level gains are illustrated in the figure. These frame-level gains may include a third frame-level gain (206-3) for the third audio frame and a fourth frame-level gain (206-4) for the fourth audio frame. The third audio frame may correspond to a frame interval starting at the third playback time point (202-3), while the fourth audio frame may correspond to a frame interval starting at the fourth playback time point (202-4).

[0128] The audio metadata received in the audio bitstream (102) can specify different ramp lengths (e.g., encoder-sent, bitstream-transmitted, etc.) for applying gain smoothing to the third-frame-level gain and the fourth-frame-level gain (206-3 and 206-4).

[0129] An audio decoding device (100), including an audio renderer (108) and / or operating using an audio renderer, can determine (e.g., based on a threshold, based on the inequality of the first and second gains, based on additional determining factors, etc.) whether subframe gain smoothing should be performed for the third and fourth gains. In response to determining that subframe gain smoothing should be performed for the third and fourth gains, the audio decoding device (100) generates timing control data for applying subframe gain smoothing to the third and fourth frame-level gains (206-3 and 206-4), such as the (decoder-generated) ramp length of ramp 216-2. Additionally, optionally, or alternatively, the audio decoding device (100) can set a final or target gain 212-2 at the end of ramp (216-2). The final or target gain (212-2) can, but is not limited to, be the same as the fourth frame-level gain (206-4).

[0130] The ramp length of ramp (216-2) can be specified as a (gain change / conversion) time interval in the object audio metadata input to the audio renderer (108) during which subframe gain smoothing is to be performed. The ramp length or time interval of ramp (216-2) can be input to or used by the audio renderer (108) to determine the final or target time point 208-2 representing the end of ramp (216-2). The final or target time point (208-2) of ramp (216-2) may or may not be the same as the fourth time point (202-4). The final or target time point (208-2) of ramp (216-2) may or may not be aligned with the frame boundary separating two adjacent audio frames. For example, the final or target time point (208-2) of ramp (216-2) may be aligned with a subframe unit such as a QFM slot or a PCM sample.

[0131] In response to receiving object audio metadata, the audio renderer (108) performs a gain smoothing operation to calculate or obtain the individual subframe gain on the ramp (216-2). For example, these individual subframe gains may include different gains (or different gain values) for different subframe units (e.g., subframe units corresponding to subframe time points 210-2 in the ramp (216-2), such as subframe gain 214-2.

[0132] While built-in ramps can be used by the audio renderer (108) for audio objects (e.g., dynamic audio objects within object content), simply modifying the ramp length may alter the audio rendering of these audio objects for the purpose of performing gain smoothing operations (e.g., duck-related gain smoothing). Therefore, in some operational scenarios, the amount of gain smoothing corresponding to ducking can be achieved by simply integrating the duck-related gain (e.g., frame-level gain specified in the audio metadata of the audio bitstream (102)) into the overall object gain applied to the audio object by the audio renderer (108)—integrated or implemented using subframe gains interpolated or smoothed by the audio renderer (108)—this overall object gain is used to drive the audio speakers in the output audio channel configuration that operate with the audio renderer (108) in the audio rendering environment. For audio objects that do not have a ramp length for bitstream transmission (e.g., in channel content, etc.), the audio decoding device (100) can generate a ramp length to be input to the audio renderer (108) and implemented by the audio renderer, such as... Figure 3A As illustrated, for an audio object with a ramp length for bitstream transmission (e.g., in channel content, etc.), an audio decoding device (100) can input the transmitted ramp length to an audio renderer (108) for performing a subframe gain smoothing operation.

[0133] The generation and application of timing control data related to gain smoothing operations can take into account the update rates of both frame-level gain (e.g., ducking) and audio metadata received by the audio decoding device (100). For example, the ramp length as described herein can be set, generated, and / or used at least in part based on the update rates of gain information and audio metadata received by the audio decoding device (100). The ramp length may or may not be optimally determined for the audio object. However, the ramp length can be selected (e.g., as a sufficiently long time interval) to prevent or reduce the generation of audible artifacts (e.g., the "zipper" effect in ducking operations, etc.) during gain change / conversion operations.

[0134] In some operational scenarios, the gain smoothing operation described in this paper may or may not be optimal because some intermediate gains (e.g., intermediate ducking gains or values) may drop. For example, the upstream encoder may send more updates in the ramp determined by the audio decoding device. The ramp may be designed or specified using a ramp length longer than the update time of the gain sent by the encoder. Figure 3CAs illustrated, intermediate (e.g., frame-level, subframe-level, etc.) gains 218 can be received in the audio bitstream (102) to update the ducking gain of the audio object for internal time points of the ramp (216-2). In some operating scenarios, this intermediate gain (218) may decrease. The decrease in intermediate gain may or may not change the perceived quality of the ducking gain application.

[0135] In some operational scenarios, further improvements to subframe gain smoothing operations can be implemented in the audio decoding device (100) or its audio renderer (108). For example, the audio decoding device (100) can internally generate intermediate audio metadata (e.g., intermediate OAMD payloads or portions) such that all intermediate gain values ​​emitted or received in the audio bitstream (102) are applied by the audio decoding device (100) or its audio renderer (108) to obtain better gain smoothing curves (e.g., one or more linear segments, etc.). The audio decoding device (100) can generate those internal OAMD payloads or portions in such a way that audio objects, including but not limited to dynamic audio objects, are correctly rendered according to the intent of the content creator of the audio content represented by the audio objects.

[0136] For example, Figure 3C The slope (216-2) can be modified as follows: Figure 3D The illustration shows different slopes 216-3. Figure 3D The slope (216-3) can be utilized with, for example Figure 3C The same target gain (e.g., 212-2, etc.) and the same ramp length (e.g., between time points 208-2 and 202-3, etc.) are set as illustrated. However, Figure 3D The slope (216-3) and Figure 3C The difference of the ramp (216-2) is that the intermediate gain (218) received for the internal time points within the time interval covered by the ramp (216-3) is implemented or enforced by the audio decoding device (100) or the audio renderer (108) therein.

[0137] Under the techniques described herein, subframe gain smoothing of channel content and / or object content in response to time-varying gain (e.g., ducking gain) can be performed near the end of the media content delivery pipeline and can be performed by an audio renderer operating with (actual) output audio channel configuration (e.g., a set of audio speakers, etc.) to generate sound from channel content and / or object content.

[0138] This solution is not limited to any particular audio processing system (such as an AC-4 audio system), but can be applied to a variety of audio processing systems in which audio renderers, etc., at or near the end of the audio content delivery and consumption pipeline process or handle audio objects that represent channel audio and / or object audio that vary over time (or remain constant over time). Example audio processing systems implementing the techniques described herein may include, but are not limited to, those implementing one or more of the following: Dolby Digital + Joint Object Coding (DD+JOC), MPEG-H, etc.

[0139] Additionally, alternatively, or as a substitute, some or all of the techniques described herein may be implemented in an audio processing system in which an audio renderer operating with an output audio channel configuration is separate from a device that processes user input, which may be used to change object or channel characteristics (e.g., a ducking gain to be applied to received audio content in an audio bitstream).

[0140] Figure 2B and Figure 2C The illustration shows two example audio processing devices 100-1 and 100-2 that can cooperate with each other to render audio content received from an audio bitstream (e.g., 102, etc.) (or generate corresponding sounds from that audio content).

[0141] In some operational scenarios, the first audio processing device (100-1) may be a set-top box that receives an audio bitstream (102) comprising a set of audio objects and audio metadata of the audio objects. Additionally, optionally, or alternatively, the first audio processing device (100-1) may receive user input (e.g., 118, etc.) that can be used to adjust rendering aspects and / or characteristics of the audio objects. For example, the audio bitstream (102) may include a "main audio" program and an "associated audio" program to which a ducking gain specified in the audio metadata will be applied.

[0142] The first audio processing device (100-1) can adjust the audio metadata to generate new or modified audio metadata or OAMD that will be input to an audio renderer implemented by the second audio processing device (100-2). The second audio processing device (100-1) can be an audio / video receiver (AVR) that operates using an output audio channel configuration or its audio speaker to generate sound from audio data encoded in the audio bitstream (102).

[0143] In some operational scenarios, the first audio processing device may perform decoding of the audio bitstream (102) and generate a subframe gain at least in part based on a time-varying frame-level gain (e.g., ducking gain) specified in the audio metadata. The subframe gain may be included as part of the OAMD to be output by the first audio processing device (100-1) to the second audio processing device (100-2). A new or modified OAMD generated at least in part by the first audio processing device (100-1) for the audio object and its audio data received by the first audio processing device (100-1) may be encoded or included in the media signal encoder 110 of the first audio processing device (100-1) in the output audio / video signal 112 (e.g., an HDMI signal). The A / V signal (112) may be delivered or transmitted from the first audio processing device (100-1) to the second audio processing device (100-2), for example, via an HDMI connection, e.g., wirelessly, via a wired connection.

[0144] The media signal decoder 114 in the second audio processing device (100-2) receives the A / V signal (112) and decodes it into audio data of an audio object and an OAMD including subframe gains (such as those generated to duck the audio object and its audio data). The audio renderer (108) in the second audio processing device (100-2) uses the input OAMD from the first audio processing device (100-1) to perform audio rendering operations, including but not limited to applying subframe gains to the audio object and driving the audio speakers in the output audio channel configuration to generate sound depicting the sound source represented by the audio object.

[0145] For illustrative purposes only, it has been described that time-varying gain can be associated with ducking operations. It should be noted that, in various embodiments, some or all of the techniques described herein can be used to implement or perform subframe gain operations related to other audio processing operations besides ducking operations (e.g., audio processing operations related to applying dialogue enhancement gain, downmixing gain, etc.).

[0146] 5. Example Process Flow

[0147] Figure 4 The illustration shows an example process flow that can be implemented by an audio decoding device as described herein. In box 402, the downstream audio system (e.g., the audio decoding device (e.g., Figure 2A 100 Figure 2B 100-1 and Figure 2CThe 100-2 protocol decodes the audio bitstream into a set of one or more audio objects and audio metadata for that set of audio objects. The set of one or more audio objects includes specific audio objects. The audio metadata specifies a first set of frame-level gains, which includes a first gain and a second gain for the first and second audio frames in the audio bitstream, respectively.

[0148] In box 404, the downstream audio system determines, at least in part, whether to generate a subframe gain for a particular audio object based on the first gain and the second gain for the first audio frame and the second audio frame.

[0149] In box 406, the downstream audio system determines the ramp length of the ramp used to generate the subframe gain for the specific audio object in response to determining, at least in part, the subframe gain to be generated for the specific audio object based on the first gain and the second gain for the first audio frame and the second audio frame.

[0150] In box 408, the downstream audio system uses a ramp with a ramp length to generate a second set of gains, wherein the second set of gains includes subframe gains for a specific audio object.

[0151] In box 410, the downstream audio system causes a set of audio speakers operating in a specific playback environment to render a sound field represented by a set of audio objects with a second set of gains applied.

[0152] In an embodiment, the set of audio objects includes: a first subset of audio objects representing a main audio program; and a second subset of audio objects representing an associated audio program; a particular audio object is included in one of the following: the first subset of audio objects or the second subset of audio objects.

[0153] In an embodiment, the first audio frame and the second audio frame are one of the following: two consecutive audio frames in a particular audio object, or two discontinuous audio frames in a particular audio object separated by one or more intermediate audio frames in the particular audio object.

[0154] In the embodiments, the first gain and the second gain are associated with one of the following: ducking operation, dialogue enhancement operation, user-controlled gain shift operation, downmixing operation, gain smoothing operation applied to music and effects (M&E), gain smoothing operation applied to dialogue, gain smoothing operation applied to M&E and dialogue (M&E+dialogue), or other gain shift operation.

[0155] In one embodiment, the built-in ramp used to handle the spatial motion of audio objects is reused as a ramp for generating subframe gain for a particular audio object.

[0156] In an embodiment, the first audio frame includes a first audio data portion of a specific audio object, and the second audio frame includes a second audio data portion of the specific audio object, wherein the second audio data portion of the specific audio object is different from the first audio data portion of the specific object.

[0157] In this embodiment, the audio metadata is not affected by the ramp length specification.

[0158] In this embodiment, the audio metadata specifies a ramp length that the encoder sends, which is different from the ramp length.

[0159] In an embodiment, the set of gains includes intermediate gains corresponding to time points within the time interval represented by the ramp; these intermediate gains are excluded from a second set of gains to be applied to the set of audio objects in the decoded presentation.

[0160] In one embodiment, the set of gains includes intermediate gains corresponding to time points within the time interval represented by the ramp; these intermediate gains are included in a second set of gains to be applied to the set of audio objects in the decoded presentation.

[0161] In one embodiment, the set of audio objects includes a second audio object; wherein the ramp length sent by the encoder is specified in the audio metadata received along with the audio stream; the ramp length sent by the encoder is used as the ramp length for generating subframe gain for the second audio object.

[0162] In this embodiment, the second set of gains is generated by the first audio processing device; the sound field is rendered by the second audio processing device.

[0163] In this embodiment, the second set of gains is generated by interpolation.

[0164] In one embodiment, a non-transitory computer-readable storage medium includes software instructions that, when executed by one or more processors, cause to perform any of the methods described herein. Note that although individual embodiments are discussed herein, any combination of the embodiments and / or portions of the embodiments discussed herein can be combined to form further embodiments.

[0165] 6. Implementation Mechanism – Hardware Overview

[0166] According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to perform these techniques, or may include digital electronic devices (e.g., one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs)) that are persistently programmed to perform these techniques, or may include one or more general-purpose hardware processors programmed to execute the techniques according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that incorporates hardwired and / or program logic to implement the techniques.

[0167] For example, Figure 5 This is a block diagram illustrating a computer system 500 on which embodiments of the present invention may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 for processing information. The hardware processor 504 may, for example, be a general-purpose microprocessor.

[0168] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When stored in a non-transitory storage medium accessible to processor 504, such instructions cause computer system 500 to become a device-specific dedicated machine for performing the operations specified in the instructions.

[0169] Computer system 500 further includes a read-only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions of processor 504. Storage device 510 (such as a magnetic disk or optical disk) is provided and coupled to bus 502 for storing information and instructions.

[0170] Computer system 500 can be coupled to display 512, such as a liquid crystal display (LCD), via bus 502 for displaying information to the computer user. Input device 514, including alphanumeric keys and other keys, is coupled to bus 502 for transmitting information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, trackball, or cursor arrow keys, for transmitting directional information and command selections to processor 504 and for controlling cursor movement on display 512. Typically, this input device has two degrees of freedom on two axes (a first axis (e.g., x-axis) and a second axis (e.g., y-axis)), allowing the device to specify a position in a plane.

[0171] Computer system 500 may implement the techniques described herein using device-specific hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that are combined with or program the computer system 500 to be a dedicated machine. According to one embodiment, the techniques described herein are executed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium (such as storage device 510). Execution of the sequence of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0172] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, floppy hard disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, flash EPROMs, NVRAMs, any other memory chips or memory cartridges.

[0173] Storage media differ from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including conductors containing bus 502. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.

[0174] Various forms of media can involve loading one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 500 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 502. Bus 502 loads the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 before or after execution by processor 504.

[0175] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides bidirectional data communication coupled to network link 520, which connects to local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity to a corresponding type of telephone line. As another example, communication interface 518 may be a local area network (LAN) card for providing data communication connectivity to a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying streams of digital data representing various types of information.

[0176] Network link 520 typically provides data communication to other data devices via one or more networks. For example, network link 520 may provide a connection to host computer 524 or to data devices operated by Internet Service Provider (ISP) 526 via local network 522. ISP 526, in turn, provides data communication services via a global packet data communication network now commonly referred to as the “Internet” 528. Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals carrying digital data streams. Signals through various networks, as well as signals on network link 520 and through communication interface 518 (which carries digital data to and from computer system 500), are example forms of transmission media.

[0177] Computer system 500 can send messages and receive data, including program code, through one or more networks, network links 520, and communication interfaces 518. In the Internet example, server 530 can transmit application request codes through the Internet 528, ISP 526, local network 522, and communication interface 518.

[0178] The received code can be executed by processor 504 and / or stored in storage device 510 or other non-volatile memory for later execution upon receipt.

[0179] 6. Equivalents, extensions, substitutes and others

[0180] In the foregoing description, embodiments of the invention have been described with reference to numerous specific details that may vary depending on the implementation. Therefore, the sole and exclusive indicator of the invention and the applicant's inventive intent is the set of claims (including any subsequent amendments) published under this application, in the specific form of such claims. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, characteristics, features, advantages, or attributes not expressly recited in the claims shall not in any way limit the scope of such claims. Therefore, this specification and the drawings should be viewed in an illustrative rather than restrictive sense.

Claims

1. A method for processing audio, comprising: The audio bitstream is decoded into a set of one or more audio objects and audio metadata of the set of audio objects, the set of one or more audio objects including a specific audio object, the audio metadata specifying a first set of frame-level gains, the first set of frame-level gains including a first gain and a second gain for a first audio frame and a second audio frame in the audio bitstream, respectively. Whether to generate a subframe gain for the specific audio object is determined at least in part based on the first gain and the second gain for the first audio frame and the second audio frame; In response to determining, at least in part, the subframe gain to be generated for the specific audio object based on the first gain and the second gain for the first audio frame and the second audio frame: Determine the ramp length used to generate the subframe gain for the specific audio object; A second set of gains is generated using the ramp having the ramp length, wherein the second set of gains includes the subframe gains for the specific audio object; This causes a set of audio speakers operating in a specific playback environment to render a sound field represented by the set of audio objects, and the second set of gains is applied to the set of audio objects.

2. The method as described in claim 1, wherein, The set of audio objects includes: Represents the first subset of audio objects in the main audio program; and This represents a second subset of audio objects associated with a specific audio program; and The specific audio object is included in one of the following: the first subset of audio objects or the second subset of audio objects.

3. The method as described in claim 1 or 2, wherein, The first audio frame and the second audio frame are one of the following: two consecutive audio frames in the specific audio object, or two non-consecutive audio frames in the specific audio object separated by one or more intermediate audio frames in the specific audio object.

4. The method as described in claim 1 or 2, wherein, The first gain and the second gain are associated with one of the following: ducking operation, dialogue enhancement operation, user-controlled gain shift operation, downmixing operation, gain smoothing operation applied to music and effects, gain smoothing operation applied to dialogue, and gain smoothing operation applied to music and effects as well as dialogue.

5. The method as described in claim 1 or 2, wherein, The built-in ramp used to handle the spatial motion of audio objects is reused as a ramp for generating the subframe gain for the specific audio object.

6. The method of claim 2, wherein, The first gain and the second gain are ducking gains, which are used to reduce the loudness level of the first audio object subset representing the main audio program relative to the loudness level of the second audio object subset representing the associated audio program, wherein a built-in ramp for handling the spatial motion of audio objects is reused to generate subframe ducking gains for the main audio program or the associated audio program, respectively.

7. The method as described in claim 1 or 2, wherein, The first audio frame includes a first audio data portion of the specific audio object, and the second audio frame includes a second audio data portion of the specific audio object, wherein the second audio data portion of the specific audio object is different from the first audio data portion of the specific audio object.

8. The method as claimed in claim 1 or 2, wherein, The audio metadata is not affected by the specification of the ramp length.

9. The method as claimed in claim 1 or 2, wherein, The audio metadata specifies a ramp length sent by an encoder that is different from the ramp length.

10. The method as claimed in claim 1 or 2, wherein, The first set of frame-level gains includes intermediate gains corresponding to time points within the time interval represented by the ramp; wherein the intermediate gains are excluded from the second set of gains to be applied to the set of audio objects in the decoded presentation.

11. The method as claimed in claim 1 or 2, wherein, The first set of frame-level gains includes intermediate gains corresponding to time points within the time interval represented by the ramp; wherein the intermediate gains are included in the second set of gains to be applied to the set of audio objects in the decoded presentation.

12. The method as claimed in claim 1 or 2, wherein, The set of audio objects includes a second audio object; wherein the ramp length sent by the encoder is specified in the audio metadata received along with the audio stream; wherein the ramp length sent by the encoder is used as the ramp length for generating subframe gain for the second audio object.

13. The method as claimed in claim 1 or 2, wherein, The second set of gains is generated by the first audio processing device; wherein the sound field is rendered by the second audio processing device.

14. The method as claimed in claim 1 or 2, wherein, The second set of gains is generated by interpolation.

15. The method as claimed in claim 1 or 2, wherein, The determination of whether to generate a subframe gain for the specific audio object, based at least in part on the first gain and the second gain for the first audio frame and the second audio frame, includes: If the difference between the first gain and the second gain exceeds a minimum gain difference threshold, it is determined that a subframe gain should be generated for the specific audio object; and / or if the difference between the first gain and the second gain does not exceed the minimum gain difference threshold, it is determined that a subframe gain will not be generated for the specific audio object.

16. The method of claim 15, wherein, A different minimum gain difference threshold is used for positive gain changes than for negative gain changes, wherein for the positive gain change, the second gain value is greater than the first gain, and for the negative gain change, the second gain is less than the first gain.

17. The method as claimed in claim 1 or 2, wherein, The determination of whether to generate a subframe gain for the specific audio object, based at least in part on the first gain and the second gain for the first audio frame and the second audio frame, includes: If the absolute value of the rate of change between the first gain and the second gain exceeds a minimum gain rate of change threshold, then it is determined that a subframe gain should be generated for the specific audio object; and / or If the absolute value of the rate of change between the first gain and the second gain does not exceed the minimum rate of change threshold, then it is determined that no subframe gain will be generated for the specific audio object.

18. The method of claim 17, wherein, A different minimum gain rate of change threshold is used for positive rates of change than for negative rates of change.

19. An electronic device comprising one or more processors and a memory, the memory storing one or more programs including instructions that, when executed by the one or more processors, cause the electronic device to perform the method as claimed in any one of claims 1 to 18.

20. A non-transitory computer-readable storage medium comprising software instructions that, when executed by one or more processors, cause the method of any one of claims 1 to 18 to be performed.

21. A computer program product comprising one or more programs, said one or more programs including instructions that, when executed by one or more processors, cause a device to perform the method as claimed in any one of claims 1 to 18.

Citation Information

Patent Citations

  • Systems, methods, apparatus, and computer-readable media for audio object clustering

    US20140025386A1

  • Processing of time-varying metadata for lossless resampling

    WO2015006112A1