Method for object-based audio crossfading
Automated object-based audio crossfading techniques address the challenge of aligning audio and metadata in object-based audio clips by remapping channels and setting dynamic metadata transitions, resulting in efficient and artifact-free crossfading.
Patent Information
- Application Number
- PCT/US2025/038744
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-16
- Filing Date
- 2025-07-22
- Publication Date
- 2026-01-29
AI Technical Summary
Crossfading between object-based audio clips is challenging due to the need to align both audio and metadata, leading to spatial artifacts and requiring laborious manual adjustments in existing techniques.
Automated object-based audio crossfading techniques that remap channels and dynamically set metadata transition points to minimize spatial artifacts, using channel mapping and metadata-guided crossfading to create artifact-free transitions.
Reduces time and computing resources required for crossfading, enabling efficient and artifact-free transitions between object-based audio clips.
Smart Images

Figure US2025038744_29012026_PF_FP_ABST
Abstract
Description
METHOD FOR OBJECT-BASED AUDIO CROSSFADING CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from Spanish Patent Application Nos. P202430623 and P202430946, filed on 23 July 2024 and 14 November 2024, respectively, and United States Provisional Patent Application Nos.63 / 719,759 and 63 / 734,429, filed on 13 November 2024 and 16 December 2024, respectively, each of which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0002] This disclosure relates generally to audio signal processing and, more specifically, handling of object-based audio crossfades. BACKGROUND
[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.
[0004] Crossfading is an audio editing technique that creates a transition, such as a smooth transition, between two audio clips. Crossfading often includes causing a first audio clip to fade out and decrease in volume, while simultaneously causing a second audio clip to fade in and increase in volume. Crossfading avoids abrupt transitions between two audio clips.
[0005] Object-based audio keeps audio channels independent during mixing. For example, instead of mixing sounds to specific output channels, each channel is kept independent and is associated with its own metadata, such as position data, size, or movement. These independent channels are then distributed independently and may be assigned to specific outputs. Some object-based audio systems support up to, for example, 128 independent channels (including 118 audio objects) assigned to 64 speakers, or the like.
[0006] It is with respect to these and other considerations that the disclosure made herein is presented. BRIEF SUMMARY OF THE DISCLOSURE
[0007] Examples, aspects, and instances described herein provide systems and methods for automating object audio crossfades. For example, after a crossfade overlap range has beendefined for two object audio clips, each clip’s audio and metadata are analyzed for each object channel. Metadata may be static metadata or dynamic metadata that changes over time. This information may be used to, for one or more channels, remap the object channel’s order to minimize spatial artifacts resulting from the crossfade. A metadata transition point is also selected for each channel, based on the audio and metadata of each channel, which further minimizes spatial artifacts. A crossfade is then performed across the two object audio clips, where the audio is crossfaded based on the new channel mapping and dynamic metadata transitions at the selected metadata transition point. The resulting crossfaded object-based audio signals can be archived, encoded, attributed, or rendered for audition.
[0008] In some aspects, techniques described herein relate to a method for crossfading between object-based content channels, the object-based content channels including at least a first and a second object-based content channel, the method including: determining a crossfade overlap range for the object-based content channels; mapping a channel order for the object-based content channels based on metadata and audio associated with the object-based content channels; selecting a crossfade transition point for each of the object-based content channels based on the metadata and the audio associated with the object-based content channels; and crossfading the object-based content channels based on the channel order and the crossfade transition point for each of the object-based content channels.
[0009] In some aspects, techniques described herein relate to a method for crossfading between a plurality of object-based content channels, the method including: determining, with an electronic processor, a crossfade overlap range for the plurality of object-based content channels; matching, with the electronic processor, a first object-based content channel comprising audio during the crossfade overlap range with a second object-based content channel comprising silence during the crossfade overlap range; determining, with the electronic processor, a difference metric measuring differences between each of remaining object-based content channels over the crossfade overlap range; matching, with the electronic processor, the remaining object-based content channels based on minimizing the difference metric; and crossfading, with the electronic processor, the object-based content channels based on the matched object-based content channels.
[0010] In some aspects, techniques described herein relate to a method for crossfading between object-based content channels, the object-based content channels comprising at least a first object-based content channel and a second object-based content channel, the method comprising: determining, with an electronic processor, a crossfade overlap range for the first object-basedcontent channel and the second object-based content channel; selecting, with the electronic processor, a crossfade transition point for the first object-based content channel and the second object-based content channel based on the metadata and the audio associated with the first object-based content channel and the second object-based content channel; and crossfading, with the electronic processor, the first object-based content channel and the second object-based content channel based on the selected crossfade transition point, wherein crossfading includes transitioning between first metadata associated with the first object-based content channel to second metadata associated with the second object-based content channel at the selected crossfade transition point.
[0011] Various aspects of the present disclosure provide for processing of audio signals, and effect improvements in at least the technical fields of audio processing, audio encoding, audio decoding, virtual reality, object-based audio, and the like.
[0012] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.
[0013] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associate drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims. DESCRIPTION OF THE DRAWINGS
[0014] These and other more detailed and specific features of various embodiments are more fully disclosed in the following description, reference being had to the accompanying drawings, in which:
[0015] FIGS.1A, 1B, 1C and 1D illustrate different examples of a crossfading between two object-based content channels according to some aspects of the present disclosure.
[0016] FIG.2 illustrates a system for crossfading a plurality of object audio assets according to some aspects of the present disclosure.
[0017] FIG.3 illustrates example Pseudocode for matching objects according to some aspects of the present disclosure.
[0018] FIGS.4A-4B illustrate an object mapping scheme for mapping and matching object- based content channels according to some aspects of the present disclosure.
[0019] FIG.5 illustrates a block diagram of an example method for crossfading object audio assets according to some aspects of the present disclosure.
[0020] FIG.6 illustrates a block diagram of another example method for crossfading object audio assets according to some aspects of the present disclosure.
[0021] FIG.7 illustrates a block diagram of another example method for crossfading object audio assets according to some aspects of the present disclosure.
[0022] FIG.8 illustrates a block diagram of an example apparatus according to some aspects of the present disclosure. DETAILED DESCRIPTION
[0023] In the following description, numerous details are set forth, such as audio device configurations, timings, operations, and the like, in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to one skilled in the art that these specific details are merely examples and not intended to limit the scope of this application.
[0024] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0025] In digital audio content creation, crossfading two sections of audio is frequently employed to create a smooth transition between the two sections. In a traditional, channel-based workflow, crossfading is performed by applying a gain fade out to the first section, a gain fade in to the second section, and then mixing the two sections together such that the faded regions overlap. Common use cases include gapless album assembly, masking of discontinuous audio edits, and creative blending of two pieces of content.
[0026] With the introduction of object-based audio formats, such as Dolby Atmos™, crossfading between two pieces of content becomes more difficult. Unlike the traditional channel-based case, each object channel consists of both audio and metadata that each need to be resolved in the crossfaded region. Even if audio in the object-based case can be crossfaded as in the channel- based case, the separate metadata cannot be “mixed” in the same manner as the audio. For example, for the duration of the crossfade, the two object channels being crossfaded must essentially share the same metadata.
[0027] As this metadata drives rendering of the object channel, differences in metadata between the two crossfaded objects may result in metadata loss and spatial artifacts in the rendering. The severity of these artifacts is dependent on the degree of difference between the metadata in each object channel, as well as the audio contained within that object channel in the crossfade region.
[0028] To overcome this problem using known techniques, mixing and mastering engineers rely on either carefully planning and constructing the layouts of the object audio clips to be crossfaded or rely on performing laborious track re-ordering and edits in a digital audio workstation to minimize spatial artifacts. However, known techniques are time consuming and waste computing resources by requiring numerous trials and edits.
[0029] Examples, aspects, and instances described herein provide for object-based audio crossfading techniques that handle both the independent object channels and their respective metadata and address issues with known techniques in audio processing technology as previously noted. Object-based audio crossfading techniques described herein may consist of two parts. First, audio and metadata-guided channel remapping may be performed. Next, the metadata transition point may be dynamically set for each channel, which minimizes spatial artifacts in the final output rendered audio. In many situations, a completely artifact free, or nearly-completely artifact free, “perfect” crossfade may be created for all object audio channels using the two parts separately or in conjunction. Additionally, the techniques described herein provide an automated, novel approach that reduces time and computing resources required by mixing engineers for crossfading between two object-based audio clips.
[0030] A given channel may be considered artifact free, with respect to a given rendered output configuration, if the following is true: the rendered output of the crossfaded object-audio channel is perceptually identical to an equivalent traditional crossfade of the rendered input object audio channels. For a given channel, an artifact free output may be provided in the following cases, which may be maximized during processing: (i) audio is only present on one side of the crossfade in the channel and relevant static per-channel metadata matches exactly, and (ii) dynamic metadata matches exactly for the duration of the crossfade range and relevant static per- channel metadata matches exactly.
[0031] These are so-called “perfect” cases, and, in many instances, it is possible to map the channels of two audio-object clips such that one of these conditions is met for each pair. When this is not possible, a new channel can be added that does meet these conditions, assuming a maximum channel limit for the object audio format has not been reached. When the channel limit has been reached or when adding an additional channel is otherwise undesirable, an optimal possible mapping between the available channels may be determined. This mapping may be based on a difference metric, which is evaluated between possible channel pairs.
[0032] As described herein, audio objects, channels, and object-based channel content may each refer to an audio object and may be referenced interchangeably. An object-audio asset may henceforth refer to a clip of audio that includes multiple audio objects. As used herein, the term "audio object" may refer to a stream of audio signals and associated metadata. The metadata may indicate at least the position and apparent size of the audio object. However, the metadata also may indicate rendering constraint data, content type data (e.g. dialog, effects, etc.), gain data, trajectory data, etc. Some audio objects may be static, whereas others may have time- varying metadata: such audio objects may move, may change size and / or may have other properties that change over time. When audio objects are monitored or played back in a reproduction environment, the audio objects may be rendered according to at least the position and size metadata. The rendering process may involve computing a set of audio object gain values for each channel of a set of output channels. Each output channel may correspond to one or more reproduction speakers of the reproduction environment.
[0033] FIGS.1A, 1B, 1C and 1D illustrate different example embodiments of a crossfading between two object-based content channels, including a first object-based content channel 100 (e.g., 100A, 100B, 100C, 100D) and a second object-based content channel 102 (e.g., 102A, 102B, 102C, 102D). Generally, each object-based content channel may include both audio and metadata (e.g., dynamic metadata, static metadata, static per-channel metadata, etc.). Acrossfading of the audio and metadata occurs during a crossfade overlap range 106, where the audio is crossfaded across a portion of the whole range of the crossfade overlap range 106. Metadata is switched from the first object-based content channel 100 to the second object-based content channel 102 (or vice versa) at a crossfade transition point 104 (e.g., 104A, 104B, 104C, 104D). The crossfade transition point 104 is contained within the crossfade overlap range 106.
[0034] FIG.1A illustrates a crossfading embodiment where the first object-based content channel 100A contains both audio and metadata and the second object-based content channel 102A contains metadata and no audio (e.g., or audio outside of the perceptual range such as silent audio, which is considered “no audio” as this phrase is used in the present disclosure), either for the duration of playback or during one or more segments of time (e.g., during the crossfade overlap range 106). The crossfading between the audio of the first object-based content channel 100A and the second object-based content channel 102A occurs over the crossfade overlap range 106. The dynamic metadata of the resulting crossfaded object-based content channel transitions from the dynamic metadata of the first object-based content channel 100A to the dynamic metadata of the second object-based content channel 102A at crossfade transition point 104A. The crossfade transition point 104A may be determined as described with respect to FIG.2, FIG.5, and / or FIG.7. For example, the crossfade transition point 104A may be in the middle of the crossfade overlap range 106.
[0035] FIG.1B illustrates a crossfading embodiment similar to FIG.1A, but where crossfade transition point 104B is at / near the end of the crossfade overlap range 106 to capture the dynamic metadata of the audio of the first object-based content channel 100B for the entirety of the crossfade overlap range in the resulting crossfaded object-based channel output.
[0036] FIG.1C illustrates a crossfading embodiment where the first object-based content channel 100C contains metadata and no audio, either for the duration of playback or during one or more segments of time (e.g., during the crossfade overlap range 106). The second object- based content channel 102C contains both metadata and audio. The crossfading between the audio of the second object-based content channel 102C and the first object-based content channel 100C occurs over the crossfade overlap range 106, and the dynamic metadata of the resulting crossfaded object-based audio channel transitions from the dynamic metadata of the first object- based content channel 100C to the dynamic metadata of the second object-based content channel 102C at crossfade transition point 104C. The crossfade transition point 104C may be determined as described with respect to FIG.2, FIG.5, and / or FIG.7. For example, the crossfade transition point 104C may be at the beginning of the crossfade overlap range 106 to capture the dynamicmetadata of the audio of the second object-based content channel 102C for the entirety of the crossfade overlap range in the resulting crossfaded object-based channel output.
[0037] FIG.1D illustrates a crossfading embodiment where both the first object-based content channel 100D and the second object-based content channel 102D contain audio and metadata. The crossfading between the audio of the first object-based content channel 100D and the second object-based content channel 102D occurs over the crossfade overlap range 106. The dynamic metadata of the resulting crossfaded object-based audio channel transitions from the dynamic metadata of the first object-based content channel 100D to the dynamic metadata of the second object-based content channel 102D at crossfade transition point 104D. The crossfade transition point 104A-104D may be determined as described with respect to FIG.2, FIG.5, and / or FIG.7. In some embodiments, the crossfade transition point 104A-104D may be determined based on minimizing a difference metric associated with the two channels.
[0038] FIG.2 illustrates a system 200 for crossfading a plurality of object audio assets according to some aspects of the present disclosure. The object audio assets in the example of FIG.2 comprise both audio and metadata. In the example of FIG.2, the system 200 includes a first object audio asset 202 and a second object audio asset 204 connected to a first channel analysis module 206 and a second channel analysis module 208, respectively. The system 200 also include a processing module 210 and a crossfading module 212. The crossfading module 212 generates a crossfaded asset 214.
[0039] The first channel analysis module 206 receives the first object audio asset 202 and analyzes the first object audio asset 202 to identify audio and metadata (e.g., first audio and first metadata) associated with each channel (e.g., each object channel) of the first object audio asset 202 for further analysis by the processing module 210. The second channel analysis module 208 receives the second object audio asset 204 and analyzes the second object audio asset 204 to identify audio and metadata (e.g., second audio and second metadata) associated with each channel of the second object audio asset 204 for further analysis by the processing module 210. Accordingly, the processing module 210 receives audio and metadata associated with each object channel of the first object audio asset 202 and audio and metadata associated with each object channel of the second object audio asset 204.
[0040] The processing module 210 includes a difference calculation module 216, a channel mapping module 218, and a metadata transition module 220. The difference calculation module 216 determines (e.g., calculates or estimates) a difference metric between an object channel of the first object audio asset 202 and an object channel of the second object audio asset 204. Thedifference metric is a measure of the difference between two object channels, such as, for example, over the time span of the crossfade. The difference metric may be calculated based on a weighted combination of several characteristics of the object channels. For example, the difference metric may be based on differences in pre-rendered object audio contained in the object channels. The pre-rendered object audio may be faded out and in on the first object audio asset 202 and the second object audio asset 204, respectively, to simulate the energy present in the desired crossfade.
[0041] In another example, the difference metric may be based on dynamic metadata. Dynamic metadata as described herein refers to rendering parameters that can change over time and may include spatial position (XYZ), size (for example, width, height, depth, object spread), snap (for example, channel lock, object snap), zone mask (for example, zone exclusion, object zone), and the like. In yet another example, the difference metric may represent the difference in any static, per-channel metadata between the object channels (for example, headphone render modes). The difference metric may be some metadata that indicates the perceived spatial distance from a listener to a rendered object. Static metadata as described herein refers to rendering parameters that remain fixed over time. Per-channel metadata as described herein refers to metadata that may vary from channel to channel.
[0042] In a further example, the difference metric may be based on differences in per-object rendered output between the object channels. The per-object rendered output may be faded out and in on the first object audio asset 202 and the second object audio asset 204, respectively, to simulate the energy present in the desired crossfade. Any rendering layout or combination of layouts may be implemented, including speaker layouts (for example, stereo, 5.1, 7.1, 7.0.4, 7.1.4, etc.) and headphone layouts (for example, binaural rendering with a given head related transfer function [HRTF]). In another example, the difference metric may be based on differences in per-object object-to-speaker gains, which may be calculated as a first step during the rendering process. The difference metric may also be based on higher level audio features or metrics derived from any of the prior-noted characteristics of the object channels.
[0043] The difference metric may account for the audio energy on each side of the crossfade (either pre-render or post-render), differences in dynamic metadata (derived either directly from the channel’s metadata or indirectly from per-channel rendered output), static per-channel metadata, or a combination thereof. To account for objects that are dynamically moving during the crossfade, the difference metric may be framed over the duration of crossfade and either a maximum, average, or sum of differences taken over those frames.
[0044] In some instances, the difference metric is only calculated by the difference calculation module 216 when the channel limit has been reached or when adding an additional channel is otherwise undesirable. For example, suppose that the system 200 is implemented for two songs, Song A and Song B, where crossfading is occurring from Song A to Song B. In some instances, the difference calculation module 216 (or, in other instances, the processing module 210) only determines the difference metric when a specific condition is satisfied, such as if there are no sufficient silent objects in Song A. Suppose, for example, there are ^^and ^^number of non- silent objects in Song A and Song B, respectively. In this example, crossfading two non-silent objects will happen if and only if: ^^^^^^ − ^^ < ^^ Equation (1)where ^^^^^^ − ^^ is the equal to the number of silent objects in Song A. When the condition ofEquation (1) is satisfied, ^^ − (^^^^^^ − ^^) non-silent objects in Song B map to non-silentobjects in Song A. Meanwhile, the(^^^^^^ − ^^) objects in Song B map to the silentobjects in Song A.
[0045] The difference calculation module 216 may calculate the metadata difference for all non- silent object pairs (e.g., each non-silent object from Song A mapped to each non-silent object from Song B). For example, suppose there is only one metadata transition for both Song A and Song B. For two objects i and j belonging to Song A and Song B, respectively, the overallmetadata difference may be measured by the overall distance, denoted by ^^^^(^, ^). The overalldistance may be defined as a weighted sum of several terms, such as a combination of any of the terms provided by the weighted sum of Equation 2: ^^^^(^, ^) = ^^^^^^^^(^, ^) + ^^^^^^^^(^, ^) + ^^^^^^^^(^, ^)+^^^^^^^^(^, ^) + ^^^^^^^^(^, ^) Equation (2)where ^^^^is a weighting value for the difference between the position ^^^^of objects i and j, ^^^^is a weighting value for the difference between the size ^^^^of objects i and j, ^^^^is a weighting value for the difference between a zone mask ^^^^of objects i and j, ^^^^is a weighting value for the difference between a snap ^^^^of objects i and j, and ^^^is a weighting value for the difference between the headphone rendering mode ^^^^of objects i and j.
[0046] In some instances, the four headphone rendering modes (HRMs) (mid (default), near, far, and bypass (off)) may be categorized into two groups: (i) near / mid / far and (ii) bypass, according to how the headphone render consumes these metadata. The difference between anynear / mid / far objects is smaller than the difference between bypass objects and near / mid / farobjects. Therefore, the HRM distance ^^^^(^, ^) may be further divided into two terms withdifferent weights.
[0047] Additionally, metadata clean-up may be initially applied to simplify the metadata. Potential outputs of the metadata clean-up may be snap overrides size, snap is kept only if the object is proscenium, and zone-mask possibly being set to zero.
[0048] As another example, multiple metadata transitions (e.g., multiple metadata updates) may exist during the crossfade region. In such an instance, the whole crossfade region may be divided into several chunks, where each chunk contains only one metadata transition. If there are two or more metadata transitions or updates within an audio block or chunk, metadata resampling may be applied to reduce the metadata to one update per chunk. The metadata resampling aligns the metadata transitions between the two channels being compared. In a chunk k, the overall distance for the chunk, denoted by ^[^]^^^, may be calculated via the Equation(2). The overall distance for the whole region may be obtained by ∑ [^]^ ^^^^ .
[0049] In another instance, the difference calculation module 216the difference metric based on the energy difference between the first object audio asset 202 and the second object audio asset 204. For example, the energy of the two objects during the crossfade region are denoted as "^and "#, respectively. The energy difference can be represented by their product "^"#. In a further example, a combined difference may be calculated by the product of the difference and the energy difference, as provided by Equation (3): ^(^, ^) = $"^"#%^^^^(^, ^) Equation (3)In yetlinear function &("^"#) may replace "^"#to control the contribution amount for the energy aspect.
[0050] The channel mapping module 218 receives the difference metric from the difference calculation module 216. The channel mapping module 218 is configured to map (or re-map) the content channels from the first object audio asset 202 to the content channels of the second object audio asset 204 for the crossfading. For example, the channel mapping module 218 may prioritize mapping channels containing audio with channels that are fully silent for the duration of the crossfade range. When all channels are capable of being mapped in this manner or by adding additional silent channels, no further calculations are necessary. However, there may be a limited number of silent channels and a limit on the number of channels that may be added (forexample, a limit based on a maximum file size). As one example, to limit the file size, a user may set a maximum number of the final output channels. Accordingly, in cases where it is not possible to map all channels with audio to silent channels, the channel mapping module 218 maps the content channels from the first object audio asset 202 to the content channels of the second object audio asset 204 using the difference metric.
[0051] In some instances, the channel mapping module 218 implements a difference matrix to determine how to match non-silent content channels. For example, the channel mapping module 218 may construct a difference matrix, denoted by D, with the calculated combined difference for all non-silent object pairs, as provided by Equation (4): ^(1,1) ⋯ ^(1, ^^)' - Equation (4)
[0052] object matching using a “global optimized” approach, a “worst case first” approach, or similar max-min approaches. In the global optimized approach, the channel mapping module 218 determines the best matching solution for the object pairs by minimizing the overall difference (for example, using the Hungarian algorithm).
[0053] In the “worst case first” approach, the channel mapping module 218 aims to prevent the worst case where the two objects with a large difference (as indicated by the difference matrix or the respective difference metric) are mapped. For example, the channel mapping module 218 may identify the object in Song B that would cause the worst case and map it to the object in Song A with minimum cost. In some instances, to obtain the worst object j*, the channel mapping module 218 calculates the minimum distance for the given column j, i.e., identifies the minimum value for every column, as provided by Equation (5): .# = m∀i^n ^(^, ^) ^ = 1, … , ^^ Equation (5)and picks the object j* with the maximum .#, as provided by Equation (6): ^∗ = argmax .# Equation (6)#
[0054] When silent objects are available in Song A, the channel mapping module 218 maps the object j* to the silent object. Otherwise, the channel mapping module 218 maps object j* to the object i* that results in .#. After the mapping, the channel mapping module 218 may remove the column j* (and the corresponding row i* when there are no remaining silent objects in Song A)from the difference matrix D and repeat the process until all nBobjects are mapped to objects in Song A.
[0055] As another example, consider a crossfade segment comprising a Song A and a Song B, where Song A has a fade-out function and Song B has a fade-in function. For a given object of either Song A or Song B, the object’s fade region may be rendered into a 7.0.4 speaker layout having 11 channels, and each channel is multiplied by the corresponding fade-in or fade-out function. FIG.3 illustrates an example pseudocode for matching the objects in Song A with objects in Song B. The variable lAi in the pseudocode of FIG.3 is defined as the root-mean- square (RMS) of the rendered and faded audio of a speaker i of the Song A, where the RMS is implemented to approximate the perceived volume or perceived loudness of a speaker (e.g., the loudness of speaker i in Song A). The RMS of the rendered and faded audio may determine continuous power of the audio or sound intensity of the audio over time. The variable lBi is defined as the RMS of the rendered and faded audio of speaker i of the Song B. The variable LAis the sum of the loudness of each object in Song A (e.g., 9^ = ∑::^;: .^^ ). The variable LB is thesum of the loudness of each object in Song B (e.g., 9^ =. τAi is the relative loudness ofspeaker i in Song A (e.g., < ^^^ = =>?=). The variable distdistance between two objects a,b (e.g., ^^@A(B, C) = ∑::^;: ^^ ∙ 9^ ∙ 9^ ∙ (<^^ − <^^)E ), where wi is a pre-determined weight givento each speaker. A wimay be, for example, 2 for the left (L), center (C), and right (R)speakers and 1 for other speakers.
[0056] In the example pseudocode of FIG.3, the channel mapping module 218 computes the speakers’ loudness lAi and lBi. The channel mapping module 218 creates a list with all objects in Song A and all objects in Song B, sorts the list of objects in Song A by increasing values of LA, and sorts the list of objects in Song B by decreasing values of LB. The channel mapping module 218 may then loop through each object in Song A and assign the object in Song A to an object in Song B with the minimum distance dist(a,b) (e.g., the object that minimizes the distance function).
[0057] FIGS.4A-4B illustrate an object mapping scheme for mapping and matching object- based content channels according to some aspects of the present disclosure. In the example of FIGS.4A-4B, a shaded box represents an object-based content channel that contains both audio and metadata, and an unshaded box represents an object-based content channel that does not contain audio or contains silent audio.
[0058] FIG.4A illustrates an object audio asset A (for example, the first object audio asset 202) that includes a first plurality of object-based content channels 402 (for example, eight object- based content channels) and an object audio asset B (for example, the second object audio asset 204) that includes a second plurality of object-based content channels 404. The arrows between the first plurality of object-based content channels 402 and the second plurality of object-based content channels 404 illustrate a mapping scheme 406 that maps (or matches) each channel in object audio asset A with a corresponding channel in object audio asset B.
[0059] FIG.4B illustrates the resulting object audio asset after crossfading, where the first plurality of object-based content channels 402 in object audio asset A have been crossfaded with the second plurality of object-based content channels 404 in object audio asset B based on the mapping shown in FIG.4A (and as determined by the channel mapping module 218).
[0060] Returning to FIG.2, after the channel mapping is defined by the channel mapping module 218, the metadata transition module 220 determines a metadata transition point for each channel. The metadata transition point is a point in time in the samples where metadata transitions from the first object audio asset 202 to the second object audio asset 204. The metadata transition point may be adjusted by the metadata transition module 220 per channel to minimize spatial artifacts, based on the audio energy (e.g., the perceived loudness) on each side of the crossfade (e.g., the energy of the first object audio asset 202 and the energy of the second object audio asset 204).
[0061] As one example, let d be the distance of the crossfade segment. The metadata transition point c from the start of the crossfade segment may be set according to Equation (6): F= ?=?=G?H ∙ ^ ^& (9^ + 9^) > 0Equation (6)
[0062] A or Song B) may retain the metadata longer thanthe quieter clip. In particular, when B is silent and A is not silent, the metadata transitions at the end of the crossfade. When A is silent and B is not, the metadata transitions at the beginning of the crossfade. When both clips are similar in loudness, the metadata transitions at the middle of the crossfade segment. As the previously-described dist function maps two clips having similar speaker rendering, the metadata transition is smooth. In some instances, the metadata transition point is selected within the distance d of the crossfade segment and weighted by the amount of audio energy present on each side of the crossfade. The amount of audio energy may be deriveddirectly from the pre-rendered object channels, or from the rendered output of each object channel. In either case, a fade may be applied before analysis to simulate the audio contained in the final crossfaded object audio presentation.
[0063] In some instances, the transition of metadata associated with the first object audio asset 202 to metadata associated with the second object audio asset 204 occurs over a timespan (for example, in instances where channel mapping and metadata transitions are abrupt and jarring). In such instances, the metadata transition point is instead a metadata transition timespan during which transition occurs smoothly around some point. The metadata transition timespan may be implemented by resampling the metadata in this range to a common framing and then crossfading between each available continuous parameter. In some implementations, dynamic metadata parameters with a float range are transitioned over a metadata transition timespan (for example, x / y / z position and size), while parameters without a continuous range are transitioned at the metadata transition point (for example, transitioned at the middle of the metadata transition timespan).
[0064] The crossfading module 212 includes a first remapping module 222 and a metadata transition module 226 associated with the first object audio asset 202. The crossfading module 212 also includes a second remapping module 224 and a metadata transition module 228 associated with the second object audio asset 204. The first remapping module 222 and the second remapping module 224 remap the channels of the first object audio asset 202 and the second object audio asset 204 as indicated by the channel mapping module 218. The metadata transition module 226 and the metadata transition module 228 apply the metadata transitioning as indicated by the metadata transition module 220. Crossfading is performed (at block 230) to generate the crossfaded asset 214. The crossfaded asset 214 may be stored in a memory (for example, the memory 820 of FIG.8). In another instance, the crossfaded asset 214 may be output via an output device, such as one or more speakers.
[0065] The illustrated blocks of modules in the system 200 of FIG.2 are merely examples for crossfading object audio assets. In other examples, the system 200 may include additional blocks, may omit blocks, may combine the functions of blocks, or may divide portions of the blocks into additional blocks.
[0066] FIG.5 illustrates a block diagram of an example method 500 for crossfading object audio assets, which may be performed by the system 200 of FIG.2. The method 500 may be performed by an electronic processor (for example, the electronic processor 810 of FIG.8), which may be configured to perform method 500 via execution of machine-executableinstructions. The various process blocks illustrated in FIG.5 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure.
[0067] At block 502, the method 500 includes determining a crossfade overlap range for the object-based content channels. For example, the processing module 210 may determine a range of the crossfade overlap range 106 for the first object audio asset 202 and the second object audio asset 204.
[0068] At block 504, the method 500 includes mapping a channel order for the object-based content channels based on metadata and audio associated with the object-based content channels. For example, the channel mapping module 218 may map the content channels from the first object audio asset 202 to the content channels of the second object audio asset 204, as previously described with respect to FIG.2 and FIGS.4A-4B.
[0069] At block 506, the method 500 includes selecting a crossfade transition point for each of the object-based content channels based on the metadata and the audio associated with the object-based content channels. For example, the metadata transition module 220 determines a metadata transition point for each channel, as previously described with respect to FIG.2.
[0070] At block 508, the method 500 includes crossfading the object-based content channels based on the channel order and the crossfade transition point for each of the object-based content channels. For example, the processing module 210 crossfades the first object audio asset 202 and the second object audio asset 204 to generate the crossfaded asset 214, as previously described with respect to FIG.2.
[0071] FIG.6 illustrates a block diagram of an example method 600 for crossfading object audio assets, which may be performed by the system 200 of FIG.2. The method 600 may be performed by an electronic processor (for example, the electronic processor 810 of FIG.8), which may be configured to perform method 600 via execution of machine-executable instructions. The various process blocks illustrated in FIG.6 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure.
[0072] At block 602, the method 600 includes determining a crossfade overlap range for the plurality of object-based content channels. For example, the processing module 210 may determine a range of the crossfade overlap range 106 for the first object audio asset 202 and the second object audio asset 204.
[0073] At block 604, the method 600 includes matching a first object-based content channel comprising audio during the crossfade overlap range with a second object-based content channel comprising silence during the crossfade overlap range. For example, the channel mapping module 218 maps a first object-based content channel that includes audio with a second object- based content channel that is silent during the crossfade overlap range 106, as previously described with respect to FIG.2. The first object-based content channel may be associated with the first object audio asset 202, and the second object-based content channel may be associated with the second object audio asset 204.
[0074] At block 606, the method 600 includes determining a difference metric measuring differences between each of the remaining object-based content channels over the crossfade overlap range. For example, the difference calculation module 216 determines (e.g., calculates or estimates) a difference metric between the first object-based content channel of the first object audio asset 202 and the second object-based content channel of the second object audio asset 204, as previously described with respect to FIG.2.
[0075] At block 608, the method 600 includes matching the remaining object-based content channels based on minimizing the difference metric. For example, after mapping object-based content channels that include audio to object-based content channels that are silent, the channel mapping module 218 matches the remaining object-based content channels by minimizing the difference metric determined by the difference calculation module 216, as previously described with respect to FIG.2.
[0076] At block 610, the method 600 includes crossfading the object-based content channels based on the matched object-based content channels. For example, the crossfading module 212 (implementing the first remapping module 222 and the second remapping module 224) crossfades the first object audio asset 202 and the second object audio asset 204 to generate the crossfaded asset 214.
[0077] FIG.7 illustrates a block diagram of an example method 700 for crossfading object audio assets, which may be performed by the system 200 of FIG.2. The method 700 may be performed by an electronic processor (for example, the electronic processor 810 of FIG.8), which may be configured to perform method 700 via execution of machine-executable instructions. The various process blocks illustrated in FIG.7 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure.
[0078] At block 702, the method 700 includes determining a crossfade overlap range for the first object-based content channel and the second object-based content channel. For example, the processing module 210 may determine a range of the crossfade overlap range 106 for the first object audio asset 202 and the second object audio asset 204.
[0079] At block 704, the method 700 includes selecting a crossfade transition point for the first object-based content channel and the second object-based content channel based on the metadata and the audio associated with the first object-based content channel and the second object-based content channel. For example, the metadata transition module 220 may determine a metadata transition point for each channel.
[0080] At block 706, the method 700 includes crossfading the first object-based content channel and the second object-based content channel based on the selected crossfade transition point. For example, the crossfading module 212 (implementing the metadata transition module 226 and the metadata transition module 228) crossfades the first object audio asset 202 and the second object audio asset 204 to generate the crossfaded asset 214.
[0081] Examples, aspects, and instances described herein may relate to an apparatus (e.g., computer-implemented apparatus or apparatus having processing capability in general) for performing or implementing methods and techniques described throughout the present disclosure. For example, the apparatus may be an audio renderer or a mixer.
[0082] FIG.8 illustrates a block diagram of an example apparatus 800. In particular, apparatus 800 includes an electronic processor 810 and a memory 820 coupled to the electronic processor 810. The memory 820 may store instructions for the electronic processor 810. The electronic processor 810 may also receive, among others, suitable input data 830 (e.g., input frames of time-frequency coefficients, etc.), depending on use cases and / or implementations. The electronic processor 810 may be adapted to carry out or implement the methods / techniques described throughout the present disclosure and to generate corresponding output data 840, depending on use cases and / or implementations.
[0083] In some examples, the memory 820 may be located internal to the electronic processor 810, such as for an internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, memory 820 may be located external to the electronic processor 810, such as a ROM, a RAM, flash memory or a removable medium, or another non-transitory computer readable medium. The memory 820 may store instructions implemented by theelectronic processor 810 to perform the methods described throughout the present disclosure (for example, the method 500, the method 600, and the method 700).
[0084] The present disclosure likewise relates to corresponding computer programs, computer program products, and computer-readable storage media storing such computer programs or computer program products. Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0085] Aspects of the methods and apparatus / systems described herein may be implemented in an appropriate computer-based audio processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of the audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0086] One or more of the components, blocks, processes or other functional components (modules) may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer- readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
[0087] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented insoftware (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, the apparatus (e.g., encoders) described above can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0088] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0089] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, andcouplings.
[0090] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims. Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.
[0091] EEE1. A method for crossfading between object-based content channels, the object-based content channels comprising at least a first object-based content channel and a second object- based content channel, the method comprising: determining, with an electronic processor, a crossfade overlap range for the object-based content channels; mapping, with the electronic processor, a channel order for the object-based content channels based on metadata and audio associated with the object-based content channels; selecting, with the electronic processor, acrossfade transition point for each of the object-based content channels based on the metadata and the audio associated with the object-based content channels; and crossfading, with the electronic processor, the object-based content channels based on the channel order and the crossfade transition point for each of the object-based content channels.
[0092] EEE2. The method according to EEE1, wherein the crossfading further comprises: transitioning, with the electronic processor, between first dynamic metadata associated with the first object-based content channel to second dynamic metadata associated with the second object- based content channel at the crossfade transition point for the first object-based content channel.
[0093] EEE3. The method according to EEE2, wherein the transition between the first dynamic metadata and the second dynamic metadata takes place over a range of time centered around the crossfade transition point for the first object-based content channel.
[0094] EEE4. The method according to any one of EEE1 to EEE3, wherein mapping the channel order further comprises: determining, with the electronic processor, a difference metric measuring the differences between each of the object-based content channels over the crossfade overlap range.
[0095] EEE5. The method according to any one of EEE1 to EEE4, wherein mapping the channel order further comprises: matching, with the electronic processor, the first object-based content channel with the second object-based content channel, wherein the first object-based content channel comprises audio during the crossfade overlap range, and the second object-based content channel is silent during the crossfade overlap range.
[0096] EEE6. The method according to EEE4, wherein mapping the channel order further comprises: matching, with the electronic processor, the first object-based content channel with the second object-based content channel based on minimizing the difference metric, wherein the first and second object-based content channels comprise audio during the crossfade overlap range.
[0097] EEE7. The method according to any one of EEE1 to EEE6, wherein mapping the channel order further comprises creating a third object-based content channel of the object-based content channels.
[0098] EEE8. The method according to any one of EEE4 to EEE7, wherein the determining the difference metric comprises at least one of: determining differences in pre-rendered object audio comprised in the object-based content channels; determining differences based on dynamicmetadata associated with the object-based content channels; determining differences in static, per-channel metadata between the object-based content channels; determining differences in per- object rendered output between the object-based content channels; determining differences in per-object object-to-speaker gains of the object-based content channels; and determining differences in audio features of the object-based content channels.
[0099] EEE9. The method according to any one of EEE4 to EEE8, wherein mapping the channel order further comprises minimizing the difference metric.
[0100] EEE10. The method according to any one of EEE1 to EEE9, wherein the crossfade transition point for each of the object-based content channels is comprised within the crossfade overlap range.
[0100] EEE11. The method according to any one of EEE1 to EEE10, wherein the crossfade transition point for each of the object-based content channels comprises a time in samples where dynamic metadata transitions from the first object-based content channel to the second object- based content channel.
[0101] EEE12. The method according to any one of EEE1 to EEE11, wherein selecting the crossfade transition point for each of the object-based content channels comprises minimizing spatial artifacts associated with the object-based content channels.
[0102] EEE13. The method according to any one of EEE1 to EEE12, wherein the crossfading further comprises: crossfading the object-based content channels based on the channel order and the crossfade transition point for each of the object-based content channels to output a crossfaded object-based content channel with dynamic metadata.
[0103] EEE14. The method according to EEE13, wherein the crossfaded object-based content channel is perceptually identical to an equivalent traditional crossfade of the object-based content channels.
[0104] EEE15. The method according to any one of EEE1 to EEE14, wherein the metadata associated with the object-based content channels comprises at least one of dynamic metadata, static metadata, and / or per-channel static metadata.
[0105] EEE16. The method according to EEE15, wherein when the metadata comprises dynamic metadata, the dynamic metadata comprising at least one of rendering parameters that change over time, a spatial position, a headphone rendering mode, a headphone rendering metadata, a size, a snap, and a zone mask.
[0106] EEE17. The method according to any one of EEE15 to EEE16, wherein when the metadata comprises static metadata, the static metadata comprising of rendering parameters that remain fixed over time.
[0107] EEE18. The method according to any one of EEE15 to EEE17, wherein when the metadata comprises per-channel static metadata, the per-channel static metadata comprising metadata that varies between channels of the object-based content channels.
[0108] EEE19. A method for crossfading between a plurality of object-based content channels, the method comprising: determining, with an electronic processor, a crossfade overlap range for the plurality of object-based content channels; matching, with the electronic processor, a first object-based content channel comprising audio during the crossfade overlap range with a second object-based content channel comprising silence during the crossfade overlap range; determining, with the electronic processor, a difference metric measuring differences between each of remaining object-based content channels over the crossfade overlap range; matching, with the electronic processor, the remaining object-based content channels based on minimizing the difference metric; and crossfading, with the electronic processor, the object-based content channels based on the matched object-based content channels.
[0109] EEE20. The method according to EEE19, wherein the determining the difference metric comprises at least one of: determining differences in pre-rendered object audio contained in the object-based content channels; determining differences based on dynamic metadata associated with the object-based content channels; determining differences in static, pre-channel metadata between the object-based content channels; determining differences in per-object rendered output between the object-based content channels; determining differences in per-object object-to- speaker gains of the object-based content channels; and determining differences in audio features of the object-based content channels.
[0110] EEE21. The method according to EEE19, wherein determining the difference metric includes calculating a matrix of differences between non-silent channel pairs included in the remaining object-based content channels.
[0111] EEE22. The method according to EEE19, wherein determining the difference metric includes determining a difference in energy between the remaining object-based content channels during a crossfade region.
[0112] EEE23. A method for crossfading between object-based content channels, the object- based content channels comprising at least a first object-based content channel and a secondobject-based content channel, the method comprising: determining, with an electronic processor, a crossfade overlap range for the first object-based content channel and the second object-based content channel; selecting, with the electronic processor, a crossfade transition point for the first object-based content channel and the second object-based content channel based on first metadata and first audio associated with the first object-based content channel and second metadata and second audio associated with the second object-based content channel; and crossfading, with the electronic processor, the first object-based content channel and the second object-based content channel based on the selected crossfade transition point, wherein crossfading includes transitioning between the first metadata associated with the first object- based content channel to the second metadata associated with the second object-based content channel at the selected crossfade transition point.
[0113] EEE24. The method according to EEE23, wherein the transition between the first metadata and the second metadata takes place over a range of time centered around the selected crossfade transition point.
[0114] EEE25. The method according to any one of EEE23 to EEE24, wherein selecting the crossfade transition point comprises minimizing spatial artifacts associated with the first object- based content channel and the second object-based content channel.
[0115] EEE26. The method according to any one of EEE23 to EEE25, wherein the metadata associated with the object-based content channels comprises at least one of dynamic metadata, static metadata, and / or per-channel static metadata.
[0116] EEE27. The method according to EEE26, wherein when the metadata comprises dynamic metadata, the dynamic metadata comprises at least one of rendering parameters that change over time, a spatial position, a size, a snap, and a zone mask.
[0117] EEE28. The method according to EEE26, wherein when the metadata comprises static metadata, the static metadata comprises rendering parameters that remain fixed over time.
[0118] EEE29. The method according to EEE26, wherein when the metadata comprises per- channel static metadata, the per-channel static metadata comprises metadata that varies between the first object-based content channel and the second object-based content channel.
[0119] EEE30. The method according to any one of EEE1 to EEE29, further comprising: storing, with the electronic processor, a crossfaded audio object in a memory.
[0120] EEE31. The method according to any one of EEE1 to EEE29, further comprising: outputting, via one or more speakers, a crossfaded audio object.
[0121] EEE32. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to any one of EEE1 to EEE31.
[0122] EEE33. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE1 to EEE31.
[0123] EEE34. A non-transitory computer-readable storage medium storing the program according to EEE33.
[0124] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.
[0125] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0126] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.
[0127] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.
Claims
CLAIMS What is claimed is:
1. A method for crossfading between object-based content channels, the object-based content channels comprising at least a first object-based content channel and a second object- based content channel, the method comprising: determining, with an electronic processor, a crossfade overlap range for the object-based content channels; mapping, with the electronic processor, a channel order for the object-based content channels based on metadata and audio associated with the object-based content channels; selecting, with the electronic processor, a crossfade transition point for each of the object- based content channels based on the metadata and the audio associated with the object-based content channels; and crossfading, with the electronic processor, the object-based content channels based on the channel order and the crossfade transition point for each of the object-based content channels.
2. The method of claim 1, wherein the crossfading further comprises: transitioning, with the electronic processor, between first dynamic metadata associated with the first object-based content channel to second dynamic metadata associated with the second object-based content channel at the crossfade transition point for the first object-based content channel.
3. The method of claim 2, wherein the transition between the first dynamic metadata and the second dynamic metadata takes place over a range of time centered around the crossfade transition point for the first object-based content channel.
4. The method of any one of claims 1-3, wherein mapping the channel order further comprises: determining, with the electronic processor, a difference metric measuring the differences between each of the object-based content channels over the crossfade overlap range.
5. The method of any one of claims 1-4, wherein mapping the channel order further comprises: matching, with the electronic processor, the first object-based content channel with the second object-based content channel, wherein the first object-based content channel comprisesaudio during the crossfade overlap range, and the second object-based content channel is silent during the crossfade overlap range.
6. The method of claim 4, wherein mapping the channel order further comprises: matching, with the electronic processor, the first object-based content channel with the second object-based content channel based on minimizing the difference metric, wherein the first and second object-based content channels comprise audio during the crossfade overlap range.
7. The method of any one of claims 1-6, wherein mapping the channel order further comprises creating a third object-based content channel of the object-based content channels.
8. The method of any one of claims 4-7, wherein the determining the difference metric comprises at least one of: (a) determining differences in pre-rendered object audio comprised in the object-based content channels; (b) determining differences based on dynamic metadata associated with the object-based content channels; (c) determining differences in static, per-channel metadata between the object-based content channels; (d) determining differences in per-object rendered output between the object-based content channels; (e) determining differences in per-object object-to-speaker gains of the object-based content channels; and (f) determining differences in audio features of the object-based content channels.
9. The method of any one of claims 4-8, wherein mapping the channel order further comprises minimizing the difference metric.
10. The method of any one of claims 1-9, wherein the crossfade transition point for each of the object-based content channels is comprised within the crossfade overlap range.
11. The method of any one of claims 1-10, wherein the crossfade transition point for each of the object-based content channels comprises a time in samples where dynamic metadata transitions from the first object-based content channel to the second object-based content channel.
12. The method of any one of claims 1-11, wherein selecting the crossfade transition point for each of the object-based content channels comprises minimizing spatial artifacts associated with the object-based content channels.
13. The method of any one of claims 1-12, wherein the crossfading further comprises: crossfading the object-based content channels based on the channel order and the crossfade transition point for each of the object-based content channels to output a crossfaded object-based content channel with dynamic metadata.
14. The method of claim 13, wherein the crossfaded object-based content channel is perceptually identical to an equivalent traditional crossfade of the object-based content channels.
15. The method of any one of claims 1-14, wherein the metadata associated with the object- based content channels comprises at least one of dynamic metadata, static metadata, and / or per- channel static metadata.
16. The method of claim 15, wherein when the metadata comprises dynamic metadata, the dynamic metadata comprising at least one of rendering parameters that change over time, a spatial position, a headphone rendering mode, a headphone rendering metadata, a size, a snap, and a zone mask.
17. The method of claims 15 or 16, wherein when the metadata comprises static metadata, the static metadata comprising of rendering parameters that remain fixed over time.
18. The method of any one of claims 15-17, wherein when the metadata comprises per- channel static metadata, the per-channel static metadata comprising metadata that varies between channels of the object-based content channels.
19. A method for crossfading between a plurality of object-based content channels, the method comprising: determining, with an electronic processor, a crossfade overlap range for the plurality of object-based content channels; matching, with the electronic processor, a first object-based content channel comprising audio during the crossfade overlap range with a second object-based content channel comprising silence during the crossfade overlap range;determining, with the electronic processor, a difference metric measuring differences between each of remaining object-based content channels over the crossfade overlap range; matching, with the electronic processor, the remaining object-based content channels based on minimizing the difference metric; and crossfading, with the electronic processor, the object-based content channels based on the matched object-based content channels.
20. The method of claim 19, wherein the determining the difference metric comprises at least one of: (a) determining differences in pre-rendered object audio contained in the object-based content channels; (b) determining differences based on dynamic metadata associated with the object-based content channels; (c) determining differences in static, pre-channel metadata between the object-based content channels; (d) determining differences in per-object rendered output between the object-based content channels; (e) determining differences in per-object object-to-speaker gains of the object-based content channels; and (f) determining differences in audio features of the object-based content channels.
21. The method of claim 19, wherein determining the difference metric includes calculating a matrix of differences between non-silent channel pairs included in the remaining object-based content channels.
22. The method of claim 19, wherein determining the difference metric includes determining a difference in energy between the remaining object-based content channels during a crossfade region.
23. A method for crossfading between object-based content channels, the object-based content channels comprising at least a first object-based content channel and a second object- based content channel, the method comprising: determining, with an electronic processor, a crossfade overlap range for the first object- based content channel and the second object-based content channel;selecting, with the electronic processor, a crossfade transition point for the first object- based content channel and the second object-based content channel based on first metadata and first audio associated with the first object-based content channel and second metadata and second audio associated with the second object-based content channel; and crossfading, with the electronic processor, the first object-based content channel and the second object-based content channel based on the selected crossfade transition point, wherein crossfading includes transitioning between the first metadata associated with the first object- based content channel to the second metadata associated with the second object-based content channel at the selected crossfade transition point.
24. The method of claim 23, wherein the transition between the first metadata and the second metadata takes place over a range of time centered around the selected crossfade transition point.
25. The method of claim any one of claims 23-24, wherein selecting the crossfade transition point comprises minimizing spatial artifacts associated with the first object-based content channel and the second object-based content channel.
26. The method of any one of claims 23-25, wherein the metadata associated with the object- based content channels comprises at least one of dynamic metadata, static metadata, and / or per- channel static metadata.
27. The method of claim 26, wherein when the metadata comprises dynamic metadata, the dynamic metadata comprises at least one of rendering parameters that change over time, a spatial position, a size, a snap, and a zone mask.
28. The method of claim 26, wherein when the metadata comprises static metadata, the static metadata comprises rendering parameters that remain fixed over time.
29. The method of claim 26, wherein when the metadata comprises per-channel static metadata, the per-channel static metadata comprises metadata that varies between the first object-based content channel and the second object-based content channel.
30. The method of any one of claims 1-29, further comprising: storing, with the electronic processor, a crossfaded audio object in a memory.
31. The method of any one of claims 1-29, further comprising: outputting, via one or more speakers, a crossfaded audio object.
32. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to any one of claims 1-31.
33. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1-31.
34. A non-transitory computer-readable storage medium storing the program according to claim 33.
Citation Information
Patent Citations
Insertion of Sound Objects Into a Downmixed Audio Signal
US20170251321A1
Intelligent Crossfade With Separated Instrument Tracks
US20180277076A1
Using metadata to aggregate signal processing operations
US20210005211A1