Multi-feature upmix detection
A control system extracts audio features and applies a classifier model to detect upmixed audio data, addressing the lack of spatial variation in upmixer-generated audio and enhancing detection accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2025-11-18
- Publication Date
- 2026-05-28
AI Technical Summary
There is a need for a detection system and method that can accurately determine when object-based audio is produced using an upmixer tool, as upmixers often fail to provide an immersive experience due to lack of spatial variation.
Implementing a control system that extracts audio object and channel features from audio data, applies a classifier model to these features, and generates a result indicating whether the audio data was created with an upmixer, using a multi-pass tiered workflow for increased accuracy and reduced computational overhead.
The system provides increased accuracy in detecting upmixed audio data, ensuring higher processing speed and reduced computational requirements while maintaining high accuracy levels.
Smart Images

Figure US2025055976_28052026_PF_FP_ABST
Abstract
Description
Docket No.: D24183WO01MULTI-FEATURE UPMIX DETECTIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application is related to United States Provisional Patent Application No. 63 / 724,201, filed on 22 November 2024, and United States Provisional Patent Application No. 63 / 907,963, filed on 30 October 2025, the entire contents of each of which are incorporated herein by reference.TECHNICAL FIELD
[0002] The present disclosure is generally related to audio processing techniques that may be used to determine when audio data in an audio object-based audio format is produced using an upmixer process.BACKGROUND
[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.
[0004] Sound production can be manifested in a number of interesting ways, using a variety of different audio formats. In a simple example, a stereo format is used to separate left and right channels of audio, providing some spatial separation to the listener. In more complex examples, multi-channel or audio object based formats can be used to provide an immersive three- dimensional experience to the listener.
[0005] Multi-channel audio formats typically include more than two channels of audio, where each channel is intended for a different speaker. These audio formats may find their way into a number of listening scenarios, including but not limited to home theatre entertainment, professional audio production, film or theatre production, and gaming. The specific multichannel audio format is designated by two or three numbers separated by a decimal point. The first number indicates the number of main audio channels for full-range audio. The second number indicates the number of subwoofer channels that carry low-frequency effect (LFE) audio. The third number, when included, indicates the number of height channels that carry the over-head audio. For example, in 5.1.2 multi-channel audio there is a total of eight channels of audio, with five full-range audio channels for left-front, right-front, center, left-rear and rightrear, one low-frequency effect channel for a subwoofer, and two height channels for left and right over-head audio. Some example multichannel formats include stereo (2.0), surround soundDocket No.: D24183WO01 without height channels (5.1, 7.1, 9.1, 9.2, 11.1, etc.), Dolby Digital or AC-3 (5.1), DigitalTheatre System or DTS (5.1 or 7.1), and Dolby Atmos or surround sound with height channels (5.1.2, 7.2.4, 9.1.6, etc.).
[0006] Audio object-based audio formats offer more flexibility and spatial accuracy than channel-based audio formats. Audio data in an audio object-based audio format may be referred to herein as “audio object-based audio data,” as “object-based audio data,” or simply as “objectbased audio.” In object-based audio, each sound source is treated as an independent audio object. Each audio object has metadata to define the object’s location in three-dimensional (3D) space and to determine how it should be rendered for headphones and for loudspeakers. The audio objects are mixed dynamically, meaning that audio objects can be moved anywhere in the 3D space during playback. For example, an audio object corresponding to a helicopter or airplane can move around the user as they are listening. Object-based audio is contrasted to multi-channel audio in that at least some audio objects, which may be referred to herein as “dynamic” audio objects, are not “static” or fixed to a single channel, but instead are moveable. The audio objects are processed or rendered to present the final audio perceived by the user during playback. Object-based audio can be thus rendered for playback on any number of different audio systems based on the capabilities available to the listener (e.g., headphones, soundbars, surround sound system, etc.). Examples of object-based audio include Dolby Atmos ®, Digital Theater Systems (DTS):X, and Moving Picture Experts Group (MPEG)-H.
[0007] It is with respect to these and other considerations that the disclosure made herein is presented.SUMMARY
[0008] Techniques are generally described for evaluating audio data to determine when audio data in an audio object-based audio format that is produced using an upmixer process.
[0009] Briefly stated, devices, systems and methods are disclosed that involve detecting use of an upmixer to create upmixed audio data in an audio object-based audio format. Some example methods involve obtaining, by a control system, audio data in an audio object-based audio format and extracting, by the control system, audio object features from the audio data, to produce extracted audio object features. Some example methods involve applying a classifier model to one or more of the extracted audio object features to generate a result that indicates whether the audio data was created with an upmixer.Docket No.: D24183WO01
[0010] There is a need for a detection system and method that is capable of determining when object-based audio is produced using an upmixer tool.
[0011] Example contemplated methods described herein may include processes to extract features from an input object-based audio and a rendered multi-channel audio signal that is based on the input object-based audio, where analysis of the extracted features indicates that the objectbased audio was generated by an upmix process.
[0012] Some embodiments described herein are directed to using machine learning techniques to perform analysis of the extracted features.
[0013] Some examples described herein include applying a classifier mode to extracted features from input object-based to generate a result that indicates whether the object-based audio was created with an upmixer process, or an upmixer tool.
[0014] Some example benefits achieved include increased accuracy of detecting whether objectbased audio was created with an upmixer process, or an upmixer tool. This increased accuracy may be obtained, at least in part, by implementing one or more of the disclosed methods involving audio object-based signal feature analysis, including but not limited to determining the number of perceptible dynamic audio objects in audio data. Some disclosed methods involve a multi-pass tiered workflow based on a confidence score (e.g., a score in a range from 0% to 100%), which indicates a level of confidence that a result corresponds to an upmixer detection. For example, a first pass with a first classifier model that has a short processing time (such as one of the summation-based methods that are described with reference to Figure 2) may be used to evaluate a batch of audio object based deliverables. The results that are in particular confidence score ranges (e.g., in a range of 0%-20% or 80%- 100%) may not require a second round evaluation with a second classifier (such as one of the ME-based methods that are described herein), whereas the second round of evaluation may be applied to audio data corresponding to the remaining scores (e.g., in a range of 21 %-79%). Such methods provide increased processing speed and reduced computational overhead requirements for a portion of the audio data being evaluated, as well as higher accuracy levels for the remaining audio data being evaluated.
[0015] According to some examples, an apparatus is described for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format.
[0016] Some example apparatus may include an interface system and a control system, where the control system is configured to implement: a control system configured to implement: anDocket No.: D24183WO01 audio object feature extractor configured to receive audio data in an audio object-based audio format, and extract audio object features from the audio data; a tenderer configured to receive the audio data, and to generate a rendered multi-channel audio signal responsive to the audio data; a channel feature extractor configured to receive the rendered multi-channel audio signal, and to extract channel features from the rendered multi-channel audio signal; and a classifier model. The control system may further be configured to: receive extracted channel features from the channel feature extractor; receive extracted audio object features from the audio object feature extractor; apply the classifier model to one or more of the extracted channel features and one or more of the extracted audio object features, and generate a result, wherein the result indicates whether the audio data was created with an upmixer.
[0017] According to some examples, methods are described for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format. Some such methods may involve obtaining, by a control system, audio data in an audio object-based audio format and extracting, by the control system, audio object features from the audio data, to produce extracted audio object features. Some such methods may involve applying a classifier model to one or more of the extracted audio object features to generate a result that indicates whether the audio data was created with an upmixer.
[0018] Some additional disclosed examples also involve methods for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format. Some such methods may involve obtaining, by a control system, audio data in an audio object-based audio format and rendering, by the control system, a multi-channel audio signal responsive to the audio data, to produce a rendered multi-channel audio signal. Some such methods may involve extracting, by the control system, channel features from the rendered multi-channel audio signal. The extracted channel features may include one or more of : (a) a measure of similarity between pairs of rendered channels as a cross-correlation or covariance; (b) a measure of individual signal level associated with each rendered channel; (c) a measure of individual energy level associated with each rendered channel; (d) a measure of individual loudness level associated with each rendered channel; (e) a measure of relative signal level associated with different rendered channel combinations; or (f) a measure of relative loudness level associated with different rendered channel combinations.
[0019] Some such methods may involve extracting, by the control system, audio object features from the audio, to produces extracted audio object features. The extracted audio object features may include one or more of: (g) a number of perceptible static objects; (h) a number ofDocket No.: D24183WO01 perceptible dynamic objects; (i) a 3D map with quantized coordinates that accumulates a short- time level of each audio object at a current location of the audio object over an entire program; or (j) histograms of the short-time level of each audio object. Some such methods may involve applying a classifier model to the extracted audio object features and / or the extracted channel features to generate a result that indicates whether the audio data was created with an upmixer.
[0020] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.
[0021] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associated drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims.DESCRIPTION OF DRAWINGS
[0022] Figure 1A illustrates a schematic block diagram of an example device architecture 101 (in this example, an apparatus 101) that may be used to implement various aspects of the present disclosure.
[0023] Figure IB illustrates a schematic block diagram of an example CPU 141 implemented in the device architecture 101 of Figure 1 A that may be used to implement various aspects of the present disclosure.
[0024] Figure 1C illustrates a block diagram of a system for detecting the use of an upmixer, in accordance with aspects of the present disclosure.
[0025] Figure 2 illustrates another block diagram of a system for detecting the use of an upmixer, in accordance with additional aspects of the present disclosure.
[0026] Figures 3 A, 3B, 3C and 3D illustrate graphical depictions of data and features that are relevant to a process of detecting the use of an upmixer, in accordance with some aspects of the present disclosure.
[0027] Figures 4A, 4B, 4C and 4D illustrate graphical depictions of data and features of a conventional Atmos mix.
[0028] Figures 5 A, 5B, 5C and 5D illustrate graphical depictions of data and features of another conventional Atmos mix.Docket No.: D24183WO01
[0029] Figures 6 A, 6B, 6C and 6D illustrate graphical depictions of data and features of another conventional Atmos mix.
[0030] Figures 7A, 7B, 7C and 7D illustrate graphical depictions of data and features of an upmixed Atmos mix.
[0031] Figures 8 A, 8B, 8C and 8D illustrate graphical depictions of data and features of another upmixed Atmos mix.
[0032] Figures 9 A, 9B, 9C and 9D illustrate graphical depictions of data and features of another upmixed Atmos mix.
[0033] Figure 10 is a flow diagram that outlines various example methods 1000 according to some disclosed implementations.
[0034] Figure 11 is a detailed flow diagram that outlines various example methods 1100 according to some disclosed implementations.
[0035] In the drawings, specific arrangements or orderings of schematic elements, such as those representing devices, units, instruction blocks and data elements, are shown for ease of description. However, it should be understood by those skilled in the art that the specific ordering or arrangement of the schematic elements in the drawings is not meant to imply that a particular order or sequence of processing, or separation of processes, is required. Further, the inclusion of a schematic element in a drawing is not meant to imply that such element is required in all embodiments or that the features represented by such element may not be included in or combined with other elements in some implementations.
[0036] Further, in the drawings, where connecting elements, such as solid or dashed lines or arrows, are used to illustrate a connection, relationship, or association between or among two or more other schematic elements, the absence of any such connecting elements is not meant to imply that no connection, relationship, or association can exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the disclosure. In addition, for ease of illustration, a single connecting element is used to represent multiple connections, relationships or associations between elements. For example, where a connecting element represents a communication of signals, data, or instructions, it should be understood by those skilled in the art that such element represents one or multiple signal paths, as may be needed, to effect the communication.
[0037] The same reference symbol used in various drawings indicates like elements.DETAILED DESCRIPTIONDocket No.: D24183WO01
[0038] The present disclosure relates to techniques for determining when object-based audio is produced using an upmixer tool. These techniques may be implemented as any variety of systems, devices, and methods. In some examples, the techniques may be described in terms of a particular set of steps, operations, or functions that may occur in a certain order, which is merely provided for convenience and clarity. Moreover, the specific details are provided as examples and not intended to limit the scope of this application.Nomenclature
[0039] In this document, the terms “and”, “or” and “and / or” are used. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”).
[0040] The term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0041] This document describes various processing functions that are associated with structures such as blocks, elements, components, circuits, etc. In general, these structures may be implemented by a processor that is controlled by one or more computer programs.
[0042] AcronymsADM - Audio Definition ModelADM BWF - Audio Definition Model Broadcast Wave FormatASIC - Application- Specific Integrated Circuit CD-ROM - Compact Disc Read-Only Memory CNN - Convolutional Neural NetworkCPU - Central Processing UnitDocket No.: D24183WO01DAW - Digital Audio WorkstationDSP - Digital Signal ProcessorEPROM - Erasable Programmable Read-Only Memory FPGA - Field-Programmable Gate Array GUI - Graphical User Interface I / O - Input / OutputITU-R - International Telecommunication Union Radiocommunication SectorLUFS - Eoudness Units Full ScaleMP4 - Moving Picture Experts Group (MPEG) -4 Part 14RAM - Random Access MemoryROM - Read Only MemorySNR - Signal-to-Noise RatioSVM - Support Vector MachineWAV - Waveform Audio File FormatXME - Extensible Markup Eanguage
[0043] Audio mixing or sound engineers are tasked with creating a “mix”, which includes a combination of sounds from multiple sound sources. The mix is typically developed on a digital audio workstation or DAW. The DAW can include a software program and / or an electronic device, where the DAW provides a comprehensive set of tools and features for the entire music production process. An example DAW includes tools for: recording audio tracks, cutting and pasting portions of audio tracks, arranging and modifying audio tracks, applying effects (e.g., compression, reverb, delay, chorus, etc.) and / or equalization, and mixing multiple tracks into a final audio file.
[0044] In object-based audio, sound sources are treated as independent audio objects, each with their own spatial position that can be moved independently in three-dimensional (3D) space. Since the audio objects can be moved independently, the process of creating the mix for a final audio file that includes audio objects can be quite complex.
[0045] According to some examples, the “mix” for the final audio file in object-based audio includes two parts. The first part of the mix is the audio data itself, which may, for example, be a WAV file. The second part of the mix is an ADM or Audio Definition Model file, which is typically an XME file. It is important to note that the ADM file is not the audio data itself. Instead, the ADM file is metadata, which describes the technical properties of the audio. Examples of the technical properties found in the ADM file include how the audio is placed orDocket No.: D24183WO01 positioned in the 3D space, what language the audio is in, diffuseness and size of the audio, panning of the audio, grouping of audio in packs, etc.
[0046] One example way to create a mix is for the audio mixing engineer to use a panning tool in the DAW, where the audio engineer manually positions each of the audio objects in their 3D spatial location in the mix. Such panning tools are often used in scoring audio used in movies, where some audio objects in a particular scene may remain stationary (e.g., an actor speaking from a fixed location), while other audio objects in that particular scene may be in motion moving around the listener (e.g., a helicopter flying around the listener from front to back, right to left, etc.).
[0047] Another way to create a mix in the DAW in the format of object-based audio is for the audio mixing engineer to use an upmix tool, which automatically places audio objects in one or more locations. The upmix tools typically do not involve manual panning of individual audio objects (e.g., sound sources). Instead, the upmix tool receives a previously-generated input audio signal, such as a stereo (2 channel) audio signal, processes the input audio signal using an algorithm, and then automatically outputs a multi-channel audio signal having more channels than the number of channels in the input audio signal (e.g., 5.1.2, 7.1.4, etc.). In some examples, this multi-channel output audio signal may then be used as the channel bed signal for the objectaudio mix, or the individual channels of the multi-channel audio signal may be assigned to audio objects, which may be panned to fixed locations.
[0048] Although there may be some control over the upmix process, an upmix tool typically does not require any manual panning of sound sources to produce a multi-channel signal. The resulting multi-channel signal can be used within a mix to create an object-enabled deliverable such as an Audio Definition Model Broadcast Wave Format (ADM BWF). The Broadcast Wave Format (BWF) is an extension of the Microsoft® WAV audio format, which is a recording format often used in motion picture, radio and television production (see ITU-R BS.1352-3, Annex 1).
[0049] ADM BWF is a file format that can be used to deliver a Dolby® Atmos mix. In Dolby® Atmos, the term “bed channels” is used to describe the channels that correspond to the specific speaker locations of a multichannel audio file. Dolby® Atmos masters are typically stored in ADM BWF files that are ready for encoding then distribution, and include bed channels. For example, 7.1.4 can be used as the bed channels within the ADM BWF. The mix may also contain audio objects, which may be static / stationary audio objects or dynamic / moveable audio objects.Docket No.: D24183WO01
[0050] Using an upmixer tool to produce an audio mix is often not preferred. When an audio mixing engineer manually adjusts the spatial position of objects in the final mix, the engineer exercises a type of creativity that can result in a highly immersive experience for the listener. In contrast, an upmixer does not exercise creativity, which may result in a mix that does not have enough spatial variation, and thus fails to provide an immersive experience. Thus, the present disclosure describes systems, methods and device that process audio files for a mix to determine whether an upmixer was used in creating the mix.
[0051] Figure 1A illustrates a schematic block diagram of an example device architecture 101 (in this example, an apparatus 101) that may be used to implement various aspects of the present disclosure. Architecture 101 may include, but is not limited to, servers and client devices, standalone devices, systems, etc., which may be configured to perform the methods that are described with reference to any or all of the disclosed figures. In some examples, the architecture 101 may be, or may include, a digital audio workstation (DAW). As shown, the architecture 101 includes central processing unit (CPU) 141, which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 142 or a program loaded from, for example, storage unit 148 to random access memory (RAM) 143. The CPU 141 may be, for example, an electronic processor 141. The CPU 141 is an instance of what may be referred to herein as a “control system” or an element of a control system. The ROM 142 and RAM 143 are instances of what may be referred to herein as a “memory system” or an element of a memory system. In RAM 143, the data required when CPU 141 performs the various processes is also stored, as required. In this example, CPU 141, ROM 142, and RAM 143 are connected to one another via bus 144. Input / output (I / O) interface 145 is also connected to bus 144. The bus 144 and the I / O interface 145 are instances of “interface system” elements as that term is used in this disclosure.
[0052] According to this example, the following components are connected to I / O interface 145: input unit 146, which may include a keyboard, a mouse, or the like; output unit 147, which may include a display system including one or more displays, a loudspeaker system including one or more loudspeakers, etc.; storage unit 148 including a hard disk, or another suitable storage device; and communication unit 149 including a network interface card such as a network card (e.g., wired or wireless). The communication unit 149 may be referred to herein as being part of an interface system.
[0053] In some implementations, the input unit 146 may include a microphone system that includes one or more microphones. In some examples, the microphone system may include two,Docket No.: D24183WO01 three or more microphones in different positions, enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0054] According to some implementations, the output unit 147 may include systems with various numbers of loudspeakers. Output unit 147 — and / or another component of the apparatus 101, such as the CPU 141 — may be capable of rendering audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0055] In some embodiments, communication unit 149 may be configured to communicate with other devices (e.g., via a network). In this example, drive 150 is also connected to I / O interface 145, as required. Removable medium 151, such as a magnetic disk, an optical disk, a magnetooptical disk, a flash drive or another suitable removable medium is mounted on drive 150, so that a computer program read therefrom may be installed into storage unit 148, as required. A person skilled in the art would understand that although apparatus 101 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.
[0056] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs and / or on a computer- readable, non-transitory storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 149, and / or installed from the removable medium 151, as shown in Figure 1A.
[0057] According to some examples, the CPU 141 may be, or may be part of, a control system that is configured by machine executable instructions to perform some or all of the methods that are disclosed herein. In some examples, the control system may be configured to implement an audio object feature extractor configured to receive audio data in an audio object-based audio format, and extract object features from the audio data. According to some other examples, the control system may be configured to implement a tenderer configured to receive the audio data, and to generate a rendered multi-channel audio signal responsive to the audio data. In some additional examples, the control system may be configured to implement a channel feature extractor configured to receive the rendered multi-channel audio signal, and to extract channel features from the rendered multi-channel audio signal.
[0058] In some examples, the control system may be configured to implement a classifier model. The control system may be configured to receive extracted object features from the audio objectDocket No.: D24183WO01 feature extractor. The control system may be configured to apply the classifier model to the extracted audio object features. The control system may be configured to generate a result based on applying the classifier model to the extracted audio object features. The result may indicate whether the audio data was created with an upmixer.
[0059] According to some examples, the control system may be configured to receive extracted channel features from the channel feature extractor and to receive extracted audio object features from the audio object feature extractor. The control system may be configured to apply the classifier model to one or more of the extracted channel features and one or more of the extracted audio object features and to generate a result. The result may indicate whether the audio data was created with an upmixer.
[0060] In some alternative examples, the control system may not be configured to implement a classifier model. The control system may be configured for communication — e.g., via a network interface — with a second device or system that is configured to implement the classifier model. The control system may be configured to provide one or more of the extracted audio object features to the second device or system. The second device or system may be configured to generate a result based on applying the classifier model to the extracted audio object features, and to provide the result to the control system. The result may indicate whether the audio data was created with an upmixer.
[0061] In some examples, the control system may be configured to provide one or more of the extracted channel features and one or more of the extracted audio object features to the second device or system. The second device or system may be configured to apply the classifier model to one or more of the extracted channel features and one or more of the extracted audio object features, to generate a result and to provide the result to the control system. The result may indicate whether the audio data was created with an upmixer.
[0062] Figure IB illustrates a schematic block diagram of an example CPU 141 implemented in the device architecture 101 of Figure 1A that may be used to implement various aspects of the present disclosure. The CPU 141 includes an electronic processor 160 and a memory 161. The electronic processor 160 is electrically and / or communicatively connected to the memory 161 for bidirectional communication. The memory 161 stores encoding software 162 and decoding software 163. The memory 161 may be, for example, a ROM, a RAM, or another non-transitory computer readable medium. The electronic processor 160 may implement the encoding software 162 stored in the memory 161 to perform one, some or all of the disclosed methods. Additionally, the electronic processor 160 may implement the decoding software 163 stored in the memory 161 to perform one, some or all of the disclosed methods.Docket No.: D24183WO01
[0063] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 141 in combination with other components of Figure 1A), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0064] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0065] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine- readable signal medium or a machine -readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0066] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that theDocket No.: D24183WO01 program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.
[0067] Figure 1C illustrates a block diagram of a system for detecting the use of an upmixer, in accordance with aspects of the present disclosure. As with other disclosed examples, the types and numbers of elements that are shown in Figure 1C and described herein are merely presented by way of example. Other examples may include one or more different types of elements, different numbers of elements, or both.
[0068] The illustrated system 100 of Figure 1C includes a set of blocks to illustrate various functional partitions. A first block is a Data Storage System 102, which may include one or more storage devices having audio data in one or more audio object-based audio formats stored thereon. In this example, the Data Storage System 102 is providing an object-based audio file (“Object Audio”) 103. A second block is an Audio Object Feature Extractor 104, which is configured to receive the Object Audio 103 and responsively extract Audio Object Features 105. A third block is for a Renderer 106, which is configured to receive the Object Audio 103 and responsively render a Rendered Output Audio Signal 107. A fourth block is for a Channel Feature Extractor 108, which receives the Rendered Output Audio Signal 107 and responsively extracts Channel Features 109. A fifth block is for a Classifier 110, which receives and analyzes the Audio Object Features 105 and the Channel Features 109, and responsively provides a Result 112 indicating whether the Object Audio 103 was created with an upmixer. In some examples, the Result 112 may indicate a probability (e.g. a percentage score in a range of 0%- 100%) that the Object Audio 103 has been created with an upmixer, whereas in other examples the Result 112 may be a binary result (e.g., yes or no) indicating the Classifier 110’s assessment of whether the Object Audio 103 was created with an upmixer.
[0069] According to some implementations, the Audio Object Feature Extractor 104, the Renderer 106 and the Channel Feature Extractor 108 may be implemented by an instance of a control system as disclosed herein, which may reside in a first device. In some such implementations, the Classifier 110 may be implemented by one or more other devices, such as by a server of a cloud-based service. However, in some implementations the Audio Object Feature Extractor 104, the Renderer 106, the Channel Feature Extractor 108 and the Classifier 110 may be implemented by a control system of a single device. In some examples, the DataDocket No.: D24183WO01Storage System 102 may reside in the same device that includes the control system, whereas in other examples the Data Storage System 102 may reside in one or more other devices.
[0070] In some examples, the Object Audio 103 may correspond to an Dolby Atmos deliverable such as an ADM BWF file. This deliverable may, for example, include a collection of audio objects with an audio data file for each of the audio objects (e.g., a WAV file), and metadata (e.g., ADM) for each audio object. The metadata may, for example, indicate where the audio object is to be positioned and / or panned over time. The deliverable further may also include a bed, which is audio data that corresponds to a set of fixed channels. In some examples, it is possible for a Dolby Atmos deliverable to include no actual audio objects and to only include the bed, which is not an ideal situation since the immersive experience will not be as enjoyable to the listener.
[0071] As shown in Figure 1C, the Object Audio 103 flows over two paths. In a first (upper) path 116, the Object Audio 103 is provided to the Audio Object Feature Extractor 104, which processes the Object Audio 103 and directly extracts a first set of features — the Audio Object Features 105 — from the Object Audio 103. In the second (lower) path 118, the Object Audio 103 is provided to a Renderer 106, which applies a tenderer process to the Object Audio 103 that generates a Rendered Output Audio Signal 107 that is suitable for playback according to the channel format required by the bed (e.g., 5.2.2, 7.1.4, 9.1.6, etc.). The Channel Feature Extractor 108 in the second path receives the Rendered Output Audio Signal 107 and applies an extraction process to find a second set of features, the Channel Features 109. The Classifier 110 receives the first and second set of features from the first and second paths, respectively, and determines if the Object Audio 103 was generated by an upmixer.Feature Extraction
[0072] This section describes examples of features that may be extracted by the System 100 of Figure 1C. The first set of features may be derived by the Audio Object Feature Extractor 104, via direct inspection of the Object Audio 103. The second set of features may be derived by the Channel Feature Extractor 108 from the Rendered Output Audio Signal 107, which is a rendered version of the Object Audio 103.Docket No.: D24183WO01Audio -Based Si Features
[0073] The features described in the bullet points of this paragraph may be derived by the Audio Object Feature Extractor 104 by analyzing the Object Audio 103. The analysis may involve analyzing the audio data itself (e.g., analyzing a WAV file), analyzing the corresponding timebased metadata (e.g., ADM metadata), or both.• The number of perceptible static audio objects. Whether an audio object is “perceptible” may be determined by applying a threshold, which may be a peak threshold, an RMS threshold, a Loudness Units Full Scale (LUFS) threshold, or another loudness or level threshold. In some examples, -50 or -60 on the Loudness K-weighted Full Scale (LKFS) may be selected as a threshold. Other examples may involve applying other thresholds. o A static audio object is an audio object that does not change its spatial position over time. o Perceptible static audio objects are counted, while non-perceptible static audio objects are not counted (or ignored for this feature).• The number of perceptible dynamic audio objects o A dynamic audio object is an object that does change its spatial position over time. o Perceptible dynamic audio objects are counted, while non-perceptible dynamic objects are not counted (or ignored for this feature).• A 3D map with quantized coordinates that accumulates the short-time level of each audio object at the current location of the audio object over the entire program. o The measurements can be either weighted (e.g., according to a model of human sound perception) or unweighted. o The 3D map gives an indication of the locations in the render space where the audio object produced energy. o The 3D map itself may be a useful input to a classifier. Alternatively, or additionally, derived features from the map (e.g. identification of the number of audio objects which are stationary vs. non-stationary for the entire duration of a program) may be used as features.• Histograms of the short-time level of each audio object o In some examples, “short-time” may, for example, be a time_on the order of tens of milliseconds, such as the duration of a single frame of audio data. In otherDocket No.: D24183WO01 examples, “short-time” may be a time corresponding to more than one audio frame, e.g., 2, 3, 4 or 5 audio frames. o These measurements could be weighted or unweighted RMS, or a suitable weighted or unweighted loudness measurement. o The histograms give an indication of how much content each audio object had and at roughly what level. They can also be used to identify if an audio object is silent for the duration of the entire program. Some upmixed audio data in an audio object-based audio format included silent audio objects, apparently with the intention of making the audio data seem to be legitimate (in other words, not produced by an upmixing process). o The level histograms may, in some examples, be input to a classifier. In other examples, the level histograms may be converted into another form useful for a classifier (e.g. separation into Gaussians or finding the peak level).Render-Based Signal Features
[0074] The features described below are examples of features that may be derived by the Channel Feature Extractor 108, by analyzing the Rendered Output Audio Signal 107 that is produced by the Renderer 106 from the Object Audio 103.• Cross-correlation or covariance between channels o This is a measure of similarity between all channel pair combinations of the Rendered Output Audio Signal 107• Individual channel signal level, energy level, or loudness level (e.g. RMS level or a loudness measurement made according to the ITU-R BS.1170 loudness standard)• Relative signal level, energy level, or loudness level of different channel combinations (e.g. front channels compared to back channels)Classifier
[0075] The Classifier 110 may be configured to receive as input one or more types of features described above and to use one or more different methods to classify the Object Audio 103 as either upmixed or not upmixed. For some approaches described herein, configuring a classifier may involve a training process based, at least in part, on a set of example audio deliverables in an audio object-based audio format, but known to have been created with an upxmixer tool, as well as a set of example audio object-based audio deliverables known to have not been created with an upmixer (e.g., a conventional Dolby Atmos mix created by an audio engineer).Docket No.: D24183WO01
[0076] Figure 2 illustrates another block diagram of a system for detecting the use of an upmixer, in accordance with additional aspects of the present disclosure. According to this example, Figure 2 illustrates a feature extraction module 203 and an instance of the Classifier 110 of Figure 1C. As with other disclosed examples, the types and numbers of elements that are shown in Figure 2 and described herein are merely presented by way of example. Other examples may include one or more different types of elements, different numbers of elements, or both.
[0077] In some examples, the feature extraction module 203 may include an instance of the Audio Object Feature Extractor 104 of Figure 1C, an instance of the Channel Feature Extractor 108 of Figure 1C, or both. Accordingly, the feature extraction module 203 may receive as input the Object Audio 103, the Rendered Output Audio Signal 107, or both, depending on the particular implementation. According to this example, the feature extraction module 203 is shown outputting features Fl through Fn, with feature F2 also explicitly shown, and providing features Fl through Fn to an instance of the Classifier 110. In this example, n represents an integer of three or more. Accordingly, in this example the feature extraction module 203 is outputting at least three features and possibly more. In some alternative examples, the feature extraction module 203 may output fewer features, e.g., only two features.
[0078] According to this example, the Classifier 110 includes a weighting module 205 and a summation module 209. In this example, the weighting module 205 is configured to apply one of the weights (Wi, W2, ... Wn) to a corresponding one of the features (Fi, F2, ... Fn), to produce the weighted features 207a through 207n, including the weighted feature 207b. According to this example, the weighted feature 207a corresponds to WiFi, the weighted feature 207b corresponds to W2F2 and the weighted feature 207n corresponds to WnFn. In some examples, the weights (Wi, W2, ... Wn) applied by the weighting module 205 may have been previously selected by a person or by an algorithm.
[0079] In this example, the summation module 209 is configured to produce the Result 112, which in this instance is the following weighted sum:In one example, weights Wi through W55 equal 0.1, weights W56 through WHO equal 0.8, features Fi through F55 equal max(CG,0) and features F56 through Fno equal abs(min(CCi,0)), where CG represents the cross-correlation for the ithrendered channel pair. Accordingly, max(CG,0) represents the amount of positive correlation and abs(min(CG,0)) represents the amount of negative correlation.Docket No.: D24183WO01
[0080] In some examples, a control system may be configured to compare the Result 112 to a threshold level to make a determination about the source audio. In some examples, when Result is greater than the threshold level, the classifier may determine that an instance of the Object Audio 103 was created by an upmixer. In such examples, the selected features may be those that tend to show that an upmix was created by an upmixer, such as the amount of off-axis correlation between channels, the amount of negative correlation between channels, etc. In some alternative implementations, when Result is less than the threshold level, the classifier may determine that an instance of the Object Audio 103 was created by an upmixer. In such examples, the selected features may be those that tend to show that an upmix was not created by an upmixer, such as the number of detectable dynamic audio objects, the number of detectable static audio objects in non-canonical locations, etc.
[0081] According to some examples, at least some aspects of the Classifier 110 may be based on one or more machine learning (ML) techniques. An example ML approach may involve supervised learning, where labeled features of a training set are inputted to the Classifier 110 in the learning phase to build a model, and the resulting trained Classifier 110 may be operationally used to process new instances of instance of the Object Audio 103 and provide Results 112 indicating whether the Object Audio 103 was created by an upmixer, such as Results 112 indicating a probability that the input Object Audio 103 was upmixed. In some examples, the training set may include labeled features for verified examples of both the upmixed and non- upmixed instances of Object Audio 103. In some examples, the Classifier 110 may be based on one or more supervised learning algorithms including, but not limited to, linear regression, polynomial regression, Naive Bayes, decision tree, support vector machine (SVM) and / or K- nearest neighbor. Instances of the Classifier 110 that were trained according to SVM and K- nearest neighbor supervised learning algorithms were found to produce good results.
[0082] In some examples, when extensive sets of data are available, the ML model may be based on Deep Leaning. Deep learning often requires less human intervention, since the features of a dataset can be extracted automatically, versus simpler ML techniques that often require a person to manually identify features and adjust the algorithm accordingly. Deep learning is a subset of machine learning that is modelled after the human brain as a multi-layer neural network that can process information. The neural network includes a large number of computational nodes, arranged in layers of connectivity. The neural network may include an input layer, an output layer, and a hidden layer. When a neural network includes three or more layers, the network is considered “deep.” After extensive training, a deep learning model is capable of providing analysis of complex datasets.Docket No.: D24183WO01Workflow
[0083] In an example workflow, a set of audio object based deliverables (e.g., multiple ADM BWF files) can be provided for evaluation on a batch basis, where the results of processing the entire batch are provided in an output file that can be reviewed by a user at a later time. This type of workflow is useful when the audio data files to be analyzed may be quite large (for example, hundreds of audio data files, thousands of audio data files, etc.) and the processing time may be very long.
[0084] In another example workflow, an iterative process may be employed where during each iteration of the process a different classifier model is used. Such as workflow may be useful when there is less confidence in a first classifier model that has a very short processing time (such as one of the summation-based methods that are described with reference to Figure 2) when compared to a second classifier model with a relatively longer processing time (such as one of the ML-based methods that are described herein).
[0085] In still another example, some classifier models may be configured to generate a confidence score (e.g., a score in a range from 0% to 100%), which indicates a level of confidence that a result corresponds to an upmixer detection. This is in contrast to the various described examples, where a classifier may generate a binary value of upmixer detected or no upmixer detected. For such a confidence score type of model, a multi-pass tiered approach can be applied in the workflow. For example, initially a first pass with a first classifier (such as one of the summation-based methods that are described with reference to Figure 2) can be used to evaluate a batch of audio object based deliverables, where the results in a particular range (e.g., in a range from 0%-20% or 80%- 100%) are considered to be confident scores and the remaining scores (e.g., in a range from 21 %-79%) may need a second round evaluation with a second classifier (such as one of the ML-based methods that are described herein).of Data and Features
[0086] Figures 3 A, 3B, 3C and 3D illustrate graphical depictions of data and features that are relevant to a process of detecting the use of an upmixer, in accordance with some aspects of the present disclosure. Figures 3A-3D are intended as guides to Figures 4A-9D, in the sense that the types of data depicted and the arrangement of graphical depictions of the data in Figures 3A- 3D are consistent with those illustrated in Figures 4A-9D. Figure 3A, as well as Figures 4A, 5A, 6A, 7A, 8A and 9A, show examples of perceptible static and / or dynamic audio object positions throughout the duration of an audio asset (for example, throughout a musical composition such as a song). Accordingly, these examples show instances of what is referred to herein as a “3D map” of perceptible audio objects. In these examples, each of the audio assets is in a DolbyDocket No.: D24183WO01Atmos format, and may be referred as an “Atmos mix” or an “Atmos asset.” Figure 3B, as well as Figures 4B, 5B, 6B, 7B, 8B and 9B, show examples of the correlation between channels of a 7.1.4 render of the Atmos mix. Figure 3C, as well as Figures 4C, 5C, 6C, 7C, 8C and 9C, show examples of how various metadata fields were used when creating the Atmos mix. Figure 3D, as well as Figures 4D, 5D, 6D, 7D, 8D and 9D, show examples of energy distribution for channels and audio objects over the duration of the audio asset. Accordingly, these examples show instances of what is referred to herein as a “level histogram.” As with other disclosed examples, the types and numbers of elements that are shown in Figures 3A-9D and described herein are merely presented by way of example. Other examples may involve one or more different types of data, different depictions of the data, or combinations thereof.
[0087] Figures 4A, 4B, 4C and 4D illustrate graphical depictions of data and features of a conventional Atmos mix. As used herein, the phrase “conventional Atmos mix” refers to an Atmos mix that generated, at least in part, according to personal input by an audio engineer. For example, during the process of producing a conventional Atmos mix, the audio engineer may specify — e.g., according to audio object metadata selected by the audio engineer — one or more audio object locations, one or more audio object sizes, one or more audio object trajectories, the diffusiveness of one or more audio objects etc.
[0088] Referring to Figure 4A, one may observe that this instance of a conventional Atmos mix included more than 15 audio objects. In this example, as well as in the examples shown in Figures 5A, 6A, 7A and 8A, each audio object is shown with a different shade of grayscale. In these examples, the size of each audio object corresponds with the total energy or level at a position, summed throughout the duration of the audio asset. These audio objects were all static, there is a large number of them, many of which are in non-canonical locations. As used herein the phrase “non-canonical location” in the context of an Atmos mix refers to a location that does not correspond to a canonical loudspeaker layout for Atmos, such as 5.1.2, 7.2.4, 9.1.6, etc. The presently disclosed invention identifies that having a relatively large (e.g., 10 or more, 12 or more, 14 or more, 16 or more, etc.) number of perceptible audio objects, at least some of which are in non-canonical locations, is one of the characteristics of a conventional Atmos mix.
[0089] In Figure 4B, one may observe that there is little off-axis correlation and no negative correlation between channels. In this context, the “axis” is the diagonal from the upper left comer to the lower right corner, corresponding to each channel’s correlation with itself. Having little off-axis correlation and having no negative correlation between channels are both characteristics of a conventional Atmos mix.
[0090] Referring to Figure 4C, one may observe that when producing this instance of a conventional Atmos mix, the audio engineer did not use many types of metadata features. TheDocket No.: D24183WO01 presently disclosed invention identifies that the use of the headphone-related metadata can sometimes indicate the types of customization one might expect to find in a conventional Atmos mix. However, in this example, the audio engineer intermittently used the headphone render off (hp_render_off) metadata feature, but did not use the headphone render far (hp_render_far) or headphone render near (hp_render_near) metadata features at all. Accordingly, in this example, the use of metadata features did not provide a strong indication that the mix being evaluated was a conventional Atmos mix.
[0091] The histogram of Figure 4D shows wide energy distribution in individual channels and a varied distribution of energy across channels. The presently disclosed invention identifies that both of these features are characteristics of a conventional Atmos mix.
[0092] Figures 5 A, 5B, 5C and 5D illustrate graphical depictions of data and features of another conventional Atmos mix. Referring to Figure 5A, one may observe that this instance of a conventional Atmos mix included numerous perceptible dynamic audio objects. As in Figure 4A, each audio object in Figure 5A is shown with a different shade of greyscale. One may also observe that several of the dynamic audio objects, including audio objects 501a, 501b and 501c, traverse significant portions of the reference environment. Audio object 501a, for example, traverses the entire length of the x axis for values of x=0 and x=l. Audio object 501c traverses throughout most of the x,y plane in the upper half of the reference environment, at various z values between z=0 and z=l. The presently disclosed invention identifies that having 1 or more perceptible dynamic audio objects is a characteristic of a conventional Atmos mix. The amount of movement of a dynamic audio object and the complexity of the dynamic audio object’s movement(s) are further indications of audio engineer customization and therefore of a conventional Atmos mix.
[0093] In Figure 5B, one may observe that there is more off-axis correlation than in Figure 4B, but still not a great deal, and only a small amount of negative correlation between channels. Having little off-axis correlation and little negative correlation between channels both indicate a conventional Atmos mix.
[0094] Referring to Figure 5C, one may observe that when producing this instance of a conventional Atmos mix, the audio engineer used the headphone render off (hp_render_off) and the headphone render far (hp_render_far) metadata features. Accordingly, in this example, the use of metadata features provides some indication that the mix being evaluated was a conventional Atmos mix.
[0095] The histogram of Figure 5D shows wide energy distribution in individual channels and a varied distribution of energy across channels, both of which are characteristics of a conventional Atmos mix.Docket No.: D24183WO01
[0096] Figures 6 A, 6B, 6C and 6D illustrate graphical depictions of data and features of another conventional Atmos mix. Referring to Figure 6A, one may observe that this instance of a conventional Atmos mix included numerous perceptible audio objects, each of which is shown with a different grayscale shade. One may also observe that several of the audio objects are in non-canonical locations and some of the audio objects seem to be dynamic: several of the audio objects are shown in more than one location, although a range of intermediate locations is not shown. Having a relatively large number of perceptible audio objects, at least some of which are in non-canonical locations and some of which appear to be dynamic, are all characteristics of a conventional Atmos mix.
[0097] In Figure 6B, one may observe that there is not a great deal of off-axis correlation and no negative correlation between channels. Having little off-axis correlation and no negative correlation between channels both indicate a conventional Atmos mix.
[0098] Referring to Figure 6C, one may observe that when producing this instance of a conventional Atmos mix, the audio engineer used the headphone render off (hp_render_off), the headphone render near (hp_render_near) and the headphone render far (hp_render_far) metadata features. Accordingly, in this example, the use of metadata features provides indications that the mix being evaluated was a conventional Atmos mix.
[0099] The histogram of Figure 6D shows wide energy distribution in individual channels and a varied distribution of energy across channels, both of which are characteristics of a conventional Atmos mix.
[0100] Figures 7 A, 7B, 7C and 7D illustrate graphical depictions of data and features of an upmixed Atmos mix. As used herein the term “upmixed Atmos mix” refers to audio data in an Atmos format that has been created using an upmix tool, and is not a conventional Atmos mix. Referring to Figure 7A, one may observe that this instance of an upmixed Atmos mix included only two perceptible audio objects, each of which seems to be in a canonical location. Having only two static perceptible audio objects, both of which are in canonical locations, are characteristics of an upmixed Atmos mix.
[0101] In Figure 7B, one may observe that there is a great deal of off-axis correlation between channels. In this example, the center channel correlates with all channels except the LFE channel, all of the left surround channels correlate with each other and all of the right surround channels correlate with each other. Having a lot of off-axis correlation indicates an upmixed Atmos mix.
[0102] Referring to Figure 7C, one may observe that when producing this instance of a conventional Atmos mix, the audio engineer used the headphone render off (hp_render_off), the headphone render near (hp_render_near) and the headphone render far (hp_render_far) metadataDocket No.: D24183WO01 features. Accordingly, in this example, the use of metadata features provides some indications that the mix being evaluated might not be an upmixed Atmos mix.
[0103] The histogram of Figure 7D shows little energy distribution in individual channels and little variation in the distribution of energy across channels. For example, the distributions of energy in channels 5, 7 and 9 seems to be identical. Likewise, the distributions of energy in channels 6, 8 and 10 seems to be identical, as well as the distributions of energy in channels 1 and 2, and the distributions of energy in channels 11 and 12. The small amount of energy distribution in individual channels and the little variation in the distribution of energy across channels are indications of an upmixed Atmos mix.
[0104] Figures 8 A, 8B, 8C and 8D illustrate graphical depictions of data and features of another upmixed Atmos mix. Referring to Figure 8A, one may observe that this instance of an upmixed Atmos mix included only two perceptible audio objects, each of which seems to be in a canonical location. Having only two static perceptible audio objects, both of which are in canonical locations, are characteristics of an upmixed Atmos mix.
[0105] In Figure 8B, one may observe that there is a great deal of off-axis correlation between channels and a substantial amount of negative correlation between channels. In this example, all of the left channels correlate with each other and all of the right channels correlate with each other. According to this example, the RTF and RRR channels negatively correlate with all of the left channels, and the LTF and LRR channels negatively correlate with all of the right channels. Having a lot of off-axis and a lot of negative correlation indicates an upmixed Atmos mix.
[0106] Referring to Figure 8C, one may observe that when producing this instance of a conventional Atmos mix, the audio engineer used the headphone render off (hp_render_off), the headphone render near (hp_render_near) and the headphone render far (hp_render_far) metadata features. Accordingly, in this example, the use of metadata features provides some indications that the mix being evaluated might not be an upmixed Atmos mix.
[0107] The histogram of Figure 8D shows little energy distribution in individual channels and a relatively small amount variation in the distribution of energy across channels. The small amount of energy distribution in individual channels and across channels are indications of an upmixed Atmos mix.
[0108] Figures 9A, 9B, 9C and 9D illustrate graphical depictions of data and features of another upmixed Atmos mix. Referring to Figure 9A, one may observe that this instance of an upmixed Atmos mix included no perceptible audio objects. Having no perceptible audio objects indicates an upmixed Atmos mix.
[0109] In Figure 9B, one may observe that there is a great deal of off-axis correlationDocket No.: D24183WO01 between channels. In this example, the left, right and center channels correlate with each other and with the LTF, RTF, LTR and RTR channels. Having a lot of off-axis correlation indicates an upmixed Atmos mix.
[0110] Referring to Figure 9C, one may observe that when producing this instance of a conventional Atmos mix, the audio engineer used the headphone render off (hp_render_off) and the headphone render near (hp_render_near) metadata features. Accordingly, in this example, the use of metadata features provides some indications that the mix being evaluated might not be an upmixed Atmos mix.
[0111] The histogram of Figure 9D shows a substantial amount of energy distribution in individual channels and some variation in the distribution of energy across channels. However, the energy distribution in channels 1, 2 and 3 is nearly identical, the energy distribution in channels 5 and 6 is nearly identical, the energy distribution in channels 7 and 8 is nearly identical and the energy distribution in channels 9 and 10 is nearly identical. The matching energy distributions in many sets of individual channels are indications of an upmixed Atmos mix.Example Embodiments
[0112] In some embodiments, a system is deployed as part of a Quality Assurance (QA) System that processes object-based audio and provides the output for interactive access by a QA professional.
[0113] In some variations, the QA professional can manually select a first classifier model to run a first QA test on an object based audio file, and then manually select a second (alternative) classifier model to run a second QA test on the same object based audio file. By comparing the results of various classifier models, the QA professional may be able to gather further insights on the results provided.
[0114] In some additional embodiments, a QA system is deployed within a cloud based architecture. Optionally, the object-based audio can be batched processed through the QA system and the outputs may be tabulated in a human readable report.
[0115] Figure 10 is a flow diagram that outlines various example methods 1000 according to some disclosed implementations. The example methods 1000 may be partitioned into blocks, such as blocks 1005, 1010 and 1015. The various blocks may be described as operations, processes, methods, steps, acts or functions. The blocks of methods 1000, like other methods described herein, are not necessarily performed in the order indicated. In some implementations, one or more of the blocks of methods 1000 may be performed concurrently. Moreover, some implementations of methods 1000 may include more or fewer blocks thanDocket No.: D24183WO01 shown and / or described. The blocks of methods 1000 may be performed by one or more devices, for example, the device that is shown in Figure 1A, Figure IB or Figure 1C.
[0116] Processing may commence at block 1005. Block 1005 involves “obtaining, by a control system, audio data in an audio object-based audio format.” In some examples, block 1005 may involve obtaining the audio data from a memory system of the same device that includes the control system. According to some examples, block 1005 may involve obtaining the audio data from a memory system of another device. Processing may continue to block 1010.
[0117] Block 1010 involves “extracting, by the control system, audio object features from the audio data, to produce extracted audio object features.” According to some examples, block 1010 may involve extracting a number of perceptible static objects, a number of perceptible dynamic objects, a 3D map with quantized coordinates that accumulates a short-time level of each audio object at a current location of the audio object over an entire program, one or more histograms of the short-time level of each audio object, or combinations thereof. Processing may continue to block 1015.
[0118] Block 1015 involves “applying a classifier model to one or more of the extracted audio object features to generate a result that indicates whether the audio data was created with an upmixer.” According to some examples, block 1015 may involve applying a classifier model that is implemented by the control system. In some examples, method 1000 may involve providing, by the control system, one or more of the extracted audio object features to another device that is implementing the classifier model. In some instances, the result may be a binary “yes or no” result that indicates whether the audio data was created with an upmixer. In some examples, the result may be an estimated probability of whether the audio data was created with an upmixer.
[0119] According to some examples, method 1000 may involve rendering, by the control system, a multi-channel audio signal responsive to the audio data, to produce a rendered multichannel audio signal and extracting, by the control system, channel features from the rendered multi-channel audio signal. In various examples, the extracted channel features may include one or more of: a measure of similarity between pairs of rendered channels as a cross-correlation or covariance, a measure of individual signal level associated with each rendered channel, a measure of individual energy level associated with each rendered channel, a measure of individual loudness level associated with each rendered channel, a measure of relative signal level associated with different rendered channel combinations, a measure of relative energy level associated with different rendered channel combinations, a measure of relative loudness level associated with different rendered channel combinations, or combinations thereof. In someDocket No.: D24183WO01 examples, method 1000 may involve applying the classifier model to the extracted channel features to generate the result.
[0120] Figure 11 is a detailed flow diagram that outlines various example methods 1100 according to some disclosed implementations. The example methods 1100 may be partitioned into blocks, such as blocks 1105, 1110, 1115, 1120 and 1125. The various blocks may be described as operations, processes, methods, steps, acts or functions. The blocks of methods 1100, like other methods described herein, are not necessarily performed in the order indicated. In some implementations, one or more of the blocks of methods 1100 may be performed concurrently. Moreover, some implementations of methods 1100 may include more or fewer blocks than shown and / or described. The blocks of methods 1100 may be performed by one or more devices, for example, the device that is shown in Figure 1A, Figure IB or Figure 1C.
[0121] Processing may commence at block 1105. Block 1105 involves “obtaining, by a control system, audio data in an audio object-based audio format.” In some examples, block 1105 may involve obtaining the audio data from a memory system of the same device that includes the control system. According to some examples, block 1105 may involve obtaining the audio data from a memory system of another device. In some examples, block 1105 may involve obtaining at least some of the audio data 102 of Figure 1C. Processing may continue to block 1110.
[0122] Block 1110 involves “rendering, by the control system, a multi-channel audio signal responsive to the audio data, to produce a rendered multi-channel audio signal.” In some examples, block 1110 may be performed by the tenderer 106 of Figure 1C. Processing may continue to block 1115.
[0123] Block 1115 involves “extracting, by the control system, channel features from the rendered multi-channel audio signal, wherein the extracted channel features include one or more of: a measure of similarity between pairs of rendered channels as a cross-correlation or covariance; a measure of individual signal level associated with each rendered channel; a measure of individual energy level associated with each rendered channel; a measure of individual loudness level associated with each rendered channel; a measure of relative signal level associated with different rendered channel combinations; a measure of relative energy level associated with different rendered channel combinations; or a measure of relative loudness level associated with different rendered channel combinations.” According to some examples, block 1115 may be performed by the channel feature extractor 108 of Figure 1C. Processing may continue to block 1120.
[0124] Block 1120 involves “extracting, by the control system, audio object features from the audio data, to produce extracted audio object features, wherein the extracted audioDocket No.: D24183WO01 object features include one or more of: a number of perceptible static objects; a number of perceptible dynamic objects; a 3D map with quantized coordinates that accumulates a short-time level of each audio object at a current location of the audio object over an entire program; or one or more histograms of the short-time level of each audio object.” In some examples, block 1120 may be performed by the audio object feature extractor 104 of Figure 1C. Processing may continue to block 1125.
[0125] Block 1125 involves “applying a classifier model to the extracted audio object features and / or the extracted channel features to generate a result that indicates whether the audio data was created with an upmixer.” In some examples, block 1125 may be performed by the classifier 110 of Figure 1C. According to some examples, block 1125 may involve applying a classifier model that is implemented by the control system. In some examples, method 1100 may involve providing, by the control system, one or more of the extracted audio object features to another device that is implementing the classifier model. In some instances, the result may be a binary “yes or no” result that indicates whether the audio data was created with an upmixer. In some examples, the result may be an estimated probability of whether the audio data was created with an upmixer.
[0126] According to some examples, block 1125 may involve one or more of the following: calculating the result as a weighted sum of a number (n) of extracted features (Fi, F2, ... Fn) and a set of corresponding weights (Wi, W2, ... Wn); calculating the result by applying one or more of a linear regression, a polynomial regression, a Naive Bayes, decision tree, or a K- nearest neighbor process; comparing the calculated result to a threshold level to determine if the object-based audio was created with the upmixer; and / or generating a confidence score to indicate a level of confidence for a calculated result.
[0127] In some examples, method 1100 may involve applying a first classifier model in a first pass to achieve a first result for a first range of confidence scores and applying a second classifier model in a second pass to achieve a second result for a second range of confidence scores, where the first classifier model is different from the second classifier model. In some instances, the first range of confidence scores may be different from the second range of confidence scores. According to some examples, if the first range of confidence scores is sufficiently high (e.g., in a range from 80% to 100%), or sufficiently low (e.g., in a range from 0% to 20%), method 1100 may involve only applying the first classifier model.
[0128] According to some examples, the audio data received by the control system may include a bed of channels with no audio objects.
[0129] While this document contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions ofDocket No.: D24183WO01 features that may be specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can, in some cases, be excised from the combination, and the claimed combination may be directed to a sub combination or variation of a sub combination. Logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
[0130] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):EEE 1. An apparatus for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format, the apparatus comprising: an interface system; and a control system configured to implement: an audio object feature extractor configured to receive audio data in an audio object-based audio format, and extract audio object features from the audio data; a tenderer configured to receive the audio data, and to generate a rendered multichannel audio signal responsive to the audio data; a channel feature extractor configured to receive the rendered multi-channel audio signal, and to extract channel features from the rendered multi-channel audio signal; and a classifier model, wherein the control system is further configured to receive extracted channel features from the channel feature extractor; receive extracted audio object features from the audio object feature extractor; apply the classifier model to one or more of the extracted channel features and one or more of the extracted audio object features, and generate a result, wherein the result indicates whether the audio data was created with an upmixer.Docket No.: D24183WO01EEE 2. An apparatus for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format, the apparatus comprising: an interface system; and a control system configured to implement: an audio object feature extractor configured to receive audio data in an audio object-based audio format, and extract audio object features from the audio data; and a classifier model, wherein the control system is further configured to: receive extracted audio object features from the audio object feature extractor; apply the classifier model to the extracted audio object features; and generate a result based on applying the classifier model to the extracted audio object features, wherein the result indicates whether the audio data was created with an upmixer.EEE 3. The apparatus of EEE 2, wherein the control system is further configured to implement: a tenderer configured to receive the audio data, and generate a rendered multi-channel audio signal responsive to the audio data; and a channel feature extractor configured to receive the rendered multi-channel audio signal, and extract channel features from the rendered multi-channel audio signal.EEE 4. The apparatus of EEE 3, wherein the control system is further configured to: obtain extracted channel features from the channel feature extractor; and apply the classifier model to one or more of the extracted channel features, and generate the result based in part on the extracted channel features.EEE 5. The apparatus of any one EEEs 1, 3 or 4, wherein the extracted channel features include one or more of: a measure of similarity between pairs of rendered channels as a cross-correlation or covariance; a measure of individual signal level associated with each rendered channel; a measure of individual energy level associated with each rendered channel; a measure of individual loudness level associated with each rendered channel; a measure of relative signal level associated with different rendered channel combinations;Docket No.: D24183WO01 a measure of relative energy level associated with different rendered channel combinations; or a measure of relative loudness level associated with different rendered channel combinations.EEE 6. The apparatus of any one of the preceding EEEs, wherein the audio data comprises a bed of channels with no audio objects.EEE 7. The apparatus of any one of the preceding EEEs, wherein the extracted audio object features include one or more of: a number of perceptible static audio objects, a number of perceptible dynamic audio objects, a 3D map with quantized coordinates that accumulates a short-time level of each audio object at a current location of the audio object over an entire program, or histograms of short- time level of each audio object.EEE 8. The apparatus of any one of the preceding EEEs , where the classifier model calculates the result as a weighted sum of a number (n) of extracted features (Fi, F2, ... Fn) and a set of corresponding weights (Wi, W2, ... Wn).EEE 9. The apparatus of any one of the preceding EEEs, where the classifier model includes a machine learning model that applies one or more of: linear regression, polynomial regression, Naive Bayes, decision tree, or K-nearest neighbor.EEE 10. The apparatus of any one of the preceding EEEs, where the classifier model is a deep learning model that includes a multi-layer neural network.EEE 11. The apparatus of any one of the preceding EEEs, where the classifier model is a support vector model.EEE 12. The apparatus of any one of the preceding EEEs, where the extracted features used by the classifier model include extracted audio object features and extracted channel features.EEE 13. The apparatus of any one of the preceding EEEs, where the control system is further configured to compare the result to a threshold level to determine if the object-basedDocket No.: D24183WO01 audio was created with the upmixer.EEE 14. The apparatus of any one of the preceding EEEs, where the result corresponds to a binary value or a value over a range.EEE 15. The apparatus of any one of the preceding EEEs, where the control system includes a processor that is configured to execute non-transitory machine -readable instructions.EEE 16. A method for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format, the method comprising: obtaining, by a control system, audio data in an audio object-based audio format; extracting, by the control system, audio object features from the audio data, to produce extracted audio object features; and applying a classifier model to one or more of the extracted audio object features to generate a result that indicates whether the audio data was created with an upmixer.EEE 17. The method of EEE 16, further comprising: rendering, by the control system, a multi-channel audio signal responsive to the audio data, to produce a rendered multi-channel audio signal; extracting, by the control system, channel features from the rendered multi-channel audio signal; and applying the classifier model to the extracted channel features to generate the result.EEE 18. A method for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format, the method comprising: obtaining, by a control system, audio data in an audio object-based audio format; rendering, by the control system, a multi-channel audio signal responsive to the audio data, to produce a rendered multi-channel audio signal; extracting, by the control system, channel features from the rendered multi-channel audio signal, wherein the extracted channel features include one or more of: a measure of similarity between pairs of rendered channels as a cross-correlation or covariance; a measure of individual signal level associated with each rendered channel; a measure of individual energy level associated with each rendered channel; a measure of individual loudness level associated with each rendered channel;Docket No.: D24183WO01 a measure of relative signal level associated with different rendered channel combinations; a measure of relative energy level associated with different rendered channel combinations; or a measure of relative loudness level associated with different rendered channel combinations; extracting, by the control system, audio object features from the audio data, to produce extracted audio object features, wherein the extracted audio object features include one or more of: a number of perceptible static objects; a number of perceptible dynamic objects; a 3D map with quantized coordinates that accumulates a short-time level of each audio object at a current location of the audio object over an entire program; or one or more histograms of the short-time level of each audio object; and applying a classifier model to the extracted audio object features and / or the extracted channel features to generate a result that indicates whether the audio data was created with an upmixer.EEE 19. The method of any one of the preceding EEEs, wherein the audio data comprises a bed of channels with no audio objects.EEE 20. The method of any one of the preceding EEEs, wherein applying the classifier model includes one or more of: calculating the result as a weighted sum of a number (n) of extracted features (Fi, F2, ... Fn) and a set of corresponding weights (Wi, W2, ... Wn); calculating the result by applying one or more of a linear regression, a polynomial regression, a Naive Bayes, decision tree, or a K-nearest neighbor process; comparing the calculated result to a threshold level to determine if the object-based audio was created with the upmixer; and / or generating a confidence score to indicate a level of confidence for a calculated result.EEE 21. The method of any one of the preceding EEEs, wherein applying the classifier model comprises: applying a first classifier model in a first pass to achieve a first result for a first range of confidence scores;Docket No.: D24183WO01 applying a second classifier model in a second pass to achieve a second result for a second range of confidence scores; wherein the first classifier model is different from the second classifier model; and wherein the first range of confidence scores is different from the second range of confidence scores.
Claims
Docket No.: D24183WO01CLAIMSWhat is claimed is:
1. An apparatus for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format, the apparatus comprising: an interface system; and a control system configured to implement: an audio object feature extractor configured to receive audio data in an audio object-based audio format, and extract audio object features from the audio data; a tenderer configured to receive the audio data, and to generate a rendered multichannel audio signal responsive to the audio data; a channel feature extractor configured to receive the rendered multi-channel audio signal, and to extract channel features from the rendered multi-channel audio signal; and a classifier model, wherein the control system is further configured to receive extracted channel features from the channel feature extractor; receive extracted audio object features from the audio object feature extractor; apply the classifier model to one or more of the extracted channel features and one or more of the extracted audio object features, and generate a result, wherein the result indicates whether the audio data was created with an upmixer.
2. An apparatus for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format, the apparatus comprising: an interface system; and a control system configured to implement: an audio object feature extractor configured to receive audio data in an audio object-based audio format, and extract audio object features from the audio data; and a classifier model, wherein the control system is further configured to: receive extracted audio object features from the audio object feature extractor; apply the classifier model to the extracted audio object features; and generate a result based on applying the classifier model to the extracted audio object features, wherein the result indicates whether the audio data was created with an upmixer.Docket No.: D24183WO013. The apparatus of claim 2, wherein the control system is further configured to implement: a tenderer configured to receive the audio data, and generate a rendered multi-channel audio signal responsive to the audio data; and a channel feature extractor configured to receive the rendered multi-channel audio signal, and extract channel features from the rendered multi-channel audio signal.
4. The apparatus of claim 3, wherein the control system is further configured to: obtain extracted channel features from the channel feature extractor; and apply the classifier model to one or more of the extracted channel features, and generate the result based in part on the extracted channel features.
5. The apparatus of any one claims 1, 3 or 4, wherein the extracted channel features include one or more of: a measure of similarity between pairs of rendered channels as a cross-correlation or covariance; a measure of individual signal level associated with each rendered channel; a measure of individual energy level associated with each rendered channel; a measure of individual loudness level associated with each rendered channel; a measure of relative signal level associated with different rendered channel combinations; a measure of relative energy level associated with different rendered channel combinations; or a measure of relative loudness level associated with different rendered channel combinations.
6. The apparatus of any one of the preceding claims, wherein the audio data comprises a bed of channels with no audio objects.
7. The apparatus of any one of the preceding claims, wherein the extracted audio object features include one or more of: a number of perceptible static audio objects, a number of perceptible dynamic audio objects, a 3D map with quantized coordinates that accumulates a short-time level of each audio object at a current location of the audio object over an entire program, or histograms of short- time level of each audio object.Docket No.: D24183WO018. The apparatus of any one of the preceding claims , where the classifier model calculates the result as a weighted sum of a number (n) of extracted features (Fi, F2, ... Fn) and a set of corresponding weights (Wi, W2, ... Wn).
9. The apparatus of any one of the preceding claims, where the classifier model includes a machine learning model that applies one or more of: linear regression, polynomial regression, Naive Bayes, decision tree, or K-nearest neighbor.
10. The apparatus of any one of the preceding claims, where the classifier model is a deep learning model that includes a multi-layer neural network.
11. The apparatus of any one of the preceding claims, where the classifier model is a support vector model.
12. The apparatus of any one of the preceding claims, where the extracted features used by the classifier model include extracted audio object features and extracted channel features.
13. The apparatus of any one of the preceding claims, where the control system is further configured to compare the result to a threshold level to determine if the object-based audio was created with the upmixer.
14. The apparatus of any one of the preceding claims, where the result corresponds to a binary value or a value over a range.
15. The apparatus of any one of the preceding claims, where the control system includes a processor that is configured to execute non-transitory machine-readable instructions.Docket No.: D24183WO0116. A method for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format, the method comprising: obtaining, by a control system, audio data in an audio object-based audio format; extracting, by the control system, audio object features from the audio data, to produce extracted audio object features; and applying a classifier model to one or more of the extracted audio object features to generate a result that indicates whether the audio data was created with an upmixer.
17. The method of claim 16, further comprising: rendering, by the control system, a multi-channel audio signal responsive to the audio data, to produce a rendered multi-channel audio signal; extracting, by the control system, channel features from the rendered multi-channel audio signal; and applying the classifier model to the extracted channel features to generate the result.
18. A method for detecting use of an upmixer to create upmixed audio data in an audio object-based audio format, the method comprising: obtaining, by a control system, audio data in an audio object-based audio format; rendering, by the control system, a multi-channel audio signal responsive to the audio data, to produce a rendered multi-channel audio signal; extracting, by the control system, channel features from the rendered multi-channel audio signal, wherein the extracted channel features include one or more of: a measure of similarity between pairs of rendered channels as a cross-correlation or covariance; a measure of individual signal level associated with each rendered channel; a measure of individual energy level associated with each rendered channel; a measure of individual loudness level associated with each rendered channel; a measure of relative signal level associated with different rendered channel combinations; a measure of relative energy level associated with different rendered channel combinations; or a measure of relative loudness level associated with different rendered channel combinations; extracting, by the control system, audio object features from the audio data, to produce extracted audio object features, wherein the extracted audio object features include one or moreDocket No.: D24183WO01 of: a number of perceptible static objects; a number of perceptible dynamic objects; a 3D map with quantized coordinates that accumulates a short-time level of each audio object at a current location of the audio object over an entire program; or one or more histograms of the short-time level of each audio object; and applying a classifier model to the extracted audio object features and / or the extracted channel features to generate a result that indicates whether the audio data was created with an upmixer.
19. The method of any one of the preceding claims, wherein the audio data comprises a bed of channels with no audio objects.
20. The method of any one of the preceding claims, wherein applying the classifier model includes one or more of: calculating the result as a weighted sum of a number (n) of extracted features (Fi, F2, ... Fn) and a set of corresponding weights (Wi, W2, ... Wn); calculating the result by applying one or more of a linear regression, a polynomial regression, a Naive Bayes, decision tree, or a K-nearest neighbor process; comparing the calculated result to a threshold level to determine if the object-based audio was created with the upmixer; and / or generating a confidence score to indicate a level of confidence for a calculated result.
21. The method of any one of the preceding claims, wherein applying the classifier model comprises: applying a first classifier model in a first pass to achieve a first result for a first range of confidence scores; applying a second classifier model in a second pass to achieve a second result for a second range of confidence scores; wherein the first classifier model is different from the second classifier model; and wherein the first range of confidence scores is different from the second range of confidence scores.
Citation Information
Patent Citations
Adaptive Audio Processing Based on Forensic Detection of Media Processing History
US20140336800A1
Multi-Channel Audio Content Analysis Based Upmix Detection
US20150243289A1
Apparatus and method for providing a measure of spatiality associated with an audio stream
US20200021934A1
US202463724201P
US202563907963P