Multiscale attention-based spatial phase features for audio classification

Multiscale attention-based spatial phase features address the challenge of classifying input audio types by analyzing inter-channel phase differences, improving classification accuracy and maintaining content integrity across diverse devices and styles.

WO2026156016A1PCT designated stage Publication Date: 2026-07-23DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2026-01-13
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately classify input audio types, such as binauralized and non-binauralized audio, due to unknown rendering techniques, diverse devices, and varying content creator styles, leading to potential double processing and misinterpretation of audio content.

Method used

Implementing multiscale attention-based spatial phase features to analyze band split inter-channel phase differences, compute statistical averages, and perform cross-band normalization to classify audio types, including binaural and stereo audio, using electronic processors in wearable devices, external devices, or cloud-based systems.

Benefits of technology

Enhances audio classification accuracy by distinguishing between different audio types, ensuring proper rendering and maintaining content creator intent, while being robust to diverse devices and content styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026011151_23072026_PF_FP_ABST
    Figure US2026011151_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A method for classifying audio includes receiving input audio, and computing band split inter-channel phase differences (ICPDs) based on the input audio. The method also includes computing statistical averages of the band split ICPDs based on the input audio. The method also includes performing cross-band normalization of the band split ICPDs, and calculating multiscale attention-based spatial features in discriminative band clusters. The multiscale attention-based spatial features may be used to perform classification of the input audio.
Need to check novelty before this filing date? Find Prior Art

Description

D24019WO01MULTISCALE ATTENTION-BASED SPATIAL PHASE FEATURES FOR AUDIO CLASSIFICATION CROSS-REFERENCE TO RELATED APPLICATIONSThis application claims the benefit of priority from International Patent Application No.PCT / CN2025 / 073134, filed January 17, 2025, International Patent Application No.PCT / CN2025 / 122531, filed September 19, 2025, U. S. Provisional Patent Application No.63 / 886,690, filed September 23, 2025, and U. S Provisional Patent Application No. 63 / 899,087, filed October 14, 2025, each of which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0001] The present disclosure relates generally to identifying and extracting features from input audio that are used to classify one or more types of the input audio, for example, when the type of input audio is unknown.SUMMARY

[0002] Wearable devices (e.g., earbuds, headphones, smart glasses, other head-mounted displays, etc.) are often equipped with one or more sensors and one or more electronic processors. For example, a wearable device may include an inertial measurement unit (IMU) configured to provide real-time data from a gyroscope(s) and / or accelerometer(s) to indicate a status / orientation of the wearable device. This real-time data may be processed / analyzed to determine the orientation and / or physical movement of a head of a user wearing the wearable device. Using such orientation and movement information of the user, spatial audio (e.g., immersive audio signals conveying complex three-dimensional sound environments) can be output by the wearable device, for example, by executing one or more algorithms in the electronic processor(s) of the wearable device.

[0003] Executing such an algorithm(s) improves the user’s experience by providing a virtual listening space that is relatively static to the real-world environment surrounding the user rather than being static to the wearable device itself. In other words, when the user uses a wearable device equipped with such an algorithm(s), the sound will be perceived by the user as if the sound is generated from the real-world environment, which enhances the listening experience provided to the user by making the listening experience more immersive.

[0004] One or more electronic processors executing such algorithms may be located within the wearable device itself, within an external device (e.g., a smart phone) communicativelyD24019WO01coupled to the wearable device, within a remote device (e.g., a server) communicatively coupled to the wearable device, or combinations thereof. Alternatively, similar algorithms may be executed in situations that do not include a wearable device. For example, an external device such as a smart phone, laptop computer, desktop computer, smart television, and / or the like may receive input audio and may output spatial audio to built-in speakers of the external device and / or external speakers communicatively coupled to the external device.

[0005] In the situations explained above and in similar situations, binauralization (e.g., binaural rendering) is a common technique for rendering and / or processing audio (e.g., spatial audio) to provide an immersive experience in headphone playback, external device playback, external speaker playback, etc. while maintaining an intention of a content creator of input audio that is provided for output and consumption by a user / listener. Binauralization is also known as binaural rendering and involves processing an input audio signal with a set of filters to simulate how sounds reach a listeners’ ears from different directions. These filters may be head related transfer functions (HRTFs), which when applied to an input signal, introduce directional cues (e.g., level and phase modifications) which during playback create a perception that a sound source is arriving from a specific direction relative to the listener (e.g., directly in front of the listener, front-left and above the listener, etc.). Many different types of non-binauralized input audio may be binauralized as explained herein. For example, such non-binauralized input audio may include mono audio, stereo audio, multi-channel audio, and spatial audio such as objectbased audio. Spatial audio, in formats such as Dolby Atmos, is often stored on over-the-top (OTT) media service platforms for consumption by end-user devices (e.g., wearable devices, external devices and / or external speakers, etc.). In practice, binauralized audio may be generated from spatial audio at multiple points in a media delivery system, including (1) on OTT media platforms, (2) on consumer smart devices like smartphones, personal computers (PCs), laptops, and / or (3) on wearable devices like headphones, earbuds, headsets, head-mounted displays, etc. Classification of received audio at any point in the media delivery system may prevent audio from being double processed and may aid in maintaining the intention of the content creator of the audio. For example, classification of binaural / binauralized audio versus non-binaural / non-binauralized audio (e.g., stereo audio, etc.) in input audio received by a downstream device may be used to avoid redundant binauralization (i.e., applying binauralization to existing binaural audio signals) thereby maintaining content creators' intention for content playback in the end-user devices.

[0006] However, there is a technological problem that the type of received input audio is often not known by downstream devices such as end-user devices. Additionally, there is aD24019WO01technological problem that several factors impede robust classification of input audio (e.g., binauralized and non-binauralized audio classification). For example, when the audio classification is implemented in consumer external devices (e.g., smartphones, etc.), the audio rendering and pre- and / or post-processing of OTT media platforms are a full or partial black box (i.e., they are unknown). As another example, when the audio classification is implemented in consumer wearable devices (e.g., headphones, etc.), the audio rendering and pre- and / or postprocessing of the OTT media platforms and external devices (e.g., smartphones) are a full or partial black box (i.e., they are unknown). As another example, audio is often encoded with a variety of codecs with varying bitrates, and content streaming from OTT media platforms to smart external devices and wearable devices includes multiple distinct codec processes. The details of such upstream codec processing are often not known to downstream devices that receive input audio from upstream devices. As yet another example, different content creators have unique styles for content creation of audio. Downstream devices that receive the input audio of different content creators are often not specifically aware of a style of content creation that was used to create the input audio.

[0007] To address the above-noted technological problems of classifying input audio (e.g., unknown input audio), the systems, methods, and devices described herein utilize multiscale attention-based spatial phase features to achieve classification of input audio (e.g., classification / discrimination between binauralized audio and non-binauralized audio (e.g., stereo audio, etc.). The systems, methods, and devices described herein address the above-noted technological problems by being robust to different rendering techniques that include binauralization and stereo downmixing, such as those used by Dolby and / or other technology ecosystems (e.g., Apple iOS, Android OS, etc.). Additionally, the systems, methods, and devices described herein are robust to different / diverse end-user devices with unknown pre- and post-processing. Furthermore, the systems, methods, and devices described herein are content independent to be robust to different styles of different content creators.

[0008] In one aspect of the present disclosure, there is provided a method for classifying audio. The method may include receiving, with an electronic processor, input audio. The method may also include computing, with the electronic processor, band split inter-channel phase differences (ICPDs) based on the input audio. The method may also include computing, with the electronic processor, statistical averages of the band split ICPDs based on the input audio. The method may also include performing, with the electronic processor, cross-band normalization of the band split ICPDs. The method may also include calculating, with the electronic processor, multiscale attention-based spatial features in discriminative band clusters.D24019WO01The multiscale attention- based spatial features may be used to perform classification of the input audio.

[0009] In addition to any combination of features described above, the method may include using, with the electronic processor, the multiscale attention-based spatial features to perform classification indicative of whether the input audio is binaural audio or non-binaural audio.

[0010] In addition to any combination of features described above, the method may include using, with the electronic processor, the multiscale attention-based spatial features to perform classification indicative of whether the input audio is professional binaural audio or stereorendered binaural audio.

[0011] In addition to any combination of features described above, the electronic processor may include one or more electronic processors that are located in at least one of a group consisting of a cloud-based computing device, an external device, a wearable device, and combinations thereof.

[0012] In addition to any combination of features described above, the band split ICPDs may include a plurality of frames of the input audio. The statistical averages of the band split ICPDs may be calculated in each band based on multiple adjacent frames.

[0013] In addition to any combination of features described above, computing band split ICPDs may be based on an equivalent rectangular bandwidth (ERB).

[0014] In addition to any combination of features described above, the multiscale attentionbased spatial features may be independent values across each ERB.

[0015] In addition to any combination of features described above, performing the crossband normalization of the band split ICPDs may include calculating, with the electronic processor, a ratio of an absolute value of each band of the band split ICPDs to an average absolute value of all frequency bands of the band split ICPDs.

[0016] In addition to any combination of features described above, calculating the multiscale attention-based spatial features may include calculating, with the electronic processor, shortterm scales, middle-term scales, and long-term scales of aggregation of the band split ICPDs using an aggregation function that is based on a cross-band normalized ICPD.D24019WO01

[0017] In addition to any combination of features described above, the short-term scales, the middle -term scales, and the long-term scales of aggregation may include an exponentially weighted average of an aggregation function. In addition to any combination of features described above, weights of the exponentially weighted average may be preset to indicate different time scales.

[0018] In addition to any combination of features described above, attention of the multiscale attention-based spatial features may capture bands from approximately 400 Hz to approximately 1500 Hz.

[0019] In addition to any combination of features described above, the method may include providing, with the electronic processor, the multiscale attention-based spatial features to an audio classification model. The audio classification model may be configured to be executed to perform classification of the input audio to generate classification results.

[0020] In addition to any combination of features described above, the method may include processing, with the electronic processor, the input audio in accordance with the classification results.

[0021] In another aspect of the present disclosure, the techniques described herein relate to an apparatus including the electronic processor and a memory storing instructions, which when executed by the electronic processor, cause the apparatus to perform the method.

[0022] In another aspect of the present disclosure, the techniques described herein relate to a non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method.

[0023] Other aspects of the embodiments will become apparent by consideration of the detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] FIG. 1 illustrates a wearable device being worn on a head of a user, according to some example instances.

[0025] FIG. 2 is a hardware block diagram of the wearable device of FIG. 1, according to some example instances.D24019WO01

[0026] FIG. 3 is a hardware block diagram of an external device of FIG. 1, according to some example instances.

[0027] FIG. 4 is a hardware block diagram of a cloud-based computing device of FIG. 1, according to some example instances.

[0028] FIG. 5 illustrates a block diagram of audio classification and playback implemented by an electronic processor of a downstream device such as the wearable device of FIG. 1, according to some example instances.

[0029] FIG. 6 illustrates a flowchart of a method for controlling the wearable device to output audio, according to some example instances.

[0030] FIGS. 7A, 7B, and 7C illustrate graphs of spatial phase features in each of a plurality of frequency bands of input audio, according to some example instances.DETAILED DESCRIPTION

[0031] The following description sets forth exemplary methods, parameters, and the like. It should be recognized, however, that such description is not intended as a limitation on the scope of the present disclosure but is instead provided as a description of example embodiments.

[0032] Although the following description uses terms “first,” “second,” etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, a first touch could be termed a second touch, and, similarly, a second touch could be termed a first touch, without departing from the scope of the various described embodiments. The first touch and the second touch are both touches, but they arc not the same touch.

[0033] The terminology used in the description of the various described embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various described embodiments and the appended claims, the singular forms “a," “an," and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “includes,” “including,” “comprises,” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude theD24019WO01presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0034] The term “if is, optionally, construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” is, optionally, construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event],” depending on the context. The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” arc to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0035] Relative terminology, such as, for example, “about,” “approximately,” “substantially,” etc., used in connection with a quantity or condition would be understood by those of ordinary skill to be inclusive of the stated value and has the meaning dictated by the context (e.g., the term includes at least the degree of error associated with the measurement accuracy, tolerances [e.g., manufacturing, assembly, use, etc.] associated with the particular value, etc.). Such terminology should also be considered as disclosing the range defined by the absolute values of the two endpoints. For example, the expression “from about 2 to about 4” also discloses the range “from 2 to 4”. The relative terminology may refer to plus or minus a percentage (e.g., 1%, 5%, 10%, or more) of an indicated value.

[0036] FIG. 1 illustrates a wearable device 105 being worn on a head 110 of a user 115 according to some example instances. The wearable device 105 is illustrated as a set of headphones. In other instances, the wearable device 105 may alternatively be embodied as a set of earbuds, smart glasses, other head-mounted devices such as virtual reality or augmented reality headsets, etc. As shown in FIG. 1, the wearable device 105 may be communicatively coupled to an external device 120 such as a smart phone, a tablet, a laptop computer, a desktop computer, or the like. As shown in FIG. 1, the wearable device 105 may be communicatively coupled to the external device 120 via a communicative connection 125 such as a wired communicative connection 125. In other instances, the wearable device 105 may beD24019WO01communicatively coupled to the external device 120 via a wireless communicative connection (e.g., via a short-range wireless connection such as Bluetooth™ and / or a longer-range wireless connection). In some instances, the wearable device 105 and / or the external device 120 is communicatively coupled to additional and / or alternative devices such as other external devices located nearby the wearable device 105 and / or other devices located remotely from the wearable device 105. For example, and as illustrated in FIG. 1, a cloud-based computing device 130 (e.g., including one or more servers) may be communicatively coupled to the external device 120 and / or the wearable device 105 via a wireless (or wired) connection.

[0037] The wearable device 105 and the external device 120 may be considered end-user devices that are operated by a user. Such end-user devices may output audio for consumption by the user. End-user devices may also be considered downstream devices because they receive input audio from upstream devices for output and / or for processing and then output. The cloudbased computing device 130 is upstream of the end-user devices and may not be considered an end-user device. Nevertheless, in some instances, the cloud-based computing device 130 may be considered a downstream device because it may receive input audio from upstream devices to be forwarded (with or without additional processing) to the end-user devices. As explained herein, the disclosed methods may be implemented on any one or a combination of downstream devices to perform classification of received audio.

[0038] In some instances, the devices shown in FIG. 1 may vary. For example, multiple external devices 120 and / or cloud-based computing devices 130 may be included. In some instances, one or more of the devices 105, 120, and 130 may not be present. In some instances, the external device 120 may be communicatively coupled to external speakers 335 (see FIG. 3) such as a surround sound speaker system in the user’s residence.

[0039] FIG. 2 is a hardware block diagram of the wearable device 105 according to some example instances. In the example illustrated in FIG. 2, the wearable device 105 includes an electronic processor 205 (for example, a microprocessor or other electronic device). The electronic processor 205 includes input and output interfaces (not shown) and is electrically coupled to a memory 210, a network interface 215, speakers 220 (e.g., one or more audio output devices), and one or more orientation and / or movement sensors 225. In some instances, the wearable device 105 includes fewer or additional components in configurations different from that illustrated in FIG. 2 and / or optionally combines two or more components shown in FIG. 2. For example, the wearable device 105 may additionally include one or a combination of a microphone(s), a camera(s), and a display screen. As another example, the electronic processorD24019WO01205 may include multiple electronic processors 205 within the wearable device 105 (e.g., a main electronic processor 205 of the wearable device 105 and an electronic processor included in an IMU) that together function to control various aspects of the wearable device 105. In some instances, the wearable device 105 is implemented within a distributed system including one or more components located in different devices. For example, in some embodiments, the wearable device 105 (e.g., the electronic processor 205) includes local hardware components and one or more external hardware components (e.g., one or more electronic processors of the external device 120). In other words, the electronic processor 205 may include any one or a combination of electronic processors located within a single device (e.g., the wearable device 105) or distributed among various devices and / or systems.

[0040] In some instances, the wearable device 105 performs functionality other than the functionality described below. The various components of the wearable device 105 described herein are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0041] The memory 210 may include read only memory (ROM), random access memory (RAM), other non-transitory computer-readable media, or a combination thereof. The electronic processor 205 is configured to receive instructions and data from the memory 210 and execute, among other things, the instructions. In particular, the electronic processor 205 executes instructions stored in the memory 210 to perform the methods described herein.

[0042] The network interface 215 sends data to and receives data from other devices (e.g., the external device 120, the cloud-based computing device 130, etc.). In some instances, the network interface 215 includes one or more transceivers for wirelessly communicating with the other devices. Alternatively or in addition, the network interface 215 may include a connector or port for receiving a wired connection to one or more of the other devices, such as connector to receive a coaxial cable. The electronic processor 205 may receive data (for example, audio data and / or the like) from one or more other devices (e.g., the external device 120, the cloud-based computing device 130, and / or another device) via the network interface 215. The electronic processor 205 may output the data (e.g., audio / audio data) via the speakers 220 (and / or via a display screen in instances where the wearable device 105 includes the display screen). In some instances, the speakers 220 include a left speaker and right speaker. In some instances, the speakers 220 may include additional and / or alternative speakers 220. In some instances, the wearable device 105 may also receive power from another device (e.g., the external device 120) via the network interface 215 to power components of the wearable device 105 (e.g., theD24019WO01electronic processor 205). In some instances, the wearable device 105 may include a separate power connection to receive power from another device and / or may include its own power supply such as a rechargeable or replaceable battery.

[0043] The orientation and / or movement sensor(s) 225 may include one or more sensors that are configured to collect and / or determine data indicative of an orientation (e.g., three-dimensional orientation including pitch, roll, yaw, and / or the like) and / or a movement of the head 110 of the user 115 of the wearable device 105 (i.e., head orientation / movement data). For example, the orientation and / or movement sensor(s) 225 may include an inertial measurement unit (IMU) that may include one or more accelerometers and / or gyroscopes. The head orientation / movement data may be an output of any one or a combination of sensors of the IMU (e.g., raw data from one or more sensors). Additionally or alternatively, the head orientation / movement data may be an output of the IMU itself. For example, the IMU may include an electronic processor configured to analyze raw data from one or more sensors and make one or more determinations to generate head orientation / movement data that is then provided to the electronic processor 205. In some instances, the head orientation / movement data includes absolute orientation data, velocity data, acceleration data, magnetic field data, and / or the like. In some instances, the head orientation / movement data includes linear and angular data, acceleration data from one or more accelerometers, angular velocity data from one or more gyroscopes, magnetometer data, and / or the like. In some instances, the electronic processor 205 is configured to receive the data indicative of the orientation and / or the movement of the head 110 of the user 115 of the wearable device 105 (i.e., the head orientation / movement data) from the orientation and / or movement sensor(s) 225. In some instances, the electronic processor 205 is configured to receive data indicative of the orientation and / or the movement of the head 110 of the user 115 of the wearable device 105 (i.e., the head orientation / movement data) from one or more sensors external to device 105 (e.g., camera, infrared, radar, lidar, sonar, millimeter wave (mmWave), or other sensors capable capturing data associated with user pose). In such implementations, the one or more external sensors (not illustrated) may be used in place of or in conjunction with the one or more orientation and / or movement sensors 225 of device 105 to provide more robust determination of the orientation and / or movement of the head 110 of the user 115 of the wearable device 105.

[0044] FIG. 3 is a hardware block diagram of the external device 120 according to some example instances. As illustrated, in some instances, the external device 120 may include the same or similar components shown in FIG. 2 with respect to the wearable device 105 (e.g., an electronic processor 305, a memory 310, a network interface 315, speakers 320, and orientationD24019WO01and / or movement sensors 325). Such components of the external device 120 may be similar to the components described above with respect to the wearable device 105 and may perform similar general functions. In some instances, the external device 120 includes fewer or additional components in configurations different from that illustrated in FIG. 3 and / or optionally combines two or more components shown in FIG. 3. For example, as shown in FIG.3, the external device 120 may include a display 330 such as a touchscreen configured to display a user interface and / or receive user inputs from a user. As another example, the external device 120 may be optionally communicatively coupled to external speakers 335 (e.g., surround sound speakers) to output audio in addition to or as an alternative to the built-in speakers 320 of the external device 120. As another example, the external device 120 may include one or more microphones. As another example, the external device 120 may not include the orientation and / or movement sensors 325 or such sensors may not be used when outputting audio from the external device 120. In some instances, one or more components of the external device 120 (e.g., the speaker(s) 320, the network interface 315, and / or the like) may be different than that of the wearable device 105.

[0045] FIG. 4 is a hardware block diagram of the cloud-based computing device 130 according to some example instances. As illustrated, in some instances, the cloud-based computing device 130 may include the same or similar components shown in FIG. 2 with respect to the wearable device 105 (e.g., an electronic processor 405, a memory 410, and a network interface 415). Such components of the cloud-based computing device 130 may be similar to the components described above with respect to the wearable device 105 and may perform similar general functions. In some instances, the cloud-based computing device 130 includes fewer or additional components in configurations different from that illustrated in FIG.4 and / or optionally combines two or more components shown in FIG. 4. For example, the cloud-based computing device 130 may include more or fewer instances of the components shown in FIG. 4 and / or one or more of the components may be combined with each other. In some instances, one or more components of the cloud-based computing device 130 (e.g., the network interface 315 and / or the like) may be different than that of the wearable device 105 and / or the external device 120.

[0046] As mentioned previously herein, the methods described herein may be performed by any one or a combination of electronic processors 205, 305, 405 located in a single downstream device 105, 120, 130 or in a combination of downstream devices 105, 120, 130. In other words, an electronic processor 205, 305, 405 that performs the methods disclosed herein may include one or more electronic processors 205 within the wearable device 105 that together function toD24019WO01control various aspects of the wearable device 105. In some instances, the electronic processor 205, 305, 405 is implemented within a distributed system including one or more components located in different devices. For example, in some instances, the wearable device 105 (e.g., the electronic processor 205) includes local hardware components and one or more external hardware components (e.g., one or more electronic processors 305, 405 of the external device 120 and / or the cloud-based computing device 130). In other words, the electronic processor 205, 305, 405 that performs the methods described herein may include any one or a combination of electronic processors located within a single downstream device (e.g., the wearable device 105, the external device 120, or the cloud-based computing device 130) or distributed among various downstream devices and / or systems. Thus, in the claims, if an apparatus or system is claimed, for example, as including an electronic processor or other element configured in a certain manner, for example, to make multiple determinations (or if a method step is claimed as being performed by an electronic processor), the claim or claim element should be interpreted as meaning one or more electronic processors (or other element) where any one of the one or more electronic processors (or other element) is configured as claimed, for example, to make some or all of the multiple determinations. To reiterate, those electronic processors and processing may be distributed within a single device or across multiple devices.

[0047] FIG. 5 illustrates a block diagram 500 of audio classification and playback implemented by the electronic processor 205, 305, 405 of a downstream device 105, 120, 130 according to some example instances. As shown in FIG. 5, the downstream device 105, 120, 130 receives input audio 505 (e.g., two-channel input audio 505), for example, from an upstream device such as the external device 120, the cloud-based computing device 130, an audio server, or the like. The input audio 505 may include any type of audio including binaural / binauralized audio or non-binaural / non-binauralized audio. The non-binaural / non-binauralized audio may include, for example, mono audio, stereo audio, multi-channel audio, and / or spatial audio such as object-based audio. In some instances, binaural audio refers to audio captured with a specific technique (e.g., binaural microphones in ear canals of a subject head or dummy head).However, binaural audio or binauralized audio may also refer to audio captured in another manner that is processed (e.g., via binauralization) into a pair of ear signals, for example, to create a set of signals the simulates the signals received by the binaural microphones). In some instances, binauralization may refer to a process of generating binaural signals (e.g., ear signals) from other audio. For example, as explained previously herein, binauralization may involve processing an input audio signal with a set of filters to simulate how sounds reach a listeners’ ears from different directions. These filters may be head related transfer functions (HRTFs),D24019WO01which when applied to an input signal, introduce directional cues (e.g., level and phase modifications) which during playback create a perception that a sound source is arriving from a specific direction relative to the listener (e.g., directly in front of the listener, front-left and above the listener, etc.).

[0048] The downstream device 105, 120, 130 is configured to perform multiscale attentionbased spatial phase feature identification and / or extraction 510 on the input audio 505 to be used for classification 515 of the input audio. In some instances, the multiscale attention-based spatial phase feature identification and / or extraction 510 includes performing a band split interchannel phase difference (ICPD) statistical averages calculation 520. The band split ICPD statistical averages calculation 520 may include computing, with the electronic processor 205, 305, 405, band split ICPDs based on the input audio 505. The band split ICPD statistical averages calculation 520 may also include computing, with the electronic processor 205, 305, 405, statistical averages of the band split ICPDs based on the input audio 505.

[0049] In some instances, at block 520, a one-channel audio signal waveform s[t] (e.g., each channel of the two-channel input audio 505) is divided into windowed, overlapped frames in a time-frequency (TF) domain as S [t][f], where f represents the frequency index, and t represents the time index. S[t][f] may be referred to as “bins.” In some instances, a group of filter banks, such as Equivalent-Rectangular-Bandwidth (ERB), is used to reduce computational resources used during the audio analysis and / or classification process (e.g., at block 520). The statistical averages of the ICPDs may be calculated in each band based on multiple adjacent frames.

[0050] In some instances, a Hybrid-complex quadrature mirror filter (HCQMF) audio transformation is used to convert the audio signal to bands in the frequency domain. In some instances, a final output HCQMF domain signal of each frame is a 77-bin complex signal. In some instances, the 77 bins are further split and grouped into 20 bands based on ERB. In other words, the electronic processor 205, 305, 405 may compute the band split ICPDs based on ERB. In some instances, the multiscale attention-based spatial features calculated at block 530 are independent values across each ERB. The center frequency for 20 ERBs, in accordance with some example instances, is listed in Table 1 below. The center frequency may refer to the frequency where peak frequency response of each band is located.Table 1:Band Index Center Frequency (Hz)D24019WO011 472 1413 2344 3055 4696 6567 8448 10319 121910 131311 187512 262513 337514 431315 543816 675017 862518 1087519 1350020 19500

[0051] In some instances, the band split ICPD statistical averages may be calculated (at block 520) in accordance with Equations 1-3 below. In Equations 1-3 below, L[t] [ ] is one frame time- frequency transformation of the left channel audio, R[t][ ] is the time-frequency representation of the right channel audio, conj[ ] represents the conjunction of a complex value, real{ } represents a calculation of the real part of a complex value, and imag{ } represents a calculation of the imaginary part of a complex value. A^ is the phase difference or angle between left channel audio and right channel audio. A<p[t][f] is the phase difference or angle per t-f bin of left channel audio and right channel audio.Equation 1:fLRReal[band] = real[L[t][f] * conj[R[t][f]]]}ebandD24019WO01= ∑|L[t][f]| * |R[t][f]|* cos(Δφ[t][f])Equation 2:fLRImag[band] = ∑f∈band ∑t imag{L[t][f] * conj[R[t][f]]}‘f eband ^—‘t ^|L[t][ / ]| * |R[t][ / ]|I1f eband* sin(Δφ[t][f])Equation 3: ICPD[band] = atan2(fLRImag[band], fLRReal[band])

[0052] Once the time-frequency bins are calculated, Equations 1 and 2 are respectively used to calculate a real part and an imaginary part of the complex correlation between the left audio channel and the right audio channel. In some instances, the correlation between the left audio channel and the right audio channel within a band is calculated based on bins within a band. The phase difference within a band can be calculated using Equation 3. ICPD[band] that is calculated using Equation 3 is the average phase difference or angle within a band of the current frame and several past frames, which is a time frequency statistic measure for the phase difference. ICPD[band] can be derived from / ? as well as the amplitudes of left channel audio and right channel audio using Equations 1-3 above. In some instances, the calculated phase difference may be treated as an average phase difference within a band, which reduces overfitting and / or noise.

[0053] In some instances, the multiscale attention-based spatial phase feature identification and / or extraction 510 may then include performing, with the electronic processor 205, 305, 405, cross-band normalization 525 of the ICPDs. In some instances, the electronic processor 205, 305, 405 performs cross-band normalization 525 by using the phase difference for each band to calculate a ratio of the phase difference in each band to the total average phase difference in all the bands (i.e., all 20 bands). In some instances, performing the cross-band normalization 525 of the ICPDs includes calculating, with the electronic processor 205, 305, 405, a ratio of an absolute value of each band of the ICPDs to an average absolute value of all frequency bands (i.e., all 20 bands) of the ICPDs. In some instances, the electronic processor 205, 305, 405 may perform the cross-band normalization 525 according to Equation 4 below. In some instances, performing the cross-band normalization 525 avoids overfitting of left or right directions of audio objects. In other words, the cross-band normalization 525 aids the electronic processorD24019WO01205, 305, 405 in capturing relative high or low fluctuations of the ICPDs without overfitting of left or right directions of audio objects.Equation 4: ICPD^band] =

[0054] In some instances, the multiscale attention-based spatial phase feature identification and / or extraction 510 may then include calculation 530, with the electronic processor 205, 305, 405, of the multiscale attention-based spatial features in discriminative band clusters as explained herein. In some instances, different scales of attention / aggregation are used to try to capture different scales of phase difference. For example, short-term, middle-term, and longterm scales of aggregation may be used and then combined together. In some instances, calculating the multiscale attention-based spatial features includes calculating, with the electronic processor 205, 305, 405, short-term scales, middle-term scales, and long-term scales of aggregation of the ICPDs using an aggregation function that is based on a cross-band normalized ICPD. In some instances, these scales of aggregation are of low computational complexity and low memory consumption. In some instances, the combination of these scales of aggregation is an exponentially weighted average of an aggregation function where the weights are preset to indicate different time scales, for example, according to Equation 5 below with an example aggregation function shown in Equation 6 below. In other words, in some instances, Equations 5 and 6 below may be used to perform the calculation 530 of the multiscale attention-based spatial features.Equation 5: ICPD^f[band][t] = ascale■ JCPD^[band][t - 1]+(1 - ■ f(JCPDn(’rm[band][t]) Equation 6: f(ICPDnorm[band][t]) = log10(ICPDnorm[band][t] + 1.0)

[0055] In some instances, different a values may be calculated using Equation 7 below, where n represents a stride of samples per frame, fsis the sample rate, and r is a time factor (i.e., in seconds) determining different scales of attention. Typical values of the r for short-term, middle-term, and long-term scales can be r=2 seconds, 4 seconds, and 8 seconds.Equation 7: α = e−n / (f*τ)D24019WO01

[0056] Using the short-term, middle-term, and long-term scales of aggregation explained above (i.e., multiscale aggregation or multiscale attention-based spatial feature identification / extraction) captures short-term patterns of phase differences, middle-term patterns of phase differences, and long-term patterns of phase differences in the input audio signal 505. The electronic processor 205, 305, 405 may then combine each of the different scales of attention / aggregation to capture the difference between different types of audio (e.g., binauralized audio versus non-binauralized audio such as stereo audio, etc.) included in the input audio 505.

[0057] FIGS. 7A, 7B, and 7C illustrate graphs of spatial phase features in each of the 20 frequency bands listed in Table 1 in accordance with one example. Each of FIGS. 7A, 7B, and 7C has a different scale of time on its x-axis such that FIG. 7A represents a short-term scale of aggregation (e.g., 0.02 seconds), FIG. 7B represents a middle-term scale of aggregation (e.g., 0.2 seconds), and FIG. 7C represents a long-term scale of aggregation (e.g., four seconds). In other words, the time window of data in FIG. 7A is the smallest in this example (i.e., the least of amount of phase difference data is being analyzed). On the other hand, the time window of data in FIG. 7C is the largest in this example (i.e., the most amount of phase difference data is being analyzed). In some instances, three scales are respectively preset as 0.0, 0.9, and 0.995 in FIGS.7A, 7B, and 7C to represent the contextual time 0.0s, 0.2s, and 4s separately under a sample rate of 48000 Hz. The y-axis in each graph represents frequency with the center frequency from Table 1 labeled in each graph.

[0058] As illustrated in FIGS. 7A, 7B, and 7C, different aggregation scales capture different contextual patterns of the ICPD. FIGS. 7A, 7B, and 7C show the discriminative ability of the disclosed spatial phase features with different scales under 20 ERBs. The dashed line histogram with top-right-to-bottom-left cross-hatching represents the binauralized audio, and the solid line histogram with top-left-to-bottom-right cross-hatching represents the non-binauralized audio (e.g., stereo audio, etc.). The percentage shown in each graph is a discriminative metric that calculates the classification accuracy of the two distributions under the best threshold in each example.

[0059] Based on observation and as indicated in FIGS. 7A, 7B, and 7C, the larger the preset scale, the more discriminative ability (i.e., higher accuracy) since longer contextual phase structure information will be captured with a larger preset scale (i.e., the larger time window of FIG. 7C generally results in more accurate classification than the short time windows of FIGS.7A and 7B). However, a larger preset scale increases latency. Accordingly, a tradeoff betweenD24019WO01latency and accuracy may be balanced to keep latency lower than any latency requirements (e.g., less than or equal to two seconds of latency) while maximizing accuracy given a latency requirement. Also as shown in FIGS. 7A, 7B, and 7C, the comparison across the bands with different center frequencies demonstrates that the most discriminative / accurate bands range from approximately 400 Hz to approximately 1500 Hz, which corresponds to a band index from 5 to 10 (see Table 1). For example, the highest accuracy percentage and the largest difference / discrimination between the histograms in the graphs of each of FIGS. 7A, 7B, and 7C occurs in the 656 Hz and 844 Hz graphs. Therefore, in some instances, attention to the features in the input audio may focus on the most discriminative / accurate frequency bands of for example, approximately 400 Hz to approximately 1500 Hz, approximately 650 Hz to approximately 850 Hz, approximately 400 Hz to approximately 900 Hz, approximately 650 Hz to approximately 1350 Hz, approximately 800 Hz to approximately 1350 Hz, or the like. In some instances, attention of the multiscale attention-based spatial features captures bands in the above-noted ranges, similar ranges, larger ranges, and / or smaller ranges. In some instances, attention of the multiscale attention-based spatial features may focus on a single frequency band (e.g., 656 Hz, 844 Hz, or the like). For example, depending on processing and / or memory capacity of a device performing the methods described herein and / or depending on latency requirements of the device, the multiscale attention-based spatial features may focus on a single frequency band, a subset of the 20 bands listed in Table 1, different bands than those listed in Table 1, etc. In other words, to reduce complexity and computation, less than 20 frequency bands could be analyzed, for example, by focusing on the most discriminative / accurate bands (e.g., a certain frequency band range) identified above.

[0060] In some instances, short-term, middle-term, and long-term scales are set on a case-by-case basis to control (e.g., optimize) a tradeoff of accuracy and latency (e.g., typical values are over 0.9). In some instances, a set of experience values are 0.99, 0.995, and 0.9973 corresponding to T = 2, 4, and 8 as indicated above.

[0061] In some instances, the percentages shown in the graphs of FIGS. 7A, 7B, and 7C indicate the accuracy of classifying different types of audio (e.g., binaural versus non-binaural) for a feature within the specific band of each graph. Accordingly, combining the calculations for multiple bands increases the accuracy of classifying different types of audio since independent classifications of types of audio are being performed in each frequency band.

[0062] In some instances, the multiscale attention-based spatial features calculated at block 530 are used to perform classification 515 of the input audio 505. In some instances, theD24019WO01electronic processor 205, 305, 405 may be configured to use the multiscale attention-based spatial features to perform classification 515 indicative of whether the input audio 505 is binaural audio or non-binaural audio (e.g., stereo audio, etc.). In some instances, there may be different types of binaural audio such as professional binaural audio and stereo binaural audio. For example, professional binaural audio includes binaural audio that was rendered from an object-based file (e.g., Dolby Atmos™ audio). On the other hand, stereo binaural audio may be rendered by up-mixing stereo audio (i.e., non-binaural audio) and is not professional binaural audio because it was not rendered from an object-based file. In some instances, the electronic processor 205, 305, 405 may be configured to use the multiscale attention-based spatial features to perform classification 515 indicative of whether the input audio 505 is professional binaural audio or stereo-rendered binaural audio. In some instances, the electronic processor 205, 305, 405 may additionally or alternatively be configured to use the multiscale attention-based spatial features to perform classification 515 indicative of whether the input audio includes other types of audio such as other types of binaural audio (e.g., audio captured directly using binaural recording techniques such as using dummy head recording, where a mannequin head is fitted with a microphone in each ear).

[0063] In some instances, the electronic processor 205, 305, 405 may provide the multiscale attention-based spatial features identified at block 530 to an audio classification model during classification 515. The audio classification model may be configured to be executed to perform classification of the input audio 505 to generate classification results. In some instances, the classification model may use the multiscale attention-based spatial features themselves / directly to generate the classification results. In some instances, the classification model may use the multiscale attention-based spatial features in combination with other features of the input audio 505 (e.g., identified using other methods and / or with other devices) to generate the classification results. In some instances, the electronic processor 205, 305, 405 may execute the audio classification model. In other instances, the audio classification model may be executed by a different electronic processor on another device. In some instances, the classification model is a machine learning or deep learning model. In some instances, the machine learning or deep learning model may be trained using training data of pre-classified audio.

[0064] After the classification 515 of the input audio 505 has been executed, the electronic processor 205, 305, 405 may be configured to process the input audio 505 in accordance with classification results. As shown in the example instance of FIG. 5, the electronic processor 205, 305, 405 may execute rendering and post-processing(s) (at block 535) of the input audio 505 based on the classification results from the classification 515 of the input audio 505. TheD24019WO01electronic processor 205, 305, 405 may then execute playback 540 of output audio 545 (e.g., two-channel output audio) on the speakers 220, 320 of the downstream device 105, 120 and / or on the external speakers 335 communicatively coupled to the downstream device 105, 120. In some instances, when the downstream device is the cloud-based computing device 130, the downstream device may provide rendered and post-processed output audio 545 to an end-user downstream device 105, 120 for playback. Accordingly, in such instances, the cloud-based computing device 130 may not itself execute playback 540 of the output audio 545.

[0065] FIG. 6 illustrates a flowchart of a method 600 for classifying audio according to some example instances. The method 600 is described as being performed by the electronic processor 205, 305, 405 as described previously herein. While a particular order of processing steps, message receptions, and / or message transmissions is indicated in FIG. 6 as an example, timing and ordering of such steps, receptions, and transmissions may vary where appropriate without negating the purpose and advantages of the examples set forth in detail throughout the remainder of this disclosure.

[0066] At block 605, the electronic processor 205, 305, 405 receives the input audio 505 as described previously herein with respect to FIG. 5. At block 610, the electronic processor 205 computes band split ICPDs based on the input audio 505 as described with respect to block 520 of FIG. 5. At block 615, the electronic processor 205 computes statistical averages of the band split ICPDs based on the input audio 505 as explained with respect to block 520 of FIG. 5.

[0067] At block 620, the electronic processor 205 performs cross-band normalization of the ICPDs as explained with respect to block 525 of FIG. 5. At block 625, the electronic processor 205 calculates multiscale attention-based spatial features in discriminative band clusters as explained with respect to block 530 of FIG. 5. In some instances, the multiscale attention-based spatial features are used to perform classification 515 of the input audio 505 as explained with respect to block 515 of FIG. 5.

[0068] As indicated in FIG. 6, in some instances, the method 600 may repeat as additional input audio 505 is received to continue classifying the input audio 505 (e.g., classifying one or more additional input audio files for use during playback of the input audio on an end-user device 105, 120).

[0069] The foregoing description, for purpose of explanation, has been described with reference to specific embodiments. However, the illustrative discussions above are not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications andD24019WO01variations are possible in view of the above teachings. The embodiments were chosen and described in order to best explain the principles of the techniques and their practical applications. Others skilled in the art are thereby enabled to best utilize the techniques and various embodiments with various modifications as are suited to the particular use contemplated.

[0070] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):EEE 1. A method for classifying audio, the method comprising: receiving, with an electronic processor, input audio: computing, with the electronic processor, band split inter-channel phase differences (ICPDs) based on the input audio; computing, with the electronic processor, statistical averages of the band split ICPDs based on the input audio; performing, with the electronic processor, cross-band normalization of the band split ICPDs; and calculating, with the electronic processor, multiscale attention-based spatial features in discriminative band clusters, wherein the multiscale attention-based spatial features are used to perform classification of the input audio.EEE 2. The method of EEE 1, further comprising using, with the electronic processor, the multiscale attention-based spatial features to perform classification indicative of whether the input audio is binaural audio or non-binaural audio.EEE 3. The method of EEE 1, further comprising using, with the electronic processor, the multiscale attention-based spatial features to perform classification indicative of whether the input audio is professional binaural audio or stereo-rendered binaural audio.EEE 4. The method of any one of the preceding EEEs, wherein the electronic processor includes one or more electronic processors that are located in at least one of a group consisting of a cloud-based computing device, an external device, a wearable device, and combinations thereof.EEE 5. The method of any one of the preceding EEEs, wherein the band split ICPDs include a plurality of frames of the input audio; and wherein the statistical averages of the band split ICPDs are calculated in each band based on multiple adjacent frames.EEE 6. The method of any one of the preceding EEEs, wherein computing band split ICPDs is based on an equivalent rectangular bandwidth (ERB).D24019WO01EEE 7. The method of EEE 6, wherein the multiscale attention-based spatial features are independent values across each ERB.EEE 8. The method of any one of the preceding EEEs, wherein performing the cross-band normalization of the band split ICPDs includes calculating, with the electronic processor, a ratio of an absolute value of each band of the band split ICPDs to an average absolute value of all frequency bands of the band split ICPDs.EEE 9. The method of any one of the preceding EEEs, wherein calculating the multiscale attention-based spatial features includes calculating, with the electronic processor, short-term scales, middle-term scales, and long-term scales of aggregation of the band split ICPDs using an aggregation function that is based on a cross-band normalized ICPD.EEE 10. The method of EEE 9, wherein the short-term scales, the middle-term scales, and the long-term scales of aggregation include an exponentially weighted average of an aggregation function, and wherein weights of the exponentially weighted average are preset to indicate different time scales.EEE 11. The method of any one of the preceding EEEs, wherein attention of the multiscale attention-based spatial features captures bands from approximately 400 Hz to approximately 1500 Hz.EEE 12. The method of any one of the preceding EEEs, further comprising providing, with the electronic processor, the multiscale attention-based spatial features to an audio classification model, wherein the audio classification model is configured to be executed to perform classification of the input audio to generate classification results.EEE 13. The method of EEE 12, further comprising processing, with the electronic processor, the input audio in accordance with the classification results.EEE 14. An apparatus including the electronic processor and a memory storing instructions, which when executed by the electronic processor, cause the apparatus to perform the method of any one of EEEs 1-13.D24019WO01EEE 15. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any one of EEEs 1-13.

[0071] Although the disclosure and examples have been fully described with reference to the accompanying drawings, it is to be noted that various changes and modifications will become apparent to those skilled in the art. Such changes and modifications are to be understood as being included within the scope of the disclosure and examples as defined by the claims.

[0072] Various features and advantages are set forth in the following claims.

Claims

D24019WO01CLAIMSWhat is claimed is:

1. A method for classifying audio, the method comprising:receiving, with an electronic processor, input audio;computing, with the electronic processor, band split inter-channel phase differences (ICPDs) based on the input audio;computing, with the electronic processor, statistical averages of the band split ICPDs based on the input audio:performing, with the electronic processor, cross-band normalization of the band split ICPDs; andcalculating, with the electronic processor, multiscale attention-based spatial features in discriminative band clusters, wherein the multiscale attention-based spatial features are used to perform classification of the input audio.

2. The method of claim 1, further comprising using, with the electronic processor, the multiscale attention-based spatial features to perform classification indicative of whether the input audio is binaural audio or non-binaural audio.

3. The method of claim 1, further comprising using, with the electronic processor, the multiscale attention-based spatial features to perform classification indicative of whether the input audio is professional binaural audio or stereo-rendered binaural audio.

4. The method of any one of the preceding claims, wherein the electronic processor includes one or more electronic processors that are located in at least one of a group consisting of a cloud-based computing device, an external device, a wearable device, and combinations thereof.

5. The method of any one of the preceding claims, wherein the band split ICPDs include a plurality of frames of the input audio; andwherein the statistical averages of the band split ICPDs are calculated in each band based on multiple adjacent frames.

6. The method of any one of the preceding claims, wherein computing band split ICPDs is based on an equivalent rectangular bandwidth (ERB).D24019WO017. The method of claim 6, wherein the multiscale attention-based spatial features are independent values across each ERB.

8. The method of any one of the preceding claims, wherein performing the cross-band normalization of the band split ICPDs includes calculating, with the electronic processor, a ratio of an absolute value of each band of the band split ICPDs to an average absolute value of all frequency bands of the band split ICPDs.

9. The method of any one of the preceding claims, wherein calculating the multiscale attention-based spatial features includes calculating, with the electronic processor, short-term scales, middle-term scales, and long-term scales of aggregation of the band split ICPDs using an aggregation function that is based on a cross-band normalized ICPD.

10. The method of claim 9, wherein the short-term scales, the middle-term scales, and the long-term scales of aggregation include an exponentially weighted average of an aggregation function, and wherein weights of the exponentially weighted average are preset to indicate different time scales.

11. The method of any one of the preceding claims, wherein attention of the multiscale attention-based spatial features captures bands from approximately 400 Hz to approximately 1500 Hz.

12. The method of any one of the preceding claims, further comprising providing, with the electronic processor, the multiscale attention-based spatial features to an audio classification model, wherein the audio classification model is configured to be executed to perform classification of the input audio to generate classification results.

13. The method of claim 12, further comprising processing, with the electronic processor, the input audio in accordance with the classification results.

14. An apparatus including the electronic processor and a memory storing instructions, which when executed by the electronic processor, cause the apparatus to perform the method of any one of claims 1-13.D24019WO0115. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any one of claims 1-13.