Metadata-based audio signal processing

Metadata-based audio signal processing adapts audio output to content type, addressing the inadequacies of conventional systems by ensuring optimal settings are applied automatically, enhancing user experience and device compatibility.

WO2026161219A1PCT designated stage Publication Date: 2026-07-30BOSE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BOSE CORP
Filing Date
2026-01-08
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Conventional audio systems fail to effectively control signal processing based on audio data, leading to inadequate adaptation of audio output to different content types.

Method used

Implementing metadata-based audio signal processing that evaluates metadata to determine appropriate output capabilities and applies equalization, spatialization, or anchoring adjustments to enhance audio output based on content type, using processors and metadata transport mechanisms.

Benefits of technology

Enhances user experience by ensuring audio output aligns with content creator intent, providing optimized audio settings automatically without user intervention, and improving compatibility across various audio devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026010572_30072026_PF_FP_ABST
    Figure US2026010572_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Various implementations include audio devices and methods for controlling audio output. Certain implementations include a method of controlling audio device output, including: receiving an audio file or audio stream for output at an audio device, the audio file or audio stream provided with metadata defining audio output settings, wherein the audio file or audio stream is in a standard format, evaluating the metadata for corresponding output capabilities at the audio device, and if the audio device does not have corresponding output capabilities defined by the metadata, outputting the audio file or audio stream in the standard format.
Need to check novelty before this filing date? Find Prior Art

Description

METADATA-BASED AUDIO SIGNAL PROCESSINGPRIORITY CLAIM

[0001] This application claims priority to US Patent Application No. 19 / 238,649 filed on June 16, 2025 and US Provisional Application No. 63 / 749,306 filed January 24, 2025, the entire contents of each of which are hereby incorporated by reference.TECHNICAL FIELD

[0002] This disclosure generally relates to audio systems. More particularly, the disclosure relates to metadata-based audio signal processing in audio systems.BACKGROUND

[0003] Certain types of audio content can benefit from particular signal processing, for example, equalization, volume control, mode control, etc. However, certain conventional systems do not effectively control signal processing based on audio data.SUMMARY

[0004] All examples and features mentioned below can be combined in any technically possible way.

[0005] Various implementations include audio devices and methods for controlling audio output. Certain implementations include a method of controlling audio device output, including: receiving an audio file or audio stream for output at an audio device, the audio file or audio stream provided with metadata defining audio output settings, wherein the audio file or audio stream is in a standard format, evaluating the metadata for corresponding output capabilities at the audio device, and if the audio device does not have corresponding output capabilities defined by the metadata, outputting the audio file or audio stream in the standard format.

[0006] Particular implementations include a device having a processor configured to: receive an audio file or audio stream for output at an audio device, the audio file or audio stream provided with metadata defining audio output settings, wherein the audio file or audio stream is in a standard format, evaluate the metadata for corresponding output capabilities at the audio device, and if the audio device does not have corresponding output capabilities defined by the metadata, output the audio file or audio stream in the standard format. In some cases, the device includes an audio device having an electro-acoustic transducer coupled with the processor.PN-25-007-WO Page 1 of 22

[0007] Additional particular implementations include a method that includes: classifying a library of audio files or audio streams according to content type; and appending the library with metadata classifiers for subsequent playback according to the metadata in a standard audio format.

[0008] Implementations may include one of the following features, or any combination thereof.

[0009] In some aspects, the metadata includes at least one of descriptive metadata or prescriptive metadata.

[0010] In some aspects, the metadata includes both descriptive metadata and prescriptive metadata. In some examples, where the metadata includes both descriptive metadata and prescriptive metadata and a conflict exists between the descriptive metadata and prescriptive metadata, the prescriptive metadata controls over the descriptive metadata.

[0011] In some aspects, the method further includes, if the metadata includes prescriptive metadata, applying specified signal processing instructions to audio output at the audio device according to the prescriptive metadata.

[0012] In some aspects, the prescriptive metadata specifies how to render content in the audio file or audio stream with the specified signal processing instructions that are configured to override conflicting default audio output settings for the audio device. In one example, prescriptive metadata includes: {FREQCE} = Apply Flat EQ and Center Channel Extraction mode regardless of content type.

[0013] In some aspects, the method further includes, if the metadata includes descriptive metadata and not prescriptive metadata, selecting an output mode for the audio output at the audio device from a predefined set of output modes.

[0014] In some aspects, the descriptive metadata describes the nature of content in the audio file or the audio stream. In some examples, selecting an output mode is based on heuristics or predefined rules according to the nature of the content. In some examples, the nature of the content can include one or more of: Music (MU), Podcast (PO), Movie (MO), Spoken Word (SW), Episodic (EP), or Gaming (GA). In a particular example, where descriptive metadata includes music (MU), the output mode can include applying default tuning, and avoiding pausing during interruptions.

[0015] In some aspects, if the audio device has corresponding output capabilities defined by the metadata, outputting the audio file or audio stream according to the metadata is performed by applying at least one of: an equalization adjustment, a spatialization adjustment, a virtualization adjustment, or an anchoring adjustment to the audio file or audio stream.PN-25-007-WO Page 2 of 22

[0016] In some aspects, the standard format includes at least one of a stereo format or a stereo transport format. In particular examples, the standard format may explicitly exclude objectbased audio formats. In such cases, the processor is configured to process only channel-based stereo or surround signals, or to ignore object metadata that conflicts with defined metadata.

[0017] In some aspects, the metadata includes a set of indicators of audio output settings based on audio output device type.

[0018] In some aspects, the set of indicators includes two or more indicators for assigning audio output settings in response to detecting two or more audio output devices, where the set of indicators is applied differently to distinct audio output devices. In some examples, the audio output devices include stereo paired speakers, grouped speakers, speakers paired with an openear wearable audio device, a soundbar paired with a wearable audio device, etc. In additional implementations, the metadata may include synchronization markers and / or role designations, e.g., to enable coordinated playback across a set of two or more devices. For example, a wearable audio device may be instructed to render only spatial effects while a paired soundbar renders dialog or low-frequency effects, with both synchronized to a shared clock or frame reference.

[0019] In some aspects, the method further includes receiving an input from a user interface to disable the metadata-based evaluation for corresponding output capabilities at the audio device, and outputting the audio file or audio stream in the standard format in response to receiving the input.

[0020] In some aspects, the evaluating is performed at the audio device.

[0021] In some aspects, the audio device has limited processing capability. In certain examples, the audio device with limited processing capability includes a wearable audio device or a portable audio device.

[0022] In some aspects, the metadata is provided with a metadata transport mechanism.

[0023] In some aspects, the metadata is provided via at least one of: file-embedded tags, Bluetooth transport protocols, Wi-Fi, general wireless protocols, non-audible transports such as ultrasound, encoded carrier signals in the audio content, or a companion data channel associated with the audio stream. In some examples, the metadata is appended to the title of the content. In certain examples, the content type is indicated by a set of content-based features represented by a unique two-letter code. In some examples, the content type is one of a group of predefined content types.

[0024] In some aspects, evaluating the metadata is performed automatically without user input.PN-25-007-WO Page 3 of 22

[0025] In some aspects, the method further includes providing a digital audio workstation (DAW) plug-in with a menu of metadata assignment options.

[0026] In some aspects, the menu of metadata assignment options includes options for assigning at least one of prescriptive metadata or descriptive metadata to each audio file or audio stream. For example, the menu of metadata assignment options can include options for assigning prescriptive metadata to each audio file or stream such as defining equalization, distribution, virtual speakers, level, spatialization, or anchoring.

[0027] In some aspects, the method further includes: classifying a library of audio files or audio streams according to content type; and appending the library with metadata classifiers for subsequent playback according to the metadata in a standard audio format.

[0028] In some aspects, the audio device includes at least one of: an audio headset, a fixed speaker, a portable speaker, or a vehicle audio system.

[0029] In some cases, the group of content types is predefined.

[0030] In some aspects, the controller is further configured to select the audio output mode based on at least one secondary factor including, user movement, proximity to a multimedia device, proximity to an external speaker, proximity to another wearable audio device, or presence in a vehicle.

[0031] Two or more features described in this disclosure, including those described in this summary section, may be combined to form implementations not specifically described herein.

[0032] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features, objects and advantages will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG. 1 is a data flow diagram illustrating aspects of metadata-based audio control according to various implementations.

[0034] FIG. 2 is a flow diagram illustrating processes in a method according to various implementations.

[0035] FIG. 3 is a flow diagram illustrating processes in a method according to various additional implementations.

[0036] It is noted that the drawings of the various implementations are not necessarily to scale. The drawings are intended to depict only typical aspects of the disclosure, and therefore should not be considered as limiting the scope of the implementations. In the drawings, like numbering represents like elements between the drawings.PN-25-007-WO Page 4 of 22DETAILED DESCRIPTION

[0037] This disclosure is based, at least in part, on the realization that adaptively controlling audio output based on content can enhance the user experience.

[0038] Additional details of content-based audio control are described for example in US Patent Application No. 18 / 238,668 (Content-Based Audio Spatialization, filed August 28, 2023) and US Patent Application No. 19 / 027,307 (Audio Device with Machine Learning (ML) Based Content Detection, filed January 17, 2025), the entire contents of each of which is incorporated by reference herein.

[0039] Particular aspects are described in the context of wearable devices such as those provided by Bose Corporation (Framingham, MA, USA), for example, Bose Ultra Open Earbuds, Bose QuietComfort Earbuds, Bose QuietComfort Ultra Headphones, Bose QuietComfort Headphones; as well as speakers provided by Bose Corporation, for example, Bose Smart Soundbar varieties, television speakers (e.g., the Bose TV Speaker), home theater speakers (e.g., the Bose Surround Speaker varieties and / or Bose Bass Module varieties), portable speakers such as a portable smart speaker (e.g., one of the Bose Soundlink varieties or the Bose Portable Smart Speaker) or a portable professional speaker such as the Bose SI Pro Portable Speaker. In further implementations, audio devices as described herein can include vehicle audio systems such as speaker systems in an automobile, electric vehicle, or other transport vehicle. Additional audio devices can include pass-through or control devices such as amplifiers and / or switches.

[0040] Certain conventional audio systems employ audio control and other audio output adjustments without consideration for content type. Because different types of content can benefit from audio control (e.g., mode control, volume control, spatialization, virtualization, orchestration, etc.) in distinct ways, these conventional systems can have deficiencies.

[0041] The systems and methods disclosed according to various implementations use content metadata to adaptively control audio output. A particular approach includes automatically selecting signal processing for audio output based on content metadata. In some aspects, automatically selecting the output mode is performed without user input.

[0042] Commonly labeled components in the FIGURES are considered to be substantially equivalent components for the purposes of illustration, and redundant discussion of those components is omitted for clarity.

[0043] Various implementations include audio devices and methods for controlling audio output. Certain implementations include a method of controlling audio device output. FIG. 1 isPN-25-007-WG Page 5 of 22a schematic data flow diagram illustrating an audio device 10 interfacing with an additional device 20 according to various implementations. In some cases, the additional device 20 is referred to as a source device and / or a smart device, e.g., an electronic device having network communication capabilities and processing capabilities. It is understood that in some cases, the audio device 10 can include functions described relative to the device 20, e.g., where an audio file or audio stream is stored or otherwise supplied by an integrated device 20 at the audio device 10. In other cases, the device 20 is separate (e.g., physically separate) from the audio device 10, e.g., where device 20 is a smart device such as a smart phone, tablet, computing device, amplifier unit, etc. In particular examples, the audio device 10 includes at least one electro-acoustic transducer 30 (e.g., a single driver or an array of drivers) and a processor 40 coupled with the transducer(s) 30. In additional, optional implementations, the audio device 10 includes one or more microphones 50 coupled with the processor 40. In some optional implementations, the audio device 10 includes an interface 60, e.g., a user interface enabling inputs and / or outputs such as one or more buttons, touch screens, voice command interfaces, etc. Additional implementations include a communications unit (or interface) 70 configured to communication with additional devices (e.g., device 20, and additional audio devices 10a, 10b, 10c, etc.) via one or more conventional communication protocols, e.g., Wi-Fi, Bluetooth (BT), BT Low Energy (BLE), broadcast, SimpleSync (developed by Bose Corporation, Framingham, MA, USA), general wireless protocols, non-audible transports such as ultrasound, via encoded carrier signals, direct (i.e., wired) connection, etc. Additional, optional electronics 80 at the audio device 10 can include sensors such as orientation sensors, optical sensors, capacitive sensors, etc., as well as communications equipment. In some cases where the (e.g., source) device 20 is paired with audio device 10, those devices are referred to as a source and sink, respectively. In some cases, the audio device 10 can be directly connected (e.g., via a network connection) to an audio service 100, e.g., an audio file and / or streaming platform 100, or can be connected to the audio service 100 via device 20. Additional audio devices 10A, 10B are also illustrated as examples of further devices that could be present in a given space, e.g., multiple speakers in a room, home, office, house of worship, or other venue. Further, a set of audio devices 10 can form a system such as a vehicle audio system or an installed audio system in a venue.

[0044] In particular cases, the processor 40 includes a chip or chipset that is configured to run a metadata-based audio control program (cont. program) 90 to control output functions at one or more audio devices 10, e.g., in a space such as in a room, a home, an office, meeting space, etc. As shown in FIG. 1, the processor 40 (and control program 90) functions can be performed at the audio device(s) 10 and / or at source device 20.PN-25-007-WO Page 6 of 22

[0045] FIG. 2 shows a flow diagram illustrating processes in a method performed by the processor 40 (e.g., running control program 90) to control output at audio device(s) 10 based on audio file or stream metadata. With reference to FIGS. 1 and 2, the control program 90 is configured to:

[0046] Pl : receive an audio file or audio stream (file or stream) 120 for output at an audio device 10. As noted herein, in certain cases, the audio file or audio stream 120 is stored locally at the audio device 10 and / or source device 20, or can be accessed via an audio service 100 (e.g., over a network connection). In particular implementations, the audio file or audio stream 120 is provided with metadata 130 (e.g., via one or more metadata transport mechanisms described herein) defining audio output settings. In some cases, the audio file or audio stream 120 is in a standard format. In some aspects, the standard format includes a stereo format and / or a stereo transport format. In particular cases, the standard format excludes object-based formatting, or the processor 40 is otherwise configured to ignore object-based formatting. In particular examples, the standard format may explicitly exclude object-based audio formats such as Dolby Atmos (which can be carried by Dolby Digital, or Dolby Digital Plus), or MPEG-H. In such cases, the processor is configured to process only channel-based stereo or surround signals, or to ignore object metadata that conflicts with (e.g., SonicSync-defined) metadata.

[0047] In further non-limiting example implementations, such as where the audio device 10 is part of an out-loud speaker system, the processor 40 can be configured to act in a multi-channel rendering configuration at least in part using guided metadata, e.g., if such metadata overlaps with the description of an object-based Tenderer.

[0048] D2: evaluate the metadata 130 for corresponding output capabilities at the audio device 10. In particular cases, evaluating the metadata 130 is performed automatically without user input. In some examples, the metadata 130 is evaluated in response to receiving the audio file or stream 120. In further examples, metadata 130 is automatically evaluated when one or more operating modes is enabled and / or active, e.g., a metadata-based control operating mode. In some cases, the metadata-based audio output control can be disabled, e.g., via a user interface command. For example, in response to receiving an input from a user interface (e.g., interface 70) to disable the metadata-based evaluation for corresponding output capabilities at the audio device 10, the processor 40 is configured to output the audio file or audio stream 120 in the standard format. In still further implementations, select aspects of metadata-based control operating modes can be disabled, e.g., based on user preferences. For example, a user may wish to override one or more aspects of an output mode without necessarily disabling the metadatabased audio output control altogether. In a non-limiting example, a content creator may usePN-25-007-WO Page 7 of 22metadata 130 to specify an immersion mode (FR) with a Flat Response EQ, but the user may wish to only override the Flat Response EQ, e.g., with their Custom EQ.

[0049] Returning to FIG. 2, following evaluation of the metadata 130 for corresponding output capabilities at the audio device 10, process P3 includes: if the audio device 10 does not have corresponding output capabilities defined by the metadata 130 (No to D2), output the audio file or audio stream 120 in the standard format; and

[0050] P4: if the audio device 10 does have corresponding output capabilities defined by the metadata 130 (Yes to D2), output the audio file or audio stream 120 according to the metadata 130. In some aspects, if the audio device 10 has corresponding output capabilities defined by the metadata, outputting the audio file or audio stream 120 according to the metadata 130 includes applying at least one of: an equalization adjustment, a spatialization adjustment, a virtualization adjustment, or an anchoring adjustment to the audio file or audio stream 120.

[0051] In some cases, evaluating the metadata 130 in decision D2 includes evaluating the metadata 130 based on a type of metadata present (if any). FIG. 3 shows a flow diagram illustrating processes in evaluating the metadata 130 according to various implementations. In some aspects, the metadata 130 includes descriptive metadata 140 and / or prescriptive metadata 150. In particular aspects, the metadata 130 includes both descriptive metadata 140 and prescriptive metadata 150.

[0052] In a first process (P100), the processor 40 (e.g., running control program 90) evaluates the metadata 130 for descriptive metadata 140 and / or prescriptive metadata 150.

[0053] In some aspects, the descriptive metadata 140 describes the nature of content in the audio file or the audio stream 120. In some non-limiting examples, the nature of the content can include one or more of: Music (MU), Podcast (PO), Movie (MO), Spoken Word (SW), Episodic (EP), or Gaming (GA). Distinct from descriptive metadata 140, prescriptive metadata 150 specifies how to render content in the audio file or audio stream 120 with specified signal processing instructions that are configured to override conflicting default audio output settings for the audio device 10.

[0054] Based on the evaluation of the metadata 130, the processor 40 is configured to take one or more actions (at decision DI 10) to control audio output of the audio file or stream 120. As indicated in FIG. 2, if the metadata 130 does not include descriptive 140 or prescriptive 150 metadata, the processor 40 is configured to output the audio file or audio stream 120 in the standard format, as shown in process (P3).

[0055] If the metadata 130 includes prescriptive metadata 150, in process Pl 20, specified signal processing instructions are applied to audio output at the audio device 10 according to thePN-25-007-WO Page 8 of 22prescriptive metadata. In one example, prescriptive metadata 150 includes: {FREQCE} = Apply Flat EQ and Center Channel Extraction mode regardless of content type. Additional examples include {FREQLR} to apply Flat EQ and Large Room simulation (prescriptive), with an optional {MO} tag to indicate Movie content (descriptive). In certain examples where prescriptive metadata takes precedence over descriptive metadata, the processor 40 will follow the rendering behavior defined by {FREQLR}, regardless of any default behavior typically associated with the {MO} content type. In another example, {DCEQPO} would apply Default EQ (prescriptive) while also signaling that the content is a podcast (descriptive). In this case, the prescriptive EQ setting will control playback.

[0056] In some cases, if the metadata 130 includes descriptive metadata 140 and not prescriptive metadata 150, process Pl 30 includes selecting an output mode for the audio output at the audio device 10 from a predefined set of output modes. In some examples, selecting the output mode is based on heuristics or predefined rules according to the nature of the content. In a particular example, where descriptive metadata includes music (MU), the output mode can include applying default tuning, and avoiding pausing during interruptions.

[0057] In some aspects, the metadata 130 includes both descriptive metadata 140 and prescriptive metadata 150. In such cases, prescriptive metadata 150 can control over descriptive metadata 140, e.g., if a conflict exists, as illustrated in optional process P140 in FIG. 3. In such cases, the prescriptive metadata is interpreted as explicit creator intent and is configured to override any conflicting default or descriptive-based behaviors. For example, if descriptive metadata indicates the content is a podcast (suggesting pause-on-interruption behavior), but prescriptive metadata includes {DS} for immersive disabled, the prescriptive setting takes precedence, disabling immersive features regardless of the default podcast profile.

[0058] In further implementations, the metadata 130 includes a set of indicators of audio output settings based on audio output device type, e.g., whether the audio device(s) 10 includes open-ear (non-occluding) headphones, occluding headphones (e.g., earbuds or on-ear headphones), a portable speaker, a soundbar, a paired speaker, an entertainment audio system, a vehicle audio system, etc. In some aspects, the set of indicators includes two or more indicators for assigning audio output settings in response to detecting two or more audio output devices 10. For example, the set of indicators is applied differently to distinct audio output devices 10, e.g., different indicators are applied to occluding headphones, as compared with stereo paired speakers, as compared with a paired soundbar and non-occluding headset. In some examples, the audio output devices 10 include stereo paired speakers, grouped speakers, speakers paired with an open-ear wearable audio device, a soundbar paired with a wearable audio device, anPN-25-007-WO Page 9 of 22open-ear wearable audio device paired with a vehicle audio system, etc. In additional implementations, the metadata 130 may include synchronization markers and / or role designations to enable coordinated playback across a set of two or more devices 10. For example, a wearable audio device 10 may be instructed to render only spatial effects while a paired soundbar renders dialog or low-frequency effects, with both synchronized to a shared clock or frame reference.

[0059] As noted herein, particular implementations include evaluating metadata 130 for characteristics indicative of audio device output. In some cases, the metadata 130 is evaluated by the processor(s) 40 at audio device 10. In particular examples, such as where the audio device 10 includes a wearable audio device or portable audio device, that audio device 10 can have limited processing capability. In such cases, the control program 90 enables metadatabased analysis of the audio file or stream 120 with relatively limited processing capability. In other cases, the metadata 130 is evaluated by the processor(s) 40 at device 20, e.g., a smart device and / or source device.

[0060] Further, as noted herein, metadata 130 about the audio file or stream 120 can be provided in any of a number of mechanisms. For example, the metadata 1 0 can be provided with a metadata transport mechanism including one or more of: file-embedded tags, Bluetooth transport protocols. Wi-Fi, general wireless protocols, non-audible transports such as ultrasound, encoded carrier signals in the audio content, or a companion data channel associated with the audio file or stream 120. In some examples, the metadata 130 is appended to the title of the content (e.g., content in the audio file or stream 120). In certain examples, the content type is indicated by a set of content-based features represented by a unique two-letter code (e.g., MU, PO, SW, etc.). In some examples, the content type is one of a group of predefined content types.

[0061] In particular cases, the processor 40 is configured to analyze metadata 130 across multiple metadata transport paths, including both embedded and out-of-band mechanisms. In further non-limiting examples, the processor 40 is configured to analyze metadata 130 across at least the following metadata transport (or, delivery) paths:

[0062] (I) Classic Bluetooth (BR / EDR): (a) AVRCP (e.g., title, artist, album), which can include standard media metadata, with limited customization; (b) A2DP, which supports ID3 tag passthrough in MP3; and metadata embedded in audio stream; and (c) HFP, which includes call-related metadata (e.g., name, number); and is relevant for hands-free mode.

[0063] (II) Bluetooth Low Energy (BLE): (a) GATT, which includes custom services and characteristics for spatial and / or audio effect metadata; (b) LE Audio + MCS, which includesPN-25-007-WO Page 10 of 22low-latency control over playback metadata and commands; and (c) Periodic Advertising w / AUX Data, which includes venue broadcast scenarios (e.g., theaters, public spaces).

[0064] (ITT) Embedded Metadata in Audio Files: (a) ID3 tags (MP3 / AAC), which supports descriptive and prescriptive fields; (b) Ogg / Vorbis comments, which is beneficially customizable for FLAC / Opus; and (c) ADTS (AAC), including in-band signaling in live broadcasts.

[0065] (IV) Hybrid / Advanced Formats: (a) BEE + AVRCP, including mixed transport, for example, BEE for advanced metadata, AVRCP for core fields; (b) BLE + LE Audio ISO Channels, for example using synchronized metadata via isochronous channels; and (c) Codecnative metadata (e.g., Dolby Atmos) embedded spatialization data in the audio bitstream.

[0066] While various implementations describe selecting audio output settings (e.g., signal processing settings and / or audio output modes) based on detected characteristics of metadata 130, additional implementations include selecting the audio output settings and / or modifying the audio output settings based on at least one secondary factor detectable at the audio device 10 and / or device 20. In some cases, the secondary factor includes one or more of: user movement (e.g., as indicated by sensors in electronics 80 at audio device 10), proximity to a multimedia device (e.g., as indicated by communications-based proximity detection such as BT signal strength, common Wi-Fi network detection, etc.), proximity to an external speaker (e.g., as indicated by communications-based proximity detection such as BT signal strength, BT Channel Sounding, common Wi-Fi network detection, etc.), proximity to another wearable audio device (e.g., as indicated by communications-based proximity detection such as BT signal strength, previous device pairing, etc.), or presence in a vehicle (e.g., as indicated by communications-based proximity detection such as BT signal strength, previous device pairing, etc.).

[0067] In further implementations, a digital audio workstation (DAW) plug-in 200 is provided (FIG. 1), e.g., via a device 20 and / or in connection with service 100, enabling a user to assign metadata to audio files or streams 120. In particular cases, the DAW plug-in 200 includes a menu of metadata assignment options. In some aspects, the menu of metadata assignment options includes options for assigning at least one of prescriptive metadata or descriptive metadata to each audio file or audio stream 120. For example, the menu of metadata assignment options can include options for assigning prescriptive metadata to each audio file or stream 120 such as defining equalization, distribution, virtual speakers, level, spatialization, or anchoring. In certain cases, the DAW plug-in 200 can be provided as an audio development and / or editing tool for a user creating, categorizing, or otherwise packaging audio files or streams 120, e.g., for distribution.PN-25-007-WO Page 11 of 22

[0068] In still further implementations, a method includes: (A) classifying a library of audio files or audio streams 120 according to content type (e.g., Music (MU), Podcast (PO), Movie (MO), Spoken Word (SW), Episodic (EP), or Gaming (GA)); and (B) appending the library with metadata classifiers for subsequent playback according to the metadata 130 in a standard audio format. In some cases, the group of content types is predefined, e.g., according to a limited list of content types such Music (MU), Podcast (PO), Movie (MO), Spoken Word (SW), Episodic (EP), and Gaming (GA).

[0069] Example Modes

[0070] As noted herein, particular implementations include metadata-based audio control, e.g., signal processing. In particular implementations, content creators and / or distributors can select or otherwise designate predefined character strings (e.g., in one or more metadata transport paths) to provide context for signal processing of content at an audio device 10 (e.g., wearable audio device such as headphones). Signal processing can be applied based on one or more modes, and can be device 10 (e.g., wearable device type, model, etc.) specific. Modes can be invoked automatically based on metadata 130. In some cases, the selected mode (or other audio output preference) can be overridden, e.g., via a user command. The disclosed approaches can aid in delivering audio content as the creator intended. Further, the disclosed approaches can ensure that users do not receive an unintended degraded audio experience (e.g., double spatialization).

[0071] In particular cases, as noted herein, output modes are predefined. In other examples, modes can be created or assigned on content-by-content basis, and / or created as hybrids of existing predefined modes. Example modes can include (among others), immersive, loudspeaker, and EQ. Other modes are also possible.

[0072] Immersive Modes: In some cases, immersive modes describe targeted processing modes to be invoked on an audio device 10 for a desirable listening experience. These modes enable content creators to tailor the playback experience according to the type and preparation of the content. Immersive modes can be disabled, or enabled to provide immersive content. When enabled, examples of immersive content can include (among others):

[0073] Full Rotation: Beneficial for content that has already been binaurally prepared, such as binaural audio recordings, immersive headphone mixes, or other mixes intended for headphone playback. This mode allows user interactivity by keeping the audio image oriented to the space around the user as she turns her head.

[0074] Center Channel Extraction: Separates the center-image content (such as dialogue) and anchors it to the space around the listener while allowing all other content outside the centerPN-25-007-WO Page 12 of 22image to move with the listener's head movements. This hybrid experience can be beneficial for pre-binaurally encoded content, like object-based (e.g., ATMOS) headphone-rendered movies with dialog elements in the center channel.

[0075] Loudspeaker Presentation: Virtualizes stereo to speakers placed + / - 30 degrees from center, maintaining the listener in desired zone (or “sweet spot(s)”) with IMU inputs (e.g., from the wearable audio device electronics). This can be beneficial for content intended to be rendered on loudspeakers. Loudspeaker Presentation can include (among others):

[0076] (I) Small room: Can be beneficial for content mixed for playback in an intimate setting. The room size setting determines the length of reflections in the simulated virtual room.

[0077] (II) Large room: This can be beneficial for content mixed for playback in a larger, open setting. The room size setting determines the length of reflections in the simulated virtual room.

[0078] EQ Modes: EQ modes can allow content creators to set the desired equalization for a “reference” listening experience, which can include (in non-limiting examples):

[0079] Flat Response: EQ settings providing a flat frequency response, ensuring that the audio is as accurate as possible and closely reflects the original studio mix; and

[0080] Default / User Controlled: This mode defaults to either the device (e.g., wearable) EQ tuning or the user-selected EQ tuning.

[0081] As noted herein, content types may include but are not limited to: Music (MU), Podcast (PO), Movie (MO), Spoken Word (SW), and Episodic (EP). In some cases, the content types describe the nature of the content and determine if the audio device pauses or attenuates audio during interruptions. This may impact SpeakEasy signal processing, impacting how content is treated when the listener is interrupted. In some examples, for MU, the content is typically attenuated but not paused during interruptions. In further examples, for PO, content is likely paused during interruptions. In additional examples, for MO, content is likely paused during interruptions. In further examples, for SW, content is likely paused during interruptions. In still further examples, for EP, content is likely paused during interruptions.

[0082] Particular implementations can beneficially provide structured encoding for audio content, which can be selected by content creators, distributors, or intermediaries. In some examples, a tagging stmcture is used, e.g., with specific encoding. In one particular, nonlimiting example, ID3 tagging is used, with two-letter encoding.

[0083] Certain examples rely on a structured encoding in a text field using multi-letter representation. In certain of these cases, each feature combination is represented by a unique two-letter code, with a short signature (e.g., device maker signature such as BOS) at the start toPN-25-007-WO Page 13 of 22identify the metadata as specific to a device maker (e.g., Bose Corporation). The metadata can be parsed from left to right, allowing future features to be added at the end of the string.

[0084] Examples of Multi-Letter Encoding Scheme:

[0085] BOS Signature: starts the metadata string for a particular device type (e.g., indicating a Bose Corporation device)

[0086] Immersive Disabled: DS

[0087] Immersive Full Rotation: FR

[0088] Immersive Center Channel Extraction: CE

[0089] Loudspeaker Small Room: SR

[0090] Loudspeaker Large Room: LR

[0091] Flat Response EQ: EQ

[0092] EQ Default / User Controlled: DC

[0093] Music: MU

[0094] Podcast: PO

[0095] Movie: MO

[0096] Spoken Word: SW

[0097] Episodic: EP

[0098] In particular implementations, features can be combined, e.g., two or more features can be combined in a string. For example: Podcast (PO) + Immersive Full Rotation (FR) + Flat Response EQ (EQ): Combined Code: BOSPOFREQ.

[0099] Various tags are possible for different content types, e.g., ID3 Tags. Non-limiting examples of ID3 Tags are shown below, with tags encapsulated in {brackets} and appended at the end of the title field in the ID3 data:

[0100] Song example (Immersive Small Room, Default EQ): {BOSSRDCMU}, which can be appended to the Title as:

[0101] Title: “Song Title {BOSSRDCMU}”.

[0102] Podcast example (Immersive Disabled, Default EQ): {BOSDSPODC}, which can be appended to the Title as:

[0103] Title: “Podcast Title {BOSDSPODC}”.

[0104] Movie example (Immersive Center Channel Extraction, Flat Response EQ):{BOSCEEQMO}, which can be appended to the Title as:

[0105] Title: “Movie Title {BOSCEEQMO}”.

[0106] Additional ExamplesPN-25-007-WO Page 14 of 22

[0107] In further implementations, processor 40 can be configured to analyze metadata 130 for instructions on synchronization and / or coordination of playback across multiple devices, e.g., across multiple audio devices 10. In certain cases, metadata 130 can include time alignment indicators and / or role -based rendering indicators (e.g., wearable audio device plus soundbar in shared environments) that can be evaluated to adjust or otherwise assign output modes.

[0108] Further implementations enable source-side metadata generation, e.g., authoring and / or appending metadata 130 for audio files or streams 120 using playback sources (e.g., device(s) 20) such as mobile applications on smart devices, audio-visual (AV) receivers, and / or venue sound systems (e.g., soundboards). In some cases, the control program 90 can be run at any processor at a playback source to enable authoring and / or appending metadata 130 to audio files or streams 120. Further, DAW 200 can be run as a program and / or application at any device described herein to enable source-side metadata generation.

[0109] Additional implementations can enhance cross-device compliance and schema enforcement, for example, by standardizing interpretation of metadata 130 across third-party devices (e.g., devices 20). For example, metadata 130 can be authored and / or appended to audio files or streams 120 in a consistent manner to enable certification and compliance across a plurality of distinct device types (e.g., devices varying by manufacturer, operating system, platform, etc.).

[0110] Further implementations enable distribution of metadata 130 in a scalable platform, e.g., via broadcast or another communications protocol. In some cases, metadata 130 can be broadcast simultaneously to multiple devices using a communications protocol such as BLE, LE Audio, or other transport layers. In particular cases, this can enhance audio output control across a plurality of devices, e.g., with grouped speakers, multiple audio systems, paired audio devices 10, etc.

[0111] Various implementations can enhance the user experience, e.g., by applying desirable audio signal processing (or modes) based on content type. Further, these implementations can beneficially enable content creators, distributors, etc., to assign content indicators to content in metadata. These metadata tags enable content creators to preselect the desirable signal processing for their content, enabling a premium listening experience as the content creator intended. In particular cases, users can override these settings by disabling automatic features (or select features of such modes) on their audio devices (e.g., via device interface and / or via connected smart device). Certain of the disclosed structured tagging approaches, can enable low-friction adoption of content identification and enhance the creator and user experiences.PN-25-007-WO Page 15 of 22

[0112] In any case, the approaches described according to various implementations have the technical effect of enhancing audio output for a user based on the detected type of audio content. For example, the approaches described according to various implementations efficiently identify audio content type to tailor audio output at one or more speaker systems. Further, the approaches described according to various implementations can effectively identify types of audio content with specific metadata, making adoption of these approaches by various content providers more efficient. Users of the disclosed systems and methods experience an enhanced audio experience when compared with conventional systems.

[0113] Particular implementation described herein are configured to be deployed at an audio device and / or at a connected smart device. One or both devices can include a controller including one or more microcontrollers or processors having a digital signal processor (DSP). In some cases, the controller is referred to as control circuit(s). The controller(s) may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The controller may provide, for example, for coordination of other components of the audio device, such as control of user interfaces (not shown) and applications run by the audio device. In various implementations, controller includes a metadata-based audio control module (or program), which can include software and / or hardware for performing audio control processes described herein. For example, controller can include metadata-based audio control module in the form of a software stack having instructions for controlling functions in outputting audio to one or more audio devices in a system according to any implementation described herein. As described herein, the controller, as well as other controller(s) described herein, is configured to control functions in audio output according to various implementations. In addition to a controller as described herein, an audio device can further include at least one transducer for providing an audio output (e.g., electro-acoustic transducers), at least one microphone (e.g., a microphone array), a communications unit (e.g., including a wireless communication system such as a BT module), along with additional electronics such as orientation sensors, optical sensors, capacitive sensors, user interfaces, etc.

[0114] Additional aspects of audio control (e.g., spatialization), for example, in open ear, on ear or in-ear audio devices, are described in US Patent Nos. 10,972,857 (“Directional Audio Selection”), 10,929,099 (“Spatialized Virtual Personal Assistant”) and US 11,036,464 (“Spatialized Augmented Reality (AR) Audio Menu”), each of which is incorporated here by reference in its entirety.

[0115] The above description provides embodiments that are compatible with BLUETOOTH SPECIFICATION Version 5.2 [Vol 0], 31 Dec. 2019, as well as any previous version(s), e.g.,PN-25-007-WO Page 16 of 22version 4.x and 5.x devices. Additionally, the connection techniques described herein could be used for Bluetooth LE Audio, such as to help establish a unicast connection. Further, it should be understood that the approach is equally applicable to other wireless protocols (e.g., nonBluetooth, future versions of Bluetooth, and so forth) in which communication channels are selectively established between pairs of stations.

[0116] In some implementations, the host-based elements of the approach are implemented in a software module (e.g., an “App”) that is downloaded and installed on the source / host (e.g., a “smartphone,” television, soundbar, or smart speaker), in order to provide the spatialized audio output aspects according to the approaches described above.

[0117] While the above describes a particular order of operations performed by certain implementations of the invention, it should be understood that such order is illustrative, as alternative embodiments may perform the operations in a different order, combine certain operations, overlap certain operations, or the like. References in the specification to a given embodiment indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic.

[0118] The functionality described herein, or portions thereof, and its various modifications (hereinafter “the functions”) can be implemented, at least in part, via a computer program product, e.g., a computer program tangibly embodied in an information carrier, such as one or more non-transitory machine-readable media, for execution by, or to control the operation of, one or more data processing apparatus, e.g., a programmable processor, a computer, multiple computers, and / or programmable logic components.

[0119] A computer program can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a network.

[0120] Actions associated with implementing all or part of the functions can be performed by one or more programmable processors executing one or more computer programs to perform the functions of the calibration process. All or part of the functions can be implemented as, special purpose logic circuitry, e.g., an FPGA and / or an ASIC (application-specific integrated circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind ofPN-25-007-WO Page 17 of 22digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory' or both. Components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.

[0121] In various implementations, unless otherwise noted, electronic components described as being “coupled” can be linked via conventional hard-wired and / or wireless means such that these electronic components can communicate data with one another. Additionally, subcomponents within a given component can be considered to be linked via conventional pathways, which may not necessarily be illustrated.

[0122] A number of implementations have been described. Nevertheless, it will be understood that additional modifications may be made without departing from the scope of the inventive concepts described herein, and, accordingly, other embodiments are within the scope of the following claims.PN-25-007-WO Page 18 of 22

Claims

CLAIMSWe claim:

1. A method of controlling audio device output, comprising:receiving an audio file or audio stream for output at an audio device, the audio file or audio stream provided with metadata defining audio output settings, wherein the audio file or audio stream is in a standard format,evaluating the metadata for corresponding output capabilities at the audio device, and if the audio device does not have corresponding output capabilities defined by the metadata, outputting the audio file or audio stream in the standard format.

2. The method of claim 1 , wherein the metadata includes at least one of descriptive metadata or prescriptive metadata.

3. The method of claim 2, wherein the metadata includes both descriptive metadata and prescriptive metadata.

4. The method of claim 2, further comprising, if the metadata includes prescriptive metadata, applying specified signal processing instructions to audio output at the audio device according to the prescriptive metadata.

5. The method of claim 4, wherein the prescriptive metadata specifies how to render content in the audio file or audio stream with the specified signal processing instructions that are configured to override conflicting default audio output settings for the audio device.

6. The method of claim 2, further comprising, if the metadata includes descriptive metadata and not prescriptive metadata, selecting an output mode for the audio output at the audio device from a predefined set of output modes.

7. The method of claim 2, wherein the descriptive metadata describes the nature of content in the audio file or the audio stream, wherein selecting output mode is based on heuristics or predefined rules according to the nature of the content.PN-25-007-WO Page 19 of 228. The method of claim 1, wherein if the audio device has corresponding output capabilities defined by the metadata, outputting the audio file or audio stream according to the metadata by applying at least one of: an equalization adjustment, a spatialization adjustment, a virtualization adjustment, or an anchoring adjustment to the audio file or audio stream.

9. The method of claim 1, wherein the standard format includes at least one of a stereo format or a stereo transport format.

10. The method of claim 1, wherein the metadata includes a set of indicators of audio output settings based on audio output device type.

11. The method of claim 10, wherein the set of indicators includes two or more indicators for assigning audio output settings in response to detecting two or more audio output devices, wherein the set of indicators is applied differently to distinct audio output devices.

12. The method of claim 1 , further comprising:receiving an input from a user interface to disable the metadata-based evaluation for corresponding output capabilities at the audio device, andoutputting the audio file or audio stream in the standard format in response to receiving the input.

13. The method of claim 1, wherein the evaluating is performed at the audio device.

14. The method of claim 13, wherein the audio device has limited processing capability.

15. The method of claim 1, wherein the metadata is provided with a metadata transport mechanism.

16. The method of claim 1, wherein the metadata is provided via at least one of: file-embedded tags, Bluetooth transport protocols, or a companion data channel associated with the audio stream.

17. The method of claim 1, wherein evaluating the metadata is performed automatically without user input.PN-25-007-WO Page 20 of 2218. The method of claim 1, further comprising providing a digital audio workstation (DAW) plug-in with a menu of metadata assignment options,wherein the menu of metadata assignment options includes options for assigning at least one of prescriptive metadata or descriptive metadata to each audio file or audio stream.

19. The method of claim 1, further comprising:classifying a library of audio files or audio streams according to content type; and appending the library with metadata classifiers for subsequent playback according to the metadata in a standard audio format.

20. The method of claim 1, wherein the audio device includes at least one of: an audio headset, a fixed speaker, a portable speaker, or a vehicle audio system.PN-25-007-WO Page 21 of 22