Content-based audio spatialization
By combining audio output devices and controllers, and automatically selecting spatialization modes based on content characteristics, the problem of existing audio systems being unable to adaptively adjust is solved, resulting in a better user experience and audio output effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-10
- Publication Date
- 2026-04-24
AI Technical Summary
Existing audio systems fail to adapt to content type during spatialization, resulting in a poor user experience.
By combining audio output devices and controllers, spatialization modes are automatically selected based on content characteristics, enabling personalized audio output on wearable audio devices. The spatialization modes are dynamically adjusted using machine learning content classifiers and orientation sensors.
It enhances the user's immersive and personalized audio experience, improves the spatial effect of audio output, and adapts to different types of audio content.
Smart Images

Figure CN121925629A_ABST
Abstract
Description
Priority Statement
[0001] This application claims priority to U.S. Patent Application No. 18 / 238,668, filed August 28, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates generally to audio systems. More specifically, this disclosure relates to content-based audio spatialization within audio systems. Background Technology
[0003] Certain types of audio content can benefit from spatialization, for example, to provide users with an immersive experience. However, it may not be desirable to spatialize all types of content or to spatialize all types of content in a similar way. Conventional spatialization methods may limit the user experience. Summary of the Invention
[0004] All examples and features mentioned below can be combined in any technically possible way.
[0005] Various specific implementations include methods for spatializing audio output based on content characteristics. Additional specific implementations include devices configured to provide spatialized audio output based on content characteristics, such as wearable audio devices.
[0006] In some specific aspects, an audio system includes: at least one audio output device for providing audio output based on data; and at least one controller coupled to the at least one audio output device, the controller being configured to: determine the content type of the audio output from a group of content types using the data; automatically select a spatialization mode for the audio output from a plurality of spatialization modes based on the determined content type; and apply the selected spatialization mode to the audio output.
[0007] In an additional specific aspect, a method for controlling audio output at an audio device includes: determining the content type of the audio output from a group of content types using data, automatically selecting a spatialization mode of the audio output at the audio device from a plurality of spatialization modes based on the determined content type, and applying the selected spatialization mode to the audio output.
[0008] Specific implementations may include one of the following features, or any combination thereof.
[0009] In some cases, the data associated with the audio output includes audio data, such as audio metadata.
[0010] In some cases, the spatialization mode for audio output is automatically selected without user input.
[0011] In some respects, content type groups are predefined.
[0012] In a specific implementation, a content classifier is used to determine the content type.
[0013] In some respects, a content classifier is a machine learning content classifier.
[0014] In some cases, content type is determined based on the dynamic range of samples from the source audio. In specific implementations, the content classifier is updated over time through comparisons with more classifier / audio data and / or user feedback.
[0015] In some respects, analyzing data includes analyzing metadata to determine the content type.
[0016] In some respects, determining the content type of data is based on confidence intervals.
[0017] In some respects, confidence intervals include an indicator of the probability that the data belongs to one of the content types in the content type group.
[0018] In some respects, the content type group includes at least three different content types, and the plurality of spatialization patterns includes at least three different spatialization patterns.
[0019] In certain contexts, this content type group may include subgroups and / or subgenres, such as action films, dramas, comedies, news, musicals, etc.
[0020] In some respects, one of the multiple spatialization modes includes non-spatialized audio output.
[0021] In some respects, this content type group includes two or more of the following: i) voice-based audio content, ii) music-based audio content, iii) theater audio content, and iv) live event audio content. In some examples, voice-based audio content includes podcasts and voice programs, and music-based audio content includes mixed music content, such as content that includes music and dialogue.
[0022] In some respects, the controller is further configured to receive user commands at the wearable audio device after the selected spatialization mode has been applied to the audio output, and to modify the selected spatialization mode based on the user commands.
[0023] In some respects, the spatialization mode selected by the application needs to be confirmed first via user interface commands.
[0024] In some respects, the choice of spatialization mode for audio output is further based on the type of wearable audio device.
[0025] In some respects, the controller is configured to select the default spatialization mode for audio output based on the type of wearable audio device.
[0026] In some aspects, wearable audio devices include open-ear wearable audio devices, and in response to the controller detecting movement of the wearable audio device, it determines whether the data is related to video output. If the data is related to video output, the spatialization of the audio output is fixed to the video device providing the video output or spatialization is disabled; otherwise, if the data is not related to video output, the spatialization of the audio output is fixed to the user's head position. In certain cases, spatialization is disabled based on a hysteresis factor.
[0027] In some respects, wearable audio devices can be categorized as in-ear, on-ear, over-ear, or near-ear.
[0028] In some respects, the controller is further configured to coordinate the selected spatialization mode with the audio output at one or more speakers outside the wearable audio device.
[0029] In some respects, the one or more speakers include at least one of a soundbar, a home entertainment speaker, a smart speaker, a portable speaker, or a vehicle speaker.
[0030] In some respects, wearable audio devices also include orientation sensors.
[0031] In some examples, the orientation sensor includes a magnetometer, gyroscope, and / or accelerometer. In one example, the orientation sensor includes an inertial measurement unit (IMU). In additional examples, the orientation sensor may include a vision-based sensor, such as a camera or lidar. In other examples, the orientation sensor may include Wi-Fi and / or Bluetooth connectivity and is configured to calculate angle of arrival (AoA) and / or angle of departure (AoD) data. In a particular example, the orientation sensor is configured to communicate via a local area network (LAN) standard such as IEEE 802.11.
[0032] In some respects, the controller is further configured to determine the orientation of the user of the wearable audio device based on data from the orientation sensor, and to adjust the selected spatialization mode based on the determined user orientation.
[0033] In various specific implementations, the user's orientation is determined based on one or more of the following: the user's gaze direction, the direction the wearable audio device is pointing in space, or the orientation of the wearable audio device relative to external audio devices (such as speakers in the area).
[0034] In some respects, adjusting the selected spatialization mode includes disabling the spatialization of the audio output in response to detecting a sudden change in user orientation.
[0035] In certain circumstances, a sudden change in orientation is detected when a user's orientation and / or location changes significantly over a period of time (e.g., meeting an orientation / location threshold). In some cases, a significant change in orientation or location over a period of time (e.g., a few seconds or less) is considered sudden.
[0036] In some respects, the selected spatialization mode is adjusted based on at least one head tracking (HT) algorithm selected from the following: head-fixed HT, external device-fixed HT, or hysteresis HT.
[0037] In some respects, the controller is further configured to select a spatialization mode based on at least one secondary factor, including user movement, proximity to a multimedia device, proximity to an external speaker, proximity to another wearable audio device, or presence in the vehicle.
[0038] In some respects, at least a portion of the controller is configured to run on a processor at a connected speaker or wearable smart device.
[0039] In some respects, the controller is further configured to adjust the selected spatialization mode when additional data is received.
[0040] In some respects, the spatialization model considers at least one of first-order reflections, second-order reflections, or post-reverberation.
[0041] In some respects, the data includes information about at least one of the following: the number of channels in the audio output, encoding information about the audio output, or the audio type (e.g., DOLBY Atmos with 5.1 baseline tracks and object data such as xyz coordinate data).
[0042] Two or more features described in this disclosure, including those described in the content section of this invention, may be combined to form specific embodiments not specifically described herein.
[0043] Details of one or more specific embodiments are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. Attached Figure Description
[0044] Figure 1 It is a block diagram of a system including a loudspeaker assembly, based on various disclosed specific implementations.
[0045] Figure 2 It is a schematic diagram of a space including various loudspeakers, based on various specific implementations.
[0046] Figure 3 It is a flowchart illustrating the process in various specific implementation methods.
[0047] Figure 4 This is a signal flow diagram illustrating various specific implementations of methods for content-based spatialization.
[0048] Figure 5 This is a data flow diagram illustrating aspects categorized according to various specific implementations.
[0049] It should be noted that the accompanying drawings for various specific embodiments are not necessarily drawn to scale. The drawings are intended only to illustrate typical aspects of this disclosure and should not be considered as limiting the scope of the specific embodiments. In the drawings, similar numbers indicate similar elements between the figures. Detailed Implementation
[0050] This disclosure is based, at least in part, on the understanding that content-adaptive spatialization of audio output can enhance the user experience.
[0051] Conventional audio systems employ spatialization and other audio output adjustments without considering content type. Because different types of content can benefit from spatialization and arrangement in different ways, these conventional systems are flawed.
[0052] Various specific implementations of the disclosed systems and methods adaptively control the spatialization mode used for audio output using content data. One particular method includes automatically selecting a spatialization mode for audio output from multiple spatialization modes based on a determined content type, and applying the selected spatialization mode to the audio output.
[0053] For illustrative purposes, the parts usually labeled in the accompanying drawings are considered to be substantially equivalent, and redundant discussion of those parts is omitted for clarity.
[0054] Figure 1 Examples of spaces 5 including system 10 according to various embodiments are shown, the system comprising a set of devices. In various embodiments, the devices shown in system 10 include at least one far-field (FF) speaker 20 and a pair of near-field (NF) speakers 30A, 30B. In some embodiments, the NF speakers 30A, 30B are part of a wearable audio device such as two earphones in a wearable audio device. In some examples, the NF speakers 30A, 30B are corresponding earbuds in a pair of wireless earphones. In other examples, the NF speakers 30A, 30B are part of an on-ear or other headphone assembly that may be obstructed (designed to cover the user's ear canal opening during use) or non-obstructed (designed not to cover the user's ear canal during use).
[0055] One or more additional devices 40 are shown, which are optional in some specific embodiments. The additional devices 40 may be configured to communicate with the FF speaker 20, NF speaker 30A, 30B, and / or other electronic devices in space 5 using any of the communication protocols or methods described herein. In some respects, system 10 is located in or around space 5, such as enclosed or partially enclosed rooms in homes, offices, theaters, sports or entertainment venues, religious places, etc. In some cases, space 5 has one or more walls and a ceiling. In other cases, space 5 comprises an open space without walls and / or a ceiling.
[0056] In various embodiments, at least one far-field loudspeaker 20 includes audio devices not intended to be worn by a user, such as a standalone amplifier or a group of amplifiers, such as a soundbar, portable loudspeaker, hardwired (i.e., semi-permanent or mounted) loudspeaker, etc. While loudspeaker 20 is described as a “far-field” device, loudspeaker 20 is not necessarily located within a recognized “far-field” acoustic distance relative to any other device in system 10. That is, far-field loudspeaker 20 does not need to be located in the far field to function according to the various embodiments described herein. In various embodiments, the minimum far-field distance is defined as approximately 0.5 meters in some cases, approximately two (2) meters or more in additional cases, and approximately three (3) meters or more in certain cases. It should be understood that the minimum far-field distance can vary based on the environment, for example, inside a vehicle (approximately 0.5 meters to approximately 2 meters), in a room within a home (e.g., approximately 2 meters to approximately 5 meters), or in an entertainment venue such as a concert hall (e.g., approximately 5 meters to approximately 50 meters).
[0057] In some cases, speaker 20 includes a controller 50 and a communication unit 60 coupled to the controller 50. In some examples, the communication unit 60 includes a Bluetooth module 70 (e.g., including a Bluetooth radio) enabling communication with other devices via the Bluetooth protocol. In some example implementations, speaker 20 may also include one or more microphones 80 (e.g., a single microphone or a microphone array) and at least one electroacoustic transducer 90 for providing audio output. Speaker 20 may also include additional electronics 100, such as a power manager and / or power supply (e.g., a battery or power connector), memory, sensors (e.g., an IMU, an accelerometer / gyroscope / magnetometer, an optical sensor, a voice activity detection system), etc. In some cases, the memory may include flash memory and / or non-volatile random access memory (NVRAM). In certain circumstances, the memory stores: microcode for the program used to process and control the controller 50, as well as various reference data; data generated during the execution of any of the various programs executed by the controller 50; the Bluetooth connection process; and / or various updatable data for proper safekeeping, such as paired device data, connection data, device contact information, etc. Figure 1 Some of the components shown above are optional and are indicated by dashed lines.
[0058] In some cases, controller 50 may include one or more microcontrollers or processors having a digital signal processor (DSP). In some cases, controller 50 is referred to as control circuitry. Controller 50 may be implemented as a chipset of chips that includes multiple independent analog and digital processors. Controller 50 may provide coordination of other components, such as speaker 20, such as control of a user interface (not shown) and applications operated by speaker 20. In various specific implementations, controller 50 includes one or more spatial (audio) rendering control modules, which may include software and / or hardware for performing the audio control processes described herein. For example, controller 50 may include a spatial audio rendering control module in the form of a software stack having instructions for controlling the function of outputting audio to one or more speakers in system 10 according to any specific implementation described herein. As described herein, controller 50, as well as other controllers described herein, are configured to control functions in spatial audio output control according to various specific implementations.
[0059] Communication unit 60 may include a BT module 70 configured to employ a wireless communication protocol such as Bluetooth, and additional network interfaces such as those employing one or more additional wireless communication protocols such as IEEE 802.11, Bluetooth Low Energy, or other local area network (LAN) or personal area network (PAN) protocols such as WiFi. In a particular embodiment, communication unit 60 is particularly adapted to communicate with speakers 30A, 30B and other communication units 60 in device 40 via Bluetooth. In another particular embodiment, communication unit 60 is configured to communicate with the BT devices described herein using broadcast audio via BLE or similar connections (e.g., including proxy connections). In a further embodiment, communication unit 60 is configured to wirelessly communicate with any other device in system 10 via one or more of the following: Bluetooth (BT); BT Low Energy (LE) audio; broadcast (e.g., to one or both NF speakers 30A, 30B, and / or additional device 40), such as via synchronized unicast; synchronized mixing audio connections (also known as SimpleSync) via BT or other wireless connections. TM The proprietary connectivity protocol is from Bose Corporation, Framingham, MA, USA; multiple transmit streams, such as those used in broadcasting, allow different devices with different pairs of non-blocking near-field loudspeakers (e.g., similar to NF loudspeakers 30A, 30B) to simultaneously output different portions of the audio signal in sync with a first portion of the audio signal. In a further embodiment, the communication unit 60 is configured to communicate with any other device in system 10 via, for example, a hardwired connection between any two or more devices.
[0060] As noted herein, controller 50 controls the general operation of FF speaker 20. For example, controller 50 performs processing of audio and data communication with auxiliary devices (e.g., NF speakers 30A, 30B), as well as audio output and signal processing at FF speaker 20. In addition to general operation, controller 50 initiates communication functions implemented in communication module 60 upon detecting certain triggers (or events) described herein. Controller 50 initiates operations (e.g., spatialization and / or synchronization of audio output) between FF speaker 20 and NF speakers 30A, 30B based on the characteristics of the audio content and / or data associated with the audio content.
[0061] In some examples, Bluetooth module 70 uses radio frequency (RF) communication between FF speaker 20 and NF speakers 30A, 30B (and, in some specific embodiments, additional device 40) to achieve wireless connectivity. Bluetooth module 70 exchanges radio signals, including data input / output, via an antenna (not shown). For example, in transmit mode, Bluetooth module 70 processes data through channel coding and spreading, converting the processed data into an RF signal and transmitting it. In receive mode, Bluetooth module 70 converts the received RF signal into a baseband signal, processes the baseband signal through despreading and channel decoding, and restores the processed signal to data. Additionally, Bluetooth module 70 can ensure secure communication between devices and uses encryption to protect data.
[0062] As noted in this article, Bluetooth-enabled devices include Bluetooth radios or other Bluetooth-specific communication systems that connect via the Bluetooth protocol. Figure 1 In the example shown, FF speaker 20 is a BT source device (also referred to as an "input device" or "host device"), and NF speakers 30A and 30B are part of a single BT destination device (also referred to as an "output device," "destination device," or "peripheral device") or different BT destination devices. Example Bluetooth-enabled source devices include, but are not limited to, smartphones, tablets, personal computers, laptops, notebook computers, netbooks, radios, audio systems (e.g., portable and / or fixed), Internet Protocol (IP) phones, communication systems, entertainment systems, headphones, smart speakers, a piece of sports and / or fitness equipment, portable media players, audio storage and / or playback systems, smartwatches, or other smart wearable devices, etc. Example Bluetooth-enabled destination devices include, but are not limited to, headphones, audio speakers (e.g., portable and / or fixed, with or without "smart" device functionality), entertainment systems, communication systems, smartphones, vehicle audio systems, a piece of sports and / or fitness equipment, amplified (or outdoor) audio equipment, wearable personal audio devices, etc. Additional BT devices may include portable game consoles, portable media players, audio gateways, BT gateway devices (used to bridge BT connections between other BT-enabled devices), audio / video (A / V) receivers as part of a home entertainment or home theater system, etc. Bluetooth-enabled devices, as described herein, can change their role from source to destination or from destination to source depending on the specific application.
[0063] In various specific embodiments, a first speaker in the NF loudspeaker (e.g., NF loudspeaker 30A) is configured to output audio to the user's left ear, and a second speaker in the NF loudspeaker (e.g., NF loudspeaker 30B) is configured to output audio to the user's right ear. In certain embodiments, NF loudspeakers 30A and 30B are housed in a common device (e.g., contained within a common housing) or otherwise form part of a common loudspeaker system. For example, NF loudspeakers 30A and 30B may include speakers within a seat or headrest, such as left / right speakers in the headrest and / or seat back portion of an entertainment seat, gaming seat, theater seat, car seat, etc. In some cases, NF loudspeakers 30A and 30B are positioned in the near field relative to the user's ear, for example, up to approximately 30 cm. In some of these cases, NF loudspeakers 30A and 30B may include headrest speakers or speakers worn on the body or shoulder at a distance of approximately 30 cm or less from the user's ear. In other specific cases, such as when the NF speakers 30A and 30B include headrest speakers, the near-field distance is approximately 15 cm or less. In further specific cases, such as when the NF speakers 30A and 30B include over-ear or near-ear wearable audio devices, the near-field distance is approximately 10 cm or less. In additional specific cases, such as when the NF speakers 30A and 30B include on-ear wearable audio devices, the near-field distance is approximately 5 cm or less. These example near-field ranges are merely illustrative, and various form factors can be considered within one or more of the example near-field ranges indicated herein.
[0064] In other specific embodiments, NF speakers 30A and 30B are part of a wearable audio device such as a wired or wireless wearable audio device. For example, NF speakers 30A and 30B may include the earpiece in a wearable headset with wireless coupling or hardwired connection. NF speakers 30A and 30B may also be part of a wearable audio device of any form factor, such as a pair of audio glasses, an on-ear or near-ear audio device, or an audio device placed on or around the user's head and / or shoulder area. In a particular specific embodiment, NF speakers 30A and 30B are non-blocking near-field speakers, meaning that when worn, the speaker 30 and its housing do not completely block (or obstruct) the user's ear canal. That is, at least some ambient acoustic signals can be transmitted to the user's ear canal without obstruction from NF speakers 30A and 30B. In additional specific implementations, the NF speakers 30A, 30B may include blocking devices (e.g., a pair of over-ear headphones or earplugs with an ear canal sealing feature) that can enable a transparent (or “perceived”) mode to transmit ambient acoustic signals to the user’s ears during playback.
[0065] like Figure 1 As shown, NF speakers 30A and 30B may include controllers 50a and 50b and communication units 60a and 60b (e.g., with BT modules 70a and 70b) enabling communication between FF speaker 20 and NF speakers 30A and 30B. The auxiliary device 40 may include one or more components described relative to FF speaker 20; in some embodiments, each component is shown as optional, indicated by dashed lines. The symbols “a” and “b” indicate that components in the device (e.g., NF speakers 30A, NF speakers 30B, auxiliary device 40) are physically separate from similarly labeled components in FF speaker 20, but may take on similar forms and / or functions as their corresponding labeled components in FF speaker 20. For brevity, additional descriptions of these similarly labeled components are omitted. Furthermore, as described herein, the additional NF speakers 30A, 30B and the additional device 40 may differ from the FF speaker 20 in form factor, intended use and / or capabilities, but in various specific implementations, they are configured to communicate with the FF device 20 in accordance with one or more communication protocols described herein (e.g., Bluetooth, BLE, Broadcast, SimpleSync, etc.).
[0066] Generally, Bluetooth modules 70, 70a, and 70b include a Bluetooth radio and additional circuitry. More specifically, Bluetooth modules 70, 70a, and 70b include both a Bluetooth radio and a Bluetooth LE (BLE) radio. In various embodiments, the presence of a BLE radio in Bluetooth module 70 is optional. That is, as noted herein, various embodiments use only the (classic) Bluetooth radio for connectivity. In embodiments including a BLE radio, the Bluetooth radio and the BLE radio are typically on the same integrated circuit (IC) and share a single antenna, while in other embodiments, the Bluetooth radio and the BLE radio are implemented as two separate ICs sharing a single antenna or as two separate ICs with two separate antennas. The Bluetooth specification, Bluetooth 5.2 (Low Power), provides forty channels spaced 2MHz apart to the FF speaker 20. The forty channels are labeled 0 to 39, including 3 broadcast channels and 37 data channels. Channels labeled 37, 38, and 39 are designated as broadcast channels in the Bluetooth specification, while the remaining channels 0-36 are designated as data channels in the Bluetooth specification. Certain example methods for Bluetooth-related pairing are described in U.S. Patent No. 9,066,327 (published June 23, 2015), the entire contents of which are incorporated herein by reference. Additionally, methods for selecting connections between paired devices and / or prioritizing connections between paired devices are described in U.S. Patent Application No. 17 / 314,270 (filed May 7, 2021), the entire contents of which are incorporated herein by reference.
[0067] As described herein, various specific implementations are particularly well-suited for spatializing (or otherwise controlling externalization and / or orchestration) audio output at one or more sets of speakers (e.g., NF speakers 30A, 30B and / or FF speakers 20). In some cases, the spatialized audio output is also coordinated at one or more additional devices 40. In specific cases, controllers 50 at one or more speakers are configured to coordinate the spatialized audio output to enhance the user experience, for example, by customizing the audio spatialization based on content type. For example, users of NF speakers 30A, 30B may have a more immersive and / or more personalized audio experience compared to listening to the audio output without considering content type.
[0068] Figure 2An embodiment of an audio system 10 in a space 105 (e.g., a room in a home, office, or entertainment venue) is shown. This space 105 is merely one example of various spaces that may benefit from the disclosed embodiment. In this example, a first user 110 is in a first seating position (e.g., seat 120), and a second user 130 is in a second seating position (e.g., seat 140). User 110 is wearing a wearable audio device 150, which in this example includes a set of audio glasses, such as Bose Frames audio glasses from Perth Ltd., Framingham, Massachusetts, USA. In other cases, wearable audio device 150 may include another open-ear audio device, such as a set of on-ear or near-ear headphones. In any case, wearable audio device 150 includes a set (e.g., two) of non-blocking NF speakers 30A, 30B. User 130 is located in seat 140, which includes a set of non-blocking NF speakers 30A, 30B in the headrest and / or neck / backrest portion 160. The space also shows an FF speaker 20, which may include a standalone speaker, such as a soundbar (e.g., one of the Bose Smart Soundbar series from Bose Ltd.) or a television speaker (e.g., the Bose TV Speaker from Bose Ltd.). In additional cases, the FF speaker 20 may include a home theater speaker (e.g., the Bose Surround Speaker series and / or the Bose Bass Module series) or a portable speaker such as a portable smart speaker (e.g., one of the Bose Soundlink series or the Bose Portable Smart Speaker from Bose Ltd.) or a portable professional speaker such as the Bose S1 Pro Portable Speaker. In this non-limiting example, additional devices 40A and 40B are present in space 105. For example, additional device 40A may include a television and / or visual display system (e.g., a projector-based video system or a smart monitor), and additional device 40B may include a smart device (e.g., a smartphone, tablet computing device, surface computing device, laptop computer, etc.). The devices in space 105 and the interactions between the devices are intended only to illustrate some aspects of the various aspects of this disclosure.
[0069] refer to Figure 2 As an illustrative example, according to certain specific implementations, the audio system 10 is configured to control the spatialization of audio outputs in the FF speaker 20 and / or NF speakers 30A, 30B. In specific cases, the processes performed according to various specific implementations are controlled by controllers at one or more speakers in the audio system 10 (e.g., controller 50 in the FF speaker 20 and / or controller 50 in the NF speakers 30A, 30B). Figure 1 The controllers 50a and 50b) control the audio output. In some cases, one or more controllers 50 are configured to coordinate the synchronous audio output at two or more types of speakers, for example, to help spatialize the audio output. Certain features of the synchronous audio control are further described in U.S. Patent Application No. 17 / 835,223 (filed June 8, 2022), the entire contents of which are incorporated herein by reference.
[0070] Figure 3 Example flowcharts illustrating processes executed by one or more controllers 50 according to various specific implementations are shown. Figure 4 An example is shown. Figure 3 A sample data flow diagram providing further details of the process flow. (See reference) Figure 3 In the first process (P1), controller 50 uses data associated with the audio output to determine the content type of the audio from a content type group. In a specific example, the data associated with the audio output includes audio data and / or metadata associated with the audio to be rendered or output. In certain cases, the data associated with the audio output is obtained along with the audio data, for example, as metadata in a file or file stream, or metadata tags within an audio stream or audio file. In other cases, for example, a decoder is used to extract the data associated with the audio output from the audio data. In some cases, algorithmic detection (e.g., a machine learning classification algorithm that identifies the style and / or type of audio content in the audio output) is used to identify the data associated with the audio output in the audio data. In some examples, such as Figure 4 As shown, audio (e.g., a file or stream) 210 is obtained from an audio source (e.g., a stored or transmitted audio file and / or a cloud-based audio streaming platform), and an audio decoder 220 extracts data 230 regarding the audio output to determine the spatialization settings at the speakers. In one example, the audio 210 is forwarded to an audio distributor 240 for post-processing and distribution to one or more speakers, such as NF speakers 30A, 30B and / or FF speakers 20. In some cases, the data 230 is sent to a content classifier 250 to determine the type of content in the audio 210. In a particular example, the content classifier 250 is configured to determine the type of content in the audio 210 based on the dynamic range of samples of the audio. In certain cases, the content classifier 250 may detect whether the audio content contains characteristics that indicate the content type, such as whether it includes dialogue, music, crowd noise, etc. In addition, in some cases, data 230 includes data about at least one of the following: the number of channels in the audio output, encoding information about the audio output, or the type of audio (e.g., DOLBY Atmos with 5.1 baseline tracks and object data such as xyz coordinate data).
[0071] Additionally, where available, metadata 260 about audio 210 from an audio source (e.g., television, streaming music service, audio file, etc.) can be sent to metadata parser 270 to help identify the content type of audio 210. Metadata parser 270 can be configured to separate metadata related to the content type of audio 210, such as general content types like movies, television, podcasts, sports broadcasts (e.g., live sports broadcasts) or genres like action films, dramas, comedies, news, music videos, sports, etc., and send the content type data to content profiler 280.
[0072] Depending on the specific implementation, the content classifier 250 may be updated over time, for example, through comparison and / or adjustment of more classifiers with audio data based on user feedback. For example, the content classifier 250 may be updated when additional data 230 is received and the corresponding spatialization pattern is applied to the audio output. In some cases, the controller 50 is configured to update the content classifier 250 based on user feedback, such as user adjustments to audio settings after audio is output in a specific spatialization pattern. In other cases, the controller 50 is configured to update the content classifier 250 after receiving user feedback, for example, in response to cues (such as audio cues, visual cues, or haptic cues) at one or more devices (such as wearable audio devices). In still other cases, the controller 50 monitors audio adjustments made by the user (e.g., via an interface at the audio device) for a period of time after the initiation of spatialized audio output (e.g., within a few minutes after the initiation of spatialized audio output). In a specific example, content classifier 250 is a machine learning (ML) content classifier that is trained on data such as audio data (and audio metadata) and spatialization setting data, and is configured to determine the content type of the audio output from a group of content types.
[0073] In some cases, the content type of data 230 is determined based on confidence intervals. For example, the content type of data 230 may not always be directly related to a predefined content type, and / or may include data indicating multiple content types. In such cases, content classifier 250 may be configured to apply confidence intervals to the determination of one or more content types and select the content type of the data based on the largest (or relatively highest) confidence interval. In specific cases, the confidence interval includes a probability indicator that data 230 belongs to one content type in a group of content types (e.g., one content type in two, three, four, or more content type groups). In some examples, the group of content types includes two or more of the following: i) voice-based audio content, ii) music-based audio content, iii) theater audio content, and iv) live event audio content. In some examples, voice-based audio content includes podcasts and voice programs, and music-based audio content includes mixed music content, such as content that includes music and dialogue (which may include descriptive language, such as director's clips or descriptions to help visually impaired people).
[0074] Figure 5 A non-limiting example of a data stream in a content classifier 250 is illustrated, which may include an ML content classifier as described herein. In these cases, the content classifier 250 may be trained using training data 500, including audio data and / or metadata about the audio, to form distinctions in terms of audio characteristics, content type, and / or audio file information. It should be understood that the content classifier 250 may be updated over time with additional training data 500 and feedback data 510 (as described herein), such as user feedback on the selection of content type for audio output. Furthermore, new data 230 may be used to update the content classifier 250 when making future selections of content type. In various specific implementations, the content classifier is configured to analyze the following of the data 230 (e.g., audio data): dynamic range 520; sub-content in the (audio) data 230, such as dialogue 530, music 540, crowd noise 550; and / or data file / stream characteristics, such as encoding 560, number of channels 570, or audio file type 580. Based on these characteristics of data 230, content classifier 250 selects one or more possible content types, such as content type 1 and content type 2, and assigns confidence intervals (e.g., confidence interval 1, confidence interval 2) to each content type. As described herein, content classifier 250 may apply probabilistic methods to determine the content type of data 230, for example, using relative confidence intervals. In some cases, the content type with the highest (maximum) confidence interval (e.g., confidence interval 2) is selected as the determined content type.
[0075] In some examples, data 230 may have a wide dynamic range, including crowd noise, and include encoding and / or channel numbers (e.g., multi-channel surround). This data 230 may primarily indicate live events such as sports broadcasts (e.g., content type 1), and secondarily indicate theater content (e.g., content type 2). Based on the weighting of the above factors, content classifier 250 may assign confidence intervals to each content type (e.g., content type 1, confidence interval 1; content type 2, confidence interval 2), such that sports broadcasts have the highest confidence interval and are selected as that content type. In other examples, data 230 may have a narrow dynamic range, including music, and include encoding and / or channel numbers indicating surround sound output. This data 230 may primarily indicate music-based audio content, such as classical music content (e.g., content type 3), and secondarily indicate speech-based audio content (e.g., content type 4). Based on the weighting of the above factors, the content classifier 250 can assign a confidence interval to each content type (e.g., content type 3, confidence interval 3; content type 4, confidence interval 4), such that music-based audio content has the highest confidence interval and is selected as that content type. In certain cases, the confidence interval may be based on one or more characteristics of the audio content. For example, the energy distribution in the frequency domain (within and outside the range of human speech) can represent speech, singing, or crowd noise. The confidence interval (or simply, confidence level) can be determined in part by the difference between the energy within that range and the average value on the audible spectrum. In other examples, traffic noise can be characterized as infrequent and outside the sound range (e.g., low energy), dialogue can be characterized as distinct from music (e.g., by energy differences), live events such as concerts can be characterized by crowd noise accompanying the music content, and / or film features can include the priority of content subtypes, such as sound effects or music taking precedence over dialogue.
[0076] return Figure 3 and Figure 4 The outputs of content classifier 250 and metadata parser 270 (if metadata 260 is available) are combined at content parser 280 and processed in process P2 ( Figure 3 For example, spatial mode selection 290 is performed based on content type. In certain cases, controller 50 automatically performs spatial mode selection based on the determined content type, i.e., without user input. After spatial mode selection 290 is completed, in process P3 ( Figure 3 The selected spatialization mode is applied in ( ). (See reference) Figure 4 The selected spatialization mode can be applied, for example, using digital signal processing tuning profile 330 to the audio output signal from audio distributor 240, for example, to the post-processed output at NF speaker 30 (post-processing signal 340) and / or FF speaker 20 (post-processing signal 350).
[0077] In certain cases, the content type group is predefined and includes two or more content types. In further embodiments, the content type group includes at least three different content types. In some cases, the multiple spatialization modes include at least two spatialization modes. In further embodiments, the multiple spatialization modes include at least three different spatialization modes. In a particular example, one of the multiple spatialization modes includes non-spatialized audio output.
[0078] In some respects, spatialization patterns consider at least one of first-order reflections, second-order reflections, or post-reverberation. One aspect of the audio experience controlled by the tuning of the speaker system is the sound field. "Sound field" refers to the listener's perception of where sound is coming from. Specifically, a wide (sound from both sides of the listener), deep (sound from near and far), and precise (the listener can identify where a particular sound appears to be coming from) sound is generally desired. Furthermore, the sound field is typically limited by the physical space in which the user is located (e.g., a room). In an ideal system, someone listening to recorded music can close their eyes, imagine watching a live performance, and pinpoint the location of each musician. A related concept is "surround effect," which we use to refer to the perception of sound coming from all directions (including from behind the listener), regardless of whether the sound can be precisely located. The perception of sound field and surround effect (and generally, sound location) is based on the difference in level and arrival time (phase) between the sounds reaching the listener's two ears, and the sound field can be controlled by manipulating the audio signals produced by the speakers to control these inter-early level and time differences. As described in U.S. Patent No. 8,325,936 (“Directionally Radiating Sound in a Vehicle”), not only near-field speakers but also stationary speakers can be used collaboratively to control spatial perception, which is incorporated herein by reference. Additional aspects of spatialization, such as in open-ear, on-ear, or in-ear audio devices, are described in U.S. Patent Nos. 10,972,857 (“Directional Audio Selection”), 10,929,099 (“Spatialized Virtual Personal Assistant”), and 11,036,464 (“Spatialized Augmented Reality (AR) Audio Menu”), each of which is incorporated herein by reference in its entirety.
[0079] In some cases, at least one secondary or additional factor is considered when selecting the spatialization mode. Secondary (or additional) factors may include user movement, proximity to multimedia devices, proximity to external speakers, proximity to another wearable audio device, or presence within a vehicle. Figure 4 In one example shown, activity or proximity detection 300 or speaker type 310 (e.g., the type of wearable audio device) can be used to assist in spatialization mode selection. Activity or proximity detection 300 may consider whether one or more additional devices (e.g., FF speaker 20, NF speaker 30) are active, and / or the proximity of a main output device such as NF speaker 30 to FF speaker 20. Information about speaker type 310 may include whether the speaker is a Bluetooth-connected speaker, amplifier, open-ear speaker, on-ear speaker (e.g., headphones), or in-ear speaker (e.g., earbuds). Rule set 320 assigns a spatialization mode based on the determined content type, and in some cases, based on activity or proximity detection 300 and / or information about speaker type 310. In one example, rule set 320 may assign synchronization settings (e.g., indicating whether to synchronize audio output) to the audio output at FF speaker 20 (e.g., a soundbar) and NF speaker 30 based on activity or proximity detection 300, such as based on the detection of a power cycle of NF speaker 30, the proximity of NF speaker 30 to FF speaker 20, and / or the identification of the type of NF speaker 30.
[0080] As described herein, in certain situations, controller 50 is configured to select a default spatialization mode for audio output based on the type of wearable audio device (e.g., NF speaker 30). For example, in some aspects, the types of wearable audio devices include in-ear wearable audio devices, on-ear wearable audio devices, over-ear wearable audio devices, or near-ear wearable audio devices, and the default spatialization mode is selected based on the type of wearable audio device (e.g., having at least two different spatialization modes and at least two different device types).
[0081] In additional specific implementations, users can customize or otherwise configure the spatialization patterns for audio output, for example, by utilizing feedback from prompts from controller 50 and / or via interfaces at FF speaker 20, NF speaker 30, and / or device 40 (in electronic device 100). In some examples, the application (e.g., running at device 40) may enable users to configure and / or customize spatialization patterns and related settings according to any of the methods described herein. For example, controller 50 may present one or more configurable parameters to the user (e.g., via an interface at device 40 or other devices), including: a list of available spatialization patterns (e.g., the user can remove a particular spatialization pattern from the list if they do not like it), how spatialization patterns are automatically selected (e.g., whether there is an immediate transition or a smooth transition, such as via a lag period), or other aspects of the techniques described herein. In some implementations, such user customizations (and / or other user customizations) are stored and linked to a user profile accessible by controller 50. In some such implementations, companion software may be used to perform user customization (and optionally user profile settings), such as a companion application accessible from a peripheral device such as device 40 (e.g., a smartphone or tablet computer to which a wearable audio device is connected). Furthermore, user configuration and / or customization may be performed during the device's startup or setup phase and may subsequently be edited via any of the interfaces described herein.
[0082] In some respects, the wearable audio device (NF speaker 30) includes open-ear wearable audio devices, and in response to the controller 50 detecting movement of the wearable audio device, the controller 50 determines whether data 230 is related to video output (e.g., at attachments 40A and / or 40B). If data 230 is related to video output, the controller 50 either fixes the spatialization of the audio output to the video device providing the video output (e.g., attachment 40A) or disables spatialization. Further, if data 230 is not related to video output, the controller 50 fixes the spatialization of the audio output to the user's head position (e.g., as determined using orientation sensors in attachments 100a, 100b at the NF speaker 30). In certain cases, spatialization is disabled based on a hysteresis factor (e.g., a delay of a few seconds or less). The hysteresis factor can be used to mitigate false triggering when adjusting or otherwise disabling spatialization based on the user's momentary, short-term, or unintentional head movement.
[0083] The orientation of a wearable audio device (e.g., NF speaker 30) can be determined using electronic devices 100a, 100b, including orientation sensors. In some examples, the orientation sensor includes a magnetometer, gyroscope, and / or accelerometer. In one example, the orientation sensor includes an inertial measurement unit (IMU). In additional examples, the orientation sensor may include a vision-based sensor, such as a camera or lidar. In other examples, the orientation sensor may include Wi-Fi and / or Bluetooth connectivity and is configured to calculate angle of arrival (AoA) and / or angle of departure (AoD) data. In a particular example, the orientation sensor is configured to communicate via a local area network (LAN) standard such as IEEE 802.11. In some aspects, the controller 50 is further configured to determine the orientation of the user of the wearable audio device (e.g., NF speaker 30) based on data from the orientation sensor and to adjust a selected spatialization mode based on the determined user orientation.
[0084] In various implementations, the user's orientation is determined based on one or more of the following: the user's gaze direction, the direction the wearable audio device is pointing in space, or the orientation of the wearable audio device relative to external audio devices (such as speakers in the area). In some aspects, adjusting the selected spatialization mode includes disabling the spatialization of the audio output in response to detecting a sudden change in the user's orientation. A sudden change in orientation can be detected when the user's orientation and / or position changes significantly over a period of time (e.g., meeting an orientation / position threshold). In some cases, a significant change in orientation or position over a period of time (e.g., a few seconds or less) is considered sudden. In further implementations, adjusting the selected spatialization mode is based on at least one head tracking (HT) algorithm selected from: head-fixed HT, external device-fixed HT, or hysteresis HT.
[0085] return Figure 3 In some additional implementations, the controller is configured to apply the selected spatialization mode (P3) only after receiving a user interface command. For example, applying the selected spatialization mode (P3) first requires confirmation via a user interface command (e.g., voice command, touch command, gesture-based command, etc.). In some examples, controller 50 prompts the user to confirm the change in spatialization mode using audio cues (e.g., at NF speaker 20) and / or visual cues (e.g., on the screens of devices 40A, 40B).
[0086] In further specific implementation, for example, Figure 3As depicted in the optional process (shown by dashed lines), after applying the selected spatialization mode (P3), the controller 50 is further configured to: receive a user command at the wearable audio device (P4) and modify the selected spatialization mode based on the user command (P5). In some cases, the user command includes audible commands (e.g., voice commands), tactile commands (e.g., via a touch interface or button), or gesture-based commands such as head shaking or looking down (e.g., detected via an orientation sensor).
[0087] In an additional embodiment, controller 50 is further configured to adjust the selected spatialization mode upon receiving additional data (e.g., additional data 230 from audio 210, such as the progress of audio stream 210). The additional data may also include data from one or more sensors in the additional electronics 100 and / or data from communication unit 60. For example, controller 50 may adjust the selected spatialization mode in response to detecting the addition of a speaker (e.g., FF speaker 20) to a speaker group connected to NF speakers 30A, 30B. In these examples, BT module 70a may detect the connection or disconnection of an FF speaker 20 to NF speakers 30A, 30B, or the WiFi module in communication unit 60 may detect the addition or removal of another FF speaker 20 to or from a speaker group in space 5 via a WiFi connection. In response to these detected changes, controller 50 may further adjust the selected spatialization mode to enhance the audio output at speakers 30A, 30B that are connected to (or not connected to) speaker 20. In addition, additional data may include detecting the presence of another NF speaker in space 5, such as detecting another BT device nearby or another WiFi device on the network via communication unit 60, for example, to adjust the spatialization of the audio output of all users in space 5.
[0088] In various specific implementations, the controller 50 at the wearable audio device (e.g., controllers 50a, 50b at the NF speaker 30) is configured to coordinate the selected spatialization mode with the audio output at one or more speakers external to the wearable audio device, for example, coordinating the output with the controller 50 at the FF speaker 20 and / or the controller 50c at the additional device 40. In some cases, the controllers 50a, 50b at the wearable audio device are configured to coordinate the spatialized audio output with an additional speaker, such as the FF speaker 20, including a soundbar, home entertainment speaker, smart speaker, portable speaker, or vehicle speaker. In these or other cases, at least a portion of the controller 50 (e.g., having functionality for coordinating the spatialized audio output) is configured to operate on a processor at the connected speaker (e.g., the FF speaker 20, such as a soundbar, smart speaker, or vehicle speaker) or on a processor at the wearable smart device (e.g., device 40, such as a smartwatch, smart ring, smart glasses, etc.).
[0089] While this article describes various example configurations of devices and sources, it should be understood that environments (e.g., environment 105) may vary. Figure 2 Any device in the audio device group can act as a source device, destination device, and / or connection device. For example, FF speaker 20 and NF speakers 30A, 30B can be connected to a common source device, such as one of the additional devices (e.g., television, audio gateway device, smartphone, tablet computing device, etc.) 40 described herein. For example, the source device may include a television system, a smartphone, or a tablet computer. In additional embodiments, the source device includes network-based and / or cloud-based devices, such as network-connected audio systems. In further embodiments, FF speaker 20 and / or NF speakers 30A, 30B act as source devices, for example, having integrated network and / or cloud communication capabilities. In this case, FF speaker 20 and / or NF speakers 30A, 30B receive audio signals from a network (or cloud) connected gateway device (e.g., a wireless or hardwired internet router). Figure 2In one example shown, the source device may include an auxiliary device 40A or 40B, which is a network and / or cloud-connected device running a software program or application (also referred to as an "app") configured to manage audio output to FF speaker 20 and / or NF speakers 30A, 30B. In some examples, the source device is connected to both FF speaker 20 and / or NF speaker 30A and is configured to coordinate the audio output between those speakers, for example, by sending signals to one or both of FF speaker 20 and NF speakers 30A, 30B. In an additional example, source device 210 sends signals to FF speaker 20 or NF speakers 30A, 30B (one or both), which are forwarded between those speaker connections. In some specific implementations, NF speakers 30A, 30B forward signals via a "spying" type method or otherwise synchronize their output. While this document describes specific example scenarios, the FF speaker 20 and NF speakers 30A, 30B can forward or otherwise transmit signals in any technically feasible manner, and the examples described herein (e.g., SimpleSync, broadcast, Bluetooth, etc.) should not be considered as limitations on various specific implementations.
[0090] Implementation varies depending on specific examples, such as Figure 2 In the specific embodiments shown, the FF speaker 20 is housed in a soundbar (such as a Bose Soundbar as described herein). In further specific embodiments, the FF speaker 20 includes a plurality of speakers configured to output at least the left, right, and center channels of the audio signal 220 to a space (e.g., space 105). For example, the FF speaker 20 may include a group of stereo-paired portable speakers, such as portable speakers configured to operate individually and in stereo pairs. In additional examples, the FF speaker 20 may include a plurality of speakers in a group of stereo and / or surround sound speakers, such as two, three, four, or more speakers arranged in a space (e.g., space 105).
[0091] Furthermore, according to certain examples, NF speakers 30A and 30B are configured to coordinate audio output at FF speaker 20 in response to a trigger. In specific cases, a trigger may include one or more of the following: a detected connection between two devices housing different speakers (e.g., a wired connection, or a wireless connection such as a BT or Wi-Fi connection, such as a Wi-Fi RTT connection); a detected directional alignment between two devices housing different speakers (e.g., via BT angle of arrival (AoA) and / or angle of departure (AoD) data); a user grouping request (e.g., from an application or voice command, such as via an attached device 40B or a wearable audio device (e.g., wearable audio device 150)); proximity or location detection (e.g., devices identified as being close to each other or in the same area or space, such as via BT AoA and / or AoD data, and / or via Wi-Fi RTT); or a user-initiated command or user response to a prompt following a trigger described herein. Any trigger described herein may enable controller 50 at a device to coordinate spatialized output across multiple speakers (e.g., in a space such as space 105) based on content type.
[0092] While various embodiments include descriptions of non-blocking variants of the NF speakers 30A and 30B, in additional embodiments, the NF speakers 30A and 30B may include blocked near-field speakers, such as over-ear or in-ear headphones, operating in a transparent (or transparent) mode. For example, a pair of headphones with passive and / or active noise cancellation capabilities may replace the non-blocking variants of the NF speakers 30A and 30B described herein. In these cases, the blocked near-field speakers may operate in a shared experience (or social) mode, which may be enabled via user interface commands and / or any triggers described herein. In a particular example, the transparent (or transparent) mode allows a user to experience ambient audio output from the FF speaker 20 while also experiencing spatialized playback of audio from the NF speakers 30A and 30B.
[0093] Additionally, while various embodiments are described as beneficially enhancing the user audio experience without necessarily knowing the user's head position (e.g., via user head tracking capabilities), these embodiments can be used in conjunction with systems configured to track the user's head position. In such cases, data regarding the user's head position (e.g., as indicated by an IMU, optical tracking system, and / or proximity detection system) can be used as input to one or more processing units (e.g., at controller 50) to further enhance the user audio experience, for example, by adjusting the audio signal output to the NF speakers 30A, 30B and / or FF speakers 20 (e.g., in terms of spatialization, externalization, etc.). However, data regarding the user's head position is not necessary for the beneficial deployment of methods and systems according to the various embodiments.
[0094] In any case, the methods described according to various specific embodiments have the technical effect of enhancing the spatialization of the user's audio output based on the detected audio content type. For example, the methods described according to various specific embodiments spatialize the audio output at one or more speaker systems based on the identified type of audio content output at those speakers. In some cases, these methods select from multiple spatialization profiles to provide the best match for a type of audio content, thereby enhancing both individual and group user experiences. Furthermore, the methods described according to various specific embodiments can effectively identify the type of audio content with or without content metadata, thereby enhancing the adaptability of systems deploying such methods. Compared to conventional systems, the disclosed systems and methods provide users with an enhanced immersive audio experience.
[0095] This article describes various wireless connectivity scenarios. It should be understood that any number of wireless connections and / or communication protocols can be used to couple space (e.g., space 105 ( Figure 2 The device in the document. Examples of wireless connection scenarios and triggering for connecting wireless devices are described in more detail in U.S. Patent Applications No. 17 / 714,253 (filed April 4, 2022) and No. 17 / 314,270 (filed May 7, 2021), the entire contents of each of which are incorporated herein by reference.
[0096] It should also be understood that any RF protocol can be used to communicate between devices, depending on the specific implementation, including Bluetooth, Wi-Fi, or other proprietary or non-proprietary protocols. The NF speakers 30A and 30B are housed in wearable audio devices (e.g., Figure 2 In specific implementations of the present invention, such implementations may advantageously use wireless protocols otherwise used by wearable audio devices to receive audio data other than the technologies described herein (e.g., Bluetooth), thereby eliminating the need for wearable audio devices to include additional components and costs.
[0097] In specific implementations utilizing Bluetooth LE Audio, a unicast topology can be used for a one-to-one connection between the FF speaker 20 and the NF speakers 30A, 30B. In some implementations, an LE Audio broadcast topology (such as Broadcast Audio) can be used to send one or more sets of audio data to multiple groups of NF speakers 30 (although a broadcast topology can still be used for only one group of NF speakers 30A, 30B). For example, in some such implementations, the broadcast audio data is the same for all groups of NF speakers 30 within range, such that all groups of NF speakers 30 receive the same audio content. However, in other such implementations, different audio data is broadcast to the groups of NF speakers 30 within range, such that some NF speakers 30 can select a first audio data, and other NF speakers 30 can select a second audio data different from the first audio data. Different audio data can allow for differences in audio personalization (e.g., EQ settings), dialogue language selection, dialogue intelligibility enhancement, externalization / spatialization enhancement, and / or other differences as understood based on this disclosure. Furthermore, the use of the LE audio broadcast topology (and other specific implementations described differently herein) allows for local adjustment of the volume level of the received audio content at each group of NF speakers 30 (as opposed to a single global volume level from the FF speakers 20 in system 10).
[0098] The above description provides an implementation scheme compatible with Bluetooth specification version 5.2 [Vol 0] as of December 31, 2019, and any previous versions (e.g., version 4.x and 5.x devices). Additionally, the connectivity techniques described herein can be used for Bluetooth LE audio, such as to facilitate the establishment of unicast connections. Furthermore, it should be understood that this method is equally applicable to other wireless protocols (e.g., non-Bluetooth, future versions of Bluetooth, etc.) in which communication channels are selectively established between pairs of stations.
[0099] In some specific implementations, the host-based elements of the method are implemented in a software module (e.g., an “App”) that is downloaded and installed on the source / host (e.g., a “smartphone”, television, soundbar, or smart speaker) to provide spatialized audio output aspects according to the method described above.
[0100] Although a particular order of operations performed by certain specific embodiments of the invention has been described above, it should be understood that such an order is exemplary, as alternative embodiments may perform these operations in a different order, combine certain operations, overlap certain operations, etc. References to a given embodiment in the specification indicate that the embodiment may include a particular feature, structure, or characteristic, but each embodiment may not necessarily include that particular feature, structure, or characteristic.
[0101] The functions or portions thereof described herein, and their various modifications (hereinafter referred to as "functions") may be implemented at least in part by computer program products, such as computer programs tangibly implemented in an information carrier, such as one or more non-transitory machine-readable media, for performing or controlling the operation of one or more data processing devices, such as programmable processors, computers, multiple computers and / or programmable logic components.
[0102] Computer programs can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form (including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment). Computer programs can be deployed on a single computer, distributed across a site or multiple sites, or executed on multiple computers interconnected by a network.
[0103] The actions associated with all or part of the functions implemented in the calibration process can be performed by one or more programmable processors executing one or more computer programs. All or part of the functions can be implemented as special-purpose logic circuits, such as FPGAs and / or ASICs (Application-Specific Integrated Circuits). Processors suitable for executing computer programs include, for example, both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, the processor will receive instructions and data from read-only memory or random access memory, or both. The components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.
[0104] In various specific implementations, unless otherwise stated, electronic components described as "coupled" may be linked via conventional hardwired and / or wireless means, enabling these electronic components to transmit data to each other. Additionally, sub-components within a given component may be considered linked via conventional paths, which may not necessarily be illustrated.
[0105] Several specific embodiments have been described. However, it should be understood that additional modifications may be made without departing from the scope of the inventive concept described herein, and therefore, other embodiments are within the scope of the following claims.
Claims
1. A wearable audio device, comprising: At least one audio output device, the at least one audio output device being used to provide audio output based on data; as well as At least one controller, coupled to the at least one audio output device, is configured to: The data is used to determine the content type of the audio output from the content type group. Based on the determined content type, the spatialization mode for the audio output is automatically selected from multiple spatialization modes, and The selected spatialization mode is applied to the audio output.
2. The wearable audio device of claim 1, wherein the spatialization mode for the audio output is automatically selected without user input.
3. The wearable audio device of claim 1, wherein the content type group is predefined.
4. The wearable audio device of claim 1, wherein the content type is determined using a content classifier, wherein the content classifier is a machine learning content classifier.
5. The wearable audio device of claim 1, wherein analyzing the data includes: Analyze the metadata to determine the content type.
6. The wearable audio device of claim 1, wherein determining the content type of the data is based on a confidence interval, wherein the confidence interval includes a probability indicator that the data belongs to a content type in the group of content types.
7. The wearable audio device of claim 1, wherein the content type group comprises at least three different content types, and wherein the plurality of spatialization modes comprises at least three different spatialization modes, wherein one of the plurality of spatialization modes comprises non-spatialized audio output.
8. The wearable audio device of claim 1, wherein the group of content types includes two or more of the following: i) Voice-based audio content, ii) Music-based audio content, iii) Theater audio content, and iv) Live event audio content.
9. The wearable audio device of claim 1, wherein the controller is further configured to, after applying the selected spatialization mode to the audio output, Receive user commands at the wearable audio device, and The selected spatialization mode is modified based on the user command.
10. The wearable audio device of claim 1, wherein the selection of the spatialization mode for the audio output is further based on the type of the wearable audio device, wherein the controller is configured to select a default spatialization mode for the audio output based on the type of the wearable audio device, and The wearable audio device of the aforementioned type includes an open-ear wearable audio device, and in response to the controller detecting movement of the wearable audio device, Determine whether the data is related to the video output. If the data is related to video output, then the spatialization of the audio output is fixed to the video device providing the video output, or spatialization is disabled, or... If the data is unrelated to the video output, then the spatialization of the audio output is fixed at the user's head position.
11. The wearable audio device of claim 1, wherein the controller is further configured to coordinate the selected spatialization mode with audio output at one or more speakers external to the wearable audio device, wherein the one or more speakers include at least one of the following: Soundbars, home entertainment speakers, smart speakers, portable speakers, or vehicle speakers.
12. The wearable audio device of claim 1, further comprising an orientation sensor, wherein the controller is further configured to: The orientation of the user of the wearable audio device is determined based on data from the orientation sensor, and The selected spatialization pattern is adjusted based on the user's determined orientation. The adjustments to the selected spatialization mode include: In response to detecting a sudden change in the user's orientation, the spatialization of the audio output is disabled, and The spatialization mode selected for adjustment is based on at least one head-tracking HT algorithm chosen from the following: Head-fixed HT, external device-fixed HT, or delayed HT.
13. The wearable audio device of claim 1, wherein the controller is further configured to select the spatialization mode based on at least one secondary factor, the at least one secondary factor comprising: The user is moving, approaching a multimedia device, approaching an external speaker, approaching another wearable audio device, or is in a vehicle.
14. The wearable audio device of claim 1, wherein the data includes data concerning at least one of: The number of channels in the audio output, encoding information about the audio output, or the type of audio.
15. A method for controlling audio output at an audio device, the method comprising: Using data, determine the content type of the audio output from the content type group; Based on the determined content type, the spatialization mode for audio output at the audio device is automatically selected from multiple spatialization modes; as well as The selected spatialization mode is applied to the audio output.
16. The method of claim 15, wherein the automatic selection of the spatialization mode for the audio output is performed without user input, and wherein the content type group is predefined.
17. The method of claim 15, wherein the data includes data concerning at least one of: The number of channels in the audio output, encoding information about the audio output, or the type of the audio, wherein the content type is determined using a content classifier.
18. The method of claim 15, wherein analyzing the data comprises: Analyze the metadata to determine the content type.
19. The method of claim 15, wherein determining the content type of the data is based on a confidence interval, wherein the confidence interval includes a probability indicator that the data belongs to a content type in the content type group, wherein the content type group includes at least three different content types, and wherein the plurality of spatialization patterns includes at least three different spatialization patterns.
20. The method of claim 15, wherein the group of content types comprises two or more of the following: i) Voice-based audio content, ii) Music-based audio content, iii) Theater audio content, and iv) Live event audio content. The spatialization mode selected for the audio output is also based on at least one of the following: Types of wearable audio devices, or At least one secondary factor, said at least one secondary factor including: The user is moving, approaching a multimedia device, approaching an external speaker, approaching another wearable audio device, or is in a vehicle.
Citation Information
Patent Citations
Spatialized virtual personal assistant
US10929099B2
Directional audio selection
US10972857B2
Spatialized augmented reality (AR) audio menu
US11036464B2
Proximity-based connection for Bluetooth devices
US11937159B2
Proximity-based connection for bluetooth devices
US20220361264A1