Improved audio system with mixed rendering audio

The audio system uses mixed rendering with far-field and near-field speakers to enhance personalization and clarity for multiple users, overcoming limitations of conventional systems by synchronizing distinct audio outputs for improved user experience.

JP7835903B2Active Publication Date: 2026-03-25BOSE CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-06-08
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Conventional loud audio systems fail to provide personalized audio experiences for multiple users in a listening environment, limiting user experience due to insufficient personalization.

Method used

An audio system employing mixed rendering techniques, utilizing a far-field speaker and a pair of non-occluding near-field speakers to output synchronized, distinct portions of an audio signal, enhancing audio quality and personalization without sacrificing social interaction.

Benefits of technology

Improves audio clarity, personalization, and spatialization for individual users while maintaining a shared audio experience, addressing dialogue intelligibility and frequency range limitations of conventional systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007835903000001
    Figure 0007835903000001
  • Figure 0007835903000002
    Figure 0007835903000002
  • Figure 0007835903000003
    Figure 0007835903000003
Patent Text Reader

Abstract

Various embodiments include an audio system and method for mixed rendering to improve audio output. Certain embodiments include at least one far-field speaker configured to output a first portion of an audio signal and a pair of non-enclosed near-field speakers configured to output a second portion of the audio signal in synchronization with the output of the first portion of the audio signal, wherein the first portion of the audio signal is different from the second portion of the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to audio systems. More particularly, this disclosure relates to improving mixed rendering audio in an audio system.

Background Art

[0002] In conventional loud audio systems, when multiple users are located within the listening environment, personalization becomes insufficient. That is, these conventional systems provide the same loud audio content to all users within the environment. This conventional approach may limit the user experience.

Summary of the Invention

Means for Solving the Problems

[0003] All examples and features mentioned below can be combined in any technically possible way.

[0004] Various embodiments include techniques for improving audio output using mixed rendering. Additional embodiments include an audio system configured to employ a mixed rendering technique to improve audio output.

[0005] In some particular aspects, the audio system includes at least one far-field speaker configured to output a first portion of an audio signal, and a pair of non-occluding near-field speakers configured to output a second portion of the audio signal in synchronization with the output of the first portion of the audio signal, wherein the first portion of the audio signal is different from the second portion of the audio signal.

[0006] In additional specific embodiments, a method for controlling an audio system includes outputting a first portion of an audio signal to at least one far-field speaker, and outputting a second portion of the audio signal to a pair of unbound near-field speakers in synchronization with the output of the first portion of the audio signal, wherein the first portion of the audio signal is distinct from the second portion of the audio signal.

[0007] The embodiments may include one of the following features, or any combination thereof.

[0008] In some cases, synchronizing and outputting the first and second parts of the audio signal can be done within approximately 100 milliseconds (ms).

[0009] In certain embodiments, the first portion of the audio signal and the second portion of the audio signal include separate channels of common content.

[0010] In certain embodiments, a second portion of the audio signal is binaurally encoded, and when the second portion of the audio signal is output in sync with the output of the first portion of the audio signal, the externalization of the content in the audio signal is increased.

[0011] In some cases, the second part of the audio signal enhances the spoken content within the audio signal.

[0012] In certain embodiments, a second portion of the audio signal increases the intelligibility of the speech content within the audio signal. In some examples, the speech content is removed from the audio track and then added back to it. In certain cases, a near-field device (e.g., a wearable audio device) adds a level of audio detail (i.e., clarity) that is not readily achievable from a far-field audio device (e.g., a loudspeaker located several meters or more away), thereby improving the impulse response of the audio output. In certain embodiments, a far-field audio device can supplement the bass output in an unobstructed near-field device (e.g., an open-ear wearable audio device), which may struggle to provide the desired bass without obstruction. In such examples, the near-field and far-field devices have the advantage of complementary audio output.

[0013] In certain embodiments, the first portion of the audio signal does not contain spoken content, while the second portion of the audio signal does. In one example, the language of the spoken content in the second portion of the audio signal is selectable.

[0014] In certain embodiments, the first and second portions of the audio signal include different frequency ranges than the audio signal. In a particular example, the first portion of the audio signal is output in a frequency range where the sound field diffuses (or is not modal), e.g., above approximately 200 Hz. In additional examples, the frequency limit in each device is adjustable to improve the output.

[0015] In certain cases, i) the first part of the audio signal has a wider frequency range than the second part of the audio signal, or ii) the second part of the audio signal excludes frequencies below a predetermined threshold.

[0016] In some embodiments, a pair of non-obstructed near-field speakers are housed within a wearable audio device. In certain examples, the wearable audio device includes a pair of wirelessly coupled audio devices, such as wireless earphones.

[0017] In additional embodiments, a pair of non-obstructed near-field speakers are housed in a seat, such as in the neck and / or head support portion of the seat.

[0018] In a particular embodiment, the first speaker in a pair of non-occlusive near-field speakers is configured to output audio to the user's left ear, and the second speaker in a pair of non-occlusive near-field speakers is configured to output audio to the user's right ear.

[0019] In some cases, i) at least one far-field speaker is housed within a soundbar, or ii) at least one far-field speaker includes multiple speakers configured to output at least left, right, and center channels of an audio signal.

[0020] In certain embodiments, at least one far-field speaker is configured to output a third portion of the audio signal when it does not provide output in synchronization with a pair of unbound near-field speakers, the third portion of the audio signal being different from the first portion of the audio signal. In some examples, the third portion of the audio signal is the entire audio signal.

[0021] In certain cases, at least one far-field speaker is configured to operate differently in different modes, including, for example, Mode 1: outputting audio in synchronization with a pair of unbound near-field speakers, and Mode 2: outputting audio only without synchronization with a pair of unbound near-field speakers. In additional embodiments, a pair of unbound near-field speakers is configured to operate differently in different modes, including, for example, Mode 1: outputting a subset of the frequency range of the audio signal when playing in synchronization with the far-field speakers, and Mode 2: outputting an audio signal with a wider frequency range than in Mode 1, up to the full frequency range of the audio signal, when outputting audio only (or not synchronized with at least one far-field speaker).

[0022] In some cases, a second portion of the audio signal can be user-configured to enhance at least one of the following: external localization, spatialization, or dialogue. In certain cases, the second portion of the audio signal can be user-configured via an interface, such as an audio interface and / or application interface on a computing device, e.g., a smart device. In certain cases, the second portion of the audio signal can be configured to adjust tone, equalization, volume, and other parameters to different unoccluded near-field speakers. In additional embodiments, the adjustment of the audio output to at least one far-field speaker is configured to perform the same or proportional adjustment as the audio output to an unoccluded near-field speaker. In even further embodiments, the adjustment of the audio output in a given device (e.g., an unoccluded near-field speaker and / or a far-field speaker) is limited so that at least one parameter of the audio signal is limited to a certain range. In some of these embodiments, volume adjustment may be limited to a certain range, tone adjustment may be limited to a certain range, and / or equalization adjustment may be limited to a certain range.

[0023] In certain embodiments, user-configurable parameters (e.g., a second portion of an audio signal) can be controlled using an interface on the non-occluded near-field speaker, such as an adjustment interface including buttons and switches, a touch interface including a capacitive touch interface, or an audio interface via a microphone in the non-occluded near-field speaker. In some embodiments, the adjustment interface on the non-occluded near-field speaker may include a scroll-through or cycle-through type adjustment mechanism that allows the user to cycle through or scroll through modes using each interface interaction, such as a button touch or a switch flip.

[0024] In additional embodiments, non-enclosed near-field speakers can be replaced with enclosed near-field speakers, such as over-ear or in-ear headphones, operating in transparent (or hear-through) mode. In these cases, the enclosed near-field speakers can operate in shared experience (or social) mode.

[0025] In certain cases, a second portion of the audio signal is automatically customized based on the detection of the type of device housing at least one of a pair of unoccluded near-field speakers. In some embodiments, the device type is detected using a device identifier, and the second portion of the audio signal is automatically customized based on the device identifier, for example, to adapt the audio output to a different device type (e.g., different wearable audio devices). In certain embodiments, the customized audio output is triggered based on at least one of a Bluetooth® (BT) connection (or a variant of a BT connection), a previous BT pairing scenario, or a previously defined coupling between devices.

[0026] In some embodiments, a pair of unobstructed near-field speakers are configured to output a second portion of an audio signal in sync with the output of a first portion of an audio signal in response to a trigger. In certain cases, the trigger may include one or more of the following: a detected connection between two devices housing different speakers (such as a wired or wireless connection, like a BT or Wi-Fi connection); detected orientation matching between two devices housing different speakers (such as via BT angle of arrival (AoA) and / or angle of departure (AoD) data); a user grouping request (such as from an application or voice command); proximity or location detection (such as devices identified as being in close proximity to each other or in the same zone or space, via BT AoA and / or AoD, or Wi-Fi RTT, etc.); or a user-initiated trigger or command for synchronous output.

[0027] In certain embodiments, the system further includes a first device that houses at least one far-field speaker, and the first device is configured to transmit either an audio signal or a second portion of the audio signal to a second device that includes a pair of non-occluding near-field speakers, the transmission including playback timing data to enable the second portion of the audio signal to be output in synchronization with the first portion of the audio signal. In certain cases, the audio signal or the second portion of the audio signal is transmitted wirelessly via one or more of: broadcast (to the second device and at least one additional device) via Bluetooth (BT), BT Low Energy (LE) audio, synchronous unicast, etc.; synchronous downmix audio connection via BT or other wireless connection; for example, different pairs of non-occluding near-field speakers having different devices that output different portions of the audio signal simultaneously in synchronization with the first portion of the audio signal, such as when one portion focuses on spatialization and / or improved out-of-head localization and another portion focuses on improved dialog, or when one portion includes dialog output in a first language and another portion includes dialog output in a second different language, etc., via multiple transmission streams such as broadcast.

[0028] In some cases, the system further includes a first device comprising at least one far-field speaker and a second device comprising a pair of unbound near-field speakers, wherein i) the first and second devices receive an audio signal from a common source device, or ii) the first device receives a first portion of an audio signal from the first source device and the second device receives a second portion of an audio signal from the second source device. In certain embodiments, in scenario (i), the common source device may include a television or streaming video player (via wired and / or wireless transmission, for example), where the first and second devices wirelessly receive an audio signal from the television or streaming video player, or a soundbar receives an audio signal via a hardwired connection to the television or streaming video player and the unbound near-field speakers wirelessly receive an audio signal from the television or streaming video player. In certain cases, in scenario (ii), a first device, such as a soundbar, receives a first portion of an audio signal from a first source device, such as a television or streaming video player, and a second device, such as a wearable audio device, receives a second portion of an audio signal from a second source device, such as a smartphone, tablet, or computing device.

[0029] In certain embodiments, the first device and the second device receive an audio signal from a common source device, the audio signal includes audio content that takes into account surround and height effects, and outputting the audio signal in the first device and the second device causes the room housing the first device and the second device to be evoked in order to improve the out-of-head localization for the user. In some of these cases, the audio output is perceived by the user as object-based (compared to channel-based) so that the user perceives the audio output as an object in space. In some aspects, the system provides an audio signal having audio content that takes into account surround and height effects in response to determining that a non-occluded near-field speaker is likely to be located at a minimum distance, such as a distance of several meters or more, from at least one far-field speaker. In some of these aspects, the surround and height effects are beneficial in large spaces such as religious services, concert halls, arenas, conference rooms, etc., where the distance between the far-field speaker and the non-occluded near-field speaker meets a threshold.

[0030] Two or more features described in this disclosure, including the features described in the Summary section of the present invention, may be combined to form embodiments not specifically described herein.

[0031] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.

Brief Description of the Drawings

[0032] [Figure 1] FIG. is a block diagram of a system including a far-field speaker and a set of non-occluded near-field speakers, according to various disclosed embodiments. [Figure 2] FIG. is a schematic diagram of a space including a far-field speaker and a non-occluded near-field speaker, according to various embodiments. [Figure 3]These are audio signal flow diagrams based on various embodiments. [Figure 4] This is an audio signal flow diagram with various additional embodiments. [Modes for carrying out the invention]

[0033] Please note that the drawings of various embodiments are not necessarily to scale. The drawings are intended to illustrate only typical aspects of the present disclosure and should not be considered to limit the scope of the embodiments. In the drawings, similar numbering indicates similar elements between drawings.

[0034] This disclosure is at least in part based on the recognition that coordinating audio outputs from separate speaker systems can improve individual and group user experiences.

[0035] Conventional audio systems that use dialogue enhancement rely on a mix of spectral correction and dynamic range compression. While this mixed approach has its advantages, dialogue intelligibility is still negatively affected by one or more speech transfer index (STI) factors, such as background noise, the quality of the audio playback equipment, reverberation in the space (e.g., a room), and masking effects from other audio content in the mix. Conventional systems may fail to adequately account for these STI factors, potentially reducing dialogue intelligibility.

[0036] In addition, conventional audio systems employing binaural rendering in occluded wearable audio devices (e.g., over-ear headphones) rely on advanced room models and head position (i.e., tracking) data to localize virtual sound sources outside the head. These conventional methods are costly and consume a considerable amount of power. Furthermore, processing head position data can introduce latency-related drawbacks to these conventional methods.

[0037] Systems and methods disclosed according to various embodiments utilize at least one far-field speaker to output a first portion of an audio signal and a pair of non-occlusive (e.g., open-ear) near-field speakers to output a second, distinct portion of the audio signal in synchronization with the output of the first portion. In these cases, users of non-occlusive near-field speakers (e.g., in-seat speakers or wearable audio devices that do not occlude the ear canal) can have a personalized audio experience without sacrificing the social aspects of an open-ear audio environment. For example, users of non-occlusive near-field speakers can experience improved audio quality, clarity, personalization / customization, etc., compared to listening to far-field speakers alone, without sacrificing social interaction with other users in the same space (due to non-occlusive near-field speakers). Furthermore, the blending of audio outputs from the hybrid systems and technologies described herein can provide numerous audio consumption benefits, including improved personalization, enhanced out-of-head localization / spatialization, increased dialogue intelligibility (e.g., for audio for video content or podcasts), and / or the ability to better experience the full frequency spectrum of audio content (e.g., experiencing full-body bass compared to bass received from open-ear wearable audio devices).

[0038] Commonly labeled components in the diagram are considered substantially equivalent for illustrative purposes, and redundant descriptions of those components are omitted for clarity.

[0039] Figure 1 shows an example of a space 5 including a system 10 that includes a set of devices according to various embodiments. In various embodiments, the devices shown in system 10 include at least one far-field (FF) speaker 20 and a pair of unenclosed near-field (NF) speakers 30A, 30B. One or more additional devices 40 are shown, which are optional in some embodiments. The additional devices 40 can be configured to communicate with the FF speaker 20, NF speakers 30A, 30B, and / or other electronic devices in space 5 using any communication protocol or method described herein. In certain embodiments, system 10 is located in or around a space 5, such as an enclosed or partially enclosed room, like a home, office, theater, sports or entertainment venue, or religious venue. In some cases, space 5 has one or more walls and a ceiling. In other cases, space 5 includes an outdoor location without walls and / or a ceiling.

[0040] In various embodiments, at least one far-field speaker 20 includes a standalone loudspeaker or set of loudspeakers, such as an audio device not intended to be worn by the user, e.g., a soundbar, a portable speaker, or a hardwired (i.e., semi-permanent or fixed) speaker. Although the speaker 20 is described as a “far-field” device, the speaker 20 does not need to be located within a generally acceptable “far-field” acoustic distance to any other device in the system 10. That is, the far-field speaker 20 does not need to be located in the far field in order to function according to the various embodiments described herein. In various embodiments, the minimum far-field distance is defined as about 0.5 meters in some cases, about 2 meters or more in additional cases, and about 3 meters or more in specific cases. Please understand that the minimum far-field distance can vary depending on the environment, for example, inside a vehicle (approximately 0.5 meters to 2 meters), inside a room in a home (e.g., approximately 2 meters to 5 meters), or inside an entertainment facility such as a concert hall (e.g., approximately 5 meters to 50 meters).

[0041] In certain cases, the speaker 20 includes a controller 50 and a communication (comm.) unit 60 coupled to the controller 50. In certain examples, the communication unit 60 includes a Bluetooth module 70 (e.g., including a Bluetooth radio) that enables communication with other devices via the Bluetooth protocol. In certain exemplary embodiments, the speaker 20 may also include one or more microphones 80 (e.g., a single microphone or a microphone array) and at least one electroacoustic transducer 90 for providing audio output. The speaker 20 may also include additional electronics 100 such as a power manager and / or power supply (e.g., a battery or power connector), memory, and sensors (e.g., an IMU, an accelerometer / gyroscope / magnetometer, an optical sensor, a voice activity detection system). In some cases, the memory may include flash memory and / or non-volatile random access memory (NVRAM). In certain cases, the memory stores microcode and various reference data for programs that process and control the controller 50, data generated during the execution of any of the various programs performed by the controller 50, the Bluetooth connection process, and / or various updatable data for storage such as paired device data, connection data, device contact information, etc. Certain of the above components shown in Figure 1 are optional and are shown by dashed lines.

[0042] In certain cases, the controller 50 may include one or more microcontrollers or processors having digital signal processors (DSPs). In some cases, the controller 50 is referred to as a control circuit. The controller 50 may be implemented as a chipset of chips including separate and multiple analog and digital processors. The controller 50 can provide coordination of other components of the speaker 20, such as, for example, control of a user interface (not shown) and applications performed by the speaker 20. In various embodiments, the controller 50 includes a (audio) mixed rendering control module (or more modules) which may include software and / or hardware for performing the audio control processes described herein. For example, the controller 50 may include a mixed rendering control module in the form of a software stack having instructions for controlling the function of outputting audio to one or more speakers in the system 10 according to any embodiment described herein. As described herein, the controller 50, and other controllers described herein, are configured to control the function in mixed rendering audio output techniques according to various embodiments.

[0043] The communication unit 60 may include a BT module 70 configured to use a wireless communication protocol such as Bluetooth, along with an additional network interface that uses one or more additional wireless communication protocols such as IEEE 802.11, Bluetooth Low Energy, or other local area network (LAN) or personal area network (PAN) protocols such as Wi-Fi. In certain embodiments, the communication unit 60 is particularly suited to communicating with other communication units 60 in BT devices 30, 40 via Bluetooth. In additional specific embodiments, the communication unit 60 is configured to communicate with the BT devices described herein using broadcast audio via BLE or a similar connection (including, for example, a proxy connection). In yet another embodiment, the communication unit 60 is configured to communicate wirelessly with any other device in the system 10 via Bluetooth (BT), BT Low Energy (LE) Audio, Synced Unicast, etc., via broadcast (e.g., to one or both of the NF speakers 30A, 30B and / or additional device 40), Synced Downmix Audio Connection via BT or other wireless connections (also known as SimpleSync®, a proprietary connection protocol of Bose Corporation, Framingham, MA, USA), for example, via one or more transmission streams such as broadcast, to enable different devices having different pairs of unbound near-field speakers (e.g., similar to the NF speakers 30A, 30B) to simultaneously output different parts of an audio signal in synchronization with a first part of the audio signal. In yet another embodiment, the communication unit 60 is configured to communicate with any other device in the system 10, for example, via a hardwired connection between any two or more devices.

[0044] As described herein, the controller 50 controls the general operation of the FF speaker 20. For example, the controller 50 performs processing when controlling audio and data communication with other devices (e.g., NF speakers 30A, 30B), as well as audio output and signal processing in the FF speaker 20. In addition to general operation, the controller 50 activates communication functions implemented in the communication module 60 when it detects certain triggers (or events) described herein. The controller 50 activates operations between the FF speaker 20 and the NF speakers 30A, 30B (e.g., synchronization of audio output) when certain conditions are met.

[0045] In certain examples, the Bluetooth module 70 enables a wireless connection using radio frequency (RF) communication between the FF speaker 20 and the NF speakers 30A, 30B (and, in some embodiments, an additional device 40). The Bluetooth module 70 transmits and receives radio signals containing data that is input and output via an antenna (not shown). For example, in transmit mode, the Bluetooth module 70 processes the data by channel coding and spreading, converts the processed data into a radio frequency (RF) signal, and transmits the RF signal. In receive mode, the Bluetooth module 70 converts the received RF signal into a baseband signal, processes the baseband signal by despreading and channel decoding, and restores the processed signal to data. In addition, the Bluetooth module 70 can ensure secure communication between devices and protect data using encryption.

[0046] As referred to herein, a Bluetooth-enabled device includes a Bluetooth radio or other Bluetooth-specific communication system that enables connectivity via the Bluetooth protocol. In the example shown in Figure 1, the FF speaker 20 is a BT source device (also referred to as the “input device” or “host device”), and the NF speakers 30A and 30B are either part of a single BT sink device (also referred to as the “output device”, “destination device”, or “peripheral device”) or separate BT sink devices. Exemplary Bluetooth-enabled source devices include, but are not limited to, smartphones, tablet computers, personal computers, laptop computers, notebook computers, netbook computers, radios, audio systems (e.g., portable and / or fixed), Internet Protocol (IP) phones, communication systems, entertainment systems, headsets, smart speakers, sports and / or fitness equipment, portable media players, audio storage and / or playback systems, smartwatches or other smart wearable devices. Exemplary Bluetooth-enabled sink devices include, but are not limited to, headphones, headsets, audio speakers (portable and / or fixed, with or without “smart” device capabilities), entertainment systems, communication systems, smartphones, vehicle audio systems, sports and / or fitness equipment, outdoor audio devices, and wearable private audio devices. Additional BT devices may include portable game players, portable media players, audio gateways, BT gateway devices (for bridging BT connections between other BT-enabled devices), audio / video (A / V) receivers as part of home entertainment or home theater systems, etc. Bluetooth-enabled devices as described herein may change their role from source to sink or sink to source depending on the specific application.

[0047] In various specific embodiments, a first speaker in the NF speaker (e.g., NF speaker 30A) is configured to output audio to the user's left ear, and a second speaker in the NF speaker (e.g., NF speaker 30B) is configured to output audio to the user's right ear. In certain embodiments, the NF speakers 30A and 30B are housed in a common device (e.g., included in a common housing) or otherwise form part of a common speaker system. For example, the NF speakers 30A and 30B may include in-seat or in-headrest speakers, such as left / right speakers in the headrest and / or seatback portion of entertainment seats, gaming seats, theater seats, or car seats. In certain cases, the NF speakers 30A and 30B are positioned in the near field relative to the user's ears, for example, at a maximum distance of about 30 centimeters. In some of these cases, the NF speakers 30A and 30B may include headrest speakers or speakers mounted on the body or shoulder, located approximately 30 centimeters or less from the user's ears. In other specific cases, the near-field distance is approximately 15 centimeters or less, for example, when the NF speakers 30A and 30B include headrest speakers. In yet another specific case, the near-field distance is approximately 10 centimeters or less, for example, when the NF speakers 30A and 30B include on-head or near-ear wearable audio devices. In yet another specific case, the near-field distance is approximately 5 centimeters or less, for example, when the NF speakers 30A and 30B include on-ear wearable audio devices. These exemplary near-field ranges are illustrative only, and various form factors may be considered within one or more of these exemplary near-field ranges, as described herein.

[0048] In yet another embodiment, the NF speakers 30A, 30B are part of a wearable audio device, such as a wired or wireless wearable audio device. For example, the NF speakers 30A, 30B may include earphones in a wearable headset that are wirelessly coupled or have a hardwired connection. The NF speakers 30A, 30B may also be part of a wearable audio device of any form factor, such as a pair of audio glasses, an on-ear or near-ear audio device, or an audio device placed on or around the user's head and / or shoulder area. As described herein according to a particular embodiment, the NF speakers 30A, 30B are non-obstructive near-field speakers, meaning that when worn, the speaker 30 and its housing do not completely block (or obstruct) the user's ear canal. That is, at least some ambient acoustic signals can pass into the user's ear canal without interference from the NF speakers 30A, 30B. In additional embodiments further described herein, the NF speakers 30A, 30B may include an occlusion device (e.g., a pair of over-ear headphones or earphones having canal-sealing features) that allows the hear-through (or "recognition") mode to pass ambient acoustic signals as playback to the user's ears.

[0049] As shown in Figure 1, the NF speakers 30A and 30B may include controllers 50a and 60b and communication units 60a and 60b (e.g., having BT modules 70a and 70b) to enable communication between the FF speaker 20 and the NF speakers 30A and 30B. The additional device 40 may include one or more components described with reference to the FF speaker 20, each of which is shown with a dashed line as optional in certain embodiments. The notations "a" and "b" indicate that the components within the device (e.g., NF speaker 30A, NF speaker 30B, additional device 40) are physically separate from similarly labeled components within the FF speaker 20, but can take on similar forms and / or functions to their labeled counterparts within the FF speaker 20. Further descriptions of these similarly labeled components are omitted for brevity. Furthermore, as described herein, the additional NF speakers 30A, 30B and the additional device 40 may differ from the FF speaker 20 in terms of form factor, intended use, and / or capabilities, but in various embodiments they are configured to communicate with the FF device 20 according to one or more communication protocols described herein (e.g., Bluetooth, BLE, broadcast, SimpleSync, etc.).

[0050] Generally, Bluetooth modules 70, 70a, and 70b include a Bluetooth radio and additional circuitry. More specifically, Bluetooth modules 70, 70a, and 70b include both a Bluetooth radio and a Bluetooth LE (BLE) radio. In various embodiments, the presence of a BLE radio within Bluetooth module 70 is optional. That is, as referred to herein, various embodiments utilize only a (classic) Bluetooth radio for connectivity. In embodiments including a BLE radio, the Bluetooth radio and the BLE radio are typically on the same integrated circuit (IC) and share a single antenna, while in other embodiments, the Bluetooth radio and the BLE radio are implemented as two separate ICs sharing a single antenna, or as two separate ICs having two separate antennas. The Bluetooth specification, namely Bluetooth 5.2: Low Energy, provides 40 channels to the FF speaker 20 at 2 MHz intervals. The 40 channels are labeled 0-39 and include three advertising channels and 37 data channels. Channels labeled 37, 38, and 39 are designated as advertising channels in the Bluetooth specification, while the remaining channels 0 through 36 are designated as data channels in the Bluetooth specification. Certain exemplary methods of Bluetooth-related pairing are described in U.S. Patent No. 9,066,327 (published June 23, 2015), which is incorporated in its entirety by reference. Furthermore, methods for selecting and / or prioritizing connections between paired devices are described in U.S. Patent Application Publication No. 17 / 314,270 (filed May 7, 2021), which is incorporated in its entirety by reference.

[0051] As described herein, various embodiments are particularly suitable for synchronized audio output in both NF speakers 30A, 30B and FF speaker 20. In certain cases, the synchronized audio output is tuned in one or more additional devices 40. In certain cases, a controller 50 in one or more of the speakers is configured to tune the synchronized audio output to enhance the user experience, enabling a personalized audio experience without sacrificing the social aspects of an open-ear audio environment, for example. For example, users of NF speakers 30A, 30B can experience improved audio quality, clarity, personalization / customization, etc., compared to listening to only the FF speaker 20, without sacrificing social interaction with other users in the same space (due to the non-occlusive nature of the NF speaker 30).

[0052] Figure 2 shows one embodiment of the audio system 10 in a space 105, such as a room in a home, office, or entertainment facility. This space 105 is just one example of various spaces from which the disclosed embodiment can be beneficial. In this example, a first user 110 is present in a first seating position (e.g., seat 120), and a second user 130 is present in a second seating position (e.g., seat 140). User 110 is wearing a wearable audio device 150, which in this example includes a set of audio glasses, such as Bose Frames audio glasses by Bose Corporation (Framingham, MA, USA). In other cases, the wearable audio device 150 may include another open-ear audio device, such as a pair of on-ear or near-ear headphones. In either case, the wearable audio device 150 includes a set of (e.g., two) non-occlusive NF speakers 30A, 30B. User 130 is positioned in a seat 140 which includes a set of non-enclosed NF speakers 30A, 30B within the headrest and / or neck / backrest portion 160. FF speakers 20 are also shown in the space, which may include standalone speakers such as a soundbar (e.g., one of the Bose Smart Soundbar types by Bose Corporation) or a television speaker (e.g., a Bose TV Speaker by Bose Corporation). In additional cases, FF speakers 20 may include portable speakers such as home theater speakers (e.g., Bose Surround Speaker types and / or Bose Bass Module types), or portable smart speakers (e.g., one of the Bose Soundlink types or Bose Portable Smart Speaker by Bose Corporation), or portable professional speakers such as a Bose S1 Pro Portable Speaker. In this non-limiting example, additional devices 40A and 40B are present in space 105.For example, additional device 40A may include a television and / or visual display system (e.g., a projector-based video system or a smart monitor), and additional device 40B may include a smart device (e.g., a smartphone, tablet computing device, surface computing device, laptop, etc.). The devices and device interactions within space 105 are merely illustrative of some of the various aspects of the present disclosure.

[0053] Referring to an exemplary example in Figure 2, according to a particular embodiment, the audio system 10 is configured to control synchronous audio outputs in both the FF speaker 20 and the NF speakers 30A, 30B. In a particular case, the process performed according to various embodiments is controlled by controllers in one or more of the speakers in the audio system 10, for example, controller 50 in the FF speaker 20 and / or controllers 50a, 50b in the NF speakers 30A, 30B (Figure 1). In a particular embodiment, the FF speaker 20 is configured to output a first portion of the audio signal, and the NF speakers 30A, 30B are configured to output a second portion of the audio signal in synchronization with the output of the first portion of the audio signal. Figure 3 shows a signal flow diagram illustrating an exemplary audio signal flow in relation to Figure 2. As shown in this example, the FF speaker 20 and the NF speakers 30A, 30B can be connected to a common source device 210. In certain embodiments, the source device 210 may include one of the additional devices 40 described herein (e.g., televisions, audio gateway devices, smartphones, tablet computing devices, etc.). For example, the source device 210 may include a television system, a smartphone, or a tablet. In additional embodiments, the source device 210 may include a network-based and / or cloud-based device, such as a network-attached audio system. In further embodiments, the FF speaker 20 and / or NF speakers 30A, 30B may function as a source device having, for example, integrated network and / or cloud communication capabilities. In such cases, the FF speaker 20 and / or NF speakers 30A, 30B receive audio signals from a network (or cloud)-attached gateway device, such as a wireless or hardwired internet router. In one example shown in Figure 3, the source device 210 is a network and / or cloud-attached device running a software program or software application (also referred to as an "app") configured to manage audio output to the FF speaker 20 and / or NF speakers 30A, 30B.In certain examples, source device 210 transmits signals to both FF speaker 20 and NF speakers 30A, 30B. In additional examples, source device 210 transmits signals to either FF speaker 20 or NF speakers 30A, 30B (one or both), and these signals are transmitted between their speaker connections. In certain embodiments, NF speakers 30A, 30B transmit signals via a “snooping” type technique or otherwise synchronize their outputs. While certain exemplary scenarios are described herein, FF speaker 20 and NF speakers 30A, 30B can transmit signals or transmit them in any technically feasible way, and the examples described herein (e.g., SimpleSync, broadcast, BT, etc.) should not be considered limiting to various embodiments.

[0054] In a specific example as shown in Figure 3, source device 210 transmits the audio signal 220 to FF speaker 20 and NF speakers 30A, 30B. In a further embodiment, as shown in Figure 4, two separate source devices 210A and 210B transmit signals to FF speaker 20 and NF speakers 30A, 30B, respectively. For example, using space 105 in Figure 2 as an example, an additional device 40A (e.g., a video system) can transmit a first portion 230 of the audio signal 220 to FF speaker 20, and an additional device 40B (e.g., a smartphone or tablet) can transmit a second portion 240 of the audio signal 220 to one or both sets of NF speakers 30A, 30B.

[0055] Continuing with reference to Figures 2, 3, and 4, in an exemplary embodiment, the FF speaker 20 is configured to output a first portion of the audio signal 220 (received, for example, from a source device 210), and the NF speakers 30A, 30B are configured to output a second, separate portion of the audio signal 220 (received, for example, from a source device 210) in synchronization with the output of the first portion from the FF speaker 20. In some cases, the synchronized output of the first portion 220 and the second portion 220 of the audio signal includes outputting those portions in separate speakers within approximately 100 milliseconds (ms) of each other.

[0056] In some cases, the first and second portions of the audio signal 220 contain separate channels of common content. In certain embodiments, the second portion of the audio signal 220 is binaurally encoded, for example, to improve the output in the NF speakers 30A and 30B. In these cases, the second portion of the audio signal 220 increases the out-of-head localization of the content within the audio signal 220 when output in the NF speaker 20 in synchronous with the output of the first portion of the audio signal 220. That is, the controllers 50a and 50b in the NF speakers 30A and 30B can be configured to output the second portion of the audio signal 220 to improve the out-of-head localization of the audio signal 220. This improved out-of-head localization can generate or increase the user's perception that the audio signal is output at one or more locations in three-dimensional space that are further from the user's head than the NF speakers 30A and 30B.

[0057] According to certain embodiments, a second portion of the audio signal 220 enhances the speech content within the audio signal 220. For example, the second portion of the audio signal 220 can increase the intelligibility of the speech content within the audio signal 220. In some examples, the speech content is removed from the audio track and then added back to the audio track in the second portion of the audio signal 220. As described herein, the signal processing (e.g., removal and addition of speech content) can be performed in one or more devices in space (e.g., space 105), for example, in FF speaker 20, NF speaker 30A, 30B, and / or additional device 40. In further embodiments, one or more portions of the signal processing are performed in a distributed computing system and / or a cloud computing system. In certain cases, as shown in Figure 2, near-field speakers 30A, 30B (e.g., wearable audio devices) add a level of audio detail (i.e., clarity) that is not easily achieved by far-field audio devices such as FF speakers 20 (e.g., loudspeakers located several meters or more away), thereby improving the impulse response of the audio output. In certain embodiments, far-field speakers 20 can supplement the bass output of non-occlusive (e.g., open-ear wearable audio devices) NF speakers 30A, 30B, which may struggle to provide desirable bass without occlusion. In such examples, NF speakers 30A, 30B and far-field speakers 20 have complementary audio output benefits (e.g., for users 110, 130 in Figure 2).

[0058] In a further embodiment, a first portion of the audio signal 220 (output by the FF speaker 20) does not contain spoken content, and a second portion of the audio signal 220 (output by the NF speakers 30A and 30B) contains spoken content. In one example, the language of the spoken content in the second portion of the audio signal 220 is selectable. For example, users 110 and 130 in the same space 105 may have different language preferences for audio playback (e.g., one has English as the primary language and the other has Spanish as the primary language). In such a case, the system 10 can enable one or both users 110 and 130 to select the language of the spoken content output to the NF speakers 30A and 30B. This language selection can be performed via a device interface, for example, via an interface on an additional device 40, or via any user interface on an audio device such as a wearable audio device 150. In these cases, since the first portion of the audio signal 220 output to the FF speaker 20 does not contain speech, language selection in the NF speakers 30A and 30B can be personalized without affecting other users in the space 105.

[0059] In certain embodiments, a first portion 230 and a second portion 240 of the audio signal 220 include different frequency ranges than the audio signal 220. In a particular example, the first portion 230 is output in a frequency range where the sound field diffuses (or is not modal), e.g., above about 200 Hz. In additional examples, the frequency limit in each device (e.g., FF speaker 20, and NF speakers 30A, 30B) is adjustable to improve output. In some cases, the adjustability of the frequency limit is based on the device type and / or user selectable (or user adjustable).

[0060] In further embodiments, the frequencies of the first portion 230 and the second portion 240 differ in at least one embodiment. In some examples, the first portion 230 of the audio signal has a wider frequency range than the second portion 240 of the audio signal. In some such embodiments, a portion of the frequency range of the first portion 230 overlaps with a portion of the frequency range of the second portion 240. In further examples, the entire frequency range of the second portion 240 is within the frequency range of the first portion 230. In additional embodiments, the second portion 240 of the audio signal excludes frequencies below a predetermined threshold.

[0061] According to a particular exemplary embodiment as shown in Figure 2, the FF speaker 20 is housed in a soundbar, such as one of the Bose Soundbars described herein. In further embodiments, the FF speaker 20 includes a plurality of speakers configured to output at least the left, right, and center channels of an audio signal 220 into space (e.g., space 105). For example, the FF speaker 20 may include a stereo pair set of portable speakers, e.g., portable speakers configured to operate separately as a stereo pair. In additional examples, the FF speaker 20 may include a plurality of speakers in a stereo and / or surround sound speaker set, such as two, three, four, or more speakers placed in space (e.g., space 105).

[0062] Specific examples of coordinated audio outputs between FF speaker 20 and NF speakers 30A, 30B are described herein, but in some cases, FF speaker 20 is configured to operate in a manner that is not synchronized with NF speakers 30A, 30B. Using Figure 2 as an example, in some cases, FF speaker 20 is configured to output a third portion of audio signal 220 when it does not provide output in synchronization with NF speakers 30A, 30B. In various embodiments, the third portion of audio signal 220 is different from the first portion of the audio signal. For example, the third portion of audio signal 220 may include the entire audio signal. In such cases, NF speakers 30A, 30B may be deactivated (e.g., in sleep mode and / or powered off) or may output audio different from audio signal 220.

[0063] In certain cases, the FF speaker 20 is configured to operate in different modes, for example, two or more modes. For example, in a first mode (mode 1), the FF speaker 20 is configured to output audio in synchronization with a pair of NF speakers 30A and 30B, as described in the scenarios with reference to Figures 2 to 4. In a second mode (mode 2), the FF speaker 20 is configured to output only its own audio, without outputting audio in synchronization with the pair of NF speakers 30A and 30B. In some of these cases, the NF speakers 30A and 30B are configured to operate differently when the FF speaker 20 is operating in a different mode. For example, when the FF speaker 20 is operating in mode 1, the NF speakers 30A and 30B are configured to output a subset of the frequency range of the audio signal (e.g., audio signal 220) when outputting audio in synchronization with the FF speaker 20. In an additional example, when the FF speaker 20 operates in mode 2, the NF speakers 30A and 30B are configured to output an audio signal (e.g., audio signal 220) with a wider frequency range than when the FF speaker 20 operates in mode 1, up to the full frequency range of the audio signal 220. That is, when the NF speakers 30A and 30B are not outputting audio in sync with the FF speaker 20, for example, when they are outputting only audio, the NF speakers 30A and 30B can output the full range of the audio signal 220, or at least a wider frequency range of the audio signal than when they are outputting in sync with the FF speaker 20.

[0064] As described herein, the user experience of the tuned output in the FF speaker 20 and NF speakers 30A, 30B may be tuned or otherwise modified in response to user commands, such as specific triggers and / or user interface commands. For example, as seen in Figures 2 to 4, a second portion 240 of the audio signal 220 is user-configurable to enhance out-of-head localization, spatialization, and / or dialogue. That is, the second portion of the audio signal 240 may be tuned (e.g., via user interface commands) to enhance the perception of sound outside the NF speakers 30A, 30B (e.g., within space 105), to enhance the perception of sound localized from a position around the user (e.g., within space 105), and / or to enhance the audio output of dialogue, such as dialogue in an entertainment program like a television program. In a particular example, the second portion of the audio signal 240 may be configurable to enhance specific acoustic effects in audio playback. For example, a second portion of the audio signal 240 may be configured to emphasize a portion of the audio track to provide a dramatic effect (e.g., to emphasize the playback of an instrument during a horror movie, to emphasize an actor's voice during a tense scene, to improve the clarity of dialogue in the audio signal 240, and / or to improve the spatialization of the audio output).

[0065] As described herein, the second portion of the audio signal 240 is configured to be user-adjustable according to, for example, user interface commands, user profiles or settings, the type of device configured to output the second audio signal 240 (e.g., NF speakers 30A, 30B), and / or the characteristics of the space (e.g., space 105). In a particular example, a separate user (e.g., user 110 versus user 130) may want to experience the output of the second portion of the audio signal 240 according to specific output settings (e.g., separate volume levels, separate equalization settings, etc.). These settings may be adjustable during output or at any other time via user interface adjustments and / or user profile adjustments. In certain cases, NF speakers 30A and 30B are configured to adjust the audio output settings of a second portion of the audio signal 240 according to known or otherwise detected characteristics of the space 105 and / or characteristics of the audio signal 240, for example, when the audio playback includes a podcast or talk show and the space 105 is a vehicle or other small cabin, the spatialization and / or dialogue clarity of the second portion of the audio signal 240 is improved. In further examples, separate users in a large space such as a concert hall, arena, or worship hall may want to adjust the audio output settings in NF speakers 30A and 30B to improve a particular portion of the audio output. For example, a user further from the soundstage may want separate equalization and / or volume settings for a user closer to the soundstage.

[0066] In some examples, a second portion 240 of the audio signal 220 is user-configurable via an interface such as a voice interface and / or application interface on a computing device, for example, a smart device such as the additional device 40B in Figure 2. In certain examples, the voice interface may include a voice interface such as a virtual personal assistant (VPA) that receives and / or processes commands in a wearable audio device such as the wearable audio device 150.

[0067] As described herein, according to certain examples, NF speakers 30A, 30B are configured to output a second portion 240 of the audio signal 220 in synchronization with the output of a first portion 230 of the audio signal 220 in response to a trigger. In certain cases, a trigger may include one or more of the following: a detected connection between two devices housing different speakers (such as a wired or wireless connection, such as a Wi-Fi connection, such as a BT or Wi-Fi RTT connection); a detected directional alignment between two devices housing different speakers (such as via BT angle of arrival (AoA) and / or angle of departure (AoD) data); a user grouping request (e.g., from an application or voice command, via an additional device 40B or a wearable audio device such as a wearable audio device 150); proximity or location detection (e.g., devices identified as being in close proximity to each other or in the same zone or space, via BT AoA and / or AoD data and / or via Wi-Fi RTT); or a user-initiated command or user response to a prompt following a trigger described herein. Any trigger described herein may cause the controller 50 in the device to prompt a user to initiate a synchronized audio output mode, for example, using a user interface prompt. The user can then take action in response to the prompt, for example, by accepting, rejecting, or failing to respond to the prompt through any user interface described herein. In certain embodiments, the user can accept a prompt to initiate a synchronized audio output mode using a gesture, such as a gesture detectable by a device housing the NF speakers 30A, 30B.

[0068] In certain cases, a second portion 240 of the audio signal 220 can be configured to adjust tone, equalization, volume and / or other parameters for different non-closed near-field speakers (e.g., NF speakers 30A, 30B). In additional embodiments, the adjustment of the audio output to the NF speakers 30A, 30B is configured to provide the same or proportional adjustment as the audio output to the FF speaker 20. In yet another example, the audio output adjustment in a given device (e.g., NF speakers 30A, 30B and / or far-field speaker 20) is limited so that at least one parameter of the audio signal is limited to a certain range. In some of these examples, volume adjustment may be limited to a certain range, tone adjustment to a certain range, and / or equalization adjustment to a certain range. In certain examples, user-configurable parameters (e.g., a second portion 240 of the audio signal 220) can be controlled using interfaces on the NF speakers 30A and 30B, such as adjustment interfaces like buttons and switches, touch interfaces like capacitive touch interfaces, or audio interfaces via microphones in the NF speakers 30A and 30B. In some specific examples, the adjustment interfaces on the NF speakers 30A and 30B may include scroll-through or cycle-through type adjustment mechanisms, allowing the user to cycle through or scroll through modes with each interface interaction, such as a button touch or a switch flip. In such cases, the user-configurable parameters (e.g., volume, tone, equalization, etc.) are adjustable within a defined range of predetermined (or lock-step) changes, where pressing a button once initiates a first change in the parameter, pressing the button again initiates a second (e.g., incremental) change in the parameter, and so on.

[0069] In certain cases, a second portion 240 of the audio signal 220 is automatically customized based on the detection of the type of device housing one or more of the NF speakers 30A, 30B, for example, whether the device includes audio glasses, on-ear headphones, near-ear headphones, or in-seat speakers. In some examples, the device type is detected using a device identifier, and the second portion 240 of the audio signal 220 is automatically customized based on the device identifier, for example, to adapt the audio output to different device types (e.g., different wearable audio devices). For example, the second portion 240 of the audio signal 220 can be output with distinct parameters (e.g., volume or equalization) based on the type of NF speakers 30A, 30B, and as a result, the output of the second portion 240 of the audio signal 220 is distinct in a pair of audio glasses compared to in-seat speakers. In certain embodiments, the customized audio output is triggered based on at least one of a Bluetooth (BT) connection (or a variant of a BT connection), a previous BT pairing scenario, or a previously defined coupling between devices. This connection trigger may be based on a detected device identifier and / or a connection between any devices within communication range in a space (e.g., space 105 in Figure 2). For example, detecting a BT connection, a previous BT pairing, or another previously defined coupling between FF speaker 20 and NF speakers 30A, 30B (e.g., via communication unit 60 in Figure 1) can trigger a customized audio output based on the device type of NF speakers 30A, 30B.

[0070] Returning to the exemplary configuration of Figure 2, still referring to Figure 1, in certain cases, the system 10 includes a first device, such as a soundbar or speaker housing the FF speaker 20. In such cases, the first device is configured to transmit either the audio signal 220 or a second portion 240 of the audio signal 220 to a second device, which includes NF speakers 30A, 30B. For example, the second device may include a seat 140 or a head / neck rest portion of the seat 140, or the second device may include a wearable audio device, such as a wearable audio device 150. The transmission of the audio signal 220 (or the second portion 240 of the audio signal 220) to the second device may include transmitting playback timing data to enable the second portion 240 of the audio signal 220 to be output in sync with the first portion 230 of the audio signal 220. In certain cases, the audio signal 220 or a second portion 240 of the audio signal 220 is transmitted wirelessly from the first device to the second device using one or more wireless transmission protocols. In certain cases, the audio signal 220 or a second portion 240 of the audio signal is transmitted from the first device to the second device via Bluetooth (BT) or BT Low Energy (LE) audio. In additional cases, the audio signal 220 or a second portion 240 of the audio signal is transmitted from the first device to the second device using broadcast (to the second device and at least one additional device), such as via synchronous unicast. In further cases, the audio signal 220 or a second portion 240 of the audio signal is transmitted from the first device to the second device via a synchronous downmix audio connection via BT or other wireless connection.In additional examples, an audio signal 220 or a second portion 240 of an audio signal is transmitted from a first device to a second device via multiple transmission streams, such as a broadcast, allowing different devices having different pairs of unoccluded near-field speakers (e.g., NF speakers 30A, 30B) to simultaneously output different portions of the audio signal 220 in synchronization with the first portion 230 of the audio signal 220. In some such examples, one portion focuses on improving spatialization and / or out-of-head localization, another portion focuses on improving dialogue, or one portion includes dialogue output in a first language and another portion includes dialogue output in a second different language. This exemplary configuration allows two users 110, 130 in the same space 105 to experience common content in different languages ​​and / or with different output parameters (e.g., different volume levels, equalization levels, etc.).

[0071] In addition, different sets of NF speakers 30 can receive the same audio signal (e.g., the same second portion 240 of the audio signal), but can process the received audio signal locally in different ways to provide audio output in separate sets of NF speakers 30. For example, a first set of NF speakers 30 and a second set of NF speakers 30 (e.g., in the same space 105 in Figure 2) can receive the same second portion 240 of the audio signal and process that second portion 240 of the audio signal locally in different ways to output separate audio to their respective users (e.g., the first set processes the second portion 240 to focus on improving dialogue intelligibility, and the second set processes the second portion 240 to focus on improving out-of-head localization / spatialization). Furthermore, different sets of NF speakers 30 can receive different audio signals, such as different second parts 240 or separate additional parts of the audio signal (for example, the first set of NF speakers 30 receives the second part, and the second set of NF speakers 30 receives the third part). In some of these examples, the separate second part, or separate additional parts (for example, the second and third parts), may be processed (for example, by the FF speaker 20) before being transmitted to the separate NF speaker 30.

[0072] In some cases, as described herein with respect to Figures 3 and 4, i) a first device (e.g., including an FF speaker 20) and a second device (e.g., including NF speakers 30A and 30B) receive an audio signal 220 from a common source device 210 (Figure 3), or ii) the first device receives a first portion 230 of the audio signal 220 from the first source device 210A, and the second device receives a second portion 240 of the audio signal 220 from the second source device 210B (Figure 4). In a particular example, in scenario (i), the common source device 210 may include a television or streaming video player (via wired and / or wireless transmission, for example), where the first and second devices wirelessly receive the audio signal 220 from the television or streaming video player, or the soundbar receives the audio signal 220 via a hardwired connection to the television or streaming video player, and the NF speakers 30A and 30B wirelessly receive the audio signal 220 from the television or streaming video player. In a particular example, in scenario (ii), the first device, such as a soundbar, receives the first portion 230 of the audio signal 220 from the first source device 210A, such as a television or streaming video player, and the second device, such as a wearable audio device, receives the second portion of the audio signal 220 from the second source device 210B, such as a smartphone, tablet, or computing device.

[0073] In certain embodiments, such as the example shown in Figure 3, the first and second devices receive an audio signal 220 from a common source device 210, the audio signal 220 including audio content that takes surround and height effects into account. That is, the audio signal 220 includes audio content that takes surround sound (e.g., radial) and / or height (e.g., vertical) effects into account in the audio output. Certain embodiments of audio content having surround and / or height effects are described in U.S. Patent Application No. 16 / 777,404 (US PG PUB No. 2021 / 0243544, filed January 30, 2020), which is incorporated herein by reference in its entirety. In some examples, outputting the audio signal 220 in the first and second devices evokes a room (e.g., space 105, Figure 2) housing the first and second devices in order to improve out-of-head localization for the user. In some of these cases, the audio output is perceived by the user (e.g., user 110 and / or user 130) as object-based (compared to channel-based), and therefore the user perceives the audio output as an object in space. In some examples, the system 10 provides an audio signal 220 having audio content that takes surround and height effects into account, in response to determining that NF speakers 30A, 30B are likely to be located at a minimum distance, e.g., several meters or more, from at least one FF speaker 20. In some of these examples, surround and height effects are beneficial in large spaces such as religious services, concert halls, arenas, conference rooms, and others where the distance between the FF speaker 20 and the NF speakers 30A, 30B meets a threshold.

[0074] Various embodiments include descriptions of non-occluded variants of the NF speakers 30A, 30B, but in additional embodiments, the NF speakers 30A, 30B may include occluded near-field speakers such as over-ear or in-ear headphones operating in transparent (or hear-through) mode. For example, a pair of headphones having passive and / or active noise-canceling capabilities may be substituted for the non-occluded variants of the NF speakers 30A, 30B described herein. In these cases, the occluded near-field speakers may operate in shared experience (or social) mode, which can be activated via user interface commands and / or any triggers described herein. In certain examples, the transparent (or hear-through) mode allows the user to experience ambient audio output from the FF speaker 20 while also experiencing coordinated playback of audio from the NF speakers 30A, 30B.

[0075] Furthermore, various embodiments are described as beneficially improving the user audio experience without knowledge of the user's head position (e.g., via user head tracking capability), but these embodiments may be used in conjunction with a system configured to track the user's head position. In such cases, data on the user's head position (e.g., as indicated by an IMU, optical tracking system, and / or proximity detection system) may be used as input to one or more processing components (e.g., in a controller 50) to further improve the user audio experience by adjusting the output of a second portion of the audio signal 240 (e.g., with respect to spatialization, out-of-head localization, etc.). However, data on the user's head position is not necessary for the beneficial development of the methods and systems according to the various embodiments.

[0076] In all cases, the methods described according to various embodiments have the technical effect of improving audio reproduction for users in an environment by utilizing both near-field (unobstructed) speakers and far-field speakers. For example, the methods described according to various embodiments adjust the audio output in separate speaker systems to improve individual and group experiences. The systems and methods described according to various embodiments enable users of unobstructed near-field speakers (e.g., in-seat speakers or wearable audio devices that do not obstruct the ear canal) to have a personalized audio experience without sacrificing the social aspects of an open-ear audio environment. For example, users of unobstructed near-field speakers can experience improved audio quality, clarity, personalization / customization, etc., compared to listening to far-field speakers only, without sacrificing social interaction with other users in the same space (due to the non-obstructive nature of near-field speakers). Furthermore, the systems and methods described herein enable users in the same space to share a common audio experience, i.e., audio content output via far-field speakers, while still allowing customization of audio content output in near-field speakers. Moreover, far-field speakers can improve the spatialization and / or out-of-head localization of audio output to the user in non-occluded scenarios, giving a greater impression of three-dimensional sound in space.

[0077] Various wireless connectivity scenarios are described herein. It should be understood that any number of wireless connectivity and / or communication protocols can be used to connect devices in a space, for example, space 105 (Figure 2). Examples of wireless connectivity scenarios and triggers for connecting wireless devices are described in further detail in U.S. Patent Applications No. 17 / 714,253 (filed April 4, 2022) and No. 17 / 314,270 (filed May 7, 2021), each of which is incorporated herein by reference in whole.

[0078] It is further understood that any RF protocol, including Bluetooth, Wi-Fi, or other proprietary or non-proprietary protocols, may be used to communicate between devices according to the embodiment. In embodiments in which the NF speakers 30A, 30B are housed within a wearable audio device (e.g., Figure 2), such embodiments can advantageously use wireless protocols other than those described herein (such as Bluetooth) that are used separately by the wearable audio device to receive audio data, thereby eliminating the need for the wearable audio device to include additional components and costs.

[0079] In embodiments utilizing Bluetooth LE Audio, a unicast topology may be used for a one-to-one connection between the FF speaker 20 and the NF speakers 30A and 30B. In some embodiments, an LE audio broadcast topology (such as broadcast audio) can be used to transmit one or more sets of audio data to multiple sets of NF speakers 30 (however, the broadcast topology can still be used for only one set of NF speakers 30A and 30B). For example, in some such embodiments, the broadcasted audio data is the same for all sets of NF speakers 30 within range so that all sets of NF speakers 30 receive the same audio content. However, in other such embodiments, different audio data is broadcast to sets of NF speakers 30 within range so that some NF speakers 30 can select first audio data and other NF speakers 30 can select second audio data different from the first audio data. Different audio data may enable differences in audio personalization (e.g., EQ settings), dialogue language selection, improved dialogue intelligibility, improved out-of-head localization / spatialization, and / or other differences, as can be understood based on this disclosure. Furthermore, using the LE audio broadcast topology (and other embodiments described in various ways herein) allows the volume level of the received audio content to be locally adjusted in each set of NF speakers 30 (as opposed to a single global volume level from the FF speakers 20 in system 10).

[0080] The above description provides embodiments compatible with BLUETOOTH SPECIFICATION version 5.2 [Vol 0], December 31, 2019, and any earlier versions, e.g., versions 4.x and 5.x devices. In addition, the connection techniques described herein may be used for Bluetooth LE audio, for example, to help establish a unicast connection. Furthermore, it should be understood that this technique is equally applicable to other wireless protocols (e.g., non-Bluetooth, future versions of Bluetooth, etc.) where a communication channel is selectively established between a pair of stations. Furthermore, while certain embodiments are described above as not requiring manual intervention to initiate pairing, in some embodiments, manual intervention may be required to complete pairing, for example, to provide an additional security aspect to the technique (e.g., a "Are you sure?" prompt presented to the user of the source / host device).

[0081] In some embodiments, the host-based elements of the method are implemented in a software module (e.g., an "app") that is downloaded and installed on a source / host (e.g., a "smartphone") in order to provide a coordinated audio output mode according to the method described above.

[0082] While the above describes a specific sequence of operations performed by a particular embodiment of the present invention, alternative embodiments may perform operations in a different order, combine certain operations, or duplicate certain operations; therefore, such an order is illustrative. References to given embodiments in this specification indicate that the embodiments described may include certain features, structures, or characteristics, but not all embodiments necessarily include such features, structures, or characteristics.

[0083] The functionalities or parts thereof described herein, and various modifications thereof (hereinafter referred to as "the Functionalities") may be implemented, at least in part, through computer program products (for example, computer programs tangibly embodied in information carriers such as non-temporary machine-readable media for execution by or control of the operation of one or more data processing devices (e.g., programmable processors, computers, multiple computers, and / or programmable logical components, etc.)).

[0084] Computer programs can be written in any form of programming language, including compiled or interpreted languages, and can be deployed as standalone programs or in any form, including modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs can be deployed to run on one computer or on multiple computers at one location, or they can be distributed across multiple locations and interconnected by a network.

[0085] The operations associated with performing all or part of the functions may be performed by one or more programmable processors that execute one or more computer programs to perform the functions of the calibration process. All or part of the functions may be implemented as special-purpose logic circuits, such as FPGAs and / or ASICs (Application-Specific Integrated Circuits). Suitable processors for executing computer programs include, by example, both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, the processor will receive instructions and data from read-only memory, random-access memory, or both. The components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.

[0086] In various embodiments, unless otherwise specified, electronic components described as "coupled" can be linked via conventional wired and / or wireless means so that these electronic components can communicate data with one another. Furthermore, subcomponents within a given component can be considered to be linked via conventional paths, although these are not necessarily illustrated.

[0087] Several embodiments have been described. Nevertheless, additional modifications can be made without departing from the scope of the concept of the present invention as described herein, and it will be understood that other embodiments also fall within the scope of the following claims. [Explanation of symbols]

[0088] 5 Space 10 Systems 20 Far-field (FF) speakers 30A, 30B Unclosed Near-Field (NF) Speakers 40, 40A, 40B Additional devices 50, 50a, 50b, 50c controllers 60, 60a, 60b, 60c communication units 70, 70a, 70b, 70c Bluetooth (BT) Modules 80, 80a, 80b, 80c microphones 90, 90a, 90b, 90c Electroacoustic Transducers 100, 100a, 100b, 100c Additional electronic equipment 105 Space 110 First User 120 seats 130 Second User 140 seats 150 Wearable Audio Devices 160 Headrest and / or neck / backrest section 210, 210A, 210B Source Devices 220 audio signals 230 (The first part of the audio signal) 240 (Second part of the audio signal)

Claims

1. It is an audio system, A far-field speaker configured to output a first portion of an audio signal, A pair of unclosed near-field speakers configured to output a second portion of the audio signal in synchronization with the output of a first portion of the audio signal in response to a trigger, The first portion of the audio signal differs from the second portion of the audio signal, The system is configured such that when the at least one far-field speaker does not provide output in synchronization with the pair of unbound near-field speakers, it outputs a third portion of the audio signal, the third portion of the audio signal being different from the first portion of the audio signal.

2. The system according to claim 1, wherein the first portion of the audio signal and the second portion of the audio signal include separate channels of common content.

3. The system according to claim 1, wherein the second portion of the audio signal is binaurally encoded, and the second portion of the audio signal is output in synchronization with the output of the first portion of the audio signal, thereby increasing the out-of-head localization of the content in the audio signal.

4. The system according to claim 1, wherein the second portion of the audio signal enhances the spoken content within the audio signal.

5. The system according to claim 4, wherein the second portion of the audio signal increases the intelligibility of the speech content in the audio signal.

6. The system according to claim 1, wherein the first portion of the audio signal does not include spoken content, and the second portion of the audio signal includes spoken content.

7. The system according to claim 1, wherein the first portion and the second portion of the audio signal include a different frequency range from the audio signal.

8. i) The first portion of the audio signal has a wider frequency range than the second portion of the audio signal, or ii) The system according to claim 7, wherein the second portion of the audio signal excludes frequencies below a predetermined threshold, at least one of these.

9. The system according to claim 1, wherein the pair of non-obstructed near-field speakers are housed within a wearable audio device.

10. The system according to claim 1, wherein the first speaker in the pair of non-obstructed near-field speakers is configured to output audio to the user's left ear, and the second speaker in the pair of non-obstructed near-field speakers is configured to output audio to the user's right ear.

11. i) The at least one far-field speaker is housed within a soundbar, or ii) The system according to claim 1, wherein the at least one far-field speaker includes a plurality of speakers configured to output at least the left channel, right channel, and center channel of the audio signal.

12. The system according to claim 1, wherein the second portion of the audio signal is user-configurable to enhance at least one of out-of-head localization, spatialization, or dialogue.

13. The system according to claim 1, wherein the second portion of the audio signal is automatically customized based on the detection of the type of device housing at least one of the pair of non-occluded near-field speakers.

14. The system according to claim 1, further comprising a first device housing at least one far-field speaker, the first device configured to transmit either the audio signal or the second portion of the audio signal to a second device including the pair of unbound near-field speakers, the transmission including playback timing data to enable the second portion of the audio signal to be output in sync with the first portion of the audio signal.

15. A first device including at least one far-field speaker, The device further comprises a second device including the pair of unobstructed near-field speakers, The first device and the second device receive the audio signal from a common source device, or The system according to claim 1, wherein the first device receives the first portion of the audio signal from the first source device, and the second device receives the second portion of the audio signal from the second source device.

16. The system according to claim 15, wherein the first device and the second device receive the audio signal from a common source device, the audio signal includes audio content that takes surround and height effects into consideration, and the output of the audio signal by the first device and the second device evokes a room housing the first device and the second device in order to improve out-of-head localization for the user.

17. A method for controlling an audio system, Outputting a first portion of the audio signal to at least one far-field speaker, In response to a trigger, the second portion of the audio signal is output to a pair of unclosed near-field speakers in synchronization with the output of the first portion of the audio signal. Includes, The first portion of the audio signal differs from the second portion of the audio signal, The aforementioned method, When not providing output in synchronization with the pair of non-closed near-field speakers, the system further includes outputting a third portion of the audio signal to at least one far-field speaker. The third portion of the audio signal is provided in a manner different from the first portion of the audio signal.

18. The method according to claim 17, wherein the first portion and the second portion of the audio signal include separate channels of common content, and the second portion of the audio signal is user-configurable to enhance at least one of out-of-head localization, spatialization, or dialogue.

19. The method according to claim 17, wherein the second portion of the audio signal is binaurally encoded, and the second portion of the audio signal is output in synchronization with the output of the first portion of the audio signal, thereby increasing the out-of-head localization of the content in the audio signal.

20. The method according to claim 17, wherein the second portion of the audio signal enhances the speech content within the audio signal, and the second portion of the audio signal increases the intelligibility of the speech content within the audio signal.

21. The method according to claim 17, wherein the first portion of the audio signal does not include spoken content, and the second portion of the audio signal includes spoken content.

22. The first portion and the second portion of the audio signal include a different frequency range from the audio signal. i) The first portion of the audio signal has a wider frequency range than the second portion of the audio signal, or ii) The method according to claim 17, wherein the second portion of the audio signal excludes frequencies below a predetermined threshold, at least one of these.

Citation Information

Patent Citations

  • Multi-dimensional sound playback apparatus and audio-signal playback method

    JP1994165285A