An audio system with improved mixed rendering audio
The integration of far-field and non-occluding near-field speakers in an audio system addresses the challenge of personalized audio delivery in shared environments, enhancing user experience through improved clarity and spatialization.
Patent Information
- Application Number
- JP2024572463
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-08
- Filing Date
- 2023-06-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-06-08
AI Technical Summary
Conventional audio systems fail to provide personalized audio experiences for multiple users in a shared listening environment, leading to insufficient personalization and reduced user experience due to factors like background noise, reverberation, and masking effects.
An audio system utilizing a combination of far-field and non-occluding near-field speakers, where the far-field speaker outputs a first portion of the audio signal and the near-field speakers output a synchronized second portion, tailored to enhance user experience with improved clarity, personalization, and spatialization.
The system provides enhanced audio quality, clarity, and personalization for individual users while maintaining social interaction, improving dialog intelligibility and out-of-head localization without sacrificing the shared audio experience.
Smart Images

Figure 2025521233000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to audio systems. More particularly, the present disclosure relates to improving mixed rendering audio in an audio system.
Background Art
[0002] In conventional loud audio systems, when multiple users are located within the listening environment, personalization becomes insufficient. That is, these conventional systems provide the same loud audio content to all users within the environment. This conventional approach may limit the user experience.
Summary of the Invention
Means for Solving the Problems
[0003] All examples and features mentioned below can be combined in any technically possible way.
[0004] Various embodiments include techniques for improving audio output using mixed rendering. Additional embodiments include an audio system configured to employ a mixed rendering technique to improve audio output.
[0005] In some particular aspects, the audio system includes at least one far-field speaker configured to output a first portion of an audio signal and a pair of non-occluding near-field speakers configured to output a second portion of the audio signal in synchronization with the output of the first portion of the audio signal, wherein the first portion of the audio signal is different from the second portion of the audio signal.
[0006] In an additional specific aspect, a method of controlling an audio system includes outputting a first portion of an audio signal to at least one remote field speaker and, in synchronization with the output of the first portion of the audio signal, outputting a second portion of the audio signal to a pair of non-enclosed near field speakers, wherein the first portion of the audio signal is different from the second portion of the audio signal.
[0007] Embodiments may include one or any combination of the following features.
[0008] Optionally, outputting the first portion of the audio signal and the second portion of the audio signal in synchronization is performed within about 100 milliseconds (ms).
[0009] In certain embodiments, the first portion of the audio signal and the second portion of the audio signal include separate channels of common content.
[0010] In a particular aspect, the second portion of the audio signal is binaurally encoded and increases the externalization of content in the audio signal when the second portion of the audio signal is output in synchronization with the output of the first portion of the audio signal.
[0011] Optionally, the second portion of the audio signal improves the speech content within the audio signal.
[0012] In certain embodiments, the second portion of the audio signal increases the intelligibility of the spoken content within the audio signal. In some examples, the spoken content is removed from the audio track and then added back to the audio track. In certain cases, a near-field device (e.g., a wearable audio device) adds a level of audio detail (i.e., clarity) that is not easily achieved from a far-field audio device (e.g., a loudspeaker several meters or more away), thereby improving the impulse response of the audio output. In certain aspects, the far-field audio device can supplement the bass output in an unoccluded near-field device (e.g., an open-ear wearable audio device), which may struggle to provide desirable bass without occlusion. In such examples, the near-field device and the far-field device have complementary audio output advantages.
[0013] In certain aspects, the first portion of the audio signal does not include spoken content and the second portion of the audio signal includes spoken content. In one example, the language of the spoken content in the second portion of the audio signal is selectable.
[0014] In certain embodiments, the first portion of the audio signal and the second portion of the audio signal include different frequency ranges than the audio signal. In certain examples, the first portion of the audio signal is output in a frequency range where the sound field is diffuse (or non-modal), e.g., above about 200 Hz. In additional examples, the frequency limits in each device are adjustable to improve the output.
[0015] In certain cases, i) the first portion of the audio signal has a wider frequency range than the frequency range of the second portion of the audio signal, or ii) the second portion of the audio signal excludes frequencies below a predetermined threshold.
[0016] In some embodiments, a pair of non-occluding near-field speakers are housed within a wearable audio device. In a particular example, the wearable audio device includes a pair of wirelessly coupled audio devices such as wireless earphones.
[0017] In additional embodiments, a pair of non-occluding near-field speakers are housed within a seat, such as the seat's neck and / or head support portion.
[0018] In certain embodiments, a first speaker in a pair of non-occluding near-field speakers is configured to output audio to a user's left ear, and a second speaker in the pair of non-occluding near-field speakers is configured to output audio to the user's right ear.
[0019] In some cases, i) at least one far-field speaker is housed within a soundbar, or ii) at least one far-field speaker includes a plurality of speakers configured to output at least the left channel, right channel, and center channel of an audio signal.
[0020] In a particular embodiment, at least one far-field speaker is configured to output a third portion of an audio signal when not providing output in synchronization with a pair of non-occluding near-field speakers, and the third portion of the audio signal is different from a first portion of the audio signal. In some examples, the third portion of the audio signal is the entire audio signal.
[0021] In certain cases, at least one far-field speaker is configured to operate differently in different modes, including, for example, Mode 1: outputting audio in synchronization with a pair of non-enclosed near-field speakers, and Mode 2: outputting only audio without outputting in synchronization with a pair of non-enclosed near-field speakers. In an additional aspect, a pair of non-enclosed near-field speakers is configured to operate differently in different modes, including, for example, Mode 1: when playing in synchronization with a far-field speaker, outputting a subset of the frequency range of the audio signal, and Mode 2: when outputting only audio (or not in synchronization with at least one far-field speaker), outputting an audio signal with a frequency range wider than that in Mode 1, up to the full frequency range of the audio signal.
[0022] In some cases, a second portion of the audio signal is user-configurable to improve at least one of out-of-head localization, spatialization, or dialog. In certain cases, the second portion of the audio signal is user-configurable via an interface on a computing device, such as a voice interface and / or application interface on a smart device. In certain cases, the second portion of the audio signal is configurable to adjust tones, equalization, volume, and other parameters to different non-enclosed near-field speakers. In an additional embodiment, the adjustment of the audio output to at least one far-field speaker is configured to make the same adjustment or a proportional adjustment as the audio output to the non-enclosed near-field speakers. In a further aspect, the adjustment of the audio output in a given device (e.g., a non-enclosed near-field speaker and / or a far-field speaker) is limited such that at least one parameter of the audio signal is limited to a certain range. In some of these aspects, the volume adjustment may be limited to a certain range, the tone adjustment may be limited to a certain range, and / or the equalization adjustment may be limited to a certain range.
[0023] In certain embodiments, user-configurable parameters (e.g., a second portion of an audio signal) can be controlled using an interface on the non-occlusive near-field speaker via an adjustment interface such as buttons, switches, a touch interface such as a capacitive touch interface, or an audio interface via a microphone in the non-occlusive near-field speaker. In some aspects, the adjustment interface in the non-occlusive near-field speaker can include a scroll-through or cycle-through type adjustment mechanism that enables a user, for example, to cycle or scroll through modes using each interface interaction such as a button touch or switch flip.
[0024] In additional embodiments, the non-occlusive near-field speaker can be replaced with an occlusive near-field speaker such as over-ear or in-ear headphones operating in a transparent (or see-through) mode. In these cases, the occlusive near-field speaker can operate in a shared experience (or social) mode.
[0025] In certain cases, the second portion of the audio signal is automatically customized based on detection of the type of device that houses at least one of a pair of non-occlusive near-field speakers. In some aspects, the type of device is detected using a device identifier, and the second portion of the audio signal is automatically customized based on the device identifier, for example, to adapt the audio output to a different device type (e.g., a different wearable audio device). In certain embodiments, the customized audio output is triggered based on at least one of a Bluetooth® (BT) connection (or a variant of a BT connection), a previous BT pairing scenario, or a previously defined association between devices.
[0026] In some embodiments, a pair of non-occlusive near-field speakers is configured to output a second portion of an audio signal in synchronization with the output of a first portion of the audio signal in response to a trigger. In certain cases, the trigger can include one or more of a detected connection between two devices housing different speakers (such as a wired or wireless connection like a BT or Wi-Fi connection), a detected alignment of directions between two devices housing different speakers (such as via BT angle of arrival (AoA) and / or angle of departure (AoD) data), a grouping request by a user (from an application or voice command, etc.), proximity or position detection (such as via BT AoA and / or AoD, or Wi-Fi RTT, etc., for devices identified as being in proximity to each other or in the same zone or space), or a user-initiated trigger or command for synchronous output.
[0027] In certain embodiments, the system further includes a first device that houses at least one far-field speaker, and the first device is configured to transmit either an audio signal or a second portion of the audio signal to a second device that includes a pair of non-occlusive near-field speakers, the transmission including playback timing data to enable the second portion of the audio signal to be output in synchronization with the first portion of the audio signal. In certain cases, the audio signal or the second portion of the audio signal is transmitted wirelessly via one or more of: broadcast (to the second device and at least one additional device) via Bluetooth (BT), BT Low Energy (LE) audio, synchronized unicast, etc.; synchronous downmix audio connection via BT or other wireless connection; for example, when one portion focuses on spatialization and / or improved out-of-head localization and another portion focuses on improved dialog, or when one portion includes dialog output in a first language and another portion includes dialog output in a second different language, etc., different devices having different pairs of non-occlusive near-field speakers are enabled to output different portions of the audio signal simultaneously in synchronization with the first portion of the audio signal via multiple transmission streams such as broadcast.
[0028] In some cases, the system further includes a first device including at least one long-range field speaker and a second device including a pair of non-occluding near-field speakers, where either i) the first device and the second device receive an audio signal from a common source device, or ii) the first device receives a first portion of the audio signal from a first source device and the second device receives a second portion of the audio signal from a second source device. In a particular aspect, in scenario (i), the common source device can include a television or a streaming video player (e.g., via wired and / or wireless transmission), for example, the first device and the second device wirelessly receive the audio signal from the television or the streaming video player, or the soundbar receives the audio signal via a hardwired connection to the television or the streaming video player, and the non-occluding near-field speaker wirelessly receives the audio signal from the television or the streaming video player. In a particular case, in scenario (ii), a first device such as a soundbar receives a first portion of the audio signal from a first source device such as a television or a streaming video player, and a second device such as a wearable audio device receives a second portion of the audio signal from a second source device such as a smartphone, a tablet, or a computing device.
[0029] In certain embodiments, the first device and the second device receive an audio signal from a common source device, the audio signal including audio content that takes into account surround effects and height effects, and outputting the audio signal in the first device and the second device causes the room that houses the first device and the second device to come to mind in order to improve the out-of-head localization for the user. In some of these cases, the audio output is perceived by the user as object-based (as compared to channel-based) such that the user perceives the audio output as an object in space. In some aspects, the system determines that the non-occluded near-field speaker is likely to be located at a minimum distance, e.g., a distance of several meters or more, from at least one far-field speaker, and in response, provides an audio signal having audio content that takes into account surround effects and height effects. In some of these aspects, the surround effects and height effects are beneficial in large spaces such as religious services, concert halls, arenas, conference rooms, etc., where the distance between the far-field speaker and the non-occluded near-field speaker meets a threshold.
[0030] Two or more features described in this disclosure, including the features described in the Summary section of the present invention, may be combined to form embodiments not specifically described herein.
[0031] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.
Brief Description of the Drawings
[0032]
Figure 1
Figure 2
Figure 3
Figure 4
[0033] Note that the drawings of the various embodiments are not necessarily to scale. The drawings are intended to show only typical aspects of the present disclosure and should not therefore be regarded as limiting the scope of the embodiments. In the drawings, like numerals represent like elements between the drawings.
[0034] The present disclosure is based at least in part on the recognition that coordinating audio output in separate speaker systems can improve individual user experiences as well as group experiences.
[0035] Conventional audio systems that use dialog enhancement rely on a mix of spectral modification and dynamic range compression. While this mixed approach has advantages, dialog intelligibility is still adversely affected by one or more speech transmission index (STI) factors, such as background noise, the quality of the audio playback device, reverberation in the space (e.g., the room), and masking effects from other audio content in the mix. Conventional systems are unable to adequately account for these STI factors and may reduce dialog intelligibility.
[0036] In addition, conventional audio systems that employ binaural rendering in closed wearable audio devices (e.g., over-ear headphones) rely on sophisticated room models and head position (i.e., tracking) data to localize virtual sound sources outside the head. These conventional approaches are costly and consume a significant amount of power. Further, processing head position data can introduce latency-related drawbacks to these conventional approaches.
[0037] The systems and methods disclosed in accordance with various embodiments use at least one remote field speaker to output a first portion of an audio signal and a pair of non-occluding (e.g., open ear) near field speakers to output a second, distinct portion of the audio signal in synchronization with the output of the first portion. In these cases, users of the non-occluding near field speakers (e.g., in-seat speakers or wearable audio devices that do not occlude the ear canal) can have a personalized audio experience without sacrificing the social aspects of an open ear audio environment. For example, a user of the non-occluding near field speakers can experience improved audio quality, clarity, personalization / customization, etc. compared to the case of listening to only the remote field speakers without sacrificing social engagement with other users in the same space (due to the non-occluding near field speakers). Further, the blend of audio output from the hybrid systems and techniques described herein can provide a number of audio consumption benefits such as improved personalization, improved out-of-head localization / spatialization, increased dialog intelligibility (e.g., for audio for video content or podcasts), and / or the ability to better experience the full frequency spectrum of the audio content (e.g., experiencing bass throughout the body compared to bass received from an open ear wearable audio device).
[0038] Commonly labeled components in the figures are considered to be substantially equivalent components for illustrative purposes, and duplicate descriptions of those components are omitted for clarity.
[0039] FIG. 1 shows an example of a space 5 that includes a system 10 that includes a set of devices according to various embodiments. In various embodiments, the devices shown in system 10 include at least one far-field (FF) speaker 20 and a pair of non-obstructive near-field (NF) speakers 30A, 30B. One or more additional devices 40 are shown, which are optional in some embodiments. The additional devices 40 can be configured to communicate with the FF speaker 20, the NF speakers 30A, 30B, and / or other electronic devices within the space 5 using any communication protocol or technique described herein. In certain aspects, the system 10 is disposed within or around a space 5, such as an enclosed or partially enclosed room, e.g., a home, office, theater, sports or entertainment venue, religious venue, etc. In some cases, the space 5 has one or more walls and a ceiling. In other cases, the space 5 includes an outdoor location without walls and / or a ceiling.
[0040] In various embodiments, at least one far-field speaker 20 includes a stand-alone loudspeaker or a set of loudspeakers, such as a soundbar, portable speaker, hard-wired (i.e., semi-permanent or installed) speaker, etc., that are not intended to be worn by a user. The speaker 20 is described as a "far-field" device, but the speaker 20 does not necessarily need to be located within the generally accepted "far-field" acoustic distance relative to any other device within the system 10. That is, the far-field speaker 20 does not need to be located in a far field in order to function according to the various embodiments described herein. In various embodiments, the minimum far-field distance is defined as being, in some cases, about 0.5 meters away, in additional cases as being about 2 meters or more away, and in certain cases as being about 3 meters or more away. It should be understood that the minimum far-field distance can vary based on the environment, e.g., within a vehicle (about 0.5 meters to about 2 meters), within a room in a home (e.g., about 2 meters to about 5 meters), or within an entertainment facility such as a concert hall (e.g., about 5 meters to about 50 meters).
[0041] In certain cases, speaker 20 includes a controller 50 and a communication (comm.) unit 60 coupled to the controller 50. In a particular example, the communication unit 60 includes a Bluetooth module 70 (e.g., including a Bluetooth radio) that enables communication with other devices via the Bluetooth protocol. In a particular exemplary embodiment, the speaker 20 can also include one or more microphones (mics) 80 (e.g., a single microphone or a microphone array) and at least one electroacoustic transducer 90 for providing an audio output. The speaker 20 can also include additional electronic devices 100 such as a power manager and / or a power source (e.g., a battery or a power connector), memory, sensors (e.g., an IMU, an accelerometer / gyroscope / magnetometer, an optical sensor, a voice activity detection system), etc. In some cases, the memory can include flash memory and / or non-volatile random access memory (NVRAM). In certain cases, the memory stores microcode of a program for processing and controlling the controller 50 and various reference data, data generated during the execution of any of the various programs executed by the controller 50, the Bluetooth connection process, and / or various updatable data for storage such as paired device data, connection data, device contact information, etc. Certain ones of the above components shown in FIG. 1 are optional and are shown in phantom lines.
[0042] In certain cases, the controller 50 can include one or more microcontrollers or processors having a digital signal processor (DSP). In some cases, the controller 50 is referred to as a control circuit. The controller 50 can be implemented as a chipset of chips including discrete and multiple analog and digital processors. The controller 50 can provide adjustment of other components of the speaker 20, such as, for example, control of a user interface (not shown) and applications executed by the speaker 20. In various embodiments, the controller 50 can include software and / or hardware for performing the audio control processes described herein, including an (audio) mix-drendering control module (or modules). For example, the controller 50 can include a mix-drendering control module in the form of a software stack having instructions for controlling the functionality when outputting audio to one or more speakers within the system 10 according to any of the embodiments described herein. As described herein, the controller 50, as well as other controllers described herein, are configured to control the functionality in the mix-drendering audio output techniques according to various embodiments.
[0043] The communication unit 60 can include a BT module 70 configured to use a wireless communication protocol such as Bluetooth, along with one or more additional network interfaces that use one or more additional wireless communication protocols such as other local area network (LAN) or personal area network (PAN) protocols like IEEE 802.11, Bluetooth Low Energy, or Wi-Fi. In certain embodiments, the communication unit 60 is particularly suitable for communicating with other communication units 60 within the BT devices 30, 40 via Bluetooth. In additional specific embodiments, the communication unit 60 is configured to communicate with the BT devices described herein using broadcast audio via BLE or a similar connection (including, for example, a proxy connection). In yet another embodiment, the communication unit 60 communicates wirelessly with any other device within the system 10 via one or more of: Bluetooth (BT), BT Low Energy (LE) audio, synchronized unicast, etc. (e.g., to one or both of the NF speakers 30A, 30B and / or an additional device 40), a synchronized downmix audio connection via BT or another wireless connection (also referred to as the proprietary connection protocol SimpleSync™ of Bose Corporation, Framingham, MA, USA), e.g., for different devices having different pairs of non-blocking near-field speakers (similar to, e.g., the NF speakers 30A, 30B) to be able to output different portions of an audio signal simultaneously in synchronization with a first portion of the audio signal, via one or more of a plurality of transmission streams such as broadcast. In yet another embodiment, the communication unit 60 is configured to communicate with any other device within the system 10 via, for example, a hardwired connection between any two or more devices.
[0044] As described in this specification, the controller 50 controls the general operation of the FF speaker 20. For example, the controller 50 performs processes such as audio and data communication with other devices (e.g., NF speakers 30A, 30B) and audio output and signal processing in the FF speaker 20. In addition to the general operation, when the controller 50 detects a specific trigger (or event) described in this specification, it starts the communication function implemented in the communication module 60. The controller 50 starts the operation (e.g., synchronization of audio output) between the FF speaker 20 and the NF speakers 30A, 30B when a specific condition is satisfied.
[0045] In a specific example, the Bluetooth module 70 enables a wireless connection that uses radio frequency (RF) communication between the FF speaker 20 and the NF speakers 30A, 30B (and, in some embodiments, the additional device 40). The Bluetooth module 70 transmits and receives wireless signals including data input and output via an antenna (not shown). For example, in the transmission mode, the Bluetooth module 70 processes the data by channel coding and spreading, converts the processed data into a radio frequency (RF) signal, and transmits the RF signal. In the reception mode, the Bluetooth module 70 converts the received RF signal into a baseband signal, processes the baseband signal by despreading and channel decoding, and restores the processed signal to data. In addition, the Bluetooth module 70 can ensure secure communication between devices and protect the data using encryption.
[0046] As referred to herein, a Bluetooth-enabled device includes a Bluetooth radio or other Bluetooth-specific communication system that enables a connection via the Bluetooth protocol. In the example shown in FIG. 1, the FF speaker 20 is a BT source device (also referred to as an "input device" or "host device"), and the NF speakers 30A, 30B are either part of a single BT sink device (also referred to as an "output device", "destination device", or "peripheral device") or separate BT sink devices. Exemplary Bluetooth-enabled source devices include, but are not limited to, smartphones, tablet computers, personal computers, laptop computers, notebook computers, netbook computers, radios, audio systems (e.g., portable and / or fixed), Internet Protocol (IP) phones, communication systems, entertainment systems, headsets, smart speakers, sports and / or fitness devices, portable media players, audio storage and / or playback systems, smartwatches or other smart wearable devices, etc. Exemplary Bluetooth-enabled sink devices include, but are not limited to, headphones, headsets, audio speakers (e.g., portable and / or fixed, with or without "smart" device capabilities), entertainment systems, communication systems, smartphones, vehicle audio systems, sports and / or fitness devices, outdoor (or, out-of-doors) audio devices, wearable private audio devices, etc. Additional BT devices can include portable game players, portable media players, audio gateways, BT gateway devices (for bridging BT connections between other BT-enabled devices), audio / video (A / V) receivers as part of a home entertainment or home theater system, etc. A Bluetooth-enabled device as described herein can change its role from source to sink, or from sink to source, depending on the particular application.
[0047] In various specific embodiments, the first speaker (e.g., NF speaker 30A) in the NF speaker is configured to output audio to the user's left ear, and the second speaker (e.g., NF speaker 30B) in the NF speaker is configured to output audio to the user's right ear. In certain embodiments, the NF speakers 30A, 30B are housed in a common device (e.g., included in a common housing) or otherwise form part of a common speaker system. For example, the NF speakers 30A, 30B can include in-seat or headrest speakers such as left / right speakers within the headrest and / or seatback portions of an entertainment seat, a gaming seat, a theater seat, an automotive seat, etc. In certain cases, the NF speakers 30A, 30B are positioned within the near field with respect to the user's ears, e.g., up to about 30 centimeters maximum. In some of these cases, the NF speakers 30A, 30B can include headrest speakers or speakers worn on the body or shoulders that are less than about 30 centimeters from the user's ears. In other specific cases, the near-field distance is, for example, up to about 15 centimeters when the NF speakers 30A, 30B include headrest speakers. In further specific cases, the near-field distance is, for example, up to about 10 centimeters when the NF speakers 30A, 30B include on-head or near-ear wearable audio devices. In additional specific cases, the near-field distance is, for example, up to about 5 centimeters when the NF speakers 30A, 30B include on-ear wearable audio devices. These exemplary near-field ranges are merely illustrative, and as described herein, various form factors can be considered within one or more of the exemplary ranges of the near field.
[0048] In yet another embodiment, the NF speakers 30A, 30B are part of a wearable audio device such as a wired or wireless wearable audio device. For example, the NF speakers 30A, 30B can include earphones within a wirelessly coupled or hard-wired connected wearable headset. The NF speakers 30A, 30B can also be part of a wearable audio device of any form factor, such as a pair of audio glasses, on-ear or near-ear audio devices, or an audio device placed on or around the user's head and / or shoulder area. As described herein according to certain embodiments, the NF speakers 30A, 30B are non-occluding near-field speakers, which means that when worn, the speakers 30 and their housing do not completely occlude (or block) the user's ear canal. That is, at least some ambient acoustic signals can pass through the user's ear canal without interference from the NF speakers 30A, 30B. In additional embodiments described further herein, the NF speakers 30A, 30B can include an occluding device (e.g., a pair of over-ear headphones or earphones having canal-sealing features) that allows the hear-through (or "awareness") mode to pass ambient acoustic signals as playback to the user's ear.
[0049] As shown in FIG. 1, the NF speakers 30A and 30B can include controllers 50a and 60b and communication units 60a and 60b (e.g., having BT modules 70a and 70b), enabling communication between the FF speaker 20 and the NF speakers 30A and 30B. The additional device 40 can include one or more of the components described with reference to the FF speaker 20, each of which is shown in dashed lines as optional in a particular embodiment. The notations "a" and "b" indicate that the components within the devices (e.g., NF speaker 30A, NF speaker 30B, additional device 40) are physically separate from the similarly labeled components within the FF speaker 20, but can take a similar form and / or function as their labeled counterparts within the FF speaker 20. Further description of these similarly labeled components is omitted for brevity. Additionally, as described herein, the additional NF speakers 30A and 30B and the additional device 40 can differ from the FF speaker 20 in terms of form factor, intended use, and / or capabilities, but in various embodiments are configured to communicate with the FF device 20 according to one or more of the communication protocols described herein (e.g., Bluetooth, BLE, broadcast, SimpleSync, etc.).
[0050] Generally, the Bluetooth modules 70, 70a, 70b include a Bluetooth radio and additional circuitry. More specifically, the Bluetooth modules 70, 70a, 70b include both a Bluetooth radio and a Bluetooth Low Energy (BLE) radio. In various embodiments, the presence of the BLE radio within the Bluetooth module 70 is optional. That is, as referred to herein, various embodiments utilize only the (classic) Bluetooth radio for connection capabilities. In embodiments that include a BLE radio, the Bluetooth radio and the BLE radio are typically on the same integrated circuit (IC) and share a single antenna, although in other embodiments, the Bluetooth radio and the BLE radio are implemented as two separate ICs that share a single antenna or as two separate ICs with two separate antennas. The Bluetooth specification, namely, Bluetooth 5.2: Low Energy, provides 40 channels at 2 MHz intervals to the FF speaker 20. The 40 channels are labeled 0 to 39 and include three advertising channels and 37 data channels. The channels labeled 37, 38, and 39 are designated as advertising channels in the Bluetooth specification, while the remaining channels 0 to 36 are designated as data channels in the Bluetooth specification. A particular exemplary technique for Bluetooth-related pairing is described in U.S. Patent No. 9,066,327 (issued June 23, 2015), which is incorporated herein by reference in its entirety. Further, a technique for selecting and / or prioritizing connections between paired devices is described in U.S. Patent Application Publication No. 17 / 314,270 (filed May 7, 2021), which is incorporated herein by reference in its entirety.
[0051] As described herein, various embodiments are particularly suitable for synchronous audio output with both the NF speakers 30A, 30B and the FF speaker 20. In certain cases, the synchronous audio output is adjusted in one or more additional devices 40. In certain cases, the controller 50 in one or more of the speakers is configured to adjust the synchronous audio output to improve the user experience, for example, to enable a personalized audio experience without sacrificing the social aspects of an open-ear audio environment. For example, a user of the NF speakers 30A, 30B can experience improved audio quality, clarity, personalization / customization, etc. compared to the case of listening only to the FF speaker 20 without sacrificing the social interaction with other users in the same space (due to the non-occlusive nature of the NF speakers 30).
[0052] Figure 2 shows an embodiment of an audio system 10 within a space 105 such as a room in a home, office, or entertainment facility. This space 105 is merely an example of various spaces from which the disclosed embodiments can benefit. In this example, a first user 110 is present at a first seating position (e.g., seat 120), and a second user 130 is present at a second seating position (e.g., seat 140). The user 110 is wearing a wearable audio device 150, which in this example includes a set of audio glasses such as Bose Frames audio glasses by Bose Corporation (Framingham, MA, USA). In other cases, the wearable audio device 150 can include another open-ear audio device such as a set of on-ear or near-ear headphones. In any case, the wearable audio device 150 includes a set of (e.g., two) non-occluding NF speakers 30A, 30B. The user 130 is positioned within a seat 140 that includes a set of non-occluding NF speakers 30A, 30B within a headrest and / or a neck / backrest portion 160. The space also shows an FF speaker 20, which can include a stand-alone speaker such as a soundbar (e.g., one of the Bose Smart Soundbar types by Bose Corporation) or a TV speaker (e.g., Bose TV Speaker by Bose Corporation). In additional cases, the FF speaker 20 can include a home theater speaker (e.g., of the Bose Surround Speaker type and / or Bose Bass Module type), or a portable speaker such as a portable smart speaker (e.g., of the Bose Soundlink type or one of the Bose Portable Smart Speaker by Bose Corporation), or a portable professional speaker such as the Bose S1 Pro Portable Speaker. In this non-limiting example, additional devices 40A and 40B are present within the space 105.For example, additional device 40A can include a television and / or visual display system (e.g., a projector-based video system or a smart monitor), and additional device 40B can include a smart device (e.g., a smartphone, a tablet computing device, a surface computing device, a laptop, etc.). The devices within space 105 and the interactions of the devices are merely intended to illustrate some of the various aspects of the present disclosure.
[0053] Referring to the exemplary example of FIG. 2, according to certain embodiments, the audio system 10 is configured to control synchronous audio output in both the FF speaker 20 and the NF speakers 30A, 30B. In certain cases, the processes executed in accordance with various embodiments are controlled by a controller in one or more of the speakers within the audio system 10, e.g., controller 50 within the FF speaker 20 and / or controllers 50a, 50b within the NF speakers 30A, 30B (FIG. 1). In certain embodiments, the FF speaker 20 is configured to output a first portion of an audio signal, and the NF speakers 30A, 30B are configured to output a second portion of the audio signal in synchronization with the output of the first portion of the audio signal. FIG. 3 shows a signal flow diagram illustrating an exemplary audio signal flow in relation to FIG. 2. As shown in this example, the FF speaker 20 and the NF speakers 30A, 30B can be connected to a common source device 210. In certain embodiments, the source device 210 can include one of the additional devices (e.g., television, audio gateway device, smartphone, tablet computing device, etc.) 40 described herein. For example, the source device 210 can include a television system, a smartphone, or a tablet. In additional embodiments, the source device 210 includes network-based and / or cloud-based devices such as network-connected audio systems. In further embodiments, the FF speaker 20 and / or the NF speakers 30A, 30B function as a source device having, for example, integrated network and / or cloud communication capabilities. In such a case, the FF speaker 20 and / or the NF speakers 30A, 30B receive an audio signal from a network (or cloud) connected gateway device such as a wireless or hardwired internet router. In an example shown in FIG. 3, the source device 210 is a network and / or cloud connected device that executes a software program or software application (also referred to as an "app") configured to manage the audio output to the FF speaker 20 and / or the NF speakers 30A, 30B.In a particular example, the source device 210 transmits signals to both the FF speaker 20 and the NF speakers 30A, 30B. In an additional example, the source device 210 transmits signals to the FF speaker 20 or the NF speakers 30A, 30B (one or both of them), and these signals are transferred between their speaker connections. In a particular embodiment, the NF speakers 30A, 30B transfer signals via a "snooping" type of approach or otherwise synchronize the outputs. While particular exemplary scenarios are described herein, the FF speaker 20 and the NF speakers 30A, 30B can transfer signals or otherwise transmit in any technically feasible way, and the examples described herein (e.g., SimpleSync, broadcast, BT, etc.) should not be considered to limit the various embodiments.
[0054] In a particular example as shown in FIG. 3, the source device 210 transmits the audio signal 220 to the FF speaker 20 and the NF speakers 30A, 30B. In a further embodiment, as shown in FIG. 4, two separate source devices 210A and 210B transmit signals to the FF speaker 20 and the NF speakers 30A, 30B, respectively. For example, using the space 105 of FIG. 2 as an example, an additional device 40A (e.g., a video system) can transmit a first portion 230 of the audio signal 220 to the FF speaker 20, and an additional device 40B (e.g., a smartphone or a tablet) can transmit a second portion 240 of the audio signal 220 to one or both sets of the NF speakers 30A, 30B.
[0055] Continuing to refer to FIGS. 2, as well as FIGS. 3 and 4, in an exemplary embodiment, the FF speaker 20 is configured to output a first portion of an audio signal 220 (e.g., received from the source device 210), and the NF speakers 30A, 30B are configured to output a second distinct portion of the audio signal 220 (e.g., received from the source device 210) in synchronization with the output of the first portion from the FF speaker 20. In some cases, the synchronized output of the first portion 220 of the audio signal and the second portion 220 of the audio signal includes outputting those portions at separate speakers within approximately 100 milliseconds (ms) of each other.
[0056] In some cases, the first portion of the audio signal 220 and the second portion of the audio signal 220 include separate channels of common content. In certain aspects, the second portion of the audio signal 220 is binaurally encoded, for example, to improve the output at the NF speakers 30A, 30B. In these cases, the second portion of the audio signal 220 increases the out-of-head localization of the content within the audio signal 220 when output in synchronization with the output of the first portion of the audio signal 220 at the NF speaker 20. That is, the controllers 50a, 50b at the NF speakers 30A, 30B can be configured to output the second portion of the audio signal 220 to improve the out-of-head localization of the audio signal 220. This improved out-of-head localization can create or increase the user's perception that the audio signal is being output at one or more positions within a three-dimensional space that is farther from the user's head than the NF speakers 30A, 30B.
[0057] According to certain embodiments, the second portion of the audio signal 220 improves the speech content within the audio signal 220. For example, the second portion of the audio signal 220 can increase the intelligibility of the speech content within the audio signal 220. In some examples, the speech content is removed from the audio track and then added back to the audio track in the second portion of the audio signal 220. As described herein, the signal processing (e.g., removal and re-addition of speech content) can be performed in one or more of the devices within the space (e.g., space 105), such as in the FF speaker 20, NF speakers 30A, 30B, and / or additional device 40. In yet further embodiments, one or more portions of the signal processing are performed in a distributed computing system and / or a cloud computing system. In a particular case as shown in FIG. 2, the near-field speakers 30A, 30B (e.g., wearable audio devices) add an audio detail (i.e., clarity) at a level that is not easily achieved from a far-field audio device such as the FF speaker 20 (e.g., a loudspeaker several meters away), thereby improving the impulse response of the audio output. In a particular aspect, the far-field speaker 20 can supplement the bass output in the NF speakers 30A, 30B, which are non-occlusive (e.g., open-ear wearable audio devices), which may struggle to provide desirable bass without occlusion. In such an example, the NF speakers 30A, 30B and the far-field speaker 20 have complementary audio output benefits (e.g., for users 110, 130 in FIG. 2).
[0058] In a further embodiment, a first portion of the audio signal 220 (output in the FF speaker 20) does not include speech content, and a second portion of the audio signal 220 (output in the NF speakers 30A, 30B) includes speech content. In one example, the language of the speech content in the second portion of the audio signal 220 is selectable. For example, users 110 and 130 within the same space 105 may have different language preferences for audio playback (e.g., one has English as a primary language and one has Spanish as a primary language). In such a case, the system 10 can enable one or both of the users 110, 130 to select the language of the speech content output to the NF speakers 30A, 30B. This language selection can be performed via a device interface, e.g., via an interface on an additional device 40, or via any user interface on an audio device such as the wearable audio device 150. In these cases, since the first portion of the audio signal 220 output to the FF speaker 20 does not include speech, the language selection in the NF speakers 30A, 30B can be personalized without affecting other users within the space 105.
[0059] In certain embodiments, the first portion 230 of the audio signal 220 and the second portion 240 of the audio signal 220 include different frequency ranges than the audio signal 220. In a particular example, the first portion 230 is output in a frequency range where the sound field is diffuse (or non-modal), e.g., above about 200 Hz. In a further example, the frequency limits in each device (e.g., the FF speaker 20, and the NF speakers 30A, 30B) are adjustable to improve the output. In some cases, the adjustability of the frequency limits is based on the device type and / or is user-selectable (or user-adjustable).
[0060] In a further embodiment, the frequencies of the first portion 230 and the second portion 240 are different in at least one aspect. In some examples, the first portion 230 of the audio signal has a wider frequency range than the frequency range of the second portion 240 of the audio signal. In some such embodiments, a portion of the frequency range of the first portion 230 overlaps with a portion of the frequency range of the second portion 240. In a further example, the entire frequency range of the second portion 240 is within the frequency range of the first portion 230. In an additional embodiment, the second portion 240 of the audio signal excludes frequencies below a predetermined threshold.
[0061] According to certain exemplary embodiments such as shown in FIG. 2, the FF speaker 20 is housed in a soundbar such as one of the Bose Soundbars described herein. In a further embodiment, the FF speaker 20 includes a plurality of speakers configured to output at least the left channel, right channel, and center channel of the audio signal 220 into a space (e.g., space 105). For example, the FF speaker 20 can include a stereo pair set of portable speakers, e.g., portable speakers configured to operate separately in the same manner as a stereo pair. In an additional example, the FF speaker 20 can include a plurality of speakers within a stereo and / or surround sound speaker set, such as two, three, four, or more speakers disposed within a space (e.g., space 105).
[0062] Specific examples of adjusted audio output between the FF speaker 20 and the NF speakers 30A, 30B are described herein, but in some cases, the FF speaker 20 is configured to operate in a manner that is not synchronized with the NF speakers 30A, 30B. Using FIG. 2 as an example, in some cases, when the FF speaker 20 does not provide output in synchronization with the NF speakers 30A, 30B, it is configured to output a third portion of the audio signal 220. In various embodiments, the third portion of the audio signal 220 is different from the first portion of the audio signal. For example, the third portion of the audio signal 220 can include the entire audio signal. In such a case, the NF speakers 30A, 30B may be deactivated (e.g., in a sleep state and / or powered off), or may output audio that is different from the audio signal 220.
[0063] In a specific example, the FF speaker 20 is configured to operate in different modes, for example, two or more modes. For example, in the first mode (Mode 1), the FF speaker 20 is configured to output audio in synchronization with a pair of NF speakers 30A and 30B as in the scenario described with reference to FIGS. 2 to 4. In the second mode (Mode 2), the FF speaker 20 is configured to output only its audio without outputting audio in synchronization with the pair of NF speakers 30A and 30B. In some of these cases, the NF speakers 30A and 30B are configured to operate differently when the FF speaker 20 is operating in different modes. For example, when the FF speaker 20 is operating in Mode 1, the NF speakers 30A and 30B are configured to output a subset of the frequency range of the audio signal (for example, audio signal 220) when outputting audio in synchronization with the FF speaker 20. In an additional example, when the FF speaker 20 is operating in Mode 2, the NF speakers 30A and 30B are configured to output an audio signal (for example, audio signal 220) with a wider frequency range than when the FF speaker 20 is operating in Mode 1, up to the full frequency range of the audio signal 220. That is, when the NF speakers 30A and 30B are not outputting audio in synchronization with the FF speaker 20, for example, when outputting only audio, the NF speakers 30A and 30B can output up to the full range of the audio signal 220, or at least a wider frequency range of that audio signal than when outputting in synchronization with the FF speaker 20.
[0064] As described herein, the user experience of the adjusted output in the FF speaker 20 and the NF speakers 30A, 30B can be adjustable or otherwise modifiable in response to user commands, such as certain triggers and / or user interface commands. For example, referring to FIGS. 2-4, the second portion 240 of the audio signal 220 is user configurable to improve out-of-head localization, spatialization, and / or dialog. That is, the second portion of the audio signal 240 is to improve the perception of sound outside the NF speakers 30A, 30B (e.g., within the space 105), to improve the perception of localized sound from positions around the user (e.g., within the space 105), and / or to improve the audio output of dialog, such as dialog in an entertainment program like a television program, and can be adjustable (e.g., via user interface commands). In a particular example, the second portion of the audio signal 240 can be configurable to improve certain acoustic effects in audio playback. For example, the second portion of the audio signal 240 can be configured to emphasize a portion of an audio track to provide a dramatic effect (e.g., to emphasize the playback of an instrument during a scary movie, to emphasize an actor's voice during a tense scene, to improve the clarity of dialog in the audio signal 240, and / or to improve the spatialization of the audio output).
[0065] As described herein, the second portion of the audio signal 240 is configured to be user-adjustable according to, for example, user interface commands, user profiles or settings, the type of device (e.g., NF speakers 30A, 30B) configured to output the second audio signal 240, and / or the characteristics of the space (e.g., space 105). In certain examples, different users (e.g., user 110 vs. user 130) may desire to experience the output of the second portion of the audio signal 240 according to specific output settings (e.g., different volume levels, different equalization settings, etc.). These settings may be adjustable during output or at any other time via user interface adjustments and / or user profile adjustments. In certain examples, the NF speakers 30A, 30B are configured to adjust the audio output settings of the second portion of the audio signal 240 according to the characteristics of the space 105 detected in a known or other manner and / or the characteristics of the audio signal 240. For example, when the audio playback includes a podcast or talk show and the space 105 is a vehicle or other small cabin, the spatialization and / or dialog clarity of the second portion of the audio signal 240 is improved. In a further example, different users within a large space such as a concert venue, arena, house of worship, etc. may desire to adjust the audio output settings in the NF speakers 30A, 30B to improve a particular portion of the audio output. For example, users further away from the sound stage may desire different equalization and / or volume settings than users closer to the sound stage.
[0066] In some examples, the second portion 240 of the audio signal 220 may be user-configurable via an interface such as a voice interface and / or an application interface on a computing device, such as the additional device 40B of FIG. 2, etc., a smart device. In a particular example, the voice interface may include a voice interface such as a virtual personal assistant (VPA) that receives and / or processes commands in a wearable audio device such as the wearable audio device 150.
[0067] As described herein, according to certain examples, the NF speakers 30A, 30B are configured to output a second portion 240 of the audio signal 220 in synchronization with the output of a first portion 230 of the audio signal 220 in response to a trigger. In certain cases, the trigger can include a detected connection between two devices housing different speakers (such as a wired connection like a BT or Wi-Fi connection like a Wi-Fi RTT connection), a detected direction alignment between two devices housing different speakers (such as via BT angle of arrival (AoA) and / or angle of departure (AoD) data), a grouping request by the user (such as via an application or voice command, for example, through a wearable audio device like an additional device 40B or a wearable audio device 150), proximity or location detection (such as devices identified as being in proximity to each other or in the same zone or space via BT AoA and / or AoD data and / or via Wi-Fi RTT), or one or more of a user-initiated command or user response to a prompt following a trigger described herein. Any trigger described herein can prompt the user, via a controller 50 in the device, for example, using a user interface prompt, to initiate a synchronized audio output mode. The user can then take an action in response to the prompt by, for example, accepting, rejecting, or failing to respond to the prompt via any user interface described herein. In certain embodiments, the user can accept a prompt to initiate a synchronized audio output mode using a gesture, such as a gesture detectable by a device housing the NF speakers 30A, 30B.
[0068] In certain cases, the second portion 240 of the audio signal 220 can be configured to adjust tones, equalization, volume, and / or other parameters for different non-occluded near-field speakers (e.g., NF speakers 30A, 30B). In additional embodiments, the adjustment of the audio output to the NF speakers 30A, 30B is configured to provide the same adjustment or a proportional adjustment as the audio output to the FF speaker 20. In yet further examples, the adjustment of the audio output in a given device (e.g., NF speakers 30A, 30B and / or far-field speaker 20) is limited such that at least one parameter of the audio signal is restricted to a range. In some of these examples, the volume adjustment can be limited to a range, the tone adjustment can be limited to a range, and / or the equalization adjustment can be limited to a range. In a particular example, user-configurable parameters (e.g., the second portion 240 of the audio signal 220) can be controlled using an interface on the NF speakers 30A, 30B via an adjustment interface such as a button, switch, etc., a touch interface such as a capacitive touch interface, or an audio interface via a microphone in the NF speakers 30A, 30B. In some particular examples, the adjustment interface in the NF speakers 30A, 30B can include a scroll-through or cycle-through type of adjustment mechanism, for example, enabling the user to cycle or scroll through modes with each interface interaction such as a button touch or switch flip. In such cases, user-configurable parameters (e.g., volume, tone, equalization, etc.) can be adjusted within a defined range of predetermined (or lock-step) changes, such that pressing a button once initiates a first change in the parameter and pressing the button again initiates a second (e.g., incremental) change in the parameter until the range of the adjustment range is reached.
[0069] In certain cases, the second portion 240 of the audio signal 220 is automatically customized based on the type of device that houses one or more of the NF speakers 30A, 30B, e.g., whether the device includes audio glasses, on-ear headphones, near-ear headphones, or in-seat speakers. In some examples, the type of device is detected using a device identifier, and the second portion 240 of the audio signal 220 is automatically customized based on the device identifier, e.g., to adapt the audio output to different device types (e.g., different wearable audio devices). For example, the second portion 240 of the audio signal 220 can be output with different parameters (e.g., volume or equalization) based on the type of the NF speakers 30A, 30B, such that the output of the second portion 240 of the audio signal 220 is distinct in a pair of audio glasses as compared to in-seat speakers. In certain embodiments, the customized audio output is triggered based on at least one of a Bluetooth (BT) connection (or a variant of a BT connection), a previous BT pairing scenario, or a previously defined association between devices. This connection trigger can be based on the detected device identifier and / or a connection between any devices within the communication range in the space (e.g., space 105 of FIG. 2). For example, detecting a BT connection, a previous BT pairing, or another previously defined association between the FF speakers 20 and the NF speakers 30A, 30B (e.g., via the communication unit 60 of FIG. 1) can trigger a customized audio output based on the device type of the NF speakers 30A, 30B.
[0070] Returning to the exemplary configuration of FIG. 2 while continuing to refer to FIG. 1, in certain cases, system 10 includes a first device such as a sound bar or speaker that houses FF speaker 20. In such cases, the first device is configured to send either audio signal 220 or a second portion 240 of audio signal 220 to a second device that includes NF speakers 30A, 30B. For example, the second device can include seat 140 or the head / neck rest portion of seat 140, or the second device can include a wearable audio device such as wearable audio device 150. Sending audio signal 220 (or the second portion 240 of audio signal 220) to the second device can include sending playback timing data to enable the second portion 240 of audio signal 220 to be output in synchronization with the first portion 230 of audio signal 220. In a particular example, audio signal 220 or the second portion 240 of audio signal 220 is wirelessly transmitted from the first device to the second device using one or more wireless transmission protocols. In certain cases, audio signal 220 or the second portion 240 of the audio signal is transmitted from the first device to the second device via Bluetooth (BT) or BT Low Energy (LE) audio. In additional cases, audio signal 220 or the second portion 240 of the audio signal is transmitted from the first device to the second device using a broadcast (to the second device and at least one additional device), such as via synchronous unicast. In further cases, audio signal 220 or the second portion 240 of the audio signal is transmitted from the first device to the second device via a synchronous downmix audio connection via BT or other wireless connection.In an additional example, the audio signal 220 or the second portion 240 of the audio signal is transmitted from a first device to a second device via a plurality of transmission streams such as, for example, a broadcast, and different devices having different pairs of non-blocking near-field speakers (e.g., NF speakers 30A, 30B) are enabled to simultaneously output different portions of the audio signal 220 in synchronization with the first portion 230 of the audio signal 220. In some such examples, one portion focuses on improving spatialization and / or out-of-head localization, another portion focuses on improving dialog, or one portion includes dialog output in a first language and another portion includes dialog output in a second different language. This exemplary configuration enables two users 110, 130 within the same space 105 to experience common content in different languages and / or using different output parameters (e.g., at different volumes, equalization levels, etc.).
[0071] In addition, different sets of the NF speakers 30 can receive the same audio signal (e.g., the same second portion 240 of the audio signal), but the received audio signal can be locally processed in different ways to provide an audio output in a separate set of the NF speakers 30. For example, a first set of the NF speakers 30 and a second set of the NF speakers 30 (e.g., within the same space 105 of FIG. 2) can receive the same second portion 240 of the audio signal and locally process that second portion 240 of the audio signal differently to output separate audio to respective users (e.g., the first set processes the second portion 240 to focus on improving dialog intelligibility, and the second set processes the second portion 240 to focus on improving out-of-head localization / spatialization). Further, different sets of the NF speakers 30 can receive different second portions 240 of the audio signal, or different audio signals such as separate additional portions (e.g., a first set of the NF speakers 30 receives a second portion and a second set of the NF speakers 30 receives a third portion, etc.). In some of these examples, the separate second portion, or separate additional portions (e.g., the second portion and the third portion) can be processed (e.g., by the FF speakers 20) before being transmitted to separate NF speakers 30.
[0072] In some cases, as described herein with respect to FIGS. 3 and 4, i) a first device (e.g., including the FF speaker 20) and a second device (e.g., including the NF speakers 30A, 30B) receive an audio signal 220 from a common source device 210 (FIG. 3), or ii) the first device receives a first portion 230 of the audio signal 220 from a first source device 210A, and the second device receives a second portion 240 of the audio signal 220 from a second source device 210B (FIG. 4). In a particular example, in scenario (i), the common source device 210 can include a television or a streaming video player (such as via wired and / or wireless transmission), for example, the first device and the second device wirelessly receive the audio signal 220 from a television or a streaming video player, or the soundbar receives the audio signal 220 via a hard-wired connection with a television or a streaming video player, and the NF speakers 30A, 30B wirelessly receive the audio signal 220 from a television or a streaming video player. In a particular example, in scenario (ii), a first device such as a soundbar receives a first portion 230 of the audio signal 220 from a first source device 210A such as a television or a streaming video player, and a second device such as a wearable audio device receives a second portion of the audio signal 220 from a second source device 210B such as a smartphone, a tablet, or a computing device.
[0073] In certain embodiments, such as those shown in the example of FIG. 3, the first device and the second device receive an audio signal 220 from a common source device 210, and the audio signal 220 includes audio content that takes into account surround and height effects. That is, the audio signal 220 includes audio content that takes into account surround sound (e.g., radial direction) and / or height (e.g., vertical direction) effects in the audio output. Particular aspects of audio content having surround and / or height effects are described in U.S. Patent Application No. 16 / 777,404 (US PG PUB No. 2021 / 0243544, filed Jan. 30, 2020), which is hereby incorporated by reference in its entirety. In some examples, outputting the audio signal 220 in the first device and the second device causes the room (e.g., space 105, FIG. 2) housing the first device and the second device to come to mind in order to improve the out-of-head localization for the user. In some of these cases, the audio output is perceived by the user (e.g., user 110 and / or user 130) as object-based (as compared to channel-based), and thus the user perceives the audio output as an object within the space. In some examples, the system 10 provides an audio signal 220 having audio content that takes into account surround and height effects in response to determining that the NF speakers 30A, 30B are likely to be located at a minimum distance, e.g., several meters or more, from at least one FF speaker 20. In some of these examples, the surround and height effects are beneficial in large spaces such as religious services, concert halls, arenas, conference rooms, and others where the distance between the FF speaker 20 and the NF speakers 30A, 30B meets a threshold.
[0074] While various embodiments include descriptions of unoccluded deformable forms of the NF speakers 30A, 30B, in additional embodiments, the NF speakers 30A, 30B can include occluded near-field speakers such as over-ear or in-ear headphones that operate in a transparent (or through) mode. For example, a pair of headphones having passive and / or active noise cancellation capabilities can be substituted for the unoccluded deformable forms of the NF speakers 30A, 30B described herein. In these cases, the occluded near-field speakers can operate in a shared experience (or social) mode, which can be enabled via user interface commands and / or any of the triggers described herein. In a particular example, the transparent (or through) mode enables the user to experience the ambient audio output by the FF speaker 20 while also experiencing the coordinated playback of audio from the NF speakers 30A, 30B.
[0075] Furthermore, while various embodiments are described as beneficially improving the user audio experience without knowledge of the user head position (e.g., via user head tracking capabilities), these embodiments can be used with a system configured to track the user head position. In such cases, data regarding the user's head position (such as indicated by, e.g., an IMU, an optical tracking system, and / or a proximity detection system) can be used as an input to one or more processing components (e.g., in the controller 50) to further improve the user audio experience, e.g., by adjusting the output of a second portion of the audio signal 240 (with respect to, e.g., spatialization, external head localization, etc.). However, data regarding the user head position is not necessary to beneficially deploy the techniques and systems according to the various embodiments.
[0076] In any case, the techniques described according to various embodiments have the technical effect of improving audio playback for users in the environment by utilizing both near-field (non-occlusive) speakers and far-field speakers. For example, the techniques described according to various embodiments adjust the audio output in a separate speaker system to improve individual user experiences as well as group experiences. The systems and methods described according to various embodiments enable a user of a non-occlusive near-field speaker (e.g., an in-seat speaker or a wearable audio device that does not occlude the ear canal) to have a personalized audio experience without sacrificing the social aspects of an open-ear audio environment. For example, a user of a non-occlusive near-field speaker can experience improved audio quality, clarity, personalization / customization, etc., compared to the case of only listening to far-field speakers, without sacrificing the social interaction with other users in the same space (due to the non-occlusive nature of the near-field speaker). Further, the systems and methods described herein enable users in the same space to share a common audio experience, i.e., audio content output via far-field speakers, while still allowing for customization of the audio content output in the near-field speakers. Further, far-field speakers can improve the spatialization and / or externalization of audio output to the user in a non-occlusive scenario and can give a greater impression of the three-dimensional sound in the space.
[0077] Various wireless connection scenarios are described herein. It should be understood that any number of wireless connections and / or communication protocols can be used to couple devices within a space, e.g., within space 105 (FIG. 2). Examples of wireless connection scenarios and triggers for connecting wireless devices are described in more detail in U.S. Patent Application No. 17 / 714,253, filed Apr. 4, 2022, and No. 17 / 314,270, filed May 7, 2021, each of which is hereby incorporated by reference in its entirety.
[0078] It is further understood that any RF protocol, including Bluetooth, Wi-Fi, or other proprietary or non-proprietary protocols, can be used to communicate between devices in accordance with embodiments. In embodiments where the NF speakers 30A, 30B are housed within a wearable audio device (e.g., FIG. 2), such embodiments can advantageously use a wireless protocol that is otherwise used by the wearable audio device to receive audio data other than the techniques described herein (such as Bluetooth), thereby eliminating the need for the wearable audio device to include additional components and costs.
[0079] In embodiments that utilize BluetoothLE Audio, a unicast topology may be used for one-to-one connections between the FF speaker 20 and the NF speakers 30A, 30B. In some embodiments, an LE Audio broadcast topology (such as broadcast audio) may be used to transmit one or more sets of audio data to multiple sets of the NF speakers 30 (however, the broadcast topology can still only be used for one set of the NF speakers 30A, 30B). For example, in some such embodiments, the broadcast audio data is the same for all sets of the NF speakers 30 within range such that all sets of the NF speakers 30 receive the same audio content. However, in other such embodiments, different audio data is broadcast to sets of the NF speakers 30 within range so that some of the NF speakers 30 can select the first audio data and some of the NF speakers 30 can select second audio data different from the first audio data. The different audio data may enable audio personalization (e.g., EQ settings), dialog language selection, improved dialog intelligibility, differences in improved out-of-head localization / spatialization, and / or other differences, as can be understood based on the present disclosure. Further, using the LE Audio broadcast topology (as well as other embodiments described variously herein) enables the volume level of the received audio content to be locally adjusted at each set of the NF speakers 30 (as opposed to a single global volume level from the FF speaker 20 within the system 10).
[0080] The above description provides embodiments that are compatible with BLUETOOTH SPECIFICATION version 5.2 [Vol 0], December 31, 2019, as well as any previous versions, such as version 4.x and 5.x devices. Additionally, the connection techniques described herein can be used for Bluetooth LE Audio, such as to assist in establishing a unicast connection. Further, it should be understood that this approach is equally applicable to other wireless protocols (e.g., non-Bluetooth, future versions of Bluetooth, etc.) where communication channels are selectively established between pairs of peers. Further, while certain embodiments have been described above as not requiring manual intervention to initiate pairing, in some embodiments, manual intervention may be required to complete pairing (e.g., "Are you sure?") presented to the user of the source / host device, for example, to provide an additional security aspect to the approach.
[0081] In some embodiments, the host-based elements of the present approach are implemented in a software module (e.g., an "app") that is downloaded and installed on a source / host (e.g., a "smartphone") to provide a cooperative audio output mode according to the approach described above.
[0082] Above, a specific order of operations performed by certain embodiments of the present invention has been described. However, it should be understood that alternative embodiments may perform the operations in a different order, combine certain operations, repeat certain operations, etc., and thus such order is illustrative. References to a given embodiment herein indicate that while the described embodiment may include certain features, structures, or characteristics, not all embodiments necessarily include the specific features, structures, or characteristics.
[0083] The functionality described in this specification, or parts thereof, and its various modifications (hereinafter referred to as "this functionality") can be implemented, at least in part, via a computer program product (e.g., a computer program tangibly embodied in an information carrier such as one or more non-transitory machine-readable media for execution by, or to control the operation of, one or more data processing devices (e.g., programmable processors, computers, multiple computers, and / or programmable logic components, etc.)).
[0084] The computer program can be written in any form of programming language, including compiled languages or interpreted languages, and it can be arranged in any form, either as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. The computer program can be arranged to be executed on one computer or on multiple computers at one location, or it can be distributed across multiple locations and interconnected by a network.
[0085] Operations associated with performing all or part of the functionality can be performed by one or more programmable processors executing one or more computer programs to perform the functions of a calibration process. All or part of the functionality can be implemented as special purpose logic circuitry, e.g., FPGA and / or ASIC (application specific integrated circuit). Examples of processors suitable for the execution of a computer program include any one or more processors of both general purpose microprocessors and special purpose microprocessors, as well as any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory, or a random access memory, or both. The components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.
[0086] In various embodiments, unless otherwise stated, electronic components described as "coupled" can be linked via conventional wired and / or wireless means such that these electronic components can communicate data with each other. Further, sub-components within a given component can be considered to be linked via conventional paths, although not necessarily shown.
[0087] A number of embodiments have been described. Nevertheless, additional modifications can be made without departing from the scope of the inventive concept described herein, and accordingly, other embodiments are understood to be within the scope of the following claims.
Description of the Reference Numerals
[0088] 5 Space 10 System 20 Far-Field (FF) Speaker 30A, 30B Non-Enclosed Near-Field (NF) Speaker 40, 40A, 40B Additional Device 50, 50a, 50b, 50c Controller 60, 60a, 60b, 60c Communication Unit 70, 70a, 70b, 70c Bluetooth (BT) Module 80, 80a, 80b, 80c Microphone 90, 90a, 90b, 90c Electroacoustic Transducer 100, 100a, 100b, 100c Additional Electronic Device 105 Space 110 First User 120 Seat 130 Second User 140 Seat 150 Wearable Audio Device 160 Headrest and / or Neck / Backrest Portion 210, 210A, 210B Source Device 220 Audio Signal 230 First Portion (of the Audio Signal) Second part (of the audio signal) 240
Claims
1. An audio system comprising: at least one long-distance field speaker configured to output a first portion of an audio signal; and a pair of non-enclosed near-field speakers configured to output a second portion of the audio signal in synchronization with the output of the first portion of the audio signal, wherein the first portion of the audio signal is different from the second portion of the audio signal.
2. The system of claim 1, wherein the first portion of the audio signal and the second portion of the audio signal comprise separate channels of common content.
3. The system of claim 1, wherein the second portion of the audio signal is binaurally encoded and increases the out-of-head localization of the content in the audio signal when the second portion of the audio signal is output in synchronization with the output of the first portion of the audio signal.
4. The system of claim 1, wherein the second portion of the audio signal enhances the spoken content within the audio signal.
5. The system of claim 4, wherein the second portion of the audio signal increases the intelligibility of the spoken content within the audio signal.
6. The system of claim 1, wherein the first portion of the audio signal does not include spoken content and the second portion of the audio signal includes spoken content.
7. The system of claim 1, wherein the first portion of the audio signal and the second portion of the audio signal comprise different frequency ranges from the audio signal.
8. i) the first portion of the audio signal has a wider frequency range than the frequency range of the second portion of the audio signal, or ii) the second portion of the audio signal excludes frequencies below a predetermined threshold, The system of claim 7, wherein at least one of the above is satisfied.
9. The system of claim 1, wherein the pair of non-enclosed near-field speakers are housed within a wearable audio device.
10. The first speaker in the pair of non-occluding near-field speakers is configured to output audio to the user's left ear, and the second speaker in the pair of non-occluding near-field speakers is configured to output audio to the user's right ear. The system according to claim 1.
11. i) The at least one far-field speaker is housed within a soundbar, or ii) The at least one far-field speaker includes a plurality of speakers configured to output at least the left channel, right channel, and center channel of the audio signal. The system according to claim 1.
12. When the at least one far-field speaker does not provide output in synchronization with the pair of non-occluding near-field speakers, the at least one far-field speaker is configured to output a third portion of the audio signal, and the third portion of the audio signal is different from the first portion of the audio signal. The system according to claim 1.
13. The second portion of the audio signal is user-configurable to improve at least one of extra-aural localization, spatialization, or dialog. The system according to claim 1.
14. The second portion of the audio signal is automatically customized based on detection of the type of device that houses at least one of the pair of non-occluding near-field speakers. The system according to claim 1.
15. The pair of non-occluding near-field speakers is configured to output the second portion of the audio signal in synchronization with the output of the first portion of the audio signal in response to a trigger. The system according to claim 1.
16. The system further includes a first device that houses the at least one far-field speaker, and the first device is configured to transmit either the audio signal or the second portion of the audio signal to a second device that includes the pair of non-occluding near-field speakers. The transmission includes playback timing data to enable the second portion of the audio signal to be output in synchronization with the first portion of the audio signal. The system according to claim 1.
17. A first device that includes the at least one far-field speaker, and A second device that includes the pair of non-occluding near-field speakers. The first device and the second device receive the audio signal from a common source device, or The first device receives the first portion of the audio signal from a first source device, and the second device receives the second portion of the audio signal from a second source device, the system according to claim 1.
18. The first device and the second device receive the audio signal from a common source device, the audio signal includes audio content considering a surround effect and a height effect, and outputting the audio signal in the first device and the second device reminds a room accommodating the first device and the second device in order to improve out-of-head localization for a user, the system according to claim 1.
19. A method of controlling an audio system, comprising: Outputting a first portion of an audio signal to at least one remote field speaker; and Outputting a second portion of the audio signal to a pair of non-occluding near-field speakers in synchronization with the output of the first portion of the audio signal, the method comprising: The first portion of the audio signal is different from the second portion of the audio signal.
20. The first portion of the audio signal and the second portion of the audio signal include separate channels of common content, and the second portion of the audio signal is user-configurable to improve at least one of out-of-head localization, spatialization, or dialogue, the method according to claim 19.
21. The second portion of the audio signal is binaurally encoded, and when the second portion of the audio signal is output in synchronization with the output of the first portion of the audio signal, it increases the out-of-head localization of the content in the audio signal, the method according to claim 19.
22. The second portion of the audio signal improves the spoken content in the audio signal, and the second portion of the audio signal increases the intelligibility of the spoken content in the audio signal, the method according to claim 19.
23. The method according to claim 19, wherein the first portion of the audio signal does not include speech content, and the second portion of the audio signal includes speech content. **Claim 24** The first portion of the audio signal and the second portion of the audio signal include frequency ranges different from the audio signal, i) the first portion of the audio signal has a frequency range wider than the frequency range of the second portion of the audio signal, or ii) the second portion of the audio signal excludes frequencies below a predetermined threshold, and the method according to claim 19 is at least one of them.
Citation Information
Patent Citations
Multi-dimensional sound playback apparatus and audio-signal playback method
JP1994165285A