Transitioning between audio playback modes for short range wireless communications
By adjusting playback buffer levels at the audio output system without reconfiguring the wireless link, the method ensures continuous audio playback during mode transitions in wireless audio communication, addressing the issue of perceptible gaps in conventional protocols.
Patent Information
- Application Number
- PCT/US2024/031850
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-01
- Filing Date
- 2024-05-31
- Publication Date
- 2025-10-09
AI Technical Summary
Conventional wireless audio communication protocols like Bluetooth LE require reconfiguration of the audio link when switching between playback modes, leading to perceptible gaps in audio output, degrading the user experience.
A method for transitioning between high robustness and low latency modes without reconfiguring the wireless communication link, by adjusting the playback buffer levels at the audio output system while maintaining the audio source device in a high robustness mode, allowing continuous audio output.
Enables seamless transitions between audio modes with minimal or no interruption in audio delivery, improving user experience by maintaining continuous audio playback.
Smart Images

Figure US2024031850_09102025_PF_FP_ABST
Abstract
Description
Transitioning Between Audio Playback Modes for Short Range Wireless CommunicationsCROSS REFERENCES TO RELATED APPLICATIONS
[0001] This Application claims priority to US Provisional Application Number 63 / 572,678, entitled “Efficient and Autonomous Transitions Between High Robustness and Low Latency Audio Playback Modes,” filed on April 1, 2024. The entire disclosure of which is hereby incorporated by reference for all purposes.BACKGROUND
[0002] When a mode of a wireless audio communication connection changes, a user may be able to perceive the mode change. Conventionally, for some wireless communication protocols, such as Bluetooth™ Low Energy (LE), changing playback modes requires that the wireless communication connection be reconfigured. This reconfiguration can result in a perceptible gap in audio output to the user. Accordingly, the user experience is degraded by the user experiencing a gap in audio output whenever the mode of the wireless audio communication connection is changed.SUMMARY
[0003] Various embodiments are described related to a method for transitioning between wireless audio output modes. In some embodiments, a method for transitioning between wireless audio output modes is described. The method may comprise configuring an audio link between an audio output system and an audio source device to be set to a high robustness mode. The high robustness mode may comprise a transmit buffer having a first buffer size at the audio source device. The method may comprise transitioning, by the audio output system, from the high robustness mode to a low latency mode. The method may comprise, in response to transitioning from the high robustness mode to the low latency mode, decreasing, by the audio output system, a playback buffer level of a receive buffer maintained by the audio output system while the transmit buffer of the audio source device may be the first buffer size. The method may comprise, after transitioning from the high robustness mode to the low latency mode and while operating in the low latency mode: receiving, by the audio output system, via a wireless device-to-device communication link, audio packets. The method may comprise, after transitioning from the high robustness mode to the low latency mode and while operating in the low latency mode: storing, by the audio output system, audio data from the audio packets to the receive buffer. The method may comprise, after transitioning from the high robustness mode to the low latency mode and while operating in the low latency mode: based on the decreased playback buffer level of the receivebuffer maintained by the audio output system, outputting, by the audio output system via a speaker, audio based on the stored audio data from the receive buffer.
[0004] Embodiments of such a method may comprise one or more of the following features: transitioning from the high robustness mode to the low latency mode may be performed without performing a codec reconfiguration or a quality of service (QoS) reconfiguration. The audio packets may be transmitted by the audio source device to the audio output system via a connected isochronous (CIS) link and transitioning from the high robustness mode to the low latency mode does not involve the CIS link being reconfigured. Audio may be output continuously by the audio output system via the speaker while transitioning from the high robustness mode to the low latency mode. The method may further comprise determining, by the audio output system, to transition from the high robustness mode to the low latency mode based on a type of audio media being output. Transitioning, by the audio output system, from the high robustness mode to the low latency mode may be performed in response to determining to transition based on the type of audio media being output. The method may further comprise determining, by the audio output system, to transition from the high robustness mode to the low latency mode based on a wireless link quality. Transitioning, by the audio output system, from the high robustness mode to the low latency mode may be performed in response to determining to transition based on the wireless link quality. The audio output system may be a pair of true wireless earbuds that wirelessly communicate with the audio source device and wirelessly communicate with each other. The method may further comprise, in response to transitioning from the high robustness mode to the low latency mode and in response to the type of audio media being indicative of corresponding video output, transmitting, by the audio output system to the audio source device, an indication of the playback buffer level. The CIS link may be part of a Bluetooth Low Energy (LE) communication link. The method may further comprise performing a handshake between the audio output system and the audio source device to establish availability of the low latency mode and the high robustness mode at the audio output system and the audio source device. Transitioning, by the audio output system, from the high robustness mode to the low latency mode occurs without causing the audio source device to transition to the low latency mode.
[0005] In some embodiments, a method for transitioning between wireless audio output modes is described. The method may comprise configuring an audio link between an audio output system and an audio source device to be set to a high robustness mode at the audio source device.The high robustness mode may comprise a transmit buffer having a first buffer size at the audio source device. The first buffer size may be greater than a second buffer size of a low latency mode. The method may comprise transitioning, by the audio output system, from the low latency mode tothe high robustness mode. The method may comprise, in response to transitioning from the low latency mode to the high robustness mode, increasing a playback buffer level of a receive buffer maintained by the audio output system while the audio source device maintains the transmit buffer as the first buffer size. The method may comprise, while operating in the high robustness mode, receiving, by the audio output system, via a wireless device-to-device communication link, audio packets. The method may comprise storing, by the audio output system, audio data from the audio packets to the receive buffer. The method may comprise, based on the playback buffer level of the receive buffer, outputting, by the audio output system via a speaker, audio based on the stored audio data from the receive buffer.
[0006] Embodiments of such a method may comprise one or more of the following features: transitioning from the low latency mode to the high robustness mode may be performed without performing a codec reconfiguration or a quality of service (QoS) reconfiguration. The audio packets may be transmitted by the audio source device to the audio output system via a connected isochronous (CIS) link and transitioning from the low latency mode to the high robustness mode does not involve the CIS link being reconfigured. The method may further comprise determining, by the audio output system, to transition from the low latency mode to the high robustness mode based on a type of audio media being output. Transitioning, by the audio output system, from the low latency mode to the high robustness mode may be performed in response to determining to transition based on the type of audio media being output. The method may further comprise determining, by the audio output system, to transition from the low latency mode to the high robustness mode based on a wireless link quality. Transitioning, by the audio output system, from the low latency mode to the high robustness mode may be performed in response to determining to transition based on the wireless link quality.
[0007] In some embodiments, a wireless earbud is described. The device may comprise a speaker. The device may comprise a wireless communication interface. The device may comprise a processing system, comprising one or more processors, in communication with the speaker and the wireless communication interface. The processing system may be configured to configure an audio link with an audio source device to be set to a high robustness mode. The high robustness mode may comprise a transmit buffer having a first buffer size at the audio source device. The processing system may be configured to transition from the high robustness mode to a low latency mode. The processing system may be configured to, in response to transitioning from the high robustness mode to the low latency mode, decrease a playback buffer level of a receive buffer maintained by the wireless earbud while the transmit buffer of the audio source device may be the first buffer size. The processing system may be configured to, after transitioning from the highrobustness mode to the low latency mode and while operating in the low latency mode: receive, via a wireless device-to-device communication link, audio packets; store audio data from the audio packets to the receive buffer; based on the decreased playback buffer level of the receive buffer, output via the speaker, audio based on the stored audio data from the receive buffer.
[0008] Embodiments of such a device may comprise one or more of the following features: audio may be output continuously by the wireless earbud via the speaker while transitioning from the high robustness mode to the low latency mode. Transitioning from the high robustness mode to the low latency mode may be performed without performing a codec reconfiguration or a quality of service (QoS) reconfiguration. The audio packets may be transmitted by the audio source device to the wireless earbud via a connected isochronous (CIS) link and transitioning from the high robustness mode to the low latency mode may not involve the CIS link being reconfigured. The processing system may be further configured to determine to transition from the high robustness mode to the low latency mode based on a type of audio media being output. Transitioning from the high robustness mode to the low latency mode may be performed in response to determining to transition based on the type of audio media being output.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] A further understanding of the nature and advantages of various embodiments may be realized by reference to the following figures. In the appended figures, similar components or features may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.
[0010] FIG. 1 illustrates an embodiment of an audio system in which audio can be output continuously irrespective of mode changes.
[0011] FIG. 2 illustrates a block diagram of an embodiment of an audio system.
[0012] FIG. 3 illustrates an embodiment of an audio system in which true wireless earbuds communicate with each other in addition to communicating with an audio source.
[0013] FIG. 4 illustrates a state diagram illustrating how audio can be output continuously irrespective of mode changes.
[0014] FIG. 5 illustrates an embodiment of mode transitions between a high robustness mode and a low latency mode.
[0015] FIG. 6 illustrates an embodiment of a method for transitioning from a high robustness mode to a low latency mode.
[0016] FIG. 7 illustrates an embodiment of a method for transitioning from a low latency mode to a high robustness mode.DETAILED DESCRIPTION
[0017] In various wireless short range communication protocols, different operating modes can be present. For example, when Bluetooth™ Low Energy (LE) is used for transmitting audio from an audio source device to an audio output device, different codecs and Quality of Service (QoS) parameters can be set. This configurability allows for a high robustness mode and alternatively a low latency mode to be used. Each of these modes have advantages and drawbacks. A high robustness mode is more tolerant of attenuation and interference because a larger buffer of audio data is maintained, thus allowing for a greater number of retransmissions of audio. However, by using a larger buffer of audio data, the latency in audio output is increased. In comparison, a low latency mode uses a smaller buffer of audio data. Thus, the low latency mode allows for a lower latency of when the audio data is created compared to when it is output, but is less tolerant of attenuation and interference.
[0018] In conventional arrangements, switching between audio output modes can require the audio link to be renegotiated between the audio output system (e.g., earbuds, headphones, wireless speaker) and the audio source device (e.g., laptop, gaming device, smartphone, television, computer). For example, in Bluetooth LE, when changing between modes, the codec and QoS parameters used need to be renegotiated and the connected isochronous (CIS) link needs to be reestablished. If audio is being output, mode change results in a noticeable interruption in the audio being output. Therefore, when conventional techniques are applied, this outcome involves a break, such as between 500 ms and 1 s in duration, in audio delivery, which is not preferable for a positive user experience. Rather, an arrangement that allows for mode changing while permitting continuous audio output is preferable.
[0019] In embodiments detailed herein, during an initial negotiation between an audio source device and an audio output system, a CIS link is established according to a high robustness mode at both the audio source device and the audio output system. Going forward, regardless of mode changes made at the audio output system, the audio source device is maintained in the high robustness mode. The audio output system, however, switches between the high robustness modeand a low latency mode as conditions dictate. For example, since latency is not a significant issue when music is being output, the audio output system may operate in the high robustness mode. While the music is still being output, a user may activate a game on the audio source device (e.g., smartphone). For gaming, latency may be a much larger concern in order to keep the audio synchronized with the gameplay. While the music is continuously output, the audio output system can transition to the low latency mode without interrupting music playback because the CIS link itself does not need to be reestablished or reconfigured. Rather, only the playback buffer level of the buffer maintained by the audio output system is decreased.
[0020] The “playback buffer level” defines the amount of audio data that is buffered within the audio output system buffer at mode initiation before the audio output system starts playing back audio stored in the buffer. Therefore, a lower playback buffer level results in less audio data being buffered upon initiation of a mode before audio is output, whereas a higher playback buffer level results in more audio data being buffered upon initiation of a mode before audio is output. Once playback begins, while the amount of audio data stored in the buffer can vary due to factors such as interference, attenuation, and clock drift, an amount of audio equal to the playback buffer level is attempted to be buffered within the buffer.
[0021] Similarly, when the audio output system is operating in the low latency mode (and the audio source device remains in the high robustness mode), the audio output system may determine that it should transition back to the high robustness mode. For example, the need for the low latency mode may have ended or the audio link quality may have degraded to a point where the low robustness mode is needed to avoid dropped packets. The audio output device can then transition back to the high robustness mode without the CIS link needing to be reestablished since no changes are made to the configuration of the CIS link itself or the audio source device.
[0022] Throughout this document, reference is made to a “high robustness mode” and a “low latency mode.” Various QoS parameters may differ between these modes. Namely, when activated, the low latency mode may use a smaller transmit buffer and playback buffer level of the receive buffer at the audio source device and the audio output system. (In addition to adjusting the playback buffer level, the size of the receive buffer can be decreased for the low latency mode.) The adjusted buffers can allow for lower latency between when audio is transmitted / received and when output. Other names can be employed for these modes, a key differentiator being the buffer size and / or playback buffer level used. Further, while these two modes are detailed herein, other modes may be additionally available at the audio source device, audio output system, or both.
[0023] Further, embodiments detailed herein may specifically refer to Bluetooth LE. Specifically, embodiments detailed herein can prevent a CIS link, used for streaming audio packets to an audio output device, from needing to be reestablished in response to a mode change. However, it should be understood that the same principles detailed herein can be applied to other forms of personal area network (PAN) communications outside of the Bluetooth family of protocols. For example, arrangements detailed herein can be applicable to other short-range device-to-device wireless communication protocols for streaming audio.
[0024] Further detail regarding these and other embodiments is provided in relation to the figures. FIG. 1 illustrates an embodiment of audio system 100 (“system 100”) in which audio can be output continuously, irrespective of mode changes. System 100 can include: true wireless earbuds 110 (110-1, 110-2) and audio source device 120. “True wireless earbuds” refer to earbuds which are not physically connected together when in use nor are they physically connected with the audio source. In this embodiment, the audio output system is illustrated as earbuds 110; however, in other embodiments other forms of audio output system can be used, such as: one or more wireless speakers, headphones, hearing aids, a speakerphone, etc. Similarly, audio source device 120 is illustrated as a smartphone; however in other embodiments, other forms of audio source device 120 can be used, such as: a gaming device, a computer (e.g., laptop, desktop, tablet, etc.), a television, or any computerized audio source. During an initial audio session negotiation between earbuds 110 and audio source device 120, multiple links may be created. These links can include: downstream CIS link 130 (that is, from audio source device 120 to earbuds 110); downstream ACL link 140; upstream CIS link 160; and upstream ACL link 150. More generally, outside the context of Bluetooth LE, CIS links can be understood as audio communication sublinks and ACL links can be understood as control sublinks being used for all other data.
[0025] In embodiments where true wireless earbuds 110 are used, a single earbud may function as the primary earbud (for example, earbud 110-1). From the perspective of audio source device 120, audio source device 120 may only communicate with the primary earbud. Audio to be output by the secondary earbud (for example, earbud 110-2) may be obtained: by receiving and decrypting communication transmitted by audio source device 120 to the primary earbud, by direct communication with the primary earbud, or by some combination thereof. For example, the secondary earbud could attempt to receive the transmission of audio from audio source device 120 to the primary earbud. If it does not successfully receive the audio, the secondary earbud may transmit a request for audio directly to the primary earbud.
[0026] ACL link 140 may be used by audio source device 120 to transmit control data to earbuds 110. While audio data (e.g., packets containing encoded audio) may be transmitted using CIS link 130, all other data may be transmitted using ACL link 140. To complement downstream CIS link 130 and downstream ACL link 140, upstream CIS link 160 and upstream ACL link 150, respectively, are established. In situations where audio data is only being transmitted downstream, such as music playback, upstream CIS link 160 can still be necessary (e.g., such as in Bluetooth LE), such as in order to transmit acknowledgements and negative acknowledgements in response to packets received on CIS link 130.
[0027] When CIS link 130 and CIS link 160 are negotiated between earbuds 110 and audio source device 120, state transitions as detailed in relation to FIG. 4 can occur in order to perform initial codec and QoS parameter configurations. During this initial session configuration, as detailed in relation to FIG. 4, both transmit buffer 121 of audio source device 120 and receive buffer 111 of earbuds 110 are negotiated to be a same size for a high robustness mode, thus having a larger buffer size than compared to a low latency mode. Transmit buffer 121 is where audio packets are temporarily stored prior to transmission to earbuds 110. The number of packets buffered in transmit buffer 121 may grow if earbuds 110 are not acknowledging transmitted packets as successfully received, thus resulting in the same audio packet being retransmitted one or more times. Receive buffer 111 is where data from successfully received packets are stored prior to being covered to audio for output via a speaker by earbuds 110. Accordingly, if receive buffer 111 becomes fully empty, audio playback at least temporarily stops due to a lack of audio data to create audio for output.
[0028] Earbuds 110 can switch to the low latency mode without altering the configuration of CIS link 130 or CIS link 160. Therefore, audio source device 120 remains in the high robustness mode and continues operating using the larger transmit buffer 121. Earbuds 110, however, in response to switching to the low latency mode, reduces at least the playback buffer level of receive buffer 111. (The actual size of receive buffer 111 can also be decreased or, alternatively, can be maintained in size.) Therefore, the amount of audio data (or, said another way, the number of packets) that can be stored in receive buffer 111 is decreased. Earbuds 110 make this determination based on the type of audio being output. For example, spatial audio and gaming may require the low latency mode. Other forms of audio, such as music playback, video playback (with sound), and voice conversations, may be more tolerant of latency and thus the high robustness mode may be sufficient. Throughout the transition to the low latency mode by earbuds 110, audio can continue to be output from receive buffer 111 via the one or more speakers of earbuds 110.
[0029] The parameters of CIS link 130 and CIS link 160 can also be maintained when earbuds 110 transition from operating in the low latency mode to the high robustness mode. Transmit buffer 121 of audio source device 120 is already operating in the high robustness mode and therefore already has a commensurate buffer size. The playback buffer level and, possibly, the buffer size of receive buffer 111 can be increased to match the size of transmit buffer 121. Thus, without reinitializing CIS link 130 or CIS link 160, earbuds 110 can transition back to the high robustness mode. Throughout this transition, audio can continuously be output by earbuds 110.
[0030] FIG. 2 illustrates an embodiment of a block diagram of an audio system 200 (“system 200”). System 200 can include earbuds 110 and audio source device 120. In other embodiments, as detailed in relation to FIG. 1 , other forms of audio output systems (and audio source devices) can be used. Referring to earbuds 110, components of earbud 110-1 can include: antenna 210; wireless communication interface 220; processing system 230; microphone 240; speaker 250; and inertial measurement unit (IMU) 260. Receive buffer 111 can be incorporated as part of wireless communication interface 220. Some of these components are optional; for example, earbud 110-1 does not necessarily have IMU 260 or microphone 240. Earbud 110-2 may have the same components or a subset. For example, some pairs of earbuds may include only one earbud that has microphone 240, IMU 260, or both. Antenna 210 can be used for receiving and transmitting device-to-device short-range communications, such as Bluetooth-family communications, including basic rate / extended data rate (BR / EDR), and LE (including LE Audio which uses LE). Wireless communication interface 220 can be implemented as a system on a chip (SOC). Wireless communication interface 220 can include a Bluetooth radio and componentry necessary to convert raw incoming data (e.g., audio data, other data) to Bluetooth packets for transmission via antenna 210. A single radio may be present on each of earbuds 110, thus requiring transmissions, even on different frequencies, to occur at different times. Wireless communication interface 220 may also include componentry to enable one or more alternative or additional forms of wireless communication, both with an audio source and between earbuds.
[0031] Processing system 230 may include one or more special-purpose or general -purpose processors. Such special-purpose processors may include processors that are specifically designed to perform the functions of the components detailed herein. Such special-purpose processors may be ASICs or FPGAs which are general-purpose components that are physically and electrically configured to perform the functions detailed herein. Such general-purpose processors may execute special-purpose software that is stored locally using one or more non-transitory processor-readable mediums, such as random-access memory (RAM), and / or flash memory. In some embodiments,processing system 230 and wireless communication interface 220 may be part of a same circuit or soc.
[0032] In some earbuds, microphone 240 may be present. In some embodiments, each of earbuds 110 has a microphone. In other embodiments, only one of earbuds 110 has a microphone. In still other embodiments, no microphone may be present in either of earbuds 110. Audio captured using the one or more microphones of earbuds 110 can be transmitted to audio source device 120. This audio, which can be referred to as “upstream” audio, may include voice, such as for use in a telephone call, video conference, gaming, etc. Various componentry (not illustrated) may be present between wireless communication interface 220, processing system 230, and microphone 240, such as an analog to digital converter (ADC) and an amplifier.
[0033] Speaker 250 converts received analog signals to audio. Various componentry (not illustrated) may be present between wireless communication interface 220, processing system 230, and speaker 250, such as a digital to analog converter (DAC) and an amplifier.
[0034] Optional IMU 260 can be in the form of an accelerometer, gyroscope, or some other form of sensor that can detect movement or acceleration. IMU 260 may measure a direction of gravity, which can be used to determine the IMU’s orientation with respect to the direction of gravity. Side-to-side movement can be detected based on acceleration. Data from IMU 260 can be used for dynamic spatial audio. When dynamic spatial audio is actively being used, a low latency mode may be used to ensure that the audio being output matches the user’s head movement with decreased latency compared to other modes.
[0035] Various components of earbud 110-1 are not illustrated. In addition to the ADC, DAC, and amplifiers previously mentioned, earbud 110-1 also includes a power storage component, such as one or more batteries, and associated componentry to allow for recharging of the power storage component. Also present is a housing and componentry to hold earbud 110-1 within a user’s ear. One or more non-transitory processor readable mediums can be understood as present and accessible by wireless communication interface 125, processing system 230, or both. For instance, such mediums may be used for temporary storage of data (e.g., buffers) and storing data necessary for Bluetooth communication (e.g., encryption keys).
[0036] Audio source device 120 can include: antenna 262; wireless communication interface 125; processing system 280; and data storage 290. Antenna 262 can be used for receiving and transmitting Bluetooth-family communications, including BR / EDR, and LE. Wireless communication interface 125 can be implemented as a system on a chip (SOC). Wireless communication interface 125 can include a Bluetooth radio and componentry necessary to convertraw incoming data (e.g., audio data, other data) to Bluetooth packets for transmission via antenna 262. Wireless communication interface 125 can additionally or alternatively be used for one or more other forms of wireless communications. Processing system 280 may include one or more special-purpose or general -purpose processors. Such special-purpose processors may include processors that are specifically designed to perform the functions of the components detailed herein. Such special-purpose processors may be ASICs or FPGAs, which are general-purpose components that are physically and electrically configured to perform the functions detailed herein. Such general-purpose processors may execute special-purpose software that is stored locally using one or more non-transitory processor-readable mediums via data storage 290, which can include random access memory (RAM), flash memory, a hard disk drive (HDD) and / or a solid-state drive (SSD). In some embodiments, processing system 280 and wireless communication interface 125 may be part of a same circuit or SOC.
[0037] Audio source device 120 can include various other components. For example, if audio source device 120 is a smartphone, various components such as: one or more cameras, a display screen or touch screen, volume control buttons, or other wireless communication interfaces can be present. Examples of audio source device 120 include: a smartphone; a media player; a gaming device; a computer system (e.g., laptop, desktop, server); a smartwatch; any other computerized device or system which uses short-range device-to-device wireless communication to output audio or output dynamic spatial audio.
[0038] In embodiments where true wireless earbuds 110 are used, one earbud, such as earbud 110-1, may function as the primary earbud. From the perspective of audio source device 120, audio source device 120 is only communicating with the primary earbud and, thus, only transmissions 124 are present. Audio to be output by the secondary earbud (for example, earbud 110-2) may be obtained: by receiving and decrypting communication transmitted by audio source device 120 to the primary earbud, by direct communication with the primary earbud via transmissions 123, or by some combination thereof. In other embodiments, transmissions 122 are performed directly between audio source device 120 and earbud 110-2.
[0039] FIG. 3 illustrates an embodiment of an audio system 300 in which true wireless earbuds communicate with each other in addition to communicating with audio source device 120. Earbud 110-1 can perform wireless communications using cross-link 310 with earbud 110-2 and, similarly, earbud 110-2 can perform wireless communications using cross-link 311 with earbud 110-1. This communication may occur via a proprietary link specific to earbuds 110 and therefore can be outside of any Bluetooth family protocol specification; alternatively, a Bluetooth-familycommunication protocol can be used. The path between earbuds 110, when in use by user 301, is predictable because the distance and the object through which the signals pass (the head of user 301) remain constant. This path can be expected to produce insufficient attenuation to negatively impact communication between earbuds. The path, however, from audio source device 120 to the earbuds is harder to predict since the position of audio source device 120 relative to earbuds 110 can vary substantially and can result in significantly different attenuation at one earbud compared to the other, such as due to cross-body attenuation. Cross-body attenuation is depicted in FIG. 3 by having audio source device 120 closer to earbud 110-1 than earbud 110-2.
[0040] Cross-links 310 and 311 can use Bluetooth LE 2M, LE HDT (pending standardization), LE proprietary high data rate modes, classic BR / EDR, or some other standard-based or proprietary communication scheme. Therefore, while Bluetooth-compliant wireless communications occur between earbuds 110 and audio source device 120, communications directly between earbuds do not necessarily need to be compliant with Bluetooth or any other particular communication protocol.
[0041] FIG. 4 illustrates a state diagram 400 illustrating how audio can be output continuously irrespective of mode changes. State diagram 400 is based on the Bluetooth LE communication protocol states. If another protocol is used for wireless communication, variations may be present in the state transitions. When a session between two Bluetooth LE devices is initially established, such as in response to audio being about to be output from an audio source device to an audio output system (e.g., earbuds), as part of the session configuration, a CIS link may be established between the audio source device and the audio output device. Referring to FIG. 1, such a CIS link can include downlink CIS link 130 and uplink CIS link 160.
[0042] As part of initial establishment 401 the CIS link, a codec to use for audio encoding can be negotiated at state 410. After the codec is negotiated, QoS parameters can be configured at state 420. Such QoS parameters can include, for example, the size of the transmit and receive buffers used. Additional parameters that are defined for a Bluetooth LE CIS link include: ISO interval; sub-interval; subevent length; maximum protocol data unit (PDU) size; maximum service data unit (SDU) size; maximum number of subevents in a CIS event; and flush timeout control values. At state 430, the CIS link can be enabled. In Bluetooth LE, once these parameters have been established, they are to remain constant for the duration of the CIS link. After state 430, streaming of audio can occur at state 440.
[0043] In a conventional arrangement, if mode changes need to be made, such to a low latency mode from a high robustness mode or the reverse, state change path 402 must be followed. That is,a releasing state 450 is entered; then the CIS link is reestablished at state 410, with the same states of initial establishment 401 being repeated. During this reestablishment, a break in audio output occurs, which can potentially be one second in duration.
[0044] In contrast, embodiments detailed herein follow path 403 in which once streaming state 440 is entered, the audio output system can adjust the mode locally without needing to enter any of states 410, 420, 430, or 450. From the perspective of the audio source device, the QoS parameters configured at state 420 do not change. The playback buffer level and, possibly, the buffer size maintained by the audio output system can change. Changing the mode using path 403 can result in no interruption in audio output or only a minimal interruption, such as 100 ms or less, 75 ms or less, or 50 ms or less.
[0045] FIG. 5 illustrates an embodiment 500 of mode transitions between a high robustness mode and a low latency mode. Two transitions are illustrated: 1) from high robustness mode 501 to low latency mode 502; and from low latency mode 502 to high robustness mode 501. In some embodiments, these mode changes are performed on an establishing CIS Link of a Bluetooth LE wireless communication link between audio source device 510 and an audio output system 520 (e.g., earbuds). Audio source device 510 can refer to audio source device 120 of FIGS. 1-3 and audio output system 520 can refer to earbuds 110 or some other form of audio output system or device. When a session between an audio source device and an audio output system is first negotiated, a high robustness mode may always be used as the initial mode for the CIS link at state 420 at both the audio source device and the audio output system. Use of the high robustness mode results in a higher playback buffer level, meaning more audio needing to be buffered (compared to the low latency mode) being instantiated at the audio source device and the audio output system. As shown in high robustness mode 501, audio data transmit buffer 511 is set to have a playback buffer level of 260 ms; similarly, audio data receive buffer 521 is initially set to a playback buffer level of 260 ms. A same buffer size may be used for high robustness mode; for example, between 260 ms and 500 ms may be used for the buffer size.
[0046] The audio output system may then determine that it is to transition to low latency mode. This determination can be based on the audio output system detecting a type of audio that requires or benefits from a low latency mode, such as dynamic spatial audio or gaming. The type of audio can be determined based on a context type value transmitted to the audio output system for an audio stream. In addition to context type, other factors can be included in the type of audio determination, such as whether head-tracking is activated. For example, a combination of a context type indicative of stereo audio output combined with head-tracking being activated can beindicative of dynamic spatial audio. The determination can also be based on audio already being output that would benefit from low latency mode and the link quality having improved such that use of the low latency mode is less likely to result in interrupted playback. Audio benefiting from low latency mode can be in addition to a current use or can be the only audio being output. For example, music may already be being output; gaming or dynamic audio can be output with the music. This determination is made at the audio output system. As shown in low latency mode 502, the playback buffer level of audio data receive buffer 521 is reduced, such as to 80 ms. Therefore, once 80 ms worth of audio is stored within audio data receive buffer 521, audio begins being output from the buffer via one or more speakers of audio output system 520. If more than 80 ms worth of audio was already buffered in audio data receive buffer 521 when audio data receive buffer 521 was reduced in size, various techniques can be used for accelerating playback to bring the playback buffer level of audio data to 80 ms or less (in this example). While operating in the low latency mode, audio output system 520 maintains a buffer of audio data smaller than when operating in the high robustness mode. Therefore, on average, less time elapses between when audio data is received and when output as audible sound to a user.
[0047] After operating in low latency mode 502 for a time, audio output system 520 determines to transition back to the high robustness mode. There can be multiple reasons for this, such as the low latency use case may have ended. In such a situation, audio that may still be being output that is not significantly affected by latency, such as music playback, can continue through the transition of modes. As another example, the transition to high robustness mode may be performed due to link quality. Based on dropped packets, signal strength, or audio data receive buffer 521 becoming empty, audio output system 520 determines to transition back to the high robustness mode. Low latency mode 502 is transitioned to high robustness mode 503 by increasing the playback buffer level of audio data receive buffer 521. Audio data receive buffer 521 may then buffer 260 ms (in this example) of audio before resuming audio output.
[0048] After a time, if the link quality improves and there remains a need to operate in low latency mode (e.g., based on the type of audio being output), the audio output system transitions back to the low latency mode as described in relation to the transition from high robustness mode 501 to low latency mode 502.
[0049] Throughout the transitions between high robustness mode and low latency mode, audio output system 520 transmits data (e.g., via an ACL link) indicative of the playback buffer level of audio data receive buffer 521. The playback latency can be used for adjusting output video to ensure proper synchronization between the video and audio. This data may only be transmitted ifthe use case involves video playback. For example, music and podcasts have no audio / video synchronization issues. For gaming, this information is optional, as gaming applications may not prioritize synchronization over decreasing latency as much as possible.
[0050] Whenever a change in mode is made by the audio output system, if the audio output system is a pair of true wireless earbuds (or wireless speakers), communication directly between the earbuds may be necessary in order to ensure synchronization. The earbuds, via earbud-to- earbud communication link as detailed in relation to FIGS. 2 and 3, can determine an updated flush timeout (FT), which can be used to compute a SDU synchronization reference point to ensure that audio output via each earbud’s speaker is synchronized. In Bluetooth LE, Equation 1 is used to calculate the SDU synchronization reference point:
[0051] In Equation 1, SDUSR refers to the SDU Synchronization Reference, CISpefAnchorPt refers to the CIS reference anchor point, and FT refers to the flush timeout.
[0052] Various methods may be performed using the systems and arrangements detailed in FIGS. 1-5. FIG. 6 illustrates an embodiment of a method 600 for transitioning from a high robustness mode to a low latency mode. Method 600 can be performed using the arrangements detailed in relation to FIGS. 1-5.
[0053] At block 610, a session configuration is performed between an audio source device and an audio output system. This session configuration, such as in Bluetooth LE, may be used to create a CIS link for streaming audio from the audio source device to the audio output system. The session configuration may be performed in response to the initialization of audio output when no audio is currently being output. For example, a user may initiate music playback or dynamic audio at a smartphone that is wirelessly connected with a pair of earbuds.
[0054] The session configuration can proceed as detailed in relation to FIG. 4, such as for Bluetooth LE. As part of block 610, the session configuration results in the CIS link being configured for a high robustness mode. This results in the buffer size at the audio source device and the playback buffer level of the buffer of the audio output system being set to a larger value compared to a low latency mode.
[0055] At block 620, the audio output system, which may be a pair of earbuds, may determine to transition to a low latency mode. The determination to make this transition can occur while audio is being transmitted via the CIS link from the audio source device and output by the audio outputsystem. For example, music may be currently being received via the CIS link and output via the audio output system.
[0056] The determination to transition at block 620 may be based on various factors. A first factor is the type of audio being output. Dynamic spatial audio and gaming may require or benefit from the use of the low latency mode. As such, if either of these types of audio are detected by the audio output system, the determination of block 620 may be made. A second factor may be the quality of the communication link between the audio source device and the audio output system. The link quality may need to be high enough to be eligible for using the low latency mode. If the link quality was previously not high enough, even though the type of audio being output would benefit from the low latency mode, the high robustness mode may be used instead at least until the link quality improves sufficiently. Block 620 may be performed in response to the link quality having improved sufficiently (e.g., based on measured signal strength, number of dropped packets, etc.) to allow for the low latency mode to be used.
[0057] At block 630, the transition may be made by the audio output system to the low latency mode based on the determination of block 620. The transition of block 630 may be made exclusively by the audio output system without any changes being made to the CIS link itself or to the audio source device.
[0058] At block 640, in response to the transition block 630, the playback buffer level within the audio data receive buffer may be decreased from the high robustness mode. For example, the playback buffer level may be reduced from 260 ms to 80 ms. If the buffer has sufficient stored audio that the playback time of the audio exceeds the playback buffer level, various techniques may be employed to accelerate playback until the lowered payback level is reached.
[0059] In some embodiments, depending on the type of audio being output, data may be sent, such as via an ACL link to the audio source device indicating the playback buffer level of the audio data receive buffer. This information may be used to ensure that audio generated by the audio source device is synchronized with the corresponding video being output.
[0060] At block 650, audio packets may be received by the audio output system from the audio source device while the audio source device has continued to operate in the high robustness mode as configured at block 610 during the initial session configuration. The audio output system, however, is operating in the low latency mode by virtue of its reduced playback buffer level of the audio data receive buffer. At block 660, audio data from the received audio packets are buffered in the audio data receive buffer of the audio output system. At block 670, audio is output by the audiooutput system based on the buffered audio data when the playback buffer level of the buffer is reached.
[0061] FIG. 7 illustrates an embodiment of a method 700 for transitioning from a low latency mode to a high robustness mode. For method 700, method 600 has already been performed. Thus a session configuration has already been performed in which the audio source device and the audio output system were configured to a high robustness mode, with the audio output system then transitioning to operating in a low latency mode.
[0062] At block 720, the audio output system, which may be a pair of earbuds, may determine to transition back to the high robustness mode from the low latency mode. The determination to make this transition can occur while audio is being transmitted via the CIS link from the audio source device and output by the audio output system. For example, gaming audio may have ceased being streamed but music continues to be streamed via the CIS link and output via the audio output system.
[0063] The determination to transition at block 720 may be based on various factors. A first factor is the type of audio being output. Dynamic spatial audio and gaming may require or benefit from the use of the low latency mode. As such, if either of these types of audio are no longer detected or determined to be output by the audio output system, the determination of block 720 may be made. A second factor may be the quality of the communication link between the audio source device and the audio output system. The link quality may need to be high enough to be eligible for using the low latency mode. If the link quality drops sufficiently (e.g., based on measured signal strength, number of dropped packets, etc.), the determination may be made to use the high robustness mode regardless of the type of audio being output.
[0064] At block 730, the transition may be made by the audio output system to the high robustness mode from the low latency mode based on the determination of block 720. The transition of block 730 may be made exclusively by the audio output system without any changes being made to the CIS link itself or to the audio source device; the audio source device is already functioning in the high robustness mode.
[0065] At block 740, in response to the transition block 730, the playback buffer level within the audio data receive buffer may be increased from the level used for the low latency mode. For example, the playback buffer level may be increased to 260 ms from 80 ms. A short break in audio playback may occur as the buffer is filled until the playback buffer level of the high robustness mode is reached.
[0066] In some embodiments, depending on the type of audio being output, data may be sent, such as via an ACL link to the audio source device indicating the updated increased playback buffer level of the audio data receive buffer. This information may be used to ensure that audio generated by the audio source device is synchronized with the corresponding video being output.
[0067] At block 750, audio packets may be received by the audio output system from the audio source device while the audio source device and the audio output system operate in the high robustness mode. At block 760, audio data from the received audio packets are buffered in the audio data receive buffer of the audio output system. At block 770, audio is output by the audio output system based on the buffered audio data when the playback buffer level of the buffer is reached.
[0068] It should be noted that the methods, systems, and devices discussed above are intended merely to be examples. It must be stressed that various embodiments may omit, substitute, or add various procedures or components as appropriate. For instance, it should be appreciated that, in alternative embodiments, the methods may be performed in an order different from that described, and that various steps may be added, omitted, or combined. Also, features described with respect to certain embodiments may be combined in various other embodiments. Different aspects and elements of the embodiments may be combined in a similar manner. Also, it should be emphasized that technology evolves and, thus, many of the elements are examples and should not be interpreted to limit the scope of the invention.
[0069] Specific details are given in the description to provide a thorough understanding of the embodiments. However, it will be understood by one of ordinary skill in the art that the embodiments may be practiced without these specific details. For example, well-known, processes, structures, and techniques have been shown without unnecessary detail in order to avoid obscuring the embodiments. This description provides example embodiments only, and is not intended to limit the scope, applicability, or configuration of the invention. Rather, the preceding description of the embodiments will provide those skilled in the art with an enabling description for implementing embodiments of the invention. Various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the invention.
[0070] Also, it is noted that the embodiments may be described as a process which is depicted as a flow diagram or block diagram. Although each may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. A process may have additional steps not included in the figure.
[0071] Having described several embodiments, it will be recognized by those of skill in the art that various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the invention. For example, the above elements may merely be a component of a larger system, wherein other rules may take precedence over or otherwise modify the application of the invention. Also, a number of steps may be undertaken before, during, or after the above elements are considered. Accordingly, the above description should not be taken as limiting the scope of the invention.
Claims
WHAT IS CLAIMED IS:
1. A method for transitioning between wireless audio output modes, the method comprising: configuring an audio link between an audio output system and an audio source device to be set to a high robustness mode, wherein: the high robustness mode comprises a transmit buffer having a first buffer size at the audio source device, and transitioning, by the audio output system, from the high robustness mode to a low latency mode; in response to transitioning from the high robustness mode to the low latency mode, decreasing, by the audio output system, a playback buffer level of a receive buffer maintained by the audio output system while the transmit buffer of the audio source device is maintained as the first buffer size; after transitioning from the high robustness mode to the low latency mode and while operating in the low latency mode: receiving, by the audio output system, via a wireless device-to-device communication link, audio packets; storing, by the audio output system, audio data from the audio packets to the receive buffer; and based on the decreased playback buffer level of the receive buffer maintained by the audio output system, outputting, by the audio output system via a speaker, audio based on the stored audio data from the receive buffer.
2. The method for transitioning between wireless audio output modes of claim 1, wherein transitioning from the high robustness mode to the low latency mode is performed without performing a codec reconfiguration or a quality of service (QoS) reconfiguration.
3. The method for transitioning between wireless audio output modes of claim 2, wherein the audio packets are transmitted by the audio source device to the audio outputsystem via a connected isochronous (CIS) link and transitioning from the high robustness mode to the low latency mode does not involve the CIS link being reconfigured.
4. The method for transitioning between wireless audio output modes of claim 1, wherein audio is output continuously by the audio output system via the speaker while transitioning from the high robustness mode to the low latency mode.
5. The method for transitioning between wireless audio output modes of claim 1, further comprising: determining, by the audio output system, to transition from the high robustness mode to the low latency mode based on a type of audio media being output, wherein transitioning, by the audio output system, from the high robustness mode to the low latency mode is performed in response to determining to transition based on the type of audio media being output.
6. The method for transitioning between wireless audio output modes of claim 1, further comprising: determining, by the audio output system, to transition from the high robustness mode to the low latency mode based on a wireless link quality, wherein transitioning, by the audio output system, from the high robustness mode to the low latency mode is performed in response to determining to transition based on the wireless link quality.
7. The method for transitioning between the wireless audio output modes of claim 1, wherein the audio output system is a pair of true wireless earbuds that wirelessly communicate with the audio source device and wirelessly communicate with each other.
8. The method for transitioning between the wireless audio output modes of claim 1, further comprising: in response to transitioning from the high robustness mode to the low latency mode and in response to a type of audio media being indicative of corresponding video output, transmitting, by the audio output system to the audio source device, an indication of the playback buffer level.
9. The method for transitioning between wireless audio output modes of claim 3, wherein the CIS link is part of a Bluetooth Low Energy (LE) communication link.
10. The method for transitioning between wireless audio output modes of claim 1, further comprising: performing a handshake between the audio output system and the audio source device to establish availability of the low latency mode and the high robustness mode at the audio output system and the audio source device, wherein transitioning, by the audio output system, from the high robustness mode to the low latency mode occurs without the audio source device transitioning to the low latency mode.
11. A method for transitioning between wireless audio output modes, the method comprising: configuring an audio link between an audio output system and an audio source device to be set to a high robustness mode at the audio source device, wherein: the high robustness mode comprises a transmit buffer having a first buffer size at the audio source device, and transitioning, by the audio output system, from a low latency mode to the high robustness mode; in response to transitioning from the low latency mode to the high robustness mode, increasing a playback buffer level of a receive buffer maintained by the audio output system while the audio source device maintains the transmit buffer as the first buffer size; while operating in the high robustness mode, receiving, by the audio output system, via a wireless device-to-device communication link, audio packets; storing, by the audio output system, audio data from the audio packets to the receive buffer; and based on the playback buffer level of the receive buffer, outputting, by the audio output system via a speaker, audio based on the stored audio data from the receive buffer.
12. The method for transitioning between wireless audio output modes of claim 11, wherein transitioning from the low latency mode to the high robustness mode is performed without performing a codec reconfiguration or a quality of service (QoS) reconfiguration.
13. The method for transitioning between wireless audio output modes of claim 12, wherein the audio packets are transmitted by the audio source device to the audio output system via a connected isochronous (CIS) link and transitioning from the low latency mode to the high robustness mode does not involve the CIS link being reconfigured.
14. The method for transitioning between wireless audio output modes of claim 11, wherein transitioning, by the audio output system, from the low latency mode to the high robustness mode is performed in response to determining to transition based on a type of audio media being output.
15. The method for transitioning between wireless audio output modes of claim 11, further comprising: determining, by the audio output system, to transition from the low latency mode to the high robustness mode based on a wireless link quality, wherein transitioning, by the audio output system, from the low latency mode to the high robustness mode is performed in response to determining to transition based on the wireless link quality.
16. A wireless earbud, comprising: a speaker; a wireless communication interface; a buffer; and a processing system, comprising one or more processors, in communication with the speaker, the buffer, and the wireless communication interface, wherein the processing system is configured to: configure an audio link with an audio source device to be set to a high robustness mode, wherein:the high robustness mode sets the buffer to a first buffer size, and transition from the high robustness mode to a low latency mode; in response to transitioning from the high robustness mode to the low latency mode, decrease a playback buffer level of the buffer; after transitioning from the high robustness mode to the low latency mode and while operating in the low latency mode: receive, from a wireless device-to-device communication link, audio packets; store audio data from the audio packets to the buffer; and based on the decreased playback buffer level of the buffer, output via the speaker, audio based on the stored audio data from the buffer.
17. The wireless earbud of claim 16, wherein audio is output continuously by the wireless earbud via the speaker while transitioning from the high robustness mode to the low latency mode.
18. The wireless earbud of claim 17, wherein transitioning from the high robustness mode to the low latency mode is performed without performing a codec reconfiguration or a quality of service (QoS) reconfiguration.
19. The wireless earbud of claim 18, wherein the audio packets are transmitted by the audio source device to the wireless earbud via a connected isochronous (CIS) link and transitioning from the high robustness mode to the low latency mode does not involve the CIS link being reconfigured.
20. The wireless earbud of claim 16, wherein the processing system is further configured to: determine to transition from the high robustness mode to the low latency mode based on a type of audio media being output, whereintransitioning from the high robustness mode to the low latency mode is performed in response to determining to transition based on the type of audio media being output.
Citation Information
Patent Citations
Low latency mode for wireless communication between devices
US10705793B1
System having device-mount audio mode
US20190349662A1
Method and electronic system for outputting video data and audio data
US20230051068A1