Method and system for switching wireless audio connection during call

By introducing bidirectional and unidirectional wireless audio connection switching technology into smart devices, the problem of low audio signal processing efficiency during calls is solved, achieving efficient transmission and synchronous processing of audio signals and supporting audio data exchange between multiple devices.

CN122018847APending Publication Date: 2026-05-12APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
APPLE INC
Filing Date
2022-05-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, smart devices have difficulty efficiently switching wireless audio connections during calls, resulting in low efficiency in audio signal processing and making it impossible to achieve synchronized playback and audio data exchange and processing among multiple devices.

Method used

By introducing bidirectional and unidirectional wireless audio connection switching technology into smart devices, audio signal switching and audio signal processing are achieved. By establishing and switching audio connections between devices, and by using Bluetooth and wireless technologies, audio signal switching and audio signal transmission are realized.

Benefits of technology

It improves the efficiency and synchronization of audio signal transmission between devices, enhances the audio signal processing capability, and supports audio data exchange and processing between multiple devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018847A_ABST
    Figure CN122018847A_ABST
Patent Text Reader

Abstract

The invention relates to a method and system for switching a wireless audio connection during a call. A method performed by a first electronic device, the first electronic device communicatively coupled to a wireless headset, the method comprising: communicating with the wireless headset via a bidirectional wireless audio connection while participating in a call with a second electronic device; determining that a joint media playback session has been initiated in which the first electronic device and the second electronic device will stream media content independently for individual playback when participating in the call; and switching to communicate with the wireless headset via a unidirectional wireless audio connection based on the determination of one or more capabilities of the wireless headset, wherein a mixed content of 1) one or more signals associated with the call and 2) an audio signal of the media content is transmitted to the wireless headphone over the one-way wireless audio connection.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application No. 202210521389.9, filed on May 13, 2022, entitled "Method and System for Switching Wireless Audio Connections During a Call". Technical Field

[0002] One aspect of this disclosure relates to an audio system that switches between wireless audio connections during a call based on certain standards. Other aspects are also described. Background Technology

[0003] Many devices today, such as smartphones, are capable of various types of telecommunications with other devices. For example, a smartphone can make a phone call to another device. In this case, when dialing a phone number, the smartphone connects to a cellular network, which then connects the smartphone to another device (e.g., another smartphone or a landline). Additionally, smartphones are capable of video conferencing, where video and audio data are exchanged with another device. Summary of the Invention

[0004] One aspect of this disclosure is a method performed by a first electronic device (e.g., a local device) communicatively coupled to an audio output device (e.g., a wireless headset). When participating in a call with a second electronic device (e.g., a remote device), the local device communicates with the wireless headset via a two-way wireless audio connection (e.g., where audio data can be exchanged between the two devices). The local device determines that a joint media playback session has been initiated, in which the local device and the remote device will independently stream media content for individual playback by both devices while participating in the call. Based on the determination of one or more capabilities of the wireless headset, the local device switches to communicating with the wireless headset via a one-way wireless audio connection (e.g., where audio data can only be transmitted from the local device to the wireless headset), wherein a mixture of 1) one or more signals associated with the call and 2) audio signals of the media content is transmitted from the local device to the wireless headset via the one-way wireless audio connection.

[0005] In one aspect, determining one or more capabilities of the wireless headphones includes determining whether the wireless headphones support exchanging audio signals with a local device via a two-way wireless audio connection at a sampling rate of at least a threshold sampling rate (e.g., 24 kHz). In some aspects, the local device transmits a signal to mute the microphone of the wireless headphones and activate the microphone of the local device to capture user speech. In one aspect, the local device displays a pop-up notification on a display screen indicating that the microphone of the wireless headphones is muted and requesting user authentication to activate the microphone of the local device, wherein the microphone of the local device is activated in response to receiving user input at the local device.

[0006] In one aspect, the local device may receive sensor data from at least one sensor indicating whether an object is within a threshold distance of the local device, and in response to the object being within the threshold distance, switch to communicating with the wireless headphones via a two-way wireless audio connection. In another aspect, the local device determines the signal strength of a one-way wireless audio connection, and in response to the determined signal strength being below a threshold, switch to communicating with the wireless headphones via a two-way audio connection.

[0007] In some aspects, the local device may receive a control signal from the wireless headset indicating that user voice has been detected, and in response to the control signal, switch to communicating with the wireless headset via a two-way wireless audio connection. In one aspect, the control signal is a first control signal, and in response to receiving a second control signal indicating that user voice is no longer detected, switches back to communicating with the wireless headset via a one-way wireless audio connection.

[0008] In one aspect, after switching to communication with the wireless headset via a unidirectional wireless audio connection, the local device applies a scalar gain to the mixed content for at least a period of time. In another aspect, when communicating with the wireless headset via a bidirectional wireless audio connection, the local device applies noise reduction to one or more signals associated with the call. In yet another aspect, the local device may cease applying noise reduction to one or more signals associated with the call in response to switching to communication with the wireless headset via a unidirectional wireless audio connection. In some aspects, when the local device communicates with the wireless headset via a unidirectional wireless audio connection, the local device obtains microphone signals from its microphone that will be transmitted to the remote device when both the local device and the remote device participate in the call, generates an estimate of a portion of one or more signals associated with the call, and performs echo cancellation on the microphone signals using this estimate.

[0009] Another aspect of this disclosure is a method performed by a wireless headset, the method comprising communicating with a local device via a two-way wireless audio connection during a call between a local device and a remote device. The headset determines that a unidirectional wireless audio connection will be established between the local device and the wireless headset during the call, instead of a two-way wireless audio connection. In response to determining that a unidirectional wireless audio connection will be established, the microphone of the wireless headset is muted and switched from the two-way wireless audio connection to the unidirectional wireless audio connection. The wireless headset receives an audio signal via the unidirectional wireless audio connection, the audio signal comprising a mixture of signals associated with the call and signals associated with media content being played back by the local device and the remote device in a joint media playback session. The wireless headset uses the audio signals to drive a speaker.

[0010] In one aspect, the sampling rate of audio signal exchange supported by a two-way wireless audio connection is lower than the sampling rate of audio signal transmission supported by a one-way wireless audio connection. In some aspects, determining to establish a one-way wireless audio connection includes receiving a control signal from a local device to establish the one-way wireless audio connection. In one aspect, the wireless headset uses an accelerometer to detect user voice and switches from a one-way wireless audio connection to a two-way wireless audio connection in response to the detection of user voice. In some aspects, in response to the detection of user voice, the microphone of the wireless headset is activated and the microphone signal generated by the microphone is transmitted to the local device via the two-way wireless audio connection for use in making a call. In another aspect, in response to stopping the detection of user voice, the wireless headset mutes the microphone and switches from a two-way wireless audio connection to a one-way wireless audio connection.

[0011] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed embodiments below and specifically pointed out in the claims. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description

[0012] Multiple aspects are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals in the drawings indicate similar elements. It should be noted that references to "a" or "an" aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a single drawing may be used to illustrate features of more than one aspect, and for a particular aspect, not all elements in that drawing may be necessary.

[0013] Figure 1An audio system according to one aspect is shown, which includes a local device and one or more remote devices participating in a call during a joint media playback session.

[0014] Figure 2 A block diagram is shown of a local device and an audio output device according to one aspect, wherein the local device initiates a joint playback media session when participating in a call with one or more remote devices, and the audio output device communicates wirelessly with the local device.

[0015] Figure 3 It illustrates several stages according to one aspect, in which a local device and a remote device initiate a joint playback media session to synchronously play back musical works while participating in a telephone call.

[0016] Figure 4 It illustrates several stages according to one aspect, in which local and remote devices initiate a joint playback media session to synchronously play back the video while participating in a video call.

[0017] Figure 5 A block diagram of a local device is shown, which performs audio signal processing operations on the audio signal of media content based on whether speech is detected in the signal of a telephone call performed between the local device and a remote device.

[0018] Figure 6 A block diagram of a local device is shown, which performs audio signal processing operations on the audio signal of media content based on whether an audio output device detects speech.

[0019] Figure 7 A block diagram of a local device is shown, which performs audio signal processing operations based on whether speech is detected within the signal of a video call.

[0020] Figure 8 This is a flowchart of one aspect of the process for processing audio signals of media content based on whether speech is detected within the downlink audio signal.

[0021] Figure 9 This is a flowchart of one aspect of the process of displaying closed captions that represent audio content in media.

[0022] Figure 10 This is a flowchart of one aspect of the process for reversing the playback of media content when it is determined that the downlink audio signal has stopped including speech.

[0023] Figure 11A block diagram is shown according to one aspect, wherein a local device 2 is communicatively coupled to an audio output device 6 via a two-way wireless audio connection to exchange audio data when the local device and a remote device 3 participate in a call together.

[0024] Figure 12 A block diagram is shown according to one aspect, wherein during a joint media playback session and a call with a remote device 3, the local device 2 is communicatively coupled to the audio output device 6 via a two-way wireless audio connection.

[0025] Figure 13a and Figure 13b Several block diagrams are shown according to one aspect, wherein a local device 2, which is communicatively coupled to an audio output device 6 for exchanging audio data, switches between wireless audio connections based on the initiation of a joint media playback session.

[0026] Figure 14 This is a flowchart of one aspect of the process used to switch between wireless audio connections.

[0027] Figure 15 This is a flowchart of another aspect of the process used to switch between wireless audio connections.

[0028] Figure 16 It is a flowchart of one aspect of the process used to determine whether to switch between wireless audio connections based on one or more criteria.

[0029] Figure 17 This is a flowchart of one aspect of the process performed by the audio output device for switching between wireless audio connections.

[0030] Figure 18 This is a flowchart of one aspect of the process performed by the audio output device to switch from a one-way wireless audio connection to a two-way wireless audio connection based on whether voice is detected. Detailed Implementation

[0031] Various aspects of this disclosure will now be explained with reference to the accompanying drawings. Unless the shape, relative position, and other aspects of the components described in any aspect are explicitly defined, the scope of this disclosure is not limited to the components shown, which are for illustrative purposes only. Furthermore, while numerous details have been set forth, it should be understood that some embodiments may be implemented without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of the description. Moreover, unless the meaning is explicitly contrary, all scopes shown herein are to be considered to include the endpoints of each scope.

[0032] Figure 1An audio system 1 according to one aspect is illustrated, comprising a local device and one or more remote devices participating in a call during a joint media playback session. As described herein, this allows users of the devices to listen to (and / or watch) media content (e.g., on one or more devices) while participating in a session with each other. The audio system includes a local (or first electronic) device 2, a remote (or second electronic) device 3, a network 4 (e.g., a computer network, such as the Internet), a media content server 5, and an audio output device 6. In one aspect, the system may include more or fewer elements. For example, the system may have one or more remote devices, all of which participate in a call and joint media playback session with each other and with the local device, as described herein. In another aspect, the audio system may include one or more remote (electronic) servers communicatively coupled to at least some of the devices in audio system 1, and may be configured to perform at least some of the operations described herein. In yet another aspect, the system may not include an audio output device. In this case, the local device may perform audio output operations (e.g., using one or more signals to drive one or more speakers).

[0033] In one aspect, the local device (and / or remote device) can be any electronic device (e.g., having electronic components such as a processor, memory, etc.) capable of participating in a call such as a telephone (or “voice-only” call) or video (conference) call while performing a joint media playback session with one or more other devices (e.g., one or more remote devices), wherein (at least some) of the devices simultaneously play back media content (e.g., musical works, movies, etc.). Further details regarding the simultaneous playback of media content are described herein. For example, the local device can be a desktop computer, a laptop computer, a digital media player, etc. In one aspect, the device can be a portable electronic device (e.g., capable of handheld operation), such as a tablet computer, a smartphone, etc. In another aspect, the device can be a head-mounted device, such as smart glasses, or a wearable device, such as a smartwatch. In one aspect, the remote device can be a device of the same type as the local device (e.g., both devices are smartphones). In another aspect, at least some of the remote devices can be different, such as some being desktop computers and others being smartphones.

[0034] As shown in the figure, local device 2 is coupled to remote device 3 and / or media content server 5 via a computer network (e.g., the Internet) 4 (e.g., communicationally). Specifically, the local device and the remote device can be configured to establish and participate in telephone (or voice-only) calls, where the participating devices exchange audio data. For example, each device transmits at least one microphone signal as an uplink audio signal to the other participating devices and receives at least one audio signal from the other devices as a downlink audio signal for playback by one or more speakers. In one aspect, the network may include a Public Switched Telephone Network (PSTN), through which the local device and the remote device are able to make outgoing calls and / or receive incoming calls. In another aspect, the local device can be configured to establish Internet Protocol (IP) telephone (or Voice over IP (VoIP)) calls with one or more remote devices via a network (e.g., the Internet). Specifically, the local device may use any signaling protocol (e.g., Session Initiation Protocol (SIP)) to establish a communication session and any communication protocol (e.g., Transmission Control Protocol (TCP), Real-Time Transport Protocol (RTP), etc.) to exchange audio data during the call. For example, when a call is initiated (e.g., by a telephone application running within the local device), the local device may transmit one or more microphone signals (e.g., as uplink audio signals) captured by one or more microphones as audio data (e.g., IP packets) to one or more remote devices, and receive one or more (e.g., downlink audio) signals from the remote devices via the network to drive one or more speakers of the local device. On the other hand, the local device may be configured to establish a wireless (e.g., cellular) call. In this case, network 4 may include one or more cell towers, which may be part of a communication network (e.g., a 4G Long Term Evolution (LTE) network) that supports data transmission (and / or voice calls) of electronic devices such as mobile devices (e.g., smartphones).

[0035] On the other hand, local and remote devices can be configured to establish and participate in video calls with one or more remote devices. In this case, the local device can establish a video call (e.g., similar to VoIP, using SIP to initiate the session and RTP to transmit data) and exchange video and / or audio data with one or more remote devices while establishing the video call. For example, the local device may include one or more cameras that capture video, which is encoded using any video codec (e.g., H.264) and transmitted to the remote device for decoding and display on one or more displays. Further details regarding the call are described herein.

[0036] In some aspects, the media content server 5 can be a standalone server computer or a cluster of server computers configured to stream media content to electronic devices such as local and remote devices. In this case, the server can be part of a cloud computing system capable of streaming data as a cloud-based service to one or more subscribers. In some aspects, the server can be configured to stream any type of media (or multimedia) content, such as audio content (e.g., musical works, audiobooks, podcasts, etc.), still images, video content (e.g., movies, television productions, etc.), etc. In one aspect, the server can use any audio and / or video encoding format and / or any method used to stream content to one or more devices.

[0037] In one aspect, media content server 5 can be configured to simultaneously stream media content to one or more devices to allow these devices to participate in a joint media playback session. For example, the server may receive a request from a device (e.g., local device 2) to stream media content segments, which may include audio content (e.g., musical works) and / or video content (e.g., video signals associated with a movie), together with another device (e.g., remote device 3). In one aspect, a transmission request may be sent by the local device (and / or the remote device) in response to the device receiving user input to begin playing back media content, such as... Figure 3 and Figure 4 As shown. In this scenario, the server can establish communication links with local and remote devices already involved in a call (e.g., telephone and / or video). Once the communication links are established, the server can encode audio content using any codec (e.g., MP3, AAC, etc.) and / or encode video content using any codec, then transmit the encoded content to each device for decoding and output. Alternatively, the local device can transmit a message to the remote device requesting the initiation of a joint media playback session. In response, the remote device can communicate with the media content server to retrieve media content and play it synchronously with the local device. In one aspect, devices participating in the joint media playback session can synchronously output media content, allowing users to simultaneously output and experience the content. In some aspects, any timing synchronization method can be used (e.g., by the participating devices and / or the server) to ensure simultaneous and synchronized streaming of media. This document describes more about joint media playback sessions.

[0038] As shown, the audio output device 6 can be any electronic device including at least one speaker and configured to output sound by driving the speaker. For example, as shown, the device is a wireless headset (e.g., in-ear headphones or earbuds) designed to be positioned on (or in) a user's ear and designed to output sound into the user's ear canal. In some aspects, the headphones can be a sealed type with flexible earpiece ends designed to acoustically seal the entrance to the user's ear canal relative to the surrounding environment by blocking or occluding it within the ear canal. As shown, the output device includes a left earpiece for the user's left ear and a right earpiece for the user's right ear. In this case, each earpiece can be configured to output at least one audio channel of media content (e.g., the right earpiece outputs the right audio channel of a stereo recording (such as a musical work) with two-channel input, and the left earpiece outputs the left audio channel). In another aspect, the output device can be any electronic device including at least one speaker and arranged for wear by a user and arranged to output sound by driving the speaker with an audio signal. For example, the output device can be any type of headset, such as over-ear (or on-ear) headphones that at least partially cover the user's ears and are arranged to direct sound into the user's ears.

[0039] In some respects, the audio output device can be a headset, as illustrated herein. In other respects, the audio output device can be any electronic device arranged to output sound to the surrounding environment. Examples may include standalone speakers, smart speakers, home theater systems, or infotainment systems integrated into vehicles.

[0040] In one aspect, the output device can be a wireless device communicatively coupled to the local device for exchanging audio data. For example, the local device can be configured to establish a wireless connection with the audio output device via a wireless communication protocol, such as Bluetooth or any other wireless communication protocol. During the established wireless connection, the local device can exchange (e.g., transmit and receive) data packets (e.g., Internet Protocol (IP) packets) with the audio output device, which can include audio digital data of any audio format. Specifically, the local device can be configured to establish and communicate with the audio output device via a two-way wireless audio connection, such as making hands-free calls or using voice commands. Examples of two-way wireless communication protocols include, but are not limited to, Hands-Free Mode (HFP) and Headset Mode (HSP), both of which are Bluetooth communication protocols. In another aspect, the local device can be configured to establish and communicate with the output device via a one-way wireless audio connection, such as the Advanced Audio Distribution Profile (A2DP) protocol, which allows the local device to transfer audio data to one or more audio output devices. Further details regarding these wireless audio connections are described herein.

[0041] On the other hand, local device 2 can be communicatively coupled to audio output device 6 via other methods. For example, both devices can be coupled via a wired connection. In this case, one end of the wired connection can be (e.g., fixedly) connected to the audio output device, while the other end can have a connector, such as a media jack or a Universal Serial Bus (USB) connector, that inserts into the jack of the audio source device. Once connected, the local device can be configured to drive one or more speakers of the audio output device with one or more audio signals via the wired connection. For example, the local device can transmit the audio signals as digital audio (e.g., PCM digital audio). On the other hand, the audio can be transmitted in an analog format.

[0042] In some respects, local device 2 and audio output device 6 may be different (separate) electronic devices, as illustrated herein. In other respects, the local device may be a component of the audio output device (or integrated with the audio output device). For example, as described herein, at least some components of the local device (such as a controller) may be part of the audio output device, and / or at least some components of the audio output device may be part of the local device. In this case, each device may be communicatively coupled via traces that are part of one or more printed circuit boards (PCBs) within the audio output device.

[0043] Figure 2 A block diagram is shown of a local device 2 that initiates a joint playback media session when participating in a call (e.g., voice or video) with one or more remote devices 3, according to one aspect, and an audio output device 6 that communicates wirelessly with the local device is also shown. The local device 2 includes a controller 20, a network interface 21, a speaker 22, a microphone 23, a camera 24, a display screen 25, and (optionally) one or more additional sensors 40. In one aspect, the local device may include more or fewer elements as described herein. For example, the device may include two or more of at least some of the elements (e.g., having two or more microphones 23).

[0044] Controller 20 may be a dedicated processor such as an application-specific integrated circuit (ASIC), a general-purpose microprocessor, a field-programmable gate array (FPGA), a digital signal controller, or a set of hardware logic structures (e.g., filters, arithmetic logic units, and dedicated state machines). The controller is configured to perform audio signal processing operations and / or networking operations. For example, controller 20 may be configured to participate in a call and simultaneously perform a joint media playback session to stream media content to one or more remote devices via network interface 21. Alternatively, the controller may be configured to perform audio signal processing operations on audio data of the media content and / or audio data associated with the participating call (e.g., downlink signals). Further details regarding the operations performed by controller 20 are described herein.

[0045] In one aspect, one or more sensors 40 are configured to detect the environment (e.g., where a local device is located) and generate sensor data based on the environment. In some aspects, the controller may be configured to perform operations based on the sensor data generated by the one or more sensors 40. For example, the local device may include a proximity sensor (e.g., optical) designed to generate sensor data indicating a specific distance between an object and the sensor (and / or the local device). As another example, the local device may include an inertial measurement unit (IMU) designed to measure the position and / or orientation of the local device. In one aspect, the sensor may be a component of the local device (or integrated into the local device). In another aspect, the sensor may be a separate electronic device (e.g., via network interface 21) communicatively coupled to the controller. For example, the audio output device 6 may include one or more sensors whose data may be provided to the local device via a wireless connection.

[0046] The speaker 22 may be, for example, an electrically driven driver specifically designed for sound output in a particular frequency band, such as a woofer, tweeter, or midrange driver. In one aspect, the speaker 22 may be a “full-range” (or “full-frequency”) electrically driven driver that reproduces as much of the audible frequency range as possible. The microphone 23 may be any type of microphone (e.g., a differential pressure gradient microelectromechanical system (MEMS) microphone) configured to convert acoustic energy caused by sound waves propagating in an acoustic environment into an input microphone signal.

[0047] In one aspect, camera 24 is a complementary metal-oxide-semiconductor (CMOS) image sensor capable of capturing digital images including image data representing the field of view of camera 24, wherein the field of view includes the scene of the environment in which device 2 is located. In some aspects, the camera may be a charge-coupled device (CCD) camera type. The camera is configured to capture still digital images and / or video represented by a series of digital images. In one aspect, the camera may be located anywhere near the local device. In some aspects, the device may include multiple cameras (e.g., where each camera may have a different field of view).

[0048] Display screen 25 is designed to present (or display) digital image or video (or image) data. In one aspect, the display screen may use liquid crystal display (LCD) technology, light-emitting polymer display (LPD) technology, or light-emitting diode (LED) technology, although other display technologies may be used in other aspects. In some aspects, the display screen may be a touch-sensitive display screen configured to sense user input as an input signal. In some aspects, the display screen may use any touch sensing technology, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies.

[0049] Audio output device 6 includes a controller 75, a network interface 76, a speaker 77, a microphone 78, and an accelerometer 79. In one aspect, the device may include more or fewer elements. For example, the output device may include one or more microphones and / or one or more speakers. In some aspects, the output device may include a microphone that is an "external" (or reference) microphone arranged to capture sound from the acoustic environment, while having at least one other "internal" (or error) microphone arranged to capture sound (and / or sense pressure changes) within the user's ear (or ear canal). In the case of in-ear headphones, the internal microphone can sense the inside of the user's ear when the headphones are placed on (or in) the user's ear.

[0050] Accelerometer 79 is arranged and configured to receive (detect or sense) speech vibrations generated when a user (e.g., a user who may be wearing an output device) speaks, and to generate an accelerometer signal representing (or including) the speech vibrations. Specifically, the accelerometer is configured to sense bone conduction vibrations transmitted from the user's vocal cords to the user's ear (ear canal) during speaking and / or humming. For example, when the audio output device is a wireless headset, the accelerometer may be located on or inside the headset anywhere it can contact a part of the user's body to sense vibrations.

[0051] In one aspect, controller 75 is configured to perform audio signal processing operations and / or networking operations, as described herein. For example, the controller may be configured to acquire (or receive) audio data of media content (as analog or digital audio signals) or media content desired by the user (e.g., music, etc.) for playback via speaker 77. In some aspects, the controller may acquire audio data from local memory, or the controller may acquire audio data from network interface 76, thereby acquiring data from an external source such as local device 2 (via its network interface 21). For example, the output device may stream audio signals from the local device (e.g., via Bluetooth connection) for playback via speaker 77. The audio signal may be a signal input audio channel (e.g., mono). In another aspect, the controller may acquire two or more input audio channels (e.g., stereo channels) for output through two or more speakers. In one aspect, where the output device includes two or more speakers, the controller may perform additional audio signal processing operations. For example, the controller can spatially render the input audio channels (e.g., by applying spatial filters, such as the head-related transfer function (HRTF)) to generate binaural output audio signals for driving at least two speakers (e.g., a left speaker and a right speaker).

[0052] In one aspect, controller 75 may be configured to perform (additional) audio signal processing operations based on elements coupled to the controller. For example, when the output device includes two or more “outside-ear” speakers arranged to output sound into an acoustic environment rather than speakers arranged to output sound into a user’s ears (e.g., as speakers in in-ear headphones), the controller may include an audio output beamformer configured to generate speaker driver signals that produce spatially selective sound output when driving two or more speakers. Thus, when used to drive speakers, the output device may generate a directional beam pattern that can be directed to a location within the environment.

[0053] In some aspects, the controller 75 may include a sound pickup beamformer configured to process audio (or microphone) signals generated by two or more external microphones of the output device to form a directional beam pattern (as one or more audio signals) for spatially selective sound pickup in certain directions, thereby increasing sensitivity to the location of one or more sound sources. In some aspects, the controller may perform audio processing operations (e.g., perform spectral shaping) on ​​the audio signal containing the directional beam pattern and / or transmit the audio signal to a local device.

[0054] On the other hand, controller 75 may perform other functions. For example, controller 75 may be configured to perform active noise cancellation (ANC) to make speaker 77 produce noise immunity in order to reduce ambient noise from the environment leaking into the user's ears. The ANC function may be implemented as feedforward ANC, feedback ANC, or a combination thereof. Thus, controller 75 may receive a reference microphone signal from a microphone such as microphone 78 that captures external ambient sound. On the other hand, controller may perform any ANC method to produce noise immunity. On the other hand, controller 75 may perform a transparency function, wherein the sound reproduced by audio output device 6 is a reproduction of ambient sound captured by the device's external microphone in a "transparent" manner (e.g., as if the headphones were not being worn by the user). Controller 75 processes at least one microphone signal captured by at least one external microphone 78 and filters the signal through a transparency filter, which reduces acoustic blockage caused by the audio output device being located on, in, or above the user's ears, while also preserving the spatial filtering effect of the wearer's anatomical features (e.g., head, auricle, shoulders, etc.). The filter also helps to preserve the timbre and spatial cues associated with the actual ambient sound. In one respect, the filter for the transparency function can be user-specific, depending on specific measurements of the user's head. For example, the controller 75 can determine the transparency filter based on the head-related transfer function (HRTF) or the equivalent head-related impulse response (HRIR) based on the user's anthropometric measurements.

[0055] As described herein, both the local device and the audio output device are configured to establish a wireless audio connection (e.g., a Bluetooth connection) for exchanging audio data. In one aspect, controller 75 (and / or controller 20) can be configured to switch between a two-way wireless audio connection (e.g., an HFP connection) and a one-way wireless audio connection (e.g., an A2DP connection) to communicatively couple the two devices together for exchanging (and transmitting) audio data. Further details regarding the switching between audio connections are described herein.

[0056] In one respect, the operations performed by the controller may be implemented in software (e.g., as instructions stored in memory and executed by the controller) and / or may be implemented by hardware logic structures as described herein.

[0057] On the other hand, at least some of the operations performed by the audio system 20 as described herein may be performed by the local device 2 and / or the audio output device 6. For example, the local device may include two or more speakers and may be configured to perform sound output beamforming operations (e.g., when the local device includes two or more speakers). On the other hand, at least some of the operations may be performed by a remote server communicatively coupled to either device, for example via a network (e.g., the Internet).

[0058] In one aspect, at least some elements of the local device 2 and / or the audio output device 6 may be integrated with (or as a component of) each respective device. For example, when the audio output device is an over-ear headphone, the microphone, speaker, and accelerometer may be part of at least one earcup of the headphone, which is placed on the user's ear. In another aspect, at least some elements may be separate electronic devices communicatively coupled to the device. For example, the display screen 25 may be a separate device (e.g., a display monitor or television) communicatively coupled to the local device (e.g., a wired or wireless connection) to receive image data for display. As another example, the camera 24 may be a component of a separate electronic device coupled to the local device to provide captured image data (e.g., a webcam).

[0059] As described herein, the local device 2 and remote device 3 of the audio system 1 can perform a joint media playback session while participating in a call, allowing users of the devices to communicate while experiencing simultaneous media content playback. In one aspect, the local device can initiate a joint media playback session while already participating in a call. Figure 3 and Figure 4 Graphical examples of local and remote devices initiating joint media playback when participating in telephone and video conferencing calls are shown respectively.

[0060] Figure 3The diagram illustrates three phases 26 to 28 according to one aspect, in which local device 2 and remote device 3 initiate a joint playback media session to synchronously play back musical works while participating in a telephone call. Phase 26 illustrates the main (or master) screen user interface (UI) displayed on the display screen of each corresponding device when the devices participate in a telephone call. In one aspect, either device can initiate the telephone call, as described herein. Specifically, the local device's master screen UI 11 displays caller ID information of the remote device over several optional UI items, each associated with an application (e.g., applications 1 through 4), including a media application 29 that streams media content (e.g., from a media content server 5) to the local device when executed by the local device. Specifically, media application 29 could be a music streaming application that, when executed, streams music for playback by speaker 22 (and / or speaker 77 of an audio output device). Similarly, the remote device's master screen UI 12 displays caller ID information of the local device over several (similar) UI items, which are those displayed for the local device. In one respect, any device can initiate a telephone call using any known method. For example, a user of a local device might initiate a phone application stored on the local device and dial the phone number of a remote device. Once dialed, the local device can connect to the remote device via a cellular network of Network 4 (e.g., a 4G Long Term Evolution (LTE) network), as described herein.

[0061] This phase also shows the user of local device 2 pressing a UI item associated with media application 29. For example, the display screen of the local device (e.g., Figure 2 The display screen 25 shown may be a touch-sensitive display screen, as described herein. The local device may receive user input in response to the user pressing a UI item of the media application 29. The second stage 27 shows the result of the user pressing a UI item of the media application 29. Specifically, this stage shows the media application's UI 30 displayed on the local device's display screen, which shows the title of the musical work (e.g., "The Music"), and playback control UI items including a play button, a rewind button, and a fast forward button. This stage also shows that the user has pressed the "play" button.

[0062] Phase 3, 28, illustrates the result of a user selecting the play button on the local device. Specifically, once the play button is selected, the local device sends a request to the media content server 5 to begin streaming media content to both the remote and local devices. In one aspect, when multiple devices are making a call together (e.g., a conference call), the media content server 5 can stream media content to each device participating in the conference call. Therefore, both the remote and local devices play back the media content (e.g., by driving the corresponding speakers with audio data from the media content received from the media content server). Thus, the two devices play back the content simultaneously and synchronously, as indicated by the progress indicators 39 for the two devices shown in the respective media application UI at the midpoint. Further details regarding simultaneous playback of media content are described herein.

[0063] Figure 4 The diagram illustrates three phases 31 to 33 according to one aspect, in which local device 2 and remote device 3 initiate a joint playback media session to synchronously play back video while participating in a video call. Phase 31 illustrates the main screen UI displayed on the display screen of each respective device when the devices participate in a video call. Specifically, overlaid on the main screen UI 11 of the local device is the video call UI 14, which shows the video representation of the local user 38 located in the upper right corner of the UI and the video representation of the remote user 37 (which is larger than the local user's representation) located in the center of the video call UI. Similarly, overlaid on the main screen UI 12 of the remote device is the video call UI 15, which shows the video representation of the remote user located in the center of the UI and the video representation of the local user located in the upper right corner. In one aspect, video representations can be generated using video data captured by one or more cameras of each device. For example, when the local user is in the field of view of camera 24, the camera can capture the local user's video data, which is then displayed on the local device and transmitted (e.g., via network 4) to the remote device for display on the remote device's display screen.

[0064] This phase also shows the local user selecting an optional UI item associated with a media application 35 within the main screen UI 11, which could be a video streaming application. The second phase 32 shows the result of the user pressing a UI item of the media application 35. Specifically, this phase shows the media application 35's UI 18 displayed on the local device's display screen, showing the movie's title (e.g., "The Movie"), the playback duration of one hour and thirty minutes, and the play button pressed by the local user.

[0065] Phase 33 illustrates the result of the local user selecting the play button in the media application UI 18. Specifically, once the play button is selected, the local device sends a request to the media content server 5 to begin streaming media content (e.g., audio and video data of a movie) to the devices participating in the video call. Thus, both devices synchronously play back the video of the media content 36 (and output the audio of the media content) while still participating in the video call.

[0066] As these examples illustrate, when a device participates in a phone call, audio content can be played back during a joint media playback session, while when a device participates in a video call, both video and audio content can be played back during the session. On the other hand, when both local and remote devices participate in a phone or video call, any type of media content can be played back during a joint media playback session. For example, when a device participates in a phone call, video can be played back during a joint media playback session.

[0067] While participating in joint media playback sessions during a call can provide participants with a better user media experience than media content being played back through their devices (e.g., by allowing participants to discuss the media content in real time), there can be some drawbacks. For example, conversations between participants may drown out or mask the audio of the media content. For instance, when participants are watching a video, conversations between them may be indistinguishable from the dialogue in the simultaneously playing video. Therefore, participants engaged in these one-way conversations may find it difficult to speak while the video is playing. Furthermore, this can also degrade the overall user experience for those not participating in these conversations, as the conversations may distract them, causing them to focus entirely on the audio of the video. Therefore, it is essential to maintain the quality of media audio playback when participants are engaged in joint media playback sessions during a call.

[0068] To overcome these shortcomings, this disclosure describes an audio system capable of maintaining audio quality for media content playback during a media playback session by processing remotely active voice during a call. Specifically, in a joint media playback session involving a call and in which a local device and (at least one) remote device independently stream media content for synchronous playback, the audio system determines, based on the output of a voice activity detector (VAD), that the downlink (audio) signal from the remote device includes voice. If so, the audio system applies a scalar gain to the audio signal of the media content to reduce the signal level of the audio signal. The audio system then drives a speaker with the mixture of the downlink signal and the audio signal. Thus, the system can manage the signal level of the media content while the participants at the remote device are speaking.

[0069] Figure 5A block diagram of a local device 2 is shown, which performs audio signal processing operations on the audio signal of media content based on whether voice is detected in the signal of a telephone call performed between the local device 2 and at least one remote device 3. Specifically, the diagram shows a controller 20 having multiple operation blocks for performing audio signal processing operations to process remote active voice during a call and joint media playback session. As shown, the controller includes a call manager 46, a joint media playback session manager 47, a voice digital signal processor (DSP) 41, a voice activity detector (VAD) 42, a scalar gain 43, a mixer (e.g., a matrix) 44, and an optional additional DSP 45.

[0070] Call manager 46 is configured to initiate (and perform) calls between local device 2 and one or more remote devices 3. In one aspect, the call manager may initiate a call in response to user input. For example, the call manager may be part of (or receive instructions from) a telephone application executed by the local device (e.g., the controller 20 of the local device). For example, the telephone application may display a UI on the display screen 25 of the local device, which may provide the user of the local device with the ability to initiate a call (e.g., a keyboard, a contact list, etc.). Once the UI receives user input (e.g., dialing a remote user's phone number using the keyboard), the call manager may communicate with the network interface 21 of local device 2 to establish a call, as described herein. In one aspect, the telephone call may be made over any network, such as via the PSTN and / or via the Internet (e.g., for VoIP calls). In some aspects, the call manager may initiate a call as described herein and / or using any method.

[0071] Once a call is initiated, the call manager can exchange call data between the remote devices used by the local device to participate in the call. For example, the call manager can receive one or more downlink audio signals from each of the remote devices. In one aspect, the call manager can mix the downlink signals into (at least one) downlink audio signals (e.g., via matrix mixing). Furthermore, the call manager can receive microphone signals (e.g., which may include the local user's voice) from microphone 23 and can transmit the microphone signals as uplink audio signals to each remote device. In some aspects, when the local device includes two or more microphones, the call manager can transmit sound pickup beamformer signals that include a directional beam pattern.

[0072] The joint media playback session manager 47 is configured to initiate a joint media playback session between a local device and one or more remote devices, where the two devices independently stream media content for synchronized playback. For example, in response to receiving an instruction to initiate a session, the playback session manager may transmit a request to initiate a session to a media content server, as described herein. Specifically, a media application running on the local device may transmit instructions to the session manager in response to receiving user input (e.g., based on the user selecting a play button in the media application, such as...). Figure 3 and Figure 4 (As shown). On the other hand, the session manager may request user authorization before initiating a session. For example, once a user initiates media playback in a media application, the session manager may provide a notification (e.g., a pop-up notification displayed on display 25) requesting user authorization to initiate a joint media playback session with at least some of the call participants. When user authorization is received (e.g., by receiving the user's selection of a UI item within the pop-up notification), the session manager may process the request to initiate the session, as described herein.

[0073] In one aspect, the joint media playback session manager 47 is configured to receive media content data (e.g., once a session is initiated). In this case, the session manager receives at least one audio signal (or audio channel) associated with the media content. For example, the received audio signal may be associated with a musical work that a local user has requested to be played back, such as... Figure 3 As shown. In one aspect, a session manager can receive two or more audio signals from a segment of media content. For example, when streaming a musical work from a media content server, the session manager can receive two audio channels (e.g., the left and right channels of a stereo recording of the musical work). In another aspect, a session can receive two or more audio channels, such as the entire audio track of a 5.1 surround sound video.

[0074] The voice DSP 41 is configured to receive a downlink audio signal from a call manager and is configured to perform voice processing operations on the signal. In one aspect, the voice DSP may perform a noise reduction algorithm on the downlink signal to reduce (or eliminate) the noise contained therein (e.g., to produce a voice signal that primarily contains the voice of the remote user). In another aspect, the algorithm may apply a high-pass filter to process the signal, since most noise (or non-speech noise) is likely low-frequency content. In another aspect, the algorithm may improve the signal-to-noise ratio (SNR) of the signal. To do this, the voice DSP may perform spectral shaping of the downlink signal by applying one or more filters (e.g., low-pass filter, band-pass filter, high-pass filter, etc.). Alternatively, the DSP may apply a scalar gain value to the signal. In one aspect, the voice DSP may perform any method to process the downlink signal to reduce the noise contained therein.

[0075] VAD 42 is configured to receive (e.g., processed) a downlink audio signal and is configured to perform sound activity detection (or speech detection) operations to detect the presence (or absence) of user voice (speech). For example, the VAD may determine whether at least a portion of the spectral content of the downlink signal is associated with human speech. Alternatively, the VAD may determine the presence of speech based on whether the signal level of the downlink signal exceeds a threshold. In some aspects, the VAD may use any method to determine the presence of speech within the signal. The VAD is configured to generate an output based on the downlink signal. Specifically, the VAD may generate a VAD signal indicating whether speech is contained within the downlink signal. For example, the VAD signal may have a high signal level (e.g., 1) when speech is detected, and a low signal level (e.g., 0) when no speech is detected (or at least not within a threshold level). Alternatively, the VAD signal need not be a binary decision (speech / non-speech); as described herein, it can be a speech presence probability based on a scalar gain to be adjusted. In some respects, the VAD signal can also indicate the signal level of the detected speech (e.g., sound pressure level (SPL)).

[0076] As described herein, the VAD can receive a mixture of two or more downlink audio signals (e.g., mixed by call manager 46), each downlink signal received from a remote device participating in a call (e.g., a conference) with the local device. In one aspect, the VAD can receive each individual downlink signal to determine whether at least one of the downlink signals contains voice. Once voice is detected in at least one of the downlink signals, the VAD can generate a VAD signal to indicate voice detection. In some aspects, the voice DSP can process each individual downlink signal before it is received by the VAD.

[0077] On the other hand, in addition to (or instead of) generating a VAD signal, the local device may optionally receive a VAD signal from a remote device (e.g., at least one of them). Specifically, each remote device may include its own VAD and may be configured to generate a VAD signal as the output of the VAD, indicating whether at least one microphone signal generated by the microphone of the remote device (and / or its uplink signal transmitted to local device 2 during a call) includes the active voice of the remote user. Once generated, each remote device may transmit the VAD signal to the local device via network 4. Once received, scalar gain 43 may apply a scalar gain value to the audio signal of the media content based on the VAD signal received from the remote device.

[0078] Scalar gain 43 is configured to receive an audio signal from the joint media playback session manager 47 and a VAD signal from VAD 42 (and / or from at least one remote device), and is configured to process the audio signal based on the VAD signal. Specifically, the scalar gain is configured to adjust the signal level (e.g., at least a portion thereof) of the audio signal by applying one or more scalar gain values ​​based on whether the VAD signal indicates the detection of speech within the downlink audio signal. Specifically, the gain adjustment may reduce the volume level of the audio signal of media content associated with (e.g., streamed by) the joint media playback session. In one aspect, the applied scalar gain value may be a predetermined value. In another aspect, the value may be based on the VAD signal. For example, as described herein, the VAD signal may indicate the signal level of the downlink audio signal (or more specifically, the signal level of speech contained therein). In this case, the scalar gain may be configured to adjust the applied scalar gain value based on the signal. For example, when the voice detected in the downlink audio signal is at a certain signal level, a scalar gain can be applied to reduce the signal level of the audio signal to below the certain signal level of the downlink signal, so as to ensure that the sound of the media content is lower than the voice in the call.

[0079] Mixer 44 is configured to receive a processed audio signal from scalar gain 43 and a processed downlink audio signal from voice DSP 41, and is configured to perform matrix mixing operations, for example, to produce a mixed content of the two signals. The controller can use the mixed signal to drive speaker 22 to play back the sound of a call and to play back the media content of a session. Alternatively, the mixer may receive one or more unprocessed downlink audio signals. For example, the mixer may receive downlink audio signals from call manager 46 instead of the processed downlink audio signals from voice DSP 41.

[0080] In one aspect, the controller may optionally have an additional DSP 45 that can be configured to perform one or more audio signal processing operations on the mixed content. For example, the additional DSP may perform at least some of the operations described herein, such as spatially rendering the mixed content (e.g., by applying spatial filters, such as head-related transfer functions (HRTF)) to generate binaural audio signals for driving one or more speakers (e.g., left and right speakers), as described herein. The controller 20 may then use the processed mixed content to drive the speaker 22, as described herein. Thus, in response to determining that a remote user has begun (and / or is actively) speaking during a call with a local user, the controller may perform the operations described herein to reduce the volume level of the media content.

[0081] As described so far, in response to the detection of sound (or speech) in one or more downlink signals from one or more remote devices, controller 20 applies a scalar gain. Alternatively, this determination may be based on whether a local user on the local device is speaking. Specifically, a VAD signal generated by the VAD can indicate whether one or more remote users and / or a local user are speaking. To determine this, voice DSP 41 may optionally acquire a microphone signal generated by microphone 23 to perform the noise reduction operation described herein. The VAD may receive processed downlink audio signals and / or processed microphone signals from voice DSP 41 and may generate a VAD signal based on either (or both) of these signals. Thus, when a local or remote user is speaking, the local device may reduce the signal level of the audio signal of the media content.

[0082] In one aspect, when the media content includes two or more audio signals, the controller performs at least some of the operations on at least one of the audio signals. For example, when the media content includes two audio channels for stereo recording, the controller 20 may perform at least some of the operations on the two audio channels to reduce the signal level of each audio channel output by two or more speakers from the local device.

[0083] In some respects, controller 20 can process audio signals of media content, while the VAD signal indicates that the downlink signal includes remote active voice. Specifically, when the VAD signal indicates the presence of voice (e.g., as long as a remote or local user is speaking), scalar gain 43 can continue to apply a scalar gain value. Once the VAD signal indicates that there is no more voice, the controller can stop applying scalar gain 43, in which case the audio signal can enter mixer 44 without scalar gain adjustment. In one respect, once there is no more voice, the applied scalar gain value can be gradually decreased to gradually increase the signal level of the audio signal.

[0084] Figure 6 A block diagram of a local device 2 is shown, which performs audio signal processing operations on the audio signal of media content based on whether an audio output device 6 detects voice. Specifically, the diagram illustrates the local device communicatively coupled to the audio output device to perform (e.g., "hands-free") calls and such... Figure 5 The joint media playback session described above. For example, two devices may connect via a two-way wireless audio connection (e.g., according to the HFP protocol), where the two devices exchange audio data of a telephone call and media content being played back during the joint media playback session. For example, the audio output device may be a hands-free device, such as a wireless headset configured to transmit a microphone signal generated by microphone 78 to controller 20 (e.g., the controller's call manager 46), which then transmits the microphone signal as an uplink signal for the call to one or more remote devices. Furthermore, the local device transmits a mixture of audio signals and (processed) downlink signals (e.g., processed) to the audio output device via the two-way audio connection, which uses the mixed content to drive speaker 77 (instead of...). Figure 5 The speaker 22 is driven using mixed content, as shown.

[0085] This figure also illustrates that scalar gain 43 can apply a gain value based on the output of VAD 82 of the local device. Specifically, a gain value can be applied in response to the audio output device detecting the speech of a local user. For example, the audio device includes VAD 82, which is configured to receive an accelerometer signal generated by accelerometer 79 and is configured to generate a VAD signal based on the received signal. Specifically, the VAD determines whether the energy level of the accelerometer signal is higher than an accelerometer signal threshold (or energy threshold), which could be an indication that a user is speaking. In response to determining that the energy level is higher than the energy threshold, the VAD signal can be set to a high signal level, as described herein. When generating the VAD signal, audio output device 6 transmits the signal to local device 2, and scalar gain 43 receives the signal to apply a gain value based on the signal, as described herein.

[0086] In one aspect, in addition to VAD 82 receiving accelerometer signals (or alternatively), VAD may (optionally) receive microphone signals generated by microphone 78 to generate a VAD signal, as described herein. In another aspect, instead of generating a VAD, the audio output device may transmit accelerometer signals (and / or microphone signals) to VAD 42 of a local device, which may then use these signals to generate a VAD signal, as described herein. Thus, a local device (e.g., VAD 42 of the local device) may generate a VAD signal based on accelerometer signals generated by accelerometer 79.

[0087] Figure 7A block diagram of a local device 2 according to one aspect is shown, which performs audio signal processing operations based on whether voice is detected within the signal of a video call. Specifically, the diagram illustrates a controller 20 simultaneously conducting video call and joint media playback sessions with one or more remote devices while performing audio signal processing to handle remote active voice and / or performing video processing operations.

[0088] In one respect, local device 2 can perform video calls and joint media playback sessions, such as Figure 4 As shown. Specifically, call manager 46 can be configured to initiate (and conduct) video calls between local device 2 and one or more remote devices 3. In this case, in addition to transmitting the microphone signal captured by microphone 23 as an uplink audio signal, call manager can receive camera (e.g., video) signals from camera 24 and transmit the video signals as uplink video signals along with (or instead of) the uplink audio signals to the remote devices participating in the video call. For example, as described herein, call manager (e.g., in response to receiving a user request in a telephone or video conferencing application running within the local device) can establish a communication session with the remote devices, encode the microphone and camera signals, and transmit the encoded signals (as uplink signals) to the remote devices. In addition to transmitting the uplink signals, call manager can receive at least one downlink audio signal and at least one downlink video signal from each remote device participating in the video call for output to speaker 22 and display 25, respectively. In one aspect, any method can be used to initiate and conduct video calls. In some respects, the joint media playback session manager 47 can be configured to receive media content data including at least one audio signal and at least one video signal associated with a media content segment. For example, the received audio and video signals may be associated with a movie that a local user has requested to play back, such as... Figure 4 As shown.

[0089] In one respect, controller 20 can perform operations related to video calls and joint media playback sessions simultaneously. Figure 5 and Figure 6 The operation performed by the controller is similar to that described above. For example, the controller (e.g., VAD 42 of the controller) may determine whether a remote user on a remote device has started speaking (and / or is actively speaking) based on the downlink audio signal (e.g., audio content) of a video conferencing call. In response, the controller may use scalar gain 43 to apply a scalar gain value to reduce the volume level of the audio signal when output by speaker 22.

[0090] Furthermore, controller 20 includes additional operation blocks for performing audio signal processing operations and / or video processing operations based on whether the remote user's voice is active. For example, the controller includes a closed caption generator 48 and a video processor 49. The closed caption generator is configured to generate closed captions representing audio content contained within the audio signal of the media content based on the VAD signal output of VAD 42. Specifically, the caption generator can be configured to generate closed captions in response to controller 20 determining, based on the VAD signal (e.g., a VAD signal having a high signal level indicating that the downlink signal includes speech, as described herein), that the downlink signal (or at least one downlink signal) includes speech, and can be configured to display the closed captions. Thus, closed captions can be generated and displayed when the remote user begins to speak (and while the user is speaking). In one aspect, the caption generator can stop generating and displaying closed captions once the VAD signal indicates that the downlink signal no longer includes speech. In another aspect, the closed caption generator can continue generating and displaying closed captions for a period of time after the remote user has stopped speaking.

[0091] On the other hand, the closed caption generator 48 can be configured to generate closed captions for display in response to determining that the output sound level of the local device is below a threshold sound level. For example, the caption generator can determine whether the local user has lowered the volume of the local device (e.g., by adjusting the volume control of the local device to detect whether the user has lowered the volume). If so, the caption generator can automatically generate and display the captions. On the other hand, captions can be displayed based on the signal level of the audio signal associated with the media content. For example, the caption generator can generate and display captions in response to the processed audio signal of the media content with a scalar gain having a signal level below a threshold.

[0092] In one aspect, to generate closed captions, the closed caption generator is configured to receive an audio signal associated with media content streaming during a session from a session manager 47, and can be configured to generate captions based on the audio content contained therein. In some aspects, the generator may perform a speech-to-text algorithm to recognize speech contained within the audio signal and may generate a text representation of the recognized speech. Therefore, the captions may include a transcription of the audio content. In another aspect, the captions may include a text description of non-speech audio, such as a description of the current scene. In another embodiment, captions may be obtained from media content data rather than generated. In this case, the caption generator may receive the captions from the session manager. In some aspects, the caption generator may use any method to generate the captions.

[0093] In one aspect, the video processor 49 is configured to receive image data, such as downlink video signals from the call manager 46, video signals from the session manager 47, and (optionally) closed captions from the caption generator 48 (e.g., when a VAD signal indicates active remote voice), and is configured to present data for display on the display screen 25 for playback of media content (e.g., such as...) during a video call. Figure 4 (As shown). For example, a video processor can overlay closed captions onto the displayed video signal of media content. In some aspects, a video processor can perform other video processing operations on one or more of the video signals, such as image resizing, image compositing, etc.

[0094] In one aspect, the controller can adjust the playback of media content based on whether VAD 42 detects remote active voice. Specifically, once (e.g., via VAD) it is determined that the remote voice is no longer active, the joint media playback session 47 can rewind the media content to a point before the initial detection of active voice. For example, the joint media playback session manager can receive the VAD signal from VAD 42 and determine a first timestamp along the playback duration of the media content, at which the VAD signal begins to indicate that the downlink signal includes voice (e.g., the moment the VAD signal transitions from a low signal level to a high signal level). At this time, the remote user and the local user may have started talking. Once the conversation ends, the media content can be rewound to start playback at (or before) the first timestamp along the playback duration. For example, once the session manager determines a second subsequent timestamp, at which it is determined that the VAD signal indicates that the downlink signal has stopped including voice (e.g., the moment the VAD signal level transitions from a high signal level to a low signal level), the session manager can pause the playback of the media content (at or after the second timestamp). In one aspect, pausing video playback may include pausing the display of media content at a point in time during the playback duration. Additionally, audio playback of the audio signal can be paused by stopping the mixing of the downlink signal and audio signal to drive speaker 22. In another aspect, audio playback of the audio signal may be paused while playback of the downlink audio signal continues. In this case, once it is determined that audio playback will be paused, mixer 44 can stop mixing the two signals and can pass the downlink signal to drive speaker 22. Therefore, local and remote users can participate in conversations and resume experiencing the media content playback when finished.

[0095] In one aspect, playback adjustments can occur on at least some of the remote devices participating in the call and the joint media playback session with the local device. For example, in response to a remote voice no longer being active, controller 20 can transmit a control signal to the remote device, instructing the device to rewind the playback to a point along the playback duration.

[0096] Figures 8 to 10 These are flowcharts for processes 50, 60, and 70, respectively, for performing one or more operations in response to the detection of remote active voice. In one aspect, the processes may be performed by one or more devices of the audio system 1, such as... Figure 1 As shown in the diagram. For example, at least some of the operations of these processes may be performed by local device 2 (e.g., its controller 20) and / or by audio output device 6 (e.g., its controller 75).

[0097] about Figure 8 This diagram is a flowchart of one aspect of a process 50 for processing audio signals of media content based on whether speech is detected within the downlink audio signal. Process 50 begins with controller 20 initiating a call (e.g., a telephone call or video call) between local device 2 and one or more remote devices 3 (at box 51). As described herein, the call may be initiated by call manager 46 in response to receiving a request from a local user. In one aspect, the call may be initiated in response to receiving an incoming call from one or more remote devices. In this case, the call may be initiated by call manager in response to a user accepting the call (e.g., via the user selecting a UI item in a telephone application for answering the call displayed on display 25 when an incoming call signal is received from a remote device).

[0098] During a call, controller 20 initiates a joint media playback session as local device 2, where the local device and one or more remote devices independently stream media content for synchronous playback (at box 52). For example, joint media playback session manager 47 can initiate playback based on user input. In one aspect, the playback session can be between all devices making the call. In another aspect, the playback session can be initiated between the local device and at least some of the remote devices. In this case, the local user can define which remote devices will participate when initiated. In some aspects, initiating a joint media playback session may be in response to controller 20 receiving an initiation request from one or more remote devices and / or media content server 5.

[0099] As described herein, once initiated, controller 20 may receive at least one audio signal and / or at least one video signal associated with media content, and may be configured to play back media content and simultaneously output downlink audio signals and / or downlink video signals, as described herein.

[0100] Controller 20 determines, based on the output from VADs (such as VAD 42 of controller 20 and / or VAD 82 of audio output device 6), whether the downlink signal from one or more remote devices includes (e.g., remote activity) speech (at decision box 53). Specifically, the controller may determine whether the VAD signal is at a high signal level, occurring when a remote user begins or has already begun speaking. If so, controller 20 applies a scalar gain to the audio signal associated with the media content to reduce the signal level of the audio signal (at box 54). For example, upon detecting speech, the controller may apply a scalar gain 43 to the audio signal from session manager 47. Controller 20 mixes the (gain-adjusted) audio signal and the downlink signal (at box 55). Controller 20 uses the mixed content to drive the speaker (at box 56). In one aspect, the speaker may be a component of a local device, such as speaker 22. In another aspect, the speaker may be a component of a separate electronic device communicatively coupled to the local device, such as speaker 77 of audio output device 6.

[0101] Figure 9 This is a flowchart of one aspect of a process 60 for displaying closed captions representing audio content of media content. In one aspect, the process can be performed when local device 2 and one or more remote devices 3 are simultaneously engaged in a call and joint media playback session, as described herein. Process 60 begins with controller 20 receiving a downlink signal (at box 61). The controller receives an output from a VAD (e.g., VAD 42) indicating whether the downlink signal includes speech (at box 62). The controller determines whether the output from the VAD indicates that the downlink signal includes speech (at decision box 63). Specifically, the controller determines whether a user of a remote device has started (or has started) speaking. If so, the controller generates closed captions representing audio content contained within one or more audio signals of media content (at box 64). The controller then displays the closed captions (at box 65). Thus, in response to determining that a remote user is speaking, local device 2 displays the closed captions on display screen 25.

[0102] Figure 10This is a flowchart of one aspect of a process 70 for reversing playback of media content when it is determined that the downlink audio signal has stopped including speech. Process 70 begins with controller 20 determining a first timestamp along the playback duration of the media content, at which the output from the VAD begins to indicate that the downlink signal includes speech (in box 71). Controller 20 determines a second timestamp after the first timestamp along the playback duration of the media content, at which the output from the VAD indicates that the downlink signal has stopped including speech (in box 72). Specifically, the first timestamp may be determined in response to determining that the VAD signal generated by the VAD is at a high signal level, and the second timestamp may be determined in response to determining that the VAD signal changes from a high signal level to a low signal level. Controller 20 reverses the playback of the media content by pausing playback at or after the second timestamp and resuming playback of the media content from (or before) the first timestamp along the playback duration (in box 73).

[0103] Some aspects can be addressed in Figures 8 to 10 Processes 50, 60, and / or 70 described herein may be performed in variations. For example, at least some of the specific operations in these processes may not be performed in the exact order shown and described. The specific operation may not be performed in a consecutive series of operations, and different specific operations may be performed in different aspects. For example, in Figure 8 In this context, a joint media playback session can be initiated before the call is made. In this case, the local user can (e.g., in a media application such as...) Figure 3 and Figure 4 Within the media application's UI (displayed in the image), users select media content for playback and choose one or more remote devices (e.g., selecting contact information associated with the remote devices, such as phone numbers). Once the selection is complete, the local user can initiate playback by selecting the play button, for example... Figure 3 and Figure 4 As shown.

[0104] Furthermore, controller 20 may perform one or more operations in response to the detection of remote active voice. For example, when remote voice is detected to have started, controller 20 may perform the operations in processes 50 and 60 to reduce the volume level of the audio signal and display closed captions.

[0105] In one aspect, controller 20 may stop performing at least some of the operations described in processes 50, 60, and / or 70 in response to the output of a VAD indicating that the downlink signal does not include voice. For example, when the output of the VAD indicates that voice is not within the downlink signal, the controller may... Figure 8At box 54, the application of scalar gain to the audio signal is stopped. Therefore, the sound level of the media content can be restored to its previous level before the volume level was reduced (e.g., before the remote user's voice was detected). Similarly, once the remote voice is no longer determined to be active, the controller can stop generating and displaying closed captions at boxes 64 and 65.

[0106] In one aspect, the operation performed by the controller to maintain the audio quality of the media content based on the detection of remote active speech can be automatic (e.g., without user intervention). For example, the closed caption generator 48 can automatically generate and display captions based on the output of the VAD, as described in process 60. In another aspect, in response to receiving user authorization, at least some of the operations can be performed (e.g., adjusting the signal level of the audio signal by applying scalar gain, generating and displaying closed captions, and / or rewinding playback, etc.). Specifically, in response to determining that the output of the VAD indicates that the downlink signal has stopped including speech, the controller can provide a notification to the local user requesting authorization to perform at least one of the operations described herein. For example, when a second timestamp indicating that remote speech no longer exists is determined at box 72 of process 70, the controller can provide a notification to the user requesting authorization to rewind playback at box 73. In one aspect, the notification can be a pop-up notification displayed on display screen 25. Once authorization is received (e.g., by user selection of a UI item), the controller can perform at least one of the operations described herein. On the other hand, if (for example, within a certain period of time) no user authorization is received, the controller may abandon performing at least some of the operations described herein. For example, if no authorization to rewind playback is received, the controller may continue playing media content after that period of time.

[0107] As described herein, operations performed by the controller to maintain media quality during media content playback (e.g., application of scalar gain, generation and display of closed captions, and / or rewinding of media content playback, etc.) may be based on the presence of remote active voice during concurrent calls. Furthermore, at least some of these operations may be performed in response to the controller determining the presence of local active voice. For example, controller 20 may apply scalar gain to the audio signal in response to determining that the VAD output indicates 1) a microphone signal generated by a microphone from a local device or audio output device contains the voice of a local user and / or 2) an accelerometer signal generated by an accelerometer has an energy level indicating speech.

[0108] As described so far, the operations performed by the controller to maintain the audio quality of the media content may be in response to the detection of remote active speech and / or local active speech. In other words, these operations may be performed when a local or remote user is speaking. On the other hand, at least some of the operations for maintaining audio quality may be performed in response to the downlink signal level and / or the noise level of the microphone signal generated by a microphone (such as microphone 23) coupled to the local device exceeding a threshold level. Specifically, these operations may be performed when a loud sound is present at the remote or local device. Thus, for example, in response to the downlink signal or microphone signal exceeding the signal level, the controller may generate and display closed captions, as described in process 60. Furthermore, when the noise decreases (e.g., the signal level drops below the threshold), the controller 20 may rewind playback, as described in process 70.

[0109] When using audio output devices (such as wireless headphones) that connect wirelessly to media source devices, streaming content such as music and movies requires the source device to wirelessly transmit high-quality audio streams to the audio output device for output (e.g., to drive one or more speakers) to provide a good listener experience. For streaming high-quality audio, most wireless headphones establish a unidirectional wireless audio connection with the source device that supports high bit rates and sampling rates. For example, two devices can establish a Bluetooth connection using a wireless profile that provides high-quality audio, such as A2DP. A2DP allows stereo audio to be streamed from the source device to the wireless headphones using the SBC codec at sampling rates up to 48kHz.

[0110] When communicating with a source device that has initiated a call to another device and set up a joint media playback session for streaming media content, some audio output devices may not support high-quality audio. For example, to allow wireless communication between an audio output device and a source device, the two devices may establish a two-way wireless audio connection to exchange audio signals associated with the call. However, these two-way wireless audio connections only provide a low-quality audio stream to the audio output device. For example, both devices may use wireless profiles (such as HFP or HSP) that allow the exchange of audio data between multiple devices to establish a Bluetooth connection. These profiles only support “voice quality” or low-quality audio exchanged between the two devices. For example, HFPs traditionally use only codecs with sampling rates from 8kHz to 16kHz and are only capable of transmitting mono audio signals. While such a low-quality stream may be sufficient for pure voice communication, such wireless connections may not provide adequate audio quality when streaming media content in addition to making a call. However, in one respect, other audio output devices can be designed to support high-quality wireless audio transmission. For example, audio output devices can use wireless profiles with codecs having a higher sampling rate (e.g., 24kHz) to support “high-quality” two-way wireless audio connections. Therefore, switching between wireless audio connections is necessary when initiating a joint media playback session during a call based on the capabilities of the audio output device.

[0111] To overcome these shortcomings, this disclosure describes a method and audio system for switching wireless audio connections during a call. Specifically, the method can be performed by a local device 2 (e.g., using hands-free communication) communicatively coupled to an audio output device 6. For example, when participating in a call (e.g., a telephone call or video call) with a remote device, the local device communicates with the audio output device via a two-way wireless audio connection. The local device determines that a joint media playback session has been initiated, wherein the local device and the remote device independently stream media content for individual playback by both devices while participating in the call. Based on a determination of one or more capabilities of the audio output device (e.g., determining that the output device only supports low-quality audio streaming), the local device switches to communicating with a wireless headset via a one-way wireless audio connection, wherein a mixture of 1) one or more signals associated with the call and 2) audio signals of the media content is transmitted to the wireless headset via the one-way wireless audio connection. Thus, the audio output device can provide high-quality audio when participating in both a call and a joint media playback session.

[0112] Figure 11A block diagram according to one aspect is shown, wherein a local device 2 is communicatively coupled to an audio output device 6 via a two-way wireless audio connection to exchange audio data when the local device and a remote device 3 participate in a call together. Specifically, the diagram illustrates a local device communicating with an audio output device via a two-way wireless audio connection while participating in a (e.g., hands-free) call with a remote device to exchange audio data of the call between the local device and the audio output device. This is illustrated by the local device's deactivated microphone 23 (e.g., shown as a strikethrough) and the audio device's microphone 78 for capturing sound (e.g., as shown by sound waves). In one aspect, the diagram illustrates two devices before (or after) the joint media playback session has been initiated.

[0113] As shown in the figure, two devices are communicatively coupled via a bidirectional wireless audio connection 80 that allows the two devices to exchange audio data, as described herein. In one aspect, the bidirectional connection can be any type of wireless connection that allows the two devices to exchange audio data, such as an HFP connection. In another aspect, the bidirectional connection can be a “low-quality” bidirectional wireless audio connection (low-quality wireless connection) or a “high-quality” bidirectional wireless audio connection (high-quality wireless connection). In one aspect, a low-quality wireless connection can be designed to support mono audio and / or transmit audio streams at a sampling rate less than a threshold sampling rate (e.g., 24 kHz). In some aspects, a low-quality bidirectional connection can be a conventional HFP or HSP connection, as described herein. In some aspects, a high-quality audio connection can be designed to support stereo audio and / or transmit audio streams at a sampling rate at least the threshold sampling rate. In one aspect, a high-quality audio connection can be a Bluetooth connection using a wireless profile (e.g., an HFP) with a codec that transmits stereo audio streams at or above the threshold sampling rate.

[0114] In one aspect, the audio quality of a wireless connection can be based on the capabilities (or characteristics) of the audio output device (and / or the local device). For example, during the initiation of a two-way wireless audio connection, the audio output device can transmit device characteristics to the local device. In another aspect, these characteristics can indicate what type of wireless audio connection the audio output device can establish with the local device. For example, these characteristics can indicate which wireless profiles and / or audio codecs the audio output device supports. In yet another aspect, based on these characteristics, the local device can establish a two-way wireless audio connection.

[0115] For hands-free communication, both controllers 20 and 75 of the local device and the audio output device include one or more operation blocks. For example, controller 20 includes an audio call manager 46 and a voice DSP 41, while controller 75 includes an (optional) echo canceller 83. Controller 20 also includes a media playback manager 47, but this operation block is inactive (as shown with dashed boundaries) because neither device performs a joint media playback session.

[0116] As described herein, the audio call manager is configured to initiate (and perform) calls (e.g., by exchanging audio data of the call) between local device 2 and one or more remote devices 3. Specifically, the manager receives downlink audio signals from the remote devices and transmits microphone signals received from the audio output device as uplink audio signals to the remote devices. Voice DSP 41 is configured to receive the downlink audio signals from the audio call manager and is configured to perform audio signal processing (e.g., voice processing) operations on the signals to reduce (or eliminate) noise contained therein. As described herein, the voice DSP can apply noise reduction to the downlink audio signals associated with the call. The audio output device transmits the (processed) downlink audio signals to the audio output device via a two-way wireless audio connection 80 (via network interfaces 21 and 76) to drive speaker 77.

[0117] In one aspect, the audio output device may include an optional echo canceller 83 configured to receive a microphone signal captured by microphone 78 and configured to perform echo cancellation to remove linear echo from the microphone signal. Specifically, the canceller may determine a linear filter based on the transmission path between microphone 78 and speaker 77 and apply that filter to the downlink audio signal to generate an echo estimate to be subtracted from the microphone signal. In some aspects, the echo canceller may use any echo cancellation method. The (echo-cancelled) microphone signal is then transmitted via two-way wireless audio connection 80 to audio call manager 46 for transmission as an uplink audio signal to a remote device.

[0118] Figure 12 A block diagram is shown according to one aspect, wherein during a joint media playback session and a call with remote device 3, local device 2 is communicatively coupled to audio output device 6 via a two-way wireless audio connection. Specifically, the diagram illustrates the result of local device 2 initiating a joint media playback session, with both the local device and the audio output device participating in a hands-free call, such as... Figure 5As shown, the initiation of a playback session is illustrated by a media playback manager 47 that receives media content (e.g., as an audio signal) from a media content server 5. In one aspect, the diagram can be analogous to a local device coupled to an audio output device while simultaneously conducting hands-free calling and a combined media playback session. Figure 6 The diagram also shows that the controller includes one or more additional operational blocks, such as mixer 44, wireless audio connection switch decision logic 13, and scalar gain 86 (which is optional).

[0119] In one aspect, decision logic 13 is configured to determine whether to switch to a unidirectional wireless audio connection or (e.g., maintain) a bidirectional wireless audio connection to maximize the audio quality of media content and calls, thereby providing the best user experience. Specifically, the decision logic determines that a joint media playback session has been initiated by receiving a control signal from the joint media playback session manager indicating (e.g., will) to establish (e.g., a new) media session between the local device and one or more remote devices. In another aspect, the decision logic determines whether to switch based on the capabilities of the audio output device (e.g., capabilities that may have been received during the initialization of the bidirectional wireless audio connection 80), as described herein. For example, if the audio output device is determined not to support obtaining high-quality audio using a bidirectional connection (e.g., based on available audio codecs with a sampling rate below a threshold rate, as described herein), the decision logic may switch the wireless connection to a unidirectional connection. Figure 13a and Figure 13b More details are described regarding unidirectional connections. However, in this diagram, the decision logic has determined that the audio output device supports high-quality audio. In this case, the local device has established a (e.g., high-quality) bidirectional wireless audio connection 81 for streaming high-quality audio. In one aspect, this connection may be established when a hands-free call is initiated (e.g., in...). Figure 11 (In this case, once it is determined that the existing connection (e.g., between the local device and the audio output device during a hands-free call) provides high-quality audio, the local device can maintain a bidirectional connection with the audio output device. Therefore, connections 80 and 81 can be the same connection.)

[0120] On the other hand, instead of receiving features from the audio output device, decision logic 13 can retrieve one or more features based on the audio output device. Specifically, during the initialization of a hands-free call, the audio output device can transmit a device identifier to the local device. The decision logic can then use this identifier to perform a table lookup on the data structure that associates features with the device identifier.

[0121] In one aspect, when initiating a joint media playback session, the local device can determine whether to switch to a unidirectional wireless audio connection or (e.g., maintain) a bidirectional wireless audio connection to maximize the audio quality of the media content and the call, thereby providing the best user experience. In another aspect, this determination can be based on the capabilities of the audio output device, as described herein. For example, if the audio output device does not support obtaining high-quality audio through the use of a bidirectional connection (e.g., based on available audio codecs with a sampling rate below a threshold rate, as described herein), the local device can switch the wireless connection to a unidirectional connection. Figure 13a and Figure 13b More details are described regarding unidirectional connections. However, in this diagram, the local device has determined that the audio output device supports high-quality audio. In this case, the local device has established a (e.g., high-quality) bidirectional wireless audio connection 81 for streaming high-quality audio. In one aspect, this connection may be established when a hands-free call is initiated (e.g., in...). Figure 11 (In this case, once it is determined that the existing connection (e.g., between the local device and the audio output device during a hands-free call) provides high-quality audio, the local device can maintain a bidirectional connection with the audio output device. Therefore, connections 80 and 81 can be the same connection.)

[0122] In one aspect, during joint media playback sessions and calls, the local device may cease performing one or more operations and begin performing one or more audio processing operations on the downlink signal of the call and / or the audio signal of the media content. For example, controller 20 includes mixer 44 and scalar gain 86 (which is optional), wherein mixer 44 receives the audio signal of the media content from media playback manager 47 and the downlink audio signal from call manager 46, instead of voice DSP 41 receiving the downlink audio signal. In one aspect, the controller may cease performing voice DSP operations (e.g., cease applying noise reduction to the downlink audio signal) in response to a switch to communicate with the audio output device via a unidirectional connection, in order to provide a more complete spectral content of both the media content and the audio content of the downlink signal. As described herein, the mixer is configured to perform matrix mixing operations to generate mixed content of the signal. Scalar gain 86 is configured to receive the mixed content and is configured to apply scalar gain to the mixing to reduce the mixed signal level. In one aspect, the scalar gain may be applied for a period of time after initiating a joint media playback session (or after controller 20 switches to communicate with the audio output device via a one-way wireless audio connection). After this period of time, the scalar gain may be reduced (or removed) so that the gain is no longer applied to the mixed content. In another aspect, the scalar gain may be incrementally reduced for a second period of time to provide an attenuation effect. The mixed content is then transmitted via a two-way wireless audio connection 81 to the audio output device for driving speaker 77, as described herein.

[0123] Figure 13a and Figure 13b Several block diagrams are shown according to one aspect, wherein a local device 2, communicatively coupled to an audio output device 6 for exchanging audio data, switches between wireless audio connections based on the initiation of a joint media playback session. Specifically, Figure 13a A block diagram is shown in which a local device and an audio output device are coupled via a unidirectional wireless audio connection 85. Specifically, the diagram illustrates the result of local device 2 initiating a joint media playback session when participating in a call. However, this differs from a diagram in which a bidirectional wireless audio connection is maintained between the local device and the audio output device. Figure 12 The diagram shows that the local device has switched to a one-way wireless audio connection 85 in order to stream high-quality audio data to an audio output device for output (e.g., via speaker 77).

[0124] In one aspect, switching (or conversion) from a bidirectional connection to a unidirectional wireless audio connection can be based on the audio output device, as described herein. For example, decision logic 13 can determine (e.g., in response to receiving a control signal from session manager 47) that the audio output device does not support exchanging audio signals via the bidirectional wireless audio connection at a sampling rate of at least a threshold sampling rate. As described herein, this determination can be based on characteristics received from the audio output device or on performing a table lookup on a data structure using a device identifier. In another aspect, the decision logic can determine to switch to a unidirectional wireless audio connection based on the absence of characteristics received from the device and / or the failure to identify the device within the data structure (e.g., the decision to switch could be the default decision of the decision logic).

[0125] In one aspect, local device 2 and audio output device can perform one or more operations to switch from a two-way connection 80 to a one-way wireless audio connection 85. For example, local device 2 (or audio output device 6) can disconnect (or terminate) the two-way wireless audio connection 80. Once disconnected, the local device can establish a one-way wireless audio connection with the audio output device (e.g., a Bluetooth A2DP connection). In one aspect, since the two-way connection is disconnected for a one-way connection in which audio data can only be transmitted from the local device to the audio output device, the controller can be configured to activate one or more other microphones to capture local user voice for uplink audio signals. Specifically, the controller can signal to the audio output device to mute microphone 78 (as indicated by strikethrough) and can activate microphone 23 of the local device to capture local user voice. In one aspect, the activated microphone can be a component of a different electronic device. Thus, the microphone signal of microphone 23 can be transmitted as an uplink audio signal to a remote device. More details are described herein regarding the switching of wireless audio connections performed by the controller.

[0126] In one aspect, controller 20 may (optionally) perform an echo cancellation estimation operation on the microphone signal generated by microphone 23. Specifically, controller 20 includes an echo cancellation estimator 87 configured to perform an echo cancellation operation to remove echo from the microphone signal. In one aspect, the estimator may perform an echo cancellation estimation operation on the microphone signal generated by microphone 23. Figure 11 The estimator operates similarly to the canceller 83 described herein. For example, when both devices are involved in a call, the estimator can obtain the microphone signal from the local device that will be transmitted to the remote device. The estimator is configured to generate an estimate of a portion of one or more signals (e.g., downlink audio) associated with the call. For example, the estimator can determine a linear filter based on the transmission path between microphone 23 and speaker 77. In one aspect, unlike the definable transmission path between microphone 78 and speaker 77 (e.g., based on both the microphone and speaker being integrated into the audio output device at predefined locations), the transmission path between the local microphone 23 and the speaker 77 of the audio output device may not be predefined. Therefore, the estimator can estimate the transmission path. For example, the estimator can determine the distance between microphone 23 and speaker 77 based on the arrival time of the sound captured by microphone 23 and generated by speaker 77. In another aspect, the estimator can estimate the path based on the received signal strength (RSSI) of the wireless audio connection. In some aspects, the estimator can use any sound localization method to determine the location of speaker 77 and thus determine the path from speaker to microphone. In another aspect, the transmission path may be predefined (e.g., a path determined in a controlled environment such as a laboratory). Using the transmission path estimation, a linear filter is determined, which is applied to the downlink audio signal to generate an echo estimate that will be subtracted from the microphone signal, as described in this paper.

[0127] In one aspect, the wireless audio connection switching decision logic 13 can be configured to switch between a unidirectional wireless audio connection 85 and a bidirectional wireless audio connection during joint media playback sessions and calls. In another aspect, the decision logic can switch to a high-quality bidirectional wireless audio connection (e.g., Figure 12 (Connection 81 in the document). On the other hand, when the audio output device does not support a high-quality two-way wireless audio connection, the decision logic may switch the unidirectional wireless audio connection to a low-quality two-way wireless audio connection to provide hands-free communication with the audio output device, as described herein. Although less preferred than a unidirectional wireless audio connection due to its lower audio quality, this functionality may be required or needed in some cases based on one or more standards. Figure 13b The switching to a low-quality bidirectional connection is described.

[0128] In one aspect, switching to a two-way wireless audio connection may be based on the location of the local device 2 and / or the audio output device 6. For example, as described herein, when switching to a one-way wireless audio connection, the location of the microphone used during a call and before initiating a joint media playback session may be at the audio output device, which may be a wireless headset worn on the user's head. However, once a one-way connection is initiated, the location of the microphone (e.g., active) may change to a different microphone (e.g., microphone 23 of the local device) that may be separated from the audio output device. Thus, the microphone and speaker used during the call and joint media playback session may be components of different electronic devices, each located in a different location. Therefore, in order to participate in the call and joint session, the local user may be required to bring the local device and the audio output device close together (e.g., so that the microphone can capture the user's voice and so that the user can hear the sound produced by the speaker of the audio output device). In one aspect, decision logic may receive sensor data from one or more sensors 40 and may be configured to determine whether the local device and the audio output device are separated by a threshold distance. For example, the decision logic may receive image data from one or more cameras (e.g., camera 24) and use the image data to determine the location of the audio output device by using an image recognition algorithm. On the other hand, decision logic can determine the location of the audio output device based on the RSSI of a unidirectional connection. For example, in response to determining that the RSSI is below a threshold, the decision logic can execute a switch to a bidirectional connection. This is because the user may be too far from the new active microphone to clearly pick up the local user's voice.

[0129] On the other hand, the decision can be based on whether the local user is in front of (or next to) the display screen 25 of the local device. For example, a camera 24 may be positioned adjacent to the display screen and have a field of view in front of the display screen. The decision logic may receive image data from the camera and perform an image recognition algorithm to determine whether the user is present (e.g., in front of the display screen). If not, the decision logic may perform a switch. In some aspects, the decision logic may make this determination based on other sensor data, such as proximity sensor data. In this case, one or more proximity sensors may be arranged to determine whether an object is within a threshold distance from the display screen 25. If not, indicating that the local user is not in front of the display screen, the decision logic may perform a switch.

[0130] On the other hand, decision logic 13 can perform a switch based on whether an object is within a threshold distance of a local device (e.g., the microphone 23 of the local device). For example, when the local device is a smartphone, the user may place the smartphone in their pocket. In this case, the microphone may capture a muffled user voice. Thus, the decision logic can receive sensor data indicating whether an object is within the threshold distance. For example, the sensor could be a proximity sensor. In response to the object being within that distance, the decision logic can perform a switch.

[0131] In some respects, the decision logic may perform a switch based on whether the local user is speaking. For example, when the local user is not speaking, a microphone may not be necessary, so a one-way wireless connection can be established to provide high-quality audio. However, in response to determining that the local user is speaking, the decision logic may perform a switch. For example, the decision logic may receive a control signal from the audio output device in response to the local user speaking, and may perform a switch based on the received control signal. For example, when the control signal is a VAD signal generated by the VAD 82 of the audio output device in response to the detection of a high energy level of the accelerometer signal from the accelerometer 79, the decision logic may determine that the local user is speaking. On the other hand, the switch may be performed from the VAD of the local device (e.g., as shown in the image). Figure 5 The VAD 42 shown receives a VAD signal, which can be configured to detect the local user's speech based on signals received from the audio output device, such as one or more accelerometer signals and / or one or more microphone signals. Once the user speaks, the decision logic can switch to a two-way wireless audio connection and activate the microphone 78 of the output device to capture the user's speech. Once the user has finished speaking (e.g., the VAD signal indicates that the user's speech is no longer detected), the decision logic can switch back to a one-way audio connection.

[0132] Figure 13b A block diagram is shown in which the local device and audio output device have switched to a two-way wireless audio connection while conducting a joint media playback session and a call, as described herein. Specifically, the diagram illustrates the result of decision logic 13 switching to a two-way wireless audio connection (e.g., based on one or more standards) during a call and playback session. As shown, the two-way wireless audio connection 89 is a low-quality connection, which may be due to the fact that the audio output device does not support high-quality connections, as described herein. In addition to switching to a two-way connection, the local device and audio output device have restored the microphone's (active) position from the local device to the audio output device.

[0133] like Figure 12 , Figure 13a and Figure 13bThe local device may participate in a joint media playback session in which one or more audio signals of media content (e.g., a musical work) are received for playback. In one aspect, the operations performed in these figures may occur while the local device participates in a joint playback session in which multimedia content is being played back, such as when video is displayed on display 25 and audio is output by speaker 77. Furthermore, controller 20 and / or controller 75 may also perform at least some of the other operations described herein.

[0134] Figures 14 to 18 The flowcharts are for processes 90, 100, 110, 130, and 120, which perform one or more operations during a call to switch the wireless audio connection. In one aspect, at least some of the processes may be performed by one or more devices of the audio system 1, such as... Figure 1 As shown in the diagram. For example, at least processes 90, 100, and 110 are executed by local device 2 (e.g., its controller 20), and processes 130 and 120 are executed by audio output device 6 (e.g., its controller 75). On the other hand, any of the devices can perform any of the operations described herein.

[0135] Figure 14This is a flowchart of one aspect of a process 90 for switching between wireless audio connections. In one aspect, this process may be performed by a controller 20 of local device 2. Process 90 begins with the controller initiating a call between the local device and a remote device (at block 91). For example, a call manager 46 may initiate a call (e.g., a telephone or video call) between the local device and one or more remote devices, as described herein. When participating in a call with a remote device, the controller 20 communicates with the audio output device via a two-way wireless audio connection (at block 92). Specifically, local device 2 may establish a wireless connection with the audio output device via a wireless communication link (e.g., via Bluetooth or any other wireless communication protocol). For example, the local device may communicate with the audio output device to configure a Bluetooth stack, which executes within the audio output device, to exchange audio data between the devices via the two-way wireless audio connection (e.g., by negotiating codecs for decoding and encoding audio signals exchanged between the devices). During this time, the audio output device may transmit messages indicating its capabilities (e.g., the audio codecs it supports, etc.). In one aspect, based on these capabilities, the local device may establish a two-way wireless audio connection. Specifically, if high-quality audio streaming (e.g., at a sampling rate of at least a threshold sampling rate) can be supported, the local device can establish a high-quality bidirectional wireless audio connection, as described herein. Once established, the local device can transmit one or more signals associated with the call (e.g., downlink audio) to the audio output device and receive one or more microphone signals used for the call via the bidirectional connection. On the other hand, the device can establish a low-quality wireless audio connection regardless of the capabilities of the audio output device, since only voice data is exchanged between the devices.

[0136] Controller 20 determines that a joint media playback session has been initiated, in which the local and remote devices independently stream media content for separate playback by each device during a call (at box 93). Specifically, the joint media playback session manager 47 may have received a user request from the local user (e.g., via a UI displayed on display 25), or may have received a request from the media content server 5 indicating that one or more remote devices have requested to initiate a playback session.

[0137] Controller 20 determines whether the audio output device supports exchanging audio signals for calls and media content with a local device via (e.g., high-quality) two-way wireless audio connections. (At decision box 94) Specifically, wireless audio connection switching decision logic 13 may, for example, switch from (e.g., currently established) two-way wireless audio connections to one-way wireless audio connections based on one or more capabilities of the audio output device 6. For example, the decision logic may determine whether the audio output device supports high-quality audio based on a table lookup performed on a data structure that associates characteristics with device identifiers. In one aspect, since a two-way wireless audio connection has been established, the decision logic may determine the type of connection that already exists between the two devices (e.g., whether the connection is an HFP connection using a codec with a sampling rate above a threshold rate and / or whether the HFP connection supports stereo audio). If so, the controller communicates with the audio output device via (e.g., high-quality) two-way wireless audio connections when participating in a call and during a joint media playback session (at box 95). In one aspect, if the initial wireless audio connection is a low-quality connection, the controller may disconnect that connection and establish a high-quality two-way wireless audio connection. However, if the initially established two-way wireless audio connection is a high-quality connection, the controller can maintain the existing connection.

[0138] However, if the audio output device does not support a high-quality two-way wireless audio connection, controller 20 switches to communicating with the audio output device via a one-way wireless audio connection (e.g., based on one or more capabilities of the audio output device, as described herein), wherein a mixture of audio signals of one or more signals associated with the call and the media content is transmitted to the audio output device via the one-way wireless audio connection (at box 96). Specifically, as described herein, controller 20 may disconnect the two-way wireless audio connection and establish a one-way connection. Once established, the controller can stream the media content and the downlink audio signal of the call to the audio output device for playback. Figure 15 More details are described regarding the operation for switching wireless audio connections.

[0139] Figure 15 This is a flowchart of another aspect of the process 100 for switching between wireless audio connections. In one aspect, at least some of the operations performed in process 100 may be performed by controller 20 when (and / or after) switching to communication with an audio output device via a unidirectional wireless audio connection, such as... Figure 14As described in box 96. Process 100 begins with the controller transmitting a signal to mute the microphone (e.g., microphone 78) of the audio output device (at box 101). Specifically, the controller may transmit a control signal to the audio output device via a two-way wireless audio connection, causing controller 75 to mute microphone 78. In one aspect, muting controller 75 can be achieved by stopping the transmission of microphone signals generated by the microphone to the local device. In this case, microphone 78 may continue to generate microphone signals, which controller 75 can use to perform one or more operations (e.g., perform ANC functions, transparency functions, etc.). Controller 20 switches from a two-way wireless audio connection to a one-way wireless audio connection (at box 102). As described herein, a one-way wireless audio connection can be any wireless connection that provides high-quality audio (e.g., an A2DP connection). In one aspect, the one-way connection may be based on the capabilities of the audio output device.

[0140] Controller 20 provides notifications (at box 103) indicating that the microphone of the audio output device is muted and / or requesting user authorization to activate a different microphone. For example, the controller may display the notification as a pop-up notification on the display screen 25 of the local device 2, thereby alerting the local user that the microphone is muted. In one aspect, this will remind the user not to begin speaking until the microphone is activated. In some aspects, the notification may also indicate the new location of the microphone. Specifically, the notification may indicate that the microphone is located at the local device. In one aspect, the notification may also request user authorization to activate a different microphone (e.g., by displaying a UI item within the pop-up notification).

[0141] Controller 20 begins playback of media from the joint media playback session (at box 104). Specifically, controller 20 may begin transmitting one or more audio signals of the media content to an audio output device via a unidirectional connection, which can use these signals to drive one or more speakers. Additionally, when the media content includes video, the controller may display the video signal on display screen 25. The controller determines whether the user has authorized microphone switching (at decision box 105). For example, the controller may determine whether the user has selected a UI item displayed in a pop-up notification. If not, the controller may continue playback of the media content without any microphone on the local device and / or audio output device being active to capture the user's voice for the uplink signal used for the call. However, if the controller receives user authorization, it activates a different microphone and begins receiving microphone signals for transmission to a remote device (e.g., as an uplink signal) for the call (at box 106).

[0142] In one respect, the controller can provide the user with a choice of which microphone to activate for a call. For example, a pop-up notification could display a list of microphones and their locations, allowing the local user to decide which microphone to use during the call. In another respect, the controller can provide the user with the option to continue communicating with the audio output device via a two-way wireless audio connection. For example, the controller could provide a notification requesting user authorization to switch from a two-way wireless audio connection to a one-way wireless audio connection. If the user fails to provide a response (and / or fails to provide authorization by selecting a UI item), the controller can continue communication within the two-way wireless audio connection, which, as described herein, can be a low-quality connection based on the capabilities of the audio output device.

[0143] Figure 16 This is a flowchart of one aspect of a process 110 for determining whether to switch between wireless audio connections based on one or more criteria. Specifically, the process is used to determine whether to switch from communicating with an audio output device via a unidirectional wireless audio connection to communicating with the device via (e.g., a low-quality) bidirectional wireless audio connection. Process 110 begins with controller 20 communicating with the audio output device via a unidirectional wireless audio connection, for example during a call and joint media playback session, as described herein (at box 111). Controller 20 receives sensor data from at least one sensor (at box 112). For example, the controller may receive sensor data from a proximity sensor, a light sensor, a microphone (e.g., microphone 23), a camera (e.g., camera 24), etc. Controller 20 determines, based on the sensor data, whether to switch to communicating with the audio output device via a bidirectional wireless audio connection (at decision box 113). As described herein, the controller may use sensor data, such as proximity data from a proximity sensor, to determine whether an object is within a threshold distance. In response to being within the threshold distance, controller 20 switches to communicating with the audio output device via a bidirectional wireless audio connection (at box 114). As described in this article, bidirectional connections can be low-quality (e.g., traditional 8kHz HFP) connections, depending on the capabilities of the audio output device.

[0144] However, if the controller determines not to switch based on sensor data, it then determines whether the local device has received a user request to switch to a two-way wireless audio connection (at decision box 115). For example, the local device may display a UI item on display 25 allowing the local user to switch to a two-way wireless audio connection. On the other hand, for various reasons, a user may wish to switch to a two-way connection. For example, when the user's environment is noisy, the user may wish to use the onboard microphone of the audio output device. If so, the controller continues to switch connections.

[0145] If not, the controller determines the signal strength of the unidirectional wireless audio connection (at box 116). For example, the controller may determine the RSSI of the connection. The controller determines whether the signal strength is above a threshold (at decision box 117). If not, the controller may continue to switch connections. In one aspect, the signal strength may be low because the user moves away from the local device while continuing to wear the audio output device. For example, when the local device is a desktop computer with an onboard microphone for picking up the user's voice for a call, the controller may perform a switch to keep the active microphone within the user's distance if the user moves away. If the signal strength is above the threshold, the controller may continue to communicate with the audio output device via the unidirectional wireless audio connection (at box 118).

[0146] In one aspect, the controller may switch back to a unidirectional wireless audio connection when at least one of the conditions that cause the controller to switch ends. For example, when communicating with an audio output device via a bidirectional wireless audio connection, the controller may switch back to a unidirectional wireless audio connection when it is determined that the signal strength is above a threshold. Continuing the previous example, when the signal strength is above the threshold, it can be determined that the user is now in front of the desktop computer.

[0147] Figure 17This is a flowchart of one aspect of a process 130 performed by an audio output device 6 (e.g., its controller 75) for switching between wireless audio connections. Process 130 begins with the controller 75 communicating with the local device via a two-way wireless audio connection during a call between the local device 2 and the remote device 3 (at block 131). For example, the audio output device may perform hands-free communication with the local device during a call, as described herein. The controller 75 determines that a one-way wireless audio connection will be established between the local device and the audio output device instead of a two-way wireless audio connection during a call (at block 132). For example, this determination may be based on whether a two-way connection can support high audio quality. In one aspect, an existing two-way connection may support exchanging audio signals at a sampling rate lower than that supported by a one-way connection. For example, a two-way connection may be an HFP connection supporting sampling rates from 8 kHz to 16 kHz, while a one-way connection may be an A2DP connection supporting a sampling rate of 48 kHz. In one aspect, the audio output device may (e.g., from the local device) receive a control signal indicating that the two-way wireless audio connection will be disconnected. The controller 75 mutes the microphone of the audio output device (at block 133). As described herein, controller 75 may disable the microphone and / or stop transmitting microphone signals to the local device. Controller 75 switches from a two-way wireless audio connection to a one-way wireless audio connection (at box 134). For example, the audio output device may disconnect the two-way connection and transmit an acknowledgment message indicating that the connection has been disconnected to the local device. Subsequently, the audio output device may receive communication from the local device to establish a two-way wireless audio connection. In response, the audio output device may establish a connection. Controller 75 receives audio signals via the one-way wireless audio connection, which include a mixture of signals associated with a call and signals associated with media content played back by the local and remote devices in a joint media playback session (at box 135). The controller may then use the audio signals to drive the speaker (e.g., speaker 77) of the audio output device (at box 136).

[0148] Figure 18This is a flowchart of one aspect of a process 120 performed by audio output device 6 to switch from a one-way wireless audio connection to a two-way wireless audio connection based on whether voice is detected. In one aspect, prior to performing process 120, audio output device 6 may be communicatively coupled to a local device via a one-way connection to receive audio data of media content played back by the local device during a joint media playback session concurrent with a call, as described herein. For example, the audio output device may receive audio signals via a one-way connection, which include a mixture of 1) signals from a telephone (or video) call and 2) signals associated with media content, wherein the local device and the remote device are simultaneously involved in the call and joint media playback session. Furthermore, the audio output device may use the audio signals to drive a speaker. Process 120 begins with controller 75 receiving accelerometer signals (e.g., accelerometer 79) from the audio output device (at block 121). Controller 75 generates a VAD signal (e.g., as the output of VAD 82) based on the accelerometer signals (at block 122). As described herein, the VAD signal may indicate that user voice has been detected based on the energy level of the accelerometer. Controller 75 determines whether the VAD signal is above a threshold, thereby indicating that user voice has been detected (at decision box 123). If not, the audio output device continues to communicate with the local device via a one-way wireless audio connection (at box 124).

[0149] Otherwise, controller 75 switches to communicating with the local device via a two-way wireless audio connection (at box 125). Controller 75 receives microphone signals from the microphone of the audio output device (at box 126). Controller 75 then transmits the microphone signals to the local device via the two-way wireless audio connection as uplink signals to the remote device, as described herein (at box 127).

[0150] Some aspects can be addressed in Figures 14 to 18 Processes 90, 100, 110, 130, and 120 described herein are performed in variations. For example, at least some of the specific operations in these processes may not be performed in the exact order shown and described. The specific operation may not be performed in a consecutive series of operations, and different specific operations may be performed in different aspects. For example, the operation within the dashed box may be an optional operation that may not be performed when the corresponding process is executed. For example, in Figure 15 During process 100, no notification is required. Instead, playback of media content can begin (at box 104), and different microphones can be activated in response to a connection switch (at box 106).

[0151] As is widely recognized, the use of personally identifiable information should comply with privacy policies and practices that are generally accepted to meet or exceed industry or governmental requirements for protecting user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use, and the nature of authorized use should be clearly explained to users.

[0152] As previously described, one aspect of this disclosure may be a non-transitory machine-readable medium (such as microelectronic memory) storing instructions thereon that program one or more data processing units (generally referred to herein as a "processor") to perform network operations and audio signal processing operations, as described herein. In other aspects, some of these operations may be performed by specific hardware components containing hard-wired logic. Alternatively, those operations may be performed by any combination of programmed data processing units and fixed hard-wired circuit components.

[0153] While certain aspects have been described and illustrated in the accompanying drawings, it should be understood that such aspects are merely illustrative of the broad disclosure and not limiting, and that this disclosure is not limited to the specific structures and arrangements shown and described, as various other modifications will be apparent to those skilled in the art. Therefore, the description is to be regarded as exemplary and not restrictive.

[0154] In some aspects, this disclosure may include the language "[element A] and [element B] at least one". This language may refer to one or more of these elements. For example, "at least one of A and B" may refer to "A", "B", or "A and B". Specifically, "at least one of A and B" may refer to "at least one of A and at least one of B" or "at least either A or B". In some aspects, this disclosure may include the language "[element A], [element B], and / or [element C]". This language may refer to any of these elements or any combination thereof. For example, "A, B, and / or C" may refer to "A", "B", "C", "A and B", "A and C", "B and C", or "A, B, and C".

Claims

1. A method comprising: During a call between the first electronic device and the second electronic device, communication is made with the first electronic device via a one-way audio connection at the wireless headset; Voice of the user of the wireless headphones was detected at the wireless headphones. as well as At the wireless headset, in response to the detection of voice, communication is conducted with the first electronic device via a two-way audio connection by transmitting the user's voice and receiving sound, wherein the sound includes the telephone content of the call and the audio content of a joint media playback session between the first electronic device and the second electronic device.

2. The method of claim 1, further comprising: Receive accelerometer signals from the accelerometer of a wireless headset; as well as Voice activation detection (VAD) signals are generated based on accelerometer signals. The detection of a user's voice using wireless headphones includes determining if the VAD signal is greater than a threshold.

3. The method of claim 1, wherein communicating with the first electronic device via a bidirectional audio connection includes switching from communicating with the first electronic device via the unidirectional audio connection to communicating with the first electronic device via the bidirectional audio connection.

4. The method of claim 1, further comprising: In response to communicating with a first electronic device via a two-way audio connection, the microphone of the wireless headset is activated to capture speech as a microphone signal, wherein transmitting the user's speech includes transmitting the microphone signal via the two-way audio connection.

5. The method of claim 4, further comprising at least one of the following: Perform echo cancellation to remove linear echoes from the microphone signal; or Apply noise reduction to the microphone signal.

6. The method as described in claim 1, Communication via a one-way audio connection includes transmitting the user's voice, captured by a microphone, at a first sampling rate. Communication via a two-way audio connection includes exchanging speech and sound at a second sampling rate lower than the first sampling rate.

7. The method of claim 1, wherein the unidirectional audio connection includes an Advanced Audio Distribution Profile (A2DP) connection, and the bidirectional audio connection includes a Hands-Free Mode (HFP) connection or a Headphone Mode (HSP) connection.

8. A wireless headset, comprising: At least one processor; as well as A memory storing instructions that, when executed by the at least one processor, cause the wireless headset to: During a call between the first electronic device and the second electronic device, communication is made with the first electronic device via a one-way audio connection; Voice input from the user of the wireless headphones was detected. as well as In response to the detection of voice, communication is made with a first electronic device via a two-way audio connection by transmitting the user's voice and receiving sound, wherein the sound includes the telephone content of the call and the audio content of a joint media playback session between the first electronic device and a second electronic device.

9. The wireless headset of claim 8, further comprising an accelerometer, wherein the memory includes further instructions to: Receive accelerometer signals from the accelerometer of a wireless headset; and Voice activation detection (VAD) signals are generated based on accelerometer signals. The method involves detecting the user's voice in wireless headphones by determining that the VAD signal is greater than a threshold.

10. The wireless headset of claim 8, wherein the instructions for communicating with the first electronic device via a bidirectional audio connection include instructions for switching from communicating with the first electronic device via the unidirectional audio connection to communicating with the first electronic device via the bidirectional audio connection.

11. The wireless headset of claim 8, wherein the memory includes further instructions to: activate the microphone of the wireless headset to capture speech as a microphone signal in response to communication with the first electronic device via a two-way audio connection, wherein the user's speech is transmitted by transmitting the microphone signal via the two-way audio connection.

12. The wireless headset of claim 11, wherein the memory includes further instructions to perform at least one of the following: Perform echo cancellation to remove linear echoes from the microphone signal; or Apply noise reduction to the microphone signal.

13. The wireless headphones as described in claim 8, The instructions for communication via a one-way audio connection include instructions for transmitting the user's voice captured by the microphone at a first sampling rate. The instructions for communication via the bidirectional audio connection include instructions for exchanging speech and sound at a second sampling rate lower than the first sampling rate.

14. The wireless headset of claim 13, wherein the second sampling rate is 24 kHz.

15. A non-transitory machine-readable medium having instructions that, when executed by at least one processor of a wireless headset, cause the wireless headset to: During a call between the first electronic device and the second electronic device, communication is made with the first electronic device via a one-way audio connection; Voice input from the user of the wireless headphones was detected. as well as In response to the detection of voice, communication is made with a first electronic device via a two-way audio connection by transmitting the user's voice and receiving sound, wherein the sound includes the telephone content of the call and the audio content of a joint media playback session between the first electronic device and a second electronic device.

16. The non-transitory machine-readable medium of claim 15, further comprising the following instructions: Receive accelerometer signals from the accelerometer of a wireless headset; and Voice activation detection (VAD) signals are generated based on accelerometer signals. The method involves detecting the user's voice in wireless headphones by determining that the VAD signal is greater than a threshold.

17. The non-transitory machine-readable medium of claim 15, wherein the instructions for communicating with the first electronic device via a bidirectional audio connection include instructions for switching from communicating with the first electronic device via the unidirectional audio connection to communicating with the first electronic device via the bidirectional audio connection.

18. The non-transitory machine-readable medium of claim 15, further comprising instructions to: activate the microphone of a wireless headset to capture speech as a microphone signal in response to communication with a first electronic device via a bidirectional audio connection, wherein the user's speech is transmitted by transmitting the microphone signal via the bidirectional audio connection.

19. The non-transitory machine-readable medium of claim 18, further comprising instructions to perform at least one of the following: Perform echo cancellation to remove linear echoes from the microphone signal; or Apply noise reduction to the microphone signal.

20. The non-transitory machine-readable medium as described in claim 15, The instructions for communication via a one-way audio connection include instructions for transmitting the user's voice captured by the microphone at a first sampling rate. The instructions for communication via the bidirectional audio connection include instructions for exchanging speech and sound at a second sampling rate lower than the first sampling rate.

21. A method comprising: It is determined that a joint media playback session will be executed, in which the first electronic device and the second electronic device will independently stream media content for separate audio playback of the two devices when participating in a call, while the first electronic device communicates with the wireless headset via a first bidirectional wireless audio connection including a first sampling rate; Determine whether the wireless headphones support exchanging audio signals with the first electronic device via a second bidirectional wireless audio connection including a second sampling rate greater than the first sampling rate; as well as In response to the wireless headphones supporting the exchange of audio signals via a second bidirectional wireless audio connection, communication with the wireless headphones is conducted via the second bidirectional wireless audio connection instead of the first bidirectional wireless audio connection.

22. The method of claim 21, wherein the second sampling rate comprises a sampling rate of at least 24 kHz.

23. The method of claim 21, wherein the first bidirectional wireless audio connection is designed to support mono audio transmission, and the second bidirectional wireless audio connection is designed to support stereo audio transmission, wherein the communication includes transmitting at least two stereo channels via the second bidirectional wireless audio connection to drive at least two speakers of the wireless headphones.

24. The method of claim 21, wherein determining whether to receive one or more device characteristics of the wireless headphones during the initialization of a first bidirectional wireless audio connection between the first electronic device and the wireless headphones.

25. The method of claim 21, further comprising: In response to the fact that the wireless headphones do not support the exchange of audio signals via a second bidirectional wireless audio connection, communication with the wireless headphones is conducted via a unidirectional wireless audio connection instead of the first bidirectional wireless audio connection.

26. The method of claim 25, further comprising: In response to the fact that the wireless headset does not support the exchange of audio signals via a second bidirectional wireless audio connection, the microphone of the first electronic device is activated during the call to capture the user's voice for transmission to the second electronic device.

27. The method of claim 26, further comprising: In response to determining that the user is moving away from the first electronic device, Switch to communicating with the wireless headphones via the first two-way wireless audio connection; as well as The user's voice is captured in response to the call using the onboard microphone of a wireless headset.

28. A first electronic device, comprising: At least one processor; as well as A memory storing instructions that, when executed by the at least one processor, cause the first electronic device to: It is determined that a joint media playback session will be executed, in which the first electronic device and the second electronic device will independently stream media content for separate audio playback of the two devices when participating in a call, while the first electronic device communicates with the wireless headset via a first bidirectional wireless audio connection including a first sampling rate; Determine whether the wireless headphones support exchanging audio signals with the first electronic device via a second bidirectional wireless audio connection including a second sampling rate greater than the first sampling rate; as well as In response to the wireless headphones supporting the exchange of audio signals via a second bidirectional wireless audio connection, communication with the wireless headphones is conducted via the second bidirectional wireless audio connection instead of the first bidirectional wireless audio connection.

29. The first electronic device of claim 28, wherein the second sampling rate includes a sampling rate of at least 24 kHz.

30. The first electronic device of claim 28, wherein the first bidirectional wireless audio connection is designed to support mono audio transmission, and the second bidirectional wireless audio connection is designed to support stereo audio transmission, wherein the instructions for communication include instructions for transmitting at least two stereo channels via the second bidirectional wireless audio connection to drive at least two speakers of a wireless headset.

31. The first electronic device of claim 28, wherein the instruction for determining whether includes instructions for receiving one or more device characteristics of the wireless headphones during the initialization of a first bidirectional wireless audio connection between the first electronic device and the wireless headphones.

32. The first electronic device of claim 28, wherein the memory includes further instructions to: in response to the wireless headphones not supporting the exchange of audio signals via a second bidirectional wireless audio connection, communicate with the wireless headphones via a unidirectional wireless audio connection instead of the first bidirectional wireless audio connection.

33. The first electronic device of claim 32, wherein the memory includes further instructions to: in response to the wireless headset not supporting the exchange of audio signals via a second bidirectional wireless audio connection, activate the microphone of the first electronic device during the call to capture the user's voice for transmission to the second electronic device.

34. The first electronic device of claim 33, wherein the memory includes further instructions to: in response to determining that a user is moving away from the first electronic device, Switching to communicate with wireless headphones via a first two-way wireless audio connection; and The user's voice is captured in response to the call using the onboard microphone of a wireless headset.