Method and system for processing remote active voice during a call

By using a voice activity detector to detect voice activity on the remote device during a call and adjusting the audio signal processing of the media content, the problem of voice activity interfering with media playback is solved, improving the user experience.

CN115348411BActive Publication Date: 2025-09-23APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210521384.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-15
Filing Date
2022-05-13
Publication Date
2025-09-23
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

During a call, voice activity on the remote device may mask or interfere with the playback of media content, resulting in a degraded user experience. This is especially true in video calls or video conferences, where conversations may drown out or mask the sound of the media content, affecting the user's media experience.

Method used

By using a voice activity detector (VAD) in the local device to detect whether the downlink signal of the remote device contains voice, the signal level of the audio signal of the media content is adjusted based on the detection result, and the playback of the media content is paused or adjusted when necessary, such as by lowering the volume or displaying closed captions.

Benefits of technology

It effectively reduces the interference of voice activities on media content playback, improves the user's media experience quality during calls, ensures that media content is not covered or overwhelmed by voice activities, and enhances the user's concentration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115348411B_ABST
    Figure CN115348411B_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods and systems for processing remote active voice during a call. The present invention provides a method performed by a first device, the method comprising: performing an audio call with a second device by transmitting a microphone signal as an uplink signal and receiving a downlink signal for driving a first speaker and, while performing the audio call, performing a joint media playback session, wherein the two devices independently stream media content segments for synchronized playback such that the two devices simultaneously receive audio signals of the media content segments for driving respective speakers; in response to determining that a voice activity detection (VAD) signal indicates that the downlink signal includes voice, determining that the VAD signal indicates that the downlink signal includes voice; processing the audio signal of the media content segment by applying a scalar gain; and driving the first speaker with a mixture of the downlink signal and the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present disclosure relates to methods and systems for processing remote active voice during a call. Other aspects are also described. Background Art

[0002] Many devices today, such as smartphones, are capable of performing various types of telecommunication activities with other devices. For example, a smartphone can initiate a phone call with another device. In this case, when a phone number is dialed, the smartphone connects to a cellular network, which then connects the smartphone to the other device (e.g., another smartphone or a landline). Furthermore, smartphones are also capable of conducting video conference calls, in which video and audio data are exchanged with another device. Summary of the Invention

[0003] One aspect of the present disclosure is a method performed by a first electronic device (e.g., a local device) that is communicatively coupled to an audio output device (such as a wireless headset or a head-mounted device including at least one speaker). For example, the first electronic device may initiate a call (e.g., a voice call or a video call) between the local device and a second electronic device (e.g., a remote device). During the call and at the first device, a joint media playback session is initiated, in which the first and second devices independently stream media content (e.g., a musical composition, a movie, etc.) for synchronized playback. The first device determines that a downlink signal from the second device includes speech based on output from a voice activity detector (VAD). For example, the VAD may be an algorithm running locally on the first device that performs a noise reduction algorithm on the downlink signal and generates an output of the VAD based on the downlink signal. On the other hand, the output of the VAD may be received from the second device. In response to determining that the downlink signal includes speech, a scalar gain is applied to the audio signal of the media content to reduce the signal level of the audio signal, and the speaker may be driven with a mixture of the downlink signal and the audio signal. As a result, when the user of the second device is speaking, the sound level of the media content may be reduced.

[0004] In one aspect, a first device is communicatively coupled to a wireless headset for a call and a joint media playback session. In this case, the first device may generate a VAD output based on an accelerometer signal generated by an accelerometer of the wireless headset. In another aspect, the first device may receive the VAD output from the wireless headset, the wireless headset generating the VAD based on the accelerometer signal.

[0005] In some aspects, the media content includes a video signal and an audio signal, such that initiating the joint media playback session includes displaying the video signal on a display screen and driving a speaker with a mixture of the downlink signal and the audio signal. In another aspect, the first device determines a signal level of the downlink signal, and in response to the signal level being above a threshold level or in response to determining based on an output of the VAD that the downlink signal includes speech, the first device displays closed captions on the display screen representing audio content contained within the audio signal of the media content.

[0006] In one aspect, first equipment determines the first timestamp along the playback duration of media content, at this first timestamp place, the output from VAD begins to indicate downlink signal and comprises voice, and determines the second timestamp after the first timestamp along the playback duration of media content, at this second timestamp place, makes the determination that wherein the output indication downlink signal from VAD has stopped comprising voice.In response, first equipment is by at the second timestamp place or afterwards pausing the playback of media content and from the first timestamp along the playback duration, begins the playback of media content and rewinds the playback of media content.On the other hand, in response to determining that the output indication downlink signal from VAD has stopped comprising voice, first equipment can provide the notification (for example, the pop-up notification that shows on the display screen of first equipment) of the user authorization of requesting to rewind the playback of media content.

[0007] In one aspect, the call initiated with the second device can be a telephone (e.g., voice-only) call. Another aspect of the present disclosure is a method performed by a first device, wherein the first device, together with the second device, simultaneously conducts a video conference call and a joint media playback session. The first device determines that a user of the second device begins speaking based on the audio content of the video conference call, and in response to determining that the user begins speaking, reduces the volume level of the audio content of the media content associated with the joint media playback session. In one aspect, in response to determining that the user of the second device stops speaking (e.g., based on the audio content of the video conference call), the first device may increase the volume level of the audio content of the media content to the previous level before the volume level was reduced.

[0008] The above summary does not include an exhaustive list of all aspects of the present disclosure. It is contemplated that the present disclosure includes all systems and methods that can be practiced by all suitable combinations of the various aspects summarized above and disclosed in the detailed description below and particularly pointed out in the claims. Such combinations may have specific advantages not specifically set forth in the above summary. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Various aspects are shown in the accompanying drawings by way of example and not limitation, and similar reference numerals indicate similar elements in the drawings. It should be noted that references to "one" or "an" aspect in this disclosure are not necessarily to the same aspect, and are intended to mean at least one aspect. In addition, for the sake of brevity and to reduce the total number of drawings, a particular drawing may be used to illustrate features of more than one aspect, and not all elements in a drawing may be required for a particular aspect.

[0010] Figure 1 An audio system according to one aspect is shown that includes a local device and one or more remote devices participating in a call while performing a joint media playback session.

[0011] Figure 2 A block diagram is shown of a local device initiating a joint playback media session while participating in a call with one or more remote devices and an audio output device in wireless communication with the local device according to one aspect.

[0012] Figure 3 Several stages are shown in accordance with one aspect in which a local device and a remote device initiate a joint playback media session to synchronously play back a musical composition while participating in a telephone call.

[0013] Figure 4 Several stages are shown according to one aspect in which a local device and a remote device initiate a joint playback media session to synchronously play back a movie while participating in a video call.

[0014] Figure 5 A block diagram of a local device that performs audio signal processing operations on audio signals of media content based on whether speech is detected within the signals of a telephone call being conducted between the local device and a remote device is shown according to one aspect.

[0015] Figure 6 A block diagram of a local device that performs audio signal processing operations on an audio signal of media content based on whether speech is detected by an audio output device is shown according to one aspect.

[0016] Figure 7 A block diagram of a local device that performs audio signal processing operations based on whether speech is detected within a signal of a video call is shown according to one aspect.

[0017] Figure 8 is a flow chart of one aspect of a process for processing an audio signal of media content based on whether speech is detected within a downlink audio signal.

[0018] Figure 9is a flow chart of one aspect of a process for displaying closed captions representing audio content of media content.

[0019] Figure 10 is a flow chart of one aspect of a process for rewinding playback of media content upon determining that a downlink audio signal has ceased to include speech.

[0020] Figure 11 A block diagram according to one aspect is shown in which a local device 2 is communicatively coupled with an audio output device 6 via a two-way wireless audio connection to exchange audio data when the local device participates in a call with a remote device 3 .

[0021] Figure 12 A block diagram according to one aspect is shown in which a local device 2 is communicatively coupled to an audio output device 6 via a two-way wireless audio connection during a joint media playback session and call with a remote device 3 .

[0022] Figure 13a and Figure 13b Several block diagrams are shown according to one aspect in which a local device 2 communicatively coupled with an audio output device 6 for exchanging audio data switches between wireless audio connections based on initiation of a joint media playback session.

[0023] Figure 14 is a flow chart of one aspect of a process for switching between wireless audio connections.

[0024] Figure 15 is a flow chart of another aspect of a process for switching between wireless audio connections.

[0025] Figure 16 is a flow chart of one aspect of a process for determining whether to switch between wireless audio connections based on one or more criteria.

[0026] Figure 17 is a flow chart of one aspect of a process performed by an audio output device for switching between wireless audio connections.

[0027] Figure 18 is a flow chart of one aspect of a process performed by an audio output device for switching from a one-way wireless audio connection to a two-way wireless audio connection based on whether speech is detected. DETAILED DESCRIPTION

[0028] The various aspects of the present disclosure will now be explained with reference to the accompanying drawings. As long as the shape, relative position and other aspects of the components described in a certain aspect are not clearly defined, the scope of the present disclosure is not limited to the components shown here, and the components shown are only for illustrative purposes. In addition, although many details have been set forth, it should be understood that some embodiments can be implemented without these details. In other cases, well-known circuits, structures and technologies are not shown in detail to avoid blurring the understanding of the description. In addition, unless the meaning is clearly contrary, all ranges shown herein are considered to include the end values ​​of each range.

[0029] Figure 1 An audio system 1 according to one aspect is shown, which includes a local device and one or more remote devices that participate in a call when performing a joint media playback session. As described herein, this allows users of the devices to listen to (and / or watch) media content (e.g., on one or more devices) while participating in a session with each other. The audio system includes a local (or first electronic) device 2, a remote (or second electronic) device 3, a network 4 (e.g., a computer network such as the Internet), a media content server 5, and an audio output device 6. In one aspect, the system may include more or fewer elements. For example, the system may have one or more remote devices, all of which participate in calls and joint media playback sessions with each other and with the local device, as described herein. On the other hand, the audio system may include one or more remote (electronic) servers communicatively coupled to at least some of the devices in the audio system 1, and may be configured to perform at least some of the operations described herein. On the other hand, the system may not include an audio output device. In this case, the local device may perform audio output operations (e.g., using one or more signals to drive one or more speakers).

[0030] In one aspect, a local device (and / or a remote device) can be any electronic device (e.g., having electronic components such as a processor, memory, etc.) that can participate in a call such as a phone call (or "voice-only" call) or a video (conference) call while performing a joint media playback session with one or more other devices (e.g., one or more remote devices), wherein (at least some) of the devices simultaneously play back media content (e.g., a musical composition, a movie, etc.). This document describes more about the simultaneous playback of media content. For example, a local device can be a desktop computer, a laptop computer, a digital media player, etc. In one aspect, a device can be a portable electronic device (e.g., capable of handheld operation), such as a tablet computer, a smart phone, etc. In another aspect, a device can be a head-mounted device, such as smart glasses, or a wearable device, such as a smart watch. In one aspect, a remote device can be a device of the same type as the local device (e.g., both devices are smart phones). In another aspect, at least some of the remote devices can be different, such as some are desktop computers and others are smart phones.

[0031] As shown in the figure, local device 2 is coupled to remote device 3 and / or media content server 5 via computer network (for example, the Internet) 4 (for example, communicationally). Specifically, local device and remote device can be configured to establish and participate in telephone (or only voice) call, wherein the equipment participating in the call exchanges audio data. For example, each device transmits at least one microphone signal as an uplink audio signal to the other equipment participating in the call, and receives at least one audio signal as a downlink audio signal from the other equipment for playback by one or more loudspeakers. In one aspect, the network may include a public switched telephone network (PSTN), and local device and remote device can send outgoing calls and / or receive incoming calls through the public switched telephone network. On the other hand, local device can be configured to establish an Internet Protocol (IP) phone (or Voice over IP (VoIP)) call with one or more remote devices via a network (for example, the Internet). Specifically, local device can use any signaling protocol (for example, Session Initiation Protocol (SIP)) to establish a communication session and use any communication protocol (for example, Transmission Control Protocol (TCP), Real-time Transport Protocol (RTP) etc.) to exchange audio data during the call. For example, when a call is initiated (e.g., by a phone application executing within a local device), the local device may transmit one or more microphone signals captured by one or more microphones (e.g., as uplink audio signals) as audio data (e.g., IP packets) to one or more remote devices, and receive one or more (e.g., downlink audio) signals from the remote devices via the network for driving one or more speakers of the local device. In another aspect, the local device may be configured to establish a wireless (e.g., cellular) call. In this case, the network 4 may include one or more cell towers, which may be part of a communication network (e.g., a 4G Long Term Evolution (LTE) network) that supports data transmission (and / or voice calls) for electronic devices such as mobile devices (e.g., smartphones).

[0032] In another aspect, the local device and the remote device can be configured to establish and participate in a video call with one or more remote devices 3. In this case, the local device can establish a video call (e.g., similar to VoIP, using SIP to initiate a session and RTP to transmit data) and exchange video and / or audio data with the one or more remote devices while the video call is established. For example, the local device can include one or more cameras that capture video, which is encoded using any video codec (e.g., H.264) and transmitted to the remote device for decoding and display on one or more display screens. More information about calls is described herein.

[0033] In some aspects, media content server 5 can be a stand-alone server computer or server computer cluster configured to stream media content to electronic devices such as local devices and remote devices. In this case, the server can be a part of a cloud computing system that can stream data as a cloud-based service provided to one or more subscribers. In some aspects, the server can be configured to stream any type of media (or multimedia) content, such as audio content (e.g., musical works, audio books, podcasts, etc.), still images, video content (e.g., films, television productions, etc.). In one aspect, the server can use any audio and / or video encoding format and / or any method for content streaming to one or more devices.

[0034] In one aspect, the media content server 5 can be configured to stream media content to one or more devices simultaneously to allow these devices to participate in a joint media playback session. For example, the server can receive a request from a device (e.g., local device 2) to stream a media content segment that can include audio content (e.g., a musical composition) and / or video content (e.g., a video signal associated with a movie) with another device (e.g., remote device 3). In one aspect, the request can be transmitted by the local device (and / or remote device) in response to the device receiving a user input to start playback of media content, such as Figure 3 and Figure 4 As shown. In this case, the server can establish a communication link with the local device and the remote device that have participated in (for example, telephone and / or video) call. Once the communication link is established, the server can use any codec (for example, MP3, AAC etc.) to encode audio content and / or can use any codec to encode video content, and then transmit the encoded content to each device for decoding and output. On the other hand, the local device can transmit a message to the remote device request to initiate a joint media playback session. In response, the remote device can communicate with the media content server to retrieve media content and play it back synchronously with the local device. In one aspect, the equipment participating in the joint media playback session can output media content synchronously so that the user outputs and experiences content simultaneously. In some aspects, any timing synchronization method can be used (for example, by the equipment and / or server participating in the session) to ensure that media is streamed simultaneously and synchronously. More content about the joint media playback session is described herein.

[0035] As shown in the figure, audio output device 6 can be any electronic device including at least one loudspeaker and configured to perform output sound by driving loudspeaker.For example, as shown in the figure, device is a wireless headset (for example, in-ear headphones or earplugs), which is designed to be positioned on the user's ear (or in) and is designed to output sound to the user's auditory canal. In some respects, earphone can be a sealing type with a flexible earphone end, which is used to acoustically seal the entrance of the user's auditory canal relative to surrounding environment by blocking or being blocked in the auditory canal. As shown in the figure, output device includes a left earphone for the user's left ear and a right earphone for the user's right ear. In this case, each earphone can be configured to output at least one audio channel (for example, right earphone outputs the right audio channel of the binaural input of stereo recording (such as musical work) and left earphone outputs the left audio channel) of media content. On the other hand, output device can be any electronic device including at least one loudspeaker and arranged to be worn by the user and arranged to output sound by driving loudspeaker with audio signal. As another example, the output device may be any type of headphone, such as circum-aural (or supra-aural) headphones that at least partially cover the user's ears and are arranged to direct sound into the user's ears.

[0036] In some aspects, the audio output device can be a head-mounted device, as described herein. In another aspect, the audio output device can be any electronic device configured to output sound into the surrounding environment. Examples may include standalone speakers, smart speakers, home theater systems, or infotainment systems integrated into a vehicle.

[0037] In one aspect, the output device can be a wireless device that is communicatively coupled to a local device in order to exchange audio data. For example, the local device can be configured to establish a wireless connection with the audio output device via a wireless communication protocol (e.g., Bluetooth protocol or any other wireless communication protocol). During the established wireless connection, the local device can exchange (e.g., transmit and receive) data packets (e.g., Internet Protocol (IP) packets) with the audio output device, and the audio output device can include audio digital data in any audio format. Specifically, the local device can be configured to establish and communicate with the audio output device via a two-way wireless audio connection (e.g., which allows two devices to exchange audio data), such as making a hands-free call or using voice commands. Examples of two-way wireless communication protocols include, but are not limited to, hands-free profile (HFP) and headset profile (HSP), both of which are Bluetooth communication protocols. On the other hand, the local device can be configured to establish and communicate with the output device via a one-way wireless audio connection (such as (e.g., Advanced Audio Distribution Profile (A2DP) protocol), which allows the local device to transmit audio data to one or more audio output devices. This article describes more about these wireless audio connections.

[0038] On the other hand, the local device 2 can be communicatively coupled to the audio output device 6 via other methods. For example, both devices can be coupled via a wired connection. In this case, one end of the wired connection can be (e.g., fixedly) connected to the audio output device, while the other end can have a connector that plugs into a socket of the audio source device, such as a media jack or a universal serial bus (USB) connector. Once connected, the local device can be configured to drive one or more speakers of the audio output device with one or more audio signals via the wired connection. For example, the local device can transmit the audio signal as digital audio (e.g., PCM digital audio). On the other hand, the audio can be transmitted in an analog format.

[0039] In some aspects, the local device 2 and the audio output device 6 can be different (separate) electronic devices, as described herein. In another aspect, the local device can be a component of (or integrated with) the audio output device. For example, as described herein, at least some of the components of the local device (such as a controller) can be part of the audio output device, and / or at least some of the components of the audio output device can be part of the local device. In this case, each device can be communicatively coupled via traces that are part of one or more printed circuit boards (PCBs) within the audio output device.

[0040] Figure 2 A block diagram of a local device 2 initiating a joint playback media session while participating in a call (e.g., voice or video) with one or more remote devices 3 is shown, according to one aspect, and an audio output device 6 in wireless communication with the local device is shown. The local device 2 includes a controller 20, a network interface 21, a speaker 22, a microphone 23, a camera 24, a display 25, and (optionally) one or more additional sensors 40. In one aspect, the local device may include more or fewer elements, as described herein. For example, the device may include two or more of at least some of the elements (e.g., having two or more microphones 23).

[0041] The controller 20 can be a dedicated processor such as an application specific integrated circuit (ASIC), a general-purpose microprocessor, a field programmable gate array (FPGA), a digital signal controller, or a set of hardware logic structures (e.g., filters, arithmetic logic units, and a dedicated state machine). The controller is configured to perform audio signal processing operations and / or networking operations. For example, the controller 20 can be configured to participate in a call and simultaneously perform a joint media playback session to stream media content with one or more remote devices via the network interface 21. On the other hand, the controller can be configured to perform audio signal processing operations on audio data of the media content and / or audio data associated with the participating call (e.g., downlink signals). More information about the operations performed by the controller 20 is described herein.

[0042] In one aspect, one or more sensors 40 are configured to detect an environment (e.g., in which the local device is located) and generate sensor data based on the environment. In some aspects, the controller may be configured to perform operations based on the sensor data generated by one or more sensors 40. For example, the local device may include a (e.g., optical) proximity sensor designed to generate sensor data indicating that an object is at a specific distance from the sensor (and / or the local device). As another example, the local device may include an inertial measurement unit (IMU) designed to measure the position and / or orientation of the local device. In one aspect, the sensor may be a component of the local device (or integrated into the local device). On the other hand, the sensor may be a separate electronic device communicatively coupled to the controller (e.g., via the network interface 21). For example, the audio output device 6 may include one or more sensors whose data may be provided to the local device via a wireless connection.

[0043] Speaker 22 can be, for example, an electrodynamic driver that can be specifically designed for sound output in a specific frequency band, such as a woofer, tweeter, or midrange driver. In one aspect, speaker 22 can be a "full-range" (or "full-frequency") electrodynamic driver that reproduces as much of the audible frequency range as possible. Microphone 23 can be any type of microphone configured to convert acoustic energy caused by sound waves propagating in an acoustic environment into an input microphone signal (e.g., a differential pressure gradient microelectromechanical system (MEMS) microphone).

[0044] In one aspect, camera 24 is a complementary metal oxide semiconductor (CMOS) image sensor capable of capturing digital images including image data representing the field of view of camera 24, wherein the field of view includes the scene of the environment in which device 2 is located. In some aspects, the camera can be a charge coupled device (CCD) camera type. The camera is configured to capture still digital images and / or video represented by a series of digital images. In one aspect, the camera can be positioned anywhere near the local device. In some aspects, the device can include multiple cameras (e.g., each camera can have a different field of view).

[0045] Display screen 25 is designed to present (or display) a video of a digital image or video (or image) data. In one aspect, the display screen can use liquid crystal display (LCD) technology, light emitting polymer display (LPD) technology, or light emitting diode (LED) technology, although other display technologies can be used in other aspects. In some aspects, the display can be a touch-sensitive display screen configured to sense user input as an input signal. In some aspects, the display can use any touch sensing technology, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies.

[0046] The audio output device 6 includes a controller 75, a network interface 76, a speaker 77, a microphone 78, and an accelerometer 79. In one aspect, the device may include more or fewer elements. For example, the output device may include one or more microphones and / or one or more speakers. In some aspects, the output device may include a microphone that is an "external" (or reference) microphone arranged to capture sound from the acoustic environment, while having at least one other "internal" (or error) microphone arranged to capture sound (and / or sense pressure changes) within the user's ear (or ear canal). In the case of in-ear headphones, the internal microphone can sense the inside of the user's ear when the headphone is on (or in) the user's ear.

[0047] The accelerometer 79 is arranged and configured to receive (detect or sense) voice vibrations generated when a user (e.g., a user who may be wearing the output device) speaks, and to generate an accelerometer signal representing (or containing) the voice vibrations. Specifically, the accelerometer is configured to sense bone conduction vibrations transmitted from the user's vocal cords to the user's ears (ear canals) when speaking and / or humming. For example, when the audio output device is a wireless headset, the accelerometer may be located anywhere on or within the headset that can contact a portion of the user's body to sense vibrations.

[0048] In one aspect, controller 75 is configured to perform audio signal processing operations and / or networking operations, as described herein. For example, controller may be configured to obtain (or receive) the audio data (as analog or digital audio signal) of media content or the media content (for example, music etc.) that user expects, to play back by loudspeaker 77. In some aspects, controller may obtain audio data from local storage, or controller may obtain audio data from network interface 76, thereby may obtain data from external source such as local device 2 (via its network interface 21). For example, output device may stream the audio signal from local device (for example, via Bluetooth connection) to play back by loudspeaker 77. Audio signal may be a signal input audio channel (for example, mono). On the other hand, controller may obtain two or more input audio channels (for example, stereo channels) for outputting by two or more loudspeakers. In one aspect, when output device includes two or more loudspeakers, controller may perform additional audio signal processing operations. For example, the controller may spatially render the input audio channels (e.g., by applying a spatial filter such as a head-related transfer function (HRTF)) to produce binaural output audio signals for driving at least two speakers (e.g., a left speaker and a right speaker).

[0049] In one aspect, the controller 75 can be configured to perform (additional) audio signal processing operations based on the elements coupled to the controller. For example, when the output device includes two or more "out-of-the-ear" speakers arranged to output sound into an acoustic environment rather than speakers arranged to output sound into the user's ears (e.g., as speakers of in-ear headphones), the controller can include a sound output beamformer configured to generate speaker driver signals that produce spatially selective sound output when driving the two or more speakers. Thus, when used to drive the speakers, the output device can produce a directional beam pattern that can be directed to a location within the environment.

[0050] In some aspects, the controller 75 may include a sound pickup beamformer that can be configured to process audio (or microphone) signals generated by two or more external microphones of the output device to form a directional beam pattern (as one or more audio signals) for spatially selective sound pickup in certain directions to be more sensitive to one or more sound source locations. In some aspects, the controller can perform audio processing operations (e.g., perform spectrum shaping) on ​​the audio signal containing the directional beam pattern and / or transmit the audio signal to a local device.

[0051] On the other hand, the controller 75 may perform other functions. For example, the controller 75 may be configured to perform an active noise cancellation (ANC) function to cause the speaker 77 to generate anti-noise in order to reduce ambient noise from the environment that leaks into the user's ears. The ANC function may be implemented as one of feedforward ANC, feedback ANC, or a combination thereof. Therefore, the controller 75 may receive a reference microphone signal from a microphone that captures external ambient sound, such as microphone 78. On the other hand, the controller may perform any ANC method to generate anti-noise. On the other hand, the controller 75 may perform a transparency function, in which the sound played back by the audio output device 6 is a reproduction of the ambient sound captured by the device's external microphones in a "transparent" manner (e.g., as if the headphones were not being worn by the user). The controller 75 processes at least one microphone signal captured by at least one external microphone 78 and filters the signal through a transparency filter. This can reduce acoustic obstruction caused by the audio output device being positioned directly on, in, or above the user's ear, while also preserving the spatial filtering effects of the wearer's anatomical features (e.g., head, auricle, shoulders, etc.). The filter also helps preserve the timbre and spatial cues associated with the actual ambient sound. In one aspect, the filter for the transparency function can be user-specific based on specific measurements of the user's head. For example, the controller 75 can determine the transparency filter based on a head-related transfer function (HRTF) or equivalently a head-related impulse response (HRIR) based on the user's anthropometric measurements.

[0052] As described herein, both the local device and the audio output device are configured to establish a wireless audio connection (e.g., a Bluetooth connection) to exchange audio data. In one aspect, the controller 75 (and / or the controller 20) can be configured to switch between a two-way wireless audio connection (e.g., an HFP connection) and a one-way wireless audio connection (e.g., an A2DP connection) to communicatively couple the two devices together to exchange (and transmit) audio data. More information about switching between audio connections is described herein.

[0053] In one aspect, the operations performed by the controller may be implemented in software (eg, as instructions stored in a memory and executed by the controller) and / or may be implemented by hardware logic structures as described herein.

[0054] In another aspect, at least some of the operations performed by the audio system 20 as described herein can be performed by the local device 2 and / or the audio output device 6. For example, the local device can include two or more speakers and can be configured to perform sound output beamformer operations (e.g., when the local device includes two or more speakers). In another aspect, at least some of the operations can be performed by a remote server communicatively coupled to either device, for example, via a network (e.g., the Internet).

[0055] In one aspect, at least some elements of the local device 2 and / or the audio output device 6 can be integrated with (or be a component of) each respective device. For example, when the audio output device is an on-ear headset, the microphone, speaker, and accelerometer can be part of at least one ear cup of the headset that is placed on the user's ear. In another aspect, at least some elements can be separate electronic devices communicatively coupled to the device. For example, the display screen 25 can be a separate device (e.g., a display monitor or television) that is communicatively coupled (e.g., wired or wirelessly connected) to the local device to receive image data for display. As another example, the camera 24 can be a component of a separate electronic device (e.g., a webcam) that is coupled to the local device to provide captured image data.

[0056] As described herein, a local device 2 and a remote device 3 of an audio system 1 can perform a joint media playback session while participating in a call to allow users of the devices to communicate while experiencing simultaneous media content playback. In one aspect, a local device can initiate a joint media playback session while already participating in a call. Figure 3 and Figure 4 Graphical examples of a local device and a remote device initiating joint media playback while participating in a telephone call and a video conference call, respectively, are shown.

[0057] Figure 3Shown are three stages 26 to 28 according to one aspect, wherein a local device 2 and a remote device 3 initiate a joint playback media session to synchronously play back a musical work when participating in a phone call. The first stage 26 shows the main (or home) screen user interface (UI) displayed on the display screen of each respective device when the device participates in a phone call. In one aspect, any device can initiate a phone call, as described herein. Specifically, the home screen UI 11 of the local device shows the caller ID information of the remote device overlaid on several optional UI items, each of which is associated with an application (e.g., application 1 to application 4), including a media application 29 that streams media content (e.g., from a media content server 5) to the local device when executed by the local device. Specifically, the media application 29 can be a music streaming application that streams music for playback by speakers 22 (and / or speakers 77 of an audio output device) when executed. Similarly, the home screen UI 12 of the remote device shows the caller ID information of the local device overlaid on several (similar) UI items, which are those shown for the local device. In one aspect, either device can initiate a phone call using any known method. For example, a user of the local device may have initiated a phone application stored on the local device and dialed the phone number of the remote device. Once dialed, the local device can connect to the remote device via a cellular network (e.g., a 4G Long Term Evolution (LTE) network) of Network 4, as described herein.

[0058] This stage also shows that the user of the local device 2 presses a UI item associated with the media application 29. For example, the display screen of the local device (e.g., Figure 2 The display screen 25 shown may be a touch-sensitive display screen, as described herein. The local device may receive user input in response to the user pressing a UI item of the media application 29. The second stage 27 shows the result of the user pressing the UI item of the media application 29. Specifically, this stage shows the UI 30 of the media application displayed on the display screen of the local device, which displays the title of the musical piece (e.g., "The Music") and playback control UI items including a play button, a rewind button, and a fast forward button. This stage also shows that the user has pressed the "Play" button.

[0059] Phase III 28 shows the result of the user selection play button of local device.Particularly, in case selected play button, local device just to media content server 5 transmission request, to begin media content streaming to remote device and local device.In one aspect, when a plurality of equipment calls (for example, conference call) together, media content server 5 can be with media content streaming to each equipment that participates in conference call.Therefore, remote device and local device all playback media content (for example, by driving corresponding loudspeaker with the audio data of the media content that receives from the media content server).Therefore, two equipment playback content simultaneously and synchronously, this is illustrated by the process indicator 39 of two equipment shown in the corresponding media application program UI that is in the midway mark place.This paper has described more about playback media content simultaneously.

[0060] Figure 4 Three stages 31 through 33 are shown, according to one aspect, in which a local device 2 and a remote device 3 initiate a joint playback media session to synchronously play back a movie while participating in a video call. The first stage 31 illustrates the home screen UI displayed on the display screen of each respective device while the devices are participating in a video call. Specifically, overlaid on the home screen UI 11 of the local device is a video call UI 14, which shows a video representation of the local user 38 located in the upper right corner of the UI and a video representation of the remote user 37 (which is larger than the local user's representation) located in the center of the video call UI. Similarly, overlaid on the home screen UI 12 of the remote device is a video call UI 15, which shows a video representation of the remote user located in the center of the UI and a video representation of the local user located in the upper right corner. In one aspect, the video representations can be generated using video data captured by one or more cameras of each device. For example, when the local user is within the field of view of camera 24, the camera can capture video data of the local user, which is then displayed on the local device and transmitted (e.g., via network 4) to the remote device for display on the remote device's display screen.

[0061] This stage also shows the local user selecting a selectable UI item associated with a media application 35 within the home screen UI 11, which may be a video streaming application. The second stage 32 shows the result of the user pressing the UI item for the media application 35. Specifically, this stage shows the UI 18 for the media application 35 being displayed on the display screen of the local device, which displays the title of the movie (e.g., "The Movie"), a playback duration of one hour and thirty minutes, and a play button pressed by the local user.

[0062] The third stage 33 shows the result of the local user selecting the play button in the media application UI 18. Specifically, once the play button is selected, the local device transmits a request to the media content server 5 to begin streaming the media content (e.g., the audio and video data of the movie) to the device participating in the video call. As a result, the two devices synchronously play back the video of the media content 36 (and output the audio of the media content) while still participating in the video call.

[0063] As shown in these examples, when a device is involved in a phone call, audio content can be played back in a joint media playback session, and when a device is involved in a video call, both video and audio content can be played back during the session. On the other hand, when a local device and a remote device are involved in a phone call or video call, any type of media content can be played back during a joint media playback session. For example, a movie can be played back during a joint media playback session while the device is involved in a phone call.

[0064] While participating in a joint media playback session during a call can provide participants with a better user media experience than the media content being played back through the participants' devices (e.g., by allowing participants to discuss the media content of the playback session in real time), there may be some disadvantages. For example, conversations between participants may drown out or mask the sound of the media content. For example, when participants are watching a movie, conversations between participants may be indistinguishable from the dialogue of the simultaneously output movie. As a result, participants participating in these one-sided conversations may find it difficult to speak while the movie is playing. Furthermore, this may also reduce the overall user experience for those participants not participating in these conversations, as the conversations may distract them, causing them to focus their full attention on the sound of the movie. Therefore, there is a need to maintain media audio playback quality when participants participate in a joint media playback session during a call.

[0065] In order to overcome these defects, the present disclosure describes an audio system that can maintain the audio quality of media content playback during a media playback session by processing remote active voice during a call. Specifically, when participating in a call and a joint media playback session in which a local device and (at least one) remote device independently stream media content for synchronous playback, the audio system determines that a downlink (audio) signal from the remote device includes voice based on the output from a voice activity detector (VAD). If so, the audio system applies a scalar gain to the audio signal of the media content to reduce the signal level of the audio signal. The audio system then drives a loudspeaker with a mixed content of the downlink signal and the audio signal. Like this, when the participant of the remote device is speaking, the system can manage the signal level of the media content.

[0066] Figure 5A block diagram of a local device 2 is shown that performs audio signal processing operations on audio signals of media content based on whether voice is detected within the signals of a telephone call being conducted between the local device 2 and at least one remote device 3, according to one aspect. Specifically, the figure shows a controller 20 having multiple operational blocks for performing audio signal processing operations to handle remote active voice during a call and a joint media playback session. As shown, the controller includes a call manager 46, a joint media playback session manager 47, a voice digital signal processor (DSP) 41, a voice activity detector (VAD) 42, a scalar gain 43, a (e.g., matrix) mixer 44, and (optionally) an additional DSP 45.

[0067] The call manager 46 is configured to initiate (and conduct) calls between the local device 2 and one or more remote devices 3. In one aspect, the call manager can initiate calls in response to user input. For example, the call manager can be part of (or receive instructions from) a phone application executed by the local device (e.g., the controller 20 of the local device). For example, the phone application can display a UI on the display 25 of the local device, which can provide the user of the local device with the ability to initiate calls (e.g., a keyboard, a contact list, etc.). Once the UI receives user input (e.g., dialing the remote user's phone number using a keyboard), the call manager can communicate with the network interface 21 of the local device 2 to establish the call, as described herein. In one aspect, the phone call can be made over any network, such as over the PSTN and / or over the Internet (e.g., for VoIP calls). In some aspects, the call manager can initiate the call as described herein and / or using any method.

[0068] Once a call is initiated, the call manager may exchange call data between the remote devices that the local device uses to participate in the call. For example, the call manager may receive one or more downlink audio signals from each of the remote devices. In one aspect, the call manager may mix the downlink signals into (at least one) downlink audio signal (e.g., via a matrix mixing operation). In addition, the call manager may receive a microphone signal from microphone 23 (e.g., which may include the voice of the local user) and transmit the microphone signal as an uplink audio signal to each remote device. In some aspects, when the local device includes two or more microphones, the call manager may transmit a sound pickup beamformer signal for sound that includes a directional beam pattern.

[0069] The joint media playback session manager 47 is configured to initiate a joint media playback session between a local device and one or more remote devices, wherein the two devices independently stream media content for synchronized playback. For example, in response to receiving an instruction to initiate a session, the playback session manager may transmit a request to initiate a session to a media content server, as described herein. Specifically, a media application executing within the local device may transmit instructions to the session manager in response to receiving user input (e.g., based on a user selecting a play button in the media application, such as Figure 3 and Figure 4 ). In another aspect, the session manager can request user authorization before initiating a session. For example, once a user initiates media playback in a media application, the session manager can provide a notification (e.g., a pop-up notification displayed on display screen 25) requesting user authorization to initiate a joint media playback session with (at least some of) the call participants. Upon receiving user authorization (e.g., by receiving a user selection of a UI item within the pop-up notification), the session manager can proceed to request initiation of the session, as described herein.

[0070] In one aspect, the joint media playback session manager 47 is configured to receive media content data (e.g., once a session is initiated). In this case, the session manager receives at least one audio signal (or audio channel) associated with the media content. For example, the received audio signal may be associated with a piece of music that the local user has requested playback, such as Figure 3 In one aspect, the session manager can receive two or more audio signals for a piece of media content. For example, when streaming a piece of music from a media content server, the session manager can receive two audio channels (e.g., the left and right channels of a stereo recording of the piece of music). In another aspect, the session manager can receive two or more audio channels, such as the entire audio soundtrack of a 5.1 surround format film.

[0071] The voice DSP 41 is configured to receive a downlink audio signal from the call manager and to perform voice processing operations on the signal. In one aspect, the voice DSP may perform a noise reduction algorithm on the downlink signal to reduce (or eliminate) the noise contained therein (e.g., to produce a voice signal primarily containing the voice of the remote user). In one aspect, the algorithm may apply a high-pass filter to process the signal, as most noise (or non-voice noise) may be low-frequency content. In another aspect, the algorithm may improve the signal-to-noise ratio (SNR) of the signal to process the signal. To this end, the voice DSP may perform spectral shaping on the downlink signal by applying one or more filters (e.g., a low-pass filter, a band-pass filter, a high-pass filter, etc.) to the signal. For another example, the DSP may apply a scalar gain value to the signal. In one aspect, the voice DSP may perform any method to process the downlink signal to reduce the noise contained therein.

[0072] The VAD 42 is configured to receive (e.g., processed) a downlink audio signal and to perform a voice activity detection (or voice detection) operation to detect the presence (or absence) of a user's voice (speech) therein. For example, the VAD may determine whether (at least a portion of) the spectral content of the downlink signal is associated with human speech. In another aspect, the VAD may determine the presence of speech based on whether the signal level of the downlink signal exceeds a threshold. In some aspects, the VAD may use any method to determine whether speech is present within the signal. The VAD is configured to generate an output based on the downlink signal. Specifically, the VAD may generate a VAD signal indicating whether speech is contained within the downlink signal. For example, when speech is detected, the VAD signal may have a high signal level (e.g., 1), and when speech is not detected (or at least not detected within a threshold level), the VAD signal may have a low signal level (e.g., 0). In another aspect, the VAD signal need not be a binary decision (speech / no speech); as described herein, it may be a probability of speech presence based on a scalar gain to be adjusted. In some aspects, the VAD signal may also indicate the signal level (eg, sound pressure level (SPL)) of the detected speech.

[0073] As described herein, the VAD may receive a mix of two or more downlink audio signals (e.g., mixed by call manager 46), each downlink signal received from a remote device participating in a call (e.g., a conference) with a local device. In one aspect, the VAD may receive each individual downlink signal to determine whether at least one of the downlink signals contains speech. Upon detecting speech in at least one of the downlink signals, the VAD may generate a VAD signal indicating speech detection. In some aspects, a voice DSP may process each individual downlink signal prior to receipt by the VAD.

[0074] In another aspect, in addition to (or instead of) generating a VAD signal, the local device may optionally receive a VAD signal from a remote device (e.g., at least one of them). Specifically, each remote device may include its own VAD and may be configured to generate a VAD signal as an output of the VAD that indicates whether at least one microphone signal generated by a microphone of the remote device (and / or its uplink signal transmitted to the local device 2 during a call) includes active speech of the remote user. Once generated, each remote device may transmit the VAD signal to the local device via the network 4. Upon receipt, the scalar gain 43 may apply a scalar gain value to the audio signal of the media content based on the VAD signal received from the remote device.

[0075] The scalar gain 43 is configured to receive an audio signal from the joint media playback session manager 47 and a VAD signal from the VAD 42 (and / or from at least one remote device), and is configured to process the audio signal based on the VAD signal. Specifically, the scalar gain is configured to adjust the signal level of the audio signal (e.g., at least a portion thereof) by applying one or more scalar gain values ​​based on whether the VAD signal indicates the presence of voice detected within the downlink audio signal. Specifically, the gain adjustment can reduce the volume level of the audio signal of the media content associated with the joint media playback session (e.g., streamed by the joint media playback session). In one aspect, the applied scalar gain value can be a predetermined value. In another aspect, the value can be based on the VAD signal. For example, as described herein, the VAD signal can indicate the signal level of the downlink audio signal (or more specifically, the signal level of the voice contained therein). In this case, the scalar gain can be configured to adjust the applied scalar gain value based on the signal. For example, when speech detected in a downlink audio signal is at a certain signal level, the scalar gain may apply a gain value to reduce the signal level of the audio signal below the certain signal level of the downlink signal to ensure that the media content sounds below the speech within the call.

[0076] The mixer 44 is configured to receive the processed audio signal from the scalar gain 43 and the processed downlink audio signal from the voice DSP 41, and is configured to perform a matrix mixing operation, for example, to produce a mixed content of the two signals. The controller can use the mixed signal to drive the speaker 22 to play back the sound of the call and the media content of the playback session. On the other hand, the mixer can receive one or more unprocessed downlink audio signals. For example, the mixer can receive the downlink audio signal from the call manager 46 instead of receiving the processed downlink audio signal from the voice DSP 41.

[0077] In one aspect, the controller may optionally have an additional DSP 45 that can be configured to perform one or more audio signal processing operations on the mixed content. For example, the additional DSP can perform at least some of the operations described herein, such as spatially rendering the mixed content (e.g., by applying a spatial filter, such as a head-related transfer function (HRTF)) to generate a binaural audio signal for driving one or more speakers (e.g., a left speaker and a right speaker), as described herein. The controller 20 can then use the processed mixed content to drive the speaker 22, as described herein. Thus, in response to determining that a remote user has begun (and / or is actively) speaking during a call with a local user, the controller can perform the operations described herein to reduce the volume level of the media content.

[0078] As described so far, in response to detecting the presence of sound (or voice) in one or more downlink signals from one or more remote devices, the controller 20 applies a scalar gain. On the other hand, this determination can be based on whether the local user of the local device is speaking. Specifically, the VAD signal generated by the VAD can indicate whether one or more remote users and / or local users are speaking. To determine this, the voice DSP 41 can optionally obtain a microphone signal generated by the microphone 23 to perform the noise reduction operation described herein. The VAD can receive the processed downlink audio signal and / or the processed microphone signal from the voice DSP 41 and can generate the VAD signal based on either (or both) signals. Therefore, when the local user or the remote user is speaking, the local device can reduce the signal level of the audio signal of the media content.

[0079] In one aspect, when the media content includes two or more audio signals, the controller 20 performs at least some of the operations on at least one of the audio signals. For example, when the media content includes two audio channels for stereo recording, the controller 20 may perform at least some of the operations on the two audio channels to reduce the signal level of each audio channel output by two or more speakers of the local device.

[0080] In some aspects, controller 20 may process the audio signal of the media content while the VAD signal indicates that the downlink signal includes active remote speech. Specifically, while the VAD signal indicates the presence of speech (e.g., as long as the remote user or the local user is speaking), scalar gain 43 may continue to apply the scalar gain value. Once the VAD signal indicates that speech is no longer present, the controller may stop applying scalar gain 43, in which case the audio signal may enter mixer 44 without scalar gain adjustment. In one aspect, once speech is no longer present, the applied scalar gain value may be gradually reduced to gradually increase the signal level of the audio signal.

[0081] Figure 6 A block diagram of a local device 2 is shown according to one aspect, which performs audio signal processing operations on an audio signal of media content based on whether an audio output device 6 detects speech. Specifically, the figure shows that the local device is communicatively coupled to the audio output device to perform (e.g., "hands-free") calls and such. Figure 5 The joint media playback session described. For example, two devices may be connected via a two-way wireless audio connection (e.g., according to the HFP protocol), wherein the two devices exchange audio data of the telephone call and the media content being played back during the joint media playback session. For example, the audio output device may be a hands-free device, such as a wireless headset configured to transmit a microphone signal generated by the microphone 78 to the controller 20 (e.g., the controller's call manager 46), which then transmits the microphone signal to one or more remote devices as an uplink signal for the call. In addition, the local device transmits a (e.g., processed) mixed content of the audio signal and the (processed) downlink signal to the audio output device via the two-way audio connection, which uses the mixed content to drive the speaker 77 (rather than the speaker 77 as shown in FIG. Figure 5 The mixed content is shown used to drive speakers 22).

[0082] This figure also illustrates that the scalar gain 43 can apply a gain value based on the output of the VAD 82 of the local device. Specifically, the gain value can be applied in response to the audio output device detecting the voice of the local user. For example, the audio device includes a VAD 82, which is configured to receive an accelerometer signal generated by the accelerometer 79 and is configured to generate a VAD signal based on the received signal. Specifically, the VAD determines whether the energy level of the accelerometer signal is above an accelerometer signal threshold (or energy threshold), which can be an indication that the user is speaking. In response to determining that the energy level is above the energy threshold, the VAD signal can be set to a high signal level, as described herein. When the VAD signal is generated, the audio output device 6 transmits the signal to the local device 2, and the scalar gain 43 receives the signal to apply a gain value based on the signal, as described herein.

[0083] In one aspect, in addition to (or instead of) the VAD 82 receiving the accelerometer signal, the VAD may (optionally) receive a microphone signal generated by the microphone 78 to generate a VAD signal, as described herein. In another aspect, rather than generating a VAD, the audio output device may transmit the accelerometer signal (and / or microphone signal) to the local device's VAD 42, which may then use the signal to generate a VAD signal, as described herein. Thus, the local device (e.g., the local device's VAD 42) may generate a VAD signal based on the accelerometer signal generated by the accelerometer 79.

[0084] Figure 7A block diagram of a local device 2 is shown that performs audio signal processing operations based on whether voice is detected within a signal of a video call, according to one aspect. Specifically, the diagram shows a controller 20 conducting a video call and a joint media playback session simultaneously with one or more remote devices while performing audio signal processing to handle remote active voice and / or performing video processing operations.

[0085] In one aspect, the local device 2 can perform a video call and a joint media playback session, such as Figure 4 As shown. Specifically, the call manager 46 can be configured to initiate (and conduct) a video call between the local device 2 and one or more remote devices 3. In this case, in addition to transmitting the microphone signal captured by the microphone 23 as an uplink audio signal, the call manager can receive a camera (e.g., video) signal from the camera 24 and transmit the video signal as an uplink video signal together with the uplink audio signal (or in place of the uplink audio signal) to the remote devices participating in the video call. For example, as described herein, the call manager (e.g., in response to receiving a user request in a telephone or video conferencing application executed within the local device) can establish a communication session with the remote device, encode the microphone and camera signals, and transmit the encoded signals (as uplink signals) to the remote device. In addition to transmitting the uplink signals, the call manager can receive at least one downlink audio signal and at least one downlink video signal from each remote device participating in the video call for output by the speaker 22 and the display 25, respectively. In one aspect, any method can be used to initiate and conduct a video call. In some aspects, the joint media playback session manager 47 may be configured to receive media content data including at least one audio signal and at least one video signal associated with a media content segment. For example, the received audio signal and video signal may be associated with a movie that a local user has requested to be played back, such as Figure 4 shown.

[0086] In one aspect, the controller 20 may perform the same operations as described above while simultaneously conducting a video call and a joint media playback session. Figure 5 and Figure 6 4. For example, a controller (e.g., VAD 42 of the controller) may determine, based on the downlink audio signal (e.g., audio content) of the video conference call, whether a remote user of the remote device has begun speaking (and / or is actively speaking). In response, the controller may use scalar gain 43 to apply a scalar gain value to reduce the volume level of the audio signal when output by speaker 22.

[0087] In addition, the controller 20 includes additional operation blocks for performing audio signal processing operations and / or video processing operations based on whether the voice of the remote user is active. For example, the controller includes a closed caption generator 48 and a video processor 49. The closed caption generator is configured to generate closed captions representing the audio content contained in the audio signal of the media content based on the VAD signal output of the VAD 42. Specifically, the caption generator can be configured to generate closed captions in response to the controller 20 determining that the downlink signal (or at least one downlink signal) includes voice based on the VAD signal (e.g., a VAD signal with a high signal level indicating that the downlink signal includes voice, as described herein), and can be configured to display closed captions. Therefore, when the remote user starts speaking (and when the user speaks), closed captions can be generated and displayed. In one aspect, once the VAD signal indicates that the downlink signal no longer includes voice, the caption generator can stop generating and displaying closed captions. On the other hand, the closed caption generator can continue to generate and display closed captions for a period of time after the remote user stops speaking.

[0088] On the other hand, the closed caption generator 48 can be configured to generate closed captions for display in response to determining that the output sound level of the local device is below a threshold sound level. For example, the caption generator can determine whether the local user has lowered the volume of the local device (for example, by adjusting the volume control of the local device to detect whether the user has lowered the volume). If so, the caption generator can automatically generate and display captions. On the other hand, captions can be displayed based on the signal level of the audio signal associated with the media content. For example, the caption generator can generate and display captions in response to the processed audio signal of the media content by having a scalar gain with a signal level below the threshold.

[0089] In one aspect, to generate closed captions, a closed caption generator is configured to receive an audio signal associated with the media content being streamed during a session from a session manager 47, and can be configured to generate captions based on the audio content contained therein. In some aspects, the generator can execute a speech-to-text algorithm to identify the speech contained within the audio signal and can generate a text representation of the identified speech. Therefore, captions can include a transcription of the audio content. On the other hand, captions can include a text description of non-speech audio, such as a description of the current scene. In another embodiment, captions can be obtained from media content data rather than generated. In this case, the caption generator can receive captions from the session manager. In some aspects, the caption generator can use any method to generate captions.

[0090] In one aspect, the video processor 49 is configured to receive image data, such as a downlink video signal from the call manager 46, a video signal from the session manager 47, and (optionally) closed captions from the caption generator 48 (e.g., when the VAD signal indicates active remote voice), and is configured to render the data for display on the display screen 25 for playback of media content during a video call (e.g., as Figure 4 For example, the video processor may overlay closed captions on a displayed video signal of the media content. In some aspects, the video processor may perform other video processing operations on one or more of the video signals, such as image resizing, image compositing, etc.

[0091] In one aspect, the controller can adjust the playback of media content based on whether VAD 42 detects remote active voice. Specifically, once it is determined (e.g., by the VAD) that the remote voice is no longer active, the joint media playback session 47 can rewind the media content to the moment before the active voice was initially detected. For example, the joint media playback session manager can receive a VAD signal from VAD 42 and determine a first timestamp along the playback duration of the media content at which the VAD signal begins to indicate that the downlink signal includes voice (e.g., the moment the VAD signal transitions from a low signal level to a high signal level). At this point, the remote user and the local user may have begun a conversation. Once the conversation ends, the media content can be rewound to begin playback at (or before) the first timestamp along the playback duration. For example, once the session manager determines a second subsequent timestamp at which the VAD signal indicates that the downlink signal has stopped including voice (e.g., the moment the VAD signal level transitions from a high signal level to a low signal level), the session manager can pause playback of the media content (at or after the second timestamp). In one aspect, pausing video playback can include pausing the display of the media content at a moment along the playback duration. Additionally, audio playback of the audio signal can be paused by ceasing to drive the speaker 22 with the mixed content of the downlink signal and the audio signal. In another aspect, audio playback of the audio signal can be paused while playback of the downlink audio signal can continue. In this case, once it is determined that audio playback will be paused, the mixer 44 can stop mixing the two signals and can pass the downlink signal to drive the speaker 22. Thus, the local user and the remote user can engage in a conversation and, when complete, can continue to experience playback of the media content.

[0092] In one aspect, playback adjustments can occur on at least some of the remote devices participating in the call and the joint media playback session with the local device. For example, in response to the remote voice no longer being active, the controller 20 can transmit a control signal to the remote device instructing the device to rewind playback to a moment along the playback duration.

[0093] Figures 8 to 10 are flow charts of processes 50, 60, and 70, respectively, for performing one or more operations in response to detecting remote active voice. In one aspect, the processes may be performed by one or more devices of the audio system 1, such as Figure 1 For example, at least some of the operations of these processes may be performed by local device 2 (eg, controller 20 thereof) and / or by audio output device 6 (eg, controller 75 thereof).

[0094] about Figure 8 , which is a flow chart of one aspect of a process 50 for processing an audio signal of media content based on whether voice is detected within a downlink audio signal. Process 50 begins with controller 20 initiating a call (e.g., a phone call or a video call) between a local device 2 and one or more remote devices 3 (at block 51). As described herein, the call may be initiated by call manager 46 in response to receiving a request from a local user. In one aspect, the initiation of the call may be in response to receiving an incoming call from one or more remote devices. In this case, the call may be initiated by the call manager in response to the user accepting the call (e.g., via a user selecting a UI item of a phone application that is displayed on display screen 25 to answer the call when an incoming call signal is received from a remote device).

[0095] During a call, the controller 20, as a local device 2, initiates a joint media playback session in which the local device and one or more remote devices independently stream media content for synchronous playback (at block 52). For example, a joint media playback session manager 47 may initiate playback based on user input. In one aspect, the playback session may be between all devices making a call. In another aspect, a playback session may be initiated between the local device and at least some of the remote devices. In this case, when initiated, the local user may define which remote devices will participate. In some aspects, initiating a joint media playback session may be responsive to the controller 20 receiving an initiation request from one or more remote devices and / or the media content server 5.

[0096] Once initiated, controller 20 may receive at least one audio signal and / or at least one video signal associated with media content, as described herein, and may be configured to play back the media content and simultaneously output a downlink audio signal and / or a downlink video signal, as described herein.

[0097] Controller 20 determines whether the downlink signal from one or more of the remote devices includes (e.g., remotely active) speech (at decision block 53) based on output from a VAD (such as VAD 42 of controller 20 and / or VAD 82 of audio output device 6). Specifically, the controller may determine whether the VAD signal is at a high signal level, which occurs when the remote user begins to speak or has begun speaking. If so, controller 20 applies a scalar gain to the audio signal associated with the media content to reduce the signal level of the audio signal (at block 54). For example, when speech is detected, controller 20 may apply scalar gain 43 to the audio signal from session manager 47. Controller 20 mixes the (gain-adjusted) audio signal with the downlink signal (at block 55). Controller 20 drives a speaker with the mixed content (at block 56). In one aspect, the speaker may be a component of the local device, such as speaker 22. In another aspect, the speaker may be a component of a separate electronic device communicatively coupled to the local device, such as speaker 77 of audio output device 6.

[0098] Figure 9 6 is a flow chart of one aspect of a process 60 for displaying closed captions representing audio content of media content. In one aspect, the process can be performed when a local device 2 and one or more remote devices 3 are simultaneously conducting a call and a joint media playback session, as described herein. Process 60 begins when the controller 20 receives a downlink signal (at box 61). The controller receives an output from a VAD (e.g., VAD 42) indicating whether the downlink signal includes voice (at box 62). The controller determines whether the output from the VAD indicates that the downlink signal includes voice (at decision box 63). Specifically, the controller determines whether the user of the remote device begins (or has begun) to speak. If so, the controller generates closed captions representing the audio content contained in one or more audio signals of the media content (at box 64). The controller then displays the closed captions (at box 65). Like this, in response to determining that the remote user is speaking, the local device 2 displays the closed captions on the display screen 25.

[0099] Figure 107 is a flow chart of one aspect of a process 70 for rewinding the playback of media content when it is determined that a downlink audio signal has stopped including voice. Process 70 begins with the controller 20 determining a first timestamp along the playback duration of the media content at which the output from the VAD begins to indicate that the downlink signal includes voice (at block 71). The controller 20 determines a second timestamp after the first timestamp along the playback duration of the media content at which the output from the VAD stops including voice (at block 72). Specifically, the first timestamp can be determined in response to determining that a VAD signal generated by the VAD is at a high signal level, and the second timestamp can be determined in response to determining that the VAD signal changes from a high signal level to a low signal level. The controller 20 rewinds the playback of the media content (at block 73) by pausing the playback of the media content at or after the second timestamp and starting the playback of the media content from (or before) the first timestamp along the playback duration.

[0100] Some aspects can be Figures 8 to 10 Variations may be performed on processes 50, 60, and / or 70 described in the foregoing. For example, specific operations in at least some of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in a continuous series of operations, and different specific operations may be performed in different aspects. For example, in Figure 8 In this case, the local user can (for example, in a media application such as Figure 3 and Figure 4 ) selects media content for playback and selects one or more remote devices (e.g., selects contact information associated with the remote devices, such as a phone number). Once the selection is complete, the local user can initiate playback by selecting a play button, such as Figure 3 and Figure 4 shown.

[0101] In addition, the controller 20 may perform one or more operations in response to detecting a remote active voice. For example, upon detecting that a remote voice has begun, the controller 20 may perform the operations in processes 50 and 60 to reduce the volume level of the audio signal and display closed captions.

[0102] In one aspect, controller 20 may cease performing at least some of the operations described in processes 50, 60, and / or 70 in response to the output of the VAD indicating that the downlink signal does not include speech. For example, when the output of the VAD indicates that speech is not within the downlink signal, the controller may Figure 8The controller may stop applying the scalar gain to the audio signal at block 54 of the control. Consequently, the sound level of the media content may be restored to its previous level before the volume level was reduced (e.g., before the remote user's voice was detected). Similarly, once the remote voice is no longer determined to be active, the controller may stop generating and displaying closed captions at blocks 64 and 65.

[0103] In one aspect, the operations performed by the controller to maintain the audio quality of the media content based on the detection of remote active speech can be automatic (e.g., without requiring user intervention). For example, closed caption generator 48 can automatically generate and display captions based on the output of the VAD, as described in process 60. In another aspect, at least some of the operations (e.g., adjusting the signal level of the audio signal by applying a scalar gain, generating and displaying closed captions, and / or rewinding playback, etc.) can be performed in response to receiving user authorization. Specifically, in response to determining that the output of the VAD indicates that the downlink signal has ceased to include speech, the controller can provide a notification to the local user requesting authorization to perform at least one of the operations described herein. For example, upon determining at block 72 of process 70 that the second timestamp of remote speech is no longer present, the controller can provide a notification to the user requesting authorization to rewind playback at block 73. In one aspect, the notification can be a pop-up notification displayed on display screen 25. Once authorization is received (e.g., by the user selecting a UI item), the controller can perform at least one of the operations described herein. On the other hand, if (for example, within a period of time) user authorization is not received, the controller may abandon at least some of the operations described herein. For example, if authorization to rewind playback is not received, the controller may continue to play back media content after the period of time.

[0104] As described herein, operations performed by the controller to maintain media quality of media content playback (e.g., application of a scalar gain, generation and display of closed captions, and / or rewinding of media content playback, etc.) can be based on whether remote active speech is present during a concurrent call. Furthermore, at least some of the operations can be performed in response to the controller determining that local active speech is present. For example, the controller 20 can apply a scalar gain to the audio signal in response to determining that the output of the VAD indicates 1) a microphone signal generated by a microphone of a local device or audio output device contains the voice of a local user and / or 2) an accelerometer signal generated by an accelerometer has an energy level indicative of speech.

[0105] As described so far, the operation performed by the controller to maintain the audio quality of media content can be in response to detecting remote active voice and / or local active voice. In other words, these operations can be performed when a local user or a remote user is speaking. On the other hand, at least some operations in the operation to maintain audio quality can be performed in response to the signal level of a downlink signal and / or the noise level of a microphone signal produced by a microphone (such as microphone 23) coupled to a local device exceeding a threshold level. Specifically, these operations can be performed when a loud sound occurs at a remote device or a local device. Like this, for example, in response to a downlink signal or a microphone signal exceeding a signal level, the controller can generate and display closed captions, as described in process 60. In addition, when noise abates (for example, the signal level drops below the threshold), the controller 20 can rewind playback, as described in process 70.

[0106] When using an audio output device (e.g., a wireless headset) that is wirelessly connected to a media source device, streaming content such as music or movies requires the source device to stream high-quality audio to the audio output device over a wireless connection for output (e.g., to drive one or more speakers) in order to provide a good listener experience. In order to stream high-quality audio, most wireless headsets establish a one-way wireless audio connection with the source device that supports high bit rates and sampling rates. For example, two devices can establish a Bluetooth connection using a wireless profile that provides high-quality audio, such as A2DP. A2DP allows stereo audio to be streamed from a source device to a wireless headset and uses the SBC codec at a sampling rate of up to 48kHz.

[0107] When communicating with a source device that has initiated a call with another device and initiated a joint media playback session to stream media content, some audio output devices may not be able to support high-quality audio. For example, in order to allow wireless communication between an audio output device and a source device, two devices can establish a two-way wireless audio connection to exchange the audio signal associated with the call. However, these two-way wireless audio connections only provide low-quality audio streams to the audio output device. For example, both devices can use a wireless profile (such as HFP or HSP) that allows audio data to be exchanged between multiple devices to establish a Bluetooth connection. These profiles only support "voice quality" or low-quality audio exchanged between the two devices. For example, HFP traditionally only uses a codec with a sampling rate of 8kHz to 16kHz and is only capable of transmitting a mono audio signal. Although such low-quality streams may be sufficient for pure voice communication, when streaming media content in addition to making a call, such wireless connections may not provide enough audio quality. However, in one aspect, other audio output devices can be designed to support high-quality audio wireless transmission. For example, an audio output device may support a "high quality" two-way wireless audio connection using a wireless profile with a codec having a higher sampling rate (e.g., 24kHz). Therefore, it is necessary to switch between wireless audio connections when initiating a joint media playback session during a call based on the capabilities of the audio output device.

[0108] To overcome these drawbacks, the present disclosure describes a method and audio system for switching a wireless audio connection during a call. Specifically, the method can be performed by a local device 2 that is communicatively coupled to an audio output device 6 (e.g., using hands-free communication). For example, when participating in a call (e.g., a phone call or a video call) with a remote device, the local device communicates with the audio output device via a two-way wireless audio connection. The local device determines that a joint media playback session has been initiated, in which the local device and the remote device independently stream media content for playback by the two devices alone when participating in the call. Based on a determination of one or more capabilities of the audio output device (e.g., determining that the output device only supports low-quality audio streaming), the local device switches to communicating with a wireless headset via a one-way wireless audio connection, in which 1) a mixed content of one or more signals associated with the call and 2) an audio signal of the media content is transmitted to the wireless headset via the one-way wireless audio connection. Therefore, when participating in both a call and a joint media playback session, the audio output device can provide high-quality audio.

[0109] Figure 11A block diagram is shown according to one aspect in which a local device 2 is communicatively coupled to an audio output device 6 via a two-way wireless audio connection to exchange audio data while the local device is participating in a call with a remote device 3. Specifically, the figure shows the local device communicating with the audio output device via the two-way wireless audio connection while participating in a (e.g., hands-free) call with the remote device to exchange audio data for the call between the local device and the audio output device. This is illustrated by the local device's microphone 23 being deactivated (e.g., shown as struck-through) and the audio device's microphone 78 capturing sound (e.g., as shown by sound waves). In one aspect, the figure shows the two devices before (or after) a joint media playback session has been initiated.

[0110] As shown in the figure, two devices are coupled via the two-way wireless audio connection 80 communication ground that allows two devices to exchange audio data, as described herein.In one aspect, two-way connection can be any type of wireless connection that allows two devices to exchange audio data, such as HFP connection.In one aspect, two-way connection may be " low-quality " two-way wireless audio connection (low-quality wireless connection) or " high-quality " two-way wireless audio connection (high-quality wireless connection).In one aspect, low-quality wireless connection can be designed to support mono audio and / or transmit audio stream at a sampling rate less than a threshold sampling rate (for example, 24kHz).In some aspects, low-quality two-way connection can be traditional HFP or HSP connection, as described herein.In some aspects, high-quality audio connection can be designed to support stereo audio and / or transmit audio stream at a sampling rate of at least a threshold sampling rate.In one aspect, high-quality audio connection can be to use the Bluetooth connection with the wireless profile (for example, HFP) of the codec that transmits stereo audio stream at a threshold sampling rate or higher than a threshold sampling rate.

[0111] In one aspect, the audio quality of the wireless connection can be based on the capabilities (or characteristics) of the audio output device (and / or the local device). For example, during initiation of a two-way wireless audio connection, the audio output device can transmit device characteristics to the local device. In one aspect, these characteristics can indicate what type of wireless audio connection the audio output device can establish with the local device. For example, these characteristics can indicate which wireless profiles and / or audio codecs the audio output device supports. In one aspect, based on these characteristics, the local device can establish a two-way wireless audio connection.

[0112] To facilitate hands-free communication, the controllers 20 and 75 of the local device and the audio output device, respectively, both include one or more operational blocks. For example, the controller 20 includes an audio call manager 46 and a voice DSP 41, while the controller 75 includes an (optional) echo canceller 83. The controller 20 also includes a media playback manager 47, but since neither device is engaged in a joint media playback session, this operational block is inactive (as shown by the dashed border).

[0113] As described herein, the audio call manager is configured to initiate (and conduct) a call (e.g., by exchanging audio data of the call) between a local device 2 and one or more remote devices 3. Specifically, the manager receives a downlink audio signal from the remote device and transmits the microphone signal received from the audio output device to the remote device as an uplink audio signal. The voice DSP 41 is configured to receive the downlink audio signal from the audio call manager and is configured to perform audio signal processing (e.g., voice processing) operations on the signal to reduce (or eliminate) the noise contained therein. As described herein, the voice DSP can apply noise reduction to the downlink audio signal associated with the call. The audio output device transmits the (processed) downlink audio signal to the audio output device to drive the speaker 77 via a two-way wireless audio connection 80 (via network interfaces 21 and 76).

[0114] In one aspect, the audio output device may include an optional echo canceller 83 that is configured to receive a microphone signal captured by microphone 78 and to perform an echo cancellation operation to eliminate linear echo from the microphone signal. Specifically, the canceller may determine a linear filter based on the transmission path between microphone 78 and speaker 77 and apply the filter to the downlink audio signal to generate an echo estimate that will be subtracted from the microphone signal. In some aspects, the echo canceller may use any echo cancellation method. The (echo-canceled) microphone signal is then transmitted to the audio call manager 46 via the two-way wireless audio connection 80 for transmission to the remote device as an uplink audio signal.

[0115] Figure 12 A block diagram is shown according to one aspect in which a local device 2 is communicatively coupled to an audio output device 6 via a two-way wireless audio connection during a joint media playback session and a call with a remote device 3. Specifically, the diagram shows the result of the local device 2 initiating the joint media playback session while the local device and the audio output device participate in a hands-free call, as shown in FIG. Figure 5The initiation of a playback session is illustrated with a media playback manager 47 receiving media content (e.g., as an audio signal) from a media content server 5. In one aspect, the diagram may be similar to the diagram describing a local device communicatively coupled to an audio output device while simultaneously conducting a hands-free call and a joint media playback session. Figure 6 The figure also shows that the controller includes one or more additional operational blocks, such as a mixer 44, wireless audio connection switch decision logic 13, and a scalar gain 86 (which is optional).

[0116] In one aspect, the decision logic 13 is configured to determine whether to switch to a one-way wireless audio connection or (e.g., maintain) a two-way wireless audio connection to maximize the audio quality of the media content and the call, thereby providing an optimal user experience. Specifically, the decision logic determines that a joint media playback session has been initiated by receiving a control signal from the joint media playback session manager indicating that a (e.g., new) media session is to be established (e.g., between the local device and one or more remote devices). In one aspect, the decision logic determines whether to switch based on the capabilities of the audio output device (e.g., capabilities that may have been received during initialization of the two-way wireless audio connection 80), as described herein. For example, if the audio output device is determined not to support high-quality audio obtained using a two-way connection (e.g., based on an available audio codec having a sampling rate below a threshold rate, as described herein), the decision logic may switch the wireless connection to a one-way connection. Figure 13a and Figure 13b More is described about unidirectional connections. However, in this figure, the decision logic has determined that the audio output device supports high-quality audio. In this case, the local device has established a (e.g., high-quality) two-way wireless audio connection 81 for streaming high-quality audio. In one aspect, this connection may be established when a hands-free call is initiated (e.g., Figure 11 In this case, once a determination is made that an existing connection (e.g., between the local device and the audio output device during a hands-free call) provides high-quality audio, the local device can maintain a two-way connection with the audio output device. Thus, connections 80 and 81 can be the same connection.

[0117] In another aspect, rather than receiving characteristics from the audio output device, decision logic 13 may retrieve one or more characteristics based on the audio output device. Specifically, during initialization of a hands-free call, the audio output device may transmit a device identifier to the local device. The decision logic may use this identifier to perform a table lookup in a data structure that associates characteristics with device identifiers.

[0118] In one aspect, when initiating a joint media playback session, the local device can determine whether to switch to a one-way wireless audio connection or (e.g., maintain) a two-way wireless audio connection to maximize the audio quality of the media content and the call, thereby providing an optimal user experience. In one aspect, this determination can be based on the capabilities of the audio output device, as described herein. For example, if the audio output device does not support high-quality audio using a two-way connection (e.g., based on an available audio codec having a sampling rate below a threshold rate, as described herein), the local device can switch the wireless connection to a one-way connection. Figure 13a and Figure 13b More details are described regarding one-way connections. However, in this figure, the local device has determined that the audio output device supports high-quality audio. In this case, the local device has established a (e.g., high-quality) two-way wireless audio connection 81 for streaming high-quality audio. In one aspect, this connection may be established when a hands-free call is initiated (e.g., Figure 11 In this case, once a determination is made that an existing connection (e.g., between the local device and the audio output device during a hands-free call) provides high-quality audio, the local device can maintain a two-way connection with the audio output device. Thus, connections 80 and 81 can be the same connection.

[0119] In one aspect, when carrying out joint media playback session and calling, local equipment can stop performing one or more operations and start to perform one or more audio processing operations to the downlink signal of calling and / or the audio signal of media content.For example, controller 20 comprises mixer 44 and scalar gain 86 (it is optional), wherein mixer 44 receives the audio signal of the media content from media playback manager 47 and the downlink audio signal from call manager 46, rather than voice DSP 41 receiving the downlink audio signal.In one aspect, controller can be in response to switching to stop performing voice DSP operation (for example, stop applying noise reduction to downlink audio signal) with audio output device communication via unidirectional connection, so that the media content of downlink signal and the more complete spectrum content of audio content are provided.As described herein, mixer is configured to perform matrix mixing operation to generate the mixed content of signal.Scalar gain 86 is configured to receive mixed content, and is configured to scalar gain is applied to mixing, so that the signal level of mixing is reduced. In one aspect, the scalar gain may be applied for a period of time after the joint media playback session is initiated (or after the controller 20 switches to communicate with the audio output device via the one-way wireless audio connection). After this period of time, the scalar gain may be reduced (or removed) so that the gain is no longer applied to the mixed content. In one aspect, the scalar gain may be incrementally reduced for a second period of time to provide an attenuation effect. The mixed content is then transmitted to the audio output device via the two-way wireless audio connection 81 for driving the speaker 77, as described herein.

[0120] Figure 13a and Figure 13b Several block diagrams are shown according to one aspect in which a local device 2 communicatively coupled to an audio output device 6 for exchanging audio data switches between wireless audio connections based on initiation of a joint media playback session. Specifically, Figure 13a 85 is a block diagram showing a local device and an audio output device coupled via a one-way wireless audio connection 85. Specifically, the diagram shows the result of a local device 2 initiating a joint media playback session while participating in a call. However, unlike the example of ... Figure 12 , which shows that the local device has switched to a one-way wireless audio connection 85 in order to stream high-quality audio data to the audio output device for output (e.g., via speaker 77).

[0121] In one aspect, the switch (or conversion) from a two-way connection to a one-way wireless audio connection can be based on the audio output device, as described herein. For example, the decision logic 13 can determine (e.g., in response to receiving a control signal from the session manager 47) that the audio output device does not support the exchange of audio signals via the two-way wireless audio connection at a sampling rate of at least a threshold sampling rate. As described herein, this determination can be based on characteristics received from the audio output device, or based on performing a table lookup in a data structure using a device identifier. In one aspect, the decision logic can determine to switch to the one-way wireless audio connection based on not receiving characteristics from the device and / or not identifying the device within the data structure (e.g., the decision to switch can be the default decision of the decision logic).

[0122] In one aspect, local device 2 and audio output device can perform one or more operations to convert from two-way connection 80 to one-way wireless audio connection 85. For example, local device 2 (or audio output device 6) can disconnect (or terminate) two-way wireless audio connection 80. Once disconnected, local device can establish one-way wireless audio connection (e.g., Bluetooth A2DP connection) with audio output device. In one aspect, due to disconnecting the two-way connection for the one-way connection where audio data can only be transmitted from local device to audio output device, the controller can be configured to activate one or more other microphones to capture the local user voice for uplink audio signal. Specifically, the controller can transmit a signal to the audio output device to mute microphone 78 (as shown by the deleted line), and can activate the microphone 23 of the local device to capture the voice of the local user. In one aspect, the activated microphone can be a component of different electronic devices. Therefore, the microphone signal of microphone 23 can be transmitted to the remote device as an uplink audio signal. This article describes more about being performed by the controller for switching wireless audio connection.

[0123] In one aspect, the controller 20 may (optionally) perform an echo cancellation estimation operation on the microphone signal generated by the microphone 23. Specifically, the controller 20 includes an echo cancellation estimator 87 configured to perform an echo cancellation operation to cancel the echo from the microphone signal. In one aspect, the estimator may perform the same Figure 11 . For example, when two devices are involved in a call, the estimator may obtain a microphone signal from the local device that is to be transmitted to the remote device. The estimator is configured to generate an estimate of a portion of one or more (e.g., downlink audio) signals associated with the call. For example, the estimator may determine a linear filter based on the transmission path between microphone 23 and speaker 77. In one aspect, unlike the definable transmission path between microphone 78 and speaker 77 (e.g., based on the integration of both the microphone and speaker into an audio output device at predefined locations), the transmission path between the local microphone 23 and speaker 77 of the audio output device may not be predefined. Therefore, the estimator may estimate the transmission path. For example, the estimator may determine the distance between microphone 23 and speaker 77 based on the arrival time of sound produced by speaker 77 captured by microphone 23. In another aspect, the estimator may estimate the path based on the received signal strength (RSSI) of the wireless audio connection. In some aspects, the estimator may use any sound localization method to determine the location of speaker 77 and, therefore, the path from the speaker to the microphone. In another aspect, the transmission path may be predefined (e.g., a path determined in a controlled environment such as a laboratory). Using the estimate of the transmission path, a linear filter is determined that is applied to the downlink audio signal to generate an echo estimate that is subtracted from the microphone signal, as described herein.

[0124] In one aspect, the wireless audio connection switching decision logic 13 can be configured to switch between a one-way wireless audio connection 85 and a two-way wireless audio connection when conducting a joint media playback session and call. In one aspect, the decision logic can switch to a high-quality two-way wireless audio connection (e.g., Figure 12 81 in the example embodiment). On the other hand, when the audio output device does not support a high-quality, two-way wireless audio connection, the decision logic may switch the one-way wireless audio connection to a low-quality, two-way wireless audio connection to provide hands-free communication with the audio output device, as described herein. While less preferred than a one-way wireless audio connection due to lower audio quality, in some cases, such functionality may be required or desired based on one or more standards. Figure 13b A switch to a low-quality bidirectional connection is described.

[0125] In one aspect, switching to a two-way wireless audio connection can be based on the location of the local device 2 and / or audio output device 6. For example, as described herein, when switching to a one-way wireless audio connection, the location of the microphone used during the call and before initiating the joint media playback session can be located at the audio output device, which can be a wireless headset worn on the user's head. However, once the one-way connection is initiated, the location of the (e.g., active) microphone can be changed to a different microphone (e.g., microphone 23 of the local device) that may be separated from the audio output device. Therefore, the microphone and speaker used during the call and joint media playback session can be a component of different electronic devices, each device being located at a different location. Therefore, in order to participate in the call and joint session, the local user may be required to bring the local device and the audio output device into close proximity (e.g., in order for the microphone to capture the user's voice and in order for the user to hear the sound produced by the speaker of the audio output device). In one aspect, the decision logic can receive sensor data from one or more sensors 40 and can be configured to determine whether the local device is separated from the audio output device by a threshold distance. For example, the decision logic can receive image data from one or more cameras (e.g., camera 24) and use the image data to determine the location of the audio output device by using an image recognition algorithm. In another aspect, the decision logic may determine the location of the audio output device based on the RSSI of the unidirectional connection. For example, in response to determining that the RSSI is below a threshold, the decision logic may switch to a bidirectional connection. This is because the user may be too far away from the new active microphone to clearly pick up the local user's voice.

[0126] In another aspect, the decision may be based on whether the local user is in front of (or next to) the display screen 25 of the local device. For example, the camera 24 may be positioned adjacent to the display screen and have a field of view in front of the display screen. The decision logic may receive image data from the camera and execute an image recognition algorithm to determine whether the user is present (e.g., in front of the display screen). If not, the decision logic may perform the switch. In some aspects, the decision logic may make this determination based on other sensor data (such as proximity sensor data). In this case, one or more proximity sensors may be arranged to determine whether an object is within a threshold distance from the display screen 25. If not, this indicates that the local user is not in front of the display screen, and the decision logic may perform the switch.

[0127] On the other hand, decision logic 13 may perform a switch based on whether the object is within a threshold distance from the local device (e.g., the local device's microphone 23). For example, when the local device is a smartphone, the user may place the smartphone in their pocket. In this case, the microphone may capture the user's muffled voice. Thus, decision logic may receive sensor data indicating whether the object is within the threshold distance. For example, the sensor may be a proximity sensor. In response to the object being within this distance, decision logic may perform a switch.

[0128] In some aspects, the decision logic may perform the switch based on whether the local user is speaking. For example, at times when the local user is not speaking, a microphone may not be necessary, and therefore a one-way wireless connection may be established to provide high quality audio. However, in response to determining that the local user is speaking, the decision logic may perform the switch. For example, the decision logic may receive a control signal from the audio output device in response to the local user speaking, and may perform the switch based on the received control signal. For example, when the control signal is a VAD signal generated by the VAD 82 of the audio output device in response to detecting a high energy level of the accelerometer signal from the accelerometer 79, the decision logic may determine that the local user is speaking. On the other hand, the decision logic may receive a control signal from the VAD of the local device (e.g., such as Figure 5 The VAD 42 shown in FIG ) receives a VAD signal that can be configured to detect the local user's voice based on signals received from the audio output device (such as one or more accelerometer signals and / or one or more microphone signals). Once the user speaks, the decision logic can switch to a two-way wireless audio connection and activate the microphone 78 of the output device to capture the user's voice. Once the user finishes speaking (e.g., the VAD signal indicates that the user's voice is no longer detected), the decision logic can switch back to a one-way audio connection.

[0129] Figure 13b A block diagram is shown in which a local device and an audio output device have switched to a two-way wireless audio connection while conducting a joint media playback session and call as described herein. Specifically, the figure shows the result of decision logic 13 switching to a two-way wireless audio connection (e.g., based on one or more standards) during the call and playback session. As shown, the two-way wireless audio connection 89 is a low-quality connection, which may be due to the fact that the audio output device does not support high-quality connections, as described herein. In addition to switching to a two-way connection, the local device and the audio output device have restored the (active) position of the microphone from the local device to the audio output device.

[0130] like Figure 12 、 Figure 13a and Figure 13bAs described, the local device may participate in a joint media playback session in which one or more audio signals of media content (e.g., a musical composition) are received for playback. In one aspect, the operations performed in these figures may occur when the local device participates in a joint playback session in which multimedia content is being played back, e.g., video is displayed on the display screen 25 and audio is output by the speaker 77. In addition, the controller 20 and / or the controller 75 may also perform at least some of the other operations described herein.

[0131] Figures 14 to 18 Flowcharts of processes 90, 100, 110, 130, and 120, respectively, that perform one or more operations for switching a wireless audio connection during a call. In one aspect, at least some of the processes may be performed by one or more devices of the audio system 1, such as Figure 1 For example, at least processes 90, 100, and 110 are performed by local device 2 (e.g., its controller 20), and processes 130 and 120 are performed by audio output device 6 (e.g., its controller 75). In another aspect, any of the devices may perform any of the operations described herein.

[0132] Figure 14is a flow chart of one aspect of a process 90 for switching between wireless audio connections. In one aspect, the process may be performed by the controller 20 of the local device 2. Process 90 begins with the controller initiating a call between the local device and a remote device (at block 91). For example, the call manager 46 may initiate a call (e.g., a phone or video call) between the local device and one or more remote devices, as described herein. When participating in a call with the remote device, the controller 20 communicates with the audio output device via a two-way wireless audio connection (at block 92). Specifically, the local device 2 may establish a wireless connection with the audio output device via a wireless communication link (e.g., via the Bluetooth protocol or any other wireless communication protocol). For example, the local device may communicate with the audio output device to configure a Bluetooth stack, which is executed within the audio output device, to exchange audio data between the devices via the two-way wireless audio connection (e.g., by negotiating a codec for decoding and encoding audio signals exchanged between the devices). During this time, the audio output device may transmit a message indicating its capabilities (e.g., the audio codecs it supports, etc.). In one aspect, based on these capabilities, the local device may establish a two-way wireless audio connection. Specifically, if capable of supporting high-quality audio streaming (e.g., at a sampling rate of at least a threshold sampling rate), the local device can establish a high-quality, two-way wireless audio connection, as described herein. Once established, the local device can transmit one or more (e.g., downlink audio) signals associated with the call to the audio output device and receive one or more microphone signals for the call via the two-way connection. On the other hand, the device can establish a low-quality wireless audio connection regardless of the capabilities of the audio output device, as only voice data is exchanged between the devices.

[0133] The controller 20 determines that a joint media playback session has been initiated, in which the local device and the remote device independently stream media content for separate playback by the two devices while participating in the call (at block 93). Specifically, the joint media playback session manager 47 may have received a user request from a local user (e.g., via a UI displayed on the display screen 25), or may have received a request from the media content server 5 indicating that one or more remote devices have requested to initiate a playback session.

[0134] The controller 20 determines whether the audio output device supports exchanging audio signals of calls and media content with local devices via (e.g., high-quality) two-way wireless audio connections. (At decision block 94) Specifically, the wireless audio connection switching decision logic 13 can, for example, switch from (e.g., currently established) two-way wireless audio connections to one-way wireless audio connections based on one or more capabilities of the audio output device 6. For example, the decision logic can determine whether the audio output device supports high-quality audio based on a table lookup of a data structure that associates characteristics with device identifiers. In one aspect, since a two-way wireless audio connection has been established, the decision logic can determine the type of connection already present between the two devices (e.g., whether the connection is an HFP connection using a codec with a sampling rate higher than a threshold rate and / or whether the HFP connection supports stereo audio). If so, the controller communicates with the audio output device via (e.g., high-quality) two-way wireless audio connections (at block 95) when participating in a call and during a joint media playback session. In one aspect, if the initial wireless audio connection is a low-quality connection, the controller can disconnect the connection and establish a high-quality two-way wireless audio connection. However, if the initially established two-way wireless audio connection is a high quality connection, the controller may maintain the existing connection.

[0135] However, if the audio output device does not support a high-quality two-way wireless audio connection, the controller 20 switches to communicating with the audio output device via a one-way wireless audio connection (e.g., based on one or more capabilities of the audio output device, as described herein), wherein a mixed content of the audio signal of the one or more signals associated with the call and the media content is transmitted to the audio output device via the one-way wireless audio connection (at block 96). Specifically, as described herein, the controller 20 may disconnect the two-way wireless audio connection and establish a one-way connection. Once established, the controller may stream the media content and the downlink audio signal of the call to the audio output device for playback. Figure 15 More details are described regarding operations for switching wireless audio connections.

[0136] Figure 15 1 is a flow chart of another aspect of a process 100 for switching between wireless audio connections. In one aspect, at least some of the operations performed in process 100 may be performed by controller 20 when (and / or after) switching to communicating with an audio output device via a one-way wireless audio connection, such as Figure 1496 of . Process 100 begins with the controller transmitting a signal to mute the microphone (e.g., microphone 78) of the audio output device (at box 101). Specifically, the controller can transmit a control signal to the audio output device via a two-way wireless audio connection, allowing the controller 75 to mute the microphone 78. In one aspect, muting the controller 75 can mute the microphone 78 by stopping the transmission of the microphone signal generated by the microphone to the local device. In this case, the microphone 78 can continue to generate microphone signals, and the controller 75 can use these microphone signals to perform one or more operations (e.g., perform ANC functions, transparency functions, etc.). The controller 20 switches from the two-way wireless audio connection to the one-way wireless audio connection (at box 102). As described herein, the one-way wireless audio connection can be any wireless connection that provides high-quality audio (e.g., an A2DP connection). In one aspect, the one-way connection can be based on the capabilities of the audio output device.

[0137] The controller 20 provides a notification (at box 103) indicating that the microphone of the audio output device is muted and / or requests user authorization to activate a different microphone. For example, the controller may display the notification as a pop-up notification on the display screen 25 of the local device 2, thereby reminding the local user that the microphone has been muted. In one aspect, this will remind the user so that the user does not start speaking before the microphone is activated. In some aspects, the notification may also indicate the new location of the microphone. Specifically, the notification may indicate that the location of the microphone may be at the local device. In one aspect, the notification may also request user authorization to activate a different microphone (e.g., by displaying a UI item within the pop-up notification).

[0138] Controller 20 starts to play back the media of the joint media playback session (at box 104). Specifically, controller 20 can start to transmit one or more audio signals of media content to audio output device via unidirectional connection, and audio output device can use these signals to drive one or more speakers. In addition, when media content includes video, controller can display video signal on display screen 25. Controller determines whether user has authorized to switch microphone (at decision block 105). For example, controller can determine whether user has selected UI item displayed in pop-up notification. If not, then controller can continue to play back media content, and the microphone of no local device and / or audio output device is in active state to capture the user's voice of the uplink signal for calling. However, if controller receives user authorization, then controller activates different microphones and starts to receive microphone signals to be transmitted to remote device (for example, as uplink signal) for calling (at box 106).

[0139] In one aspect, the controller may provide the user with a selection of microphones that the user can activate for the call. For example, a pop-up notification may display a list of microphones and their locations so that the local user can make a decision about which microphone to use during the call. In another aspect, the user may be provided with the option of having the local device continue to communicate with the audio output device via a two-way wireless audio connection. For example, the controller may provide a notification requesting user authorization to switch from a two-way wireless audio connection to a one-way wireless audio connection. If the user fails to provide a response (and / or does not provide authorization by selecting a UI item), the controller may continue to communicate within the two-way wireless audio connection, which, as described herein, may be a low-quality connection based on the capabilities of the audio output device.

[0140] Figure 16 1 is a flow chart of one aspect of a process 110 for determining whether to switch between wireless audio connections based on one or more criteria. Specifically, the process is used to determine whether to switch from communicating with an audio output device via a one-way wireless audio connection to communicating with the device via a (e.g., low-quality) two-way wireless audio connection. Process 110 begins with controller 20 communicating with the audio output device via a one-way wireless audio connection, such as during a call and joint media playback session, as described herein (at block 111). Controller 20 receives sensor data from at least one sensor (at block 112). For example, the controller may receive sensor data from a proximity sensor, a light sensor, a microphone (e.g., microphone 23), a camera (e.g., camera 24), and the like. Based on the sensor data, controller 20 determines whether to switch to communicating with the audio output device via a two-way wireless audio connection (at decision block 113). As described herein, the controller may use sensor data, such as proximity data from a proximity sensor, to determine whether an object is within a threshold distance. In response to being within the threshold distance, controller 20 switches to communicating with the audio output device via the two-way wireless audio connection (at block 114). As described herein, the bidirectional connection may be a low quality (eg, legacy 8kHz HFP) connection, based on the capabilities of the audio output device.

[0141] However, if the controller determines not to switch based on the sensor data, the controller then determines whether the local device has received a user request to switch to a two-way wireless audio connection (at decision block 115). For example, the local device may display a UI item on the display screen 25 that allows the local user to switch to a two-way wireless audio connection. In one aspect, the user may wish to switch to a two-way connection for various reasons. For example, when the user's environment has ambient noise, the user may wish to use the onboard microphone of the audio output device. If so, the controller proceeds with switching the connection.

[0142] If not, the controller determines the signal strength of the one-way wireless audio connection (at box 116). For example, the controller may determine the RSSI of the connection. The controller determines whether the signal strength is above a threshold (at decision box 117). If not, the controller may continue to switch the connection. In one aspect, the signal strength may be lower due to the user walking away from the local device while continuing to wear the audio output device. For example, when the local device is a desktop computer with an onboard microphone for picking up the user's voice for a call, if the user walks away, the controller may perform a switch so as to keep the active microphone within the user's distance. If the signal strength is above the threshold, the controller may continue to communicate with the audio output device via the one-way wireless audio connection (at box 118).

[0143] In one aspect, the controller may switch back to a one-way wireless audio connection when at least one of the conditions that caused the controller to switch ends. For example, when communicating with an audio output device via a two-way wireless audio connection, the controller may switch back to the one-way wireless audio connection upon determining that the signal strength is above a threshold. Continuing with the previous example, when the signal strength is above the threshold, it may be determined that the user is now in front of the desktop computer.

[0144] Figure 17Flowchart of one aspect of a process 130 for switching between wireless audio connections performed by audio output device 6 (e.g., its controller 75). Process 130 begins with the controller 75 communicating with the local device via a two-way wireless audio connection during a call between local device 2 and remote device 3 (at block 131). For example, the audio output device may perform hands-free communication with the local device during the call, as described herein. The controller 75 determines whether a one-way wireless audio connection will be established between the local device and the audio output device during the call, instead of a two-way wireless audio connection (at block 132). For example, this determination may be based on whether the two-way connection can support high audio quality. In one aspect, an existing two-way connection may support exchanging audio signals at a sampling rate lower than that supported by a one-way connection. For example, a two-way connection may be an HFP connection supporting a sampling rate of 8kHz to 16kHz, while a one-way connection may be an A2DP connection supporting a sampling rate of 48kHz. In one aspect, the audio output device may receive a control signal (e.g., from the local device) indicating that the two-way wireless audio connection is to be disconnected. The controller 75 mutes the microphone of the audio output device (at block 133). As described herein, controller 75 can deactivate microphone and / or stop transmitting microphone signal to local device.Controller 75 switches to unidirectional wireless audio connection (at frame 134) from two-way wireless audio connection.For example, audio output device can disconnect two-way connection and transmit confirmation message indicating that connection has been disconnected to local device.Subsequently, audio output device can receive communication from local device to set up two-way wireless audio connection.In response, audio output device can set up connection.Controller 75 receives audio signal by unidirectional wireless audio connection, and this audio signal comprises the mixed content (at frame 135) of the signal associated with call and the media content that is played back by local device and remote device in joint media playback session.Controller can use audio signal subsequently to drive the loudspeaker (for example, loudspeaker 77) of audio output device (at frame 136).

[0145] Figure 181 is a flow chart of one aspect of a process 120 performed by an audio output device 6 for switching from a one-way wireless audio connection to a two-way wireless audio connection based on whether voice is detected. In one aspect, prior to performing process 120, the audio output device 6 may be communicatively coupled to a local device via a one-way connection to receive audio data of media content played back by the local device during a joint media playback session concurrent with a call, as described herein. For example, the audio output device may receive an audio signal via a one-way connection that includes a mixed content of 1) a signal of a phone (or video) call and 2) a signal associated with the media content, wherein the local device and the remote device are simultaneously participating in the call and the joint media playback session. In addition, the audio output device may use the audio signal to drive a speaker. Process 120 begins with the controller 75 receiving an accelerometer signal from an accelerometer (e.g., accelerometer 79) of the audio output device (at block 121). The controller 75 generates a VAD signal (e.g., as an output of VAD 82) based on the accelerometer signal (at block 122). As described herein, the VAD signal may indicate that user voice has been detected based on the energy level of the accelerometer. The controller 75 determines whether the VAD signal is above a threshold, thereby indicating that user speech is detected (at decision block 123). If not, the audio output device continues to communicate with the local device via the one-way wireless audio connection (at block 124).

[0146] Otherwise, the controller 75 switches to communicating with the local device via the two-way wireless audio connection at block 125. The controller 75 receives a microphone signal from the microphone of the audio output device at block 126. The controller 75 then transmits the microphone signal to the local device via the two-way wireless audio connection for transmission to the remote device as an uplink signal, as described herein at block 127.

[0147] Some aspects can be Figures 14 to 18 The processes 90, 100, 110, 130, and 120 described in the foregoing may be performed in various ways. For example, certain operations in at least some of these processes may not be performed in the exact order shown and described. Certain operations may not be performed in a continuous series of operations, and different certain operations may be performed in different aspects. For example, operations within dashed boxes may be optional operations that may not be performed when performing the corresponding process. For example, in Figure 15 In the process 100 of , no notification needs to be provided. Instead, playback of the media content may begin (at block 104 ), and a different microphone may be activated (at block 106 ) in response to the connection being switched.

[0148] It is understood that the use of personally identifiable information should be subject to privacy policies and practices that are generally recognized to meet or exceed industry or government requirements for maintaining user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use, and the nature of authorized use should be clearly stated to users.

[0149] As previously mentioned, one aspect of the present disclosure may include a non-transitory machine-readable medium (such as a microelectronic memory) having stored thereon instructions that program one or more data processing components (generally referred to herein as "processors") to perform network operations and audio signal processing operations, as described herein. In other aspects, some of these operations may be performed by specific hardware components containing hard-wired logic. Alternatively, those operations may be performed by any combination of programmed data processing components and fixed hard-wired circuit components.

[0150] Although certain aspects have been described and illustrated in the accompanying drawings, it should be understood that such aspects are merely illustrative of the broad disclosure and not restrictive, and that the disclosure is not limited to the exact constructions and arrangements shown and described, as various other modifications may occur to those skilled in the art. Accordingly, the description is to be regarded as illustrative and not restrictive.

[0151] In some aspects, the present disclosure may include language such as “at least one of [element A] and [element B]”. The language may refer to one or more of these elements. For example, “at least one of A and B” may refer to “A”, “B”, or “A and B”. Specifically, “at least one of A and B” may refer to “at least one of A and at least one of B” or “at least either A or B”. In some aspects, the present disclosure may include language such as “[element A], [element B], and / or [element C]”. The language may refer to any one of these elements or any combination thereof. For example, “A, B, and / or C” may refer to “A”, “B”, “C”, “A and B”, “A and C”, “B and C”, or “A, B, and C”.

Claims

1. A method, performed by a first electronic device, for processing remote active voice during a call, comprising: Initiating a call with a second electronic device; During the call, initiating a joint media playback session wherein the first electronic device and the second electronic device independently stream media content for synchronized playback; determining, based on an output from a voice activity detector (VAD), that a downlink signal from the second electronic device includes voice; In response to determining that the downlink signal includes speech, applying a scalar gain to an audio signal of the media content to reduce a signal level of the audio signal; driving a speaker with a mixture of the downlink signal and the audio signal; as well as In response to determining, based on subsequent output from the VAD, that the downlink signal has ceased to include the speech, rewinding playback of the media content by: Pause playback of the media content, and Playback of the media content is initiated at a time along the playback duration of the media content at which the output of the VAD begins to indicate that the downlink signal includes the speech.

2. The method according to claim 1, further comprising executing a noise reduction algorithm on the downlink signal to reduce noise contained therein; and An output of the VAD is generated based on the downlink signal. 3 . The method of claim 1 , further comprising receiving an output of the VAD from the second electronic device.

4. The method of claim 1 , wherein the first electronic device is communicatively coupled to a wireless headset to conduct the call and the joint media playback session, wherein the method further comprises generating an output of the VAD based on an accelerometer signal generated by an accelerometer of the wireless headset.

5. The method of claim 1 , wherein the media content comprises a video signal and the audio signal, wherein initiating the joint media playback session comprises displaying the video signal on a display screen and driving the speaker with the mixture of the downlink signal and the audio signal.

6. The method according to claim 5, further comprising: determining a signal level of the downlink signal; as well as In response to the signal level being above a threshold level or in response to determining based on the output from the VAD that the downlink signal includes speech, closed captions representing audio content contained within the audio signal of the media content are displayed on the display screen.

7. The method according to claim 1, wherein the time is a first timestamp along the playback duration of the media content at which an output from the VAD begins to indicate that the downlink signal includes the speech; wherein the method further comprises determining a second time stamp subsequent to the first time stamp along the playback duration of the media content, at which second time stamp, a determination is made that an output from the VAD indicates that the downlink signal has ceased to include the speech, Wherein the playback is paused at or after the second timestamp and the playback is started at or before the first timestamp.

8. The method according to claim 1, wherein the time is a first timestamp along the playback duration of the media content at which the output from the VAD begins to indicate that the downlink signal includes the speech, wherein the method further comprises: determining a second time stamp subsequent to the first time stamp along the playback duration of the media content, at which second time stamp a determination is made that an output from the VAD indicates that the downlink signal has ceased to include the speech; as well as Responsive to the determination that the output from the VAD indicates the downlink signal has ceased to include speech, providing a notification requesting user authorization to rewind playback of the media content.

9. The method of claim 8, wherein the notification is a pop-up notification displayed on a display screen of the first electronic device.

10. A first electronic device, comprising: at least one processor; as well as A memory storing instructions that, when executed by the at least one processor, cause the first electronic device to Initiating a call with a second electronic device; During the call, initiating a joint media playback session wherein the first electronic device and the second electronic device independently stream media content for synchronized playback; determining, based on an output from a voice activity detector (VAD), that a downlink signal from the second electronic device includes voice; In response to determining that the downlink signal includes speech, applying a scalar gain to an audio signal of the media content to reduce a signal level of the audio signal; driving a speaker with a mixture of the downlink signal and the audio signal; as well as In response to determining, based on subsequent output from the VAD, that the downlink signal has ceased to include the speech, rewinding playback of the media content by: Pause playback of the media content, and Playback of the media content is initiated at a time along the playback duration of the media content at which the output of the VAD begins to indicate that the downlink signal includes the speech.

11. The first electronic device of claim 10, wherein the memory has further instructions to: executing a noise reduction algorithm on the downlink signal to reduce noise contained therein; and An output of the VAD is generated based on the downlink signal.

12. The first electronic device of claim 10, wherein the memory has further instructions to receive an output of the VAD from the second electronic device.

13. The first electronic device of claim 10 , wherein the first electronic device is communicatively coupled to a wireless headset to conduct the call and the joint media playback session, wherein the memory has further instructions to generate an output of the VAD based on an accelerometer signal generated by an accelerometer of the wireless headset.

14. The first electronic device of claim 10 , further comprising a display screen, wherein the media content comprises a video signal and the audio signal, wherein initiating the joint media playback session comprises displaying the video signal on the display screen and driving the speaker with the mixture of the downlink signal and the audio signal.

15. The first electronic device of claim 14, wherein the memory has further instructions to: determining a signal level of the downlink signal; and In response to the signal level being above a threshold level or in response to determining based on the output from the VAD that the downlink signal includes speech, closed captions representing audio content contained within the audio signal of the media content are displayed on the display screen.

16. The first electronic device according to claim 10, wherein the time is a first timestamp along the playback duration of the media content at which an output from the VAD begins to indicate that the downlink signal includes the speech; wherein the memory has further instructions to determine a second time stamp subsequent to the first time stamp along the playback duration of the media content, at which second time stamp a determination is made that an output from the VAD indicates that the downlink signal has ceased to include the speech, Wherein the playback is paused at or after the second timestamp and the playback is started at or before the first timestamp.

17. The first electronic device according to claim 10, wherein the time is a first timestamp along the playback duration of the media content at which the output from the VAD begins to indicate that the downlink signal includes the speech and wherein the memory has further instructions to: determining a second time stamp subsequent to the first time stamp along the playback duration of the media content, at which second time stamp a determination is made that an output from the VAD indicates that the downlink signal has ceased to include the speech; and Responsive to the determination that the output from the VAD indicates the downlink signal has ceased to include speech, providing a notification requesting user authorization to rewind playback of the media content.

18. The first electronic device of claim 17, wherein the notification is a pop-up notification displayed on a display screen of the first electronic device.

19. A method performed by a first electronic device, the method comprising: conducting a video conference call and a joint media playback session simultaneously with a second electronic device; determining, based on audio content of the video conference call, that a user of the second electronic device begins speaking; In response to determining that the user begins speaking, reducing a volume level of audio content of media content associated with the joint media playback session; as well as In response to determining that the user of the second electronic device has stopped speaking based on the audio content of the video conference call, rewinding playback of the media content to restart playback at a moment along the playback duration before the user of the second electronic device began speaking.

20. The method of claim 19, further comprising displaying closed captions representing the audio content of the media content associated with the joint media playback session on a display screen of the first electronic device in response to determining that the user of the second electronic device begins to speak.

21. The method of claim 19, further comprising: determining, based on the audio content of the video conference call, that the user of the second electronic device has stopped speaking; as well as In response to determining that the user has stopped speaking, a volume level of the audio content of the media content is increased to a previous level before the decrease in the volume level.

22. The method of claim 19, wherein the user begins speaking at a first moment along the playback duration of the media content, wherein the method further comprises determining, based on the audio content of the video conference call, that the user of the second electronic device stopped speaking at a second time along the playback duration of the media content, the second time being subsequent to the first time; and Wherein the playback is restarted at or before a first moment along the playback duration.

23. The method of claim 19, wherein the user begins speaking at a first moment along the playback duration of the media content, wherein the method further comprises determining, based on the audio content of the video conference call, that the user of the second electronic device stopped speaking at a second time along the playback duration of the media content, the second time being subsequent to the first time; and In response, a notification is provided requesting user authorization to rewind playback of the media content.

24. The method of claim 23, wherein the notification is a pop-up notification displayed on a display screen of the first electronic device.

Citation Information

Patent Citations

  • Method and apparatus for displaying words service in case of mute audio

    CN101112082A

  • Method and device for achieving synchronous film watching at different places, and intelligent device

    CN105872835A

  • System and method for performing automatic gain control using an accelerometer in a headset

    US20170263267A1