Method and system for volume control
By implementing single volume control and gain adjustment in smart devices, combined with volume-gain curves and voice activity detectors, the challenge of volume control in joint media playback sessions is solved, improving the consistency of sound output and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-11
- Publication Date
- 2026-03-17
AI Technical Summary
In existing technologies, smart devices struggle to effectively control volume levels during joint media playback sessions, resulting in inconsistent sound output and a poor user experience.
By implementing a single volume control in the local device, combined with gain adjustment, the volume levels of the downlink signal and audio signal are dynamically adjusted. Volume control is optimized using volume-gain curves and a voice activity detector (VAD), enabling mixed-signal driving of the speaker.
It enables precise volume control in joint media playback sessions, improving the user experience and ensuring the consistency and adaptability of sound output.
Smart Images

Figure CN115622980B_ABST
Abstract
Description
Technical Field
[0001] One aspect of this disclosure relates to an audio system for controlling volume. Other aspects are also described. Background Technology
[0002] Many devices today, such as smartphones, can engage in various types of telecommunications activities with other devices. For example, a smartphone can make a phone call to another device. When a phone number is dialed, the smartphone connects to a cellular network, which can then connect the smartphone to another device (such as another smartphone or a landline). Furthermore, smartphones can also make video conferencing calls, in which video and audio data are exchanged with another device. Summary of the Invention
[0003] One aspect of this disclosure is a method performed by a first electronic device (e.g., a local device). For example, while participating in a call (e.g., a telephone call (or "audio only") or video call) with a second electronic device (e.g., a remote device), the local device initiates a joint media playback session in which the two devices independently stream media content for synchronized playback. The local device drives a speaker at a total volume level using a mixture of the downlink signal of the call and the audio signal of the media content. The local device receives a user adjustment to reduce the total volume level at a single volume control (e.g., a master volume control) of the local device, and in response to the user adjustment, applies a first gain adjustment to the downlink signal, a second gain adjustment to the audio signal, and drives the speaker at a reduced volume level using a mixture of these signals.
[0004] In one aspect, the individual volume control is the master volume control of the first electronic device, which is configured to provide bidirectional control to gradually increase or decrease the overall volume level. In another aspect, the master volume control is a physical control that is part of the first electronic device. In some aspects, the master volume control is a user interface (UI) item displayed on the screen of the first electronic device. In one aspect, the individual volume control includes input that includes gestures made by the user of the first electronic device.
[0005] In one aspect, a single volume control includes a plurality of volume settings, each volume setting defining a different total volume level of a first electronic device, wherein a user adjustment at the single volume control changes the current volume setting of the signal volume control to a new volume setting associated with a reduced total volume level. In some aspects, a downlink signal is associated with a first volume-gain curve that associates the plurality of volume settings with a first plurality of gains, and an audio signal of media content is associated with a second volume-gain curve that associates the plurality of volume settings with a second plurality of gains. The method further includes: in response to receiving a user adjustment, determining a first gain adjustment based on a first gain associated with the new volume setting using the first volume-gain curve, and determining a second gain adjustment based on a second gain associated with the new volume setting using the second volume-gain curve. In some aspects, the first volume-gain curve and the second volume-gain curve are linear functions of the gain relative to the plurality of volume settings of the single volume control, wherein the slope of the first volume-gain curve is greater than the slope of the second volume-gain curve, such that at each volume setting, the gain on the first volume-gain curve is lower than the gain on the second volume-gain curve. In one aspect, the first volume-gain curve and the second volume-gain curve are non-linear functions of the gain relative to the volume settings of the single volume control.
[0006] In one aspect, prior to the application of the first gain adjustment and the second gain adjustment, the signal level of the audio signal is greater than the signal level of the downlink signal, and the second gain adjustment is greater than the first gain adjustment, such that 1) the signal level of the gain-adjusted audio signal and the signal level of the gain-adjusted downlink signal are lower than the signal level of the downlink signal, and 2) the signal level of the gain-adjusted downlink signal is greater than the signal level of the gain-adjusted audio signal. In another aspect, the user adjustment is a first user adjustment, the first gain adjustment is a first attenuation, and the second gain adjustment is a second attenuation. The method further includes receiving a second user adjustment for a single volume control of the first electronic device that increases the reduced total volume level back to the total volume level; 1) applying the first gain to the gain-adjusted downlink signal, and 2) applying the second gain to the gain-adjusted audio signal, the first gain and the second gain respectively increasing the signal levels of the gain-adjusted downlink signal and the audio signal.
[0007] In one aspect, the first gain is proportional to the first attenuation, and the second gain is proportional to the second attenuation. In another aspect, when applied in response to a received first user adjustment, the second gain increases the signal level of the gain-adjusted audio signal by a greater amount than the second attenuation decreases the signal level of the audio signal. In some aspects, when a second user adjustment to a single volume control increases the total volume level of the first electronic device to the maximum volume level, the applied second gain increases the signal level of the gain-adjusted audio signal by a greater amount than the applied first gain increases the signal level of the gain-adjusted downlink signal. In one aspect, the method further includes determining the first gain adjustment and the second gain adjustment based on the streaming media content of a joint media playback session. In some aspects, determining the first gain adjustment and the second gain adjustment includes performing a table lookup on a data structure at a single volume control using the user adjustment, which correlates the gain for the downlink signal with the gain for the audio signal of the streaming media content for different user adjustments.
[0008] In one aspect, the method further includes determining, based on the output of a Voice Activity Detector (VAD), whether a microphone signal generated by the microphone of the first electronic device includes the voice of a user of the first electronic device; and in response to determining that the microphone signal includes voice, applying a first gain adjustment and a second gain adjustment to the downlink signal and the audio signal, respectively. In one aspect, the first gain adjustment is different from the second gain adjustment. In some aspects, the method further includes applying the same first gain adjustment and the second gain adjustment in response to determining that the microphone signal does not include voice. In another aspect, before receiving a user adjustment at the volume control, the first gain adjustment and the second gain adjustment are applied for a period of time in response to the microphone including voice.
[0009] In one aspect, the application of a first gain adjustment and a second gain adjustment reduces the signal levels of the downlink signal and the audio signal, respectively. The method further includes determining, before receiving a user adjustment, whether the downlink signal of a call includes voice based on the output of a voice activity detector (VAD); and in response to determining that the downlink signal includes voice, applying a third gain adjustment to the audio signal to reduce its signal level. In another aspect, the second gain adjustment reduces the signal level of the audio signal more when the downlink signal includes voice than when the downlink signal does not include voice. In some aspects, the downlink signal is a first downlink signal. The method further includes, while participating in a call and joint media playback session with a second and a third electronic device, receiving a first downlink signal from the second electronic device, a second downlink signal from the third electronic device, and an audio signal of media content; in response to a user adjustment, applying a first gain adjustment to the first downlink signal, a second gain adjustment to the audio signal, and a third gain adjustment to the second downlink signal, the third gain adjustment adjusting the signal level of the second downlink signal differently than the first gain adjustment adjusting the signal level of the first downlink signal.
[0010] According to another aspect of this disclosure, a method performed by a local device includes initiating a call with a remote device and, during the call, initiating a joint media playback session in which the two devices independently stream media content displayed on the local device's screen. The local device receives 1) a downlink signal associated with the call and 2) an audio signal associated with the media content. The local device receives user adjustments for volume control and, based on these user adjustments, 1) applies a first gain adjustment to the downlink signal and 2) applies a second gain adjustment, different from the first gain adjustment, to the audio signal. The local device uses the downlink signal and the audio signal to drive a speaker.
[0011] In one aspect, the audio signal is a first audio signal, and the method further includes displaying extended reality (XR) presented visual content on a display of a first electronic device; driving a speaker with a mixed signal comprising a downlink signal, a first audio signal associated with the media content, and a second audio signal of an object presented in the XR; and determining a first gain adjustment and a second gain adjustment based on the object within the XR presentation. In another aspect, in response to receiving a user adjustment for volume control, a third gain adjustment is applied to the second audio signal of the object presented in the XR. In some aspects, determining the first gain adjustment and the second gain adjustment includes: using sensor data from one or more sensors of the first electronic device to determine that the user desires the sound of the object to be more prominent than the sound contained in the downlink signal and the first audio signal, wherein the first gain adjustment and the second gain adjustment attenuate the downlink signal and the first audio signal more than the third gain adjustment attenuates the second audio signal.
[0012] In one aspect, determining the user's intention to make the sound of an object more prominent includes determining that the user's gaze at least one eye is focused on the object within the XR presentation. In another aspect, sensor data is motion data generated by the motion sensor of the first electronic device, and determining the user's intention to make the sound of an object more prominent includes determining, based on the motion data, that the user of the first electronic device is tilting the display screen about a central axis that extends through the display screen in a direction relative to the central axis toward the object displayed on the display screen.
[0013] In one aspect, the call is initiated by a telephone application being executed by a first electronic device, and the audio signal is a first audio signal from a media application being executed by the first electronic device. The method further includes receiving a second audio signal from a separate application being executed by the first electronic device; and determining a first gain adjustment, a second gain adjustment, and a third gain adjustment to be applied to the downlink signal, the first audio signal, and the second audio signal, respectively, based on the order in which the first electronic device begins executing the telephone application, the media application, and the separate application. In another aspect, the second audio signal is attenuated less than at least one of the first audio signal and the downlink signal.
[0014] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed embodiments below and specifically pointed out in the claims. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description
[0015] Multiple aspects are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals in the drawings indicate similar elements. It should be noted that references to "a" or "an" aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a single drawing may be used to illustrate features of more than one aspect, and for a particular aspect, not all elements in that drawing may be necessary.
[0016] Figure 1 An audio system according to one aspect is shown, which includes a local device and one or more remote devices participating in a call during a joint media playback session.
[0017] Figure 2 A block diagram is shown of a local device and an audio output device according to one aspect, wherein the local device initiates a joint playback media session when participating in a call with one or more remote devices, and the audio output device communicates wirelessly with the local device.
[0018] Figure 3 An example is shown of a local and remote device participating in a video call, performing a joint playback media session to synchronously play video and audio content, based on one aspect.
[0019] Figure 4 It is a block diagram of a local device that performs volume control operations based on one aspect.
[0020] Figure 5 An example of a volume-gain curve based on one aspect is shown.
[0021] Figure 6 This is a flowchart of one aspect of the process of using volume control to adjust the overall volume level of an audio system. Detailed Implementation
[0022] Various aspects of this disclosure will now be explained with reference to the accompanying drawings. Unless the shape, relative position, and other aspects of the components described in any aspect are explicitly defined, the scope of this disclosure is not limited to the components shown, which are for illustrative purposes only. Furthermore, while numerous details have been set forth, it should be understood that some embodiments may be implemented without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of the description. Moreover, unless the meaning is explicitly contrary, all scopes shown herein are to be considered to include the endpoints of each scope.
[0023] In one aspect, an extended reality (XR) environment (set or presented) refers to a fully or partially simulated environment that people perceive and / or interact with via electronic systems. In XR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system may detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the properties of virtual objects in the XR environment may be done in response to a representation of physical motion (e.g., a voice command).
[0024] Humans can use any of their senses to sense and / or interact with XR objects, including sight, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. As another example, audio objects can enable audio transparency, which selectively introduces ambient sound from the physical environment, with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with audio objects.
[0025] Examples of XR include virtual reality and mixed reality. A virtual reality (VR) environment is a simulated environment designed to be based entirely on computer-generated sensory input for one or more senses. In contrast to VR environments, which are designed to be based entirely on computer-generated sensory input, a mixed reality (MR) environment is a simulated environment designed to incorporate sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). Examples of mixed reality include augmented reality and augmented virtuality. An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are overlaid on the physical environment or its representation.
[0026] Many different types of electronic systems enable people to sense and / or interact with a variety of XR environments. Examples include head-mounted systems (or head-mounted devices (HMDs)), projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped like lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headsets / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology, which projects graphic images onto the human retina. Projection systems can also be configured to project virtual objects onto a physical environment, such as as holograms or on a physical surface.
[0027] Figure 1An audio system 1 according to one aspect is illustrated, which includes a local device and one or more remote devices participating in a call during a joint media playback session. As described herein, this allows users of the devices to listen to (and / or watch) media content (e.g., on multiple devices) while participating in a session with each other. The audio system includes a local (or first electronic) device 2, a remote (or second electronic) device 3, a network 4 (e.g., a computer network, such as the Internet), a media content server 5, and (e.g., optional) an audio output device 6. In one aspect, the system may include more or fewer elements. For example, the audio system may not include an audio output device (and / or the output device may be part of (or integrated into) a local device). In the absence of an audio output device, the local device may perform audio signal processing and / or audio output operations (e.g., driving one or more speakers of the local device to output sound), as described herein. In one aspect, the system may have one or more remote devices, all of which are involved in (e.g., conference) calling and joint media playback sessions with each other and with the local device, as described herein. In another aspect, the audio system may include one or more remote (electronic) servers communicatively coupled to at least some of the devices in the audio system 1, and may be configured to perform at least some of the operations described herein.
[0028] In one aspect, the local device (and / or remote device) can be any electronic device (e.g., having electronic components such as processors, memory, etc.) capable of participating in a call such as a telephone (“voice-only” or “audio-only” call) or a video (e.g., conference) call while performing a joint media playback session with one or more other devices (e.g., one or more remote devices), where (at least some) of the devices (e.g., simultaneously) play media content. For example, the media content may include musical works, movies, etc., which the local device and one or more remote devices can play simultaneously while already participating in the call. Thus, a user of the device may be able to hear the sound of the media (and / or see images or videos), and / or (e.g., simultaneously) hear the sound of the (video) call (and / or see images or videos). In some aspects, the media content can be interactive content, such as a video game in which users of both the local and remote devices participate. In another aspect, the media content can include an XR environment in which each device participating in the joint media playback session can participate. For example, a local device can participate in an XR environment by displaying image data of the XR environment on one or more displays and driving one or more speakers of the local device using one or more audio signals including the sound of the XR environment.
[0029] As described herein, an XR environment (or presentation) refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment can include AR content, MR content, VR content, etc. Many different types of electronic systems enable people to sense and / or interact with a variety of XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped as lenses designed to be placed over a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers.
[0030] In some respects, the local device can be a desktop computer, laptop computer, digital media player, etc. In another respect, the device can be a portable electronic device (e.g., capable of handheld operation), such as a tablet computer, smartphone, etc. In yet another respect, the device can be a head-mounted device, such as smart glasses, or a wearable device, such as a smartwatch. In one respect, the remote device can be the same type of device as the local device (e.g., both devices are smartphones). In yet another respect, at least some of the remote devices can be different, such as some being desktop computers and others being smartphones.
[0031] As shown in the figure, local device 2 is coupled to remote device 3 and / or media content server 5 via a computer network (e.g., the Internet) 4 (e.g., communicatively). Specifically, the local device and the remote device can be configured to establish and participate in telephone (or voice-only) calls, where the participating devices exchange audio data. For example, each device transmits at least one microphone signal as an uplink audio signal to the other participating devices and receives at least one audio signal from the other devices as a downlink audio signal for playback by one or more speakers. In one aspect, the network may include a Public Switched Telephone Network (PSTN), through which the local device and the remote device are able to make outgoing calls and / or receive incoming calls. In another aspect, the local device can be configured to establish Internet Protocol (IP) telephone (or Voice over IP) calls with one or more remote devices via a network (e.g., the Internet). Specifically, the local device may use any signaling protocol (e.g., Session Initiation Protocol (SIP)) to establish a communication session and any communication protocol (e.g., Transmission Control Protocol (TCP), Real-Time Transport Protocol (RTP), etc.) to exchange audio data during the call. For example, when (e.g., by executing a telephone application within a local device, e.g., Figure 2When application 29) shows initiating a call, the local device may transmit one or more microphone signals (e.g., as uplink audio signals) captured by one or more microphones as audio data (e.g., in IP packets) to one or more remote devices, and receive one or more (e.g., downlink audio) signals from the remote devices via the network to drive one or more speakers of the local device. Alternatively, the local device may be configured to establish a wireless (e.g., cellular) call. In this case, network 4 may include one or more cell towers, which may be part of a communication network (e.g., a 4G Long Term Evolution (LTE) network) supporting data transmission (and / or voice calls) of electronic devices such as mobile devices (e.g., smartphones).
[0032] On the other hand, local and remote devices can be configured to establish and participate in video calls with one or more remote devices. In this case, the local device can establish a video call (e.g., similar to VoIP, using SIP to initiate the session and RTP to transmit data) and exchange video and / or audio data with one or more remote devices while establishing the video call. For example, the local device may include one or more cameras that capture video, which is encoded using any video codec (e.g., H.264) and transmitted to the remote device for decoding and display on one or more displays. Further details regarding the call are described herein.
[0033] In some aspects, the media content server 5 can be a standalone server computer or cluster of server computers configured to stream media content to electronic devices such as local and remote devices. In this case, the server can be part of a cloud computing system capable of streaming data as a cloud-based service to one or more subscribers (e.g., local or remote devices). In some aspects, the server can be configured to stream any type of media (or multimedia) content, such as audio content (e.g., musical works, audiobooks, podcasts, etc.), still images, video content (e.g., movies, television productions, etc.), etc. In one aspect, the server can use any audio and / or video encoding format and / or any method used to stream content to one or more devices.
[0034] In one aspect, media content server 5 can be configured to simultaneously stream media content to one or more devices to allow these devices to participate in a joint media playback session. For example, the server may receive a request from a device (e.g., local device 2) to stream media content segments, which may include audio content (e.g., musical works) and / or video content (e.g., video signals associated with a movie), together with another device (e.g., remote device 3). In another aspect, a request may be transmitted by a local device (and / or a remote device) in response to a device receiving user input to start playing media content. In this case, the server may establish a communication link with the local and remote devices already involved in (e.g., telephone and / or video) calls. Once the communication link is established, the server may encode the audio content using any codec (e.g., MP3, AAC, etc.) and / or encode the video content using any codec, and then transmit the encoded content to each device for decoding and output. In another aspect, a local device may transmit a message to a remote device requesting the initiation of a joint media playback session. In response, the remote device may communicate with the media content server to retrieve media content and play it synchronously with the local device. In one respect, devices participating in a joint media playback session can synchronously output media content, allowing users to simultaneously output and experience the content. In other respects, any timing synchronization method can be used (e.g., by the devices and / or servers participating in the session) to ensure simultaneous and synchronous streaming of media. This article describes more about joint media playback sessions.
[0035] As shown, the audio output device 6 can be any electronic device including at least one speaker and configured to output sound by driving the speaker. For example, as shown, the device is a wireless headset (e.g., in-ear headphones or earbuds) designed to be positioned on (or in) a user's ear and designed to output sound into the user's ear canal. In some aspects, the headset can be a sealed type with flexible earpiece ends designed to acoustically seal the entrance to the user's ear canal relative to the surrounding environment by blocking or occluding it within the ear canal. As shown, the output device includes a left earpiece for the user's left ear and a right earpiece for the user's right ear. In this case, each earpiece can be configured to output at least one audio channel of media content (e.g., the right earpiece outputs the right audio channel of a stereo recording (such as a musical work) with two-channel input and the left earpiece outputs the left audio channel). In another aspect, the output device can be any electronic device including at least one speaker and arranged for wear by a user and arranged to output sound by driving the speaker with an audio signal. For example, the output device can be any type of headset, such as over-ear (or over-ear) headphones that at least partially cover the user's ears and are arranged to direct sound into the user's ears.
[0036] In some respects, the audio output device can be a headset, as illustrated herein. In other respects, the audio output device can be any electronic device arranged to output sound to the surrounding environment. Examples may include standalone speakers, smart speakers, home theater systems, or infotainment systems integrated into vehicles.
[0037] In one aspect, the output device can be a wireless device communicatively coupled to the local device for exchanging audio data. For example, the local device can be configured to establish a wireless connection with the audio output device via a wireless communication protocol (e.g., Bluetooth or any other wireless communication protocol). During the established wireless connection, the local device can exchange (e.g., transmit and receive) data packets (e.g., Internet Protocol (IP) packets) with the audio output device, which can include audio digital data of any audio format. Specifically, the local device can be configured to establish and communicate with the audio output device via a two-way wireless audio connection (e.g., one that allows two devices to exchange audio data), such as making hands-free calls or using voice commands. Examples of two-way wireless communication protocols include, but are not limited to, Hands-Free Mode (HFP) and Headset Mode (HSP), both of which are Bluetooth communication protocols. In another aspect, the local device can be configured to establish and communicate with the output device via a one-way wireless audio connection (e.g., the Advanced Audio Distribution Profile (A2DP) protocol), which allows the local device to transfer audio data to one or more audio output devices. Further details regarding these wireless audio connections are described herein.
[0038] On the other hand, local device 2 can be communicatively coupled to audio output device 6 via other methods. For example, both devices can be coupled via a wired connection. In this case, one end of the wired connection can be (e.g., fixedly) connected to the audio output device, while the other end can have a connector, such as a media jack or a Universal Serial Bus (USB) connector, that inserts into the jack of the audio source device. Once connected, the local device can be configured to drive one or more speakers of the audio output device with one or more audio signals via the wired connection. For example, the local device can transmit the audio signals as digital audio (e.g., PCM digital audio). On the other hand, the audio can be transmitted in an analog format.
[0039] In some respects, local device 2 and audio output device 6 may be different (separate) electronic devices, as illustrated herein. In other respects, the local device may be a component of the audio output device (or integrated with the audio output device). For example, as described herein, at least some components of the local device (such as a controller) may be part of the audio output device, and / or at least some components of the audio output device may be part of the local device. In this case, each device may be communicatively coupled via traces that are part of one or more printed circuit boards (PCBs) within the audio output device.
[0040] Figure 2 A block diagram is shown of a local device 2 that initiates a joint playback media session when participating in a call (e.g., voice or video) with one or more remote devices 3, according to one aspect, and an audio output device 6 that communicates wirelessly with the local device is also shown. The local device 2 includes a controller 20, a network interface 21, a speaker 22, a display (or monitor) 25, a memory 26, a volume control 12, and one or more sensors 10, including a microphone 23, a camera 24, and an inertial measurement unit (IMU) 11. In one aspect, the local device may include more or fewer elements as described herein. For example, the device may include two or more of at least some of these elements (such as having two or more microphones 23 and / or two or more speakers 22).
[0041] In one aspect, one or more sensors 10 are configured to detect an environment (e.g., in which a local device is located) and generate sensor data based on the environment. For example, camera 24 is a complementary metal-oxide-semiconductor (CMOS) image sensor capable of capturing digital images including image data representing the field of view of camera 24, where the field of view includes the scene of the environment in which device 2 is located. In some aspects, the camera may be a charge-coupled device (CCD) camera type. The camera is configured to capture still digital images and / or video represented by a series of digital images. In one aspect, the camera may be located near the local device or anywhere on the local device. In some aspects, the device may include multiple cameras (e.g., where each camera may have a different field of view).
[0042] Microphone 23 can be any type of microphone (e.g., a differential pressure gradient microelectromechanical system (MEMS) microphone) configured to convert acoustic energy caused by sound waves propagating in an acoustic environment into an input microphone signal. In some aspects, the microphone can be an "external" (or reference) microphone arranged to capture sound from the acoustic environment. In other aspects, the microphone can be an "internal" (or error) microphone arranged to capture sound (and / or sense pressure changes) within the user's ear (or ear canal). The IMU is configured to generate motion data indicating the position and / or orientation of the local device. In one aspect, the local device may include additional sensors, such as (e.g., optical) proximity sensors, designed to generate sensor data indicating that an object is at a specific distance from the sensor (and / or the local device).
[0043] In one respect, sensor 10 may be a component of a local device (or integrated into a local device). In another respect, sensor may be a separate electronic device (e.g., via network interface 21) communicatively coupled to the controller.
[0044] The speaker 22 may be, for example, an electrically driven driver specifically designed for sound output in a particular frequency band, such as a woofer, tweeter, or midrange driver. In one aspect, the speaker 22 may be a “full-range” (or “full-frequency”) electrically driven driver that reproduces as much of the audible frequency range as possible. In some aspects, the local device may include one or more speakers, wherein at least some of these speakers may be the same or different (e.g., one is a woofer and the other is a tweeter).
[0045] Display screen 25 is designed to present (or display) digital image or video (or image) data. In one aspect, the display screen may use liquid crystal display (LCD) technology, light-emitting polymer display (LPD) technology, or light-emitting diode (LED) technology, although other display technologies may be used in other aspects. In some aspects, the display screen may be a touch-sensitive display screen configured to sense user input as an input signal. In some aspects, the display screen may use any touch sensing technology, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies.
[0046] Volume control 12 is configured to adjust the volume level of the sound output of the local device in response to a user adjustment (e.g., user input) received at the control. In one aspect, the volume control may be a “master” volume control configured to control the overall volume level of the local device (e.g., the sound output level of speaker 22). In another aspect, the volume control may be a “hardware” volume control, which may be a dedicated volume input control, such as one or more buttons, a rotary knob, or a physical slider. In some aspects, the volume control may be any type of physical input device capable of adjusting the overall volume level. In one aspect, the volume control may be a single volume control comprising several volume settings (or positions), where each setting defines a different volume level of the local device (e.g., a different sound output level (e.g., dB SPL)). Specifically, the volume control may (e.g., in response to a user adjustment) gradually increase or decrease the volume level based on the volume setting or position adjusted by the user. For example, when the volume control is a rotary volume knob, the control may have several (e.g., 18) volume settings, where each successive volume setting may correspond to a degree of rotation and may increase the overall volume by a specific gain value. In this configuration, each volume setting may correspond to a 20° rotation around a central axis. For example, a first volume setting could be 0°, where the overall volume is muted, a second volume setting could be 20° (e.g., where the overall volume increases by a specific gain value), and so on. Thus, the knob generates a control signal that gradually increases or decreases the volume based on the degree and direction of the knob's rotation (e.g., clockwise rotation increases volume, counter-clockwise rotation decreases volume). In one aspect, the volume control may be a master volume control configured to provide bidirectional control to gradually increase or decrease the overall volume level of the device (e.g., sound output). In another aspect, the control may be part of a local device (e.g., integrated into the device). In yet another aspect, the volume control may be part of an electronic device communicatively coupled to a local device.
[0047] On the other hand, volume control can be "software" volume control, such as a user interface (UI) item (e.g., a graphical user interface (GUI)) displayed on the display screen 25 of the local device. For example, volume control can be a slider capable of translating (e.g., moving in at least one direction) along a predefined sliding range. When user input is received to adjust (or translate) the position of the slider (e.g., by the user touching the slider on the display screen and dragging it in one or more directions), the volume control adjusts the overall volume level based on the slider's position. In one aspect, similar to an example of physical control, the UI item can include several volume settings, where each position of the slider can correspond to a different volume level of the device. In this case, since the slider has a predefined sliding range, or a sliding distance from one side, each volume setting can correspond to a distance along the sliding range. On the other hand, the volume setting may correspond to a percentage (e.g., from 0% to 100%), which may correspond to the distance the slider travels along the sliding range from its starting position. On the other hand, volume control can include several volume settings as numerical values (e.g., 1-10), where these volume settings can be changed by user adjustments to the volume control (e.g., selecting, dragging, twisting, etc.).
[0048] In some respects, volume control can be any input from the user of the device. For example, this input can include gestures (e.g., hand gestures, finger gestures, head gestures, etc.) made by the user and detected by the device (e.g., movement of the local device caused by hand gestures detected by IMU 11). In other respects, volume control can be a voice command received via microphone 23. More details regarding volume control 12 are described herein.
[0049] Controller 20 may be a dedicated processor such as an application-specific integrated circuit (ASIC), a general-purpose microprocessor, a field-programmable gate array (FPGA), a digital signal controller, or a set of hardware logic structures (e.g., filters, arithmetic logic units, and dedicated state machines). The controller is configured to perform audio signal processing operations and / or networking operations. For example, controller 20 may be configured to participate in a call and simultaneously perform a joint media playback session to stream (e.g., exchange) media content with one or more remote devices via network interface 21. Alternatively, the controller may be configured to perform audio signal processing operations on audio data of the media content and / or audio data associated with the participating call (e.g., downlink signals). Further details regarding the operations performed by controller 20 are described herein.
[0050] Memory 26 can be any type of storage medium (e.g., non-transitory machine-readable), such as random access memory, CD-ROM, DVD, magnetic tape, optical data storage devices, flash memory devices, and phase-change memory. In one aspect, the memory can be a component of a local device (e.g., integrated into the local device). In another aspect, the memory can be part of controller 20. In some aspects, the memory can be a separate device, such as a data storage device. In this case, the memory can be communicatively coupled to controller 20 (e.g., via network interface 21) so that the controller performs one or more of the operations described herein.
[0051] As shown, the memory stores an operating system (OS) 27, media applications 28, and telephone applications 29, which, when executed by the controller, enable the local device to perform one or more operations as described herein. In one aspect, the memory may include more or fewer applications. OS 27 is a software component responsible for the management and coordination of activities and the sharing of resources of the local device 2 (e.g., controller resources, memory, etc.). In one aspect, the OS acts as a host for applications (e.g., applications 28 and 29) running on the device. In some aspects, these applications may run on top of the OS. In one aspect, the OS provides an interface to the hardware layer (not shown) of the local device and may include one or more software drivers that communicate with the hardware layer. For example, these drivers may receive and process data packets received by one or more other devices (e.g., user input devices, such as display 25, which may be a touch-sensitive display, one or more sensors in sensor 10, etc.) that are communicatively coupled to the device through the hardware layer.
[0052] In one aspect, media application 28 may be an application that streams media content (e.g., from media content server 5) to a local device when executed by the local device. Specifically, the media application may be a music streaming application that, when executed, streams music for playback on speaker 22 (and / or speaker 83 of the audio output device). Alternatively, media application 28 may be a multimedia (e.g., video and / or audio) streaming application that streams multimedia content (e.g., movies, etc.) for playback on the local device (e.g., for video playback via display screen 25 and / or audio playback via speaker 22). In another aspect, the application may retrieve media content from local storage (e.g., storage 26) and / or from a remote source (such as media content server 5), as described herein. In one aspect, for streaming media content, the media application may display a graphical user interface (GUI) on display screen 25, through which a user can navigate the application to select one or more media contents for streaming to the local device (and / or audio output device).
[0053] In one aspect, the telephone application 29 may be an application that, when executed by a local device, allows the local device to initiate and make telephone calls with one or more remote devices. For example, when initiated (e.g., when a user selects a UI item of the application displayed on display screen 25), the application may display a GUI through which the local user can dial a telephone number. Once dialed, the local device can connect to the remote device via the cellular network of network 4 (e.g., a 4G Long Term Evolution (LTE) network), as described herein. In some aspects, the telephone application may be an audio-only (or voice-only) telephone application capable of performing audio calls (e.g., where the local device and one or more remote devices exchange audio data captured by one or more microphones used to drive one or more speakers of the respective devices). In another aspect, the telephone application may be a video calling (or video conferencing) application that allows the local device 2 to make video calls with one or more remote devices, as described herein. In yet another aspect, the local device may include a video calling application separate from the telephone application.
[0054] On the other hand, local device 2 may include one or more other applications that, when executed, cause the local device (and / or audio output device) to play (or output) audio and / or video (or image) content. For example, memory may include an XR rendering application that allows local users of the device to participate in the XR environment when executed by local device 2 (e.g., its controller 20).
[0055] Audio output device 6 includes a controller 80, a network interface 81, a speaker 83, a microphone 84, an accelerometer 85, and a volume control 82. In one aspect, the device may include more or fewer elements, such as having memory. In some aspects, the microphone may be an external or internal microphone, as described herein. In the case of in-ear headphones, the internal microphone senses the inside of the user's ear when the headphones are on (or in) the user's ear. The accelerometer is arranged and configured to receive (detect or sense) speech vibrations generated when the user (e.g., a user who may be wearing the output device) speaks, and to generate an accelerometer signal representing (or containing) the speech vibrations. Specifically, the accelerometer is configured to sense bone conduction vibrations transmitted from the user's vocal cords to the user's ear (ear canal) when speaking and / or humming. For example, when the audio output device is a wireless headset, the accelerometer may be located on or inside the headset anywhere it can contact a part of the user's body to sense vibrations.
[0056] In one aspect, controller 80 is configured to perform audio signal processing operations and / or networking operations, as described herein. For example, the controller may be configured to acquire (or receive) audio data of media content (as analog or digital audio signals) or media content desired by the user (e.g., music, etc.) for playback via speaker 83. In some aspects, the controller may acquire audio data from local memory, or the controller may acquire audio data from network interface 81, thereby acquiring data from an external source such as local device 2 (via its network interface 21). For example, the output device may stream audio signals from the local device (e.g., via Bluetooth connection) for playback via speaker 83. The audio signal may be a signal input audio channel (e.g., mono). In another aspect, the controller may acquire two or more input audio channels (e.g., stereo channels) for output through two or more speakers. In one aspect, where the output device includes two or more speakers, the controller may perform additional audio signal processing operations. For example, the controller can spatially render the input audio channels (e.g., by applying spatial filters, such as the head-related transfer function (HRTF)) to generate binaural output audio signals for driving at least two speakers (e.g., a left speaker and a right speaker).
[0057] In one aspect, volume control 82 can perform operations similar to those of control 12 on the local device. For example, upon receiving user input, this control can adjust the (e.g., total) volume of the audio output device 6 (e.g., speaker 83). In other aspects, volume control 82 can be used to adjust the volume at the local device 2. In this case, upon receiving user input, audio output device 6 can transmit a control signal to the local device instructing the user to adjust the volume control 82, which the local device can use to adjust the volume of one or more audio signals. Further details regarding volume adjustment are described herein.
[0058] As described herein, controller 20 may be configured to perform (e.g., additional) audio signal processing operations based on elements coupled to the controller. For example, when the local device includes two or more “outside-ear” speakers arranged to output sound to an acoustic environment rather than speakers arranged to output sound to a user’s ears (e.g., as speakers in in-ear headphones), the controller may include a sound output beamformer configured to generate speaker driver signals that produce spatially selective sound output when driving two or more speakers. Thus, when used to drive speakers, the local device may generate a directional beam pattern that can be directed to a location within the environment.
[0059] In some aspects, controller 20 may include a sound pickup beamformer configured to process audio (or microphone) signals generated by two or more external microphones of the output device to form a directional beam pattern (as one or more audio signals) for spatially selective sound pickup in certain directions, thereby increasing sensitivity to the location of one or more sound sources. In some aspects, the controller may perform audio processing operations (e.g., perform spectral shaping) on the audio signals containing the directional beam pattern.
[0060] On the other hand, controller 80 can perform one or more functions. For example, controller 80 can be configured to perform an active noise cancellation (ANC) function to make speaker 83 produce noise immunity in order to reduce ambient noise from the environment leaking into the user's ears. The ANC function can be implemented as feedforward ANC, feedback ANC, or a combination thereof. Thus, the controller can receive a reference microphone signal from a microphone such as microphone 84 that captures external ambient sound. On the other hand, the controller can perform any ANC method to produce noise immunity. On the other hand, controller 80 can perform a transparency function, in which the sound played by the device is a reproduction of ambient sound captured by the device's external microphone in a "transparent" manner (e.g., as if the user were not wearing headphones). The controller processes at least one microphone signal captured by at least one external microphone 84 and filters the signal through a transparency filter, which reduces acoustic blockage caused by the audio output device being located on, in, or above the user's ears, while also preserving the spatial filtering effect of the wearer's anatomical features (e.g., head, auricle, shoulders, etc.). The filter also helps to preserve the timbre and spatial cues associated with the actual ambient sound. In one respect, the filter for the transparency function can be user-specific, depending on specific measurements of the user's head. For example, the controller can determine the transparency filter based on the head-related transfer function (HRTF) or the equivalent head-related impulse response (HRIR) based on the user's anthropometric measurements.
[0061] In one aspect, the local device (e.g., its controller 20) can perform (or control) at least some of the functions of the audio output device 6 (e.g., its controller 80). For example, the controller 20 can perform an ANC function, thereby generating an anti-noise signal based on a reference microphone signal (e.g., from the audio output device and / or the local device). When generated, the local device can transmit the anti-noise signal to the audio output device for audio playback via the speaker 83.
[0062] As described herein, both the local device and the audio output device are configured to establish a wireless audio connection (e.g., a Bluetooth connection) to exchange audio data. Therefore, audio data and / or control signals can be exchanged between the two devices via a wireless connection.
[0063] In one respect, the operations performed by the controller may be implemented in software (e.g., as instructions stored in memory and executed by the controller) and / or may be implemented by hardware logic structures as described herein.
[0064] On the other hand, at least some of the operations performed by the audio system 1 as described herein may be performed by the local device 2 and / or the audio output device 6. For example, the local device may include two or more speakers and may be configured to perform sound output beamforming operations (e.g., when the local device includes two or more speakers). On the other hand, at least some of the operations may be performed by a remote server communicatively coupled to either device, for example via a network (e.g., the Internet).
[0065] In one aspect, at least some elements of the local device 2 and / or the audio output device 6 may be integrated with (or as a component of) each respective device. For example, when the audio output device is an over-ear headphone, the microphone, speaker, and accelerometer may be part of at least one earcup of the headphone, which is placed on the user's ear. In another aspect, at least some elements may be separate electronic devices communicatively coupled to the device. For example, the display screen 25 may be a separate device (e.g., a display monitor or television) communicatively coupled to the local device (e.g., a wired or wireless connection) to receive image data for display. As another example, the camera 24 may be a component of a separate electronic device coupled to the local device to provide captured image data (e.g., a webcam).
[0066] As described herein, the local device 2 and remote device 3 of the audio system 1 can perform a joint media playback session while participating in a call, allowing users of the devices to communicate while experiencing simultaneous media content playback. In one aspect, the local device can initiate a joint media playback session while already participating in a call. Figure 3 A graphical example is shown of local and remote devices participating in joint media playback when they have already joined a video conference call.
[0067] Figure 3 An example is shown of a local device 2 and a remote device 3, both already involved in a video call, performing a joint media playback session to synchronously play video and audio content. Specifically, the figure illustrates a local user 30 and a remote user 31 (who may be in the same or different locations) participating in a joint media playback session while already involved in a video call. Displayed on the monitor 25 of the local device 2 is a video call user interface (UI) 44 (e.g., a video call application being executed by the local device). Figure 2The telephone application 29 shown displays an interface that shows the local user's video 46 (e.g., a video representation) and the remote user's video 45 (which is larger than the local user's video) positioned in the middle of UI 44. Similarly, a video call UI 47 is displayed on the display 32 of the remote device 3, which is displayed by a video call application (which may be the same as or different from the application running on the local device), showing the videos of the remote user and the local user (their positions are swapped relative to their positions displayed on the local device's display 25).
[0068] In one aspect, video representations can be generated using video data captured by one or more cameras on each device. For example, when a local user 30 is in the field of view of camera 24, the camera can capture video data of the local user, which is then displayed on the local device and transmitted (e.g., via network 4) to a remote device for display on display 32.
[0069] Video 49 of media content is also shown on both devices in a joint media playback session in which both devices are involved. In one aspect, the two devices can stream media content (e.g., using a video streaming application) to play the media content synchronously (e.g., displaying video on their displays and outputting audio content of the media content through one or more speakers), while both devices are involved in a video call. Thus, two users can interact with each other (e.g., have a conversation) via video call while simultaneously watching video 49 of the media content (and hearing the audio).
[0070] In one aspect, the video call UI 44 (and / or video call UI 47) may include additional video representations based on the number of remote users participating in the video call. In this example, the local device's UI 44 includes one video representation 45 for the remote users because the local device is only participating in the video call with one remote user. As more remote users join the video call, UI 44 may include additional video representations, one for each remote user. For example, when a video call with three remote devices has been participated in, the video call UI 44 may include three video representations, one for each remote device (and one video representation 46 for the local user). In one aspect, each video representation in the video representations may be positioned around UI 44. Continuing with the previous example, the three video representations may be positioned in three rows. In another aspect, the video representations may be positioned differently.
[0071] As described herein, local devices can participate in a joint media playback session while making a telephone (audio-only) call. In this scenario, both devices can display video (and / or play audio) of media content while the user engages in a voice-only conversation. In one respect, either of these devices can initiate a telephone (or video) call using any known method. For example, local user 30 might have initiated a phone application and dialed the phone number of a remote device.
[0072] As shown in this example, audio and video content can be played during a joint media playback session when the device is involved in a video call. In some aspects, any type of media content can be played during a joint media playback session when both the local and remote devices are involved in a telephone (or voice-only) call or a video call. For example, the media content may include XR presentations that can be participated in by users of both the local and remote devices. Specifically, the media content may include video (or visual) content of the XR environment as image data that can be displayed on the display screen 25 of the local device and the display screen of the remote device 3. Furthermore, the media content may include audio data (or one or more audio signals) of sound within the XR environment. For example, the audio data may include the sound of objects within the XR environment (e.g., a dog barking).
[0073] As described in this article, users of a device can participate in an XR environment. Specifically, each device can present a different (or similar) perspective of the XR environment. For example, each device can present a first-person perspective of the environment (e.g., through the viewpoint of a virtual avatar associated with the device positioned within the XR environment). In this case, sounds within the XR environment (e.g., the sounds of objects) may be perceived differently by each participant. This article describes more about the sounds of objects within an XR environment.
[0074] As described herein, a local device can participate in a joint media playback session with a remote device, where both devices are involved in a call, such as an audio-only call. In this scenario, the local device can receive downlink audio signals, including the voice of the remote user on the remote device, while simultaneously transmitting uplink audio signals, including the voice of the local user. Furthermore, the local device can receive media content, such as audio signals of musical works associated with the playback session. Therefore, the local device can drive one or more speakers with the downlink audio signals and the mixed audio signals, allowing the local device to participate in a conversation with a remote user while experiencing media content.
[0075] When participating in a joint media playback session, local and remote devices can play various types of media content, such as music, movies, and XR environments, as described herein. The audio data associated with each type of media content can be controlled differently. For example, music may have a dynamic range of 96 dB, while movie audio may have a dynamic range of 144 dB. Furthermore, media content can be controlled at different volume levels, where their signal levels can be higher (or greater) than the signal level of the downlink signal being played. Therefore, when the media content signal and the downlink signal are mixed together, the local user may perceive the media content as louder than the downlink signal (e.g., speech), thus potentially drowning out the remote user's speech contained within the downlink signal. In one aspect, to address this issue, multiple volume controls can be provided to the local user, one for the signal output by the local device. However, this solution has drawbacks. For example, to be effective, each audio signal being mixed and played requires its own volume control. Therefore, as the number of audio signals being played increases, this solution can become difficult to manage. Therefore, a single volume control (e.g., for the master volume control of a local device) is needed to adjust the overall volume level of audio playback by applying different gains to the signal.
[0076] To overcome these shortcomings, this disclosure describes an audio system including a single volume control capable of applying different volume control behaviors to different audio signals. Specifically, in response to receiving a user adjustment to adjust the overall volume level at a single (e.g., master) volume control of a local device, the local device applies a first gain adjustment to the downlink signal, a second gain adjustment to the audio signal, and drives the speaker 22 of the local device at the adjusted volume level with a mixture of these signals. For example, returning to the previous example, since the volume level of the audio signal of the media content can be controlled to be higher (or louder) than the downlink signal, when the local user lowers the volume at the volume control, the second gain adjustment can lower the signal level of the audio signal more (or at a faster rate) than the first gain adjustment lowers the signal level of the downlink signal. Therefore, as the overall volume is lowered, the level of the audio signal is lowered more than the level of the downlink signal.
[0077] Figure 4This is a block diagram of a local device 2 performing volume control operations according to one aspect. Specifically, the diagram shows a controller 20 having several operation blocks for performing audio signal processing operations to control the audio output volume of the local device. As shown, the controller includes a call manager 52, a joint media playback session manager 51, a voice activity detector (VAD) 53, a volume-gain curve selector 54, an audio signal gain selector 55, a downlink signal gain selector 56, scalar gains 90, 57, and 58, and (e.g., a matrix) mixer 59. In one aspect, the controller may have more or fewer operation blocks. For example, the controller may include additional pairs of gain selectors and scalar gains, each pair for additional audio signals to be mixed and played during call and playback sessions. In one aspect, at least some of these operation blocks may be optional, and therefore the operation of optional blocks may be omitted or combined. For example, scalar gain 90 may be optional, as described herein.
[0078] Call manager 52 is configured to initiate (and perform) calls between local device 2 and one or more remote devices 3. In one aspect, the call manager may initiate a call in response to user input. For example, the call manager may be part of (or receive instructions from) a telephone application executed by the local device (e.g., the controller 20 of the local device). For example, the telephone application may display a UI on the display screen 25 of the local device, providing the local user with the ability to initiate a call, such as a keyboard, contact list, etc. Once the UI receives user input (e.g., the local user dials a remote user's phone number using the keyboard), the call manager may communicate with the network interface 21 of local device 2 to establish a call, as described herein. In one aspect, the telephone call may be made over any network, such as via the PSTN and / or via the Internet (e.g., for VoIP calls). In some aspects, the call manager may initiate a call as described herein and / or using any method.
[0079] Once a call is initiated, the call manager can exchange call data between the remote devices that the local device uses to participate in the call. For example, the call manager can receive one or more downlink audio signals from each of the remote devices. In one aspect, the call manager can mix the downlink signals into (at least one) downlink audio signals (e.g., via matrix mixing operation). In some aspects, the call manager can receive audio and / or video (or image) data from sensor 10 and can transmit these data to each remote device that has participated in the call with the local device. For example, the call manager can receive microphone signals (which may include the voice of the local user) from one or more microphones 23 and / or can receive image data captured by one or more cameras 24 and can transmit the microphone signals and / or image data to each remote device. In some aspects, when the local device includes two or more microphones, the call manager can transmit sound pickup beamformer signals that include a directional beam pattern.
[0080] The joint media playback session manager 51 is configured to initiate a joint media playback session between a local device and one or more remote devices (e.g., devices calling from the local device), in which the devices can independently stream media content for (e.g., synchronous) playback. For example, in response to receiving an instruction to initiate a session, the playback session manager can transmit a request to initiate a session to the media content server 5, as described herein. Specifically, a media application executing within the local device can transmit an instruction to the session manager in response to receiving user input (e.g., based on the user selecting a play button in the media application's UI, which may be displayed on the display 25 of the local device 2). On the other hand, the session manager can request user authorization before initiating a session. For example, once a user initiates media playback in the media application, the session manager can provide a notification (e.g., a pop-up notification displayed on the display 25) requesting user authorization to initiate a joint media playback session with at least some of the call participants. When user authorization is received (e.g., by receiving a user selection of a UI item within the pop-up notification), the session manager can process the request to initiate a session, as described herein.
[0081] In one aspect, the joint media playback session manager 51 is configured to receive media content data (e.g., once a session is initiated). In this case, the session manager receives at least one audio signal (or audio channel) associated with the media content. For example, the received audio signal may be associated with a musical piece that a local user has requested to play (e.g., via the UI of a media application). In another aspect, the session manager may receive two or more audio signals from a segment of media content. For example, when streaming a musical piece from a media content server, the session manager may receive two or more audio channels (e.g., the left and right channels of a stereo recording of the musical piece). In yet another aspect, the session may receive two or more channels, such as the entire audio track of a 5.1 surround sound format video.
[0082] On the other hand, media content data may include audio data of at least some audio channels of a sound (or audio) object within a sound space, such as the sound space of an XR environment for the media content. The audio data of the sound object may include 1) an audio signal comprising the sound of an object within the XR environment (e.g., a virtual object), and 2) spatial data representing the sound source of the sound object in space. In one aspect, these sound objects may correspond to objects within the XR environment. For example, a sound object may include the sound of a virtual dog (e.g., a bark) and the location of the sound in the XR environment to be displayed on display screen 25. In one aspect, the spatial data may be an angle / parameter representation of a sound source in the XR environment (e.g., relative to a local device). In some aspects, the spatial data may indicate the three-dimensional (3D) position of the sound source relative to the device (e.g., located on a virtual sphere surrounding the device) as positional data (e.g., elevation, azimuth, distance, etc.). In one aspect, any method may be performed to generate an angle / parameter representation of the sound source, such as encoding the sound source into HOA B format by translating and / or upmixing the at least one ambient signal in the ambient signal to generate a high-order stereo reverberation (HOA) representation of the sound source. On the other hand, sound objects can be in any audio format (e.g., including the object's audio signal and spatial information).
[0083] VAD 53 is configured to receive microphone signals from microphone 23 of sensor 10 and / or one or more downlink audio signals from each remote device involved in the call with the local device, and is configured to perform a voice activity detection (or voice detection) operation to detect the presence (or absence) of user voice contained therein. Specifically, the controller determines whether any of these signals includes voice based on the output of the VAD. For example, the VAD may determine whether (at least a portion) of the spectral content of the signal is associated with human voice. On the other hand, the VAD may determine the presence of voice based on whether the signal level (e.g., a portion of the spectral content) of the signal exceeds a threshold. In some aspects, the VAD may use any method to determine whether voice is contained within the signal. The VAD is configured to generate an output based on the received downlink signals. Specifically, the VAD may generate a VAD signal indicating whether voice is contained within the microphone signal of the local device, and / or may generate a VAD signal indicating whether voice is contained within the downlink audio signal. For example, the VAD signal may have a high signal level (e.g., 1) when voice is detected, and a low signal level (e.g., 0) when no voice is detected (or at least not within a threshold level). On the other hand, the VAD signal does not need to be a binary decision (speech / non-speech); rather, it can be the probability of speech presence. In some respects, the VAD signal can also indicate the signal level of the detected speech (e.g., sound pressure level (SPL)).
[0084] As shown in the figure, VAD 53 can generate a VAD signal indicating whether speech is contained within one or more microphone signals and / or one or more downlink audio signals. In one aspect, the VAD may have multiple outputs. Specifically, the VAD can generate one VAD signal indicating the presence (or absence) of speech contained within the microphone signals, and generate another VAD signal for the downlink audio signals.
[0085] In some respects, the VAD can perform additional audio signal processing operations. For example, the VAD can perform speech digital signal processing (DSP) operations on downlink audio signals and / or microphone signals to reduce (or eliminate) the noise contained therein (e.g., to produce a speech signal that primarily contains speech). In one respect, to process the signal, since most noise (or non-speech noise) has low-frequency content, the VAD can apply a high-pass filter. In another respect, the VAD can improve the signal-to-noise ratio (SNR) of the signal by applying one or more filters (e.g., low-pass filter, band-pass filter, high-pass filter, etc.). In some respects, the VAD can perform any operation to reduce noise within the signal.
[0086] Scalar gain 90 is configured to receive one or more audio signals of media content from session manager 51 and is configured to process the audio signals based on VAD signals received from VAD 53. Specifically, scalar gain is configured to adjust (e.g., at least a portion) the signal level of the audio signal by applying one or more scalar gain values (e.g., as gain adjustment) to the audio signal to produce a gain-adjusted audio signal based on whether the VAD signal indicates the presence of speech detected within the downlink audio signal (and / or microphone signal). Specifically, gain adjustment may reduce the signal level of the audio signal of media content associated with (e.g., streamed by) the joint media playback session. Thus, in response to (e.g., controller 20) determining that the downlink signal includes speech, scalar gain may apply gain adjustment to the audio signal to reduce the signal level of the audio signal. In one aspect, the applied scalar gain may be a predefined value. In some aspects, the application of this scalar gain may be performed prior to other audio signal processing operations, such as the application of other scalar gains 57, as described herein.
[0087] On the other hand, the applied gain can be based on the VAD signal. For example, as described herein, the VAD signal may indicate the signal level of the downlink audio signal (or more specifically, the signal level of the speech contained therein). In this case, the scalar gain can be configured to adjust the applied scalar gain value based on the signal. For example, when speech detected in the downlink audio signal is at a certain signal level, the scalar gain may apply a gain value to reduce the signal level of the audio signal below the certain signal level of the downlink signal, in order to ensure that the sound of the media content is lower than the speech within the call.
[0088] Volume-gain curve selector 54 is configured to perform contextual analysis of call and / or media content data to determine the priority of audio signals in a call and / or media playback session, indicating which audio signals should be more prominent than others when played by audio system 1 (e.g., its local devices). Specifically, the curve selector determines whether a local user intends (or desires) for one or more sounds of the call and / or media content to be more prominent and / or heard relative to other sounds during the call and playback session, and in response, prioritizes those sounds to be more prominent relative to other sounds. For example, the curve selector determines whether to prioritize the sound of the call (e.g., downlink audio signal) and / or one or more sounds of the media content being played during the media playback session (e.g., one or more audio signals) based on one or more criteria. As described herein, after determining which audio signals should be prioritized over other audio signals (and therefore emphasized), the controller may apply different scalar gains (e.g., and / or vector gains) in response to user adjustments received at volume control 12 (and / or volume control 82 of the audio output device). Therefore, when using these gain-adjusted signals to drive speaker 22, the volume level of the sound from the prioritized audio signal can be higher (e.g., a larger output sound level) than that of the less prioritized audio signal. More details on the application of scalar gain are described herein.
[0089] In one aspect, audio signal priority can be determined based on whether a local user is speaking (e.g., one or more remote users on a remote device participating in a call with a local device). Specifically, the curve selector determines whether the microphone signal generated by microphone 23 includes the local user's voice based on the output of VAD 53. For example, the curve selector receives a VAD signal generated by VAD 53 and determines whether the VAD signal indicates that the microphone signal includes voice (e.g., whether the signal has a high signal level). In response to determining that the output of VAD indicates that voice has been detected, the curve selector can prioritize downlink audio signals over audio signals of media content.
[0090] In one respect, priority is based on whether the local user's voice is detected within a certain period of time. Specifically, once the local user has been speaking for a period of time, the curve selector can prioritize the downlink audio signal of the call. For example, when participating in a call, the device can play media, such as a movie, during the playback session. Sometimes, the local user may be talking to one or more remote users. In this case, when the local user (and the remote user) have been engaged in (e.g., for a long time) the conversation, the VAD signal can detect voice for a period of time. In response, the curve selector can prioritize the downlink signal because the local user may want to hear the remote user rather than the media content. However, in other cases, the local user may want to only comment or make a brief statement without engaging in the conversation. In this case, the VAD signal may detect voice for less than that period of time. Therefore, the curve selector can choose not to prioritize the downlink audio signal. In one respect, since the local user does not intend to have a long conversation with the remote user, instead of prioritizing the downlink audio signal, the curve selector can prioritize the media content over the downlink signal. In another respect, the curve selector can choose not to prioritize either signal. As described in this article, when two signals are not prioritized (or can be prioritized equally), the controller can apply a similar scalar gain after the volume control receives a user adjustment. This article describes more about applying scalar gain.
[0091] On the other hand, the curve selector can prioritize signals based on sensor data received from one or more sensors in sensor 10. For example, the selector can prioritize downlink audio signals based on the voice of a local (and / or remote) user. Specifically, the curve selector can perform speech recognition (e.g., by using a speech recognition algorithm) to analyze audio signals (e.g., microphone signals and / or downlink audio signals) to find (or recognize) speech within them. Specifically, the controller can analyze the audio data of the signal according to an algorithm to identify words or phrases contained therein. The curve selector can determine priority based on the identified words or phrases. For example, the selector can use the identified words or phrases to perform a table lookup on a data structure that associates the priority value of the downlink signal with one or more words or phrases. After determining that the priority value is above a threshold, the curve selector can prioritize the downlink signal.
[0092] On the other hand, a curve selector can prioritize downlink audio signals based on the audio data contained within them. For example, similar to the prioritization of VAD signals based on the microphone signals described above, a curve selector can prioritize downlink signals based on whether the VAD signal detects speech within the downlink audio signal (e.g., over a period of time). In one aspect, a curve selector can prioritize two or more downlink audio signals from two or more remote devices differently. For example, after detecting speech in one downlink signal, the curve selector can prioritize that downlink signal over other downlink signals and / or audio signals of the media content. Thus, when a remote user is speaking, speech from that user can take precedence over other remote users who may or may not be speaking at that time.
[0093] In some respects, the curve selector can prioritize one or more downlink signals and / or one or more audio signals of media content based on gestures performed by a local user. The curve selector can receive sensor data indicating whether the user is performing a gesture associated with the local user wanting the sound of one or more signals to be more prominent than other sounds being played. Specifically, the curve selector can use the sensor data to determine whether the user is gesturing toward an object displayed on the display screen 25 of the local device to emphasize the sound associated with that object. As described herein, the local device can display a video representation of a remote user on its display screen. The curve selector can determine whether the local device wants to emphasize the sound of one or more remote users. For example, the curve selector can determine whether the user intends to emphasize the sound of the remote user by determining whether the local user is viewing (or focusing on) the video representation of the remote user. Specifically, the curve selector can receive image data captured by camera 24, where the camera's field of view includes at least a portion of the local user (e.g., the local user's face). Using the image data, the curve selector can determine whether the user's gaze of at least one eye is focused on the video representation of the remote user. If so, the curve selector can prioritize the downlink audio signal associated with the remote user's device.
[0094] On the other hand, the curve selector can determine whether a user intends to emphasize the voice of a remote user based on motion data (e.g., generated by IMU 11). Specifically, the curve selector can determine whether a local user is moving at least a portion of the local device, indicating that the local user is gesturing toward the video representation of the remote user. For example, when the local device is an electronic device in which a display screen is integrated, the curve selector can determine whether the local user is tilting the display screen toward the direction of the video representation. For example, the selector determines whether the screen is tiled around a central axis that extends through the display screen in a direction relative to the central axis toward the video representation displayed on the display screen. For example, see reference... Figure 3 When a call is made, the display 25 of local device 2 faces local user 30. A curve selector can determine, in response to detecting motion data indicating that display 25 (or local device 2) is tilting away from the local user around a lateral axis (e.g., the X-axis) that passes laterally through the center point of the display, that the local user wishes to prioritize the voice of remote user 31. Alternatively, gestures can be user selections of a video representation. For example, a user might perform a touch selection at the location where a video representation is displayed on the display (possibly a touchscreen).
[0095] In some aspects, the curve selector can determine, based on sensor data, whether a user intends to emphasize one or more sounds associated with media content. As described herein, media content may include several audio signals, each associated with a sound source. For example, when the media content is an XR environment, a local device may display visual content of the XR environment on display screen 25 and may drive speaker 22 with one or more audio signals that include the sound of the XR environment. Specifically, these audio signals may each be associated with one or more objects displayed in the XR environment. In some aspects, the controller may spatially render (e.g., by applying HRTF) these audio signals to provide a 3D sound experience to the user, wherein the sound sources are perceived at different locations within the sound space. In one aspect, the curve selector can determine, based on a user gesture toward a displayed object associated with the sound, whether a user wishes to emphasize a sound within the XR environment. Specifically, the curve selector can perform this determination in a manner similar to that described with respect to a remote user. For example, the curve selector may determine that the user's gaze is focused on the object and / or determine, based on motion data, that the user is tilting the display screen in the direction toward the object. Based on this determination, the curve selector may prioritize that sound over one or more other sounds within the XR environment. Although the priority is based on objects within the XR environment, this determination can be performed on any type of media content (such as movies or video games) that includes image data and one or more audio signals.
[0096] On another front, the curve selector can determine which signals to prioritize based on the specific operational function the controller is performing. As described herein, the controller may perform an ANC function to enable noise immunity for speaker 22 (and / or speaker 83). In addition to (or instead of) this function, the controller may perform a transparency function, in which the device plays ambient sounds. In one aspect, sound priority can be determined based on which function the controller is performing. For example, if the ANC function is performed to block out ambient sounds, the controller may prioritize the sound of media content. Conversely, if the controller is performing a transparency function, the controller may prioritize downlink audio signals because the local user may have activated the transparency function to hear their own voice during a conversation.
[0097] As described above, a curve selector can prioritize sounds associated with a playback session and / or call based on one or more criteria. On the other hand, a curve selector can prioritize other sounds (or groups of sounds) being played by a local device. For example, a curve selector can prioritize sounds generated by one or more software applications executing on a local device (e.g., its controller 20) based on one or more criteria. Software applications may include a telephone application performing a call between the local device and one or more remote devices, one or more media applications performing a media playback session, and / or other applications that may be executing within the local device (e.g., messaging applications, alarm applications, etc.). In one aspect, the priority of audio signals associated with an application (e.g., including the application's sounds) can be based on the order in which the local device begins executing the application. For example, refer to... Figure 3 The curve selector can prioritize video calling applications over media applications that are playing media content, because both the local and remote devices have already participated in the video call before the local device initiates the media application and joins the playback session. In one respect, subsequently executed applications can be given lower priority than applications already executed by the local device.
[0098] In one aspect, a curve selector can prioritize several sounds on a sliding scale based on the criteria mentioned herein. For example, a curve selector can prioritize some sounds as "high priority," some as "medium priority," and others as "low priority." In some aspects, the selector can digitally prioritize audio signals, where high-priority audio signals have high values (e.g., "10"), and lower-priority audio signals have smaller values (e.g., "1"). In other aspects, some audio signals can have the same priority. For example, when a call has been initiated with multiple (e.g., three or more) remote users, the curve selector can prioritize the downlink audio signals of two remote devices to the same (e.g., high) priority in response to determining that speech is detected in both signals.
[0099] In one aspect, curve selector 54 is configured to determine (or select) the volume-gain curve of at least one audio signal that the local device is playing (e.g., mixed and played) during a call and / or playback session. Specifically, once the selector prioritizes the audio signals (e.g., determining which sounds the local user wants to be more prominent than others), the selector can determine a curve for each audio signal that indicates the amount of gain (or attenuation) to be applied to the signal based on user adjustments to volume control.
[0100] In one aspect, a volume-gain curve can associate each of several volume settings of a volume control with (e.g., different) scalar gain values that can be applied to their respective audio signals (e.g., in response to a user adjustment of the volume control). In other words, the curve is a function of gain relative to a volume setting. For example, a volume-gain curve can associate a scalar gain value with a volume setting of a volume control that indicates the position and / or orientation of the volume control, as described herein. In another aspect, a volume-gain curve can associate one or more vector gains with a volume setting, which can allow the control to select one or more frequency bands where one or more selected gains will be applied to the audio signal. In one aspect, the volume setting can be a numerical value (e.g., 1-10), where each value is associated with an adjustment to the volume control, and each numerical value can be associated with a scalar gain value. In another aspect, a volume-gain curve can indicate the signal level (e.g., in dB) of the audio signal at a particular volume setting. Specifically, the curve can indicate the desired signal level (or gain) of the audio signal at each volume setting. This is in Figure 5 As shown in the diagram. Based on the required signal level, the controller can determine the gain adjustment to be applied to each specific audio signal.
[0101] In one aspect, the curve selector can select the curve of an audio signal based on the priority of the signal. Specifically, the curve selected for a high-priority audio signal may have a lower rate of change than the curve selected for another audio signal with lower priority. For example, after determining that a local user is speaking (e.g., based on a VAD signal), the curve selector can prioritize the downlink audio signal and select a first curve for the downlink signal and a second curve for the audio signal of the media content, wherein the first curve has a lower rate of change than the second curve. As described herein, upon receiving a user adjustment to volume control, such as reducing the overall volume level (e.g., decreasing the current volume setting by one), the controller can use the first curve to determine a first gain adjustment based on the (first) gain of the first curve associated with the reduced volume setting, and can use the second curve to determine a second gain adjustment based on the (second) gain of the second curve associated with the reduced volume setting, wherein the first gain is lower than the second gain. Once determined, the controller can reduce (or attenuate) the signal level of the downlink signal according to the first gain from the first curve, and reduce the signal level of the audio signal according to the second gain from the second curve. Therefore, the signal level of the audio signal can be reduced to a greater level than the downlink audio signal, which, when mixed and played by a local device, results in the downlink audio signal having a higher volume level than the audio signal. This volume variation allows the downlink signal to stand out more than the audio signal (because the downlink signal will have a higher sound output level), thus allowing the local user to participate in conversations with the remote user without being distracted by the sound of the media content. This article describes more about applying scalar gain values.
[0102] As described herein, a volume-gain curve can be a function of gain (e.g., a scalar gain value) relative to the volume setting of volume control 12. In one aspect, a volume-gain curve can be a linear function of gain relative to a volume setting, where different curves can have different slopes based on their associated audio signal priorities, as described herein. Figure 5 An example of a linear volume-gain curve is shown and described, where curve 61 for the downlink signal has a smaller slope than curve 62 for the audio signal of the media content. More details about these curves are described herein. On the other hand, a curve can be a non-linear function. On the other hand, a curve can be any type of function of gain relative to a volume setting.
[0103] In one aspect, the curves can be stored in the memory of a local device (e.g., the memory of controller 20). Specifically, the curves can be stored in a data structure (e.g., a lookup table), where each curve is stored in a specific lookup table. In some aspects, the curves can be predefined curves determined in a controlled environment (e.g., in a laboratory). In another aspect, at least some curves can be learned through machine learning operations or through user input. For example, the curves can be user-defined (e.g., by a local user of the local device). In some aspects, the curves can be stored in memory.
[0104] As described above, curve selector 54 can select volume-gain curves for at least some audio signals based on a determined priority of the signals. In one aspect, the selector can select multiple (two or more) curves for at least one audio signal. As shown in the previous example, when the downlink signal has a higher priority, the curve selector can select a first curve with a lower rate of change for the downlink audio signal and a second curve with a higher rate of change for the audio signal. In one aspect, these selections can correspond to a user adjustment received from volume control to reduce the overall volume level of the local device, because when the user reduces the volume, the output sound level of the downlink signal will be greater than the output sound level of the audio signal to emphasize the downlink signal. In some aspects, the curve selector can select different curves for the audio signal corresponding to a user adjustment to the volume control to increase the overall volume level. For example, the curve selector can select a third curve for the downlink signal and a fourth curve for the audio signal, where the third curve has a higher rate of change than the fourth curve. As the volume level of the local device increases, this will cause the output sound level of the downlink signal to increase beyond the volume of the audio signal.
[0105] On the other hand, instead of choosing different curves, the selected curves can be used interchangeably between signals. For example, in response to a user adjustment to the volume control that lowers the overall volume level, the controller can use a first curve to determine a first gain of the downlink audio signal and a second curve to determine a second gain of the audio signal, as described herein. Alternatively, in response to a user adjustment to the volume control that increases the overall volume level, the controller can use the second curve to determine the gain of the downlink audio signal and the first curve to determine the gain of the audio signal. Further details regarding the use of curves to determine gain are described herein.
[0106] In some respects, the selection of a curve can be based on one or more audio signal processing operations that have been (or will be) performed on one or more of these signals. For example, in response to VAD53 detecting speech (e.g., within a downlink audio signal and / or microphone signal), curve selector 54 can adjust the selection of a volume-gain curve for the audio signal based on whether a scalar gain value has been applied to the audio signal to reduce the signal level, as described herein. Specifically, when scalar gain 90 is not applied (e.g., in response to a low signal level in the VAD signal), the selected curve for the audio signal may have a lower rate of change than the selected curve for the audio signal. Further details regarding the selection of different curves by the curve selector based on audio processing operations are described herein.
[0107] In one respect, instead of selecting a volume-gain curve (or otherwise), the selector can determine different rates of change of gain values for audio signals based on priority. Specifically, the selector can choose a low rate of change for audio signals with high priority and one or more high rates of change for audio signals with lower priority. For example, typically, a controller can adjust the signal level in a similar way as the volume level increases or decreases (e.g., applying gain, such as 6dB, when the volume increases, and applying a similar gain or attenuation reduction, such as -6dB, when the volume decreases). However, when different rates of change are selected, the attenuation can vary depending on the rate. For example, if the high rate of change is twice the normal rate of change, the signal level of a lower-priority audio signal can be attenuated to twice the specific attenuation (e.g., -12dB).
[0108] In some respects, these selections can correspond to a user adjustment to the volume control that reduces the overall volume. As described herein, the selector can select different (e.g., inversely proportional) rates of change for the audio signal to increase the user adjustment to the overall volume. As described herein, these rates of change can be used by a gain selector to determine a new gain based on the user adjustment to the volume control.
[0109] Audio signal gain selector 55 is configured to receive one or more volume-gain curves selected from one or more audio signals for media content by curve selector 54, and downlink signal gain selector 56 is configured to receive one or more volume-gain curves selected from one or more downlink audio signals for a call by curve selector. In one aspect, each of these gain selectors is configured to determine one or more scalar gain values using its corresponding received gain curve based on user adjustments to volume control 12.
[0110] In one aspect, each of these gain selectors is configured to receive a control signal generated by volume control 12, which can be in response to a user adjustment received by the control. Specifically, the control signal can indicate the (current or) adjusted volume setting of the volume control. For example, when the volume control is a slider in the initial (or mute) position (e.g., UI), and the user adjusts the slider to the midpoint of its sliding range, the control signal may indicate a 50% volume setting. In another aspect, the volume setting can be a numerical value associated with a user adjustment, as described herein. For example, when the control signal is a rotatable knob currently oriented at 300°, the volume setting may be at setting 15 out of 18, which may correspond to approximately 80% of the total volume level. Upon receiving a user adjustment to rotate the knob down to 280°, the control signal may indicate that the volume setting has decreased by one, to 14.
[0111] In one aspect, each of these gain selectors can select a gain adjustment (e.g., a gain for increasing at least a portion of the signal level or an attenuation for decreasing at least a portion of the signal level) as a scalar gain value from its corresponding received volume-gain curve associated with the (adjusted) volume setting of the volume control. For example, when the curve is a linear function of gain relative to the volume setting, the gain selector can select the gain along the curve mapped to the received volume setting. In another aspect, the gain selector can select the gain difference between the original volume setting and the new volume setting. For example, when decreasing the total volume level, the volume setting can be decreased by one, where the previous volume setting and the decreased volume setting both correspond to different scalar gain values. Thus, when selecting a scalar gain value, the gain selector can select the difference between the gain value associated with the previous volume setting and the gain value associated with the new (or decreased) volume setting. In some aspects, the gain selector can indicate whether the selected scalar gain value is to be applied as gain to increase the signal level or as attenuation to decrease the signal level.
[0112] On the other hand, a gain selector can choose gain adjustment based on the gain of a volume-gain curve associated with a volume setting. Alternatively, the selection of scalar gain can be based on the desired signal level of the curve associated with the volume setting. As described herein, a volume-gain curve can indicate the desired signal level. In this case, after determining the desired signal level associated with the volume setting (e.g., -30 dBFS), the gain selector can choose to increase (or decrease) the signal level of the audio signal to the desired level using scalar gain. In this case, when the desired signal level is -30 dBFS and the audio signal is at -20 dBFS, the gain selector can apply -10 dB to decrease the signal level.
[0113] On the other hand, a gain selector can determine a scalar gain value based on user adjustments to volume control in other ways. For example, as described herein, instead of a curve selector determining a curve (or otherwise), the selector can determine the rate of change of gain for one or more signals. In one aspect, a gain selector can utilize the rate of change to determine a new gain value based on an adjusted volume control. For example, when adjusting the volume control to reduce the overall volume by decreasing the volume setting by one, the gain selector can use the rate of change to adjust the current gain based on the changed volume setting. On the other hand, when a curve is in a lookup table, the gain selector can perform a table lookup on a data structure storing the lookup table using the adjusted volume setting. Specifically, gain selector 55 can perform a table lookup on a data structure that associates one or more gains for audio signals of streaming media content with different user adjustments (e.g., volume settings) to select the gain associated with the (current) volume setting, and gain selector 56 can perform (e.g., similar) a table lookup on a data structure to select another gain for a downlink audio signal, where the two gains can be different (or the same). In some respects, the gain selector can use any method to determine the appropriate scalar gain value for the signal of the call and / or playback session.
[0114] Scalar gains 57 and 58 receive the audio signal of the media content and the downlink audio signal associated with the call, respectively, and are configured to process their respective signals based on a selected scalar gain value. Specifically, each of these scalar gains is configured to apply gain adjustment to its respective signal to produce an adjusted signal, which may have a higher and / or lower signal level based on the adjustment. For example, scalar gain 57 may apply gain adjustment to attenuate the signal level of the audio signal based on the scalar gain value selected by gain selector 55, and scalar gain 58 may apply gain adjustment to attenuate the signal level of the downlink signal based on the scalar gain value selected by gain selector 56, different from the gain adjustment of scalar gain 57 (e.g., in response to a user lowering the volume at a volume control). In some aspects, each of scalar gains 90, 57, and / or 58 may perform similar operations to decrease or increase the signal level of the associated signal.
[0115] Mixer 59 is configured to receive processed (e.g., gain-adjusted) signals from scalar gains 55 and 58, and is configured to perform matrix mixing operations, such as to produce a mixed content of the two signals. The controller can use the mixed signal to drive speaker 22 to play call audio, as well as play media content for a session, where the audio of the signal has an output audio level at or below the current total volume level of the local device. Alternatively, the mixer can receive one or more unprocessed signals (e.g., ungain-adjusted signals). For example, the mixer may receive a downlink audio signal from call manager 52 instead of a processed downlink audio signal from scalar gain 58.
[0116] In one aspect, the controller may optionally perform additional DSP operations. For example, the controller may perform spatial rendering operations (by applying spatial filters, such as Head Related Transfer Function (HRTF)) on one or more of these signals (and / or the mixed content) to generate binaural audio signals for driving one or more speakers (e.g., left and right speakers), as described herein. As another example, when the local device (and / or audio output device) includes one or more speakers, the controller may render HOA representations of the audio signals and downlink signals to generate one or more speaker driver signals (e.g., based on a predefined speaker configuration). Controller 20 may then use the processed mixed content to drive speaker 22, as described herein.
[0117] As illustrated herein, the controller includes an audio signal gain selector 55 and a scalar gain 57 for processing audio signals of media content, and a downlink signal gain selector 56 and a scalar gain 58 for processing downlink audio signals. In one aspect, the controller may include more or fewer gain selectors and / or scalar gains. For example, when the media content comprises two or more audio signals, the controller may utilize the respective gain selectors and scalar gains to process each of these signals in order to adjust the signal level of each of these signals based on user adjustments to volume control, as described herein.
[0118] As described above, the controller can prioritize signals for call and / or playback sessions to determine gain adjustment based on user adjustments to volume control. On one hand, the volume-gain curve selector may not prioritize one or more of these signals. In this case, the controller can apply the same (or similar) gain adjustment to the unprioritized signals. On the other hand, signals with the same priority can also be similarly gain-adjusted.
[0119] In one respect, the order of operations described herein can be different. For example, the volume-gain curve selector 54 can be configured to determine the curve in response to a user adjustment received by the volume control (e.g., prioritizing the signal and determining the curve based on the priority). In this case,
[0120] Figure 5 An example of a volume-gain curve based on one aspect is shown. Specifically, the graph shows volume-gain curve 61 for the downlink signal and volume-gain curve 62 for the audio signal of the media content. As shown, each of these curves is a linear function of gain (or gain adjustment) relative to a volume setting, where each curve represents the desired signal level for each corresponding signal. As shown, the curves are plotted in graph 60, where the Y-axis represents the gain (or gain adjustment) relative to the dynamic range of the audio system 1 in the digital domain (e.g., relative to full-scale decibels (dBFS)). Thus, the top of the Y-axis represents the maximum signal level, where each -10 dB step below 0 dB represents the gain (or attenuation) to be applied to the signal to attenuate it below the maximum level. In one aspect, the dynamic range of the audio system can be varied based on the bit depth of the digital audio data. The X-axis of the graph includes the volume setting for volume control, where the zeroth setting represents the lowest total volume output level of the local device (e.g., mute), and the tenth setting represents the highest total volume output level.
[0121] As shown in the figure, curves 61 and 62 have different slopes. For example, the audio signal curve 62 has a slope of 1, starting at -100dB at the zero volume setting and reaching 0dB (maximum permissible signal level) at the tenth volume setting. Conversely, the downlink signal curve 61 has a lower slope of 2 / 5, starting at -40dB at the zero volume setting and extending to 0dB at the tenth setting. In one aspect, the difference in slope can correspond to a determined signal priority, as described herein. For example, due to having a higher priority than the audio signal, the downlink curve 61 may have a lower slope than the audio signal curve 62. Therefore, as the overall volume level decreases, the signal level of the downlink signal decreases less than that of the audio signal.
[0122] Graph 60 also illustrates the gain change in response to a decrease of one in the volume setting. For example, the volume control might have a current volume setting of 7. In this case, the audio signal is -30dB and the downlink signal is -12dB. In some respects, the controller can apply one or more scalar gain values to these signals so that both signals have those desired levels. In response to receiving a user adjustment, the volume control can decrease its volume setting by one, up to six. Therefore, the controller can determine the gain adjustment to the signal based on the new volume setting using curves 61 and / or 62. In this case, as the gain decreases from -30dB to -40dB, the controller can reduce the downlink signal level by 10dB, and as the gain decreases from approximately -12dB to -16dB, the controller can reduce the downlink signal level by 4dB.
[0123] Curves 61 and 62 are also shown to intersect at the maximum volume setting. In this case, the gain applied to both signals will increase their respective signal levels to their respective maximum signal levels. In one respect, at the maximum signal level, the signal level of the audio signal of the media content may be higher than the signal level of the downlink signal because the audio signal may have initially been controlled to be higher than the signal level of the downlink signal, as described herein.
[0124] On one hand, the curves can have different slopes depending on whether the volume control receives user adjustments to increase or decrease the overall volume level, as described herein. On the other hand, any of these curves can be a function of a different type. For example, downlink curve 61 could be a linear function of gain relative to volume setting, while audio signal curve 62 could be a non-linear function of gain relative to volume setting.
[0125] Figure 6 This is a flowchart of one aspect of a process 70 for adjusting the overall volume level of audio system 1 using (e.g., master) volume control 12. In one aspect, this process can be performed by a local device 2 of audio system 1 (e.g., its controller 20). Specifically, at least some of the operations described herein can be performed by... Figure 4 At least some of the operation boxes described in the text are executed.
[0126] Process 70 begins when controller 20 initiates a call (e.g., a telephone call or video call) between local device 2 and one or more remote devices 3 (at box 71). As described herein, the call may be initiated by call manager 52 in response to receiving a request from a local user. In one aspect, the call may be initiated in response to receiving an incoming call from one or more remote devices. In this case, the call may be initiated by call manager in response to a user accepting the call (e.g., via the user selecting a UI item in the phone application for answering the call displayed on display 25 when an incoming call signal is received from a remote device).
[0127] During a call, controller 20, acting as local device 2, initiates a joint media playback session in which the local device and one or more remote devices independently stream media content for synchronous playback (at box 72). For example, joint media playback session manager 47 can initiate playback based on user input. In one aspect, the playback session can be initiated between all devices making the call. In another aspect, a playback session can be initiated between the local device and at least some of the remote devices. In this case, the local user can define which remote devices will participate when initiated. In some aspects, initiating a joint media playback session may be in response to controller 20 receiving an initiation request from one or more remote devices and / or media content server 5.
[0128] The controller receives at least one downlink (audio) signal associated with the call, and at least one audio signal associated with the media content (at box 73). For example, the local device receives downlink audio signals from at least some of the remote devices involved in the call with the local device. Furthermore, the local device may receive audio data and / or image (or video) data associated with the media content of the playback session. For example, the media content may consist only of audio data (as one or more audio signals), such as a musical piece to be played simultaneously by the local device and remote devices. Alternatively, the media content may include both audio and image (or video) data, such as for a movie. Or, the media content may include other types of content, such as XR presentations (e.g., virtual reality environments) or video games (e.g., user-interactive content). In this case, initiating a joint media playback session may include independently streaming image data of the XR presentation for display on the local device's screen.
[0129] Controller 20 drives a speaker (e.g., speaker 22) (at box 74) at a total volume level using a mixture of the downlink signal of the call and the audio signal of the media content. In one aspect, the controller is capable of driving the speaker at a specific total volume level with the mixture. For example, the speaker can be driven when the volume control is at a specific volume setting (e.g., at 7 out of 10 volume settings). In some aspects, the controller can drive the speaker with the mixture of signals before applying gain adjustments based on a determined signal priority, as described herein. Therefore, the audio signal can have a signal level greater than that of the downlink signal, as described herein.
[0130] In one aspect, the controller can spatially render the signal by applying one or more spatial filters based on spatial characteristics (e.g., elevation, azimuth, distance, etc.) so that when output through one or more speaker drivers, 3D sound is produced (e.g., giving the user the perception that sound is emanating from a specific location within the acoustic space). In another aspect, the controller can transmit signals (e.g., its mixed content) to the audio output device 6 to drive one or more speakers of the output device (e.g., speaker 83).
[0131] The controller receives user adjustments to the volume control (e.g., control 12) at box 75 to adjust the overall volume level. For example, a user can turn a control on the device (e.g., a digital crown or knob) or make a specific gesture to lower the volume. In response, the volume control can transmit a control signal to controller 20 indicating that the volume setting of the volume control has been lowered.
[0132] Controller 20 determines a first gain adjustment for the downlink signal and a second gain adjustment for the audio signal (at box 76) based on user adjustments to volume control. Specifically, the controller may determine a scalar gain value to apply to one or both of these signals, as described herein. In one aspect, the first gain adjustment may differ from the second gain adjustment. For example, a volume-gain curve selector determines the priority between the downlink signal and the audio signal based on one or more criteria. For example, the curve selector determines whether VAD 53 detects speech within the downlink audio signal (e.g., over a period of time). In response, the curve selector prioritizes the downlink audio signal over the audio signal and selects (e.g., a different) volume-gain curve for each signal based on this priority. In another aspect, the gain adjustment may be determined based on the streaming media content of a joint media playback session. For example, when the audio signal is associated with an object displayed on the display screen 25 of a local device, the controller may determine whether the local user is focusing on (e.g., viewing) the object on the display screen. If so, the curve selector may determine that the local user wants to prioritize that sound over the sound of the downlink audio signal.
[0133] Using the selected curve, the controller determines a first gain adjustment and a second gain adjustment of the signal associated with the adjusted (or varied) volume setting for volume control. For example, refer to... Figure 5 When the volume setting changes from 7 to 6, the controller can determine a first gain adjustment as a first attenuation of -4dB and a second gain adjustment as a second attenuation of -10dB. Therefore, in this case, the second gain adjustment can be greater than the first gain adjustment, so that when applied, the signal level of the audio signal is reduced more than the signal level of the downlink signal.
[0134] Therefore, based on the user's adjustment to the volume control, the controller 1) applies a first gain adjustment to the downlink signal of the call, and 2) applies a second gain adjustment (at box 77) to the audio signal associated with the media content. In this case, the gain-adjusted audio signal may be reduced more than the gain-adjusted downlink signal. Therefore, the signal level of the gain-adjusted downlink signal may be greater than the signal level of the gain-adjusted audio signal, making the downlink signal sound more prominent than the audio signal when both signals are used to drive the speaker. On the other hand, gain adjustment can cause the signal levels of both gain-adjusted signals to be lower than the highest signal level of the signal before adjustment, because the volume adjustment is reducing the overall volume level. Therefore, in this example, the signal levels of both signals may be lower than the original signal level of the audio signal.
[0135] The controller then drives speaker 22 (at box 78) with a mixture of the (gain-adjusted) downlink signal and the (gain-adjusted) audio signal at an adjusted total volume level. In one aspect, the controller can perform spatial rendering of the signal by applying one or more spatial filters, as described herein. Spatial rendering of the signal can produce one or more driver signals that the controller can use to drive one or more speakers of the local device (and / or audio output device 6).
[0136] Some aspects are capable of performing variations of process 70. For example, at least some specific operations in these processes may not be performed in the exact order shown and described. The specific operation may not be performed in a consecutive series of operations, and different specific operations may be performed in different aspects. For example, at least some of these operations may be omitted. For example, the operation described at box 74 may be omitted because the controller may not drive the speaker until it receives user adjustment for volume control. In this case, the controller may begin driving the speaker after receiving user input.
[0137] As described in this process, the controller can determine different gain adjustments for the downlink signal and audio signal based on the decrease in the overall volume level. In some aspects, the controller can perform a similar operation upon receiving an increase in the overall volume level. For example, the controller can receive a second user adjustment to the volume control that increases the previously decreased overall volume level (e.g., the volume setting from 7 to 6). For example, the volume control can receive user input to increase the volume setting back to 7. In response, the controller can apply a gain to the downlink audio signal and audio signal that is proportional to the attenuation previously applied to the signal, so as to return their respective signal levels to the levels before the original gain reduction.
[0138] On the other hand, the controller may perform at least some of the operations in process 70 to determine additional gain adjustment based on the increase in the total volume level. In some aspects, the controller may determine different volume-gain profiles when the total volume level increases, as opposed to when the total volume level decreases. Thus, the controller may determine different gain adjustments for at least one signal (e.g., when the volume is increased back to a previous level). For example, after a user adjusts to increase the volume level, the controller may apply gain adjustment to the audio signal that, when applied in response to a first user adjustment, increases the signal level of the audio signal by a greater amount than the previous attenuation decreased the signal level of the audio signal.
[0139] In one aspect, process 70 determines the gain adjustment of two signals: the downlink audio signal of the call and the audio signal of the media content. In some aspects, at least some operations can be performed to determine (and apply) the gain adjustment of two or more signals. As described herein, the media content may contain an XR presentation with multiple sounds, each of which is contained within an audio signal. For example, the controller may receive two audio signals, wherein the first audio signal contains ambient sound of the XR presentation and the second audio signal is associated with an object within the XR presentation (e.g., including the sound of the object). In some aspects, the controller may determine the gain adjustment based on the object within the XR presentation. For example, the controller may (e.g., using sensor data) determine that the sound of the object that the user wants to be displayed on the screen is more prominent than other sounds. In response to receiving a user adjustment to volume control that lowers the overall volume level, the controller may determine three gain adjustments: a first gain adjustment of the downlink audio signal, a second gain adjustment of the ambient sound, and a third gain adjustment of the second audio signal of the object. In one aspect, the first and second gain adjustments may attenuate their respective signals more than the third gain adjustment attenuates the second audio signal, such that when used to drive a speaker, the local user can (primarily) hear the sound of the object.
[0140] As described above, the controller can be configured to determine different gain adjustments for the downlink signal and the audio signal of the media content. In some aspects, the controller can determine the same gain adjustment for two or more signals to be used to drive the speaker. Specifically, the controller can apply similar (or identical) gain adjustments when it is determined that the local user does not want any sound to be more prominent than others. For example, as described herein, the controller can determine different gain adjustments in response to the VAD detecting speech contained within the downlink signal. In response to determining that the microphone signal does not contain speech (e.g., the VAD output is in a low signal state), the controller can apply the same gain adjustment to the downlink signal and the audio signal of the media content. In some aspects, when the controller determines that the downlink signal should be made more prominent, the same gain adjustment may be less than (or reduce the signal level of the signal by less) the gain adjustment applied to the audio signal. However, in some aspects, once speech is detected (or the controller determines that the user wants the downlink signal to be more prominent), the controller can then determine and apply different gain adjustments.
[0141] As described in process 70, operations can be performed to determine gain adjustments for downlink signals associated with remote devices and audio signals for streaming media content. In one aspect, operations can be performed to determine multiple (e.g., two or more) gain adjustments for multiple downlink audio signals received from two or more remote devices. In some aspects, the controller may apply the same gain adjustment to downlink signals after determining that one or more downlink signals have a high priority. In another aspect, the controller may apply different gain adjustments to one or more downlink signals.
[0142] As previously described, the controller can prioritize the sounds of some software applications based on the order in which they are executed (or currently being executed) on the local device, and thus determine different gains to apply to different sounds. For example, an audio signal may be associated with media content from a media application currently being executed by the local device. Subsequently, the local user may execute another separate application (e.g., a messaging application). Upon receiving a user adjustment to the volume control, the controller may attenuate the sound of the separate application more than the sound applied to the media application (e.g., by applying a higher gain value to the audio signal of the separate device and a lower gain value to the media application). On the other hand, the controller may attenuate the sound of an application that has been executed for a longer period of time (e.g., over a period of time) and then attenuate the sound of an application that has been executed for a shorter period of time (e.g., within that period of time).
[0143] On the other hand, the master volume control is a physical control that is part of the first electronic device. In some aspects, the master volume control is a user interface (UI) item displayed on the screen of the first electronic device. In other aspects, individual volume control includes input via gestures made by the user of the first electronic device.
[0144] In one aspect, a single volume control includes a plurality of volume settings, each volume setting defining a different total volume level of a first electronic device, wherein a user adjustment at the single volume control changes the current volume setting of the signal volume control to a new volume setting associated with a reduced total volume level. In some aspects, a downlink signal is associated with a first volume-gain curve that associates the plurality of volume settings with a first plurality of gains, and an audio signal of media content is associated with a second volume-gain curve that associates the plurality of volume settings with a second plurality of gains. The method further includes: in response to receiving a user adjustment, determining a first gain adjustment based on a first gain associated with the new volume setting using the first volume-gain curve, and determining a second gain adjustment based on a second gain associated with the new volume setting using the second volume-gain curve. In some aspects, the first volume-gain curve and the second volume-gain curve are linear functions of the gain relative to the plurality of volume settings of the single volume control, wherein the slope of the first volume-gain curve is greater than the slope of the second volume-gain curve, such that at each volume setting, the gain on the first volume-gain curve is lower than the gain on the second volume-gain curve. In one aspect, the first volume-gain curve and the second volume-gain curve are non-linear functions of the gain relative to the volume settings of the single volume control.
[0145] In another aspect, the user adjustment is a first user adjustment, the first gain adjustment is a first attenuation, and the second gain adjustment is a second attenuation. The method further includes receiving a second user adjustment for a single volume control of the first electronic device that increases the reduced total volume level back to the total volume level; 1) applying a first gain to a gain-adjusted downlink signal, and 2) applying a second gain to a gain-adjusted audio signal, the first gain and the second gain respectively increasing the signal levels of the gain-adjusted downlink signal and the audio signal. In one aspect, the first gain is proportional to the first attenuation, and the second gain is proportional to the second attenuation. In another aspect, when applied in response to receiving the first user adjustment, the second gain increases the signal level of the gain-adjusted audio signal by a greater amount than the second attenuation decreases the signal level of the audio signal. In some aspects, when the second user adjustment for a single volume control increases the total volume level of the first electronic device to a maximum volume level, the applied second gain increases the signal level of the gain-adjusted audio signal by a greater amount than the applied first gain increases the signal level of the gain-adjusted downlink signal.
[0146] In some aspects, determining the first and second gain adjustments includes performing a table lookup on a data structure at a single volume control using a user adjustment that correlates the gain for the downlink signal with the gain for the audio signal of the streaming media content for different user adjustments. In some aspects, the method also includes, in response to determining that the microphone signal does not include speech, that the first and second gain adjustments are identical.
[0147] In one aspect, the application of a first gain adjustment and a second gain adjustment reduces the signal levels of the downlink signal and the audio signal, respectively. The method further includes determining, before receiving a user adjustment, whether the downlink signal of a call includes voice based on the output of a voice activity detector (VAD); and in response to determining that the downlink signal includes voice, applying a third gain adjustment to the audio signal to reduce its signal level. In another aspect, the second gain adjustment reduces the signal level of the audio signal more when the downlink signal includes voice than when the downlink signal does not include voice. In some aspects, the downlink signal is a first downlink signal. The method further includes, while participating in a call and joint media playback session with a second and a third electronic device, receiving a first downlink signal from the second electronic device, a second downlink signal from the third electronic device, and an audio signal of media content; in response to a user adjustment, applying a first gain adjustment to the first downlink signal, a second gain adjustment to the audio signal, and a third gain adjustment to the second downlink signal, the third gain adjustment adjusting the signal level of the second downlink signal differently than the first gain adjustment adjusting the signal level of the first downlink signal.
[0148] In one aspect, the call is initiated by a telephone application being executed by a first electronic device, and the audio signal is a first audio signal from a media application being executed by the first electronic device. The method further includes receiving a second audio signal from a separate application being executed by the first electronic device; and determining a first gain adjustment, a second gain adjustment, and a third gain adjustment to be applied to the downlink signal, the first audio signal, and the second audio signal, respectively, based on the order in which the first electronic device begins executing the telephone application, the media application, and the separate application. In another aspect, the second audio signal is attenuated less than at least one of the first audio signal and the downlink signal.
[0149] As is widely recognized, the use of personally identifiable information should comply with privacy policies and practices that are generally accepted to meet or exceed industry or governmental requirements for protecting user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use, and the nature of authorized use should be clearly explained to users.
[0150] As previously described, one aspect of this disclosure may be a non-transitory machine-readable medium (such as microelectronic memory) on which instructions are stored, programming one or more data processing units (generally referred to herein as a "processor") to perform network operations and audio signal processing operations, as described herein. In other aspects, some of these operations may be performed by specific hardware components containing hard-wired logic. Alternatively, those operations may be performed by any combination of programmed data processing units and fixed hard-wired circuit components.
[0151] While certain aspects have been described and illustrated in the accompanying drawings, it should be understood that such aspects are merely illustrative of the broad disclosure and not limiting, and that this disclosure is not limited to the specific structures and arrangements shown and described, as various other modifications will be apparent to those skilled in the art. Therefore, the description is to be regarded as exemplary and not restrictive.
[0152] In some aspects, this disclosure may include the language "[element A] and [element B] at least one". This language may refer to one or more of these elements. For example, "at least one of A and B" may refer to "A", "B", or "A and B". Specifically, "at least one of A and B" may refer to "at least one of A and at least one of B" or "at least either A or B". In some aspects, this disclosure may include the language "[element A], [element B], and / or [element C]". This language may refer to any of these elements or any combination thereof. For example, "A, B, and / or C" may refer to "A", "B", "C", "A and B", "A and C", "B and C", or "A, B, and C".
Claims
1. A method performed by a first electronic device, the method comprising: while engaged in a call with a second electronic device, initiating a joint media playback session in which the first electronic device and the second electronic device independently stream media content for synchronized playback; driving a loudspeaker with a mixed content of a downlink signal of the call and an audio signal of the media content at a total volume level; receiving a user adjustment at a single volume control for the first electronic device to lower the total volume level; and in response to the user adjustment, based on a determination that a microphone signal produced by a microphone of the first electronic device prior to receiving the user adjustment at the single volume control includes speech of a user of the first electronic device for a period of time, determining a first volume-gain function for the downlink signal and a second volume-gain function for the audio signal; based on the user adjustment at the single volume control, 1) determining a first gain adjustment using the first volume-gain function and 2) determining a second gain adjustment using the second volume-gain function; applying the first gain adjustment to the downlink signal and the second gain adjustment to the audio signal; and driving the loudspeaker with a mixed content of the downlink signal and the audio signal at a reduced volume level, wherein the audio signal has a signal level greater than a signal level of the downlink signal prior to the application of the first gain adjustment and the second gain adjustment, wherein the second gain adjustment is greater than the first gain adjustment such that 1) a signal level of a gain-adjusted audio signal and a signal level of a gain-adjusted downlink signal are lower than the signal level of the downlink signal and 2) the signal level of the gain-adjusted downlink signal is greater than the signal level of the gain-adjusted audio signal.
2. The method of claim 1, wherein the single volume control is a master volume control of the first electronic device, the master volume control configured to provide bi-directional control to incrementally increase or decrease the total volume level.
3. The method of claim 1, further comprising determining the first gain adjustment and the second gain adjustment based on a streamed media content of the joint media playback session.
4. The method of claim 1, further comprising: based on a determination that the microphone signal includes the speech for less than the period of time prior to receiving the user adjustment at the single volume control, determining a fifth volume-gain function for the downlink signal and a sixth volume-gain function for the audio signal; based on the user adjustment at the single volume control, 1) determining a fifth gain adjustment using the fifth volume-gain function and 2) determining a sixth gain adjustment using the sixth volume-gain function; and applying the fifth gain adjustment to the downlink signal and the sixth gain adjustment to the audio signal.
5. The method of claim 1, wherein the first electronic device is executing a plurality of software applications, the plurality of software applications including at least a phone application that is executing the call and a media application that is executing the joint media play session, wherein the first volume-gain function and the second volume-gain function are determined based on the phone application and the media application.
6. A first electronic device, comprising: a sensor; a speaker; a processor; and a non-transitory machine-readable medium having instructions that, when executed by the processor, cause the first electronic device to: initiate a joint media play session with a second electronic device while engaged in a call with the second electronic device, in which the first electronic device and the second electronic device independently stream media content for synchronized play, drive the speaker with a total volume level of a mix of downlink signals of the call and audio signals of the media content, receive sensor data from the sensor of the first electronic device, determine, based on the sensor data, that the downlink signals include a first priority and the audio signals of the media content include a second priority different from the first priority, determine, based on the first priority, a first rate of change for a downlink volume level of the downlink signals and, based on the first priority and the second priority, a second rate of change for a media content volume level of the audio signals, receive a user adjustment to lower the total volume level at a single volume control for the first electronic device, and in response to the user adjustment, decrease the downlink volume level of the downlink signals at the first rate of change and decrease the media content volume level of the audio signals at the second rate of change, wherein, prior to the user adjustment, the media content volume level is higher than the downlink volume level, wherein the first rate of change is lower than the second rate of change such that 1) the decreased downlink volume level and the decreased media content volume level are lower than the downlink volume level of the downlink signals, and 2) the decreased downlink volume level is higher than the decreased media content volume level.
7. The first electronic device of claim 6, wherein the single volume control is a master volume control of the first electronic device that is configured to provide bi-directional control to gradually increase or decrease the total volume level.
8. The first electronic device of claim 6, wherein the non-transitory machine-readable medium has further instructions to, in response to the user adjustment, determine a first gain adjustment based on the first rate of change and a second gain adjustment based on the second rate of change, wherein the decreasing includes instructions to apply the first gain adjustment and the second gain adjustment to the downlink signals and audio signals, respectively.
9. The first electronic device of claim 6, wherein the sensor comprises a microphone and the sensor data comprises a microphone signal, wherein the non-transitory machine- readable medium has further instructions for determining, based on output from a voice activity detector (VAD), whether the microphone signal produced by the microphone of the first electronic device comprises speech of a user of the first electronic device, wherein the reducing is performed in response to determining that the microphone signal comprises the speech.
10. The first electronic device of claim 9, wherein the first priority and the second priority are determined based on the microphone comprising speech for a period of time prior to the user adjustment being received at the volume control.
11. The first electronic device of claim 6, wherein the sensor comprises a camera and the sensor data comprises one or more images, wherein the instructions for determining that the downlink signal comprises the first priority and the audio signal comprises the second priority are based on determining, from the images, that a user of the first electronic device is performing a gesture.
12. A method performed by a first electronic device, the method comprising: initiating a call with a second electronic device; initiating, during the call, a joint media play session in which the first electronic device and the second electronic device both independently stream visual content for an extended reality (XR) presentation on respective displays; receiving 1) a downlink signal associated with the call and 2) an audio signal associated with the visual content, wherein the audio signal comprises sounds of an object of the XR presentation; driving a loudspeaker with the downlink signal and the audio signal; receiving sensor data from a sensor of the first electronic device; determining, based on the sensor data, that a user of the first electronic device is performing a gesture toward the object of the XR presentation displayed on a display of the first electronic device; receiving a user adjustment of a volume control for adjusting an overall volume level of the first electronic device; based on the user performing the gesture toward the displayed object, determining a first rate of change for a downlink volume level of the downlink signal and a second rate of change for an XR presentation volume level of the audio signal; and based on the user adjustment, adjusting the downlink volume level at the first rate of change and adjusting the XR presentation volume level at the second rate of change, wherein the method further comprises: determining, based on the first rate of change and the user adjustment, a first gain adjustment for the downlink signal; and determining, based on the second rate of change and the user adjustment, a second gain adjustment for the audio signal that is different than the first gain adjustment, wherein the adjusting comprises applying the first gain adjustment to the downlink signal and applying the second gain adjustment to the audio signal, wherein determining that the user is performing a gesture toward the object displayed includes determining that the user desires the sound of the object to be more prominent than a sound contained within the downlink signal, wherein the first gain adjustment attenuates the downlink signal more than the second gain adjustment attenuates the audio signal.
13. The method of claim 12, wherein the sensor comprises a camera and the sensor data comprises image data, wherein determining that the user is performing a gesture toward the object displayed comprises: determining, based on the image data, that a gaze of at least one eye of the user is focused on the object displayed.
14. The method of claim 12, wherein the sensor comprises a motion sensor and the sensor data comprises motion sensor data, wherein determining that the user is performing a gesture toward the object displayed comprises: determining, based on the motion sensor data, that a user of the first electronic device is tilting the display of the first electronic device about a central axis that extends through the display of the first electronic device in a direction relative to the central axis toward the object displayed on the display of the first electronic device.
Citation Information
Patent Citations
Unified communications system and method
US20170026509A1