Audio adjustment based on user electrical signals

JP2024533078A5Pending Publication Date: 2025-06-06QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024513226
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-07
Filing Date
2022-06-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing audio playback devices often misalign the perceived direction of sound sources with the user's perspective when the device's orientation changes relative to the user, leading to incorrect spatial audio localization.

Method used

A system that adjusts audio playback based on electrical activity data from user head sources, such as EEG data, to dynamically reposition sound sources in the sound field according to the user's preferred location or estimated orientation.

Benefits of technology

Ensures that sound sources are perceived from the intended direction relative to the user, maintaining spatial audio accuracy even when the playback device changes position or orientation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The device includes a memory and one or more processors. The memory is configured to store instructions. The one or more processors are configured to execute the instructions to obtain electrical activity data corresponding to electrical signals from one or more electrical sources in the user's head. The one or more processors are also configured to execute instructions to render audio data and adjust locations of sound sources in a sound field during playback of the audio data based on the electrical activity data.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001]

[0001] This application claims the benefit of priority to commonly owned U.S. non-provisional patent application Ser. No. 17 / 467,883, filed Sep. 7, 2021, the entire contents of which are expressly incorporated by reference into this specification. [Technical field]

[0002]

[0002] The present disclosure relates generally to adjusting audio based on a user electrical signal. [Background technology]

[0003]

[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices, including wireless telephones, such as mobile phones and smartphones, tablet and laptop computers, that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. In addition, many such devices incorporate additional features, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Such devices can also process executable instructions, including software applications, such as web browser applications that can be used to access the Internet. Thus, these devices can include significant computing power.

[0004]

[0004] Such computing devices often incorporate functionality for playing spatial audio, with sounds that can be perceived as coming from the direction of an audio source. The direction of the audio source is typically mapped to the playback device. As an example, the audio may represent a person's speech that is perceived as coming from in front of a user looking at the playback device. However, if the user places the playback device on a desk, the speech will be perceived as coming from the desk rather than in front of the user. Summary of the Invention

[0005]

[0005] According to one implementation of the present disclosure, a device includes a memory and one or more processors. The memory is configured to store instructions. The one or more processors are configured to execute instructions to obtain electrical activity data corresponding to electrical signals from one or more electrical sources in a user's head. The one or more processors are also configured to execute instructions to render audio data based on the electrical activity data to adjust a location of the sound source in a sound field during playback of the audio data.

[0006]

[0006] According to another implementation of the present disclosure, a method includes obtaining, at a device, electrical activity data corresponding to electrical signals from one or more electrical sources in a user's head, and rendering audio data based on the electrical activity data to adjust a location of the sound source in a sound field during playback of the audio data.

[0007] According to another implementation of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to obtain electrical activity data corresponding to electrical signals from one or more electrical sources in a user's head. The instructions, when executed by the one or more processors, cause the one or more processors to render audio data based on the electrical activity data to adjust a location of the sound source in a sound field during playback of the audio data.

[0008]

[0008] According to another implementation of the present disclosure, an apparatus includes means for acquiring electrical activity data corresponding to electrical signals from one or more electrical sources in a user's head, the apparatus also including means for rendering audio data based on the electrical activity data to adjust a location of the sound source in a sound field during playback of the audio data.

[0009]

[0009] Other aspects, advantages, and features of the present disclosure will become apparent after consideration of the entire application, including the following sections: brief description of the drawings, form for implementing the invention, and claims. [Brief description of the drawings]

[0010] [Figure 1]

[0010] A block diagram of a particular illustrative aspect of a system operable to adjust audio based on a user electrical signal, in accordance with certain examples of the present disclosure. [Figure 2A]

[0011] 2 is a diagram of an example aspect of operations associated with relative location estimation that may be performed by the system of FIG. 1 in accordance with some examples of the present disclosure. [Figure 2B] 2 is a diagram of an example aspect of operations associated with relative location estimation that may be performed by the system of FIG. 1 in accordance with some examples of the present disclosure. [Diagram 3]

[0012] 2 is a diagram of an exemplary aspect of the operation of components of the system of FIG. 1 in accordance with some examples of the present disclosure. [Figure 4]

[0013] FIG. 2 illustrates an example of an integrated circuit operable to adjust audio based on a user electrical signal, according to some examples of the present disclosure. [Diagram 5]

[0014] FIG. 1 is a diagram of a mobile device operable to adjust audio based on a user electrical signal, according to some examples of the present disclosure. [Figure 6]

[0015] 1 is a diagram of a headset operable to adjust audio based on a user electrical signal, according to some examples of the present disclosure. [Figure 7]

[0016] FIG. 1 illustrates a diagram of a wearable electronic device operable to adjust audio based on a user electrical signal, according to some examples of the present disclosure. [Figure 8]

[0017] FIG. 1 illustrates a diagram of a voice-controlled speaker system operable to adjust audio based on a user electrical signal, according to some examples of the present disclosure. [Figure 9]

[0018] FIG. 1 is a diagram of a headset, such as a virtual reality, mixed reality, or augmented reality headset, operable to adjust audio based on a user electrical signal, in accordance with some examples of the present disclosure. [Figure 10]

[0019] 1 is a diagram of a vehicle operable to adjust audio based on a user electrical signal, according to some examples of the present disclosure. [Figure 11]

[0020] FIG. 1 is a diagram of an earphone operable to adjust audio based on a user electrical signal, according to some examples of the present disclosure. [Figure 12]

[0021] 2 is a diagram of a particular implementation of a method for adjusting audio based on a user electrical signal that may be implemented by the device of FIG. 1 in accordance with some examples of the present disclosure. [Figure 13]

[0022] 1 is a block diagram of a particular illustrative example of a device operable to adjust audio based on a user electrical signal, in accordance with some examples of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011]

[0023] During spatial audio playback, sounds may be perceived as coming from the direction of the audio source mapped to the playback device. As an example, the audio may represent a person's speech that is perceived as coming from in front of a user looking at the playback device. However, if the user or the playback device repositions so that the playback device is no longer in front of the user, the speech will be perceived as coming from the playback device that is no longer in front of the user.

[0012]

[0024] A system and method for adjusting audio based on a user electrical signal is disclosed. An audio player can render audio data during a first playback operation to include multiple locations of a sound source in a sound field. For example, during the first playback operation, a user listening to playback of the rendered audio data will simultaneously perceive sounds from the same sound source coming from each of the multiple locations in the sound field. Illustratively, the user perceives sounds as if the same sound source is replicated at each of the multiple locations. The audio player acquires electrical activity data corresponding to electrical signals generated from electrical sources (e.g., brain cells) in the user's head during the first playback operation. As an example, the electrical activity data includes electroencephalogram (EEG) data received from an in-ear sensor. The audio player identifies one of the multiple locations as a user-preferred location of the sound source based on the electrical activity data. During a second playback operation, the audio player renders the audio data to adjust the location of the sound source based on the user-preferred location. For example, a user listening to a playback of the audio data rendered during the second playback operation perceives sound from a sound source coming from a user-preferred location in the sound field. Thus, the audio player allows the direction of the sound source to be adjusted based on user preferences instead of being mapped to the playback device.

[0013]

[0025] Certain aspects of the present disclosure are described below with reference to the drawings. In this description, common features are designated by common reference numerals. Various terms used herein are used only for the purpose of describing particular implementations and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some implementations and plural in other implementations. To illustrate, FIG. 1 illustrates a device 102 that includes one or more processors ("processor(s)" 190 in FIG. 1), indicating that in some implementations the device 102 includes a single processor 190 and in other implementations the device 102 includes multiple processors 190.

[0014]

[0026] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference number is used for each, and the different instances are distinguished by the addition of a letter to the reference number. When features as a group or type are referred to herein (e.g., when no particular one of the features is referred to), the reference number is used without the distinguishing letter. However, when one particular feature of multiple features of the same type is referred to herein, the reference number is used with the distinguishing letter. For example, referring to FIG. 1, multiple locations are illustrated and associated with reference numbers 150A and 150B. When referring to a particular one of these locations, such as location 150A, the distinguishing letter "A" is used. However, when referring to these locations as any one or group of these locations, the reference number 150 is used without the distinguishing letter.

[0015]

[0027] As used herein, the terms "comprise", "comprises", and "comprising" may be used interchangeably with "include", "includes", or "including". Additionally, the term "wherein" may be used interchangeably with "where". As used herein, "exemplary" refers to an example, implementation, and / or aspect, and should not be construed as limiting, or as indicating a preferred or preferred implementation. As used herein, orthogonal terms (e.g., "first", "second", "third", etc.) used to modify an element, such as a structure, component, operation, etc., do not in themselves indicate a priority or order of the element with respect to another element, but merely distinguish the element from another element having the same name (apart from the use of orthogonal terms). As used herein, the term "set" refers to one or more of a particular element, and the term "plurality" refers to a multiple (e.g., two or more) of a particular element.

[0016]

[0028] As used herein, "coupled" may include "communicatively coupled," "electrically coupled," or "physically coupled," as well as (or alternatively) any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), and the like. Two devices (or components) that are electrically coupled may be included in the same device or in different devices, and may be connected via electronic circuits, one or more connectors, or inductive coupling, as illustrative and non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, and the like. As used herein, "directly coupled" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) with no intervening components.

[0017]

[0029] In this disclosure, terms such as "determining," "calculating," "estimating," "shifting," "adjusting," and the like may be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, "generating," "calculating," "estimating," "using," "selecting," "accessing," and "determining" may be used interchangeably. For example, "generating," "calculating," "estimating," or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated, such as by another component or device.

[0018]

[0030] 1, a particular exemplary embodiment of a system configured to adjust audio based on a user electrical signal is disclosed and generally designated 100. System 100 includes a device 102 configured to be coupled to one or more speakers 106 via an output interface 124.

[0019]

[0031] The device 102 is configured to be coupled to one or more sensors 104 via the input interface 114. In certain aspects, the one or more sensors 104 include an in-ear sensor, an electrode cap, a neural implant, a conductive screen, a non-wearable sensor, or a combination thereof. The device 102 is configured to be coupled to one or more spatial sensors 176 configured to generate spatial data 177 indicative of spatial information (e.g., at least one of a movement, a position, or an orientation) of a user 180. In certain aspects, the one or more spatial sensors 176 include an inertial measurement unit (IMU), a camera, a global positioning system (GPS) sensor, or a combination thereof.

[0020]

[0032] In some aspects, spatial data 177 (e.g., IMU data, image data, or both) is indicative of a change in location of user 180, a change in orientation of user 180, or both. For example, an IMU of one or more spatial sensors 176 integrated within a headset worn by user 180 generates IMU data indicative of movement of the headset corresponding to movement of head 182 of user 180. Spatial data 177 includes IMU data.

[0021]

[0033] In some aspects, spatial data 177 (e.g., GPS data, image data, or both) indicates a location, orientation, or both of user 180. For example, a camera of one or more spatial sensors 176 captures a first image of user 180 at a first time and a second image of user 180 at a second time. The first image indicates a first orientation of head 182 at the first time and the second image indicates a second orientation of head 182 at the second time. Spatial data 177 indicates the first image and a second image indicating the first orientation at the first time and the second orientation at the second time, as well as a change from the first orientation to the second orientation.

[0022]

[0034] The device 102 is configured to be coupled to one or more spatial sensors 178 configured to generate spatial data 179 indicative of spatial information (e.g., at least one of movement, position, or orientation) of the fiducial 188. In certain aspects, the one or more spatial sensors 178 include an IMU, a camera, a GPS sensor, or a combination thereof. In some aspects, the spatial data 179 (e.g., IMU data, image data, or both) indicates a change in location of the fiducial 188, a change in orientation of the fiducial 188, or both. In some aspects, the spatial data 179 (e.g., GPS data, image data, or both) indicates a location, orientation, or both of the fiducial 188. In some aspects, the fiducial 188 has a fixed location, a fixed orientation, or both. In these aspects, the spatial data 179 can indicate a fixed location, a fixed orientation, or both. For example, spatial data 179 may be based on configuration settings, default data, user input, or a combination thereof, as compared to that generated by one or more spatial sensors 178, and indicates a fixed location, a fixed orientation, or both. Device 102 is configured to adjust audio based on the user electrical signal using audio player 140.

[0023]

[0035] In some implementations, one or more components of system 100 are included in device 102 and one or more components of system 100 are included in a second device configured to be coupled with device 102. In an illustrative, non-limiting example, audio player 140 is included in device 102 (e.g., a phone, a tablet, a game console, a computing device, etc.) and one or more spatial sensors 176, one or more speakers 106, one or more sensors 104, or a combination thereof are included in the second device (e.g., a user head-mounted device such as a headset for user 180).

[0024]

[0036] The one or more sensors 104 are configured to generate electrical activity data 105 corresponding to electrical signals (e.g., brain waves) from one or more electrical sources 184 (e.g., brain cells) in a head 182 of a user 180. In certain aspects, the electrical activity data 105 includes electro-oculogram (EOG) data, EEG data, or both. The input interface 114 is configured to receive the electrical activity data 105 from the one or more sensors 104. In certain aspects, the input interface 114 includes at least one of an Ethernet interface, a universal serial bus (USB) interface, a Wi-Fi interface, a Bluetooth interface (a registered trademark of Bluetooth SIG, Inc., Washington, D.C.) interface, a serial port interface, a parallel port interface, or other type of data interface.

[0025]

[0037] The device 102 includes one or more processors 190. In certain aspects, the input interface 114, the output interface 124, or both are coupled to the one or more processors 190. The one or more processors 190 include an audio player 140. In certain aspects, the audio player 140 includes an audio adjustment unit 170 configured to adjust audio data. In certain aspects, the audio data 141A corresponds to sounds captured by one or more microphones. In certain aspects, the audio data 141A corresponds to audio generated by a game engine, an audio application, etc. In certain aspects, the audio data 141A corresponds to a combination of captured sounds and virtual sounds. The audio data 141A represents a sound field 142 (e.g., a three-dimensional (3D) sound field). During playback of the audio data 141A, the sound field 142 (e.g., the 3D sound field) can be reconstructed in a manner that allows a listener to distinguish a position and / or distance between the listener and one or more sound sources of the 3D sound field.

[0026]

[0038] In an illustrative, non-limiting example, the audio data 141A is based on or converted into one of the following formats: (i) traditional channel-based audio, meant to be played through loudspeakers in pre-specified positions; (ii) object-based audio, including discrete pulse-code-modulation (PCM) data for a single audio object with associated metadata including location coordinates (among other information); or (iii) scene-based audio, which involves representing the sound field using coefficients of spherical harmonic basis functions (also called "spherical harmonic coefficients" or SHC, "higher-order Ambisonics" or HOA, and "HOA coefficients").

[0027]

[0039] The audio player 140 is configured to perform a multi-location audio generation 164. For example, the audio adjustment unit 170 is configured to generate the audio data 141B by rendering the audio data 141A to have the sound of a sound source 186 corresponding to the multiple locations 150 of the sound field 142. The audio player 140 is configured to output the audio data 141B to one or more speakers 106 during the audio playback operation 144A. The one or more sensors 104 are configured to generate electrical activity data 105 during the audio playback operation 144A. The audio player 140 is configured to determine a user preferred location 167 by performing a preferred location estimation (preferred location estimation) 166 based on the electrical activity data 105. The audio player 140 is configured to perform a single location audio generation (single location audio generation) 168 based on the user preferred location 167. For example, audio adjustment unit 170 is configured to generate audio data 141C by rendering audio data 141A based on user preferred location 167 to adjust the location of sound source 186 in sound field 142.

[0028]

[0040] In some implementations, the device 102 corresponds to or is included in one of a variety of types of devices. In an illustrative example, the one or more processors 190 are integrated into a headset device that includes one or more speakers 106 and includes or is coupled to one or more sensors 104, as further described with reference to FIG. 6. In other examples, the one or more processors 190 are integrated into at least one of a mobile phone or tablet computing device described with reference to FIG. 5, a wearable electronic device described with reference to FIG. 7, a voice-controlled speaker system described with reference to FIG. 8, a virtual reality, mixed reality, or augmented reality headset described with reference to FIG. 9, or one or more earphones described with reference to FIG. 11. In another illustrative example, the one or more processors 190 are integrated into a vehicle coupled to one or more speakers 106 and one or more sensors 104, as further described with reference to FIG. 10.

[0029]

[0041] During operation, a user 180 activates or initiates operation of the audio player 140 to play audio data 141A. The audio data 141A corresponds to spatial audio data (e.g., human speech, sounds from birds, music from instruments, etc.) representing at least sounds of a sound source 186 from a location 150A in the sound field 142. For example, during playback of the audio data 141A, the sound field 142 may be reconstructed such that sounds from the sound source 186 are perceived by a listener as coming from the location 150A in the sound field 142 (e.g., 3D space). By way of illustration, speech from an actor (e.g., movie audio or an audiobook) may be perceived as coming from in front of the listener, and sounds from a passing car may be perceived as traveling from right to left behind the listener. The audio data 141A may represent additional sounds from one or more additional sources in the sound field 142.

[0030]

[0042] The audio player 140 is configured to adjust the location of the sound source 186 in the sound field 142 based on the user preferred location 167 or based on the relative location 161 (e.g., an estimated location of the user 180 relative to a reference 188). For example, the relative location 161 corresponds to an estimate (or proxy) of the user preferred location 167. Determining the user preferred location 167 includes playing audio to the user 180 with the sound of the sound source 186 perceptible from multiple locations 150 in the sound field 142. Meanwhile, the relative location 161 can be estimated in the background without the user 180 being aware of it. In some implementations, the audio player 140 adjusts the location of the sound source 186 based on the relative location 161 if the location confidence level 181 of the relative location 161 is equal to or greater than a confidence threshold (confidence threshold) 163. However, if the location confidence level 181 is less than the confidence threshold 163, the audio player 140 determines a user preferred location 167 by playing audio along with the sound of the sound source 186 from the multiple locations 150 and adjusts the location of the sound source 186 based on the user preferred location 167. If the location confidence level 181 does not meet the confidence threshold 163, then the sound from the sound source 186 is selectively played from the multiple locations 150 to determine the user preferred location 167.

[0031]

[0043] In certain aspects, audio player 140 performs relative location estimation 160 to determine relative location 161 and location confidence level 181 based on spatial data 177, spatial data 179, electrical activity data 105, or a combination thereof, as further described with reference to FIGS. 2A-2B. For example, relative location 161 corresponds to an estimated location of user 180 relative to reference 188, and location confidence level 181 indicates an estimated confidence associated with relative location 161. In some implementations, relative location 161 corresponds to an estimated location (e.g., location, orientation, or both) of user 180 relative to the estimated location (e.g., location, orientation, or both) of reference 188.

[0032]

[0044] In certain aspects, the reference 188 includes one or more of the device 102, a display device, a playback device, one or more speakers 106, a physical reference, a virtual reference, a fixed location reference, or a mobile reference. For example, the reference 188 can include a virtual reference (e.g., a building) that has a fixed location in the virtual scene. As another example, the reference 188 can include a virtual reference (e.g., a mobile virtual person) that can change location in the virtual scene. In some examples, the reference 188 can include a physical reference (e.g., an advertising display) that has a fixed location in the physical space (e.g., mounted on a wall). In other examples, the reference 188 can include a physical reference (e.g., a mobile device) that can change location in the physical space.

[0033]

[0045] The fiducials 188 are shown separate from the device 102 as an illustrative example. In other examples, the fiducials 188 can be integrated within the device 102. In some implementations, the fiducials 188 refer to reference points (e.g., specific locations). In other implementations, the fiducials 188 can have a multi-dimensional (e.g., two-dimensional or three-dimensional) shape, such as a square, cube, rectangle, plane, prism, triangle, pyramid, circle, sphere, ellipse, oval, etc.

[0034]

[0046] In certain implementations, audio player 140 initializes relative location 161 to correspond to fiducial 188 (e.g., a mobile phone screen) that is a predetermined distance (e.g., 12 inches) from user 180 and is oriented in front of (e.g., facing) user 180. In certain aspects, the predetermined distance is based on a configuration setting, a default value, user input, or a combination thereof. In certain aspects, audio player 140 initializes place confidence level 181 to be less than confidence threshold 163. In certain aspects, confidence threshold 163 is based on a configuration setting, a default value, user input, or a combination thereof. In some aspects, audio player 140 updates relative location 161 and place confidence level 181 based on movement of fiducial 188, movement of user 180, or both, as further described with reference to FIGS. 2A-2B.

[0035]

[0047] The audio player 140 performs a comparison 162 to determine whether to use the relative location 161 for single-location audio generation 168 or to determine a user-preferred location 167. For example, the audio player 140 compares the location confidence level 181 with a confidence threshold 163. In response to determining that the location confidence level 181 is equal to or greater than the confidence threshold 163, the audio player 140 proceeds to single-location audio generation 168 based on the relative location 161, as will be further described with reference to FIGS. 2A-2B. For example, the audio adjuster 170 generates audio data 141C based on the audio data 141A and the relative location 161. By way of example, generating the audio data 141C includes rendering the audio data 141A based on the relative location 161. Alternatively, in response to determining that the location confidence level 181 is less than the confidence threshold 163, the audio player 140 performs a multi-location audio generation 164 to determine a user-preferred location 167.

[0036]

[0048] In some aspects, the comparison 162 includes comparing the relative location 161 to a previous determination of the relative location 161. For example, in response to the audio player 140 determining that the difference between the relative location 161 and the previous determination of the relative location 161 is less than the location change threshold and the location confidence level 181 is greater than a second confidence threshold, the audio player 140 reverts to the relative location estimate 160 without adjusting the position of the sound source 186.

[0037]

[0049] Audio adjustment unit 170 generates audio data 141B based on audio data 141A during multi-location audio generation 164. For example, generating audio data 141B may include rendering audio data 141A to have multiple locations of sound source 186. Illustratively, audio data 141B may represent sound of sound source 186 from location 150A in sound field 142, sound of sound source 186 from location 150B in sound field 142, sound of sound source 186 from one or more additional locations in sound field 142, or a combination thereof.

[0038]

[0050] In a particular aspect, audio data 141A represents a sound of sound source 186 from location 150A, and generating audio data 141B includes adding sounds of sound source 186 from location 150B, one or more additional locations, or a combination thereof. In an alternative aspect, audio data 141A does not include sounds of any of sound sources 186, and generating audio data 141B includes adding sounds of sound source 186 from each of multiple locations 150.

[0039]

[0051] The audio player 140 performs the preferred location estimation 166 based on the audio data 141B. For example, the audio player 140 initiates an audio playback operation 144A of the audio data 141B via one or more speakers 106. For example, the audio player 140 provides the audio data 141B to the one or more speakers 106 via the output interface 124. In certain aspects, the output interface 124 includes at least one of an Ethernet interface, a Universal Serial Bus (USB) interface, a Wi-Fi interface, a Bluetooth® (registered trademark of Bluetooth SIG, Inc., Washington) interface, a serial port interface, a parallel port interface, or another type of data interface.

[0040]

[0052] The audio data 141B includes multiple locations 150 of sound sources 186 in the sound field 142 during the audio playback operation 144A. The audio player 140 acquires electrical activity data 105 from one or more sensors 104 via the input interface 114 during the audio playback operation 144A. The electrical activity data 105 corresponds to electrical signals from one or more electrical sources 184 in the head 182 of the user 180 during the audio playback operation 144A. For example, electrical signals are generated by the one or more electrical sources 184 (e.g., brain cells) while the audio data 141B is played to the user 180. In certain aspects, the audio player 140 outputs an alert (e.g., a visual alert) during the audio playback operation 144A indicating that an audio composition is being performed.

[0041]

[0053] The audio player 140 determines the user preferred location 167 of the sound source 186 based on the electrical activity data 105. For example, the audio player 140 processes the electrical activity data 105 using a preferred location model 174 and an output (e.g., an artificial neural network, a machine learning model, or both) of the preferred location model 174 that indicates that location 150B corresponds to the user preferred location 167 of the sound source 186.

[0042]

[0054] In some implementations, the audio player 140 determines the user preferred location 167 based on performing preferred source estimation, as further described with respect to FIG. 3. For example, the audio adjuster 170 generates the audio data 141B by rendering the audio data 141A to have a first location of a speech source and a second location of a non-speech source (e.g., a car) during the audio playback operation 144A. The audio adjuster 170 determines the user preferred location 167 based on the electrical activity data 105 acquired during the audio playback operation 144A. For example, the audio player 140 determines that the first location of the speech source corresponds to the user preferred location 167 in response to determining that the electrical activity data 105 indicates that a single source is tracked by the user 180. Alternatively, the audio player 140 determines that the second location of the non-speech source corresponds to the user preferred location 167 in response to determining that the electrical activity data 105 indicates that multiple sources are tracked by the user 180. For example, the human brain tracks speech even when the user 180 is listening to non-speech sounds (e.g., a car passing by), so electrical activity data 105 indicating multiple sources are being tracked by the user 180 corresponds to the user 180 listening to non-speech sounds (e.g., a car).

[0043]

[0055] The audio player 140 performs single location audio generation 168 based on the user preferred location 167. For example, the audio adjuster 170 generates audio data 141C based on the user preferred location 167 and the audio data 141A. To illustrate, generating the audio data 141C includes rendering the audio data 141A to have the user preferred location 167 (e.g., location 150B) of the sound source 186 in the sound field 142. As an example, the audio player 140 adjusts the location of the sound source 186 from location 150A in the sound field 142 (represented by the audio data 141A) to location 150B in the sound field 142 (represented by the audio data 141C).

[0044]

[0056] The audio player 140 initiates an audio playback operation 144B of audio data 141C through one or more speakers 106. For example, the audio player 140 provides the audio data 141C to the one or more speakers 106 via the output interface 124. The audio data 141C includes a user-preferred location 167 (e.g., location 150B) of a sound source 186 in the sound field 142 during the audio playback operation 144B. For example, the location of the sound source 186 is adjusted to the user-preferred location 167 (e.g., location 150B) during the audio playback operation 144B. Illustratively, the sound source 186 is perceived as coming from a single location in the sound field 142 during the audio playback operation 144B. In certain aspects, the single location of the sound source 186 is fixed to the user-preferred location 167 (e.g., location 150B) during the audio playback operation 144B. In an alternative embodiment, a single location of sound source 186 is initialized to a user preferred location 167 (e.g., location 150B) and changes during audio playback operation 144B. For example, sound source 186 may correspond to a flying bird and the location of the sound from the bird moves within sound field 142.

[0045]

[0057] System 100 thus enables rendering audio with sound that can be perceived as coming from the direction of a sound source 186 mapped to user preferred location 167 or an estimate of user preferred location 167 (e.g., relative location 161). As an example, if user 180 places a playback device (e.g., fiducial 188) on a desk, the location of sound source 186 can be adjusted to continue to be perceived as coming from in front of user 180 (e.g., user preferred location 167).

[0046]

[0058] 2A, a diagram 200 of an example aspect of the operations associated with relative location estimation 160 is shown. Relative location estimation 160 may be implemented by audio player 140 of FIG.

[0047]

[0059] The audio player 140 performs relative location estimation 160 to determine a relative location 161. The relative location 161 includes a distance 212, a relative orientation 263, or both, of the user 180 relative to the reference 188. For example, the distance 212 indicates the distance between the user location 220 (e.g., an estimated location of the user 180) and the reference location 230 (e.g., an estimated location of the reference 188). The relative orientation 263 indicates the user orientation 222 (e.g., an estimated orientation of the user 180) relative to the reference orientation 232 (e.g., an estimated orientation of the reference 188).

[0048]

[0060] Audio player 140 performs relative location estimation 160 based on location data 270. For example, audio player 140 initializes relative location 161 corresponding to a fiducial 188 (e.g., a mobile phone screen) that is a predetermined distance (e.g., 12 inches) from user 180 and is oriented in front of (e.g., facing) user 180. Audio player 140 updates relative location 161 based on updates to location data 270.

[0049]

[0061] Examples 202A-202C illustrate top views of a horizontal plane in three-dimensional space. In a particular aspect, the horizontal plane is defined by the X-axis and the Y-axis in the three-dimensional space, and the vertical plane is defined by the X-axis and the Z-axis in the three-dimensional space. Example 202A corresponds to audio player 140 initializing relative orientation 263 and distance 212 to relative orientation 263A and distance 212A, respectively. Examples 202B-202C correspond to audio player 140 updating relative orientation 263 and distance 212 based on updates to location data 270.

[0050]

[0062] In example 202A, the audio player 140 initializes the reference location 230 of the reference 188 to reference location 230A, the reference orientation 232 of the reference 188 to reference orientation 232A, the user location 220 of the user 180 to user location 220A, and the user orientation 222 of the user 180 to user orientation 222A.

[0051]

[0063] In some implementations, spatial data 179 (e.g., GPS data, configuration data, image data, etc.) indicates that fiducial 188 is detected at a reference location 230A having a reference orientation 232A in three-dimensional space, and audio player 140 initializes reference location 230 and reference orientation 232 to reference location 230A and reference orientation 232A, respectively. In an alternative implementation, audio player 140 initializes reference location 230 to reference location 230A corresponding to the origin of three-dimensional space (e.g., 0 inches along the X-axis, 0 inches along the Y-axis, and 0 inches along the Z-axis), and reference orientation 232A corresponds to fiducial 188 facing a predetermined direction in three-dimensional space (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane).

[0052]

[0064] In some implementations, spatial data 177 (e.g., GPS data, image data, etc.) indicates that user 180 is detected in three-dimensional space at user location 220A with user orientation 222A, and audio player 140 initializes user location 220 and user orientation 222 to user location 220A and user orientation 222A, respectively. In an alternative implementation, audio player 140 initializes user location 220 to user location 220A that corresponds to a predetermined point (e.g., a predetermined distance and a predetermined direction) from reference location 230A in three-dimensional space. For example, user location 220A corresponds to a point (e.g., 12 inches along the X-axis, 0 inches along the Y-axis, and 0 inches along the Z-axis) that is a distance 212A (e.g., a predetermined distance such as 12 inches) from an origin in three-dimensional space in a relative direction 265A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane). Audio player 140 initializes user orientation 222A to a predetermined direction in three-dimensional space (e.g., 180 degrees in the horizontal plane (e.g., XY plane) and 0 degrees in the vertical plane (e.g., XZ plane)) to correspond to user 180 facing reference 188. Audio player 140 therefore initializes distance 212 to distance 212A and relative orientation 263 to relative orientation 263A at time T0. Relative orientation 263A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane) is based on reference orientation 232A of user location 220A relative to reference location 230A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane), user orientation 222A (e.g., 180 degrees in the horizontal plane and 0 degrees in the vertical plane), and relative direction 265A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane).

[0053]

[0065] The audio player 140 updates the relative location 161 based on the location data 270. For example, the location data 270 may indicate a change in location of the user 180, a change in orientation of the user 180, a change in location of the reference 188, a change in orientation of the reference 188, or a combination thereof, and the audio player 140 updates the relative location 161 based on the changes indicated by the location data 270. In certain aspects, the location data 270 includes spatial data 177 of the user 180, spatial data 179 of the reference 188, a user gaze estimate 275, or a combination thereof.

[0054]

[0066] The electrical activity data 105 can indicate the direction of the user's gaze (e.g., user gaze estimate 275). In certain aspects, the direction of the user's gaze is relative to the orientation of the head 182. For example, if the user 180 is looking at a particular gaze target and continues to look at the same gaze target while changing the orientation of the head 182 by a certain amount (e.g., 10 degrees), the direction of the user's gaze changes by a certain amount (e.g., 10 degrees). In certain aspects, the audio player 140 estimates the change in orientation of the user 180 based on the user gaze estimate 275. For example, the audio player 140 determines the change in orientation of the head 182 of the user 180 based on the spatial data 177. In some implementations, the change in orientation of the head 182 corresponds to a broad estimate of the change in the user orientation 222, and the audio player 140 refines the estimate of the change in the user orientation 222 based on the user gaze estimate 275. For example, if the user 180 moves their head 182 but keeps their gaze directed in the same location, there may be no change in the user orientation 222.

[0055]

[0067] In certain aspects, the audio player 140 processes the electrical activity data 105 (e.g., EOG data) using a gaze estimation model 274 (e.g., an artificial neural network, a machine learning model, or both) to determine a user gaze estimate 275. Determining the user orientation 222 (e.g., head orientation) based on the spatial data 177 and updating the user orientation 222 based on the user gaze estimate 275 (e.g., user gaze direction) are provided as illustrative, non-limiting examples. In some examples, the audio player 140 can process the spatial data 177 and the electrical activity data 105 (indicative of the direction of the user's gaze) to determine the user orientation 222.

[0056]

[0068] According to some studies, EOG data (e.g., electrical activity data 105) is modeled to detect saccades (e.g., rapid eye movements between fixation states), and EOG signal fluctuations indicative of a saccade reflect the direction of the gaze shift. For example, an increase in the EOG signal fluctuation indicates a rightward gaze shift, and a decrease indicates a leftward gaze shift. The amplitude of the EOG signal indicates the angle of the gaze shift. For example, a larger absolute value of the amplitude indicates a larger gaze shift. In certain aspects, the user gaze estimate 275 is determined in response to detecting a gaze. For example, a gaze is detected based on determining that the direction of the user's gaze is unchanged for at least a threshold duration.

[0057]

[0069] In some implementations, instead of the audio player 140 estimating the user location 220, the user orientation 222, or both based on the changes, the spatial data 177 (e.g., GPS data, image data, or both) can directly indicate the user location 220, the user orientation 222, or both. In some implementations, instead of the audio player 140 estimating the reference location 230, the reference orientation 232, or both based on the changes, the spatial data 179 (e.g., GPS data, image data, or both) can directly indicate the reference location 230, the reference orientation 232, or both.

[0058]

[0070] In some implementations, the fiducial 188 has a fixed location (e.g., reference location 230A). In some aspects, the spatial data 179 indicates changes (if any) to the reference orientation 232, the fixed location (e.g., reference location 230A), or both. In alternative aspects, the fiducial 188 has a fixed orientation (e.g., reference orientation 232A). In some examples, the spatial data 179 indicates a fixed location (e.g., reference location 230A), a fixed orientation (e.g., reference orientation 232A), or both. In some examples, the spatial data 179 indicates no change to the reference location 230, no change to the reference orientation 232, or both. In yet other examples, the location data 270 may not include spatial data 179. For example, the audio player 140 estimates (e.g., updates) the user location 220 and user orientation 222 based on the spatial data 177, the user gaze estimate 275, or both, and performs the relative location estimation 160 based on the fixed location and fixed orientation of the reference 188 and the estimated location and estimated orientation of the user 180.

[0059]

[0071] Examples 202B and 202C illustrate examples of the same relative orientation 263 corresponding to different user orientations 222, different reference orientations 232, and different reference locations 230. In some examples, the same relative orientation 263 may correspond to different user orientations 222, different reference orientations 232, different user locations 220, different reference locations 230, or combinations thereof. In example 202B, audio player 140 acquires location data 270 at time T1, which is later than time T0. Audio player 140 determines, based on spatial data 177, that user 180 has user orientation 222B (e.g., 135 degrees in the horizontal plane and 0 degrees in the vertical plane) and is at user location 220B. Audio player 140 determines, based on spatial data 179, that reference 188 is at reference location 230B and has reference orientation 232A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane). The audio player 140 determines the distance 212B based on the difference between the user location 220B and the reference location 230B.

[0060]

[0072] The audio player 140 determines that the user location 220B (e.g., a first point in 3D space) has a relative direction 265A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane) from the reference location 230B (e.g., a second point in 3D space). For example, the relative direction 265 is the same in example 202B as in example 202A.

[0061]

[0073] The relative orientation 265 (e.g., the orientation of a first point relative to a second point in 3D space) is based on the orientation of the user location 220 (e.g., the first point) relative to the reference location 230 (e.g., the second point) and is independent of the user orientation 222 and the reference orientation 232. In comparison, the relative orientation 263 is based on the user orientation 222 and the reference orientation 232 in addition to the relative orientation 265 of the user location 220 relative to the reference location 230. For example, the relative orientation 263 indicates the orientation of at least a first plane (e.g., including the first point) corresponding to the user 180 relative to at least a second plane (e.g., including the second point) corresponding to the reference 188. In an illustrative, non-limiting example, the first plane corresponds to a vertical cross section of the head 182 of the user 180 and the second plane corresponds to a display screen of the reference 188 (e.g., a mobile device).

[0062]

[0074] The audio player 140 determines a relative orientation 263B (e.g., 45 degrees in the horizontal plane and 0 degrees in the vertical plane) based on the user orientation 222B (e.g., 135 degrees in the horizontal plane and 0 degrees in the vertical plane) of the user location 220B relative to the reference location 230B, the reference orientation 232A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane), and the relative direction 265B (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane). In some examples, the relative orientation 263 may be different for different relative directions 265 having the same user orientation 222 and the same reference orientation 232, as will be further described with reference to FIG. 2B.

[0063]

[0075] In example 202C, audio player 140 obtains location data 270 at time T2, which is later than time T0. Audio player 140 determines, based on spatial data 177, that user 180 has a user orientation 222A (e.g., 180 degrees in the horizontal plane and 0 degrees in the vertical plane) and is at user location 220B. Audio player 140 determines, based on spatial data 179, that reference 188 is at reference location 230C and has a reference orientation 232C (e.g., 45 degrees in the horizontal plane and 0 degrees in the vertical plane). Audio player 140 determines distance 212B based on the difference between user location 220B and reference location 230C.

[0064]

[0076] The audio player 140 determines a relative orientation 265C based on a comparison between the user location 220B (e.g., a first point in 3D space) and the reference location 230C (e.g., a third point in 3D space). For example, the user location 220B has a relative orientation 265C (e.g., 45 degrees in the horizontal plane and 0 degrees in the vertical plane) from the reference location 230C. The audio player 140 determines a relative orientation 263B (e.g., 45 degrees in the horizontal plane and 0 degrees in the vertical plane) based on the user orientation 222A (e.g., 180 degrees in the horizontal plane and 0 degrees in the vertical plane) of the user location 220B relative to the reference location 230C, the reference orientation 232C (e.g., 45 degrees in the horizontal plane and 0 degrees in the vertical plane), and the relative orientation 265C (e.g., 45 degrees in the horizontal plane and 0 degrees in the vertical plane). The relative orientation 263 is the same (e.g., relative orientation 263B) in example 202C as in example 202B for a different reference location 230, a different reference orientation 232, a different user orientation 222, a different relative direction 265, the same distance 212, and the same user location 220. For example, in example 202C compared to example 202B, at least a first plane corresponding to user 180 (e.g., a vertical cross section of head 182) has the same orientation relative to at least a second plane corresponding to reference 188 (e.g., a display screen).

[0065]

[0077] 2B, a diagram 250 of an example aspect of the operation associated with relative location estimation 160 is shown. Relative location estimation 160 may be implemented by audio player 140 of FIG 1. Example 202D illustrates an example of the same user orientation 222 and the same reference orientation 232 corresponding to different relative orientations 263 due to different relative directions 265 of user location 220 with respect to reference point location 230.

[0066]

[0078] Example 202D corresponds to audio player 140 updating relative orientation 263 and distance 212 based on updates to location data 270. In example 202D, audio player 140 obtains location data 270 at time T3, which is later than time T0. Audio player 140 determines, based on spatial data 177, that user 180 has a user orientation 222A (e.g., 180 degrees in the horizontal plane and 0 degrees in the vertical plane) and is at a user location 220D. Audio player 140 determines, based on spatial data 179, that reference 188 has a reference orientation 232A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane) and is at a reference location 230D. Audio player 140 determines distance 212A based on the difference between user location 220D and reference location 230D.

[0067]

[0079] The audio player 140 determines a relative orientation 265D based on a comparison between the user location 220D (e.g., a third point in 3D space) and the reference location 230D (e.g., a fourth point in 3D space). For example, the user location 220D has a relative orientation 265D (e.g., 39 degrees in the horizontal plane and 0 degrees in the vertical plane) from the reference location 230D. The audio player 140 determines a relative orientation 263D (e.g., 39 degrees in the horizontal plane and 0 degrees in the vertical plane) based on the user orientation 222A (e.g., 180 degrees in the horizontal plane and 0 degrees in the vertical plane) of the user location 220D relative to the reference point location 230D, the reference orientation 232A (e.g., 0 degrees in the horizontal plane and 0 degrees in the vertical plane), and the relative orientation 265D (e.g., 39 degrees in the horizontal plane and 0 degrees in the vertical plane).

[0068]

[0080] The relative orientation 263D of example 202D is different from the relative orientation 263A of example 202A for the same user orientation 222 (e.g., user orientation 222A), the same reference orientation 232 (e.g., reference orientation 232A), a different relative direction 265, a different user location 220, and a different reference location 230. For example, in example 202D compared to example 202A, a first plane (e.g., a vertical cross section of head 182) corresponding to user 180 has the same user orientation 222, and a second plane (e.g., a display screen) corresponding to reference 188 has the same reference orientation 232. The first plane has a different relative orientation 263 with respect to the second plane in example 202D compared to example 202A because a relative orientation 265D of the third point with respect to the fourth point is different from a relative orientation 265A of the first point with respect to the second point.

[0069]

[0081] 3, a diagram of a system operable to adjust audio based on a user electrical signal is shown and generally designated 300. In certain aspects, the system 100 of FIG.

[0070]

[0082] 3, an example implementation of multi-location audio generation 164 and preferred location estimation 166 is illustrated. For example, multi-location audio generation 164 includes multi-source audio generation 364. To illustrate, audio player 140 generates audio data 141B by rendering audio data 141A to have a location 350 of a speech source 386 and a location 352 of a non-speech source 388 in sound field 142.

[0071]

[0083] The audio player 140 initiates an audio playback operation 144A of the audio data 141B through one or more speakers 106. The one or more sensors 104 generate electrical activity data 105 during the audio playback operation 144A. For example, the electrical activity data 105 is based on electrical signals from one or more electrical sources 184 during the audio playback operation 144A.

[0072]

[0084] In some examples, the preferred place location estimate 166 includes a preferred source estimate (preference source estimate) 366. For example, the audio player 140 identifies one of a speech source 386 or a non-speech source 388 as the user preferred source 367 based on the electrical activity data 105. The audio player 140 processes the electrical activity data 105 using a preferred source model 374 (e.g., an artificial neural network, a machine learning model, or both) to generate a count of tracked sound sources.

[0073]

[0085] A "tracked sound source" corresponds to a sound source that the auditory system of the user 180 focuses on (e.g., attends to) as the sound source moves in the sound field 142 during audio playback operation 144A. According to some studies, a linear mapping (e.g., a temporal response function (TRF)) can be derived between the EEG data and the attended and unattended sound source trajectories. From both the delta phase and alpha power of the EEG, the trajectory (e.g., path) of the attended sound source can be reliably reconstructed, even in the presence of distracting stimuli. Tracking of unattended non-speech sources (e.g., noise) is below the detection level, and unattended speech is weakly tracked (e.g., by the delta phase of the EEG).

[0074]

[0086] During audio playback operation 144A, if user 180 directs attention to speech source 386, electrical activity data 105 (e.g., delta phase and alpha power of EEG data) tracks speech source 386 and tracks non-speech source 388 below detection levels. During audio playback operation 144A, if user 180 directs attention to non-speech source 388, electrical activity data 105 tracks non-speech source 388 and weakly tracks speech source 386. For example, EEG delta phase and alpha power tracks non-speech source 388 and EEG delta phase weakly tracks speech source 386.

[0075]

[0087] In some implementations, the preferred source model 374 is trained to determine a count of tracked sound sources indicated by the electrical activity data. For example, the training electrical activity data is generated by playing audio corresponding to sound sources (e.g., one or more speech sources, one or more non-speech sources, or a combination thereof) that move (e.g., change location) in a sound field using one or more speakers 106, requesting the user 180 to focus (e.g., pay attention to) the particular sound sources during playback, collecting training electrical activity data from one or more sensors 104, and tagging the training electrical activity data with the count of the attended sound sources. The preferred source model 374 is used to process the training electrical activity data to generate an estimated count of tracked sound sources, a loss metric based on a comparison between the estimated count and the tagged count, and configuration settings (e.g., weights, biases, or a combination thereof) of the preferred source model 374 are adjusted based on the loss metric.

[0076]

[0088] In response to determining that the count of tracked sound sources has a first value (e.g., 1) indicating that a single sound source is tracked, the audio player 140 determines that the speech source 386 corresponds to the user preferred source 367 and that the location 350 corresponds to the user preferred location 167. Alternatively, in response to determining that the count of tracked sound sources has a second value (e.g., greater than 1) indicating that multiple sound sources (e.g., the speech source 386 and the non-speech source 388) are tracked by the user 180, the audio player 140 determines that the non-speech source 388 corresponds to the user preferred source 367 and that the location 352 corresponds to the user preferred location 167. By way of example, the human brain tracks speech to some extent when a listener (e.g., the user 180) is attending to non-speech audio and does not track non-speech when the listener is attending to speech audio. The audio player 140 performs single location audio generation 168 based on a user preferred location 167 (eg, one of location 350 or location 352), as described with reference to FIG.

[0077]

[0089] FIG. 4 illustrates an implementation 400 of the device 102 as an integrated circuit 402 that includes one or more processors 190. The integrated circuit 402 also includes an input interface 114, such as one or more bus interfaces, to allow electrical activity data 105 to be received for processing. The integrated circuit 402 also includes an output interface 124, such as a bus interface, to allow transmission of audio data 141. The integrated circuit 402 allows for adjusting the audio based on a user electrical signal. In some examples, the integrated circuit 402 corresponds to a component in a system coupled to one or more sensors 104, one or more speakers 106, or a combination thereof, such as a mobile phone or tablet as shown in FIG. 5, a headset as shown in FIG. 6, a wearable electronic device as shown in FIG. 7, a voice-controlled speaker system as shown in FIG. 8, a virtual reality, mixed reality, or augmented reality headset as shown in FIG. 9, a vehicle as shown in FIG. 10, or one or more earphones as shown in FIG. 11.

[0078]

[0090] 5 illustrates an implementation 500 in which the device 102 includes a mobile device 502, such as a phone or tablet, as an illustrative, non-limiting example. The mobile device 502 includes one or more speakers 106, a display screen 504, or a combination thereof. Components of one or more processors 190, including an audio player 140, are integrated within the mobile device 502 and are shown using dashed lines to indicate internal components that are not typically visible to a user of the mobile device 502.

[0079]

[0091] The mobile device 502 is coupled to one or more sensors 104. In some implementations, the mobile device 502 corresponds to the reference 188 and includes one or more motion sensors 178. The one or more motion sensors 178 are shown using dashed lines to indicate internal components that are not normally visible to a user of the mobile device 502. In some implementations, the one or more sensors 104, the one or more speakers 106, the one or more spatial sensors 176, or a combination thereof are integrated into a user head mounted device (e.g., a headset or earphones) and the audio player 140 is integrated into the mobile device 502. In some implementations, the one or more spatial sensors 176 (e.g., a camera) are integrated into the mobile device 502.

[0080]

[0092] In particular examples, the audio player 140 operates to adjust the audio based on the user electrical signal, which may also be processed to perform one or more actions on the mobile device 502, such as launching a graphical user interface or otherwise displaying information associated with adjusting the audio or information associated with speech detected in the audio on the display screen 504 (e.g., an integrated “smart assistant” application).

[0081]

[0093] 6 illustrates an implementation 600 in which the device 102 includes a headset device 602. The headset device 602 includes or is coupled to one or more sensors 104, one or more speakers 106, one or more spatial sensors 176, one or more spatial sensors 178, or a combination thereof. One or more components of the processor 190, including an audio player 140, are integrated within the headset device 602. In a particular example, the audio player 140 operates to adjust the audio based on a user electrical signal, thereby causing the headset device 602 to perform one or more operations at the headset device 602 and transmit the adjusted audio data to a second device (not shown) for further processing, or a combination thereof.

[0082]

[0094] In some implementations, the headset device 602 includes one or more sensors 104, one or more speakers 106, one or more spatial sensors 176, or a combination thereof, and is coupled to a second device that includes an audio player 140. One or more spatial sensors 178 can be included in the headset device 602, the second device, or both. In some aspects, the second device includes a vehicle, a mobile device, a phone, a game console, a communication device, a wearable electronic device, a voice-controlled speaker system, an unmanned vehicle, or a combination thereof.

[0083]

[0095] 7 illustrates an implementation 700 in which the device 102 includes a wearable electronic device 702 shown as a “smart watch.” The audio player 140 and one or more speakers 106 are integrated into or coupled to the wearable electronic device 702.

[0084]

[0096] The wearable electronic device 702 is coupled to one or more sensors 104. In some implementations, the wearable electronic device 702 corresponds to the reference 188 and includes one or more motion sensors 178. The one or more motion sensors 178 are shown using dashed lines to indicate internal components that are not normally visible to a user of the wearable electronic device 702. In some implementations, the one or more sensors 104, the one or more speakers 106, the one or more spatial sensors 176, or a combination thereof are integrated into a user head-mounted device (e.g., a headset or earphones) and the audio player 140 is integrated into the wearable electronic device 702. In some implementations, the one or more spatial sensors 176 (e.g., a camera) are integrated into the wearable electronic device 702.

[0085]

[0097] In a particular example, the audio player 140 operates to adjust the audio based on the user electrical signal, which is then processed to perform one or more operations at the wearable electronic device 702, such as launching a graphical user interface or otherwise displaying information associated with adjusting the audio or information associated with speech detected in the audio on a display screen 704 of the wearable electronic device 702. To illustrate, the wearable electronic device 702 may include a display screen configured to display notifications during an audio playback operation 144A by the wearable electronic device 702. In a particular example, the wearable electronic device 702 includes a haptic device that provides haptic notifications (e.g., vibrates) during an audio playback operation 144A. For example, the haptic notification may cause the user to look at the wearable electronic device 702 to see a displayed notification indicating an audio composition in progress. Thus, the wearable electronic device 702 may alert a user who is hearing impaired or who is wearing a headset that an audio composition is being performed.

[0086]

[0098] 8 is an implementation 800 in which the device 102 includes a wireless speaker and voice-activated device 802. The wireless speaker and voice-activated device 802 can have wireless network connectivity and is configured to perform assistant operations. One or more processors 190, including an audio player 140, one or more speakers 106, or a combination thereof, are included in the wireless speaker and voice-activated device 802.

[0087]

[0099] The wireless speaker and voice-activated device 802 is coupled to one or more sensors 104. In some implementations, the wireless speaker and voice-activated device 802 corresponds to the criteria 188 and includes one or more motion sensors 178. The one or more motion sensors 178 are shown using dashed lines to show internal components that are not normally visible to a user of the wireless speaker and voice-activated device 802. In some implementations, the one or more sensors 104, the one or more speakers 106, the one or more spatial sensors 176, or a combination thereof are integrated into a user head-mounted device (e.g., a headset or earphones) and the audio player 140 is integrated into the wireless speaker and voice-activated device 802. In some implementations, the one or more spatial sensors 176 (e.g., a camera) are integrated into the wireless speaker and voice-activated device 802.

[0088]

[0100] In operation, in response to receiving a verbal command identified as a user utterance via operation of the audio player 140, the wireless speaker and voice-activated device 802 can execute an assistant action (e.g., an integrated assistant application). The assistant action can include adjusting the temperature, playing music, turning on the lights, etc. For example, an assistant action is performed in response to receiving a command followed by a keyword or key phrase (e.g., "Hello, assistant").

[0089]

[0101] 9 illustrates an implementation 900 in which the device 102 includes a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset 902. The audio player 140, one or more sensors 104, one or more speakers 106, one or more spatial sensors 176, one or more spatial sensors 178, or combinations thereof, are included in or coupled to the headset 902.

[0090]

[0102] Audio adjustments based on the user electrical signals can be performed, and the adjusted audio signals can be output via one or more speakers 106 of the headset 902. The visual interface device is positioned in front of the user's eyes to enable augmented, mixed, or virtual reality images or scenes to be displayed to the user while the headset 902 is worn. In some implementations, the fiducials 188 correspond to virtual fiducials that can be displayed by the visual interface device. In certain examples, the visual interface device is configured to display a notification indicating that audio adjustment is in progress.

[0091]

[0103] 10 illustrates an implementation 1000 in which a device 102 corresponds to or is integrated within a vehicle 1002, shown as a manned or unmanned aerial device (e.g., a delivery drone). An audio player 140, one or more speakers 106, or a combination thereof, are integrated into the vehicle 1002.

[0092]

[0104] The vehicle 1002 is coupled to one or more sensors 104. In some implementations, the vehicle 1002 includes one or more motion sensors 178. In some implementations, the vehicle 1002 includes a fiducial 188. The one or more motion sensors 178 are shown using dashed lines to indicate internal components that are not typically visible to a user of the vehicle 1002. In some implementations, the one or more sensors 104, the one or more speakers 106, the one or more spatial sensors 176, or a combination thereof are integrated into a user head-mounted device (e.g., a headset or earphones) and the audio player 140 is integrated into the vehicle 1002. In some implementations, the one or more spatial sensors 176 (e.g., a camera) are integrated into the vehicle 1002. Audio conditioning based on the user electrical signal can be performed and the conditioned audio signal can be output via one or more speakers 106 of the vehicle 1002.

[0093]

[0105] 11 is a diagram of an earphone 1100 (e.g., another particular example of device 102 of FIG. 1) operable to perform audio adjustments based on a user electrical signal. In FIG. 11, a first earphone 1102 includes at least one of the one or more spatial sensors 176, and a second earphone 1104 includes at least one of the one or more spatial sensors 176. Each of the first earphone 1102 and the second earphone 1004 also includes at least one of the one or more speakers 106. One or both of the earphones 1200 may also include an audio player 140, one or more spatial sensors 178, or a combination thereof.

[0094]

[0106] 12, a particular implementation of a method 1200 for adjusting audio based on a user electrical signal is shown. In a particular aspect, one or more operations of the method 1200 are performed by at least one of the audio player 140, the audio adjustment unit 170, one or more processors 190, the device 102, the system 100 of FIG. 1, or a combination thereof.

[0095]

[0107] Method 1200 includes acquiring 1202 electrical activity data corresponding to electrical signals from one or more electrical sources in a user's head. For example, audio player 140 of FIG. 1 acquires electrical activity data 105 corresponding to electrical signals from one or more electrical sources 184 in head 182 of user 180.

[0096]

[0108] Method 1200 also includes, at 1204, rendering audio data based on the electrical activity data to adjust a location of a sound source in a sound field during playback of the audio data. For example, audio player 140 renders audio data 141 based on electrical activity data 105 to adjust a location of sound source 186 in sound field 142 during playback of audio data 141. To illustrate, audio player 140 renders audio data 141A to adjust a location of sound source 186 in sound field 142 from location 150A to location 150B to generate audio data 141C.

[0097]

[0109] System 1200 thus enables rendering audio with sound that can be perceived as coming from a direction of sound source 186 (e.g., location 150B) that corresponds to user preferred location 167 or an estimate of user preferred location 167 (e.g., relative location 161). As an example, if user 180 places a playback device (e.g., fiducial 188) on a desk, the location of sound source 186 can be adjusted to continue to be perceived as coming from in front of user 180 (e.g., user preferred location 167).

[0098]

[0110] The method 1200 of Figure 12 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a digital signal processor (DSP), a controller, another hardware device, a firmware device, or any combination thereof. As an example, the method 1200 of Figure 12 may be performed by a processor executing instructions such as those described with reference to Figure 13.

[0099]

[0111] 13, a block diagram of a particular example implementation of a device is shown, generally designated 1300. In various implementations, device 1300 may have more or fewer components than shown in FIG. 13. In an example implementation, device 1300 may correspond to device 102. In an example implementation, device 1300 may perform one or more of the operations described with reference to FIGS. 1-12.

[0100]

[0112] In particular implementations, device 1300 includes a processor 1306 (e.g., a CPU). Device 1300 may include one or more additional processors 1310 (e.g., one or more DSPs). In particular aspects, one or more of the processors 190 of FIG. 1 correspond to processor 1306, processor 1310, or a combination thereof. Processor 1310 may include a speech and music coder-decoder (codec) 1308, including a voice coder ("vocoder") encoder 1336, a vocoder decoder 1338, an audio player 140, or a combination thereof.

[0101]

[0113] The device 1300 may include a memory 1386 and a codec 1334. The memory 1386 may include instructions 1356 executable by one or more additional processors 1310 (or processor 1306) to implement functions described with reference to the audio player 140. The device 1300 may include a modem 1348 coupled to an antenna 1352 via a transceiver 1350. The device 1300 may include or be coupled to one or more spatial sensors 176, one or more spatial sensors 178, one or more sensors 104, or a combination thereof.

[0102]

[0114] The device 1300 may include a display 1328 coupled to a display controller 1326. The one or more speakers 106 and one or more microphones 1390 may be coupled to a codec 1334. The codec 1334 may include a digital-to-analog converter (DAC) 1302, an analog-to-digital converter (ADC) 1304, or both. In certain implementations, the codec 1334 may receive analog signals from the one or more microphones 1390, convert the analog signals to digital signals (e.g., audio data 141A) using the analog-to-digital converter 1304, and provide the digital signals to a speech and music codec 1308. The speech and music codec 1308 may process the digital signals, which may be further processed by the audio player 140. In certain implementations, the speech and music codec 1308 may provide a digital signal (e.g., audio data 141C) to the codec 1334. The codec 1334 may convert the digital signal to an analog signal using the digital-to-analog converter 1302 and provide the analog signal to one or more speakers 106. For example, in response to determining that the user 180 is tracking a sound source 186 at a location 150A that is far from or behind the user 180, the audio player 140 may generate audio data 141C to adjust the location of the sound source 186 to a location 150B closer to or in front of the user 180. The analog signal corresponding to the audio data 141C may be played through one or more speakers 106 integrated within a headset or earphones of the user 180.

[0103]

[0115] In certain implementations, the device 1300 may be included in a system-in-package or system-on-chip device 1322. In certain implementations, the memory 1386, the processor 1306, the processor 1310, the display controller 1326, the codec 1334, and the modem 1348 are included in the system-in-package or system-on-chip device 1322. In certain implementations, the input device 1330 and the power supply 1344 are coupled to the system-in-package or system-on-chip device 1322. Additionally, in certain implementations, the display 1328, the input device 1330, the one or more speakers 106, the one or more microphones 1390, the antenna 1352, and the power supply 1344 are external to the system-in-package or system-on-chip device 1322 as shown in FIG. In particular implementations, each of the display 1328, the input device 1330, the one or more speakers 106, the one or more microphones 1390, the antenna 1352, and the power source 1344 may be coupled to a component of the system-in-package or system-on-chip device 1322, such as an interface (e.g., the input interface 114 or the output interface 124) or a controller.

[0104]

[0116] Device 1300 may include a smart speaker, a speaker bar, a mobile communications device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, an automobile, a computing device, a communications device, an internet-of-thing (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0105]

[0117] In accordance with the described implementation, the apparatus includes means for acquiring electrical activity data corresponding to electrical signals from one or more electrical sources in the user's head. For example, the means for acquiring can correspond to one or more sensors 104, input interface 114, audio player 140, one or more processors 190, device 102, system 100 of FIG. 1, system 300 of FIG. 3, processor 1306, processor 1310, device 1300, one or more other circuits or components configured to acquire electrical activity data, or any combination thereof.

[0106]

[0118] The apparatus also includes means for rendering the audio data based on the electrical activity data to adjust the location of the sound source in the sound field during playback of the audio data. For example, the means for rendering may correspond to audio player 140, one or more processors 190, device 102, system 100 of FIG. 1, system 300 of FIG. 3, processor 1306, processor 1310, device 1300, one or more other circuits or components configured to obtain electrical activity data, or any combination thereof.

[0107]

[0119] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 1386) includes instructions (e.g., instructions 1356) that, when executed by one or more processors (e.g., one or more processors 190, one or more processors 1310, or processor 1306), cause the one or more processors to obtain electrical activity data (e.g., electrical activity data 105) corresponding to electrical signals from one or more electrical sources (e.g., one or more electrical sources 184) in a user's head (e.g., head 182). The instructions, when executed by the one or more processors, cause the one or more processors to render audio data (e.g., audio data 141A) and adjust a location (e.g., location 150A) of a sound source (e.g., sound source 186) in a sound field (e.g., sound field 142) during playback of the audio data (e.g., audio data 141C) based on the electrical activity data.

[0108]

[0120] Certain aspects of the disclosure are described below in a set of interrelated clauses.

[0109]

[0121] A device comprising a memory configured to store instructions according to clause 1 and one or more processors, the one or more processors configured to execute the instructions to obtain electrical activity data corresponding to electrical signals from one or more electrical sources in a user's head, and to render audio data based on the electrical activity data to adjust a location of the sound source in a sound field during playback of the audio data.

[0110]

[0122] Clause 2 includes the device of clause 1, wherein the one or more processors are further configured to execute instructions to output audio data through one or more speakers.

[0111]

[0123] Clause 3 includes the device of clause 1 or 2, wherein the electrical activity data includes electrooculogram (EOG) data, electroencephalogram (EEG) data, or both.

[0112]

[0124] Clause 4 includes the device of any of clauses 1-3, further including an interface configured to receive electrical activity data from one or more sensors.

[0113]

[0125] Clause 5 includes the device of clause 4, wherein the interface includes at least one of an Ethernet interface, a Universal Serial Bus (USB) interface, a Wi-Fi interface, a Bluetooth interface, a serial port interface, or a parallel port interface.

[0114]

[0126] Clause 6 includes the device of clause 4 or 5, wherein the one or more sensors includes an in-ear sensor.

[0115]

[0127] Clause 7 includes the device of clause 4 or 5, wherein the one or more sensors include an electrode cap, a neural implant, a conductive screen, a non-wearable sensor, or a combination thereof.

[0116]

[0128] Clause 8 includes the device of any of clauses 1-7, wherein the one or more processors are further configured to execute instructions to initiate a first playback operation of audio data through one or more speakers, where the audio data is rendered to include a plurality of locations of a sound source in a sound field during the first playback operation, and the electrical activity data is based on electrical signals from the one or more electrical sources during the first playback operation of the audio data, and to determine a user-preferred location of the sound source based on the electrical activity data, and the audio data is rendered based on the user-preferred location to adjust the location during a second playback operation of the audio data.

[0117]

[0129] Clause 9 includes the device of clause 8, wherein one or more processors are configured to execute instructions to determine whether to render the audio data to include a plurality of locations based on determining whether a location confidence level is less than a confidence threshold, the location confidence level being associated with an estimated location of the user relative to a reference.

[0118]

[0130] Clause 10 includes the device of clause 9, wherein the estimated location of the user relative to the reference includes an estimated orientation of the user relative to an estimated orientation of the reference, an estimated distance of the user relative to the reference, or both.

[0119]

[0131] Clause 11 includes the device of clauses 9 or 10, wherein the reference includes one or more of a device, a display device, a physical reference, or a virtual reference.

[0120]

[0132] Clause 12 includes the device of any of clauses 9-11, wherein the one or more processors are configured to execute instructions to initialize the estimated location to correspond to a reference directed in front of the user.

[0121]

[0133] Clause 13 includes the device of any of clauses 9-12, wherein the one or more processors are configured to execute instructions to initialize the location trust level below the location threshold.

[0122]

[0134] Clause 14 includes the device of any of clauses 9 to 13, wherein the one or more processors are configured to execute instructions to update the user's estimated location relative to a reference, and the estimated location is updated based on inertial measurement unit (IMU) data of the user's headset, spatial data of the reference, a user gaze estimate, or a combination thereof.

[0123]

[0135] Clause 15 includes the device of any of clauses 1-14, wherein the one or more processors are configured to execute instructions to determine a user gaze estimate based on the electrical activity data and to update an estimated location of the user relative to a reference based at least in part on the user gaze estimate.

[0124]

[0136] Clause 16 includes the device of any of clauses 1-15, wherein one or more processors are configured to execute instructions to process the electrical activity data using a machine learning model to determine a user gaze estimate, and to update an estimated location of the user relative to a reference based on the user gaze estimate.

[0125]

[0137] Clause 17 includes the device of any of clauses 1-16, wherein the one or more processors are further configured to execute instructions to process the electrical activity data using a machine learning model to determine a user-preferred location of the sound source.

[0126]

[0138] Clause 18 includes the device of any of clauses 1-17, wherein the one or more processors are further configured to execute instructions to render the audio data to include a first location of a speech source and a second location of a non-speech source in the sound field; initiate a first playback operation of the audio data through one or more speakers, where the electrical activity data is based on electrical signals from the one or more electrical sources during the first playback operation; and determine a user-preferred location of the sound source based on the electrical activity data; and wherein the audio data is rendered based on the user-preferred location to adjust the location of the sound source during a second playback operation of the audio data.

[0127]

[0139] Clause 19 includes the device of clause 18, wherein the one or more processors are further configured to execute instructions to determine that the user preferred location corresponds to a first location of the speech source in response to determining that the electrical activity data indicates that a single sound source is being tracked.

[0128]

[0140] Clause 20 includes the device of clause 18 or 19, wherein the one or more processors are further configured to execute instructions to determine that the user preferred location corresponds to a second location of the non-speech source in response to determining that the electrical activity data indicates that a speech source and a non-speech source are being tracked.

[0129]

[0141] 2. A method according to clause 21, comprising: acquiring, in a device, electrical activity data corresponding to electrical signals from one or more electrical sources in the user's head; and rendering audio data based on the electrical activity data to adjust a location of the sound source in a sound field during playback of the audio data.

[0130]

[0142] Clause 22 includes the method of clause 21, further including outputting the audio data through one or more speakers.

[0131]

[0143] Clause 23 includes the method of clauses 21 or 22, wherein the electrical activity data includes electrooculogram (EOG) data, electroencephalogram (EEG) data, or both.

[0132]

[0144] Clause 24 includes the method of any of clauses 21-23, further including receiving electrical activity data from the one or more sensors via the interface.

[0133]

[0145] Clause 25 includes the method of clause 24, wherein the interface includes at least one of an Ethernet interface, a Universal Serial Bus (USB) interface, a Wi-Fi interface, a Bluetooth interface, a serial port interface, or a parallel port interface.

[0134]

[0146] Clause 26 includes the method of clause 24 or 25, wherein the one or more sensors include an in-ear sensor.

[0135]

[0147] Clause 27 includes the method of clause 24 or 25, wherein the one or more sensors include an electrode cap, a neural implant, a conductive screen, a non-wearable sensor, or a combination thereof.

[0136]

[0148] Clause 28 includes the method of any of clauses 21-27, further comprising initiating a first playback operation of audio data through one or more speakers, where the audio data is rendered to include a plurality of locations of a sound source in a sound field during the first playback operation, and the electrical activity data is based on electrical signals from one or more electrical sources during the first playback operation of the audio data, and determining a user preferred location of the sound source based on the electrical activity data, wherein the audio data is rendered based on the user preferred location to adjust the location during a second playback operation of the audio data.

[0137]

[0149] Clause 29 includes the method of clause 28, further including determining whether to render the audio data to include multiple locations based on determining whether a location confidence level associated with the user's estimated location relative to the reference is less than a confidence threshold.

[0138]

[0150] Clause 30 includes the method of clause 29, wherein the estimated location of the user relative to the reference includes an estimated orientation of the user relative to an estimated orientation of the reference, an estimated distance of the user relative to the reference, or both.

[0139]

[0151] Clause 31 includes the method of clause 29 or 30, wherein the reference includes one or more of a device, a display device, a physical reference, or a virtual reference.

[0140]

[0152] Clause 32 includes the method of any of clauses 29-31, further including initializing the estimated location to correspond to a reference oriented in front of the user.

[0141]

[0153] Clause 33 includes the method of any of clauses 29-32, further including initializing the location confidence level below the location threshold.

[0142]

[0154] Clause 34 includes the method of any of clauses 29-33 further comprising updating the estimated location of the user relative to the reference, wherein the estimated location is updated based on inertial measurement unit (IMU) data of the user's headset, spatial data of the reference, a user gaze estimate, or a combination thereof.

[0143]

[0155] Clause 35 includes the method of any of clauses 21-34, further including determining a user gaze estimate based on the electrical activity data; and updating an estimated location of the user relative to a reference based on the user gaze estimate.

[0144]

[0156] Clause 36 includes the method of any of clauses 21-35, further including processing the electrical activity data using a machine learning model to determine a user gaze estimate, and updating an estimated location of the user relative to a reference based on the user gaze estimate.

[0145]

[0157] Clause 37 includes the method of any of clauses 21-36, further including processing the electrical activity data using a machine learning model to determine user-preferred locations of the sound sources.

[0146]

[0158] Clause 38 includes any of the methods of clauses 21-37 further including: rendering the audio data to include a first location of a speech source and a second location of a non-speech source in the sound field; initiating a first playback operation of the audio data through one or more speakers, where the electrical activity data is based on electrical signals from the one or more electrical sources during the first playback operation; and determining a user preferred location of the sound source based on the electrical activity data, wherein the audio data is rendered based on the user preferred location to adjust the location of the sound source during a second playback operation of the audio data.

[0147]

[0159] Clause 39 includes the method of clause 38, further including, in response to determining that the electrical activity data indicates that a single sound source is being tracked, determining that the user preferred location corresponds to a first location of the speech source.

[0148]

[0160] Clause 40 includes the method of clause 38 or 39, further including, in response to determining that the electrical activity data indicates that a speech source and a non-speech source are being tracked, determining that the user preferred location corresponds to a second location of the non-speech source.

[0149]

[0161] A device comprising: a memory configured to store instructions according to clause 41; and a processor configured to execute instructions to implement the method of any of clauses 21 to 40.

[0150]

[0162] A non-transitory computer readable medium storing instructions according to clause 42, the instructions, when executed by a processor, causing the processor to perform the method of any of clauses 21 to 40.

[0151]

[0163] Apparatus according to clause 43, comprising means for carrying out the method of any of clauses 21 to 40.

[0152]

[0164] A non-transitory computer readable medium storing instructions according to clause 44 that, when executed by one or more processors, cause the one or more processors to obtain electrical activity data corresponding to electrical signals from one or more electrical sources in the user's head, and render audio data based on the electrical activity data to adjust locations of the sound sources in a sound field during playback of the audio data.

[0153]

[0165] Clause 45 includes the non-transitory computer-readable medium of clause 44, wherein the instructions, when executed by one or more processors, cause the one or more processors to update an estimated location of the user relative to a reference, the estimated location being updated based on inertial measurement unit (IMU) data of the user's headset, spatial data of the reference, a user gaze estimate, or a combination thereof.

[0154]

[0166] Clause 46 includes the non-transitory computer-readable medium of any of clauses 44-45, the instructions, when executed by one or more processors, causing the one or more processors to determine a user gaze estimate based on the electrical activity data and update an estimated location of the user relative to a reference based on the user gaze estimate.

[0155]

[0167] Clause 47 includes an apparatus including means for acquiring electrical activity data corresponding to electrical signals from one or more electrical sources within a user's head, and means for rendering audio data based on the electrical activity data to adjust a location of the sound sources in a sound field during playback of the audio data.

[0156]

[0168] Clause 48 includes the apparatus of clause 47, where at least one of the means for obtaining or the means for rendering is integrated into a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, a vehicle, a communications device, a display device, a television, a games console, a music player, a radio, a digital video player, a camera, a navigation device, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, an automobile, a computing device, an Internet of Things (IoT) device, a mobile device, or any combination thereof.

[0157]

[0169] Those skilled in the art will further appreciate that the various exemplary logical blocks, configurations, modules, circuits, and algorithmic steps described with respect to the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various exemplary components, blocks, configurations, modules, circuits, and steps have been described above generally with respect to their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0158]

[0170] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or user terminal In the alternative, the processor and the storage medium may reside as discrete components in a computing device or user terminal.

[0159]

[0171] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosed aspects. Various modifications of these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest possible scope consistent with the principles and novel features defined by the following claims.

Claims

1. A memory configured to store instructions; one or more processors; a device comprising: acquiring electrical activity data corresponding to electrical signals from one or more electrical sources within the user's head; Rendering audio data based on said electrical activity data to adjust a location of a sound source in a sound field during playback of the audio data. configured to execute the instructions; device.

2. The device of claim 1 , wherein the one or more processors are further configured to execute the instructions to output the audio data through one or more speakers.

3. 10. The device of claim 1, wherein the electrical activity data includes electrooculogram (EOG) data, electroencephalogram (EEG) data, or both, and / or further comprises an interface configured to receive the electrical activity data from one or more sensors.

4. The one or more processors: commencing a first playback operation of the audio data through one or more speakers, wherein the audio data is rendered to include a plurality of locations of the sound sources in the sound field during the first playback operation, and wherein the electrical activity data is based on the electrical signals from the one or more electrical sources during the first playback operation of the audio data. determining a user preferred location of the sound source based on the electrical activity data; and further configured to execute the instructions, wherein the audio data is rendered based on the user preferred location to adjust the location during a second playback operation of the audio data. The device of claim 1 .

5. 5. The device of claim 4, wherein the one or more processors are configured to execute the instructions to determine whether to render the audio data to include the multiple locations based on determining whether a location confidence level is below a confidence threshold, the location confidence level being associated with an estimated location of the user relative to a reference.

6. The device of claim 5 , wherein the estimated location of the user relative to the reference comprises an estimated orientation of the user relative to an estimated orientation of the reference, an estimated distance of the user relative to the reference, or both.

7. 6. The device of claim 5, wherein the reference includes one or more of the device, a display device, a physical reference, or a virtual reference, and / or the one or more processors are configured to execute the instructions to initialize the estimated location to correspond to the reference pointed in front of the user.

8. 6. The device of claim 5, wherein the one or more processors are configured to initialize the location confidence level below a location threshold, and / or the one or more processors are configured to execute the instructions to update the user's estimated location relative to the reference, wherein the estimated location is updated based on inertial measurement unit (IMU) data of the user's headset, spatial data of the reference, a user gaze estimate, or a combination thereof.

9. The one or more processors: determining a user gaze estimate based on the electrical activity data; updating an estimated location of the user relative to a reference based at least in part on the user gaze estimate; The device of claim 1 configured to execute the instructions.

10. The one or more processors: processing the electrical activity data using a machine learning model to determine a user gaze estimate; updating an estimated location of the user relative to a reference based on the user gaze estimate; The device of claim 1 configured to execute the instructions.

11. 13. The device of claim 1, wherein the one or more processors are further configured to execute the instructions to process the electrical activity data using a machine learning model to determine a user-preferred location of the sound source.

12. The one or more processors: Rendering the audio data to include a first location of a speech source and a second location of a non-speech source in the sound field; commencing a first playback operation of the audio data via one or more speakers, wherein the electrical activity data is based on the electrical signals from the one or more electrical sources during the first playback operation; determining a user preferred location of the sound source based on the electrical activity data; and further configured to execute the instructions, wherein the audio data is rendered based on the user preferred location to adjust the location of the sound source during a second playback operation of the audio data. The device of claim 1 .

13. 13. The device of claim 12, wherein the one or more processors are further configured to execute the instructions to determine that the user preferred location corresponds to the first location of the speech source in response to determining that the electrical activity data indicates that a single sound source is being tracked, and / or to determine that the user preferred location corresponds to the second location of the non-speech source in response to determining that the electrical activity data indicates that the speech source and the non-speech source are being tracked.

14. acquiring, at the device, electrical activity data corresponding to electrical signals from one or more electrical sources within the user's head; rendering the audio data based on the electrical activity data to adjust a location of a sound source in a sound field during playback of the audio data; A method comprising:

15. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: acquiring electrical activity data corresponding to electrical signals from one or more electrical sources within the user's head; rendering the audio data based on the electrical activity data to adjust a location of a sound source in a sound field during playback of the audio data; Non-transitory computer-readable medium.