Attention state evaluation based on sound
Patent Information
- Application Number
- CN202180057588.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-03
- Filing Date
- 2021-05-26
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2041-05-26
AI Technical Summary
此外,内容可能不以对特定用户有意义的方式进行呈现
[0004]在一些具体实施中,可基于用户的环境的特性(例如,真实世界物理环境、虚拟环境或每一者的组合)来选择听觉刺激。设备(例如,手持式设备、膝上型电脑、台式电脑或头戴式设备(HMD))向用户提供真实世界物理环境、扩展现实(XR)环境或每一者的组合(例如,混合现实环境)的体验(例如,视觉体验和/或听觉体验)。设备用传感器获得与用户对听觉刺激的响应相关联的生理数据(例如,脑电图(EEG)振幅、瞳孔调制、眼睛注视扫视等)。基于所获得的生理数据,本文描述的技术可在体验(例如,冥想体验)期间确定用户的注意力状态(例如,专心、走神等)。基于生理数据和相关联生理响应,该技术可向用户提供当前注意力状态与体验的预期注意力状态不同的反馈,推荐该体验的类似内容或类似部分,并且/或者调整对应于该体验的内容或反馈机制。
Smart Images

Figure CN116133594B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to presenting content via electronic devices, and more specifically, to systems, methods, and devices for determining a user’s attentional state during and / or based on the presentation of visual or auditory content. Background Technology
[0002] A user's attentional state while viewing and listening to content on electronic devices can significantly impact their experience. For example, meaningful experiences may require focus and engagement, such as meditation, learning new skills, watching educational or entertainment content, or reading documents. Technologies that improve the attentional state of users viewing and interacting with content can enhance their enjoyment, comprehension, and learning from the content. Furthermore, content may not be presented in a way that is meaningful to a particular user. Content creators and systems may be able to leverage attentional state information to provide better and more customized user experiences that users are more likely to enjoy, understand, and learn from. Summary of the Invention
[0003] The various embodiments disclosed herein include devices, systems, and methods for evaluating a user's attentional state based on physiological responses to auditory stimuli. Auditory stimuli can be selected to both elicit responses suitable for evaluating attentional state and possess additional appropriate properties. Stimuli can be selected to blend, for example, with the user's current environment, natural scene, and surrounding environment. The auditory characteristics, spatial location, and timing of the stimuli can be selected. For example, a specific auditory stimulus (e.g., bird chirping) can be selected for use in a natural soundscape during meditation based on the fact that a bird chirping sound is both an appropriate sound for eliciting responses suitable for evaluating attentional state and an expected sound consistent with or otherwise consistent with the natural soundscape environment. Furthermore, attributes of the auditory stimulus, such as the volume, spatial location, and timing of the stimulus, can be selected. Some embodiments improve the accuracy of attentional state evaluation, for example, thereby improving the evaluation of a user's attention to a task (e.g., focusing on breathing techniques during a meditation experience). Some embodiments improve the user experience by providing cognitive evaluation that minimizes or avoids interruptions or interference with the user experience, for example, without significantly disrupting the user's attention or ability to perform the task.
[0004] In some implementations, auditory stimuli can be selected based on the characteristics of the user's environment (e.g., a real-world physical environment, a virtual environment, or a combination of both). Devices (e.g., handheld devices, laptops, desktop computers, or head-mounted displays (HMDs)) provide users with experiences of real-world physical environments, extended reality (XR) environments, or combinations of both (e.g., mixed reality environments) (e.g., visual and / or auditory experiences). The devices use sensors to acquire physiological data (e.g., electroencephalogram (EEG) amplitude, pupillary modulation, eye fixation saccades, etc.) associated with the user's response to auditory stimuli. Based on the acquired physiological data, the techniques described herein can determine the user's attentional state (e.g., focused, distracted, etc.) during an experience (e.g., a meditation experience). Based on the physiological data and associated physiological responses, the techniques can provide feedback to the user that their current attentional state differs from the expected attentional state of the experience, recommend similar content or similar portions of the experience, and / or adjust the content or feedback mechanisms corresponding to the experience.
[0005] In one exemplary implementation, the integration of meditation and mindfulness practices with the techniques described herein can enhance the meditation experience by providing individuals with real-time feedback on their meditation performance. Maintaining focus and engagement during meditation can improve a user's meditation practice and help them reap the benefits associated with meditation. For example, beginners interested in meditation may have difficulty staying focused on a task during a meditation session, and they may benefit from accurate feedback on their performance. The techniques described herein can present naturalized sounds consistent with the user's environment to detect when a user becomes distracted during meditation based on the user's physiological response (or lack thereof) to auditory stimuli. Identifying defined signs of attention loss during meditation and providing performance feedback enhances the user experience, thereby providing additional benefits from the meditation session and offering guided and supported teaching methods (e.g., via scaffolded teaching methods) to guide the user through their meditation practice.
[0006] Physiological response data, such as EEG amplitude / frequency, pupillary modulation, and eye fixation saccades, can depend on an individual's attentional state and the characteristics of the scene in front of them, as well as the auditory stimuli presented within it. Physiological response data can be obtained when a user performs mindfulness tasks requiring different levels of attention, such as focused attention on breathing meditation, using devices with eye-tracking technology. In some implementations, other sensors, such as EEG sensors, may be used to obtain physiological response data. Observing repeated measurements of physiological response data to auditory stimuli can provide insights into the user's underlying attentional state at different time scales. These attentional measures can be used to provide feedback during the meditation experience.
[0007] Beyond meditation experiences, the techniques described in this paper for assessing states of attention can be utilized. For example, an educational experience could notify a student to stay focused when they appear to be drifting off. Another example could be a workplace experience where a worker is notified that they need to concentrate on their current task. This could include providing feedback to a surgeon who may be getting tired during a long surgical procedure, or reminding a truck driver who has been driving for a long time that they are losing focus and may need to pull over for a nap. The techniques described in this paper can be tailored to any user and experience that may require some type of feedback mechanism to enter or maintain one or more specific states of attention.
[0008] Generally speaking, an innovative aspect of the subject matter described in this specification can be embodied in a method that includes the following actions: selecting auditory stimuli based on environmental characteristics; presenting auditory stimuli to a user; using sensors to obtain first physiological data associated with the user's physiological response to the auditory stimuli; and evaluating the user's attentional state based on the user's physiological response to the auditory stimuli.
[0009] These and other implementation schemes may optionally include one or more of the following features.
[0010] In some specific implementations, selecting auditory stimuli involves classifying the environment into a certain environment type and selecting auditory stimuli based on that environment type.
[0011] In some specific implementations, the selection of auditory stimuli includes: classifying one or more objects in the environment; and selecting auditory stimuli based on the classified one or more objects.
[0012] In some specific implementations, the selection of auditory stimuli includes: identifying one or more auditory stimuli from an auditory stimulus database that elicit a response used to evaluate the user's attention; and selecting auditory stimuli from one or more auditory stimuli based on the environment.
[0013] In some specific implementations, auditory stimuli are discrete sounds, a series of sounds, or spatialized sounds.
[0014] In some implementations, the environment is the physical surroundings of the user. In other implementations, the environment is an extended reality (XR) experience presented to the user.
[0015] In some implementations, obtaining first physiological data associated with a user's physiological response to auditory stimuli includes monitoring a response or lack thereof occurring within a predetermined timeframe following the presentation of the auditory stimulus. In some implementations, the first physiological data includes electroencephalogram (EEG) amplitude data associated with the user.
[0016] In some implementations, the primary physiological data includes pupil movement associated with the user. In some implementations, statistical or machine learning-based classification techniques are used to evaluate attentional states. In some implementations, the method also includes providing notifications to the user based on attentional states.
[0017] In some implementations, the method also includes identifying the portion of the content associated with the attention state.
[0018] In some implementations, the method also includes customizing content based on the user's attention state.
[0019] In some implementations, the method also includes aggregating attention states determined for multiple users viewing content to provide feedback on the content.
[0020] In some implementations, the device is a head-mounted display (HMD), and the environment includes an extended reality (XR) environment.
[0021] In some embodiments, a non-transitory computer-readable storage medium stores instructions that are computer-executable to perform or cause to perform any of the methods described herein. In some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. Attached Figure Description
[0022] Therefore, this disclosure will be understood by those skilled in the art, and a more detailed description can be made with reference to some exemplary embodiments, some of which are shown in the accompanying drawings.
[0023] Figure 1 The device is shown to display a visual experience and obtain physiological data from the user, based on some specific implementations.
[0024] Figure 2 It shows some specific implementations Figure 1 The user's pupils, where the diameter of the pupil changes over time.
[0025] Figure 3 This is a flowchart representation of a method for evaluating a user's attentional state based on physiological responses to auditory stimuli selected based on environment-based characteristics.
[0026] Figure 4 It demonstrates the selection of auditory stimuli based on environmental characteristics and the evaluation of the user's attentional state based on the physiological response to the auditory stimuli.
[0027] Figure 5 This is a flowchart representation of a method for evaluating a user's attentional state based on their physiological response to auditory stimuli associated with a virtual location in a three-dimensional (3D) coordinate system.
[0028] Figure 6A and Figure 6B This paper demonstrates how to evaluate a user's attentional state based on their physiological response to auditory stimuli associated with their virtual location in a 3D coordinate system.
[0029] Figure 7 Device components of an exemplary device according to some specific implementations are shown.
[0030] Figure 8 An exemplary head-mounted device (HMD) according to some specific implementations is shown.
[0031] As is customary, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Additionally, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation
[0032] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will recognize that other effective aspects or variations do not include all the specific details set forth herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.
[0033] Figure 1 A real-world environment 5 is illustrated, which includes a device 10 with a display 15. In some embodiments, the device 10 displays content 20 and visual characteristics 30 associated with the content 20 to a user 25. For example, the content 20 may be a button, a user interface icon, a text box, a graphic, etc. In some embodiments, the visual characteristics 30 associated with the content 20 include visual characteristics such as hue, saturation, size, shape, spatial frequency, motion, highlighting, etc. For example, the content 20 may be displayed as a visual characteristic 30 with a green highlighting that covers or surrounds the content 20.
[0034] In some implementations, content 20 may be a visual experience (e.g., a meditation experience), and the visual characteristics 30 of the visual experience may continuously change during the visual experience. As used herein, the phrase "experience" refers to a period of time during which a user uses an electronic device and has one or more states of attention. In one example, a user has an experience in which the user perceives a real-world environment while holding, wearing, or being near an electronic device that includes one or more sensors that acquire physiological data to evaluate eye characteristics indicative of the user's state of attention. In another example, a user has an experience in which the user perceives content displayed by an electronic device while the same or another electronic device acquires physiological data (e.g., pupil data, EEG data, etc.) to evaluate the user's state of attention. In yet another example, a user has an experience in which the user holds, wears, or is near an electronic device that provides a series of audible or visual instructions for a guided experience. For example, the instructions may instruct the user to have a specific state of attention during a particular period of the experience, such as instructing the user to focus on his or her breathing for the first 30 seconds, to stop focusing on his or her breathing for the next 30 seconds, and to refocus on his or her breathing for the next 45 seconds, etc. During such an experience, the same or another electronic device may acquire physiological data to evaluate the user's state of attention.
[0035] In some implementations, visual feature 30 is a user-specific feedback mechanism (e.g., visual or audio cues regarding focusing on a specific task (such as breathing during a meditation experience) during the experience). In some implementations, the visual experience (e.g., content 20) may occupy the entire display area of display 15. For example, during a meditation experience, content 20 may be a meditation video or image sequence, which may include visual and / or audio cues presented to the user as visual feature 30 regarding focusing on breathing. Other visual experiences that may be displayed for content 20 and visual and / or audio cues regarding visual feature 30 will be discussed further herein.
[0036] Device 10 acquires physiological data (e.g., EEG amplitude / frequency, pupil modulation, eye fixation saccades, etc.) from user 25 via sensor 35. For example, device 10 acquires pupil data 40 (e.g., eye fixation characteristic data). Although this example and other examples discussed herein illustrate a single device 10 in a real-world environment 5, the techniques disclosed herein are applicable to multiple devices and multiple sensors, as well as other real-world environments / experiences. For example, the functionality of device 10 may be performed by multiple devices.
[0037] In some specific implementations, such as Figure 1As shown, device 10 is a handheld electronic device (e.g., a smartphone or tablet). In some embodiments, device 10 is a laptop computer or desktop computer. In some embodiments, device 10 has a touchpad, and in some embodiments, device 10 has a touch-sensitive display (also referred to as a "touchscreen" or "touchscreen display"). In some embodiments, device 10 is a wearable head-mounted display ("HMD").
[0038] In some embodiments, device 10 includes an eye-tracking system for detecting eye position and eye movement. For example, the eye-tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye-tracking camera (e.g., a near-infrared (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) towards the eyes of user 25. Furthermore, the illumination source of device 10 may emit NIR light to illuminate the eyes of user 25, and the NIR camera may capture images of the eyes of user 25. In some embodiments, the images captured by the eye-tracking system may be analyzed to detect the position and movement of the eyes of user 25, or to detect other information about the eyes such as pupil dilation or pupil diameter. Furthermore, the gaze point estimated from the eye-tracking images enables gaze-based interaction with content displayed on the near-eye display of device 10.
[0039] In some embodiments, device 10 has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing multiple functions. In some embodiments, user 25 interacts with the GUI through finger contact and gestures on a touch-sensitive surface. In some embodiments, these functions include image editing, drawing, rendering, word processing, web page creation, disk editing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, fitness support, digital photography, digital video recording, web browsing, digital music playback, and / or digital video playback. Executable instructions for performing these functions may be included in a computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0040] In some implementations, device 10 employs various physiological sensors, detection, or measurement systems. The detected physiological data may include, but is not limited to: EEG, electrocardiogram (ECG), electromyography (EMG), functional near-infrared spectroscopy (fNIRS), blood pressure, skin conductance, or pupillary response. Furthermore, device 10 can simultaneously detect multiple forms of physiological data to benefit from synchronized acquisition of physiological data. Additionally, in some implementations, the physiological data represents involuntary data, i.e., responses not controlled by conscious effort. For example, pupillary response may represent involuntary movement.
[0041] In some embodiments, one or both eyes 45 of user 25 (including one or both pupils 50 of user 25) present physiological data (e.g., pupil data 40) in the form of pupillary responses. The pupillary responses of user 25 result in changes in the size or diameter of the pupils 50 via the optic nerve and the oculomotor cranial nerve. For example, pupillary responses may include a constriction response (pupil constriction), i.e., the pupil narrows, or a dilation response (pupil dilation), i.e., the pupil widens. In some embodiments, device 10 may detect patterns representing physiological data indicating time-varying pupil diameter.
[0042] In some implementations, the pupil response may be in response to auditory stimuli detected by one or both ears 60 of the user 25. For example, device 10 may include a speaker 12 that projects sound via sound waves 14. Device 10 may include other audio sources, such as a headphone jack for headphones, wireless connectivity to external speakers, etc.
[0043] Figure 2 It shows Figure 1 The pupil diameters of user 25 were measured from 50a to 50b, with the diameters of pupils 50a to 50b varying over time. Pupil diameter tracking may potentially indicate the user's physiological state. Figure 2 As shown, the current physiological state (e.g., current pupil diameter 55a to 55b) may differ from past physiological states (e.g., past pupil diameter 57a to 57b). For example, the current physiological state may include the current pupil diameter and the past physiological state may include the past pupil diameter.
[0044] Physiological data can vary over time, and device 10 can use the physiological data to measure one or both of a user's physiological response to visual feature 30 or a user's intention to interact with content 20. For example, when device 10 presents a list of content 20, such as content experiences (e.g., meditation environments), user 25 can select an experience without user 25 pressing a physical button. In some implementations, the physiological data includes the physiological response to auditory stimuli at the radius of the pupil 50 after user 25 has scanned content 20, measured via eye-tracking technology. In some implementations, the physiological data includes the physiological response to auditory stimuli as EEG amplitude / frequency data measured via EEG technology or EMG data measured from an EMG sensor or motion sensor.
[0045] Return to Figure 1According to some specific implementations, device 10 can generate and present an extended reality (XR) environment to its corresponding user. An extended reality (XR) environment refers to a fully or partially simulated environment that a person can interact with and / or sense using electronic devices. For example, an XR environment may include virtual reality (VR) content, augmented reality (AR) content, mixed reality (MR) content, etc. Using an XR system, a portion of a person's body movement or its representation can be tracked. In response, one or more characteristics of virtual objects simulated in the XR environment can be adjusted such that they conform to one or more physical laws. For example, an XR system can detect the user's head movement and, in response, adjust the graphical and auditory content presented to the user in a manner similar to how views and sounds change in a physical environment. In another example, an XR system can detect the movement of an electronic device (e.g., a laptop, mobile phone, tablet, etc.) presenting the XR environment and, in response, adjust the graphical and auditory content presented to the user in a manner similar to how views and sounds change in a physical environment. In some cases, an XR system can adjust one or more characteristics of the graphical content in the XR environment in response to a representation of physical movement (e.g., a voice command).
[0046] Various electronic systems enable a person to interact with and / or sense an XR environment. Examples include projection-based systems, head-mounted systems, head-up displays (HUDs), windows with integrated displays, vehicle windshields with integrated displays, displays designed to be placed on a user's eyes (e.g., similar to contact lenses), speaker arrays, headsets / earpieces, input systems (e.g., wearable or handheld controllers with or without haptic feedback), tablets, smartphones, and desktop / laptop computers. A head-mounted system may include an integrated opaque display and one or more speakers. In other examples, the head-mounted system may accept an external device with an opaque display (e.g., a smartphone). The head-mounted system may include one or more image sensors and / or one or more microphones to capture images or video and / or audio of the physical environment. In other examples, the head-mounted system may include a transparent or translucent display. The medium through which light representing the image is guided may be included within the transparent or translucent display. The display may utilize OLED, LED, uLED, digital light projection, laser scanning light sources, liquid crystal on silicon, or any combination of these technologies. The medium can be a holographic medium, an optical combiner, an optical waveguide, an optical reflector, or a combination thereof. In some examples, transparent or translucent displays can be configured to selectively become opaque. Projection-based systems can use retinal projection techniques to project graphic images onto a user's retina. Projection systems can also be configured to project virtual objects into a physical environment, such as onto a physical surface or as a hologram.
[0047] Figure 3 This is a flowchart illustrating an exemplary method 300. In some specific implementations, the device is such as device 10 ( Figure 1 The techniques of method 300 are used to evaluate a user's attentional state based on physiological responses to auditory stimuli selected based on environmental characteristics (e.g., visual and / or auditory electronic content that may be a real-world physical environment, virtual content, or a combination of both). In some embodiments, the techniques of method 300 are performed on mobile devices, desktop computers, laptops, HMDs, or server devices. In some embodiments, method 300 is performed on processing logic components (including hardware, firmware, software, or combinations thereof). In some embodiments, method 300 is performed on a processor executing code stored in a non-transitory computer-readable medium (e.g., memory).
[0048] At box 302, method 300 selects an auditory stimulus (e.g., sound) based on characteristics of the environment. For example, a bird chirping sound might be selected as the auditory stimulus based on the determination that the user is viewing a lush, tree-lined environment where bird chirping can occur naturally. In some implementations, the environment is a real-world environment presented to the user. For example, the experience could include a live video of the physical environment (e.g., a live view of nature for meditation) or a live view via an HMD (e.g., the user being in a real-world view of nature for meditation, such as a quiet park). In some implementations, the environment is an XR environment presented to the user. Alternatively, the environment could be a mixed reality (MR) experience presented to the user, where virtual reality images can be overlaid on a live view of the physical environment (e.g., augmented reality (AR)).
[0049] In some implementations, the characteristics of the environment can be determined by classifying the environment into a certain environmental type (e.g., forest, park, school, beach, crowded events, etc.) and selecting the corresponding sound as the auditory stimulus. Alternatively, the characteristics of the environment can be determined by classifying one or more specific objects in the environment (e.g., trees, birds, waves, etc.) and selecting the corresponding sound as the auditory stimulus based on the classified one or more objects.
[0050] In some implementations, the system may compile a library of sounds determined to elicit appropriate user responses for evaluating attention and select one of those sounds based on the user's environment. For example, method 300 may also include determining one or more auditory stimuli from an auditory stimulus database to elicit responses for evaluating user attention and selecting an auditory stimulus from the one or more auditory stimuli based on the environment.
[0051] At box 304, method 300 presents an auditory stimulus to the user. For example, the auditory stimulus can be a discrete sound (e.g., a bird chirp) or a sequence of sounds (e.g., beep-beep-beep-beep). In some implementations, the auditory stimulus is a spatialized sound. The auditory stimulus can be a sensory stimulus related to a natural event, which is largely inattentive in the sense that the stimulus can be mixed with the user's natural scene and surrounding environment. In particular, the spatial location and timing of such stimuli can be controlled so that their statistics match the specific sensory environment the user might experience. For example, a natural soundscape can be used during meditation, where bird sounds are distributed spatially and temporally, but capable of producing sensory-evoked neural responses without unpleasant or unnatural ones. Furthermore, using spatial audio, these specific events (e.g., auditory stimuli) can be spatially (e.g., along azimuth) varying to elicit a unilateral brain response (e.g., EEG / EMP amplitude data).
[0052] At box 306, method 300 uses sensors to acquire first physiological data (e.g., EEG amplitude / frequency, pupillary modulation, eye fixation saccades, etc.) associated with a user's physiological response (or lack thereof) to an auditory stimulus. For example, acquiring the physiological data may involve monitoring a response or lack thereof that occurs within a predetermined time after the presentation of the auditory stimulus.
[0053] In some implementations, obtaining initial physiological data associated with a user's physiological response to an auditory stimulus includes monitoring a response or lack thereof that occurs within a predetermined timeframe after the presentation of the auditory stimulus. For example, the system may wait up to five seconds to see if a spatialized bird chirping outside the user's field of vision causes the user to look in that direction (e.g., a physiological response).
[0054] In some implementations, obtaining physiological data (e.g., pupil data 40) may be associated with the gaze of a user who obtains images or electrocardiogram (EOG) signals of the eye from which the gaze direction and / or movement can be determined.
[0055] At box 308, method 300 evaluates a user's attentional state based on the user's physiological response to auditory stimuli. For example, the response can be compared to the user's own previous responses or typical user responses to similar auditory stimuli. In some implementations, statistical or machine learning-based classification techniques can be used to determine the attentional state. The determined attentional state can be used to provide feedback to the user, redirect the user, provide statistical data to the user, and / or help content creators improve the experience.
[0056] In some specific implementations, one or more pupil or EEG features can be identified, aggregated, and used to classify a user's attentional state using statistical or machine learning techniques. For example, physiological data can be classified based on comparing the variability of physiological data with a threshold. For instance, if a baseline of a user's EEG data is established during an initial period (e.g., 30 to 60 seconds), and during a subsequent period following auditory stimulation (e.g., 5 seconds), the EEG data deviates from the EEG baseline by more than + / - 10% during the subsequent period, the techniques described herein can classify the user as transitioning from a first attentional state (e.g., meditation) to a second attentional state (e.g., daydreaming).
[0057] In some implementations, physiological responses can be compared based on the statistical frequency of the sounds. For example, a number of natural sounds can be presented to the user, some of which are common (e.g., appearing 80% of the time) and some of which are less common (e.g., appearing 20% of the time). In some implementations, the novelty of the less common sounds relative to the more common sounds may amplify the physiological response to the less common sounds, and the physiological response to the less common sounds can be measured approximately 300 ms to 800 ms after presentation to the user.
[0058] In some implementations, machine learning models can be used to classify a user's attention state. For example, labeled training data about a user can be fed to the machine learning model. In some implementations, the machine learning model is a neural network (e.g., an artificial neural network), decision tree, support vector machine, Bayesian network, etc. These labels can be collected in advance from the user or from a group of people and fine-tuned later for individual users. Creating the labeled data might require many users to experience (e.g., a meditation experience) in which the user can listen to natural sounds (e.g., auditory stimulation) using mixed natural probes, and then randomly ask the user how focused or relaxed they were shortly after the probes were presented. The answers to these questions can generate labels in time before the questions are asked, and deep neural networks or deep long short-term memory (LSTM) networks may learn combinations of features specific to that user or task given those labels.
[0059] In some specific implementations, use cases for evaluating attentional state based on the user's physiological response to auditory stimuli may include meditation experiences, educational experiences, career experiences, etc.
[0060] In some implementations, feedback may be provided to the user based on determining that the first state of attention (e.g., distraction) differs from the expected state of attention for the experience (e.g., focused attention). In some implementations, method 300 may also include presenting feedback (e.g., audio feedback such as "Control your breathing," visual feedback, etc.) during the experience in response to determining that the first state of attention differs from the expected second state of attention for the experience. In one example, during a part of the meditation experience instructing the user to focus on his or her breathing, the method determines to present feedback to remind the user to focus on breathing based on detecting that the user is instead in a distracted state of attention.
[0061] In some implementations, content recommendations for content developers can be provided based on determining the user's state of attention during the presented experience and changes in the experience or content presented therein. For example, a user may be highly focused when a particular type of content is presented. In some implementations, method 300 may also include identifying content based on the similarity between the content and the experience, and providing content recommendations to the user based on determining that the user has a primary state of attention (e.g., distraction) during the experience.
[0062] In some implementations, the content of the experience can be adjusted to correspond with the experience based on an attention state that differs from the expected attention state of the experience. For example, an experienced developer can adjust the content to improve the recorded content for subsequent use by the user or other users. In some implementations, method 300 may also include adjusting the content corresponding to the experience in response to determining that a first attention state differs from a second attention state intended for use in the experience.
[0063] In some implementations, the techniques described herein obtain physiological data from the user (e.g., pupil data 40, EEG amplitude / frequency data, pupil modulation, eye fixation saccades, etc.) based on identifying typical user interactions with the experience. For example, the techniques can determine the variability of a user's eye fixation characteristics in relation to interactions with auditory stimuli presented within the experience. Furthermore, the techniques described herein can then adjust the visual characteristics of the experience, or adjust / change the sound associated with the auditory stimuli, to enhance physiological response data associated with the experience and / or future interactions with the auditory stimuli presented within the experience. Additionally, in some implementations, changing the auditory stimuli after the user interacts with them within the experience informs the user's physiological response in subsequent interactions with the experience or a specific segment of the experience. For example, before the auditory stimuli are changed within the experience, the user may present an anticipated physiological response associated with the change in the auditory stimuli. Therefore, in some implementations, the techniques identify the user's intention to interact with the auditory stimuli based on anticipated physiological responses. For example, technology can adapt or train instruction sets by capturing or storing users’ physiological data (including users’ responses to enhanced / updated auditory stimuli) based on users’ interactions with experiences and auditory stimuli, and can detect users’ future intentions to interact with experiences and auditory stimuli by recognizing users’ physiological responses in the expected presentation of enhanced / updated auditory stimuli.
[0064] In some implementations, estimators or statistical learning methods are used to better understand or predict physiological data (e.g., pupillary data characteristics, EEG data, etc.). For example, statistical data of EEG data can be estimated by sampling the dataset with replacement data (e.g., bootstrap method).
[0065] Figure 4 This is a system flowchart based on an exemplary environment 400 of some specific implementations, wherein the attention state evaluation system can select auditory stimuli based on environmental characteristics and evaluate the user's attention state based on the physiological response to the auditory stimuli. In some specific implementations, the system flow of exemplary environment 400 is in the device (e.g., Figure 1 Device 10) such as a mobile device, desktop computer, laptop computer, or server device. The content of exemplary environment 400 may be displayed on a device having a screen for displaying images (e.g., monitor 15) and / or a screen for viewing stereoscopic images (e.g., [device name missing]). Figure 1 On a device 10, such as a head-mounted display (HMD). In some embodiments, the system processes of exemplary environment 400 are executed on processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, the system processes of exemplary environment 400 are executed on a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).
[0066] The system processes in the exemplary environment 400 originate from the physical environment (e.g., Figure 1 The physical environment (5) uses sensors to acquire image and / or sound data, analyzes and classifies the environmental image and / or sound data, selects auditory stimuli based on environmental characteristics, presents the auditory stimuli to the user, obtains first physiological data associated with the user's physiological response to the auditory stimuli, and evaluates the user's attentional state based on the user's physiological response to the auditory stimuli. For example, the attentional state evaluation technique described herein determines the user's attentional state (e.g., focused, distracted, etc.) during an experience (e.g., a meditation experience) by providing auditory stimuli based on the user's environment (e.g., birdsong while meditating outdoors, school bells while studying at school, workplace noises while working, etc., such as the rolling sound of a file cart in the workroom) based on the obtained physiological data.
[0067] In one exemplary embodiment, environment 400 includes an image and sound synthesis pipeline that acquires or obtains data about the physical environment (e.g., image data from an image source such as a camera on device 402). Exemplary environment 400 is an example of acquiring image sensor data (e.g., light intensity data, depth data, and location information) and sound data from one or more image frames of the current environment. For example, a user acquires physical environment data (e.g., image data from a camera on device 402). Figure 1 The physical environment 5) includes image data 404 and sound data 406. Image sources may include a depth camera acquiring depth data of the physical environment, a light intensity camera (e.g., an RGB camera) acquiring light intensity image data (e.g., a sequence of RGB image frames), and a position sensor for acquiring positioning information. Sound sources may include a microphone on device 402 (e.g., a microphone on the device). Figure 1 Equipment 10).
[0068] In some embodiments, a position sensor may be used to acquire positioning information, which can be used to acquire additional information about the device's position relative to the environment during the acquisition of image data 404 and / or sound data 406. For the positioning information, some embodiments include a visual inertial ranging (VIO) system that uses a sequence of camera images (e.g., image data 404) to determine equivalent ranging information to estimate the distance traveled. Alternatively, some embodiments of this disclosure may include a SLAM system (e.g., a position sensor). This SLAM system may include a GPS-independent, multi-dimensional (e.g., 3D) laser scanning and range measurement system that provides real-time simultaneous localization and mapping. This SLAM system can generate and manage highly accurate point cloud data resulting from reflections of laser scans from objects in the environment. Accurately tracking the movement of any points in the point cloud over time allows the SLAM system to use points in the point cloud as reference points for its position, maintaining an accurate understanding of its position and orientation as it travels through the environment. This SLAM system may also be a visual SLAM system that relies on light intensity image data to estimate the position and orientation of the camera and / or device.
[0069] In one exemplary embodiment, environment 400 includes an environment classifier instruction set 410 configured to be executable by a processor to generate classified environmental data from image data and / or sound data of the environment. For example, the environment classifier instruction set 410 draws from sensors on device 402 and the physical environment (e.g., Figure 1The system acquires image data 404 (e.g., real-time camera footage such as RGB images from a light intensity camera) and / or sound data 406 from other sources of physical environment information (e.g., camera positioning information such as pose data from a position sensor) and classifies the environment as classified environmental data 414. The classified environmental data 414 is determined using one of several class-specific neural networks: Environment-Class 1 Neural Network 412A, Environment-Class 2 Neural Network 412B, Environment-Class 3 Neural Network 412C, and Environment-Class N Neural Network 412N (collectively referred to herein as Environment-Class Neural Network 412). For example, a first network (e.g., Environment-Class 1 Neural Network 412A) is trained to analyze specific objects or features of the environment to determine its classification. For example, Environment-Class Neural Network 412 may detect trees, animals, etc., to determine that the user's current environment is outside a natural area. Environment-Class Neural Network 412 may detect cars, buildings, etc., to determine that the user's current environment is outside an urban area. The environmental neural network 412 can detect desktops, books, students, etc., to determine if the user's current environment is within a classroom. Each category may also include subcategories. For example, a quiet nature trail could be in a city park or in a more distant area outside the city. Each category (or subcategory) can also enhance the auditory stimulus selection process to select auditory stimuli that are largely unnoticed sensory stimuli related to natural events in a way that blends with the user's natural scene and surrounding environment. In particular, the spatial location and timing of such stimuli can be controlled so that their statistics match the specific sensory environment the user might experience. For example, a natural soundscape could be used during meditation, where bird sounds are distributed spatially and temporally but are able to generate sensory-evoked neural responses without producing unpleasant or unnatural ones.
[0070] In one exemplary embodiment, environment 400 also includes an auditory stimulation instruction set 420 configured with instructions executable by a processor to select auditory stimuli based on environmental data. For example, auditory stimulation instruction set 420 acquires physical environment data (e.g., from environment classifier instruction set 410) Figure 1The system uses classified environmental data 414 of the physical environment 5) and determines auditory stimuli 422 based on the classification of the environment and selects the auditory stimulus from the auditory stimulus database 425. Alternatively, the auditory stimulus instruction set 420 selects auditory stimuli 422 based on identified characteristics of the environment. For example, if a specific object (e.g., a bird) is identified in the environment, the auditory stimulus instruction set 420 may select bird sounds as auditory stimuli 422. In particular, the spatial location and timing of such stimuli can be controlled so that their statistics match the specific sensory environment that the user may experience. For example, natural soundscapes can be used during meditation, where bird sounds are distributed in space and time but are able to produce sensory-evoked neural responses without unpleasant or unnatural neural responses. Furthermore, using spatial audio, these specific events (e.g., auditory stimuli) can be spatially (e.g., along azimuth) varied to elicit unilateral brain responses. In some specific implementations, cognitive appraisal techniques using auditory stimulation can be employed throughout the entire duration of the experience without significantly disrupting the user's attention or ability to perform the task, but at the same time producing measures that improve the user's attention to the task (e.g., focusing on breathing techniques during a meditation experience).
[0071] In one exemplary embodiment, environment 400 also includes a content instruction set 430 configured to be executable by a processor to provide and / or track content to be displayed on the device. For example, content instruction set 430 may acquire auditory stimuli 422 from auditory stimulus instruction set 420 and provide content 432 to user 25. Content 432 may include background images and sound data 434. Content 432 may be an XR experience (e.g., a meditation experience), or it may be an MR experience comprising some XR content and images of a physical environment. Alternatively, the user may wear an HMD and view the real physical environment via a live camera view, or the HMD may allow the user to view a display, such as smart glasses that the user can view through, but still present visual and / or audio cues. During the experience, while user 25 is viewing and listening to background images and sound data 434, pupil data 435 of the user's eyes (e.g., pupil data 40 such as eye fixation characteristic data) may be monitored and sent as physiological data 444. Additionally or alternatively, the user 25 wears a sensor 440 (e.g., an EEG sensor) that generates sensor data 442 (e.g., EEG data) as physiological data. Therefore, when auditory stimulation 422 is presented to the user, physiological data 444 (e.g., pupil data 435) and / or sensor data 442 are sent to a physiological tracking instruction set 450 using one or more of the techniques discussed herein or other potentially suitable techniques to track the user's physiological attributes as physiological tracking data 452.
[0072] In one exemplary embodiment, environment 400 also includes an attention state instruction set 460 configured to be executable by a processor to evaluate a user's attention state (e.g., attention states such as daydreaming, meditation, etc.) based on physiological responses (e.g., eye fixation responses) using one or more of the techniques discussed herein or other techniques that may be appropriate. For example, attention state instruction set 460 acquires physiological tracking data 452 from physiological tracking instruction set 450 and determines the user 25's attention state (e.g., attention states such as daydreaming, meditation, etc.) before, during, and / or after the presentation of auditory stimulus 422. In some embodiments, attention state instruction set 460 may then provide feedback data 464 to content instruction set 430 based on cognitive evaluation. For example, identifying defined signs of attention loss during meditation and providing performance feedback can enhance the user experience, thereby providing additional benefits from the meditation session and providing guided and supported teaching methods (e.g., scaffolded teaching methods) to guide the user through their meditation practice.
[0073] In some implementations, the content state instruction set 430 may utilize feedback data 464 to present audio and / or visual feedback cues or mechanisms to the user 25 to help them relax and focus on their breathing during a meditation session. In educational experiences, based on an assessment from the attention state instruction set 460 indicating that the user 25 is distracted by auditory stimuli 422, the feedback cues to the user may be gentle reminders (e.g., soothing or calming visual and / or audio alerts) to re-engage the learning task. As discussed herein, auditory stimuli 422 are intended to be selected as natural sounds from the user's current environment, such that the user should not be distracted by auditory stimuli 422 if they are focused on the task at hand. For example, in a meditation experience with tranquil natural content by a lake, the user should not be distracted by auditory stimuli 422 that sound like birdsong in the background. In another example, for users identified as being in a workplace environment, they should not be distracted by auditory stimuli that sound like someone walking through their office / workspace (such as a colleague pushing a file cart through their workspace).
[0074] Figure 5 This is a flowchart illustrating an exemplary method 500. In some specific implementations, devices such as device 10 ( Figure 1The technique of executing method 500 evaluates a user's attentional state based on a physiological response to auditory stimuli associated with a virtual location in a 3D coordinate system. The auditory stimuli may be presented during the presentation of content in the environment (e.g., visual and / or auditory electronic content that may be a real-world physical environment, virtual content, or a combination of both). In some embodiments, the technique of method 500 is executed on a mobile device, desktop computer, laptop computer, HMD, or server device. In some embodiments, method 500 is executed on processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, method 500 is executed on a processor executing code stored in a non-transitory computer-readable medium (e.g., memory).
[0075] At box 502, method 500 presents auditory stimuli during the presentation of an XR environment, wherein the auditory stimuli are associated with virtual locations in a 3D coordinate system. For example, the auditory stimulus may be a sound generated by manipulating sound produced by stereo speakers, speaker arrays, or headphones to virtually place a sound source in 3D space (e.g., to the user's left, right, behind, above, below, etc.). In some implementations, the sound may be a discrete sound at a single virtual location (e.g., a bird's chirp). Alternatively, the sound may be an audio segment occurring over time, and the 3D coordinates associated with the virtual location remain stationary during the time period. The auditory stimulus may be a sensory stimulus related to a natural event, which is largely inattentive in the sense that the stimulus can blend with the user's natural scene and surrounding environment. For example, an auditory stimulus (e.g., a bird's chirp) may occur at the same 3D location relative to the user. In particular, the spatial location and timing of such stimuli can be controlled such that their statistics match the specific sensory environment the user may experience. For example, natural soundscapes can be used during meditation, where bird sounds are distributed spatially and temporally, yet capable of generating sensory-evoked neural responses without producing unpleasant or unnatural ones. Furthermore, using spatial audio, these specific events (e.g., auditory stimuli) can be spatially varied (e.g., along azimuth) to elicit a unilateral brain response. Alternatively, the sound can be a time-varying audio segment, and the 3D coordinates associated with the virtual location change during the time period. For example, the sound can spatially change over time, and the user is presented with the impression that the sound is getting closer (e.g., a bird flying towards you, and the auditory stimulus is getting closer because it appears to be approaching the user).
[0076] In some implementations, the system may compile a library of sounds determined to elicit appropriate user responses for evaluating attention and select one of those sounds based on the user's environment. For example, method 500 may also include determining one or more auditory stimuli from an auditory stimulus database to elicit responses for evaluating user attention and selecting an auditory stimulus from the one or more auditory stimuli during the presentation of the XR environment.
[0077] At box 504, method 500 uses sensors to acquire first physiological data (e.g., EEG amplitude / frequency, pupillary modulation, eye fixation saccades, etc.) associated with a user's physiological response (or lack thereof) to an auditory stimulus. For example, acquiring the physiological data may involve monitoring a response or lack thereof that occurs within a predetermined time after the presentation of the auditory stimulus.
[0078] In some implementations, obtaining initial physiological data associated with a user's physiological response to auditory stimuli includes determining possible responses that correspond to a virtual location. For example, a sound to the user's left produces a response in which the user looks to the left.
[0079] In some implementations, obtaining initial physiological data associated with a user's physiological response to an auditory stimulus includes monitoring a response or lack thereof that occurs within a predetermined timeframe after the presentation of the auditory stimulus. For example, the system may wait up to five seconds to see if a spatialized bird chirping outside the user's field of vision causes the user to look in that direction (e.g., a physiological response).
[0080] In some implementations, obtaining physiological data (e.g., pupil data 40) may be associated with the gaze of a user who obtains images or electrocardiogram (EOG) signals of the eye from which the gaze direction and / or movement can be determined.
[0081] At box 506, method 500 evaluates a user's attentional state based on the user's physiological response to auditory stimuli. For example, the response can be compared to the user's own previous responses or typical user responses to similar auditory stimuli. In some implementations, statistical or machine learning-based classification techniques can be used to determine the attentional state. The determined attentional state can be used to provide feedback to the user, redirect the user, provide statistical data to the user, and / or help content creators improve the experience.
[0082] In some implementations, one or more pupil or EEG features may be identified, aggregated, and used to classify a user's attentional state using statistical or machine learning techniques. In some implementations, physiological data is classified based on comparing the variability of the physiological data to a threshold. For example, if a baseline of a user's EEG data is established during an initial period (e.g., 30 to 60 seconds), and the EEG data deviates from the EEG baseline by more than + / - 10% during a subsequent period (e.g., 5 seconds) following auditory stimulation, the techniques described herein may classify the user as transitioning from a first attentional state (e.g., meditation) to a second attentional state (e.g., daydreaming).
[0083] In some implementations, machine learning models can be used to classify a user's attention state. For example, labeled training data about a user can be fed to the machine learning model. In some implementations, the machine learning model is a neural network (e.g., an artificial neural network), decision tree, support vector machine, Bayesian network, etc. These labels can be collected in advance from the user or from a group of people and fine-tuned later for individual users. Creating the labeled data might require many users to experience (e.g., a meditation experience) in which the user can listen to natural sounds (e.g., auditory stimulation) using mixed natural probes, and then randomly ask the user how focused or relaxed they were shortly after the probes were presented. The answers to these questions can generate labels in time before the questions are asked, and deep neural networks or deep long short-term memory (LSTM) networks may learn combinations of features specific to that user or task given those labels.
[0084] In some specific implementations, use cases for evaluating attentional state based on the user's physiological response to auditory stimuli may include meditation experiences, educational experiences, career experiences, etc.
[0085] In some implementations, feedback may be provided to the user based on the determination that the initial state of attention (e.g., distraction) differs from the expected state of attention for the experience (e.g., focused attention). In some implementations, method 500 may also include presenting feedback (e.g., audio feedback, such as "Control your breathing," visual feedback, etc.) during the experience in response to the determination that the initial state of attention differs from the expected second state of attention for the experience. In one example, during a part of the meditation experience instructing the user to focus on his or her breathing, the method determines to present feedback to remind the user to focus on breathing based on detecting that the user is instead in a distracted state of attention.
[0086] In some implementations, content recommendations for content developers can be provided based on determining the user's state of attention during the presented experience and changes in the experience or content presented therein. For example, a user may be highly focused when a particular type of content is presented. In some implementations, method 500 may also include identifying content based on the similarity between the content and the experience, and providing content recommendations to the user based on determining that the user has a primary state of attention (e.g., distraction) during the experience.
[0087] In some implementations, the content of the experience can be adjusted to correspond with the experience based on an attention state that differs from the expected attention state of the experience. For example, an experienced developer can adjust the content to improve the recorded content for subsequent use by the user or other users. In some implementations, method 500 may also include adjusting the content corresponding to the experience in response to determining that a first attention state differs from a second attention state intended for use in the experience.
[0088] In some implementations, estimators or statistical learning methods are used to better understand or predict physiological data (e.g., pupillary data characteristics, EEG data, etc.). For example, statistical data of EEG data can be estimated by sampling the dataset with replacement data (e.g., bootstrap method).
[0089] Figure 6A and Figure 6B This paper demonstrates how to evaluate a user's attentional state based on their physiological response to auditory stimuli associated with their virtual location in a 3D coordinate system. Figure 6A The illustration shows an auditory stimulus being presented to a user at a 3D location of the presented content during content presentation, wherein the user has a physiological response to the auditory stimulus (e.g., the user looking towards the 3D location of the spatialized sound) via acquired physiological data. For example, a user (e.g., user 25) is being presented with content 610a, which includes background sound and visual content (e.g., a natural scene for meditation), and the user's pupil data 612a is monitored as a baseline. Content 620a is then presented as an auditory stimulus, as any physiological response (e.g., EEG amplitude / frequency, pupil modulation, eye fixation saccades, etc.) is monitored in the user's pupil data 622a. After a period of time (e.g., 0 to 5 seconds) following the onset of the auditory stimulus, content 630a is presented again with the same auditory stimulus, and the user's pupil data 632a indicates that the user's eye fixation is drawn to the 3D location of the auditory stimulus. Therefore, the user has a physiological response to the auditory stimulus, and thus the evaluation of attentional state would be that the user is inattentive and may be distracted (e.g., not focused on the task at hand, such as meditation). In some implementations, if a user is deemed distracted, a feedback mechanism or prompt can be presented along with the content to help the user refocus on the task related to the content.
[0090] Figure 6B The illustration shows an auditory stimulus being presented to a user during content presentation, where the user shows no physiological response to the stimulus based on acquired physiological data. For example, a user (e.g., user 25) is being presented with content 610b, which includes background sound and visual content (e.g., a natural scene for meditation), and the user's pupil data 612b is monitored as a baseline. Content 620b is then presented as an auditory stimulus, as any physiological response (e.g., EEG amplitude / frequency, pupil modulation, eye fixation saccades, etc.) is being monitored in the user's pupil data 622b. After a certain period (e.g., 0 to 5 seconds) following the onset of the auditory stimulus, content 630b is presented again with the same auditory stimulus, and the user's pupil data 632b indicates that the user's eye fixation is not drawn to the 3D location of the auditory stimulus. Therefore, the user shows no physiological response to the auditory stimulus, and thus the evaluation of attentional state would be that the user is attentive and not distracted (e.g., focused on the task at hand, such as meditation).
[0091] In some implementations, the technology can be trained on multiple sets of user physiological data and then adapted individually for each user. For example, content creators can customize the meditation experience based on user physiological data, such as allowing users to request background music for meditation or more or less audio or visual cues to maintain the meditation.
[0092] In some implementations, the customization of the experience can be controlled by the user. For example, a user can choose the meditation experience he or she desires, such as selecting the surrounding environment, background scene, music, etc. Additionally, the user can change the threshold for providing feedback mechanisms in response to auditory stimuli. For example, a user can customize the sensitivity of triggering the feedback mechanism based on previous experiences in responding to auditory stimuli. For instance, a user might expect less feedback notification and allow for a certain degree of distraction (e.g., eye position deviation) before triggering notification. Therefore, a specific experience can be customized at the trigger threshold when higher standards are met. For example, in some experiences (such as educational experiences), a user may not want to be disturbed during a learning session, even if he or she briefly stares at the task or loses focus due to briefly looking at the auditory stimulus for a moment (e.g., less than 30 seconds) to think about what he or she has just read. However, a student / reader will want to be notified if he or she loses focus for a longer period (e.g., longer than or equal to 30 seconds) in response to an auditory stimulus.
[0093] In some specific implementations, the techniques described herein can interpret the user 25’s real-world environment 5 (e.g., visual qualities such as brightness, contrast, semantic context) when assessing how much the presented content or feedback mechanism is modulated or adjusted to enhance the user 25’s physiological response to visual characteristics 30 (e.g., feedback mechanism).
[0094] In some implementations, physiological data (e.g., pupil data 40) may vary over time, and the techniques described herein can use physiological data to detect patterns. In some implementations, a pattern is a change in physiological data from one time point to another, and in some other implementations, a pattern is a series of changes in physiological data over a period of time. Based on the detected patterns, the techniques described herein can identify changes in a user's attentional state (e.g., daydreaming), and then provide the user 25 with feedback mechanisms (e.g., visual or auditory cues about focusing on breathing) during an experience (e.g., a meditation session) to return to the desired state (e.g., meditation). For example, the user 25's attentional state can be identified by detecting patterns in the user's gaze characteristics, the visual or auditory cues associated with the experience can be adjusted (e.g., a voice feedback mechanism indicating "focus on breathing" may also include visual cues or changes in the surrounding environment of the scene), and the user's gaze characteristics compared to the adjusted experience can be used to confirm the user's attentional state.
[0095] In some implementations, the techniques described herein can utilize training or calibration sequences to adapt to the specific physiological characteristics of a particular user 25. In some implementations, the technique presents a training scenario to the user 25, instructing the user 25 to interact with screen items (e.g., feedback objects). By providing the user 25 with known intentions or regions of interest (e.g., via instructions), the technique can record the user's physiological data (e.g., pupil data 40) and identify patterns associated with the user's gaze. In some implementations, the technique can modify the visual characteristics 30 associated with content 20 (e.g., feedback mechanisms) to further adapt to the unique physiological characteristics of the user 25. For example, the technique can instruct the user to subjectively select a button associated with an auditory stimulus at the center of the screen while counting to three, and record the user's physiological data (e.g., pupil data 40) to identify patterns associated with the user's attentional state. Furthermore, the technique can alter or modify the visual characteristics associated with auditory stimuli to identify patterns associated with the user's physiological response to the modified visual characteristics. In some implementations, patterns associated with user 25's physiological responses are stored in a user profile associated with that user, and this user profile can be updated or recalibrated at any time in the future. For example, the user profile can be automatically modified over time during the user experience to provide a more personalized user experience (e.g., a personal meditation experience).
[0096] In some specific implementations, machine learning models (e.g., trained neural networks) are applied to identify patterns in physiological data, including patterns in content (e.g., Figure 1 The physiological responses to auditory stimuli during the presentation of content 20. Furthermore, this machine learning model can be used to match these patterns with indications of interest or intent corresponding to user 25's interaction with the auditory stimuli. In some specific implementations, the techniques described herein can learn patterns specific to a particular user 25. For example, the technique may begin learning by identifying peak patterns that indicate user 25's interest or intent in response to a specific visual feature 30 within the content, and use this information to subsequently identify similar peak patterns as another indication of user 25's interest or intent. This learning can take into account the user's relative interactions with multiple visual features 30 in order to further tailor the visual features 30 and enhance the user's physiological responses to auditory stimuli and the presented content.
[0097] In some implementations, the position and features of the user 25's head 27 (e.g., the edges of the eyes, nose, or nostrils) are extracted by the device 10 and used to find the coarse position coordinates of the user 25's eyes 45, thereby simplifying the determination of precise eye 45 features (e.g., position, gaze direction, etc.) and making gaze characteristic measurements more reliable and robust. Furthermore, the device 10 can easily combine the position of the 3D components of the head 27 with gaze angle information obtained through eye component image analysis to identify a given screen object viewed by the user 25 at any given time. In some implementations, using 3D mapping combined with gaze tracking allows the user 25 to freely move his or her head 27 and eyes 45, while reducing or eliminating the need for active tracking of the head 27 using sensors or transmitters on the head 27.
[0098] By tracking the eyes 45, some implementations reduce the need for recalibrating the user 25 after the user 25 moves his or her head 27. In some implementations, device 10 uses depth information to track the movement of the pupil 50, thereby enabling the calculation of a reliably presented pupil diameter 55 based on a single calibration of the user 25. Utilizing techniques such as central pupillary corneal reflection (PCCR), pupil tracking, and pupil shape, device 10 can calculate the pupil diameter 55 and the gaze angle of the eyes 45 from a fixed point on the head 27, and use the positional information of the head 27 to recalculate the gaze angle and other gaze characteristic measurements. In addition to reduced recalibration, further beneficial effects of tracking the head 27 may include reducing the number of light projection sources and the number of cameras used to track the eyes 45.
[0099] In some implementations, the techniques described herein can identify specific objects within content presented on the display 15 of device 10 at a location in the user's gaze direction. Furthermore, the techniques can modify the state of visual characteristics 30 associated with a specific object or overall content experience in response to verbal commands received from user 25 and the user's perceived attentional state. For example, a specific object within the content could be an icon associated with a software application, which the user 25 could gaze at, say the word "select" to select the application, and have a highlighting effect applied to the icon. The techniques can then use additional physiological data (e.g., pupil data 40) in response to visual characteristics 30 (e.g., feedback mechanisms) to further identify the user 25's attentional state as confirmation of the user's verbal command. In some implementations, the techniques can identify a given interactive item in response to the user's gaze direction and manipulate that given interactive item in response to physiological data (e.g., variability in gaze characteristics). The techniques can then confirm the user's gaze direction based on further identifying the user's attentional state using physiological data in response to auditory stimuli. In some implementations, the technology can remove interactive items or objects based on identified interests or intentions. In other implementations, upon determining the user's interests or intentions (e.g., in response to auditory stimuli), the technology can automatically capture images of the content.
[0100] As a power-saving feature, the technology described herein can detect when user 25 is not looking at the display and can activate power-saving technologies, such as disabling physiological sensors when user 25 looks away for more than a certain threshold time period. Furthermore, in some embodiments, the technology can dim or completely black the display when user 25 is not looking at it (e.g., reduce brightness). When user 25 looks at the display again, the technology can deactivate the power-saving technologies. In some embodiments, the technology can use a first sensor to track physiological attributes and then activate a second sensor based on that tracking to obtain physiological data. For example, the technology can use a camera (e.g., a camera on device 10) to identify that user 25 is looking in the direction of device 10, and then activate an eye sensor when it is determined that user 25 is looking in the direction of device 10.
[0101] Figure 7This is a block diagram of an exemplary device 700. Device 700 illustrates an exemplary device configuration of device 10. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 10 includes one or more processors 702 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 706, one or more communication interfaces 708 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, SPI, I2C and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 710, one or more displays 712, one or more sensor systems 714, memory 720, and one or more communication buses 704 for interconnecting these components and various other components.
[0102] In some embodiments, the one or more communication buses 704 include circuitry for communication between interconnecting system components and control system components. In some embodiments, the one or more I / O devices and sensors 706 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, or one or more depth sensors (e.g., structured light, time-of-flight, etc.) and / or similar devices.
[0103] In some embodiments, one or more displays 712 are configured to present a view of a physical or graphical environment to a user. In some embodiments, one or more displays 712 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 712 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, device 10 includes a single display. As another example, device 10 includes displays for each of the user's eyes.
[0104] In some embodiments, the one or more sensor systems 714 are configured to acquire sensor data corresponding to at least a portion of the physical environment 5. For example, the one or more sensor systems 714 include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, etc. In various embodiments, the one or more sensor systems 714 also includes an illumination source emitting light, such as a flash. In various embodiments, the one or more sensor systems 714 also includes an on-camera image signal processor (ISP) configured to perform multiple processing operations on the image data.
[0105] Memory 720 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 720 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 720 optionally includes one or more storage devices remotely located to one or more processors 702. Memory 720 includes a non-transitory computer-readable storage medium.
[0106] In some embodiments, memory 720 or a non-transitory computer-readable storage medium of memory 720 stores an optional operating system 730 and one or more instruction sets 740. Operating system 730 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, instruction set 740 includes executable software defined by binary information stored in charge. In some embodiments, instruction set 740 is software that can be executed by one or more processors 702 to implement one or more of the techniques described herein.
[0107] Instruction set 740 includes content instruction set 742, physiological tracking instruction set 744, and attention state instruction set 746. Instruction set 740 can be represented as a single software executable file or multiple software executable files.
[0108] In some implementations, the content instruction set 742 may be executed by the processor 702 to provide and / or track content for display on the device. The content instruction set 742 may be configured to monitor and track content over time (e.g., during an experience such as a meditation session) and / or identify change events occurring within the content. In some implementations, the content instruction set 742 may be configured to add change events to the content (e.g., a feedback mechanism) using one or more of the techniques discussed herein or other techniques that may be appropriate. For these purposes, in various implementations, the instructions include instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.
[0109] In some implementations, the physiological tracking instruction set 744 may be executed by processor 702 to track a user's physiological attributes (e.g., EEG amplitude / frequency, pupil modulation, eye fixation saccades, etc.) using one or more of the techniques discussed herein or other techniques that may be appropriate. For these purposes, in various implementations, the instructions include instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0110] In some implementations, the attention state instruction set 746 may be executed by processor 702 to evaluate a user's attention state (e.g., daydreaming, concentrating, meditating, etc.) based on physiological responses (e.g., eye gaze responses) using one or more of the techniques discussed herein or other techniques that may be appropriate. For these purposes, in various implementations, the instructions include instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0111] Although instruction set 740 is shown as residing on a single device, it should be understood that in other specific implementations, any combination of elements may reside on separate computing devices. Furthermore, Figure 7 This is used more as a functional description of various features present in a particular implementation, and differs from the structural diagrams of the specific implementations described herein. As will be recognized by those skilled in the art, items shown individually can be combined, and some items can be separate. The actual number of instruction sets and how features are allocated therein will vary depending on the specific implementation and can depend in part on the specific combination of hardware, software, and / or firmware chosen for that particular implementation.
[0112] Figure 8A block diagram of an exemplary head-mounted device 800 according to some specific embodiments is shown. The head-mounted device 800 includes a housing 801 (or encapsulation) that houses various components of the head-mounted device 800. The housing 801 includes (or is coupled to) eye pads (not shown) disposed at a proximal (user 25) end of the housing 801. In various specific embodiments, the eye pads are plastic or rubber components that comfortably and snugly hold the head-mounted device 800 in a proper position on the face of the user 25 (e.g., around the eyes of the user 25).
[0113] Housing 801 houses display 810, which displays images, emits light toward or onto the eyes of user 25. In various embodiments, display 810 emits light through an eyepiece having one or more lenses 805 that refract the light emitted by display 810, causing the display to appear to user 25 at a virtual distance greater than the actual distance from the eyes to display 810. In order for user 25 to focus on display 810, in various embodiments, the virtual distance is at least greater than the minimum focal length of the eye (e.g., 8 cm). Furthermore, to provide a better user experience, in various embodiments, the virtual distance is greater than 1 meter.
[0114] The housing 801 also houses a tracking system comprising one or more light sources 822, a camera 824, and a controller 880. The one or more light sources 822 emit light onto the eyes of user 25, which is reflected as a light pattern (e.g., a flash) detectable by the camera 824. Based on this light pattern, the controller 880 can determine the eye-tracking characteristics of user 25. For example, the controller 880 can determine the gaze direction and / or blinking state (open or closed eyes) of user 25. Also, the controller 880 can determine the pupil center, pupil size, or point of focus. Thus, in various embodiments, light is emitted by the one or more light sources 822, reflected from the eyes of user 25, and detected by the camera 824. In various embodiments, light from the eyes of user 25 is reflected from a heat mirror or passes through an eyepiece before reaching the camera 824.
[0115] The housing 801 also houses an audio system including one or more audio sources 826, which the controller can utilize to deliver audio to the user's ear 60 via sound waves 14, according to the techniques described herein. For example, the audio source 826 can provide sound for both background noise and auditory stimulation that can be spatially represented in a 3D coordinate system. The audio source 826 may include a speaker, a connection to an external speaker system (such as headphones), or an external speaker connected via a wireless connection.
[0116] Display 810 emits light within a first wavelength range, and the one or more light sources 822 emit light within a second wavelength range. Similarly, camera 824 detects light within the second wavelength range. In various specific embodiments, the first wavelength range is the visible wavelength range (e.g., a wavelength range of approximately 400 nm to 700 nm within the visible spectrum), and the second wavelength range is the near-infrared wavelength range (e.g., a wavelength range of approximately 700 nm to 1400 nm within the near-infrared spectrum).
[0117] In various specific implementations, eye tracking (or specifically, a defined gaze direction) is used to enable user interaction (e.g., user 25 selects an option by looking at display 810), provide punched rendering (e.g., presenting a higher resolution in the area of display 810 that user 25 is looking at and a lower resolution elsewhere on display 810), or correct distortion (e.g., for an image to be presented on display 810).
[0118] In various specific implementations, the one or more light sources 822 emit light toward the eyes of the user 25, and the light is reflected in multiple flashes.
[0119] In various implementations, camera 824 is a frame / shutter-based camera that generates images of the user 25's eye at one or more time points at a frame rate. Each image includes a matrix of pixel values corresponding to the pixels in the image, the pixels corresponding to the positions of the camera's light sensor matrix. In specific implementations, each image is used to measure or track pupil dilation by measuring changes in pixel intensity associated with one or both of the user's pupils.
[0120] In various specific implementations, camera 824 is an event camera that includes multiple light sensors (e.g., a light sensor matrix) at multiple corresponding locations, which generates an event message indicating a specific location of a particular light sensor in response to a particular light sensor detecting a change in light intensity.
[0121] It should be understood that the specific embodiments described above are cited by way of example, and this disclosure is not limited to what has been specifically shown and described above. Rather, the scope includes both combinations and sub-combinations of the various features described above, as well as variations and modifications of the various features that would occur to those skilled in the art upon reading the foregoing description and which are not disclosed in the prior art.
[0122] As described above, one aspect of the present invention is the collection and use of physiological data to improve the user experience of electronic devices in interacting with electronic content. In some cases, the collected data may include personal information. For example, such information may include data that uniquely identifies a particular person or can be used to identify an individual's interests, characteristics, or behaviors. Such information data may include physiological data, demographic data, location data, device characteristics of a personal device, or any other personal information. Such information may be used for the benefit of the user. For example, personal information data may be used to improve the interactivity and control capabilities of electronic devices. Any personal information and / or physiological data should be used in accordance with well-known privacy policies and / or privacy practices. Such policies and practices should meet or exceed industry or government information privacy and data requirements. The collection of such information should be based on user consent and should only be used for lawful and reasonable purposes. Furthermore, the collected personal information should not be used or shared outside of those lawful purposes, and reasonable measures should be taken to protect access to the information and ensure that such access is secure.
[0123] In some implementations, users can selectively block access to and / or use of their personal information. Hardware or software components may be provided to prevent or block access to such information. For example, a system may be configured to allow users to opt in or out of the collection of personal information. In another example, users may opt out of providing personal information for a specific purpose, such as targeted content delivery.
[0124] While this disclosure broadly covers the use of personal information, various specific implementations can be carried out without access to such personal information. These implementations will not be rendered inoperable without all or part of such personal information. For example, preferences or settings can be inferred from non-personal information data or a minimal amount of personal information, such as content requested by a user's associated device, other non-personal information available to the content delivery service, or publicly available information, thereby selecting content and delivering it to the user.
[0125] In some implementations, data is stored in a manner that allows access only to the data's owner. For example, a public / private key system can be used to encrypt data such as personal information. In other implementations, data can be stored anonymously (e.g., without needing to identify personal information about the user, such as legal name, username, time, and location data). This makes it impossible for others to determine the identity of the user associated with the stored data.
[0126] This document sets forth numerous specific details to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter can be practiced without these specific details. In other instances, methods, apparatus, or systems known to a person of ordinary skill have not been described in detail so as not to obscure the claimed subject matter.
[0127] Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as “processing,” “calculating,” “computing,” “determining,” and “identifying” refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, which manipulate or convert data representing physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of a computing platform.
[0128] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific embodiments of the subject matter of this invention. The teachings contained herein can be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.
[0129] Specific implementations of the methods disclosed herein can be performed in the operation of such computing devices. The order of the boxes presented in the above examples can be changed; for example, the boxes can be reordered, grouped, or divided into sub-boxes. Some boxes or procedures can be executed in parallel.
[0130] The use of “applies to” or “configured to” in this document implies open and inclusive language, which does not exclude applicability to or configuration to devices performing additional tasks or steps. Similarly, the use of “based on” implies openness and inclusivity, as processes, steps, calculations, or other actions “based on” one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbering included in this document are for illustrative purposes only and are not intended to be restrictive.
[0131] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various objects, these objects should not be limited by these terms. These terms are merely used to distinguish one object from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.
[0132] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and the appended claims, the singular forms “a” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “comprising” or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, objects, or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, objects, components, or groups thereof.
[0133] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when the prerequisite is detected" or "in response to detection" that the prerequisite is true, depending on the context.
[0134] The foregoing description and summary of the present invention should be understood as illustrative and exemplary in every respect, and not restrictive, and the scope of the invention disclosed herein is determined not only by the detailed description of the illustrative specific embodiments, but also by the full extent permitted by patent law. It should be understood that the specific embodiments shown and described herein are merely illustrative of the principles of the invention, and various modifications can be made by those skilled in the art without departing from the scope and spirit of the invention.
Claims
1. A method for evaluating attentional state, the method comprising: In devices that include processors and speakers: Determine the type of the user's environment based on sensor data; Auditory stimuli are selected based on choosing a sound from a plurality of sounds that corresponds to a specific type of environment, wherein the auditory stimuli are selected based on a virtual position in a 3D coordinate system associated with the environment and are selected to both elicit a response for evaluating the user’s attentional state and be consistent with the environment, and wherein the spatial position and timing of the auditory stimuli are controlled to match the environment. The auditory stimulus is presented to the user via the speaker at the virtual location located in the 3D coordinate system. Using sensors to obtain first physiological data associated with the user's physiological response to the auditory stimulus, wherein obtaining the first physiological data includes determining that the user's physiological response to the auditory stimulus is consistent with the virtual location; and The user's attention state is evaluated based on the user's physiological response to the auditory stimulus.
2. The method of claim 1, wherein selecting the auditory stimulus based on selecting a sound corresponding to a determined type of environment comprises: The environment is classified into a certain environment type; as well as The auditory stimulus is selected based on the environment type.
3. The method of claim 1, further comprising selecting the auditory stimulus based on selecting a sound corresponding to a determined type of object within the environment, including: To classify one or more objects in the environment; as well as The auditory stimulus is selected based on one or more objects classified.
4. The method of claim 1, wherein selecting the auditory stimulus comprises: Identify one or more auditory stimuli from an auditory stimulus database that elicit a response used to evaluate the user's attention; as well as The auditory stimulus is selected from one or more auditory stimuli based on the environment.
5. The method of claim 1, wherein the environment is the physical environment surrounding the user.
6. The method of claim 1, wherein the environment is an extended reality (XR) experience presented to the user.
7. The method according to claim 1, wherein the auditory stimulus is a discrete sound, a series of sounds, or a spatialized sound.
8. The method of claim 1, wherein obtaining the first physiological data associated with the user's physiological response to the auditory stimulus includes monitoring a response or lack of response occurring within a predetermined time period after the presentation of the auditory stimulus.
9. The method of claim 1, wherein the first physiological data includes electroencephalogram (EEG) amplitude data associated with the user.
10. The method of claim 1, wherein the first physiological data includes pupil movement associated with the user.
11. The method of claim 1, wherein statistical or machine learning-based classification techniques are used to evaluate the attention state.
12. The method of claim 1, further comprising providing a notification to the user based on the attention state.
13. The method of claim 1, further comprising identifying a portion of the content associated with the attention state.
14. The method of claim 1, further comprising customizing content based on the user's attention state.
15. The method of claim 1, further comprising aggregating attention states determined for multiple users viewing the content to provide feedback on the content.
16. The method of claim 1, wherein the device is a head-mounted device (HMD) and the environment includes an extended reality (XR) environment.
17. An apparatus, the apparatus comprising: speaker; Non-transitory computer-readable storage medium; and One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the device to perform operations including: Determine the type of the user's environment based on sensor data; Auditory stimuli are selected based on choosing a sound from a plurality of sounds that corresponds to a specific type of environment, wherein the auditory stimuli are selected based on a virtual position in a 3D coordinate system associated with the environment and are selected to both elicit a response for evaluating the user’s attentional state and be consistent with the environment, and wherein the spatial position and timing of the auditory stimuli are controlled to match the environment. The auditory stimulus is presented to the user via the speaker at the virtual location located in the 3D coordinate system. Using sensors to obtain first physiological data associated with the user's physiological response to the auditory stimulus, wherein obtaining the first physiological data includes determining that the user's physiological response to the auditory stimulus is consistent with the virtual location; and The user's attention state is evaluated based on the user's physiological response to the auditory stimulus.
18. The device of claim 17, wherein selecting the auditory stimulus based on selecting a sound corresponding to a determined type of environment comprises: The environment is classified into a certain environment type; as well as The auditory stimulus is selected based on the environment type.
19. The device of claim 17, further comprising selecting the auditory stimulus based on selecting a sound corresponding to a determined type of object within the environment, including: To classify one or more objects in the environment; as well as The auditory stimulus is selected based on one or more objects classified.
20. A non-transitory computer-readable storage medium storing program instructions executable on a device to perform operations, said operations including: Determine the type of the user's environment based on sensor data; Auditory stimuli are selected based on choosing a sound from a plurality of sounds that corresponds to a specific type of environment, wherein the auditory stimuli are selected based on a virtual position in a 3D coordinate system associated with the environment and are selected to both elicit a response for evaluating the user’s attentional state and be consistent with the environment, and wherein the spatial position and timing of the auditory stimuli are controlled to match the environment. The auditory stimulus is presented to the user via a speaker at the virtual location located in the 3D coordinate system; Using sensors to obtain first physiological data associated with the user's physiological response to the auditory stimulus, wherein obtaining the first physiological data includes determining that the user's physiological response to the auditory stimulus is consistent with the virtual location; and The user's attention state is evaluated based on the user's physiological response to the auditory stimulus.
Citation Information
Patent Citations
Human performance optimization and training methods and systems
CN107427716A