Attention tracking to reinforce focus shifts

By tracking user attention and generating focus transition markers, the method assists users in resuming virtual content engagement after interruptions, addressing the challenge of cognitive deficits in conventional display devices.

JP7765616B2Active Publication Date: 2025-11-06GOOGLE LLC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024517522
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-09-20
Filing Date
2022-09-21
Publication Date
2025-11-06
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

Conventional display devices fail to consider cognitive deficits when users are interrupted while engaging with virtual content, making it difficult for them to resume their tasks after disruptions.

Method used

A method to track user attention, detect defocus events, predict next focus events, and generate focus transition markers such as visual, audio, or haptic cues to re-engage users with the virtual content.

Benefits of technology

Minimizes cognitive transition costs by quickly guiding users back to their previous activity, reducing the time and effort required to regain focus on virtual screens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007765616000001
    Figure 0007765616000001
  • Figure 0007765616000002
    Figure 0007765616000002
  • Figure 0007765616000003
    Figure 0007765616000003
Patent Text Reader

Abstract

The system and method relate to tracking a user's attention to content presented on a virtual screen, detecting a defocus event associated with a first region of the content, and determining a next focus event associated with a second region of the content. The determination may be based at least in part on the defocus event and the tracked attention of the user. The system and method may include generating a marker for distinguishing the second region of the content from a remainder of the content based on the determined next focus event, and triggering execution of the marker associated with the second region of the content in response to detecting a refocus event associated with the virtual screen.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of and claims priority to U.S. Provisional Patent Application No. 63 / 261,449, filed September 21, 2021, entitled "ATTENTION TRACKING TO AUGMENT FOCUS TRANSITIONS," which claims priority to U.S. Non-Provisional Patent Application No. 17 / 933,631, filed September 20, 2022, entitled "ATTENTION TRACKING TO AUGMENT FOCUS TRANSITIONS," the disclosures of which are incorporated herein by reference in their entireties.

[0002] This application also claims priority to U.S. Provisional Patent Application No. 63 / 261,449, filed September 21, 2021, the disclosure of which is incorporated herein by reference in its entirety. [Background technology]

[0003] Technical Field This description relates generally to methods, devices, and algorithms used to process user attention to content.

[0004] background Display devices may enable user interaction with virtual content presented within a field of view on the display device. A user of the display device may be interrupted by interactions and / or other content. After such an interruption occurs, it may be difficult for the user to resume a task associated with the virtual content. Conventional display devices do not consider the cognitive deficits that occur when attempting to re-immerse oneself in the virtual content after encountering an interruption. Therefore, it may be beneficial to provide a mechanism for tracking tasks associated with virtual content viewed on the display device and, for example, triggering the display device to respond to the user based on such interruptions. Summary of the Invention

[0005] overview Wearable display devices can be configured as augmented reality devices that can view both the physical world and augmented content. A user wearing such a device may have their field of view disrupted by events occurring in the physical world. Such disruption can arise from external stimuli of one or more senses (e.g., auditory, visual, tactile, olfactory, gustatory, vestibular, or proprioceptive stimuli). For example, a voice call to the user of the wearable display device, a person appearing in the surroundings, a person touching the user, detecting a smell or taste, a balance shift (e.g., while moving), or discovering an obstacle. Such disruption can also arise from internal shifts, such as the user remembering or thinking about something that causes a shift in focus. Technical constraints also exist, such as screen timeouts, where power and heat can trigger intermittent display sleep or off modes. Any of these disruptions can pose cognitive challenges to the user when attempting to resume or return to a previous activity. For example, after a disruption or interruption occurs, the user may struggle to resume engagement with the content on the virtual screen of the wearable display device. The systems and methods described herein may be configured, for example, to analyze a user's focus, analyze the type of stimulus, and generate specific focus transition markers (e.g., visual content, visualization, audio content) to re-engage the user with previously analyzed content.

[0006] One or more computer systems may be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination thereof installed on the system that, during operation, causes the system to perform the actions. One or more computer programs may be configured to perform particular operations or actions by containing instructions that, when executed by a data processing device, cause the device to perform the actions.

[0007] In one general aspect, a method is proposed that includes tracking a user's attention to content presented on a virtual screen, detecting a defocus event associated with a first region of the content, and determining a next focus event associated with a second region of the content, the determination being based at least in part on the defocus event and the user's tracked attention. The method further includes generating a marker for distinguishing the second region of the content from the remainder of the content based on the determined next focus event, and triggering execution of the marker associated with the second region of the content in response to detecting a refocus event associated with the virtual screen. Executing the marker with the control may involve displaying a marking with a virtual control element on the virtual screen, where the user can act to trigger at least an action on the virtual screen. For example, the user may be able to interact with the presented content via the control of the marker.

[0008] In one example embodiment, detecting a defocus event associated with the first region of content may be electronically detecting that a user was previously focused on the first region of content and is no longer focused on the first region (e.g., as determined by the gaze trajectory of one or both of the user's eyes, and thus the user's visual focus). For example, the defocus event may be detected by determining that the user looks away from the first region.

[0009] Determining the next focus event associated with the second region of the content may be determining a second region of the content that the user is likely to view after a defocus event, and the likelihood that the user will focus on the second region, in particular the likelihood score, and / or the second region, is determined using a predictive model, to which information about the detected defocus event (e.g., the time the defocus event occurred, sounds around the user when the defocus event occurred, etc.) and information about the user's tracked attention (e.g., the duration the user focuses on a first or another region of the presented content, a certain amount of time before the defocus event, etc.) are provided as input parameters, and the predictive model outputs an indication as to which second region of the content presented on the virtual screen the user will focus on within a threshold amount of time after the defocus event, given the detected defocus event and its circumstances and previously tracked user attention to the presented content.

[0010] Detecting a refocus event may be a determination that a user refocuses on content presented on the virtual screen. The user's refocusing on content occurs after a defocus event associated with the same content. For example, the refocus event may include a detected gaze focusing on previously viewed content and / or content the user engaged with before the defocus event. Thus, the refocus event occurs an amount of time after the defocus event that is less than a threshold.

[0011] The marker for distinguishing the second region may be associated with a focus transition marker that is visually presented when the marker is executed. Thus, a corresponding marker may be generated after determining the next focus event. However, the execution of the marker on the virtual screen, for example, for visualizing the difference between the second region and the rest of the content and / or for visually, audibly, and / or tactilely highlighting the second region, and therefore, for example, activation and / or presentation, may only be triggered upon detection of a refocus event, i.e., after detecting that the user refocuses on the content presented on the virtual screen.

[0012] Implementations may include any of the following features, either alone or in combination. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium. Details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 illustrates an example of a wearable computing device for tracking user attention and generating focus transition markers to assist the user in returning to content provided by the wearable computing device, in accordance with implementations described throughout this disclosure. [Figure 2] FIG. 1 illustrates a system for modeling and tracking user attention and generating content based on user attention, according to implementations described throughout this disclosure. [Figure 3A] A figure illustrating an example of tracking a user's focus on a display of a wearable computing device and generating content based on the tracked focus, according to implementations described throughout this disclosure. [Figure 3B]A figure illustrating an example of tracking a user's focus on a display of a wearable computing device and generating content based on the tracked focus, according to implementations described throughout this disclosure. [Figure 3C] A figure illustrating an example of tracking a user's focus on a display of a wearable computing device and generating content based on the tracked focus, according to implementations described throughout this disclosure. [Figure 3D] A figure illustrating an example of tracking a user's focus on a display of a wearable computing device and generating content based on the tracked focus, according to implementations described throughout this disclosure. [Figure 4A] 10A-10C illustrate examples of user interfaces depicting focus transition markers, according to implementations described throughout this disclosure. [Figure 4B] 10A-10C illustrate examples of user interfaces depicting focus transition markers, according to implementations described throughout this disclosure. [Figure 4C] 10A-10C illustrate examples of user interfaces depicting focus transition markers, according to implementations described throughout this disclosure. [Figure 4D] 10A-10C illustrate examples of user interfaces depicting focus transition markers, according to implementations described throughout this disclosure. [Figure 5] 1 is a flowchart illustrating an example of a process for tracking user focus and generating focus-transitioning content in response to user focus, according to implementations described throughout this disclosure. [Figure 6] 1 illustrates examples of computing devices and mobile computing devices that may be used with the techniques described herein. DETAILED DESCRIPTION OF THE INVENTION

[0014] Like reference symbols in the various drawings indicate like elements. Detailed Description This disclosure describes systems and methods for generating user interface (UI) content, audio content, and / or haptic content based on modeled and / or tracked user attention (e.g., focus, gaze, and / or interaction). UI content or effects for existing display content may be generated to assist a user in regaining focus (e.g., attention) from the physical world to content on a display screen (e.g., a virtual screen of a wearable computing device). For example, when a user's attention is distracted from the virtual screen to the physical world, the systems and methods described herein can predict a refocusing event, and in response to the occurrence of the refocusing event, the system may generate and / or modify content for display on the virtual screen to help the user re-immerse themselves in the content on the virtual screen.

[0015] The UI content, audio content, and / or haptic content described herein may be described as focus transition markers, which may include any or all of UI markers, UI changes, audio outputs, haptic feedback, presented UI content augmentations, audio content augmentations, haptic event augmentations, and the like. The focus transition markers may assist a user in viewing and / or accessing content through a virtual screen, for example, to quickly transition from the physical world back to the virtual screen. In some implementations, the focus transition markers may be part of a computing interface configured for the virtual screen of a battery-powered computing device. The computing interface may analyze and / or detect certain aspects of attentional (e.g., focus) behavior associated with a user accessing the computing interface. In some implementations, certain sensors associated with the virtual screen (or the computing device housing the virtual screen) may be used to analyze and / or detect certain aspects of attentional behavior associated with the user.

[0016] In some implementations, the haptic effect may include a haptic cue to indicate changes that may have occurred during and after the defocus event, but before the detected refocus event. In some implementations, the haptic effect may be provided by a headphone driver (e.g., a bone conduction signal based on a speaker signal). In some implementations, the haptic effect may be provided by vibrations generated by components of the device. For example, the haptic effect may be provided by a linear actuator (not shown) located within the device and configured to generate a haptic signal.

[0017] In some implementations, a user's attention (e.g., focus) may be modeled and / or tracked with respect to one or more interactions with, for example, a display screen (e.g., a virtual screen on an AR device (e.g., an optical see-through wearable display device, a video see-through wearable display device)) and with the physical world. The modeled user attention may be used to determine when to display such UI content, audio content, and / or haptic content during interactions with the virtual screen.

[0018] Modeling and / or tracking a user's attention (e.g., focus) may enable the systems and methods described herein to predict actions and / or interactions the user may perform next. The user's attention and predicted actions may be used to generate one or more focus transition markers, which may represent UI content enhancements (e.g., visualizations), audio content enhancements, haptic event enhancements, and the like. Focus transition markers may be generated as specific visualization techniques to indicate differences between regions of content within a virtual screen. In some implementations, the presentation of focus transition markers may provide the advantage of minimizing the time used to switch content in portions of the virtual screen.

[0019] In some implementations, focus transition markers may serve to assist a user in switching between a focus area on a virtual screen and a focus area in the physical world. For example, if a user is wearing (or looking at) an electronic device having a virtual screen and moves away from the screen to look at the physical world, a system described herein may detect an obstruction to viewing the screen, determine the next potential task to be performed and / or viewed by the user within the screen, and generate focus transition markers to assist the user in returning to viewing the screen content.

[0020] Generating and presenting a focus transition marker to a user who has decided to look away from (and return to) the virtual screen may provide the advantage of minimizing the time and / or cognitive transition costs incurred by the user when the user is distracted and / or loses focus from the virtual screen to the physical world. The loss of focus may include a change in the user's attention. The focus transition marker may be designed to minimize the amount of time utilized to regain attention (e.g., focus, interaction, gaze, etc.) from viewing the physical world to, for example, viewing an area of ​​the virtual screen through an augmented reality (AR) wearable device.

[0021] In general, focus transition markers (e.g., content) may be generated for display and / or execution in the user interface or associated audio interface of the wearable computing device. In some implementations, the focus transition markers may include visual, audio, and / or tactile content intended to direct the user's focus to draw the user's attention to a particular visual effect, audio effect, movement effect, or other content item being displayed or executed.

[0022] As used herein, focus transition markers may include, but are not limited to, any one or more (or any combination) of UI content, UI controls, changes in UI content, audio content, haptic content, and / or haptic events, animations, content rewind and replay, spatial audio / video effects, highlight effects, de-emphasis effects, text and / or image enhancement effects, three-dimensional effects, overlay effects, and the like.

[0023] The focus-transition marker may be triggered for display (or execution) based on several signals. Example signals may include any one or more gaze signals (e.g., gaze), head-tracking signals, explicit user input, and / or remote user input (or any combination thereof). In some implementations, the systems described herein may detect specific changes in one or more of the above signals to trigger the focus-transition marker to be displayed or removed from the screen. Executing a marker with a control may involve displaying a marking with a virtual control element on a virtual screen, where a user may act to trigger at least an action on the virtual screen. For example, a user may be able to interact with presented content via the marker's control.

[0024] In some implementations, the systems and methods described herein may utilize one or more sensors onboard the wearable computing device. The sensors may detect signals that may trigger a focus transition marker to be displayed or removed from the display. Such sensor signals may also be used to determine the form and / or functionality of a particular focus transition marker. In some implementations, the sensor signals may be used as input to perform machine learning and / or other modeling tasks that may determine the form and / or functionality of a focus transition marker for particular content.

[0025] 1 is an example of a wearable computing device 100 for tracking user attention and generating focus transition markers to assist the user in returning to content provided by the wearable computing device 100, according to implementations described throughout this disclosure. In this example, the wearable computing device 100 is depicted (e.g., presented) in the form of AR smart glasses (e.g., an optical see-through wearable display device, a video see-through wearable display device). However, a battery-powered device of any form factor may be substituted for and combined with the systems and methods described herein. In some implementations, the wearable computing device 100 includes a system-on-chip (SOC) architecture (not shown) that is onboard and in communication with (or has communication access to) one or more sensors, attention inputs, machine learning (ML) models, processors, encoders, and the like.

[0026] In operation, the wearable computing device 100 may provide a view of content (e.g., content 102a, 102b, and 102c) within the virtual screen 104, as well as a view of the physical world view 106 behind the virtual screen 104. The content depicted on or behind the screen 104 may include physical content as well as augmented reality (AR) content. In some implementations, the wearable computing device 100 may be communicatively coupled to other devices, such as mobile computing devices, server computing devices, tracking devices, and the like.

[0027] The wearable computing device 100 may include one or more sensors (not shown) to detect a user's attention input 108 (e.g., focus, gaze, attention) to content (e.g., 102a, 102b, 102c) depicted on the virtual screen 104. The detected attention 108 (e.g., focus) may be used as a basis for determining which content 102a, 102b, 102c is being viewed. Any or all of the determined content 102a, 102b, 102c may be modified in some manner to generate a focus transition marker 110 (e.g., a content effect), for example, when the device 100 detects a defocus event (e.g., a change in the user's attention) and / or when it detects a defocus event and then a refocus event on the screen 104.

[0028] Attention input 108 may be generated by device 100 or provided to any number of models acquired by device 100. The models may include an attention / focus model 112, a real-time user model 114, and a content recognition model 116. Each model 112-116 may be configured to analyze specific aspects of user attention and / or content accessed by device 100 to generate specific focus transition markers that may assist the user in returning to an interrupted task or content viewing.

[0029] The attention / focus model 112 may represent a real-time model of a user's attention (e.g., focus) and / or interactions, for example, based on real-time sensor data captured by the device 100. The model 112 may detect and track user focus (e.g., attention 108) on the virtual screen 104 and / or track interactions with the virtual screen 104, as well as focus and / or attention and / or interactions with the physical world. In some implementations, the model 112 may receive or detect specific signals (e.g., eye tracking or user input) to track such focus and / or interactions. The signals may be utilized to determine which portion of the content within the virtual screen 104 is currently the user's focus.

[0030] Briefly, attention / focus model 112 may be used to determine whether a user of device 100 is gaze-gaze on user interface content in virtual screen 104, or alternatively, whether the user is gaze-gaze on the physical world beyond the user interface. For example, model 112 may analyze gaze signals provided to a user of device 100. The gaze signals may be used to generate a model of the user's attention. For example, the gaze signals may be used to generate a heat map 118 of gaze positions within a predetermined time period. Device 100 may determine an amount (e.g., a percentage of screen 104) that heat map 118 overlays on virtual screen 104 over the predetermined time period. If this amount is determined to be greater than a predetermined threshold, device 100 may determine that the map indicates that the user is paying attention to the particular user interface indicated by the heat map. If this amount is determined to be less than a predetermined threshold, device 100 may determine that the map indicates that the user is distracted from the user interface.

[0031] The content recognition model 116 may represent the current and expected user attention to specific content on the virtual display. Similar to the attention / focus model, the user's focus and / or attention may be detected and / or tracked based on gaze. The model 116 may determine (or estimate) specific movement trajectories as a basis for predicting actions the user may intend to perform on the content and / or virtual screen 104. Predictions may be provided to the real-time user model 114, which is generated, for example, to model the user's current state of attention / focus and interaction with the content. The model 116 may be used to determine and predict the state of the user interface depicting the content in order to predict tasks the user may perform upon returning focus / attention to the user interface (e.g., content).

[0032] In some implementations, the content recognition model 116 may predict motion within continuous content. For example, the content recognition model 116 may model content that includes elements that can be described as a sequence of motions. Using such content, the model 116 may determine a user's predicted focus and attention to guide the user in returning focus and attention to the content. For example, the model 116 may estimate a gaze or motion trajectory to highlight (or otherwise mark) one or more locations where the user might have looked if focus had remained on the content on the virtual screen 104. For example, when a user performs an action that triggers an object to move, the user may follow the object's trajectory with their eyes. The user looks away (i.e., loses focus on the content), but the object continues to move within its trajectory. Thus, when the user refocuses on the content, the model 116 may generate and render a focus transition marker to guide the user to the object's current location.

[0033] In some implementations, the content recognition model 116 may predict transitions in sequential actions. That is, for content having elements that can be described as a sequence of actions (e.g., triggered within a user interface), focus transition markers may be generated to highlight the results of a particular action. For example, if clicking a button exposes an icon at a certain location on the virtual screen 104, the model 116 may determine, based on the content and / or location within that sequence of actions, that the icon may be highlighted when the user refocuses from the distraction to the virtual screen. Such highlighting may guide the user's gaze toward the change.

[0034] In some implementations, the models 112-116 may function as ML models with pre-training based on user input, attention input, etc. In some implementations, the models 112-116 may function based on real-time user attention data. In some implementations, the models 112-116 may function as a combination of ML models and real-time models.

[0035] In some implementations, wearable computing device 100 may detect defocus and refocus events that may occur near a particular detected focus or attention of a user. The defocus and refocus events may trigger the use of any one or more of models 112-116 to generate focus transition markers and associated content. For example, wearable computing device 100 may detect such events and determine how to generate focus transition content to, for example, assist the user in refocusing on content within screen 104. This determination may include analyzing content items and / or user interactions to match user focus with a particular predicted next action and generate and provide updated and user-relevant content. In some implementations, the focus transition content includes zooming in on rendered content or zooming out on rendered content to assist the user in refocusing on particular content without generating additional content as a focus transition marker.

[0036] In a non-limiting example, a user may be viewing content 102a on a virtual screen 104. The wearable computing device 100 may track the user's focus on the content 102a (e.g., region 120 of the content 102a) on the screen 104. As shown in FIG. 1 , the tracked focus may be depicted as a heat map 118, with the primary focus being in the central region 120. Thus, in this example, the device 100 may determine that the user is focusing on a musical score near region 120, as shown by a composition application.

[0037] At some point, the user may be interrupted or otherwise disturbed by other content or by an external physical world view 106. The device 100 may determine that a defocus event has occurred with respect to the previously focused content 102a. When the user is interrupted (e.g., the content 102a is determined to lose focus), the device 100 may determine a next focus event associated with, for example, another region 122 of the content 102a (or other content or region). This determination may be based at least in part on the defocus event and the user's continuously tracked focus corresponding to a gaze trajectory that occurred during a predetermined time period prior to the detected defocus event.

[0038] For example, in musical score content 102a, device 100 may determine that when the user returns focus onto content 102a (or generally to virtual screen 104), the user may likely perform a next focus event or action that includes at least engaging with a composition application near region 120 (e.g., after region 120 content that may have already been viewed and / or created by the user). Device 100 may then generate at least one focus transition marker in the form of focus transition content based on the determined next focus event to distinguish second region 122 of the content from the remainder of the content shown on virtual screen 104.

[0039] In another example, if a user is looking at the sheet music content 102a and playing an instrument while following the sheet music content 102a, the user may be able to move the cursor back to another area of ​​the sheet music, allowing the user to play a section again or skip a section.

[0040] In this example, the focus-transition content includes an expanded version of a second region 122b indicating the highlight region where the next focus event should be performed based on the most recent known focus region 120b, along with a menu 124 for the next step, all of which may be presented as output content 102b. In this example, the focus-transition marker includes an expanded content region 122b to allow the user to begin editing the next phrase. The next phrase is selected based on the content recognition model 116, which determined that the user's most recent focus region was 120a, and is depicted in the output content 102b as marker 120b. The menu 124 shows the focus-transition marker (e.g., content) to allow the user to begin the next task in the most recent application used before the focus-loss event. Here, the menu includes options that allow the user to open composition mode, play music to marker 120b (e.g., bookmark), and open the most recent entry. Other options are possible. Simply put, the output content 102b shows the content 102a with focus transition markers to help the user regain focus on the content on the virtual screen 104.

[0041] As used herein, a defocus event may refer to the detection of a loss of attention to content that was previously the user's focus (e.g., attention). For example, a defocus event may include a user looking away detected by wearable computing device 100. In some implementations, a defocus event may include a user-related movement such as a head turn, head tilt, device removal from the user (e.g., removing the wearable computing device), and / or any combination thereof. In some implementations, a defocus event may include an external event detected by device 100 that triggers the user to move away from content and switch to other content, a different set of content presented on a virtual screen, and / or any combination thereof.

[0042] As used herein, a refocus event may refer to the detection of continued attention (or focus) on content that occurs after a defocus event associated with the same content. For example, a refocus event may include a detected gaze focused on previously viewed content and / or content that the user engaged with before the defocus event. Thus, a refocus event may occur after a threshold amount of time after the defocus event.

[0043] As used herein, a next focus event may refer to any next step, action, focus, or attention that a user may be expected to perform after a defocus event. The next focus event may be determined by the model described herein.

[0044] During operation, device 100 may detect attention input 108 and may use any or all of attention / focus model 112, real-time user model 114, and / or content recognition model 116 (or other user input or device 100 input) when device 100 detects 126 a change in user attention / focus. Once a change in attention / focus is determined, device 100 may trigger focus transition marker generation 128 and focus transition marker execution 130 to trigger display of output 102b on virtual screen 104 in response to detecting a refocus event associated with screen 104.

[0045] In some implementations, the systems and methods described herein may implement content-aware focus transition markers (e.g., augmentations in a user interface) to track and guide user attention (e.g., focus, interaction, gaze, etc.) within or beyond the display (e.g., virtual screen) of a wearable computing device. The content-aware focus transition markers may include UI, audio, and / or haptic content to assist the user when transitioning between two or more focus regions. The focus transition markers may include content augmentations that reduce confusion in switching between virtual world content and physical world content.

[0046] In some implementations, the systems and methods described herein can generate and use content-aware models of current and anticipated user attention (e.g., focus, interaction, gaze, etc.) to estimate gaze and / or movement trajectories to anticipate what actions the user is about to take. This information can be fed into a real-time user model of the user's state, which can be combined with the current UI model and UI state. Thus, the systems and methods described herein can predict where and what area / content the user is expected to return to after looking away for a period of time.

[0047] Although multiple models 112-116 are depicted, a single model or multiple additional models may be utilized to perform all modeling and prediction tasks on device 100. In some implementations, each model 112, 114, and 116 may represent a different algorithm or code snippet that may be executed on one or more processors, for example, onboard device 100 and / or communicatively coupled to device 100 via a mobile device.

[0048] Wearable computing device 100 is depicted as AR glasses in this example. In general, device 100 may include any or all components of systems 100 and / or 200 and / or 600. Wearable computing device 100 may also be referred to as smart glasses, which represent an optical head-mounted display device designed in the shape of glasses. For example, device 100 may represent hardware and software that can add information (e.g., project a display) along with what the wearer sees through device 100.

[0049] Although device 100 is shown as a wearable computing device as described herein, other types of computing devices are possible. For example, wearable computing device 100 may include any battery-powered device, including, but not limited to, a head-mounted display (HMD) device such as an optical head-mounted display (OHMD) device, a transparent head-up display (HUD) device, an augmented reality (AR) device, or other devices such as goggles or headsets having sensors, a display, and computing capabilities. In some examples, wearable computing device 100 may instead be a wristwatch, a mobile device, jewelry, a ring controller, or other wearable controller.

[0050] 2 illustrates a system 200 for modeling and tracking user attention and generating content based on user attention associated with wearable computing device 100, according to implementations described throughout this disclosure. In some implementations, the modeling and / or tracking may be performed on wearable computing device 100. In some implementations, the modeling and / or tracking may be shared among one or more devices. For example, the modeling and / or tracking may be completed partially on wearable computing device 100 and partially on mobile computing device 202 (e.g., a mobile companion device communicatively coupled to device 100) and / or server computing device 204. In some implementations, the modeling and / or tracking may be performed on wearable computing device 100, with output from such processing provided to mobile computing device 202 and / or server computing device 204 for further analysis.

[0051] In some implementations, the wearable computing device 100 includes one or more computing devices, at least one of which is a display device that can be worn on or near human skin. In some examples, the wearable computing device 100 is or includes one or more wearable computing device components. In some implementations, the wearable computing device 100 may include a head-mounted display (HMD) device, such as an optical head-mounted display (OHMD) device, a transparent head-up display (HUD) device, a virtual reality (VR) device, an AR device, or other devices, such as goggles or a headset, that have sensors, a display, and computing capabilities. In some implementations, the wearable computing device 100 includes AR glasses (e.g., smart glasses). AR glasses refer to optical head-mounted display devices designed in the shape of eyeglasses. In some implementations, the wearable computing device 100 is or includes a smartwatch. In some implementations, the wearable computing device 100 is or includes a piece of jewelry. In some implementations, wearable computing device 100 is or includes a ring controller device or other wearable controller. In some implementations, wearable computing device 100 is or includes earbuds / headphones or smart earbuds / headphones.

[0052] 2, system 200 includes a wearable computing device 100 that is communicatively coupled to a mobile computing device 202 and, optionally, to a server computing device 204. In some implementations, the communicative coupling may occur over a network 206. In some implementations, the communicative coupling may occur directly between the wearable computing device 100, the mobile computing device 202, and / or the server computing device 204.

[0053] The wearable computing device 100 includes one or more processors 208, which may be formed in a substrate configured to execute one or more machine-readable instructions, or software, firmware, or a combination thereof. The processors 208 may be semiconductor-based and may include semiconductor materials capable of implementing digital logic. The processors 208 may include, for example, a CPU, a GPU, and / or a DSP, to name a few.

[0054] The wearable computing device 100 may also include one or more memory devices 210. The memory devices 210 may include any type of storage device that stores information in a format that can be read and / or executed by the processor 208. The memory devices 210 may store applications 212 and modules that, when executed by the processor 208, perform certain operations. In some examples, the applications 212 and modules may be stored on an external storage device and loaded into the memory device 210.

[0055] The wearable computing device 100 includes a sensor system 214. The sensor system 214 includes one or more image sensors 216 configured to detect and / or acquire image data of displayed content and / or of content being viewed by a user. In some implementations, the sensor system 214 includes multiple image sensors 216. The image sensors 216 may capture and record images (e.g., pixels, frames, and / or portions of an image) and video.

[0056] In some implementations, the image sensor 216 is a red, green, blue (RGB) camera. In some examples, the image sensor 216 includes a pulsed laser sensor (e.g., a LiDAR sensor) and / or a depth camera. For example, the image sensor 216 may be a camera configured to detect and transmit information used to create an image. In some implementations, the image sensor 216 is an eye-tracking sensor (or camera), such as, for example, an eye gaze / gaze tracker 218 that captures the eye movements of a user accessing the device 100.

[0057] Gaze / gaze tracker 218 includes instructions stored in memory 210 that, when executed by processor 208, cause processor 208 to perform the gaze detection operations described herein. For example, gaze / gaze tracker 218 may determine a location on virtual screen 226 to which the user's gaze is directed. Gaze / gaze tracker 218 may make this determination based on identifying and tracking the location of the user's pupils in images captured by imaging devices (e.g., sensors 216 and / or cameras 224) of sensor system 214.

[0058] During operation, the image sensor 216 is configured to acquire (e.g., capture) image data (e.g., optical sensor data) continuously or periodically while the device 100 is powered on. In some implementations, the image sensor 216 is configured to operate as an always-on sensor. In some implementations, the image sensor 216 may be activated in response to detecting an attention / focus change or gaze change associated with a user. In some implementations, the image sensor 216 may track an object or area of ​​interest.

[0059] The sensor system 214 may also include an inertial motion unit (IMU) sensor 220. The IMU sensor 220 may detect movement, motion, and / or acceleration of the wearable computing device 100. The IMU sensor 220 may include a variety of different types of sensors, such as, for example, an accelerometer, a gyroscope, a magnetometer, and other such sensors. In some implementations, the sensor system 214 may include embedded-screen sensors that may detect certain user actions, movements, etc. directly from the virtual screen.

[0060] In some implementations, sensor system 214 may also include an audio sensor 222 configured to detect audio received by wearable computing device 100. Sensor system 214 may include other types of sensors, such as optical sensors, distance and / or proximity sensors, contact sensors such as capacitive sensors, timers, and / or other sensors and / or different combinations of sensors. Sensor system 214 may be used to obtain information associated with the position and / or orientation of wearable computing device 100. In some implementations, sensor system 214 also includes or has access to an audio output device (e.g., one or more speakers) that can be triggered to output audio content.

[0061] The sensor system 214 may also include a camera 224 capable of capturing still and / or video images. In some implementations, the camera 224 may be a depth camera that may collect data related to the distance of external objects from the camera 224. In some implementations, the camera 224 may be a point-tracking camera that may detect or track, for example, one or more optical markers on an external device, such as, for example, an optical marker on an input device or a finger on a screen. The input may be detected, for example, by the input detector 215.

[0062] The wearable computing device 100 includes a virtual screen 226 (e.g., a display). The virtual screen 226 may include a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting display (OLED), an electrophoretic display (EPD), or a micro-projection display employing an LED light source. In some examples, the virtual screen 226 is projected onto the user's field of view. In some examples, in the case of AR glasses, the virtual screen 226 may provide a transparent or translucent display so that a user wearing the AR glasses can see not only the image provided by the virtual screen 226 but also information located in the field of view of the AR glasses behind the projected image. In some implementations, the virtual screen 226 represents a virtual monitor that generates a screen image larger than is physically present.

[0063] Wearable computing device 100 also includes a UI renderer 228. UI renderer 228 may work in conjunction with screen 226 to render user interface objects or other content to a user of wearable computing device 100. For example, UI renderer 228 may receive images captured by device 100 and generate and render additional user interface content, such as focus transition markers 230.

[0064] The focus transition marker 230 may include any or all of content 232, UI effects 234, audio effects 236, and / or haptic effects 238. The content 232 may be provided by a UI renderer 228 and displayed on the virtual screen 226. The content 232 may be generated using a content generator 252 according to a model 248 and one or more threshold conditions 254. For example, the content 232 may be generated by the content generator 252 based on the model 248 and rendered on the virtual screen 226 by the UI renderer 228 in response to satisfying the threshold condition 254. For example, the content 232 may be generated in response to detecting a refocus event associated with the virtual screen 226 (after detecting a defocus event for a period of time). The device 100 may trigger the execution (and / or rendering) of a focus transition marker 230 (e.g., content 232) in a determined area of ​​the virtual screen 226 based on detecting a refocus event occurring a predetermined time period after the defocus event.

[0065] The content 232 may include any one or more of the new content that is rendered over existing displayed content in the screen 226. For example, the content 232 may include decorations, highlights, or font changes to existing content or text, symbols, arrows, outlines, underlines, bolding, or other emphasis on existing content. In a similar manner, the absence of a change in content may indicate to the user that no change occurred during a disruption or interruption experienced by the user (e.g., during a loss-of-focus event).

[0066] The UI effects 234 may include any one or more of animations, lighting effects, flashing content on / off, text or image movement, text or image replacement, video snippets to play content, shrinking / expanding content to highlight new or missed content, and the like. In a similar manner, the absence of a displayed UI effect may indicate to the user that no change occurred during the interruption or disruption experienced by the user (e.g., during a focus loss event). That is, a focus transition marker associated with a focus loss event may be suppressed (i.e., during the focus loss event and before the next focus event) in response to determining that no change (e.g., a substantial change) to the content depicted on the virtual screen occurred.

[0067] The audio effects 236 may include audio cues to indicate that changes may have occurred during and after the out-of-focus event, but before the detected refocus event. The audio effects 236 may include audio added to content already available for display on the virtual screen 226, such as an audio cue preset by the user to indicate that content was missed during the out-of-focus time. The preset audio cue may also be provided with visual or tactile data associated with any changes that may have occurred while the user was out of focus. In some implementations, the focus transition marker 230 includes a video or audio playback of a detected change in the presented content from the time associated with the out-of-focus event to the time associated with the refocus event. The playback of the detected change may include a visual or audio cue corresponding to the detected change. In a similar manner, the absence of an audio effect may indicate to the user that no change occurred during the interruption or disruption experienced by the user (e.g., during the out-of-focus event).

[0068] The haptic effect 238 may include a haptic cue to indicate that a change may have occurred during and after the defocus event, but before the detected refocus event. In some implementations, the haptic effect 238 may be provided by a headphone driver (e.g., a bone conduction signal based on a speaker signal). In some implementations, the haptic effect 238 may be provided by vibrations generated by components of the device 100. For example, the haptic effect 238 may be provided by a linear actuator (not shown) located within the device 100 and configured to generate a haptic signal.

[0069] Haptic effects may include haptic content, such as vibrations from speakers associated with the screen 226 or device 100, or wearable earbuds. Haptic effects 238 may also include gestures presented on the screen 226, animations indicating selections to be made or changes missed during out-of-focus time periods, playback events indicating changes missed during out-of-focus time periods, etc.

[0070] In some implementations, the focus transition marker 230 includes a video or audio playback of a detected change in the presented content from a time associated with a defocus event to a time associated with a refocus event, and the playback of the detected change includes a haptic cue (e.g., an effect) corresponding to the detected change.

[0071] In some implementations, haptic effects may provide an indication as to a change that occurred during a loss-of-focus event. For example, haptic effects 238 may include vibrations to suggest motion replays, menu presentations and selections, and / or other types of selectable actions the user can take to review missed content. In a similar manner, the absence of a haptic effect may indicate to the user that no change occurred during the interruption or disruption the user experienced (e.g., during a loss-of-focus event).

[0072] Wearable computing device 100 also includes a control system 240 that includes various control system devices to facilitate operation of wearable computing device 100. Control system 240 may utilize processor 208, sensor system 214, and / or any number of on-board CPUs, GPUS, DSPs, and the like operably coupled to components of wearable computing device 100.

[0073] Wearable computing device 100 also includes a communications module 242. Communications module 242 may enable wearable computing device 100 to communicate to exchange information with another computing device within range of device 100. For example, wearable computing device 100 may be operably coupled to another computing device via antenna 244, e.g., through a wired connection, a wireless connection such as via Wi-Fi or Bluetooth, or other type of connection, to facilitate communication.

[0074] Wearable computing device 100 may also include one or more antennas 244 configured to communicate with other computing devices via wireless signals. For example, wearable computing device 100 may receive one or more wireless signals and use the wireless signals to communicate with other devices, such as mobile computing device 202 and / or server computing device 204, or other devices within range of antenna 244. The wireless signals may be triggered via a wireless connection, such as a short-range connection (e.g., a Bluetooth connection or a near-field communication (NFC) connection) or an Internet connection (e.g., Wi-Fi or a mobile network).

[0075] Wearable computing device 100 may also include or have access to user permissions 246. Such permissions 246 may be pre-configured based on user-provided and authentication permissions for content. For example, permissions 246 may include permissions for camera, content access, habits, past input and / or behavior, and the like. A user may provide such permissions to enable device 100 to identify contextual cues about the user that may be used, for example, to anticipate the next user action. If the user has not configured permissions 246, the systems described herein may not generate focus transition markers and content.

[0076] In some implementations, the wearable computing device 100 is configured to communicate with a server computing device 204 and / or the mobile computing device 202 over a network 206. The server computing device 204 may represent one or more computing devices that may take the form of several different devices, such as a standard server, a group of such servers, or a rack server system. In some implementations, the server computing device 204 is a single system that shares components such as a processor and memory. The network 206 may include the Internet and / or other types of data networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, or other types of data networks. The network 206 may also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within the network 206.

[0077] The input detector 215 may, for example, process input received from the sensor system 214 and / or overt user input received on the virtual screen 226. The input detector 215 may detect a user's attention input (e.g., focus, gaze, attention) to content depicted on the virtual screen 226. The detected attention (e.g., focus) may be used as a basis for determining which content is being viewed. Any or all of the determined content may be modified in some manner to generate focus transition markers 230 (e.g., content effects), for example, when the device 100 detects a defocus event and / or detects a defocus event and then a refocus event on the screen 226.

[0078] Attention input may be provided to any number of models 248 generated by or acquired by device 100. Models 248 shown in FIG. 2 include, but are not limited to, attention model 112, real-time user model 114, and content recognition model 116, each of which is described in detail above. Any or all of the models may invoke or access a neural network (NN) 250 to perform machine learning tasks onboard device 100. The NN may be used in conjunction with the machine learning models and / or actions to predict and / or detect specific attentional focus and defocus, as well as to generate focus transition markers 230 (e.g., content 232, UI effects 234, audio effects 236, and / or haptic effects 238) to assist a user of device 100 in regaining focus on an area of ​​virtual screen 226.

[0079] In some implementations, processing performed on sensor data acquired by sensor system 214 is referred to as a machine learning (ML) inference operation. An inference operation may refer to an image processing operation, step, or sub-step involving an ML model that makes (or derives) one or more predictions. Certain types of processing performed by wearable computing device 100 may make predictions using ML models. For example, machine learning may use statistical algorithms that learn data from existing data to make decisions about new data, a process called inference. In other words, inference refers to the process of taking an already trained model and using the trained model to make predictions.

[0080] 3A-3D illustrate an example of tracking a user's focus on a display of a wearable computing device and generating content based on the tracked focus, according to implementations described throughout this disclosure. FIG. 3A depicts a virtual screen 302A with a physical world view 304A in the background. In some implementations, the virtual screen 302A is shown, but the physical world view 304A is not. In this example, the physical world view 304A is shown for reference, but in operation, the user may be detected as looking at content on the virtual screen 302A, and therefore the physical world view 304A may be removed from view effects, blur effects, transparency effects, or other effects to allow the user to focus on the content depicted on the screen 302A.

[0081] In this example, wearable computing device 100 may determine, via gaze / gaze tracker 218 (or other device associated with sensor system 214), that a user operating device 100 is directing their attention (e.g., focusing) on ​​content 306, shown here as a music application. Gaze / gaze tracker 218 may track the user's gaze / gaze and generate map 308 in which gazes of the highest intensity are shown as gaze targets 310. Map 308 is not generally provided for viewing by a user operating device 100 and is shown here for illustrative purposes. However, map 308 may depict a representative gaze analysis performed by sensor system 214 to determine a user's gaze (e.g., attention, focus) on content of screen 302A and / or content of material world view 304A.

[0082] Because device 100 detects that the user is focusing on content 306, device 100 may modify content 306 to place it in a more central portion of screen 302A. In some implementations, device 100 may magnify content 306 to allow the user to interact with additional content associated with content 306. Gaze target 310 may change over time as the user interacts with content 310, 312, 314, and / or other content configured for display within screen 302A. Additionally, gaze target 310 may change if the user experiences a distraction from physical world view 304A. For example, at some point, the user may shift their attention / focus to another area of ​​screen 302A or content within view 304A.

[0083] As shown in FIG. 3B , the user may have experienced a distraction from the physical world. The distraction may be visual, auditory, or tactile in nature. The distraction may trigger the user's gaze to shift to another focus or away from their original focus (e.g., gaze target). For example, as indicated by indicator 318, the user may hear another user call the user's name. The user may then shift focus from the content 306 in screen 302B, and the focus (e.g., attention, gaze target) may change, as indicated by map 316 depicting the user's gaze change. In this example, gaze / gaze tracker 218 may detect, for example, that the user is looking beyond screen 302B and may begin to trigger a change in the presentation of content on device 100. For example, device 100 may begin to transparently show physical world view 304B through screen 302B. In some implementations, device 100 may remove screen 302B from view or may blur or present other effects to allow the user to focus on the content depicted in physical view 304B.

[0084] When device 100 (e.g., gaze / gaze tracker 218) detects a change in gaze from 310 (FIG. 3A) to 316 (FIG. 3B), device 100 may turn off the display of screen 302B after a predetermined amount of time, as shown in FIG. 3C. In this example, device 100 may detect a change in user focus, as indicated by the user changing their gaze / gaze to region 320 as indicated by map 322. Additionally, device 100 may detect that no user is being spoken, as indicated by indicator 324. Indicator 324 may not be shown to the user of device 100, but instead may be determined and / or detected by device 100 and used as an input (e.g., via input detector 215) to indicate that the user from physical environment 304C has stopped speaking. This audio cue, combined with the user's change in gaze / focus / attention (e.g., to region 320), may trigger device 100 to trigger focus transition marker 230 based on the input to bring the user back to the focus position associated with the virtual screen rather than the physical world.

[0085] For example, gaze / gaze tracker 218 may detect that the user is no longer focused on physical world view 304C and may also detect that the user is attempting to refocus (e.g., or apparently access) a virtual screen, such as that shown in FIG. 3D as virtual screen 302D. In this example, device 100 may have determined the next logical item that the user may attend to, content 312, from the previous interface (302A). Such a decision to select content 312 may be based on the user previously focusing on content 310 to modify a music application, but the user's gaze has begun to shift (as shown by map 308) toward content 312, which is a map application with ongoing directions. This shift may be interpreted as a gaze trajectory that occurred during a predetermined period of time before the detected out-of-focus event.

[0086] During operation, device 100 may trigger model 248 to determine a next step or action based on a previous step or action. The trigger may occur, for example, at a time of loss of focus from screen 302A. Once device 100 determines that content 312 is a likely next gaze target or action, the system may generate focus transition marker 230 to help the user quickly focus on content 312 when focus / attention is detected returning to the virtual screen, such as shown in virtual screen 302D in FIG. 3D .

[0087] In response to detecting that the user is expected to pay immediate attention to content 312, content generator 252 may generate a focus transition marker to help the user immerse their attention / focus on content 312 on screen 302D. The focus transition marker, here, includes, for example, content 326 indicating missed directions and messages about directions that the user may have received while paying attention to 304C. Here, content 326 includes missed destination directions 328 and a missed message 330 from another user. Content 326 includes a control 332 that may allow the user to rewind time 334, in this example, 35 seconds, to view any other missed content that may have occurred in the last 35 seconds. In some implementations, such a control 332 may be provided when content 326 is video, audio, instructions, or other continuous content that the user may have been viewing on screen 302C before the interruption. In some implementations, the time 334 may be adjusted to the precise amount of time between the defocus and refocus events to allow the user to catch up and / or replay any or all missed content. In some implementations, the control 332 may be highlighted or otherwise emphasized to indicate selection, as indicated by highlight 336. Of course, other visualizations are possible.

[0088] In some implementations, the playback and / or controls 332 may provide options for playing missed content at a standard or accelerated speed. In some implementations, playback may be triggered based on refocus events and detecting that the user is focusing on a particular object, control, or area of ​​the UI. For example, a user may trigger playback of a musical score when gazing at the beginning of the score. In another example, a user may trigger a lesson summary by triggering playback after looking away from the teacher and then back at the teacher. In this manner, a user may have an instant playback feature that can be used in real time when the user missed the previous seconds or minutes of content.

[0089] In some implementations, device 100 may use model 248 to identify what the user is currently attending to (e.g., focusing on) by measuring or detecting characteristics such as user input, user presence, user proximity to content, user or content or device orientation, speech activity, gaze activity, content interaction, and the like. Model 248 may be used to infer (e.g., determine, predict) knowledge regarding priorities for managing a user's attention. This knowledge may be used to statistically model attention and other interactive behaviors performed by the user. Model 248 may determine the relevance of information or actions in the context of current user activity. The relevance of the information may be used to present any number of options and / or focus transition markers (e.g., content) that the user may actively select or passively engage with. Thus, device 100 may generate focus transition markers that may enhance data and / or content presented in the area determined to be the user's focus, while diluting other surrounding or background details.

[0090] 4A-4D illustrate example user interfaces depicting focus transition markers according to implementations described throughout this disclosure. To generate the focus transition markers, device 100 may generate and use a real-time user model 114 representing the user's attention. For example, device 100 may track the user's focus (e.g., attention) and / or interaction across a virtual screen (e.g., a user interface in an HMD or AR glasses) and the physical world (e.g., a conversation partner in a meeting). Device 100 may leverage detection signals obtained from gaze tracking or explicit user input to predict with high accuracy which portion (or region) of the screen content the user is interacting with.

[0091] As shown in FIG. 4A , a video conference is being hosted on device 100's virtual screen 402A, with a physical world view 404A shown in the background. Here, the user may be participating in a new call with several other users. At some point, the user may be interrupted, and device 100 may detect a defocus event associated with a particular region of screen 402A. During the defocus event, device 100 may determine a next focus event associated with a second region of screen 402A (or a next focus event associated with the same first region). The next focus event may be based at least in part on the defocus event and the user's tracked focus. The user's tracked focus may correspond to a gaze trajectory that occurred during a time period prior to the detected defocus event. For example, the time period may include a period of several seconds before the defocus event to the time of the defocus event.

[0092] The device 100 may generate a focus transition marker 230 (e.g., and / or content) based on the determined next focus event to distinguish the second region of the content from the remainder of the content. In some implementations, the focus transition marker 230 may be to distinguish between the first region and updates to the first region that may have occurred over time.

[0093] As shown in Figure 4B, defocus and refocus events may have occurred. Device 100 may have generated content in virtual screen 402B to act as marker 230 to help the user transition to the previous and ongoing content from Figure 4A.

[0094] In some implementations, focus transition marker 230 may be used to indicate new content within a real-time stream. For example, when a user loses focus during a live stream, transcription, translation, or video (e.g., a focus loss event), device 100 may highlight the new content that appears upon detecting a refocus event back to the live stream, etc. When a refocus event occurs, the system may distinguish the new content from the old content. Thus, the highlight may be a focus transition marker that functions to direct the user's focus / attention to useful and actionable locations within the screen. In some implementations, the highlight may gradually disappear after a short timeout. In some implementations, the transcription may also scroll up (e.g., scroll back 20 seconds) to the location where the user's attention was last detected.

[0095] In this example, while the user was out of focus, Ella joined the call at 12:32 406, triggered recording of the meeting 408, and Slavo made further remarks 410. Device 100 generated marker 230 for bolded new content 406, 408, and 410 to distinguish the new content 406, 408, and 410 from the already-seen (non-bolded) content on screen 402B. Of course, other effects may be substituted for the text decoration. In addition to decorating the text, device 100 resized the content on screen 402B to show a related set of content to make the change more apparent to a user who is refocusing. The content shown on screen 402B may be generated, executed, and triggered for display (e.g., rendering) within a UI in response to detecting a refocus event associated with virtual screen 402B (or 402A), for example.

[0096] In some implementations, after a predetermined amount of time, the live stream is again fully rendered to the screen 402C and physical world view 404C, as shown in Figure 4C, in which the bolded text is gradually shown back to the original text without the markers / content highlighting.

[0097] At some point, the user may lose focus and regain focus. FIG. 4D depicts another example of a refocus event, in which a focus transition marker 230 may be generated to assist the user. In this example, the user may have missed several minutes of content due to a disturbance and may want to view the missed items. Here, the user may not want to read the meeting updates, which are presented as marker 412 depicting a box of scrollable content 414. If device 100 detects that the amount of content may be significant and that it may be difficult for the user to find the most recent known content before the loss-of-focus event, device 100 may generate a marker, such as marker 416, that includes controls for downloading and playing audio of the missed information. The user may be presented with marker 416, and the user may select marker 416 to start the audio from the point of loss of focus.

[0098] In some implementations, device 100 may generate content (e.g., focus transition markers) for content having elements that can be described as a movement sequence. For example, device 100 may utilize sensor system 214 to detect a movement sequence and anticipate user attention based on the sequence and the state of the UI / content / user actions within the sequence. For example, a gaze or movement trajectory may be detected by sensor system 214 and used to generate focus transition marker 230 for the movement sequence. By estimating the gaze or movement trajectory, device 100 may highlight (e.g., based on predictive model 248) the area or location the user would have looked at if the user had remained focused on the virtual screen. For example, if the user performs an action that triggers an object to move, device 100 (and the user) may follow the object's trajectory. While the user loses focus (e.g., looks away), the object continues to move, and when the user regains focus (e.g., looks back), device 100 may use content, shapes, highlights, or other such markers to guide the user to the object's current location.

[0099] In some implementations, device 100 may predict transitions in sequential actions. That is, for content with elements that can be described as a sequence of actions (e.g., triggered in a user interface), focus transition markers may direct the user to the outcome of a particular action. For example, if selecting a UI button exposes an icon in an area on the virtual screen, the icon may be highlighted when the user returns focus to the virtual screen to guide the user's gaze toward the change.

[0100] In some implementations, focus transition markers may indicate differences in sequential content. For example, without performing prediction, device 100 may generate marker 230 to simply highlight what has changed on the virtual screen. In another example, if a user is watching a recording and some words are corrected while the user was out of focus (e.g., looking away), the corrected words may be highlighted as markers for display to the user when the user regains focus. In the case of spatial differences or if movement is a cue, device 100 may highlight using different visualization techniques, such as illustrating motion vectors with streamlines and the like.

[0101] 5 is a flowchart illustrating an example of a process 500 for performing image processing tasks on a computing device, according to implementations described throughout this disclosure. In some implementations, the computing device is a wearable computing device that is battery-powered. In some implementations, the computing device is a non-wearable computing device that is battery-powered.

[0102] Process 500 may utilize a processing system on a computing device having at least one processing device, a speaker, optional display capabilities, and memory storing instructions that, when executed, cause the processing device to perform multiple operations and computer-implemented steps recited in the claims. Generally, computing device 100, system 200, and / or 600 may be used in describing and performing process 500. The combination of device 100 and system 200 and / or 600 may represent a single system in some implementations. Generally, process 500 utilizes the systems and algorithms described herein to detect user attention and use it to generate and display focus transition markers (e.g., content) to assist the user in quickly focusing when returning attention / focus to the virtual screen.

[0103] At block 502, process 500 includes tracking a user's focus on content presented on the virtual screen. The tracked focus may include gaze, attention, interaction, and the like. For example, sensor system 214 and / or input detector 215 may track a user's focus and / or attention on content (e.g., content 310, 320, etc.) presented for display on virtual screen 226. In some implementations, the tracking may be performed by gaze / gaze tracker 218. In some implementations, the tracking may be performed by image sensor 216 and / or camera 224. In some implementations, the tracking may be performed by IMU 220. In some implementations, the tracking may include analyzing and tracking signals including any one or more gaze signals (e.g., gaze), head tracking signals, overt user input, and / or remote user input (or any combination thereof). In some implementations, model 248 may be used to analyze the tracked focus or related signals to determine the next step and / or the next focusing event. Any combination of the above may be used to identify and track a user's focus.

[0104] At block 504, process 500 includes detecting a defocus event associated with the first region of content. For example, sensor system 214 may detect that the user has moved their gaze away from the first region toward another region, or that the user has moved their focus or attention away from the screen to view the physical world.

[0105] At block 506, process 500 includes determining a next focus event associated with the second region of the content. This determination may be based at least in part on the defocus event and the user's tracked focus corresponding to a gaze trajectory that occurred during a predetermined time period prior to the detected defocus event. For example, gaze / gaze tracker 218 may track focus as the user engages with content on virtual screen 226 and / or physical world view 304C. The movement or motion trajectory tracked by tracker 218 may be used in combination with one or more threshold conditions 254 and / or model 248 to predict what action (e.g., event) the user may perform next. For example, real-time user model 114 and the user's state of attention / focus may be used in combination with the detected UI model and UI state to predict which event / action the user may perform next.

[0106] In some implementations, the user's focus is further determined based on detected user interaction with content (e.g., content 310). The next focus event may be determined according to the detected user interaction and a predetermined next action in a sequence of actions associated with the content and / or associated with a UI depicted on virtual screen 226. That is, device 100 may predict that if user focus returns to content 310, the next action may be to work within the composition application for the next measure in the sequence based on a focus loss event occurring for the previous measure. In some implementations, the second region may be a different portion of content than the first region if the user is predicted to change applications based on permission-based acquired knowledge about previous user behavior. In some implementations, the second region may be a different portion of content than the first region if a general movement trajectory is used by model 248 to determine the next event to occur in other content.

[0107] At block 508, process 500 includes generating a focus transition marker to distinguish the second region of the content from the remainder of the content based on the determined next focus event. For example, content generator 252 may work with sensor system 214 and model 248 to determine the configuration and / or appearance of focus transition marker 230 to place in the second region. For example, device 100 may determine whether to generate content 232, UI effects 234, sound effects 236, haptic effects 238, and / or other visualizations to distinguish the second region from other content depicted on virtual screen 226.

[0108] In some implementations, the focus transition marker 230 represents a video or audio replay of the detected change in content from a time associated with the defocus event and a time associated with the refocus event. Generally, the replay of the detected change may include a visual or audio cue corresponding to the detected change. For example, the focus transition marker 230 may be a snippet of video, audio, or generated UI transition based on content that may have been missed between the defocus event and the refocus event performed by the user and detected by the tracker 218.

[0109] In some implementations, the focus transition marker 230 represents a playback of the detected changes in content from the time associated with the defocus event and the time associated with the refocus event, and the playback of the changes includes a haptic cue corresponding to the changes. For example, the playback may include content indicating that the user should select the focus transition marker 230 to review changes that may have occurred between the defocus event and the refocus event.

[0110] At block 510, process 500 includes triggering execution of a focus transition marker in or associated with the second region in response to detecting a refocus event associated with the virtual screen. For example, device 100 may generate focus transition marker 230 but wait to execute and / or display the marker in or associated with the second region until a refocus event is detected.

[0111] In some implementations, a focus transition marker 230 associated with a particular defocus event may be configured to be suppressed in response to determining that no substantial changes to the content depicted on the virtual screen have occurred. For example, device 100 may detect that no changes (or only a minimum threshold level of changes) have occurred since the user defocus event occurred to the content depicted on virtual screen 226. In response, device 100 may not display focus transition marker 230. In some implementations, device 100 may instead soften the effect of focus transition marker 230 by dimming and then brightening the determined content that was most recently in focus before the defocus event.

[0112] In some implementations, the focus transition marker 230 includes at least one control and a highlighted indicator that is superimposed over a portion of the second region, as shown in FIG. 3D by highlight 336 and control 332. The focus transition marker may be gradually removed from the display after a predetermined period of time after the refocusing event. For example, the refocus content / marker 326 may be shown for several seconds to a minute after the user refocuses on the content of the screen 302D.

[0113] In some implementations, the content depicts a sequence of movements. For example, the content may be a sequence in a sports game involving a moving ball. A first region may represent a first sequence of movements that the user focused their attention / focus on during the game. At some point, the user may lose focus (e.g., look away) from the virtual screen 226. A second region may represent a second sequence of movements that is configured to occur after the first sequence. If the user loses focus after the first sequence of movements, the device 100 may continue to track the content (e.g., the moving ball) during the second sequence of movements.

[0114] When the user refocuses, device 100 may have generated focus transition marker 230 based on content tracking while the user was not paying attention to the game on screen 226. Transition marker 230 may depict the ball in a second region, but may also show the ball's trajectory between the first region and the second region (e.g., between and beyond a first sequence of movements and a second sequence of movements) to quickly show the user what occurred in the game during the time between the defocus and refocus events. Transition marker 230 may be triggered for execution (and / or rendering) at a location associated with the second region of content. In some implementations, triggering the execution of focus transition marker 230 associated with the second region includes triggering the second sequence to execute with focus transition marker 230 depicted in the second sequence of movements.

[0115] In some implementations, content is presented within a user interface (e.g., as illustrated by content 310, 312, and 314) of virtual screen 302A. Sensor system 214 may detect user focus / attention and generate a model of the user's attention based on the user's tracked focus over a first period of time engaged with the UI depicted on screen 302A. Device 100 may then acquire a model for the UI. The model for the UI may define multiple states of the UI and user interactions associated with the UI. Based on the user's tracked focus, the model of attention, and the determined state from the multiple states of the UI, device 100 may trigger rendering at least one focus transition marker (e.g., marker 326) overlaid on at least a portion of a second region of screen 302D and the content therein over a second period of time.

[0116] In some implementations, the virtual screen 226 is associated with an augmented reality device (e.g., wearable computing device 100) configured to provide a field of view that includes an augmented reality view (e.g., screen 302A) and a physical world view (e.g., 304A). In some implementations, tracking the user's focus includes determining whether the user's focus is a detected gaze associated with the augmented reality view (e.g., screen 302A) or a detected gaze associated with the physical world view (e.g., view 304A).

[0117] In some implementations, triggering the execution of a focus transition marker associated with the second region includes resuming the content from the time associated with the defocus event (e.g., performing sequential playback of events / audio / visuals associated with the content) when the user's focus is a detected gaze associated with the augmented reality view (e.g., 302A), and pausing and fading the content until a refocus event is detected when the user's focus is a detected gaze associated with the physical world view (e.g., 304A).

[0118] 6 illustrates an example of a computer device 600 and a mobile computer device 650 that may be used with the techniques described herein (e.g., to implement a client computing device 600, a server computing device 204, and / or a mobile computing device 202). Computing device 600 includes a processor 602, a memory 604, a storage device 606, a high-speed interface 608 that connects to memory 604 and a high-speed expansion port 610, and a low-speed interface 612 that connects to a low-speed bus 614 and storage device 606. Each of components 602, 604, 606, 608, 610, and 612 may be interconnected using various buses and mounted on a common motherboard or in any other suitable manner. Processor 602 may process instructions for execution within computing device 600, such as instructions stored in memory 604 or on storage device 606, to display graphical information for a GUI on an external input / output device, such as a display 616, coupled to high-speed interface 608. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and memory types, where appropriate. Also, multiple computing devices 600 may be connected, each providing a portion of the required operations (e.g., as a bank of servers, a group of blade servers, or a multiprocessor system).

[0119] The memory 604 stores information within the computing device 600. In one implementation, the memory 604 is a volatile memory unit(s). In another implementation, the memory 604 is a non-volatile memory unit(s). The memory 604 may also be another form of computer-readable medium, such as a magnetic or optical disk.

[0120] The storage device 606 can provide mass storage for the computing device 600. In one implementation, the storage device 606 can be or include a computer-readable medium such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including devices in a storage area network or other configuration. A computer program product can be tangibly embodied on an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 604, the storage device 606, or the memory on the processor 602.

[0121] The high-speed controller 608 manages bandwidth-intensive operations for the computing device 600, while the low-speed controller 612 manages less bandwidth-intensive operations. Such an allocation of functionality is merely exemplary. In one implementation, the high-speed controller 608 is coupled to the memory 604, the display 616 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 610 that may accept various expansion cards (not shown). In an implementation, the low-speed controller 612 is coupled to the storage device 606 and the low-speed expansion port 614. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, or networking device such as a switch or router, for example, through a network adapter.

[0122] Computing device 600, as shown in the figure, may be implemented in several different forms. For example, it may be implemented as a standard server 620, or multiple times within a group of such servers. It may also be implemented as part of a rack server system 624. Additionally, it may be implemented in a personal computer, such as a laptop computer 622. Alternatively, components from computing device 600 may be combined with other components in a mobile device (not shown), such as device 650. Each such device may include one or more computing devices 600, 650, and the entire system may be made up of multiple computing devices 600, 650 in communication with each other.

[0123] Computing device 650 includes, among other components, a processor 652, memory 664, an input / output device such as a display 654, a communication interface 666, and a transceiver 668. Device 650 may also be provided with a storage device such as a microdrive or other device to provide additional storage. Each of components 650, 652, 664, 654, 666, and 668 are interconnected using various buses, and some of the components may be mounted on a common motherboard or in other suitable manner.

[0124] The processor 652 may execute instructions in the computing device 650, such as instructions stored in the memory 664. The processor may be implemented as a chipset of chips including separate and multiple analog and digital processors. For example, the processor may provide coordination of other components of the device 650, such as controls for a user interface, applications run by the device 650, and wireless communications by the device 650.

[0125] The processor 652 may communicate with a user through a control interface 658 and a display interface 656 coupled to a display 654. The display 654 may be, for example, a TFT LCD (thin film transistor liquid crystal display), an LED (light emitting diode) or an OLED (organic light emitting diode) display, or other suitable display technology. The display interface 656 may include appropriate circuitry for driving the display 654 to present graphical and other information to the user. The control interface 658 may receive commands from the user and convert them for submission to the processor 652. Additionally, an external interface 662 may be provided in communication with the processor 652 to enable near-field communication of the device 650 with other devices. For example, the external interface 662 may provide for wired communication in some implementations or wireless communication in other implementations, and multiple interfaces may also be used.

[0126] Memory 664 stores information within computing device 650. Memory 664 may be implemented as one or more of a computer-readable medium(s), a volatile memory unit(s), or a non-volatile memory unit(s). Expansion memory 674 may also be provided and connected to device 650 through expansion interface 672, which may include, for example, a SIMM (Single In-Line Memory Module) card interface. Such expansion memory 674 may provide additional storage space for device 650 or may also store applications or other information for device 650. Specifically, expansion memory 674 may include instructions to perform or supplement the processes described above and may also include secure information. Thus, for example, expansion memory 674 may be provided as a security module for device 650 and may be programmed with instructions to enable secure use of device 650. Additionally, secure applications can be provided via the SIMM card along with additional information, such as placing identifying information on the SIMM card in an unhackable manner.

[0127] The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one implementation, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is, for example, a computer- or machine-readable medium, such as memory 664, expansion memory 674, or memory on processor 652, which may be received via transceiver 668 or external interface 662.

[0128] Device 650 may communicate wirelessly through communication interface 666, which may include digital signal processing circuitry if necessary. Communication interface 666 may provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communication may occur, for example, through radio frequency transceiver 668. In addition, short-range communication may occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, GPS (Global Positioning System) receiver module 670 may provide additional navigation-related and location-related wireless data to device 650, which may be used as appropriate by applications executing on device 650.

[0129] Device 650 may also communicate vocally using audio codec 660, which may receive speech information from a user and convert it into usable digital information. Audio codec 660 may also generate audible sounds for the user, such as through a speaker in a headset of device 650. Such sounds may include sounds from a telephone voice call, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications running on device 650.

[0130] The computing device 650 may be implemented in several different forms, as shown in the figure. For example, it may be implemented as a mobile phone 680. It may also be implemented as part of a smartphone 682, personal digital assistant, or other similar mobile device.

[0131] Various implementations of the systems and techniques described herein may be realized in digital electronic circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0132] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0133] To provide for interaction with a user, the systems and techniques described herein may be implemented on a computer that has a display device (such as an LED (light-emitting diode), or OLED (organic LED), or LCD (liquid crystal display) monitor / screen) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices may be used to provide for interaction with a user as well; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and input from the user may be received in any form, including audio, speech, or tactile input.

[0134] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as data servers), or includes middleware components (e.g., as application servers), or includes front-end components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.

[0135] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0136] In some implementations, the computing device depicted in the figure may include sensors that interface with the AR headset / HMD device 690 to generate an augmented environment for viewing inserted content in a physical space. For example, one or more sensors included on the computing device 650 depicted in the figure or another computing device may provide input to the AR headset 690 or generally provide input to the AR space. The sensors may include, but are not limited to, a touchscreen, an accelerometer, a gyroscope, a pressure sensor, a biometric sensor, a temperature sensor, a humidity sensor, and an ambient light sensor. The computing device 650 may use the sensors to determine the absolute position and / or detected rotation of the computing device in the AR space, which can then be used as input to the AR space. For example, the computing device 650 may be incorporated into the AR space as a virtual object such as a controller, a laser pointer, a keyboard, a weapon, etc. The user's positioning of the computing device / virtual object when incorporated into the AR space may allow the user to position the computing device to view the virtual object in a certain manner in the AR space.

[0137] In some implementations, the AR headset / HMD device 690 represents the device 600 and includes a display device (e.g., a virtual screen) that may include a see-through near-eye display, such as one that uses a water basin or wave-guiding optics. For example, such an optical design may project light from a display source onto a portion of teleprompter glass that acts as a beam splitter positioned at a 45-degree angle. The beam splitter may allow for reflection and transmission values, where light from the display source is partially reflected while the remaining light is transmitted. Such an optical design may allow a user to see both real-world physical items next to digital images generated by the display (e.g., UI elements, virtual content, focus transition markers, etc.). In some implementations, wave-guiding optics may be used to render content on the virtual screen of the device 600.

[0138] In some implementations, one or more input devices included on or connected to the computing device 650 can be used as input to the AR space. The input devices can include, but are not limited to, a touchscreen, a keyboard, one or more buttons, a trackpad, a touchpad, a pointing device, a mouse, a trackball, a joystick, a camera, a microphone, an earphone or earbuds with input capabilities, a game controller, or other connectable input devices. When the computing device is integrated into the AR space, a user interacting with an input device included on the computing device 650 can cause certain actions to occur within the AR space.

[0139] In some implementations, the touchscreen of the computing device 650 can be rendered as a touchpad in the AR space. A user can interact with the touchscreen of the computing device 650. The interaction is rendered in the AR headset 690, for example, as movement on the rendered touchpad in the AR space. The rendered movement can control a virtual object in the AR space.

[0140] In some implementations, one or more output devices included on the computing device 650 may provide output and / or feedback to a user of the AR headset 690 within the AR space. The output and feedback may be visual, tactile, or audio. The output and / or feedback may include, but is not limited to, vibration, turning one or more lights or strobes on and off or blinking and / or flashing, sounding an alarm, playing a chime, playing a song, and playing an audio file. Output devices may include, but are not limited to, vibration motors, vibration coils, piezoelectric devices, electrostatic devices, light-emitting diodes (LEDs), strobes, and speakers.

[0141] In some implementations, the computing device 650 may appear as a separate object in the computer-generated 3D environment. User interactions with the computing device 650 (e.g., rotating, shaking, touching, or swiping a finger across the touchscreen) may be interpreted as interactions with an object in the AR space. In the example of a laser pointer in the AR space, the computing device 650 appears as a virtual laser pointer in the computer-generated 3D environment. As the user manipulates the computing device 650, the user in the AR space sees the laser pointer moving. The user receives feedback from their interaction with the computing device 650 within the AR environment on the computing device 650 or on the AR headset 690. The user's interaction with the computing device may be translated into interactions with a user interface generated within the AR environment for the controllable device.

[0142] In some implementations, the computing device 650 may include a touchscreen. For example, a user may interact with the touchscreen to interact with a user interface for the controllable device. For example, the touchscreen may include user interface elements such as sliders that may control properties of the controllable device.

[0143] Computing device 600 is intended to represent various forms of digital computers and devices, including, but not limited to, laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 650 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are intended to be examples only and are not intended to limit the implementation forms of the subject matter described and / or claimed herein.

[0144] Although several embodiments have been described, various modifications may be made without departing from the spirit and scope of the present invention.

[0145] Additionally, the logic flow depicted in the figures does not require the particular order or sequential numbering shown to achieve desirable results. Additionally, other steps may be provided or steps may be excluded from the described flow, and other components may be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the following claims.

[0146] In addition to the above, a user may be provided with controls that allow the user to make choices regarding whether and when the systems, programs, or features described herein may enable collection of user information (e.g., information about the user's social networks, social actions or activities, occupation, user preferences, or the user's current location) and whether content or communications are sent to the user from the server. Additionally, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, a user's identity may be processed so that personally identifiable information cannot be determined about the user, or a user's geographic location may be generalized (e.g., to the city, zip code, or state level) if location information is obtained so that the user's specific location cannot be determined. Thus, a user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.

[0147] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover all such modifications and changes as fall within the scope of the present implementations. It is to be understood that they have been presented by way of example only, and not limitation, and that various changes in form and detail may be made. Any portions of the apparatus and / or methods described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein may include various combinations and / or subcombinations of the functions, components, and / or features of the different implementations described.

Claims

1. Identifying a user's attention to content presented on a virtual screen; Detecting a defocus event associated with a first region of the content; determining a next focus event associated with a second region of the content; the out-of-focus event includes the user's attention being taken away from the virtual screen; the next focus event is a predicted action for the user after the defocus event; the second region is a portion of the content that is different from the first region; The determination is based on the out-of-focus event and the identified attention of the user, and the method further comprises: generating a marker to distinguish the second region of the content from the remainder of the content based on the determined next focus event; triggering execution of the marker associated with the second region of the content in response to detecting a refocus event associated with the virtual screen; and further comprising The method, wherein the refocus event includes the user's attention returning to the content after the defocus event.

2. the marker includes at least one control and a highlighted indicator overlaid on a portion of the second region; The method of claim 1 , wherein the marker is gradually removed from the display according to a predetermined time period after the refocusing event.

3. the attention of the user is further identified based on detected user interaction with the content; The method of claim 1 or claim 2, wherein the next focus event is further determined according to the detected user interaction and a predetermined next action in a sequence of actions associated with the content.

4. the content depicts a sequence of movements; the first region represents a first sequence of movements; the second region represents a second sequence of movements configured to occur after the first sequence of movements; 3. The method of claim 1 or claim 2, wherein triggering the execution of the marker associated with the second region of the content includes triggering the execution of the second sequence of movements using the marker depicted in the second sequence of movements.

5. 3. The method of claim 1 or claim 2, wherein the marker associated with the defocus event is suppressed in response to determining that no substantial change to the content depicted on the virtual screen has occurred.

6. 3. The method of claim 1 or claim 2, wherein the markers represent a replay of a detected change in the content from a time associated with the defocus event and from a time associated with the refocus event, the replay of the change including a tactile cue corresponding to the change.

7. The content is presented within a user interface on the virtual screen, the method comprising: generating a model of the user's attention based on the identified attention of the user over a first period of time; and obtaining a model for the user interface, the model defining a plurality of states and interactions associated with the user interface, the method further comprising:

3. The method of claim 1, further comprising triggering the rendering of at least one additional marker superimposed over at least a portion of the second region for a second time period based on the identified attention of the user, the model of the attention, and the determined state of the user interface from the plurality of states.

8. the virtual screen is associated with an augmented reality device configured to provide a field of view including an augmented reality view and a physical world view; identifying the attention of the user includes determining whether the attention of the user is associated with the augmented reality view or the physical world view; Triggering the execution of the marker associated with the second region comprises: resuming the content from a time associated with the out-of-focus event if the attention of the user is associated with the augmented reality view; or The method of claim 1 or claim 2, comprising pausing and fading the content when the user's attention is associated with the physical world view until the refocus event is detected.

9. at least one processing device; at least one sensor; a memory storing instructions that, when executed, the at least one sensor identifying a user's focus on a virtual screen; the at least one sensor detecting a defocus event associated with a first region of the virtual screen; the at least one sensor determining a next focus event associated with a second region of the virtual screen; the out-of-focus event includes the user losing focus from the virtual screen; the next focus event is a predicted action for the user after the defocus event; the second area is a portion of the virtual screen that is different from the first area; the determination is based on the out-of-focus event and the identified focus of the user; triggering generation of a focus transition marker to distinguish the second region of the virtual screen from the remainder of the virtual screen based on the determined next focus event; in response to the at least one sensor detecting a refocus event associated with the virtual screen, triggering execution of the focus transition marker associated with the second region of the virtual screen; The wearable computing device, wherein the refocus event includes the user returning the focus to the virtual screen after the defocus event.

10. the focus transition marker includes at least one control and a highlighted indicator overlaid on a portion of the second region; The wearable computing device of claim 9 , wherein the focus transition markers are progressively removed from the display according to a time period after the refocusing event.

11. the focus of the user is further determined based on a detected user interaction with the virtual screen; 11. The wearable computing device of claim 9 or claim 10, wherein the next focus event is further determined according to the detected user interaction and a predetermined next action in a sequence of actions associated with the virtual screen.

12. the focus of the user on the virtual screen corresponds to the focus of the user on content depicted on the virtual screen; the content depicts a sequence of movements; the first region represents a first sequence of movements; the second region represents a second sequence of movements configured to occur after the first sequence of movements; 11. The wearable computing device of claim 9 or claim 10, wherein triggering execution of the focus transition marker associated with the second region comprises triggering execution of a second sequence with the focus transition marker depicted in the second sequence of movements.

13. 13. The wearable computing device of claim 12, wherein the focus transition marker represents a video or audio playback of a detected change in the content from a time associated with the defocus event and a time associated with the refocus event, the playback of the detected change including a visual or audio cue corresponding to the detected change.

14. 13. The wearable computing device of claim 12, wherein the focus transition marker represents a replay of a detected change in the content from a time associated with the defocus event and a time associated with the refocus event, the replay of the detected change including a tactile cue corresponding to the detected change.

15. The content is presented within a user interface on the virtual screen, and the action is: generating a model of the user's attention based on the identified focus of the user over a first period of time; obtaining a model for the user interface, the model defining a plurality of states and interactions associated with the user interface; and triggering the rendering of at least one focus transition marker superimposed on at least a portion of the second region for a second time period based on the identified focus of the user, the model of attention, and the determined state of the user interface from the plurality of states.

16. the virtual screen is associated with an augmented reality device configured to provide a field of view including an augmented reality view and a physical world view; 11. The wearable computing device of claim 9 or claim 10, wherein identifying the focus of the user includes determining whether the focus of the user is associated with the augmented reality view or the physical world view.

17. 1. A program comprising instructions that, when executed by processing circuitry of a wearable computing device, Identifying a user's focus on content presented on a virtual screen; Detecting a defocus event associated with a first region of the content; determining a next focus event associated with a second region of the content; and the out-of-focus event includes the user's focus moving away from the virtual screen; the next focus event is a predicted action for the user after the defocus event; the second region is a portion of the content that is different from the first region; the determination is based on the out-of-focus event and the identified focus of the user; generating a focus transition marker to distinguish the second region of the content from the remainder of the content based on the determined next focus event; triggering execution of the focus transition marker associated with the second region of the content in response to detecting a refocus event associated with the virtual screen; The refocus event includes the user returning the focus to the content after the defocus event.

18. the focus transition marker includes at least one control and a highlighted indicator overlaid on a portion of the second region; 18. The program of claim 17, wherein the focus transition markers are progressively removed from the display according to a time period after the refocus event.

19. the focus of the user is further determined based on detected user interaction with the content; 19. The program of claim 17 or claim 18, wherein the next focus event is further determined according to the detected user interaction and a predetermined next action in a sequence of actions associated with the content.

20. 19. The program of claim 17 or claim 18, wherein the focus transition marker represents a video or audio playback of a detected change in the content from a time associated with the defocus event and a time associated with the refocus event, and the playback of the detected change includes a visual or audio cue corresponding to the detected change.

Citation Information

Patent Citations

  • Low temperature rapid dyeing of polyamide coating and films

    JP1977017585A

  • Microprogram controller

    JP1985015743A

  • JP1987038381A

  • Read help image display device

    JP2003345335A

  • Display control system and display control program

    JP2018063322A