A head-mounted device for tracking screen time
A head-mounted device with eye tracking and neural network analysis effectively tracks screen time across multiple devices, addressing measurement challenges and optimizing power usage through adaptive modes.
Patent Information
- Application Number
- JP2024568289
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-16
- Filing Date
- 2023-05-05
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2043-05-05
AI Technical Summary
Existing technologies face challenges in accurately measuring and aggregating screen time across multiple devices, due to differences in operating systems and limited capabilities in older devices.
A head-mounted device equipped with an eye tracking camera and a processor that captures eye images to identify screen reflections, uses neural networks for analysis, and switches between high power, low power, and camera-less modes to track screen time efficiently while managing power consumption.
The solution enables accurate tracking of screen time across multiple devices, provides alerts for excessive screen use, and optimizes power usage by adapting to different modes based on battery levels.
Smart Images

Figure 2025517367000001_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of, and claims the benefit of, U.S. Patent Application No. 17 / 663,444, filed May 16, 2022, the disclosure of which is incorporated herein by reference in its entirety.
[0002] The present disclosure relates to head-mounted devices, and more particularly, to head-mounted devices configured to measure screen time across multiple devices. [Background technology]
[0003] Screen time is the amount of time a person spends using a device screen. Excessive screen time can have negative effects on sleep, physical health, brain development, and / or behavior. Summary of the Invention
[0004] In some aspects, the technologies described herein relate to a head mounted device comprising an eye tracking camera pointed at a user's eye and a processor in communication with the eye tracking camera, the processor being configured with instructions to capture an eye image of the user's eye using the eye tracking camera, analyze the eye image to identify a screen reflex, start a screen timer after the screen reflex is identified, periodically capture subsequent eye images to track the screen reflex over time, stop the screen timer when the screen reflex can no longer be tracked, and record a first instance of screen time based on the screen timer.
[0005] In some aspects, the techniques described herein relate to a head mounted device, and analyzing the eye images to identify screen reflections is performed by a neural network.
[0006] In some aspects, the techniques described herein relate to a head-mounted device, and the processor is further configured to classify the screen time of the first instance as relating to the first device based on a machine learning model configured with device information captured by the head-mounted device, and record the screen time of the first instance of the first device in a database.
[0007] In some aspects, the technology described herein relates to a head mounted device, and the device information is any combination of a device identifier, the content of the screen reflection, the characteristics of the screen reflection, and the position of the user.
[0008] In some aspects, the techniques described herein relate to a head mounted device, wherein the processor is further configured, by the instructions, to detect that a mode of the head mounted device for screen time tracking is a low power mode.
[0009] In some aspects, the techniques described herein relate to a head mounted device further comprising a front-facing camera, wherein the processor is further configured to detect that a mode of the head mounted device for screen time tracking is a high power mode, capture a field of view image of a user's field of view using the front-facing camera, analyze the field of view image to identify a screen, start a screen timer after the screen is identified in the field of view image, periodically capture subsequent field of view images to track the screen over time, stop the screen timer when the screen can no longer be tracked, and record screen time of a second instance based on the screen timer.
[0010] In some aspects, the techniques described herein relate to a head mounted device further comprising at least one orientation sensor configured to sense orientation data of the head mounted device and at least one position sensor configured to sense relative position data of the device, wherein the processor is further configured to detect that a mode of the head mounted device for screen time tracking is a camera-less mode for screen time tracking, wherein in the camera-less mode, the processor is configured to capture head mounted device orientation data and device relative position data, analyze the orientation data and the relative position data to identify a viewing state, start a screen timer after the viewing state is identified, periodically capture subsequent orientation data and subsequent relative position data to track the viewing state over time, stop the screen timer when the viewing state can no longer be tracked, and record a third instance of screen time based on the screen timer.
[0011] In some aspects, the technology described herein relates to a head-mounted device, wherein the at least one orientation sensor includes an inertial measurement unit (IMU) and the at least one position sensor includes an ultra-wideband (UWB) sensor.
[0012] In some aspects, the techniques described herein relate to a head mounted device, wherein the processor is further configured to generate an alert based on the screen time of the first instance, the screen time of the second instance, and / or the screen time of the third instance, and display the alert on a display of the head mounted device.
[0013] In some aspects, the techniques described herein relate to a computer-implemented method that includes detecting that a mode of a head mounted device for screen time tracking is a low power mode, capturing an eye image of an eye using an eye tracking camera of the head mounted device, analyzing the eye image to identify a screen reflection, starting a screen timer after the screen reflection is identified, periodically capturing subsequent eye images to track the screen reflection over time, and stopping the screen timer when the screen reflection can no longer be tracked.
[0014] In some aspects, the techniques described herein are computer-implemented and include periodically capturing subsequent eye images to track screen reflection over time, repeating the capturing and analyzing at a cycle period greater than 2 seconds.
[0015] In some aspects, the techniques described herein relate to a computer-implemented method, where detecting that a mode of a head mounted device for screen time tracking is a low power mode includes detecting a battery level of the head mounted device and determining that the battery level is below a threshold.
[0016] In some aspects, the techniques described herein relate to a computer-implemented method, where analyzing an eye image to identify a screen reflection includes applying an eye image, including an eye and a screen reflection, to an input of a neural network, and receiving a screen image, including the screen reflection, at an output of the neural network.
[0017] In some aspects, the techniques described herein relate to a computer-implemented method, further including detecting that a mode of a head mounted device for screen time tracking is a high power mode, capturing a field of view image using a front camera of the head mounted device, analyzing the field of view image to identify a screen, starting a screen timer after the screen is identified, periodically capturing subsequent field of view images to track the screen over time, and stopping the screen timer when the screen can no longer be tracked.
[0018] In some aspects, the techniques described herein relate to a computer-implemented method, where detecting that a mode of a head mounted device for screen time tracking is a high power mode includes detecting a battery level of the head mounted device and determining that the battery level exceeds a threshold.
[0019] In some aspects, the techniques described herein relate to a computer-implemented method, further including detecting that a mode of the head-mounted device for screen time tracking is a camera-less mode, capturing orientation data of the head-mounted device using at least one orientation sensor of the head-mounted device, capturing relative position data of the device from an ultra-wideband (UWB) signal received by the head-mounted device, analyzing the orientation data and the relative position data to determine a viewing state, starting a screen timer after the viewing state is identified, periodically capturing subsequent orientation data and relative position data to track the viewing state over time, and stopping the screen timer when the viewing state can no longer be tracked.
[0020] In some aspects, the techniques described herein relate to a computer-implemented method, further including recording the screen time in a database based on the screen timer and generating an alert when the screen time meets a criterion.
[0021] In some aspects, the techniques described herein relate to a computer-implemented method that further includes classifying the instance of screen time as corresponding to a device based on device information captured by the head-mounted device.
[0022] In some aspects, the technology described herein relates to a head mounted device comprising an eye tracking camera directed at a user's eyes, a front camera directed at the user's field of view, at least one orientation sensor configured to sense orientation data of the head mounted device worn by the user, and a processor configured by instructions to detect a low power mode based on a battery level of the head mounted device, in which a screen time of the device is determined based on eye images from the eye tracking camera applied to a neural network, and a processor configured by instructions to detect a high power mode based on a battery level of the head mounted device, in which a screen time of the device is determined based on images from the front camera applied to an image recognition algorithm.
[0023] In some aspects, the technology described herein relates to a head-mounted device, wherein the processor is configured by the instructions to further detect a camera-less mode based on user input, wherein in the camera-less mode, screen time of the device is captured from an ultra-wideband (UWB) signal received by the head-mounted device based on orientation data of the head-mounted device and relative position data of the device and is applied to a machine learning model.
[0024] In some aspects, the technology described herein relates to a head-mounted device, where the head-mounted device is augmented reality glasses (i.e., AR glasses).
[0025] The foregoing summary, as well as other illustrative objects and / or advantages of the present disclosure, and the manner in which they are accomplished, are further described in the following detailed description and accompanying drawings. [Brief description of the drawings]
[0026] [Figure 1] 1 illustrates a graph of a user's screen time over a period of time, according to a possible embodiment of the present disclosure. [Diagram 2] FIG. 2 is a perspective view of a head-mounted device according to a possible embodiment of the present disclosure. [Diagram 3] FIG. 1 is a system block diagram of a head-mounted device according to a possible embodiment of the present disclosure. [Figure 4] 1 is a flowchart illustrating a multi-mode screen time tracking method according to a possible implementation of the present disclosure. [Diagram 5] FIG. 5 is a block diagram of a possible neural network for the low power mode of the multi-mode screen time tracking method shown in FIG. [Figure 6] 13 shows a screen image discerned in reflection in an eye image according to a possible embodiment of the present disclosure. [Figure 7] 1 is a flowchart of a method for recording screen time according to a possible embodiment of the present disclosure. [Figure 8] 1 illustrates an alert generated based on screen time according to a possible embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0027] The elements in the drawings are not necessarily to scale relative to each other and like reference numerals indicate corresponding parts throughout the several views.
[0028] It may be desirable to measure a person's total screen time over a period of time. However, if the person uses multiple devices, it may not be practical to accurately measure a person's total screen time. For example, in the course of a day, a person may watch movies on a television, play video games on a computer, read a book on a tablet, and video chat on a mobile phone. One or more of these devices may be configured to record (i.e., measure) screen time, but for various reasons, there is no way to measure and aggregate screen time across multiple devices. For example, some of the devices may not be able to easily exchange screen time information due to different operating systems, while other devices (e.g., older televisions) may not be able to measure or communicate this information at all.
[0029] The present disclosure addresses the technical problem of measuring and recording (i.e., tracking) screen time across multiple devices by using a head mounted device configured to monitor a user to determine the time the user is using the screen of the device and to calculate the screen time of the device based on these times. The head mounted device may be further configured to track multiple screen times for each of the multiple devices. Additionally, the head mounted device may be configured to combine the screen times of the devices to measure an overall screen time of the person. Additionally, the head mounted device may be configured to take an action when the screen time of one or more devices meets a criterion. For example, the criterion may be a threshold related to a maximum screen time over a period of time and the action may be related to limiting the user's screen time. The criterion and the period of time may be configurable and suitable for various applications.
[0030] Screen time tracking may require periodic measurements using a head mounted device. Periodic measurements may consume power and may limit the usage time of the head mounted device. For example, augmented reality glasses (i.e., AR glasses) may have a limited amount of power due to small battery size. As a result, regular screen time tracking may cause technical issues of consuming more power than is practical for the head mounted device. The present disclosure further addresses the technical issue of power consumption associated with screen time measurement by disclosing a multi-mode approach, where each mode has a different power consumption. Selecting a mode of screen time tracking may be based on the power state of the head mounted device. For example, in a low battery state, a screen time tracking mode that consumes less power may be selected to extend the operation time of the head mounted device. Furthermore, the multi-mode approach may enable additional modes that provide further benefits to the user. For example, a non-image mode for tracking screen time is disclosed, which may improve the user's privacy.
[0031] A user may not be aware of the screen time accumulated over a period of time. For example, a user may be too distracted by screen content to be aware of the time spent observing the screen. The present disclosure further addresses the need to provide a screen time alert to a user. For example, when a certain amount of total screen time across devices is reached, the head mounted device may be configured to generate an alert and / or temporarily limit the functionality of the device. These actions may help prevent or reduce symptoms associated with excessive device use and may help the user follow rules regarding screen time. Additionally, these actions may help the user or those responsible for the user (e.g., parents, managers, etc.) to understand the user's behavior regarding screen time in general. For example, screen time tracking may be used to verify viewing time (e.g., track the user's work).
[0032] FIG. 1 illustrates a graph of screen time for multiple devices. The screen time for each device is plotted on a timeline for a period of time (e.g., 24 hours). The screen time is shown as a line for each instance of screen engagement, with the length of each line corresponding to the screen time for the device. For example, an instance 110 of screen time for a desktop computer is shown in a first graph 101. The screen time for the instance 110 may be calculated as the elapsed time between the beginning 111 (i.e., start) of the instance 110 and the end 112 (i.e., stop) of the instance 110. The accumulated screen time for all devices (i.e., across devices) is also plotted as the total screen time for the period 100. The graph is presented for ease of understanding. It should be understood that the illustrated devices and screen times are hypothetical and represent one possible implementation of many possible variations. For example, screen time for additional and / or different devices may be included.
[0033] As shown, a first graph 101 shows instances of screen time corresponding to a first device (i.e., a desktop computer), a second graph 102 shows instances of screen time corresponding to a second device (i.e., a television), a third graph 103 shows instances of screen time associated with a third device (i.e., a mobile phone), and a fourth graph 104 shows instances of screen time associated with a fourth device (i.e., a tablet). A fifth graph 105 shows instances of screen time across all devices (i.e., across all devices). The instances of screen time may be summed to calculate a total screen time over a period of time.
[0034] Total screen time may be calculated in a variety of ways. In a possible implementation, the total screen time may be calculated as a running sum of accumulated screen time. In another possible implementation, the total screen time may be the sum of all instances of screen time measured during that period. Thus, screen time may be measured as a time (e.g., minutes) or as a percentage (e.g., minutes / period).
[0035] In possible implementations, measuring screen time may include a screen timer. For example, a screen timer may be started and stopped for each instance of screen time during the period. The screen timer may then be reset when the period ends and restarted at the beginning of a subsequent period. The period 100 may be set by a user (e.g., a parent) or may be set at the factory based on requirements (e.g., health guidelines). The period 100 may also be adjustable. For example, the period may end when a total screen time is reached. For example, screen time may accumulate until a threshold is reached, at which point the period ends.
[0036] The total screen time tracking may be customized to calculate a running total of screen time based on device, time period (e.g., time of day), and / or location. For example, screen time on a user's work device may be tracked at times corresponding to a work day, and in possible implementations, a daily, weekly, or monthly total screen time for the work device may be calculated. Additionally, a time period corresponding to a particular screen time tracking may be triggered based on the location of the head mounted device (i.e., the user). For example, a time period corresponding to a work day may begin when the user arrives at a location corresponding to work. The user may track screen time based on device, time, and / or location.
[0037] 2 is a perspective view of a head-mounted device according to a possible embodiment of the present disclosure. As shown, the head-mounted device 200 can be implemented as smart glasses (e.g., AR glasses) configured to be worn on a user's head. The head-mounted device 200 includes a left lens and a right lens coupled to the user's ears by a left arm and a right arm, respectively. The user (i.e., the wearer) can see the world through the left lens and the right lens coupled to each other by a bridge configured to rest on the wearer's nose.
[0038] The head mounted device 200 may include sensing devices configured to help determine where the user's focus is directed. For example, the head mounted device 200 may include at least one front-facing camera. The front-facing camera 230 may include an image sensor capable of detecting intensity levels of visible and / or near-infrared light (i.e., NIR light). The front-facing camera 230 may be directed toward a forward field of view (i.e., forward FOV 235) or may include optics that route light from the forward FOV 235 to an image sensor. For example, the front-facing camera 230 may be disposed on the front of the head mounted device and may include at least one lens for creating an image of the forward field of view (forward FOV 235) on the image sensor. Because the forward FOV 235 includes all (or a portion) of the user's field of view, an image or video of the world as seen from the user's point of view (POV) may be captured by the front-facing camera 230.
[0039] The head mounted device 200 may further include at least one eye tracking camera. The eye tracking camera 220 may be directed at a field of view of the eye (i.e., eye FOV 225) or may include optics that route light from the eye FOV 225 to an eye image sensor. For example, the eye tracking camera 220 may be directed at the user's eye and may include at least one lens for generating an image of the eye FOV 225 on the eye image sensor. The eye FOV 225 may include all (or a portion) of the user's eye field of view to include an image or video of the eye. The eye image may be analyzed by a processor (not shown) of the head mounted device to determine where the user is looking. For example, the relative position of the pupils in the eye image may correspond to the user's gaze direction.
[0040] The head mounted device may further include at least one orientation sensor 250. The orientation sensor(s) may be implemented combining an accelerometer, a gyroscope, and a magnetometer to form an inertial measurement unit (i.e., IMU) to determine the orientation of the head mounted device. The IMU may be configured to provide multiple measurements that describe the orientation and movement of the head mounted display. For example, the IMU may have six degrees of freedom (6-DOF) that can describe three translational motions (i.e., x-direction, y-direction, or z-direction) along the axes of the world coordinate system 260 and three rotational motions (i.e., pitch, yaw, roll) about the axes of the world coordinate system 260. Data from the IMU may be combined with information about the Earth's magnetic field using sensor fusion to determine the orientation of the head mounted device coordinate system 270 relative to the world coordinate system 260. Information from the front camera 230, eye FOV 225, and IMU 250 may be combined to determine where the user's focus is directed, enabling augmented reality applications. The head mounted display may further comprise an interface device for these applications.
[0041] Head mounted device 200 may further include a human interface subsystem configured to present information to a user. For example, head mounted device 200 may include a display configured to display information (e.g., text, graphics, images) in a display area 240 in one or both lenses. The display area may be all or a portion of the lens and may be visually transparent or translucent, allowing a user to see through the display area when not in use. Head mounted device 200 may further include one or more speakers (e.g., earphones 255) configured to play sounds (e.g., voice, music, tones). This disclosure describes systems and methods for configuring sensing and interface devices of a head mounted device for screen time measurement and response (e.g., alerts).
[0042] 3 is a system block diagram of a head-mounted device according to a possible implementation of the present disclosure. As previously described, the head-mounted device 200 includes a front-facing camera 230 that is pointed toward the user's viewpoint and configured to capture images of a forward FOV 235. The head-mounted device 200 further includes an eye-tracking camera 220 that is pointed toward the user's eye (or eyes) and configured to capture images of an eye FOV 225. The head-mounted device 200 further includes at least one orientation sensor 250 (e.g., an IMU) configured to sense the movement, position, and / or orientation of the head of a person wearing the head-mounted device 200.
[0043] The head mounted device 200 may further comprise at least one position sensor 310. The at least one position sensor may be configured to determine a position of the head mounted device (i.e., of the user). The position sensor 310 may comprise an ultra-wideband (UWB) sensor. The position sensor 310 may communicate with the screen device 317 via a communication link 315. For example, the head mounted device 200 and the screen device 317 may exchange packets of information over the UWB communication link to determine the relative positions of the devices. For example, the position sensor may be configured to determine a round trip time (RTT) of packets communicated between the devices. The range between the head mounted device and the screen device 317 may be sensed based on the round trip time. Furthermore, the position sensor may comprise multiple receivers configured to receive packets communicated from the screen device 317. The position sensor may be configured to determine the time of arrival of the packets at the receivers to determine the angle between the screen device 317 and the position sensor 310. The location sensor(s) may further include a Global Positioning System (GPS) sensor that may be used to determine the geographic location of the head-mounted device (i.e., the user). The geographic location may be further determined by a sensor fusion approach where information from a local area network (e.g., a WiFi network) and / or a cellular network may further refine the geographic location.
[0044] The head mounted device 200 further comprises at least one processor 330. The processor communicates with the cameras, sensors, and other modules and electronics of the head mounted device. The processor is configured to execute the multi-mode method for tracking screen time across devices with instructions (e.g., software, applications, etc.). The instructions may be non-transitory computer readable instructions stored in and recalled from the memory 340. Alternatively, the instructions may be communicated to the processor from a computer coupled to the network 370 via the communication interface 350.
[0045] The processor 330 may be in communication with the battery 360. The processor 330 may be configured, via instructions, to monitor the level of the battery and determine a mode of the head mounted device based on the level of the battery. For example, if the level of the battery is below a threshold, the head mounted device may be determined to be in a low power mode, and if the level of the battery is above the threshold, the head mounted device may be determined to be in a high power mode.
[0046] The processor may be in communication with the display 320 of the head mounted device 200. The processor 330 may be configured, upon instruction, to transmit text, graphics, video, images, etc. to the display 320. For example, the processor 330 may be configured to present an alert to the user on the display 320 in response to the measured screen time.
[0047] FIG. 4 is a flowchart of a multi-mode screen time tracking method according to a possible embodiment of the present disclosure. The method 400 includes detecting 410 a mode (i.e., a screen time mode) that the head mounted device uses to measure screen time. Detecting the screen time mode (i.e., a mode) may be implemented in various ways. In one possible embodiment, detecting 410 the mode may include receiving an input from a user (i.e., a user input 402) specifying the mode. In another possible embodiment, detecting 410 the mode may include receiving the mode from a device setting 403 of the head mounted device, where the device setting (i.e., a system setting) may be configurable by an application (e.g., an AR application) running on the head mounted device. In another possible embodiment, detecting 410 the mode may include determining the mode based on a battery level 401 of the head mounted device to balance the quality of tracking and the operating time of the head mounted device.
[0048] In possible implementations, the detected mode may be high power mode 490, low power mode 600, or no camera mode 900. Each mode may have different accuracy and power consumption. For example, the head mounted device 200 may consume more power when in high power mode 490 than when in low power mode 600. The high power mode 490 may generate screen time with higher accuracy (i.e., confidence) than the low power mode 600. The no camera mode 900 may be used without any imaging, providing privacy to the user and providing the lowest power consumption of the modes. However, the no camera mode 900 may be optional since it may depend on position and orientation sensors that may or may not be present in all screen devices.
[0049] In possible implementations, screen time tracking may alternate between modes depending on the need for either accuracy or battery life. For example, a device's screen time tracking may begin in a high power mode 490 and then move to a low power mode 600, either between instances or within the same instance. This may help to recognize the device or screen content at the start of an instance of screen time with a higher degree of confidence before transitioning to a low power mode where screen time can be tracked more efficiently. In another example, a high power mode may be triggered when a screen is detected during a low power mode. This may help to increase accuracy only if it is determined with some confidence that a screen is being displayed.
[0050] As shown in the embodiment of Figure 4, the high power mode 490 uses the head mounted device's front camera 230 and image recognition algorithms 424 to determine the beginning (i.e., start) and end (i.e., stop) of an instance of screen time. The low power mode 600 uses the head mounted device's eye tracking camera 220 and a neural network 434 (e.g., a U-net neural network) to determine the beginning / end of an instance of screen time. The camera-less mode 900 uses the head mounted device's position / orientation sensor 443 and a machine learning model (e.g., sensor fusion) to determine the beginning / end of an instance of screen time.
[0051] Each mode may trigger the on / off of the screen timer 430 according to the start / stop of an instance. In other words, the screen timer 430 may be used to determine screen time regardless of the mode used to determine the start / stop of an instance of screen time. For example, the screen timer may be activated (i.e., started) when it is determined that the screen is within the user's field of view, and then deactivated (i.e., stopped) when it is determined that the screen is no longer within the user's field of view. The determination of the screen viewing state may be performed in different ways depending on the mode. The screen timer 430 may be reset at the end of an instance. Alternatively, the screen timer 430 may be turned on / off without resetting to accumulate screen time for multiple instances during a given period, and then reset at the end of the given period.
[0052] As shown in FIG. 4, the method 400 may include a high power mode 490. The high power mode 490 uses the front camera 230 of the head mounted device to consume more power than the eye tracking camera 220. In the low power mode 600, the front camera 230 may be configured on (i.e., enabled, powered on, etc.) to capture an image(s) corresponding to the user's field of view. For example, the front camera 230 captures a video (e.g., high resolution color video) of items in the user's field of view. The image(s) may be analyzed by an image recognition algorithm 424 to identify (i.e., detect) the screen(s) the user is looking at. Detection of the screen may trigger a screen timer 430 to begin timing.
[0053] The high power mode 490 process described above may be repeated periodically (i.e., cyclically) to track the screen in subsequent captured images to determine the screen time of the instance. Determining the screen time of the instance may include determining when the screen timer 430 should be stopped (i.e., deactivated or turned off). For example, if it is determined that no screen has been detected in the images for several cycles, the screen timer may be turned off. Each cycle may be separated by a cycle period. For example, the high power mode 490 may be repeated with a cycle period of every few seconds (e.g., 2 seconds). When the instance ends, the screen timer may output the screen time of the instance for recording and / or monitoring.
[0054] As shown in FIG. 4, the method 400 may include a low power mode 600. The low power mode 600 consumes less power using the eye tracking camera 220 of the head mounted device than a front camera. In the low power mode 600, the eye tracking camera 220 may be configured on (i.e., enabled, powered on, etc.) to capture an image(s) corresponding to one (or both) eyes of a user. For example, the eye tracking camera 220 may capture a video of the user's eyes. The image of the eyes may be analyzed by a neural network to detect a reflection and identify a screen in the reflection. Identification of the screen may trigger the screen timer 430 to begin timing.
[0055] 5 is a flow chart block diagram of a possible neural network for the low power mode of the multi-mode screen time tracking method shown in FIG. 4. The neural network is a U-net neural network 500 configured to receive an eye image 510 (i.e., an input image) from the eye tracking camera 220. The U-net neural network 500 can be configured to detect a reflection in the eye image and identify a screen in the reflection. The identification can include generating a mask including the reflection and applying the mask to the eye image 510 of the eye to generate a screen image 520 (i.e., an output image). In other words, the U-net neural network can be configured to predict whether a screen will appear in the eye image, and if so, to predict where the screen will be located in the eye image (i.e., the center position of the screen).
[0056] 6 illustrates a screen image identified in a reflection in an eye image according to a possible embodiment of the present disclosure. As shown, a reflection can be detected in the eye image 610. The screen may be identified in the reflection by a neural network to generate a screen image 620. In this possible embodiment, the entire screen is reflected. In other words, the screen image 620 can be generated based on the results of a U-net neural network 500 (see FIG. 5) configured to identify the location of the screen in the image by masking (i.e., segmenting) the eye image 610.
[0057] Returning to FIG. 4, the low power mode 600 process described above may be repeated periodically (i.e., cyclically) to track the screen in subsequent captured images to determine the screen time of an instance of screen viewing. Determining the screen time of an instance may include determining when the screen timer 430 should be stopped (i.e., deactivated or turned off). For example, if it is determined that no screen has been detected in the image for several cycles, the screen timer may be turned off. Each cycle may be separated by a cycle period. For example, the low power mode may be repeated with a cycle period of every few seconds (e.g., 5 seconds). The cycle period of the low power mode may be longer than the cycle period of the high power mode to conserve power. The cycle period may be selected depending on a balance between screen time accuracy and power consumption. In a possible implementation, the low power mode 600 may be the default mode of the method.
[0058] As shown in FIG. 4, the method 400 may include a camera-less mode 900. The camera-less mode 900 uses a position / orientation sensor 443 of the head-mounted device to determine user orientation data and relative position data of the device. The user orientation data and the device (i.e., screen) relative position data may be analyzed by a machine learning model 444 to determine that a screen viewing state (i.e., viewing state) exists. For example, a viewing state may exist when the orientation of the head-mounted device is pointing toward a device having a screen. Detection of the viewing state may trigger the screen timer 430 to begin timing. As previously described, the relative position data may include ranges and angles calculated based on using ultra-wideband (UWB) communications. For example, the range between UWB devices may be calculated based on the round trip time of UWB packets communicated between the UWB devices. Additionally, the angle between the UWB devices may be calculated based on the difference in arrival times of the UWB packets at spaced UWB receivers on the UWB devices receiving the UWB packets.
[0059] The above process of the no-camera mode 900 may be repeated periodically (i.e., cyclically) to track viewing states and determine screen time for an instance of screen viewing. Determining the screen time for an instance may include determining when the screen timer 430 should be stopped (i.e., deactivated or turned off). For example, the screen timer 430 may be turned off when it is determined that a viewing state has not been detected in the position / orientation data for several cycles. Each cycle may be separated by a cycle period. For example, the low power mode may be repeated with a cycle period of every few seconds (e.g., 5 seconds).
[0060] As shown in FIG. 4, after the screen time is calculated, the screen time may be output for a recording / monitoring 450 process of the method 400. The recording / monitoring process 450 may include recording the screen time determined from the screen timer in a database 750. The recording / monitoring process 450 may further include linking 455 the screen time with a device (i.e., classifying the screen time for the device). The recording / monitoring process 450 may further include generating 465 an alert based on the screen time(s) stored in the database 750. For example, the screen time of an instance may be output from the screen timer in a high power mode, a low power mode, or a no camera mode, and classified, recorded, evaluated, and an alert may be generated.
[0061] 7 is a flowchart of a method for recording screen time according to a possible implementation of the present disclosure. The method 700 includes receiving screen time (e.g., of an instance) captured using a head-mounted device in a high power mode, a low power mode, or a camera-less mode. The method 700 further includes classifying 710 the screen time as screen time corresponding to the device.
[0062] The classification 710 may utilize a machine learning model 715. The machine learning model 715 may receive device information 720 captured during an instance to link the screen time 701 with a type of device (e.g., laptop, desktop, TV, etc.) or a specific device (e.g., John's phone, TV in the living room). The device information may affect the confidence of the classification. For example, if a hand holding a screen is recognized in the image, the model may decrease confidence that the screen is a TV and increase confidence that the screen is a phone or tablet.
[0063] Device information 720 may include screen characteristics 721 derived from a captured image (e.g., FOV image). In a possible implementation, screen characteristics 721 include physical or electronic attributes of the screen derived from the image, such as screen size (e.g., length, width), screen resolution, or screen refresh rate. In another possible implementation, screen characteristics 721 may include attributes of the periphery of the screen derived from the image, such as the housing of the screen, or the hand holding the screen.
[0064] Device information 720 may further include screen content 722 derived from the captured image (e.g., eye image). In possible implementations, screen content may include motion, color, images, etc. that are indicative of the type of screen being viewed. For example, a particular user interface may be recognized and used to determine the device.
[0065] The device information 720 may further include a device identification (Device ID 723) derived from the location data. In possible implementations, the device ID 723 may include identification information (e.g., an address) communicated via UWB packets. For example, an address communicated in a packet transmitted from the device may be used to determine the device.
[0066] The device information 720 can further include a user location 724. The user location can be a location derived from a location service (e.g., GPS, indoor positioning, etc.). The user location 724 can be a room. The room knowledge can inform the determination the machine learning model makes. For example, a TV is more likely to be in the living room than the bathroom. In other words, the method can include receiving a location of a head mounted device and selecting a likely device based on the location.
[0067] The method for recording screen time may further include verifying 730 whether the classification made by the machine learning model 715 is correct or incorrect. The verification may be done in various ways. In one possible implementation, an image from a front camera may be captured (e.g., during a low power mode). The FOV image may be analyzed to determine whether the classification of the device was correct. In another possible implementation, the verification may include querying the user to classify the device (e.g., when no results from the model are returned) or to confirm that the classification was correct. The results of this verification may be used to update 740 the model. For example, updating 740 the model may include adjusting the model based on instances of verification or based on a rate of verification (e.g., false identification rate, true identification rate) such that the confidence of the classification attributed to the particular device information 720 changes.
[0068] The method for recording screen time may further include recording the (categorized) screen time in a database 750. The recording may include updating (i.e., increasing) the screen time of the device with the screen time 701 received (or derived) from the screen timer. Thus, recording 745 may include querying the database for an entry corresponding to the device and then adding the screen time of the device to the entry or creating a new entry for the device.
[0069] Returning to FIG. 4 , the multi-modal screen time tracking method 400 may further include generating an alert based on the screen time recorded in the database 750 and transmitting the alert to a display of the head mounted device. Generating the alert may include querying the database based on one or more rules corresponding to the one or more alerts. For example, a rule may specify that the screen time of a particular device must not exceed a value during a certain period of time. Generating the alert may include querying the database 750 to determine the screen time of the particular device and then comparing the screen time to a threshold. If the comparison indicates that the screen time of the particular device is equal to or greater than the threshold, an alert may be generated. Various alerts may be generated.
[0070] The alert may include a message presented to the user regarding screen time. For example, the message may instruct the user to "take a break" from using the device. The alert may include a report of screen time usage over a period (or periods). The alert may include a comparison of a first screen time from a first period and a second screen time from a second period. These types of alerts may be presented on a display of the head-mounted device for the user to view.
[0071] FIG. 8 is an example of an alert based on screen time. The alert may include a daily screen time for all devices (e.g., desktop, phone, tablet, etc.) shown as a bar graph for each day of the week. The alert may further include the accumulated screen time for each of the devices over the week. The alert may further include the total screen time for all devices for the week. The alert may further include an analysis of the screen time, such as a daily average of screen time. The alert may further include the change in screen time from a previous period (i.e., last week). The alert may be triggered to display at the end of a period (e.g., week) or may be triggered by the user. In some implementations, an alert may be generated when the screen time meets one or more criteria.
[0072] In the present specification and / or drawings, exemplary embodiments are disclosed. The present disclosure is not limited to such exemplary embodiments. The use of the term "and / or" includes any and all combinations of one or more of the associated listed items. The figures are schematic and, therefore, are not necessarily drawn to scale. Unless otherwise indicated, certain terms are used in a generic and descriptive sense and not for purposes of limitation.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein may be used in the practice or testing of the present disclosure. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. As used herein, the term "comprising" and variations thereof are used synonymously with the term "including" and variations thereof and are open-ended terms. As used herein, the term "optional" or "optionally" means that the subsequently described feature, event, or circumstance may or may not occur, and the description includes instances when the feature, event, or circumstance occurs and instances when the feature, event, or circumstance does not occur. As used herein, ranges may be expressed as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, the embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another aspect. It will be further understood that the endpoints of each of the ranges are meant both in relation to the other endpoint, and independently of the other endpoint.
[0074] Some embodiments may be implemented using various semiconductor processing and / or packaging technologies. Some embodiments may be implemented using various types of semiconductor processing technologies associated with semiconductor substrates, including, but not limited to, silicon (Si), gallium arsenide (GaAs), gallium nitride (GaN), silicon carbide (SiC), etc.
[0075] As described herein, certain features of the described embodiments have been illustrated, but numerous modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It should therefore be understood that the appended claims are intended to cover all such modifications and changes that fall within the scope of the embodiments. They are presented only by way of example, not limitation, and it should be understood that various changes in form and details may be made. Any part of the apparatus and / or method described herein may be combined in any combination, except in mutually exclusive combinations. The embodiments described herein may include various combinations and / or subcombinations of the functions, components, and / or features of the different embodiments described.
[0076] In the foregoing description, when an element is referred to as being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it will be understood that the element may be directly on, directly connected to, or directly coupled to another element, or there may be one or more intervening elements. In contrast, when an element is referred to as being directly on, directly connected to, or directly coupled to another element, there are no intervening elements. Throughout the detailed description, elements that are shown to be directly on, directly connected to, or directly coupled may be so referenced, even if the terms directly on, directly connected, or directly coupled are not used. The claims of this application may be amended to recite the exemplary relationships, if any, described in the specification or shown in the drawings.
[0077] As used herein, the singular may include the plural unless the context clearly indicates otherwise in a particular instance. Spatially relative terms (e.g., on, above, upper of, below, below, below, and below) are intended to encompass various orientations of the device during use or operation in addition to the orientation depicted in the drawings. In some embodiments, the relative terms "on" and "below" may include vertically above and vertically below, respectively. In some embodiments, the term "adjacent" may include laterally adjacent or horizontally adjacent.
Claims
1. A head-mounted device, an eye tracking camera pointed at the user's eyes; a processor in communication with the eye tracking camera, the processor being configured by instructions to: capturing an eye image of the user's eye using the eye tracking camera; analyzing the eye image to identify a screen reflection; starting a screen timer after the screen reflection is identified; periodically capturing subsequent eye images to track the screen reflection over time; stopping the screen timer when the screen reflection can no longer be tracked; and recording a screen time of the first instance based on the screen timer. Head-mounted device.
2. The head mounted device of claim 1 , wherein analyzing the eye image to identify screen reflections is performed by a neural network.
3. The processor further comprises: classifying the screen time of the first instance as relating to a first device based on a machine learning model configured with device information captured by the head-mounted device; and recording the screen time of the first instance of the first device in a database. A head mounted device according to any preceding claim.
4. The head mounted device of claim 3 , wherein the device information is any combination of a device identifier, content of the screen reflection, characteristics of the screen reflection, and a position of the user.
5. The processor, under instructions, further configured to detect that a mode of the head mounted device for screen time tracking is a low power mode; A head mounted device according to any preceding claim.
6. It also has a front camera, The processor further comprises: Detecting that the mode of the head mounted device for screen time tracking is a high power mode; capturing a field of view image of the user's field of view using the front camera; analyzing the field of view image to identify a screen; starting the screen timer after the screen is identified in the field of view image; periodically capturing subsequent field of view images to track said screen over time; stopping the screen timer when the screen is no longer tracked; and recording a screen time of the second instance based on the screen timer. The head mounted device of claim 5 .
7. at least one orientation sensor configured to sense orientation data of the head mounted device; at least one position sensor configured to sense relative position data of the device; The processor further comprises: and configured to detect that the mode of the head mounted device for screen time tracking is a camera-less mode for screen time tracking, in which in the camera-less mode, the processor: capturing the orientation data of the head mounted device and the relative position data of the device; analyzing the orientation data and the relative position data to identify a viewing state; starting the screen timer after the viewing state is identified; periodically capturing subsequent orientation data and subsequent relative position data to track said viewing state over time; stopping the screen timer when the viewing state is lost; and recording a screen time of a third instance based on the screen timer. The head mounted device of claim 6.
8. the at least one orientation sensor includes an inertial measurement unit (IMU); the at least one position sensor includes an ultra-wideband (UWB) sensor; The head mounted device of claim 7.
9. The processor further comprises: generating an alert based on the screen time of the first instance, the screen time of the second instance, and / or the screen time of the third instance; displaying the alert on a display of the head mounted device. A head-mounted device according to claim 7 or 8.
10. Detecting that a mode of the head mounted device for screen time tracking is a low power mode; capturing an eye image of an eye using an eye tracking camera of the head mounted device; analyzing the eye image to identify a screen reflection; starting a screen timer after the screen reflection is identified; periodically capturing subsequent eye images to track the screen reflection over time; stopping the screen timer when the screen reflection can no longer be tracked; A computer-implemented method comprising:
11. periodically capturing subsequent eye images to track said screen reflection over time; 11. The computer-implemented method of claim 10, comprising repeating the capturing and analyzing at a cycle period greater than two seconds.
12. Detecting that the mode of the head mounted device for screen time tracking is a low power mode includes: Detecting a battery level of the head mounted device; determining that the battery level is below a threshold; 12. A computer implemented method according to claim 10 or claim 11, comprising:
13. Analyzing the eye image to identify a screen reflection includes: applying the eye image, including the eye and the screen reflection, to an input of a neural network; receiving a screen image including the screen reflection at an output of the neural network; A computer-implemented method according to any one of claims 10 to 12, comprising:
14. Detecting that the mode of the head mounted device for screen time tracking is a high power mode; capturing a field of view image using a front camera of the head mounted device; analyzing the field of view image to identify a screen; starting the screen timer after the screen is identified; periodically capturing subsequent field of view images to track said screen over time; stopping the screen timer when the screen is no longer tracked; The computer-implemented method of any of claims 10 to 13, further comprising:
15. Detecting that the mode of the head mounted device for screen time tracking is a high power mode includes: Detecting a battery level of the head mounted device; determining that the battery level exceeds a threshold; 15. The computer implemented method of claim 14, comprising:
16. Detecting that the mode of the head mounted device for screen time tracking is a camera-less mode; capturing orientation data of the head mounted device using at least one orientation sensor of the head mounted device; Capturing relative position data of the head-mounted device from ultra-wideband (UWB) signals received by the device; analyzing the orientation data and the relative position data to determine a viewing state; starting a screen timer after the viewing state is identified; periodically capturing subsequent orientation and relative position data to track said viewing state over time; stopping the screen timer when the viewing state is lost; 16. The computer implemented method of claim 14 or claim 15, further comprising:
17. recording a screen time in a database based on the screen timer; generating an alert when the screen time meets a criterion; 20. The computer implemented method of claim 16, further comprising:
18. 20. The computer-implemented method of claim 17, further comprising classifying an instance of screen time as corresponding to a device based on device information captured by the head-mounted device.
19. A head-mounted device, an eye tracking camera pointed at the user's eyes; a front-facing camera directed toward the user's field of view; at least one orientation sensor configured to sense orientation data of the head mounted device worn by the user; a processor, the processor being configured to: configured to detect a low power mode based on a battery level of the head mounted device; In the low power mode, The screen time of the device is determined based on eye images from the eye tracking camera applied to a neural network, and the processor is configured to: configured to detect a high power mode based on the battery level of the head mounted device; In the high power mode: A head mounted device, wherein the screen time of the device is determined based on images from the front camera applied to an image recognition algorithm.
20. The processor, under instructions, further configured to detect a camera-free mode based on user input; In the no-camera mode, The screen time of the device is captured from an ultra-wideband (UWB) signal received by the head mounted device based on the orientation data of the head mounted device and the relative position data of the device and applied to a machine learning model.
20. A head mounted device as claimed in claim 19.
21. 21. A head mounted device as claimed in claim 19 or claim 20, wherein the head mounted device is an augmented reality pair of glasses.
Citation Information
Patent Citations
Method and device for simulating image by using AR glasses, AR glasses and medium
CN114063778A
Three dimensional display processing apparatus and three dimensional display processing method
JP2008078846A
Information processing device, information processing method and program
JP2016149660A
System And Methods for Providing Information
US20150268719A1
Tethering type head mounted display and method for controlling the same
US20170123744A1