Attention monitoring method, apparatus, device, and storage medium
By monitoring the eye opening and gaze information of VR/AR device users through head-mounted devices, the problem of attention distraction caused by prolonged use is solved, enabling timely intervention and attention concentration, and improving work efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING GOERTEK ACOUSTICS TECH CO LTD
- Filing Date
- 2023-09-22
- Publication Date
- 2026-05-01
AI Technical Summary
Prolonged use of VR/AR devices can lead to distraction, affecting work efficiency and accuracy.
By collecting facial videos of users through head-mounted devices, extracting eye information, monitoring eye opening and closing and gaze information, determining attention status, and taking intervention measures when attention is distracted.
It improves the accuracy of attention monitoring, and timely intervention measures ensure that users maintain concentration while working, thereby improving work efficiency.
Smart Images

Figure CN117315535B_ABST
Abstract
Description
Attention monitoring methods, devices, equipment and storage media Technical Field
[0001] This invention relates to the field of attention monitoring technology, and in particular to an attention monitoring method, apparatus, device, and storage medium. Background Technology
[0002] VR / AR devices are used in many important work scenarios, such as meetings, driving, and remote operation of equipment, where user engagement is crucial. However, VR / AR devices can interfere with the user's interaction with reality to some extent. Prolonged concentration can lead to rapid sensory and mental fatigue, causing the user's attention to quickly wander, reducing work accuracy and efficiency. At this point, the user is no longer suitable to continue working in this state. Summary of the Invention
[0003] The main objective of this invention is to provide an attention monitoring method, device, equipment, and storage medium, aiming to solve the technical problem in the prior art where users' attention is distracted and their work efficiency is affected when they wear VR / AR devices for extended periods.
[0004] To achieve the above objectives, the present invention provides an attention monitoring method, the method comprising the following steps:
[0005] When the user's usage time exceeds the preset time, the system collects the user's facial video within the preset time period through the head-mounted device and obtains the facial image in the facial video.
[0006] Determine the position of the eyes in the face image, and extract the eye information from the face image based on the eye position;
[0007] Based on the eye information, determine the user's eye opening and closing information within the preset time period, and determine the user's gaze information within the preset time period;
[0008] Based on the eye opening and closing information and the gaze information, the user's attention monitoring results are determined, and intervention measures are taken according to the attention monitoring results.
[0009] Optionally, determining the eye position in the face image includes:
[0010] Noise interference is removed from the face image to obtain the first face image;
[0011] The first face image is normalized based on the grayscale difference to obtain the second face image;
[0012] The grayscale distribution and gradient distribution of local regions in the second face image were statistically analyzed to obtain statistical results;
[0013] Based on the statistical results, non-eye information in the second face image is filtered out to obtain the third face image;
[0014] The position of the eyes in the third face image is determined by coarse localization.
[0015] Optionally, determining the user's eye opening and closing information within the preset time period based on the eye information includes:
[0016] Determine the number of frames of the face image in the face video;
[0017] The number of target frames of the target face image of the user within a preset time period is determined based on the eye information, wherein the target face image is a face image with an eye opening degree less than the opening threshold;
[0018] Based on the number of frames and the number of target frames, the user's eye opening and closing information within the preset time period is determined.
[0019] Optionally, determining the user's eye opening and closing information within the preset time period based on the eye information includes:
[0020] Based on the eye information, determine the user's single eye-closing duration and total eye-closing duration within a preset time period;
[0021] Set a duration threshold;
[0022] Based on the duration of a single eye closure, the total duration of eye closure, and the duration threshold, the user's eye opening and closing information within the preset time period is determined.
[0023] Optionally, determining the user's gaze information within the preset time period includes:
[0024] The user's EOG signal is collected within the preset time period using a head-mounted device;
[0025] Endpoint detection and frame segmentation are performed on the acquired EOG signals to obtain the energy of each frame;
[0026] Based on the energy of each frame, the signal features of the EOG signal are extracted;
[0027] The user's eye movements are identified based on the signal characteristics, and the user's gaze information within a preset time period is determined based on the eye movements.
[0028] Optionally, determining the user's gaze information within a preset time period based on the eye movement includes:
[0029] After the user is illuminated by an infrared illuminator built into the head-mounted device, the infrared photosensitive device built into the head-mounted device receives the target reflected light, wherein the target reflected light is the reflected light from the edge of the iris and sclera of the human eye.
[0030] The reflected light from the target is converted into an eye movement signal, and the direction and amplitude of eye movement are determined based on the eye movement signal;
[0031] The user's gaze point position is determined based on the direction, the amplitude, and the user's head position;
[0032] Based on the fixation point location and the eye movement, the user's fixation information within a preset time period is determined.
[0033] Optionally, the intervention based on the attention monitoring results includes:
[0034] Determine the user's usage scenario at the current moment;
[0035] When the usage scenario is a highly concentrated scenario, the user's work status is determined to be abnormal based on the attention monitoring results;
[0036] When the user's working state is determined to be abnormal, intervention measures are taken through the head-mounted device to restore the user's attention. The intervention measures include viewpoint enhancement, animation effects, surround sound effects, vibration alerts, and danger warnings.
[0037] Furthermore, to achieve the above objectives, the present invention also proposes an attention monitoring device, the attention monitoring device comprising:
[0038] The acquisition module is used to acquire facial videos of the user within a preset time period through a head-mounted device when the user's usage time is detected to be longer than a preset time period, and to acquire facial images in the facial videos.
[0039] A determining module is used to determine the position of the eyes in the face image and extract eye information from the face image based on the eye position;
[0040] The determining module is further configured to determine the user's eye opening and closing information within the preset time period based on the eye information, and to determine the user's gaze information within the preset time period;
[0041] The determining module is further configured to determine the user's attention monitoring results based on the eye opening and closing information and the gaze information, and to take intervention measures based on the attention monitoring results.
[0042] Furthermore, to achieve the above objectives, the present invention also proposes an attention monitoring device, which includes: a memory, a processor, and an attention monitoring program stored in the memory and executable on the processor, the attention monitoring program being configured to implement the steps of the attention monitoring method described above.
[0043] In addition, to achieve the above objectives, the present invention also proposes a storage medium storing an attention monitoring program, which, when executed by a processor, implements the steps of the attention monitoring method described above.
[0044] The attention monitoring method, apparatus, device, and storage medium proposed in this invention, when the user's usage time exceeds a preset time, acquires facial video of the user within a preset time period using a head-mounted device and obtains facial images from the facial video; determines the eye positions in the facial images and extracts eye information from the facial images based on the eye positions; determines the user's eye opening and closing information and gaze information within the preset time period based on the eye information; determines the user's attention monitoring result based on the eye opening and closing information and the gaze information, and takes intervention measures based on the attention monitoring result. By collecting the user's dynamic eye information in this way and combining eye opening and closing information and gaze information to determine the user's attention status, the accuracy of the monitoring results can be effectively improved. Timely intervention measures can be taken when the user's attention is detected to be distracted, further ensuring that the user's attention is focused when working. Attached Figure Description
[0045] Figure 1 is a schematic diagram of the structure of the attention monitoring device in the hardware operating environment involved in the embodiment of the present invention;
[0046] Figure 2 is a flowchart illustrating the first embodiment of the attention monitoring method of the present invention;
[0047] Figure 3 is a schematic diagram of the process for determining eye opening and closing information in the first embodiment of the attention monitoring method of the present invention;
[0048] Figure 4 is a flowchart illustrating the second embodiment of the attention monitoring method of the present invention;
[0049] Figure 5 is a flowchart illustrating the process of determining the fixation point position in the second embodiment of the attention monitoring method of the present invention;
[0050] Figure 6 is a schematic diagram of the specific process for monitoring the user's public state in the second embodiment of the attention monitoring method of the present invention;
[0051] Figure 7 is a structural block diagram of the first embodiment of the attention monitoring device of the present invention.
[0052] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0053] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0054] Referring to Figure 1, which is a schematic diagram of the attention monitoring device structure in the hardware operating environment of the embodiment of the present invention.
[0055] As shown in Figure 1, the attention monitoring device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0056] Those skilled in the art will understand that the structure shown in Figure 1 does not constitute a limitation on the attention monitoring device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0057] As shown in Figure 1, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and an attention monitoring program.
[0058] In the attention monitoring device shown in Figure 1, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the attention monitoring device of the present invention can be set in the attention monitoring device, and the attention monitoring device calls the attention monitoring program stored in the memory 1005 through the processor 1001 and executes the attention monitoring method provided in the embodiment of the present invention.
[0059] Based on the above hardware structure, an embodiment of the attention monitoring method of the present invention is proposed.
[0060] Referring to Figure 2, which is a flowchart of a first embodiment of an attention monitoring method according to the present invention.
[0061] In this embodiment, the attention monitoring method includes the following steps:
[0062] Step S10: When the user's usage time exceeds the preset time, the user's face video is captured by the head-mounted device within the preset time period, and the face image in the face video is obtained.
[0063] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a mobile phone, tablet computer, or personal computer, or an electronic device or attention monitoring device capable of performing the above functions. The following description uses the attention monitoring device as an example to illustrate this embodiment and the subsequent embodiments.
[0064] It should be noted that the user refers to the user who wears a head-mounted device to perform the work. The head-mounted device can be a VR / AR device. The preset time can be set in advance. When the user's usage time exceeds the preset time, the user may experience distraction (i.e., sensory and mental fatigue). Therefore, when the user's usage time exceeds the preset time, the user's attention can be monitored. When distraction is detected, timely intervention measures can be taken to help the user quickly adjust their state and refocus their attention. Specifically, after wearing the VR / AR device, the user can choose whether to activate the VR / AR device working status intervention mode. Once the intervention mode is activated, the user's attention can be monitored, and intervention measures can be taken when distraction is detected.
[0065] It should be noted that facial videos can be captured using the built-in camera of a VR / AR device. During the camera capture process, because the camera used by the VR / AR device is very close to the person's face, it can effectively eliminate the interference of different environments on the image. Facial images refer to all frame images obtained from facial videos.
[0066] In one embodiment, determining the eye position in the face image includes:
[0067] Noise interference is removed from the face image to obtain the first face image;
[0068] The first face image is normalized based on the grayscale difference to obtain the second face image;
[0069] The grayscale distribution and gradient distribution of local regions in the second face image were statistically analyzed to obtain statistical results;
[0070] Based on the statistical results, non-eye information in the second face image is filtered out to obtain the third face image;
[0071] The position of the eyes in the third face image is determined by coarse localization.
[0072] It should be noted that nonlinear median filtering can be used to remove noise interference in face images. To reduce the impact of factors such as uneven brightness of the image acquisition environment and equipment projection, the grayscale differences between face images need to be normalized beforehand, stretching the grayscale range of different eye images to the same area. After corresponding recognition and matching processing, a second face image is obtained. Next, the grayscale distribution, gradient distribution, and symmetry of the local area can be statistically analyzed to filter out non-eye image information in the second face image, resulting in a third face image. Finally, after obtaining a roughly stable eye position through coarse localization, the eye position is searched within a given range. Specifically, valley point search, directional projection, and eye symmetry can be used. A combination of methods can be used to locate the position information of the eyes in a face image. For example, valley point search (valley point is the minimum value point of a local area in a grayscale image) can be performed in the third face image. Then, low grayscale value points around the pupil can be found (i.e., the pupil position in the eye region can be determined). Then, through directional projection (directional projection refers to the accumulation of pixel values in the horizontal and vertical directions based on the direction information of the image gradient) to obtain a projection map, the position of large pixel change in the eye region can be found by observing the projection map (i.e., the position of the eyeball edge in the eye region). Finally, by comparing the pupil positions of the left and right eyes through eyeball symmetry, the pupil center position in the eye region can be determined, thus obtaining the final eye position (i.e., pupil position, eyeball edge position, and pupil center position).
[0073] In the specific implementation, as shown in Figure 3, the human eye image needs to be processed sequentially to remove noise, normalize, and analyze grayscale gradient information before a roughly stable human eye position can be obtained. Then, the precise position of the human eye is determined by various eye positioning analysis methods in order to obtain the eye opening and closing information.
[0074] In this embodiment, noise interference can be avoided by removing noise from the face image, and the impact of factors such as uneven projection of the acquisition environment and equipment can be reduced by normalizing the face image.
[0075] Step S20: Determine the position of the eyes in the face image, and extract the eye information in the face image based on the eye position.
[0076] It should be noted that eye information can include the distance between the upper and lower eyelids. When a person is tired or distracted, their eyes become heavy and dull, and symptoms such as eyelids drooping may occur, which can lead to a decrease in the distance between the upper and lower eyelids or even cause the eyes to close.
[0077] Step S30: Determine the user's eye opening and closing information within the preset time period based on the eye information, and determine the user's gaze information within the preset time period.
[0078] It should be noted that eye opening and closing information includes whether the user is focused or not; fixation information refers to the duration and orientation of the user's eye focus within a preset time period.
[0079] In one embodiment, determining the user's eye opening and closing information within the preset time period based on the eye information includes:
[0080] Determine the number of frames of the face image in the face video;
[0081] The number of target frames of the target face image of the user within a preset time period is determined based on the eye information, wherein the target face image is a face image with an eye opening degree less than the opening threshold;
[0082] Based on the number of frames and the number of target frames, the user's eye opening and closing information within the preset time period is determined.
[0083] It should be noted that the eye opening and closing threshold can be preset. When the eye opening and closing degree of a person is less than the threshold, it may be that the eye is about to close or the eye is tired. The eye opening and closing information can be determined by the ratio of the target number of frames to the number of frames. When the ratio of the target number of frames to the number of frames is greater than the ratio threshold, the eye opening and closing information can indicate that the user's attention is not focused enough. When the ratio of the target number of frames to the number of frames is less than or equal to the ratio threshold, the eye opening and closing information can indicate that the user's attention is focused.
[0084] In one embodiment, determining the user's eye opening and closing information within the preset time period based on the eye information includes:
[0085] Based on the eye information, determine the user's single eye-closing duration and total eye-closing duration within a preset time period;
[0086] Set a duration threshold;
[0087] Based on the duration of a single eye closure, the total duration of eye closure, and the duration threshold, the user's eye opening and closing information within the preset time period is determined.
[0088] It should be noted that the duration of a single eye closure refers to the duration of a user's continuous eye closure within a preset time period, while the total duration of eye closure refers to the sum of the durations of all single eye closures within the preset time period.
[0089] In the specific implementation, the duration threshold includes a first duration threshold and a second duration threshold. The first duration threshold is less than the second duration threshold. First, the duration of each single eye closure within a preset time period can be determined based on eye information and compared with the first duration threshold. If there is a single eye closure duration greater than the first duration threshold, the eye opening and closing information can be directly used to indicate that the user's attention is not focused enough. If there is no single eye closure duration greater than the first duration threshold, the total eye closure duration needs to be compared with the second duration threshold. If the total eye closure duration is greater than the second duration threshold, the eye opening and closing information can be used to indicate that the user's attention is not focused enough. If the total eye closure duration is less than or equal to the second duration threshold, the eye opening and closing information can be used to indicate that the user's attention is focused.
[0090] Step S40: Based on the eye opening and closing information and the gaze information, determine the user's attention monitoring results, and take intervention measures according to the attention monitoring results.
[0091] It should be noted that eye opening and closing information includes whether the user's attention is focused or not; fixation information refers to the duration and location of the user's eye focus within a preset time period, including whether the fixation time is short or long.
[0092] Understandably, when a person's attention is not focused, the focus of their eyes becomes unstable, often randomly jumping to surrounding objects or areas, indicating scattered thoughts and reduced sensitivity to the environment; when attention is not focused, the time the eyes can focus is shortened, usually only a few hundred milliseconds, and the eyes move quickly from one area to another without staying focused for long.
[0093] In practice, when eye opening / closing information indicates insufficient user attention and fixation information shows a short duration of eye focus, the attention monitoring result can be determined as very inattentive; when eye opening / closing information indicates focused user attention and fixation information shows a short duration of eye focus, the attention monitoring result can be determined as inattentive; when eye opening / closing information indicates insufficient user attention and fixation information shows a long duration of eye focus, the attention monitoring result can be determined as inattentive; when eye opening / closing information indicates focused user attention and fixation information shows a long duration of eye focus, the attention monitoring result can be determined as focused; intervention measures are required when the monitoring result is either very inattentive or inattentive.
[0094] In practice, intervention measures can be taken through VR / AR devices. These measures can include one or more of the following: highlighting moving points on the screen, converting screen information to speech, special effects animation, surround sound effects, magnifying local images, and issuing danger warnings.
[0095] In one embodiment, taking intervention measures based on the attention monitoring results includes:
[0096] Determine the user's usage scenario at the current moment;
[0097] When the usage scenario is a highly concentrated scenario, the user's work status is determined to be abnormal based on the attention monitoring results;
[0098] When the user's working state is determined to be abnormal, intervention measures are taken through the head-mounted device to restore the user's attention. The intervention measures include viewpoint enhancement, animation effects, surround sound effects, vibration alerts, and danger warnings.
[0099] It is understandable that the usage scenarios include general scenarios and highly focused scenarios. General scenarios can be ordinary work scenarios where VR / AR devices are used for tasks that do not require high attention. When the user's current usage scenario is a general scenario, the intervention measures taken when the user's attention is diverted or fatigue is detected can be to remind the user to stop working and take a break. Highly focused scenarios can be usage scenarios such as meetings and driving. When the user's attention is diverted or fatigue is detected, the intervention measures taken can be one or more of the following: viewpoint highlighting, animation effects, surround sound effects, vibration reminders, and danger warnings.
[0100] In this embodiment, by first determining the user's usage scenario and then monitoring the user's work status, intervention measures are taken in a timely manner when the user's work status is abnormal, so as to avoid affecting work efficiency due to the user's distraction.
[0101] This embodiment, when detecting that a user's usage time exceeds a preset time, acquires facial video of the user within a preset time period using a head-mounted device and obtains facial images from the facial video; determines the eye positions in the facial images and extracts eye information based on the eye positions; determines the user's eye opening and closing information and gaze information within the preset time period based on the eye information; determines the user's attention monitoring result based on the eye opening and closing information and the gaze information, and takes intervention measures based on the attention monitoring result. By collecting the user's dynamic eye information in this way, and combining eye opening and closing information and gaze information to determine the user's attention status, the accuracy of the monitoring results can be effectively improved. Timely intervention measures can be taken when the user's attention is detected to be distracted, further ensuring that the user's attention is focused when working.
[0102] Referring to Figure 4, Figure 4 is a flowchart illustrating a second embodiment of an attention monitoring method according to the present invention.
[0103] Based on the first embodiment described above, the attention monitoring method of this embodiment, in determining the user's gaze information within the preset time period, includes:
[0104] Step S301: Collect the user's EOG signal within the preset time period using a head-mounted device.
[0105] It should be noted that the EOG signal is an electrical signal caused by the potential difference between the retina and cornea of the human eye.
[0106] Step S302: Perform endpoint detection and frame segmentation on the acquired EOG signal to obtain the energy of each frame.
[0107] In practical implementation, an endpoint detection algorithm based on short-time energy can be used. The acquired EOG signal is processed by Hamming windowing and then divided into frames to obtain the energy of each frame (i.e., the square of the amplitude).
[0108] S w (n) = S(n) * W(n)
[0109]
[0110] Step S303: Extract the signal features of the EOG signal based on the energy of each frame.
[0111] It should be noted that the energy E of each frame is calculated sequentially starting from the beginning of the EOG signal. i (i = 1, 2, 3, ..., n), the energy of each frame is compared with the threshold. If it is continuously lower than the threshold setting, it is determined that the signal ends. The EOG signal characteristics are obtained after the extracted signal segment is filtered.
[0112] Step S304: Identify the user's eye movement based on the signal characteristics, and determine the user's gaze information within a preset time period based on the eye movement.
[0113] It should be noted that after extracting the EOG signal features, the SVM classification algorithm is used to classify eye movements, ultimately identifying the user's eye movements. Then, based on the identified eye movements, gaze information is determined. Specifically, a convolutional neural network model for eye movement recognition can be built based on the Tensorflow computing framework, and the model can be trained with the collected eye movement image dataset to improve the accuracy of recognition.
[0114] In one embodiment, determining the user's gaze information within a preset time period based on the eye movement includes:
[0115] After the user is illuminated by the infrared illuminator built into the head-mounted device, the infrared photosensitive device built into the VR / AR device receives the target reflected light, wherein the target reflected light is the reflected light from the edge of the iris and sclera of the human eye;
[0116] The reflected light from the target is converted into an eye movement signal, and the direction and amplitude of eye movement are determined based on the eye movement signal;
[0117] The user's gaze point position is determined based on the direction, the amplitude, and the user's head position;
[0118] Based on the fixation point location and the eye movement, the user's fixation information within a preset time period is determined.
[0119] It should be noted that while infrared light is invisible to the human eye when it shines on it, it can be reflected back to an infrared camera (i.e., an infrared photosensitive device). This device captures the reflected infrared light around the eye (i.e., target reflected light), pinpoints the eye's position within the eye socket, and tracks its movement. Based on changes in the reflected light received by the infrared photosensitive device, the position and orientation of the eye can be determined. In VR / AR devices, the eye-side infrared photosensitive device receives infrared light reflected from the edges of the iris and sclera and converts it into electrical signals. Based on the electrical signals generated by the infrared photosensitive device, a differential signal reflecting eye movement can be calculated. When the eye moves to one side, the differential signal changes. By analyzing these changes, the direction and amplitude of the eye movement can be determined. Based on the direction and amplitude of the eye movement, combined with the user's head posture while using the VR / AR device, the user's gaze point position—the position of the point the user is looking at while using the VR / AR device—can be calculated.
[0120] Understandably, the fixation point location is used to correct fixation information.
[0121] In the specific implementation, as shown in Figure 5, it is necessary to locate the eye position by using the face image and extract the pupil feature from the infrared image to fit the pupil position, and then perform gaze point mapping modeling to obtain the final gaze point position.
[0122] In practice, as shown in Figure 6, after wearing the VR / AR device, the user can choose whether to enable the VR / AR device working status intervention mode. Once the intervention mode is enabled, the user's attention can be monitored. By determining the user's eye opening and closing information and gaze information, the user's attention status can be comprehensively analyzed, and intervention measures can be taken when the user's attention is detected to be distracted.
[0123] This embodiment uses a VR / AR device to collect the user's EOG signals within a preset time period; it performs endpoint detection and frame segmentation on the collected EOG signals to obtain the energy of each frame; based on the energy of each frame, it extracts the signal features of the EOG signals; it identifies the user's eye movements based on the signal features, and determines the user's gaze information within the preset time period based on the eye movements. Through this method, the user's gaze information can be quickly determined using EOG signals, thereby effectively improving the user's monitoring speed.
[0124] Furthermore, this embodiment of the invention also proposes a storage medium storing an attention monitoring program, which, when executed by a processor, implements the steps of the attention monitoring method described above.
[0125] Referring to Figure 7, which is a structural block diagram of the first embodiment of the attention monitoring device of the present invention.
[0126] As shown in Figure 7, the attention monitoring device proposed in this embodiment of the invention includes:
[0127] The acquisition module 10 is used to acquire the user's facial video within a preset time period through a head-mounted device when the user's usage time is detected to be longer than a preset time period, and to acquire the facial image in the facial video.
[0128] The determining module 20 is used to determine the position of the eyes in the face image and extract the eye information in the face image based on the position of the eyes.
[0129] The determining module 20 is further configured to determine the user's eye opening and closing information within the preset time period based on the eye information, and to determine the user's gaze information within the preset time period.
[0130] The determining module 20 is further configured to determine the user's attention monitoring results based on the eye opening and closing information and the gaze information, and to take intervention measures based on the attention monitoring results.
[0131] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0132] This embodiment, when detecting that a user's usage time exceeds a preset duration, uses a VR / AR device to capture facial video of the user within a preset time period and obtains facial images from the video. It determines the eye positions within the facial images and extracts eye information based on these positions. Based on this eye information, it determines the user's eye opening and closing information and gaze information within the preset time period. Based on the eye opening and closing information and the gaze information, it determines the user's attention monitoring result and takes intervention measures accordingly. By collecting the user's dynamic eye information and combining eye opening and closing information with gaze information to determine the user's attention status, the accuracy of the monitoring results can be effectively improved. Furthermore, by taking timely intervention measures when the user's attention is detected to be distracted, it can further ensure that the user's attention is focused when working.
[0133] In one embodiment, the determining module 20 is further configured to:
[0134] Noise interference is removed from the face image to obtain the first face image;
[0135] The first face image is normalized based on the grayscale difference to obtain the second face image;
[0136] The grayscale distribution and gradient distribution of local regions in the second face image were statistically analyzed to obtain statistical results;
[0137] Based on the statistical results, non-eye information in the second face image is filtered out to obtain the third face image;
[0138] The position of the eyes in the third face image is determined by coarse localization.
[0139] In one embodiment, the determining module 20 is further configured to:
[0140] Determine the number of frames of the face image in the face video;
[0141] The number of target frames of the target face image of the user within a preset time period is determined based on the eye information, wherein the target face image is a face image with an eye opening degree less than the opening threshold;
[0142] Based on the number of frames and the number of target frames, the user's eye opening and closing information within the preset time period is determined.
[0143] In one embodiment, the determining module 20 is further configured to:
[0144] Based on the eye information, determine the user's single eye-closing duration and total eye-closing duration within a preset time period;
[0145] Set a duration threshold;
[0146] Based on the duration of a single eye closure, the total duration of eye closure, and the duration threshold, the user's eye opening and closing information within the preset time period is determined.
[0147] In one embodiment, the determining module 20 is further configured to:
[0148] The user's EOG signal is collected within the preset time period using a head-mounted device;
[0149] Endpoint detection and frame segmentation are performed on the acquired EOG signals to obtain the energy of each frame;
[0150] Based on the energy of each frame, the signal features of the EOG signal are extracted;
[0151] The user's eye movements are identified based on the signal characteristics, and the user's gaze information within a preset time period is determined based on the eye movements.
[0152] In one embodiment, the determining module 20 is further configured to:
[0153] After the user is illuminated by the infrared illuminator built into the head-mounted device, the infrared photosensitive device built into the VR / AR device receives the target reflected light, wherein the target reflected light is the reflected light from the edge of the iris and sclera of the human eye;
[0154] The reflected light from the target is converted into an eye movement signal, and the direction and amplitude of eye movement are determined based on the eye movement signal;
[0155] The user's gaze point position is determined based on the direction, the amplitude, and the user's head position;
[0156] Based on the fixation point location and the eye movement, the user's fixation information within a preset time period is determined.
[0157] In one embodiment, the determining module 20 is further configured to:
[0158] Determine the user's usage scenario at the current moment;
[0159] When the usage scenario is a highly concentrated scenario, the user's work status is determined to be abnormal based on the attention monitoring results;
[0160] When the user's working state is determined to be abnormal, intervention measures are taken through the head-mounted device to restore the user's attention. The intervention measures include viewpoint enhancement, animation effects, surround sound effects, vibration alerts, and danger warnings.
[0161] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0162] In addition, for technical details not described in detail in this embodiment, please refer to the attention monitoring method provided in any embodiment of the present invention, which will not be repeated here.
[0163] Furthermore, it should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0164] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0165] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0166] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. An attention monitoring method, characterized in that, The attention monitoring method includes: when the user's usage time exceeds a preset time, acquiring a facial video of the user within a preset time period using a head-mounted device, and obtaining a facial image from the facial video; determining the eye position in the facial image, and extracting eye information from the facial image based on the eye position; determining the user's eye opening and closing information within the preset time period based on the eye information; determining the user's gaze information within the preset time period, including acquiring the user's EOG signal within the preset time period using a head-mounted device; performing endpoint detection and frame segmentation processing on the acquired EOG signal to obtain the energy of each frame; extracting the signal features of the EOG signal based on the energy of each frame; and identifying the user's eye movement based on the signal features. The process involves: 1) Illuminating the user with an infrared illuminator built into the head-mounted device; 2) Receiving the reflected light from the target area, specifically the reflection from the edges of the iris and sclera of the human eye, via an infrared photosensitive device built into the head-mounted device; 3) Converting the reflected light into an eye movement signal and determining the direction and amplitude of the eye movement based on the signal; 4) Determining the user's fixation point position based on the direction, amplitude, and head position; 5) Determining the user's fixation information within a preset time period based on the fixation point position and eye movement, whereby the fixation information includes the dwell time and orientation of the eye's focal point within the preset time period; 6) Determining the user's attention monitoring result based on the eye opening / closing information and the fixation information, and then taking intervention measures accordingly.
2. The method as described in claim 1, characterized in that, Determining the eye position in the face image includes: removing noise interference from the face image to obtain a first face image; normalizing the first face image based on grayscale differences to obtain a second face image; statistically analyzing the grayscale distribution and gradient distribution of local regions in the second face image to obtain statistical results; filtering out non-eye information in the second face image based on the statistical results to obtain a third face image; and determining the eye position in the third face image through coarse localization.
3. The method as described in claim 1, characterized in that, The step of determining the user's eye opening and closing information within the preset time period based on the eye information includes: determining the number of frames of the face image in the face video; determining the number of target frames of the target face image of the user within the preset time period based on the eye information, wherein the target face image is a face image with an eye opening and closing degree less than an opening and closing threshold; and determining the user's eye opening and closing information within the preset time period based on the number of frames and the number of target frames.
4. The method as described in claim 3, characterized in that, The step of determining the user's eye opening and closing information within the preset time period based on the eye information includes: determining the duration of a single eye closure and the total duration of eye closure within the preset time period based on the eye information; setting a duration threshold; and determining the user's eye opening and closing information within the preset time period based on the duration of a single eye closure, the total duration of eye closure, and the duration threshold.
5. The method as described in claim 1, characterized in that, The intervention measures based on the attention monitoring results include: determining the user's usage scenario at the current moment; when the usage scenario is a highly focused scenario, determining whether the user's working state is abnormal based on the attention monitoring results; when the user's working state is determined to be abnormal, taking intervention measures on the user through the head-mounted device to restore the user's attention, wherein the intervention measures include viewpoint highlighting, animation effects, surround sound effects, vibration reminders, and danger warnings.
6. An attention monitoring device, characterized in that, The attention monitoring device includes: an acquisition module, used to acquire a facial video of the user within a preset time period via a head-mounted device when the user's usage time is detected to exceed a preset time period, and to acquire a facial image in the facial video; a determination module, used to determine the eye position in the facial image, and to extract eye information from the facial image based on the eye position; the determination module is further used to determine the user's eye opening and closing information within the preset time period based on the eye information; to determine the user's gaze information within the preset time period, including acquiring the user's EOG signal within the preset time period via the head-mounted device; performing endpoint detection and frame segmentation processing on the acquired EOG signal to obtain the energy of each frame; and extracting the signal features of the EOG signal based on the energy of each frame; and further... The system identifies the user's eye movements based on the signal characteristics; after illuminating the user with an infrared illuminator built into the head-mounted device, it receives the reflected light from the target area, specifically the reflected light from the edges of the iris and sclera of the human eye, via an infrared photosensitive device built into the head-mounted device; it converts the reflected light into an eye movement signal and determines the direction and amplitude of the eye movement based on the signal; it determines the user's fixation point position based on the direction, amplitude, and the user's head position; it determines the user's fixation information within a preset time period based on the fixation point position and the eye movements, wherein the fixation information is the dwell time and orientation information of the eye's focal point within the preset time period; the determining module is further configured to determine the user's attention monitoring results based on the eye opening and closing information and the fixation information, and to take intervention measures based on the attention monitoring results.
7. An attention monitoring device, characterized in that, The device includes: a memory, a processor, and an attention monitoring program stored in the memory and executable on the processor, the attention monitoring program being configured to implement the steps of the attention monitoring method as described in any one of claims 1 to 5.
8. A storage medium, characterized in that, The storage medium stores an attention monitoring program, which, when executed by a processor, implements the steps of the attention monitoring method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method for extracting and identifying characteristics of electro-ocular signal
CN101599127A
Attention judgment method, attention judgment device, attention judgment system, attention judgment equipment and storage medium
CN109902630A
CNN-based attention detection system and method
CN112836630A
Eye-tracking calibration
US20190204913A1