Attention state display method, device and equipment based on eye movement tracking equipment

By embedding eye-tracking devices and light source components compatible with a large field of view into virtual reality head-mounted display devices, and combining attention fixation network models and blink detection, the problem of small field of view in existing technologies has been solved, achieving more accurate and real-time display of attention status, and improving device battery life and user experience.

CN120973223APending Publication Date: 2025-11-18COMMUNICATION UNIVERSITY OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511035784.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In existing technologies, attention status display methods based on eye-tracking devices suffer from low-quality eye image acquisition with a field of view of less than 120 degrees, making it difficult to extract subtle features of eye movement. This results in high device wear and short battery life. Furthermore, the accuracy and comprehensiveness of determining attention status solely through pupil and gaze prediction information are insufficient, impacting user experience.

Method used

By embedding an eye-tracking device compatible with a large field of view into a virtual reality head-mounted display device, and combining it with a light source component to improve the quality of eye images, the system utilizes a trained attention fixation network model and blink detection to generate fixation point and blink frequency information, comprehensively determines the attention state, and then visualizes it.

Benefits of technology

It improves the accuracy and real-time performance of attention state detection, extends device battery life, enhances user experience, and reduces waste of data transmission resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973223A_ABST
    Figure CN120973223A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an attention state display method, device and equipment based on eye movement tracking equipment. A specific embodiment of the method comprises the following steps: controlling a virtual reality head-mounted display device embedded in an eye movement tracking device, and collecting an eye image sequence; performing pupil detection on the eye image sequence to obtain a pupil information sequence; inputting the eye image sequence into an attention gazing network model to obtain a gazing sight line information sequence; generating a fixation point information sequence; performing blink detection processing on the eye image sequence to obtain a blink frequency information sequence; generating attention state information; and visually displaying the attention state information, and sending and visually displaying the eye image sequence, the fixation point information sequence and the user attention image. According to the embodiment, by means of the low-power-consumption head-mounted display device compatible with the large field angle and the experience feeling, the accuracy and the real-time performance of attention state detection can be improved, the endurance time of the device is prolonged, and the user experience feeling is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically to a method, apparatus, and device for displaying attention states based on eye-tracking devices. Background Technology

[0002] Attention state detection technology involves capturing and analyzing human eye movement trajectories using low-power, portable virtual reality head-mounted displays to infer the user's gaze point and direction of gaze in real time, thereby detecting the user's attention status. For displaying attention states based on eye-tracking devices, the typical approach is as follows: Eye image sequences are acquired using a virtual reality head-mounted display with a small field of view (less than 120 degrees). Pupil detection and gaze prediction are then performed on these eye image sequences to obtain pupil information sequences and gaze prediction information sequences. Next, the pupil-corneal reflex method is used to estimate the gaze point from the pupil information sequences and gaze prediction information sequences, resulting in a gaze point information sequence. Finally, the attention state information is determined using the gaze point information sequence, and then visualized and transmitted.

[0003] However, in practice, it has been found that when displaying attention states based on eye-tracking devices using the above methods, the following technical problems often arise: Firstly, because it can only acquire eye image sequences with a field of view of less than 120 degrees, and image acquisition is only performed using the light source of the video played by the device, it is difficult to adapt to and detect eye images with a field of view exceeding 120 degrees. The quality of the acquired eye images is low, making it difficult to extract subtle features of eye movement. This forces low-power virtual reality headsets to frequently acquire images, increasing device wear and reducing battery life. Secondly, determining attention state information solely based on acquired pupil information and gaze prediction information is too simplistic in considering the factors influencing user attention state information, resulting in low accuracy and comprehensiveness of the generated attention state information. This further leads to low accuracy of the generated attention state images, requiring frequent data transmission, wasting transmission resources, and reducing user experience.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure propose a method, apparatus, and device for displaying attention state based on an eye-tracking device to solve one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a method for displaying attention state based on an eye-tracking device, comprising: controlling a virtual reality head-mounted display device embedded with an eye-tracking device to acquire a sequence of eye images for a target user; performing pupil detection on the eye image sequence to obtain a pupil information sequence; inputting the eye image sequence into a trained attention gaze network model to obtain a gaze line information sequence; generating a gaze point information sequence based on the pupil information sequence and the gaze line information sequence; performing blink detection processing on the eye image sequence to obtain a blink frequency information sequence; generating attention state information for the target user based on the pupil information sequence, the gaze point information sequence, and the blink frequency information sequence; visualizing the attention state information to obtain a user attention image; and sending the eye image sequence, the gaze point information sequence, and the user attention image to a terminal page for visual display.

[0008] Secondly, some embodiments of this disclosure provide an attention state display device based on an eye-tracking device, comprising: a control unit configured to control a virtual reality head-mounted display device embedded with an eye-tracking device to acquire an eye image sequence for a target user; a pupil detection unit configured to perform pupil detection on the eye image sequence to obtain a pupil information sequence; an input unit configured to input the eye image sequence into a trained attention gaze network model to obtain a gaze line information sequence; a first generation unit configured to generate a gaze point information sequence based on the pupil information sequence and the gaze line information sequence; a blink detection unit configured to perform blink detection processing on the eye image sequence to obtain a blink frequency information sequence; a second generation unit configured to generate attention state information for the target user based on the pupil information sequence, the gaze point information sequence, and the blink frequency information sequence; and a visualization display unit configured to visualize the attention state information to obtain a user attention image, and to send the eye image sequence, the gaze point information sequence, and the user attention image to a terminal page for visualization display.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0011] The various embodiments of this disclosure have the following beneficial effects: The attention state display method based on eye-tracking devices in some embodiments of this disclosure, by being compatible with low-power head-mounted display devices with large field of view and good user experience, can improve the accuracy and real-time performance of attention state detection, increase device battery life, and enhance user experience. Specifically, the reasons for increased device wear and tear, reduced battery life, lower accuracy and comprehensiveness of attention state information, and increased waste of transmission resources are as follows: Since it can only acquire eye image sequences with a field of view less than 120 degrees, and image acquisition is only performed using the light source of the video played by the device, it is difficult to adapt to and detect eye images with a large field of view exceeding 120 degrees. The quality of the acquired eye images is low, making it difficult to extract subtle features of eye movement. This causes the low-power virtual reality head-mounted display device to frequently acquire images, leading to increased device wear and reduced battery life. Meanwhile, determining attention state information solely based on collected pupil information and gaze prediction information is too simplistic, resulting in low accuracy and comprehensiveness of the generated attention state information. This further leads to low accuracy of the generated attention state images, necessitating frequent data transmission, wasting transmission resources, and reducing user experience. Therefore, some embodiments of this disclosure employ an eye-tracking device-based attention state display method. First, it controls a virtual reality head-mounted display device embedded with an eye-tracking device to collect a sequence of eye images for the target user. Here, since the eye-tracking device is adaptable to virtual reality head-mounted display devices with large field of view and includes a light source component, it can improve the quality and comprehensiveness of the collected eye images, enhance the adaptability of the virtual reality head-mounted display device, facilitate subsequent extraction of subtle eye movement features, reduce the number of repeated acquisitions, and decrease device wear and battery life. Second, pupil detection is performed on the aforementioned eye image sequence to obtain a pupil information sequence. Here, pupil information detection can extract the pupil position information of the target user, facilitating subsequent determination of the target user's gaze point location and improving the accuracy of pupil information detection. Next, the aforementioned eye image sequence is input into the trained attention-based gaze network model to obtain a gaze information sequence. Here, the powerful learning ability of the attention-based gaze network model improves the accuracy of gaze information detection, enabling subsequent determination of the target user's gaze location. Then, based on the aforementioned pupil information sequence and gaze information sequence, a gaze point information sequence is generated. Here, accurate pupil and gaze information improves the accuracy of the gaze point information. Finally, blink detection processing is performed on the aforementioned eye image sequence to obtain a blink frequency information sequence.Here, blink frequency detection can detect the user's eye movement information in real time, and blink frequency can further reflect the target user's attention state information, facilitating subsequent determination of the target user's attention state information. Then, based on the aforementioned pupil information sequence, fixation point information sequence, and blink frequency information sequence, attention state information for the target user is generated. Here, combining multiple influencing factors to determine the target user's attention state information can more comprehensively determine the target user's attention state information, improving the accuracy and real-time performance of the attention state information. Finally, the aforementioned attention state information is visualized to obtain a user attention image, and the aforementioned eye image sequence, fixation point information sequence, and user attention image are sent to the terminal page for visualization. This improves the accuracy and quality of the user attention image, reduces the number of data transmissions and the amount of data, reduces the waste of transmission resources, and improves the efficiency of real-time display. Therefore, this attention state display method based on eye-tracking devices, by being compatible with low-power head-mounted display devices with large field of view and good user experience, can improve the accuracy and real-time performance of attention state detection, increase device battery life, and improve user experience. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 These are flowcharts of some embodiments of the attention state display method based on an eye-tracking device according to the present disclosure;

[0014] Figure 2 This is a schematic diagram of the structure of the image acquisition component included in the eye-tracking device in some embodiments of the attention state display method based on the eye-tracking device according to the present disclosure;

[0015] Figure 3 This is a schematic diagram of the circuit of the light source component included in the eye-tracking device in some embodiments of the attention state display method based on the eye-tracking device according to the present disclosure;

[0016] Figure 4 This is a schematic diagram of the top structure in a double-layer structure of the housing structure component included in some embodiments of the attention state display method based on the eye-tracking device according to the present disclosure;

[0017] Figure 5This is a schematic diagram of the bottom structure in a double-layer structure of the housing structure component included in some embodiments of the attention state display method based on the eye-tracking device according to the present disclosure;

[0018] Figure 6 This is an overall schematic diagram of the housing structure components included in some embodiments of the attention state display method based on an eye-tracking device according to the present disclosure;

[0019] Figure 7 This is a schematic diagram of the structure of some embodiments of an attention state display device based on an eye-tracking device according to the present disclosure;

[0020] Figure 8 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0021] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0022] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0023] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0024] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0025] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0026] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0027] Figure 1A flow 100 of some embodiments of an attention state display method based on an eye-tracking device according to the present disclosure is shown. This attention state display method based on an eye-tracking device includes the following steps:

[0028] Step 101: Control the virtual reality head-mounted display device with embedded eye-tracking equipment to acquire eye image sequences for the target user.

[0029] In some embodiments, the execution entity (e.g., an electronic device) of the above-described attention state display method based on an eye-tracking device can control a virtual reality head-mounted display device embedded with an eye-tracking device via a wired or wireless connection to acquire a sequence of eye images for a target user. The virtual reality head-mounted display device can be a device that displays a constructed three-dimensional virtual scene in front of the target user, providing an immersive experience with a field of view exceeding 120 degrees. For example, the virtual reality head-mounted display device could be the Pimax 5K Plus virtual reality headset. The eye-tracking device can be a device embedded in the virtual reality head-mounted display device that uses an infrared light source for supplemental lighting to acquire reflected light from the target user's eyes in low-light scenes and a camera to acquire a sequence of eye images for the target user. The target user can be a user wearing a virtual reality head-mounted display device to acquire eye images of a single eye. The eye images in the eye image sequence can be images that only include the eye area. The eye image sequence can be a sequence of images with temporal information captured within a preset time period. The preset time can be a pre-set time, for example, 5 minutes.

[0030] In some optional implementations of certain embodiments, the eye-tracking device includes: an image acquisition component, multiple light source components, and a housing structure component. The image acquisition component may be a component that acquires images using a camera sensor. Specifically, the image acquisition component may employ an OV7251-2B global shutter image sensor, acquiring 120 frames per second. Figure 2 The diagram shows a schematic of the image acquisition component. The light source components mentioned above can be used to supplement the lighting of a virtual reality head-mounted display device. Due to the enclosed structure of the virtual reality head-mounted display device, the target user's eyes are completely surrounded by light, and the image brightness is determined solely by the brightness of the virtual reality scene. This results in low brightness in the video captured by the image acquisition component, affecting the accuracy of subsequent pupil detection. Therefore, supplementary lighting is needed. These light source components can be eight near-infrared LEDs (Light-Emitting Diodes) with a wavelength of 850nm and a power of 0.2W. Figure 3 As shown, Figure 3 A circuit diagram of a light source assembly is shown. The circuit diagram may include: a light source driving unit, a voltage regulation unit, and a power supply interface unit. The light source driving unit may be a unit used to provide power to the LEDs to ensure consistent output light intensity from each light source. The voltage regulation unit may be a unit using an AP1117 chip to precisely regulate the power supply voltage. The power supply interface unit may be a unit using a Micro USB interface to introduce power and achieve stable operation of eight LEDs. The housing structure assembly may be a component for stably embedding the aforementioned virtual reality head-mounted display device in a double-layered housing structure. The housing structure assembly may include: a top housing assembly and a bottom housing assembly. The top housing assembly may have openings for the lens of the image acquisition assembly and the housing mechanism assembly to ensure image acquisition clarity and data transmission accuracy. Figure 4 The diagram shows a schematic of the top assembly of the housing. The bottom assembly of the housing may be a component that contacts the Fresnel lens of the virtual reality headset to ensure that the virtual reality headset does not shift due to movement during use. Figure 5 The diagram shown illustrates the bottom components of the housing. Figure 6 The diagram shown is a schematic of the shell structure components.

[0031] It should be noted that the dual-layer structure of the aforementioned shell components not only enhances the stability of the virtual reality headset but also improves the wearing comfort of the target user, reducing the pressure on the user's eyes caused by the eye-tracking device. The aforementioned light source components not only ensure uniform illumination of the eye area but also effectively reduce shadow interference from the target user's face, thereby improving the quality of the acquired eye image sequence.

[0032] Optionally, the virtual reality head-mounted display device described above, which controls the embedded eye-tracking device, may acquire eye image sequences for a target user by including the following steps:

[0033] The first step is to determine the set of mounting area location information for the image acquisition component on the housing structure component. The mounting area location information in this set can be information about the location of the area on the housing structure component where the image acquisition component can be mounted. For example, the set of mounting area location information may include, but is not limited to, at least one of the following: location information for the area directly above the eyes, location information for the area to the left front of the eyes, location information for the area to the right front of the eyes, and location information for the bridge of the nose.

[0034] The second step involves filtering the installation area location information set to select at least one installation area location that meets the installation structure conditions. These installation structure conditions can include meeting anatomical analysis requirements for comfortable human wear and ensuring no obstruction by objects. The at least one installation area location can be any of the following: the area directly above the eyes, where dynamic occlusion by eyelashes can lead to unstable image quality when the camera is positioned directly above the user, and the area near the bridge of the nose, which can cause pressure and negatively impact user experience.

[0035] Thirdly, based on the field-of-view information of the virtual reality head-mounted display device and the offset constraint information corresponding to the image acquisition component, the location information of at least one installation area is further filtered to obtain the target installation location information. The target installation location information can be specific two-dimensional spatial coordinates obtained by mapping the field-of-view information of the virtual reality head-mounted display device. The field-of-view information can be the maximum angle range that a target user can view a virtual reality scene through the virtual reality head-mounted display device. The offset constraint information can be the maximum offset information that the image acquisition component allows for the maximum pupil movement of the target user.

[0036] As an example, the aforementioned execution entity can use a field-of-view constraint function to perform secondary filtering on at least one of the installation location information based on the field-of-view information and offset constraint information to obtain the target installation location information. The aforementioned field-of-view constraint function can be:

[0037]

[0038] Where Δx represents the horizontal axis position information with the pupil of the target user's eye as the origin and the direction of the right eye as the positive direction. max This represents the maximum horizontal axis offset of the image acquisition component relative to the center point of the pupil. h The attenuation coefficient, representing the length of the Fresnel lens in the virtual reality head-mounted display device, is 0.02. FOV represents the field of view information of the aforementioned virtual reality head-mounted display device. Δy represents the vertical axis position information with the pupil of the target user's eye as the origin and the direction of the right eye as the positive direction. max This represents the maximum vertical axis offset of the image acquisition component relative to the center point of the pupil. v The attenuation coefficient represents the width of the Fresnel lens and has a value of 0.01.

[0039] The fourth step involves uniformly installing the multiple light source components onto the housing structure component, and, based on the target installation position information, installing the image acquisition component onto the housing structure component to obtain the eye-tracking device. The uniform installation can be achieved by using the ratio of the perimeter of the housing structure component to the number of light source components as the spacing between the multiple light source components.

[0040] The fifth step involves controlling the eye-tracking device to be embedded into the virtual reality head-mounted display device, and controlling the virtual reality head-mounted display device to collect eye-tracking videos of the target user. The eye-tracking video can be a video recording the movement trajectory information of a single eye of the target user. The eye-tracking video can be a video captured at 120 frames per second.

[0041] Step 6: Crop eye images from the eye-tracking video frame sequence corresponding to the aforementioned eye-tracking video to obtain an eye image sequence. Each eye-tracking video frame in the sequence can be a video frame included in the aforementioned eye-tracking video. Each eye image in the eye image sequence can be an image where each frame only includes the eye region. In practice, the execution entity can first perform grayscale processing on the aforementioned eye-tracking video frame sequence using a weighted average algorithm to obtain a grayscale eye-tracking video frame sequence. The weighted average algorithm can be the sum of the products of 0.114 and the blue channel pixel value, 0.587 and the green channel pixel value, and 0.299 and the red channel pixel value, used as the grayscale pixel value. Then, an eye region localization model is used to localize the eye region in the aforementioned eye-tracking video frame sequence to obtain an eye image sequence. The eye region localization model can be a pre-trained deep neural network model that performs eye region localization and cropping on the input eye-tracking video frames. For example, the eye region localization model described above could be a Haar Cascade Classifier for eye detection. The scale factor of this Haar Cascade Classifier could be set to 1.1, meaning a scaling factor of 10% each time the image pyramid is constructed. The minimum neighborhood number of this Haar Cascade Classifier could be set to 5 to ensure each detection window has sufficient neighborhood support, thereby reducing the false detection rate. The minimum detection region size of this Haar Cascade Classifier could be set to (30, 30) pixels to filter out excessively small non-target regions.

[0042] Step 102: Perform pupil detection on the eye image sequence to obtain a pupil information sequence.

[0043] In some embodiments, the executing entity may perform pupil detection on the aforementioned eye image sequence to obtain a pupil information sequence. The pupil information in the pupil information sequence may be geometric information describing the pupil. This pupil information may include: the position information of the pupil's center point and the pupil's diameter.

[0044] In some optional implementations of certain embodiments, the above-described pupil detection of the eye image sequence to obtain a pupil information sequence may include the following steps:

[0045] The first step is to perform the following filtering steps for each eye image in the above eye image sequence:

[0046] Sub-step 1 involves performing edge detection processing on the aforementioned eye image to obtain an eye edge contour set. The eye edge contours in this set can be images representing the edge information of the eye image. In practice, the executing entity can first perform Gaussian filtering on the eye image to obtain a filtered eye image. Then, it retains at least one pixel value less than or equal to a binarization threshold from all the pixel values ​​in the eye image, and sets multiple pixel values ​​greater than the binarization threshold to 255, resulting in a thresholded eye image. The binarization threshold can be obtained through threshold statistical analysis of the grayscale histograms of a large number of collected eye images. For example, the binarization threshold could be 62 pixels. Next, morphological closing operations are used to optimize the thresholded eye image, resulting in an optimized eye image. Finally, the Canny algorithm is used to perform edge detection on the optimized eye image to obtain an eye edge contour set.

[0047] Sub-step 2 involves performing elliptical contour fitting on the aforementioned eye edge contour set to obtain an eye ellipse information set. This set includes: an eye ellipse image set, an eye ellipse center coordinate information set, an eye ellipse major axis information set, and an eye ellipse minor axis information set. The eye ellipse information in this set can be the information of the ellipse corresponding to the pupil. The eye ellipse images in the eye ellipse image set can be images that characterize the pupil, obtained by fitting the ellipse shape of the pupil. The eye ellipse center coordinate information in the eye ellipse center coordinate information set can be the position information of the center point of the pupil ellipse.

[0048] Sub-step 3: Based on the aforementioned set of information on the major axis and minor axis of the eye ellipse, determine the set of eye ellipse area and the set of eye ellipse aspect ratios from the aforementioned set of eye ellipse information. In practice, the executing entity can first input half of the set of information on the major axis and half of the set of information on the minor axis of the eye ellipse into the ellipse area formula to obtain the set of eye ellipse area. Then, determine the ratio of each major axis information of the eye ellipse in the aforementioned set of information on the major axis of the eye ellipse to the corresponding minor axis information in the aforementioned set of information on the minor axis of the eye ellipse, and use this ratio as the eye ellipse aspect ratio, thus obtaining the set of eye ellipse aspect ratios.

[0049] Sub-step 4: Based on the aforementioned set of eye ellipse areas and the aforementioned set of eye ellipse aspect ratios, the aforementioned set of eye ellipse images is filtered to obtain a filtered set of eye ellipse images. Specifically, the filtered eye ellipse images in the aforementioned set of filtered eye ellipse images can be eye ellipse images that satisfy the following conditions: the eye ellipse area ranges from [160, 3000] pixels, and the eye ellipse aspect ratio ranges from [0.5, 2].

[0050] Sub-step 5, in response to determining that the filtered set of eye elliptical images includes at least two filtered eye elliptical images, and that the eye images do not meet the initial eye image conditions, filters the set of eye elliptical center coordinate information corresponding to the filtered set of eye elliptical images based on the previous pupil center information corresponding to the previous eye image, to obtain pupil center information. The initial eye image conditions may be the condition that the eye image is the first eye image in the eye image sequence.

[0051] As an example, the aforementioned executing entity may first, in response to determining that the filtered set of eye elliptical images includes at least two filtered eye elliptical images, and that the eye images do not meet the initial eye image conditions, determine the center distance information between each eye ellipse center coordinate information in the eye ellipse center coordinate information set and the center distance information of the previous pupil center information, thus obtaining a center distance information set. Then, the eye ellipse center coordinate information corresponding to the center distance information with the smallest value is selected from the aforementioned center distance information set as the pupil center information.

[0052] Sub-step 6: In response to determining that the above-filtered set of elliptical eye images is empty and that the above eye images do not meet the conditions of the initial eye images, the pupil center information of the previous pupil is determined as the pupil information of the eye image.

[0053] Sub-step 7: Generate pupil information based on the aforementioned pupil center information, the aforementioned set of major axis information for the eye ellipse, and the aforementioned set of minor axis information for the eye ellipse. In practice, the executing entity can first determine the sum of the major axis information and minor axis information of the eye ellipse corresponding to the aforementioned pupil center information, and the ratio of this sum to 2, as the diameter information of the eye ellipse. Then, the aforementioned diameter information of the eye ellipse and the aforementioned pupil center information are used to determine the pupil information.

[0054] Optionally, in response to determining that the aforementioned eye images satisfy the initial eye image conditions, the elliptical eccentricity, the distance from the corner of the eye to the eye, and the mean grayscale pixel value of each filtered elliptical eye image in the filtered eye image set are determined, resulting in an elliptical eccentricity set, a corner-of-eye distance information set, and a set of mean grayscale pixel values. The elliptical eccentricity characterizes whether the filtered elliptical eye image is close to a circle; the larger the value, the closer it is to a circle. The elliptical eccentricity can be the arithmetic square root of the difference between 1 and the square of the ratio of the major axis to the minor axis of the filtered elliptical eye image. Then, the elliptical eccentricity set, the corner-of-eye distance information set, the set of mean grayscale pixel values, and the elliptical area set are summed with fixed weights to obtain an elliptical weight set. The fixed weights can be pre-set weight values. For example, the fixed weights could be 0.3 for the elliptical eccentricity, 0.4 for the mean grayscale pixel value, 0.2 for the corner-of-eye distance information, and 0.1 for the elliptical area. Finally, the eye ellipse center coordinates and pupil diameter information corresponding to the ellipse weight with the largest value in the above ellipse weight set are selected as pupil information.

[0055] In addressing the technical problems mentioned in the background section, the following technical challenge often arises: how to accurately extract pupil information to precisely identify the user's gaze direction within a wide field of view, thereby determining the user's attention state. A conventional solution to this problem typically involves using the least squares algorithm to fit an ellipse to the pupil area, obtaining pupil information to determine the user's state. However, this conventional solution still suffers from the following issues: the least squares method is susceptible to outliers during pupil fitting, and the edge points involved in the fitting may not be true edge points, leading to low accuracy in the fitted pupil and consequently, low accuracy in generating the user attention image, wasting significant transmission resources. We have decided to adopt the following solution:

[0056] Optionally, the above-mentioned pupil detection of the eye image sequence to obtain a pupil information sequence may include the following steps:

[0057] First, for each eye image in the above eye image sequence, perform the following first determination step:

[0058] Sub-step 1: In response to determining that the eye image is the eye image located at the initial position in the aforementioned eye image sequence, pupil region detection is performed on the aforementioned eye image to obtain pupil detection bounding box information. The pupil detection bounding box information can be information about the rectangular box that selects the location of the pupil. The pupil region detection can be performed using the YOLO model.

[0059] Sub-step 2: Based on the pupil detection frame information, perform the following second determination step:

[0060] The first sub-step involves determining the pupil center seed point information based on the pupil detection box information. This pupil center seed point information can be the information of the initial edge points of the ellipse corresponding to the fitted pupil. As an example, the execution entity can first determine the detection box center point and half the detection box width as the detection box center point and target detection box width, respectively. Secondly, it can determine the region extending downwards from the detection box center point, including the neighboring pixel set. Thirdly, it can determine the average gray value of the gray value set of the neighboring pixel set. Subsequently, it can determine the average gray value of at least one neighboring pixel less than the average gray value from the neighboring pixel set, as the target average gray value. Then, in response to determining that the gray value corresponding to the detection box center point is less than the target average gray value, the detection box center point is determined as the pupil center seed point. Finally, in response to determining that the gray value corresponding to the center point of the detection box is greater than or equal to the target gray value, the neighboring pixel located at the median position among the at least one neighboring pixel is determined as the pupil center seed point.

[0061] The second sub-step involves determining the set of connected region outer contour points within the connected region set containing the pupil center seed point. The connected region outer contour point information in this set can be the contour point information of the pupil contour where the pupil center seed point is located. In practice, the executing entity can first binarize the eye image using the difference and sum of the grayscale values ​​corresponding to the pupil center seed point and a preset grayscale threshold as the grayscale value range, obtaining a binarized eye image. In this binarized eye image, grayscale values ​​within the grayscale value range can be set to 255, and grayscale values ​​outside the grayscale value range can be set to 0. The preset grayscale threshold can be a pre-set value. This preset grayscale threshold can be determined based on specific circumstances and is not limited here. Next, an opening operation is performed on the binarized eye image to obtain a processed eye image. Finally, the outer contour of the processed eye image is extracted to obtain an eye contour set. Next, the set of contour points intersecting the horizontal line passing through the aforementioned pupil center seed point and the aforementioned eye contour set is determined. Subsequently, the aforementioned eye contour set is sorted from largest to smallest contour length to obtain an eye contour sequence. Following the order of the aforementioned eye contour sequence, the sequence of intersection contour points and the sequence of contour point distances of the aforementioned pupil center seed point are determined sequentially. When the eye contour point set corresponding to the smallest contour point distance is selected from the aforementioned contour point distance sequence, it is determined as the contour point information set outside the connected region, and the execution step of determining the contour point distance ends.

[0062] The third sub-step involves performing pupil ellipse fitting on the set of outer contour points of the connected region to obtain pupil ellipse information. This pupil ellipse information can be the information of the ellipse corresponding to the pupil in the aforementioned eye image. This pupil ellipse information can include the center of the pupil ellipse, its major axis, and its minor axis. The pupil ellipse fitting process can be performed using the least squares method.

[0063] The fourth sub-step involves determining the set of outline point information for the connected region's outer contour points and the set of outline point weighted distance information for the pupil ellipse information. The outline point weighted distance information in this set can be the distance between the outline point information for the connected region's outer contour points and the center point of the pupil ellipse information after assigning different weights. In practice, the executing entity can first determine the absolute values ​​of the outline point distance information sets for the connected region's outer contour points and the pupil ellipse information, using these as the set of absolute outline point distance values. Then, it can filter out at least one outline point information for the connected region from the set of outline point information for the connected region whose absolute outline point distance value is less than or equal to a distance fluctuation factor, using this as the target outline point information set. The distance fluctuation factor can be a threshold used to determine outliers. For example, the distance fluctuation factor can be 0.5. Finally, the set of outline point weighted distance information is determined by the difference between a preset value and the square of the ratio of the outline point distance information corresponding to each target outline point information in the target outline point information set to the distance fluctuation factor. The preset value can be 1. Finally, the weight of the outer contour point information set of the connected region after removing the above target outer contour point information set is set to 0.

[0064] The fifth sub-step involves generating a pupil detection objective function for the contour point information set outside the connected region, based on the contour point weight distance information set. This pupil detection objective function can be the product of a weight function corresponding to the contour point weight distance information set and a least squares algorithm, or an elliptic constraint function. The elliptic constraint function can be the product of four times the major axis of the ellipse and half the distance to the focal point, with the difference between the product and the square of the minor axis equal to 1. This elliptic constraint function indicates that the fit is an ellipse, not a hyperbola or parabola.

[0065] The sixth sub-step involves iteratively optimizing the pupil detection objective function based on the pupil ellipse information to obtain pupil center information, pupil diameter information, and pupil deflection angle information, which are then used as pupil information. As an example, the execution entity can use the center point, major axis, and minor axis of the pupil ellipse information as the initial values ​​of the pupil detection objective function, and utilize the L-BFGS-B (Limited-memory Broyden-Fletcher-Goldfarb-Shanno with Boxconstraints) algorithm to iteratively optimize the pupil detection objective function to obtain the pupil information. In the iterative optimization process, the set of contour point weight distance information dynamically changes with the number of iterations. The iteration condition for the iterative optimization process can be that the difference between the function value of the pupil detection function in the current iteration and the function value in the previous iteration is less than a preset convergence threshold. For example, the preset convergence threshold could be 10 to the power of -6.

[0066] Sub-step 2, in response to determining that the eye image is not located at the initial position in the aforementioned eye image sequence, determines the minimum bounding rectangle information corresponding to the connected region outer contour point information set of the previous frame eye image as the pupil detection box information of the next frame eye image, so as to execute the aforementioned second determination step again, and generate and send a user attention image based on the obtained pupil information. The minimum bounding rectangle information can be the rectangle information containing the connected region outer contour point information set and having the smallest area.

[0067] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem mentioned in the background art: "Because the least squares method is easily affected by outliers when fitting pupils, and the edge points involved in the fitting may not be real edge points, the accuracy of the fitted pupils and the accuracy of the generated user attention images are low, thus wasting a lot of transmission resources." The factors leading to low accuracy of the fitted pupils and the generated user attention images, and thus wasting a lot of transmission resources, are often as follows: the least squares method is easily affected by outliers when fitting pupils, and the edge points involved in the fitting may not be real edge points. If these factors are resolved, the accuracy of the fitted pupils and the accuracy of the generated user attention images can be improved, reducing the waste of a large amount of transmission resources. To achieve this effect, this disclosure first determines whether the current eye image is located at the initial position. If it is, pupil region detection is performed on the eye image, which improves the accuracy of initial position detection and thus enhances the efficiency and accuracy of pupil region detection in subsequent eye images. If it is not located at the initial position, the minimum bounding rectangle information of the previous frame's eye image is used as the pupil detection box information for the current frame. This improves reusability and the real-time and continuous processing of eye image sequences, while reducing computational load. Second, a downward search is performed based on the center point of the pupil detection information to determine the pupil center seed point. This can, to some extent, prevent the seed point from falling into non-pupil regions such as white spots or eyelids in the eye image, ensuring the authenticity of the seed point and improving the accuracy of the subsequent connected region outer contour point information set. Third, the connected region outer contour point information set is determined by the length order of each eye contour after binarization of the eye image. This removes a large number of outliers, preserves the pupil contour, and determines the connected region outer contour points sequentially by ellipse length order, which can terminate the calculation early, reducing computational load and improving retrieval efficiency. Subsequently, pupil ellipse fitting is performed to obtain pupil ellipse information, which can be used as the initial value for subsequent iterative optimization. Then, using the pupil ellipse information, the contour point weight distance information set and the pupil detection objective function are iteratively optimized. The contour point weight distance information set can effectively reduce the interference of outliers and suppress the influence of false edges. The pupil detection objective function can highlight the importance of valid points and ensure the correctness of the geometric model. Iterative optimization can avoid the occurrence of unreasonable parameter values ​​to a certain extent, further improving the accuracy of pupil detection and the accuracy of user attention images, and reducing the waste of a large amount of transmission resources.

[0068] Step 103: Input the eye image sequence into the trained attention gaze network model to obtain the gaze information sequence.

[0069] In some embodiments, the execution entity may input the aforementioned eye image sequence into a trained attention-gaze network model to obtain a gaze information sequence. The gaze information in this sequence may include the direction of the target user's gaze and its position on the video. The attention-gaze network model may be a deep neural network model that estimates gaze vectors from the input eye image sequence and outputs gaze information. For example, the attention-gaze network model may be MobileViT (Mobile Vision Transformer).

[0070] In some optional implementations of certain embodiments, the process of inputting the aforementioned eye image sequence into a trained attention-focusing network model to obtain a gaze information sequence may include the following steps:

[0071] The first step involves inputting the aforementioned eye image sequence into the convolutional pooling network included in the trained attention-focusing network model to obtain the first eye feature map set. The trained attention-focusing network model further includes: a first-stage residual attention network, a second-stage residual attention network, a third-stage residual attention network, a fourth-stage residual attention network, and a fully connected output network. The convolutional pooling network can be a network consisting of convolutional layers with a 7x7 kernel and a stride of 2, and max-pooling layers with a 2x2 kernel and a stride of 2.

[0072] The second step involves inputting the first eye feature map set into the first-stage residual attention network to obtain the second eye feature map set. The first-stage residual attention network can be a deep neural network comprising three cascaded residual convolutional networks and a gaze attention mechanism network. The cascaded residual convolutional networks can be a first convolutional layer with a 1x1 kernel and a stride of 1, a second convolutional layer with a 3x3 kernel and a stride of 1, and a third convolutional layer with a 1x1 kernel and a stride of 1, wherein the residual convolutional network is output after feature fusion of the first eye image feature map and the feature map output from the third convolutional layer. The aforementioned gaze attention mechanism network can be implemented as follows: First, a first eye feature map is compressed using a one-dimensional average pooling layer and a one-dimensional max pooling layer to obtain a pooled eye feature map. Then, a 1*1 convolutional layer further compresses the pooled eye feature map to obtain an eye channel feature map with one channel. Next, a global average pooling layer is used sequentially to pool the eye channel feature map, obtaining a class weight vector representing the importance of the class. Subsequently, the class weight vector is input into a multilayer perceptron consisting of two fully connected layers to obtain a mapped feature vector. Then, a sigmoid function is used to compress each element in the mapped feature vector to a feature vector between (0, 1), and this feature vector is then reconstructed with the first eye image feature map to output a second eye feature map. It should be noted that the gaze attention mechanism network can enhance the features most useful for class classification at the current stage in the feature map set, while suppressing redundant or irrelevant features, thereby improving the network's representation ability and generalization performance.

[0073] The third step involves inputting the second eye feature map set into the second-stage residual attention network to obtain the third eye feature map set. The second-stage residual attention network can be a deep neural network comprising an eight-stage residual convolutional network and a gaze attention mechanism network.

[0074] The fourth step involves inputting the aforementioned third eye feature map set into the aforementioned third-stage residual attention network to obtain the fourth eye feature map set. The aforementioned third-stage residual attention network can be a deep neural network comprising a 36-stage residual convolutional network and a gaze attention mechanism network.

[0075] Fifth, the fourth eye feature map set is input into the fourth-stage residual attention network to obtain the fifth eye feature map set. The fourth-stage residual attention network can be a deep neural network comprising a three-stage residual convolutional network and a gaze attention mechanism network.

[0076] Step 6: Input the fifth eye feature map set into the fully connected output network to obtain the gaze information sequence. The fully connected output network can be a deep neural network comprising a cascaded global average pooling layer and two fully connected layers. The first fully connected layer can map the 2048-dimensional eye image feature map to 512 dimensions and introduce non-linear feature information through a ReLU (Linear Rectification Function) activation function. The second fully connected layer can map the 512-dimensional eye feature map to 2 dimensions, where the 2 dimensions represent the gaze position coordinates in the horizontal and vertical directions, respectively.

[0077] In addressing the technical problems mentioned in the background section, the following technical challenge often arises: how to accurately identify and track the gaze of users wearing virtual reality headsets to determine if they are currently focused. A conventional solution to this problem is to train a conventional convolutional neural network using a large number of eye-tracking images to accurately identify the user's gaze. However, this conventional solution still suffers from the following issues: accurately identifying gaze using a conventional convolutional neural network requires a large number of convolutional layers to build a deep network, leading to the vanishing gradient problem and a large number of model parameters. Training the model consumes significant hardware resources, while virtual reality headsets require portability, limiting hardware resources and resulting in higher wear and tear on the display device, shortening its battery life. Therefore, we have decided to adopt the following solution:

[0078] Optionally, the above attention-focusing network model is trained through the following steps:

[0079] The first step is to acquire an eye-tracking image set. Each eye-tracking image in the set includes an eye-tracking sample image and gaze direction sample information. The eye-tracking images in the set can be training data for training the attention-gaze network model. For example, the eye-tracking image set can be an image set from the UnityEyes dataset. The gaze direction sample information can be information recording the gaze direction of the eyes in the eye-tracking image.

[0080] The second step involves batch processing the aforementioned eye-tracking image set to obtain a batch set of eye-tracking images. The batch sets of eye-tracking images in this set can be training data from each training session of the attention network model. The number of batch sets of eye-tracking images included in the set can be the same as the maximum number of executions of the determination step.

[0081] The third step is to randomly select an eye-tracking batch image group from the above-mentioned eye-tracking batch image group set as the target eye-tracking batch image group.

[0082] Fourth, based on the target eye-tracking batch image group, perform the following determination steps:

[0083] Sub-step 1 involves inputting the batch of eye-tracked images into the attention-focusing network model to obtain a gaze prediction information set. The gaze prediction information in this set can be the gaze vector predicted and output by the attention-focusing network model.

[0084] Sub-step 2 involves inputting the gaze prediction information group and the gaze sample information group corresponding to the eye-tracking batch image group into the gaze loss function to obtain the gaze loss function value, wherein the gaze loss function is the loss function corresponding to the attention gaze network model.

[0085] The above-mentioned gaze loss function can be:

[0086]

[0087] Where Loss represents the gaze loss function value. N1 represents the number of target eye-tracking batch images included in the above target eye-tracking batch image group. This represents the i-th gaze prediction information in the aforementioned gaze prediction information group. 1i Let represent the i-th gaze sample in the above gaze sample information group. α represents the hyperparameter of the attention gaze network model, used to control the degree of influence of the attention loss corresponding to the gaze gaze attention mechanism network on the loss. i λ represents the attention weight corresponding to the gaze attention mechanism network for the i-th eye-tracking sample image. A higher value indicates a more important region in the target eye-tracking batch of images. L2 This represents the regularization coefficient. w i This represents the weight of the i-th parameter. It represents the square of the L2 norm of the weight of the i-th parameter.

[0088] It should be noted that the above gaze loss function comprises three parts. The first part is the mean squared error loss, used to measure the difference between the gaze prediction information group and the gaze sample information. The second part is the weighted attention loss of the gaze attention mechanism network. By using the self-attention weights of the gaze attention mechanism network, attention to important eye areas can be ensured. The third part is L2 regularization, which can prevent overfitting and improve the generalization ability of the gaze attention network model.

[0089] Sub-step 3 involves determining the gaze precision and gaze recall of the gaze prediction information set and the gaze sample information set corresponding to the eye-tracking batch image group. The gaze precision represents the proportion of gaze samples predicted as positive within the gaze prediction information set. The gaze recall represents the proportion of gaze samples that are positive and their corresponding gaze predictions that are also positive. The gaze precision helps prevent over-reporting by the attention-based gaze network model, while the gaze recall helps prevent under-reporting.

[0090] Sub-step 4, in response to determining that the fixation precision, fixation recall, and fixation loss function values ​​meet the model training conditions, identifies the attention fixation network model as a trained attention fixation network model, and embeds the trained attention fixation network model into the virtual reality head-mounted display device. The aforementioned model training conditions can be that the fixation precision is greater than or equal to a preset precision threshold, the fixation recall is greater than or equal to a preset recall threshold, and the fixation loss function value is less than or equal to a preset loss threshold. The preset precision threshold, preset recall threshold, and preset loss threshold can all be pre-set values ​​and can be set according to specific scenarios; no limitation is made here.

[0091] In the fifth step, in response to the determination that the fixation precision, fixation recall and fixation loss function values ​​do not meet the model training conditions, a new set of eye-tracking batch images is randomly selected from the set of eye-tracking batch images after removing the target eye-tracking batch image set, and used as the target eye-tracking batch image set.

[0092] Step 6 involves performing hybrid parameter tuning on the attention-focusing network model to obtain the tuned model, which is then used as the attention-focusing network model for re-execution of the aforementioned determination steps. This hybrid parameter tuning can be achieved by combining adaptive momentum optimization algorithms (e.g., the Adam optimizer), mixed precision training, early stopping, and learning rate scheduling. For example, this hybrid parameter tuning could involve using learning rate scheduling, adaptive momentum optimization, and mixed precision training to accelerate training during initialization; if the difference between adjacent gaze loss function values ​​is less than a preset loss difference during iterative training, an early stopping strategy is executed to terminate training and roll back to the optimal weights.

[0093] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem mentioned in the background: "Due to the need for a large number of convolutional layers to construct a deep convolutional network for accurate gaze recognition in conventional convolutional neural networks, there is a gradient vanishing problem and a large number of model parameters. Training the model requires a large amount of hardware resources, while virtual reality head-mounted display devices need to be portable, and hardware resources are limited, resulting in high wear and tear on the display device and shortening its battery life." If these factors are resolved, the high wear and tear on the display device can be reduced, and its battery life extended. To achieve this effect, this disclosure first trains the eye-tracking image set in batches, obtaining each batch of eye-tracking images used to train the attention network, for subsequent model training. Secondly, batches of eye-tracking images are sequentially input into the attention-focusing network model. The gaze prediction information and the corresponding gaze sample information from the batches of eye-tracking images are input into the loss function for model training. The attention-focusing network model includes a multi-stage residual attention network, and the loss function includes mean squared error loss, weighted attention loss from the gaze-focusing attention mechanism network, and L2 regularization. This can, to some extent, avoid the gradient vanishing problem during deep network training, improving training efficiency and stability, and constructing a lightweight model. The gaze-focusing attention mechanism network can automatically focus on detailed eye information, improving the feature extraction capability of the eye region. Furthermore, training the model using a loss function composed of three loss components improves the accuracy and efficiency of model training. Finally, when the loss function does not meet the model training conditions, multiple parameter adjustment algorithms are used to perform hybrid model parameter adjustments, further improving the accuracy and adaptability of model parameter adjustments, reducing the number of model training iterations, reducing the hardware resources required for model training, adapting to the portability of display devices, reducing the high wear and tear on display devices, and extending the battery life of display devices.

[0094] Step 104: Generate a gaze point information sequence based on the pupil information sequence and the gaze line information sequence.

[0095] In some embodiments, the execution entity can generate a gaze point information sequence based on the pupil information sequence and the gaze line information sequence. The gaze point information in the gaze point information sequence can be the two-dimensional coordinates of the point on the terminal page where the target user's gaze falls. As an example, the execution entity can use a three-dimensional gaze point prediction algorithm to generate the gaze point information sequence based on the pupil information sequence and the gaze line information sequence. The three-dimensional gaze point prediction algorithm can be a Gaussian regression algorithm.

[0096] In some optional implementations of certain embodiments, generating a gaze point information sequence based on the pupil information sequence and the gaze line information sequence may include the following steps:

[0097] The first step is to control the eye-tracking device to acquire a set of eye calibration images for a preset calibration point set. This preset calibration point set can be nine pre-defined calibration points located on the terminal page.

[0098] The second step is to generate a calibration pupil information sequence and a calibration gaze information sequence for the aforementioned eye calibration image set. The implementation method for generating these sequences can refer to the implementation methods corresponding to steps 102 and 103 above.

[0099] The third step involves inputting the aforementioned calibration pupil information sequence, calibration gaze information sequence, and preset calibration point set into the gaze point mapping model to obtain a gaze point mapping model with coefficient parameters, which serves as the coefficient parameter gaze point mapping model. This gaze point mapping model can be a regression function comprising six parameters. The gaze point mapping model can be expressed as follows:

[0100]

[0101] Among them, (x g y g (x) represents the x and y coordinates in the eye calibration image coordinate system, with the pupil center corresponding to the calibration pupil information as the starting point and the corresponding calibration gaze information as the direction vector. p y p Let A = [a0, a1, a2, a3, a4, a5] represent the x and y coordinates of any preset calibration point in the preset calibration point set. Let B = [b0, b1, b2, b3, b4, b5] represent the vector composed of the first coefficients of the above gaze point mapping model.

[0102] The fourth step is to solve for the model coefficients of the above-mentioned gaze point mapping model to obtain the solved gaze point mapping model. The model coefficients can be solved using the least squares method.

[0103] The fifth step is to perform model validation on the solved gaze point mapping model to obtain the model validation results. These results represent the average difference between the two-dimensional coordinate sequences corresponding to the calibrated pupil information sequence, the calibrated gaze line information sequence, and the two-dimensional coordinates corresponding to the preset calibration point set.

[0104] Step 6: In response to the determination that the above model verification results meet the model verification conditions, the above pupil information sequence and the above gaze line information sequence are input into the solved gaze point mapping model to obtain the gaze point information sequence. The above model verification conditions can be that the average difference corresponding to the above model verification results is less than or equal to a preset coordinate verification threshold. The above preset coordinate verification threshold can be a pre-set value, which can be determined according to specific circumstances and is not limited here.

[0105] Step 105: Perform blink detection processing on the eye image sequence to obtain a blink frequency information sequence.

[0106] In some embodiments, the execution entity may perform blink detection processing on the aforementioned eye image sequence to obtain a blink frequency information sequence. The blink frequency information in the blink frequency information sequence may be temporal information of the blink frequencies corresponding to the opening and closing of the eyes of the target user within the acquisition time corresponding to the aforementioned eye video sequence.

[0107] In some optional implementations of certain embodiments, the above-described blink detection processing of the eye image sequence to obtain a blink frequency information sequence may include the following steps:

[0108] The first step is to perform spatial normalization on the above eye image sequence to obtain a spatially normalized eye image sequence. This spatial normalization can be achieved by scaling the eye image sequence to 64*64 pixels.

[0109] The second step involves inputting the spatially normalized eye image sequence into the first convolutional pooling feature extraction network of the blink detection model to obtain the first blink feature map. The blink detection model further includes a second convolutional pooling feature extraction network, a third convolutional pooling feature extraction network, and a blink classification output network. The first convolutional pooling feature extraction network can be a deep neural network that first expands the number of channels to 32 using a 3x3 convolutional layer, then uses the ReLU activation function to enhance non-linear expression, and finally compresses the blink feature map to a 31x31 deep neural network using a 2x2 max-pooling layer with a stride of 2. The loss function of the blink detection model can be expressed as:

[0110]

[0111] Where L represents the loss function of the blink detection model. N² represents the number of blink detection training data points used to train the blink detection model. 2i p represents the true label of the i-th blink detection training data, i.e., the classification label of either "eye open" or "blink". i This represents the predicted probability value for the i-th blink detection training data.

[0112] The third step involves inputting the first blink feature map set into the second convolutional pooling feature extraction network to obtain the second blink feature map set. This second convolutional pooling feature extraction network can be a deep neural network that first expands the number of channels to 64 using a 3x3 convolutional layer, then uses the ReLU activation function to enhance non-linear expressive power, and finally compresses the blink feature map to a 14x14 deep neural network using a 2x2 max pooling layer with a stride of 2.

[0113] The fourth step involves inputting the second blink feature map set into the third convolutional pooling feature extraction network to obtain the third blink feature map set. This third convolutional pooling feature extraction network can be a deep neural network with a 3x3 convolutional kernel to expand the number of channels to 128, followed by a ReLU activation function to enhance non-linear expression, and finally a 2x2 max-pooling layer with a stride of 2 to compress the blink feature map to 6x6, forming a 6x6x128 third blink feature map.

[0114] The fifth step is to perform feature flattening on the aforementioned third blink feature map to obtain a blink feature vector set. This feature flattening can be achieved by performing a Flatten operation on the aforementioned 6*6*128 dimensional third blink feature map, flattening it into a 1*1*4068 dimensional vector.

[0115] Step 6: Input the aforementioned blinking feature vector set into the aforementioned blinking classification output network to obtain the open / closed eye classification information sequence. The open / closed eye classification information in this sequence ensures the target user's eye-opening / closing state in the current eye image. The aforementioned blinking classification output network can be a deep neural network that includes a fully connected layer to reduce the dimensionality from 4068 to 128, a ReLU activation function to enhance the non-linear feature information of the feature map, a Dropout layer to alleviate overfitting, a fully connected layer to reduce the dimensionality to 1, and a Sigmoid function to restrict the output value to [0, 1], representing the classification confidence of whether the target user's eyes are open or closed.

[0116] Step 7: Detect the above-mentioned open / closed eye classification information sequence to obtain the blink frequency information sequence. This blink frequency information sequence can be a sequence that determines whether a blink has occurred, using a state queue of 3 frames in the form of "1-0-1". "1" indicates open eyes, and "0" indicates closed eyes.

[0117] It should be noted that the aforementioned blink detection model features a lightweight structure with fewer parameters, making it suitable for real-time inference on low-power devices with limited computing resources, such as virtual reality headsets. Furthermore, by employing three convolutional pooling operations, the size of the blink feature map is compressed to 6*6, focusing on key regions and reducing redundant computation. Simultaneously, Dropout regularization is used to enhance the robustness of the blink detection model, ensuring its generalization ability.

[0118] Step 106: Generate attention state information for the target user based on the pupil information sequence, fixation point information sequence, and blink frequency information sequence.

[0119] In some embodiments, the executing entity can generate attentional state information for a target user based on the pupil information sequence, the fixation point information sequence, and the blink frequency information sequence. The attentional state information in the aforementioned attentional state information set can characterize the temporal information of the target user's concentration and inattention within the time corresponding to the eye image sequence. This attentional state information can be information closely related to the eye movement feature information set (e.g., pupil diameter, fixation point information, and blink frequency). For example, if the target user's pupil diameter decreases, the fixation point sequence becomes irregular, or frequent blinking occurs over a period of time, it indicates that the target user is inattentive during that period.

[0120] As an example, the aforementioned execution entity can input the aforementioned pupil information sequence, fixation point information sequence, and blink frequency information sequence into the attention state detection model to obtain attention state information. The attention state detection model can be a model fusing reinforcement learning and a bidirectional long short-term memory neural network. The reinforcement learning can be a TD3 (Twin Delayed Deep Deterministic Policy Gradient) model used to determine the length of the pupil information sequence, fixation point information sequence, and blink frequency information sequence input to the bidirectional long short-term memory neural network, i.e., the window size. The reinforcement learning environment can be the complete pupil information sequence, fixation point information sequence, and blink frequency information sequence. The reinforcement learning state can be the time window size. The reinforcement learning action can be the adjustment range of the time window size. The reinforcement learning reward can be the prediction performance (e.g., mean squared error) of the bidirectional long short-term memory neural network. The loss function of the attention state detection model can be expressed as:

[0121]

[0122] Among them, L totalThis represents the loss function of the attention state detection model. N3 represents the amount of attention state training data used to train the attention state detection model. 3i This represents the true label of the i-th attention state training data, i.e. whether it is a state of focused attention. The label for a focused state is 1, and the label for a state of unfocused attention is 0. This represents the predicted probability value of the attention state in the attention training data. T represents the total number of steps within the time window learned by reinforcement learning. γ t R represents the reward obtained at time step t based on the time window selected according to reinforcement learning. t This represents the discount factor, used to adjust the importance of rewards at different time steps.

[0123] It should be noted that the above attention state detection model, through a composite loss function, can simultaneously improve the classification effect and the adaptability of the time window selection, thereby achieving the accuracy of attention state detection.

[0124] Step 107: Visualize the attention state information to obtain the user attention image, and send the eye image sequence, fixation point information sequence and user attention image to the terminal page for visualization.

[0125] In some embodiments, the execution entity may visualize the attention state information set to obtain a user attention image, and send the eye image sequence, the fixation point information sequence, and the user attention image to a terminal page for visualization. The user attention image may be an image representing user attention information in the form of an image.

[0126] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an attention state display device based on an eye-tracking device. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, this attention state display device based on eye-tracking equipment can be specifically applied to various electronic devices.

[0127] like Figure 7As shown, an attention state display device 700 based on an eye-tracking device includes: a control unit 701, a pupil detection unit 702, an input unit 703, a first generation unit 704, a blink detection unit 705, a second generation unit 706, and a visualization display unit 707. The control unit 701 is configured to control a virtual reality head-mounted display device embedded with an eye-tracking device to acquire a sequence of eye images for a target user. The pupil detection unit 702 is configured to perform pupil detection on the aforementioned eye image sequence to obtain a pupil information sequence. The input unit 703 is configured to input the aforementioned eye image sequence into a trained attention gaze network model to obtain a gaze information sequence. The first generation unit 704 is configured to generate a gaze point information sequence based on the aforementioned pupil information sequence and the aforementioned gaze information sequence. The blink detection unit 705 is configured to perform blink detection processing on the aforementioned eye image sequence to obtain a blink frequency information sequence. The second generation unit 706 is configured to generate attention state information for the target user based on the pupil information sequence, the fixation point information sequence, and the blink frequency information sequence. The visualization display unit 707 is configured to visualize the attention state information to obtain a user attention image, and to send the eye image sequence, the fixation point information sequence, and the user attention image to a terminal page for visualization display.

[0128] It is understandable that the units described in the attention state display device 700 based on the eye-tracking device are similar to the reference units. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the attention state display device 700 based on the eye-tracking device and the units contained therein, and will not be repeated here.

[0129] The following is for reference. Figure 8 It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 800 suitable for implementing some embodiments of the present disclosure. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0130] like Figure 8As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0131] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 8 Each box shown can represent a device or multiple devices as needed.

[0132] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, it performs the functions defined above in the methods of some embodiments of this disclosure.

[0133] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0134] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0135] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the contents included in steps 101-107.

[0136] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0138] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a control unit, an eye image cropping unit, a pupil detection unit, an input unit, a first generation unit, a blink detection unit, a second generation unit, and a visualization display unit. The names of these units do not necessarily limit the specific unit; for example, the control unit may also be described as "a unit that controls a virtual reality head-mounted display device with an embedded eye-tracking device to acquire eye image sequences for a target user."

[0139] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0140] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for displaying attention state based on an eye-tracking device, comprising: Control a virtual reality head-mounted display device with embedded eye-tracking equipment to acquire eye image sequences for the target user; Pupil detection is performed on the eye image sequence to obtain a pupil information sequence; The eye image sequence is input into the trained attention gaze network model to obtain a gaze information sequence; Generate a gaze point information sequence based on the pupil information sequence and the gaze line information sequence; The eye image sequence is processed by blink detection to obtain a blink frequency information sequence; Based on the pupil information sequence, the fixation point information sequence, and the blink frequency information sequence, attention state information for the target user is generated. The attention state information is visualized to obtain a user attention image, and the eye image sequence, the fixation point information sequence, and the user attention image are sent to the terminal page for visualization.

2. The method according to claim 1, wherein, The eye-tracking device includes: an image acquisition component, multiple light source components, and a housing structure component; and The virtual reality head-mounted display device, which controls the embedded eye-tracking device, acquires eye image sequences for the target user, including: Determine the set of mounting area location information of the image acquisition component on the housing structure component; At least one installation area location information that meets the installation structure conditions is selected from the set of installation area location information; Based on the field of view information of the virtual reality head-mounted display device and the offset constraint information corresponding to the image acquisition component, the location information of the at least one installation area is further filtered to obtain the target installation location information; The multiple light source components are evenly installed onto the housing structure component, and the image acquisition component is installed onto the housing structure component according to the target installation position information to obtain an eye-tracking device; Controlling the eye-tracking device to be embedded in the virtual reality head-mounted display device, and controlling the virtual reality head-mounted display device to collect eye-tracking videos for the target user; The eye image sequence is obtained by cropping the eye image from the eye-tracking video frame sequence corresponding to the eye-tracking video.

3. The method according to claim 1, wherein, The step of performing pupil detection on the eye image sequence to obtain a pupil information sequence includes: For each eye image in the eye image sequence, the following filtering steps are performed: Edge detection processing is performed on the eye image to obtain an eye edge contour set; Elliptical contour fitting is performed on the eye edge contour set to obtain an eye ellipse information set, wherein the eye ellipse information set includes: an eye ellipse image set, an eye ellipse center coordinate information set, an eye ellipse major axis information set, and an eye ellipse minor axis information set; Based on the major axis information set and the minor axis information set of the eye ellipse, determine the eye ellipse area set and the eye ellipse aspect ratio set of the eye ellipse information set; Based on the set of eye ellipse areas and the set of eye ellipse aspect ratios, the set of eye ellipse images is filtered to obtain a filtered set of eye ellipse images. In response to determining that the filtered set of elliptical eye images includes at least two filtered elliptical eye images, and that the eye images do not meet the initial eye image conditions, the set of eye ellipse center coordinate information corresponding to the filtered set of elliptical eye images is filtered according to the previous pupil center information corresponding to the previous eye image to obtain pupil center information. In response to determining that the filtered set of elliptical eye images is empty and that the eye image does not meet the initial eye image conditions, the previous pupil center information is determined as the pupil center information of the eye image. Pupil information is generated based on the pupil center information, the major axis information set of the eye ellipse, and the minor axis information set of the eye ellipse.

4. The method according to claim 1, wherein, The process of inputting the eye image sequence into the trained attention-gaze network model to obtain a gaze information sequence includes: The eye image sequence is input into the convolutional pooling network included in the trained attention-focusing network model to obtain a first eye feature map set. The trained attention-focusing network model further includes: a first-stage residual attention network, a second-stage residual attention network, a third-stage residual attention network, a fourth-stage residual attention network, and a fully connected output network. The first eye feature map set is input into the first stage residual attention network to obtain the second eye image feature map set; The second eye image feature set is input into the second-stage residual attention network to obtain the third eye image feature set; The third eye image feature set is input into the third-stage residual attention network to obtain the fourth eye image feature set; The fourth eye image feature set is input into the fourth-stage residual attention network to obtain the fifth eye image feature set; The fifth eye image feature set is input into the fully connected output network to obtain a gaze information sequence.

5. The method according to claim 1, wherein, The step of generating a gaze point information sequence based on the pupil information sequence and the gaze line information sequence includes: Control the eye-tracking device to acquire a set of eye calibration images for a preset calibration point set; Generate a calibration pupil information sequence and a calibration gaze information sequence for the eye calibration image set; The calibration pupil information sequence, the calibration gaze line information sequence, and the preset calibration point set are input into the gaze point mapping model to obtain a gaze point mapping model with coefficient parameters, which is used as the coefficient parameter gaze point mapping model. The model coefficients of the gaze point mapping model with the coefficient parameters are solved to obtain the solved gaze point mapping model; The solved gaze point mapping model is validated to obtain the model validation results; In response to determining that the model validation result meets the model validation conditions, the pupil information sequence and the gaze line information sequence are input into the solved gaze point mapping model to obtain the gaze point information sequence.

6. The method according to claim 1, wherein, The blink detection processing of the eye image sequence to obtain a blink frequency information sequence includes: The eye image sequence is spatially normalized to obtain a spatially normalized eye image sequence. The spatially normalized eye image sequence is input into the first convolutional pooling feature extraction network of the blink detection model to obtain the first blink feature map set. The blink detection model further includes: a second convolutional pooling feature extraction network, a third convolutional pooling feature extraction network, and a blink classification output network. The first blink feature map is input into the second convolutional pooling feature extraction network to obtain the second blink feature map; The second blink feature map is input into the third convolutional pooling feature extraction network to obtain the third blink feature map; The third blink feature set is input into the blink classification output network to obtain the open and closed eye classification information sequence. The blink frequency information sequence is obtained by detecting the open and closed eye classification information sequence.

7. An attention state display device based on an eye-tracking device, comprising: The control unit is configured to control a virtual reality head-mounted display device with an embedded eye-tracking device to acquire eye image sequences for a target user; A pupil detection unit is configured to perform pupil detection on the eye image sequence to obtain a pupil information sequence; The input unit is configured to input the eye image sequence into the trained attention gaze network model to obtain a gaze information sequence; The first generation unit is configured to generate a gaze point information sequence based on the pupil information sequence and the gaze line information sequence; A blink detection unit is configured to perform blink detection processing on the eye image sequence to obtain a blink frequency information sequence. The second generation unit is configured to generate attention state information for the target user based on the pupil information sequence, the fixation point information sequence, and the blink frequency information sequence. The visualization display unit is configured to visualize the attention state information, obtain a user attention image, and send the eye image sequence, the fixation point information sequence, and the user attention image to a terminal page for visualization display.

8. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Virtual reality interaction method and device based on eye movement tracking

    CN110502100A

  • Personalized sight tracking method based on space-time attention mechanism and Gaussian process

    CN118692133A

  • Human eye attention positioning method and device based on pre-trained neural network

    CN119323824A

  • Apparatus control method, model training method and electronic apparatus

    EP4488803A1