An eye movement monitoring method based on near-eye infrared imaging of a head-mounted device
By constructing a lightweight eye feature detection network model, using lightweight convolution strategy and inverse residual structure, the requirements of high real-time and high precision in wearable devices are solved, and the effect of efficient and real-time monitoring of eye movement status parameters is achieved.
Patent Information
- Application Number
- CN202410827292.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-06-25
AI Technical Summary
Existing deep learning eye feature detection methods are difficult to meet the needs of high real-time and high precision in wearable devices with limited computing resources.
A lightweight eye feature detection network model is built, and a lightweight convolution strategy and inverse residual structure are adopted to reduce the amount of calculation and parameter, and to improve the nonlinear expression ability and robustness of the network.
It realizes efficient real-time monitoring of eye movement status parameters in wearable devices, balances the requirements of accuracy and real-time, and improves the robustness and generalization of detection.
Smart Images

Figure CN118918630B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of eye detection, and particularly to an eye movement monitoring method based on near-eye infrared imaging of a head-mounted device. Background Art
[0002] Eye movement monitoring technology can detect and identify important eye features (such as eyelids, pupils, irises, etc.) from eye images through advanced image processing and analysis methods. Combining temporal information, it can monitor the eye movement state parameters of users (such as fatigue state, gaze direction, etc.), and is widely used in fields such as medical diagnosis, assisted driving, and virtual reality.
[0003] Eye feature detection is the core content of eye movement monitoring, which directly determines the accuracy and robustness of advanced functions such as fatigue detection and gaze estimation. Traditional eye feature detection methods usually adopt strategies such as thresholding to extract pupil pixels or ellipse fitting to extract the pupil contour. The principle is simple, but it can only detect pupil features and has poor robustness, with low detection accuracy in environments such as light reflection, illumination, and occlusion. Compared with traditional methods, the advantage of deep learning lies in its powerful feature learning ability, which can extract higher-level and more abstract features from input images, capture important information such as eye shape and texture, thereby improving detection accuracy and robustness, and more accurately locating eye features. However, existing deep learning eye feature detection methods usually focus on improving accuracy, resulting in an increasing scale of the network model and a sharp increase in computational complexity. In practical engineering applications, such as wearable devices, the trade-off between accuracy and real-time performance is a factor that needs to be considered. Existing methods are difficult to meet the task requirements of high real-time performance. Therefore, it is still challenging and difficult to ensure high real-time performance and high accuracy under limited computing resources of wearable devices.
[0004] Compared with large model deep learning algorithms, lightweight deep learning algorithms aim to further reduce the number of model parameters and complexity while maintaining model accuracy, and have gradually become a research hotspot in computer vision. The patent with application number 202311547575.0 proposes a real-time binocular pupil inspection system. This method uses two groups of binocular infrared cameras to capture the left and right eyes of the tester in real time, uses a lightweight object detection model Yolo to detect the pupil positions of the left and right eyes of the tester in real time, and uses a semantic segmentation model FastSCNN to extract the pupil area near the detected pupil positions. Finally, the pupil diameter, pupil shape are measured, and pupil gaze tracking and blink detection are performed to achieve reliable binocular detection. However, the two-stage strategy of this method has a large computational amount, especially the semantic segmentation model takes a long time, and it is difficult to meet the real-time performance requirements under limited computing resources such as wearable devices. Summary of the Invention
[0005] The object of the present invention is to provide an eye movement monitoring method based on near-eye infrared imaging of a head-mounted device. By constructing a lightweight eye feature detection network model and using a lightweight convolution strategy to efficiently extract important information from images, the amount of calculation and the number of parameters are reduced, meeting the real-time application requirements in scenarios where wearable devices have limited computing resources.
[0006] To achieve the above object, the present invention provides an eye movement monitoring method based on near-eye infrared imaging of a head-mounted device. The steps include:
[0007] S1. Collect an existing eye movement infrared image dataset and perform label normalization processing;
[0008] S2. Construct a lightweight eye feature detection network model for the head-mounted device;
[0009] S3. Train and test the lightweight eye feature detection network model based on the eye movement infrared image dataset processed in step S1, and extract eye features;
[0010] S4. Based on the eye feature results, combine the timing information to monitor the eye movement state parameters.
[0011] Preferably, the lightweight eye feature detection network model includes a Conv module, a DepthwiseConv module, a Block module, and a fully connected layer. The Conv module includes a convolutional layer and a batch normalization layer. The DepthwiseConv module is a depth convolutional layer. The Block module is an inverted residual structure. The fully connected layer is arranged at the end of the lightweight eye feature detection network.
[0012] Preferably, in step S3, a loss function is used to train the lightweight eye feature detection network model. The training process includes:
[0013] S31. Set the network parameters of the lightweight eye feature detection network model;
[0014] S32. Input a near-infrared eye image, and gradually extract the near-infrared eye image features using convolution, depth convolution, and inverted residuals, and obtain feature maps of different scales through a branch strategy;
[0015] S33. Aggregate the feature maps of different scales, and use the fully connected layer and the inverted residual structure to map the feature maps to the x and y coordinates of the eye feature points.
[0016] Preferably, in the training process, a data augmentation strategy is adopted to increase the training data of near-infrared eye images. The data augmentation strategy includes, but is not limited to, cropping, scaling, Gaussian filtering, and median filtering methods.
[0017] Preferably, during the training process, the network parameters are optimized by adjusting parameters and loading pre-trained weights.
[0018] Preferably, the loss function is the mean squared error (MSE) of the sum of squared errors, which calculates the mean of the sum of squared differences between the predicted data and the corresponding points of the original data. The formula is:
[0019]
[0020] In the formula, x i represents the true value of the x-coordinate of the eye feature point, and y i represents the true value of the y-coordinate of the eye feature point. represents the predicted value of the x-coordinate of the eye feature point; represents the predicted value of the y-coordinate of the eye feature point, and m represents the number of eye feature detection points.
[0021] Preferably, in step S4, the eye movement state parameters include but are not limited to eyelid opening degree, PERCLOS, blink frequency, average closing speed, average opening degree, and maximum continuous closing time.
[0022] Therefore, the present invention adopts the above-mentioned eye movement monitoring method based on near-eye infrared imaging of a head-mounted device, and has the following beneficial effects:
[0023] (1) Construct a lightweight eye feature detection network, use an inverted residual structure to improve the non-linear expression ability of the network, reduce the number of channels through dilation layers and projection layers to reduce the computational amount and the number of parameters, and perform convolution operations on each input channel independently with corresponding convolution kernels through depthwise separable convolution, greatly reducing the computational amount and the number of parameters;
[0024] (2) Aggregate multi-scale feature modules, which can effectively integrate feature information and context features at different scales, improve the positioning accuracy of eye feature points, and enhance the robustness and generalization of the lightweight eye feature detection network model;
[0025] (3) In the actual engineering application scenario where the computing resources of wearable devices are limited, it can balance the requirements of accuracy and real-time performance and monitor eye movement state parameters in real time.
[0026] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flowchart of the method according to an embodiment of the present invention;
[0028] Figure 2 is a diagram of an inverted residual structure according to an embodiment of the present invention;
[0029] Figure 3The shortcut connection diagram of the inverted residual structure according to the embodiment of the present invention;
[0030] Figure 4 The structural diagram of the lightweight eye feature detection network according to the embodiment of the present invention;
[0031] Figure 5 The flowchart of blink detection according to the embodiment of the present invention;
[0032] Figure 6 The eyelid feature points according to the embodiment of the present invention;
[0033] Figure 7 The test result diagram of the lightweight eye feature detection network according to the embodiment of the present invention. Detailed implementation manners
[0034] Embodiment
[0035] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0036] Refer to Figure 1 , an eye movement monitoring method based on near-eye infrared imaging of a head-mounted device, the steps include:
[0037] S1. Collect the existing eye movement infrared image dataset and perform label normalization processing.
[0038] S2. Construct a lightweight eye feature detection network model for the head-mounted device.
[0039] Specifically, the lightweight eye feature detection network model includes multiple Conv modules, one DepthwiseConv module, multiple Block modules and three fully connected layers.
[0040] The Conv module includes a convolutional layer and a batch normalization layer. The batch normalization normalizes the input of each batch to keep the mean and variance of the input stable, accelerates the convergence of the network, improves the generalization ability of the model, and has a certain regularization effect.
[0041] The DepthwiseConv module is a depth convolutional layer. The depth convolution independently convolves each channel of the input with the corresponding convolution kernel to generate the corresponding output channels, thereby reducing the amount of calculation and the number of parameters.
[0042] The Block module has an inverted residual structure and uses the ReLU6 activation function to limit the output value between 0 and 6, improving the numerical stability of the model, enhancing the anti-noise performance, and helping to reduce the computational and storage overhead. The inverted residual structure performs dimensionality increase through 1×1 convolution and then dimensionality reduction through 1×1 convolution. As Figure 2 shown, the expansion layer and the projection layer correspond to the dimensionality increase and dimensionality reduction operations of 1×1 convolution respectively. The former enhances the network's expressive ability by increasing the number of channels, thereby providing more feature information. The latter reduces the computational amount and the number of parameters by reducing the number of channels, while helping the network to concentrate on learning and expressing the main features. The inverted residual structure performs a shortcut connection only when stride = 1 and the shapes of the input feature matrix and the output matrix are the same. As Figure 3 shown.
[0043] S3. Train and test the lightweight eye feature detection network model based on the eye movement infrared image dataset processed in step S1, and extract eye features.
[0044] The training process includes:
[0045] S31. Set the network parameters of the lightweight eye feature detection network model.
[0046] S32. Input the near-infrared eye image, and gradually extract the features of the near-infrared eye image using convolution, depth convolution, and inverted residual, and obtain four feature maps of different scales through the branch strategy.
[0047] S33. Aggregate the four feature maps of different scales, and use the fully connected layer and the inverted residual structure to map the feature maps to the x and y coordinates of 50 eye feature points, representing the positions of the eyelids, pupils, and irises. Among them, the number of feature points of the eyelids, pupils, and irises are 34, 8, and 8 respectively.
[0048] As Figure 4 shown, it is the network structure of the lightweight eye feature detection network model. In the figure, the form "3×3Conv 64,2" of the Conv module indicates that the size of the convolution kernel is 3×3, the number of convolution kernels is 64, and the stride is 2.
[0049] During training, X is the input image, which first passes through a 3×3 convolutional layer, a depth convolutional layer, and three block inverted residual convolutional modules. Then it is divided into two branches. One branch passes through a 3×3 convolutional layer and a 1×1 convolutional layer to form the first feature block; the other branch passes through four block inverted residual convolutional modules. It is divided into two branches again. One branch passes through a 3×3 convolutional layer and two 1×1 convolutional layers to form the second feature block; the other branch passes through five block inverted residual convolutional modules. Finally, one branch passes through a 3×3 convolutional layer and three 1×1 convolutional layers to form the third feature block; the other branch passes through a 3×3 convolutional layer and four 1×1 convolutional layers to form the fourth feature block. Before the fully connected layer, all four feature blocks reduce the number of channels to 1 through convolutional layers and are aggregated to combine cross-channel and cross-scale context features. Three fully connected layers are set at the end. The first fully connected layer integrates the feature information within different receptive fields extracted by the previous convolutional layer and maps this information into the sample label space so that the network can better understand and process the input data and enhance the model's representational ability. The second and third fully connected layers gradually reduce the output dimension of the network and finally map it to the x and y coordinates of 50 eye feature points. Among them, the number of feature points for the eyelid, pupil, and iris is 34, 8, and 8 respectively.
[0050] Specifically, during the training process, data augmentation strategies such as cropping, scaling, Gaussian filtering, and median filtering are adopted to increase the training data of near-infrared eye images.
[0051] Specifically, during the training process, network parameters are optimized by means of adjusting parameters and loading pre-trained weights.
[0052] Specifically, the loss function adopted is the mean squared error MSE, which calculates the mean of the sum of squares of the corresponding points between the predicted data and the original data, reflecting the distribution of the prediction error. The formula is:
[0053]
[0054] In the formula, x i represents the true value of the x coordinate of the eye feature point, y i represents the true value of the y coordinate of the eye feature point, represents the predicted value of the x coordinate of the eye feature point; represents the predicted value of the y coordinate of the eye feature point, and m represents the number of eye feature detection points.
[0055] S4. As Figure 5 shown, based on the eye feature results and combined with the temporal information, eye movement state parameters such as eyelid opening degree, PERCLOS, blink frequency, average closing speed, average opening degree, and maximum continuous closing time are monitored.
[0056] Eye opening degree: That is, the degree of eye opening. As Figure 6 shown, according to the eyelid feature points, the eye opening degree is expressed as:
[0057]
[0058] In the formula, p1 represents the upper left edge point of the eyelid, p2 represents the lower left edge point of the eyelid, p3 represents the upper right edge point of the eyelid, p4 represents the lower right edge point of the eyelid, p5 represents the inner corner of the eyelid, and p6 represents the outer corner of the eye
[0059] PERCLOS (Percentage of Eyelid Closure Over the Pupil Over Time): It refers to the percentage of time with the eyes closed over a period of time. When the eye opening degree is less than the set threshold t, it is regarded as the closed-eye state. In the fatigued state, PERCLOS will increase significantly. PERCLOS is expressed as:
[0060]
[0061] Blinking frequency: That is, the number of blinks per unit time. When the state of the eyes experiences opening - closing - opening, it is regarded as 1 blink. In the fatigued state, the blinking frequency will increase or decrease significantly. The blinking frequency is expressed as:
[0062]
[0063] Average closing speed: It is characterized by the average value of the time required for the eyes to close from the open state. The longer the time, the slower the closing speed. In the fatigued state, the closing speed will be significantly slower. The average closing speed is expressed as:
[0064] Average closing speed = mean(time experienced by the eyes from opening to closing)
[0065] Average opening degree: It refers to the percentage of time with the eyes open over a period of time. When the eye opening degree is greater than the set threshold t, it is regarded as the open-eye state. In the fatigued state, the average opening degree will be significantly reduced. The average opening degree can be expressed as:
[0066]
[0067] Maximum continuous closing time: In the fatigued state, the maximum continuous closing time will be significantly prolonged. The maximum continuous closing time is expressed as:
[0068] Maximum continuous closing time = max(continuous closing time)
[0069] To verify the effectiveness of the proposed solution of the present invention, the existing dataset of near-infrared eye images was classified at a ratio of 4:1 for network training and testing. The hardware information of the test device is as follows: GPU, model NVIDIA 2080Ti, and the memory is 12GB. The test results are as Figure 7 shown. It can be seen from the figure that the present invention can well detect the eye feature positions such as eyelids, pupils, and irises in the head-mounted near-eye scenario.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solution of the present invention, and these modifications or equivalent replacements cannot make the modified technical solution deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for monitoring eye movement based on near-eye infrared imaging of a head-mounted device, characterized in that the steps include: S1, collect existing eye movement infrared image datasets and perform label normalization processing; S2. Constructing a lightweight eye feature detection network model for a head-mounted device, wherein the lightweight eye feature detection network model includes multiple Conv modules, a DepthwiseConv module, multiple Block modules, and three fully connected layers, wherein the Conv module includes a convolution layer and a batch normalization layer, the DepthwiseConv module is a deep convolution layer, the Block module is an inverted residual structure, and the fully connected layer is arranged at the tail of the lightweight eye feature detection network; S3, training and testing a lightweight eye feature detection network model based on the eye movement infrared image dataset processed in step S1 to extract eye features, the training process includes: S31, setting network parameters of a lightweight eye feature detection network model; S32, input the near-infrared eye image, use convolution, deep convolution and inverse residual to gradually extract the features of the near-infrared eye image, and obtain feature maps of different scales through a branching strategy, specifically: The input image first passes through a 3×3 convolution layer, a deep convolution layer, and three block modules, and then is divided into two paths. One path passes through a 3×3 convolution layer and a 1×1 convolution layer to form the first feature block; the other path passes through four block modules and is divided into two paths again. One path passes through a 3×3 convolution layer and two 1×1 convolution layers to form the second feature block; the other path passes through five block modules, and finally, one path passes through a 3×3 convolution layer and three 1×1 convolution layers to form the third feature block; the other path passes through a 3×3 convolution layer and four 1×1 convolution layers to form the fourth feature block; S33, aggregate the feature maps of different scales, and use the fully connected layer and the inverted residual structure to map the feature maps to the x and y coordinates of the eye feature points, specifically: The first fully connected layer integrates the feature information of different sizes of receptive fields extracted by the previous convolutional layer and maps this information into the sample label space. The second and third fully connected layers gradually reduce the output dimension of the network and map it to the x and y coordinates of 50 eye feature points, where the number of feature points of the eyelid, pupil, and iris are 34, 8, and 8 respectively. S4. Based on the eye feature results and combined with timing information, the eye movement state parameters are monitored. The eye movement state parameters include eyelid opening and closing degree, PERCLOS, blinking frequency, average eye closing speed, average eye opening degree, and maximum continuous eye closure time.
2. The eye movement monitoring method based on near-eye infrared imaging of a head-mounted device according to claim 1, characterized in that: During the training process, a data enhancement strategy is used to increase near-infrared eye image training data, and the data enhancement strategy includes but is not limited to cropping, scaling, Gaussian filtering and median filtering.
3. The eye movement monitoring method based on near-eye infrared imaging of a head-mounted device according to claim 1, characterized in that: During the training process, network parameters are optimized by adjusting parameters and loading pre-trained weights.
4. The eye movement monitoring method based on near-eye infrared imaging of a head-mounted device according to claim 1, characterized in that: The loss function is the mean value of the sum of squared errors (MSE), which is the mean of the sum of squares of the corresponding points of the predicted data and the original data. The formula is: In the formula, x i Indicates the true value of the x coordinate of the eye feature point, y i Indicates the true value of the y coordinate of the eye feature point, Indicates the predicted value of the x coordinate of the eye feature point; represents the predicted value of the y coordinate of the eye feature point, and m represents the number of eye feature detection points.
Citation Information
Patent Citations
Real-time double-eye pupil examination system and detection method
CN117530654A
Sight line tracking model training method, and sight line tracking method and device
CN110058694A
Eye movement interaction system based on visual image information
CN114821753A