A method for establishing an eye state estimation network and a method for estimating an eye state
The Light EyeNet network addresses the issues of lighting and noise in eye state estimation by employing a multi-scale feature extraction and attention mechanism, resulting in improved accuracy and reduced computational cost.
Patent Information
- Application Number
- CN202510517556.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing methods of eye state estimation are susceptible to adverse factors such as external light, noise and face tilt, resulting in slow detection speed and insolubly robust enough.
The Light EyeNet network is adopted, which includes five sequentially connected backbone networks, attention modules, global pooling layer, fully connected layer and the first convolutional module. The number of channels is reduced through the fast feature extraction module, and the key information is extracted and ignored irrelevant information is ignored, thereby enhancing the robustness of the model.
It improves detection speed, reduces computing costs, and shows significant computing resource advantages on mobile devices with limited resources, enhancing the robustness and accuracy of the model.
Smart Images

Figure CN120047991B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of eye state detection, and particularly to a method for establishing an eye state estimation network and a method for estimating an eye state. Background Art
[0002] The eyes are regarded as an important window for human emotions and states. Eye State Estimation (ESE) has important research and application values in the fields of human-computer interaction, sleep research, and fatigue driving detection.
[0003] The current eye state estimation methods can be roughly divided into two categories: those based on handcrafted features and those based on deep features. In the methods based on handcrafted features, the Viola-Jones detector extracts Haar features around the eyes and trains a cascade classifier based on Adaboost to locate the positions of the eyes. In the methods based on deep features, the Hough transform is used to locate the iris and pupil to determine the eye region, or the Variance Projection Function (VPF) is used to locate the key points of the eyes to determine the position and shape of the eyes.
[0004] These early eye state detection methods mainly rely on handcrafted features and are easily affected by adverse factors such as external light, noise, and face tilt. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for establishing an eye state estimation network and a method for estimating an eye state that can improve the detection speed.
[0006] To achieve the above purpose, the present invention adopts the following technical solution:
[0007] A method for establishing an eye state estimation network, which establishes a Light EyeNet network. The Light EyeNet network includes five backbone networks connected in sequence, as well as an attention module, a global pooling layer, a fully connected layer, and a first convolution module;
[0008] The five backbone networks connected in sequence are used for extracting features of different scales from the input image, and the output of the previous backbone network is used as the input of the next backbone network;
[0009] The backbone network includes a second convolution module, a plurality of fast feature extraction modules, a third convolution module, and a fourth convolution module connected in sequence;
[0010] The fast feature extraction module includes two branches. One branch performs depth convolution operation and pointwise convolution operation on the input image in sequence to obtain a first feature map, and the other branch performs max pooling operation and channel adjustment operation on the input image in sequence to obtain a second feature map. The first feature map and the second feature map are concatenated and then the fused feature map is output.
[0011] The attention module is used to extract the key information of the feature map output by the fourth backbone network.
[0012] The global pooling layer is used to map the feature map output by the attention module to a low-dimensional space.
[0013] The first convolution module is used to extract features from the feature map output by the fifth backbone network.
[0014] The fully connected layer is used to perform feature connection on the feature map output by the global pooling layer and the feature map output by the first convolution module and output the eye state feature map.
[0015] Preferably, the attention module performs convolution operation, normalization processing and activation function processing on the feature map output by the fourth backbone network in sequence.
[0016] Preferably, the backbone network includes 11 fast feature extraction modules.
[0017] A method for estimating eye state includes the following steps executed in sequence:
[0018] S1: Obtain an experimental data set, perform face key point localization on each picture in the experimental data set, and obtain an eye data set.
[0019] S2: Preprocess the eye data set to obtain a data set of 128×128 pixels, and this data set includes a training image set and a test image set.
[0020] S3: Input the training image set into the Light EyeNet network established by the method for establishing an eye state estimation network as described in any one of the above to perform training, and output an eye state feature map.
[0021] S4: Calculate the loss function of the Light EyeNet network, perform backpropagation on the Light EyeNet network using this loss function, and obtain an adjusted Light EyeNet network.
[0022] S5: Input the test image set into the adjusted Light EyeNet network for testing, and output the test result.
[0023] Preferably, the steps of face key point localization in step S1 are as follows:
[0024] S1-1: Input each picture of the experimental data set into the Facemesh module of Mediapipe, and the Facemesh module outputs the predicted key points.
[0025] S1-2: Select the large eye corner point ( , ) and the small eye corner point ( , ) corresponding to one of the eyes from the key points, and calculate the width of the eye area using the following formula :
[0026] ;
[0027] S1-3: Expand the large eye corner point and the small eye corner point according to the width of the eye area to obtain the expanded coordinates of the eye area ( , ), ( , ), and the specific expansion process is shown in the following formula:
[0028] ;
[0029] ;
[0030] ;
[0031] ;
[0032] Among them, represents the expansion ratio value;
[0033] S1-4: Use the method in S1-2 to S1-3 above to obtain the expanded coordinates of the other eye area;
[0034] S1-5: Crop the eye area pictures with the expanded coordinates corresponding to the two eyes as the eye data set.
[0035] Preferably, in step S2, the preprocessing of the eye data set includes random cropping, border padding, random color jittering, and pixel normalization.
[0036] Preferably, in step S4, the loss function is calculated based on the combination of binary cross-entropy loss and Focal loss, and the specific calculation formula is as follows:
[0037] ;
[0038] ;
[0039] ;
[0040] Among them, represents binary cross-entropy loss, represents Focal loss, and is the true label, is the predicted probability, and are the adjustment parameters in
[0041] An eye state estimation system includes a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, the eye state estimation method described in any one of the above is implemented.
[0042] By adopting the foregoing design scheme, the beneficial effects of the present invention are as follows: The present application uses a fast feature extraction module for feature extraction, reduces the number of channels, lowers the computational cost, and improves the extraction speed;
[0043] The Light EyeNet network increases the breadth of the model through feature fusion between different levels, applies the attention mechanism to extract key information and ignores irrelevant information, and improves the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a schematic structural diagram of the Light EyeNet network of the present invention;
[0045] Figure 2 is a schematic structural diagram of the backbone network of the present invention;
[0046] Figure 3 is a schematic structural diagram of the fast feature extraction module of the present invention;
[0047] Figure 4 is a schematic structural diagram of the attention module of the present invention;
[0048] Figure 5 is a schematic diagram of face key point positioning of the present invention;
[0049] Figure 6 is a schematic diagram of eye state estimation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0051] The terms "first", "second", "third", etc. in the specification, claims, and the above-mentioned drawings of the present invention are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0052] A method for establishing an eye state estimation network, establishing a Light EyeNet network as shown in Figure 1 The LightEyeNet network includes five backbone networks connected in sequence, as well as an attention module, a global pooling layer, a fully connected layer, and a first convolutional module.
[0053] The five backbone networks connected in sequence are used for feature extraction of the input image at different scales, and the output of the previous backbone network is used as the input of the next backbone network; the first backbone network outputs a feature map with a pixel size of 128*128, the second backbone network outputs a feature map with a pixel size of 64*64, the third backbone network outputs a feature map with a pixel size of 32*32, the fourth backbone network outputs a feature map with a pixel size of 16*16, and the fifth backbone network outputs a feature map with a pixel size of 8*8.
[0054] As shown in Figure 2 The backbone network includes a second convolutional module, a plurality of fast feature extraction modules, a third convolutional module, and a fourth convolutional module connected in sequence; in this embodiment, the convolutional kernel of the second convolutional module is 3×3, and the convolutional kernels of the third convolutional module and the fourth convolutional module are 1×1; the backbone network includes 11 fast feature extraction (SpeedyBlock) modules, which can efficiently extract multi-level features and accurately estimate the eye state while ensuring computational efficiency.
[0055] As shown in Figure 3As shown in the figure, the fast feature extraction module includes two branches. One branch performs depth convolution operation and pointwise convolution operation on the input image in sequence to obtain the first feature map. This branch uses a depth convolution module with a convolution kernel of 3×3 (3×3 DW Conv) and a pointwise convolution module with a convolution kernel of 1×1 (1x1 Conv) to process the image, which can reduce the number of channels of the tensor between the depth convolution module and the pointwise convolution module, facilitating the reduction of computational cost and the improvement of feature extraction speed.
[0056] The other branch performs max pooling operation and channel adjustment operation on the input image in sequence to obtain the second feature map, and splices the first feature map and the second feature map to output a fused feature map.
[0057] The attention module is used to extract the key information of the feature map output by the fourth backbone network and ignore the irrelevant information, enhancing the robustness of the model; in this embodiment, as Figure 4 shown, the attention module performs convolution operation (1x1 Conv), normalization processing (BN), and activation function (Relu) processing on the feature map output by the fourth backbone network in sequence.
[0058] The global pooling layer (GP) is used to map the feature map output by the attention module to a low-dimensional space, that is, map the multi-dimensional feature to a low-dimensional space, reducing the number of parameters and retaining important features.
[0059] The first convolution module is used to extract features from the feature map output by the fifth backbone network; in this embodiment, the convolution kernel of the first convolution module is 1×1.
[0060] The fully connected layer is used to connect the feature map output by the global pooling layer and the feature map output by the first convolution module to output an eye state feature map. In this embodiment, the input of the fully connected layer is 896 neurons and the output is 1 neuron, which is used for the final eye state prediction.
[0061] In this embodiment, the Light EyeNet network is built in a Pytorch environment on a Linux system with a CPU Intel® Core™ i7-8750H 2.21GHz and 32GB of memory, and the graphics card is GeForce RTX 2080Ti.
[0062] And the Adam optimizer is used to train the Light EyeNet network, and the training process is optimized by carefully selected hyperparameters. Two key hyperparameters of the Adam optimizer and are set to 0.9 and 0.999 respectively; takes the value of 1, The value is 3; the learning rate is 0.001; the batch size is set to 64, and the epoch is set to 20; when training, if the accuracy does not increase after more than 5 epochs, the early stopping strategy is adopted to stop training.
[0063] In this embodiment, in order to reflect the test speed of the Light EyeNet network, the following ablation experiments are carried out: on the same training set, five groups of transfer learning are additionally carried out to train the eye state classifier. They are to perform transfer learning to train the eye state estimation network based on MobileNetV2, MobileNetV3, ShuffleNetV2(x1.0), ShuffleNetV2, and the proposed MobileViT(xxs). The last classification layer and the penultimate fully connected layer of these five networks are removed respectively, and then a custom fully connected layer is connected. The output is a single neuron (i.e., the predicted value of the eye state). The training image size is 224×224, and the data preprocessing method and loss function are the same as those of the Light EyeNet network. The test results of different algorithms on the test set are shown in Table 1. Among them, MAdds is the number of multiply-accumulate operations, which approximately contains 2 FLOPs (floating-point operation times).
[0064] Table 1 Test results of different algorithms
[0065]
[0066] In this experiment, for the eye state classification task, six different network models are used for transfer learning training on the same training set. By comparing the performance of these models on the test set, we can conclude that in terms of computing resource consumption: from the three indicators of the number of parameters (Params), the number of multiply-accumulate operations (MAdds), and the CPU inference time, Light EyeNet shows significant advantages. Light EyeNet has the fewest number of parameters, only 91.02K, the MAdds is also relatively low, 31.17M, and at the same time the CPU inference time is only 9.8ms, which is much smaller than other models. This shows that Light EyeNet has great advantages in computing resource consumption and is suitable for use on mobile devices with limited resources.
[0067] In summary, we can see that the Light EyeNet network is superior to the other five models in terms of CPU inference time, the number of parameters, and MAdds. Therefore, according to different application requirements, a suitable model can be selected. Especially for mobile devices with limited computing resources, it is preferred to consider using Light EyeNet.
[0068] To comprehensively evaluate the performance and effectiveness of Light EyeNet, this application designed a series of comparative experiments, and compared the method of combining Mediapipe with Light EyeNet in detail with the improved model based on the SSD network, the method based on the cascade regression framework, SeetaFace6 Open EyeStateDetector, and the third-order hourglass network. Specifically, these methods first use advanced face key point localization technology to accurately determine the position of the eye region, then intercept the eye region based on this key point information, and conduct detailed state classification on it. In contrast, the improved model based on the SSD network and the method based on the cascade regression framework directly use regression technology to simultaneously predict the eye position and eye state, which provides a more efficient and concise detection process. These different method systems have demonstrated their unique advantages and applicability in the field of eye state detection, providing us with rich perspectives and effective comparisons.
[0069] The experimental results of these five groups of algorithms are shown in Table 2:
[0070] Table 2 Comparative experimental results of Light EyeNet
[0071]
[0072] It can be seen from the experimental comparison results that Mediapipe+Light EyeNet is superior to the other four groups of algorithms in both inference speed and accuracy.
[0073] This embodiment also provides an estimation method for estimating the eye state using the Light EyeNet network established by the above method.
[0074] An estimation method for eye state, including the following steps executed in sequence:
[0075] S1: Obtain an experimental data set, perform face key point localization on each picture in the experimental data set, and obtain an eye data set;
[0076] In this embodiment, the experimental data includes two parts. One part is from the publicly available online data set CEW (Closed eyes in the wild), and the other part is crawled samples from the network and taken by a camera. After sorting, the final data set contains 10,000 pictures, and the ratio of open eyes to closed eyes is 1:1.
[0077] The steps of face key point localization in step S1 are as follows:
[0078] S1-1: Input each image of the experimental data set into the Facemesh module of Mediapipe, and the Facemesh module outputs the predicted key points; in this embodiment, the facial key points are located as follows Figure 5 As shown in the figure, the Facemesh module outputs 468 key points for each image. The Facemesh module can perform face detection and key point positioning under various unfavorable conditions such as uneven lighting, tilted faces, and faces wearing masks. It is suitable for deployment on devices with limited computing resources.
[0079] S1-2: Select one of the eye corner points corresponding to the eye from each key point ( , ) and small eye point ( , ), the eye corner here is the general name of the canthus, the inner canthus (near the bridge of the nose) is called the big canthus, and the outer canthus (near the temple) is called the small canthus. The following formula is used to calculate the width of the eye area :
[0080] ;
[0081] S1-3: According to the width of the eye area Expand the large eye corner points and small eye corner points to obtain the expanded coordinates of the eye area ( , ), , ), the specific expansion process is shown in the following formula:
[0082] ;
[0083] ;
[0084] ;
[0085] ;
[0086] in, Indicates the external expansion ratio value; in this embodiment ;
[0087] S1-4: Using the methods in S1-2 to S1-3 above to obtain the outer expansion coordinates of another eye area;
[0088] S1-5: The eye area image is intercepted using the outer expansion coordinates corresponding to the two eyes as the eye data set.
[0089] S2: Preprocess the eye dataset to obtain a dataset of 128×128 pixels, which includes a training image set and a test image set; the dataset is divided into a training set and a test set, where the training set contains 8,000 pictures and the test set contains 2,000 pictures.
[0090] In this embodiment, the preprocessing of the eye dataset includes random cropping, border padding, random color jittering, and pixel normalization.
[0091] S3: Input the training image set into the Light EyeNet network established by the method for establishing an eye state estimation network described in any of the above items for training, and output an eye state feature map;
[0092] S4: Calculate the loss function of the Light EyeNet network, and perform backpropagation on the Light EyeNet network using this loss function to obtain an adjusted Light EyeNet network;
[0093] In this embodiment, in step S4, the loss function is calculated based on the combination of binary cross-entropy loss and Focal loss, and the specific calculation formula is as follows:
[0094] ;
[0095] ;
[0096] ;
[0097] where, represents the binary cross-entropy loss, represents the Focal loss, represents is the true label, is the predicted probability, , are the adjustment parameters in; using this loss function for network adjustment accelerates the network convergence speed and improves the estimation accuracy of the Light EyeNet network at the same time.
[0098] S5: Input the test image set into the adjusted Light EyeNet network for testing, and output the test results, as Figure 6 shown.
[0099] In this embodiment, an eye state estimation system for implementing the above method is also provided.
[0100] An eye state estimation system includes a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, the eye state estimation method described in any one of the above is implemented.
[0101] In summary, the present application uses the fast feature extraction module SpeedyBlock for feature extraction, reduces the number of channels, lowers the computational cost, and improves the extraction speed. Through the carefully designed convolutional layers and the fast feature extraction module SpeedyBlock, the backbone network achieves accurate estimation of the eye state while ensuring computational efficiency. The Light EyeNet network increases the breadth of the model through feature fusion between different levels, applies the attention mechanism to extract key information and ignores irrelevant information, thereby improving the robustness of the model.
[0102] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. Method for establishing an eye state estimation network, characterized in that: Build the Light EyeNet network, which includes five sequentially connected backbone networks, as well as an attention module, a global pooling layer, a fully connected layer, and a first convolutional module; The five sequentially connected backbone networks are used to extract features of different scales from the input image, and the output of the previous backbone network is used as the input of the next backbone network; The backbone network includes a second convolutional module, multiple fast feature extraction modules, a third convolutional module, and a fourth convolutional module connected in sequence; The fast feature extraction module includes two branches. One branch performs depth convolution operations and pointwise convolution operations on the input image in sequence to obtain a first feature map. The other branch performs max pooling operations and channel adjustment operations on the input image in sequence to obtain a second feature map. The first feature map and the second feature map are concatenated and then the fused feature map is output; The attention module is used to extract the key information of the feature map output by the fourth backbone network; The global pooling layer is used to map the feature map output by the attention module to a low-dimensional space; The first convolutional module is used to extract features from the feature map output by the fifth backbone network; The fully connected layer is used to perform feature connection on the feature map output by the global pooling layer and the feature map output by the first convolutional module to output an eye state feature map.
2. The method for establishing an eye state estimation network according to claim 1, wherein: The attention module performs convolution operations, normalization processing, and activation function processing on the feature map output by the fourth backbone network in sequence.
3. The method for establishing an eye state estimation network according to claim 2, characterized in that: The backbone network includes 11 fast feature extraction modules.
4. A method for estimating the eye state, characterized in that: It includes the following steps executed in sequence: S1: Obtain an experimental dataset, perform face key point localization on each picture in the experimental dataset, and obtain an eye dataset; S2: Preprocess the eye dataset to obtain a dataset of 128×128 pixels, which includes a training image set and a test image set; S3: Input the training image set into the Light EyeNet network established by the method for establishing an eye state estimation network as described in any one of claims 1-3 above for training, and output an eye state feature map; S4: Calculate the loss function of the Light EyeNet network, perform backpropagation on the Light EyeNet network using the loss function, and obtain an adjusted Light EyeNet network; S5: Input the test image set into the adjusted Light EyeNet network for testing, and output the test results.
5. The method for estimating the eye state according to claim 4, wherein: The steps of face key point localization in step S1 are as follows: S1-1: Input each picture of the experimental dataset into the Facemesh module of Mediapipe, and the Facemesh module outputs the predicted key points; S1-2: Select one of the large eye corner points corresponding to one of the eyes from each key point ( , ) and the small eye corner points ( , ), and calculate the width of the eye region using the following formula : ; S1-3: According to the width of the eye region Expand the outer corner points of the large and small eyes, and obtain the expanded coordinates of the eye region ( , ), ( , ). The specific expansion process is shown in the following formula: ; ; ; ; Among them, represents the external expansion ratio value; S1-4: Use the method in S1-2 to S1-3 above to obtain the outer expansion coordinates of the other eye region; S1-5: Intercept the eye region pictures corresponding to the two eyes with the outer expansion coordinates as the eye dataset.
6. The method for estimating the eye state according to claim 5, wherein: In step S2, the preprocessing of the eye dataset includes random cropping, border padding, random color jitter, and pixel normalization.
7. The method for estimating the eye state according to claim 6, wherein: In step S4, the loss function is calculated based on the combination of binary cross-entropy loss and Focal loss, and the specific calculation formula is as follows: The calculation is as follows: ; ; ; Among them, represents binary cross-entropy loss, represents Focal loss, is the true label, is the predicted probability, and is the adjustment parameter in 8. An eye state estimation system, comprising a memory and a processor, wherein a computer program is stored on the memory, and characterized in that: When the computer program is executed by a processor, it implements the method for estimating the eye state according to any one of claims 4-7 above.
Citation Information
Patent Citations
Method for generating detection model and method for detecting eye state by using detection model
CN114283488A
Eye fatigue detection algorithm
CN118512150A