Establishment method of eye state estimation network and estimation method of eye state
By establishing a Light EyeNet network, using the fast feature extraction module and attention mechanism, the existing eye state detection method is easily disturbed by external interference and slow speed, and efficient and robust eye state detection is achieved.
Patent Information
- Application Number
- CN202510517556.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing eye status detection methods are susceptible to adverse factors such as external light, noise and face tilt, and the detection speed is slow.
The Light EyeNet network is adopted, which includes five sequentially connected backbone networks, attention modules, global pooling layer, fully connected layer and the first convolutional module. The key information is extracted through the rapid feature extraction module and attention mechanism to enhance the robustness of the model.
It improves detection speed, reduces computing costs, increases the breadth and robustness of the model, and is suitable for use on mobile devices with limited resources.
Smart Images

Figure CN120047991A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of eye state detection, and specifically relates to a method for establishing an eye state estimation network and a method for estimating an eye state. Background Art
[0002] The eyes are regarded as an important window for human emotions and states. Eye State Estimation (ESE) has important research and application values in the fields of human-computer interaction, sleep research, and fatigue driving detection.
[0003] Current eye state estimation methods can be roughly divided into two categories: those based on handcrafted features and those based on deep features. In the methods based on handcrafted features, the Viola-Jones detector locates the position of the eyes by extracting Haar features around the eyes and training a cascade classifier based on Adaboost. In the methods based on deep features, the Hough transform is used to locate the iris and pupil to determine the eye region, or the Variance Projection Function (VPF) is used to locate the key points of the eyes to determine the position and shape of the eyes.
[0004] These early eye state detection methods mainly rely on handcrafted features and are easily affected by adverse factors such as external light, noise, and face tilt. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for establishing an eye state estimation network and a method for estimating an eye state that can improve the detection speed.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A method for establishing an eye state estimation network, which establishes a Light EyeNet network. The Light EyeNet network includes five backbone networks connected in sequence, as well as an attention module, a global pooling layer, a fully connected layer, and a first convolutional module; The five backbone networks connected in sequence are used to extract features of different scales from the input image, and the output of the previous backbone network is used as the input of the next backbone network; The backbone network includes a second convolutional module, a plurality of fast feature extraction modules, a third convolutional module, and a fourth convolutional module connected in sequence; The fast feature extraction module includes two branches. One branch performs depth convolution operations and pointwise convolution operations on the input image in sequence to obtain a first feature map, and the other branch performs max pooling operations and channel adjustment operations on the input image in sequence to obtain a second feature map. The first feature map and the second feature map are concatenated and then the fused feature map is output; The attention module is used to extract the key information of the feature map output by the fourth backbone network; The global pooling layer is used to map the feature map output by the attention module to a low-dimensional space; The first convolution module is used to extract features from the feature map output by the fifth backbone network; The fully connected layer is used to perform feature connection on the feature map output by the global pooling layer and the feature map output by the first convolution module to output an eye state feature map.
[0007] Preferably, the attention module sequentially performs convolution operation, normalization processing, and activation function processing on the feature map output by the fourth backbone network.
[0008] Preferably, the backbone network includes 11 fast feature extraction modules.
[0009] A method for estimating eye state includes the following steps executed sequentially: S1: Obtain an experimental data set, perform face key point localization on each picture in the experimental data set, and obtain an eye data set; S2: Preprocess the eye data set to obtain a data set of 128×128 pixels, which includes a training image set and a test image set; S3: Input the training image set into the Light EyeNet network established by the method for establishing an eye state estimation network as described in any one of the above to perform training, and output an eye state feature map; S4: Calculate the loss function of the Light EyeNet network, perform backpropagation on the Light EyeNet network using the loss function, and obtain an adjusted Light EyeNet network; S5: Input the test image set into the adjusted Light EyeNet network for testing, and output the test result.
[0010] Preferably, the steps of face key point localization in step S1 are as follows: S1-1: Input each picture of the experimental data set into the Facemesh module of Mediapipe, and the Facemesh module outputs the predicted key points; S1-2: Select one large eye corner point corresponding to one eye from the key points ( , ) and one small eye corner point ( , ), and calculate the width of the eye region using the following formula : ; S1-3: According to the width of the eye region expand the inner corner point and the outer corner point of the eye to obtain the expanded coordinates of the eye region ( , ), ( , ). The specific expansion process is shown in the following formula: ; ; ; ; wherein, represents the expansion ratio value; S1-4: Use the method in S1-2 to S1-3 above to obtain the expanded coordinates of the other eye region; S1-5: Crop the eye region images corresponding to the two eyes with the expanded coordinates as the eye dataset.
[0011] Preferably, in step S2, the preprocessing of the eye dataset includes random cropping, border padding, random color jittering, and pixel normalization.
[0012] Preferably, in step S3, the loss function is calculated based on the combination of binary cross-entropy loss and Focal loss . The specific calculation formula is as follows: ; ; ; wherein, represents the binary cross-entropy loss, represents the Focal loss, is the true label, is the predicted probability, , are the adjustment parameters in
[0013] An eye state estimation system includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, the eye state estimation method described in any one of the above is implemented.
[0014] By adopting the foregoing design scheme, the beneficial effect of the present invention is that: the present application uses a fast feature extraction module for feature extraction, reduces the number of channels, reduces the calculation cost, and improves the extraction speed; The Light EyeNet network increases the breadth of the model through feature fusion between different levels, applies the attention mechanism to extract key information and ignores irrelevant information, and improves the robustness of the model. Description of the Drawings
[0015] Figure 1 Schematic diagram of the structure of the Light EyeNet network of the present invention; Figure 2 Schematic diagram of the structure of the backbone network of the present invention; Figure 3 Schematic diagram of the structure of the fast feature extraction module of the present invention; Figure 4 Schematic diagram of the structure of the attention module of the present invention; Figure 5 Schematic diagram of the face key point positioning of the present invention; Figure 6 Schematic diagram of the eye state estimation of the present invention. Detailed Description of the Invention
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0017] The terms "first", "second", "third", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0018] A method for establishing an eye state estimation network, establishing a Light EyeNet network as shown in Figure 1 The Light EyeNet network includes five backbone networks connected in sequence, as well as an attention module, a global pooling layer, a fully connected layer, and a first convolutional module.
[0019] Five sequentially connected backbone networks are used to extract features of different scales from the input image, and the output of the previous backbone network is used as the input of the next backbone network; the output of the first backbone network is a feature map with a pixel size of 128*128, the output of the second backbone network is a feature map with a pixel size of 64*64, the output of the third backbone network is a feature map with a pixel size of 32*32, the output of the fourth backbone network is a feature map with a pixel size of 16*16, and the output of the fifth backbone network is a feature map with a pixel size of 8*8.
[0020] As Figure 2 shown, the backbone network includes a second convolution module, a plurality of fast feature extraction modules, a third convolution module, and a fourth convolution module connected in sequence; in this embodiment, the convolution kernel of the second convolution module is 3×3, and the convolution kernels of the third convolution module and the fourth convolution module are 1×1; the backbone network includes 11 fast feature extraction (SpeedyBlock) modules, which can efficiently extract multi-level features and accurately estimate the eye state while ensuring computational efficiency.
[0021] As Figure 3 shown, the fast feature extraction module includes two branches. One branch performs depth convolution operation and pointwise convolution operation on the input image in sequence to obtain the first feature map. This branch uses a depth convolution module (3×3 DW Conv) with a convolution kernel of 3×3 and a pointwise convolution module (1x1 Conv) with a convolution kernel of 1×1 to process the image, which can reduce the number of channels of the tensor between the depth convolution module and the pointwise convolution module, is beneficial to reducing the computational cost, and improving the feature extraction speed.
[0022] The other branch performs max pooling operation and channel adjustment operation on the input image in sequence to obtain the second feature map, and the first feature map and the second feature map are concatenated and then the fused feature map is output.
[0023] The attention module is used to extract the key information of the feature map output by the fourth backbone network and ignore the irrelevant information, enhancing the robustness of the model; in this embodiment, as Figure 4 shown, the attention module performs convolution operation (1x1 Conv), normalization processing (BN), and activation function (Relu) processing on the feature map output by the fourth backbone network in sequence.
[0024] The global pooling layer (GP) is used to map the feature map output by the attention module to a low-dimensional space, that is, map the multi-dimensional features to a low-dimensional space, reducing the number of parameters and retaining the important features.
[0025] The first convolution module is used to extract features from the feature map output by the fifth backbone network; in this embodiment, the convolution kernel of the first convolution module is 1×1.
[0026] This fully connected layer is used to perform feature connection on the feature map output by the global pooling layer and the feature map output by the first convolutional module to output the eye state feature map. In this embodiment, the input of the fully connected layer is 896 neurons, and the output is 1 neuron, which is used for the final eye state prediction.
[0027] In this embodiment, the Light EyeNet network is built in a Pytorch environment with a CPU Intel® Core™ i7-8750H 2.21GHz and 32GB of memory under the Linux system, and the graphics card is GeForce RTX 2080Ti.
[0028] And the Adam optimizer is used to train the Light EyeNet network, and the training process is optimized by carefully selected hyperparameters. The two key hyperparameters of the Adam optimizer and are set to 0.9 and 0.999 respectively; takes the value of 1, takes the value of 3; the learning rate is 0.001; the batch size is set to 64, and the epoch is set to 20; when the accuracy does not increase after more than 5 epochs during training, the early stopping strategy is adopted to stop training.
[0029] In this embodiment, in order to reflect the test speed of the Light EyeNet network, the following ablation experiments are carried out: five groups of transfer learning are additionally carried out on the same training set to train the eye state classifier, which are based on MobileNetV2, MobileNetV3, ShuffleNetV2 (x1.0), ShuffleNetV2, and the proposed MobileViT (xxs) to transfer learning to train the eye state estimation network. The last classification layer and the penultimate fully connected layer of these five networks are removed respectively, and then a custom fully connected layer is connected. The output is all one neuron (i.e., the eye state prediction value). The training image size is 224×224, and the data preprocessing method and loss function are the same as those of the Light EyeNet network. The test results of different algorithms on the test set are shown in Table 1, where MAdds is the number of multiply-accumulate operations, which approximately contains 2 FLOPs (floating-point operation counts).
[0030] Table 1 Test results of different algorithms
[0031] In this experiment, for the eye state classification task, six different network models were used for transfer learning training on the same training set. By comparing the performance of these models on the test set, we can conclude that in terms of computational resource consumption: from the three indicators of the number of parameters (Params), the number of multiply-accumulate operations (MAdds), and the CPU inference time, Light EyeNet shows significant advantages. Light EyeNet has the fewest parameters, only 91.02K, and also has a relatively low MAdds of 31.17M. At the same time, the CPU inference time is only 9.8ms, which is much less than that of other models. This indicates that Light EyeNet has great advantages in computational resource consumption and is suitable for use on mobile devices with limited resources.
[0032] In summary, we can see that the Light EyeNet network is superior to the other five models in terms of CPU inference time, the number of parameters, and MAdds. Therefore, according to different application requirements, a suitable model can be selected. Especially for mobile devices with limited computational resources, Light EyeNet should be given priority.
[0033] To comprehensively evaluate the performance and effectiveness of Light EyeNet, this application designed a series of comparative experiments, and compared the method of combining Mediapipe with Light EyeNet in detail with the improved model based on the SSD network, the method based on the cascade regression framework, SeetaFace6 Open EyeStateDetector, and the third-order hourglass network. Specifically, these methods first use advanced face key point localization technology to accurately determine the position of the eye region, and then intercept the eye region based on this key point information and conduct detailed state classification. In contrast, the improved model based on the SSD network and the method based on the cascade regression framework directly use regression technology to simultaneously predict the eye position and eye state, which provides a more efficient and concise detection process. These different method systems have all demonstrated their unique advantages and applicability in the field of eye state detection, providing us with rich perspectives and effective comparisons.
[0034] The experimental results of these five groups of algorithms are shown in Table 2: Table 2 Comparative experimental results of Light EyeNet
[0035] From the experimental comparison results, it can be seen that Mediapipe + Light EyeNet is superior to the other four groups of algorithms in terms of inference speed and accuracy.
[0036] This embodiment also provides an estimation method for estimating the eye state using the Light EyeNet network established by the above method.
[0037] An estimation method for eye state includes the following steps executed in sequence: S1: Obtain an experimental data set, perform face key point localization on each picture in the experimental data set, and obtain an eye data set; In this embodiment, the experimental data includes two parts. One part is from the publicly available data set CEW (Closed eyes in the wild) on the Internet, and the other part is samples crawled from the network and taken by a camera. After sorting, the final data set contains 10,000 pictures, and the ratio of open eyes to closed eyes is 1:1.
[0038] The steps of face key point localization in step S1 are as follows: S1-1: Input each picture of the experimental data set into the Facemesh module of Mediapipe, and the Facemesh module outputs the predicted key points; in this embodiment, the face key point localization is as Figure 5 shown. The Facemesh module outputs 468 key points for each picture. The Facemesh module can perform face detection and key point localization under various adverse conditions such as uneven lighting, tilted face, and wearing a mask on the face, and is suitable for deployment on devices with limited computing resources.
[0039] S1-2: Select the large canthus point ( , ) and the small canthus point ( , ) corresponding to one of the eyes from the key points. Here, the canthus is the general term for the corner of the eye. The inner canthus (near the bridge of the nose) is called the large canthus, and the outer canthus (near the temple) is called the small canthus. And use the following formula to calculate the width of this eye area : ; S1-3: Expand the large canthus point and the small canthus point according to the width of this eye area to obtain the expanded coordinates ( , ) of this eye area, ( , ). The specific expansion process is shown in the following formula: ; ; ; ; Where represents the external expansion ratio value; in this embodiment ; S1-4: Obtain the external expansion coordinates of the other eye region by using the methods in S1-2 to S1-3 above; S1-5: Crop the eye region pictures according to the external expansion coordinates corresponding to the two eyes as the eye dataset.
[0040] S2: Preprocess the eye dataset to obtain a dataset of 128×128 pixels, which includes a training image set and a test image set; the dataset is divided into a training set and a test set, where the training set contains 8000 pictures and the test set contains 2000 pictures.
[0041] In this embodiment, the preprocessing of the eye dataset includes random cropping, border padding, random color jittering, and pixel normalization.
[0042] S3: Input the training image set into the Light EyeNet network established by the method for establishing the eye state estimation network described in any one of the above to perform training, and output the eye state feature map; S4: Calculate the loss function of the Light EyeNet network, and perform backpropagation on the Light EyeNet network with this loss function to obtain the adjusted Light EyeNet network; In this embodiment, in step S3, the loss function is calculated based on the combination of binary cross-entropy loss and Focal loss The specific calculation formula is as follows: ; ; ; where, represents the binary cross-entropy loss, represents the Focal loss, and represents is the true label, is the predicted probability, , is the adjustment parameter in; using this loss function for network adjustment accelerates the convergence speed of the network and improves the estimation accuracy of the Light EyeNet network at the same time.
[0043] S5: Input the test image set into the adjusted Light EyeNet network for testing, and output the test results, as Figure 6 shown.
[0044] In this embodiment, an eye state estimation system for implementing the above method is also provided.
[0045] An eye state estimation system includes a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, the eye state estimation method described in any one of the above is implemented.
[0046] In summary, this application uses the fast feature extraction module SpeedyBlock for feature extraction, reduces the number of channels, lowers the computational cost, and improves the extraction speed; the backbone network, through the carefully designed convolutional layers and the fast feature extraction module SpeedyBlock, achieves accurate estimation of the eye state while ensuring computational efficiency; the Light EyeNet network increases the breadth of the model through feature fusion between different levels, applies the attention mechanism to extract key information and ignores irrelevant information, and improves the robustness of the model.
[0047] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for establishing an eye state estimation network, characterized in that: Build a Light EyeNet network, which includes five sequentially connected backbone networks, as well as an attention module, a global pooling layer, a fully connected layer, and the first convolutional module; The five backbone networks connected in sequence are used to extract features of different scales from the input image, and the output of the previous backbone network serves as the input of the next backbone network; The backbone network includes a second convolution module, a plurality of fast feature extraction modules, a third convolution module and a fourth convolution module which are connected in sequence; The fast feature extraction module includes two branches, one of which performs a depth convolution operation and a point-by-point convolution operation on the input image in sequence to obtain a first feature map, and the other branch performs a maximum pooling operation and a channel adjustment operation on the input image in sequence to obtain a second feature map, and the first feature map is spliced with the second feature map to output a fused feature map; The attention module is used to extract the key information of the feature map output by the fourth backbone network; The global pooling layer is used to map the feature map output by the attention module to a low-dimensional space; The first convolution module is used to extract features from the feature map output by the fifth backbone network; The fully connected layer is used to perform feature connection on the feature map output by the global pooling layer and the feature map output by the first convolution module to output an eye state feature map.
2. The method for establishing an eye state estimation network as claimed in claim 1, characterized in that: The attention module sequentially performs convolution operation, normalization processing and activation function processing on the feature map output by the fourth backbone network.
3. The method for establishing an eye state estimation network as claimed in claim 2, characterized in that: The backbone network includes 11 fast feature extraction modules.
4. A method for estimating eye status, characterized in that: The process includes the following steps: S1: Acquire an experimental data set, locate facial key points on each image in the experimental data set, and acquire an eye data set; S2: preprocessing the eye data set to obtain a data set of 128×128 pixels, where the data set includes a training image set and a test image set; S3: inputting the training image set into the Light EyeNet network established by the method for establishing an eye state estimation network as described in any one of claims 1 to 3 for training, and outputting an eye state feature map; S4: Calculate the loss function of the Light EyeNet network, use the loss function to back-propagate the Light EyeNet network, and obtain an adjusted Light EyeNet network; S5: Input the test image set into the adjusted Light EyeNet network for testing and output the test results.
5. The method for estimating eye state as claimed in claim 4, characterized in that: The steps of locating facial key points in step S1 are as follows: S1-1: Input each image of the experimental data set into the Facemesh module of Mediapipe respectively, and the Facemesh module outputs the predicted key points; S1-2: Select one of the eye corner points corresponding to the eye from each key point ( , ) and small eye point ( , ), and use the following formula to calculate the width of the eye area : ; S1-3: According to the width of the eye area Expand the large eye corner points and small eye corner points to obtain the expanded coordinates of the eye area ( , ), , ), the specific expansion process is shown in the following formula: ; ; ; ; in, Indicates the expansion ratio value; S1-4: Using the methods in S1-2 to S1-3 above to obtain the outer expansion coordinates of another eye area; S1-5: The eye area image is intercepted using the outer expansion coordinates corresponding to the two eyes as the eye data set.
6. The method for estimating eye state according to claim 5, characterized in that: In step S2, preprocessing of the eye data set includes random cropping, edge filling, random color jittering and pixel normalization.
7. The method for estimating eye state according to claim 6, characterized in that: In step S3, the loss function is based on the combination of binary cross entropy loss and focal loss. The specific calculation formula is as follows: ; ; ; in, represents the binary cross entropy loss, represents Focal loss, is the true label, is the predicted probability, , for The adjustment parameters in .
8. An eye state estimation system, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that: When the computer program is executed by a processor, the method for estimating the eye state as described in any one of claims 4 to 7 is implemented.
Citation Information
Patent Citations
Method for generating detection model and method for detecting eye state by using detection model
CN114283488A
Eye key point detection method, eye state detection method and related device
CN115620383A
Line-of-sight estimation method based on local super-resolution fusion attention mechanism
CN117830783A
Model training method, face key point positioning method and device, equipment and medium
CN117912085A
Eye fatigue detection algorithm
CN118512150A