Eye feature detection and eye closing state discrimination method based on near-eye infrared image

Through the method of eye feature detection and eye-closing state judgment based on near-eye infrared images, the improved YOLOv8n-ASD network and feature fusion technology are used to solve the problem that eye movement monitoring in the prior art is difficult to take into account high real-time and high precision, and efficient and accurate eye movement monitoring on resource-constrained devices is achieved.

CN120108026APending Publication Date: 2025-06-06SHENZHEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510269347.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to take into account high real-time and high accuracy when performing eye movement monitoring, and has poor adaptability to different individual differences and imaging equipment characteristics.

Method used

The eye feature detection and eye-closing state discrimination method based on near-eye infrared images is adopted, and the key points of the eyelid, pupil, and iris are extracted through the improved lightweight YOLOv8n-ASD network, and feature fusion is combined with Gaussian attention heat map and two-dimensional spatial geometric features for high-precision eye-closing state discrimination.

Benefits of technology

It realizes efficient eye movement monitoring on resource-constrained equipment, taking into account high real-time and high accuracy, and adapts to different individual differences and imaging equipment characteristics, providing strong technical support in the fields of fatigue monitoring, eye movement tracking, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108026A_ABST
    Figure CN120108026A_ABST
Patent Text Reader

Abstract

The invention discloses an eye feature detection and eye closing state discrimination method based on a near-eye infrared image, and relates to the field of AI image processing, and the method comprises the steps: S1, obtaining a target near-eye infrared image; s2, inputting the target near-eye infrared image into the target eye feature detection and eye closing state discrimination model to obtain an eye closing state result corresponding to the target near-eye infrared image; wherein the model comprises an eye feature detection sub-model and an eye closing state discrimination sub-model; the eye feature detection sub-model is determined after a CBS module in a lower sampling layer of Backbone and Neck of the YOLOv8n network is replaced by an ADown module and a C2f module is replaced by a C2fSD module; the eye closing state discrimination sub-model is determined after a Gaussian attention mechanism model and a two-dimensional space geometric feature extraction module are added into the MobileNetV2 network; according to the invention, the positioning precision of the eye feature points and the eye closing state discrimination accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI image processing, and in particular to a method for detecting eye features and distinguishing eye-closing states based on near-eye infrared images. Background Art

[0002] In recent years, deep learning technology, with its powerful feature extraction capabilities, can effectively solve the problem of limited expression of traditional manual features. Thanks to this, eye movement monitoring technology has shown increasingly important application value in medical diagnosis, gaze tracking, fatigue monitoring and human-computer interaction.

[0003] In related technologies, most models used for eye movement monitoring rely on large-scale labeled data, which easily leads to the model falling into the dilemma of overfitting or underfitting, and it is difficult to adapt to different individual differences and imaging equipment characteristics. Secondly, models that require high-precision monitoring results usually use complex network architectures, such as deeply stacked residual modules or multi-branch intensive connections, which significantly increase the number of parameters and computing load, not only increasing the hardware resource consumption of training and reasoning, but also making it more difficult to achieve low-latency real-time detection in embedded platforms with limited computing power, such as wearable glasses, portable medical devices, etc.

[0004] Therefore, there is an urgent need for an eye movement monitoring method that can achieve both high real-time performance and high precision. Summary of the invention

[0005] In view of this, the present invention provides a method for eye feature detection and eye closure state discrimination based on near-eye infrared images to solve the technical problem in the related art that it is difficult to achieve both high real-time performance and high precision when performing eye movement monitoring.

[0006] The present invention provides a method for detecting eye features and distinguishing eye-closing state based on near-eye infrared images, the method comprising:

[0007] S1, obtaining a near-eye infrared image of the target;

[0008] S2, inputting the target near-eye infrared image into the target eye feature detection and eye-closing state discrimination model to obtain the eye-closing state result corresponding to the target near-eye infrared image;

[0009] The target eye feature detection and eye closure state discrimination model is a pre-trained neural network; the target eye feature detection and eye closure state discrimination model comprises: a target eye feature detection sub-model, and a target eye closure state discrimination sub-model cascaded with the target eye feature detection sub-model;

[0010] The target eye feature detection submodel is determined after replacing the CBS module in the downsampling layer of Backbone and Neck of the YOLOv8n network with the ADown module, and the C2f module with the C2f_SD module; the target closed eye state discrimination submodel is determined after adding a Gaussian attention mechanism model and a two-dimensional space geometric feature extraction module to the MobileNetV2 network.

[0011] In an optional implementation, the S2 includes:

[0012] S21, inputting the target near-eye infrared image into the target eye feature detection sub-model to obtain the coordinates of multiple key points corresponding to the target near-eye infrared image;

[0013] S22, inputting the target near-eye infrared image and the coordinates of multiple key points corresponding to the target near-eye infrared image into the target closed-eye state discrimination sub-model to obtain the closed-eye state result corresponding to the target near-eye infrared image.

[0014] In an optional implementation, the C2f_SD module is determined by replacing the Bottleneck unit in the C2f module with a StarBlock unit and replacing the ordinary convolution unit with a depth convolution unit.

[0015] In an optional implementation, the target eye feature detection sub-model includes a Backbone network, a Neck network and a Head network;

[0016] The Backbone network includes a CBS module, an ADown module, a C2f_SD module and an SPPF module connected in series in sequence; wherein the ADown module and the C2f_SD module are serially stacked 4 times;

[0017] The Neck network is obtained by replacing all four C2f modules of the Neck network in the YOLOv8n network with C2f_SD modules and all two CBS modules with DWCBS modules; the DWCBS module is determined by replacing the ordinary convolution unit in the CBS module with the deep convolution unit;

[0018] The Head network has the same architecture as the Head network in the YOLOv8n network.

[0019] In an optional embodiment, the target eye closing state discrimination sub-model is obtained by sequentially adding a Gaussian attention mechanism module and a two-dimensional space geometric feature extraction module between the inverted residual block sequence and the global average pooling module in the MobileNetV2 network.

[0020] In an optional implementation, the multiple key point coordinates corresponding to the target near-eye infrared image include: multiple eye face point coordinates, multiple pupil point coordinates and multiple iris point coordinates, and S22 includes:

[0021] S221, selecting a plurality of eye-face point coordinates from a plurality of key point coordinates corresponding to the target near-eye infrared image, and converting the plurality of eye-face point coordinates into a multi-peak Gaussian distribution matrix;

[0022] S222, extracting features from the target near-eye infrared image to obtain a plurality of feature maps corresponding to the target near-eye infrared image;

[0023] S223, performing spatial weighting on the multi-peak Gaussian distribution matrix and multiple feature maps by element-by-element multiplication to obtain CNN deep semantic features;

[0024] S224, calculating the two-dimensional spatial geometric features between the coordinates of multiple eyelid points, and performing a normalization operation on the two-dimensional spatial geometric features;

[0025] S225, fusing the normalized two-dimensional spatial geometric features with the CNN deep semantic features and then classifying them to obtain the closed-eye state result corresponding to the target near-eye infrared image.

[0026] In an optional implementation, the Gaussian attention mechanism module is determined based on a Gaussian heat map; and the two-dimensional space geometric feature extraction module is determined based on Euclidean features.

[0027] In an optional implementation, the S225 includes:

[0028] The normalized two-dimensional spatial geometric features are fused with the CNN deep semantic features by splicing;

[0029] The fused features are mapped to the joint latent space and then classified and predicted to obtain the eye-closing state result corresponding to the target near-eye infrared image.

[0030] The embodiment of the present invention adopts an improved lightweight YOLOv8n-ASD network to extract multiple key points of the eyelid, pupil, and iris from near-eye infrared images, realizes efficient information retention through multi-branch pooling and feature recombination, and integrates star operations and depthwise separable convolution operations to reduce redundant calculations; secondly, based on the eyelid key point coordinates output by the YOLOv8n-ASD network, the embodiment of the present invention integrates the enhanced feature expression constructed by the Gaussian attention heat map and the two-dimensional spatial geometric features into the MobileNetV2 network through a multi-feature fusion mechanism to realize high-precision eye closure state discrimination, and provide strong technical support for fatigue monitoring, eye tracking, VR (virtual reality) and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0032] Figure 1 is a flow chart of a method for detecting eye features and distinguishing eye-closed state based on near-eye infrared images according to an embodiment of the present invention;

[0033] Figure 2 is a diagram of the eye feature detection and eye-closing state discrimination model architecture according to an embodiment of the present invention;

[0034] Figure 3 2. It is a YOLOv8n-ASD network model diagram according to an embodiment of the present invention;

[0035] Figure 4 is a schematic diagram of a downsampling operation flow of an ADown module according to an embodiment of the present invention;

[0036] Figure 5 is a structural comparison diagram of a C2f module according to an embodiment of the present invention and an improved C2f_SD module;

[0037] Figure 6 is a schematic diagram of the operation flow of a StarBlock unit according to an embodiment of the present invention;

[0038] Figure 7 is a schematic diagram of a deep convolution operation flow according to an embodiment of the present invention;

[0039] Figure 8 is a schematic diagram of a point-by-point convolution operation process according to an embodiment of the present invention;

[0040] Fig. 9 is a schematic diagram of a depthwise separable convolution operation process according to an embodiment of the present invention;

[0041] Fig.10 is a diagram of the eye-closing state discrimination sub-model architecture according to an embodiment of the present invention;

[0042] Fig.11 is a visualized Gaussian heat map according to an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0044] Figure 1 It is a method for eye feature detection and eye closure state discrimination based on near-eye infrared images according to an embodiment of the present invention. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0045] like Figure 1 As shown, the process includes the following steps:

[0046] S1. Acquire a near-eye infrared image of the target.

[0047] S2. Input the target near-eye infrared image into the target eye feature detection and closed-eye state discrimination model to obtain the closed-eye state result corresponding to the target near-eye infrared image.

[0048] Among them, the eye feature detection and closed eyes state discrimination model architecture is as follows Figure 2 As shown, the cascade processing flow is composed of the eye feature detection sub-model (YOLOv8n-ASD network) and the closed eye state discrimination sub-model (EyeNet-GE network).

[0049] Specifically, the target eye feature detection and eye closure state discrimination model includes: a target eye feature detection sub-model, and a target eye closure state discrimination sub-model cascaded with the target eye feature detection sub-model. The input of the target eye feature detection and eye closure state discrimination model is a near-eye infrared image, and the output is an eye closure state result corresponding to the near-eye infrared image. Among them, the eye closure state result includes closed eyes and open eyes.

[0050] Secondly, the target eye feature detection sub-model (YOLOv8n-ASD network) is based on the YOLOv8n algorithm improved YOLOv8n-ASD, and the YOLOv8n network architecture is lightweight and feature enhanced. The target closed eye state discrimination sub-model (EyeNet-GE network) is based on the eyelid key points output by the target eye feature detection sub-model, constructs Gaussian attention heat maps and two-dimensional spatial geometric features, and injects them as prior knowledge into a lightweight classification network with MobileNetV2 as the backbone.

[0051] Specifically, the original CBS module (Conv-BatchNorm-SiLU) in the downsampling layer of the backbone network (Backbone) and the neck network (Neck) of the YOLOv8n network is replaced by the ADown module, and the standard C2f module is replaced by the C2f_SD module; the target closed eye state discrimination sub-model is determined after adding the Gaussian attention mechanism model and the two-dimensional space geometric feature extraction module to the MobileNetV2 network.

[0052] It should be noted that when training the eye feature detection and eye closure state discrimination model, the two-level sub-module can achieve end-to-end optimization through a cascade training strategy, and discriminate the eye closure state while detecting the key points of the eye features.

[0053] In an optional embodiment, if Figure 3 As shown in the figure, the target eye feature detection sub-model includes Backbone network, Neck network and Head network.

[0054] Specifically, the Backbone network includes a CBS module, an ADown module, a C2f_SD module and an SPPF module which are connected in series in sequence; wherein the ADown module and the C2f_SD module are serially stacked four times.

[0055] The Neck network is obtained by replacing all four C2f modules of the Neck network in the YOLOv8n network with C2f_SD modules and all two CBS modules with DWCBS modules; the DWCBS module is determined by replacing the ordinary convolution unit (Conv) in the CBS module with a deep convolution unit (DWConv).

[0056] The Head network has the same architecture as the Head network in the YOLOv8n network.

[0057] Among them, the CBS module in the original YOLOv8n network is replaced by the ADown module, which is aimed at the downsampling stage (such as the feature dimension reduction operation at the cross-layer connection in Backbone). The ADown module is used to replace the original CBS model to achieve spatial resolution compression, which can optimize the feature downsampling process and balance the model efficiency and detection accuracy. The following is a description of the ADown module:

[0058] The downsampling operation process of the ADown module is as follows Figure 4As shown in the figure, assuming that the dimension of the given input feature map is C, the ADown module first divides the input feature map into two sub-feature maps of equal dimensions along the channel dimension through channel splitting, and performs lightweight convolutions with different parameters, such as strided convolution, 1×1 convolution, etc. Among them, the first branch uses average pooling to compress the spatial dimension, and then performs channel alignment and nonlinear feature enhancement through the CBS module; the second branch implements spatial downsampling through maximum pooling, and also performs feature optimization through the CBS module. Finally, the output feature maps of the two branches are fused through channel splicing to obtain the final downsampled feature map. This process captures feature responses of different granularities through parallel pooling operations (average pooling retains global information, and maximum pooling focuses on significant features), and dynamically adjusts the feature distribution in combination with learnable convolution kernel parameters, while reducing the spatial resolution while maintaining the integrity of multi-scale information, thereby providing high information density feature expression for subsequent tasks.

[0059] Compared with the traditional downsampling method, the ADown module of the embodiment of the present invention adopts a multi-branch parallel processing and feature reorganization strategy in the process of reducing the spatial resolution of the feature map, which maximizes the retention of detail information and spatial correlation in the original image, and provides richer semantic clues for subsequent target detection tasks. At the same time, it can also significantly reduce redundant parameters and computational complexity, thereby alleviating the information loss problem caused by downsampling in deep networks, and showing better deployment adaptability in resource-constrained devices (such as mobile terminals or edge computing platforms).

[0060] In addition, the learnable convolution kernel parameters within the ADown module can dynamically adjust the feature extraction mode through end-to-end training, enabling it to adapt to different data distributions and optimize the fusion effect of multi-scale features in complex scenarios. It can take into account both high efficiency and high performance, enabling the ADown module to significantly improve the robustness of the YOLOv8n-ASD network model to scale changes of near-eye infrared images while maintaining a low computing cost.

[0061] In the related technology, the C2f module in the YOLOv8n network works with modules such as Conv and Bottleneck to complete the feature extraction task of the input image. The convolutional layer and other feature extraction layers in C2f have local receptive fields. In the case of occlusion, the local receptive field cannot capture enough features. Occlusion not only leads to direct loss of feature information, but also causes aliasing between features. In addition, the C2f module has a large amount of calculation and parameters, which affects the detection speed of the model.

[0062] Based on this, the embodiment of the present invention replaces the C2f module with the lightweight C2f_SD module, so as to more comprehensively capture the feature information in the image, enhance the semantic features of the occluded target and reduce the overall calculation amount and parameter amount.

[0063] In an optional implementation, the C2f_SD module is determined by replacing the Bottleneck unit in the C2f module with a StarBlock unit, and replacing the ordinary convolution unit Conv with a deep convolution unit DWConv.

[0064] Specifically, Figure 5 As shown in the figure, the left side is the YOLOv8 standard C2f module, and the right side is the improved C2f_SD module (also known as the lightweight feature fusion module). The improvements include two points: (1) The original Bottleneck unit is replaced by the StarBlock unit, which realizes dynamic calibration of features by introducing element-by-element multiplication operations and enhances the interaction ability of local features. (2) The ordinary convolution unit is replaced by the deep convolution unit, which greatly reduces the computational complexity by decoupling spatial convolution and channel convolution.

[0065] The following is an explanation of the operation process of the Starblock unit:

[0066] like Figure 6 As shown in the figure, the Starblock unit uses a 7×1 depthwise convolution unit (DWConv) to complete input feature extraction, which can effectively reduce the amount of calculation; then batch normalization (BatchNorm) is used to efficiently achieve output standardization; two 1×1 convolutions with an expansion factor of 4 are used to increase the feature dimension, and one of them is selected to use ReLU6 for activation; then element multiplication (i.e., Star Operation) is performed on the two outputs to achieve global interaction while exponentially amplifying the implicit dimension of the feature; then 1×1 Conv convolution and 7×1 depthwise convolution (DWConv) are used in turn to achieve feature dimensionality reduction and extraction.

[0067] It should be noted that DSConv (Depthwise Separable Convolution) is called depthwise separable convolution, which is a convolution operation commonly used in neural networks. It consists of two parts: depthwise convolution (Depthwise CONV, DWConv) and pointwise convolution (Pointwise Conv, PWConv). Figure 7 As shown in Figure 1, in a depth convolution, each input channel is convolved with a separate filter (kernel), which means that each input channel generates a corresponding output channel. Depth convolution is mainly used to capture the spatial information of the input data.

[0068] like Figure 8As shown in the figure, point-by-point convolution is a 1×1 convolution operation that convolves all channels of the input at each position. Point-by-point convolution can be regarded as a convolution operation performed on the channel dimension of the input data without involving spatial information. It is used to linearly combine the feature maps of each channel generated by the deep convolution.

[0069] The depth-wise separable convolution operation is as follows: Fig. 9 As shown in the figure, the depthwise separable convolution splits the convolution process into two stages: depthwise convolution and pointwise convolution. In the depthwise convolution, each input channel is independently spatially convolved with a single filter to extract local features within the channel. The main advantage of using depthwise convolution is that it reduces the amount of calculation and the number of parameters, while improving the efficiency and speed of the model; pointwise convolution uses 1×1 convolution to linearly combine the multi-channel features output by the depthwise convolution to achieve cross-channel information interaction. Pointwise convolution further reduces the number of parameters by reducing the dimension of the input channel.

[0070] In summary, the overall optimization process of the C2f_SD module includes: first, StarBlock uses the channel attention mechanism to generate weight coefficients, and adaptively weights the feature map through element-by-element multiplication to enhance the expressiveness of key features; second, DWConv uses the grouped convolution strategy to decompose the single-layer convolution into deep convolution (extracting spatial features) and point-by-point convolution (fusing channel information), thereby reducing the number of parameters while ensuring the feature extraction effect.

[0071] In some optional implementations, step S2 includes:

[0072] S21, inputting the target near-eye infrared image into the target eye feature detection sub-model to obtain the coordinates of multiple key points corresponding to the target near-eye infrared image;

[0073] In some optional implementations, the multiple key point coordinates corresponding to the target near-eye infrared image include: multiple eye face point coordinates, multiple pupil point coordinates, and multiple iris point coordinates. Preferably, there are 34 eye face point coordinates, 8 pupil point coordinates, and 16 iris point coordinates.

[0074] S22, inputting the target near-eye infrared image and the coordinates of multiple key points corresponding to the target near-eye infrared image into the target closed-eye state discrimination sub-model to obtain the closed-eye state result corresponding to the target near-eye infrared image.

[0075] In an optional embodiment, if Fig.10As shown in the figure, the target eye closure state discrimination sub-model EyeNet-GE is obtained by successively adding a Gaussian attention mechanism module (Guassian attention) and a two-dimensional spatial geometric feature extraction module (2D Euclidean feature) between the inverted residual block sequence (InvertedResidual Blocks) and the global average pooling module (AvgPool) in the MobileNetV2 network.

[0076] In some optional implementations, step S22 includes:

[0077] S221. Filter out multiple eye-face point coordinates from multiple key point coordinates corresponding to the target near-eye infrared image, and convert the multiple eye-face point coordinates into a multi-peak Gaussian distribution matrix.

[0078] Among them, the conversion of multiple eye face point coordinates into a multi-peak Gaussian distribution matrix is ​​completed based on the Gaussian attention mechanism module, and the Gaussian attention mechanism module is determined based on the Gaussian heat map.

[0079] Specifically, the Gaussian heat map is a spatial weight matrix based on the two-dimensional Gaussian distribution. Its core function is to guide the neural network to focus on the key areas in the image through the probability density distribution, thereby enhancing the local feature expression and suppressing irrelevant background interference.

[0080] For a key point in the image (c x ,c y ), whose corresponding Gaussian heat map H∈R in the feature map space of size H×W H×W It can be defined as:

[0081]

[0082] Where (x, y) represents the coordinates of any position in the heat map, and x∈[0,W-1], y∈[0,H-1]. (c x ,c y ) is the central coordinate of the key point in the feature map space, which is obtained by downsampling the original image coordinates. σ represents the standard deviation of the Gaussian distribution, controls the weight decay rate, and is usually related to the radius r of the key point coverage area.

[0083] S222. Perform feature extraction on the target near-eye infrared image to obtain multiple feature maps corresponding to the target near-eye infrared image.

[0084] Among them, after the target near-eye infrared image is input into the EyeNet-GE network, it undergoes a series of operations such as convolution and inverse residual to extract multiple feature maps.

[0085] S223. The multi-peak Gaussian distribution matrix and multiple feature maps are spatially weighted by element-by-element multiplication to obtain CNN deep semantic features.

[0086] S224, calculating the two-dimensional spatial geometric features between the coordinates of multiple eye-face points, and performing a normalization operation on the two-dimensional spatial geometric features.

[0087] The calculation of the two-dimensional spatial geometric features between the coordinates of multiple eye-face points is completed based on a two-dimensional spatial geometric feature extraction module, and the two-dimensional spatial geometric feature extraction module is determined based on Euclidean features.

[0088] Specifically, the Euclidean feature is a spatial metric representation based on the Cartesian coordinate system, and its core mathematical expression is the L2 norm calculation of the straight-line distance between two points. When judging the closed eye state, the two-dimensional spatial geometric feature calculates the absolute pixel distance between the corresponding key point pairs, takes the 34 selected eyelid key point coordinates as input, and calculates the Euclidean distance sequence between the corresponding positions of the upper eyelid (the 2nd to 17th key points) and the lower eyelid (the 18th to 33rd key points) for each sample, and constructs a model to quantify the spatial position difference of the eyelid key points. Specifically, let the key point data matrix be P∈R N×34×2 , where N is the number of samples, and the 34 key points are arranged in order, then for any i∈{2,3,...,17}, j=i+16∈{18,19,...,33}, to calculate p i With p j The generated feature vector F∈R N×16 It consists of 16 groups of cross-segment Euclidean distances, and the calculation formula for each element is:

[0089]

[0090] Among them, k=i-2=∈{0,1,...,15}.

[0091] Finally, the two-dimensional spatial geometric features between the coordinates of multiple eyelid points are normalized to obtain the normalized two-dimensional spatial geometric features.

[0092] S225, fusing the normalized two-dimensional spatial geometric features with the CNN deep semantic features and then classifying them to obtain the closed-eye state result corresponding to the target near-eye infrared image.

[0093] The embodiment of the present invention converts the coordinates of the 34 selected eyelid key points into a multi-peak Gaussian distribution matrix based on the Gaussian heat map, and spatially weights the feature map extracted by the MobileNetV2 network by element-by-element multiplication, significantly enhancing the activation response of the eyelid contour area while suppressing the background noise caused by eyelash occlusion or mirror reflection. This process not only constrains the network to learn the key point distribution prior as an auxiliary supervisory signal, but also optimizes the feature extraction direction through back propagation, so that the model can adaptively focus on the discriminative area.

[0094] In addition, the Gaussian weight decay feature gives the model strong robustness to local occlusion. Even if the key points are partially obscured, the heat map can still retain the effectiveness of gradient propagation in the central area through probability density diffusion, ensuring the continuity of feature expression. The generated Gaussian heat map visualization effect is as follows: Fig.11 shown.

[0095] In an optional implementation, step S225 includes:

[0096] The normalized two-dimensional spatial geometric features are fused with the CNN deep semantic features by splicing;

[0097] The fused features are mapped to the joint latent space and then classified and predicted to obtain the eye-closing state result corresponding to the target near-eye infrared image.

[0098] In the feature fusion stage, the embodiment of the present invention couples the two-dimensional spatial feature vector with the CNN deep semantic feature through splicing, maps it to a unified latent space through a fully connected layer, and then inputs it into the classification head. At the same time, the network is constrained to maintain the geometric consistency between the predicted key points when optimizing the cross entropy loss.

[0099] In summary, the eye feature detection and closed eye state discrimination method based on near-eye infrared images of an embodiment of the present invention, through multi-branch pooling parallel processing, element-by-element multiplication, depth-wise separable convolution and other operations, combines the fully connected layer and the activation function to map the scale features into the image coordinates of 58 eye feature key points, and uses the generated 34 eyelid key points to generate Gaussian attention heat maps and two-dimensional spatial geometric features, which are injected into a lightweight classification network to further discriminate the closed eye state.

[0100] The present invention has the following beneficial effects:

[0101] (1) The C2f_SD module uses the channel attention mechanism to generate weight coefficients and adaptively weights the feature map through element-by-element multiplication to enhance the expressiveness of key features;

[0102] (2) Based on the coordinates of 34 eyelid key points output by the lightweight YOLOv8n-ASD network, a Gaussian attention heat map is dynamically generated. Its radius is adaptively adjusted according to the eyelid opening and closing degree, and is fused with the original feature map through element-by-element multiplication to achieve dynamic calibration of features in the spatial dimension.

[0103] (3) Extract the vertical Euclidean distance sequence between the key points of the upper and lower edges of the eyelids, construct a two-dimensional spatial geometric feature, concatenate it with the deep semantic feature after normalization, and map it to the joint latent space through a fully connected layer to constrain the classification result;

[0104] (4) The lightweight deep learning network can balance the requirements of accuracy and real-time performance in actual engineering application scenarios where wearable devices have limited computing resources, providing an effective solution to meet application needs in resource-constrained environments;

[0105] (5) The lightweight YOLOv8n-ASD network adopts the method of fusing star operation and deep convolution, introduces element-by-element multiplication operation to realize dynamic calibration of features, and enhances the interactive ability of local features. In the fusion process, according to the characteristics of near-eye infrared images, the parameters, sequence, weights, etc. of star operation and deep convolution are optimized so that the two can work together and play a better effect than using them alone, thereby improving the positioning accuracy of eye feature points and making the model more robust and generalizable.

[0106] (6) The EyeNet-GE network has multi-feature guidance and constructs a cross-modal feature fusion architecture. Gaussian attention heat map and two-dimensional spatial geometric features are spliced ​​in three channels. Through the synergy of knowledge prior and deep learning features, high-precision eye closure state discrimination is achieved.

[0107] In addition, the present invention describes the training process of the eye feature detection and eye closure state discrimination model:

[0108] a: Process the near-eye infrared image dataset and perform label normalization.

[0109] b: Through multi-branch pooling parallel processing and feature fusion mechanism, efficient information retention is achieved during feature map downsampling and the number of model parameters is reduced.

[0110] c: Adaptively weight the feature maps through element-by-element multiplication to enhance the expressiveness of key features, and use deep separable convolution to optimize cross-stage feature interactions, significantly reducing the amount of computation while maintaining multi-scale feature expressiveness.

[0111] d: Use a fully connected layer and activation function to map the scale features to the image coordinates of 58 eye feature key points.

[0112] e: Extract 34 eyelid key points out of 58 key points to generate Gaussian attention heat map and two-dimensional spatial geometric features.

[0113] f: Inject Gaussian attention heatmap and two-dimensional spatial geometric features as prior knowledge into the lightweight eye classification network MobileNetV2.

[0114] g: Use the dataset to train the network and optimize the network parameters by adjusting parameters and loading pre-trained weights.

[0115] h: Perform deep learning network testing.

[0116] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for detecting eye features and distinguishing eye closure status based on near-eye infrared images, characterized in that: The method comprises: S1, obtaining a near-eye infrared image of the target; S2, inputting the target near-eye infrared image into the target eye feature detection and eye-closing state discrimination model to obtain the eye-closing state result corresponding to the target near-eye infrared image; The target eye feature detection and eye closure state discrimination model is a pre-trained neural network; the target eye feature detection and eye closure state discrimination model comprises: a target eye feature detection sub-model, and a target eye closure state discrimination sub-model cascaded with the target eye feature detection sub-model; The target eye feature detection submodel is determined after replacing the CBS module in the downsampling layer of Backbone and Neck of the YOLOv8n network with the ADown module, and the C2f module with the C2f_SD module; the target closed eye state discrimination submodel is determined after adding a Gaussian attention mechanism model and a two-dimensional space geometric feature extraction module to the MobileNetV2 network.

2. The method according to claim 1, characterized in that The S2 includes: S21, inputting the target near-eye infrared image into the target eye feature detection sub-model to obtain the coordinates of multiple key points corresponding to the target near-eye infrared image; S22, inputting the target near-eye infrared image and the coordinates of multiple key points corresponding to the target near-eye infrared image into the target closed-eye state discrimination sub-model to obtain the closed-eye state result corresponding to the target near-eye infrared image.

3. The method according to claim 1, characterized in that The C2f_SD module is determined by replacing the Bottleneck unit in the C2f module with a StarBlock unit and replacing the ordinary convolution unit with a depth convolution unit.

4. The method according to claim 3, characterized in that: The target eye feature detection sub-model includes a Backbone network, a Neck network and a Head network; The Backbone network includes a CBS module, an ADown module, a C2f_SD module and an SPPF module connected in series in sequence; wherein the ADown module and the C2f_SD module are serially stacked 4 times; The Neck network is obtained by replacing all four C2f modules of the Neck network in the YOLOv8n network with C2f_SD modules and all two CBS modules with DWCBS modules; the DWCBS module is determined by replacing the ordinary convolution unit in the CBS module with the deep convolution unit; The Head network has the same architecture as the Head network in the YOLOv8n network.

5. The method according to claim 2, characterized in that: The target eye-closing state discrimination sub-model is obtained by sequentially adding a Gaussian attention mechanism module and a two-dimensional space geometric feature extraction module between the inverted residual block sequence and the global average pooling module in the MobileNetV2 network.

6. The method according to claim 5, characterized in that The multiple key point coordinates corresponding to the target near-eye infrared image include: multiple eye face point coordinates, multiple pupil point coordinates and multiple iris point coordinates, and S22 includes: S221, selecting a plurality of eye-face point coordinates from a plurality of key point coordinates corresponding to the target near-eye infrared image, and converting the plurality of eye-face point coordinates into a multi-peak Gaussian distribution matrix; S222, extracting features from the target near-eye infrared image to obtain a plurality of feature maps corresponding to the target near-eye infrared image; S223, performing spatial weighting on the multi-peak Gaussian distribution matrix and multiple feature maps by element-by-element multiplication to obtain CNN deep semantic features; S224, calculating the two-dimensional spatial geometric features between the coordinates of multiple eyelid points, and performing a normalization operation on the two-dimensional spatial geometric features; S225. The normalized two-dimensional spatial geometric features are fused with the CNN deep semantic features and then classified to obtain the closed-eye state result corresponding to the target near-eye infrared image.

7. The method according to claim 5, characterized in that The Gaussian attention mechanism module is determined based on the Gaussian heat map; the two-dimensional space geometric feature extraction module is determined based on the Euclidean feature.

8. The method according to claim 6, characterized in that The S225 includes: The normalized two-dimensional spatial geometric features are fused with the CNN deep semantic features by splicing; The fused features are mapped to the joint latent space and then classified and predicted to obtain the eye-closing state result corresponding to the target near-eye infrared image.

Citation Information

Cited By

  • Image snow removal method based on Transform modeling and multi-scale dynamic filtering

    CN121685976A