Unmanned aerial vehicle operator fatigue state detection method, storage medium and equipment

By using YOLOv5 face detection algorithm and ERT algorithm to perform eye state detection, calculate eye aspect ratio and fatigue state coefficient, the problem of weak fatigue information extraction ability in the prior art is solved, and accurate detection and real-time alarm of the fatigue state of drone operators is achieved.

CN119942614APending Publication Date: 2025-05-06THE SECOND RES INST OF CIVIL AVIATION ADMINISTRATION OF CHINA

Patent Information

Application Number
CN202510016608.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the ability to extract important fatigue information in facial images is weak, the recognition ability and accuracy of the model are low, making it difficult to effectively detect the fatigue status of the drone operator.

Method used

A drone operator fatigue state detection method is adopted. By obtaining the video stream of face information of the person to be detected, the eye state detection is performed using YOLOv5 face detection algorithm and ERT algorithm to generate eye opening and closing state information, and the eye aspect ratio (EAR) and fatigue state coefficient PERCLOS are calculated. If PERCLOS exceeds the threshold, fatigue alarm information is generated.

Benefits of technology

By improving the calculation accuracy of eye aspect ratio and fatigue state coefficient, accurate detection and real-time alarm of fatigue state of drone operators is achieved, and the accuracy and efficiency of detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942614A_ABST
    Figure CN119942614A_ABST
Patent Text Reader

Abstract

The invention relates to the field of fatigue recognition, in particular to an unmanned aerial vehicle operator fatigue state detection method, a storage medium and equipment. Comprising the following steps: performing eye state detection on each frame of to-be-detected face image, and generating eye opening and closing state information corresponding to each frame of to-be-detected face image; generating a fatigue state coefficient of the to-be-detected person according to the eye opening and closing state information corresponding to each frame of to-be-detected face image; and if the fatigue state coefficient is greater than Y1, generating fatigue alarm information of the to-be-detected person. The fatigue degree is evaluated by using the aspect ratio of the eyes. Therefore, the fatigue state of the target person in each frame can be accurately evaluated. Besides, in order to enable the model to better identify the fatigue features, a CBAM module and a C3 module are combined in the first YOLOv5 network, the network can focus on the feature part, the ability of the model to extract important information in the target person image can be improved, and thus the identification ability and accuracy of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In the field of drone flight control, drone operators are responsible for commanding the take-off and landing of aircraft and coordinating air traffic, and their working status is directly related to aviation safety. Especially for drone operators, they need to continuously monitor the status of the drone and the surrounding environment, and make corresponding control instructions through the ground control station, which requires them to concentrate highly. Long-term concentration and human-computer interaction will cause brain fatigue. This will in turn affect work performance. Therefore, the development of effective fatigue detection technology is of great significance for preventing aviation accidents and ensuring aviation safety.

[0003] Traditional fatigue detection methods mainly rely on physiological signal monitoring, such as heart rate, brain waves, etc., but these methods have disadvantages such as strong invasiveness and poor real-time performance. In contrast, fatigue detection methods based on image recognition have the advantages of being non-invasive and easy to implement. They can monitor the facial expressions and behavioral characteristics of target personnel in real time, providing a new way to evaluate fatigue status. The research on this technology in the field of fatigue detection has evolved from traditional methods to deep learning technology. Early face detection methods mainly relied on manually extracted features and simple classifiers, such as Haar Cascades and Histogram of Oriented Gradients (HOG) features combined with support vector machines (SVM) for classification. Although these methods are effective under specific conditions, their accuracy is limited in complex environments.

[0004] With the development of deep learning technology, face detection technology has ushered in a huge change. Deep convolutional neural networks (CNNs) are widely used in face detection tasks due to their powerful feature extraction capabilities. After 2012, face detection methods based on deep learning have gradually become mainstream. These methods can automatically learn complex feature representations from images, significantly improving the accuracy of detection.

[0005] However, in the prior art, the ability to extract important fatigue information from facial images is weak, and the recognition ability and accuracy of the model are low. Summary of the invention

[0006] In view of the above technical problems, the technical solution adopted by the present invention is:

[0007] According to one aspect of the present invention, a method for detecting fatigue status of a drone operator is provided, comprising the following steps:

[0008] A facial information video stream of a person to be detected corresponding to a detection period is obtained; the facial information video stream includes multiple frames of facial images of the person to be detected.

[0009] Eye state detection is performed on each frame of the facial image to be tested, and eye opening and closing state information corresponding to each frame of the facial image to be tested is generated.

[0010] According to the eye opening and closing status information corresponding to each frame of the facial image to be tested, the fatigue state coefficient PERCLOS of the person to be tested is generated; PERCLOS meets the following conditions:

[0011]

[0012] Among them, N f 、N t They are the number of video frames with closed eyes and the total number of video frames in the facial information video stream, respectively.

[0013] If PERCLOS>Y1, a fatigue alarm message for the person to be detected is generated; Y1 is the fatigue status threshold.

[0014] Obtain the video stream of facial information of the person to be detected corresponding to the detection period, including:

[0015] Use the first face detection algorithm to obtain a video stream of facial information of a person to be detected corresponding to a detection period.

[0016] The first face detection algorithm includes: the S model and the CBAM module in the YOLOv5 face detection algorithm.

[0017] The CBAM module is added to the C3 module in the S-model Backbone.

[0018] Furthermore, the eye state detection includes:

[0019] Use the ERT algorithm to locate facial feature points on the face image to be tested, so as to generate multiple feature point coordinates P of the eye area in the face image to be tested. 37 , P 38 , P 39 , P 40 , P 41 , P 42 , P 43 , P 44 , P 45 , P 46 , P 47 and P 48 ; Among them, P 37 and P 40 are the coordinates of the outer corner of the left eye and the inner corner of the left eye respectively; P 38 and P 39 are the coordinates of the left and right uniform sampling points of the upper eyelid of the left eye; P 42 and P 41 are the coordinates of the left and right evenly sampled points of the lower eyelid of the left eye; P 43and P 46 are the coordinates of the inner and outer corners of the right eye; P 44 and P 45 are the coordinates of the left and right evenly sampled points on the upper eyelid of the right eye respectively; P 48 and P 47 are the coordinates of the left and right evenly sampled points on the lower eyelid of the right eye respectively.

[0020] Based on P 37 , P 38 , P 39 , P 40 , P 41 , P 42 , P 43 , P 44 , P 45 , P 46 , P 47 and P 48 , generate the eye opening and closing state information r corresponding to the face image to be tested EAR ; r EAR satisfies the following conditions:

[0021]

[0022] wherein, r L_EAR and r R_EAR are the opening and closing state information corresponding to the left and right eyes respectively.

[0023] If r EAR < Y2, it is determined that the eyes in the face image to be tested are in the closed state, and Y2 is the first opening and closing state threshold.

[0024] Furthermore, after generating the eye opening and closing state information r corresponding to the face image to be tested EAR , the eye state detection further includes:

[0025] If r EAR ≥ Y3, it is determined that the eyes in the face image to be tested are in the open state, Y3 is the second opening and closing state threshold, and Y3 ≥ Y2.

[0026] Furthermore, the first face detection algorithm uses a Feature Pyramid Network and a Pixel Aggregation Network for multi-scale feature fusion in the Neck part.

[0027] The first face detection algorithm also includes: a BiFPN structure.

[0028] The BiFPN structure is used to enable the Feature Pyramid Network to receive the feature information processed by the Pixel Aggregation Network while adding skip connections at the same feature scale to prevent the loss of original feature information.

[0029] At the same time, the BiFPN structure configures different feature fusion weights according to the importance of different input features.

[0030] Furthermore, the first face detection algorithm also includes: a GSConv convolution module.

[0031] Use the GSConv convolution module to replace the original ordinary convolution module in the output part of the YOLOv5 face detection algorithm.

[0032] Furthermore, obtaining a video stream of facial information of the person to be detected corresponding to the detection period includes:

[0033] An initial facial video stream of a person to be detected corresponding to a detection period is obtained; each frame of an image in the initial facial video stream includes a facial image of the person to be detected and a background image.

[0034] The YOLOv5 face detection algorithm is used to replace the first face detection algorithm, and the face image of the person to be detected in each frame image in the initial face video stream is identified to generate a video stream of face information of the person to be detected.

[0035] Furthermore, the YOLOv5 face detection algorithm is used to identify the face image of the person to be detected in each frame of the initial face video stream, including:

[0036] Use the S model in the YOLOv5 face detection algorithm to identify the face image of the person to be detected in each frame of the initial face video stream.

[0037] Furthermore, the duration of the detection period is 1 minute.

[0038] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method for detecting fatigue status of a drone operator is implemented.

[0039] According to a third aspect of the present invention, there is provided an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for detecting fatigue status of a UAV operator when executing the computer program.

[0040] The present invention has at least the following beneficial effects:

[0041] In the present invention, by using the image detection technology of face recognition, the part of the target person's face contained in the collected image can be determined, and then the ERT (Ensemble of Regresion Trees) algorithm, i.e., the regression tree set based on gradient boosting, is used to locate the facial feature points. Multiple preset points on the two eyes of the target person are mapped to the facial coordinates, and each point is represented by a linear index and configured with corresponding position coordinates. Then, the aspect ratio (Eye Aspect Ratio, EAR) of the target person's eyes can be calculated through these points. Since the evaluation of eye fatigue can be carried out by monitoring the degree of opening and closing of the eyes. The present invention uses the aspect ratio of the eyes as an indicator to evaluate the fatigue degree of the target person. Thus, the fatigue state of the target person in each frame can be accurately evaluated by calculating the EAR in each frame of the face image. At the same time, in order to distinguish from normal blinking behavior, the present application uses the percentage of eye closure time, that is, the ratio of the number of frames of the target person's eye closure state to the total number of frames PERCLOS as a fatigue measurement indicator. Thus, a non-invasive method can be used to detect the fatigue state of the target person and issue a corresponding alarm prompt.

[0042] In addition, in order to enable the model to better identify fatigue features, the channel attention mechanism and the spatial attention mechanism (convolutional block attention module, CBAM) are introduced into the YOLOv5 network. CBAM uses two modules, channel attention and spatial attention. The channel attention module mainly compresses the feature map in the spatial dimension, converts the single average pooling into average pooling and maximum pooling for processing, and is mainly responsible for focusing on the meaningful parts of the feature map. The spatial attention module is after the channel attention module and is a supplement to the channel attention. It mainly focuses on the location of the meaningful part of the feature map. Finally, the weights obtained by the above two attention mechanisms are multiplied by the input feature map for adaptive feature refinement. Therefore, combining the CBAM module with the C3 module in Backbone can make the network focus on the feature part, which can improve the ability of the first YOLOv5 model to extract important information from the target person image, thereby improving the recognition ability and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0044] Figure 1A flow chart of a method for detecting fatigue status of a drone operator provided by an embodiment of the present invention;

[0045] Figure 2 Schematic diagram of the function of the CBAM attention mechanism in an embodiment of the present invention;

[0046] Figure 3 This is a flowchart of the Bottleneck module structure before the improvement of the YOLOv5 S model in an embodiment of the present invention;

[0047] Figure 4 This is a flowchart of the Bottleneck module structure after the improvement of the YOLOv5 S model in an embodiment of the present invention;

[0048] Figure 5 It is a working structure diagram of the YOLOv5 S model FPN in an embodiment of the present invention;

[0049] Figure 6 It is a working structure diagram of the YOLOv5 S model PAN in an embodiment of the present invention;

[0050] Figure 7 This is a working structure diagram of the YOLOv5 S model after BiFPN is introduced in an embodiment of the present invention;

[0051] Figure 8 This is a network model structure diagram of the first face detection algorithm in an embodiment of the present invention;

[0052] Fig. 9 A visualization result diagram of the bounding box loss or positioning loss (also called box_loss) during the training process of the network model corresponding to the first face detection algorithm in an embodiment of the present invention;

[0053] Fig.10 This is a visualization result diagram of the classification loss (i.e., cls_loss) during the training process of the network model corresponding to the first face detection algorithm in an embodiment of the present invention;

[0054] Fig.11 This is a visualization result diagram of the model prediction accuracy (i.e., Precision) during the training process of the network model corresponding to the first face detection algorithm in the embodiment of the present invention;

[0055] Fig.12 This is a visualization result diagram of the model recall rate (ie, Recall) during the training process of the network model corresponding to the first face detection algorithm in the embodiment of the present invention;

[0056] Fig.13 This is a visualization result diagram of the mean mAP@0.5 average accuracy rate during the training process of the network model corresponding to the first face detection algorithm in an embodiment of the present invention;

[0057] Fig.14 This is a visualization result diagram of the mean mAP@0.5:0.95 average accuracy rate during the training process of the network model corresponding to the first face detection algorithm in an embodiment of the present invention;

[0058] Fig.15 4 is a distribution diagram of key points of the eye area in an embodiment of the present invention. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0060] As a possible embodiment of the present invention, Figure 1 As shown, a method for detecting fatigue status of a drone operator is provided, the method comprising the following steps:

[0061] S100: Obtain a video stream of facial information of a person to be detected corresponding to a detection period. The video stream of facial information includes multiple frames of facial images of the person to be detected. Specifically, the duration of the detection period may be 1 minute.

[0062] S100 includes:

[0063] S101: Acquire an initial facial video stream of a person to be detected corresponding to a detection period. Each frame of the initial facial video stream includes a facial image of the person to be detected and a background image.

[0064] S102: Using the YOLOv5 face detection algorithm, identify the face image of the person to be detected in each frame of the initial face video stream to generate a video stream of face information of the person to be detected.

[0065] Specifically, in actual use, a camera can be placed at one or several locations in the working area of ​​the drone operator to capture the frontal image of the drone operator. Since the camera has a larger shooting range, there will inevitably be a large range of background images in the captured image. In order to use the ERT algorithm to locate the key points of the eyes more accurately and quickly, it is necessary to remove the background noise of each frame of the captured video stream. Therefore, in this embodiment, the YOLOv5 face detection algorithm is used to identify the area corresponding to the face.

[0066] Preferably, the S model in the YOLOv5 face detection algorithm is used to identify the facial image of the person to be detected in each frame image in the initial facial video stream.

[0067] Specifically, the YOLOv5 face detection algorithm provides four models of different sizes: S, M, L, and X. Among them, the S model has a small file size (24 MiB), a simple structure, and a fast calculation speed. It is suitable for the real-time requirements of the scene in this embodiment, so the S model is selected. The model consists of an input end, a backbone network (Backbone), a neck (Neck), and an output end (Output). The input end is responsible for image preprocessing, using Mosaic data enhancement, adaptive anchor frame calculation and size adjustment, and the processed image size is 640×640 pixels. Backbone includes Focus and C3 structures. Focus enhances data through segmentation and splicing, and the C3 structure learns residual features through convolution and bottleneck layers. The Neck part combines the feature pyramid network (FPN) and the path aggregation network (PAN). The former transmits high-dimensional semantic information, and the latter transmits low-dimensional information, which jointly improves detection accuracy. The output end includes the Bounding box loss function and non-maximum suppression (NMS). The YOLOv5 face detection algorithm specifically uses GIOU_Loss to optimize the bounding box relationship and improve the detection effect.

[0068] In addition, the YOLOv5 face detection algorithm may be replaced with the first face detection algorithm. The first face detection algorithm has a faster recognition speed and higher recognition accuracy than the existing YOLOv5 face detection algorithm, and can better meet the real-time requirements in the scenario of this embodiment.

[0069] Specifically, the first face detection algorithm includes: the S model, CBAM module, BiFPN structure and GSConv convolution module in the YOLOv5 face detection algorithm. At the same time, the first face detection algorithm uses a feature pyramid network and a pixel aggregation network in the Neck part for multi-scale feature fusion.

[0070] The existing S model is improved using the CBAM module, BiFPN structure and GSConv convolution module in the following ways to further improve the recognition accuracy and efficiency of the model.

[0071] First: The CBAM module is added to the C3 module in the S model Backbone.

[0072] In some flight control scenarios where attention needs to be paid to the actual flight status of the drone, such as drone competitions or aerobatic performances or complex terrain flight scenarios, in order to improve the field of vision of the flight control personnel, most areas of the drone operator's operating room are installed with glass. In the images collected during the day, the background area is also bright white, which is close to the color of the human face. As a result, the human face and the image background pixels are similar, and the background accounts for a large proportion. In addition, for some tasks that require high concentration, such as monitoring of drone light shows, it is usually necessary to avoid direct light exposure to the display or operation interface to prevent glare and reduce eye fatigue of the flight control personnel, so the light in the flight control personnel's operating room will not be too strong. The operating room may even use a darker environment with appropriate local lighting to facilitate observation of details on the screen. Therefore, in this state, due to the dim light, it is easy to make the face image in the collected image not obvious, and it is difficult to distinguish the boundary with the dim background. Based on the limitations of the above scenes, it is more difficult to extract important feature information from the facial image of the flight control personnel in the acquired image.

[0073] Based on this, in order to make the model better identify fatigue characteristics, the channel attention mechanism and the spatial attention mechanism CBAM are introduced into the YOLOv5 network of the S model. CBAM uses two modules: channel attention and spatial attention. Its structure is shown in the figure below. Figure 2 As shown in the figure. The channel attention module mainly compresses the feature map in the spatial dimension and converts the single average pooling into average pooling and maximum pooling. Its specific calculation is shown in formula (1). It is mainly responsible for focusing on the meaningful parts of the feature map. The spatial attention module is after the channel attention module and is a supplement to the channel attention. It mainly focuses on the location of the meaningful parts of the feature map. Its specific calculation is shown in formula (2). Finally, the weights obtained by the above two attention mechanisms are multiplied by the input feature map to perform adaptive feature refinement.

[0074]

[0075] Where: F represents the input feature map; AvgPool represents the average pooling operation; MaxPool represents the maximum pooling operation; MLP represents the multi-layer perceptron. After the input feature map undergoes the average pooling and maximum pooling operations, a weighted feature map W0 (F avg c )、W0(F max c ), and then processed by the multi-layer perceptron, we get W1(W0(F avg c ))、W1(W0(F max c)) Two weighted eigenvalues; σ represents the Sigmoid function; W0∈RC / r, W1∈RC×C / r. The weights of MPL are shared by W0 and W1, and there is a RELU activation function before W0.

[0076]

[0077] Where f 7×7 Represents a convolution kernel of size 7×7.

[0078] Combine the CBAM module with the C3 module in Backbone, such as Figure 4 As shown, the purpose is to improve the Bottleneck module in the C3 module. The original Bottleneck module is as follows Figure 3 As shown in , the input features are processed by two convolutions and combined with shortcut connections to generate output. The network structure of the improved Bottleneck module is as follows: Figure 4 As shown in the figure, the two constitute the C3CBAM module. After the introduction of the CBAM module, the network focuses on the feature part, which can improve the ability of the first face detection algorithm to extract important information from the target person (that is, the flight control personnel) image, thereby improving the recognition ability and accuracy of the model.

[0079] Second: The BiFPN structure is combined with the Concat structure in YOLOv5 to form the BiFPN_Concat module.

[0080] Specifically, the first face detection algorithm also includes: a BiFPN structure;

[0081] The BiFPN structure is used to enable the feature pyramid network to add jump connections of the same feature scale while receiving feature information processed by the pixel aggregation network to prevent the loss of original feature information; at the same time, the BiFPN structure configures different feature fusion weights according to the importance of different input features.

[0082] like Figure 5 and Figure 6 As shown in the figure, the existing YOLOv5 uses feature pyramid network (FPN) and pixel aggregation network (PAN) in the Neck part for multi-scale feature fusion. The combination of the two allows different backbone layers to fuse parameters of different detection layers in different ways. However, since the information processed by the PAN structure is feature information processed by the FPN structure, the lack of processing of the original feature information can easily lead to deviations in network training and learning, thereby affecting the detection results. Figures 5 to 7 P3 to P7 in the figure represent the 3rd to 7th characteristic layers respectively.

[0083] like Figure 7As shown, the BiFPN structure in this embodiment is improved on the basis of the PAN structure. The BiFPN structure enables the PAN structure to add jump connections to the original input features of the same feature scale while accepting the feature information processed by the FPN structure, thereby preventing the loss of the original feature information. At the same time, BiFPN will also set different weights according to the importance of different input features when fusing parameters of different detection layers.

[0084] by Figure 7 Take the P6 layer feature fusion in as an example to illustrate:

[0085]

[0086] Among them, P td 6 is the intermediate feature in the P6 layer, P in 6 and P out 6 are the input features and output features in the P6 layer; Resize(P in 7) is the input feature of the P7 layer after adjusting the image size; Resize(P out 5) is the output feature of the P5 layer after adjusting the image size; w i is the learnable weight of the corresponding feature; w i , is the learned weight of the corresponding feature. The values ​​of i are 1, 2 and 3.

[0087] Both PAN and FPN can use Concatenation (also known as Concat structure) to achieve effective fusion of feature maps. Concatenation is a fusion splicing operation that splices two or more feature maps together along the channel dimension to form a new feature map with more channels. The existing Concat structure in YOLOv5 simply superimposes the input feature maps without distinguishing the feature maps added at the same time. However, since the input feature maps have different resolutions, and different resolutions contribute differently to the fused input feature maps, simple feature map superposition processing is not the best feature fusion operation. By introducing the BiFPN structure into the Neck part and combining it with the Concat structure of the neck to form the BiFPN_Concat module, it can retain the original features and adjust the importance of different input features by adding jump connections and weighted feature fusion. Feature fusion is optimized, thereby improving the recognition rate of target person fatigue.

[0088] Third: Use the GSConv convolution module to replace the original ordinary convolution module in the output part of the existing YOLOv5 face detection algorithm.

[0089] In order to make the YOLOv5 network of the S model in this embodiment more lightweight while maintaining detection accuracy, the GSConv convolution is introduced to replace the original ordinary convolution in the output part of the Neck part. GSConv combines standard convolution (SC), depthwise separable convolution (DSC) and Shuffle hybrid convolution, enhances feature extraction capabilities through SC, improves speed through DSC, and realizes efficient interaction of channel information through ShuffleNet. This design reduces the number of parameters and computational complexity of the model without sacrificing performance. Although GSConv may increase the number of network layers and inference time, by applying GSConv only in the Neck part and combining the CBAM attention mechanism with the BiFPN structure, the accuracy of the model can be maintained while improving the detection speed. The specific network model structure of the first face detection algorithm after the above three improvements is as follows. Figure 8 shown. Figure 8 The Spp in it is spatial pyramid pooling, which is a necessary operation step like C3 and CONV, and is also an existing step of the existing YOLOv5 algorithm.

[0090] Specifically, Figures 9 to 14 As shown in FIG. 1 , the visualization results of various evaluation indicators in the network model training process corresponding to the first face detection algorithm in this embodiment are shown. In this embodiment, a total of 200 test trainings are performed. Fig. 9 The training results shown in the figure show that the vertical axis is the mean of the GIoU loss function. The smaller the loss value, the more accurate the face box detection. Fig. 9 From the content, we can see that the loss value of this experiment tends to 1.2 after 100 rounds of training, which shows that the detection box of this experiment can accurately detect the face. Fig.10 In the figure, the vertical axis is the mean classification loss. The smaller the value, the more accurate the classification. The classification loss value in the experiment tends to 0 after 100 rounds of model training, which shows that the model can correctly classify the fatigue state.

[0091] In addition, the first face detection algorithm in this embodiment and the existing YOLOv5 S face detection algorithm are used to perform face recognition on the same image set, and the accuracy P (Precision), recall R (Recall) and mean average accuracy (mAP) are used to evaluate the detection results of the model. Specifically, the accuracy is the proportion of correct predictions among the targets predicted by the model; the recall rate is the proportion of correct targets predicted by the model among all real targets; the mean average accuracy is divided into mAP@0.5 and mAP@0.5:0.95. In this embodiment, the mAP@0.5 partial evaluation index is mainly used. In addition, the accuracy P, recall R and mean average accuracy mAP meet the following conditions respectively:

[0092]

[0093]

[0094] Where: I TP is the number of positive samples correctly identified as positive classes; L FP is the number of negative samples mistakenly identified as positive; T FN is the number of positive samples mistakenly identified as negative classes; C is the total number of sample categories in the target detection dataset; AP i is the average precision of the i-th classification category.

[0095] Finally, the evaluation results of the first face detection algorithm and the existing YOLOv5 S face detection algorithm are shown in Table 1 below:

[0096] Table 1

[0097]

[0098] From the results in Table 1 above, we can see that the accuracy P, mAP value, and recall rate R of the first face detection algorithm are all improved compared to the YOLOv5S face detection algorithm, increasing by 1.2%, 0.1%, and 0.5% respectively, with significant improvement effects. In summary, the improved algorithm is effective in improving fatigue state recognition.

[0099] In addition, the present invention also uses the same training set to train the existing YOLOv5 S face detection algorithm and the first face detection algorithm, respectively, to obtain a model that meets the scenario used in this embodiment. Then configure the same hardware model operating environment, run the trained YOLOv5 S face detection algorithm model and the first face detection algorithm model, and obtain the GFLOPS and calculation parameter amount of the two models respectively. GFLOPS is used to indicate the number of floating-point operations that the model can complete per unit time. The higher the GFLOPS value, the faster the calculation speed of the model itself; the calculation parameter amount is used to indicate the amount of data that the model needs to calculate and consider. The fewer the parameters, the higher the calculation efficiency of the model.

[0100] Finally, the results of the two indicators of the first face detection algorithm and the existing YOLOv5 S face detection algorithm are shown in Table 2 below:

[0101] Table 2

[0102]

[0103] From the results in Table 2 above, it can be seen that the first face detection algorithm has a smaller parameter amount and a higher calculation speed than the YOLOv5 S face detection algorithm, which shows that the model corresponding to the first face detection algorithm has higher calculation efficiency.

[0104] S200: Perform eye state detection on each frame of the facial image to be tested, and generate eye opening and closing state information corresponding to each frame of the facial image to be tested.

[0105] Eye state detection includes:

[0106] S201: Use the ERT algorithm to locate facial feature points on the face image to be tested, so as to generate coordinates P of multiple feature points of the eyes in the face image to be tested. 37 , P 38 , P 39 , P 40 , P 41 , P 42 , P 43 , P 44 , P 45 , P 46 , P 47 and P 48 Among them, P 37 and P 40 are the coordinates of the outer corner of the left eye and the inner corner of the left eye respectively; P 38 and P 39 are the coordinates of the left and right uniform sampling points of the upper eyelid of the left eye; P 42 and P 41 are the coordinates of the left and right evenly sampled points of the lower eyelid of the left eye; P 43 and P 46 are the coordinates of the inner corner and outer corner of the right eye; P 44 and P 45 are the coordinates of the left and right uniform sampling points of the upper eyelid of the right eye; P 48 and P 47 They are the coordinates of the left and right uniform sampling points of the lower eyelid of the right eye.

[0107] Through the ERT algorithm, the 68 points can be mapped to facial coordinates, each point is represented by a linear index, and each part can be represented by a fixed value range. The 68 key points mentioned in this embodiment generally refer to a standard marking scheme for locating facial feature points. This marking scheme is widely used in the fields of computer vision and image processing, especially in facial recognition and analysis. These key points cover the main structural features of the face, including eyebrows, eyes, nose, mouth, and facial contours. In this embodiment, only the key points of the eyes are used. The distribution of these points is as follows: Fig.15 shown.

[0108] S202: According to P 37 , P 38 , P 39 , P 40 , P41 , P 42 , P 43 , P 44 , P 45 , P 46 , P 47 and P 48 , generate the eye opening and closing state information r corresponding to the face image to be measured EAR . r EAR satisfies the following conditions:

[0109]

[0110] Among them, r L_EAR and r R_EAR are the opening and closing state information corresponding to the left eye and the right eye respectively.

[0111] S203: If r EAR <Y2, it is determined that the eyes in the face image to be measured are in the closed state, and Y2 is the first opening and closing state threshold.

[0112] S204: If r EAR ≥Y3, it is determined that the eyes in the face image to be measured are in the open state, Y3 is the second opening and closing state threshold, and Y3≥Y2.

[0113] In this embodiment, the evaluation of eye fatigue can be carried out by monitoring the degree of eye opening and closing. Specifically, the eye aspect ratio (EAR) is used as an index to evaluate the fatigue degree of the target person. As Fig.15 shown, the calculation of EAR is realized by marking 12 key feature points around the human eye. Y3 and Y2 can be adaptively determined according to the specific situation of the被测人员. For example, when the eye closing degree exceeds 50%, 70%, and 80% respectively, it is determined that the human eye is in the closed state.

[0114] S300: According to the eye opening and closing state information corresponding to each frame of the face image to be measured, generate the fatigue state coefficient PERCLOS of the person to be detected. PERCLOS satisfies the following conditions:

[0115]

[0116] Among them, N f , N t are the number of video frames with closed eyes and the total number of video frames in the face information video stream respectively.

[0117] S400: If PERCLOS>Y1, generate the fatigue warning information of the person to be detected. Y1 is the fatigue state threshold. Y1 can be determined according to the specific usage scenario. For example, Y1 is 50%.

[0118] For the specific determination method of S300 and S400 disclosed in this embodiment, it is more applicable to the determination situation where Y3 = Y2. That is, when judging the opening and closing state of the human eye according to EAR, there will only be two states, namely the open eye state or the closed eye state. Then, based on these two result states, the fatigue state coefficient PERCLOS is calculated to perform subsequent fatigue judgment.

[0119] In addition, in actual use, the fatigue of personnel actually has different degrees, and the corresponding opening and closing degrees on the eyes are also different. Therefore, there will be multiple opening and closing states of the eyes, such as the eyes being open, the eyes being closed, and the eyes being in a semi-open and semi-closed state. In order to be able to more precisely distinguish the opening and closing state of the eyes, at this time, it is necessary to make Y3 > Y2. Thus, when the eyes are open, it corresponds to r EAR < the situation of Y2, when the eyes are closed, it corresponds to r EAR ≥ Y3, and when the eyes are in a semi-open and semi-closed state, it corresponds to r EAR belonging to the situation of [Y2, Y3). After the refinement, it is also possible to correspondingly calculate the proportion of the number of images in the situation of [Y2, Y3) in the total number of images in S300 to generate the second fatigue state coefficient. And in S400, add the judgment of the second fatigue state coefficient and the corresponding threshold. Usually, the fatigue state of personnel is a gradual process. Therefore, by adding the judgment of the second fatigue state coefficient, the appearance of the personnel's fatigue state can be detected earlier for early warning. At the same time, the two dimensions of the fatigue state coefficient PERCLOS and the second fatigue state coefficient can also be combined to set corresponding judgment conditions to improve the accuracy of fatigue state judgment.

[0120] In addition, although the steps of the methods in this disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0121] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0122] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0123] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as a system, method or program product. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as a "circuit", "module" or "system".

[0124] The electronic device according to this embodiment of the present invention is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0125] The electronic device is presented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: the at least one processor mentioned above, the at least one storage device mentioned above, and a bus connecting different system components (including storage devices and processors).

[0126] The storage stores program codes, which can be executed by the processor, so that the processor executes the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification.

[0127] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read only memory (ROM).

[0128] The storage may also include a program / utility having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0129] The bus may represent one or more of several types of bus structures including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures.

[0130] The electronic device may also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may communicate with one or more devices that enable a user to interact with the electronic device, and / or may communicate with any device (e.g., routers, modems, etc.) that enables the electronic device to communicate with one or more other computing devices. Such communication may be performed through an input / output (I / O) interface. Furthermore, the electronic device may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through a network adapter. The network adapter communicates with other modules of the electronic device through a bus. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0131] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.

[0132] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of the present specification is stored. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of the present specification.

[0133] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0134] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0135] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0136] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0137] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to an exemplary embodiment of the present invention, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0138] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.

[0139] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A method for detecting fatigue status of a drone operator, characterized in that: The method comprises the following steps: Acquire a video stream of facial information of a person to be detected corresponding to a detection period; the video stream of facial information includes a plurality of frames of facial images of the person to be detected; Perform eye state detection on each frame of the facial image to be tested, and generate eye opening and closing state information corresponding to each frame of the facial image to be tested; According to the eye opening and closing state information corresponding to each frame of the facial image to be tested, the fatigue state coefficient PERCLOS of the person to be tested is generated; PERCLOS satisfies the following conditions: Among them, N f 、N t are the number of video frames with closed eyes and the total number of video frames in the facial information video stream, respectively; If PERCLOS>Y1, then generate fatigue warning information of the person to be detected; Y1 is the fatigue state threshold; Obtain the video stream of facial information of the person to be detected corresponding to the detection period, including: Using a first face detection algorithm, obtaining a video stream of facial information of a person to be detected corresponding to a detection period; The first face detection algorithm includes: the S model and CBAM module in the YOLOv5 face detection algorithm; The CBAM module is added to the C3 module in the S-model Backbone.

2. The method for detecting fatigue status of a drone operator according to claim 1, characterized in that: The eye state detection includes: The ERT algorithm is used to locate the facial feature points of the face image to be tested, so as to generate the coordinates P of multiple feature points of the eye part in the face image to be tested. 37 , P 38 , P 39 , P 40 , P 41 , P 42 , P 43 , P 44 , P 45 , P 46 , P 47 and P 48 ; Among them, P 37 and P 40 are the coordinates of the outer corner of the left eye and the inner corner of the left eye respectively; P 38 and P 39 are the coordinates of the left and right uniform sampling points of the upper eyelid of the left eye; P 42 and P 41 are the coordinates of the left and right evenly sampled points of the lower eyelid of the left eye; P 43 and P 46 are the coordinates of the inner corner and outer corner of the right eye; P 44 and P 45 are the coordinates of the left and right uniform sampling points of the upper eyelid of the right eye; P 48 and P 47 are the coordinates of the left and right evenly sampled points of the lower eyelid of the right eye; According to P 37 , P 38 , P 39 , P 40 , P 41 , P 42 , P 43 , P 44 , P 45 , P 46 , P 47 and P 48 , generate the eye opening and closing status information r corresponding to the face image to be tested EAR ; r EAR The following conditions must be met: Among them, r L_EAR and r R_EAR They are the opening and closing status information corresponding to the left eye and the right eye respectively; If r EAR < is less than Y2, it is determined that the eyes in the face image to be measured are in a closed state, where Y2 is the first opening / closing state threshold.

3. The method for detecting fatigue status of a drone operator according to claim 2, characterized in that: When generating the eye opening and closing state information r corresponding to the face image to be tested EAR Afterwards, the eye state detection further includes: If r EAR ≥Y3, it is determined that the eyes in the face image to be tested are in the open state, Y3 is the second open and closed state threshold, and Y3≥Y2.

4. The method for detecting fatigue status of a drone operator according to claim 1, characterized in that: The first face detection algorithm uses a feature pyramid network and a pixel aggregation network in the Neck part to perform multi-scale feature fusion; The first face detection algorithm also includes: a BiFPN structure; The BiFPN structure is used to enable the feature pyramid network to add skip connections of the same feature scale while receiving feature information processed by the pixel aggregation network to prevent the loss of original feature information; At the same time, the BiFPN structure configures different feature fusion weights according to the importance of different input features.

5. The method for detecting fatigue status of a drone operator according to claim 4, characterized in that: The first face detection algorithm also includes: a GSConv convolution module; Use the GSConv convolution module to replace the original ordinary convolution module in the output part of the YOLOv5 face detection algorithm.

6. The method for detecting fatigue status of a drone operator according to claim 1, characterized in that: Obtain the video stream of facial information of the person to be detected corresponding to the detection period, including: Acquire an initial facial video stream of a person to be detected corresponding to a detection period; each frame of the initial facial video stream includes a facial image of the person to be detected and a background image; The first face detection algorithm is replaced by the YOLOv5 face detection algorithm to identify the face image of the person to be detected in each frame image in the initial face video stream, so as to generate a face information video stream of the person to be detected.

7. The method for detecting fatigue status of a drone operator according to claim 6, characterized in that: Using the YOLOv5 face detection algorithm, identifying the face image of the person to be detected in each frame of the initial face video stream, including: The S model in the YOLOv5 face detection algorithm is used to identify the face image of the person to be detected in each frame of the initial face video stream.

8. The method for detecting fatigue status of a drone operator according to claim 1, characterized in that: The duration of the detection period is 1 minute.

9. A non-transitory computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for detecting fatigue status of a drone operator as described in any one of claims 1 to 8 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for detecting fatigue status of a drone operator as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Fall behavior detection method and system based on improved YOLOv5

    CN114898470A

  • Steel rope defect detection method based on improved YOLOv5s model

    CN117094953A

  • Driver face detection method based on lightweight improved YOLOv8 model and medium

    CN117935333A

Cited By

  • Pilot fatigue detection method and device

    CN120877257A

  • Method and apparatus for detecting pilot fatigue

    CN120877257B