A face detection method and its detection system in a classroom environment based on the YOLO deep network

By introducing smaller pooling cores, hybrid attention modules and adaptive spatial feature fusion modules into YOLO deep network, combined with EIOU loss function and transfer learning pre-training operations, the problem of face detection in classroom environments is solved, and efficient and accurate face detection is achieved.

CN115240259BActive Publication Date: 2025-06-27XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210894051.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-06-27
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

In the classroom environment, face detection tasks face difficulties in small-scale face detection, severe occlusion, atypical posture, blurred faces, large background proportion, large changes in front and back face sizes, and insufficient data sets, resulting in insufficient detection accuracy and speed of the existing technology.

Method used

The face detection method in a classroom environment based on YOLO deep network is adopted. By using a smaller pooling core in the spatial pyramid pooling structure of the network, a hybrid attention module and an adaptive spatial feature fusion module are added, and the EIOU loss function and transfer learning pre-training operations are used to improve detection accuracy and robustness.

Benefits of technology

This method can more easily detect small-scale faces in classroom environments, improve overall face detection performance, enhance model robustness and detection accuracy, while reducing calculation costs and detection time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240259B_ABST
    Figure CN115240259B_ABST
Patent Text Reader

Abstract

The present invention discloses a face detection method and its detection system in a classroom environment based on the YOLO deep network. It is improved on the original YOLOX algorithm. Smaller pooling kernels are used in the spatial pyramid pooling structure of the network, which can help the model more easily detect small-scale faces in the classroom environment and improve the overall face detection performance. A hybrid attention module is added to the network to enable the model to learn to suppress useless background information and improve the detection accuracy. An adaptive spatial feature fusion operation is added to the network to solve the inconsistency problem existing in the PAFPN structure. The EIOU loss function is used to replace the IOU loss function to minimize the width difference and height difference between the ground truth box and the predicted box, which can accelerate the convergence speed. Transfer learning pre-training operation is used to solve the problem of insufficient data and improve the accuracy of face detection by the model in the classroom environment. A module is divided to divide the face detection data set collected in the classroom environment into a training set, a validation set, and a test set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning detection, and particularly relates to a face detection method and a detection system based on a YOLO deep network in a classroom environment. Background Art

[0002] The classroom is one of the application scenarios of face detection technology. In the traditional teaching environment, teachers can only count the attendance of students by taking roll calls or signing in class. However, when there are many students, this method is very time-consuming. Introducing face detection technology into the classroom can conduct real-time detection and analysis of students' classroom attendance rate, concentration, etc., helping teachers understand the class attendance situation and the class status and learning conditions of each student, and then making corresponding adjustments in teaching methods and strategies to improve teaching quality.

[0003] In the face detection task in a classroom environment, there are the following difficulties:

[0004] 1. The face scales in the classroom environment are generally small, and the distance between faces is small, making it difficult to distinguish.

[0005] 2. The poses of students in the classroom environment cannot be predicted, and there are many serious occlusions, atypical poses, and blurred faces.

[0006] 3. The background in the classroom environment accounts for a large proportion, which will affect face detection.

[0007] 4. The scale change of the front and back faces in the classroom environment is large, which requires a high demand for the detection algorithm.

[0008] 5. The available face detection datasets in the classroom environment are few and the cost of making datasets is high, which is not enough to support the training of large and complex networks.

[0009] The existing technical solutions include face detection methods based on traditional handcrafted features and face detection methods based on deep learning.

[0010] Before the introduction of deep learning methods in the field of face detection, face detection work was mainly based on classical methods, that is, extracting manual features from images (or sliding windows on images), and then inputting the features into classifiers (or classifier sets) to detect possible face areas. The performance of these detectors depends largely on the computational efficiency and expressiveness of the features. With the continuous improvement and exploration of researchers, face detection methods based on traditional manual features have achieved good detection results. However, the features designed manually based on experience have great limitations and are easily disturbed by environmental factors (such as blur, occlusion, brightness, etc.). Therefore, the application scenarios of face detection methods based on traditional manual features are limited and the robustness in complex environments is not good. At the same time, traditional face detection algorithms cannot automatically extract features useful for detection tasks from raw images without human intervention, and due to performance limitations, traditional methods cannot process large amounts of data.

[0011] With the breakthrough work of deep neural networks in image classification in 2012, the mode of face detection has also undergone a huge change. Inspired by the rapid development of deep learning in computer vision, in the past few years, many deep learning-based frameworks have been applied to the field of face detection, and have achieved significant improvements in detection accuracy. Due to the advantages of high detection efficiency and strong stability, various face detection algorithm models based on deep learning have become the mainstream framework for face detection tasks. As researchers explore new technologies and new networks, more and more excellent deep learning-based face detection networks have been proposed, such as YOLOX, YOLO-face and YOLO5Face, etc. These algorithms have achieved very advanced results on various face detection data benchmarks. However, for a long time in the past, face detection algorithms only pursued the improvement of detection accuracy and ignored the size of the model and the detection speed of the algorithm. Even many algorithms increased the network model and reduced the detection speed in exchange for the improvement of detection progress. These algorithms are very difficult to train and have high requirements on hardware equipment. They have high computational costs and long detection time, making it difficult to apply them in practice. Summary of the invention

[0012] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to propose a face detection method and its detection system in a classroom environment based on the YOLO deep network. Using smaller pooling kernels in the spatial pyramid pooling structure of the network can help the model more easily detect small-scale faces in the classroom environment and improve the overall face detection performance; adding a hybrid attention module to the network allows the model to learn to suppress useless background information and improve the detection accuracy; adding an adaptive spatial feature fusion operation to the network solves the inconsistency problem existing in the PAFPN structure; using the EIOU loss function instead of the IOU loss function minimizes the width difference and height difference between the ground truth box and the predicted box, which can accelerate the convergence speed; using transfer learning pre-training operations to solve the problem of insufficient data and improve the accuracy of face detection in the classroom environment by the model.

[0013] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0014] A face detection method in a classroom environment based on the YOLO deep network, comprising the following steps:

[0015] S1. Divide the face detection dataset in the classroom environment into a training set, a validation set, and a test set;

[0016] S2. Read the images in the training set and validation set divided in step S1, convert them to the RGB format and adjust the size of the images, and then perform data augmentation on the training set divided in step S1;

[0017] S3. Construct a face detection convolutional neural network in the classroom environment based on the YOLOX deep network, and name it YOLOXs-face;

[0018] S4. Use the EIOU loss function and the cross-entropy loss function to construct the loss function of this method;

[0019] S5. Use the pre-training dataset to train the YOLOXs-face network to obtain a pre-trained model;

[0020] S6. Use the training set processed in step S2 to continue training the YOLOXs-face network based on the pre-trained model obtained in step S5, use the validation set processed in step S2 for validation, and save the network model with the best performance on the validation set;

[0021] S7. Use the test set divided in step S1 to test on the network model saved in step S6 to obtain the face detection result in the classroom environment;

[0022] S8. For the detection result obtained in step S7, use the F1 score and the average precision to quantitatively evaluate the detection performance of the network model.

[0023] Specifically, in step S1, the samples in the face detection dataset in the classroom environment are randomly divided into a training set, a validation set, and a test set according to the ratio of 11:4:5.

[0024] Specifically, step S2 is as follows:

[0025] S201. Preprocess the images in the validation set divided in step S1. First, convert the images to the RGB format, then use the bilinear interpolation method to scale the images in the validation set and the test set proportionally, and finally unify the image sizes by adding gray bars to the images;

[0026] S202. Preprocess the images in the training set divided in step S1. First, convert the images to the RGB format, then scale the images proportionally, and then randomly scale the aspect ratio of the images; unify the image sizes by adding gray bars to the images, and horizontally flip the images according to the probability. Finally, randomly change the hue, saturation, and brightness of the images to achieve data augmentation;

[0027] S203. Adjust the ground truth boxes for the preprocessed validation set in step S201 and the preprocessed training set in step S202 respectively.

[0028] Specifically, step S3 is as follows:

[0029] S301. Build a face detection network in the classroom environment based on the YOLO deep network, named YOLOXs-face; the YOLOXs-face network includes a feature extraction module, a feature enhancement module, and a feature point prediction module.

[0030] S302. Build a CBS module that includes a convolutional layer, a batch normalization layer, and a SiLU non-linear activation layer;

[0031] S303. Build a residual module that includes a convolutional layer, a batch normalization layer, and a SiLU non-linear activation layer;

[0032] S304. Build a Focus module based on the CBS module in step S302. This module first slices the input image, expands the input from three channels to twelve channels, and then uses a CBS module to perform a convolutional operation on the feature layer;

[0033] S305. Build an SPP module based on the CBS module in step S302. This module consists of a CBS module and a max pooling operation;

[0034] S306. Construct the CSP_N module and the CSP2_N module based on the CBS module in step S302 and the residual module in step S303. The CSP_N module includes a backbone branch and a residual branch. Its backbone branch contains one CBS module and N residual modules, and its residual branch contains one CBS module. Input the data into the backbone branch and the residual branch respectively to obtain feature layers of the same size. Stack the feature layers and then pass them through the CBS module to get the output. The CSP2_N module includes a backbone branch and a residual branch. Its backbone branch contains one CBS module and N residual modules without residual edges, and its residual branch contains one CBS module. Input the data into the backbone branch and the residual branch respectively to obtain feature layers of the same size. Stack the feature layers and then pass them through the CBS module to get the output.

[0035] S307. Construct the feature extraction module CSPDarkNet network of the face detection network YOLOXs-face in step S301 based on the CBS module in step S302, the Focus module in step S304, the SPP module in step S305, and the CSP_N module and the CSP2_N module in step S306. This structure will perform feature extraction operations on the input data. Input the data in the training set after data augmentation in step S2 into the CSPDarkNet network, and obtain three effective feature layers in the middle layer, the middle-lower layer, and the bottom layer of the CSPDarkNet structure.

[0036] S308. Construct the feature enhancement module Attention network of the face detection network YOLOXs-face in step S301. This network consists of three CBAM attention modules. Input the three effective feature layers obtained in step S307 into the three CBAM attention modules respectively to obtain three mixed attention feature layers.

[0037] S309. Construct the feature enhancement module PAFPN network of the face detection network YOLOXs-face in step S301 based on the CBS module in step S302 and the CSP2_N module in step S306. This network consists of the FPN and PAN networks. Input the three mixed attention feature layers obtained in step S308 into the PAFPN network. First, perform feature transfer and fusion in the FPN network through upsampling, and then obtain three enhanced feature layers through downsampling fusion in the FAN network.

[0038] S310. Construct the feature enhancement module ASFF network of the face detection network YOLOXs-face in step S301. This network consists of three adaptive spatial feature fusion modules. Input the three enhanced feature layers obtained in step S309 into the ASFF network to adaptively fuse different feature layers and obtain three fused feature layers.

[0039] S311. Based on the CBS module in step S302, construct the feature point prediction Yolo Head network of the face detection network YOLOXs-face in step S301. This network is composed of Yolo Head modules; input the three fused feature layers obtained in step S310 into the Yolo Head network, and perform classification and regression operations on the feature layers to obtain prediction results at three different scales;

[0040] S312. Integrate the prediction results obtained in step S311 to obtain the final face detection result in the classroom environment.

[0041] Furthermore, in step 301, the feature extraction module is composed of a CSPDarkNet network, the feature enhancement module is composed of an Attention network, a PAFAN network, and an ASFF network, and the feature point prediction is composed of a Yolo Head network.

[0042] Furthermore, the SPP module constructed in step S305 contains two CBS modules and three max-pooling operations with pooling kernel sizes of 7×7, 5×5, and 3×3 respectively;

[0043] Furthermore, in step S307, in the CSPDarkNet network, there are successively a Focus module, a CBS module, a CSP_1 module, a CBS module, a CSP_3 module, a CBS module, a CSP_3 module, a CBS module, an SPP module, and a CSP2_1 module. The outputs of the two CSP_3 modules and the CSP2_1 module are used as effective feature layers.

[0044] Furthermore, in step S311, the Yolo Head module first uses a convolution operation to adjust the number of channels of the input feature layer, and then inputs the adjusted feature layer into the classification branch and the regression branch respectively. Among them, the classification branch first uses two CBS modules to extract features, and then uses a convolution operation to predict the category. The regression branch first uses two CBS modules to extract features, and uses two 1×1 convolution operations respectively to obtain the confidence and regression parameters.

[0045] Furthermore, in step S4, the loss function L of the algorithm is:

[0046] L = 5·L EIOU +L OBJ +L CLS

[0047] Where L EIOU is the EIOU loss function, representing the loss of the prediction box, and L OBJ and L CLS are cross-entropy loss functions, representing the loss of confidence and the loss of class prediction respectively.

[0048] Further, the EIOU loss function L EIOU is as follows:

[0049]

[0050] where IOU represents the intersection over union of the predicted bounding box and the ground truth bounding box, b and b gt represent the center points of the predicted bounding box and the ground truth bounding box respectively, ρ represents the Euclidean distance between two points, c represents the diagonal distance of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box, and C w and C h represent the width and height of the minimum bounding rectangle of the predicted bounding box and the ground truth bounding box respectively.

[0051] Further, the cross-entropy loss function L OBJ and L CLS used to calculate L BWL is:

[0052] L BWL = -(y log σ(p) + (1 - y) log σ(1 - p))

[0053] where y is the label, p is the predicted value, and σ represents the sigmoid function.

[0054] Further, in steps S5 and S6, when training the face detection network YOLOXs-face, the Adam optimizer is used for optimization.

[0055] A detection system for a face detection method in a classroom environment based on the YOLOX deep network, comprising:

[0056] A partitioning module that partitions the face detection data set collected in the classroom environment into a training set, a validation set, and a test set;

[0057] A preprocessing module that adjusts the image sizes of the validation set and the training set, and then performs data augmentation on the training set;

[0058] A network module that constructs a YOLOXs-face network based on the YOLOX deep network;

[0059] A pre-training module: uses a pre-training data set to train the YOLOXs-face network to obtain a pre-trained model;

[0060] A training module that, based on the pre-trained model, uses the training set processed by the preprocessing module to train the YOLOXs-face network;

[0061] The verification module, while training, uses the verification set processed by the preprocessing module for verification and saves the network model with the best performance on the verification set.

[0062] The detection module uses the divided test set to test on the saved optimal network model to obtain the face detection results in the classroom environment.

[0063] The pre-trained dataset adopts the WIDER FACE dataset.

[0064] Compared with the prior art, the present invention has the following advantages:

[0065] 1) Use a spatial pyramid pooling structure with smaller kernels. The network structure of the spatial pyramid pooling structure is as Figure 6 shown. It uses pooling kernels of different sizes to perform max-pooling operations on the input feature layer, extracting spatial feature information of different sizes from the input feature layer. This processing can improve the detection accuracy and robustness of the model. And compared with those convolutional neural network models containing fully connected layers that can only process images of a fixed input size, the spatial pyramid pooling structure does not limit the size of the input image, making the usage scenario of the network more flexible. In the present invention, the sizes of the pooling kernels are sequentially modified to 7×7, 5×5, and 3×3. Using smaller-scale pooling kernels in the spatial pyramid pooling structure of the network can help the model more easily detect small-scale faces in the classroom environment and improve the overall face detection performance.

[0066] 2) Add a hybrid threshold attention mechanism fusion operation. The attention mechanism in computer vision draws on the human attention thinking mode. When processing visual information, humans will pay different degrees of attention to the received information, focusing on the information beneficial to result prediction and automatically ignoring irrelevant content. In computer vision, a mask is generally used to form the attention mechanism, and the model assigns different weights to each position of the input to achieve the purpose of focusing on important information and ignoring irrelevant content. Adding an attention mechanism to the network in the present invention can enable the model to learn to suppress useless background information and improve the detection accuracy. The flow chart of the attention module used in the present invention is as Figure 7 shown, and its detailed content is shown in Figure 8 .

[0067] 3) Incorporate adaptive spatial feature fusion operation. In YOLOX-s, the PAFPN network is used to perform feature fusion operations on three effective feature layers, and then high-level semantic information is used for large object detection and low-level semantic information is used for small object detection. In a classroom environment, the scales of faces in the front and back of the classroom generally vary greatly, that is, there are both large-scale and small-scale faces in the same picture. In this case, the conflicts between features on different layers often occupy the main part of PAFPN. This inconsistency will interfere with the gradient calculation during training and reduce the effectiveness of the feature pyramid. The present invention solves the inconsistency problem existing in the PAFPN structure by adding an adaptive spatial feature fusion module after the PAFPN structure. Figure 9 Taking the adaptive spatial feature fusion module-3 as an example, the structure of the adaptive spatial feature fusion module is shown. The implementation of adaptive feature fusion is simple, and the computational cost added to the model is negligible.

[0068] 4) Use EIOU to improve the loss function. Since the IOU loss used by the YOLOX-s network has limitations, the present invention uses the EIOU loss to replace the IOU loss when training the model. The EIOU loss function has three parts: IOU loss, center distance loss, and width-height loss. Among them, the width-height loss directly minimizes the width difference and height difference between the ground truth box and the predicted box, which can accelerate the convergence speed.

[0069] 5) Use transfer learning pre-training operation. For the face detection task in a classroom environment, the publicly available datasets for model training are too few, and the cost of making datasets is high. To solve the problem of insufficient data, the present invention introduces a transfer learning pre-training operation. First, the network is trained using the WIDER FACE dataset to obtain a general face detection model, and then the classroom environment face detection dataset is trained on the basis of the general model to obtain a face detection model for the classroom environment. Compared with randomly initializing network parameters, using the transfer learning pre-training operation can accelerate the convergence speed of the model during training and improve the accuracy of the model in face detection in a classroom environment. Description of the Drawings

[0070] Figure 1 It is a flowchart of the present invention.

[0071] Figure 2 It is the overall structure diagram of YOLOXs-face proposed by the present invention.

[0072] Figure 3 It is the structure diagram of the CSPDarkNet network of the present invention.

[0073] Figure 4 It is the network structure diagram of the PAFPN of the present invention.

[0074] Figure 5 It is the structural diagram of the Yolo Head module.

[0075] Figure 6 It is the network structural diagram of the spatial pyramid pooling structure.

[0076] Figure 7 It is the overall structural diagram of the hybrid attention module.

[0077] Figure 8 It is the detailed structural diagram of the channel attention mechanism and the spatial attention mechanism in the hybrid attention module. Among them, Figure (a) is the channel attention module, and Figure (b) is the spatial attention module.

[0078] Figure 9 Taking the adaptive spatial feature fusion module - 3 as an example to show the structure of the adaptive spatial feature fusion module. Specific implementation manners

[0079] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0080] The present invention provides a face detection method and its detection system in a classroom environment based on the YOLO deep network. First, the data set is divided into a training set, a validation set, and a test set; then data augmentation is performed on the training set; then a face detection convolutional neural network in a classroom environment based on the YOLOX deep network is constructed; the network model is trained using a pre - trained data set to obtain a pre - trained model; the network model is trained using the training set on the basis of the pre - trained model, and the model with the best performance on the validation set is saved; finally, the test set is tested using the best model to obtain the results of various indicators for face detection in a classroom environment. The present invention adds a series of improvement measures on the basis of the original YOLOX object detection algorithm, including using a spatial pyramid pooling structure with a smaller kernel, adding a hybrid attention module and an adaptive spatial feature fusion module, using the EIOU to improve the loss function, and using transfer learning pre - training operations, so as to effectively improve the accuracy of face detection in a classroom environment under the premise of less computing resources.

[0081] Please refer to Figure 1 , a face detection method in a classroom environment based on the YOLOX deep network, including the following steps:

[0082] S1. Divide the face detection data set in a classroom environment into a training set, a validation set, and a test set, specifically as follows:

[0083] S101. Write Python code to divide the SCUT-HEAD-PartA dataset into training set, validation set and test set. The SCUT-HEAD-PartA dataset consists of 2,000 images. Randomly select 1,100 images as the training set, 400 images as the validation set, and 500 images as the test set.

[0084] S2. Read the images in the training set and validation set divided in step S1, convert them to RGB format and adjust the size of the images. Then perform data augmentation on the training set divided in step S1, as follows:

[0085] S201. Preprocess the images in the validation set divided in step S1. First, convert the images to RGB format, then use the bilinear interpolation method to scale the images in the validation set and test set proportionally so that the long side of the image is 640. Finally, create a gray image with a size of 640×640 and place the scaled image in the center of the gray image.

[0086] S202. Preprocess the images in the training set divided in step S1. First, convert the images to RGB format, then scale the images proportionally, then randomly change the aspect ratio of the images, then create a gray image with a size of 640×640 and place the scaled image in the center of the gray image. Then, horizontally flip the images according to the probability, and finally randomly change the hue, saturation and brightness of the images to achieve data augmentation.

[0087] S203. Perform real box adjustment on the validation set preprocessed in step S201 and the training set preprocessed in step S202 respectively.

[0088] S3. Participate in Figure 2 Construct a face detection convolutional neural network based on the YOLOX deep network in the classroom environment, and name it YOLOXs-face, as follows:

[0089] S301. Construct a face detection network based on the YOLO deep network in the classroom environment, named YOLOXs-face; the YOLOXs-face network includes a feature extraction module, a feature enhancement module and a feature point prediction module. The feature extraction module consists of a CSPDarkNet network, the feature enhancement module consists of an Attention network, a PAFAN network and an ASFF network, and the feature point prediction consists of a Yolo Head network.

[0090] S302. Construct a CBS module containing 1 convolutional layer, 1 batch normalization layer and 1 SiLU non-linear activation layer.

[0091] The CBS module consists of 1 convolutional layer, 1 batch normalization layer, and 1 SiLU non-linear activation layer. The size of the convolutional kernel in the convolutional layer varies according to the usage scenario. The batch normalization layer and the SiLU activation layer follow the convolutional layer in sequence.

[0092] S303. Construct a residual module that contains 2 convolutional layers, 2 batch normalization layers, and 2 SiLU non-linear activation layers;

[0093] The residual module consists of 2 convolutional layers, 2 batch normalization layers, and 2 SiLU non-linear activation layers. The sizes of the convolutional kernels of the 2 convolutional layers are 1×1 and 3×3 in sequence. A batch normalization layer and a SiLU activation layer are added after each convolutional layer. After the output layer is connected to the input in a residual connection, it serves as the final output of this module.

[0094] S304. Construct a Focus module based on the CBS module in step S302. This module first slices the input image, expands the input from three channels to twelve channels, and then uses a CBS module to perform a convolutional operation on the feature layer;

[0095] The Focus module takes a value every other pixel in a picture, expands the input from three channels to twelve channels, and then uses a CBS module with a convolutional kernel size of 3×3 to adjust the number of channels of the feature layer.

[0096] S305. Construct an SPP module based on the CBS module in step S302. This module contains two CBS modules and three max-pooling operations with pooling kernel sizes of 7×7, 5×5, and 3×3 respectively;

[0097] The SPP module first uses a CBS module with a convolutional kernel size of 1×1 to adjust the number of channels of the input feature layer, then uses three max-pooling operations with sizes of 7×7, 5×5, and 3×3 respectively to perform max-pooling on the feature layer. Next, these three extracted feature layers and the initial feature layer are stacked, and then a CBS module with a convolutional kernel size of 1×1 is used to adjust the number of channels of the stacked feature layer to obtain the final output.

[0098] S306. Construct the CSP_N module and CSP2_N module based on the CBS module in step S302 and the residual module in step S303. The CSP_N module includes a backbone branch and a residual branch. Its backbone branch includes one CBS module and N residual modules, and its residual branch includes one CBS module. Input the data into the two branches respectively to obtain two feature layers of the same size. Stack the two feature layers and then pass them through one CBS module to obtain the output. The CSP2_N module includes a backbone branch and a residual branch. Its backbone branch includes one CBS module and N residual modules without residual edges, and its residual branch includes one CBS module. Input the data into the two branches respectively to obtain two feature layers of the same size. Stack the two feature layers and then pass them through one CBS module to obtain the output.

[0099] S307. Construct the feature extraction module CSPDarkNet network of the YOLOXs-face face detection network in step S301 based on the CBS module in step S302, the Focus module in step S304, the SPP module in step S305, and the CSP_N module and CSP2_N module in step S306. This structure will perform feature extraction operations on the input data. Input the data in the training set enhanced in step S2 into the CSPDarkNet network, and obtain three effective feature layers with sizes of 80×80×128, 40×40×256, and 20×20×512 respectively in the middle layer, middle-lower layer, and bottom layer of the CSPDarkNet structure.

[0100] In the CSPDarkNet network, there are a Focus module, a CBS module, a CSP_1 module, a CBS module, a CSP_3 module, a CBS module, a CSP_3 module, a CBS module, an SPP module, and a CSP2_1 module in sequence. Take the outputs of the two CSP_3 modules and the CSP2_1 module as effective feature layers. The convolution kernels of the above CBS modules are all 3×3.

[0101] S308. Construct the feature enhancement module Attention network of the YOLOXs-face face detection network in step S301. This network consists of three CBAM attention modules. Input the three effective feature layers obtained in step S307 into the three CBAM attention modules respectively to obtain three hybrid attention feature layers with sizes of 80×80×128, 40×40×256, and 20×20×512 respectively.

[0102] Channel attention module: First, perform global max pooling and global average pooling on the input feature layer respectively. Then, use a shared fully connected layer to process the pooled feature layer. Next, add the two resulting outputs together, and then use the Sigmod activation function to process and obtain the weights of each channel of the input feature layer. Finally, multiply the weights by the input feature layer to obtain the output;

[0103] Spatial attention module: First, take the maximum value and the average value on the channels of each feature point of the input feature layer. Then, stack these two results, use a convolution with 1 output channel to adjust the number of channels, and then use the Sigmod activation function to process and obtain the weights of each channel of the input feature layer. Finally, multiply the weights by the input feature layer to obtain the output.

[0104] S309. Construct the feature enhancement module PAFPN network of the face detection network YOLOXs-face in step S301 based on the CBS module in step S302 and the CSP2_N module in step S306. This network is composed of an FPN and a PAN network. The specific structure is as Figure 4 shown; Input the three hybrid attention feature layers obtained in step S308 into the PAFPN network. First, perform feature transfer and fusion through upsampling in the FPN network, and then obtain three enhanced feature layers with sizes of 80×80×128, 40×40×256, and 20×20×512 respectively through downsampling fusion in the FAN network;

[0105] S310. Construct the feature enhancement module ASFF network of the face detection network YOLOXs-face in step S301. This network is composed of three adaptive spatial feature fusion modules; Input the three enhanced feature layers obtained in step S309 into the ASFF network, and let different feature layers adaptively fuse to obtain three fused feature layers with sizes of 80×80×128, 40×40×256, and 20×20×512 respectively;

[0106] The adaptive feature fusion module enables the network to directly learn how to perform spatial filtering on the features of other layers, thereby only retaining useful information for combination. For a certain feature layer, the adaptive feature fusion module first integrates and adjusts other feature layers to the same size, and then trains to find the optimal fusion method. At each spatial position, different feature layers are adaptively fused, filtering the features carrying contradictory information and strengthening the discriminative features.

[0107] S311. Construct the feature point prediction Yolo Head network of the face detection network YOLOXs-face in step S301 based on the CBS module in step S302. This network is composed of three Yolo Head modules. The structure of the Yolo Head module is asFigure 5 As shown in the figure; input the three fused feature layers obtained in step S310 into the Yolo Head network, perform classification and regression operations on the feature layers, and obtain three prediction results with sizes of 80×80×6, 40×40×6, and 20×20×6 respectively;

[0108] The Yolo Head module first uses a 1×1 convolution operation to adjust the number of channels of the input feature layer, and then inputs the adjusted feature layer into the classification branch and the regression branch respectively. Among them, the classification branch first uses two CBS modules to extract features, and then uses a 1×1 convolution operation to predict the category. The regression branch first uses two CBS modules to extract features, and uses two 1×1 convolution operations to obtain the confidence and regression parameters respectively;

[0109] S312. Integrate the prediction results obtained in step S311 to obtain a result with a size of 8400 (80×80 + 40×40 + 20×20)×6, where 8400 represents the number of prediction boxes finally obtained by the network, and 6 represents the prediction result of the network, including the regression coefficients (x, y, w, h) of the prediction box, the confidence that the prediction box contains an object, and the probability that the object in the prediction box is a face;

[0110] S4. Use the EIOU loss function and the cross-entropy loss function to construct the loss function L of this method; specifically:

[0111] L = 5·L EIOU +L OBJ +L CLS

[0112] Among them, L EIOU is the EIOU loss function, representing the loss of the prediction box, and L OBJ and L CLS are the cross-entropy loss functions, representing the loss of confidence and the loss of class prediction respectively;

[0113] Furthermore, the EIOU loss function L EIOU is:

[0114]

[0115] Among them, IOU represents the intersection over union of the prediction box and the ground truth box, b and b gt represent the center points of the prediction box and the ground truth box respectively, ρ represents the Euclidean distance between two points, c represents the diagonal distance of the minimum bounding box of the prediction box and the ground truth box, C w and C h represent the widths and heights of the minimum bounding boxes of the prediction box and the ground truth box respectively;

[0116] Furthermore, used to calculate L OBJand L CLS The cross-entropy loss function L BWL is as follows:

[0117] L BWL = -(y log σ(p) + (1 - y) log σ(1 - p))

[0118] where y is the label, p is the predicted value, and σ represents the sigmoid function.

[0119] S5. Use the pre-training dataset to train the YOLOXs-face network to obtain a pre-trained model; when training the face detection network YOLOXs-face, the batch size is 24, the Adam optimizer is used for optimization, the initial learning rate is 0.001, and the learning rate is multiplied by 0.98 for each training round, and a total of 400 rounds of training are performed.

[0120] S6. Use the training set processed in step S2 to continue training the YOLOXs-face network based on the pre-trained model obtained in step S5, and use the validation set processed in step S2 for validation, and save the network model with the best performance on the validation set; in steps S5 and S6, when training the face detection network YOLOXs-face, the batch size is 24, the Adam optimizer is used for optimization, the initial learning rate is 0.001, and the learning rate is multiplied by 0.98 for each training round, and a total of 400 rounds of training are performed.

[0121] S7. Use the test set divided in step S1 to test on the network model saved in step S6 to obtain the face detection results in the classroom environment;

[0122] S8. For the detection results obtained in step S7, use the F1 score and the average precision to quantitatively evaluate the detection performance of the network model.

[0123] A detection system for a face detection method in a classroom environment based on the YOLOX deep network, comprising:

[0124] A partitioning module that partitions the face detection dataset collected in the classroom environment into a training set, a validation set, and a test set;

[0125] A preprocessing module that adjusts the image sizes of the validation set and the training set, and then performs data augmentation on the training set;

[0126] A network module that constructs a YOLOXs-face network based on the YOLOX deep network;

[0127] A pre-training module: Use the pre-training dataset to train the YOLOXs-face network to obtain a pre-trained model;

[0128] The training module trains the YOLOXs-face network using the training set processed by the preprocessing module based on the pre-trained model.

[0129] The verification module, while training, uses the verification set processed by the preprocessing module for verification and saves the network model with the best performance on the verification set.

[0130] The detection module tests using the divided test set on the saved optimal network model to obtain the face detection results in the classroom environment.

[0131] The pre-trained dataset uses the WIDER FACE dataset.

[0132] Simulation experiment

[0133] 1. Experimental conditions:

[0134] Table 1 Experimental environment configuration of the present invention

[0135]

[0136] 2. Simulation content and result analysis:

[0137] The samples in the simulation experiment of the present invention come from three parts: The first part is the SCUT-HEAD dataset, which is a face detection dataset for the classroom environment and is divided into two parts, A and B. The images in part A are from classroom surveillance videos, and the images in part B are from the Internet. In the present invention, it is used to train the network to obtain a face detection model in the classroom environment and verify the effectiveness of the proposed improvement; the second part is the WIDER FACE dataset, which is a popular face detection benchmark dataset and is used in the present invention to implement transfer learning pre-training operations; the third part is the picture data of the real classroom environment collected and annotated by us from the Internet, which is used in the present invention to test the generalization ability of the model.

[0138] The images in the data used in the present invention have inconsistent sizes, and their sizes are unified to 640×640 in the data preprocessing stage.

[0139] Next, the detection performance of the YOLOXs-face network model proposed by the present invention is quantitatively evaluated using the F1 score and the Average Precision (AP). The following are the specific meanings of each index:

[0140] TP (True Positive): True positive example, representing the positive samples that are correctly classified.

[0141] FN (False Negative): False negative example, representing the positive samples that are misclassified.

[0142] FP (False Positive): False positive example, representing a negative sample that is misclassified;

[0143] TN (True Negative): True negative example, representing a negative sample that is correctly classified.

[0144] In the face detection task, to obtain the above metrics, it is necessary to first determine whether each prediction result is correct. Different from classification problems, in the face detection task, it is necessary to calculate the Intersection over Union (IOU) between the predicted bounding box and the ground truth bounding box to determine whether the detection result is correct. The calculation method of IoU:

[0145]

[0146] Among them, A and B represent the predicted bounding box and the ground truth bounding box of the face respectively.

[0147] First, input the image into the model to obtain the predicted bounding box. For each predicted bounding box, calculate its IOU value with all the ground truth bounding boxes in the image, and take the maximum IOU value as MaxIOU. At this time, set a threshold (usually set to 0.5). When MaxIOU is greater than this threshold, the predicted bounding box is classified as a true positive example TP; otherwise, the predicted bounding box is classified as a false positive example FP.

[0148] Recall is for the original samples, which represents how many positive examples in the samples are predicted correctly. There are also two possibilities. One is to predict the original positive class as a positive class (TP), and the other is to predict the original positive class as a negative class (FN). Among them, TP + FN is equal to the number of ground truth bounding boxes:

[0149]

[0150] Precision is for the prediction results, which represents how many of the samples predicted as positive are truly positive samples. Then there are two possibilities for predicting as positive. One is to predict the positive class as a positive class (TP), and the other is to predict the negative class as a positive class (FP):

[0151]

[0152] The F1 score is a metric used to measure the accuracy of a binary classification model, taking into account both the accuracy and recall of the classification model. It can be regarded as a weighted average of accuracy and recall:

[0153]

[0154] Average Precision (AP) is a performance metric for algorithms that predict the target location and category:

[0155]

[0156] Among them, p represents Precision, r represents Recall, and p is a function with r as a parameter.

[0157] The present invention uses ablation experiments to verify the effectiveness of the improvement.

[0158] Table 2 List of ablation experiment results obtained from the simulation experiments of the present invention

[0159]

[0160] Combined with the results in Table 2, it can be seen that the improvements made by the present invention on the YOLOX deep network are all effective. Among them, EIOU means replacing the IOU loss function in YOLOX-s with the EIOU loss function, and this improvement increases the detection accuracy of the model by 1%; ASFF means adding an adaptive spatial feature fusion module to the network, and this improvement increases the detection accuracy of the model by 0.1%; Attention means adding the CBAM attention mechanism module to the network, and after adding it, the detection accuracy of the model increases by 0.05%; SPP(3,5,7) means using smaller pooling kernels in the spatial pyramid pooling structure of the backbone network, and this improvement improves the detection performance of the network by 0.07%; finally, Pretrained means using transfer learning pre-training operation, and this improvement improves the detection performance of the network by 0.83%.

[0161] The present invention compares the detection results of different networks for faces. Among them, YOLO-face is a face detection network based on YOLOv3, and Tinaface is one of the most advanced face detectors with the best detection effect currently.

[0162] Table 3 List of model comparison results obtained from the simulation experiments of the present invention

[0163]

[0164] As can be seen from the results in Table 3, compared with other algorithms, the YOLOXs-face method proposed in the present invention is very balanced and more suitable for real-time face detection tasks in the classroom environment. Compared with the YOLO-face algorithm, the YOLOXs-face method proposed in the present invention has better face detection performance, fewer model parameters and computational complexity, and faster model detection speed. Compared with the YOLOX-s algorithm, the YOLOXs-face method proposed in the present invention greatly improves the detection accuracy of the model while only increasing a small amount of model parameters, computational complexity, and detection time. Compared with the Tinaface algorithm, although the detection accuracy of the YOLOXs-face method proposed in the present invention is lower, its model parameters, computational complexity, and image detection time are much smaller than those of Tinaface. Therefore, YOLOXs-face has lower requirements for hardware devices, which is conducive to the popularization and application of the method. At the same time, the F1 score of YOLOXs-face is higher than that of Tinaface. Considering comprehensively, the performance of YOLOXs-face is better.

[0165] In summary, the present invention provides a face detection method and its detection system in the classroom environment based on the YOLO deep network, which makes a series of improvements on the original YOLOX algorithm, including using a spatial pyramid pooling structure with a smaller kernel, adding a hybrid attention module and an adaptive spatial feature fusion module, using the EIOU to improve the loss function, and using transfer learning pre-training operations, so as to improve the accuracy of face detection in the classroom environment under the premise of less computing resources.

[0166] The present invention has better detection effect on faces in the classroom environment. First, the present invention uses a smaller pooling kernel in the spatial pyramid pooling structure of the network, which can help the model more easily detect small-scale faces in the classroom environment and improve the overall face detection performance. Second, the present invention adds a hybrid threshold attention mechanism fusion operation and an adaptive spatial feature fusion operation, which reduces the influence of the environment on face detection and the influence between faces of different scales, and reduces the probability of misdetection.

[0167] The present invention has low requirements for hardware devices and good universality. The model size of the method proposed in the present invention is smaller than that of the prior art and can run well on devices with small memory.

[0168] The present invention has low computational cost and short detection time. Compared with the prior art, the computational complexity of the network proposed in the present invention is smaller and can run well on devices with poor performance.

[0169] The model in the present invention adopts a modular design concept, adding or modifying modules according to the deficiencies of the basic network in the face detection task in the classroom environment. With the development of emerging technologies and the proposal of better network modules, the present invention can be iteratively updated at any time to improve the model performance.

Claims

1. A face detection method in a classroom environment based on the YOLO deep network, characterized in that: Specifically, it includes the following steps: S1. Divide the face detection dataset in the classroom environment into a training set, a validation set, and a test set; S2. Read the images in the training set and validation set divided in step S1, convert them to the RGB format and adjust the size of the images, and then perform data augmentation on the training set divided in step S1; S3. Build a convolutional neural network for face detection in the classroom environment based on the YOLOX deep network, and name it YOLOXs-face; S4. Build a loss function using the EIOU loss function and the cross-entropy loss function; S5. Use the pre-training dataset to train the YOLOXs-face network to obtain a pre-trained model; S6. Continue to train the YOLOXs-face network on the basis of the pre-trained model obtained in step S5 using the training set processed in step S2, use the validation set processed in step S2 for validation, and save the network model with the best performance on the validation set; S7. Test using the test set divided in step S1 on the network model saved in step S6 to obtain the face detection result in the classroom environment; S8. For the detection result obtained in step S7, use the F1 coefficient and the average precision to quantitatively evaluate the detection performance of the network model.

2. The face detection method in a classroom environment based on the YOLO deep network according to claim 1, characterized in that: In step S1, the samples in the face detection dataset in the classroom environment are randomly divided and divided into a training set, a validation set, and a test set according to the ratio of 11:4:

5.

3. A face detection method in a classroom environment based on the YOLO deep network according to claim 1, characterized in that: The specific method of step S2 is as follows: S201. Preprocess the images in the validation set divided in step S1. First, convert the images to the RGB format, then use the bilinear interpolation method to scale the images in the validation set and the test set proportionally, and finally unify the size of the images by adding gray bars to the images; S202. Preprocess the images in the training set divided in step S1. First, convert the images to the RGB format, then scale the images proportionally, and then randomly scale the aspect ratio of the images; unify the size of the images by adding gray bars to the images, and horizontally flip the images according to the probability, and finally randomly change the hue, saturation, and brightness of the images to achieve data augmentation; S203. Perform real box adjustment on the validation set preprocessed in step S201 and the training set preprocessed in step S202 respectively.

4. A face detection method in a classroom environment based on the YOLO deep network according to claim 1, characterized in that: The specific method of step S3 is as follows: S301. Build a face detection network in the classroom environment based on the YOLO deep network, named YOLOXs-face; the YOLOXs-face network includes a feature extraction module, a feature enhancement module, and a feature point prediction module; S302. Build a CBS module including a convolutional layer, a batch normalization layer, and a SiLU non-linear activation layer; S303. Build a residual module including a convolutional layer, a batch normalization layer, and a SiLU non-linear activation layer; S304. Build a Focus module based on the CBS module in step S302. This module first slices the input image, expands the input from three channels to twelve channels, and then uses a CBS module to perform a convolutional operation on the feature layer; S305. Construct the SPP module based on the CBS module in step S302. This module consists of the CBS module and the max pooling operation; S306. Construct the CSP_N module and the CSP2_N module based on the CBS module in step S302 and the residual module in step S303. The CSP_N module includes a backbone branch and a residual branch. Its backbone branch contains one CBS module and N residual modules, and its residual branch contains one CBS module. Input the data into the backbone branch and the residual branch respectively to obtain feature layers of the same size. Stack the feature layers and then pass them through the CBS module to obtain the output; The CSP2_N module includes a backbone branch and a residual branch. Its backbone branch contains one CBS module and N residual modules without residual edges, and its residual branch contains one CBS module. Input the data into the backbone branch and the residual branch respectively to obtain feature layers of the same size. Stack the feature layers and then pass them through the CBS module to obtain the output; S307. Construct the feature extraction module CSPDarkNet network of the face detection network YOLOXs-face in step S301 based on the CBS module in step S302, the Focus module in step S304, the SPP module in step S305, and the CSP_N module and the CSP2_N module in step S306. This structure will perform feature extraction operations on the input data; Input the data in the training set after data augmentation in step S2 into the CSPDarkNet network, and obtain three effective feature layers in the middle layer, the middle and lower layer, and the bottom layer of the CSPDarkNet structure; S308. Construct the feature enhancement module Attention network of the face detection network YOLOXs-face in step S301. This network consists of three CBAM attention modules; Input the three effective feature layers obtained in step S307 into the three CBAM attention modules respectively to obtain three mixed attention feature layers; S309. Construct the feature enhancement module PAFPN network of the face detection network YOLOXs-face in step S301 based on the CBS module in step S302 and the CSP2_N module in step S306. This network consists of the FPN and PAN networks; Input the three mixed attention feature layers obtained in step S308 into the PAFPN network. First, perform feature transfer and fusion in the FPN network through upsampling, and then obtain three enhanced feature layers through downsampling fusion in the FAN network; S310. Construct the feature enhancement module ASFF network of the face detection network YOLOXs-face in step S301. This network consists of three adaptive spatial feature fusion modules; Input the three enhanced feature layers obtained in step S309 into the ASFF network to adaptively fuse different feature layers and obtain three fused feature layers; S311. Based on the CBS module in step S302, construct the feature point prediction Yolo Head network of the face detection network YOLOXs-face in step S301. This network consists of Yolo Head modules; input the three fused feature layers obtained in step S310 into the Yolo Head network, perform classification and regression operations on the feature layers, and obtain prediction results at three different scales. S312. Integrate the prediction results obtained in step S311 to obtain the final face detection result in the classroom environment.

5. The face detection method in a classroom environment based on the YOLO deep network according to claim 4, characterized in that: In step 301, the feature extraction module consists of a CSPDarkNet network, the feature enhancement module consists of an Attention network, a PAFAN network, and an ASFF network, and the feature point prediction consists of a Yolo Head network; in step S307, the CSPDarkNet network sequentially includes a Focus module, a CBS module, a CSP_1 module, a CBS module, a CSP_3 module, a CBS module, a CSP_3 module, a CBS module, an SPP module, and a CSP2_1 module. The outputs of the two CSP_3 modules and the CSP2_1 module are used as effective feature layers.

6. A face detection method in a classroom environment based on the YOLO deep network according to claim 4, characterized in that The SPP module constructed in step S305 includes two CBS modules and three max-pooling operations with pooling kernel sizes of 7×7, 5×5, and 3×3 respectively.

7. A face detection method in a classroom environment based on the YOLO deep network according to claim 4, characterized in that: In step S311, the Yolo Head module first uses a convolution operation to adjust the number of channels of the input feature layer, and then inputs the adjusted feature layer into the classification branch and the regression branch respectively. Among them, the classification branch first uses two CBS modules to extract features, and then uses a convolution operation to predict the category. The regression branch first uses two CBS modules to extract features, and uses two 1×1 convolution operations to obtain the confidence and regression parameters respectively.

8. A face detection method in a classroom environment based on the YOLO deep network according to claim 1, characterized in that: In step S4, the loss function L is: L = 5·L EIOU + L OBJ + L CLS Among them, L EIOU is the EIOU loss function, representing the loss of the predicted bounding box, L OBJ and L CLS are the cross-entropy loss functions, representing the loss of confidence and the loss of class prediction respectively; Furthermore, the EIOU loss function L EIOU is as follows: Among them, IOU represents the intersection over union of the predicted bounding box and the ground truth bounding box, b and b gt respectively represent the center points of the predicted bounding box and the ground truth bounding box, ρ represents the Euclidean distance between two points, c represents the diagonal distance of the minimum bounding box of the predicted bounding box and the ground truth bounding box, C w and C h respectively represent the width and height of the minimum bounding box of the predicted bounding box and the ground truth bounding box; Further, for calculating L OBJ and L CLS The cross-entropy loss function L BWL is: L BWL = -(y log σ(p) + (1 - y) log σ(1 - p)) where y is the label, p is the predicted value, and σ represents the sigmod function.

9. The face detection method in a classroom environment based on the YOLO deep network according to claim 1, wherein: In steps S5 and S6, when training the face detection network YOLOXs-face, the Adam optimizer is used for optimization.

10. A detection system for implementing any one of the detection methods of claims 1 to 9, characterized in that: It includes: A partitioning module that partitions the face detection data set collected in the classroom environment into a training set, a validation set, and a test set; A preprocessing module that adjusts the image sizes of the validation set and the training set, and then performs data augmentation on the training set; A network module that constructs the YOLOXs-face network based on the YOLOX deep network; A pre-training module: Use the pre-training data set to train the YOLOXs-face network to obtain a pre-trained model; A training module that, based on the pre-trained model, uses the training set processed by the preprocessing module to train the YOLOXs-face network; A validation module that, during training, uses the validation set processed by the preprocessing module for validation, and saves the network model with the best performance on the validation set; A detection module that uses the divided test set to test on the saved optimal network model to obtain the face detection result in the classroom environment.

Citation Information

Patent Citations

  • Student classroom behavior detection method, system and terminal based on deep learning

    CN114359606A

  • Video understanding-based non-inductive dish taking model establishment method and device, and dish taking method and device

    CN114743153A