Face detection method based on improved MY-DCF network

By improving the MY-DCF network, using MobileNetV3 to replace CSPDarkNet53 and introducing NHS activation function and feature enhancement module, the problem of large calculations on small devices is solved, and efficient real-time face detection is achieved.

CN120340094APending Publication Date: 2025-07-18NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510438992.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When existing face detection algorithms are deployed on small devices with limited computing resources, there are problems such as complex models, large computing volume and inconvenient deployment, and lack real-time performance.

Method used

Using the improved MY-DCF network, by replacing YOLOv4's backbone extraction network CSPDarkNet53 with MobileNetV3, and introducing NHS activation function, combining the depth separation convolution, feature enhancement module and convolution attention module, the network structure is optimized to reduce the amount of parameters and calculations.

Benefits of technology

On the premise of ensuring recognition accuracy, the amount of network parameters and calculations are significantly reduced, the detection rate is improved, and real-time face detection is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340094A_ABST
    Figure CN120340094A_ABST
Patent Text Reader

Abstract

The invention discloses a face detection method based on an improved MY-DCF network. The method comprises the following steps: S1, performing labeling and preprocessing operation on an obtained face data set; s2, constructing a face detection network model based on an improved MY-DCF network; s3, respectively inputting the preprocessed face data set into a face detection network model, calculating a loss function and performing training, and storing an optimal model weight to obtain a trained face detection network; and S4, inputting a detected face image into the face detection network model obtained in the step S3, and performing face detection. According to the method, a backbone extraction network CSPDarkNet53 in a YOLOv4 network is replaced by MobilenetV3, an NHS activation function is introduced, the calculation efficiency and the nonlinear characteristic are balanced by dynamically adjusting parameters, the parameter quantity and size of a face detection network model are reduced, and the detection rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and particularly to a face detection method based on an improved MY-DCF network. Background Art

[0002] Face detection, as an important task in the field of computer vision, plays a crucial role in numerous face-related applications. With the rapid development of artificial intelligence technology and the wide application of deep learning methods, significant progress has been made in face detection methods based on deep learning. However, in actual application scenarios, most face detection tasks need to be deployed on small devices with limited computing resources. Although current face detection algorithms have high accuracy, they often suffer from problems such as complex models, large computational requirements, and inconvenient deployment, and do not have good real-time performance. Summary of the Invention

[0003] Object of the Invention: The object of the present invention is to provide a face detection method based on an improved MY-DCF network, which reduces the number of network parameters and computational requirements while ensuring the recognition accuracy.

[0004] Technical Solution: A face detection method based on an improved MY-DCF network includes the following steps: S1, perform annotation and preprocessing operations on the obtained face dataset; S2, construct a face detection network model based on the improved MY-DCF network; S3, input the preprocessed face dataset into the face detection network model respectively, calculate the loss function and perform training, save the optimal model weights, and obtain the trained face detection network; S4, input the detected face image into the face detection network model obtained in step S3 for face detection.

[0005] Further, the operations of performing annotation and preprocessing on the obtained face dataset include the following steps: S11, annotate the face image using the labelImg image annotation software. The annotation includes "face rectangular area position" and "face semantic information". In the rectangular area position of the face and the surrounding shoulder area, generate an XML file after annotation; write a program in Python to convert the annotation position coordinates saved in the XML file into a corresponding text file; S12, preprocess the annotated face dataset, introduce a multi-scale training strategy to scale and crop the original dataset and perform normalization processing.

[0006] Further, the steps of constructing the face detection network model are as follows: First, replace the backbone extraction network CSPDarkNet53 of YOLOv4 with the MobileNetV3 network. At the same time, during the feature fusion process, replace the ordinary convolutions in the PANet network with depthwise separable convolutions; finally, the output network uses a target detection head. Among them, the neck structure of the MobileNetV3 network uses a two-dimensional convolutional layer for initial feature extraction, and then passes through multiple bneck modules in sequence. Some bneck modules are equipped with SE attention modules and NHS activation functions to extract multi-level features of the image; Then, delete the SPP module of the CSPDarkNet53 network and add a feature enhancement module. The feature enhancement module is located on the output side of the backbone network and is connected to the PANet network; In the PANet network, multiple depthwise separable convolutions are stacked for different resolution branches. In each scale branch, a convolutional attention module is inserted after the depthwise separable convolution stack; an upsampling operation is performed on the (52, 52) scale branch, and upsampling and downsampling operations are performed on the (26, 26) scale branch in sequence; Finally, detection heads are deployed on three scales of (52, 52), (26, 26), and (13, 13) respectively.

[0007] Furthermore, the expression of the activation function NHS is: , where, is the input, is the Sigmoid function, is the parameter controlling the degree of non-linearity, and e is the natural constant.

[0008] Furthermore, the depthwise separable convolution first uses N convolutional kernels for separate convolutions, and then convolves the 3×3 feature map with the N convolutional kernels; and different convolutional kernels are used for convolution operations according to the number of input channels. Finally, the output result channels are adjusted by convolutional kernels with a unit of 1; the expression of the depthwise separable convolution is as follows: , where, NESC() represents the depthwise separable convolution operation, z is the input, DW() is the depth convolution operation, PW() is the pointwise convolution operation, and Evo() is the evolutionary computing module, is the row index, is the column index, is the channel index, is the number of channels; is the weight parameter of the depth convolution part, is the weight parameter of the pointwise convolution part.

[0009] Furthermore, the feature enhancement module includes three branches and one bypass pruning. Each branch consists of convolutional kernels of different sizes and dilated convolutions. These branches run in parallel and generate outputs with different scales and features. Then, the outputs of each branch are cascaded through the bypass pruning to form the final feature representation; Each branch is composed of different numbers of convolutional layers, and dilated convolutions with dilation rates of 1, 2, and 3 are respectively adopted. Finally, the operation of stacking and fusing the subnet results is performed to obtain the feature map with enhanced features. Among them, the bypass pruning is a residual connection used to solve the problems of gradient disappearance and gradient explosion in the training of deep networks.

[0010] Furthermore, the expression of the convolutional attention module is as follows: , where X represents the input feature, represents the size of the convolutional kernel used in the th convolutional operation. Conv( ) represents the convolutional operation, Concat( ) represents the concatenation operation, represents the feature of the th scale of the adaptive convolutional kernel adjustment. i = 1, 2,..., a, where a represents the total number of scales; represents the feature after concatenating multiple scale features, represents Sigmoid, represents the result obtained by multiplying the channel feature map by F, represents the result obtained by multiplying the spatial feature map by . FC( ) represents the fully connected layer, and y( ) represents the feature compressed by global average pooling. represents element-wise multiplication, and F represents the final feature representation.

[0011] Compared with the prior art, the present invention has the following remarkable effects: 1. In the present invention, the backbone extraction network CSPDarkNet53 in the YOLOv4 network is replaced with MobilenetV3, and the NHS activation function is introduced. By dynamically adjusting the parameters, a balance is achieved between computational efficiency and non-linear characteristics, greatly reducing the number of parameters and the size of the face detection network model, and significantly improving the detection rate; 2. The present invention introduces depthwise separable convolutions, which can automatically adjust the structure of the depth convolutional kernel through evolutionary computation to meet the requirements of specific tasks. On the premise of ensuring the face detection accuracy, the number of parameters of the face detection network model is greatly reduced, and the model size is significantly reduced by about 80% compared with the original model. And the calculation speed is increased by nearly 67%; 3. The present invention introduces a feature enhancement module and an attention mechanism. The feature enhancement module uses dilated convolution and adopts a multi-branch convolutional layer structure with different dilation rates to capture a wider range of context information. The convolutional attention module dynamically adjusts the size of the convolutional kernel, enabling the network to adaptively adjust the convolutional kernel according to the task requirements, thereby improving the flexibility and generalization ability of the face detection network model and solving the deficiency that traditional convolutional modules cannot adapt to tasks. By using convolutional kernels of multiple scales and fusing multi-scale features, compared with existing models, the face detection network model of the present invention has improved the detection accuracy by nearly 3%, and can reduce the number of network parameters and computational complexity while ensuring the recognition accuracy, meeting the requirements of real-time detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flowchart of the present invention; Figure 2 is a structural diagram of the network model of the present invention; Figure 3 is an improved diagram of the backbone network of the network model; Figure 4 is a structural diagram of depthwise separable convolution; Figure 5 is a structural diagram of the feature enhancement module; Figure 6 is a structural diagram of the SE module; Figure 7 is a structural diagram of the DCAM module; Figure 8 In (a) of, it is a comparison diagram of the accuracy curve and loss curve before and after improving the MY-DCF network using the LFW dataset, and (b) is a comparison diagram of the accuracy curve and loss curve before and after improving the MY-DCF network using the WIDER Face dataset; Figure 9 In (a) of, it is a comparison diagram of the accuracy curve and loss curve of the present invention and different algorithms using the LFW dataset, and (b) is a comparison diagram of the accuracy curve and loss curve of the present invention and different algorithms using the WIDER Face dataset; Figure 10 is a detection result diagram of the improved MY-DCF network model algorithm. DETAILED DESCRIPTION OF THE INVENTION

[0013] The present invention will be further described in detail below with reference to the accompanying drawings of the specification and the specific embodiments.

[0014] To solve the problems of complex model, large computational volume and inconvenient deployment, and to achieve the robustness of the detection method, the present invention provides a face detection method based on an improved MY-DCF (MoblieNetV3 YOLOv4-Depthwise Separable Convolution Dynamic Convolutional Attention Module Feature Enhance Module) network. Figure 1 As the overall flowchart of the method of the present invention, it is a real-time face detection algorithm, including the following steps: Step 1, obtain face images. The present invention selects the publicly available Labeled Faces in the Wild (LFW) and WIDER Face face datasets on the Internet.

[0015] Label the face images using the labelImg image annotation software. The annotation includes "the position of the face rectangular area" and "face semantic information", that is, the rectangular area position of the face and the surrounding shoulder area. After annotation, an XML file is generated; write a program in Python to convert the annotation position coordinates saved in the XML file into a corresponding text file.

[0016] Construct a dataset from the annotated face images and perform preprocessing. Introduce a multi-scale training strategy to scale, crop and normalize the original dataset, and use images of different scales as training samples to increase the diversity of the dataset.

[0017] Both of the preprocessed datasets can be used for training and evaluating the face detection algorithm.

[0018] Step 2, construct a face detection network model based on the improved MY-DCF network, and input the preprocessed face pictures into the face detection network to calculate the loss function and perform training. The construction of the face detection network based on the improved MY-DCF network includes: Propose a new activation function NHS (New Hard Swish), which can by adjusting the size of, so that the new activation function NHS achieves a balance between computational efficiency and non-linear characteristics, so as to better adapt to different deep learning tasks and hardware environments; the activation function NHS, the expression is: , , where, is the input, is the Sigmoid function, is the parameter controlling the non-linear degree, and e is the natural constant.

[0019] Replace the backbone extraction network CSPDarkNet53 in YOLOv4 with the MobileNetV3 network, and propose a depthwise separable convolution (NESC, Neural Evolutionary Separable Convolution) to replace the traditional 3×3 convolution in the PANet (Path Aggregation Network). Introduce the idea of evolutionary computing, and adjust the structure of the depth convolution kernel through adaptive evolution, which can improve the feature extraction effect and reduce the computational complexity and the number of parameters of the model. The computational complexity of the standard convolution is Q1, the sum of the computational complexities of the DW convolution and the PW convolution is Q2, and the reduced computational complexity is Q3. The calculation formulas are as shown in Equations (3) to (5): , , , Among them, the computational complexity of the standard convolution is Q1, the sum of the computational complexities of the DW convolution and the PW convolution is Q2, the reduced computational complexity is Q3, and the input mapping size is , the Kernel of the standard convolution is , and are the width and height of the input mapping respectively, M is the number of input channels, D is the height and width of the convolution kernel, and N is the number of output channels.

[0020] Design a feature enhancement module to replace the SPP module (Spatial Pyramid Pooling). By adopting dilated convolution and using a multi-branch convolution layer structure with different dilation rates, it can capture more extensive context information.

[0021] Add a convolutional attention module to the face detection network. The SE module (Squeeze-and-Excitation) in the neck performs compression and excitation operations. Design a DCAM module (Dynamic Convolutional Block Attention Module). By dynamically adjusting the size of the convolution kernel, it can adapt to input features of different scales, fuse features of multiple scales, and input the features into the DCAM module during each upsampling process to strengthen feature extraction.

[0022] Figure 2It is the structural diagram of the face detection network model of the present invention. This network mainly consists of a MobileNetV3 backbone network, a feature enhancement module, a PANet network, and a final Head Network. The image is first input into the backbone network for initial feature extraction using two-dimensional convolutional layers, and then sequentially passes through multiple bneck modules (neck modules). Some of them contain SE modules to enhance the feature response of key channels, combined with the NHS activation function to extract multi-level features of the image. The FEM module (feature enhancement module) is located on the output side of the backbone network and is connected to the PANet network to enhance the feature expression ability. The PANet network realizes multi-scale feature fusion. Among them, multiple NESC (depthwise separable convolution) are stacked in different resolution branches. In each scale branch, a DCAM module is inserted after the NESC stacking. An upsampling operation is performed on the (52, 52) scale branch, and upsampling and downsampling operations are sequentially performed on the (26, 26) scale branch. Finally, detection heads are deployed at three scales of (52, 52), (26, 26), and (13, 13) to achieve multi-scale face detection.

[0023] The detailed implementation steps are as follows: First, replace the backbone extraction network CSPDarkNet53 of YOLOv4 with a lightweight MobileNetV3 network. At the same time, during the feature fusion process, replace the ordinary convolution in the PANet network with depthwise separable convolution (NESC) to improve the feature extraction effect, realize the lightweight of the network model, and balance speed and accuracy. Then, delete the SPP module in CSPDarkNet53, add a feature enhancement module (FEM module), and use dilated convolution. By using a multi-branch convolutional layer structure with different dilation rates, wider context information is captured. At the same time, a DCAM module (convolutional attention module) is designed. By dynamically adjusting the convolution kernel size and using convolution kernels of multiple scales to fuse multi-scale features, during each upsampling process, the features are input into the DCAM module to strengthen feature extraction, thus constituting the improved MY-DCF network model structure. By introducing only a small number of parameters, the detection accuracy of the lightweight network is improved, and the lightweight of the network structure is realized while ensuring a certain detection accuracy.

[0024] As Figure 3As shown, the MobileNetV3 network structure includes an initial two-dimensional convolution and bneck(208,208,16), which are used for the first-step feature extraction of the input image; and the bneck module is repeatedly used in the subsequent network layers to gradually extract deeper features; the SE module is introduced in the neck structure to adaptively focus on key channel information; finally, on the feature map of the deepest layer (13×13), the structure of the SE module and the NHS activation function is adopted to enhance the non-linear expression ability of the network while maintaining the lightweight design, thus taking into account both computational efficiency and detection accuracy.

[0025] The improved MY-DCF network model integrates MobileNetV3 and YOLOv4. It uses the lightweight MobileNetV3 network to replace the backbone extraction network CSPDarkNet53 in YOLOv4, reducing the number of model parameters and computational volume to achieve the lightweight of the network model and balance speed and accuracy. An SE module is provided in the neck structure of the backbone network MobileNetV3, which can perform compression and excitation operations to enhance the attention to specific feature channels while reducing the attention to irrelevant feature channels, enabling the network to more effectively learn and utilize the information in the input data.

[0026] The basic network structure of the present invention is shown in Table 1. The size of the input image of this network is 416×416×3. First, ordinary convolution operations are performed to convert the input feature map into a feature map with 16 channels. A series of bottleneck (abbreviated as bneck) modules are used in the middle layer, and the number of channels and size of the input feature map of each bottleneck module vary between different levels. Since the SE module will increase the additional computational volume and the number of parameters, it is designed to add the SE module only to some bottleneck layers, mainly used in the layers with more channels and more complex features, as shown in Table 1, to enhance the important information of the feature map; the NHS activation function is used in all bottleneck levels, increasing the linear and non-linear expression abilities. The expressions of the activation function NHS are shown in Eqs. (1) and (2).

[0027] The NHS activation function utilizes the non-linear characteristics of the Swish activation function, enabling it to better capture the complex relationships in the input data and improve the expression ability of the neural network. When =0, the NHS activation function becomes Hard-Swish, which is to use the advantages of Hard Swish in cases where computational efficiency is required. As increases, the NHS activation function gradually approaches Swish, retaining its non-linear characteristics. The NHS activation function can be adjusted by The size is such that the NHS activation function balances computational efficiency and non-linear characteristics, thus better adapting to different deep learning tasks and hardware environments.

[0028] Table 1 Basic network structure of the present invention

[0029] Figure 4 is the structural diagram of depthwise separable convolution. In the standard convolution process, N convolutional kernels are convolved with the input data. In the depthwise separable 3×3 convolution, N convolutional kernels are first used for separate convolutions respectively, and then the 3×3 feature map is convolved with N convolutional kernels, finally achieving the effect of reducing the number of parameters. It uses different convolutional kernels for convolution operations according to the number of input channels, and the adjustment of the output result channels is processed through convolutional kernels with a unit of 1. Depthwise separable convolution reduces the computational amount and the number of parameters to 1 / N of the traditional convolution, thus greatly reducing the complexity of the model and significantly reducing the computational amount and the model size.

[0030] The present invention proposes a depthwise separable convolution (NESC), and the formula is as follows: , , , where NESC() represents the depthwise separable convolution operation, z is the input, DW() is the depth convolution operation, PW() is the pointwise convolution operation, and Evo() is the evolutionary computing module. is the row index, is the column index, is the channel index, is the number of channels, is the weight parameter of the depth convolution part, is the weight parameter of the pointwise convolution part.

[0031] The input feature map first undergoes depth convolution, performing convolution operations on each input channel to generate the feature map after depth convolution. Then, adaptive evolutionary computing is introduced to dynamically adjust the structure of the depth convolution kernel, making the convolution operation more adaptable to the task requirements. Finally, pointwise convolution operations are performed on the output of the depth convolution adjusted by the evolutionary computing. The pointwise convolution integrates the output of the depth convolution to generate the final output feature map features.

[0032] Through evolutionary computation, NESC can automatically adjust the structure of the depth convolution kernel to meet the requirements of specific tasks. This adaptability makes NESC more flexible and adaptable, enabling it to optimize the design of the depth convolution kernel to better capture relevant information in the input features. At the same time, NESC maintains the lightweight design of depthwise separable convolution, which separates the depth convolution and pointwise convolution steps to reduce the computational amount and the number of parameters.

[0033] The Feature Enhance Module (FEM) is as Figure 5 shown. It includes three branches and one bypass pruning. Each branch consists of convolution kernels of different sizes and dilated convolutions. These branches run in parallel and generate outputs of different scales and features. Then, the outputs of each branch are cascaded through the bypass pruning to form the final feature representation.

[0034] Within each branch, each branch is composed of a different number of convolutional layers. After the feature map undergoes convolutional operations of different depths, different levels of semantic information can be obtained. Among them, the bypass pruning (shortcut) is a residual connection used to solve the vanishing gradient and exploding gradient problems in the training of deep networks. It can further reduce the computational burden of the network by removing some specific shortcut connections. In addition, this feature enhancement module uses dilated convolutions, which can increase the receptive field of the feature map without reducing the resolution of the feature map. The feature enhancement module uses dilated convolutions with dilation rates of 1, 2, and 3 respectively, and finally performs a stacking (concat) operation to fuse the subnet results to obtain a feature-enhanced feature map. As the depth of the channel convolutional layer increases, the dilation rate of the dilated convolution also increases, enabling different branches to have different receptive fields, which helps to capture a wider range of context information and enhance the network's feature perception and feature extraction capabilities. After the output feature maps of each branch are cascaded, the semantic information and spatial information in the deep layer of the network can be made more abundant.

[0035] The improved MY-DCF network model of the present invention adopts two attention mechanisms: the first attention mechanism is the SE module in the MobileNetV3 backbone network, and the structure of the SE module is as Figure 6 shown; the second attention mechanism is the DCAM module improved on the PANet network, and the structure of the DCAM module is as Figure 7 shown. The DCAM module is introduced during each upsampling process to strengthen feature extraction.

[0036] As Figure 6 shown, X represents the input data, whose dimension is , representing the height, width, and number of channels of the input features. Through convolutional operation, the processed vector U is obtained, whose dimension is , represents the compression operation of the SE module, represents the excitation operation of the SE module, represents the operation of scaling the features, applying the output of the SE module to the output of the convolutional layer to obtain the final output features .

[0037] The SE module performs compression and excitation operations, through which it can enhance the attention to specific feature channels while reducing the attention to irrelevant feature channels. This mechanism helps to improve the performance and generalization ability of the network, enabling the network to more effectively learn and utilize the information in the input data. The weight vector of the channels and the output expression of the SE module are shown in Equations (9) and (10): , , where y represents the input features; represents the features compressed by global average pooling; represents the output of the SE module; represents the weight vector of the channels; FC() represents the fully connected layer; represents ReLU; represents Sigmoid.

[0038] As Figure 7 shown, the channel attention module and the spatial attention module process the input features in parallel: First, the input features are merged to obtain the feature representation ; Then, the channel attention module and the spatial attention module calculate the weight coefficients in the channel dimension and the spatial dimension respectively; Finally, the attention weights in these two dimensions are multiplied (or fused) element-wise with the original features to obtain the final corrected features. This parallel processing can simultaneously highlight the attention to important channels and the emphasis on key spatial positions, achieving the comprehensive enhancement of multi-dimensional features.

[0039] The DCAM module can more comprehensively capture local and global features by dynamically adjusting the convolutional kernel size, using convolutional kernels of multiple scales, and fusing multi-scale features. And in each upsampling process, the features are input into the DCAM module to strengthen feature extraction. Through adaptive receptive field adjustment and multi-scale feature integration, the DCAM module can comprehensively perceive the information of different scales and levels of the input features, improving the generalization ability of the improved MY-DCM network model; the expressions are shown in Equations (11) to (15): , , , , , where X represents the input feature, represents the size of the convolutional kernel used in the i-th convolutional operation, Conv( ) represents the convolutional operation, and Concat( ) represents the concatenation operation, represents the feature of the th scale of the adaptive convolutional kernel adjustment, i = 1, 2,..., a, and a represents the total number of scales; represents the feature after concatenating multiple scale features, represents Sigmoid, represents the result obtained by multiplying the channel feature map by F, represents the result obtained by multiplying the spatial feature map by , FC( ) represents the fully connected layer, and y( ) represents the feature compressed by global average pooling, represents element-wise multiplication, and F represents the final feature representation.

[0040] Step 3: Input the preprocessed face dataset into the constructed face detection network structure, calculate the loss function and perform training, and retain the optimal model weights.

[0041] Step 4: Input the detected face image into the face detection network model obtained in Step 3, output the position and confidence of the detected face, and evaluate the model performance using indicators such as the Precision-Recall curve, mAP, and FPS. The specific implementation process is as follows: To further evaluate the performance of the face detection network model, accuracy (Accuracy), precision (Precision), recall (Recall), F1 Score (F1 score), number of parameters (Parameter), computational complexity (GFLOPs), and frames per second (FPS) are used as evaluation indicators for the experimental results. Among them, the number of parameters represents the number of network calculation parameters, and FPS represents the number of frames of pictures recognized per second. The calculation formulas for each evaluation indicator are defined as shown in Equations (16) to (20): , , , , , In formulas (16) to (20), FP represents the number of samples that are actually negative but predicted to be positive, TN represents the number of samples that are actually negative and predicted to be negative, TP represents the number of samples that are actually positive and predicted to be positive, FN represents the number of samples that are actually positive but predicted to be negative, M represents the total number of images, and S represents the time taken to process all images during the recognition process.

[0042] Table 2 shows the model parameters of the improved MY-DCM network. The M3-YOLOv4 network replaces the backbone extraction network CSPDarkNet53 in the YOLOv4 network with MobilenetV3 and incorporates the NHS activation function. By dynamically adjusting the parameters, a balance is achieved between computational efficiency and non-linear characteristics. The improved M3-YOLOv4-D network structure is based on the M3-YOLOv4 network, replacing traditional convolutions with the new type of depthwise separable convolution NESC. Then, the SPP module in the improved M3-YOLOv4-D network is removed, and the feature enhancement module FEM is added. Additionally, the DCAM attention module is introduced into the PANet network, which can dynamically adjust the convolution kernel size to adapt to input features of different scales, and then fuse features of multiple scales to form the improved MY-DCF network model.

[0043] Table 2 Comparison of network model parameters

[0044] As can be seen from Table 2, due to the large number of parameters in the original backbone network CSPDarknet53 of YOLOv4, after replacing the backbone network with MobilenetV3, the number of parameters of the M3-YOLOv4 model is significantly reduced, and the model size is reduced to 154.26 MB, which is approximately 37% less than the original model, and the detection rate is significantly increased by nearly twice. When the ordinary convolutions in the feature extraction network (PANet network) are replaced with depthwise separable convolution NESC, the number of parameters of the improved MY-DCF network model is greatly reduced, and the model size is significantly reduced by approximately 81% compared to the original model (YOLOv4), and the detection rate is further improved. After further introducing the FEM feature enhancement module in the MobilenetV3 network and the DCAM attention mechanism in the PANet network, compared with M3-YOLOv4-D, the frame rate decreases and the computational load increases to a certain extent, while the number of parameters and the model size only change slightly, indicating that the introduction of the FEM module and the DCAM module has a small impact on the overall parameters and model of the lightweight network. It can reduce the number of network parameters and compress the network model size while ensuring a certain detection accuracy, achieving the lightweight of the network structure.

[0045] Figure 8Among them, (a) and (b) are respectively the comparison charts of the accuracy and loss curves before and after improving the network. From the results, the accuracy verification curve of the improved MY-DCF network model of the present invention rises steadily with the increase of the number of iterations, and tends to be stable after 100 iterations. Moreover, with a significant reduction in the number of model parameters and computational complexity, it also has a high detection accuracy. At the same time, the improved MY-DCF network shows a stable downward trend in the loss curve, without large fluctuations. After 100 iterations, the curve gradually becomes stable. Compared with the M3-YOLOv4 and M3-YOLOv4-D models before improvement, the oscillation frequency of the loss curve is significantly reduced, and the model can converge quickly during the training process. The performance of each improved network model is shown in Table 3: Table 3 Performance Comparison of Different Improved Network Models

[0046] To further evaluate the detection performance of the improved MY-DCF network model, the improved MY-DCF network model was respectively compared with mainstream recognition deep learning models (YOLOv4, Faster R-CNN, VGG19, AlexNet) on the LFW and WIDER Face face datasets. The detection results of different network models are shown in Table 4. The improved MY-DCF network model of the present invention has a significant reduction in the number of parameters compared with other models, and there is a significant improvement in the detection frame rate. The improved MY-DCF network model has the highest detection accuracy on the LFW and WIDER Face face datasets. Compared with the original network model YOLOv4, the detection accuracy remains basically the same, but the number of parameters is greatly reduced compared with the original network model, and the detection frame rate is greatly improved. The accuracy curves and loss curves of the improved MY-DCF network model and different networks are as Figure 9 shown in (a) and (b) in it.

[0047] Table 4 Detection Results of Different Network Models

[0048] Observe Figure 9It can be found that, compared with other network models, the improved MY-DCF network model is more efficient in detecting face images, and the accuracy curve rises steadily with the increase of the number of iterations. After 100 epochs of training, the improved MY-DCF network model is compared with the experiments of other network models (Faster R-CNN, VGG19, AlexNet), and the accuracy of the improved MY-DCF network model has been significantly improved. There are obvious oscillation phenomena in the loss curves of AlexNet and VGG19 in the loss curve image, which indicates that the optimization of these two models is more difficult during the training process. While the oscillation frequency of the improved MY-DCF network model is smaller, and it shows a downward trend during the oscillation; and the model converges quickly. When the number of iterations is about 60, the loss value is basically stable at about 0.3, and then the network structure tends to converge.

[0049] The present invention proposes a face detection method based on an improved MY-DCF network, which can reduce the number of network parameters and the amount of calculation while ensuring the recognition accuracy, meeting the requirements of real-time detection; In summary, the improved MY-DCF network model has been greatly improved in terms of accuracy and detection speed, the number of model parameters has decreased, and the value of FPS is also in the leading position; Therefore, the overall performance advantage of the improved MY-DCF network model of the present invention is more prominent, and it can meet the requirements of both recognition accuracy and speed at the same time.

[0050] As Figure 10 shown, it can be seen that the improved MY-DCF network model designed by the present invention has achieved good results in face detection. The detection box accurately frames the face in the image, and can adapt to faces under different scales, poses, scenes and lighting conditions, and the detection accuracy is generally high. This indicates that the face detection model designed by the present invention has good generalization ability and robustness, providing a good basis for the further research and application of face recognition, face analysis and other related tasks.

[0051] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and all of them are within the protection scope of the present invention.

Claims

1. A face detection method based on an improved MY-DCF network, characterized in that, It includes the following steps: S1. Perform annotation and preprocessing operations on the obtained face dataset; S2. Construct a face detection network model based on the improved MY-DCF network; S3. Input the preprocessed face dataset into the face detection network model respectively, calculate the loss function and perform training, save the optimal model weights, and obtain the trained face detection network; S4. Input the detected face image into the face detection network model obtained in step S3 for face detection.

2. The face detection method based on the improved MY-DCF network according to claim 1, wherein The operations of annotating and preprocessing the obtained face dataset include the following steps: S11. Use the labelImg image annotation software to annotate the face images. The annotation includes "the position of the face rectangular area" and "face semantic information". In the rectangular area position of the face and the surrounding shoulder area, generate an XML file after annotation; write a program in Python to convert the annotation position coordinates saved in the XML file into a corresponding text file; S12. Preprocess the annotated face dataset, introduce a multi-scale training strategy to scale, crop and normalize the original dataset.

3. The face detection method based on the improved MY-DCF network according to claim 1, wherein, The steps of constructing the face detection network model are as follows: First, replace the backbone extraction network CSPDarkNet53 of YOLOv4 with the MobileNetV3 network. At the same time, during the feature fusion process, replace the ordinary convolution in the PANet network with depthwise separable convolution; finally, the output network uses a target detection head; among them, the neck structure of the MobileNetV3 network uses a two-dimensional convolutional layer for initial feature extraction, and then passes through multiple bneck modules in turn. Some bneck modules are equipped with SE attention modules and NHS activation functions to extract multi-level features of the image; Then, delete the SPP module of the CSPDarkNet53 network and add a feature enhancement module. The feature enhancement module is located on the output side of the backbone network and is connected to the PANet network; In the PANet network, multiple depthwise separable convolutions are stacked on different resolution branches. In each scale branch, a convolutional attention module is inserted after the depthwise separable convolution stacking; an upsampling operation is performed on the (52, 52) scale branch, and upsampling and downsampling operations are performed on the (26, 26) scale branch in turn; Finally, detection heads are deployed on the three scales of (52, 52), (26, 26), and (13, 13) respectively.

4. The face detection method based on the improved MY-DCF network according to claim 3, wherein, The expression of the activation function NHS is: , , Among them, is the input, is the Sigmoid function, is the parameter for controlling the degree of non-linearity, and e is the natural constant.

5. The face detection method based on the improved MY-DCF network according to claim 3, characterized in that, The depthwise separable convolution first uses N convolutional kernels for separate convolutions, and then convolves the 3×3 feature map with N convolutional kernels; and uses different convolutional kernels for convolution operations according to the number of input channels. Finally, the output result channels are adjusted by a convolutional kernel with a unit of 1. The expression of the depthwise separable convolution is as follows: , Among them, NESC() represents the depthwise separable convolution operation, z is the input, DW() is the depthwise convolution operation, PW() is the pointwise convolution operation, and Evo() is the evolutionary computing module. is the row index, is the column index, is the channel index, is the number of channels; is the weight parameter of the depthwise convolution part, is the weight parameter of the pointwise convolution part.

6. The face detection method based on the improved MY-DCF network according to claim 3, wherein The feature enhancement module includes three branches and one bypass pruning. Each branch consists of convolutional kernels of different sizes and dilated convolutions. These branches run in parallel and generate outputs of different scales and features. Then, the outputs of each branch are cascaded through the bypass pruning to form the final feature representation; Each branch is composed of a different number of convolutional layers, and dilated convolutions with dilation rates of 1, 2, and 3 are respectively used. Finally, the operation of stacking and fusing the subnet results is performed to obtain the feature map with enhanced features. Among them, the bypass pruning is a residual connection used to solve the vanishing gradient and exploding gradient problems in the training of deep networks.

7. The face detection method based on the improved MY-DCF network according to claim 3, characterized in that, The expression of the convolutional attention module is as follows: , Wherein, X represents the input feature, represents the convolution kernel size used in the th convolution operation, Conv( ) represents the convolution operation, and Concat( ) represents the concatenation operation, represents the feature of the th scale of the adaptive convolution kernel adjustment, i = 1, 2,..., a, and a represents the total number of scales; represents the feature after concatenating multiple scale features, represents Sigmoid, represents the result obtained by multiplying the channel feature map by F, represents the result obtained by multiplying the spatial feature map by , FC( ) represents the fully connected layer, and y( ) represents the feature compressed by global average pooling, represents element-wise multiplication, and F represents the final feature representation.