Face target detection method based on improved YOLOv8

By introducing improvement methods such as DAttention, AIFI and TSACLoss in the YOLOv8 network, the problem of insufficient detection accuracy of dense small faces in complex scenarios is solved, and higher detection accuracy and lower missed detection rate are achieved.

CN120088835AActive Publication Date: 2025-06-03NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510294889.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-03
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art has problems of poor positioning, missed detection and insufficient detection accuracy of dense small faces in complex scenarios.

Method used

Based on the original YOLOv8 network model, a deformable attention mechanism DAttention is introduced to improve the feature fusion module C2f in the backbone network, using the attention-based intra-scale feature interaction network AIFI instead of the spatial pyramid pooling module SPPF, and a translation-sensitive angle constraint loss function TSACLoss is proposed to optimize the positioning of the detection target.

Benefits of technology

Through these improvements, the accuracy of face object detection is significantly improved, especially the detection performance of dense small faces in complex scenarios, reducing the missed detection rate and missed detection rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088835A_ABST
    Figure CN120088835A_ABST
Patent Text Reader

Abstract

The invention provides a face target detection method based on improved YOLOv8, and the method comprises the steps: carrying out the preprocessing of a face data set, constructing and training an improved YOLOv8 network model, inputting a to-be-detected image into the model, and obtaining a prediction result; the improved YOLOv8 network model comprises a backbone network, a neck network and a head network; the backbone network comprises a convolution module CBS and a feature fusion module C2f of an original model, a deformable attention mechanism DAttention is introduced to improve the feature fusion module C2f in the backbone network, the C2f of the last layer is replaced, an attention-based intra-scale feature interaction network is used to replace a spatial pyramid pooling module, and the feature fusion module C2f of the original model is obtained. And optimizing the bounding box loss function through the angle constraint loss function. According to the invention, the precision of a face target detection task is effectively improved, and especially for dense small face image detection, the problems of missing detection and wrong detection of targets are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and object detection, and in particular, to a face object detection method based on improved YOLOv8. Background Art

[0002] In the field of computer vision, face object detection, as an important research direction, is dedicated to accurately identifying and locating human faces from complex image or video data. With the deep integration of security monitoring, mobile device applications, and face recognition technology, the demand for accurate face recognition and detection has increased sharply in today's era. Currently, the mainstream object detection algorithms in deep learning can be divided into two categories: two-stage detection algorithms and one-stage detection algorithms.

[0003] Two-stage algorithms can usually achieve high detection accuracy, but their detection speed is insufficient. Representative two-stage algorithms include: R-CNN, Faster R-CNN, and Mask R-CNN, etc. Representative one-stage algorithms include: YOLO, SSD, and RetinaNet. One-stage algorithms directly predict objects and bounding boxes in the network, sacrificing some accuracy to speed up the detection. YOLO is an end-to-end detection algorithm, known for its relatively fast detection speed and compact model size. Although its accuracy is not as high as that of two-stage algorithms such as Faster R-CNN, it has great advantages in scenarios with high real-time requirements and limited computing resources.

[0004] Most traditional face detection algorithms use the detection mechanism of object detection boxes. However, if the IoU value is too low, small faces will be judged as background and missed. By designing multi-scale feature fusion networks, constructing attention modules to weaken the interference of background information, and establishing context association models for assistance, etc., the face detection accuracy can be further improved. Although the above algorithms can show high performance in most face object detection tasks, for some dense small-sized faces in complex scenarios, there are still many cases of false detection and missed detection, and the face object detection model needs to be further optimized and improved. Summary of the Invention

[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a face object detection method based on improved YOLOv8 in view of the deficiencies of the prior art, so as to solve the problems such as poor localization, false detection and missed detection, and insufficient detection accuracy of the prior art for dense small face objects in complex scenarios.

[0006] Based on the original YOLOv8 network model, the method of the present invention introduces a deformable attention mechanism DAttention to improve the feature fusion module C2f in the backbone network, making the network more focused on key face information; uses an attention-based intra-scale feature interaction network AIFI to replace the spatial pyramid pooling module SPPF to improve the feature extraction efficiency; and proposes a translation-sensitive angular constraint loss function to optimize the localization of the detection target. The method includes the following steps: Step 1: Use a face dataset. If the face dataset already provides a training set and a validation set, directly select the training set and the validation set for use. Otherwise, divide the training set and the validation set in a ratio of 4:1, convert all the labels of the training set and the validation set into the YOLO format, and all classify them into the face category, that is, the human face category; Step 2: Construct and train the improved YOLOv8 network model, where the improved YOLOv8 network model includes a backbone network, a neck network, and a head network; Input the preprocessed training set images into the backbone network to obtain three feature maps of different sizes; Step 3: Input the three feature maps of different sizes into the neck network, and perform feature fusion in two ways: bottom-up and top-down based on the feature pyramid network FPN and the path aggregation network PAN to obtain the fused feature maps; Step 4: Use the detection head in the head network to predict the fused feature maps, generate the category of the target and the position of the bounding box, and use non-maximum suppression NMS (Non-Maximum Suppression, NMS) to remove overlapping prediction boxes to obtain the final detection result; during the training process, use the validation set for validation synchronously, calculate the model evaluation metrics after each round of training, and select the model weights based on the model evaluation metrics.

[0007] Step 1 includes: For the face dataset, create a folder structure in the YOLO way, put the image data and text labels in the training set and the validation set into the corresponding folders, and write a script to convert the face dataset into the YOLO format, and unify the category number of the target object as face:0.

[0008] Step 2 includes: Step 2.1: Uniformly adjust the image size in the training set to 640×640 (it can also be 480×480 or 800×800. The size of 640×640 can achieve a better balance between computing resource consumption and feature extraction accuracy), check that the channel order of the image is red, green, blue RGB, use the Mosaic data augmentation method to process the image, and then input it into the backbone network; Step 2.2, reconstruct the backbone network through the attention-weighted feature fusion module C2f_DA and the attention-based intrascale feature interaction module AIFI (Attention-based Intrascale Feature Interaction, AIFI). The reconstructed backbone network includes a convolution module, a feature fusion module, an attention-weighted feature fusion module, and an intrascale feature interaction module; There are 6 convolution modules, namely the first convolution module, the second convolution module, the third convolution module, the fourth convolution module, the fifth convolution module, and the sixth convolution module. There are 3 feature fusion modules, namely the first feature fusion module, the second feature fusion module, and the third feature fusion module; Among them, the first convolution module, the second convolution module, and the first feature fusion module are serially connected to form a basic feature extraction link. Subsequently, the third convolution module and the second feature fusion module are serially connected, and are repeatedly stacked via the fourth convolution module and the third feature fusion module to achieve step-by-step downsampling. The fifth convolution module is serially connected to the attention-weighted feature fusion module, and the output is compressed in channels by the sixth convolution module and input to the intrascale feature interaction module. The intrascale feature interaction module enhances the global semantics through multi-head self-attention and finally performs a residual connection with the output of the attention-weighted feature fusion module. Among them, the second feature fusion module and the third feature fusion module directly transfer the shallow high-resolution features to the deep neck network through cross-layer skip connections; The 6 convolution modules are used for basic feature extraction and channel dimension adjustment; The 3 feature fusion modules are used for multi-scale feature fusion; The attention-weighted feature fusion module adopts the attention heads of the deformable attention mechanism DAttention. The number of attention heads is defaulted to 8, which capture different subspace features respectively. The output formula is: , Among them, represents the output of the m-th attention head, represents the softmax function, represents the embedding for querying the m-th attention head, represents the deformed key-value, represents the dimension of each attention head, represents the sampling function, is the sampling grid coordinate, is the dynamic offset, represents the deformed value; In addition, a fixed-position encoding table is used to explicitly learn the spatial relationship between the query and the key-value, making up for the long-distance dependencies that may be ignored by the dynamic offset. The fixed-position encoding table follows a normal distribution with an expected value of 0 and a variance of 0.01 2 and satisfies: , where represents the fixed-position encoding table, is a three-dimensional tensor, is the length of the query sequence, is the length of the key-value sequence; The attention-weighted feature fusion module first preliminarily processes the input data through convolution, then splits the data and sends it into more than two residual blocks incorporating the deformable attention mechanism DAttention to achieve attention weighting and feature fusion at different levels. Finally, the outputs of the residual blocks are concatenated, which helps to solve the problem of dense small face feature confusion and high missed detection rate caused by the fixed sampling grid in the feature fusion module; Among them, a 1×1 sixth convolution module is added after the attention-weighted feature fusion module to adapt to the in-scale feature interaction module; the size of the feature map output by the attention-weighted feature fusion module is 20×20×1024. The number of channels is compressed to 256 by the 1×1 sixth convolution module. The feature map with a size of 20×20×256 is flattened by the in-scale feature interaction module to become 400×256. The spatial information is enhanced through two-dimensional sine-cosine position encoding, and the interaction features are learned using the multi-head self-attention mechanism and the feed-forward network. Finally, the size of the output feature map is restored to 20×20×1024; the attention-weighted feature fusion module and the in-scale feature interaction module are cascaded to achieve two-stage feature extraction from local deformation perception to global semantic enhancement; Step 2.3, Extract the feature information of the fifth, seventh, and eleventh layers of the backbone network to obtain three feature maps with different sizes, and the sizes are 80×80, 40×40, and 20×20 respectively.

[0009] Step 3 includes: Step 3.1, The neck network includes a feature fusion module and a convolution module; There are 4 feature fusion modules in the neck network, namely the fourth feature fusion module, the fifth feature fusion module, the sixth feature fusion module, and the seventh feature fusion module; there are 2 convolution modules in the neck network, namely the seventh convolution module and the eighth convolution module; The fourth and fifth feature fusion modules are used for the Feature Pyramid Network (FPN). Starting from the 20×20 feature map, through upsampling and cross-layer connection with the third feature fusion module, the result is output to the fourth feature fusion module. After upsampling, the fourth feature fusion module is cross-layer connected with the second feature fusion module, and the result is output to the fifth feature fusion module, completing the bottom-up feature fusion and enhancing the semantic information. The sixth and seventh feature fusion modules are used for the Path Aggregation Network (PAN). Starting from the 80×80 feature map, through downsampling by the seventh convolution module and cross-layer connection with the fourth feature fusion module, the result is output to the sixth feature fusion module. After downsampling by the eighth convolution module, the sixth feature fusion module is cross-layer connected with the in-scale feature interaction module, and the result is output to the seventh feature fusion module, completing the top-down feature fusion and refining the localization. The two convolution modules are used for feature dimension alignment. Step 3.2: During the fusion process of the four feature fusion modules, feature maps with different resolutions undergo convolution, upsampling, downsampling, and element-wise addition more than twice, and finally form three optimized feature maps with sizes of 80×80, 40×40, and 20×20, and the optimized feature maps are output.

[0010] Step 4 includes: Step 4.1: The head network includes convolution modules and convolutional layers. There are 6 convolution modules and 6 convolutional layers in the head network. The 6 convolution modules and the 6 convolutional layers correspond one by one and are serially connected. The convolution modules and convolutional layers are used to generate bounding box predictions and classification predictions. Step 4.2: According to a predefined threshold (usually 0.25), filter out the prediction boxes in the prediction feature map with a confidence level less than the threshold; then use Non-Maximum Suppression (NMS) to remove redundant prediction boxes and retain the prediction boxes most likely to contain the target object. The formula for NMS is: , where, represents the score of the i-th prediction box, represents the prediction box and the i-th prediction box the intersection over union between them, is the prediction box with the highest score, represents the i-th prediction box, represents the threshold, usually 0.7; i takes values from 1 to K; represents the number of prediction boxes in the current batch; Step 4.3, after obtaining the prediction boxes screened by non-maximum suppression (NMS), calculate the loss value between the prediction boxes and the ground truth boxes. The loss value includes object classification loss and bounding box regression loss. Among them, an angle-constrained intersection over union loss (TAIoU) for cross-space translation is designed to optimize the bounding box regression loss. The formula for TAIoU is: , The TAIoU includes a translation-sensitive angle-constrained loss function (TSACLoss, Translation-Sensitive Angular Constraint Loss) and an intersection over union metric (IoU, Intersection over Union). Among them, the formula for the translation-sensitive angle-constrained loss function TSACLoss is: , , , Among them, represents the angle loss, represents the width difference between the prediction box and the ground truth box, and represent the upper left vertex and the lower left vertex of the ground truth box respectively, and represent the upper left vertex and the lower left vertex of the prediction box respectively, is the vertical distance difference between point and point is the vertical distance difference between point and point represents the vector from point to point after translation by distance d, represents the vector from point to point after translation by distance d, represents the vector from point to point after translation by distance d, is the width of the ground truth box, and TSACLoss is the angle-constrained loss value; The penalizes the vertical offset between the prediction box and the ground truth box to ensure the alignment of the object, further constrains the width consistency. The key of TSACLoss is to break through the single-image space constraint, introduce a translation matching space, and determine the specific translation distance through experiments to achieve a more accurate loss value calculation.

[0011] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method described above.

[0012] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction runs on a computer, the steps of the method described above are executed.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention optimizes the YOLOv8 network structure. Aiming at the deficiency of YOLOv8 in detecting dense small faces, DAttention is introduced to improve the feature fusion module C2f in the backbone network. By adjusting the attention weights, the subtle features of small targets can be better captured, enabling the network to achieve precise focus; AIFI is used to replace SPPF to improve the efficiency and pertinence of feature extraction; 2. The present invention proposes a translation-sensitive angular constraint loss function TSACLoss, and combines it with IoU to obtain the TAIoU bounding box loss function to replace the CIoU of the original model; TSACLoss breaks the tradition of measuring the overlap degree of bounding boxes from the same image space, uses the translation distance as a constraint, and considers the geometric angular relationship of the vertices of the bounding box, making up for the pain points of IoU. Compared with CIoU, TAIoU can optimize the positioning of the prediction box, improve the quality of the box, and thus improve the detection accuracy of the network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a structural diagram of the improved YOLOv8 network model of the present invention.

[0015] Figure 2 It is a structural diagram of the attention-weighted feature fusion module in the backbone network of the present invention.

[0016] Figure 3 It is a structural diagram of the in-dimension feature interaction module in the backbone network of the present invention.

[0017] Figure 4 It is a schematic diagram of the principle of the translation-sensitive angular constraint loss function proposed by the present invention.

[0018] Figure 5 It is a curve comparison diagram of the original model and various improved modules of the present invention in terms of evaluation indexes during the training process. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following further specific descriptions of the present invention are made in conjunction with the drawings and specific embodiments, and the above or other advantages of the present invention will become clearer.

[0020] This embodiment provides a face target detection method based on improved YOLOv8. On the basis of the original YOLOv8 network model, the DAttention is introduced to improve the feature fusion module C2f in the backbone network, making the network more focused on key face information; the AIFI is used to replace the SPPF to improve the feature extraction efficiency; the TSACLoss is proposed to optimize the localization of the detection target by the model. The overall network model of the improved YOLOv8 is as shown in Figure 1 shown. The input image passes through the backbone network, the neck network and the head network in sequence, and the detection result is output; the structure of the attention-weighted feature fusion module (C2f_DA) is as shown in Figure 2 shown. First, the input data is preliminarily processed by the convolution module, then the data is segmented and sent into multiple attention-weighted residual blocks, and finally the outputs are concatenated; the intra-scale feature interaction module (AIFI) is as shown in Figure 3 shown. It mainly adopts two-dimensional sine-cosine position encoding and multi-head attention mechanism; the principle of the translation-sensitive angular constraint loss function (TSACLoss) is as shown in Figure 4 shown. It is not limited to the original image space, and the translation amount is used to control the optimization effect of the angular loss on the target localization; the curves of various evaluation indicators in the training process are as shown in Figure 5 shown.

[0021] The overall network model of the improved YOLOv8 is built and trained based on the Pytorch framework. The WIDER FACE dataset with higher difficulty is used as the dataset, the SGD is used as the optimizer, the initial learning rate is 0.01, the momentum optimization parameter is set to 0.005, the training batch size is 8, and a total of 250 rounds of training are performed. The method includes the following steps: Step 1: Use the publicly available large-scale face dataset WIDER FACE. WIDER FACE provides a training set containing 12,880 pictures and a test set containing 3,226 pictures. Convert all the labels of the training set and the validation set into the YOLO format and all belong to the face category, that is, the face class. The specific steps are as follows: Step 1.1: For WIDER FACE, create a folder structure in the YOLO way, put the image data and text labels in the training set and the validation set into the corresponding folders, and write a script to convert the face dataset into the YOLO format, and unify the category number of the target object as face:0.

[0022] Step 2: Build and train the improved YOLOv8 network model. The improved YOLOv8 network model includes a backbone network, a neck network and a head network; the training set images are preprocessed and then input into the backbone network to obtain three feature maps of different sizes. The specific steps are as follows: Step 2.1, uniformly adjust the image size in the training set to 640×640, check that the channel order of the image is RGB (Red, Green, Blue), process the image using the Mosaic data augmentation method, and then input it into the backbone network; Step 2.2, reconstruct the backbone network through the attention-weighted feature fusion module C2f_DA and the attention-based intra-scale feature interaction module AIFI. The reconstructed backbone network includes a convolution module, a feature fusion module, an attention-weighted feature fusion module, and an intra-scale feature interaction module; There are 6 convolution modules, namely the first convolution module, the second convolution module, the third convolution module, the fourth convolution module, the fifth convolution module, and the sixth convolution module. There are 3 feature fusion modules, namely the first feature fusion module, the second feature fusion module, and the third feature fusion module; Among them, the first convolution module, the second convolution module, and the first feature fusion module are connected in series to form a basic feature extraction link. Subsequently, the third convolution module and the second feature fusion module are connected in series, and are repeatedly stacked via the fourth convolution module and the third feature fusion module to achieve step-by-step downsampling. The fifth convolution module is connected in series with the attention-weighted feature fusion module, and its output is compressed in channels by the sixth convolution module and input into the intra-scale feature interaction module. The intra-scale feature interaction module enhances the global semantics through multi-head self-attention and finally performs a residual connection with the output of the attention-weighted feature fusion module. Among them, the second feature fusion module and the third feature fusion module directly transfer the shallow high-resolution features to the deep neck network through cross-layer skip connections; The 6 convolution modules are used for basic feature extraction and channel dimension adjustment; The 3 feature fusion modules are used for multi-scale feature fusion; The attention-weighted feature fusion module uses the attention head of the deformable attention mechanism DAttention. The number of attention heads is defaulted to 8, which capture different subspace features respectively. The output formula is: , Among them, represents the output of the m-th attention head, represents the softmax function, represents the embedding for querying the m-th attention head, represents the deformed key-value, represents the dimension of each attention head, represents the sampling function, is the sampling grid coordinate, is the dynamic offset, represents the deformed value; In addition, a fixed-position encoding table is used to explicitly learn the spatial relationship between the query and key values, compensating for the long-distance dependencies that may be ignored by the dynamic offsets. The fixed-position encoding table follows a normal distribution with an expected value of 0 and a variance of 0.01 2 and satisfies: , where represents the fixed-position encoding table, is a three-dimensional tensor, 8 corresponds to the number of attention heads, is the query sequence length, is the key-value sequence length; The attention-weighted feature fusion module first preliminarily processes the input data through convolution, then splits the data and sends it into two or more residual blocks incorporating the deformable attention mechanism DAttention to achieve attention weighting and feature fusion at different levels. Finally, the outputs of the residual blocks are concatenated, which helps to solve the problem of dense small face feature confusion and high miss detection rate caused by the fixed sampling grid in the feature fusion module; Among them, a sixth convolutional module with a size of 1×1 is added after the attention-weighted feature fusion module to adapt to the in-scale feature interaction module; the size of the feature map output by the attention-weighted feature fusion module is 20×20×1024. The number of channels is compressed to 256 by the sixth convolutional module with a size of 1×1. The feature map with a size of 20×20×256 is flattened by the in-scale feature interaction module to become 400×256. The spatial information is enhanced through two-dimensional sine-cosine position encoding, and the interaction features are learned using the multi-head self-attention mechanism and the feed-forward network. Finally, the size of the output feature map is restored to 20×20×1024; the attention-weighted feature fusion module is cascaded with the in-scale feature interaction module to achieve two-stage feature extraction from local deformation perception to global semantic enhancement; Step 2.3, Extract the feature information of the fifth, seventh, and eleventh layers of the backbone network to obtain three feature maps with different sizes, and the sizes are 80×80, 40×40, and 20×20 respectively.

[0023] Step 3, Input the three feature maps with different sizes into the neck network, and based on the Feature Pyramid Network FPN and the Path Aggregation Network PAN, perform feature fusion in two ways: bottom-up and top-down, to obtain the fused feature map. The specific steps are as follows: Step 3.1, The neck network includes a feature fusion module and a convolutional module; There are 4 feature fusion modules, namely the fourth feature fusion module, the fifth feature fusion module, the sixth feature fusion module, and the seventh feature fusion module; there are 2 convolutional modules, namely the seventh convolutional module and the eighth convolutional module; The fourth and fifth feature fusion modules are used for the Feature Pyramid Network (FPN). Starting from the 20×20 feature map, through upsampling and cross-layer connection with the third feature fusion module, the result is output to the fourth feature fusion module. After upsampling, the fourth feature fusion module is cross-layer connected with the second feature fusion module, and the result is output to the fifth feature fusion module to complete the bottom-up feature fusion and enhance the semantic information. The sixth and seventh feature fusion modules are used for the Path Aggregation Network (PAN). Starting from the 80×80 feature map, downsampling is achieved through the seventh convolutional module and cross-layer connection with the fourth feature fusion module. The result is output to the sixth feature fusion module. The sixth feature fusion module achieves downsampling through the eighth convolutional module and cross-layer connection with the intra-scale feature interaction module. The result is output to the seventh feature fusion module to complete the top-down feature fusion and refine the localization. The 2 convolutional modules are used for feature dimension alignment. Step 3.2, during the fusion process of the 4 feature fusion modules, feature maps with different resolutions undergo convolution, upsampling, downsampling, and element-wise addition more than twice, and finally form optimized feature maps of three sizes: 80×80, 40×40, and 20×20, and the optimized feature maps are output.

[0024] Step 4, use the detection head in the head network to predict the fused feature map, generate the category of the target and the position of the bounding box, and use Non-Maximum Suppression (NMS) to remove overlapping prediction boxes to obtain the final detection result. During the training process, the validation set is used for validation synchronously. After each round of training, the model evaluation metrics are calculated, and the model weights are selected based on the model evaluation metrics. The specific steps are as follows: Step 4.1, the head network includes convolutional modules and convolutional layers. There are 6 convolutional modules and 6 convolutional layers. The 6 convolutional modules and the 6 convolutional layers correspond one by one and are connected in series. The convolutional modules and convolutional layers are used to generate bounding box predictions and classification predictions. Step 4.2, according to a predefined threshold (usually 0.25), filter out the prediction boxes with confidence less than the threshold in the predicted feature map; then use Non-Maximum Suppression (NMS) to remove redundant prediction boxes and retain the prediction box most likely to contain the target object. The formula for NMS is: , where represents the score of the i-th prediction box, represents the intersection over union between the prediction box and the i-th prediction box , is the prediction box with the highest score, represents the i-th prediction box, Denote the threshold, generally 0.7; i takes values from 1 to K; Denote the number of predicted bounding boxes in the current batch; In step 4.3, after obtaining the predicted bounding boxes screened by non-maximum suppression (NMS), calculate the loss value between the predicted bounding boxes and the ground truth bounding boxes. The loss value includes object classification loss and bounding box regression loss. Among them, an angle-constrained intersection over union loss (TAIoU) for cross-space translation is designed to optimize the bounding box regression loss. The formula for TAIoU is: , The TAIoU includes a translation-sensitive angle-constrained loss function and an intersection over union metric (IoU). Among them, the formula for the translation-sensitive angle-constrained loss function (TSACLoss) is: , , , Among them, Denote the angle loss, Denote the width difference between the predicted bounding box and the ground truth bounding box, And Denote the top-left vertex and the bottom-left vertex of the ground truth bounding box respectively, And Denote the top-left vertex and the bottom-left vertex of the predicted bounding box respectively, Is The vertical distance difference between point And point Is The vertical distance difference between point And point Denote the vector from point To point After a translation distance d, Denote the vector from point To point After a translation distance d, Denote the vector from point To point After a translation distance d, Is the width of the ground truth bounding box, and TSACLoss is the angle-constrained loss value; The Penalizes the vertical offset between the predicted bounding box and the ground truth bounding box to ensure the alignment of the object, Then further constrains the width consistency. The key of TSACLoss is to break through the single-image space constraint, introduce a translation matching space, and determine the specific translation distance through experiments to achieve a more accurate loss value calculation.

[0025] The trained model is evaluated on the Easy, Medium, and Hard subsets of the WIDER FACE dataset. The evaluation metrics used in this invention are: accuracy P, recall R, mean average precision 50, and mean average precision 50-95; among them, the precision performance on the subset specifically refers to mean average precision 50.

[0026] Accuracy refers to the proportion of samples that are truly positive among those predicted as positive by the model. The formula is as follows: , Recall refers to the proportion of samples that are actually positive and are correctly predicted as positive by the model. The formula is as follows: , Mean average precision 50 refers to the mean average precision when the intersection over union (IoU) threshold is 0.5, while mean average precision 50-95 is the mean average precision with the IoU ranging from 0.5 to 0.95 in steps of 0.05.

[0027] Figure 5 Shows the comparison of training metrics between the baseline network model and each improved model. Among them, except for C2f_DA, other improved modules can significantly exceed the original model in terms of accuracy, and there are also slight improvements in other metrics; while the comprehensive improved model has achieved the best results in terms of accuracy, recall, mean average precision 50, and mean average precision 50-95.

[0028] Table 1

[0029]

[0030] To better reflect the performance of the improved YOLOv8 network model, the specific performances of the original model and the improved model on different difficulty subsets are listed below, as shown in Table 1.

[0031] Table 1 lists the accuracy performance, number of parameters, and computational complexity (floating-point operations per second) of different models. Compared with the baseline network model, the improved C2f_DA module, AIFI module, and TSACLoss loss function can all achieve better performance on the subsets of simple and medium difficulties; except for the C2f_DA module, other improvements can increase the accuracy value on the subset of difficult difficulty; among them, the TSACLoss loss function performs the best. Generally speaking, the improved YOLOv8 model of the present invention achieves accuracy performances of 93.4%, 91.4%, and 81.0% on the three subsets of different difficulties respectively, which are 0.8%, 0.6%, and 0.7% higher than the original network model respectively. Although the number of parameters of the improved model has increased, the actual computational complexity has decreased instead, and the consumed computing resources have been reduced.

[0032] The experiments of the embodiments show that the technical solution of the present invention performs well on the relatively difficult face detection dataset WIDER FACE, effectively improves the accuracy of face target detection, and provides a feasible solution for solving the problem of detecting dense small faces in complex environments.

[0033] The present invention provides a face target detection method based on improved YOLOv8. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A face target detection method based on improved YOLOv8, characterized in that: The following steps are involved: Step 1: Use a face dataset. If the face dataset has provided a training set and a validation set, directly select the training set and the validation set for use. Otherwise, divide the training set and the validation set in proportion, convert all labels of the training set and the validation set into the YOLO format, and attribute all of them to the face category, i.e., the face category. Step 2, constructing and training an improved YOLOv8 network model, wherein the improved YOLOv8 network model includes a backbone network, a neck network, and a head network; The training set images are preprocessed and then input into the backbone network to obtain feature maps of three different sizes; Step 3, inputting the feature maps of the three different sizes into the neck network, and performing feature fusion in a bottom-up and top-down manner based on the feature pyramid network FPN and the path aggregation network PAN to obtain a fused feature map; Step 4: Use the detection head in the head network to predict the fused feature map, generate the target category and the position of the bounding box, and use non-maximum suppression (NMS) to remove overlapping prediction boxes to obtain the final detection result. During the training process, the validation set is used for verification. After each round of training, the model evaluation index is calculated, and the model weight is selected based on the model evaluation index.

2. The method according to claim 1, characterized in that Step 1 includes: creating a folder structure for the face dataset in the YOLO way, putting the image data and text labels in the training set and validation set into the corresponding folders, writing a script to convert the face dataset into the YOLO format, and unifying the category number of the target object as face:

0.

3. The method according to claim 2, characterized in that Step 2 includes: Step 2.1, uniformly adjust the size of the images in the training set, check that the channel order of the images is red, green, and blue (RGB), use the Mosaic data enhancement method to process the images, and then input them into the backbone network; Step 2.2, reconstruct the backbone network through the attention-weighted feature fusion module C2f_DA and the attention-based intra-scale feature interaction module AIFI. The reconstructed backbone network includes a convolution module, a feature fusion module, an attention-weighted feature fusion module, and an intra-scale feature interaction module; There are 6 convolution modules, namely, a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a fifth convolution module and a sixth convolution module; there are 3 feature fusion modules, namely, a first feature fusion module, a second feature fusion module and a third feature fusion module; Among them, the first convolution module, the second convolution module and the first feature fusion module are connected in series to form a basic feature extraction link, and then the third convolution module and the second feature fusion module are connected in series, and are repeatedly stacked through the fourth convolution module and the third feature fusion module to achieve step-by-step downsampling; the fifth convolution module is connected in series with the attention weighted feature fusion module, and the output is compressed by the sixth convolution module and input to the intra-scale feature interaction module. The intra-scale feature interaction module enhances the global semantics through multi-head self-attention, and finally performs a residual connection with the output of the attention weighted feature fusion module; Among them, the second feature fusion module and the third feature fusion module directly transfer the shallow high-resolution features to the deep neck network through cross-layer jump connections; The six convolution modules are used for basic feature extraction and channel dimension adjustment; The three feature fusion modules are used for multi-scale feature fusion; The attention weighted feature fusion module adopts the attention head of the deformable attention mechanism DAttention, and the output formula is: , in, represents the output of the mth attention head, represents the softmax function, represents the embedding of the query m-th attention head, Indicates the transformed key value. represents the dimension of each attention head, represents the sampling function, are the sampling grid coordinates, is the dynamic offset, Indicates the value after deformation; Use a fixed position encoding table to explicitly learn the spatial relationship between the query and the key value. The fixed position encoding table has an expectation of 0 and a variance of 0.

01. 2 The normal distribution satisfies: , in, represents a fixed position encoding table, is a three-dimensional tensor, is the query sequence length, is the length of the key value sequence; The attention weighted feature fusion module first preliminarily processes the input data through convolution, then divides the data and sends it to more than two residual blocks that incorporate the deformable attention mechanism DAttention to achieve different levels of attention weighting and feature fusion, and finally splices the output of the residual blocks; Among them, a 1×1 sixth convolution module is added after the attention weighted feature fusion module to adapt to the intra-scale feature interaction module; the feature map size output by the attention weighted feature fusion module is 20×20×1024, and the number of channels is compressed to 256 by the 1×1 sixth convolution module. The feature map with a size of 20×20×256 is flattened by the intra-scale feature interaction module to 400×256, and the spatial information is enhanced by two-dimensional sine-cosine position encoding. The multi-head self-attention mechanism and feedforward network are used to learn the interactive features, and finally the output feature map size is restored to 20×20×1024; the attention weighted feature fusion module is cascaded with the intra-scale feature interaction module to realize the two-stage feature extraction from local deformation perception to global semantic enhancement; Step 2.3, extract the feature information of the fifth, seventh, and eleventh layers of the backbone network to obtain feature maps of three different sizes, with sizes of 80×80, 40×40, and 20×20 respectively.

4. The method according to claim 3, characterized in that Step 3 includes: Step 3.1, the neck network includes a feature fusion module and a convolution module; The neck network has four feature fusion modules, namely the fourth feature fusion module, the fifth feature fusion module, the sixth feature fusion module and the seventh feature fusion module; the neck network has two convolution modules, namely the seventh convolution module and the eighth convolution module; The fourth feature fusion module and the fifth feature fusion module are used for feature pyramid network FPN, starting from the 20×20 feature map, after upsampling, they are connected to the third feature fusion module across layers, and the result is output to the fourth feature fusion module. After upsampling, the fourth feature fusion module is connected to the second feature fusion module across layers, and the result is output to the fifth feature fusion module, completing the bottom-up feature fusion and enhancing the semantic information; The sixth feature fusion module and the seventh feature fusion module are used for the path aggregation network PAN, starting from the 80×80 feature map, downsampling is achieved through the seventh convolution module, and the fourth feature fusion module is cross-layer connected, and the result is output to the sixth feature fusion module. The sixth feature fusion module is downsampled through the eighth convolution module, and the scale feature interaction module is cross-layer connected, and the result is output to the seventh feature fusion module, completing the top-down feature fusion and refining the positioning; The two convolution modules are used for feature dimension alignment; In step 3.2, during the fusion process of the four feature fusion modules, feature maps of different resolutions undergo more than two convolutions, upsampling, downsampling, and element-by-element addition, and finally form optimized feature maps of three sizes: 80×80, 40×40, and 20×20, and output the optimized feature maps.

5. The method according to claim 4, characterized in that Step 4 includes: Step 4.1, the head network includes a convolution module and a convolution layer; The head network has 6 convolution modules and 6 convolution layers, and the 6 convolution modules correspond to the 6 convolution layers one by one and are connected in series; The convolutional module and convolutional layer are used to generate bounding box predictions and classification predictions; Step 4.2: Filter out the prediction boxes whose confidence is less than the threshold in the prediction feature map; then use non-maximum suppression (NMS) to remove redundant prediction boxes and retain the prediction boxes that are most likely to contain the target object. The formula of NMS is: , in, represents the score of the i-th prediction box, Represents the prediction box With the i-th prediction box The intersection ratio between is the predicted box with the highest score, represents the i-th prediction box, Represents the threshold; i ranges from 1 to K; Indicates the number of prediction boxes in the current batch; Step 4.3, after obtaining the predicted box filtered by non-maximum suppression NMS, calculate the loss value between the predicted box and the true box, which includes the target classification loss and the bounding box regression loss; among them, the angle-constrained intersection-over-union loss TAIoU across spatial translation is designed to optimize the bounding box regression loss. The formula of TAIoU is: , The TAIoU includes a translation-sensitive angle constraint loss function TSACLoss and an intersection-over-union index IoU; wherein the formula of the translation-sensitive angle constraint loss function TSACLoss is: , , , in, represents the angle loss, Indicates the width difference between the predicted box and the real box, and Respectively represent the vertex in the upper left corner and the vertex in the lower left corner of the real box, and Respectively represent the vertices of the upper left corner and the lower left corner of the prediction box, yes Point and The vertical distance difference of the point, yes Point and The vertical distance difference of the point, After the translation distance d, Click to The vector of the point, After the translation distance d, Click to The vector of the point, After the translation distance d, Click to The vector of the point, is the width of the true box, and TSACLoss is the angle constraint loss value.

6. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 5.

7. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 5 are executed.

Citation Information

Patent Citations

  • Pedestrian small target detection method based on improved YOLOv8

    CN118865444A

  • Limestone granularity detection method and equipment based on YOLO-ADM and storage medium

    CN119359735A

  • CNN and Transform-based three-dimensional coronary artery segmentation method and system

    CN119477942A

Cited By

  • Tunnel anomaly detection model establishment and detection early warning method

    CN120388294A

  • Tunnel anomaly detection model establishment and detection and early warning method

    CN120388294B

  • Cross-scale space target detection method and device

    CN120526129A

  • A cross-scale space target detection method and device

    CN120526129B

  • Cardiovascular atherosclerosis detection system based on data analysis

    CN121236417A