A face detection method based on DETR

Through the DETR-based face detection method, combined with the WiderFace data set and feature pyramid network, the problem of poor face detection effect in complex scenarios is solved, and face detection with high robustness and portability is achieved.

CN114926882BActive Publication Date: 2025-05-13SHENZHEN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210562807.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-23
Publication Date
2025-05-13
Estimated Expiration
2042-05-23

AI Technical Summary

Technical Problem

The existing face detection algorithm is not effective in complex scenarios, requiring a lot of manual setup of anchors, relying on prior knowledge, resulting in poor robustness and portability.

Method used

Using DETR-based face detection method, data preprocessing is performed by selecting the WiderFace dataset and Data Anchor Sample sampling method, combining the ResNet-50 and Transformer models, a feature pyramid network is introduced for feature fusion, reducing manual settings and hyperparameters.

Benefits of technology

It improves the robustness and portability of face detection, can better detect small-scale faces, reduces computing complexity and memory usage, and realizes end-to-end face detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926882B_ABST
    Figure CN114926882B_ABST
Patent Text Reader

Abstract

The invention discloses a face detection method based on DETR, comprising the following steps: selecting a WiderFace data set; adopting a DataAnchor Sample sampling method to randomly crop and randomly scale pictures in the WiderFace data set to obtain preprocessed pictures; adopting ResNet-50 as a backbone network to extract features of the preprocessed pictures, selecting picture feature maps of different layers and inputting them into a Transformer model; selecting a feature pyramid network to fuse the picture feature maps, and outputting a fused feature map; a decoder in the Transformer model decodes the fused feature map to obtain an optimized face detection model; adopting a training set in the WiderFace data set to train the optimized face detection model to obtain a final face detection model; and performing face detection according to the obtained final face detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a face detection method based on DETR. Background Art

[0002] Face recognition is a hot area in deep learning today and is widely used, such as access control, subway, airport and other entrances and exits; or community security, supermarket marketing and other life scenarios. Face detection refers to computer technology that can detect faces in pictures or videos. It is the basis of algorithms such as face recognition, liveness detection, and face positioning.

[0003] Although the face detection effect in simple scenes is quite good, it is still challenging to detect faces in complex scenes. Complex scenes are more common in life, such as dense crowds, occlusion, interference under different light sources, and faces at a distance. These challenges make it difficult to further apply face detection to life scenes. At present, most face detection algorithms are based on one-stage convolutional neural networks, but most of these face detection algorithms based on one-stage convolutional neural networks need to rely on manual setting of anchors according to the target. Researchers need to have a lot of prior knowledge and know the size and proportion of the target to set the corresponding anchors. The anchor design is directly related to the prediction results. Once the anchor setting is biased, the detection effect will be reduced. In addition, using a large number of anchors for prediction will cause uneven matching of positive and negative samples, which needs to be solved based on a lot of prior knowledge of researchers. Non-network learning has poor robustness and portability. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a face detection method based on DETR, so as to directly perform face detection by acquiring a model, reduce manual settings and improve the detection effect.

[0005] In order to solve the above technical problems, the object of the present invention is to achieve the following technical solutions: provide a face detection method based on DETR, comprising the following steps:

[0006] Dataset selection: Select the WiderFace dataset;

[0007] Data preprocessing: The DataAnchor Sample sampling method is used to randomly crop and randomly scale the images in the WiderFace dataset to obtain the preprocessed images;

[0008] Image feature map extraction: ResNet-50 is used as the backbone network to extract features from the preprocessed images, obtain image feature maps, and select image feature maps at different layers to input into the Transformer model;

[0009] Feature fusion: In the Transformer model, a feature pyramid network (FPN) is used to fuse the feature maps of the input images at different layers, and the fused feature maps are output to the decoder in the Transformer model.

[0010] Model acquisition: The decoder in the Transformer model decodes the fused feature map to obtain an optimized face detection model;

[0011] Model training: Use the training set in the WiderFace dataset to train the optimized face detection model to obtain the final face detection model;

[0012] Face detection: Perform face detection based on the final face detection model obtained.

[0013] Its further technical solution is: the step of data preprocessing specifically includes:

[0014] Randomly select a face from the image in the WiderFace dataset and crop it, get the size of the face after cropping the image, and get the anchor that is closest to the size of the face;

[0015] According to the index of the anchor closest to the face size, a value smaller than the index is randomly selected as the index of the selected anchor;

[0016] The size ratio between the size of the selected anchor and the size of the face is calculated, and a ratio is randomly selected between the range of the size ratio being halved and the size ratio being doubled as the image scaling ratio. The original image before cropping the face is scaled according to the image scaling ratio to obtain the preprocessed image.

[0017] Its further technical solution is: in the step of extracting the picture feature map, the step of selecting the picture feature maps of different layers and inputting them into the Transformer model is specifically: selecting the picture feature maps of the 3rd, 4th and 5th layers and inputting them into the Transformer model.

[0018] The further technical solution is: the step of feature fusion specifically includes:

[0019] The number of channels of the output image feature map of each layer of the backbone network is unified to 256 through linear mapping;

[0020] The bilinear interpolation algorithm is used to upsample the image feature map of the 5th layer; the image feature map of the 4th layer is added and fused with the upsampled image feature map of the 5th layer through a 1×1 ordinary convolution, and then updated through a 3×3 channel-by-channel convolution and a 1×1 ordinary convolution to obtain a new image feature map of the 4th layer; the bilinear interpolation algorithm is used to upsample the new image feature map of the 4th layer; the image feature map of the 3rd layer is added and fused with the upsampled new image feature map of the 4th layer through a 1×1 ordinary convolution, and then updated through a 3×3 channel-by-channel convolution and a 1×1 ordinary convolution to obtain a new image feature map of the 3rd layer as a multi-scale feature map after the fusion of the image feature maps of the 3rd, 4th and 5th layers;

[0021] The multi-scale feature maps are flattened and concatenated in the spatial dimension to obtain a single-layer two-dimensional feature map as a fused feature map, which is then output to the decoder in the Transformer model.

[0022] A further technical solution is that the decoder in the step of feature fusion and the step of model acquisition adopts a deformable decoder.

[0023] The beneficial technical effects of the present invention are as follows: a face detection method based on DETR of the present invention selects a WiderFace data set for model construction and training to improve robustness, and adopts a Data Anchor Sample sampling method to randomly crop and randomly scale images in the WiderFace data set to obtain preprocessed images to increase the proportion of small-scale faces in the data set, so that the final face detection model obtained by subsequent training can better detect small-scale faces, and introduces a feature pyramid network into the Transformer model to fuse shallow spatial information with high-level semantic information, so that the fused feature map has richer information, improves the accuracy of the network and reduces the computational complexity, reduces memory usage, and improves the computing speed. By directly acquiring the model and performing face detection according to the model, end-to-end face detection is achieved, without the need to set a large number of anchors for detection, without the need for artificial prior knowledge, reducing artificial settings and hyperparameters, reducing the possibility of a decrease in network performance and detection effect caused by artificial settings, and improving the robustness and portability of the detection network. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.

[0025] Figure 1 A flow chart of a face detection method based on DETR provided by an embodiment of the present invention;

[0026] Figure 2 A specific flow chart of data preprocessing of a DETR-based face detection method provided by an embodiment of the present invention;

[0027] Figure 3 A specific flow chart of the steps of feature fusion of a DETR-based face detection method provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0030] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0031] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0032] See also Figure 1 As shown, Figure 1 A flow chart of a DETR-based face detection method provided in an embodiment of the present invention, the DETR-based face detection method comprising the following steps:

[0033] Step S11, data set selection: select the WiderFace data set. Among them, the public WiderFace data set is a benchmark data set for face detection. The image source of this data set is the WIDER data set, which contains rich annotations, including occlusion, posture, event category and face bounding box. There are a total of 32,203 pictures in the WiderFace data set, which may include a large number of dense scenes and small-scale faces. All pictures in the WiderFace data set are annotated with faces, with a total of 393,703 faces annotated. The WiderFace data set is divided into 61 categories according to the type of event scene, and each category is divided into training set, validation set and test set according to the ratio of 40%, 10% and 50%. By selecting the WiderFace data set, it has strong robustness in subsequent model construction and training.

[0034] Step S12, data preprocessing: randomly crop and randomly scale the images in the WiderFace dataset using the Data Anchor Sample sampling method to obtain preprocessed images. The images in the WiderFace dataset are randomly cropped and randomly scaled to change the data distribution of the original WiderFace dataset, thereby increasing the proportion of small-scale faces, so that the subsequently trained model can better detect small-scale faces.

[0035] Step S13, image feature map extraction: ResNet-50 is used as the backbone network to extract features of the preprocessed image, obtain the image feature map, and select the image feature maps of different layers to input into the Transformer model.

[0036] Step S14, feature fusion: Select the feature pyramid network (FPN) in the Transformer model to fuse the input image feature maps of different layers, and output the fused feature map to the decoder in the Transformer model. Among them, the feature pyramid network is introduced into the Transformer model to fuse the shallow spatial information with the high-level semantic information, so that the fused feature map has richer information, improves the accuracy of the network, reduces the computational complexity, reduces the memory usage, and increases the computational speed.

[0037] Step S15, model acquisition: the decoder in the Transformer model decodes the fused feature map to obtain an optimized face detection model.

[0038] Step S16, model training: use the training set in the WiderFace dataset to train the optimized face detection model to obtain the final face detection model.

[0039] Step S17, face detection: performing face detection according to the obtained final face detection model.

[0040] Among them, the face detection method based on DETR selects the WiderFace data set for model construction and training to improve robustness, and adopts the Data Anchor Sample sampling method to randomly crop and randomly scale the pictures in the WiderFace data set to obtain preprocessed pictures to increase the proportion of small-scale faces in the data set, so that the final face detection model obtained by subsequent training can better detect small-scale faces. Moreover, a feature pyramid network is introduced into the Transformer model to fuse shallow spatial information with high-level semantic information, so that the fused feature map has richer information, improves the accuracy of the network and reduces the computational complexity, reduces memory usage, and increases the computing speed. By directly obtaining the model and performing face detection according to the model, end-to-end face detection is achieved, without the need to set a large number of anchors for detection, without the need for artificial prior knowledge, reducing artificial settings and hyperparameters, reducing the possibility of a decrease in network performance and detection effect caused by artificial settings, and improving the robustness and portability of the detection network.

[0041] Combination Figure 2 , the step S12 specifically includes:

[0042] Step S121, randomly select a face from the image in the WiderFace dataset for cropping, obtain the size of the face after cropping the image, and obtain the anchor closest to the size of the face. According to the obtained anchor closest to the size of the face, its index can be obtained. Among them, the index of the anchor closest to the size of the face can be based on the index of formula (1):

[0043] i anchor =argmin i [abs(s ianchor -s face )],s ianchor =2 4+i ,i=0,1,2,3,4 (1)

[0044] In the formula, i anchor Indicates the index of the anchor that is closest to the size of the face, s ianchor Indicates the size of the anchor with index i, s face Indicates the size of the face, which is the size of the face obtained after cropping the image, abs(s ianchor -s face ) means calculating the absolute value of the difference between the size of the anchor with index i and the size of the face, argmin i[abs(s ianchor -s face )] indicates calculating the index of the anchor that minimizes the absolute value of the difference between the size of the anchor and the size of the face. The size of the face obtained after cropping the image is the size of the randomly selected face, and the anchor closest to the size of the face is the anchor closest to the size of the randomly selected face.

[0045] Step S122: According to the obtained index of the anchor closest to the size of the face, a value smaller than the index of the anchor closest to the size of the face is randomly selected as the index of the selected anchor. The selected anchor refers to the anchor corresponding to the value of the randomly selected index, and the index of the selected anchor can be expressed by formula (2):

[0046] i target ∈{0,1,2,3,…,min(5,i anchor +1)} (2)

[0047] In the formula, i target Indicates the index of the selected anchor, i anchor Represents the index of the anchor that is closest to the size of the face, and min represents the function used to obtain the minimum value;

[0048] Step S123, calculate and obtain the size ratio between the size of the selected anchor and the size of the face, randomly select a ratio between the size ratio being reduced by half and the size ratio being increased by one time as the image scaling ratio, scale the original image before cropping the face according to the image scaling ratio, and obtain the preprocessed image. The size ratio between the size of the selected anchor and the size of the face can be expressed by formula (3):

[0049] s * =s itarget / s face (3)

[0050] In the formula, s itarget Indicates the size of the selected anchor, s face Indicates the size of the face, s * Indicates size ratio.

[0051] The image scaling ratio can be expressed by formula (4):

[0052] s final =random(s * / 2,s * *twenty four)

[0053] In the formula, sfinal Indicates the image scaling ratio, s * represents the size ratio, and the random function is expressed in [s * / 2,s * *2] to generate a random real number.

[0054] The index of the anchor corresponds to the level of the backbone network, and the size of the anchor corresponds to the number of channels of the feature map of the image output by each layer of the backbone network. By cropping the face in the image and randomly scaling it according to the size of the cropped face, the network can focus on the face area.

[0055] Specifically, the step of selecting the image feature maps of different layers and inputting them into the Transformer model in step S13 is specifically: selecting the image feature maps of the 3rd, 4th and 5th layers and inputting them into the Transformer model. Figure 3 , the step S14 specifically includes:

[0056] Step S141, the number of channels of the output picture feature map of each layer of the backbone network is unified to 256 after being mapped by linear mapping. Among them, the output picture feature map of each layer of the backbone network refers to the picture feature map of each layer of the selected backbone network input to the Transformer model. The original number of channels of the output picture feature map of the 3rd, 4th and 5th layers of the backbone network are 512, 1024 and 2048 respectively. The number of channels of the output picture feature map of each layer is unified by linear mapping to facilitate subsequent feature fusion operations and reduce computational complexity.

[0057] Step S142, upsampling the image feature map of the 5th layer by using a bilinear interpolation algorithm; adding and fusing the image feature map of the 4th layer with the upsampled image feature map of the 5th layer through a 1×1 ordinary convolution, and then updating it through a 3×3 channel-by-channel convolution and a 1×1 ordinary convolution to obtain a new image feature map of the 4th layer; upsampling the new image feature map of the 4th layer by using a bilinear interpolation algorithm; adding and fusing the image feature map of the 3rd layer with the upsampled new image feature map of the 4th layer through a 1×1 ordinary convolution, and then updating it through a 3×3 channel-by-channel convolution and a 1×1 ordinary convolution to obtain a new image feature map of the 3rd layer as a multi-scale feature map after the fusion of the image feature maps of the 3rd, 4th and 5th layers.

[0058] The step S142 is specifically as follows:

[0059] The bilinear interpolation algorithm is used to upsample the image feature map of the 5th layer;

[0060] The image feature map of the 4th layer is added and fused with the upsampled image feature map of the 5th layer through a 1×1 ordinary convolution to obtain a preliminary fused image feature map of the 4th layer;

[0061] The initially fused image feature map of the 4th layer is updated by a 3×3 channel-by-channel convolution and then a 1×1 common convolution to obtain a new image feature map of the 4th layer;

[0062] The bilinear interpolation algorithm is used to upsample the new 4th layer image feature map;

[0063] The image feature map of the third layer is added and fused with the new image feature map of the fourth layer after upsampling through 1×1 ordinary convolution to obtain a preliminary fused image feature map of the third layer;

[0064] The initially fused third-layer image feature map is updated through 3×3 channel-by-channel convolution and then 1×1 ordinary convolution to obtain a new third-layer image feature map. The new third-layer image feature map is used as a multi-scale feature map after the image feature maps of the third, fourth and fifth layers are fused.

[0065] The number of channels of the output image feature map of each layer is the number of channels after the mapping is unified. Information extraction can be achieved by performing channel-by-channel convolution on the image feature map after the addition and fusion of two adjacent layers, and it can be matched with the subsequent decoder for decoding operations.

[0066] Step S143, flatten the multi-scale feature map and splice it in the spatial dimension to obtain a single-layer two-dimensional feature map as a fused feature map, and output the fused feature map to the decoder in the Transformer model. Flattening the multi-scale feature map means transforming the three dimensions (W, H, C) of the multi-scale feature map into two dimensions (WH, C), where W represents width, H represents height, C represents level, and WH represents the resolution of the multi-scale feature map. Since the multi-scale feature map is obtained based on the new third-layer image feature map, C is 3.

[0067] Specifically, the decoder in step S14 and step S15 may adopt a deformable decoder. By replacing the original decoder with the deformable decoder, local attention may be increased and computational complexity may be reduced, thereby increasing the number of target queries to adapt to face detection in dense scenes.

[0068] In summary, a face detection method based on DETR of the present invention selects a WiderFace data set for model construction and training to improve robustness, and randomly crops and randomly scales the pictures in the WiderFace data set by adopting the DataAnchor Sample sampling method to obtain preprocessed pictures to increase the proportion of small-scale faces in the data set, so that the final face detection model obtained by subsequent training can better detect small-scale faces, and introduces a feature pyramid network into the Transformer model to fuse shallow spatial information with high-level semantic information, so that the fused feature map has richer information, improves the accuracy of the network and reduces the computational complexity, reduces memory usage, and improves the computing speed. By directly acquiring the model and performing face detection according to the model, end-to-end face detection is realized, and there is no need to set a large number of anchors for detection, no need for artificial prior knowledge, and reduces artificial settings and hyperparameters. The possibility of a decrease in network performance and detection effect caused by artificial settings is reduced, and the robustness and portability of the detection network are improved.

[0069] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A face detection method based on DETR, characterized in that: The following steps are involved: Dataset selection: Select the WiderFace dataset; Data preprocessing: The Data Anchor Sample sampling method is used to randomly crop and randomly scale the images in the WiderFace dataset to obtain preprocessed images; Image feature map extraction: ResNet-50 is used as the backbone network to extract features from the preprocessed images, obtain image feature maps, and select image feature maps at different layers to input into the Transformer model; Feature fusion: In the Transformer model, a feature pyramid network (FPN) is used to fuse the feature maps of different layers of the input images, and the fused feature maps are output to the decoder in the Transformer model; Model acquisition: The decoder in the Transformer model decodes the fused feature map to obtain an optimized face detection model; Model training: Use the training set in the WiderFace dataset to train the optimized face detection model to obtain the final face detection model; Face detection: Perform face detection based on the final face detection model obtained; The step of selecting the image feature maps of different layers and inputting them into the Transformer model in the step of extracting the image feature map is specifically as follows: selecting the image feature maps of the 3rd, 4th and 5th layers and inputting them into the Transformer model; The step of feature fusion specifically includes: The number of channels of the output image feature map of each layer of the backbone network is unified to 256 through linear mapping; The bilinear interpolation algorithm is used to upsample the image feature map of the 5th layer; the image feature map of the 4th layer is added and fused with the upsampled image feature map of the 5th layer through a 1×1 ordinary convolution, and then updated through a 3×3 channel-by-channel convolution and a 1×1 ordinary convolution to obtain a new image feature map of the 4th layer; the bilinear interpolation algorithm is used to upsample the new image feature map of the 4th layer; the image feature map of the 3rd layer is added and fused with the upsampled new image feature map of the 4th layer through a 1×1 ordinary convolution, and then updated through a 3×3 channel-by-channel convolution and a 1×1 ordinary convolution to obtain a new image feature map of the 3rd layer as a multi-scale feature map after the fusion of the image feature maps of the 3rd, 4th and 5th layers; The multi-scale feature maps are flattened and concatenated in the spatial dimension to obtain a single-layer two-dimensional feature map as a fused feature map, which is then output to the decoder in the Transformer model.

2. The face detection method based on DETR according to claim 1, characterized in that: The data preprocessing steps specifically include: Randomly select a face from the image in the WiderFace dataset and crop it, get the size of the face after cropping the image, and get the anchor that is closest to the size of the face; According to the index of the anchor closest to the face size, a value smaller than the index is randomly selected as the index of the selected anchor; The size ratio between the size of the selected anchor and the size of the face is calculated, and a ratio is randomly selected between the range of the size ratio being halved and the size ratio being doubled as the image scaling ratio. The original image before cropping the face is scaled according to the image scaling ratio to obtain the preprocessed image.

3. The face detection method based on DETR according to claim 1, characterized in that: The decoder in the step of feature fusion and the step of model acquisition adopts a deformable decoder.

Citation Information

Patent Citations

  • High-accuracy detection and comparison method for shielded face images

    CN111310718A

  • Real-time target detection method based on Pearson coefficient matrix and attention fusion

    CN114187569A