An unsupervised low-light face detection model training method and detection method
Through the combination of high and low-order vision, the face detection model is trained by using brightening, degradation and multi-task advanced vision model migration, which solves the problem of poor low-light face detection performance in the existing technology, and achieves significantly improved detection performance.
Patent Information
- Application Number
- CN202110320033.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-03-25
AI Technical Summary
The existing unsupervised low-light face detection method cannot effectively utilize the existing normal-light face detection training data set, and cannot effectively solve the high-dimensional semantic differences and domain gaps between low-light and normal-light images, resulting in poor detection performance.
The face detection model is trained through low-order visual migration and high-order visual migration, combining brightening, degradation and multi-task advanced visual model migration. Specific steps include collecting normal and low-light data, performing brightening and degradation processing, and performing contrast learning and self-supervised learning in multi-task advanced vision model transfer.
It significantly improves the performance of low-light face detection, and the average accuracy is increased from 16.1 to 44.4, meeting the needs of actual application.
Smart Images

Figure CN115131844B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of digital image low-light enhancement and face detection, and relates to an unsupervised low-light face detection model training method and a detection method combining low-level and high-level vision. Background Art
[0002] Low light is a common image degradation, and insufficient light is usually caused by low-light shooting environments, camera failures, incorrect parameter settings, etc. Face detection in low-light environments has always attracted the attention of the academic and industrial communities. Traditional face detection model training requires a large-scale labeled training set. However, in low-light environments, data is difficult to label, and there are already a large number of normal-light face detection training data sets in the industry. Building a new low-light face detection data set will consume human and material resources repeatedly. How to make full use of the existing labeled normal-light face detection training data set and train a low-light face detection model without additional low-light labeling, that is, train a low-light face detection model in an unsupervised manner, has broad practical significance and application value.
[0003] Traditional unsupervised low-light face detection methods can be divided into three categories. The method based on brightening brightens the low-light image, thereby improving the performance of the face detection model trained on normal-light images. The method based on darkening reduces the brightness of the normal-light face detection training data, artificially synthesizes low-light face data, and then builds and trains a detection model on the synthesized training data set. The method based on domain transfer uses classical domain transfer techniques, regards normal-light images as the source domain, regards low-light images as the target domain, and then transfers the model trained on the normal-light image domain to the low-light image domain.
[0004] However, the methods based on brightening or darkening ignore the high-dimensional semantic differences between normal-light and low-light images, and the method based on domain transfer cannot solve the huge domain gap between the normal-light and low-light image domains. The existing unsupervised low-light face detection methods have poor effects and cannot meet the requirements of actual applications. Summary of the Invention
[0005] In view of the above problems, the purpose of the present invention is to provide an unsupervised low-light face detection model training method and a detection method combining low-level and high-level vision, which comprehensively improve face detection performance by combining low-level vision transfer and high-level vision transfer. The overall process is as shown in the appendix Figure 1 shown.
[0006] The technical solution adopted by the present invention is as follows:
[0007] An unsupervised low-light face detection method includes the following steps:
[0008] 1) Collect the labeled normal illumination training data and the unlabeled low illumination training data;
[0009] 2) Brighten the low illumination training data;
[0010] 3) Obtain the noise and color bias distributions of the brightened low illumination data, and apply the noise and color bias distributions to the normal illumination data to obtain the degraded normal illumination face training data;
[0011] 4) Multitask high-order visual model transfer. Use the brightened low illumination data obtained in step 2), the degraded normal illumination face training data obtained in step 3), and the original normal illumination face training data obtained in step 1) to train a face detection model. Perform face detection learning and puzzle-based self-supervised learning on the original normal illumination face training data, perform contrastive learning and puzzle-based self-supervised learning on the brightened low illumination data, and jointly perform contrastive learning on the degraded normal illumination face training data and the original normal illumination face training data.
[0012] 5) For the low illumination face detection image to be detected, first brighten it, and then input it into the trained low illumination face detection model, and the model outputs the face detection result.
[0013] Compared with the prior art, the positive effects of the present invention are as follows:
[0014] The present invention significantly improves the low illumination face detection performance. On the Dark Face low illumination face detection benchmark test set, the mean of AveragePrecision of the general face detector Dual Shot Face Detector can be increased from 16.1 to 44.4. Description of the Drawings
[0015] Figure 1 It is a flowchart for training an unsupervised low illumination face detection model combining high and low order vision;
[0016] Figure 2 It is a flowchart for degrading normal illumination data;
[0017] Figure 3 It is a schematic diagram of multitask high-order visual model transfer. Detailed Embodiment
[0018] To make the above features and advantages of the present invention more obvious and understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings as follows.
[0019] This embodiment discloses an unsupervised low illumination face detection method, which is specifically described as follows:
[0020] Step 1: Collect and construct the normal illumination face detection training dataset H; collect low-light images to form the low-light training dataset L. Among them, the face image samples in the normal illumination face detection training dataset contain face annotation boxes, and the face image samples in the low-light training dataset do not need to contain face annotation boxes.
[0021] Step 2: Brighten the low-light training dataset L to make its brightness close to that of normal illumination images. The brightening technique can use the low-light enhancement model RetinexNet based on the retina-cortex theory. No denoising is performed during the brightening process.
[0022] Step 3: Degrade the normal illumination images in the set H, and the process is as shown in the appendix Figure 2 as follows.
[0023] First, blur the brightened low-light training dataset E(L) using Gaussian blur, and the new dataset obtained is denoted as E(L) blur . The Gaussian blur uses a convolution kernel size of 25 and a variance of 75. Similarly, blur the normal illumination training dataset H, and the new dataset obtained is denoted as H blur .
[0024] Then, use the image conversion model to learn the image mapping from E(L) blur to E(L), and apply it to H blur , and the newly obtained dataset is denoted as H noise . The image conversion model can use the deep learning image translation model Pix2Pix.
[0025] Finally, perform color perturbation on H noise . For each image, randomly adjust the brightness within the perturbation range of (-0.6, +0.2), randomly adjust the contrast within the perturbation range of (-0.4, +0.4), randomly adjust the saturation within the perturbation range of (-0.4, +0.4), and randomly adjust the hue within the perturbation range of (-0.2, +0.2). The dataset after color perturbation is denoted as D(H).
[0026] Step 4: Build a face detection model and train the model using the constructed datasets E(L), H, and D(H). For face detection, use the general face detector Dual Shot Face Detector, or it can be replaced with other detection models. The model first performs standard face detection training on H as pre-training, and then performs multi-task high-order visual model transfer learning on E(L), H, and D(H). The total loss function term for the multi-task high-order visual model transfer learning of the face detection model is:
[0027]
[0028] Among them, the loss function weight λ det is set to 1, is set to 0.05, is set to 0.05, and λ E(L)↑ is set to 0.05. The training batch size is 8. First, it is trained for 20k iterations with a learning rate of 1e-4, and then it is trained for 40k iterations with a learning rate of 1e-5.
[0029] L det is the original detection training loss function Progressive Anchor Loss of the face detection model, indicating that the face detection model conducts standard training on the normally illuminated training dataset with annotations.
[0030] The other three are high-order visual transfer loss functions, and their training methods are as shown in the appendix Figure 3 as follows.
[0031] 1) is the domain transfer loss function based on jigsaw self-supervised learning. Jigsaw self-supervised learning divides an image into 9 image patches of 3×3, shuffles the order of the image patches, and trains the model to restore the original order of the image patches. For a brightened low-light image E(L n ), after shuffling the order of its image patches, it is denoted as E(L n ) jig . Substitute E(L n ) jig into the face detection model to obtain the deep features Denote the dictionary serial number of the shuffled order of the image patches of E(L n ) jig in all permutation and combination schemes of the image patches as For the face detection model Dual Shot Face Detector with VGG as the backbone, the deep features can adopt the multi-layer features that fuse six layers of conv3_3, conv4_3, conv5_3, conv_fc7, conv6_2, and conv7_2. After feature extraction, a multi-layer perceptron is used as the jigsaw self-supervised classifier, and the multi-layer perceptron adopts the network structure of "fully connected layer - rectified linear unit - fully connected layer".
[0032] Similarly, for an image H n in the normally illuminated image dataset, after shuffling the order of its image patches, it is denoted as (H n ) jig . Substitute (H n ) jig into the face detection model to obtain the deep features Denote (H n ) jigThe lexicographical order of the shuffled order of image patches among all permutation and combination schemes of image patches is The domain transfer loss function based on jigsaw self-supervised learning simultaneously predicts the shuffled order of patches for low-light and normal-light images:
[0033]
[0034] where L c is the cross-entropy classification loss function.
[0035] 2) is the contrastive learning loss function for cross-noise and color perturbation on normal-light images:
[0036]
[0037] where L q is the InfoNCE loss function, and D * (H n ) means sampling an image D(H n ) from the dataset D(H) with a probability of 50%, and sampling an image H n from the dataset H with a probability of 50%. The plus and minus signs in the loss function represent positive and negative samples in contrastive learning. D * (H n ) + represents the positive sample for contrastive learning of D * (H n ), and D * (H n ) - represents the negative sample for contrastive learning of D * (H n ). The contrastive learning adopts the momentum-based contrastive learning strategy MoCo. Similar to the jigsaw-based self-supervised learning, the classifier in contrastive learning also uses a multi-layer perceptron of "fully connected layer - rectified linear unit - fully connected layer", and the features also use a multi-layer feature that fuses six layers of conv3_3, conv4_3, conv5_3, conv_fc7, conv6_2, and conv7_2.
[0038] 3) L E(L)↑ is the contrastive learning loss function on the brightened low-light image:
[0039] L E(L)↑ = L q (E(L n ), E(L n ) + , E(L n ) - ),
[0040] Among them, E(L n ) + represents the positive sample for contrastive learning of E(L n ), and E(L n ) - represents the negative sample for contrastive learning of E(L n ). L E(L)↑ adopts the same contrastive learning strategy, classifier structure, and feature selection method as .
[0041] Step 5: In the inference stage, for the low-light face image to be detected, first use the low-light enhancement technique in Step 2 to brighten it, and then input it to the low-light face detection model trained in Step 4.
[0042] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A method for training an unsupervised low-light face detection model, the steps of which include: 1) Collect the labeled normal-light face training data and the unlabeled low-light face training data to obtain the normal-light face detection training dataset H and the low-light training dataset L; 2) Brighten the images in the low-light training dataset L to obtain the brightened low-light training dataset E(L); 3) Obtain the noise and color deviation distribution of the low-light face training data in the set E(L), and apply the noise and color deviation distribution to the normal-light face training data in the set H to obtain the degraded normal-light face training dataset D(H); 4) Use the set E(L), the set D(H) and the set H to train the face detection model.
2. The method according to claim 1, wherein, the method for obtaining the degraded normal-light face training dataset D(H) is: 1) Fuzzify the set E(L) using Gaussian blur, and denote the new data set obtained as E(L) blur ; Fuzzify the set H using Gaussian blur, and denote the new data set obtained as H blur ; 2) Use an image conversion model to learn the image mapping from E(L) blur to E(L), and apply this image mapping to the set H blur , and the newly obtained data set is denoted as H noise ; 3) Perform color perturbation on the set H noise After the color perturbation, the data set is the degraded normal illumination face training data set D(H).
3. The method according to claim 1, wherein, during the process of training the face detection model, face detection learning and jigsaw-based self-supervised learning are performed on the set H, contrast learning and jigsaw-based self-supervised learning are performed on the set E(L), and contrast learning is jointly performed on the sets D(H) and H.
4. The method according to claim 1 or 3, wherein, The total loss function used when training the face detection model is where λ det , λ E(L)↑ is the loss function weight, and L det is the loss function used when training the face detection model using the set H; It is a domain transfer loss function for training a face detection model based on jigsaw puzzles in the set H during self-supervised learning. The method for training a face detection model based on jigsaw-based self-supervised learning is as follows: divide an image into multiple image patches, shuffle the order of the image patches of the image, and then train the face detection model to restore the original order of the image patches. is the image H in the set H n The deep features extracted by inputting the shuffled image patches into the face detection model is H n The dictionary order number of the shuffled order of the corresponding image patch in all the permutation and combination schemes of the image patches is the image E(L) in E(L n ) The deep features extracted by inputting the shuffled image patches into the face detection model is E(L n ) The dictionary order number of the shuffled order of the corresponding image patch in all the permutation and combination schemes of the image patches; It is the contrastive learning loss function when training a face detection model using set D(H) and set H. L E(L)↑ is the contrastive learning loss function when training a face detection model using the set E(L), L E(L)↑ = L q (E(L n ), E(L n ) + , E(L n ) - ); where L c is the cross-entropy classification loss function, L q is the InfoNCE loss function, D * (H n ) represents the training samples sampled from the dataset D(H) with a probability of 50% and the training samples sampled from the dataset H with a probability of 50%, D * (H n ) + represents the contrastive learning positive samples of D * (H n ), D * (H n ) - represents the contrastive learning negative samples of D * (H n ), E(L n ) represents the training samples collected from E(L), E(L n ) + represents the contrastive learning positive samples of E(L n ), E(L n ) - represents the contrastive learning negative samples of E(L n ). The contrastive learning adopts the momentum-based contrastive learning strategy MoCo.
5. The method according to claim 4, wherein, a multi-layer perceptron is used as the self-supervised classifier based on jigsaw and contrast learning, and the multi-layer perceptron adopts a network structure of "fully connected layer - rectified linear unit - fully connected layer".
6. The method according to claim 4, wherein, the face detection model is the face detector Dual ShotFaceDetector; the deep features are multi-layer features that adopt the fusion of six layers of conv3_3, conv4_3, conv5_3, conv_fc7, conv6_2, and conv7_2.
7. A face detection method based on the low-light face detection model trained according to claim 1, wherein, the low-light face detection image to be detected is brightened and then input into the trained low-light face detection model, and the face detection result is output.
Citation Information
Patent Citations
Detection and recognition method of human face in video under low-light conditions
CN106446872A
Target intelligent detection method under weak light
CN110210401A