Face detection method and device and related product

Through the face detection model, the image is extracted and the face density distribution is determined. Combined with the comparison results of prediction and real density distribution, the face samples are divided dynamically and easily and sample allocation are solved, which solves the problem of low face detection accuracy in the prior art and achieves higher detection accuracy.

CN119942624AActive Publication Date: 2025-05-06ANHUI LISTENAI CO LTD +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510437679.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing face detection technology is low in accuracy when dealing with difficult faces such as complex environments, low resolution, changing postures and occlusion, and the static difficulty evaluation method leads to overfitting the model or insufficient learning.

Method used

Feature extraction of images through the face detection model, determine the face density distribution, combine the comparison results of prediction and real density distribution, dynamically divide face samples, positive and negative samples are allocated, and the model is updated to improve detection accuracy.

Benefits of technology

It realizes an accurate assessment of the difficulty of face samples, dynamically adjusts sample allocation, avoids overfitting and insufficient learning, and improves the accuracy of face detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942624A_ABST
    Figure CN119942624A_ABST
Patent Text Reader

Abstract

The invention provides a face detection method and device and a related product, and relates to the technical field of computer vision. When the method is executed, firstly, a face detection model is utilized to perform feature extraction on images in a training set to obtain image multi-scale features, then face density distribution in the images is determined according to the image multi-scale features, and the face density distribution comprises predicted face density distribution and real face density distribution. The method comprises the following steps: firstly, carrying out face detection on a face sample, then carrying out difficulty division on the face sample according to a comparison result between real face density distribution and predicted face density distribution, then carrying out face detection positive and negative sample distribution according to the difficulty division on the face sample, updating a face detection model according to the face detection positive and negative sample distribution, and finally, carrying out face detection. And performing face detection by using the updated face detection model. In this way, the effectiveness of face detection model training can be improved, and then the accuracy of face detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a face detection method, device and related products. Background Art

[0002] Face detection is a key task in the field of computer vision and is the basic premise for various tasks such as face recognition, face attribute classification, face calibration and face tracking. The quality of face detection profoundly affects the performance of subsequent tasks, thus having a significant impact on security monitoring, human-computer interaction and virtual reality.

[0003] With the development of deep learning technology, general object detection based on neural networks, such as SSD, YOLO and RetinaNet, has greatly promoted the development of face detection. However, general face detection still faces the following challenges: 1) complex environmental factors, such as lighting and face background objects; 2) shooting distance and hardware configuration of image acquisition equipment, such as resolution, shutter speed, etc.; 3) variability of facial poses and occlusion between multiple faces. In real-world applications, these challenges often coexist, and recent methods often fail to achieve satisfactory results for faces with important features such as small size, low resolution and pose variation. Faces with one or more of these features are generally called hard faces.

[0004] In the prior art, when dealing with challenging face samples such as hard faces, the prior art usually distinguishes faces of different difficulty levels through label assignment. However, after pre-defining anchor points, these methods assign fixed labels to all faces in the image, which means that the difficulty assessment of each face remains unchanged throughout the training process. This static difficulty assessment method is not ideal because continuously increasing the loss weight or the number of positive samples for specific face samples during training may cause the model to overfit these faces and under-learn other faces, resulting in low face detection accuracy.

[0005] In summary, how to improve the accuracy of face detection is a problem that technical personnel in this field need to solve urgently. Summary of the invention

[0006] In view of this, the present application provides a face detection method, device and related products, aiming to improve the accuracy of face detection.

[0007] In a first aspect, the present application provides a face detection method, comprising: Use the face detection model to extract features from the images in the training set and obtain multi-scale features of the images; Determining a face density distribution in the image according to the multi-scale features of the image; the face density distribution includes a predicted face density distribution and a real face density distribution; According to the comparison result between the real face density distribution and the predicted face density distribution, the face samples are divided into easy and difficult categories; According to the difficulty classification of the face samples, positive and negative samples for face detection are allocated; According to the allocation of positive and negative samples of the face detection, the face detection model is updated; The updated face detection model is used to perform face detection.

[0008] Optionally, determining the face density distribution in the image according to the multi-scale features of the image includes: Each scale feature in the multi-scale features of the image is subjected to multi-layer convolution to obtain a predicted face density map; the predicted face density map is used to indicate the predicted face density distribution; When supervising the face density map branch, an all-zero input mask is given; Setting the position of the face center point in the all-zero input mask to 1; A real face density map is generated using a Gaussian filter; the real face density map is used to indicate the real face density distribution.

[0009] Optionally, after determining the face density distribution in the image according to the multi-scale features of the image, the method further includes: Classifying and regressing the multi-scale features of the image to obtain face confidence and face position; the face confidence is used for face judgment; Determining the size of the face according to the result of the face judgment; Positive and negative samples for face detection are pre-allocated according to the face position and the face size.

[0010] Optionally, after determining the face density distribution in the image according to the multi-scale features of the image, the method further includes: Storing the predicted face density map in a prediction result list; Assigning an index number in the prediction result list to the predicted face density map; Storing the real face density map in a real result list; An index number in the real result list is assigned to the real face density map.

[0011] Optionally, the dividing the face samples into easy and difficult categories according to the comparison result between the real face density distribution and the predicted face density distribution includes: Indexing a target predicted face density map from the prediction result list; Indexing a target real face density map from the real result list; According to the target predicted face density map, the predicted regional average density is calculated; Calculating the real regional average density according to the target real face density map; The human face samples are divided into easy and difficult categories according to the predicted regional average density and the real regional average density simple difference index.

[0012] In a second aspect, the present application provides a face detection device, comprising: The feature extraction module is used to extract features from images in the training set using the face detection model to obtain multi-scale features of the image; A first determination module is used to determine the face density distribution in the image according to the multi-scale features of the image; the face density distribution includes the predicted face density distribution and the real face density distribution; A classification module, used for classifying the face samples into easy and difficult categories according to the comparison result between the real face density distribution and the predicted face density distribution; A first allocation module, configured to allocate positive and negative samples for face detection according to the difficulty classification of the face samples; An updating module, used for updating the face detection model according to the allocation of positive and negative samples for face detection; The face detection module is used to perform face detection using the updated face detection model.

[0013] Optionally, the first determining module includes: A multi-layer convolution submodule, used for performing multi-layer convolution on each scale feature in the multi-scale features of the image to obtain a predicted face density map; the predicted face density map is used to indicate the predicted face density distribution; A submodule is given to give an all-zero input mask when supervising the face density map branch; A setting submodule, used to set the position of the center point of the face in the all-zero input mask to 1; The generating submodule is used to generate a real face density map using a Gaussian filter; the real face density map is used to indicate the real face density distribution.

[0014] Optionally, the device further comprises: A classification and regression module, used to classify and regress the multi-scale features of the image to obtain face confidence and face position; the face confidence is used to perform face judgment; A second determination module, used to determine the size of the face according to the result of the face determination; The pre-allocation module is used to pre-allocate positive and negative samples for face detection according to the face position and face size.

[0015] Optionally, the device further comprises: A first storage module, used for storing the predicted face density map in a prediction result list; A second allocation module, used for allocating an index number in the prediction result list to the predicted face density map; A second storage module, used for storing the real face density map into a real result list; The third allocation module is used to allocate an index number in the real result list to the real face density map.

[0016] Optionally, the division module includes: A first indexing submodule, used to index a target predicted face density map from the prediction result list; A second indexing submodule, used to index a target real face density map from the real result list; A first calculation submodule is used to calculate the predicted regional average density according to the target predicted face density map; A second calculation submodule is used to calculate the real regional average density according to the target real face density map; The division submodule is used to divide the human face samples into easy and difficult categories according to the predicted regional average density and the real regional average density simple difference index.

[0017] In a third aspect, an embodiment of the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a face detection method as described in any one of the implementation methods in the first aspect of the embodiment of the present application.

[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the terminal device executes a face detection method as described in any one of the implementation methods in the first aspect of the embodiment of the present application.

[0019] The present application provides a face detection method. When executing the method, firstly, a face detection model is used to extract features of images in a training set to obtain multi-scale features of the image, and then the face density distribution in the image is determined based on the multi-scale features of the image, wherein the face density distribution includes the predicted face density distribution and the real face density distribution, and then the face samples are divided into difficulty levels based on the comparison results between the real face density distribution and the predicted face density distribution, and then, based on the difficulty level of the face samples, positive and negative samples of face detection are allocated, and based on the allocation of positive and negative samples of face detection, the face detection model is updated, and finally, face detection is performed using the updated face detection model. In this way, by comparing the predicted face density distribution with the real face density distribution, the difficulty of the face samples can be accurately evaluated, and the positive and negative samples of face detection can be allocated based on the difficulty level, thereby improving the effectiveness of the face detection model training, and thus improving the accuracy of face detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 A flow chart of a face detection method provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a face detection device provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. The present application provides a face detection method, device and related products for use in the field of computer vision technology. The above is only an example and does not limit the application field of the method and device names provided in the present application.

[0023] Face detection is a key task in the field of computer vision and is the basic premise for various tasks such as face recognition, face attribute classification, face calibration and face tracking. The quality of face detection profoundly affects the performance of subsequent tasks, thus having a significant impact on security monitoring, human-computer interaction and virtual reality.

[0024] With the development of deep learning technology, general object detection based on neural networks, such as SSD, YOLO and RetinaNet, has greatly promoted the development of face detection. However, general face detection still faces the following challenges: 1) complex environmental factors, such as lighting and face background objects; 2) shooting distance and hardware configuration of image acquisition equipment, such as resolution, shutter speed, etc.; 3) variability of facial poses and occlusion between multiple faces. In real-world applications, these challenges often coexist, and recent methods often fail to achieve satisfactory results for faces with important features such as small size, low resolution and pose variation. Faces with one or more of these features are generally called hard faces.

[0025] In the prior art, when dealing with challenging face samples such as hard faces, the prior art usually distinguishes faces of different difficulty levels through label assignment. However, after pre-defining anchor points, these methods assign fixed labels to all faces in the image, which means that the difficulty assessment of each face remains unchanged throughout the training process. This static difficulty assessment method is not ideal because continuously increasing the loss weight or the number of positive samples for specific face samples during training may cause the model to overfit these faces and under-learn other faces, resulting in low face detection accuracy.

[0026] After research, the inventor proposed the technical solution of the present application. First, the face detection model is used to extract features from the images in the training set to obtain multi-scale features of the image. Then, based on the multi-scale features of the image, the face density distribution in the image is determined, wherein the face density distribution includes the predicted face density distribution and the real face density distribution. Then, based on the comparison result between the real face density distribution and the predicted face density distribution, the face samples are divided into difficulty levels. Then, based on the difficulty level of the face samples, positive and negative samples for face detection are allocated. Based on the allocation of positive and negative samples for face detection, the face detection model is updated. Finally, face detection is performed using the updated face detection model. In this way, by comparing the predicted face density distribution with the real face density distribution, the difficulty of the face samples can be accurately evaluated, and the positive and negative samples for face detection can be allocated based on the difficulty level, thereby improving the effectiveness of the face detection model training and thus improving the accuracy of face detection.

[0027] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. It should be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0028] See also Figure 1 , Figure 1 A flow chart of a face detection method provided in an embodiment of the present application includes: S101: Use the face detection model to extract features from images in the training set to obtain multi-scale features of the images.

[0029] The face detection model is used to extract features from the images in the training set. The face detection model is a deep learning model that can obtain multi-scale features of the image. The scale of the multi-scale features of the image is divided according to different downsampling ratios, such as downsampling 4 times, 8 times, 16 times, 32 times, etc.

[0030] S102: Determine the face density distribution in the image according to the multi-scale features of the image.

[0031] After obtaining the multi-scale features of the image in S101, the face density map branch of the face detection model is first used to obtain a predicted face density map for each scale through multi-layer convolution, and then the intermediate features of the face density map branch are given to the face detection branch of the corresponding scale. After feature alignment, the face confidence and position result predictions are obtained through multi-layer convolution. Among them, the face confidence is used to determine whether the image is a face. Among them, when supervising the face density map branch, a specific size of all-zero input mask is first given, and then the position of the face center point in the all-zero input mask is set to 1, and finally a Gaussian filter is applied to generate a real face density map, wherein the face center point is determined based on the face annotation rectangle on the detection scale, and these rectangles are obtained when performing face detection at the corresponding scale. The calculation method of the real face density map D(x) corresponding to the specific image is as follows: ; Where D(x) is the real face density map corresponding to the image, δ is the Dirac function, x is the image coordinate, and x i is the coordinate of the center point of the face, n is the number of faces, G σ is a two-dimensional Gaussian kernel.

[0032] Optimization objective function Loss corresponding to the face density map prediction task dm for: ; Where X is the predicted face density map and Y is the real face density map.

[0033] By supervising the face detection branch, the optimization objective function of the multi-task is obtained: ; Among them, Loss dm is the optimization objective function for face density map prediction, Loss det is the optimization objective function for face detection, λ1 and λ2 are weighted coefficients of two optimization objectives, where λ1 and λ2 are preset by those skilled in the art based on experience. In some implementations, , .

[0034] S103: According to the comparison result between the real face density distribution and the predicted face density distribution, the face samples are divided into easy and difficult categories.

[0035] Before executing step S103, it is necessary to first store the predicted face density map obtained in step S102 in the prediction result list, and assign an index number in the prediction result list to the predicted face density map; first store the real face density map obtained in step S102 in the real result list, and assign an index number in the real result list to the real face density map.

[0036] Specifically, the difficulty of face samples is divided as follows: First, get the corresponding predicted face density map pred_density according to index i from the prediction result list Pred. This map represents the prediction of face density on a certain level feature map, where i is the scale index. If there are 3 scales, there are 3 predicted face density maps in the prediction pred list, i represents the i-th predicted face density map, i.e. the i-th scale. Get the corresponding real face density map gt_density according to index i from the real face density map list. This map reflects the actual face distribution.

[0037] Next, for each real face rectangle, the average density on the predicted face density map is calculated through the region of interest alignment (RoIAlign) operation, that is, the predicted regional average density pred_avg_rois, and the average density on the real face density map is calculated, that is, the real regional average density gt_avg_rois.

[0038] Then, the predicted regional average density pred_avg_rois and the true regional average density gt_avg_rois are used to calculate the simple difference index diff between the predicted regional average density and the true regional average density: ; The weight gt_weights of each face rectangle is calculated based on the difference index diff. In order to ensure that the weight can reflect the importance of the difference while maintaining a positive value, the softmax function is needed to convert the difference value. The temperature coefficient T can adjust the smoothness of the softmax output. The larger the T value, the closer the output is to a uniform distribution; the smaller the T value, the more the output tends to highlight the maximum value: ; Finally, the weights gt_weights are normalized. Normalization is the process of ensuring that the weights are of appropriate scale. This is done by dividing by the average value of the weights so that they remain comparable even on datasets of different sizes: ; The final weights of different samples can be used to adjust the sample distribution and loss function during the training process. This evaluation method more accurately reflects the gap between the prediction and reality of different samples.

[0039] S104: Allocate positive and negative samples for face detection according to the difficulty classification of face samples.

[0040] Model training is a dynamic process. As training progresses, difficult samples may change. Therefore, it is necessary to evaluate sample weights in real time and adjust the distribution of positive and negative samples based on the latest training results. This dynamic adjustment can help the model continue to improve and better adapt to the changing data distribution. First, calculate the mean and variance of the weights of all face samples at a certain scale, denoted by m and v respectively, where m and v are real numbers. Then, the sample set with sample weights greater than m+v is denoted as S, and the remaining sample set is denoted as T, where S and T are both real numbers. Next, calculate the mean number of pre-allocated positive samples in set T, denoted by k, where k is a real number. Then, add the faces in set S with less than k pre-allocated positive samples to k positive samples.

[0041] Finally, for the weights gt_weights of different face samples, the loss weight of the positive samples assigned to each face is weighted according to the following formula: ; S105: updating the face detection model according to the allocation of positive and negative samples for face detection.

[0042] According to the distribution of positive and negative samples of face detection, the weights of positive samples of faces in loss calculation are updated. When the weights of positive samples of faces in loss calculation are not updated, the weights of all positive samples of faces are 1. After the weights of positive samples of faces in loss calculation are updated, the weights of some positive samples, such as simple samples, are less than 1, and the weights of some positive samples, such as difficult samples, are greater than 1. Then, the model with the redistribution of positive and negative samples and the reweighting of positive samples is used for training to obtain the updated face detection model. By specifically increasing the number of positive sample allocations and loss weights of face samples with larger weights, the model can be more focused on difficult samples, thereby improving its overall performance.

[0043] S106: Perform face detection using the updated face detection model.

[0044] In an embodiment of the present application, a face detection model is first used to extract features from images in a training set to obtain multi-scale features of the image, and then the face density distribution in the image is determined based on the multi-scale features of the image, wherein the face density distribution includes the predicted face density distribution and the real face density distribution, and then the face samples are divided into difficulty levels based on the comparison results between the real face density distribution and the predicted face density distribution, and then, based on the difficulty level of the face samples, positive and negative samples for face detection are allocated, and based on the allocation of positive and negative samples for face detection, the face detection model is updated, and finally, face detection is performed using the updated face detection model. In this way, by comparing the predicted face density distribution with the real face density distribution, the difficulty of the face samples can be accurately evaluated, and the positive and negative samples for face detection can be allocated based on the difficulty level, thereby improving the effectiveness of face detection model training and thus improving the accuracy of face detection.

[0045] The above are some specific implementations of the face detection method provided in the embodiment of the present application. Based on this, the present application also provides a corresponding device. The device provided in the embodiment of the present application will be introduced from the perspective of functional modularization.

[0046] See also Figure 2 , Figure 2 A schematic diagram of the structure of a face detection device provided in an embodiment of the present application, the face detection device 200 includes: A feature extraction module 210 is used to extract features from images in a training set using a face detection model to obtain multi-scale features of the images; A first determination module 220, configured to determine a face density distribution in the image according to the multi-scale features of the image; the face density distribution includes a predicted face density distribution and a real face density distribution; A classification module 230, configured to classify the face samples into easy and difficult categories according to the comparison result between the real face density distribution and the predicted face density distribution; A first allocation module 240, configured to allocate positive and negative samples for face detection according to the difficulty classification of the face samples; An updating module 250, configured to update the face detection model according to the allocation of positive and negative samples for face detection; The face detection module 260 is used to perform face detection using the updated face detection model.

[0047] Optionally, the first determining module 220 includes: A multi-layer convolution submodule, used for performing multi-layer convolution on each scale feature in the multi-scale features of the image to obtain a predicted face density map; the predicted face density map is used to indicate the predicted face density distribution; A submodule is given to give an all-zero input mask when supervising the face density map branch; A setting submodule, used to set the position of the center point of the face in the all-zero input mask to 1; The generating submodule is used to generate a real face density map using a Gaussian filter; the real face density map is used to indicate the real face density distribution.

[0048] Optionally, the face detection device 200 further includes: A classification and regression module, used to classify and regress the multi-scale features of the image to obtain face confidence and face position; the face confidence is used to perform face judgment; A second determination module, used to determine the size of the face according to the result of the face determination; The pre-allocation module is used to pre-allocate positive and negative samples for face detection according to the face position and face size.

[0049] Optionally, the face detection device 200 further includes: A first storage module, used for storing the predicted face density map in a prediction result list; A second allocation module, used for allocating an index number in the prediction result list to the predicted face density map; A second storage module, used for storing the real face density map into a real result list; The third allocation module is used to allocate an index number in the real result list to the real face density map.

[0050] Optionally, the step 230 includes: A first indexing submodule, used to index a target predicted face density map from the prediction result list; A second indexing submodule, used to index a target real face density map from the real result list; A first calculation submodule is used to calculate the predicted regional average density according to the target predicted face density map; A second calculation submodule is used to calculate the real regional average density according to the target real face density map; The division submodule is used to divide the human face samples into easy and difficult categories according to the predicted regional average density and the real regional average density simple difference index.

[0051] The embodiments of the present application also provide corresponding devices and computer storage media for implementing the solutions provided by the embodiments of the present application.

[0052] like Figure 3 As shown, the computer device 01 is in the form of a general computing device. The components of the computer device 01 may include but are not limited to: one or more processors or processor units 03, a system memory 08, and a bus 04 connecting different system components (including the system memory 08 and the processor unit 03).

[0053] Bus 04 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor or a local bus using any of a variety of bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus and Peripheral Component Interconnect (PCI) bus.

[0054] The computer device 01 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 01, including volatile and non-volatile media, removable and non-removable media.

[0055] The system memory 08 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 09 and / or cache memory 10. The computer device 01 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 11 may be used to read and write non-removable, non-volatile magnetic media ( Figure 3 not shown, usually called a "hard drive"). Although Figure 3Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 04 via one or more data medium interfaces. The system memory 08 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.

[0056] A program / utility 12 having a set (at least one) of program modules 13 may be stored, for example, in system memory 08, such program modules 13 including but not limited to an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 13 generally perform the functions and / or methods of the embodiments described herein.

[0057] The computer device 01 may also communicate with one or more external devices 02 (e.g., keyboard, pointing device, display 07, etc.), one or more devices that enable a user to interact with the computer device 01, and / or any device that enables the computer device 01 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed via input / output (I / O) interface 06. Furthermore, the computer device 01 may also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN) and / or public network, such as the Internet) via network adapter 05. Figure 3 As shown, the network adapter 05 communicates with other modules of the computer device 01 via the bus 04. It should be understood that although Figure 3 Not shown, other hardware and / or software modules may be used in conjunction with computer device 01, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0058] The processor unit 03 executes various functional applications and data processing by running the programs stored in the system memory 08, for example, implementing a face detection method provided in an embodiment of the present application.

[0059] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0060] Through the description of the above implementation methods, it can be known that those skilled in the art can clearly understand that all or part of the steps in the above-mentioned embodiment method can be implemented by means of software plus a general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a read-only memory (ROM) / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in each embodiment of the present application or some parts of the embodiments.

[0061] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.

[0062] The above description is merely an exemplary embodiment of the present application and is not intended to limit the protection scope of the present application.

Claims

1. A face detection method, characterized in that: include: Use the face detection model to extract features from the images in the training set and obtain multi-scale features of the images; Determining a face density distribution in the image according to the multi-scale features of the image; the face density distribution includes a predicted face density distribution and a real face density distribution; According to the comparison result between the real face density distribution and the predicted face density distribution, the face samples are divided into easy and difficult categories; According to the difficulty classification of the face samples, positive and negative samples for face detection are allocated; According to the allocation of positive and negative samples of the face detection, the face detection model is updated; The updated face detection model is used to perform face detection.

2. The method according to claim 1, characterized in that The determining, according to the multi-scale features of the image, the face density distribution in the image includes: Each scale feature in the multi-scale features of the image is subjected to multi-layer convolution to obtain a predicted face density map; the predicted face density map is used to indicate the predicted face density distribution; When supervising the face density map branch, an all-zero input mask is given; Setting the position of the face center point in the all-zero input mask to 1; A real face density map is generated using a Gaussian filter; the real face density map is used to indicate the real face density distribution.

3. The method according to claim 1, characterized in that After determining the face density distribution in the image according to the multi-scale features of the image, the method further includes: Classifying and regressing the multi-scale features of the image to obtain face confidence and face position; the face confidence is used for face judgment; Determining the size of the face according to the result of the face judgment; Positive and negative samples for face detection are pre-allocated according to the face position and the face size.

4. The method according to claim 2, characterized in that: After determining the face density distribution in the image according to the multi-scale features of the image, the method further includes: Storing the predicted face density map in a prediction result list; Assigning an index number in the prediction result list to the predicted face density map; Storing the real face density map in a real result list; An index number in the real result list is assigned to the real face density map.

5. The method according to claim 4, characterized in that The step of dividing the face samples into easy and difficult categories according to the comparison result between the real face density distribution and the predicted face density distribution includes: Indexing a target predicted face density map from the prediction result list; Indexing a target real face density map from the real result list; According to the target predicted face density map, the predicted regional average density is calculated; Calculating the real regional average density according to the target real face density map; The human face samples are divided into easy and difficult categories according to the predicted regional average density and the real regional average density simple difference index.

6. A face detection device, characterized in that: include: The feature extraction module is used to extract features from images in the training set using the face detection model to obtain multi-scale features of the image; A first determination module is used to determine the face density distribution in the image according to the multi-scale features of the image; the face density distribution includes the predicted face density distribution and the real face density distribution; A classification module, used for classifying the face samples into easy and difficult categories according to the comparison result between the real face density distribution and the predicted face density distribution; An allocation module, used for allocating positive and negative samples for face detection according to the difficulty classification of the face samples; An updating module, used for updating the face detection model according to the allocation of positive and negative samples for face detection; The face detection module is used to perform face detection using the updated face detection model.

7. The device according to claim 6, characterized in that The first determining module includes: A multi-layer convolution submodule, used for performing multi-layer convolution on each scale feature in the multi-scale features of the image to obtain a predicted face density map; the predicted face density map is used to indicate the predicted face density distribution; A submodule is given to give an all-zero input mask when supervising the face density map branch; A setting submodule, used to set the position of the center point of the face in the all-zero input mask to 1; The generating submodule is used to generate a real face density map using a Gaussian filter; the real face density map is used to indicate the real face density distribution.

8. The device according to claim 6, characterized in that The device also includes: A classification and regression module, used to classify and regress the multi-scale features of the image to obtain face confidence and face position; the face confidence is used to perform face judgment; Determining the size of the face according to the result of the face judgment; The pre-allocation module is used to pre-allocate positive and negative samples for face detection according to the face position and the face size.

9. A computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the face detection method according to any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the face detection method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Face detection model training method and device

    CN109271970A

  • Video-based real-time face tracking method and system

    CN112085764A

  • Face detection neural network, training method, face detection method and storage medium

    CN112287820A

  • Face recognition method and device, electronic equipment and readable storage medium

    CN112329619A

  • Face recognition method and device, electronic equipment and computer readable storage medium

    CN116152870A