Face detection method and device, electronic equipment, storage medium and chip system
Patent Information
- Application Number
- CN202210993664.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-08-18
AI Technical Summary
[0003]目前在智能终端设备如智能手机,人脸检测和人脸关键点检测为了获得较高的检测精度,通常出现功耗过高、无法达到实时检测30fps和检测结果出现明显时延等问题;或者为了降低功耗、提升实时性,减小算法复杂度的同时导致模型检测精度下降严重,且依旧存在检测结果时延明显的问题
[0016] According to a fifth aspect of the present disclosure, a chip system is provided, comprising: a storage medium for storing instructions; and a processing circuit for executing the instructions to implement the face detection method provided in the first aspect.
Smart Images

Figure CN117636411B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image recognition, and more particularly to a face detection method, apparatus, electronic device, storage medium, and chip system. Background Technology
[0002] Image-based face location awareness includes face detection and facial landmark detection. Face detection locates the face region in an image and outputs a bounding box and confidence score. Facial landmark detection locates the positional information of facial features, including eyebrows, eyes, nose, mouth, and facial contours. With the development of artificial intelligence technologies such as deep learning, both face detection and facial landmark detection technologies have achieved significant improvements in detection accuracy when hardware computing power is sufficient. However, for intelligent terminal devices with limited computing power, achieving high-precision, high-stability, low-power, and low-latency face detection and facial landmark detection remains a major challenge.
[0003] Currently, in smart terminal devices such as smartphones, face detection and facial landmark detection often suffer from problems such as excessive power consumption, inability to achieve real-time detection at 30fps, and significant latency in detection results in order to obtain higher detection accuracy. Alternatively, in order to reduce power consumption, improve real-time performance, and reduce algorithm complexity, the model detection accuracy is severely reduced, and the problem of significant latency in detection results still exists. Summary of the Invention
[0004] To overcome the problems existing in related technologies, this disclosure provides a face detection method, apparatus, electronic device, storage medium, and chip system.
[0005] According to a first aspect of the present disclosure, a face detection method is provided, comprising: acquiring an image to be detected through a portion of an image signal processing pipeline; identifying a face region in the image to be detected using a face detection model to obtain a face image bounding box, and determining a face detection confidence level; the face detection confidence level characterizes the accuracy of the face image bounding box; identifying facial key points in the image to be detected using a facial key point detection model, and determining the facial key point confidence level; the facial key point confidence level characterizes the certainty of alignment between the facial key points and the face region; cascading the facial key point detection model with the face detection model; filtering face image bounding boxes that do not include the face region based on the face detection confidence level and the facial key point confidence level, correcting the position of the filtered face image bounding boxes by tracking the facial key points, and obtaining a processed face image bounding box; and performing inter-frame smoothing on the processed face image bounding box to obtain a face detection result.
[0006] Optionally, the step of obtaining the image to be detected through a certain module of the image signal processing pipeline includes: performing noise reduction, scaling, gamma processing and downsampling processing on the original image through the RAW module in the image signal processing pipeline to obtain the image to be detected.
[0007] Optionally, the step of identifying the face region in the image to be detected by the face detection model and obtaining a rectangular bounding box containing the face region includes: selecting one frame at fixed intervals as the input image of the face detection model for the image to be detected; determining the face region and other regions in the input image by using a pre-set classification threshold in the face detection model; the classification threshold characterizing the probability that a region in the input image belongs to the face region; and taking the smallest rectangular region containing all the face regions as the rectangular bounding box of the face image.
[0008] Optionally, the preset classification threshold can be set to a value between 0 and 0.55.
[0009] Optionally, the step of identifying facial key points in the image to be detected by a facial key point detection model and determining the confidence level of the facial key points includes: identifying facial key points in the rectangular frame of the face image and in each frame of the image to be detected that has not been identified by the facial detection model, and determining the confidence level of the facial key points.
[0010] Optionally, the step of filtering face image rectangles that do not include the face region based on the face key point confidence score includes: obtaining the face key point confidence score corresponding to the face image rectangle in each frame; and filtering out the corresponding face image rectangle when both the face detection confidence score and the face key point confidence score are less than a predetermined threshold.
[0011] Optionally, the facial landmark detection model is trained through the following steps: acquiring sample images; the sample images include non-face images of a specified proportion; cropping the face region in the sample images by probabilistically random offset expansion of the rectangular box; marking the face region and facial landmarks in the sample images to obtain marked samples; and training the facial landmark detection model using the marked samples.
[0012] Optionally, the image to be detected is written into an on-chip memory (OCM) accessible by the network processor (NPU); the NPU reads the image to be detected from the OCM as the input image for the face detection model and the face key point detection model.
[0013] According to a second aspect of the present disclosure, a face detection apparatus is provided, comprising: an acquisition module configured to acquire an image to be detected through a portion of an image signal processing pipeline; the image to be detected containing a face image; and a recognition module configured to recognize a face region in the image to be detected using a face detection model to obtain a face image bounding box and determine a face detection confidence level; the face detection confidence level characterizing the accuracy of the face image bounding box; the recognition module is further configured to recognize face key points in the image to be detected using a face key point detection model to determine the face image. Facial landmark confidence; the facial landmark confidence represents the accuracy of alignment between the facial landmark and the face region; the facial landmark detection model is cascaded with the face detection model; a processing module is configured to filter face image rectangles that do not include the face region based on the face detection confidence and the facial landmark confidence, and correct the position of the filtered face image rectangles by tracking the facial landmarks to obtain the processed face image rectangles; a smoothing module is configured to perform inter-frame smoothing on the processed face image rectangles to obtain the face detection result.
[0014] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to implement the steps of the aforementioned face detection method.
[0015] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the steps of the face detection method provided in the first aspect of the present disclosure.
[0016] According to a fifth aspect of the present disclosure, a chip system is provided, comprising: a storage medium for storing instructions; and a processing circuit for executing the instructions to implement the face detection method provided in the first aspect.
[0017] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: acquiring a target image through a portion of the image signal processing pipeline; identifying the face region in the target image through a face detection model to obtain a face image bounding box and determining the face detection confidence; identifying face key points in the target image through a face key point detection model and determining the face key point confidence; the face key point confidence represents the accuracy of the alignment between the face key points and the face region; cascading the face key point detection model with the face detection model; filtering face image bounding boxes that do not include face regions based on the face detection confidence and face key point confidence; correcting the position of the filtered face image bounding boxes by tracking the face key points to obtain the processed face image bounding boxes; and performing inter-frame smoothing on the processed face image bounding boxes to obtain the face detection result. By cascading face detection and facial landmark detection, the system enables the filtering and correction of face detection results based on the facial landmark detection results. This ensures detection accuracy and stability while allowing deployment on intelligent terminal devices with limited computing power. It improves the performance of face-related tasks such as face focusing, face recognition, and face beautification, thereby enhancing the user experience.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0020] Figure 1 This is a schematic diagram of a terminal shown in an exemplary embodiment of this disclosure.
[0021] Figure 2 This is a flowchart illustrating an exemplary embodiment of the present disclosure of a face detection method.
[0022] Figure 3 This is a flowchart of a sub-step of step S102 as illustrated in an exemplary embodiment of this disclosure.
[0023] Figure 4 This is a schematic diagram illustrating the relationship between the confidence level and offset of facial key points, as shown in an exemplary embodiment of this disclosure.
[0024] Figure 5 This is a schematic diagram illustrating the model performance of an exemplary embodiment of this disclosure.
[0025] Figure 6 This is a block diagram illustrating a face detection device according to an exemplary embodiment.
[0026] Figure 7This is a block diagram illustrating an apparatus for face detection according to an exemplary embodiment.
[0027] Explanation of reference numerals in the attached figures
[0028] 120 - Terminal; 20 - Face detection device; 201 - Acquisition module; 202 - Recognition module; 203 - Processing module; 204 - Smoothing module; 800 - Device; 802 - Processing component; 804 - Memory; 806 - Power component; 808 - Multimedia component; 810 - Audio component; 812 - Input / output (I / O) interface; 814 - Sensor component; 816 - Communication component. Detailed Implementation
[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0030] It should be noted that all actions involving the acquisition of signals, information, or data (such as facial data used for facial recognition) in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the device is located, and with authorization from the owner of the relevant device.
[0031] In related technologies, face recognition solutions often focus on the algorithmic complexity and performance of the detection model itself, neglecting the overall algorithmic architecture and failing to effectively address the latency issue of detection results. Currently, in intelligent terminal devices with limited computing power, no effective face detection and facial landmark detection technology has been found that achieves high accuracy, high stability, low power consumption, and low latency. Similar technologies cannot simultaneously achieve the aforementioned high accuracy, high stability, low power consumption, and low latency. Their pursuit of high accuracy often results in high power consumption due to the use of complex algorithms, while low-power, low-complexity algorithms cannot guarantee high detection accuracy. Therefore, this disclosure provides a face detection method that effectively balances the above indicators, achieving high accuracy, high stability, low power consumption, and low latency face detection and facial landmark detection.
[0032] Figure 1 A schematic diagram of a terminal provided in an exemplary embodiment of this disclosure is shown.
[0033] Terminal 120 may include at least one of smartphones, laptops, desktop computers, tablets, smart speakers, and smart robots.
[0034] Terminal 120 can be used to capture or record raw images. Terminal 120 includes a display, which can be used to display the raw images or the face detection results.
[0035] Terminal 120 includes a first memory and a first processor. The first memory stores a first program; the first program is invoked and executed by the first processor to implement the face detection method provided in this disclosure. The first memory may include, but is not limited to, the following: Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), and Electrically Erasable Programmable Read-Only Memory (EEPROM).
[0036] The first processor can consist of one or more integrated circuit chips. Optionally, the first processor can be a general-purpose processor, such as a central processing unit (CPU) or a network processor (NP). Optionally, the first processor can implement the face detection method provided in this disclosure by calling a pre-trained face detection model and a facial landmark model. For example, the pre-trained face detection model and facial landmark model in the terminal can be trained by the terminal; or, trained by the server and obtained by the terminal from the server.
[0037] An exemplary embodiment of this disclosure illustrates a face detection method comprising: acquiring a target image through a portion of an image signal processing pipeline, wherein the image signal processing pipeline is an ISP pipeline used to perform denoising, scaling, gamma processing, and downsampling on the original image to obtain the target image; identifying face regions in the target image using a face detection model to obtain face image bounding boxes, and determining face detection confidence, wherein face detection confidence characterizes the accuracy of the face image bounding boxes; then identifying face key points in the target image using a face key point detection model, and determining face key point confidence, wherein face key point confidence characterizes the certainty of alignment between face key points and face regions, and the face key point detection model and the face detection model are functionally cascaded; filtering face image bounding boxes that do not include face regions based on the face detection confidence and face key point confidence, correcting the position of the filtered face image bounding boxes by tracking face key points, and obtaining processed face image bounding boxes; and performing inter-frame smoothing on the processed face image bounding boxes to obtain a face detection result. By cascading face detection and facial landmark detection, the system can filter and correct face detection results based on the facial landmark detection results, ensuring detection accuracy while achieving high stability. It can be deployed on intelligent terminal devices with limited computing power, improving the performance of face-related tasks such as face focusing, face recognition, and face beautification, thereby enhancing the user experience.
[0038] Figure 2 This is a flowchart illustrating a face detection method according to an exemplary embodiment, the method being performed by a computer device, for example, by... Figure 1 The terminal shown is used to execute. Figure 2 The face detection method shown includes the following steps:
[0039] In step S101, the image to be detected is acquired through a portion of the image signal processing pipeline modules.
[0040] An Image Signal Processor (ISP) pipeline consists of RAW, RGB, and YUV modules. These modules perform denoising, scaling, gamma processing, and downsampling on the raw image to obtain the image to be detected. It's important to note that the order in which these processes are performed does not affect the resulting image. Because image signal processing involves large amounts of data and strict real-time requirements, ISPs are typically implemented in hardware. These modules in the ISP pipeline are interconnected and operate simultaneously at high speed under clock control. Image data is continuously transferred from one module to the next until all algorithmic processing is completed, and finally, the data flows out of the ISP pipeline in YUV or RGB format.
[0041] If a complete ISP pipeline is used to output a YUV image, and then the Y channel is downsampled, and the downsampled single-channel image is sent to a face detection model and a face landmark model for face detection and landmark detection, and then the detection results are sent to other functions, such as face autofocus and FaceID, there will be a multi-frame delay in the detection results compared to real-time video display frames.
[0042] In one implementation, a simplified ISP pipeline can be used to process the Bayer Pattern format image obtained from the CMOS into a Y single-channel image, which can then be used as the image to be detected. In this disclosure, the obtained image to be detected is written into on-chip memory (OCM) accessible by the Neural-Network Processing Unit (NPU).
[0043] For example, some modules of the ISP pipeline can be used to process the original image, such as using only one of the RAW, RGB, and YUV modules to process the original image. For instance, the RAW module can be used to perform noise reduction, scaling, gamma processing, and downsampling on the original image, thereby achieving low latency.
[0044] The original image mentioned above is an unprocessed image acquired by the terminal. The image to be detected is the image used for face detection and facial landmark detection, that is, it serves as the input image for the face detection model and the facial landmark detection model, and the image to be detected contains a face image.
[0045] In step S102, the face region in the image to be detected is identified by the face detection model to obtain the face image bounding box and determine the face detection confidence.
[0046] The NPU reads the image to be detected from the OCM and uses it as the input image for the face detection model.
[0047] The face detection model is pre-trained, either by the terminal itself or by a server, from which the terminal obtains the data. This model is used to identify face regions in an image to be detected, generating face image bounding boxes that contain the face regions.
[0048] It should be noted that step S102 also includes sub-steps S1021, S1022, and S1023. The specific method for obtaining the rectangular bounding box of the face image through the face detection model will be described in detail in the sub-steps of step S102. Please refer to [link / reference]. Figure 3 , Figure 3 This is a flowchart of a sub-step of step S102 as illustrated in an exemplary embodiment of this disclosure.
[0049] In sub-step S1021, for the image to be detected, one frame is selected every fixed frame as the input image for the face detection model.
[0050] Face detection and facial landmark detection are completely independent functions. However, in intelligent terminal devices with limited computing power, to deploy face detection and balance performance and power consumption, multiple frames of images to be detected can be processed. One frame can be selected every fixed interval for face detection, and the face detection confidence score for that frame can be determined. Facial landmark detection and face tracking are performed on every frame, thereby reducing the complexity of the face detection algorithm. For example, one frame can be selected every 5 frames, or every 8 frames, or other reasonable values; this disclosure does not limit this. Assuming that one frame is selected every 5 frames for face detection, the four frames that were not used for face detection reuse the bounding boxes and face detection confidence scores of the frames that were used for face detection.
[0051] In sub-step S1022, the face region and other regions in the input image are determined by using the pre-set classification threshold in the face detection model.
[0052] The classification threshold represents the probability that a region in the input image belongs to a face region, or it can also be called the face classification threshold.
[0053] Selecting one frame from multiple images to be detected for face detection reduces the complexity of the face detection algorithm, but the face detection performance will also decrease accordingly. In order to reduce the false detection rate, the binary classification threshold for face detection will be relatively large. For the original images obtained in some extreme scenarios, such as large face pose, backlight, low light, occlusion, etc., serious face false detection will occur.
[0054] The classification threshold in related technologies in this field is usually 0.55, that is, if the probability that a certain region in the image to be detected is a face region is greater than 0.55, then the region is considered to be a face region. However, this cannot solve the problem of missed face detection in extreme scenarios.
[0055] In this disclosure, the classification threshold ranges from 0 to 0.55. In one embodiment, the classification threshold is set to 0.25 to achieve a higher face detection recall rate. Then, in subsequent steps, the problem of false face detection caused by a low classification threshold is addressed.
[0056] In sub-step S1023, the smallest rectangular region containing all face regions is taken as the face image bounding box.
[0057] For example, taking a classification threshold of 0.25 as an example, the region in the image to be detected with a probability greater than 0.25 is taken as the face region, and then the smallest rectangular region containing all face regions is taken as the face image bounding box.
[0058] In one implementation, the bounding box of the face image obtained by the face detection model can be saved to Double Data Rate Synchronous Dynamic Random-Access Memory (DDR).
[0059] In step S103, facial key points in the image to be detected are identified by the facial key point detection model, and the confidence level of the facial key points is determined.
[0060] The NPU reads the face image bounding box from the OCM and the image to be detected that has not been recognized by the face detection model as the input images for the face key point detection model.
[0061] This facial landmark detection model is pre-trained, either by the terminal itself or by a server, with the terminal obtaining the data from the server. This model is used to identify facial landmarks in an image and determine their confidence levels.
[0062] Facial landmark detection locates the positional information of facial features, including eyebrows, eyes, nose, mouth, and facial contours. The confidence score of a facial landmark ranges from 0 to 1.0, representing the certainty of alignment between the landmark and the facial region. A higher confidence score indicates better alignment, while a lower score indicates poorer alignment. For example, if both the facial landmark detection model and the face detection model identify a region as a nose, then the certainty of alignment between the landmark and the facial region is relatively high.
[0063] For each frame in the image to be detected, the facial landmarks in each frame are identified by the facial landmark detection model, and the confidence level of the facial landmarks is determined.
[0064] In one implementation, the facial landmarks detected by the facial landmark detection model and their confidence scores can be saved to DDR.
[0065] It should be noted that the facial landmark detection model and the face detection model are functionally cascaded. The facial landmark detection model is trained through the following steps:
[0066] Step 1: Obtain sample images. These sample images include a specified proportion of non-face images, i.e., images that do not contain faces. The non-face image samples are used to ensure that the confidence score of facial landmarks, while characterizing the certainty of the alignment between facial landmarks and face regions, also has the ability to distinguish non-face images. For example, for facial landmark detection in non-face images, the confidence score value is close to 0 or equal to 1.0.
[0067] Step 2: Extract the face region from the sample image by expanding the rectangular box using probabilistic random offset.
[0068] In related technologies, a centered extended rectangle is typically used to crop the face region, or region of interest, in a sample image. In this disclosure, a probabilistically random offset extended rectangle is used to crop the face region in the sample image. That is, when cropping the face region, not only is the smallest rectangle containing the entire face region cropped, but also an offset extended rectangle containing only a portion of the face is cropped using a probabilistically random offset. Specifically, the centered extended rectangle is probabilistically and randomly offset around the face, ensuring that a portion of the face is contained within the offset extended rectangle. It should be noted that the size of the offset extended rectangle is the same as the size of the centered extended rectangle. Finally, the face region and facial landmarks in the sample image are labeled to obtain a labeled sample; the face region in the sample image is labeled using the offset extended rectangle and / or the centered extended rectangle, and the offset degree phi of the offset extended rectangle is recorded. Facial landmarks are then labeled using points.
[0069] Step 3: Train the facial landmark detection model using labeled samples.
[0070] The facial landmark detection model is trained using the sample images from step two.
[0071] In step S104, face image rectangles that do not include face regions are filtered based on face detection confidence and face key point confidence. The positions of the filtered face image rectangles are corrected by tracking the face key points to obtain the processed face image rectangles.
[0072] In the aforementioned steps, a lower classification threshold (e.g., 0.25) was used to achieve a higher face detection recall rate, that is, to recall more face image bounding boxes that may not contain faces. This step addresses the false face detection problem introduced by the low classification threshold. As mentioned earlier, face detection confidence represents the accuracy of the face image bounding box, while facial landmark confidence represents the certainty of alignment between facial landmarks and the face region. Therefore, face image bounding boxes that do not include face regions can be filtered based on face detection confidence and facial landmark confidence. The lower the face detection confidence, the greater the positional deviation within the face image bounding box; the higher the face detection confidence, the smaller the positional deviation. The lower the facial landmark confidence, the worse the alignment between the facial landmarks and the face region, indicating a higher probability of false face detection within the bounding box. Therefore, the face detection confidence and facial landmark confidence corresponding to each frame's face image bounding box can be obtained. When both the face detection confidence and facial landmark confidence are less than a predetermined threshold, the corresponding face image bounding box is filtered out. This achieves high-precision face bounding box detection and facial landmark detection, effectively improving the performance of face detection and facial landmark detection in general scenes, especially in extreme scenes.
[0073] The aforementioned predetermined thresholds can be obtained based on experience or other reasonable methods, and this disclosure does not impose any restrictions on them.
[0074] It should be noted that the facial landmark detection model is also used to track facial landmarks. By tracking these landmarks, the position of the bounding box in the face image is corrected, resulting in a processed face image bounding box. Please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram illustrating the relationship between facial landmark confidence and offset, as shown in an exemplary embodiment of this disclosure. The horizontal axis represents the offset phi, and the vertical axis represents the facial landmark confidence. Assuming the face bounding box is a square with a side length denoted as box_size, the expanded face bounding box has a side length of roi_size = 1.3 * box_size; the face bounding box offset = phi * roi_size, as shown... Figure 4As shown, when the facial landmark alignment confidence threshold is set to 0.55, the facial landmark detection position offset can reach 0.5*roi_size. Since the real-time video frame rate is high (e.g., 30fps), the position offset of the moving face between video stream frames is basically less than 0.5*roi_size. Therefore, the facial landmark detection model can handle the face tracking task.
[0075] It should be noted that the two steps of filtering out the face image rectangle and correcting the position of the face image rectangle are not sequential. That is, you can filter out the face image rectangle first and then correct the position of the filtered face image rectangle, or you can correct the position of the face image rectangle first and then filter out the face image rectangle after the position is corrected.
[0076] In step S105, the processed face image bounding box is smoothed between frames to obtain the face detection result.
[0077] In the previous step, the face image bounding boxes obtained by filtering out the face image bounding boxes and correcting their positions are smoothed between frames to remove jitter, resulting in the final face detection result, which can be used for tasks such as face focusing, face recognition, and face beautification.
[0078] In one implementation, the smoothing module can be deployed in the CPU to perform smoothing processing of face detection and facial landmark detection results, thereby improving the flexibility of the face detection method of this disclosure and facilitating the scalability of face detection function maintenance in the future.
[0079] Please see Figure 5 , Figure 5This is a schematic diagram illustrating the performance of the model as shown in the exemplary embodiments of this disclosure. The model refers to both a face detection model and a facial landmark detection model. The vertical axis represents the detection accuracy of the model, and the horizontal axis represents the number of images to be detected. The figure also shows three sets of experimental data for backlit scenes at different distances (1m, 2m, 3m). Each set of data contains 180 face images; the size of the face in the image varies with distance (the farther the distance, the smaller the face size). The figure demonstrates that the facial landmark detection model has better stability and robustness than the face detection model. If conventional methods in this field are used for face detection, and the face classification threshold is set to 0.55, the face detection accuracy is only 1.67% in a 3m backlit scene, resulting in a false negative rate of up to 98.33%. However, by using the cascaded face detection and facial landmark detection functions of this disclosure, in one embodiment, the face classification threshold is set to 0.2, which achieves a high recall rate for face detection, and the facial landmark alignment confidence threshold is set to 0.55 to filter out possible false positives. This allows the face detection accuracy at 3m to reach 91.11%, and the face detection rate at 2m to increase from 96.67% to 100%, greatly improving the face detection performance in extreme scenarios.
[0080] In summary, the face detection method provided in this disclosure includes: acquiring a target image through a portion of an image signal processing pipeline; identifying face regions in the target image using a face detection model to obtain face image bounding boxes and determining face detection confidence; identifying face key points in the target image using a face key point detection model and determining face key point confidence; the face key point confidence represents the accuracy of alignment between face key points and face regions; cascading the face key point detection model and the face detection model; filtering face image bounding boxes that do not include face regions based on the face detection confidence and face key point confidence; correcting the position of the filtered face image bounding boxes by tracking face key points to obtain processed face image bounding boxes; and performing inter-frame smoothing on the processed face image bounding boxes to obtain the face detection result. By cascading face detection and facial landmark detection, the system enables the filtering and correction of face detection results based on the facial landmark detection results. This ensures detection accuracy and stability while allowing deployment on intelligent terminal devices with limited computing power. It improves the performance of face-related tasks such as face focusing, face recognition, and face beautification, thereby enhancing the user experience.
[0081] Figure 6 This is a block diagram illustrating a face detection device according to an exemplary embodiment. (Refer to...) Figure 6 The device 20 includes an acquisition module 201, an identification module 203, a processing module 205, and a smoothing module 207.
[0082] The acquisition module 201 is configured to acquire the image to be detected through a portion of the image signal processing pipeline;
[0083] The recognition module 203 is configured to identify the face region in the image to be detected through a face detection model, obtain a face image bounding box, and determine the face detection confidence level; the face detection confidence level characterizes the accuracy of the face image bounding box.
[0084] The recognition module 203 is also configured to identify facial key points in the image to be detected through a facial key point detection model, and determine the confidence level of the facial key points; the confidence level of the facial key points represents the accuracy of the alignment between the facial key points and the face region; the facial key point detection model is cascaded with the face detection model;
[0085] The processing module 205 is configured to filter face image rectangles that do not include the face region based on the face detection confidence and the face key point confidence, and to correct the position of the filtered face image rectangles by tracking the face key points to obtain the processed face image rectangles.
[0086] The smoothing module 207 is configured to perform inter-frame smoothing on the processed face image bounding box to obtain the face detection result.
[0087] Optionally, the acquisition module 201 is further configured to perform denoising, scaling, gamma processing, and downsampling on the original image through the RAW module in the image signal processing pipeline to obtain the image to be detected.
[0088] Optionally, the recognition module 203 is further configured to select one frame at fixed intervals from the image to be detected as the input image of the face detection model;
[0089] The face region and other regions in the input image are determined by using a pre-set classification threshold in the face detection model; the classification threshold represents the probability that a region in the input image belongs to the face region.
[0090] The smallest rectangular region containing all the face regions is taken as the face image rectangle.
[0091] Optionally, the preset classification threshold can be set to a value between 0 and 0.55.
[0092] Optionally, the recognition module 203 is further configured to identify facial key points through a facial key point detection model for each frame of the face image bounding box and the image to be detected that has not been recognized by the face detection model, and determine the confidence level of the facial key points.
[0093] Optionally, the processing module 205 is further configured to obtain the face detection confidence and face key point confidence corresponding to the rectangular box of the face image in each frame;
[0094] If the confidence scores of the face detection and the facial landmarks are less than a predetermined threshold, the corresponding face image bounding boxes are filtered out.
[0095] Optionally, the facial landmark detection model is trained through the following steps:
[0096] Acquire sample images; the sample images include non-face images of a specified proportion;
[0097] The face region in the sample image is extracted by probabilistically random offset extended rectangle;
[0098] Mark the face regions and facial key points in the sample image to obtain the marked sample;
[0099] The facial landmark detection model is obtained by training the labeled samples.
[0100] Optionally, the acquisition module 201 is also configured to write the image to be detected into an on-chip memory (OCM) accessible by the network processor (NPU).
[0101] The NPU reads the image to be detected from the OCM and uses it as the input image for the face detection model and the face landmark detection model.
[0102] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0103] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the face detection method provided in this disclosure.
[0104] This disclosure also provides a chip, which includes: a storage medium for storing instructions; and a processing circuit for executing the instructions to implement the aforementioned face detection method.
[0105] Figure 7 This is a block diagram illustrating an apparatus 800 for face detection according to an exemplary embodiment. For example, apparatus 800 may be an electronic device, such as a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0106] Reference Figure 7The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0107] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the aforementioned face detection method. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0108] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of such data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0109] The power supply component 806 provides power to the various components of the device 800. The power supply component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 800.
[0110] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0111] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0112] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0113] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in the position of device 800 or a component of device 800, the presence or absence of user contact with device 800, the orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0114] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0115] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the face detection method described above.
[0116] The aforementioned electronic device can be a standalone electronic device or a part of a standalone electronic device. For example, in one embodiment, the electronic device can be an integrated circuit (IC) or a chip, wherein the integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), and SoC (System on Chip). The aforementioned integrated circuit or chip can be used to execute executable instructions (or code) to implement the aforementioned face detection method. The executable instructions can be stored in the integrated circuit or chip or obtained from other devices or equipment. For example, the integrated circuit or chip includes a processor, memory, and an interface for communicating with other devices. The executable instructions can be stored in the processor, and when the executable instructions are executed by the processor, the face detection method described above is implemented; alternatively, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution to implement the face detection method described above.
[0117] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of the device 800 to complete the face detection method described above. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0118] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the face detection method described above when executed by the programmable device.
[0119] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0120] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A face detection method, characterized in that, include: The image to be detected is obtained through a portion of the image signal processing pipeline modules; For the image to be detected, a frame is selected every fixed frame as the input image for the face detection model, and the face region in the input image is identified by the face detection model to obtain the face image bounding box, and the face detection confidence is determined; the face detection confidence represents the accuracy of the face image bounding box. The facial landmark detection model is used to identify the rectangular bounding box of the face image and the facial landmarks in the image to be detected that have not been identified by the face detection model, and the confidence level of the facial landmarks is determined. The confidence level of the facial key points represents the degree of certainty that the facial key points are aligned with the facial region; The facial landmark detection model is cascaded with the facial detection model; Based on the face detection confidence and the face key point confidence, filter out face image rectangles that do not include the face region, and correct the position of the filtered face image rectangles by tracking the face key points to obtain the processed face image rectangles. Inter-frame smoothing is performed on the processed face image bounding box to obtain the face detection result.
2. The method according to claim 1, characterized in that, The step of acquiring the image to be detected through a portion of the image signal processing pipeline includes: The image to be detected is obtained by performing noise reduction, scaling, gamma processing, and downsampling on the original image through the RAW module in the image signal processing pipeline.
3. The method according to claim 1, characterized in that, The step of identifying the face region in the input image using the face detection model to obtain the face image bounding box includes: The face region and other regions in the input image are determined by using a pre-set classification threshold in the face detection model; the classification threshold represents the probability that a region in the input image belongs to the face region. The smallest rectangular region containing all the face regions is taken as the face image rectangle.
4. The method according to claim 3, characterized in that, The preset classification threshold ranges from 0 to 0.
55.
5. The method according to claim 1, characterized in that, The step of identifying the rectangular bounding box of the face image and the facial landmarks in the image to be detected that have not been identified by the face detection model, and determining the confidence level of the facial landmarks, includes: For each frame of the face image bounding box and the image to be detected that has not been identified by the face detection model, face key points are identified by the face key point detection model, and the confidence level of the face key points is determined.
6. The method according to claim 1, characterized in that, The step of filtering out rectangular bounding boxes of face images that do not include the face region based on the face detection confidence score and the face key point confidence score includes: Obtain the face detection confidence score and face landmark confidence score corresponding to the rectangular bounding box of the face image in each frame; If both the face detection confidence score and the face key point confidence score are less than a predetermined threshold, the corresponding face image rectangle is filtered out.
7. The method according to claim 1, characterized in that, The facial landmark detection model is trained through the following steps: Acquire sample images; the sample images include non-face images of a specified proportion; The face region in the sample image is extracted by probabilistically random offset extended rectangle; Mark the face regions and facial key points in the sample image to obtain the marked sample; The facial landmark detection model is obtained by training the labeled samples.
8. The method according to any one of claims 1-7, characterized in that, Following the step of acquiring the image to be detected through a portion of the image signal processing pipeline, the following is included: The image to be detected is written into the on-chip memory (OCM) accessible by the network processor (NPU); The NPU reads the image to be detected from the OCM and uses it as the input image for the face detection model and the face landmark detection model.
9. A face detection device, characterized in that, include: The acquisition module is configured to acquire the image to be detected through a portion of the image signal processing pipeline. The recognition module is configured to select one frame at fixed intervals as the input image for the face detection model from the image to be detected, and to identify the face region in the input image through the face detection model to obtain a face image bounding box and determine the face detection confidence; the face detection confidence represents the accuracy of the face image bounding box. The recognition module is further configured to identify the rectangular frame of the face image and the face key points in the image to be detected that have not been identified by the face detection model through the face key point detection model, and determine the confidence level of the face key points. The confidence level of the facial landmarks represents the accuracy of the alignment between the facial landmarks and the facial region; The facial landmark detection model is cascaded with the facial detection model; The processing module is configured to filter out face image rectangles that do not include the face region based on the face detection confidence and the face key point confidence, and to correct the position of the filtered face image rectangles by tracking the face key points to obtain the processed face image rectangles. The smoothing module is configured to perform inter-frame smoothing on the processed face image bounding box to obtain the face detection result.
10. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the executable instructions to implement the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 8.
12. A chip system, characterized in that, include: Storage medium, used to store instructions; A processing circuit is used to execute the instructions to implement the face detection method as described in any one of claims 1-8.
Citation Information
Patent Citations
Face alignment detection method and device
CN110059637A
Information output method and device
CN111325050A
Target detection method and device for driver state monitoring
CN111814568A