Face Detection Method and Device, Computer Readable Storage Medium, and Terminal Device
By combining traditional machine learning and deep learning models, traditional machine learning models are used to preprocess images and input deep learning models for face detection, solving the problems of low accuracy of traditional methods and large computing volume of deep learning methods, and achieving efficient face detection on edge devices.
Patent Information
- Application Number
- CN202111082423.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-09-15
AI Technical Summary
In the prior art, the face detection method based on traditional machine learning has low accuracy and is prone to false detection. However, the face detection method based on deep learning has a large amount of calculation and is not suitable for use on edge devices with limited computing power and low power consumption requirements.
Combining traditional machine learning models and deep learning models, the images to be detected are preprocessed through traditional machine learning models, face images are extracted and input into deep learning models for detection, and secondary detection is used for use in deep learning models to improve accuracy and reduce the amount of calculation.
It improves the accuracy of face detection, reduces the computing volume of deep learning models, and improves detection speed. It is suitable for use on edge devices with limited computing power and low power consumption requirements.
Smart Images

Figure CN113780202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a face detection method and apparatus, a computer-readable storage medium, and a terminal device. Background Art
[0002] The goal of face detection is to determine the positions, sizes, and poses of all faces in an input image. It is widely used in fields such as digital camera autofocus, video surveillance, and authentication, and has great commercial value.
[0003] Currently, the main face detection methods are divided into face detection methods based on traditional machine learning and face detection methods based on deep learning. When traditional machine learning is used for face detection, it has the characteristic of relatively fast detection speed. The detection accuracy of face detection methods based on deep learning is relatively high.
[0004] However, face detection methods based on traditional machine learning have low accuracy and are prone to false detection; the computational complexity of deep learning face detection methods is relatively large, and they are not suitable for use on edge devices with limited computing power and low power consumption requirements. Summary of the Invention
[0005] The technical problem solved by the present invention is how to balance the accuracy of face detection and reduce the computational complexity.
[0006] To solve the above technical problem, an embodiment of the present invention provides a face detection method, which includes: obtaining an image to be detected; preprocessing the image to be detected by using a traditional machine learning model to obtain a face image, where the face image is a part of the image to be detected; inputting the face image into a deep learning model to obtain a detection result, where the detection result includes face information.
[0007] Optionally, the preprocessing the image to be detected by using a traditional machine learning model includes: extracting a face in the image to be detected by using the traditional machine learning model to obtain the face image and its corresponding face confidence.
[0008] Optionally, the inputting the face image into a deep learning model includes: when the face confidence is less than a first threshold, inputting the face image into the deep learning model.
[0009] Optionally, the detection result further includes a detection confidence, and the method further includes: if the detection confidence reaches a second threshold, retaining the face in the detection result; if the detection confidence is lower than the second threshold, deleting the face in the detection result.
[0010] Optionally, the number of extracted faces is one or more, and each face corresponds to a face image.
[0011] Optionally, the deep learning model is trained in the following manner: obtaining training samples, where the training samples include a plurality of positive samples and negative samples, the positive samples are images containing only the face part, and the negative samples are images not containing a face; and training the deep learning model using the training samples.
[0012] Optionally, the traditional machine learning model includes an adaboos detector, and the deep learning model includes an onet detector.
[0013] To solve the above technical problem, an embodiment of the present invention also discloses a face detection device, which includes: an acquisition module for acquiring an image to be detected; a preprocessing module for preprocessing the image to be detected using a traditional machine learning model to obtain a face image, where the face image is a part of the image to be detected; and a detection module for inputting the face image into a deep learning model to obtain a detection result, where the detection result includes face information.
[0014] An embodiment of the present invention also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the face detection method.
[0015] An embodiment of the present invention also discloses a terminal device, including a memory and a processor, where a computer program that can run on the processor is stored on the memory, and when the processor runs the computer program, it executes the steps of the face detection method.
[0016] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:
[0017] In the technical solution of the present invention, for an image to be detected, a traditional machine learning model can be used for preprocessing, and the obtained face image after preprocessing is input into a deep learning model for face detection. Compared with using a single detection model for face detection, the detection accuracy is improved; in addition, since the image to be detected is input into the deep learning model after preprocessing, the face image input into the deep learning model is a part of the image to be detected, so the computational amount of the deep learning model becomes smaller, the complexity of the model is reduced, and the detection speed is also improved. Description of the Drawings
[0018] Figure 1 is a flowchart of a face detection method in an embodiment of the present invention;
[0019] Figure 2 is a specific flowchart of a face detection method in an embodiment of the present invention;
[0020] Figure 3 This is a schematic structural diagram of a face detection device in an embodiment of the present invention. Detailed implementation manners
[0021] As described in the background art, the accuracy of face detection methods based on traditional machine learning is relatively low, and false detection is likely to occur; the computational complexity of deep learning face detection methods is relatively large, and they are not suitable for use on edge devices with limited computing power and low power consumption requirements.
[0022] In the technical solution of the present invention, for an image to be detected, a traditional machine learning model can be used for preprocessing, and the face image obtained by the preprocessing is input into a deep learning model for face detection. Compared with using a single detection model for face detection, the detection accuracy is improved; in addition, since the image to be detected is input into the deep learning model after preprocessing, the face image input into the deep learning model is a part of the image to be detected, so the computational complexity of the deep learning model becomes smaller, the complexity of the model is reduced, and the detection speed is also improved.
[0023] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.
[0024] Figure 1 This is a flowchart of a face detection method in an embodiment of the present invention.
[0025] The face detection method in the embodiment of the present invention can be used on the terminal device side, that is, each step of the face detection method can be executed by the terminal device.
[0026] Specifically, Figure 1 The face detection method shown may include the following steps:
[0027] Step 101: Obtain an image to be detected;
[0028] Step 102: Use a traditional machine learning model to preprocess the image to be detected to obtain a face image, where the face image is a part of the image to be detected;
[0029] Step 103: Input the face image into a deep learning model to obtain a detection result, where the detection result includes face information.
[0030] It should be noted that the sequence numbers of the steps in this embodiment do not represent the limitation of the execution sequence of each step.
[0031] It can be understood that in specific implementation, the face detection method can be implemented in the form of a software program, and the software program runs in a processor integrated inside a chip or a chip module.
[0032] In the specific implementation of step 101, the image to be detected refers to the image in which it is necessary to detect whether there is a human face. The image to be detected can specifically be a photo or a frame of a video. The image to be detected can be obtained by the terminal device through shooting or retrieved by the terminal device from a database, where multiple images and / or multiple videos are stored.
[0033] In the specific implementation of step S102, a traditional machine learning model can be first used to preprocess the image to be detected. After preprocessing by the traditional machine learning model, the human face in the image to be detected can be located, that is, the position of the human face in the image is determined, which can be specifically represented by coordinates. That is to say, after preprocessing, a face image is obtained. The face image refers to the area in the image to be detected that only contains the human face and is a part of the image to be detected.
[0034] Specifically, the face image can be the area occupied by the human face in the image to be detected, which can specifically be a rectangular area.
[0035] Furthermore, in the specific implementation of step S103, the face image can be used as the input image of the deep learning model. The deep learning model outputs a detection result, and the detection result includes the detected human face information. Compared with directly inputting the image to be detected into the deep learning model, the input image of the deep learning model is smaller, the computational amount of the deep learning model becomes smaller, and the complexity of the deep learning model is also reduced, thus improving the efficiency of human face detection.
[0036] Specifically, the process of the deep learning model detecting the face image can be processes such as face recognition and face comparison after face location, and the human face information is determined, such as face feature points, human face identity, etc.
[0037] It can be understood that the traditional machine learning module and the deep learning model can be pre-constructed. Among them, the traditional machine learning model can be constructed using traditional machine learning algorithms, such as algorithms like k-nearest neighbor method, naive Bayes, decision tree, support vector machine, AdaBoost method, hidden Markov model, conditional random field, etc.; the deep learning model can be a model constructed based on a deep neural network, such as an onet detector.
[0038] Compared with using a single traditional machine learning model for human face detection, the embodiment of the present invention uses a deep learning model for human face detection, and the detection accuracy is improved. In addition, compared with using a single deep learning model for human face detection, the embodiment of the present invention preprocesses the image to be detected, reducing the computational amount of the deep learning model and improving the detection speed.
[0039] In an embodiment of the present invention, a face detection method based on deep learning is combined with a face detection method based on traditional machine learning. The face detection method based on deep learning enhances the face detection method based on traditional machine learning, and greatly improves the false detection phenomenon of the face detection method based on traditional machine learning with a very small increase in computational complexity, thus enhancing the user experience.
[0040] In a specific embodiment, in the preprocessing stage, a traditional machine learning model is used to extract the faces in the image to be detected, and a face image and its corresponding face confidence can be obtained. Among them, the face confidence of the face image represents the probability that there is a face in the face image.
[0041] It should be noted that the calculation method of the face confidence can refer to the prior art, and the embodiments of the present invention do not limit this.
[0042] Furthermore, in the preprocessing stage, the number of extracted faces is one or more, and each face corresponds to a face image. Each face also corresponds to a face confidence.
[0043] Specifically, the image to be detected may include one face or multiple faces. In the case where the image to be detected includes multiple faces, each face is extracted by the traditional machine learning model, and multiple face images are obtained, and each face image includes one face.
[0044] In a specific embodiment, Figure 1 Step 103 shown may include the following steps: when the face confidence is less than the first threshold, the face image is input into the deep learning model.
[0045] In this embodiment, the face confidence of the face image being less than the first threshold indicates that the probability of there being a face in the face image is relatively small. Considering that the false detection rate of the traditional machine learning model is relatively high and the face detection rate of the deep learning model is relatively high, the face image can be input into the deep learning model for secondary detection.
[0046] That is to say, when the face confidence of the face image is greater than or equal to the first threshold, it indicates that the probability of there being a face in the face image is relatively large, and the face image may not be input into the deep learning model. At this time, the face image output by the traditional machine learning model can be directly output as the final detection result.
[0047] In a specific embodiment, the face detection method may further include the following steps: if the detection confidence reaches the second threshold, the face in the detection result is retained; if the detection confidence is lower than the second threshold, the face in the detection result is deleted.
[0048] In this embodiment, when the deep learning model outputs the detection result, it can also output the detection confidence. The detection confidence is for the input face image, that is, each face image has a corresponding detection confidence. The detection confidence can represent the probability that a face exists in the face image.
[0049] Specifically, when the detection confidence reaches the second threshold, it indicates that the probability of a face existing in the face image is relatively high, and then the face in the detection result can be retained. Correspondingly, when the detection confidence is lower than the second threshold, it indicates that the probability of a face existing in the face image is relatively low, and then the face in the detection result can be deleted. At this time, it also means that there is no face in this face image.
[0050] In a specific embodiment, please refer to Figure 2 , Figure 2 which shows the specific process of a face detection method.
[0051] Step 201, obtain the image to be detected.
[0052] Step 202, preprocess the image to be detected by using a traditional machine learning model. The detection result output by the traditional machine learning model includes a face image and a face confidence.
[0053] Step 203, determine whether the face confidence is less than the first threshold. If so, go to step 203; otherwise, go to step 206.
[0054] Step 204, input the face image into the deep learning model. The detection result output by the deep learning model includes face information and a detection confidence.
[0055] Step 205, determine whether the detection confidence is greater than the second threshold. If so, go to step 206; otherwise, go to step 208.
[0056] Step 206, retain the detection result. When the face confidence is greater than the first threshold, the retained detection result is the face image output by the traditional machine learning model. When the detection confidence is greater than the second threshold, the retained detection result is the face information output by the deep learning model.
[0057] Step 207, output the detection result.
[0058] Step 208, delete the detection result.
[0059] The embodiment of the present invention utilizes the relatively strong feature representation ability of the traditional machine learning model to optimize the misdetection problem of the traditional machine learning model and improve the accuracy of face detection.
[0060] In a specific application scenario of the present invention, the traditional machine learning model includes an AdaBoost detector, and the deep learning model includes an ONet detector. First, use the AdaBoost detector to perform face detection on the input image; secondly, judge the AdaBoost face detection result. If the confidence of the face is lower than the first threshold (check_threshold), then use the ONet detector to perform secondary detection on the face area. If the confidence of the secondary detection by the ONet detector is greater than the second threshold (face_threshold), then retain the face, otherwise delete the face.
[0061] Please refer to Figure 3 , the embodiment of the present invention also discloses a face detection device 30. The face detection device 30 may include:
[0062] An acquisition module 301, configured to acquire an image to be detected;
[0063] A preprocessing module 302, configured to preprocess the image to be detected by using a traditional machine learning model to obtain a face image, where the face image is a part of the image to be detected;
[0064] A detection module 303, configured to input the face image into a deep learning model to obtain a detection result, where the detection result includes face information.
[0065] In a specific implementation, the above face detection device may correspond to a chip with a face detection function in a terminal device, such as a SOC (System-On-a-Chip), a baseband chip, etc.; or correspond to a chip module including a chip with a face detection function in a terminal device; or correspond to a chip module with a data processing function chip, or correspond to a terminal device.
[0066] For more content about the working principle and working mode of the face detection device 30, please refer to Figures 1 to 2 the relevant description in, which will not be elaborated here.
[0067] Regarding each device and product described in the above embodiments, each module / unit included therein can be a software module / unit, a hardware module / unit, or can be partly a software module / unit and partly a hardware module / unit. For example, for each device and product applied to or integrated into a chip, each module / unit included therein can be implemented in the form of hardware such as circuits, or at least part of the module / unit can be implemented in the form of a software program that runs on a processor integrated inside the chip, and the remaining (if any) part of the module / unit can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a chip module, each module / unit included therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module, or at least part of the module / unit can be implemented in the form of a software program that runs on a processor integrated inside the chip module, and the remaining (if any) part of the module / unit can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a terminal, each module / unit included therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal, or at least part of the module / unit can be implemented in the form of a software program that runs on a processor integrated inside the terminal, and the remaining (if any) part of the module / unit can be implemented in the form of hardware such as circuits.
[0068] An embodiment of the present invention also discloses a storage medium, which is a computer-readable storage medium, and has a computer program stored thereon. When the computer program runs, it can execute Figure 1 or Figure 2 the steps of the face detection method shown in
[0069] An embodiment of the present invention also discloses a terminal device, which may include a memory and a processor, and a computer program that can run on the processor is stored on the memory. When the processor runs the computer program, it can execute Figure 1 or Figure 2 the steps of the face detection method shown in
[0070] It should be understood that the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article represents that the associated objects before and after are in an "or" relationship.
[0071] In the embodiments of the present application, "a plurality of" means two or more.
[0072] In the embodiments of the present application, the first, second, etc. descriptions are only for indicating and distinguishing the described objects, without any order, nor do they represent special limitations on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.
[0073] In the embodiments of the present application, "connection" means various connection methods such as direct connection or indirect connection to achieve communication between devices, and the embodiments of the present application do not make any limitations on this.
[0074] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU for short), and this processor may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), field programmable gate arrays (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0075] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0076] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more collections of available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0077] It should be understood that in various embodiments of the present application, the order numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0078] In several embodiments provided in the present application, it should be understood that the disclosed methods, apparatuses, and systems can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical, or other forms.
[0079] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0080] In addition, in each embodiment of the present invention, each functional unit may be integrated into a processing unit, may be physically included separately for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware, or in the form of a hardware plus software functional unit.
[0081] The above-mentioned integrated unit implemented in the form of a software functional unit may be stored in a computer-readable storage medium. The above-mentioned software functional unit stored in a storage medium includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute some steps of the methods described in the various embodiments of the present invention.
[0082] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.
Claims
1. A face detection method, characterized in that, Including: Obtain the image to be detected; Use a traditional machine learning model to preprocess the image to be detected to obtain a face image and its corresponding face confidence, where the face image is a part of the image to be detected; Input the face image into a deep learning model to obtain a detection result, where the detection result includes face information; The step of inputting the face image into the deep learning model includes: When the face confidence is less than a first threshold, input the face image into the deep learning model; When the face confidence is greater than or equal to the first threshold, directly output the face image; The traditional machine learning model includes an adaboost detector, and the deep learning model includes an onet detector; use the adaboost detector to perform face detection on the input image; if the confidence of the face image is lower than the first threshold, use the onet detector to perform secondary detection on the face image, and if the detection confidence of the secondary detection by the onet detector is greater than a second threshold, retain the face in the detection result, otherwise delete the face in the detection result.
2. The face detection method according to claim 1, wherein The step of using a traditional machine learning model to preprocess the image to be detected includes: Use the traditional machine learning model to extract the face in the image to be detected to obtain the face image and its corresponding face confidence.
3. The face detection method according to claim 2, wherein The number of extracted faces is one or more, and each face corresponds to a face image.
4. The face detection method according to claim 1, wherein The deep learning model is trained in the following manner: Obtain training samples, where the training samples include a plurality of positive samples and negative samples, the positive samples are images containing only the face part, and the negative samples are images without a face; Use the training samples to train the deep learning model.
5. A face detection device, characterized in that, Including: An acquisition module for obtaining the image to be detected; A preprocessing module for using a traditional machine learning model to preprocess the image to be detected to obtain a face image and its corresponding face confidence, where the face image is a part of the image to be detected; A detection module for inputting the face image into a deep learning model to obtain a detection result, where the detection result includes face information; When the face confidence is less than the first threshold, the detection module inputs the face image into the deep learning model; When the face confidence is greater than or equal to the first threshold, the detection module directly outputs the face image; The traditional machine learning model includes an adaboost detector, and the deep learning model includes an onet detector; use the adaboost detector to perform face detection on the input image; if the confidence of the face image is lower than the first threshold, use the onet detector to perform secondary detection on the face image, and if the detection confidence of the secondary detection by the onet detector is greater than a second threshold, retain the face in the detection result, otherwise delete the face in the detection result.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by a processor, it executes the steps of the face detection method according to any one of claims 1 to 4.
7. A terminal device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored on the memory, characterized in that When the processor runs the computer program, it executes the steps of the face detection method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Face detection method based on two-stage detection
CN110189255A
Image AU detection method and device, electronic equipment and storage medium
CN110399788A