Face image detection method and device based on multiple models, and electronic equipment

By adopting a multi-model detection network in face detection, combining the first detection network and the second detection network, the problem of low detection accuracy of the prior art in side faces or blocking faces is solved, and more efficient and accurate face detection is achieved.

CN120088830APending Publication Date: 2025-06-03GUANGZHOU HUYA INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510141701.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing face detection methods have low detection accuracy when processing side face or obscured face images, and a single model is difficult to effectively detect in complex scenarios.

Method used

Using a multi-model face image detection method, the first detection network is used to initially detect whether there is a positive face. If it is not detected, the image is input to the second detection network, which includes multiple personal face detection models for detection in complex situations.

Benefits of technology

It improves the efficiency and accuracy of face detection, can better process complex images, reduce errors, and obtain more accurate face image information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088830A_ABST
    Figure CN120088830A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image detection, in particular to a face image detection method and device based on multiple models and electronic equipment, and the detection method comprises the steps: obtaining a to-be-detected image; the to-be-detected image is input into a first detection network for human face detection, human face image information is output if human face information is detected, the to-be-detected image is input into a second detection network if human face information is not detected, and the second detection network comprises a plurality of human face detection models; the plurality of face detection models respectively process the to-be-detected image to obtain a plurality of corresponding detection results; the face image information corresponding to the to-be-detected image is obtained according to the multiple detection results, the multiple face detection models are combined, face detection is carried out on the image, whether the image contains the face or not is determined, the specific expression form containing the face is determined, reliable reference data can be provided for face recognition, and the face recognition efficiency is improved. The application of face detection is promoted, and the accuracy and efficiency of detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image detection, and more specifically, to a multi-model-based face image detection method, device, and electronic device. Background Art

[0002] With the development needs of people's production activities, it is very important to locate and recognize images. Therefore, image detection technology has emerged. Currently, face detection has received relatively high attention. The commonly used detection methods are face detection networks or models. Due to the complexity of image data, single detection networks or models have many defects and high costs. In some special application scenarios, there will be special effects in image data, and the face is partially blocked, such as sunglasses special effects. Using a face detection network or model to detect such image data, the effect is not obvious, and it only has a good detection effect on the frontal face of the human face. For example, using the Multi-task Cascaded Convolutional Networks (MTCNN) and the Dlib library. Dlib is an open-source library containing machine learning algorithms and tools, mainly used in the fields of image processing and computer vision. The above face detection algorithms cannot detect human faces in side face images or partially occluded face image data.

[0003] Convolutional Networks, MTCNN) and the Dlib library. Dlib is an open-source library containing machine learning algorithms and tools, mainly used in the fields of image processing and computer vision. The above face detection algorithms cannot detect human faces in side face images or partially occluded face image data. Summary of the Invention

[0004] The present invention aims to overcome at least one defect (shortcoming) of the above-mentioned prior art, and provides a multi-model-based face image detection method, device, and electronic device, so as to achieve the effect of improving image detection efficiency.

[0005] According to a first aspect of the present application, there is provided a multi-model-based face image detection method, the method comprising:

[0006] Obtain an image to be detected;

[0007] Input the image to be detected into a first detection network for face detection. If face information is detected, output face image information. If face information is not detected, input the image to be detected into a second detection network. The second detection network includes a plurality of face detection models, and the plurality of face detection models respectively process the image to be detected to obtain corresponding plurality of detection results;

[0008] Obtain face image information corresponding to the image to be detected according to the plurality of detection results.

[0009] The face image detection method uses the first detection network and the second detection network to detect the image to be detected. The first detection network first performs face detection on the image to be detected. If no frontal face information is detected, the image to be detected is input into the second detection network. A number of face detection models are set in the second detection network. The set face detection models can detect various complex situations included in the image to be detected. In the actual detection process, the dual detection using a number of face detection models in the first detection network and the second detection network can improve the detection efficiency, the detection accuracy is higher, the impact on the image to be detected is smaller, and it can also reduce the error caused by the additional impact on the image to be detected during the detection process. More accurate face image information can be obtained according to the results of the dual detection.

[0010] Optionally, obtaining the face image information corresponding to the image to be detected according to the several detection results includes:

[0011] Calculating the confidence level for each of the detection results to obtain the corresponding confidence level value;

[0012] Judging whether there is a face in the image to be detected according to all the confidence level values. If there is, process according to the several detection results to obtain the face image information corresponding to the image to be detected.

[0013] A number of face detection models are set in the second detection network. One detection model corresponds to one detection result, and each detection model works independently and cooperates with each other to complete the entire image detection process. Calculate the confidence level for each detection result to obtain the corresponding confidence level value. Judge whether there is a face in the image to be detected according to the confidence level value, and then further process the detection result to obtain the detailed face image information in the image.

[0014] Optionally, judging whether there is a face in the image to be detected according to all the confidence level values includes:

[0015] Presetting a confidence level threshold;

[0016] Adding up the respective confidence level values to obtain a total confidence level;

[0017] If the total confidence level is greater than the preset confidence level threshold, it is determined that there is a face in the image to be detected.

[0018] A confidence threshold is preset. The confidence threshold is used to determine whether there is a human face in the image to be detected. The detection results are processed and analyzed. Since each detection model focuses on different detection key points, the detailed situation of the content in the image to be detected can be obtained according to the detection results of each detection model. By calculating the confidence of each detection result, the corresponding confidence value is obtained. The obtained confidence values are added together to obtain a final confidence value. The final confidence value can describe the detection situation of the image to be detected. Comparing the final confidence value with the preset confidence threshold can determine the face image information in the image to be detected.

[0019] Optionally, adding the respective confidence values to obtain a confidence sum specifically includes:

[0020] Weights are respectively set for each confidence value;

[0021] The weight is multiplied by the corresponding confidence value to obtain a multiplication result after multiplication;

[0022] All the multiplication results are added together to obtain the confidence sum.

[0023] During the detection process, the first detection network is mainly used to detect the frontal face of a human face. However, since some images do not completely contain the frontal face or are blocked, the second detection network is needed for further detection. The second detection network contains multiple face detection models. Each detection model has a detection result, and a detection result corresponds to a confidence value. Since the detection content of each detection model is different, in order to improve the accuracy of the detection result, weights need to be set for the confidence values, and the obtained detection results are weighted and then summed.

[0024] Optionally, the weights corresponding to each confidence value are set according to the corresponding face detection model.

[0025] Among the several face detection models, the detection key points of each face detection model are different. Therefore, when setting weights, the corresponding weights will be set according to the category of the detection model, which can reduce the error caused by the difference in the focus of the detection model.

[0026] Optionally, the setting method of the weights corresponding to each confidence value is:

[0027] Face factor analysis is performed on the detection result corresponding to the confidence value to obtain an analysis result, and the weight corresponding to the confidence value is set according to the analysis result.

[0028] Each of the several face detection models corresponds to a detection result. Face factor analysis is performed on the detection results. The purpose of the face factor analysis is to determine whether the image contains face factors, such as facial features and texture features. In the actual detection process, since the skin has texture, when setting weights, the confidence weight of the detection model corresponding to the detection result involving texture feature detection is set relatively low to ensure that the accuracy of the detection result is not affected. For other detection results that may affect the detection result, corresponding weights will be set for the confidence of the detection model corresponding to the detection result.

[0029] Optionally, the first detection network is an MTCNN face detection model.

[0030] The first detection network is independently composed of an MTCNN (Multi-task Cascaded Convolutional Networks) face detection model, a multi-task cascaded convolutional neural network. The MTCNN (Multi-task Cascaded Convolutional Networks) face detection model is mainly used to detect frontal faces. It can simultaneously achieve face detection and key point localization, and gradually screen and optimize the face detection results through three cascaded networks (P-Net, R-Net, O-Net), which is efficient and accurate for the recognition of frontal faces.

[0031] Optionally, the several face detection models are composed of an SSD model, a MediaPipe model, a CenterFace model, and a YOLOv8 model.

[0032] The second detection network consists of multiple models. By leveraging the advantages of individual models, a reasonable and efficient detection network is formed. Among them, the SSD (Single Shot Multibox Detector) model can simultaneously predict the positions and categories of multiple objects of different categories in a single forward propagation. It is a lenient model for face detection and will detect suspected faces. It is mainly used for detecting faces as the target. The MediaPipe model provides various real-time face processing functions, including the 3D positions of face key points and the tracking of facial contours, and is used for detecting face key points. The CenterFace model is a lightweight model that focuses on detecting faces and their key points, is suitable for edge computing devices, performs well on mobile devices, and has a certain robustness to lighting. Under lighting conditions, it further detects faces. The YOLOv8 model is the latest version of YOLO, which has enhanced the detection ability for small objects (such as distant faces) and also includes the recognition of some face factors. The four models, namely the SSD (Single Shot Multibox Detector) model, the MediaPipe model, the CenterFace model, and the YOLOv8 model, can detect faces in the image to be detected from different dimensions, and then evaluate whether a face is detected based on the corresponding confidence values, so as to complete face detection from a more comprehensive dimension, avoiding the inability to detect faces in the image to be detected due to the lack of frontal face information and improving the accuracy of face detection.

[0033] According to the second aspect of the present application, there is provided a multi-model based face image detection device, including:

[0034] A first acquisition module for acquiring an image to be detected;

[0035] A detection module for inputting the image to be detected into the first detection network for face detection. If face information is detected, it outputs face image information. If no face information is detected, it inputs the image to be detected into the second detection network. The second detection network includes several face detection models, and the several face detection models respectively process the image to be detected to obtain corresponding several detection results;

[0036] A second acquisition module for acquiring the face image information corresponding to the image to be detected according to the several detection results.

[0037] Each module in the multi-model based face image detection device cooperates with each other to implement the detection of face images. The first acquisition module acquires the image to be detected. The image to be detected enters the detection module, and the frontal face detection of the image is first performed through the first detection network. If no frontal face is detected, it enters the second detection network. The second detection network includes several face detection models, and the several face detection models are used to further detect the image to be detected to determine the relevant information of the face image in the image to be detected.

[0038] According to the third aspect of the present application, there is provided an electronic device, including:

[0039] A memory for storing one or more computer programs;

[0040] A processor, when the one or more computer programs are executed by the processor, implements the multi-model based face image detection method described in the first aspect above.

[0041] The electronic device provides a complete system for multi-model based face image detection, stores the image to be detected, and processes the stored image. The processed image is then saved by the memory.

[0042] According to the fourth aspect of the present application, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the multi-model based face image detection method described in the first aspect above when executed.

[0043] Based on any of the above aspects, the multi-model based face image detection method, device, electronic device, and computer storage medium provided by the embodiments of the present application can accurately and efficiently detect face images. By using the characteristics of a single face detection model to establish a detection network, different types of images can be detected through the established detection network, reducing the error of the detection effect caused by problems in the image itself or during image transmission. At the same time, the layer-by-layer detection of the detection network can improve the accuracy of detection and obtain relatively accurate information of the face image. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 FIG. is a schematic application scenario diagram of a multi-model based face image detection method provided in this embodiment.

[0046] Figure 2 Flowchart of a multi - model based face image detection method provided in this embodiment.

[0047] Figure 3 Schematic diagram of functional modules of a multi - model based face image detection device provided in this embodiment.

[0048] Figure 4 Schematic diagram of the structure of an electronic device provided in this embodiment. Detailed implementation manners

[0049] The accompanying drawings of this application are only for illustrative purposes and should not be construed as limitations on this application. To better illustrate the following embodiments, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well - known structures and their descriptions in the drawings may be omitted.

[0050] In order to enable those skilled in the art of this technology to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0051] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above - mentioned accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0052] In existing face detection methods, various types of face detection models are used. In the actual detection process, face detection models have certain limitations. Coupled with various complex problems in actual detection images, existing models such as Multi-task Cascaded Convolutional Networks (MTCNN) have good detection effects for clear frontal faces. However, for side faces, or when the frontal face is blocked, or when the frontal face is unclear, the detection accuracy of the model will be greatly reduced; since the obtained detection images often go through multiple compressions, the overall clarity of the compressed pictures will be reduced, which will also affect face detection. After multiple transmissions over the Internet, the clarity of the images will be further reduced. Some compression algorithms will also perform a scaling operation on the images, resulting in changes in the pixels of the images, thereby causing the key points of the face to shift; face detection is a key technology in face recognition, and face recognition plays an important role in life. Therefore, it is imperative to improve the efficiency of face detection.

[0053] This embodiment provides a technical solution that can solve the above problems. The following will combine the accompanying drawings to detail the specific implementation manners of the present application.

[0054] Exemplarily, it is a schematic diagram of an application scenario of a multi-model-based face image detection method provided by an embodiment of the present application. As Figure 1 shown, the application scenario at least includes a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has the functions of storing and processing images; the terminal device 200 has the functions of storing and recording image information, and can also have image processing functions.

[0055] It can be understood that the server 100 can be an independent electronic device or a cluster composed of multiple electronic devices; the terminal 200 can be a smart phone terminal, a personal computer, a tablet computer, a vehicle-mounted terminal, etc., but is not limited thereto.

[0056] In an implementable manner, the server 100 and the terminal 200 can respectively execute the multi-model-based face image detection method provided by the embodiment of the present application. Or, optionally, the multi-model-based face image detection method provided by the embodiment of the present application is partially executed in the server 100 and partially executed in the terminal 200.

[0057] As Figure 2 shown, this embodiment provides a multi-model-based face image detection method, which can include the following steps:

[0058] S110. Obtain the image to be detected;

[0059] In this embodiment, the image to be detected may be an image containing a human face. The human face type may be a frontal face, a profile face, or include special effects; it may also be an unknown image, and the type and source of the image are not limited.

[0060] In an alternative implementation, an image of a live broadcast cover is obtained as the image to be detected, so as to use the method of this application to detect whether there is human face information in the image of the live broadcast cover. The type and source of the photo include but are not limited to this.

[0061] S120: Input the image to be detected into a first detection network for face detection. If face information is detected, output face image information; if no face information is detected, input the image to be detected into a second detection network. The second detection network includes several face detection models, and the several face detection models respectively process the image to be detected to obtain corresponding several detection results.

[0062] In this embodiment, the first detection network may be a single face detection model or may be composed of a combination of multiple face detection models. The face information can be used to determine whether it is a human face, including frontal and profile faces, facial key points, or texture features, etc. The face image information is used to describe the manifestation form of the human face in the image. The manifestation form includes a frontal face, a profile face, or whether there is occlusion on the face, whether there are special effects, etc.

[0063] In an alternative implementation, the first detection network can be used to detect whether the image to be detected includes a frontal face. According to the result of the first detection network, it is determined whether to input the image to be detected into the second detection network again. If the image to be detected does not contain a frontal face, input the image to be detected into the second detection network.

[0064] In an alternative implementation, an image to be detected is obtained and input into a first detection network. The first detection network is independently composed of an MTCNN (Multi-task Cascaded Convolutional Networks) face detection model. The MTCNN (Multi-task Cascaded Convolutional Networks) face detection model is mainly used to detect frontal faces. It can simultaneously achieve face detection and key point localization. Through three cascaded networks (P-Net, R-Net, O-Net), the face detection results are gradually screened and optimized. According to the detection result of the first detection network, it is determined whether to perform secondary detection. If the first detection network detects face information, that is, a frontal face, output face image information; if no frontal face is detected, input the image to be detected into the second detection network to continue face detection.

[0065] In this embodiment, the second detection network includes several face detection models. Specifically, different face detection models can be used to detect the image to be detected in parallel, so that the detection results of different face detection models can be used to improve the accuracy of face detection.

[0066] In a preferred embodiment, the several face detection models may adopt the SSD (Single Shot Multibox Detector, single-stage object detection) model, the MediaPipe model, the CenterFace model, and the YOLOv8 model. By leveraging the advantages of each individual model, a reasonable and efficient detection network is formed. The image to be detected is input into the SSD model in the second detection network. The SSD (Single Shot Multibox Detector, single-stage object detection) model is a lenient model for face detection, and all suspected faces will be detected. It is mainly used for face target localization, and a detection result will be obtained based on the detection in the SSD (Single Shot Multibox Detector, single-stage object detection) model. The image to be detected is input into the MediaPipe model in the second detection network. The MediaPipe model provides various face processing functions with strong real-time performance, including the 3D positions of face key points and the tracking of facial contours, and is used for detecting face key points. A detection result will be obtained based on the detection in the MediaPipe model. The image to be detected is input into the CenterFace model in the second detection network. The CenterFace model is a lightweight model that focuses on detecting faces and their key points under lighting conditions and further stably detects faces. A detection result will be obtained based on the detection in the CenterFace model. The image to be detected is input into the YOLOv8 model. The YOLOv8 model is the latest YOLO version, which has increased the detection ability for small objects (such as distant faces), including the recognition of some face factors. A detection result will be obtained based on the detection in the YOLOv8 model. Four detection results corresponding to the four face detection models are obtained through the second detection network. Subsequently, it can be further determined whether there is face information in the image to be detected based on these detection results. In a preferred embodiment, the SSD (Single Shot Multibox Detector, single-stage object detection) model, the MediaPipe model, the CenterFace model, and the YOLOv8 model are selected for face detection. Each face detection model has different focuses on face detection and can be well combined to detect faces in the image to be detected. Especially when no face information can be detected after the image to be detected is detected by the first detection network, the second detection network is used to continue face detection, which can well exert the advantages of each face detection model, thereby improving the accuracy of face detection. It can be known that the types and numbers of detection models in the second detection network and the combination order are not specifically limited.

[0067] S130. Obtain the face image information corresponding to the image to be detected according to the several detection results.

[0068] In this embodiment, the several detection results are generated by several face detection models in the second detection network. One detection result corresponds to one detection model. The face image information can be directly obtained by processing the detection results, or the confidence of the detection results can be calculated, and the face image information can be obtained by analyzing the confidence values.

[0069] In an alternative implementation, according to the multiple detection results generated by the second detection network, the confidence values corresponding to the multiple detection results are respectively calculated. Specifically, the confidence values corresponding to each of the detection results are calculated; according to all the confidence values, it is determined whether there is a face in the image to be detected. If there is, the face image information corresponding to the image to be detected is obtained according to the several detection results.

[0070] Exemplarily, the confidence value corresponding to the SSD (Single Shot Multibox Detector) model is X 1 and the confidence value corresponding to the MediaPipe model is X 2 and the confidence value corresponding to the CenterFace model is X 3 and the confidence value corresponding to the YOLOv8 model is X 4 , according to the obtained confidence values X 1 、X 2 、X 3 、X 4 to determine whether there is a face in the image to be detected.

[0071] In a more preferred implementation, the specific steps of determining whether there is a face in the image to be detected according to all the confidence values may include:

[0072] Preset a confidence threshold;

[0073] Add all the confidence values to obtain a total confidence value;

[0074] If the total confidence value is greater than the preset confidence threshold, it is determined that there is a face in the image to be detected. In this embodiment, a weighted scheme is adopted to discriminate faces, which can completely exclude some images to be detected that contain face factors but are not face images, and can well detect the side faces in the image to be detected, including some occluded faces, etc., improving the accuracy of the face detection method of this application.

[0075] Exemplarily, the obtained confidence value X1 and X 2 and X 3 and X 4 are weighted, and after weighting, the total confidence is obtained, so that the total confidence can be compared with the preset confidence threshold.

[0076] In a more preferred embodiment, weights can also be set for the confidence values corresponding to each face detection model, and weighted based on the weights and the confidence values. Specifically, weights are set for each confidence value, and the product of the weights and the corresponding confidence values is obtained as the multiplication result after multiplication. All the multiplication results are added to obtain the total confidence. Exemplarily, the obtained confidence values X 1 and X 2 and X 3 and X 4 are respectively set with weights Y 1 and Y 2 and Y 3 and Y 4 , and each confidence value X 1 and X 2 and X 3 and X 4 is multiplied by the corresponding weight Y 1 and Y 2 and Y 3 and Y 4 , and then added together to obtain the final total confidence.

[0077] In the specific implementation process, the weights can be set according to the specific type of face detection model, and the face factor analysis can also be performed based on the detection results obtained by each face detection model for the image to be detected. According to the analysis results of the face factor analysis, the weights corresponding to the confidence levels are dynamically set. The face factor analysis can include the number of various types of faces appearing. The weighted weights are set according to the number of various types of faces appearing in the actual detected cover image, and then the corresponding weights are weighted with the face confidence levels predicted by the corresponding models. In this embodiment, the weights can be set considering the particularity of the specific face detection model, so as to give play to the advantages of each face detection model. The weights can also be set considering the detection results of the specific face detection model for the specific image to be detected, so as to realize the dynamic adjustment of the weights to achieve an accurate face detection effect. Exemplarily, the weight corresponding to the confidence level value of the SSD model is greater than the weight corresponding to the confidence level value of the MediaPipe model, and the weight corresponding to the confidence level value of the CenterFace model is greater than the weight corresponding to the confidence level value of the YOLOv8 model. Among them, the weight corresponding to the confidence level value of the SSD model is the highest, and the weight corresponding to the confidence level value of the YOLOv8 model is the lowest. After weighting each confidence level and then adding them up, a total confidence level is obtained. The weighting form includes but is not limited to this.

[0078] In a more preferred implementation manner, the preset confidence threshold can be set according to the actual situation, such as setting the confidence threshold to 0.7, 0.8 or other values. Exemplarily, if the total confidence level is greater than the preset confidence threshold of 0.7, it indicates that there is a face in the image to be detected. Then, the face image information in the image to be detected is obtained according to the detection results of each face detection model. The confidence threshold includes but is not limited to this.

[0079] This application uses a combination of a first detection network and a second detection network to perform face detection on the image to be detected. When the first detection network fails to detect face information, several face detection models in the second detection network are used to perform face detection on the image to be detected in parallel, so as to determine whether there is a face by combining the detection results of several face detection models. When it is determined that there is a face, the face information is output. This method can effectively filter out non-human images in the image to be detected. When applied to the live broadcast cover image, it can further improve the efficiency of manual review of the cover image and filter out some known abnormal cover images in advance. By automatically identifying and excluding those known abnormal or non-compliant cover images, the review team can focus on higher-risk or images that require detailed review. The application of this technology can reduce the workload of manual screening, speed up the overall review process, and improve the review quality at the same time.

[0080] As shown Figure 3 in the figure, the embodiment of the present application further provides a multi-model based face image detection device 210.

[0081] Optionally, the multi-model based face image detection device 210 may include:

[0082] A first acquisition module 211, configured to acquire an image to be detected;

[0083] In this embodiment, the first acquisition module 211 can be used to execute Figure 2 the steps S110 shown. For the specific description of the first acquisition module 211, reference can be made to the description of the steps S110.

[0084] A detection module 212, configured to input the image to be detected into a first detection network for face detection. If face information is detected, face image information is output. If no face information is detected, the image to be detected is input into a second detection network. The second detection network includes several face detection models, and the several face detection models respectively process the image to be detected to obtain corresponding several detection results.

[0085] In this embodiment, the detection module 212 can be used to execute Figure 2 the steps S120 shown. For the specific description of the detection module 212, reference can be made to the description of the steps S120.

[0086] A second acquisition module 213, configured to obtain the face image information corresponding to the image to be detected according to the several detection results.

[0087] In this embodiment, the second acquisition module 213 can be used to execute Figure 2 the steps S130 shown. For the specific description of the second acquisition module 213, reference can be made to the description of the steps S130.

[0088] It can be understood that the above device embodiments and the above method embodiments can correspond to each other. Similar descriptions of the device embodiments can refer to the method embodiments. To avoid repetition, details are not described here again. The multi-model based face image detection device provided by the embodiment of the present application can execute the multi-model based face image detection method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. The functional modules of the multi-model based face image detection device can be implemented in the form of hardware, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules.

[0089] Specifically, each step of the method embodiment of the present application can be completed by the integrated logic circuit of the hardware in the processor and / or the instructions in the form of software. The steps of the multi-model based face image detection method in the embodiments of the present application can be directly embodied as being completed by the hardware encoding processor, or completed by the combination of the hardware and software modules in the encoding processor. Optionally, the software module can be located in a random access memory, read-only memory, programmable read-only memory, flash memory, electrically erasable programmable memory, register and other storage media. The storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0090] Embodiments of the present application provide an electronic device 310, the structure of which is as Figure 4 shown. The electronic device 310 may be the server 100 or the terminal 200 shown in this embodiment Figure 1 shown.

[0091] As Figure 4 shown, the electronic device 310 includes a memory 311, a processor 312, a communication module 313, an input / output interface 314, etc. Optionally, the memory 311, the processor 312, the communication module 313, and the input / output interface 314 can be connected and communicate through a bus 315.

[0092] The memory 311 is used to store one or more computer programs and transmit the code of the computer programs to the processor 312; when the one or more computer programs are executed by the processor 312, the multi-model based face image detection method in the embodiments of the present application is implemented.

[0093] Optionally, the electronic device 310 can be connected to a network through the communication module 313 to communicate with other devices, such as terminals or servers, through the network to achieve data interaction. The electronic device 310 can be various forms of digital computers, for example, a desktop computer, a server, a workbench, a mainframe computer or other types of computers. The electronic device 310 can also be various forms of mobile terminals, for example, a smart phone, a tablet computer, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar mobile terminals.

[0094] Optionally, the electronic device 310 may be connected to required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 314. The electronic device 310 itself may have a display device, and may also externally connect other display devices through the input / output interface 314. Optionally, a storage device, such as a hard disk, etc., may also be connected through the input / output interface 314, so as to store the data in the electronic device 310 into the storage device, or read the data in the storage device, and may also store the data in the storage device into the memory 311. It can be understood that the input / output interface 314 may be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 314 may be components of the electronic device 310 or external devices connected to the electronic device 310 when needed.

[0095] Optionally, the memory 311 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory, etc.

[0096] Optionally, the computer program stored in the processor 312 may be divided into one or more modules. The one or more modules are stored in the memory 311 and executed by the processor 312 to complete the method provided by its own embodiment. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the computer program instruction segments are used to describe the execution process of the computer program in the electronic device 310.

[0097] Optionally, the processor 312 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 312 include but are not limited to a central processing unit, a graphics processing unit, a digital signal processor, various dedicated artificial intelligence computing chips, various processors running machine learning model algorithms, and may also be any suitable controller, microcontroller, processor, etc. The processor 312 executes each method and process of this embodiment. Exemplarily, such as a multi-model based face image detection method of an embodiment of the present application.

[0098] Optionally, the bus 315 may include a path for transmitting information. The bus 315 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. According to different functions, the bus 315 may be divided into an address bus, a data bus, a control bus, etc.

[0099] In an alternative implementation, an embodiment of the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments. Part or all of the computer program can be loaded and / or installed on the memory 311 of the electronic device 310. When the computer program is executed by the processor 312, one or more steps of a multi-model based face image detection method according to an embodiment of the present application can be executed.

[0100] Optionally, the computer-readable storage medium may be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.

[0101] Obviously, the above embodiments of the present application are merely examples for clearly illustrating the technical solutions of the present application, rather than limitations on the specific implementation manners of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present application shall be included in the protection scope of the claims of the present application.

Claims

1. A multi-model based face image detection method, characterized in that: The method comprises: Acquire the image to be detected; Input the image to be detected into a first detection network for face detection, and output face image information if face information is detected; if face information is not detected, input the image to be detected into a second detection network, wherein the second detection network includes a plurality of face detection models, and the plurality of face detection models respectively process the image to be detected to obtain a plurality of corresponding detection results; The facial image information corresponding to the image to be detected is obtained according to the plurality of detection results.

2. The multi-model based face image detection method according to claim 1, characterized in that: The acquiring the facial image information corresponding to the image to be detected according to the plurality of detection results includes: Performing confidence calculation on each of the detection results to obtain a corresponding confidence value; It is determined whether there is a face in the image to be detected according to all confidence values. If there is a face, the face image information corresponding to the image to be detected is obtained by processing according to several detection results.

3. The multi-model based face image detection method according to claim 2, characterized in that: The step of judging whether a face exists in the image to be detected according to all confidence values ​​includes: Preset confidence threshold; Add up the confidence values ​​to get the total confidence; If the total confidence level is greater than the preset confidence level threshold, it is determined that a face exists in the image to be detected.

4. The multi-model based face image detection method according to claim 3, characterized in that: The step of adding up the confidence values ​​to obtain the total confidence value specifically includes: Set weights for each confidence value separately; The weight is multiplied by the corresponding confidence value to obtain a multiplication result; All the multiplication results are added together to obtain the total confidence level.

5. The multi-model based face image detection method according to claim 4, characterized in that: The weights corresponding to the respective confidence values ​​are set according to the corresponding face detection model.

6. The multi-model based face image detection method according to claim 4, characterized in that: The weights corresponding to the respective confidence values ​​are set as follows: A facial factor analysis is performed based on the detection result corresponding to the confidence value to obtain an analysis result, and a weight corresponding to the confidence value is set based on the analysis result.

7. A multi-model based face image detection method according to any one of claims 1 to 6, characterized in that: The first detection network is an MTCNN face detection model.

8. The multi-model based face image detection method according to any one of claims 1 to 6, characterized in that: The several face detection models include SSD model, MediaPipe model, CenterFace model, and YoLov8 model.

9. A multi-model based face image detection device, characterized in that: include: A first acquisition module acquires an image to be detected; A detection module, inputting the image to be detected into a first detection network for face detection, outputting face image information if face information is detected, and inputting the image to be detected into a second detection network if face information is not detected, wherein the second detection network includes a plurality of face detection models, and the plurality of face detection models respectively process the image to be detected to obtain a plurality of corresponding detection results; The second acquisition module acquires the face image information corresponding to the image to be detected according to the plurality of detection results.

10. An electronic device, characterized in that: include: a memory for storing one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements a multi-model based face image detection method as described in any one of claims 1 to 8.