Face detection method based on heteronymous binocular camera, terminal and storage medium

By combining infrared and RGB images with heterogeneous binocular cameras, and utilizing face detection models and 3D supervision information, the detection of real and fake faces is achieved. This solves the problems of long detection time and high hardware requirements in existing technologies, and realizes face detection with low power consumption, low cost and high stability.

CN115690881BActive Publication Date: 2026-05-19HUNAN LANGGUO VISUAL RECOGNITION RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN LANGGUO VISUAL RECOGNITION RES INST CO LTD
Filing Date
2022-11-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Current face detection technology has shortcomings in terms of rapid response and environmental adaptability. Dynamic liveness detection is time-consuming, while silent liveness detection requires a large amount of computation and has high hardware requirements. It is also easily affected by environmental changes and is difficult to effectively distinguish between real and fake faces.

Method used

We employ a heterogeneous binocular camera-based approach to acquire infrared and RGB images. These images are then fused using a face detection model and 3D supervision information. By leveraging the imaging characteristics of different sensors, we can perform realistic face detection, thereby reducing computational load and hardware requirements.

Benefits of technology

It achieves low power consumption, low cost, and stable real face detection, improving user experience, reducing reliance on hardware, and minimizing environmental impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690881B_ABST
    Figure CN115690881B_ABST
Patent Text Reader

Abstract

The present application provides a kind of face detection method based on heterologous binocular camera, terminal and storage medium, the face detection method includes: S101: the infrared image of detection object, RGB image is acquired, the face in infrared image, RGB image is identified, and the face is intercepted and fused to form detection image;S102: detection image is input into face detection model, and the face detection classification of detection image is acquired according to face detection model and 3D supervision information.This application can effectively utilize the characteristics of different imaging characteristics of true and false faces on different sensors to detect real face, with small calculation, low power consumption, reduced hardware requirements, low cost and not easily affected by environment, good stability, improved user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of face detection technology, and in particular to a face detection method, terminal, and storage medium based on a heterogeneous binocular camera. Background Technology

[0002] With technological advancements and societal development, facial recognition is finding increasingly diverse applications. A successful fake face attack could potentially cause significant losses, making its security a top priority in payment and other scenarios. Forgery attacks targeting this scenario include, but are not limited to, using printed paper, electronic screens, or mannequin masks. Therefore, determining whether a detected face is real or fake is one of the core technologies of this system.

[0003] Current liveness detection solutions are divided into dynamic liveness detection and silent liveness detection, which have the following problems:

[0004] ① At present, dynamic liveness detection is difficult to deploy in scenarios that require fast response: because the ordinary convolutional design method tends to use long sequences as input to extract dynamic features. Its model input is dynamic multi-frame images, or it needs to complete a combination of actions such as blinking, opening mouth, shaking head, and nodding, which makes the method time-consuming.

[0005] ② At present, most silent dynamic liveness detection products are not effective in preventing fake face attacks in unknown domains: Existing products in the same field use classifiers to identify single-frame face images under visible light to determine whether the face is live or fake. They use cross-entropy or softmax loss functions to perform liveness detection on the image. This detection method is complex, computationally intensive, and products using this method have high hardware requirements, high power consumption, narrow adaptability, and are prone to failure when the environment changes (such as different illuminance), which reduces the user experience. Summary of the Invention

[0006] To overcome the shortcomings of existing technologies, this invention proposes a face detection method, terminal, and storage medium based on a heterogeneous binocular camera. During face detection, infrared and RGB images of the target object are acquired, and the face image is extracted from the infrared and RGB images and fused to obtain the detection image. The face detection model and 3D supervision information are used to classify the face in the detection image, thereby effectively utilizing the different imaging characteristics of real and fake faces on different sensors to perform real face detection. This method has low computational load, low power consumption, reduced hardware requirements, low cost, and is not easily affected by the environment, exhibiting good stability and improving the user experience.

[0007] To address the aforementioned problems, the present invention employs the following technical solution: a face detection method based on a heterogeneous binocular camera, the face detection method comprising: S101: acquiring infrared images and RGB images of the object to be detected, identifying faces in the infrared images and RGB images, and cropping and fusing the faces to form a detection image; S102: inputting the detection image into a face detection model, and obtaining a face detection classification of the detection image based on the face detection model and 3D supervision information.

[0008] Furthermore, the steps of acquiring the infrared image and RGB image of the detection object specifically include: after determining that the detection conditions are met, controlling the IR camera and RGB camera to aim at the face and capture the infrared image and RGB image of the face.

[0009] Furthermore, the step of identifying faces in infrared images and RGB images specifically includes: inputting the infrared images and RGB images into corresponding face recognition models, and identifying faces in the infrared images and RGB images through the face recognition models.

[0010] Furthermore, the step of identifying faces in the infrared and RGB images using a face recognition model specifically includes: extracting features from the infrared and RGB images using convolutional layers in the face recognition model, and identifying the region where the face is located based on the extraction results.

[0011] Furthermore, the step of cropping and fusing the face to form a detection image specifically includes: cropping the image of the area where the face is located, converting the image into an image of the same size, and performing channel stitching based on the features of the image to obtain a detection image.

[0012] Furthermore, before the step of obtaining the face detection classification of the detection image based on the face detection model and 3D supervision information, the method further includes: obtaining a supervision image of the face, extracting the 3D information of the supervision image, and normalizing the 3D information to form 3D supervision information.

[0013] Furthermore, the step of obtaining the face detection classification of the detected image based on the face detection model and 3D supervision information specifically includes: extracting features of the detected image through the convolutional layer of the face detection model, supervising the features using the 3D supervision information, and processing the supervision results and features through the fully connected layer of the face detection model to obtain the face detection classification.

[0014] Furthermore, the step of processing the supervision results and features through the fully connected layer of the face detection model to obtain the face detection classification specifically includes: obtaining the predicted probability that the detected image belongs to different detection results based on the output of the fully connected layer, and obtaining the classification of the detected image based on the predicted probability.

[0015] Based on the same inventive concept, the present invention also proposes an intelligent terminal, which includes a processor and a memory. The memory stores a computer program, and the processor is communicatively connected to the memory. The computer program executes the face detection method based on a heterogeneous binocular camera as described above.

[0016] Based on the same inventive concept, the present invention also proposes a computer-readable storage medium storing program data, which is used to execute the face detection method based on heterogeneous binocular cameras as described above.

[0017] Compared with existing technologies, the beneficial effects of this invention are as follows: during face detection, infrared and RGB images of the target object are acquired, face images are extracted from the infrared and RGB images and fused to obtain a detection image, and face detection and classification are performed on the detection image using a face detection model and 3D supervision information. This effectively utilizes the different imaging characteristics of real and fake faces on different sensors to perform real face detection. It has low computational load, low power consumption, reduced hardware requirements, low cost, and is not easily affected by the environment. It also has good stability and improves the user experience. Attached Figure Description

[0018] Figure 1 This is a structural diagram of an embodiment of the face detection method based on a heterogeneous binocular camera according to the present invention;

[0019] Figure 2 This is a flowchart illustrating an embodiment of the face detection method based on a heterogeneous binocular camera according to the present invention.

[0020] Figure 3 This is a structural diagram of an embodiment of the smart terminal of the present invention;

[0021] Figure 4 This is a structural diagram of an embodiment of the computer-readable storage medium of the present invention. Detailed Implementation

[0022] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that the various embodiments of this disclosure described and shown in the accompanying drawings can be combined with each other without conflict, and the structural components or functional modules can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. The singular terms “a,” “said,” and “the” as used in this disclosure and the appended claims are also intended to include the plural terms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0024] Please see Figure 1 , Figure 2 ,in, Figure 1 This is a structural diagram of an embodiment of the face detection method based on a heterogeneous binocular camera according to the present invention; Figure 2 This is a flowchart illustrating an embodiment of the face detection method based on a heterogeneous binocular camera according to the present invention. (Combined with...) Figure 1 , Figure 2 The face detection method based on heterogeneous binocular cameras of the present invention will be described.

[0025] In this embodiment, the device executing the face detection method based on a heterogeneous binocular camera can be a mobile phone, laptop, tablet computer, or other smart terminal capable of performing face detection based on captured images. The heterogeneous binocular camera includes an infrared camera and an RGB camera. This camera can be mounted on the smart terminal or connected to it. The smart terminal performs face detection and classification based on the images captured by the heterogeneous binocular camera.

[0026] In this embodiment, the face detection method based on heterogeneous binocular cameras executed by the smart terminal includes:

[0027] S101: Acquire infrared and RGB images of the object to be detected, identify faces in the infrared and RGB images, and extract and merge faces to form a detection image.

[0028] In this embodiment, the steps of acquiring the infrared image and RGB image of the object to be detected specifically include: after determining that the detection conditions are met, controlling the IR camera and RGB camera to aim at the face and capture the infrared image and RGB image of the face.

[0029] Specifically, after determining that the detection conditions are met, the smart terminal controls the IR camera and RGB camera to take pictures of the detected object within the detection area, thereby acquiring the face image of the detected object. The detection conditions can be pre-set conditions such as receiving a face detection command, the detected object appearing in the detection area, the current face detection time, the appearance of a specific detected object, or the recognition of a specific scene.

[0030] In this embodiment, during face detection, the smart terminal can first identify the face of the target object, determine the position of the face, and then control the IR camera and RGB camera to aim at the face to capture an image of the face. In other embodiments, the smart terminal can also guide the target object to move to a designated position or position the face in a designated position using voice, text, images, etc., and then capture an image of the face. The smart terminal can also adjust the focus and change the shooting direction of the IR camera and RGB camera to aim at the target object's face.

[0031] In this embodiment, when the smart terminal controls the IR camera and the RGB camera to capture a face, it controls the IR camera and the RGB camera to capture the face of the detected object simultaneously.

[0032] In a preferred embodiment, to facilitate image processing, the smart terminal can also control the IR camera and the RGB camera to capture images of the same size.

[0033] In this embodiment, the steps of recognizing faces in infrared images and RGB images specifically include: inputting the infrared images and RGB images into the corresponding face recognition models, and recognizing faces in the infrared images and RGB images through the face recognition models.

[0034] The steps for identifying faces in infrared and RGB images using a face recognition model specifically include: extracting features from the infrared and RGB images using convolutional layers in the face recognition model, and identifying the region where the face is located based on the extraction results.

[0035] Specifically, the face recognition model is a deep learning network with multiple convolutional layers. After the face recognition model obtains infrared and RGB images, it first uses convolutional layers to convert the infrared and RGB images into images of size 112*112*128. Then, it uses other convolutional layers to process the image into an image of size 64*64*256, which is the image of the area where the face is located.

[0036] In other embodiments, when extracting features from infrared images and RGB images through convolutional layers to obtain images of the face region, images of the face region of different sizes can also be generated. Based on preset image size information, images that do not conform to the image size information can be enlarged or reduced to be adjusted to match the image size information.

[0037] In this embodiment, the step of cropping and fusing a face to form a detection image specifically includes: cropping the image of the area where the face is located, converting the image into an image of the same size, and performing channel stitching based on the features of the image to obtain the detection image.

[0038] One approach is to use a face recognition model to fuse images of the face region together to generate a detection image. Alternatively, after obtaining images of the face region output by the face recognition model, a pre-defined image fusion tool can be used to fuse these images together to obtain the detection image.

[0039] In this embodiment, the face recognition model can be stored on a smart terminal, or it can be set on other devices such as a cloud or a server connected to the smart terminal. The smart terminal transmits the captured infrared image or RGB image to the device to obtain the detection image or the image of the area where the face is located fed back by the device.

[0040] S102: Input the detected image into the face detection model, and obtain the face detection classification of the detected image based on the face detection model and 3D supervision information.

[0041] In this embodiment, the face detection model requires the input image size to be consistent with the size of the detection image. In other embodiments, they may not be consistent. Before the detection image is input into the face detection model, the smart terminal first determines the size of the detection image to see if it meets the image detection requirements of the face detection model. If it does, the detection image is directly transmitted to the face detection model; if it does not, the detection image is scaled to generate a detection image that meets the image detection requirements.

[0042] In this embodiment, the face detection model is a deep learning convolutional network, which is trained using real face images, fake face images, and face photos.

[0043] To improve recognition accuracy, the face detection and classification process, which involves acquiring a supervision image of the face based on the face detection model and 3D supervision information, includes the following steps before the face detection and classification steps: acquiring a supervision image of the face, extracting 3D information from the supervision image, and normalizing the 3D information to form 3D supervision information. Utilizing the characteristic that the depth information of the entire 2D photograph is 0, 3D supervision information is added to identify fake face images whose depth information does not meet the requirements.

[0044] In one specific embodiment, 3D face modeling is performed based on a single static image of a face to obtain 3D information of the face. Depth normalization is then performed based on the 3D information so that the depth of the point closest to the camera is 1 and the depth of the point farthest from the camera is 0, thereby forming 3D supervision information.

[0045] In this embodiment, a static image of the face of each detection object can be acquired, and 3D face modeling can be performed to obtain the 3D information of the face of each detection object. This allows for the use of different 3D supervision information for each detection object, improving the accuracy of face detection.

[0046] In this embodiment, the steps of obtaining face detection classification of the detected image based on the face detection model and 3D supervision information specifically include: extracting features of the detected image through the convolutional layer of the face detection model, supervising the features using 3D supervision information, and processing the supervision results and features through the fully connected layer of the face detection model to obtain face detection classification.

[0047] In one specific embodiment, the face detection model is a deep learning convolutional network. The convolutional layers of this network extract features from the convolutional image, generating two 64*64*64 images. These two images are then processed by different convolutional layers to obtain another 64*64*64 image. One of these images is supervised by corresponding 3D supervision information to obtain a feature layer representing 3D information. This 3D information is used to prevent spoofing by photos or screen information. The two images are then output to a fully connected layer, which performs face detection and classification.

[0048] In this embodiment, the step of processing the supervision results and features through the fully connected layer of the face detection model to obtain the face detection classification specifically includes: obtaining the predicted probability of the detected image belonging to different detection results based on the output of the fully connected layer, and obtaining the classification of the detected image based on the predicted probability. The predicted probability includes the probability of a real face, the probability of a 3D fake face, and the probability of a 2D fake face. The classification of the detected image is obtained by processing the predicted probability through supervision information, and this classification includes real face, 2D fake face, and 3D fake face. The supervision information includes classifications corresponding to different probability ranges.

[0049] Beneficial effects: The face detection method based on heterogeneous binocular cameras of this invention acquires infrared and RGB images of the target object during face detection, extracts the face image from the infrared and RGB images and fuses them to obtain the detection image, and uses a face detection model and 3D supervision information to classify the face in the detection image. This effectively utilizes the different imaging characteristics of real and fake faces on different sensors to detect real faces. It has low computational load, low power consumption, reduces hardware requirements, low cost, is not easily affected by the environment, has good stability, and improves the user experience.

[0050] Based on the same inventive concept, this invention also proposes a smart terminal, please refer to [link / reference]. Figure 3 , Figure 3 This is a structural diagram of an embodiment of the smart terminal of the present invention, combined with... Figure 3 The smart terminal of the present invention will be described in detail below.

[0051] In this embodiment, the smart terminal includes a processor and a memory. The memory stores a computer program, and the processor is communicatively connected to the memory. The computer program executes the face detection method based on a heterogeneous binocular camera as described in the above embodiment.

[0052] In some embodiments, the memory may include, but is not limited to, high-speed random access memory and non-volatile memory. For example, one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable functional devices, discrete gate or transistor functional devices, or discrete hardware components.

[0053] Based on the same inventive concept, this invention also proposes a computer-readable storage medium, please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a structural diagram of an embodiment of the computer-readable storage medium of the present invention, in conjunction with... Figure 4 The computer-readable storage medium of the present invention will be described.

[0054] In this embodiment, a computer-readable storage medium stores program data that is used to execute the mesh generation method as described in the above embodiments.

[0055] The computer-readable storage medium may include, but is not limited to, floppy disks, optical disks, CD-ROMs (compact disc read-only memory), magneto-optical disks, ROMs (read-only memory), RAMs (random access memory), EPROMs (erasable programmable read-only memory), EEPROMs (electrically erasable programmable read-only memory), magnetic cards or optical cards, flash memory, or other types of media / machine-readable media suitable for storing machine-executable instructions. The computer-readable storage medium may be a product not connected to a computer device or a component used in a computer device.

[0056] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A face detection method based on heterogeneous binocular cameras, characterized in that, The face detection method includes: S101: Acquire infrared and RGB images of the object to be detected, identify faces in the infrared and RGB images, and extract and fuse the faces to form a detection image; S102: Input the detected image into the face detection model, and obtain the face detection classification of the detected image based on the face detection model and 3D supervision information; The step of obtaining the face detection classification of the detected image based on the face detection model and 3D supervision information specifically includes: Features of the detected image are extracted by the convolutional layer of the face detection model, the features are supervised by the 3D supervision information, and the supervision results and features are processed by the fully connected layer of the face detection model to obtain face detection classification. The step of processing the supervision results and features through the fully connected layer of the face detection model to obtain face detection classification specifically includes: Based on the output of the fully connected layer, the predicted probability of the detected image belonging to different detection results is obtained, and the classification of the detected image is obtained based on the predicted probability. The predicted probabilities include the probability of a real face, the probability of a 3D fake face, and the probability of a 2D fake face. The predicted probabilities are processed by the 3D supervision information to obtain the classification of the detected image. The classification includes real face, 2D fake face, and 3D fake face. The 3D supervision information includes classifications corresponding to different probability ranges.

2. The face detection method based on heterogeneous binocular cameras as described in claim 1, characterized in that, The steps of acquiring the infrared image and RGB image of the object to be detected specifically include: After confirming that the detection conditions are met, the IR camera and RGB camera are aimed at the face to capture infrared and RGB images of the face.

3. The face detection method based on heterogeneous binocular cameras as described in claim 1, characterized in that, The steps for identifying faces in infrared and RGB images specifically include: The infrared image and RGB image are respectively input into the corresponding face recognition model, and the face recognition model identifies the face in the infrared image and RGB image.

4. The face detection method based on heterogeneous binocular cameras as described in claim 3, characterized in that, The step of identifying faces in the infrared and RGB images using a face recognition model specifically includes: The infrared and RGB images are used to extract features through the convolutional layers in the face recognition model, and the region where the face is located is identified based on the extraction results.

5. The face detection method based on heterogeneous binocular cameras as described in claim 4, characterized in that, The step of extracting and fusing the face to form a detection image specifically includes: An image of the area containing the face is cropped, the image is converted to an image of the same size, and the detection image is obtained by channel stitching based on the features of the image.

6. The face detection method based on heterogeneous binocular cameras as described in claim 5, characterized in that, Before the step of obtaining the face detection classification of the detected image based on the face detection model and 3D supervision information, the method further includes: Obtain a supervised image of a face, extract 3D information from the supervised image, and normalize the 3D information to form 3D supervised information.

7. A smart terminal, characterized in that, The smart terminal includes a processor and a memory. The memory stores a computer program. The processor is communicatively connected to the memory and executes the face detection method based on a heterogeneous binocular camera as described in any one of claims 1-6 through the computer program.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program data, which is used to execute as claimed in claim 1. The face detection method based on heterogeneous binocular camera as described in any one of the six claims.