A method and apparatus for live face detection, a terminal device, and a storage medium
By processing near-infrared images, extracting and segmenting the structured information of facial contours, and randomly stitching them together, the problem of low detection flexibility of near-infrared images in different scenarios or modules is solved, achieving wider applicability and improving the generalization ability of the detection model.
Patent Information
- Application Number
- CN202111577808.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-12-22
AI Technical Summary
In existing technologies, face liveness detection based on near-infrared images has low flexibility in different scenarios or modules, resulting in unstable detection results.
By processing near-infrared images, the contour structure information of the target face is extracted, and then image matting and random stitching are performed to enhance the generalization ability of the live face detection model and improve the applicability of the detection results.
It improves the flexibility of live face detection, enabling the detection results to be applied to more generalized scenarios or modules, and enhances the generalization ability of the detection model.
Smart Images

Figure CN114495196B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, specifically to image processing and deep learning technologies, and in particular to a method, apparatus, terminal device, and storage medium for live face detection. Background Technology
[0002] With the increasing application of computer vision and deep learning technologies in the field of artificial intelligence, near-infrared (NIIR) images are gaining popularity in face recognition scenarios. NIIR images, formed by remote sensors receiving near-infrared spectral data reflected or emitted by objects, are unaffected by visible light, leading to their widespread use in face liveness detection. However, in NIIR-based face liveness detection, the sensitivity of NIIR images to data distribution means that the results can vary depending on the scenario or module used, resulting in limited flexibility. Therefore, finding a more flexible approach to face liveness detection is a pressing issue. Summary of the Invention
[0003] This application provides a method, apparatus, terminal device, and storage medium for live face detection. By processing near-infrared images, a face image to be processed, including the contour structure information of the target face, is obtained. Then, the contour structure information of the face is processed to obtain a target face image including spoofing features. This allows the live face detection model to focus on more local and detailed spoofing features, thereby enhancing the generalization ability of the live face detection model. Therefore, the obtained live face detection results can be applied to more generalized scenarios or modules, thereby improving the flexibility of face liveness detection.
[0004] In view of the above, the first aspect of this application provides a method for live face detection, comprising:
[0005] Acquire a first near-infrared image, wherein the first near-infrared image includes a target human face;
[0006] The first near-infrared image is processed to obtain the contour structure information of the target face;
[0007] Based on the contour structure information of the target face, the first near-infrared image is processed to obtain a face image to be processed that includes only the target face.
[0008] The image of the face to be processed is segmented according to the contour structure information of the target face, and the segmented images are randomly stitched together to obtain the target face image.
[0009] Based on the target face, a second aspect of this application provides a live face detection device, comprising:
[0010] The first acquisition module is used to acquire a first near-infrared image, wherein the first near-infrared image includes a target human face;
[0011] The first processing module is used to process the first near-infrared image to obtain the contour structure information of the target face;
[0012] The second processing module is used to perform image matting on the first near-infrared image based on the contour structure information of the target face to obtain a face image to be processed that includes only the target face.
[0013] The third processing module is used to cut the face image to be processed according to the contour structure information of the target face, and randomly stitch the cut images to obtain the target face image.
[0014] The second acquisition module is used to determine the liveness detection result of the target face based on the target face image and the preset liveness detection model.
[0015] A third aspect of this application provides a terminal device, including: a memory and a processor; wherein the memory is used to store a computer program; and the processor is used to execute the computer program in the memory to implement the methods described in the above aspects.
[0016] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods provided in the above aspects.
[0017] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0018] By processing near-infrared images, a face image to be processed, including the contour structure information of the target face, is obtained. Then, the contour structure information of the face is processed to obtain a target face image including spoof features. This allows the liveness detection model to focus on more local and detailed spoof features, thereby enhancing the generalization ability of the liveness detection model. Therefore, the obtained liveness detection results can be applied to more generalized scenarios or modules, thus improving the flexibility of face liveness detection. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a system architecture for a live face detection system as described in this application.
[0020] Figure 2 A schematic diagram of an embodiment of the live face detection method provided in this application;
[0021] Figure 3A schematic diagram of the architecture of the preset face location detection model provided in the embodiments of this application;
[0022] Figure 4 A schematic diagram of an embodiment of the face image to be processed provided in this application;
[0023] Figure 5 A schematic diagram of an embodiment of a human face image to be processed after being cut and randomly stitched together, as provided in this application;
[0024] Figure 6 A schematic diagram of an embodiment of the target face image provided in this application;
[0025] Figure 7 A schematic diagram of the architecture of a preset live face detection model provided in an embodiment of this application;
[0026] Figure 8 A schematic diagram of a preset live face detection device provided in an embodiment of this application;
[0027] Figure 9 This is a schematic diagram of the structure of one embodiment of the terminal device in this application. Detailed Implementation
[0028] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data used can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms “comprising” and “corresponding to,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] This application provides a method for live face detection. By enhancing the generalization ability of the live face detection model, the obtained live face detection results can be applied to more generalized scenarios or modules, thereby improving the flexibility of face liveness detection.
[0030] It is understood that the live face detection method in this application embodiment can be executed by a terminal device or by a server. First, the system architecture of this application embodiment will be introduced. Figure 1This is a schematic diagram of a system architecture for a live face detection system according to an embodiment of this application. The system includes a terminal device. Specifically, when the live face detection method is deployed on a server, a first near-infrared image including the target face can be acquired in real time from the terminal device, and then the live face detection result is obtained through the method provided in this embodiment. Alternatively, when the live face detection method is deployed on a terminal device, the terminal device acquires and processes the first near-infrared image of the target face, and obtains the live face detection result based on the method provided in this embodiment, thereby improving the flexibility of live face detection.
[0031] In one embodiment, the terminal device may include a data acquisition module for acquiring a second near-infrared image. The terminal device normalizes the second near-infrared image to obtain a first near-infrared image, thereby executing the live face detection method of this application. It should be noted that the data acquisition module may be a component of the terminal device or an external device independent of the terminal device; the data acquisition module may communicate with the terminal device via wired or wireless means; the data acquisition module may include any one or more combinations of infrared cameras, depth cameras, color cameras, and other cameras, without limitation herein.
[0032] It should be noted that the aforementioned terminal devices and servers communicate via wireless networks, wired networks, or removable storage media. It should also be noted that... Figure 1 The server in this context can be a single server, a server cluster consisting of multiple servers, or a cloud computing center, etc.; the specifics are not limited here. The terminal device can be... Figure 1 The tablets, laptops, door locks, and mobile phones or other terminal devices shown are described. The aforementioned wireless networks use standard communication technologies and / or protocols. The wireless network is typically the Internet, but can also be any network, including but not limited to Bluetooth, Local Area Network (LAN), Metropolitan Area Network (MAN), Wide Area Network (WAN), mobile, private networks, or any combination of virtual private networks. In some embodiments, custom or dedicated data communication technologies may be used to replace or supplement the aforementioned data communication technologies. The removable storage medium can be a Universal Serial Bus (USB) flash drive, external hard drive, or other removable storage medium.
[0033] Although Figure 1 Only four terminal devices and one server are shown, but it should be understood that... Figure 1 The examples provided are for illustrative purposes only; the actual number of terminal devices and servers should be determined flexibly based on the specific circumstances.
[0034] Based on the above introduction, the following section uses a terminal device as the execution subject to describe the method for live face detection in this application. Figure 2 A schematic diagram of an embodiment of the live face detection method provided in this application, the method comprising:
[0035] Step 201: Obtain the first near-infrared image.
[0036] In this embodiment, a first near-infrared image is acquired using a data acquisition module that has a communication connection with the terminal device. This first near-infrared image includes the target face. Specifically, a second near-infrared image is acquired, and this second near-infrared image also includes the target face. Typically, near-infrared (Infrared Radiationir) images are 10-bit, meaning the second near-infrared image contains image data from 0 to 1024. However, standard images for image processing are usually 8-bit. To improve image processing efficiency, the second near-infrared image needs to be normalized to obtain the first near-infrared image. In this case, the first near-infrared image is an 8-bit image, specifically containing image data from 0 to 256.
[0037] Step 202: Process the first near-infrared image to obtain the contour structure information of the target face.
[0038] In this embodiment, based on the first near-infrared image obtained in step 201, the contour structure information of the target face in the first near-infrared image is obtained through a face position detection model. That is, the first near-infrared image is used as the input of a preset face position detection model, and the preset face position detection model outputs the contour structure information of the target face in the first near-infrared image. The contour structure information of the target face in the first near-infrared image may include the position of the left ear, the position of the right ear, the position of the left eye, the position of the right eye, the position of the nose, and the position of the mouth, etc., which can indicate the contour information of the face. The specific location is not limited here.
[0039] Specifically, the preset face location detection model includes a first convolutional layer, a first batch of normalization layers, a first activation layer, and a first output layer. The first output layer includes a first classification layer and a first regression layer. The objective loss function of the first classification layer is the cross-entropy loss function, and the objective loss function of the first regression layer is the Smooth L1 loss function. For ease of understanding, Figure 3This is a schematic diagram of the architecture of the face location detection model provided in this application embodiment. A first near-infrared image is input to a first convolutional layer. The first convolutional layer then performs convolution processing on the first near-infrared image and inputs the convolutionally processed first near-infrared image to a first batch normalization layer. Based on this, the first batch normalization layer performs batch normalization processing on the convolutionally processed first near-infrared image, and then inputs the batch-normalized first near-infrared image to a first activation layer. Similarly, the first activation layer performs activation processing on the batch-normalized first near-infrared image, and then inputs the activated first near-infrared image to a first classification layer and a first regression layer. Finally, the location of the target face is output based on the first classification layer and the first regression layer. It should be understood that... Figure 3 The example provided is for illustrative purposes only. In practical applications, face location detection models may also include multiple first convolutional layers, multiple first-order normalization layers, and multiple first activation layers. Figure 3 The examples provided do not constitute the only limitation of this scheme.
[0040] In one embodiment, the preset face location detection model is pre-trained. The specific training method is as follows: A first set of near-infrared image samples and a set of target face locations are acquired. The target face location set includes the location of the target face in each of the first near-infrared image samples, and each of the first near-infrared image samples is an 8-bit image. Based on this, the first set of near-infrared image samples is used as input to the face location detection model to be trained. The model outputs the predicted location of the target face in each of the first near-infrared image samples. Then, based on the predicted location and the target face location in each of the first near-infrared image samples, the model parameters of the face location detection model to be trained are updated according to the target loss function to obtain the face location detection model.
[0041] The following details the process of updating the model parameters. The location of the target face in each first near-infrared image sample is used as the target for iterative training. Specifically, the loss value of the target loss function is determined based on the difference between the target face location in each first near-infrared image sample and its predicted location. The convergence condition of the loss function is then checked based on this loss value. If convergence is not achieved, the model parameters of the face location detection model are updated using the target loss function value. This process continues until the target loss function converges, resulting in the optimal face location detection model. The convergence condition for the aforementioned target loss function can be that the value of the target loss function is less than or equal to a first preset threshold. For example, the value of the first preset threshold can be 0.005, 0.01, 0.02, or other values close to 0. Alternatively, the difference between two consecutive values of the target loss function can be less than or equal to a second preset threshold. The value of the second threshold can be the same as or different from the value of the first threshold. For example, the value of the second preset threshold can be 0.005, 0.01, 0.02, or other values close to 0. The terminal device can also use other convergence conditions, which are not limited here.
[0042] Furthermore, the following formula (1) shows the Smooth L1 loss function of this embodiment:
[0043]
[0044] Here, x refers to the loss value of the Smooth L1 loss function, which is determined based on the difference between the location of the target face in each first near-infrared image sample and the predicted location of the target face in each first near-infrared image sample.
[0045] Since the first output layer in this scheme includes a first classification layer and a first regression layer, the aforementioned steps specifically involve updating the model parameters of the first classification layer and the first regression layer in the face location detection model to be trained based on the target loss function of the first classification layer and the target loss function of the first regression layer, until the target loss function of the first classification layer and the target loss function of the first regression layer reach the convergence condition, thus obtaining the optimal model parameters of the face location detection model to be trained, thereby obtaining the face location detection model in this scheme.
[0046] Step 203: Perform image matting on the first near-infrared image based on the contour structure information of the target face to obtain a face image to be processed that includes only the target face.
[0047] In this embodiment, the terminal device performs image segmentation processing on the first near-infrared image based on the location of the target face, obtaining a face image to be processed that includes only the target face. For ease of understanding, Figure 4This is a schematic diagram of an embodiment of the face image to be processed provided in this application. In this case, the face image A1 to be processed includes the contour structured information A2 of the target face. It should be understood that... Figure 4 The examples provided are for understanding this scheme only and should not be construed as limiting this application.
[0048] Step 204: Cut the face image to be processed according to the contour structure information of the target face, and randomly stitch the cut images to obtain the target face image.
[0049] In this embodiment, the terminal device processes the face image to be processed to obtain a target face image. Specifically, based on the contour structure information of the target face, the face image to be processed is averaged and divided according to a preset ratio to obtain the divided face image to be processed. For ease of understanding, Figure 5 This is a schematic diagram of an embodiment of the segmented face image to be processed provided in this application. Figure 5 Image (A) shows the face image to be processed, which includes the contour structure information of the target face. Figure 5 Image (B) shows the cropped face image to be processed.
[0050] In one embodiment, the contour structured information of the target face may include 68 or 98 key facial feature points. The face image to be processed is then segmented proportionally based on the coordinates of these facial feature points on the image to obtain the segmented image. It should be noted that the preset ratio can be 1:3, 1:4, 1:5, etc., and is not limited here.
[0051] Furthermore, the segmented face image to be processed is randomly stitched together to obtain the target face image. The target face image obtained after random stitching must maintain the same size as the face image to be processed. For ease of understanding, Figure 6 This is a schematic diagram of an embodiment of the target face image provided in this application. Figure 6 Figure (A) shows the cropped face image to be processed. Figure 6 Image (B) shows a target face image obtained after random stitching. Figure 6 Image (C) shows another target face image obtained by random splicing.
[0052] It should be understood that Figure 5 and Figure 6 The examples provided are for understanding this scheme only and should not be construed as limiting this application.
[0053] Step 205: Based on the target face image and the preset live face detection model, determine the live face detection result of the target face.
[0054] In this embodiment, the terminal device inputs the target face image into a preset live face detection model to obtain the live face detection result. The live face detection result includes: the target face image is a live face image, or the target face image is a non-live face image.
[0055] Specifically, the terminal device obtains a feature map of the target face image based on the target face image using a preset liveness detection model. In this embodiment, the feature map size is 15*15. Then, the feature map of the target face image is averaged to obtain a liveness data score for the target face in the target face image. This liveness data score indicates the probability that the target face is a real person. Based on this, if the liveness data score of the target face is greater than a preset threshold, the liveness detection result indicates that the target face image is a live face image; conversely, if the liveness data score of the target face is less than or equal to the preset threshold, the liveness detection result indicates that the target face image is a non-live face image. For example, taking a preset threshold of 0.5 as an example, if the liveness data score of the target face is greater than 0.5, it means that the probability that the target face is a real person is greater than 0.5, and in this case, the target face image is a live face image. Conversely, if the liveness score of the target face is less than or equal to 0.5, it means that the probability of the target face being a real person is less than or equal to 0.5, and the target face image is a non-live face image.
[0056] Furthermore, the live face detection model in this embodiment includes a second convolutional layer, a first pooling layer, a feature fusion layer, and a third convolutional layer. The objective loss function of the live face detection model is the mean squared error loss (MSE loss). For ease of understanding, Figure 7 This is a schematic diagram of the architecture of the live face detection model provided in this application embodiment. The target face image is input to a second convolutional layer, which then performs convolution processing on the target face image. The convolutionally processed target face image is then input to a first pooling layer. Based on this, the first pooling layer performs pooling processing on the convolutionally processed target face image, and then inputs the pooled target face image to a feature fusion layer. Similarly, the feature fusion layer performs feature fusion processing on the pooled target face image, and then inputs the feature-fused target face image to a third convolutional layer. The third convolutional layer performs convolution processing on the feature-fused target face image and outputs the live face detection result. It should be understood that... Figure 7The example is only for understanding this scheme. In practical applications, the live face detection model can also include multiple second convolutional layers and multiple first pooling layers. After performing a second convolutional layer to a first pooling layer, the result is then input to the next second convolutional layer and the next first pooling layer. That is, a second convolutional layer and a first pooling layer need to be processed continuously without going through other processing layers.
[0057] In one embodiment, the preset live face detection model is pre-trained and deployed on the terminal device. The specific training method includes: acquiring a set of target face image samples and a set of live face detection results. The set of live face detection results includes the live face detection results of each target face image sample. As can be seen from the foregoing embodiment, since the target face image is obtained by randomly recombining the contour structure information of the segmented target face, the target face image sample is obtained by randomly recombining the contour structure information of the target face in the same face image to be processed multiple times. This ensures that the target face image sample input to the live face detection model to be trained is randomly recombined each time, and the distribution of image pixel values is irregular. This greatly improves the richness of the sample data obtained for the same face image to be processed, making the live face detection model pay more attention to the more local and detailed spoof features in the face to be processed, increasing the model's generalization ability, and making the final live face detection model converge better. Based on this, the target face image sample set is used as the input to the live face detection model to be trained. The live face prediction result of each target face image sample is output by the live face detection model to be trained. Then, based on the live face prediction result of each target face image sample and the live face detection result of each target face image sample, the model parameters of the live face detection model to be trained are updated according to the target loss function to obtain the live face detection model.
[0058] The following details the process of updating the aforementioned model parameters. The terminal device uses the liveness detection result of each target face image sample as the target for iterative training. Specifically, the loss value of the target loss function is determined based on the difference between the liveness detection result and the liveness prediction result of each target face image sample. The convergence condition of the loss function is then checked based on this loss value. If convergence is not achieved, the model parameters of the liveness detection model under training are updated using the target loss function value until the target loss function converges. At this point, the optimal model parameters for the liveness detection model under training are obtained, resulting in the optimal liveness detection model.
[0059] It should be understood that the convergence condition of the aforementioned target loss function can be that the value of the target loss function is less than or equal to a first preset threshold. For example, the value of the first preset threshold can be 0.005, 0.01, 0.02, or other values close to 0. Alternatively, the difference between two consecutive values of the target loss function can be less than or equal to a second preset threshold. The value of the second threshold can be the same as or different from the value of the first threshold. For example, the value of the second preset threshold can be 0.005, 0.01, 0.02, or other values close to 0. Terminal devices can also use other convergence conditions, which are not limited here. Furthermore, since the target loss function of the live face detection model in this scheme is the mean squared error loss function, the aforementioned steps specifically involve updating the model parameters of the third convolutional layer in the live face detection model to be trained according to the target loss function of the third convolutional layer until the target loss function of the third convolutional layer reaches the convergence condition, thus obtaining the optimal model parameters of the live face detection model to be trained, thereby obtaining the live face detection model in this scheme.
[0060] The above provides a detailed introduction to the live face detection method in this solution. The following section introduces the live face detection device provided in this solution. Figure 8 This is a schematic diagram of a live face detection device provided in an embodiment of this application, as shown below. Figure 8 As shown, the live face detection device includes:
[0061] The first acquisition module 801 is used to acquire a first near-infrared image, wherein the first near-infrared image includes a target human face;
[0062] The first processing module 802 is used to process the first near-infrared image to obtain the contour structure information of the target face;
[0063] The second processing module 803 is used to perform image matting processing on the first near-infrared image based on the contour structure information of the target face to obtain a face image to be processed that includes only the target face.
[0064] The third processing module 804 is used to cut the face image to be processed according to the contour structure information of the target face, and randomly stitch the cut images to obtain the target face image.
[0065] The second acquisition module 805 is used to acquire a live face detection result based on the target face image and a preset live face detection model, wherein the live face detection result indicates that the target face image is a live face image, or that the target face image is a non-live face image.
[0066] Optionally, in the above Figure 8Based on the corresponding embodiments, in another embodiment of the live face detection device provided in this application, the first acquisition module 801 is specifically used to: acquire a second near-infrared image, wherein the second near-infrared image includes a target face; and perform normalization processing on the second near-infrared image to obtain a first near-infrared image, wherein the first near-infrared image is an 8-bit image.
[0067] Optionally, in the above Figure 8 Based on the corresponding embodiments, in another embodiment of the live face detection device provided in this application, the first processing module 802 is specifically used to: obtain the position of the target face in the first near-infrared image through a face position detection model based on the first near-infrared image; and perform image matting processing on the first near-infrared image based on the position of the target face to obtain the face image to be processed of the target face.
[0068] Optionally, in the above Figure 8 Based on the corresponding embodiments, in another embodiment of the live face detection device provided in this application, the face location detection model includes a first convolutional layer, a first batch normalization layer, a first activation layer and a first output layer. The first output layer includes a first classification layer and a first regression layer. The target loss function of the first classification layer is the cross-entropy loss function, and the target loss function of the first regression layer is the Smooth L1 loss function.
[0069] Optionally, in the above Figure 8 Based on the corresponding embodiments, in another embodiment of the live face detection device provided in this application, the second acquisition module 805 is specifically used for: acquiring feature maps of the target face image through a live face detection model based on the target face image; performing mean calculation processing on the feature maps of the target face image to obtain the live data score of the target face in the target face image; comparing the live data score with a preset threshold, and determining the live face detection result based on the comparison result; wherein, when the live data score of the target face is greater than the preset threshold, the live face detection result indicates that the target face image is a live face image; when the live data score of the target face is less than or equal to the preset threshold, the live face detection result indicates that the target face image is a non-live face image.
[0070] Optionally, in the above Figure 8 Based on the corresponding embodiments, in another embodiment of the live face detection device provided in this application, the live face detection model includes a second convolutional layer, a first pooling layer, a feature fusion layer and a third convolutional layer, and the target loss function of the live face detection model is the mean squared error loss function.
[0071] Next, the terminal device in the embodiments of this application will be described further. Specifically, Figure 9This is a schematic diagram of the structure of one embodiment of the terminal device in this application, as shown below. Figure 9 As shown, the terminal device includes a processor 910 and a memory 920 coupled to the processor 910. In some implementations, they may be coupled together via a bus. The processor 910 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. The processor may also be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The memory 920 stores a computer program 930, which can execute any of the methods described in the possible implementations above. After executing the computer program 930, the processor 910 can perform corresponding operations according to the instructions of the computer program 930. Furthermore, after executing the computer program 930 in the memory 920, the processor 910 can perform all executable operations according to the instructions of the computer program 930.
[0072] In addition, this application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the steps of any of the methods described in the foregoing embodiments.
[0073] This application also provides a computer program product including a program that, when run on a computer, causes the computer to perform the steps performed by the server in any of the methods described in the foregoing embodiments.
[0074] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0075] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, at least two units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0076] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across at least two network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0077] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0078] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal device (which may be a personal computer or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0079] The above-described embodiments are merely illustrative of the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for detecting live faces, characterized in that, include: Acquire a first near-infrared image, wherein the first near-infrared image includes a target human face; The first near-infrared image is processed to obtain the contour structure information of the target face; Based on the contour structure information of the target face, the first near-infrared image is processed to obtain a face image to be processed that includes only the target face. The image of the face to be processed is segmented according to the contour structure information of the target face, and the segmented images are randomly stitched together to obtain the target face image. The target face image is input into a preset live face detection model to obtain the feature map of the target face image; The feature map of the target face image is processed by mean calculation to obtain the liveness score of the target face in the target face image; The liveness detection result is determined based on the liveness data score.
2. The method according to claim 1, characterized in that, The acquisition of the first near-infrared image includes: Acquire a second near-infrared image, wherein the second near-infrared image includes the target face; The second near-infrared image is normalized to obtain the first near-infrared image, wherein the first near-infrared image is an 8-bit image.
3. The method according to claim 2, characterized in that, The process of processing the first near-infrared image to obtain the contour structure information of the target face includes: The first near-infrared image is input into a preset face location detection model to obtain the contour structure information of the target face in the first near-infrared image.
4. The method according to claim 3, characterized in that, The preset face location detection model includes a first convolutional layer, a first batch of normalization layers, a first activation layer, and a first output layer. The first output layer includes a first classification layer and a first regression layer. The target loss function of the first classification layer is the cross-entropy loss function, and the target loss function of the first regression layer is the Smooth L1 loss function.
5. The method according to claim 4, characterized in that, The step of determining the live face detection result based on the liveness data score includes: The liveness data score is compared with a preset threshold, and the liveness detection result is determined based on the comparison result; wherein, when the liveness data score is greater than the preset threshold, the liveness detection result indicates that the target face image is a live face image, and when the liveness data score is less than or equal to the preset threshold, the liveness detection result indicates that the target face image is a non-live face image.
6. The method according to claim 5, characterized in that, The preset live face detection model includes a second convolutional layer, a first pooling layer, a feature fusion layer, and a third convolutional layer. The target loss function of the live face detection model is the mean squared error loss function.
7. A live face detection device, characterized in that, The live face detection device includes: The first acquisition module is used to acquire a first near-infrared image, wherein the first near-infrared image includes a target human face; The first processing module is used to process the first near-infrared image to obtain the contour structure information of the target face; The second processing module is used to perform image matting on the first near-infrared image based on the contour structure information of the target face to obtain a face image to be processed that includes only the target face. The third processing module is used to cut the face image to be processed according to the contour structure information of the target face, and randomly stitch the cut images to obtain the target face image. The second acquisition module is used to input the target face image into a preset live face detection model, acquire the feature map of the target face image, perform mean calculation on the feature map of the target face image to obtain the live data score of the target face in the target face image, and determine the live face detection result based on the live data score.
8. A terminal device, characterized in that, include: Memory and processor; The memory is used to store computer programs; The processor is used to execute a computer program in the memory to implement the steps of the method according to any one of claims 1 to 6.
9. The terminal device as described in claim 8, characterized in that, It also includes a data acquisition module, which the processor uses to acquire a second near-infrared image to implement the steps in the method of any one of claims 2 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Human face living body detection method and device and computer readable storage medium
CN112613471A
Face living body detection method and device and electronic equipment
CN112836625A