Face liveness detection method, device, electronic device and computer storage medium

Through the combination of encoded convolutional neural network and decoded convolutional neural network, detection based on live face sample images is realized, solving the problems of limited detection range and low security, and improving the universality and security of face live face detection.

CN115035559BActive Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110240298.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-04
Publication Date
2025-07-08
Estimated Expiration
2041-03-04

AI Technical Summary

Technical Problem

The existing facial live detection methods have problems such as incomplete coverage and unforeseen unknown attack types when collecting large-scale non-living data, resulting in limited detection range and low security.

Method used

The original feature image is encoded with a coding convolutional neural network, and image reconstruction is carried out by decoding the convolutional neural network, and the reconstruction loss is calculated to determine the detection result. Only the live face sample images are used during the training process to avoid the use of non-living samples.

Benefits of technology

It improves the scope and security of facial live detection, can better predict unknown non-living attacks, and enhances the universality and security of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035559B_ABST
    Figure CN115035559B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a face liveness detection method, device, electronic device, and computer-readable storage medium, which relate to the field of artificial intelligence. The method includes: performing at least one-dimensional feature vector encoding on an original feature image to be detected based on a trained encoding convolutional neural network to obtain at least one-dimensional first feature vectors; performing image reconstruction based on the first feature vectors and a trained decoding convolutional neural network to obtain a first target image after reconstruction; determining a first target loss between the original feature image and the first target image, and determining a detection result of the detection based on the first target loss. The embodiment of the present application not only realizes face liveness detection, but also, since it does not need to be trained based on non-living face sample images, it does not need to make assumptions about non-living objects, and thus has better prediction ability for unknown non-living attacks, improving the security and universality of face liveness detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a face liveness detection method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common face recognition, machine learning, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0003] Face liveness detection is a specific application of face recognition, which is used to determine whether the face image data obtained from the camera is a real person, so as to resist illegal attacks such as using printed photos or photos displayed on an electronic screen to deceive the face recognition system. Since the silent liveness detection method does not require the user to perform interaction actions such as opening the mouth / blinking, the experience is better, so the silent liveness detection method has become a relatively widely used face liveness detection method currently.

[0004] This method mainly uses machine learning methods such as neural networks to classify the input pictures to determine whether the face in the current picture is a live body. This classification method often requires collecting a large number of labeled live and non-live data to train the classification model. The more training data, the better the effect of the trained model. However, in practice, collecting large-scale non-live data has the following disadvantages:

[0005] 1) There are various non-live scenarios. Due to limitations in manpower and material resources, it is difficult to cover all scenarios during the collection process, such as various lighting conditions, so the scope of application of the detection method will be limited;

[0006] 2) It is impossible to collect unforeseen unknown attack types, which weakens the security of the detection method. Summary of the Invention

[0007] This application provides a face liveness detection method, apparatus, electronic device, and computer-readable storage medium, which can solve the problems of limited application range and low security of face liveness detection. The technical solutions are as follows:

[0008] According to one aspect of this application, a face liveness detection method is provided. The method includes:

[0009] Performing at least one-dimensional feature vector encoding on the original feature image to be detected based on the trained encoding convolutional neural network to obtain at least one-dimensional first feature vector;

[0010] Performing image reconstruction based on the first feature vector and the trained decoding convolutional neural network to obtain the reconstructed first target image;

[0011] Determine the first target loss between the original feature image and the first target image, and determine the detection result of the detection based on the first target loss.

[0012] In one or more embodiments, the encoding convolutional neural network obtained by training is used to perform at least one-dimensional feature vector encoding on the original feature image to be detected, and obtain at least one-dimensional first feature vectors, including:

[0013] Scale the original size of the original feature image to a preset target size to obtain a scaled image;

[0014] Use the encoding convolutional neural network to perform convolutional layer processing, pooling layer processing, and activation layer processing on the scaled image to obtain at least one-dimensional first feature vectors.

[0015] In one or more embodiments, the first target image after reconstruction is obtained by performing image reconstruction based on the first feature vectors and the decoding convolutional neural network obtained by training, including:

[0016] Perform image reconstruction on the first feature vectors through the decoding convolutional neural network to obtain a first target image with a target size.

[0017] In one or more embodiments, the determining the first target loss between the original feature image and the first target image, and determining the detection result of the detection based on the first target loss, includes:

[0018] Calculate the reconstruction loss between the original feature image and the first target image;

[0019] When the reconstruction loss is less than the reconstruction loss threshold, determine that the detection result of the detection is a live face;

[0020] When the reconstruction loss is not less than the reconstruction loss threshold, determine that the detection result of the detection is a non-live face.

[0021] In one or more embodiments, the calculating the reconstruction loss between the original feature image and the first target image includes:

[0022] Calculate the mean square loss, absolute value loss, and structural similarity loss between the original feature image and the first target image;

[0023] Use the weighted sum of the mean square loss, absolute value loss, and structural similarity loss as the reconstruction loss.

[0024] In one or more embodiments, the encoding convolutional neural network and the decoding convolutional neural network are obtained by training in the following manner:

[0025] Training steps: encoding the sample images with annotations using the encoding convolutional neural network to obtain a second feature vector with at least one dimension;

[0026] Based on the second feature vector, performing image reconstruction using the decoding convolutional neural network to obtain a reconstructed second target image;

[0027] Determining a second target loss between the original feature image and the second target image;

[0028] When the second target loss is not less than the target loss threshold, updating the encoding convolutional neural network and the decoding convolutional neural network using a preset stochastic gradient descent algorithm to obtain an updated target encoding convolutional neural network and a target decoding convolutional neural network, and using the target encoding convolutional neural network as the current encoding convolutional neural network, and using the target decoding convolutional neural network as the current decoding convolutional neural network;

[0029] Repeating the training steps until the second target loss is less than the target loss threshold to obtain a trained encoding convolutional neural network and a decoding convolutional neural network.

[0030] According to another aspect of the present application, there is provided a face liveness detection device, the device comprising:

[0031] An encoding module, configured to encode a raw feature image to be detected into a feature vector with at least one dimension using a trained encoding convolutional neural network to obtain a first feature vector with at least one dimension;

[0032] A reconstruction module, configured to perform image reconstruction based on the first feature vector and a trained decoding convolutional neural network to obtain a reconstructed first target image;

[0033] A detection module, configured to determine a first target loss between the original feature image and the first target image, and determine a detection result of the detection based on the first target loss.

[0034] In one or more embodiments, the encoding module includes:

[0035] A scaling sub-module, configured to scale the original size of the original feature image to a preset target size to obtain a scaled image;

[0036] An encoding sub-module, configured to perform convolutional layer processing, pooling layer processing, and activation layer processing on the scaled image using the encoding convolutional neural network to obtain a first feature vector with at least one dimension.

[0037] In one or more embodiments, the reconstruction module is specifically configured to:

[0038] The first target image with a preset target size is obtained by reconstructing the first feature vector through the decoding convolutional neural network.

[0039] In one or more embodiments, the detection module includes:

[0040] A calculation sub-module, configured to calculate the reconstruction loss between the original feature image and the first target image;

[0041] A determination sub-module, configured to determine that the detection result of the detection is a live face when the reconstruction loss is less than a reconstruction loss threshold; and to determine that the detection result of the detection is a non-live face when the reconstruction loss is not less than the reconstruction loss threshold.

[0042] In one or more embodiments, the calculation sub-module includes:

[0043] A first calculation unit, configured to calculate the mean square loss, absolute value loss, and structural similarity loss between the original feature image and the first target image;

[0044] A second calculation unit, configured to use the weighted sum of the mean square loss, absolute value loss, and structural similarity loss as the reconstruction loss.

[0045] In one or more embodiments, it further includes:

[0046] An encoding module, further configured to encode a sample image with annotations through the encoding convolutional neural network to obtain a second feature vector with at least one dimension;

[0047] A reconstruction module, further configured to reconstruct an image based on the second feature vector through the decoding convolutional neural network to obtain a reconstructed second target image;

[0048] A detection module, further configured to determine a second target loss between the original feature image and the second target image;

[0049] An update module, configured to, when the second target loss is not less than a target loss threshold, update the encoding convolutional neural network and the decoding convolutional neural network by using a preset stochastic gradient descent algorithm to obtain an updated target encoding convolutional neural network and a target decoding convolutional neural network, and use the target encoding convolutional neural network as the current encoding convolutional neural network, and use the target decoding convolutional neural network as the current decoding convolutional neural network;

[0050] The encoding module, the reconstruction module, the detection module, and the update module are repeatedly called until the second target loss is less than the target loss threshold, to obtain a trained encoding convolutional neural network and a decoding convolutional neural network.

[0051] According to another aspect of the present application, there is provided an electronic device, which includes:

[0052] One or more processors;

[0053] A memory;

[0054] One or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the face liveness detection method shown in the first aspect.

[0055] According to still another aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the face liveness detection method shown in the first aspect of the present application.

[0056] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementations of any of the above aspects.

[0057] The beneficial effects brought by the technical solution provided by the present application are:

[0058] In the embodiments of the present invention, for the original feature image to be detected, a trained encoding convolutional neural network is used to perform at least one-dimensional feature vector encoding to obtain at least one-dimensional first feature vector, and then a trained decoding convolutional neural network is used to reconstruct the first feature vector to obtain a reconstructed first target image, and then the reconstruction loss between the first target image and the original feature image is calculated, and based on the calculated reconstruction loss, it is determined whether the face in the original feature image is a live body or a non-live body. In this way, not only the face liveness detection is realized, but also, since the encoding convolutional neural network and the decoding convolutional neural network are trained based on the face sample images of live bodies and do not need to be trained based on the face sample images of non-live bodies, the application scope of face liveness detection will not be restricted. At the same time, it is not necessary to make assumptions about non-live bodies, and thus has better prediction ability for unknown non-live body attacks, improving the security and universality of face liveness detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application.

[0060] Figure 1 Schematic diagram of the application environment of a face liveness detection method provided by an embodiment of the present application;

[0061] Figure 2 Schematic flow chart of a face liveness detection method provided by an embodiment of the present application;

[0062] Figure 3 Schematic flow chart of the encoding process provided by an embodiment of the present application;

[0063] Figure 4 Schematic flow chart of the decoding process provided by an embodiment of the present application;

[0064] Figure 5 Schematic logic diagram of a face liveness detection method provided by an embodiment of the present application;

[0065] Figure 6 Schematic flow chart of the training process of the encoding convolutional neural network and the decoding convolutional neural network provided by an embodiment of the present application;

[0066] Figure 7 For the present application Figure 6 Corresponding logic diagram;

[0067] Figure 8 Schematic structural diagram of a face liveness detection device provided by an embodiment of the present application;

[0068] Figure 9 Schematic structural diagram of an electronic device for face liveness detection provided by an embodiment of the present application. Detailed implementation manners

[0069] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as limiting the present application.

[0070] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of this application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0071] To make the objectives, technical solutions and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0072] First, several terms related to this application will be introduced and explained:

[0073] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.

[0074] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0075] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0076] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0077] In this application, face liveness detection can be performed based on artificial intelligence. Specifically, computer vision technology is used for face liveness detection. Face liveness detection refers to determining whether the pictures used for face recognition in a face recognition system are real live human faces or non-live human faces such as printed photos, photos displayed on electronic device screens, or face masks. An artificial neural network can be used during the detection process, and the artificial neural network can be trained based on machine learning.

[0078] Face liveness detection is a specific application of face recognition, which is used to determine whether the face image data obtained from a camera is a real person, so as to resist illegal attacks that use printed photos or photos displayed on electronic screens to deceive the face recognition system. Since the silent liveness detection method does not require users to perform interactive actions such as opening the mouth / blinking, and has a better experience, the silent liveness detection method has become a relatively widely used face liveness detection method currently.

[0079] This method mainly uses machine learning methods such as neural networks to classify the input images to determine whether the face in the current image is a live body. This classification method often requires collecting a large amount of labeled live and non-live data to train the classification model. The more training data, the better the effect of the trained model. However, in practice, collecting a large amount of non-live data has the following disadvantages:

[0080] 1) There are various non-live scenarios. Due to limitations in human and material resources, it is difficult to cover all scenarios during the collection process, such as various lighting conditions, so the scope of application of the detection method will be limited;

[0081] 2) It is impossible to collect unforeseen unknown attack types, which weakens the security of the detection method.

[0082] The face liveness detection method, device, electronic device and computer-readable storage medium provided in this application aim to solve the above technical problems in the prior art.

[0083] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0084] An application environment for a face liveness detection method is provided in an embodiment of the present invention. Refer to Figure 1 , this application environment includes: a first device 101 and a second device 102. The first device 101 and the second device 102 are connected through a network. The first device 101 is an access device, and the second device 102 is an accessed device. The first device 101 can be a terminal, and the second device 102 can be a server. The terminal can have the following characteristics:

[0085] (1) In terms of the hardware system, the device has a central processing unit, a memory, an input component, and an output component. That is to say, the device is often a microcomputer device with communication functions. In addition, it can have various input methods, such as a keyboard, a mouse, a touch screen, a microphone, and a camera, and can be adjusted according to needs. At the same time, the device often has various output methods, such as a receiver, a display screen, etc., and can also be adjusted according to needs;

[0086] (2) In terms of the software system, the device must have an operating system, such as Windows Mobile, Symbian, Palm, Android, iOS, etc. At the same time, these operating systems are becoming more and more open, and personalized application programs developed based on these open operating system platforms are emerging in an endless stream, such as an address book, a schedule, a notepad, a calculator, and various games, which greatly meet the needs of personalized users;

[0087] (3) In terms of communication capabilities, the device has flexible access methods and high-bandwidth communication performance, and can automatically adjust the selected communication method according to the selected service and the environment it is in, thus facilitating user use. The device can support mobile communications of 3GPP (3rd Generation Partnership Project), 4GPP (4th Generation Partnership Project), 5GPP (5th Generation Partnership Project), LTE (Long Term Evolution), and WIMAX (World Interoperability for Microwave Access), computer network communications based on TCP / IP (Transmission Control Protocol / Internet Protocol) and UDP (User Datagram Protocol) protocols, and short-range wireless transmission methods based on Bluetooth and infrared transmission standards. It not only supports voice services, but also supports a variety of wireless data services;

[0088] (4) In terms of function usage, the device pays more attention to humanization, personalization, and multi-functionality. With the development of computer technology, the device has entered a "human-centered" mode from a "device-centered" mode, integrating embedded computing, control technology, artificial intelligence technology, and biometric authentication technology, etc., fully reflecting the people-oriented principle. Due to the development of software technology, the device can adjust settings according to personal needs and is more personalized. At the same time, the device itself integrates a large number of software and hardware, and its functions are becoming more and more powerful.

[0089] The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0090] In this application, the artificial neural network can be deployed in the second device. In this way, the user can initiate a face liveness detection request on the first device. The first device sends the face liveness detection request to the second device. The second device performs face liveness detection, obtains the detection result, and feeds back the detection result to the first device.

[0091] The artificial neural network can also be deployed in the first device. In this way, the user can initiate a face liveness detection request on the first device, and the first device directly performs face liveness detection using the artificial neural network to obtain a detection result.

[0092] In practical applications, the deployment of the artificial neural network can be set according to actual needs, and the present application does not limit this.

[0093] In the above application environment, a face liveness detection method can be executed, as Figure 2 shown. The method includes:

[0094] Step S201: Based on the trained encoded convolutional neural network, perform at least one-dimensional feature vector encoding on the original feature image to be detected, and obtain at least one-dimensional first feature vector;

[0095] Among them, the original feature image to be detected can be an image obtained by the first device, and the image includes at least one-dimensional feature vector. For example, it can be an image selected by the user from the album (including at least one image) of the first device, or an image captured in real time by the image capture device (such as a camera) of the first device, or the original feature image to be detected can be obtained through other means. In practical applications, it can be adjusted according to actual needs, and the embodiments of the present invention do not limit this.

[0096] Furthermore, the embodiments of the present invention only take the face liveness detection of one original feature image as an example for illustration. In practical applications, the original feature image to be detected can be one, that is, perform face liveness detection on one original feature image, or multiple, that is, perform face liveness detection on multiple original feature images simultaneously. In practical applications, it can be set according to actual needs, and the embodiments of the present invention do not limit this either.

[0097] After the first device obtains the original feature image to be detected, it can use the trained encoded convolutional neural network to perform at least one-dimensional feature vector encoding on the original feature image, and the encoded output form is a one-dimensional or multi-dimensional feature vector. Among them, the encoding can be performed in the first device, or the first device first sends the original feature image to the second device, and then performs the encoding in the second device. In practical applications, it can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0098] Step S202: Based on the first feature vector and the trained decoded convolutional neural network, perform image reconstruction to obtain the reconstructed first target image;

[0099] After encoding the original feature image to obtain the first feature vector, the trained decoding convolutional neural network can be used to decode the first feature vector and reconstruct the image to obtain the reconstructed first target image.

[0100] Step S203: Determine the first target loss between the original feature image and the first target image, and determine the detection result of the detection based on the first target loss.

[0101] After obtaining the reconstructed first target image through reconstruction, the original feature image and the first target image can be compared to determine the first target loss between the two, and then the detection result of face liveness detection can be determined based on the first target loss. Among them, the detection result can be that the face in the original feature image is a live body, or the face in the original feature image is a non-live body.

[0102] In the embodiment of the present invention, at least one-dimensional feature vector encoding is performed on the original feature image to be detected by using the trained encoding convolutional neural network to obtain at least one-dimensional first feature vector, and then the trained decoding convolutional neural network is used to reconstruct the image of the first feature vector to obtain the reconstructed first target image, and then the reconstruction loss between the first target image and the original feature image is calculated, and based on the calculated reconstruction loss, it is determined whether the face in the original feature image is a live body or a non-live body. In this way, not only face liveness detection is realized, but also, since the encoding convolutional neural network and the decoding convolutional neural network are trained based on the face sample images of live bodies and do not need to be trained based on the face sample images of non-live bodies, the application range of face liveness detection will not be limited. At the same time, it is not necessary to make assumptions about non-live bodies, and thus has better prediction ability for unknown non-live body attacks, improving the security and universality of face liveness detection.

[0103] In the embodiment of the present invention, Figure 2 each of the steps in a face liveness detection method shown in

[0104] Step S201: Perform at least one-dimensional feature vector encoding on the original feature image to be detected by using the trained encoding convolutional neural network to obtain at least one-dimensional first feature vector;

[0105] Among them, the original feature image to be detected can be an image acquired by a first device, and the image includes at least one-dimensional feature vector. For example, it can be an image selected by a user from the album (including at least one image) of the first device, or it can be an image acquired in real time by the image acquisition device (such as a camera) of the first device, or the original feature image to be detected can be obtained through other means, and can be adjusted according to actual needs in practical applications, and the embodiment of the present invention does not limit this.

[0106] Furthermore, the embodiments of the present invention only take the face liveness detection of a single original feature image as an example for illustration. In practical applications, the original feature image to be detected can be one, that is, the face liveness detection is performed on a single original feature image, or can be multiple, that is, the face liveness detection is performed on multiple original feature images simultaneously. In practical applications, it can be set according to actual needs, and the embodiments of the present invention do not limit this either.

[0107] After the first device obtains the original feature image to be detected, it can use the trained encoding convolutional neural network to perform at least one-dimensional feature vector encoding on the original feature image, and the encoded output form is a one-dimensional or multi-dimensional feature vector. Among them, the encoding can be performed in the first device, or the first device first sends the original feature image to the second device, and then the encoding is performed in the second device. In practical applications, it can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0108] In a preferred embodiment of the present invention, performing at least one-dimensional feature vector encoding on the original feature image to be detected based on the trained encoding convolutional neural network to obtain at least one-dimensional first feature vector includes:

[0109] Scaling the original size of the original feature image to a preset target size to obtain a scaled image;

[0110] Performing convolutional layer processing, pooling layer processing, and activation layer processing on the scaled image through the pre-trained encoding convolutional neural network to obtain at least one-dimensional first feature vector.

[0111] Specifically, before encoding the original feature image, preprocessing can be performed on the original first. The preprocessing includes but is not limited to scaling the original feature image, scaling the original feature image from the original size to a preset target size to obtain a scaled image. For example, the preset target size is 240*240. When the size of the original feature image is 800*800, then it is scaled down to 240*240, or when the size of the original feature image is 125*125, then it is scaled up to 240*240. It should be noted that in practical applications, the size of the target size can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0112] After obtaining the scaled image, the scaled image can be encoded through the pre-trained encoding convolutional neural network. As Figure 3As shown, the pre-trained encoded convolutional neural network may include multiple layers of networks, and each layer of network includes but is not limited to a convolutional layer, a pooling layer, and an activation layer. The scaled image passes through the multiple layers of networks. Among them, each layer of network can perform convolutional layer processing, pooling layer processing, and activation layer processing. The encoded convolutional neural network can finally output the encoding of the original feature image, that is, a one-dimensional or multi-dimensional first feature vector. It should be noted that in practical applications, the number of network layers of the encoded convolutional neural network can be preset according to actual needs, and the embodiments of the present invention do not limit this either.

[0113] Step S202, based on the first feature vector and the trained decoded convolutional neural network, perform image reconstruction to obtain the reconstructed first target image;

[0114] After encoding the original feature image to obtain the first feature vector, the trained decoded convolutional neural network can be used to decode and reconstruct the first feature vector to obtain the reconstructed first target image.

[0115] In a preferred embodiment of the present invention, performing image reconstruction based on the first feature vector to obtain the reconstructed first target image includes:

[0116] Perform image reconstruction on the first feature vector through the decoded convolutional neural network to obtain the first target image with a preset target size.

[0117] Specifically, as Figure 4 shown, the pre-trained decoded convolutional neural network may also include multiple layers of networks, and each layer of network includes but is not limited to an activation layer and a transposed convolutional layer. During reconstruction, the first feature vector passes through the multiple layers of networks. Among them, each layer of network can perform activation layer processing and transposed convolutional layer processing. The decoded convolutional neural network can finally output the first target image with a preset target size. For example, continuing with the above example, if the size of the scaled image is 240*240, then the size of the reconstructed first target image is also 240*240.

[0118] It should be noted that in practical applications, the number of network layers of the decoded convolutional neural network can be preset according to actual needs, and the embodiments of the present invention do not limit this. Further, the encoded convolutional neural network and the decoded convolutional neural network can be deployed in the same face liveness detection model. In this way, when performing face liveness detection, the face liveness detection model can be directly used for detection; they can also be deployed in different detection models. In this way, when performing face liveness detection, multiple detection models can be used for joint detection; of course, other methods can also be used, which can be set according to actual needs in practical applications, and the embodiments of the present invention do not limit this either.

[0119] Step S203: Determine the first target loss between the original feature image and the first target image, and determine the detection result of the detection based on the first target loss.

[0120] After obtaining the reconstructed first target image through reconstruction, the original feature image and the first target image can be compared to determine the first target loss between the two. Then, based on the first target loss, the detection result of face liveness detection can be determined. Among them, the detection result can be that the face in the original feature image is a live body, or the face in the original feature image is a non-live body.

[0121] In a preferred embodiment of the present invention, determining the first target loss between the original feature image and the first target image, and determining the detection result of the detection based on the first target loss includes:

[0122] Calculate the reconstruction loss between the original feature image and the first target image;

[0123] When the reconstruction loss is less than the reconstruction loss threshold, determine that the detection result of the detection is a live face;

[0124] When the reconstruction loss is not less than the reconstruction loss threshold, determine that the detection result of the detection is a non-live face.

[0125] Specifically, determining the first target loss between the original feature image and the first target image includes, but is not limited to, calculating the reconstruction loss between the two. If the calculated reconstruction loss is less than the reconstruction loss threshold, it is determined that the detection result of face liveness detection is a live face; if the calculated reconstruction loss is not less than the reconstruction loss threshold, it is determined that the detection result of face liveness detection is a non-live face. For example, if the calculated reconstruction loss between the original feature image and the first target image is 0.08, which is less than the reconstruction loss threshold of 0.1, then it can be determined that the face in the original feature image is a live body. Then, the detection result is displayed, and the user can know the detection result of the original feature image.

[0126] Among them, calculating the reconstruction loss between the original feature image and the first target image includes:

[0127] Calculate the mean square loss, absolute value loss, and structural similarity loss between the original feature image and the first target image;

[0128] Take the weighted sum of the mean square loss, absolute value loss, and structural similarity loss as the reconstruction loss.

[0129] In the embodiments of the present invention, the reconstruction loss may include, but is not limited to, mean square loss, absolute value loss, and structural similarity loss. Among them, the mean square loss is to calculate the Euclidean distance between the predicted value and the true value. The smaller the Euclidean distance, the closer the predicted value and the true value are, the smaller the mean square error between the two, and the more similar the two are; the absolute value loss is to calculate the absolute value of the difference between the predicted value and the target value. The smaller the absolute value, the closer the predicted value and the true value are, and the more similar the two are; the structural similarity loss is to calculate the similarity of the structures between two images, mainly including three aspects of indicators: illumination, contrast, and structure.

[0130] After separately calculating the mean square loss, absolute value loss, and structural similarity loss between the original feature image and the first target image, the weighted sum of the three is used as the final reconstruction loss. Specifically, the mean square loss, absolute value loss, and structural similarity loss each have corresponding weights. For example, the mean square loss is denoted as L2, and its corresponding weight is denoted as W0; the absolute value loss is denoted as L1, and its corresponding weight is denoted as W1; the structural similarity loss is denoted as L S , and its corresponding weight is denoted as W2. Then the final reconstruction loss L can be calculated by the following formula:

[0131] L = W0 * L2 + W1 * L1 + W2 * L S ;

[0132] Compared with using a single loss, using the weighted sum of multiple losses in the embodiments of the present invention can enable the encoding convolutional neural network and the decoding convolutional neural network to learn more data representations of face liveness during the training process, thereby improving the accuracy of face liveness detection during the detection process.

[0133] For ease of understanding, Figure 5 a logical schematic diagram of a face liveness detection method according to an embodiment of the present invention is shown. Specifically, after obtaining the original feature image to be detected, the original feature image is input into the trained encoding convolutional neural network. The encoding convolutional neural network encodes the original feature image and outputs a one-dimensional or multi-dimensional first feature vector. Then the first feature vector is input into the trained decoding convolutional neural network. The decoding convolutional neural network decodes the first feature vector and outputs the reconstructed first target image. Among them, the size of the first target image is the same as that of the original feature image. Then the reconstruction loss between the first target image and the original feature image is calculated. If the reconstruction loss is less than the reconstruction loss threshold, it is determined that the face in the original feature image is a live body; if the reconstruction loss is not less than the reconstruction loss threshold, it is determined that the face in the original feature image is a non-live body.

[0134] In an embodiment of the present invention, for an original feature image to be detected, a trained encoding convolutional neural network is used to perform at least one-dimensional feature vector encoding to obtain at least one-dimensional first feature vector. Then, a trained decoding convolutional neural network is used to perform image reconstruction on the first feature vector to obtain a reconstructed first target image. Next, the reconstruction loss between the first target image and the original feature image is calculated, and based on the calculated reconstruction loss, it is determined whether the face in the original feature image is a live body or a non-live body. In this way, not only face liveness detection is realized, but also, since the encoding convolutional neural network and the decoding convolutional neural network are trained based on live face sample images and do not need to be trained based on non-live face sample images, the application range of face liveness detection is not restricted. At the same time, there is no need to make assumptions about non-live bodies, and thus, there is better predictive ability for unknown non-live attacks, improving the security and universality of face liveness detection.

[0135] Further, in an embodiment of the present invention, the final reconstruction loss is obtained based on mean square loss, absolute value loss, and structural similarity loss. Compared with using a single loss, using a weighted sum of multiple losses can enable the encoding convolutional neural network and the decoding convolutional neural network to learn more data representations of face liveness during training, further improving the security and universality of face liveness detection.

[0136] In an embodiment of the present invention, the training methods of the encoding convolutional neural network and the decoding convolutional neural network pre-trained in the above embodiment are described in detail, as Figure 6 shown. The training method includes:

[0137] Step S601: Use an encoding convolutional neural network to encode a sample image with annotations to obtain at least one-dimensional second feature vector;

[0138] Step S602: Based on the second feature vector, use a decoding convolutional neural network to perform image reconstruction to obtain a reconstructed second target image;

[0139] Step S603: Determine the second target loss between the original feature image and the second target image;

[0140] Specifically, Steps S601 to S603 are the same in principle as Steps S201 to 203. For details, reference can be made to Steps S201 to 203. To avoid repetition, they will not be elaborated here.

[0141] It should be noted that the difference between Step S601 and Step S201 is that during training, the input sample image has annotations, and the type of annotation is face liveness. That is to say, during training, the faces in the sample images are all live bodies, and there is no need to input non-live face sample images.

[0142] Furthermore, the difference between the first feature vector and the second feature vector is only in expression, and this difference is only for convenience of understanding, and there is no substantial difference between the two; the difference between the first target image and the second target image is only in expression, and this difference is only for convenience of understanding, and there is no substantial difference between the two; the difference between the first target loss and the second target loss is only in expression, and this difference is only for convenience of understanding, and there is no substantial difference between the two.

[0143] Step S604, when the second target loss is not less than the target loss threshold, update the encoding convolutional neural network and the decoding convolutional neural network using a preset stochastic gradient descent algorithm to obtain the updated target encoding convolutional neural network and target decoding convolutional neural network, and use the target encoding convolutional neural network as the current encoding convolutional neural network, and, use the target decoding convolutional neural network as the current decoding convolutional neural network;

[0144] Specifically, the second target loss includes but is not limited to the reconstruction loss. When the second target loss is not less than the target loss threshold, update the parameters of the encoding convolutional neural network and the decoding convolutional neural network using a preset stochastic gradient descent algorithm, so as to obtain the updated target encoding convolutional neural network and target decoding convolutional neural network, and use the target encoding convolutional neural network as the current encoding convolutional neural network, and, use the target decoding convolutional neural network as the current decoding convolutional neural network.

[0145] Step S605, repeat steps S601 to S605 until the second target loss is less than the target loss threshold to obtain the trained encoding convolutional neural network and decoding convolutional neural network.

[0146] After obtaining the current encoding convolutional neural network and the current decoding convolutional neural network, repeat steps S601 to S605 until the second target loss is less than the target loss threshold, and then the trained encoding convolutional neural network and decoding convolutional neural network can be obtained.

[0147] For ease of understanding, Figure 7A logical schematic diagram showing an embodiment of the present invention. Specifically, after obtaining a sample image with annotations, the sample image is input into an untrained encoding convolutional neural network. The encoding convolutional neural network encodes the sample image and outputs a one-dimensional or multi-dimensional second feature vector. Then, the second feature vector is input into an untrained decoding convolutional neural network. The decoding convolutional neural network decodes the second feature vector and outputs a reconstructed second target image. Here, the second target image has the same size as the sample image. Then, the reconstruction loss (i.e., the second target loss) between the second target image and the sample image is calculated. If the reconstruction loss is not less than the reconstruction loss threshold, the parameters of the encoding convolutional neural network and the decoding convolutional neural network are updated using a preset stochastic gradient descent algorithm until the reconstruction loss is less than the reconstruction loss threshold, obtaining the trained encoding convolutional neural network and decoding convolutional neural network.

[0148] In an embodiment of the present invention, an encoding convolutional neural network is used to encode a sample image with annotations to obtain at least a one-dimensional second feature vector. Then, based on the second feature vector, a decoding convolutional neural network is used for image reconstruction to obtain a reconstructed second target image. Next, the second target loss between the original feature image and the second target image is determined. When the second target loss is not less than the target loss threshold, a preset stochastic gradient descent algorithm is used to update the encoding convolutional neural network and the decoding convolutional neural network, obtaining an updated target encoding convolutional neural network and target decoding convolutional neural network. The target encoding convolutional neural network is used as the current encoding convolutional neural network, and the target decoding convolutional neural network is used as the current decoding convolutional neural network. The training steps are repeated until the second target loss is less than the target loss threshold, obtaining the trained encoding convolutional neural network and decoding convolutional neural network. Since the encoding convolutional neural network and the decoding convolutional neural network are trained based on live face sample images and do not need to be trained based on non-live face sample images, the scope of use of face liveness detection is not restricted. At the same time, there is no need to make assumptions about non-live objects, and thus it has better predictive ability for unknown non-live attacks, improving the security and universality of face liveness detection.

[0149] Furthermore, in an embodiment of the present invention, the final reconstruction loss is obtained based on mean square loss, absolute value loss, and structural similarity loss. Compared with using a single loss, using a weighted sum of multiple losses can enable the encoding convolutional neural network and the decoding convolutional neural network to learn more data representations of face liveness during training, further improving the security and universality of face liveness detection.

[0150] Figure 8 The structural schematic diagram of a face liveness detection device provided by an embodiment of the present application is as Figure 8 shown. The device in this embodiment may include:

[0151] An encoding module 801, configured to perform at least one-dimensional feature vector encoding on an original feature image to be detected based on a trained encoding convolutional neural network, so as to obtain at least one-dimensional first feature vector;

[0152] A reconstruction module 802, configured to perform image reconstruction based on the first feature vector and a trained decoding convolutional neural network to obtain a first target image after reconstruction;

[0153] A detection module 803, configured to determine a first target loss between the original feature image and the first target image, and determine a detection result of the detection based on the first target loss.

[0154] In a preferred embodiment of the present invention, the encoding module includes:

[0155] A scaling sub-module, configured to scale the original size of the original feature image to a preset target size to obtain a scaled image;

[0156] An encoding sub-module, configured to perform convolutional layer processing, pooling layer processing, and activation layer processing on the scaled image by using an encoding convolutional neural network to obtain at least one-dimensional first feature vector.

[0157] In a preferred embodiment of the present invention, the reconstruction module is specifically configured to:

[0158] Perform image reconstruction on the first feature vector through a decoding convolutional neural network to obtain a first target image with a preset target size.

[0159] In a preferred embodiment of the present invention, the detection module includes:

[0160] A calculation sub-module, configured to calculate a reconstruction loss between the original feature image and the first target image;

[0161] A determination sub-module, configured to determine that the detection result of the detection is a human face liveness when the reconstruction loss is less than a reconstruction loss threshold; and determine that the detection result of the detection is a non-human face liveness when the reconstruction loss is not less than the reconstruction loss threshold.

[0162] In a preferred embodiment of the present invention, the calculation sub-module includes:

[0163] A first calculation unit, configured to calculate the mean square loss, absolute value loss, and structural similarity loss between the original feature image and the first target image;

[0164] A second calculation unit, configured to use the weighted sum of the mean square loss, absolute value loss, and structural similarity loss as the reconstruction loss.

[0165] In a preferred embodiment of the present invention, it further includes:

[0166] The encoding module is further configured to perform encoding on the sample image with annotations by using an encoding convolutional neural network to obtain a second feature vector with at least one dimension;

[0167] The reconstruction module is further configured to perform image reconstruction on the second feature vector by using a decoding convolutional neural network to obtain a reconstructed second target image;

[0168] The detection module is further configured to determine a second target loss between the original feature image and the second target image;

[0169] The updating module is configured to, when the second target loss is not less than the target loss threshold, update the encoding convolutional neural network and the decoding convolutional neural network by using a preset stochastic gradient descent algorithm to obtain an updated target encoding convolutional neural network and an updated target decoding convolutional neural network, and use the target encoding convolutional neural network as the current encoding convolutional neural network, and use the target decoding convolutional neural network as the current decoding convolutional neural network;

[0170] Repeatedly call the encoding module, the reconstruction module, the detection module, and the updating module until the second target loss is less than the target loss threshold to obtain a trained encoding convolutional neural network and a trained decoding convolutional neural network.

[0171] The face liveness detection device in this embodiment can execute the face liveness detection method shown in the foregoing embodiments of this application, and its implementation principle is similar and will not be elaborated here.

[0172] In an embodiment of the present invention, at least one-dimensional feature vector encoding is performed on the original feature image to be detected by using the trained encoding convolutional neural network to obtain a first feature vector with at least one dimension, and then the trained decoding convolutional neural network is used to perform image reconstruction on the first feature vector to obtain a reconstructed first target image, and then the reconstruction loss between the first target image and the original feature image is calculated, and based on the calculated reconstruction loss, it is determined whether the face in the original feature image is a live body or a non-live body. In this way, not only face liveness detection is realized, but also, since the encoding convolutional neural network and the decoding convolutional neural network are trained based on live face sample images and do not need to be trained based on non-live face sample images, the application range of face liveness detection is not limited. At the same time, it is not necessary to make assumptions about non-live bodies, and thus has better prediction ability for unknown non-live attacks, improving the security and universality of face liveness detection.

[0173] Furthermore, in the embodiments of the present invention, the final reconstruction loss is obtained based on the mean square loss, the absolute value loss, and the structural similarity loss. Compared with using only one kind of loss, using the weighted sum of multiple losses can enable the encoding convolutional neural network and the decoding convolutional neural network to learn more data representations of face liveness during the training process, further improving the security and universality of face liveness detection.

[0174] In the embodiments of the present application, an electronic device is provided. The electronic device includes: a memory and a processor; at least one program stored in the memory and configured to, when executed by the processor, compared with the prior art, achieve: in the embodiments of the present invention, for the original feature image to be detected, at least one-dimensional feature vector encoding is performed using the trained encoding convolutional neural network to obtain at least one-dimensional first feature vector, then the trained decoding convolutional neural network is used to perform image reconstruction on the first feature vector to obtain the reconstructed first target image, and then the reconstruction loss between the first target image and the original feature image is calculated, and based on the calculated reconstruction loss, it is determined whether the face in the original feature image is a live body or a non-live body. In this way, not only face liveness detection is achieved, but also, since the encoding convolutional neural network and the decoding convolutional neural network are trained based on the face sample images of live bodies and do not need to be trained based on the face sample images of non-live bodies, the application range of face liveness detection will not be restricted. At the same time, there is no need to make assumptions about non-live bodies, and thus it has better prediction ability for unknown non-live body attacks, improving the security and universality of face liveness detection.

[0175] In an alternative embodiment, an electronic device is provided, as Figure 9 shown. Figure 9 The electronic device 9000 shown includes: a processor 9001 and a memory 9003. Among them, the processor 9001 and the memory 9003 are connected, such as connected through a bus 9002. Optionally, the electronic device 9000 may further include a transceiver 9004, and the transceiver 9004 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 9004 is not limited to one, and the structure of the electronic device 9000 does not constitute a limitation to the embodiments of the present application.

[0176] The processor 9001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 9001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0177] The bus 9002 may include a path for transmitting information between the above components. The bus 9002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 9002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0178] The memory 9003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0179] The memory 9003 is used to store the application program code (computer program) for executing the solution of this application, and is controlled by the processor 9001 for execution. The processor 9001 is used to execute the application program code stored in the memory 9003 to implement the content shown in the foregoing method embodiments.

[0180] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.

[0181] The embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.

[0182] It should be understood that although the steps in the flowchart of the accompanying drawings are displayed in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps is not strictly limited in order, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same moment, but can be executed at different moments, and their execution order does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0183] The above are only some embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A face liveness detection method, characterized in that, Including: Performing at least one-dimensional feature vector encoding on the original feature image to be detected based on the trained encoding convolutional neural network to obtain at least one-dimensional first feature vector; Performing image reconstruction based on the first feature vector and the trained decoding convolutional neural network to obtain the reconstructed first target image; wherein, both the encoding convolutional neural network and the decoding convolutional neural network are trained only based on the in vivo face sample images; Calculating the mean square loss, absolute value loss and structural similarity loss between the original feature image and the first target image, taking the weighted sum of the mean square loss, absolute value loss and structural similarity loss as the first target loss between the original feature image and the first target image, and determining the detection result of the detection based on the magnitude relationship between the first target loss and the reconstruction loss threshold, and the structural similarity loss includes the similarity loss between the original feature image and the first target image in terms of illumination, contrast and structure.

2. The face liveness detection method according to claim 1, wherein The performing at least one-dimensional feature vector encoding on the original feature image to be detected based on the trained encoding convolutional neural network to obtain at least one-dimensional first feature vector includes: Scaling the original size of the original feature image to a preset target size to obtain a scaled image; Performing convolutional layer processing, pooling layer processing and activation layer processing on the scaled image by using the encoding convolutional neural network to obtain at least one-dimensional first feature vector.

3. The face liveness detection method according to claim 1, wherein The performing image reconstruction based on the first feature vector and the trained decoding convolutional neural network to obtain the reconstructed first target image includes: Performing image reconstruction on the first feature vector by using the decoding convolutional neural network to obtain a first target image with a preset target size.

4. The face liveness detection method according to claim 1, wherein The determining the first target loss between the original feature image and the first target image and determining the detection result of the detection based on the first target loss includes: Calculating the reconstruction loss between the original feature image and the first target image; When the reconstruction loss is less than the reconstruction loss threshold, determining that the detection result of the detection is a face in vivo; When the reconstruction loss is not less than the reconstruction loss threshold, determining that the detection result of the detection is not a face in vivo.

5. The face live detection method according to claim 2 or 3, wherein The encoding convolutional neural network and the decoding convolutional neural network are trained in the following manner: Training step: Performing encoding on the sample image with annotation by using the encoding convolutional neural network to obtain at least one-dimensional second feature vector; Performing image reconstruction based on the second feature vector by using the decoding convolutional neural network to obtain the reconstructed second target image; Determining the second target loss between the original feature image and the second target image; When the second target loss is not less than the target loss threshold, update the encoding convolutional neural network and the decoding convolutional neural network by using a preset stochastic gradient descent algorithm to obtain an updated target encoding convolutional neural network and a target decoding convolutional neural network, and use the target encoding convolutional neural network as the current encoding convolutional neural network, and use the target decoding convolutional neural network as the current decoding convolutional neural network; Repeat the training step until the second target loss is less than the target loss threshold to obtain a trained encoding convolutional neural network and a decoding convolutional neural network.

6. A face liveness detection device, characterized in that, It includes: An encoding module, configured to encode an original feature image to be detected into at least one-dimensional first feature vectors based on the trained encoding convolutional neural network; A reconstruction module, configured to perform image reconstruction based on the first feature vectors and the trained decoding convolutional neural network to obtain a reconstructed first target image; wherein, both the encoding convolutional neural network and the decoding convolutional neural network are trained only based on in vivo face sample images; A detection module, configured to calculate the mean square loss, absolute value loss, and structural similarity loss between the original feature image and the first target image, use the weighted sum of the mean square loss, absolute value loss, and structural similarity loss as the first target loss between the original feature image and the first target image, and determine the detection result of the detection based on the magnitude relationship between the first target loss and the reconstruction loss threshold, and the structural similarity loss includes the similarity loss between the original feature image and the first target image in terms of illumination, contrast, and structure.

7. The face liveness detection device according to claim 6, wherein, The encoding module includes: A scaling sub-module, configured to scale the original size of the original feature image to a preset target size to obtain a scaled image; An encoding sub-module, configured to perform convolutional layer processing, pooling layer processing, and activation layer processing on the scaled image by using the encoding convolutional neural network to obtain at least one-dimensional first feature vectors.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory; One or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the face liveness detection method according to any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that, The computer storage medium is used to store computer instructions, and when the computer instructions run on a computer, the computer can execute the face liveness detection method according to any one of claims 1 to 5 above.

Citation Information

Patent Citations

  • Living body detection method and device, equipment and storage medium

    CN111753595A