Liveness detection method and apparatus, and device, medium and product
Patent Information
- Application Number
- PCT/CN2026/075358
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-24
- Filing Date
- 2026-01-28
- Publication Date
- 2026-08-27
Smart Images

Figure CN2026075358_27082026_PF_FP_ABST
Abstract
Description
Liveness detection methods, devices, equipment, media and products
[0001] This application claims priority to Chinese Patent Application No. 202510211650.9, filed on February 24, 2025, entitled "Liveness Identification Method, Apparatus, Device, Medium and Product", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, medium, and product for liveness detection. Background of the Invention
[0003] Liveness detection technology is used to verify whether biometric samples originate from a real, living person and are not obtained through forgery. Its core purpose is to ensure that the input to a biometric system comes from a genuine user, thereby improving the system's security and accuracy. For example, in facial recognition, it detects that the input facial image / video is a real user's face, not a mask, hood, etc. Classification is typically performed by extracting and analyzing Remote Photoplethysmography (rPPG) features from the facial video. This method relies on capturing skin color changes in the video to extract the pulse wave signal and then analyzes the rPPG signal using a specified algorithm (such as Fast Fourier Transform (FFT)) to determine whether it is a genuine user. Summary of the Invention
[0004] This application provides a method, apparatus, device, medium, and product for liveness detection.
[0005] A liveness detection method includes:
[0006] Acquire a first video including a first object, wherein the first object is the object to be identified as a liveness detection subject;
[0007] Analyze the texture features of the first object in the first video to obtain the difference feature representation corresponding to the first video. The difference feature representation is used to indicate the differences in the object texture of the first object between image frames of the first video.
[0008] Analyze the physiological characteristics of the first object in the first video to obtain the physiological signal feature representation corresponding to the first video. The physiological signal feature representation is used to indicate the changes in the biological pulse signal exhibited by the first object in the first video.
[0009] The first object in the first video is classified by combining the differential feature representation and the physiological signal feature representation to obtain the liveness recognition result corresponding to the first object. The liveness recognition result is used to indicate that the first object is a live object.
[0010] A liveness detection device includes:
[0011] The acquisition module is used to acquire a first video including a first object, wherein the first object is the object to be identified as a liveness detection object;
[0012] The first analysis module is used to analyze the texture features of the first object in the first video to obtain the difference feature representation corresponding to the first video. The difference feature representation is used to indicate the differences in the object texture of the first object between the image frames of the first video.
[0013] The second analysis module is used to analyze the physiological characteristics of the first object in the first video and obtain the physiological signal feature representation corresponding to the first video. The physiological signal feature representation is used to indicate the changes in the biological pulse signal exhibited by the first object in the first video.
[0014] The classification module is used to classify the first object in the first video by combining the differential feature representation and the physiological signal feature representation, and to obtain the liveness recognition result corresponding to the first object. The liveness recognition result is used to indicate whether the first object is a live object.
[0015] A computer device includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by the processor to implement the liveness detection method as described in any of the embodiments of this application.
[0016] A computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by a processor to implement the liveness detection method of any embodiment in the present application.
[0017] A computer program product or computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the liveness detection methods described in the above embodiments.
[0018] In the process of achieving liveness detection for a first object in a first video, the differences in the object's texture between image frames are analyzed to obtain a difference feature representation. The changes in the object's bio-pulse signal are analyzed to obtain a physiological signal feature representation. The liveness detection result is generated by combining the difference feature representation representing texture features and the physiological signal feature representation representing physiological features. In other words, during liveness detection, corresponding features are obtained from both the visual and physiological signal branches, and the features from multiple branches are fused for liveness detection. This overcomes the shortcomings of relying solely on the bio-pulse signal feature dimension for liveness detection, improving the accuracy of the liveness detection result and enhancing the robustness of the liveness detection process. (Brief description of the attached figures follows.)
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 is a structural block diagram of a computer system according to an embodiment of this application;
[0021] Figure 2 is a flowchart of a liveness detection method according to an embodiment of this application;
[0022] Figure 3 is a schematic diagram of the liveness detection process according to an embodiment of this application;
[0023] Figure 4 is a flowchart of the liveness detection method according to an embodiment of this application;
[0024] Figure 5 is a schematic diagram of the network structure of the feature extraction network according to an embodiment of this application;
[0025] Figure 6 is a schematic diagram of the network structure of the visual branching network according to an embodiment of this application;
[0026] Figure 7 is a schematic diagram of the network structure of the physiological signal branch network according to an embodiment of this application;
[0027] Figure 8 is a schematic diagram of a PhysNet based on 3D CNN according to an embodiment of this application;
[0028] Figure 9 is a schematic diagram of the PhysNet network architecture based on 3D CNN according to an embodiment of this application;
[0029] Figure 10 is a schematic diagram of PhysNet based on RNN according to an embodiment of this application;
[0030] Figure 11 is a schematic diagram of the network architecture of PhysNet based on RNN according to an embodiment of this application;
[0031] Figure 12 is a schematic diagram of a liveness detection model according to an embodiment of this application;
[0032] Figure 13 is a flowchart of the model training method according to an embodiment of this application;
[0033] Figure 14 is a schematic diagram of the model architecture of the first model according to an embodiment of this application;
[0034] Figure 15 is a flowchart of the facial liveness recognition method according to an embodiment of this application;
[0035] Figure 16 is a structural block diagram of a liveness detection device according to an embodiment of this application;
[0036] Figure 17 is a structural block diagram of a liveness detection device according to an embodiment of this application;
[0037] Figure 18 is a schematic diagram of the server structure according to an embodiment of this application. Methods for implementing the present invention.
[0038] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0039] In this application, the terms "first" and "second" are used to distinguish between identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first" and "second", nor is there any limitation on the quantity or execution order.
[0040] First, a brief introduction to the terms used in the embodiments of this application will be given.
[0041] Liveness detection is a biometric technology designed to verify whether the detected object is a real, living person, and not a forgery such as a photograph, video, or model. Its principle is based on analyzing vital characteristics, such as subtle changes in facial expressions, eye movements, skin texture, and blood flow. Liveness detection has a wide range of applications, including secure verification in payment processes, unlocking smartphones, security checks, online identity authentication, security access control, and identity verification for public transportation. Liveness detection effectively improves the security and reliability of biometric systems.
[0042] In one example, let's consider the application of liveness detection in a payment scenario. For instance, an object can make a payment using biometrics, such as facial recognition, fingerprint recognition, palm print recognition, or iris recognition. Compared to traditional methods like scanning QR codes, swiping cards, or entering payment passwords, biometric payments are more convenient and efficient, effectively preventing identity theft and unauthorized access. In biometric payment scenarios, after the device captures a video of the object, the device or server first performs liveness detection on the object within the video. If the object is confirmed to be alive, further payment authentication is performed, preventing payment verification through forged biometrics (such as photos, videos, masks, or fake fingerprints), thus enhancing the security of the payment process.
[0043] In another example, taking the application of liveness detection in identity verification scenarios, objects can be identified through biometrics, such as for unlocking mobile phones, payment verification, application login, account registration, and system login. Compared to traditional password and CAPTCHA verification, biometric identity verification is more secure and can effectively prevent identity theft and unauthorized intrusion. In biometric identity verification scenarios, after the device captures a video of the object, the device or server first performs liveness detection on the object in the video. If the object is confirmed to be alive, further identity verification is performed, thereby further preventing identity verification through forged biometrics (such as photos, videos, masks, or fake fingerprints), and improving the security of the identity verification process.
[0044] In another example, taking the application of liveness detection in ride-hailing scenarios, liveness detection combined with biometric verification can be used to verify the driver's identity, ensuring that only successfully matched individuals can access advanced vehicle functions. Alternatively, liveness detection combined with biometric verification can be applied to the passenger's contactless payment process, such as facial recognition or fingerprint payment. This often involves combining liveness detection and other technologies to further prevent bypassing driver identity verification and / or passenger payment verification by forging biometrics (such as photos, videos, masks, or fake fingerprints), thereby enhancing the security of ride-hailing scenarios.
[0045] In another example, let's consider the application of liveness detection in security access control. For instance, individuals unlock access points using biometric features to enter specific areas, such as facial recognition, palm print recognition, or iris recognition. Compared to traditional passwords or keys, biometric unlocking is more convenient and secure, effectively preventing identity theft and unauthorized intrusion. It is often combined with technologies like liveness detection to further prevent bypassing the system by forging biometric features (such as photos, videos, masks, or fake fingerprints), thus enhancing the security of the access control system.
[0046] rPPG (reactive photoplethysmography) is a non-contact physiological signal measurement technology that uses a camera to capture subtle color changes in facial skin to measure physiological indicators such as heart rate and respiratory rate. Its principle is to use reflected ambient light to measure subtle changes in skin brightness, thereby capturing and analyzing the periodic changes in skin color caused by blood flow due to heartbeats. Generally, rPPG technology can obtain signals similar to a heartbeat, which can then be used to predict heart rate.
[0047] Figure 1 shows a structural block diagram of a computer system provided in an exemplary embodiment of this application. The computer system 100 includes a terminal 120 and a server 140.
[0048] Terminal 120 has an application installed and running that supports biometric / liveness recognition. This application can be any type of application, such as an e-commerce application, a payment application, a video application, a social application, or a game application.
[0049] The device types of terminal 120 include at least one of the following: payment devices with biometric identification functions, POS (Point of Sale) machines, smartphones, laptops, desktop computers, tablets, smart speakers, and smart robots.
[0050] Terminal 120 includes an image acquisition device for acquiring color images and videos. For example, the image acquisition device can be at least one of a monocular camera, a binocular camera, a depth camera (RGB-D camera), an infrared camera, or a video camera. Exemplarily, terminal 120 also includes a display; the display is used to display a biometric interface / liveness recognition interface, or to display images or videos acquired by the image acquisition device, or to display liveness recognition results.
[0051] Terminal 120 is connected to server 140 via a wireless network or a wired network.
[0052] Those skilled in the art will understand that the number of the aforementioned devices can be more or less. For example, there may be only one device, or there may be dozens or hundreds of devices, or even more. This application does not limit the number or type of devices.
[0053] Server 140 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 140 provides functional services for implementing liveness detection. For example, server 140 undertakes the primary computing task, and terminal 120 undertakes the secondary computing task; or, server 140 undertakes the secondary computing task, and terminal 120 undertakes the primary computing task; or, server 140 and terminal 120 collaborate on computing using a distributed computing architecture.
[0054] In other examples, the aforementioned server 140 can be implemented as a physical server or as a cloud server in the cloud. The aforementioned server 140 can also be implemented as a node in a blockchain system.
[0055] For example, terminal 120 acquires a first video 122 via image acquisition device 121. The first video 122 includes a first object to be identified as a liveness detector. Terminal 120 sends the first video 122 to server 140, where a liveness detection service performs a liveness detection process on the first video 122. This service may include a visual branch 123, a physiological signal branch 124, and a classifier 125. During the liveness detection process on the first video 122, visual branch 123 analyzes the texture features of the first object in the first video 122 to obtain a difference feature representation corresponding to the first video 122. This difference feature representation indicates the differences in the texture of the first object between image frames of the first video 122. Physiological signal branch 124 analyzes the physiological features of the first object in the first video 122 to obtain a physiological signal feature representation corresponding to the first video 122. This physiological signal feature representation indicates the changes in the biological pulse signal exhibited by the first object in the first video 122. The classifier 125 combines differential feature representation and physiological signal feature representation to classify the first object in the first video 122 and obtains the liveness recognition result 126 corresponding to the first object. The liveness recognition result 126 is used to indicate that the first object is a live object.
[0056] For example, server 140 feeds back the obtained liveness detection result 126 to terminal 120, and terminal 120 displays the liveness detection result 126 to the user; or, server 140 applies the liveness detection result 126 to downstream tasks, such as providing it to modules or devices that perform downstream tasks. For example, in a payment scenario, server 140 determines whether the first object is a real living person based on the liveness detection result 126. If the first object is determined to be a real living person, biometric authentication is performed on the first object through the first video 122. When the biometric result of the first object matches the account information of the currently logged-in account of terminal 120, a payment operation is performed (through the payment module or payment server).
[0057] In other examples, the liveness detection method provided in this application embodiment can also be executed by either the terminal 120 or the server 140 independently, without limitation. In one example, the terminal 120 includes a functional device for implementing liveness detection, and the terminal 120 directly inputs the acquired first video to the aforementioned functional device to obtain a liveness detection result; in another example, the server 140 retrieves the first video from a storage module (e.g., a database) and inputs the first video to the functional service for implementing liveness detection to obtain a liveness detection result.
[0058] Based on the above-described terminology and application scenarios, the liveness detection method provided in this application will be described. This method is executed by a computer device (including a terminal and / or a server). Figure 2 shows a flowchart of the liveness detection method according to an embodiment of this application, which includes steps 210 to 240.
[0059] Step 210: Obtain the first video including the first object.
[0060] The first video includes an image of a first object. The first object is the object to be identified as a liveness detection subject. In this example, the first object is used as the target of liveness detection to determine whether the first object is a real live person.
[0061] The first type of object is an object with biometric features for biometric identification, such as a facial object for facial recognition, a palm object for palmprint recognition, a finger object for fingerprint recognition, an eye object for iris recognition, etc., without any specific limitation.
[0062] The first video is video data obtained by capturing a first object using an image acquisition device. For example, the image acquisition device includes at least one of a monocular camera, a binocular camera, a depth camera, an infrared camera, and a video camera.
[0063] Step 220: Analyze the texture features of the first object in the first video to obtain the difference feature representation corresponding to the first video.
[0064] In this embodiment of the application, the process of performing liveness detection on the first video includes a visual branch and a physiological signal branch. The visual branch is used to identify texture features in the first video, and the physiological signal branch is used to identify physiological features in the first video. Step 220 is an execution step in the visual branch.
[0065] For example, the difference feature represents the differences in the object texture of the first object between image frames of the first video. That is, the difference feature represents information that characterizes the changes in the object texture of the first object over time, reflecting the visual information conveyed by the first video.
[0066] Texture features are used to describe the regularity of the local spatial arrangement pattern of pixel values in the image frame of the first video, reflecting the texture, structure, and details of the object surface of the first object in the image frame. In one example, if the first object is implemented as a face object, then the texture feature of the face object is the skin texture feature present on the face; in another example, if the first object is implemented as a hand object, then the texture feature of the hand object is the skin texture feature present on the hand.
[0067] The feature extraction process of the differential feature representation is achieved through a pre-trained visual branch network. That is, the visual branch network can extract the texture features of the first object in the video, thereby obtaining the differential feature representation.
[0068] For example, the aforementioned visual branch network can be implemented using neural networks such as Convolutional Neural Networks (CNN), Feedforward Neural Networks (FNN), Residual Networks (ResNet), and Transformers, without specific limitations here.
[0069] For example, the input to the pre-trained visual branch network is a first video, such as [example video]. The first video is input into the pre-trained visual branch network to obtain the differential feature representation corresponding to the first video.
[0070] For example, the input to the pre-trained visual branch network is the pre-extracted shallow features of a first video. Feature extraction is performed on the first video to obtain the first shallow features of the first video. The first shallow features are then input into the pre-trained visual branch network to obtain the differential feature representation corresponding to the first video.
[0071] Shallow features refer to the features obtained by extracting features from the first video through one or more convolutional layers. They are a preliminary abstraction and representation of the input first video and can reflect the basic attributes of the first video, such as information on color, texture, edges, and corners in the video frames, as well as the temporal information between video frames.
[0072] For example, the convolutional layer used to extract shallow features includes a convolution module and a pooling module. The convolution module is used to implement convolution operations, and the pooling module is used to implement pooling operations.
[0073] Before analyzing the texture features of the first object in the first video, preprocessing operations can be performed on the first video. For example, preprocessing operations include at least one of video format conversion, video cropping, frame splitting, noise reduction, and target annotation.
[0074] The video format conversion instruction converts the data format of the first video into a data format that can be input into downstream networks (e.g., visual branch networks or convolutional layers for extracting shallow features). For example, the aforementioned data formats include MP4, Audio Video Interleave (AVI), Windows Media Video (WMV), and Flash Video (FLV).
[0075] Video cropping removes irrelevant edge portions from the first video, allowing the downstream network to focus on the Region of Interest (ROI). Taking facial recognition as an example, when the terminal records the first video for user facial recognition, it can display a designated area (e.g., a circular area) so that the user-controlled image acquisition device can keep the captured face within that area. Since the image acquisition device actually captures a rectangular area related to the lens magnification, it contains more image information (e.g., the user's torso, surrounding environment, etc.). Therefore, cropping the first video based on the designated area reduces the amount of input video data and lowers the data processing load on the downstream network.
[0076] Frame segmentation can decompose a first video into multiple image frames. For example, the first video can be extracted frame by frame based on its frame rate to obtain multiple image frames. Alternatively, to reduce data processing volume, keyframe recognition can be performed on the first video to identify multiple keyframes, which are then extracted for downstream recognition. Another example is using a preset frame extraction interval to extract image frames from the first video to obtain multiple image frames for downstream recognition, further reducing data processing volume.
[0077] Noise reduction processing can improve the video quality of a first video by removing noise from its image frames. For example, noise reduction processing of the first video can be implemented as spatial domain noise reduction (processing directly on the video frames, using the statistical characteristics of local regions to estimate and remove noise), temporal threshold noise reduction (using the similarity between multiple video frames to smooth noise), frequency domain noise reduction (converting the video signal to the frequency domain to remove high-frequency noise components), and deep learning noise reduction (using a pre-trained noise reduction model to analyze noise patterns in video frames and intelligently remove noise).
[0078] Target annotation allows for the labeling of the first object in the first video that requires liveness detection, enabling the downstream network to focus more attention on the target region corresponding to the first object. For example, based on the liveness detection type corresponding to the first video, the object type to be detected is determined, and the first object is identified from the first video based on the object type. The location region of the first object in the first video is then labeled, and the labeled information, along with the first video, is input into the downstream network. For instance, when the first video requires facial liveness detection, the location region corresponding to the facial object in the first video is identified and labeled; similarly, when the first video requires hand liveness detection, the location region corresponding to the hand object in the first video is identified and labeled.
[0079] Step 230: Analyze the physiological characteristics of the first object in the first video to obtain the physiological signal feature representation corresponding to the first video.
[0080] The process of performing liveness detection on the first video includes a visual branch and a physiological signal branch, wherein step 230 is a step in the physiological signal branch.
[0081] The physiological signal feature representation is used to indicate the changes in the biological pulse signal exhibited by the first object in the first video. In other words, the physiological signal feature representation characterizes the information on the changes in the physiological characteristics of the first object over time.
[0082] Physiological characteristics refer to the quantifiable and observable properties of an organism in terms of its physiological functions and structure, which reflect the organism's physiological functions. For example, the aforementioned physiological characteristics include heart rate, respiratory rate, and blood flow rate.
[0083] A biological pulse signal is a physiological signal that indicates the physiological characteristics of an organism. It reflects the activity of the heart and the state of the vascular system, and is used to characterize the signal generated by changes in blood vessel volume caused by heartbeats. In this embodiment, the biological pulse signal can be obtained through first video detection.
[0084] The aforementioned biological pulse signal can be realized as an rPPG signal, that is, the physiological signal characteristics are represented by the changes in the rPPG signal exhibited by the first object in the first video.
[0085] For example, the generation of physiological signal feature representations can be achieved by extracting the physiological signal feature representations corresponding to the rPPG signal from the first video using statistical methods. For instance, the aforementioned statistical methods can be implemented as ambient light-based rPPG signal extraction methods (Green's algorithm), Independent Component Analysis (ICA) algorithms, and POS (Plane Orthogonal-to-Skin) algorithms, etc.
[0086] For example, the generation of physiological signal feature representations can be achieved through artificial intelligence (AI) models. For instance, by inputting a first video into a pre-trained physiological signal branch network, the physiological signal branch network can extract the physiological features of the first object in the first video, thereby obtaining a physiological signal feature representation.
[0087] For example, the aforementioned physiological signal branch network can be implemented using neural networks such as convolutional neural networks, feedforward neural networks, residual networks, and converters, without specific limitations here.
[0088] For example, the input to the pre-trained physiological signal branch network is a first video. For instance, the first video is input into the pre-trained physiological signal branch network to obtain the physiological signal feature representation corresponding to the first video.
[0089] For example, the input to the pre-trained physiological signal branch network is the shallow features of the first video that have been extracted in advance. For instance, feature extraction is performed on the first video to obtain the second shallow features of the first video. The second shallow features are then input into the pre-trained physiological signal branch network to obtain the physiological signal feature representation corresponding to the first video.
[0090] For example, the first shallow feature used for texture feature extraction and the second shallow feature used for physiological feature extraction can be implemented as the same shallow feature representation or as different shallow feature representations, without limitation here.
[0091] Before analyzing the physiological characteristics of the first object in the first video, preprocessing operations are performed on the first video. For example, preprocessing operations include at least one of video format conversion, video cropping, frame segmentation, noise reduction, and target labeling.
[0092] Step 240: Combine differential feature representation and physiological signal feature representation to classify the first object in the first video and obtain the liveness recognition result corresponding to the first object.
[0093] The aforementioned liveness detection result is used to indicate whether the first object is a living person; that is, the liveness detection result can indicate whether the first object in the first video is a real living person.
[0094] In the process of performing liveness detection on the first video, after the visual branch and the physiological signal branch, a classification step (i.e., step 240) may be included. This step is used to classify the first object according to the differential feature representation output by the visual branch and the physiological signal feature representation output by the physiological signal branch, so as to obtain the liveness detection result corresponding to the first object.
[0095] For example, the above-mentioned liveness detection results include indicating that the first object is a live object, or indicating that the first object is a fraudulent object, wherein the fraudulent object indicates that the first object is not a real live object.
[0096] When the liveness detection result indicates that the first object belongs to a fraud target, the liveness detection result may also include the type of fraud to which the first object belongs, such as model fraud, mask fraud, headgear fraud, image fraud, etc.
[0097] The process of generating liveness detection results is achieved through a pre-trained classifier. That is, the classifier can combine differential feature representation and physiological signal feature representation to classify the first object in the first video, thereby obtaining the liveness detection result corresponding to the first object.
[0098] For example, the input to a pre-trained classifier is a differential feature representation and a physiological signal feature representation. Alternatively, both the differential feature representation and the physiological signal feature representation can be input into the pre-trained classifier to obtain a liveness detection result.
[0099] For example, the input to a pre-trained classifier is a feature representation obtained by fusing differential feature representation and physiological signal feature representation. For instance, differential feature representation and physiological signal feature representation can be fused to obtain a fused feature representation, which can then be input into a pre-trained classifier to obtain a liveness detection result.
[0100] For example, the fusion of differential feature representation and physiological signal feature representation can be achieved through one of the following methods:
[0101] • Feature splicing: By splicing the differential feature representation and the physiological signal feature representation, a fused feature representation is obtained;
[0102] • Feature addition: The differential feature representation and the physiological signal feature representation are added element by element to obtain the fused feature representation;
[0103] • Feature averaging: The element-wise mean of the differential feature representation and the physiological signal feature representation is calculated to obtain the fused feature representation;
[0104] • Feature stacking: Stacking differential feature representations and physiological signal feature representations to obtain fused feature representations;
[0105] • Based on deep learning fusion, differential feature representations and physiological signal feature representations are input into a pre-trained convolutional neural network (or convolutional layer) for feature fusion to obtain fused feature representations.
[0106] For example, the classifier described above can be implemented as at least one of the following:
[0107] • Classifiers based on linear models, such as linear regression classifiers, logistic regression classifiers, support vector machines (SVM), etc.
[0108] • Decision tree-based classifiers, such as Decision Tree Classifier, Random Forest Classifier, Gradient Boosting Tree (GBT) Classifier, etc.
[0109] • Neural network-based classifiers, such as multilayer perceptrons (MLP), convolutional neural networks, recurrent neural networks, etc.
[0110] • Nearest neighbor-based classifiers, such as Nearest Neighbor Classifier, K-Nearest Neighbors (KNN), etc.
[0111] • Reinforcement learning-based classifiers, such as Q-learning-based classifiers and deep Q-networks.
[0112] For example, Figure 3 shows a schematic diagram of the liveness detection process according to an embodiment of this application. The liveness detection process 300 includes the following steps: 301, capturing video; 302, video preprocessing; 303, liveness detection; 304, outputting the liveness detection result.
[0113] The obtained liveness detection results are then applied to downstream tasks. For example, these downstream tasks can be implemented as identity verification, target recognition, and video screening.
[0114] In one example, taking downstream task-based identity verification as an example, the decision to execute the identity verification process is determined based on the liveness detection result. For instance, in the payment process, identity verification is achieved through facial recognition. The terminal uploads a first video of the user, and the server performs liveness detection on the first video to obtain the result. When the liveness detection result indicates that the facial object appearing in the first video is a live object, the facial object in the first video is matched with the account verification information of the payment account to complete the identity verification. When the liveness detection result indicates that the facial object appearing in the first video is a fraudulent object, the identity verification is interrupted, and a verification failure message is sent back to the terminal.
[0115] In another example, taking target recognition as the downstream task, target recognition is performed on the first video based on the liveness detection results. That is, the liveness detection process is a preprocessing operation performed before the target recognition task is performed on the first video. For example, when the first video includes multiple first objects, liveness detection results are obtained for each of the multiple first objects in the first video. Based on the liveness detection results corresponding to each of the multiple first objects, the live objects among the multiple objects are determined. Target recognition is then performed on the determined live objects, such as identifying the location of the live objects in each image frame of the first video and identifying the object information of the live objects.
[0116] In the process of achieving liveness detection for a first object in a first video, differential feature representations are obtained by analyzing the differences in the object's texture between image frames; physiological signal feature representations are obtained by analyzing the changes in the object's bio-pulse signal in the first video. Liveness detection results are generated by combining the differential feature representations characterizing texture features and the physiological signal feature representations characterizing physiological features. That is, during liveness detection, corresponding features are obtained from both the visual and physiological signal branches, and the features from multiple branches are fused for liveness detection. This overcomes the shortcomings of relying solely on the single feature dimension of bio-pulse signals for liveness detection, improving the accuracy of the liveness detection results and enhancing the robustness of the liveness detection process.
[0117] The liveness detection method provided in this application does not increase the user's interaction cost during implementation. That is, the user only needs to record and upload a video, which reduces the user's operational needs in the verification process and maintains a good user experience.
[0118] Meanwhile, improving the accuracy of liveness detection results can improve the processing performance of downstream tasks in liveness detection. For example, in the case of identity verification, it can improve the accuracy of the final identity verification result.
[0119] When the feature representations output from the visual branch and the physiological signal branch need to be fused during the classification stage, having consistent data representations for the input visual branch and the physiological signal branch can improve the fusion effect of feature fusion during the classification stage. This avoids potential problems caused by different feature spaces in the output features of different branches and can also promote the overall integration of the liveness detection process for the first video. By extracting shallow features from the first video and inputting these shallow features into the visual branch and the physiological signal branch respectively, the feature representation for the classification stage can be obtained.
[0120] Figure 4 shows a flowchart of a liveness detection method according to an embodiment of this application, which includes steps 410 to 450.
[0121] Step 410: Obtain the first video including the first object.
[0122] The first object is the object to be identified as a liveness detection object. That is, the first object is used as the target of liveness detection to determine whether the first object is a real live body.
[0123] The first object is a biometric object used for biometric identification, such as a facial object for facial recognition, a palm object for palmprint recognition, a finger object for fingerprint recognition, an eye object for iris recognition, etc., without any specific limitation.
[0124] Step 420: Extract the visual feature representation of the first video.
[0125] For example, the aforementioned visual features represent visual information used to indicate the first video, and the visual features represent inputs used as visual branches and physiological signal branches.
[0126] Among them, visual feature representation is the shallow feature of the first video. Shallow features refer to the features obtained by extracting features from the first video through one or more convolutional layers. They are a preliminary abstraction and representation of the input first video and can reflect the basic attributes of the first video, such as information such as color, texture, edges, and corners in the video frames, as well as the temporal information between video frames.
[0127] For example, the convolutional layer used to extract shallow features includes a convolution module and a pooling module. The convolution module is used to implement convolution operations, and the pooling module is used to implement pooling operations.
[0128] Visual feature representations are extracted using a pre-trained feature extraction network. For example, a first video is input into the pre-trained feature extraction network to obtain visual feature representations.
[0129] For example, the feature extraction network described above can be implemented by at least one type of neural network, such as a convolutional neural network, a feedforward neural network, a residual network, or a converter, without any specific limitations.
[0130] The pre-trained feature extraction network described above includes at least one convolutional layer. Figure 5 shows a schematic diagram of the network structure of the feature extraction network 500 according to an embodiment of this application. The feature extraction network 500 includes multiple convolutional layers, including convolutional layer 501, convolutional layer 502, convolutional layer 503, convolutional layer 504, and convolutional layer 505, and multiple max pooling layers, including max pooling layer 506 and max pooling layer 507.
[0131] The convolutional layer 501 has a kernel size of 3*3 and includes 16 kernels of size 3*3. The convolutional layer 501 is used to receive the input first video and extract the feature map of the first video through convolution operation, and input it into the max pooling layer 506.
[0132] The max pooling layer 506 is used to reduce the dimensionality of the feature map output by the convolutional layer 501 while retaining important feature information. For example, the max pooling layer 506 achieves dimensionality reduction by sliding a fixed-size window (e.g., a 2x2 window) across the feature map and selecting the maximum value within the window as the output. The max pooling layer 506 then inputs the dimensionality-reduced feature map into the convolutional layer 502.
[0133] The convolutional layer 502 has a kernel size of 3*3 and includes 32 convolutional kernels of size 3*3. The convolutional layer 502 is used to receive the feature map output by the max pooling layer 506, and extract the features of the above feature map through convolution operation to obtain the feature map after further convolution, and input the obtained feature map into the convolutional layer 503.
[0134] The convolutional layer 503 has a kernel size of 3*3 and includes 64 kernels of size 3*3. The convolutional layer 503 is used to receive the feature map output by the convolutional layer 502, and extract the features of the above feature map through the convolution operation to obtain the feature map after further convolution. The obtained feature map is then input into the max pooling layer 507.
[0135] Max pooling layer 507 is used to reduce the dimensionality of the feature map output by convolutional layer 503 while retaining important feature information. Max pooling layer 507 then inputs the dimensionality-reduced feature map into convolutional layer 504.
[0136] The convolutional layer 504 has a kernel size of 3*3 and includes 64 kernels of size 3*3. The convolutional layer 504 is used to receive the feature map output by the max pooling layer 507, and extract the features of the above feature map through convolution operation to obtain the feature map after further convolution, and input the obtained feature map into the convolutional layer 505.
[0137] The convolutional layer 505 has a kernel size of 3*3 and includes 64 kernels of size 3*3. The convolutional layer 505 is used to receive the feature map output by the convolutional layer 504 and extract the features of the above feature map through convolution operation to obtain a further convolutional feature map, which serves as the visual feature representation output by the feature extraction network 500.
[0138] The Feature Extraction Network 500 extracts and compresses features by gradually increasing the number of convolutional kernels and using pooling layers, providing rich feature representations for downstream visual and physiological signal branches.
[0139] In the process of extracting shallow features from the first video, the first video can be divided into frames. Features are extracted from the obtained image frames, and the extracted features are fused together to form the shallow features of the first video. For example, the first video is divided into frames to obtain multiple image frames; image features are extracted from each of the multiple image frames to obtain image feature representations corresponding to each of the multiple image frames; and the image feature representations corresponding to the multiple image frames are fused together to obtain the visual feature representation.
[0140] For example, the fusion of multiple image feature representations can be achieved by feature concatenation, feature addition, feature averaging, feature stacking, or deep learning-based fusion, without any limitation.
[0141] In one example, feature concatenation is used to fuse multiple image feature representations. For instance, multiple image feature representations are concatenated according to the temporal relationship between their corresponding image frames to obtain a visual feature representation. That is, the concatenated visual feature representation carries the temporal relationship between image frames.
[0142] Before extracting the visual feature representation of the first video, preprocessing operations can be performed on the first video. For example, preprocessing operations include at least one of video format conversion, video cropping, frame segmentation, noise reduction, and target labeling.
[0143] Step 430: Extract the texture features of the first object in the first video based on the visual feature representation to obtain the differential feature representation.
[0144] For example, the difference feature represents the differences in the object texture of the first object between image frames of the first video. That is, the difference feature represents information that characterizes the changes in the object texture of the first object over time, reflecting the visual information conveyed by the first video.
[0145] Texture features are used to describe the regularity of the local spatial arrangement pattern of pixel values in the image frame of the first video, reflecting the texture, structure, and details of the object surface of the first object in the image frame. In one example, if the first object is implemented as a face object, then the texture feature of the face object is the skin texture feature present on the face; in another example, if the first object is implemented as a hand object, then the texture feature of the hand object is the skin texture feature present on the hand.
[0146] Differential feature representations are features extracted through a pre-trained visual branching network. For example, visual feature representations are input into a pre-trained visual branching network to obtain differential feature representations.
[0147] The feature processing of visual feature representations by the visual branch network may include: performing central difference convolution (CDC) on the visual feature representations to extract information on the temporal changes of the texture features of the first object in the first video, and obtaining the differential feature representation.
[0148] Among them, central difference convolution is a convolution operation used to extract local differences between image frames. By combining the central difference idea with traditional convolution methods to extract intensity and gradient information of image frames, the visual branch network's ability to represent fine-grained features is enhanced.
[0149] In liveness detection, the texture details of the first object change over time (e.g., light reflection), and these changes can be used to distinguish between live and fake objects. Central difference convolution exhibits strong illumination invariance in liveness detection and can incorporate finer-grained anti-spoofing cues, such as grid effects and screen reflections. Therefore, implementing a visual branch network based on central difference convolution can improve feature extraction from differential feature representations in liveness detection.
[0150] In one example, a CDC module is set up in the visual branch network to perform central difference convolutions on the features, where the CDC module is defined as shown in Equation 1. Equation 1: O=(I*K+b)-θ·(I*K) diff +b)
[0151] Where I is the feature map of the input CDC module, K is the convolution kernel, b is the bias term in the convolution operation, θ is the weight coefficient, and K diff is the differential convolution kernel, and O is the feature map output by the CDC module.
[0152] For example, a visual branch network may include one or more CDC modules.
[0153] To further optimize the temporal and spatial features extracted from the first video in the visual branch, a temporal-spatial attention module based on the temporal-spatial attention mechanism is added after the CDC module. The CDC module and the temporal-spatial attention module together form the TDC (Temporal-Spatial Attention with CDC) module.
[0154] Among them, the spatiotemporal attention mechanism is an attention mechanism that combines the time dimension and the image spatial dimension to realize feature processing. By applying the attention mechanism to different time points and spatial locations (pixel positions in the image), the degree of attention to different parts of the first video is dynamically controlled, thereby improving the accuracy and efficiency of semantic understanding of the first video.
[0155] In one example, the spatiotemporal attention module is defined as shown in Formula 2.
[0156] Formula 2: A = σ(Conv(concat(Acg)) c (X),Max c (X),Acg t (X),Max t (X))))
[0157] Where A is the attention map output by the spatiotemporal attention module, X is the feature map input to the spatiotemporal attention module, and Acg c() represents the average pooling operation on the channel dimension (spatial dimension), Max c () represents the max pooling operation in the channel dimension (spatial dimension), Acg t () represents the average pooling operation over time, Max t () represents the max pooling operation in the time dimension, Conv() represents the 3D convolution operation, concat() represents the concatenation operation along the channel dimension, and σ() is the sigmoid activation function.
[0158] For example, when fusing central difference convolution and spatiotemporal attention mechanisms to extract texture features, the feature processing of the differential feature representation by the visual branch network includes: performing central difference convolution on the visual feature representation to obtain a first feature representation. This first feature representation includes sub-features corresponding to multiple time steps of the first video, i.e., the first feature representation is the output of the CDC module; based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation, a differential feature representation is obtained through the spatiotemporal attention mechanism, i.e., the differential feature representation is the output obtained after inputting the first feature representation into the spatiotemporal attention module. A time step refers to a unit that divides the first video into frames in chronological order. For example, when the first video is divided into frames, the time node corresponding to each frame is a time step. Depending on the actual processing needs or settings, the number of video frames can be set to any number. The smaller the time step, the more refined the extracted temporal information of the video, and correspondingly, more computation is required; conversely, the larger the time step, the coarser the extracted temporal information of the video, and correspondingly, less computation is required.
[0159] For example, the correlation between sub-features extracted at different time steps by the spatiotemporal attention mechanism includes spatial correlation and temporal correlation.
[0160] Spatial correlation is used to indicate the relationship between different image regions within the image frame corresponding to each time step. For example, spatial correlation is measured by the vector distance (e.g., Hammanton distance or Euclidean distance) between the embedding vectors of each image patch in the sub-feature; or, spatial correlation is measured by the Pearson correlation coefficient between the embedding vectors of each pixel block in the sub-feature.
[0161] Temporal correlation is used to indicate the temporal relationship between image frames corresponding to different time steps, that is, the dynamic changes between image frames corresponding to different time steps, which can characterize the motion information and temporal evolution in the video. For example, temporal correlation is measured by the vector distance (e.g., Mahalanobis distance) between sub-features corresponding to different time steps; or, temporal correlation is measured by the Pearson correlation coefficient between sub-features corresponding to different time steps; or, temporal correlation is measured by the attention weights between sub-features corresponding to different time steps.
[0162] For example, a visual branch network may include one or more TDC modules.
[0163] After setting up a TDC module in the visual branch network, a self-attention module and a dimensionality reduction convolution module are added at the end of the network to enhance the representation capability of the features output by the visual branch network. For example, the feature processing process implemented by the above visual branch network includes: performing central difference convolution on the visual feature representation to obtain a first feature representation; obtaining an attention feature representation based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation through a spatiotemporal attention mechanism; performing dimensionality reduction convolution on the attention feature representation to obtain a second feature representation; performing self-attention analysis on the attention feature representation to obtain a third feature representation; and fusing the second and third feature representations to obtain a differential feature representation.
[0164] That is, the visual feature representation is input into the TDC module to obtain the attention feature representation, which includes obtaining the first feature representation by inputting the visual feature representation into the CDC module, and inputting the first feature representation into the spatiotemporal attention module to obtain the attention feature representation; inputting the attention feature representation into the dimensionality reduction convolution module to obtain the second feature representation; inputting the attention feature representation into the self-attention module to obtain the third feature representation; and fusing the second feature representation and the third feature representation to obtain the differential feature representation.
[0165] The aforementioned self-attention module can also be implemented as a multi-head attention module. The multi-head attention module is a self-attention module implemented based on a multi-head attention mechanism, used to obtain the attention feature representation of the TDC module output in different subspaces through multiple independent self-attention mechanisms running in parallel.
[0166] The aforementioned dimensionality reduction convolution module is implemented using a 1*1 convolution kernel. Because a 1*1 convolution kernel does not cover the height and width dimensions, but only operates on the channel dimension at each pixel position, it can adjust the number of channels or integrate channel information, thereby achieving feature dimensionality reduction.
[0167] Figure 6 shows a schematic diagram of the network structure of the visual branch network 600 according to an embodiment of this application. The visual branch network 600 includes multiple TDC modules (including TDC module 601, TDC module 602, and TDC module 603), a convolutional layer 604, an adaptive pooling layer 605, a self-attention module 606, a fully connected layer (FC-T) 607, and a dimension reduction convolutional module 608.
[0168] Among them, TDC modules 601, 602, and 603 are stacks of CDC modules and spatiotemporal attention modules. The input of TDC module 601 is the visual feature representation output by the feature extraction network. TDC module 601 inputs the output feature map to TDC module 602, TDC module 602 inputs the output feature map to TDC module 603, and finally TDC module 603 outputs the attention map and inputs the attention map to the convolutional layer 604.
[0169] Convolutional layer 604 is a convolutional layer with 64 3*3 kernels. Convolutional layer 604 performs further feature extraction on the attention map and inputs the output to adaptive pooling layer 605.
[0170] The adaptive pooling layer 605 is used to adjust the size of the feature map output by the convolutional layer 604 to a fixed size, output the adjusted feature map, and simultaneously input the adjusted feature map into the self-attention module 606 and the dimensionality reduction convolutional module 608. For example, the pooling operation performed by the adaptive pooling layer 605 can be max pooling, average pooling, or other forms of pooling operation, which are not limited here.
[0171] The dimension reduction convolution module 608 is a convolutional layer with a 1*1 kernel. The dimension reduction convolution module 608 reduces the dimension of the adjusted feature map output by the adaptive pooling layer 605 to obtain the second feature representation.
[0172] The self-attention module 606 extracts features from the adjusted feature map output by the adaptive pooling layer 605 based on a self-attention mechanism or a multi-head attention mechanism to obtain a self-attention map. When the self-attention module 606 is implemented as a multi-head attention mechanism, after the self-attention layer of each head outputs a self-attention map, feature transformation is performed through the fully connected layer 607 to obtain a third feature representation.
[0173] Finally, the second and third feature representations are fused to obtain the differential feature representations output by the visual branch network 600.
[0174] The fusion of the second and third feature representations can be achieved through at least one of the following methods: feature concatenation, feature addition, feature averaging, feature stacking, and deep learning-based fusion, without any limitation.
[0175] Step 440: Generate physiological signal feature representation of the first object in the first video based on visual feature representation.
[0176] The physiological signal feature representation is used to indicate the changes in the biological pulse signal exhibited by the first object in the first video. In other words, the physiological signal feature representation characterizes the information on the changes in the physiological characteristics of the first object over time.
[0177] Physiological characteristics refer to the quantifiable and observable properties of an organism in terms of its physiological functions and structure, which reflect the organism's physiological functions. For example, the aforementioned physiological characteristics include heart rate, respiratory rate, and blood flow rate.
[0178] A biological pulse signal is a physiological signal that indicates the physiological characteristics of an organism. It reflects the activity of the heart and the state of the vascular system, and is used to characterize the signal generated by changes in blood vessel volume caused by heartbeats. In this embodiment, the biological pulse signal can be obtained through first video detection.
[0179] The aforementioned biological pulse signal can be realized as an rPPG signal, that is, the physiological signal characteristics are represented by the changes in the rPPG signal exhibited by the first object in the first video.
[0180] Physiological signal feature representations are generated through a pre-trained physiological signal branch network. For example, visual feature representations are input into the physiological signal branch network to obtain physiological signal feature representations. For instance, the physiological signal branch network is used to reconstruct accurate physiological signal feature representations, such as rPPG signals, from the visual feature representations corresponding to a first video.
[0181] Figure 7 shows a schematic diagram of the network structure of the physiological signal branch network 700 according to an embodiment of this application. The physiological signal branch network 700 includes a max pooling layer 701, a convolutional layer 702, a convolutional layer 703, a transposed convolutional layer 704, a transposed convolutional layer 705, a convolutional layer 706, an adaptive pooling layer 707, and a convolutional layer 708.
[0182] The max-pooling layer 701 reduces the dimensionality of the input visual feature representation while preserving important feature information. For example, the max-pooling layer 701 achieves dimensionality reduction by sliding a fixed-size window (e.g., a 2x2 window) across the feature map and selecting the maximum value within the window as the output. The max-pooling layer 701 then inputs the dimensionality-reduced feature map into the convolutional layer 702.
[0183] Convolutional layers 702 and 703 are convolutional layers with 64 3x3 kernels, used to further extract features from the reduced feature map and input the output to the transposed convolutional layer 704.
[0184] Transposed convolutional layer 704 and transposed convolutional layer 705 are used to upsample the input feature map to increase its spatial dimension. Transposed convolutional layer 705 inputs the upsampled feature map into convolutional layer 706.
[0185] Convolutional layer 706 is a convolutional layer with 64 3*3 kernels, used to further extract features from the upsampled feature map and input the output to adaptive pooling layer 707.
[0186] The adaptive pooling layer 707 is used to adjust the size of the feature map output by the convolutional layer 706 to a fixed size, output the adjusted feature map, and simultaneously input the adjusted feature map into the convolutional layer 708.
[0187] Convolutional layer 708 is a convolutional layer with a 1x1 kernel, used to project the adjusted feature map onto the signal space to obtain a representation of physiological signal features. The signal space is a digital space with at least one dimension set up for physiological signals, used to convert physiological dynamic information in video footage (such as subtle skin changes caused by heartbeat) into quantifiable and analyzable digital signals.
[0188] The physiological signal branch network can be implemented as PhysNet (Remote Photoplethysmograph Signal Measurement from Facial Videos Using Spatio-Temporal Network). The feature processing steps implemented by the physiological signal branch network for visual feature representation include: performing convolution and pooling operations on the visual feature representation to obtain a multi-channel feature representation, which includes the physiological features of the first object in both spatial and temporal dimensions; and projecting the multi-channel feature representation onto the signal space to obtain the physiological signal feature representation.
[0189] PhysNet is an end-to-end spatiotemporal network. The input to this network architecture is a T-frame image with RGB channels. After multiple convolution and pooling operations, a multi-channel manifold is formed to represent the spatiotemporal features. The latent manifold is projected onto the signal space using channel convolution operations with 1*1*1 convolution kernels to generate a predicted rPPG signal of length T.
[0190] In this embodiment, the input to PhysNet is the visual feature representation output by the feature extraction network. That is, after PhysNet performs multiple convolution and pooling operations on the visual feature representation, it projects the features onto the signal space using a 1*1*1 convolution kernel to obtain the rPPG signal corresponding to the first video.
[0191] Figure 8 shows a schematic diagram of a PhysNet based on 3D CNN according to an embodiment of this application. The PhysNet based on 3D CNN includes an input layer 801, a 3D CNN layer 802, an aggregation layer 803, and an output layer 804.
[0192] The input layer 801 consists of a series of images arranged in chronological order, such as the first to the Tth frames of a first video for which liveness detection is required. In this embodiment, the input layer 801 can also be implemented as a visual feature representation, which includes sub-feature representations corresponding to each time step.
[0193] Layer 802 of the 3D CNN is a three-dimensional convolutional neural network with three-dimensional convolutional kernels, which can capture the spatial and temporal features in the input PhysNet data, thereby obtaining multi-channel feature representations.
[0194] Aggregation layer 803 is used to integrate the multi-channel feature representations to generate a comprehensive feature representation that reflects the rPPG signal of the first video, i.e., the rPPG signal indicated by output layer 804. In one example, aggregation layer 803 uses a 1*1*1 convolution kernel to project the features onto the signal space to obtain the rPPG signal corresponding to the first video.
[0195] Figure 9 shows a schematic diagram of a PhysNet network architecture based on 3D CNN according to an embodiment of this application. The network architecture includes a 1*5*5 convolutional layer 901, a 1*2*2 max pooling layer 902, a 3*3*3 convolutional layer 903, a spatial average pooling layer 904, and a 1*1*1 convolutional layer 905.
[0196] The system consists of several layers: a 1*5*5 convolutional layer 901 with a kernel size of 1*5*5 and 32 channels; a 1*2*2 max pooling layer 902 with a pooling window size of 1*2*2, used to reduce the spatial dimension of the feature map; a 3*3*3 convolutional layer 903 consisting of two convolutional layers with a kernel size of 3*3*3 and 64 channels; the combination of the 1*2*2 max pooling layer 902 and the 3*3*3 convolutional layer 903 is repeated four times; a global average pooling layer 904 is used to globally average the features in the spatial dimension, thereby reducing the size of the feature map; and a 1*1*1 convolutional layer 905 with a kernel size of 1*1*1 and 1 channel, used to project the features into the signal space to output the rPPG signal.
[0197] Figure 10 shows a schematic diagram of an RNN-based PhysNet provided in an embodiment of this application. The RNN-based PhysNet includes an input layer 1001, a 2D CNN layer 1002, an LSTM (Long Short-Term Memory) layer 1003, an aggregation layer 1004, and an output layer 1005.
[0198] The input layer 1001 consists of a series of images arranged in chronological order, such as the first to the Tth frames of a first video for which liveness detection is required. In this embodiment, the input layer 1001 can also be implemented as a visual feature representation, which includes sub-feature representations corresponding to each time step.
[0199] The image frame (or sub-feature representation) corresponding to each time step is input into the corresponding 2D CNN layer 1002 for feature extraction. The 2D CNN layer 1002 is a convolutional neural network with two-dimensional convolutional kernels, which can capture spatial features in the input PhysNet data, such as edges and textures.
[0200] The LSTM layer 1003 can process the feature maps output by the 2D CNN layer 1002. There is a temporal order relationship between the feature maps output by each 2D CNN layer 1002. The LSTM layer 1003 can capture the temporal dynamic characteristics in the feature map sequence, such as facial color changes.
[0201] Aggregation layer 1004 is used to integrate features from the output of LSTM layer 1003 to generate a comprehensive feature representation that reflects the rPPG signal of the first video, i.e., the rPPG signal indicated by output layer 1005. In one example, aggregation layer 1004 uses a 1*1*1 convolution kernel to project the features onto the signal space to obtain the rPPG signal corresponding to the first video.
[0202] Figure 11 shows a schematic diagram of a PhysNet network architecture based on RNN provided in an embodiment of this application. This network architecture includes a 1*5*5 convolutional layer 1101, a 1*2*2 max pooling layer 1102, a 1*3*3 convolutional layer 1103, a spatial average pooling layer 1104, an LSTM layer 1105, and a 1*1*1 convolutional layer 1106.
[0203] The system consists of several layers: a 1*5*5 convolutional layer 1101 with a kernel size of 1*5*5 and 32 channels; a 1*2*2 max pooling layer 1102 with a pooling window size of 1*2*2, used to reduce the spatial dimension of the feature map; a 1*3*3 convolutional layer 1103 with two kernel sizes of 1*3*3 and 64 channels; the combination of the 1*2*2 max pooling layer 1102 and the 1*3*3 convolutional layer 1103 is repeated 4 times; a global average pooling layer 1104 is used to globally average the features in the spatial dimension, thereby reducing the size of the feature map; an LSTM layer 1105 with 64 output channels, capable of capturing the dynamic changes in the time series corresponding to the feature map; and a 1*1*1 convolutional layer 1106 with a kernel size of 1*1*1 and 1 channel, used to project the features onto the signal space, thereby outputting the rPPG signal.
[0204] Step 450: Combine differential feature representation and physiological signal feature representation to classify the first object in the first video and obtain the liveness recognition result corresponding to the first object.
[0205] The aforementioned liveness detection result is used to indicate whether the first object is a living person; that is, the liveness detection result can indicate whether the first object in the first video is a real living person.
[0206] For example, the above-mentioned liveness detection results include indicating that the first object is a live object, or indicating that the first object is a fraudulent object, wherein the fraudulent object indicates that the first object is not a real live object.
[0207] In some examples, the input to the pre-trained classifier is a combination of differential feature representations and physiological signal feature representations. For instance, by inputting both differential feature representations and physiological signal feature representations into the pre-trained classifier, a liveness detection result is obtained.
[0208] In other examples, the input to the pre-trained classifier is a feature representation obtained by fusing differential feature representation and physiological signal feature representation. For instance, fusing differential feature representation and physiological signal feature representation yields a fused feature representation, which is then input into the pre-trained classifier to obtain the liveness detection result.
[0209] For example, the classifier described above can be implemented as at least one of the following: a linear model-based classifier; a decision tree-based classifier; a neural network-based classifier; a nearest neighbor-based classifier; a reinforcement learning-based classifier, etc.
[0210] In summary, in the process of achieving liveness detection for the first object in the first video, the difference in texture of the first object between image frames is analyzed to obtain a difference feature representation. The change in the biological pulse signal of the first object in the first video is analyzed to obtain a physiological signal feature representation. The liveness detection result is generated by combining the difference feature representation representing texture features and the physiological signal feature representation representing physiological features. That is, in the process of liveness detection, corresponding features are obtained from the visual branch and the physiological signal branch respectively, and the features of multiple branches are fused to perform liveness detection. This makes up for the shortcomings of liveness detection from a single feature dimension of biological pulse signal, improves the accuracy of liveness detection results, and also improves the robustness of the liveness detection process.
[0211] In this embodiment, a feature extraction network is used to extract features from the first video to obtain a visual feature representation. The visual feature representation is used as the input to the visual branch and the physiological signal branch. The visual feature representation can ensure that the features input to the visual branch and the physiological signal branch have consistent data representation. That is, it ensures that the two branches operate on consistent data representation, thereby promoting more effective integration and reducing the potential information related to feature space mismatch between different branches. This improves the fusion effect of feature fusion in the classification stage.
[0212] Figure 12 shows a schematic diagram of a liveness detection model 1200 according to an embodiment of this application. The liveness detection model 1200 includes a feature extraction network 1210, a visual branch network 1220, a physiological signal branch network 1230, and a classifier 1240.
[0213] The liveness recognition process implemented by the liveness recognition model 1200 includes: inputting a first video into a feature extraction network 1210 to obtain a visual feature representation; inputting the visual feature representation into a visual branch network 1220 and a physiological signal branch network 1230 respectively; in the visual branch network 1220, the visual feature representation is processed into a differential feature representation; in the physiological signal branch network 1230, the visual feature representation is processed into a physiological signal feature representation; and inputting the differential feature representation and the physiological signal feature representation into a classifier 1240 to obtain a liveness recognition result.
[0214] For example, the aforementioned liveness detection model is obtained through training a first model. Figure 13 shows a flowchart of the model training method according to an embodiment of this application. The method includes steps 1310 to 1330.
[0215] Step 1310: Obtain the sample video including the sample object, and the sample tag corresponding to the sample object.
[0216] Sample labels are used to indicate whether a sample object is a real, living organism. Sample labels are manually labeled information.
[0217] The collected sample video set can include positive samples and negative samples. Positive samples are sample videos where the sample object is a live object, and negative samples are sample videos where the sample object is a fraudulent object.
[0218] Step 1320: Input the sample video into the first model to generate the prediction detection result.
[0219] The predicted detection results are used to indicate the liveness classification results of the first model for the sample objects.
[0220] Obtain the model architecture of the first model, initialize the model parameters of the first model, and use the initialized first model for the model training process.
[0221] The model architecture of the first model is the same as that of the liveness detection model; for example, the model architecture of the first model differs from that of the liveness detection model in some aspects.
[0222] When the model architecture of the first model is the same as that of the liveness detection model, the first model includes a feature extraction network, a first branch network, a second branch network, and a classifier to be trained. The first branch network is used to extract texture features and is used to train the visual branch network in the liveness detection model. The second branch network is used to extract physiological features and is used to train the physiological signal branch network in the liveness detection model.
[0223] For example, the sample video is input into a feature extraction network to obtain a sample visual feature representation; the sample visual feature representation is input into a first branch network to obtain a predicted difference feature representation; the sample visual feature representation is input into a second branch network to obtain a first predicted signal feature representation, wherein the first predicted signal feature representation is used to indicate the changes in the biological pulse signal of the sample object predicted by the second branch network in the sample video; the predicted difference feature representation and the first predicted signal feature representation are input into a classifier to be trained to obtain a predicted detection result.
[0224] For example, in addition to the model architecture corresponding to the liveness detection model, the first model also introduces a teacher prediction model for training the physiological signal branch. That is, the second branch network learns under the guidance of the teacher prediction model to realize the training process.
[0225] The teacher prediction model is a pre-trained machine learning model used to predict the physiological signal feature representation of sample objects in the sample video.
[0226] For example, the teacher prediction model mentioned above can be implemented using neural networks such as convolutional neural networks, feedforward neural networks, residual networks, and converters, without specific limitations here.
[0227] The teacher prediction model is implemented as an rPPG-Estimate Teacher Model. The output of this teacher prediction model is an rPPG signal, which is used as a pseudo-label for training the second branch network.
[0228] Step 1330: Based on the difference between the sample detection results and the predicted detection results, iteratively train the first model to obtain the liveness detection model.
[0229] Based on the difference between the sample detection result and the predicted detection result, the first loss value is obtained; based on the first loss value, the first model is iteratively trained to obtain the liveness detection model.
[0230] The first loss value is obtained by inputting the sample detection result and the predicted detection result into a first preset loss function to obtain the first loss value. For example, the first preset loss function can be implemented as at least one of the following: Cross-Entropy Loss, Binary Cross-Entropy Loss (BCE Loss), Mean Squared Error Loss (MSE), Logarithmic Loss, Least Absolute Deviations Loss (L1 Loss), etc., without limitation.
[0231] For example, when the first model also includes a teacher prediction model, the output of the teacher prediction model is used as a pseudo-label to supervise the learning of the second branch network. That is, the teacher prediction model guides the second branch network to learn key information that combines with video features in the process of generating physiological signal feature representations, thereby improving the training effect of the second branch network and thus improving the model performance of the trained liveness detection model. At the same time, the training of the teacher prediction model can use non-liveness training data such as masks and models as supplementary data sources. In other words, the introduction of the teacher prediction model can provide diverse sample experience for the physiological signal branch, enabling the physiological signal branch to learn a wider range of data distribution features, thereby improving the generalization ability of the trained liveness detection model.
[0232] For example, based on the difference between the sample detection result and the predicted detection result, a first loss value is obtained; a teacher prediction model is obtained; the sample video is input into the teacher prediction model to obtain a second prediction signal feature representation, wherein the second prediction signal feature representation is used to indicate the changes in the biological pulse signal of the sample object predicted by the teacher prediction model in the sample video; based on the difference between the first prediction signal feature representation and the second prediction signal feature representation, a second loss value is obtained; based on the first loss value and the second loss value, the first model is iteratively trained to obtain a liveness detection model.
[0233] For example, the second loss value can be obtained by inputting the feature representation of the first predicted signal and the feature representation of the second predicted signal into a second preset loss function to obtain the second loss value. For example, the second preset loss function can be implemented as at least one of the following: cross-entropy loss function, binary cross-entropy loss function, mean squared error loss function, logarithmic loss function, minimum absolute deviation loss function, etc., without limitation.
[0234] In some cases, the model parameters of the first model can be iteratively adjusted by weighting the first and second loss values to train a liveness detection model. In other cases, the model parameters of the teacher prediction model are frozen while adjusting the model parameters of the first model using the first and second loss values.
[0235] To improve the training performance of the second branch network, multiple loss functions are combined to obtain a second loss value. During the training of the second branch network, in addition to considering the loss between the pseudo-labels output by the teacher prediction network and the feature representations output by the second branch network, the loss between the hidden layer output features of the teacher prediction network and the second branch network can also be considered.
[0236] For example, the second preset loss function includes a first loss function and a second loss function; substituting the first and second predicted signal feature representations into the first loss function yields a first sub-loss value; obtaining the first intermediate feature representation generated by the second branch network during the generation of the first predicted signal feature representation, wherein the first intermediate feature representation is the hidden layer output of the second branch network; obtaining the second intermediate feature representation generated by the teacher prediction model during the generation of the second predicted signal feature representation, wherein the second intermediate feature representation is the hidden layer output of the teacher prediction network, and the hidden layers of the second branch network and the teacher prediction network have the same dimension; substituting the first and second intermediate feature representations into the second loss function yields a second sub-loss value; and using the weighted sum of the first and second sub-loss values as the second loss value.
[0237] For example, the first loss function mentioned above can be implemented as at least one of the following: cross-entropy loss function, binary cross-entropy loss function, mean squared error loss function, logarithmic loss function, minimum absolute deviation loss function, etc., without limitation.
[0238] For example, the second loss function mentioned above can be implemented as at least one of the following: cross-entropy loss function, binary cross-entropy loss function, mean square error loss function, logarithmic loss function, minimum absolute deviation loss function, etc., without limitation.
[0239] For example, the first loss function and the second loss function can be the same or different.
[0240] In one example, the first predicted signal feature representation of the output of the second branch network is implemented as the first rPPG signal, and the second predicted signal feature representation of the output of the teacher prediction network is implemented as the second rPPG signal. This can be achieved through supervised training using the Pearson correlation coefficient and the application of mean squared error (MSE) to the hidden layer output.
[0241] The Pearson correlation coefficient is a statistic that measures the degree of linear correlation between two variables.
[0242] Since the detection performance of Pearson-based methods improves with increasing detection time, the weight data used when weighting and summing the first and second sub-loss values can vary with the video duration of the first video. For example, weight data corresponding to the video duration of the first video can be obtained, including a first weighted value and a second weighted value. The first weighted value is positively correlated with the video duration, and the second weighted value is negatively correlated with the video duration. The first sub-loss value is weighted using the first weighted value, and the second sub-loss value is weighted using the second weighted value. The sum of the weighted first and second sub-loss values yields the second loss value. In other words, the longer the video duration, the larger the weighting value for the first sub-loss value and the smaller the weighting value for the second sub-loss value; conversely, the shorter the video duration, the smaller the weighting value for the first sub-loss value and the larger the weighting value for the second sub-loss value.
[0243] In the analysis of short videos, the signals (outputs of the teacher prediction model and the second branch network) are similar and difficult to distinguish as surface information. Therefore, in the case of short videos, the first weighting value corresponding to the first sub-loss value between the first and second rPPG signals is reduced, while the second weighting value corresponding to the second sub-loss value between the hidden layer outputs of the second branch network and the teacher prediction model is increased. This allows for greater focus on the loss between hidden layer outputs, better optimization of intermediate states, and improved accuracy in obtaining the second loss value. This helps to train the second branch network more accurately even with short videos, improving the learning effect of the second branch model under the guidance of the teacher prediction model.
[0244] Correspondingly, in the analysis of long videos, the process of guiding the second branch network to predict the first rPPG signal is fully considered when the video duration is long. Therefore, in the case of long video duration, a larger first weighting value is used to give a larger proportion to the first sub-loss value. This allows the training process of the second branch network to gradually shift from focusing more on the loss between hidden layer outputs to focusing more on the loss between signals. This avoids continuously focusing too much on the loss between hidden layer outputs and ignoring the output of signal results. It helps to guide the second branch network to improve the output efficiency of signals while training accurately, and enables the second branch network to be appropriately optimized and improved at different stages, thereby improving the robustness of model training.
[0245] By combining video duration with the selection of weight data, a good coordination can be achieved between the training process for sample videos with shorter durations and the training process for sample videos with longer durations. In the stage of applying the trained physiological signal branch network, it helps the network to analyze and determine the feature representation of physiological signals more accurately, and also helps to output the feature representation of physiological signals downstream more efficiently, thus comprehensively improving the output accuracy and output efficiency of physiological signals.
[0246] In one example, a mixing factor α that grows over time is designed to reconcile the trade-off between the first and second loss functions. For instance, the second preset loss function is shown in Equation 3. Equation 3: L rPPG = (1-α)·L mse (H pre H gt )+α·L pearson (R pre ,R gt )
[0247] Among them, H pre H represents the hidden layer output of the second branch network. gt R represents the hidden layer output of the teacher prediction network.pre R represents the first rPPG signal output by the second branch network. gt L represents the second rPPG signal output by the teacher prediction network. mse () represents the second loss function, also known as the MSE loss function, L pearson () represents the first loss function, also known as the Pearson loss function.
[0248] It is evident that, although Pearson supervision can effectively learn the features of the hidden layer within a very short time interval, its efficiency in directly learning these latent features is insufficient. Therefore, when the video duration is short, the weight of the second sub-loss value is increased to ensure better optimization of intermediate states, thereby balancing the optimization of intermediate states and output results. Conversely, as the detection performance based on Pearson increases with the detection time, meaning that the Pearson loss function can effectively learn the differences between intermediate states and output results, the weight of the first sub-loss value is increased when the video duration is long to ensure a balance between the optimization of intermediate states and output results.
[0249] In another example, depending on the needs of different businesses, the above-mentioned mixing factor α can also be implemented as a value that decays over time. That is, the longer the video duration, the smaller the mixing factor α, so that the attention to the first sub-loss value can be reduced over time, while the attention to the second sub-loss value can be increased.
[0250] In this embodiment, the extracted rPPG features and signals are made more accurate by using the Pearson loss function and the MSE loss function. The Pearson loss function measures the linear relationship between the predicted rPPG signal and the real rPPG signal, thereby optimizing the temporal features. The MSE loss function further improves the accuracy of rPPG signal prediction by minimizing the squared error between the hidden layer output of the physiological signal branch and the hidden layer output of the teacher prediction model. This enables liveness detection to be completed even in very short videos, significantly improving the efficiency of liveness detection.
[0251] Figure 14 shows a schematic diagram of the model architecture of the first model 1400 according to an embodiment of this application. The model architecture of the first model 1400 includes a feature extraction network 1410, a first branch network 1420, a second branch network 1430, a classifier to be trained 1440, and a teacher prediction model 1450.
[0252] For example, a sample video is input into the first model 1400, and a visual feature representation of the sample is extracted by the feature extraction network 1410. The visual feature representation of the sample is input into the first branch network 1420 and the second branch network 1430. The first branch network 1420 outputs a predicted difference feature representation based on the visual feature representation of the sample. The second branch network 1430 outputs a first predicted signal feature representation based on the visual feature representation of the sample. The predicted difference feature representation and the first predicted signal feature representation are input into the classifier 1440 to be trained to obtain the predicted detection result.
[0253] The sample video will also be input into the teacher prediction model 1450 to obtain the second prediction signal feature representation.
[0254] The first loss value 1401 is obtained by predicting the detection results and the sample labels corresponding to the sample videos; the second loss value 1402 is obtained by using the first and second predicted signal feature representations, as well as the hidden layer outputs of the second branch network 1430 and the teacher prediction model 1450; the network parameters of the feature extraction network 1410, the first branch network 1420, the second branch network 1430, and the classifier to be trained in the first model 1400 are iteratively trained using the first loss value 1401 and the second loss value 1402. When the first loss value 1401 and the second loss value 1402 converge, a liveness detection model is obtained.
[0255] For example, the liveness recognition performance of the liveness recognition model trained by the model training method provided in this application is shown in Table 1. Table 1 indicates the training performance of existing liveness recognition models and the liveness recognition model provided in this solution on a public dataset. The evaluation metrics for training performance include the area under the curve (AUC) and the equal error rate (EER), where a higher AUC is better and a lower EER is better. The table lists different methods / models, including MS-LBP, CTA, VGG16, FBnet-RGB, and the model of this application; it also lists different cross-domain tasks, including cross-domain tasks between the 3DMAD domain, V1+, and CSMAD domains. Table 1
[0256] In summary, the liveness detection model trained in this embodiment obtains a difference feature representation by analyzing the differences in the texture of the first object in the first video between image frames, and obtains a physiological signal feature representation by analyzing the changes in the biological pulse signal of the first object in the first video. The liveness detection result is generated by combining the difference feature representation representing texture features and the physiological signal feature representation representing physiological features. That is, when performing liveness detection, corresponding features are obtained from the visual branch and the physiological signal branch respectively, and the features of multiple branches are fused to perform liveness detection, which makes up for the defects of performing liveness detection from a single feature dimension of biological pulse signal, improves the accuracy of liveness detection results, and also improves the robustness of the liveness detection process.
[0257] Figure 15 shows a flowchart of the facial liveness recognition method according to an embodiment of this application. In this example, the liveness recognition model trained above is applied to a face verification scenario. The liveness recognition model includes a feature extraction network, a visual branch network, a physiological signal branch network, and a classifier. Taking the method as being executed by an authentication server as an example, the method includes steps 1501 to 1505.
[0258] Step 1501: Obtain the face video.
[0259] In this embodiment of the application, the terminal collects a user's facial video, and the video content of the facial video includes the user's facial area.
[0260] In one example, taking an authentication scenario, when a user needs to authenticate, they turn on the terminal's camera and point it at their face, capturing a video of their face through the terminal.
[0261] For example, the terminal uploads the captured facial video to an authentication server. In other examples, when the terminal has an authentication device deployed locally, it can also send the captured facial video to the local authentication device to achieve liveness detection and authentication.
[0262] Step 1502: Input the face video into the feature extraction network to obtain a visual feature representation.
[0263] For example, after the authentication server obtains a face video from the terminal, it inputs the face video into a pre-trained feature extraction network to obtain a visual feature representation. This visual feature representation is used to indicate the visual information of the face video and represents its shallow features.
[0264] Shallow features refer to the features obtained by extracting features from the first video through one or more convolutional layers. They are a preliminary abstraction and representation of the input first video and can reflect the basic attributes of the first video, such as information on color, texture, edges, and corners in the video frames, as well as the temporal information between video frames.
[0265] For example, the network structure of the feature extraction network in this embodiment is shown in Figure 5, which will not be described in detail here.
[0266] For example, a face video can be segmented into frames to obtain multiple face image frames; image features of multiple face image frames can be extracted separately through a feature extraction network to obtain image feature representations corresponding to each face image frame; and the image feature representations corresponding to multiple face image frames can be fused to obtain a visual feature representation.
[0267] In some cases, preprocessing operations can be performed on the face video before extracting its visual feature representation. For example, preprocessing operations include at least one of video format conversion, video cropping, frame splitting, noise reduction, and target annotation.
[0268] Step 1503: Input the visual feature representation into the visual branch network to obtain the differential feature representation.
[0269] For example, difference features represent the differences in facial texture between face image frames in a face video. In other words, difference features represent information about how facial texture changes over time, reflecting the visual information conveyed by the face video.
[0270] For example, visual branching networks extract facial texture features from visual feature representations to obtain differential feature representations. Here, facial texture features refer to the skin texture features present on the face.
[0271] For example, the feature processing of visual feature representations by visual branch networks may include: performing central difference convolution on the visual feature representations, extracting information on the changes in facial texture features over time in face videos, and obtaining differential feature representations.
[0272] In some examples, to further optimize the temporal and spatial features extracted from face videos in the visual branch, a spatiotemporal attention operation based on a temporal-spatial attention mechanism is added after the CDC operation. For example, a central difference convolution is performed on the visual feature representation to obtain a first feature representation, which includes sub-features corresponding to each time step of multiple face image frames in the face video; based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation, a difference feature representation is obtained through a spatiotemporal attention mechanism.
[0273] For example, the network structure of the visual branch network in this embodiment is shown in Figure 6, which will not be described in detail here. The feature processing process implemented by the above-mentioned visual branch network includes: performing central difference convolution on the visual feature representation to obtain a first feature representation; obtaining an attention feature representation through a spatiotemporal attention mechanism based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation; performing dimensionality reduction convolution on the attention feature representation to obtain a second feature representation; performing self-attention analysis on the attention feature representation to obtain a third feature representation; and fusing the second feature representation and the third feature representation to obtain a differential feature representation.
[0274] Step 1504: Input the visual feature representation into the physiological signal branch network to obtain the facial pulse signal feature representation.
[0275] For example, facial pulse signal features are used to indicate the changes in facial pulse signals in a face video. In other words, facial pulse signal features characterize information about how the pulse characteristics of a face change over time.
[0276] Among them, the facial pulse signal can reflect the user's heart activity and the state of the vascular system, and is used to characterize the signal generated by the change in blood vessel volume caused by the heartbeat in the face.
[0277] In some examples, the aforementioned facial pulse signal can be realized as a facial rPPG signal, that is, the facial pulse signal features represent the changes in the rPPG signal of the face in the face video.
[0278] For example, the network structure of the physiological signal branch network in this embodiment is shown in Figure 7, and will not be described in detail here. For example, the feature processing process of the physiological signal branch network for the visual feature representation includes: performing convolution and pooling operations on the visual feature representation to obtain a multi-channel feature representation, which includes the physiological features of the first object in the spatial and temporal dimensions; and projecting the multi-channel feature representation onto the signal space to obtain the facial pulse signal feature representation.
[0279] Step 1505: Combine the differential feature representation and the facial pulse signal feature representation to classify the faces in the face video and obtain the face liveness recognition result.
[0280] The aforementioned face liveness recognition results are used to indicate whether a face is a live person; that is, the face liveness recognition results can indicate whether the face appearing in the face video is a real face.
[0281] For example, the above-mentioned face liveness recognition results include indicating that the face is a live face, or indicating that the face is a fraudulent face. In the case of fraudulent faces, the face in the video does not belong to a real live face, such as a headgear, mask, photo, model, etc.
[0282] For example, the input to the classifier is the differential feature representation and the facial pulse signal feature representation. For instance, by inputting the differential feature representation and the facial pulse signal feature representation together into the classifier, the result of facial liveness recognition can be obtained.
[0283] For example, the input to the classifier is a feature representation obtained by fusing the differential feature representation and the facial pulse signal feature representation. For instance, fusing the differential feature representation and the facial pulse signal feature representation yields a fused feature representation, which is then input into the classifier to obtain the facial liveness recognition result.
[0284] For example, the classifier described above can be implemented as at least one of the following: a linear model-based classifier; a decision tree-based classifier; a neural network-based classifier; a nearest neighbor-based classifier; a reinforcement learning-based classifier, etc.
[0285] In some examples, the authentication server can input the face liveness detection result output by the classifier into the authentication service. When the face liveness detection result indicates that the face in the face video is a live face, the authentication service segments the face region in the face video to obtain a face image; it then matches the face image with a pre-stored account verification face image; if the difference between the face image in the face video and the account verification face image is less than a preset threshold, it sends a successful authentication feedback to the terminal; if the difference between the face image in the face video and the account verification face image is greater than or equal to the preset threshold, it sends a failed authentication feedback to the terminal; and if the face liveness detection result indicates that the face in the face video is a fraudulent face, it sends a failed authentication feedback to the terminal.
[0286] In summary, in the process of achieving liveness detection for faces in face videos, the differences in facial texture between face image frames are analyzed to obtain differential feature representations. Furthermore, the changes in facial pulse signals in the face video are analyzed to obtain facial pulse signal feature representations. By combining the differential feature representations representing texture features and the facial pulse signal feature representations representing physiological features, liveness detection results are generated. In other words, during liveness detection, corresponding features are obtained from both the visual and physiological signal branches, and the features from multiple branches are fused to perform liveness detection. This overcomes the shortcomings of relying solely on the single feature dimension of biological pulse signals for liveness detection, improving the accuracy of liveness detection results and enhancing the robustness of the liveness detection process.
[0287] In other examples, the liveness detection method provided in this application can also be applied to the detection of other live objects, such as palms, irises, etc. Here, only the face is used as an example for illustrative purposes.
[0288] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0289] Figure 16 shows a structural block diagram of a liveness detection device according to an embodiment of this application. The device includes the following modules:
[0290] The acquisition module 1610 is used to acquire a first video including a first object, wherein the first object is the object to be identified as a liveness detection object;
[0291] The first analysis module 1620 is used to analyze the texture features of the first object in the first video to obtain the difference feature representation corresponding to the first video. The difference feature representation is used to indicate the differences in the object texture of the first object between the image frames of the first video.
[0292] The second analysis module 1630 is used to analyze the physiological characteristics of the first object in the first video to obtain the physiological signal feature representation corresponding to the first video. The physiological signal feature representation is used to indicate the changes in the biological pulse signal exhibited by the first object in the first video.
[0293] The classification module 1640 is used to classify the first object in the first video by combining the differential feature representation and the physiological signal feature representation, and to obtain the liveness recognition result corresponding to the first object. The liveness recognition result is used to indicate that the first object is a live object.
[0294] For example, the first analysis module 1620 is further configured to extract the texture features of the first object in the first video based on the visual feature representation to obtain the difference feature representation, wherein the visual feature representation is a feature extracted from the first video, and the visual feature representation is used to indicate the visual information of the first video;
[0295] The second analysis module 1630 is further configured to generate a physiological signal feature representation of the first object in the first video based on the visual feature representation.
[0296] For example, as shown in Figure 17, the device further includes:
[0297] The extraction module 1650 is further configured to divide the first video into frames to obtain multiple image frames; extract image features from the multiple image frames respectively to obtain image feature representations corresponding to the multiple image frames respectively; and fuse the image feature representations corresponding to the multiple image frames respectively to obtain the visual feature representation.
[0298] For example, the first analysis module 1620 is further configured to perform central difference convolution on the visual feature representation to extract information on the temporal variation of the texture features of the first object in the first video, thereby obtaining the difference feature representation. The central difference convolution is a convolution operation used to extract local differences between image frames.
[0299] For example, the first analysis module 1620 is further configured to perform central difference convolution on the visual feature representation to obtain a first feature representation, the first feature representation including sub-features of time steps corresponding to multiple image frames of the first video; based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation, the difference feature representation is obtained through a spatiotemporal attention mechanism, the spatiotemporal attention mechanism being an attention mechanism that combines the time dimension and the image space dimension to achieve feature processing.
[0300] For example, the first analysis module 1620 is further configured to obtain an attention feature representation based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation through the spatiotemporal attention mechanism; perform dimensionality reduction convolution on the attention feature representation to obtain a second feature representation; perform self-attention analysis on the attention feature representation to obtain a third feature representation; and fuse the second feature representation and the third feature representation to obtain the difference feature representation.
[0301] For example, the second analysis module 1630 is further configured to perform convolution and pooling operations on the visual feature representation to obtain a multi-channel feature representation, wherein the multi-channel feature representation includes the physiological features of the first object in the spatial and temporal dimensions; and project the multi-channel feature representation onto the signal space to obtain the physiological signal feature representation.
[0302] For example, the acquisition module 1610 is further configured to acquire a sample video including a sample object, and a sample tag corresponding to the sample object;
[0303] The device may further include:
[0304] The training module 1660 is used to input the sample video into the first model and generate a prediction detection result. The prediction detection result is used to indicate the liveness classification result of the first model for the sample object. Based on the difference between the sample detection result and the prediction detection result, the first model is iteratively trained to obtain the liveness recognition model.
[0305] For example, the first model includes a feature extraction network, a first branch network, a second branch network, and a classifier to be trained. The first branch network is used to extract texture features, and the second branch network is used to extract physiological features.
[0306] The training module 1660 is further configured to input the sample video into the feature extraction network to obtain a sample visual feature representation; input the sample visual feature representation into the first branch network to obtain a predicted difference feature representation; input the sample visual feature representation into the second branch network to obtain a first predicted signal feature representation, wherein the first predicted signal feature representation is used to indicate the changes in the biological pulse signal exhibited by the sample object in the sample video as predicted by the second branch network; and input the predicted difference feature representation and the first predicted signal feature representation into the classifier to be trained to obtain the predicted detection result.
[0307] For example, the training module 1660 is further configured to obtain a first loss value based on the difference between the sample detection result and the predicted detection result;
[0308] The acquisition module 1610 is further configured to acquire a teacher prediction model, wherein the teacher prediction model is a pre-trained machine learning model for predicting the physiological signal feature representation of the sample object in the sample video.
[0309] The training module 1660 is further configured to input the sample video into the teacher prediction model to obtain a second prediction signal feature representation, the second prediction signal feature representation being used to indicate the changes in the biological pulse signal of the sample object as predicted by the teacher prediction model in the sample video; to obtain a second loss value based on the difference between the first prediction signal feature representation and the second prediction signal feature representation; and to iteratively train the first model based on the first loss value and the second loss value to obtain the liveness recognition model.
[0310] For example, the training module 1660 is further configured to substitute the first predicted signal feature representation and the second predicted signal feature representation into a first loss function to obtain a first sub-loss value; obtain a first intermediate feature representation generated by the second branch network in the process of generating the first predicted signal feature representation, wherein the first intermediate feature representation is the hidden layer output of the second branch network; obtain a second intermediate feature representation generated by the teacher prediction model in the process of generating the second predicted signal feature representation, wherein the second intermediate feature representation is the hidden layer output of the teacher prediction network, and the hidden layer of the second branch network and the hidden layer of the teacher prediction network have the same dimension; substitute the first intermediate feature representation and the second intermediate feature representation into a second loss function to obtain a second sub-loss value; and use the weighted sum of the first sub-loss value and the second sub-loss value as the second loss value.
[0311] For example, the training module 1660 is further configured to acquire weight data corresponding to the video duration of the first video, the weight data including a first weighted value and a second weighted value, the first weighted value being positively correlated with the video duration and the second weighted value being negatively correlated with the video duration; the first sub-loss value is weighted by the first weighted value, and the second sub-loss value is weighted by the second weighted value, and the weighted first sub-loss value and the weighted second sub-loss value are summed to obtain the second loss value.
[0312] The examples of liveness detection devices in this application are merely illustrative of the division of the above-described functional modules. In practical applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the liveness detection devices and liveness detection method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0313] Figure 18 shows a schematic diagram of the server structure according to an embodiment of this application.
[0314] Server 1800 includes a Central Processing Unit (CPU) 1801, a system memory 1804 including Random Access Memory (RAM) 1802 and Read Only Memory (ROM) 1803, and a system bus 1805 connecting the system memory 1804 and the CPU 1801. Server 1800 also includes a mass storage device 1806 for storing the operating system 1813, application programs 1814, and other program modules 1815.
[0315] Mass storage device 1806 is connected to central processing unit 1801 via a mass storage controller (not shown) connected to system bus 1805. Mass storage device 1806 and its associated computer-readable media provide non-volatile storage for server 1800. That is, mass storage device 1806 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.
[0316] Computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1804 and mass storage device 1806 described above can be collectively referred to as memory.
[0317] Server 1800 can also connect to remote computers on a network, such as the Internet. That is, server 1800 can connect to network 1812 via network interface unit 1811 connected to system bus 1805, or it can use network interface unit 1811 to connect to other types of networks or remote computer systems (not shown).
[0318] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU.
[0319] Embodiments of this application also provide a computer device including a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The processor loads and executes the at least one instruction, at least one program, code set, or instruction set to implement the liveness detection method provided in the above-described method embodiments. For example, the computer device can be a terminal or a server.
[0320] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the liveness detection method provided in the above-described method embodiments.
[0321] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the liveness detection methods described in the above embodiments.
[0322] For example, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments described above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0323] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0324] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0325] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A liveness detection method, comprising: Acquire a first video including a first object, wherein the first object is the object to be identified as a liveness detection subject; By analyzing the texture features of the first object in the first video, a difference feature representation corresponding to the first video is obtained. The difference feature representation is used to indicate the changes in the object texture of the first object between image frames of the first video. By analyzing the physiological characteristics of the first object in the first video, a physiological signal feature representation corresponding to the first video is obtained. The physiological signal feature representation is used to indicate the changes in the biological pulse signal exhibited by the first object in the first video. By using the differential feature representation and the physiological signal feature representation to classify the first object in the first video, a liveness detection result corresponding to the first object is obtained, and the liveness detection result is used to indicate whether the first object is a live object.
2. The method according to claim 1, further comprising: Extract visual feature representations of the first video, wherein the visual feature representations are used to represent at least two basic attributes of the first video; The step of analyzing the texture features of the first object in the first video to obtain the difference feature representation corresponding to the first video includes: The difference feature representation is obtained by extracting the texture features of the first object in the first video from the visual feature representation; The step of analyzing the physiological characteristics of the first object in the first video to obtain the physiological signal feature representation corresponding to the first video includes: The physiological signal feature representation of the first object in the first video is generated based on the visual feature representation.
3. The method according to claim 2, wherein, Extracting the visual feature representation of the first video includes: The first video is divided into frames to obtain multiple image frames; By extracting the image features of the multiple image frames respectively, the image feature representations corresponding to the multiple image frames are obtained; The visual feature representation is obtained by fusing the image feature representations corresponding to the multiple image frames respectively.
4. The method according to claim 2 or 3, wherein, The difference feature representation is obtained by extracting the texture features of the first object in the first video from the visual feature representation, including: The difference feature representation is obtained by performing a central difference convolution on the visual feature representation to extract information about the temporal variation of the texture features of the first object in the first video. The central difference convolution is a convolution operation used to extract local differences between image frames.
5. The method according to claim 4, wherein, The differential feature representation is obtained by performing a central difference convolution on the visual feature representation to extract information about the temporal variation of the texture features of the first object in the first video, including: By performing central difference convolution on the visual feature representation, a first feature representation is obtained, which includes sub-features of time steps corresponding to multiple image frames of the first video. Based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation, the differential feature representation is obtained through a spatiotemporal attention mechanism, which is an attention mechanism that combines the time dimension and the image space dimension to achieve feature processing.
6. The method according to claim 5, wherein, Based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation, the differential feature representation is obtained through a spatiotemporal attention mechanism, including: Based on the correlation between the sub-features corresponding to multiple time steps in the first feature representation, an attention feature representation is obtained through the spatiotemporal attention mechanism. A second feature representation is obtained by performing a dimensionality reduction convolution on the attention feature representation; By performing self-attention analysis on the attention feature representation, a third feature representation is obtained; The difference feature representation is obtained by fusing the second feature representation and the third feature representation.
7. The method according to claim 2 or 3, wherein, The step of generating the physiological signal feature representation of the first object in the first video based on the visual feature representation includes: By performing convolution and pooling operations on the visual feature representation, a multi-channel feature representation is obtained, which includes the physiological features of the first object in the spatial and temporal dimensions. The physiological signal feature representation is obtained by projecting the multi-channel feature representation onto the signal space.
8. The method according to any one of claims 1 to 3, wherein, The method is executed by a liveness detection model, which is obtained by training a first model. The training process of the first model includes: Obtain sample videos including sample objects, and sample tags corresponding to the sample objects; By inputting the sample video into the first model, a prediction detection result is generated, which is used to indicate the liveness classification result of the first model for the sample object; The liveness detection model is obtained by iteratively training the first model based on the difference between the sample detection results and the predicted detection results.
9. The method according to claim 8, wherein, The first model includes a feature extraction network, a first branch network, a second branch network, and a classifier to be trained. The first branch network is used to extract texture features, and the second branch network is used to extract physiological features. The step of inputting the sample video into the first model to generate a prediction detection result includes: By inputting the sample video into the feature extraction network, a visual feature representation of the sample is obtained; By inputting the visual feature representation of the sample into the first branch network, a predicted difference feature representation is obtained; By inputting the visual feature representation of the sample into the second branch network, a first prediction signal feature representation is obtained. The first prediction signal feature representation is used to indicate the changes in the biological pulse signal of the sample object as predicted by the second branch network in the sample video. The predicted detection result is obtained by inputting the predicted difference feature representation and the first predicted signal feature representation into the classifier to be trained.
10. The method according to claim 9, wherein, The liveness detection model is obtained by iteratively training the first model based on the difference between the sample detection result and the predicted detection result, including: Based on the difference between the sample detection result and the predicted detection result, a first loss value is obtained; By inputting the sample video into a pre-trained teacher prediction model for predicting the physiological signal feature representation of the sample object in the video, a second prediction signal feature representation is obtained. The second prediction signal feature representation is used to indicate the changes in the biological pulse signal of the sample object as predicted by the teacher prediction model in the sample video. A second loss value is obtained based on the difference between the first predicted signal feature representation and the second predicted signal feature representation; The liveness detection model is obtained by iteratively training the first model based on the first loss value and the second loss value.
11. The method according to claim 10, wherein, Based on the difference between the first predicted signal feature representation and the second predicted signal feature representation, a second loss value is obtained, including: By substituting the first predicted signal feature representation and the second predicted signal feature representation into the first loss function, a first sub-loss value is obtained; Obtain the first intermediate feature representation generated by the second branch network during the process of generating the feature representation of the first predicted signal, wherein the first intermediate feature representation is the hidden layer output of the second branch network; Obtain the second intermediate feature representation generated by the teacher prediction model during the process of generating the second prediction signal feature representation. The second intermediate feature representation is the hidden layer output of the teacher prediction network. The hidden layer of the second branch network has the same dimension as the hidden layer of the teacher prediction network. By substituting the first intermediate feature representation and the second intermediate feature representation into the second loss function, the second sub-loss value is obtained; The weighted sum of the first sub-loss value and the second sub-loss value is used as the second loss value.
12. The method according to claim 11, wherein, The weighted sum of the first sub-loss value and the second sub-loss value is used as the second loss value, including: Obtain weight data corresponding to the video duration of the first video. The weight data includes a first weighted value and a second weighted value. The first weighted value is positively correlated with the video duration, and the second weighted value is negatively correlated with the video duration. The first sub-loss value is weighted by the first weighting value, and the second sub-loss value is weighted by the second weighting value. The weighted first sub-loss value and the weighted second sub-loss value are then summed to obtain the second loss value.
13. A liveness detection device, comprising: The acquisition module is used to acquire a first video including a first object, wherein the first object is the object to be identified as a liveness detection object; The first analysis module is used to obtain a difference feature representation corresponding to the first video by analyzing the texture features of the first object in the first video. The difference feature representation is used to indicate the changes in the object texture of the first object between image frames of the first video. The second analysis module is used to obtain a physiological signal feature representation corresponding to the first video by analyzing the physiological characteristics of the first object in the first video. The physiological signal feature representation is used to indicate the changes in the biological pulse signal exhibited by the first object in the first video. The classification module is used to classify the first object in the first video by using the differential feature representation and the physiological signal feature representation to obtain the liveness recognition result corresponding to the first object. The liveness recognition result is used to indicate whether the first object is a live object.
14. The apparatus of claim 13, further comprising: An extraction module is used to extract visual feature representations of a first video, wherein the visual feature representations are used to represent at least two basic attributes of the first video; The first analysis module is used to: extract the texture features of the first object in the first video from the visual feature representation to obtain the difference feature representation; The second analysis module is used to: generate a physiological signal feature representation of the first object in the first video based on the visual feature representation.
15. The apparatus according to claim 14, wherein, The extraction module is used for: The first video is divided into frames to obtain multiple image frames; By extracting the image features of the multiple image frames respectively, the image feature representations corresponding to the multiple image frames are obtained; The visual feature representation is obtained by fusing the image feature representations corresponding to the multiple image frames respectively.
16. The method according to claim 14 or 15, wherein, The first analysis module is used for: The difference feature representation is obtained by performing a central difference convolution on the visual feature representation to extract information about the temporal variation of the texture features of the first object in the first video. The central difference convolution is a convolution operation used to extract local differences between image frames.
17. The method according to claim 14 or 15, wherein, The second analysis module is used for: By performing convolution and pooling operations on the visual feature representation, a multi-channel feature representation is obtained, which includes the physiological features of the first object in the spatial and temporal dimensions. The physiological signal feature representation is obtained by projecting the multi-channel feature representation onto the signal space.
18. A computer device comprising a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the liveness detection method as claimed in any one of claims 1 to 12.
19. A computer-readable storage medium storing at least one piece of program code, said program code being loaded and executed by a processor to implement the liveness detection method as described in any one of claims 1 to 12.
20. A computer program product comprising a computer program or instructions that, when executed by a processor, implement the liveness detection method as described in any one of claims 1 to 12.