Face living body authentication method, device, equipment, medium and program product
By acquiring multiple frames of basic images and their physical context for physical consistency verification, the problem of insufficient accuracy and security in face liveness detection in existing technologies is solved, achieving higher accuracy and security in liveness authentication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-07
AI Technical Summary
Existing face liveness detection methods have low accuracy due to environmental and equipment factors, and the security and accuracy of the color liveness detection technology are insufficient.
By acquiring multiple frames of basic images and their physical context, a check pair is generated for physical consistency verification. A pre-trained physical consistency verification model is used to determine the physical consistency score of the face to be authenticated. If the score reaches the threshold, the face is determined to be live.
It improves the accuracy of facial liveness detection, reduces the impact of environmental and equipment factors, and enhances the security and precision of the detection.
Smart Images

Figure CN121811510A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a face liveness authentication method, device, equipment, medium, and program product. Background Technology
[0002] Facial recognition technology has been widely used in security, finance, and surveillance. However, while bringing convenience, it also requires greater attention to security. More secure liveness detection technology is crucial in facial recognition applications. Common liveness detection methods include color-coded liveness detection; however, the accuracy of color-coded liveness detection is affected by factors such as the environment, device sensors, and the image signal processor (ISP), leading to lower accuracy in facial liveness authentication. Summary of the Invention
[0003] This application provides a face liveness authentication method, apparatus, device, medium, and program product to solve the problem of low accuracy in face liveness authentication.
[0004] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a face liveness authentication method, including: The system acquires multiple base images containing the face to be authenticated and the physical context corresponding to each base image. Each base image is captured by the client based on a color sequence indicated by a dynamic color instruction. The physical context includes environmental information and device information corresponding to each base image. Multiple check pairs are generated based on multiple key images and the physical cause context corresponding to each key image in each frame. The multiple base images include the multiple key images. Physical consistency verification is performed based on all the aforementioned check pairs to obtain the physical consistency score corresponding to each frame of the key image. If all of the physical consistency scores are greater than or equal to a preset threshold, the face to be authenticated is determined to have passed the liveness authentication.
[0005] Secondly, embodiments of this application provide a face liveness authentication device, comprising: The first acquisition module is used to acquire multiple frames of basic images containing the face to be authenticated and the physical cause context corresponding to each frame of the basic images. Each frame of the basic images is obtained by the client capturing the face to be authenticated based on the color sequence indicated by the dynamic color instruction. The physical cause context includes environmental information and device information corresponding to each frame of the basic images. A generation module is used to generate multiple check pairs based on multiple key images and the physical cause context corresponding to each key image, wherein the multiple base images include the multiple key images. The verification module is used to perform physical consistency verification based on all the verification pairs to obtain the physical consistency score corresponding to each frame of the key image. The determination module is used to determine that the face to be authenticated has passed the liveness authentication if all of the physical consistency scores are greater than or equal to a preset threshold.
[0006] Thirdly, embodiments of this application provide an electronic device, including a transceiver and a processor. The transceiver is used to: acquire multiple frames of basic images containing the face to be authenticated and the physical cause context corresponding to each frame of the basic images. Each frame of the basic images is obtained by the client capturing the face to be authenticated based on a color sequence indicated by a dynamic color instruction. The physical cause context includes environmental information and device information corresponding to each frame of the basic images. The processor is used for: Multiple check pairs are generated based on multiple key images and the physical cause context corresponding to each key image in each frame. The multiple base images include the multiple key images. Physical consistency verification is performed based on all the aforementioned check pairs to obtain the physical consistency score corresponding to each frame of the key image. If all of the physical consistency scores are greater than or equal to a preset threshold, the face to be authenticated is determined to have passed the liveness authentication.
[0007] Fourthly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the face liveness authentication method as described in the first aspect above.
[0008] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the face liveness authentication method described in the first aspect above.
[0009] Sixthly, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the face liveness authentication method as described in the first aspect above.
[0010] In this embodiment, multiple base images containing the face to be authenticated and the corresponding physical context of each base image are obtained. Each base image is captured by the client based on a color sequence indicated by a dynamic color instruction. The physical context includes environmental and device information corresponding to each base image. Multiple verification pairs are generated based on multiple key images and the corresponding physical contexts. The multiple base images include the multiple key images. Physical consistency is verified based on all verification pairs to obtain a physical consistency score for each key image. If all physical consistency scores are greater than or equal to a preset threshold, the face to be authenticated is determined to have passed liveness authentication. Thus, by obtaining multiple base images containing the face to be authenticated and the corresponding physical context, the physical consistency score between the multiple base images and the physical context of the face to be authenticated can be determined. Based on the physical consistency score, the face to be authenticated is determined to have passed liveness authentication. This allows environmental and device information to be used as verification content, thereby avoiding the influence of environmental and device factors on the verification effect and improving the accuracy of face liveness authentication. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a face liveness authentication method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a face liveness authentication device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] For ease of understanding, the following describes some aspects of the embodiments of this application: In recent years, the deception methods of the black market have become increasingly sophisticated. In addition to typical face forgery methods such as photo copying, video copying, artificial masks, 3D models, and AI face swapping, they can also directly upload pre-shot videos by hijacking cameras and use illegal videos to deceive liveness detection models in order to obtain illegal profits.
[0015] Face liveness detection methods in related technologies mainly include two approaches: user interaction and color-coded liveness detection. The algorithms typically employ common deep learning networks such as UNet and ResNet. User interaction involves specifying actions and lip reading for liveness verification to correctly identify forged faces. However, this method requires user cooperation, reducing detection speed and user experience, and it cannot confirm whether the user's video has been switched. Color-coded liveness detection projects different colors of light onto the face within a specific time period, using algorithms to determine if the user in the video is "live." Because different colors of light reflect differently in attack scenarios such as photo manipulation compared to a real face, algorithms can differentiate these features to achieve liveness detection.
[0016] The relevant technical solutions have the following main drawbacks: Disadvantage 1: Common face liveness detection methods include 3D imaging and user interaction, such as specified actions and lip reading, to verify liveness and correctly identify forged faces. However, user interaction requires user cooperation, reducing detection speed and user experience, and it's impossible to confirm whether the customer's video has been switched. Therefore, before performing liveness detection, it's necessary to verify whether the video is captured in real-time by the camera. Improving the accuracy of liveness detection is a technical problem that needs to be solved, assuming the camera is not compromised. Based on this, a more secure and efficient face liveness detection solution is needed.
[0017] Disadvantage 2: The vibrant light liveness detection technology in related technologies typically separates color verification and liveness detection into two parts. That is, different small models or algorithms are used to perform the detection of these two parts, and the final liveness detection result is obtained by combining the results of both parts. Therefore, the vibrant light liveness detection technology in related technologies still needs improvement.
[0018] Disadvantage 3: During the vibrant liveness detection process, the system displays the color of the light emitted each time, resulting in lower security and vulnerability to attacks based on the emitted light color. Furthermore, the feature learning method in this liveness detection is relatively crude, leading to relatively low detection accuracy. The accuracy of vibrant liveness CAPTCHA verification is also affected by factors such as the environment, device sensors, and device ISP processing. For example, the same color light may not display consistently on the screens of different brands of mobile devices, and the same light color may result in different colors in the collected data from different brands of mobile devices, leading to low accuracy for the relevant algorithm in vibrant liveness CAPTCHA verification.
[0019] This application proposes a face liveness authentication method, apparatus, device, medium, and program product to address the problem of low accuracy in face liveness authentication.
[0020] See Figure 1 , Figure 1 This is a flowchart of a face liveness authentication method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps: Step 101: Obtain multiple frames of basic images containing the face to be authenticated and the physical context corresponding to each frame of the basic images. Each frame of the basic images is obtained by the client capturing the face to be authenticated based on the color sequence indicated by the dynamic color instruction. The physical context includes the environmental information and device information corresponding to each frame of the basic images.
[0021] Specifically, the aforementioned client can be a user's terminal device, such as a smartphone or tablet. For example, a user uses a mobile phone for facial recognition authentication.
[0022] The aforementioned multi-frame base images can be a collection of static image frames containing the face to be authenticated, continuously captured and uploaded by the client. Alternatively, these multi-frame base images can be images obtained by parsing a video file uploaded by the client. The video file can be a face video recorded according to a specific color sequence, such as a 3-second face video recorded according to a "red, green, blue" color sequence.
[0023] The aforementioned physical context can be used to indicate the physical environment parameters when capturing the face to be authenticated. For example, these may include ambient light intensity, device model, camera parameters, etc.
[0024] The aforementioned dynamic color instruction can be a random color sequence instruction sent by the server to the client. For example, the dynamic color instruction can be "display red for 0-1 seconds, green for 1-2 seconds, and blue for 2-3 seconds".
[0025] The environmental information mentioned above can include environmental parameters such as lighting conditions when the face to be authenticated is captured, such as indoor fluorescent lighting, color temperature, and brightness. The device information mentioned above can include the hardware parameters of the device used to capture the face to be authenticated, such as the device model and camera focal length.
[0026] Step 102: Generate multiple check pairs based on multiple key images and the physical cause context corresponding to each key image, wherein the multiple base images include the multiple key images.
[0027] Specifically, the aforementioned multiple key images can be a subset of image frames selected from multiple base images according to specific rules. It can be understood that these multiple key images are representative samples for core consistency verification. For example, the specific rules could be selecting a key image corresponding to different colors of light, or selecting multiple key images with the best image quality from the multiple base images.
[0028] The aforementioned check pair can refer to a data unit constructed for physical consistency verification. Each check pair consists of a key image frame and textual description information corresponding to the physical context. The textual description information can be structured text that converts the physical parameters corresponding to the physical context. For example, the textual description information could be "color is red, device is 17 Pro, ambient light is natural light".
[0029] Step 103: Perform physical consistency verification based on all the verification pairs to obtain the physical consistency score corresponding to each frame of the key image.
[0030] Specifically, the aforementioned physical consistency score can be a score used to measure the degree of matching between the key image and the physical context of the check pair. It is understood that the higher the physical consistency score, the higher the probability that the key image and its physical context conform to physical laws. For example, the score range can be 0-1, with 0.85 representing high consistency.
[0031] Step 104: If all the physical consistency scores are greater than or equal to the preset threshold, the face to be authenticated is determined to have passed the liveness authentication.
[0032] Specifically, the aforementioned preset threshold can be a pre-set threshold. For example, if the preset threshold is set to 0.8, then all physical consistency scores must be ≥0.8 to determine that the face to be authenticated has passed the liveness authentication.
[0033] Understandably, the aforementioned liveness detection is used to confirm that the face to be authenticated is a real, live face, meaning that the user undergoing identity verification is real and not using photos, videos, or other forgery methods.
[0034] In this embodiment, by acquiring multiple frames of basic images containing the face to be authenticated and the corresponding physical context, the physical consistency score between the multiple frames of basic images of the face to be authenticated and the physical context can be determined. Based on the physical consistency score, it is determined that the face to be authenticated has passed the liveness authentication. This allows environmental and device information to be used as the verification content, thereby avoiding the influence of environmental and device factors on the verification effect and improving the accuracy of face liveness authentication.
[0035] Optionally, before the step of obtaining multiple frames of base images containing the face to be authenticated and the physical context corresponding to each frame of the base images, the method further includes: Obtain a training dataset containing N sample data pairs. Each sample data pair includes a face image sample and its corresponding context sample information. The N sample data pairs include one positive sample data pair and N-1 negative sample data pairs, where N is an integer greater than 1. An initial model and a target loss function are constructed. The initial model includes an image encoder and a text encoder. The image encoder is used to encode the face image samples, and the text encoder is used to encode the context samples. The target loss function is used to characterize the negative logarithm of the ratio of a first similarity to a second similarity. The first similarity is the similarity of the positive sample data pairs, and the second similarity is the sum of the similarities of the N sample data pairs. Based on the objective loss function, the initial model is trained to obtain a pre-trained physical consistency verification model, which is used for physical consistency verification.
[0036] Specifically, the training dataset mentioned above can be a structured data set used to train the model, such as a dataset containing 50 pairs of image and text data.
[0037] The aforementioned sample data pair can be a set of associated image and text data, which may include face image samples and corresponding context sample information. The context sample information may be text information describing the shooting conditions of the face image sample.
[0038] The aforementioned positive sample data pairs can be truly matching image-text pairs, i.e., genuinely captured live human face images and text accurately describing the shooting conditions. The aforementioned negative sample data pairs can be mismatched image-text combinations, such as genuine human face images and incorrect device description text, or non-live human face images and text combinations. It is understood that genuinely captured live human face images refer to images of real, living human faces, where the subject is real and not obtained through photographs, videos, or other forgeries.
[0039] The aforementioned target loss function can be the optimization objective function for model training, and the similarity score of the aforementioned positive sample pair can be the cosine similarity between the image features and text features of the positive sample.
[0040] It is understandable that the above-mentioned training of the initial model can be a process of adjusting the parameters of the initial model through optimization algorithms.
[0041] It should be noted that any real facial image or video (visual evidence) captured by a physical device is a necessary result generated under specific physical conditions (physical causal context). Therefore, there must be a profound and unforgeable physical correlation between the cause and effect of a real captured image. The pre-trained physical consistency verification model mentioned above can be a deep learning model capable of learning, understanding, and verifying this cross-modal physical causal relationship.
[0042] For example, the pre-trained physical consistency verification model mentioned above can be a system architecture based on a two-stream encoder, the core of which lies in effectively characterizing the cause and effect and achieving deep coupling between the two.
[0043] The aforementioned image encoder, acting as a physical phenomenon decoder, receives visual evidence, i.e., image frame I, as input, aiming to decode the rich physical cues generated by the imaging process. These cues constitute a high-dimensional image feature vector V_I, which not only encodes the macroscopic activity characteristics of the organism but, more importantly, encodes the microscopic physical traces such as material surface reflection, scattering, noise patterns, and color rendering under specific lighting and the action of a specific ISP device. This vector is the digital representation of the imaging process result. This process can be formally represented as: ; Where E_I is the image encoder function, These are its trainable parameters.
[0044] The aforementioned text encoder, acting as a physical condition encoder, receives the physical context describing the imaging conditions, i.e., text T, as input. It encodes a series of physical factors declared in the text, such as device model, camera parameters, illumination type, color value, and target subject category, into a high-dimensional text feature vector V_T. This vector is the digital representation of the causes of the imaging process. This process can be formally represented as: ; Where E_T is the text encoder function, These are its trainable parameters.
[0045] A contrastive learning paradigm is adopted to deeply couple the representations of the causes and effects in a self-supervised manner within a unified shared feature space.
[0046] The objective function for training is a physical consistency constraint. For an event that actually occurs in the physical world, the textual feature vector of its cause and the image feature vector of its result are defined as a physically consistent pair, i.e., a positive sample pair, in the feature space. Conversely, any combination that is physically inconsistent is considered a physically inconsistent pair, i.e., a negative sample pair.
[0047] In a system containing N physical consistency pairs In the training batch, for any image feature vector pair Its corresponding It is its only positive sample, while all other text feature vectors within the batch are... (Where j ≠ i) are all considered negative samples. The above training method aims to narrow the distance between positive sample pairs in the feature space while widening the distance between negative sample pairs by optimizing the loss function L. A typical loss function, such as InfoNCE, can be expressed as the loss value L_i for the i-th sample pair in a batch: ; in, The cosine similarity function is used to calculate the similarity. , It is an adjustable temperature hyperparameter. The total loss function L is the average of all L_i in the batch.
[0048] Through this forced constraint, the model can learn a nonlinear mapping function between visual evidence and the physical causal context. The model is no longer memorizing appearances, but learning a set of implicit physical laws.
[0049] After the above training, the initial model evolves into a physical world consistency verification engine. In the authentication and inference phase, its workflow is as follows: (1) Input: Receive the visual evidence to be tested I and a set of physical causal contexts T; (2) Coupling verification: The results and causes are mapped to the shared feature space that has been deeply coupled through the image encoder E_I and the text encoder E_T respectively, to obtain the feature vectors v_I and v_T; (3) Consistency Verification: Calculate the similarity S between the feature vectors of the two entities. This similarity score S is no longer a simple content matching score, but rather a physical consistency score or a score indicating reasonableness of existence. This calculation can be expressed as: S = sim(v_I, v_T); (4) Output: Compare the physical consistency score S with a pre-set threshold Thresh: if S is greater than or equal to Thresh, it is determined to be a real face; otherwise, it is determined to be an attack. This decision-making process can be expressed as: Result = 1, if S >= Thresh; Result = 0, if S < Thresh; Understandably, if the physical consistency score S is greater than Thresh, the observed visual evidence is highly consistent with the physical causal context of the model and is therefore judged as real; conversely, if the physical consistency score S is less than Thresh, it indicates that there is a break in the physical causal chain and is judged as an attack.
[0050] In this implementation, an initial model and a target loss function are constructed. The target loss function is used to represent the negative logarithm of the ratio of the first similarity to the second similarity. The first similarity is the similarity of positive sample data pairs, and the second similarity is the sum of the similarities of N sample data pairs. The initial model is trained based on the target loss function to obtain a pre-trained physical consistency verification model. This allows the physical consistency verification model to learn deeper physical relationships, replacing multiple separate small models used for color verification and liveness detection in related technologies. This holistic learning approach makes feature extraction more thorough and decision boundaries clearer, thereby comprehensively improving detection accuracy and generalization ability.
[0051] Optionally, obtaining the training dataset includes: Acquire multiple frames of captured images of the object being captured, as well as multiple physical cause context metadata corresponding one-to-one with the multiple frames of captured images. The capture information corresponding to the multiple frames of captured images is different, and the capture information includes at least one of color instruction information, capture device information, attack information, and environmental information. The multiple physical cause context metadata are converted into text using a preset text template to obtain multiple context sample information that corresponds one-to-one with the multiple physical cause context metadata. The training dataset is generated based on the multi-frame acquired images and the multiple context sample information.
[0052] Specifically, the aforementioned collected objects can refer to the subjects recorded by the collected devices during the data collection process, which may include real biological individuals, i.e., living organisms, or non-living carrier attack samples used to simulate attacks.
[0053] The aforementioned multi-frame acquired images can be composed of image frames captured by the acquisition device and arranged in a time sequence, recording the complete optical response process of the acquired object under dynamic color instructions.
[0054] The aforementioned physical origin context metadata can be a structured data set recorded synchronously during the acquisition of multiple frames of images, describing the physical conditions and environment upon which the generation of the multiple frames of images depended. It is understood that the physical origin context metadata can correspond precisely to each frame of the acquired image.
[0055] The aforementioned color instruction information can be generated and issued by the system, and is used to control the screen display unit to output different colored light in a specific sequence, for a specific duration, and with a specific intensity. For example, it can be defined as "outputting a color block with an RGB value of (255,0,0) within the time interval t0 to t1".
[0056] The aforementioned data acquisition device information may be the hardware identifier and imaging parameters of the terminal device used to perform data acquisition, including but not limited to the device brand, model, image signal processor (ISP) version, and camera parameters used at the moment of acquisition, such as ISO sensitivity, shutter speed, white balance, and focal length.
[0057] The aforementioned attack information can be parameters used to describe the type and attributes of an attack method when the collected object is an attack sample. For example, the attack type can be refined to "printing attack" and further include information such as printing material (e.g., coated paper) and printing resolution; or "screen attack" and include the type of screen used (e.g., OLED).
[0058] The aforementioned environmental information can be background lighting conditions parameters in the scene being collected, independent of the system's active lighting, and may include ambient light source type (such as sunlight, fluorescent lamp), light intensity, and color temperature.
[0059] The aforementioned preset text template can be a predefined text structure framework. The preset text template can be used to convert structured metadata fields into standardized natural language descriptions according to fixed syntax and format.
[0060] The above text conversion can be a calculation process that fills the specific values in the metadata into the corresponding variable placeholders according to the preset text template, thereby generating a standardized text description.
[0061] It should be noted that the above-mentioned acquisition of multiple frames of images of the subject can be to acquire at least one video stream of the subject, and then perform image processing on at least one video stream to obtain multiple frames of images.
[0062] In this embodiment, by acquiring multiple frames of captured images of the object being collected and multiple physical cause context metadata corresponding one-to-one with the multiple frames of captured images, the acquisition information corresponding to the multiple frames of captured images is different. The acquisition information includes at least one of color instruction information, acquisition device information, attack information, and environmental information. The training dataset is generated by converting the multiple physical cause context metadata using a preset text template. This allows for the systematic acquisition of training data corresponding to multiple different dimensions such as device, environment, attack type, and color instruction. As a result, the model trained on this training dataset can learn universal physical laws that are not limited by specific devices or environments, thereby improving the generalization ability and recognition accuracy of the physical consistency verification model.
[0063] Furthermore, by using preset text templates for text conversion, the automatic generation of structured metadata into standardized text descriptions can be achieved, thereby significantly improving the efficiency of training dataset construction and reducing labor costs.
[0064] Optionally, the physical cause context metadata includes at least one of device fingerprint metadata, lighting environment metadata, and target subject metadata. The device fingerprint metadata includes at least one of the acquisition device model and camera parameters. The lighting environment metadata includes at least one of the lighting color information and ambient light information applied by the acquisition device based on the dynamic color command. The target subject metadata includes at least one of the category of the acquired object and the attack type corresponding to the acquired object. The attack type includes printing attacks, screen photo attacks, and video playback attacks.
[0065] Specifically, the aforementioned device fingerprint metadata can be a set of parameters used to uniquely identify or characterize the hardware attributes and imaging characteristics of the data acquisition device itself. The aforementioned device fingerprint metadata can be used to capture the systematic differences in imaging styles among different devices.
[0066] The aforementioned lighting environment metadata can be a set of parameters describing all light source information acting on the captured object at the moment of shooting, including the lighting actively applied by the system and the passive lighting of the environment.
[0067] The aforementioned target subject metadata can be a set of feature parameters describing the attributes and status of the collected object, used to distinguish the object's authenticity from its specific category.
[0068] The aforementioned data acquisition device model can be an identifier used to indicate the industrial design model of the data acquisition device.
[0069] The aforementioned camera parameters may refer to the specific technical parameters used by the device's image signal processor and camera module at the moment of image acquisition.
[0070] The aforementioned light color information can refer to the color attributes of the light actively projected onto the object being collected by the acquisition device through its screen display unit according to dynamic color instructions. It is usually quantified and recorded in terms of color name or RGB value.
[0071] The aforementioned ambient light information can refer to pre-existing background lighting condition parameters in the scene being collected, which are not actively controlled by the system. These parameters typically include the type of light source, ambient light intensity, and color temperature.
[0072] The category of the collected objects can be the highest-level classification of the nature of the collected objects, usually "real person" or "attack".
[0073] The above attack types can be specific technical subdivisions of the forgery method when the category of the collected object is "attack".
[0074] The aforementioned printing attack can refer to the means of attacking using facial images output from physical printing media, and can be further classified according to the printing material (such as photo paper, ordinary A4 paper) and printing technology (such as inkjet printing, laser printing).
[0075] The aforementioned screen photo attack can refer to the means of attacking by displaying static facial images on electronic display devices (such as mobile phones, tablets, and monitors), and can be classified according to screen technology (such as LCD and OLED).
[0076] The aforementioned video playback attack can refer to a method of attacking by playing a sequence of dynamic facial videos on an electronic display device.
[0077] The following is an example illustrating a specific embodiment: Step 1: Define a diverse sample collection space: To ensure the generalization ability and robustness of the subsequently trained model, a multi-dimensional sample collection space is first defined, which should at least cover: Device space: includes various brands and models of mobile devices with image acquisition capabilities, such as, but not limited to, different series of smartphones and tablets.
[0078] Attack Space: Includes various typical and atypical forgery attack methods. These attack methods include, but are not limited to: print attacks using printouts made of different materials (such as A4 plain paper, coated paper, and matte paper) for photocopying; screen attacks using screens of different display technologies (such as LCD and OLED) and devices (such as mobile phones, computers, and tablets) for photocopying; and physical attacks using artificial masks (such as silicone masks).
[0079] Environmental space: includes a variety of typical lighting environments, such as sunlight, office fluorescent lights, incandescent lights, etc.
[0080] Step 2: Perform synchronous data acquisition operation: Using a data acquisition software development kit (SDK), data acquisition operations are performed on a selected acquisition device within the sample acquisition space. These operations include: The image sensor (i.e., camera) of the acquisition device is invoked through the SDK.
[0081] On the display screen of the acquisition device, dynamic color instructions are executed, such as displaying light blocks of different colors in a preset sequence, so as to actively apply dynamically changing lighting conditions to the object being acquired.
[0082] While executing the dynamic color command, a video stream of the object being captured is recorded, and this video stream constitutes the visual evidence in this embodiment.
[0083] Step 3: Obtain and synchronize the physical causation context: While collecting the visual evidence, the SDK's built-in functional modules simultaneously acquire the physical causal context precisely corresponding to each frame of visual evidence. This step is crucial for constructing the dataset required for this embodiment. The metadata is systematically organized into the following categories: Device fingerprint metadata: Equipment Model: The specific brand and model of the data acquisition device.
[0084] Camera parameters: The specific parameters used by the image signal processor (ISP) when acquiring this frame of image, including but not limited to focal length, ISO sensitivity, shutter speed, white balance setting, etc.
[0085] Lighting environment metadata: Active lighting information: The current lighting color applied by the dynamic color command (e.g., Red, Green, Blue).
[0086] Passive ambient light information: Background ambient light information, such as color temperature and intensity, is obtained by analyzing image frames without active lighting through light sensors or by analyzing the images.
[0087] Target entity metadata: Category tags: Labels used to indicate the authenticity of the collected objects, such as "Live" or "Attack".
[0088] Attack type subdivision: If the category label is "attack", then further detailed annotations of the specific attack methods are provided (e.g., Attack_Print_A4, Attack_Screen_Playback).
[0089] Step 4: Constructing paradigmatic cross-modal data pairs: Each valid image (i.e. visual evidence) acquired in the aforementioned steps is precisely paired with its physical causal context, which is fully recorded at the same time.
[0090] Subsequently, using a preset text template, the structured metadata is deterministically converted into a text description. For example, text in the following format can be generated: "The subject is 'real person', the current challenge color is 'red', the device model is '17 pro', the camera focal length is '2.1mm', the light intensity is 'high', and the illumination angle is 'front'." Finally, visual evidence is combined with textual descriptions generated from its physical context to form cross-modal data pairs in (image, text description) format, and these data pairs are stored in a database for training the initial model of this application.
[0091] In this implementation, by combining device fingerprint metadata, lighting environment metadata, and target subject metadata, the model can perform cross-dimensional physical consistency verification, achieve multi-factor cross-validation, and improve authentication security.
[0092] Optionally, obtaining multiple frames of base images containing the face to be authenticated and the physical context corresponding to each frame of the base images includes: Receive the authentication request sent by the client, the authentication request being used to request liveness authentication of the face to be authenticated; A dynamic color instruction is generated based on the authentication request. The dynamic color instruction is used to instruct the client to capture the face to be authenticated according to the color sequence indicated by the dynamic color instruction. Send the dynamic color command to the client; The system receives a multi-frame base image containing the face to be authenticated and the physical context corresponding to each frame of the base image, sent by the client.
[0093] Specifically, the aforementioned authentication request can be a network service call instruction initiated by the client to the server, marking the start of an independent authentication session. Its purpose is to request the server to verify the authenticity and biometric activity of the face to be authenticated.
[0094] In this implementation, by having the server generate dynamic color instructions based on the authentication request, it can be ensured that the dynamic color instructions relied upon by each authentication session are unique and one-time use. Attackers cannot pass authentication by pre-recording a real video, thereby improving the anti-attack capability of the authentication process.
[0095] Optionally, the step of generating multiple check pairs based on multiple key images and the physical causal context corresponding to each key image includes: Based on the timestamp indicated by the dynamic color instruction, determine the optical reflection window corresponding to each color of illumination; Based on the optical reflection window, multiple first images are selected from the multiple base images; Using the first image in the multi-frame first image as a reference, facial key point detection and affine transformation techniques are used to spatially align the face region in the images other than the first image in the multi-frame first image to obtain multi-frame key images. Multiple check pairs are generated based on the multiple key images and the physical cause context corresponding to each key image.
[0096] Specifically, the timestamp indicated by the aforementioned dynamic color instruction can be a precise sequence of time points defined in the dynamic color instruction for the start and end of display of each specific color illumination.
[0097] The aforementioned optical reflection window can be a specific time period defined for each color of illumination on the timeline of multiple base images. It is understood that the setting of the aforementioned optical reflection window is intended to capture the range in which the optical reflection effect of the color light is most significant and stable after it has been stably projected onto the surface of the object being sampled, typically excluding the transitional state at the beginning of color switching and the unstable state before the end.
[0098] The first image mentioned above may be one or more representative static image frames extracted from the video stream within the optical reflection window. The first image may contain the clearest and most stable optical reflection features of the acquired object under specific color lighting.
[0099] The above filtering can be a process of selecting the optimal frame from all base images within an optical reflection window based on preset quality standards. For example, filtering standards may include image sharpness (avoiding motion blur), facial integrity (ensuring that key facial areas are not obscured), and lighting stability.
[0100] The aforementioned first image can refer to the first image selected in chronological order. This frame is chosen as the geometric reference for subsequent spatial alignment operations.
[0101] The aforementioned facial landmark detection refers to the technique of locating a series of predefined, anatomically significant feature points in a first image containing a human face using computer vision algorithms.
[0102] The aforementioned affine transformation technique can be a linear geometric transformation method used to achieve image translation, rotation, scaling, and cropping. Specifically, it can be used to map the face region in a subsequent first image to a position and pose that is pixel-aligned with the face region in the first first image, based on the correspondence of facial key points.
[0103] The aforementioned spatial alignment can be achieved by eliminating scale, position, and rotation differences between keyframes in the sequence caused by the free movement of the head of the captured object through the aforementioned geometric transformation.
[0104] The aforementioned key images can be a series of keyframes with a unified geometric coordinate system obtained after spatial alignment processing.
[0105] In this embodiment, by using the first image in the multi-frame first image as a reference, facial key point detection and affine transformation techniques are employed to spatially align the face regions in all images except the first image in the multi-frame first image. This eliminates interference from irrelevant geometric variables, ensuring that the differences in optical reflection features compared during consistency verification primarily stem from the interaction between the material properties of the acquired object and physical lighting conditions, rather than pose changes. This significantly reduces decision noise in the model, allowing it to focus more on essential features related to liveness detection, thereby improving authentication accuracy. Furthermore, by defining an optical reflection window and selecting keyframes, it is ensured that each selected frame most effectively represents the optical response under continuous and stable illumination of a specific color light. This avoids anomalous signals introduced by misjudging frame capture timing, ensuring the physical continuity and consistency of visual evidence across color sequences.
[0106] For example, the present application will be described below with a specific embodiment, the overall logic of which follows the following three core stages: Phase 1: Cloud-based - Generation and distribution of dynamic color commands: The goal of this phase is to create one-off, unpredictable, and secure verification tasks.
[0107] (1) Receiving requests: The cloud-side server receives authentication requests from clients (such as Apps).
[0108] (2) Dynamic color command: The server generates a verification command containing a random color sequence and time interval in real time in the background. For example, the command is: "Display red at 0 seconds, green at 1.5 seconds, red at 3.0 seconds...". This randomness is the key to defending against replay attacks and injection attacks.
[0109] (3) Encryption and distribution: In order to ensure transmission security and prevent the instruction from being stolen or tampered with during transmission, the server encrypts the dynamic color instruction and distributes it to the mobile device that initiated the request.
[0110] Phase Two: End-side - Instruction Execution and Encapsulation of Multidimensional Evidence The goal of this stage is to execute cloud-based instructions and comprehensively collect all relevant evidence in the process of completing the task.
[0111] (1) Decryption and preparation: The terminal SDK (mobile device) receives and decrypts the instructions from the cloud, then starts the front camera and adjusts the screen to the specified brightness.
[0112] (2) Execution of instructions and recording of visual evidence: The SDK strictly follows the color sequence and timestamp in the instructions to display the specified color blocks in full screen in sequence, uses screen light to illuminate the face, and simultaneously records a video of the face containing the entire light-changing process. This is the core visual evidence.
[0113] (3) Physical context collection: While recording the video, the client-side SDK will actively collect two types of auxiliary physical environment evidence: Environmental information: Analyzes and records the current ambient lighting conditions (such as color temperature and brightness) through the sensor API.
[0114] Device Information: Retrieves the specific model information of the current mobile device.
[0115] (4) Evidence Packaging and Uploading: The recorded visual evidence (video file) and the collected context metadata are packaged together and uploaded to the cloud server for verification.
[0116] Phase 3: Cloud-based Evidence Analysis and Multi-factor Unified Authentication This stage is the core of the entire authentication process, and its goal is to conduct in-depth analysis and make a final ruling on the evidence uploaded from the client side.
[0117] (1) Spatiotemporal standardization of visual evidence: Before entering the core model authentication, the system first calls a key visual evidence spatiotemporal standardization module. This module aims to proactively eliminate irrelevant variables brought about by user behavior and environment, and to "separate" and "highlight" the core visual features that reflect authenticity caused by specific physical lighting conditions with high fidelity.
[0118] The processing flow is as follows: Temporal standardization: First, the module defines a "stable optical reflection window" for each color of illumination in the video stream based on the instruction timestamp. Then, through quality screening algorithms such as sharpness assessment, the "keyframes" that best represent the effect of that illumination are accurately selected within this window.
[0119] Spatial Dimension Normalization: Subsequently, using the first keyframe as a baseline, the module employs a specially optimized (e.g., enhanced training under multi-color lighting) facial landmark detection algorithm and affine transformation technique to precisely align all subsequent keyframes spatially. This step aims to eliminate geometric noise such as user head movement and pose fine-tuning.
[0120] The final output of this module is a set of standardized image sequences with completely unified spatiotemporal references, laying a solid foundation for subsequent accurate authentication.
[0121] (2) Dynamically generate physical causal context: This is a crucial step in the logical chain. Instead of using static or generic text labels, the server dynamically generates a physical cause context for each frame based on all known information for this task.
[0122] (3) Consistency check: A pre-trained physical world consistency verification model is used to calculate a consistency score for each group (keyframe, physical causal context). A veto system is employed: authentication is only passed when the consistency scores of all keyframes in the sequence are above a threshold.
[0123] Compared with related technologies, this embodiment has the following technical advantages: 1. A significant leap forward in security: Existing technologies only determine whether color changes, making them susceptible to being fooled by pre-recorded videos. This application not only verifies color but also the unique optical reflection characteristics that color should exhibit when projected onto real skin under specific devices and environments. In other words, it verifies the authenticity of the physical event resulting from the combined effects of "color-lighting-material-device." Even if an attacker knows the color sequence, it is extremely difficult to simulate the unique optical reflection with living characteristics that should occur under specific devices and environments. This effectively defends against advanced injection attacks and addresses the aforementioned shortcoming of related technologies.
[0124] 2. Detection accuracy and generalization ability are significantly improved: This application adopts a unified end-to-end multimodal large model, which learns deeper physical relationships and replaces the multiple separate small models used for color verification and liveness detection in related technologies. This holistic learning approach makes feature extraction more thorough and decision boundaries clearer, thereby comprehensively improving detection accuracy and solving the second shortcoming of the aforementioned related technologies.
[0125] 3. Exhibits unprecedented robustness to environmental and equipment variations: This application uses variables such as phone model and ambient light as known input information (physical context) for the model, rather than as interference terms. During the training phase, the model has already learned the imaging differences under different devices and lighting conditions. Related algorithms suffer from accuracy drops due to differences in color rendering on different phone screens and camera imaging styles. This application treats these differences as known information, allowing the model to make targeted judgments. For example, the model knows that "red taken with a phone of brand one" and "red taken with a phone of brand two" are inherently different in image data, thus making the correct judgment and solving the accuracy problem caused by device and environmental factors mentioned in point three of the shortcomings of related technologies.
[0126] 4. Enhanced security strategy concealment: Even if an attacker observes the order of color changes on the screen in some way, they cannot know which deeper features (physical context) the cloud model is verifying at the same time. This asymmetry in the verification logic makes the attack target ambiguous and unpredictable, greatly increasing the difficulty of cracking the entire scheme.
[0127] It's important to note that most related technologies focus on overcoming device and environmental differences by improving algorithms (such as multimodal and multi-scale feature fusion) or upgrading hardware, aiming for a universal model. They treat different device imaging styles and ambient light variations as noise to be eliminated or obstacles to overcome. However, the core of this application does not attempt to overcome these differences, but rather treats them as valuable information. The model actively learns and utilizes information about the current device and environment to make more accurate judgments. This paradigm shift of transforming noise into information is unprecedented in related technologies.
[0128] See Figure 2 , Figure 2 This is a schematic diagram of the structure of a face liveness authentication device provided in an embodiment of this application, as shown below. Figure 2 As shown, the face liveness authentication device 200 includes: The first acquisition module 201 is used to acquire multiple frames of basic images containing the face to be authenticated and the physical cause context corresponding to each frame of the basic images. Each frame of the basic images is obtained by the client capturing the face to be authenticated based on the color sequence indicated by the dynamic color instruction. The physical cause context includes environmental information and device information corresponding to each frame of the basic images. The generation module 202 is used to generate multiple check pairs based on multiple key images and the physical cause context corresponding to each key image, wherein the multiple base images include the multiple key images. Verification module 203 is used to perform physical consistency verification based on all the verification pairs to obtain the physical consistency score corresponding to each frame of the key image. The determination module 204 is used to determine that the face to be authenticated has passed the liveness authentication if all of the physical consistency scores are greater than or equal to a preset threshold.
[0129] Optionally, the device further includes: The second acquisition module is used to acquire a training dataset, which contains N sample data pairs. Each sample data pair includes a face image sample and its corresponding context sample information. The N sample data pairs include one positive sample data pair and N-1 negative sample data pairs, where N is an integer greater than 1. A construction module is used to construct an initial model and a target loss function. The initial model includes an image encoder and a text encoder. The image encoder is used to encode the face image samples, and the text encoder is used to encode the context samples. The target loss function is used to characterize the negative logarithm of the ratio of a first similarity to a second similarity. The first similarity is the similarity of the positive sample data pairs, and the second similarity is the sum of the similarities of the N sample data pairs. The training module is used to train the initial model based on the target loss function to obtain a pre-trained physical consistency verification model, which is used for physical consistency verification.
[0130] Optionally, the second acquisition module includes: The acquisition unit is used to acquire multiple frames of acquired images of the acquired object and multiple physical cause context metadata corresponding to the multiple frames of acquired images. The acquisition information corresponding to the multiple frames of acquired images is different. The acquisition information includes at least one of color instruction information, acquisition device information, attack information, and environmental information. The conversion unit is used to convert the multiple physical cause context metadata into text using a preset text template, so as to obtain multiple context sample information that corresponds one-to-one with the multiple physical cause context metadata. The generation unit is used to generate the training dataset based on the multi-frame acquired images and the multiple context sample information.
[0131] Optionally, the first acquisition module includes: A receiving unit is configured to receive an authentication request sent by the client, the authentication request being used to request liveness authentication of the face to be authenticated; The second generation unit is used to generate a dynamic color instruction based on the authentication request. The dynamic color instruction is used to instruct the client to take a picture of the face to be authenticated according to the color sequence indicated by the dynamic color instruction. The sending unit is used to send the dynamic color command to the client; The receiving unit is used to receive multiple frames of basic images containing the face to be authenticated and the physical cause context corresponding to each frame of the basic images sent by the client.
[0132] Optionally, the generation module includes: The determining unit is used to determine the optical reflection window corresponding to each color of illumination based on the timestamp indicated by the dynamic color instruction; A filtering unit is used to filter out multiple first images from the multiple base images based on the optical reflection window; The alignment unit is used to spatially align the face regions of the images other than the first image in the multi-frame first images with the first image in the multi-frame first images as a reference, and to obtain multi-frame key images by using facial key point detection and affine transformation technology. The first generation unit is used to generate multiple check pairs based on the multiple key images and the physical cause context corresponding to each key image.
[0133] It should be noted that the face liveness authentication device provided in this application embodiment is a device capable of executing the above-described face liveness authentication method. Therefore, all implementation methods in the above-described face liveness authentication method embodiments are applicable to this device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.
[0134] For details, see Figure 3 As shown in the figure, this application embodiment also provides an electronic device, including a bus 301, a transceiver 302, an antenna 303, a bus interface 304, a processor 305, and a memory 306.
[0135] Transceiver 302 is used to: acquire multiple frames of basic images containing the face to be authenticated and the physical cause context corresponding to each frame of the basic images, wherein each frame of the basic images is obtained by the client capturing the face to be authenticated based on a color sequence indicated by a dynamic color instruction, and the physical cause context includes environmental information and device information corresponding to each frame of the basic images; Furthermore, processor 305 is used for: Multiple check pairs are generated based on multiple key images and the physical cause context corresponding to each key image in each frame. The multiple base images include the multiple key images. Physical consistency verification is performed based on all the aforementioned check pairs to obtain the physical consistency score corresponding to each frame of the key image. If all of the physical consistency scores are greater than or equal to a preset threshold, the face to be authenticated is determined to have passed the liveness authentication.
[0136] exist Figure 3In this context, a bus architecture (represented by bus 301) is used. Bus 301 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 305 and memory represented by memory 306. Bus 301 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 304 provides an interface between bus 301 and transceiver 302. Transceiver 302 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 305 is transmitted over a wireless medium via antenna 303, which further receives data and transmits it to processor 305.
[0137] Processor 305 manages bus 301 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 306 can be used to store data used by processor 305 during operation.
[0138] Alternatively, the processor 305 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).
[0139] Optionally, the processor 305 is specifically used for: Obtain a training dataset containing N sample data pairs. Each sample data pair includes a face image sample and its corresponding context sample information. The N sample data pairs include one positive sample data pair and N-1 negative sample data pairs, where N is an integer greater than 1. An initial model and a target loss function are constructed. The initial model includes an image encoder and a text encoder. The image encoder is used to encode the face image samples, and the text encoder is used to encode the context samples. The target loss function is used to characterize the negative logarithm of the ratio of a first similarity to a second similarity. The first similarity is the similarity of the positive sample data pairs, and the second similarity is the sum of the similarities of the N sample data pairs. Based on the objective loss function, the initial model is trained to obtain a pre-trained physical consistency verification model, which is used for physical consistency verification.
[0140] Optionally, the processor 305 is specifically used for: Acquire multiple frames of captured images of the object being captured, as well as multiple physical cause context metadata corresponding one-to-one with the multiple frames of captured images. The capture information corresponding to the multiple frames of captured images is different, and the capture information includes at least one of color instruction information, capture device information, attack information, and environmental information. The multiple physical cause context metadata are converted into text using a preset text template to obtain multiple context sample information that corresponds one-to-one with the multiple physical cause context metadata. The training dataset is generated based on the multi-frame acquired images and the multiple context sample information.
[0141] Optionally, the physical cause context metadata includes at least one of device fingerprint metadata, lighting environment metadata, and target subject metadata. The device fingerprint metadata includes at least one of the acquisition device model and camera parameters. The lighting environment metadata includes at least one of the lighting color information and ambient light information applied by the acquisition device based on the dynamic color command. The target subject metadata includes at least one of the category of the acquired object and the attack type corresponding to the acquired object. The attack type includes printing attacks, screen photo attacks, and video playback attacks.
[0142] Optionally, the processor 305 is specifically used for: Receive the authentication request sent by the client, the authentication request being used to request liveness authentication of the face to be authenticated; A dynamic color instruction is generated based on the authentication request. The dynamic color instruction is used to instruct the client to capture the face to be authenticated according to the color sequence indicated by the dynamic color instruction. Send the dynamic color command to the client; The system receives a multi-frame base image containing the face to be authenticated and the physical context corresponding to each frame of the base image, sent by the client.
[0143] Optionally, the processor 305 is specifically used for: Based on the timestamp indicated by the dynamic color instruction, determine the optical reflection window corresponding to each color of illumination; Based on the optical reflection window, multiple first images are selected from the multiple base images; Using the first image in the multi-frame first image as a reference, facial key point detection and affine transformation techniques are used to spatially align the face region in the images other than the first image in the multi-frame first image to obtain multi-frame key images. Multiple check pairs are generated based on the multiple key images and the physical cause context corresponding to each key image.
[0144] It should be noted that the electronic device provided in this application embodiment is a device capable of executing the above-described face liveness authentication method. Therefore, all implementation methods in the above-described face liveness authentication method embodiments are applicable to this electronic device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.
[0145] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described face liveness authentication method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0146] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described face liveness authentication method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0147] This application also provides a computer program product, including computer instructions. When executed by a processor, the computer instructions implement the various processes of the above-described face liveness authentication method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0148] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0149] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0150] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for facial liveness authentication, characterized in that, include: The system acquires multiple base images containing the face to be authenticated and the physical context corresponding to each base image. Each base image is captured by the client based on a color sequence indicated by a dynamic color instruction. The physical context includes environmental information and device information corresponding to each base image. Multiple check pairs are generated based on multiple key images and the physical cause context corresponding to each key image. The multiple base images include the key images. Physical consistency verification is performed based on all the aforementioned check pairs to obtain the physical consistency score corresponding to each frame of the key image. If all of the physical consistency scores are greater than or equal to a preset threshold, the face to be authenticated is determined to have passed the liveness authentication.
2. The method according to claim 1, characterized in that, Before the step of obtaining multiple base images containing the face to be authenticated and the physical causal context corresponding to each base image, the method further includes: Obtain a training dataset containing N sample data pairs. Each sample data pair includes a face image sample and its corresponding context sample information. The N sample data pairs include one positive sample data pair and N-1 negative sample data pairs, where N is an integer greater than 1. An initial model and a target loss function are constructed. The initial model includes an image encoder and a text encoder. The image encoder is used to encode the face image samples, and the text encoder is used to encode the context samples. The target loss function is used to characterize the negative logarithm of the ratio of a first similarity to a second similarity. The first similarity is the similarity of the positive sample data pairs, and the second similarity is the sum of the similarities of the N sample data pairs. Based on the objective loss function, the initial model is trained to obtain a pre-trained physical consistency verification model, which is used for physical consistency verification.
3. The method according to claim 2, characterized in that, The acquisition of the training dataset includes: Acquire multiple frames of captured images of the object being captured, as well as multiple physical cause context metadata corresponding one-to-one with the multiple frames of captured images. The capture information corresponding to the multiple frames of captured images is different, and the capture information includes at least one of color instruction information, capture device information, attack information, and environmental information. The multiple physical cause context metadata are converted into text using a preset text template to obtain multiple context sample information that corresponds one-to-one with the multiple physical cause context metadata. The training dataset is generated based on the multi-frame acquired images and the multiple context sample information.
4. The method according to claim 3, characterized in that, The physical cause context metadata includes at least one of device fingerprint metadata, lighting environment metadata, and target subject metadata. The device fingerprint metadata includes at least one of the acquisition device model and camera parameters. The lighting environment metadata includes at least one of the lighting color information and ambient light information applied by the acquisition device based on the dynamic color command. The target subject metadata includes at least one of the category of the acquired object and the attack type corresponding to the acquired object. The attack type includes printing attacks, screen photo attacks, and video playback attacks.
5. The method according to any one of claims 1 to 4, characterized in that, The step of obtaining multiple frames of base images containing the face to be authenticated and the physical context corresponding to each frame of the base images includes: Receive the authentication request sent by the client, the authentication request being used to request liveness authentication of the face to be authenticated; A dynamic color instruction is generated based on the authentication request. The dynamic color instruction is used to instruct the client to capture the face to be authenticated according to the color sequence indicated by the dynamic color instruction. Send the dynamic color command to the client; The system receives a multi-frame base image containing the face to be authenticated and the physical context corresponding to each frame of the base image sent by the client.
6. The method according to any one of claims 1 to 4, characterized in that, The generation of multiple check pairs based on multiple key images and the physical causal context corresponding to each key image includes: Based on the timestamp indicated by the dynamic color instruction, determine the optical reflection window corresponding to each color of illumination; Based on the optical reflection window, multiple first images are selected from the multiple base images; Using the first image in the multi-frame first image as a reference, facial key point detection and affine transformation techniques are used to spatially align the face region in the images other than the first image in the multi-frame first image to obtain multi-frame key images. Multiple check pairs are generated based on the multiple key images and the physical cause context corresponding to each key image.
7. A facial liveness authentication device, characterized in that, include: The first acquisition module is used to acquire multiple frames of basic images containing the face to be authenticated and the physical cause context corresponding to each frame of the basic images. Each frame of the basic images is obtained by the client capturing the face to be authenticated based on the color sequence indicated by the dynamic color instruction. The physical cause context includes environmental information and device information corresponding to each frame of the basic images. A generation module is used to generate multiple check pairs based on multiple key images and the physical cause context corresponding to each key image, wherein the multiple base images include the multiple key images. The verification module is used to perform physical consistency verification based on all the verification pairs to obtain the physical consistency score corresponding to each frame of the key image. The determination module is used to determine that the face to be authenticated has passed the liveness authentication if all of the physical consistency scores are greater than or equal to a preset threshold.
8. An electronic device, characterized in that, Including transceivers and processors, The transceiver is used to: acquire multiple frames of basic images containing the face to be authenticated and the physical cause context corresponding to each frame of the basic images. Each frame of the basic images is obtained by the client capturing the face to be authenticated based on a color sequence indicated by a dynamic color instruction. The physical cause context includes environmental information and device information corresponding to each frame of the basic images. The processor is used for: Multiple check pairs are generated based on multiple key images and the physical cause context corresponding to each key image in each frame. The multiple base images include the multiple key images. Physical consistency verification is performed based on all the aforementioned check pairs to obtain the physical consistency score corresponding to each frame of the key image. If all of the physical consistency scores are greater than or equal to a preset threshold, the face to be authenticated is determined to have passed the liveness authentication.
9. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the face liveness authentication method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the face liveness authentication method as described in any one of claims 1 to 6.
11. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the face liveness authentication method as described in any one of claims 1 to 6.