Liveness detection method, device, storage medium and electronic device

Through artificial intelligence image generation technology and target scene property description text, the inefficiency problem of liveness detection models in cross-scene adaptation is solved, and efficient liveness detection model training and adaptation are achieved.

CN116798129BActive Publication Date: 2025-09-12ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310604708.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-09-12
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

Existing technologies require retraining of liveness detection models when adapting across scenarios, which results in the need to recollect sample images and manually label them, resulting in low efficiency.

Method used

Through artificial intelligence image generation technology, a large number of sample object detection images are generated based on the target scene property description text, and the initial target liveness detection model is trained in combination with the reference liveness detection model to reduce the manual labeling and training time for new scenes.

Benefits of technology

It achieves efficient liveness detection in new scenarios, improves model adaptation efficiency, reduces dependence on manual labeling and long-term training, and improves liveness detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116798129B_ABST
    Figure CN116798129B_ABST
Patent Text Reader

Abstract

This specification discloses a liveness detection method, device, storage medium and electronic device, wherein the method includes: determining a target scene property description text based on a first sample object detection image in a target scene, performing artificial intelligence image generation based on the target scene property description text and the first sample object detection image to obtain multiple second sample object detection images, creating an initial target liveness detection model for the target scene based on a reference liveness detection model, and using the second sample object detection image to perform model training on the initial target liveness detection model to obtain a target liveness detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a liveness detection method, device, storage medium, and electronic device. Background Art

[0002] With the rapid development of computer technology, biometrics are widely used in our daily lives and production. For example, facial recognition for payment, facial recognition for access control, facial recognition for attendance, and facial recognition for station entry all rely on biometrics. However, with the increasing application of biometrics, the need for liveness detection in these scenarios has become increasingly prominent. While biometrics provides convenience, it also brings new risks and challenges. The most common threat to the security of biometric systems is liveness attacks, which attempt to circumvent image biometric verification by using methods such as device screens and printed photos. Therefore, liveness detection is particularly important in biometric scenarios. Summary of the Invention

[0003] This specification provides a liveness detection method, device, storage medium, and electronic device. The technical solution is as follows:

[0004] In a first aspect, this specification provides a method for detecting a living body, the method comprising:

[0005] Acquire first sample object detection images in a target scene, and determine target scene property description text based on each of the first sample object detection images;

[0006] performing artificial intelligence image generation based on the target scene property description text and the first sample object detection image to obtain a plurality of second sample object detection images, wherein the number of samples of the second sample object detection images is greater than the number of samples of the first sample object detection images;

[0007] An initial target liveness detection model for the target scene is created based on the reference liveness detection model, and the initial target liveness detection model is trained using the second sample object detection image to obtain a target liveness detection model. The reference liveness detection model is a liveness detection model under the reference scene.

[0008] In a second aspect, this specification provides a liveness detection device, comprising:

[0009] a data processing module, configured to obtain first sample object detection images in a target scene, and determine a target scene property description text based on each of the first sample object detection images;

[0010] an image generation module, configured to perform artificial intelligence image generation based on the target scene property description text and the first sample object detection image to obtain a plurality of second sample object detection images, wherein the number of samples of the second sample object detection images is greater than the number of samples of the first sample object detection images;

[0011] A model training module is used to create an initial target liveness detection model for the target scene based on the reference liveness detection model, and use the second sample object detection image to train the initial target liveness detection model to obtain a target liveness detection model. The reference liveness detection model is the liveness detection model under the reference scene.

[0012] In a third aspect, the present specification provides a computer storage medium storing at least one instruction, wherein the instruction is suitable for being loaded by a processor and executing the method steps of one or more embodiments of the present specification.

[0013] In a fourth aspect, the present specification provides a computer program product, wherein the computer program product stores at least one instruction, wherein the instruction is suitable for being loaded by a processor and executing the method steps of one or more embodiments of the present specification.

[0014] In a fifth aspect, this specification provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps of one or more embodiments of this specification.

[0015] The beneficial effects of the technical solutions provided by some embodiments of this specification include at least:

[0016] In one or more embodiments of the present specification, a target scene property description text is determined based on a first sample object detection image in a target scene, artificial intelligence image generation is performed in combination with the target scene property description text and the first sample object detection image to obtain multiple second sample object detection images, an initial target liveness detection model for the target scene is created based on a reference liveness detection model, and the initial target liveness detection model is trained using the second sample object detection image to obtain a target liveness detection model adapted to the new scene. In the new target scene, relatively good liveness detection performance can be achieved, and model training does not need to rely on a large amount of manual labeling and a long model training time, thereby improving the liveness detection efficiency and the model adaptation efficiency in the new scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is a scenario diagram of a liveness detection system provided in this specification;

[0019] Figure 2 This is a flow chart of a liveness detection method provided in this specification;

[0020] Figure 3 This is a schematic diagram of another liveness detection method provided in this manual;

[0021] Figure 4 This is a flow chart of generating a scenario property description provided in this specification;

[0022] Figure 5 This is a model training diagram of a scene property description generation model provided in this specification;

[0023] Figure 6 This is a model training diagram for the scene live sample generation model provided in this manual;

[0024] Figure 7 This is a diagram of the training process of a target liveness detection model provided in this manual;

[0025] Figure 8 This is a schematic diagram of a liveness detection device provided in this specification;

[0026] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this manual;

[0027] Figure 10 This is a schematic diagram of the structure of the operating system and user space provided in this manual;

[0028] Figure 11 yes Figure 10 The architecture diagram of the Android operating system;

[0029] Figure 12 yes Figure 10 Architecture diagram of the IOS operating system. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in this specification in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments in this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0031] In the description of this specification, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood according to the specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0032] In related technologies, various liveness detection methods have been proposed to detect liveness attacks. These methods often train liveness detection models based on reference deployment scenarios. However, when liveness detection is applied to new scenarios that differ from the reference deployment scenarios, cross-scenario adaptation issues arise. Liveness detection models trained in the reference deployment scenarios need to be retrained, sample images recollected, and samples relabeled in the new scenarios, which still poses significant challenges.

[0033] The present specification is described in detail below with reference to specific embodiments.

[0034] See Figure 1 , is a scene diagram of a living body detection system provided in this specification. Figure 1 As shown, the living body detection system may include at least a client cluster and a service platform 100 .

[0035] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.

[0036] Each client in the client cluster can be an electronic device with communication capabilities, including but not limited to wearable devices, handheld devices, personal computers, tablet computers, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may be called different names in different networks, such as user equipment, access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic devices in 5G network or future evolution network, etc.

[0037] The service platform 100 can be a separate server device, such as a rack-mounted, blade, tower, or cabinet-mounted server device, or a workstation, mainframe computer, or other hardware device with strong computing capabilities; it can also be a server cluster composed of multiple servers. The servers in the service cluster can be symmetrically composed, wherein each server has equivalent functions and status in the transaction link, and each server can provide services to the outside world independently. The independent service can be understood as not requiring the assistance of other servers.

[0038] In one or more embodiments of the present specification, the service platform 100 may establish a communication connection with at least one client in the client cluster, and based on the communication connection, complete data interaction during the liveness detection process, such as online transaction data interaction. For example, the client may collect a target object image of the target object in the environment and send it to the service platform 100, and the service platform 100 may execute the liveness detection method corresponding to one or more embodiments of the present specification to perform liveness attack detection processing; for example, the service platform 100 may not refer to at least one client in the target scenario, and any client may perform liveness detection processing in the target scenario based on the target liveness detection model;

[0039] It should be noted that the service platform 100 and at least one client in the client cluster establish a communication connection through a network for interactive communication, wherein the network can be a wireless network or a wired network, the wireless network includes but is not limited to a cellular network, a wireless local area network, an infrared network or a Bluetooth network, and the wired network includes but is not limited to Ethernet, a universal serial bus (USB) or a controller area network. In one or more embodiments of the specification, technologies and / or formats including Hypertext Markup Language (HTML) and Extensible Markup Language (XML) are used to represent data exchanged over the network (such as a target compressed package). In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.

[0040] The liveness detection system embodiments provided in this specification share the same concept as the liveness detection method described in one or more embodiments. The liveness detection method described in one or more embodiments may be executed by the service platform 100 described above; the liveness detection method described in one or more embodiments may also be executed by the electronic device corresponding to the client, depending on the actual application environment. The implementation process of the liveness detection system embodiments can be found in the method embodiments described below and will not be further described here.

[0041] based on Figure 1 The scene diagram shown is a detailed introduction to the liveness detection method provided by one or more embodiments of this specification.

[0042] See Figure 2 , provides a flow chart of a liveness detection method for one or more embodiments of this specification. This method can be implemented using a computer program and can be run on a liveness detection device based on a von Neumann architecture. The computer program can be integrated into an application or run as a standalone tool application. The liveness detection device can be an electronic device.

[0043] Specifically, the liveness detection method includes:

[0044] It is understandable that liveness detection is a detection method that determines the true physiological characteristics of an object in some identity verification scenarios. In image liveness detection applications, image liveness detection needs to verify whether the image is a real live object based on the acquired target liveness detection image. Image liveness detection needs to be able to effectively resist common liveness attack methods such as photos, face swaps, masks, occlusions, and screen re-shoots, thereby helping to identify fraudulent behavior and protect user rights.

[0045] S102: Acquire first sample object detection images in a target scene, and determine a target scene property description text based on each of the first sample object detection images;

[0046] In this specification, the image modality type of the first sample object detection image is not limited. The image modality type can be a fit of one or more image modality types such as a video modality type that carries object information, a color picture modality (RGB) type, a small video modality type, an animation modality type, a depth image (depth) modality type, an infrared image (IR) modality type, and a near-infrared modality (NIR) type.

[0047] Illustratively, in an actual target scene, a certain number of first sample object detection images may be collected based on a corresponding living body detection task through, for example, an RGB camera, a monocular camera, an infrared camera, or the like.

[0048] In one or more embodiments of this specification, an innovative artificial intelligence image generation method is used to generate high-quality samples in a new scene, and the new scene is also the target scene. The generated samples are used in the target scene to train the liveness detection model, thereby greatly reducing the dependence on manual labeling of the new scene and improving the efficiency of adaptation.

[0049] Specifically, for a target scene, i.e., a new scene to which the liveness detection model needs to be adapted, a small number of new scene images, i.e., a certain number of first sample object detection images under the target scene, are first collected. These first sample object detection images are then used to obtain a good description text of the target scene properties. The description text of the target scene properties can be regarded as a text of a scene description word (prompt);

[0050] The target scene property description text includes scene common description text and scene characteristic description text. For the same scene, different images of the scene will have common description features, such as outdoor features, low-light features, etc. However, different images of the same scene will also have certain characteristics (characteristics caused by differences in deployment angles and deployment positions), such as special buildings within the scene. When generating the second sample object detection image of the target scene: extract accurate new scene description text (i.e., target scene property description text) from a small number of new scene samples (first sample object detection images);

[0051] S104: performing artificial intelligence image generation based on the target scene property description text and the first sample object detection image to obtain a plurality of second sample object detection images, wherein the number of samples of the second sample object detection images is greater than the number of samples of the first sample object detection images;

[0052] Specifically, after extracting accurate new scene description text (that is, target scene property description text), a large number of similar scene samples (second sample object detection images) can be generated by combining the target scene property description text with a small number of collected new scene samples (first sample object detection images).

[0053] Specifically, an artificial intelligence image generation method, namely AI Generated Content (AIGC), is used to generate a large number of second sample object detection images of different types in a specific target scene. The second sample object detection images are sample object detection images in the target scene, including attack sample types and living sample types.

[0054] AIGC generally refers to content generated using artificial intelligence. In this method, it specifically refers to AI-generated image content.

[0055] S106: Creating an initial target liveness detection model for the target scene based on the reference liveness detection model, and using the second sample object detection image to train the initial target liveness detection model to obtain a target liveness detection model, where the reference liveness detection model is a liveness detection model under the reference scene.

[0056] Specifically, the initial target liveness detection model is created based on an existing or trained reference liveness detection model in a reference scene. Executing the liveness detection method of this specification implements a cross-scene liveness detection method.

[0057] Specifically, the generated sample-second sample object detection image initial target liveness detection model is efficiently trained to obtain a target liveness detection model adapted for the new scenario. This effectively applies the liveness detection model from the reference scenario to the new target scenario. This improves the problem of training a liveness detection model from scratch for a new scenario and the limited number of sample images in the new scenario. This allows for relatively good performance in the new target scenario without relying on extensive manual annotation and lengthy model training time, thus improving model adaptation efficiency in the new scenario.

[0058] It should be noted that the initial target liveness detection model and the reference liveness detection model involved in one or more embodiments of this specification rely on the training of a machine learning model, and the machine learning model includes but is not limited to the fitting of one or more machine learning models such as a convolutional neural network (CNN) model, a deep neural network (DNN) model, a recurrent neural network (RNN) model, an embedding model, a gradient boosting decision tree (GBDT) model, and a logistic regression (LR) model.

[0059] In the actual application stage, the trained target liveness detection model can be deployed and corresponding liveness detection decisions can be made. The electronic device can deploy the target liveness detection model to at least one client in the target scenario. When the user of the client performs liveness detection transactions (such as identity authentication, access control, and attendance), the user of the client can collect facial images for identity authentication;

[0060] In a specific implementation scenario, the target liveness detection model is deployed in the target scene, so as to perform liveness detection processing on the target object detection image in the target scene through the target liveness detection model.

[0061] See Figure 3 , is a schematic diagram of a liveness detection scenario provided in the embodiment of this specification. Figure 3 As shown, the client is a device for identity authentication in the target scene, which is equipped with an image acquisition device. When the user approaches the client and is within the image acquisition range of the image acquisition device, the image acquisition device will capture the target object detection image close to the person (such as Figure 3 ), and use the target liveness detection model to perform liveness detection on the collected target object detection image.

[0062] In one or more embodiments of the present specification, a target scene property description text is determined based on a first sample object detection image in a target scene, artificial intelligence image generation is performed in combination with the target scene property description text and the first sample object detection image to obtain multiple second sample object detection images, an initial target liveness detection model for the target scene is created based on a reference liveness detection model, and the initial target liveness detection model is trained using the second sample object detection image to obtain a target liveness detection model adapted to the new scene. In the new target scene, relatively good liveness detection performance can be achieved, and model training does not need to rely on a large amount of manual labeling and a long model training time, thereby improving the liveness detection efficiency and the model adaptation efficiency in the new scene.

[0063] Optional, see Figure 4 , Figure 4 This is a flow chart of generating a scene property description proposed in one or more embodiments of this specification. Specifically:

[0064] S202: Inputting each of the first sample object images into a scene property description generation model, and outputting at least one target scene common description information and at least one target scene characteristic description information;

[0065] This embodiment explains how to accurately extract new scene description text from the first sample object image, namely, target scene common description information and target scene characteristic description information, by using a scene property description generation model based on commonality and characteristic decoupling modeling to extract relevant scene text description.

[0066] The target scene property description text includes the target scene common description text and the target scene characteristic description text.

[0067] For the same scene, different images of the scene will have common descriptive features, such as outdoor features, low-light features, etc. The target scene common description text is used to represent these common descriptive features;

[0068] Different images of the same scene may have certain characteristics (characteristics caused by differences in deployment angles and deployment positions), such as special buildings in the scene. The target scene characteristic description information is used to characterize these characteristic description features.

[0069] Specifically, a scene property description generation model is obtained by creating and training, each first sample object image is input into the scene property description generation model, and at least one target scene common description information and at least one target scene characteristic description information are output.

[0070] It is understandable that, when targeting a new scene, by collecting first sample object images of the new scene, i.e., the target scene, for example, 5-10 first sample object images of each type of client device for liveness detection, assuming that a total of N first sample object images are collected, then N target scene common description information and N target scene characteristic description information can be obtained through the above method;

[0071] S204: Determine a target scene property description text based on the target scene common description information and the target scene characteristic description information.

[0072] Specifically, the target scene common description information and the target scene characteristic description information are described and summarized to obtain a target scene property description text.

[0073] In a feasible implementation, both the target scene common description information and the target scene characteristic description information may be directly used as target scene property description text.

[0074] In a feasible implementation manner, the determining of the target scene property description text based on the target scene common description information and the target scene characteristic description information may be:

[0075] A2: Sampling description text of the common description information of the target scene to obtain common description text of the target scene;

[0076] In a feasible implementation, a data sampling algorithm may be used to sample description texts of all target scene common description information;

[0077] In a feasible implementation, the text frequency corresponding to each common description text can be determined based on the common description information of the target scene, at least one reference common description text can be determined from each common description text based on the text frequency, and description text sampling can be performed on each reference common description text to obtain the common description text of the target scene.

[0078] Optionally, at least one reference common description text is determined from each of the common description texts based on the text frequency. This may be by selecting common description texts with a text frequency greater than a set frequency threshold (e.g., the frequency threshold may be N / 2) as the reference common description text, i.e., selecting common description texts with a higher frequency of occurrence as the reference common description text. This allows for text screening, and the selected reference common description text is more consistent with the common characteristics of a new scene, thereby filtering out error interference.

[0079] A4: Sampling description text of the target scene characteristic description information to obtain target scene characteristic description text;

[0080] In a feasible implementation, a data sampling algorithm may be used to sample description texts of all target scene common description information;

[0081] In a feasible implementation, the text frequency corresponding to each feature description text can be determined based on the target scene feature description information, at least one reference feature description text can be determined from each feature description text based on the text frequency, and description text sampling can be performed on each reference feature description text to obtain the target scene feature description text.

[0082] Optionally, at least one reference feature description text is determined from each feature description text based on the text frequency. This can be achieved by selecting feature description texts with a text frequency greater than a set frequency threshold (e.g., the frequency threshold can be N / 2) as reference feature description texts. In other words, feature description texts with a higher frequency of occurrence are selected as reference feature description texts. This allows for text screening, where the selected reference feature description text is more consistent with the characteristics of a new scene, thereby filtering out errors and interference.

[0083] A6: Obtain a target scene property description text based on the target scene common description text and the target scene characteristic description text.

[0084] Randomly sampling a number of target scene common description texts and a number of target scene characteristic description texts can be used as target scene property description texts used to generate object detection images in new scenes;

[0085] In one or more embodiments of the present specification, in order to improve the reliance on data collection and manual labeling in cross-scene adaptation, a small number of sample object images in a new scene can be collected to determine the scene property description text in the new scene, thereby assisting in the subsequent large-scale generation of sample images in the new scene, so as to achieve the purpose of adapting the liveness detection model to the new scene. The new scene description text can be accurately extracted based on only a small number of sample object images. Subsequently, a large number of sample images in the same scene can be generated using the description text of the new scene to ensure the effect of subsequent model training.

[0086] Optional, see Figure 5 , Figure 5 This is a flow chart of a scene property description generation model proposed in one or more embodiments of this specification.

[0087] S3002: Creating an initial scene property description generation model;

[0088] In one or more embodiments of the present specification, an initial scene property description generation model may be created based on a machine learning model in response to a liveness detection task, and model training may be performed to obtain a scene property description generation model, and the scene property description generation model may be used to generate scene property description text;

[0089] Optionally, the initial scene property description generation model may include a basic feature encoding module, a common text description generation module and a characteristic text description generation module, and the sample scene property description text includes a sample scene common description text and a sample scene characteristic description text;

[0090] S3004: Acquire a sample object detection image corresponding to at least one sample scene, wherein the sample object detection image carries a text label describing the nature of the scene;

[0091] The sample scene is based on the environmental scene configuration used for liveness detection;

[0092] Specifically, data annotation can be used in advance to annotate the description of the collected sample object detection images to obtain text labels describing the nature of the scene; for example, expert services can be introduced for data annotation to describe the characteristics and commonalities of the sample object detection images to obtain text labels describing the nature of the scene.

[0093] Schematically, the scene form description text label includes a scene common description text label and a scene characteristic description text label.

[0094] S3006: Perform at least one round of model training on the initial scene property description generation model using the sample object detection image to obtain a sample scene property description text;

[0095] Illustratively, the sample object detection image can be input into the initial scene property description generation model for at least one round of model training, as follows:

[0096] B2: Inputting the sample object detection image into the initial scene property description generation model, and performing feature extraction on the sample object detection image through the basic feature encoding module to obtain sample image features;

[0097] The input of the basic feature coding module is the sample object detection image. The basic coding module performs feature coding extraction on the sample object detection image to obtain sample image features.

[0098] The sample image features may include one or more fitting of feature types such as texture features, color features, shape features, and spatial relationship features of the object image.

[0099] B4: performing common descriptions on the sample image features by the common text description generation module to obtain common description texts of sample scenes;

[0100] The input of the common text description generation module is the sample image features of the previous step. The common text description generation module analyzes the commonalities in the scene image based on the sample image features, thereby obtaining the sample scene common description text.

[0101] B6: The characteristic text description generation module performs characteristic description on the sample image features to obtain a sample scene characteristic description text.

[0102] The input of the characteristic text description generation module is the sample image features of the previous step. The characteristic text description generation module analyzes the features in the scene image based on the sample image features to obtain the sample scene characteristic description text.

[0103] S3008: Adjusting model parameters of the initial scene property description generation model based on the scene property description text label and the sample scene property description text until the initial scene property description generation model completes model training to obtain a scene property description generation model.

[0104] Schematically, the scene property description text tag includes a scene common description text tag and a scene characteristic description text tag;

[0105] Furthermore, the adjustment process of the model parameters can be as follows:

[0106] C2: determining a common text description generation loss based on the sample scene common description text and the scene common description text label;

[0107] The common text description generation loss serves as a supervisory signal for common description generation in the process of scene property description generation, and supervises the model training effect from the common description dimension. Based on the sample scene common description text and the scene common description text label, the common text description generation loss can be obtained by using a related loss calculation function.

[0108] For example, the common text description generation loss can be calculated using the following loss function:

[0109]

[0110] Among them, LOSS common Generate loss for the common text description, the emb c-pred is the common description text of the sample scene, the emb c-label A text label describing the common characteristics of the scene;

[0111] C4: determining a characteristic text description generation loss based on the sample scene characteristic description text and the scene characteristic description text label;

[0112] The characteristic text description generation loss serves as a supervisory signal for characteristic description generation in the process of scene property description generation. From the perspective of characteristic description dimension, the model training effect is viewed. Based on the sample scene characteristic description text and the scene characteristic description text label, the characteristic text description generation loss can be obtained by using a related loss calculation function.

[0113] For example, the feature text description generation loss can be calculated using the following loss function:

[0114]

[0115] Among them, LOSS specific Generate loss for the feature text description, the emb s-pred The sample scene feature description text, the emb s-label Describing text labels for the scene characteristics;

[0116] C6: Adjust the model parameters of the initial scene property description generation model based on the common text description generation loss and the characteristic text description generation loss.

[0117] It can be understood that based on the common text description generation loss and the characteristic description text generation loss, the model comprehensive loss for the initial scene property description generation model can be obtained. Based on the model comprehensive loss, the model back propagation method is used to adjust the model parameters of the initial scene property description generation model until the model training end conditions are met, and a trained scene property description generation model can be obtained.

[0118] The model training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset number threshold, etc. The specific model training termination conditions may be determined based on actual conditions and are not specifically limited here.

[0119] In one or more embodiments of the present specification, in order to improve the dependence on data collection and manual labeling in cross-scene adaptation, a scene property description generation model is trained. Through this model, when a small number of sample object images in a new scene are collected, the scene property description generation model can determine the scene property description text in the new scene, thereby assisting in the subsequent generation of a large number of sample images in the new scene, so as to achieve the purpose of adapting the liveness detection model to the new scene. It can realize accurate extraction of new scene description text based on only a small number of sample object images, and subsequently a large number of sample images in the same scene can be generated through the description text of the new scene to ensure the effect of subsequent model training.

[0120] Optionally, when the electronic device performs the artificial intelligence image generation based on the target scene property description text and the first sample object detection image to obtain a plurality of second sample object detection images, the method may be:

[0121] The target scene property description text and the first sample object detection image are input into the scene living sample generation model, and a plurality of second sample object detection images are output.

[0122] See Figure 6 , Figure 6 This is a model training diagram of a scene living sample generation model involved in one or more embodiments of this specification.

[0123] S4002: Create an initial scene living sample generation model;

[0124] In one or more embodiments of this specification, an initial scene live sample generation model may be created based on a machine learning model in response to a liveness detection task to perform model training to obtain a scene live sample generation model, and the scene live sample generation model may be used to generate live samples in a new scene.

[0125] Schematically, in actual applications, image generation methods directly based on AIGC often directly input text and use the text to control image generation. However, the generated images are often random and uncontrollable. Therefore, for highly standardized images in liveness detection scenarios, solutions that directly generate images based on AIGC cannot meet the needs. In order to improve this phenomenon, a new scene liveness sample generation method based on material subject locking is introduced in the training stage of the scene liveness sample generation model. This method confirms the image subject by inputting reference materials (various live faces, attack materials, etc.), and uses text to fine-tune the subject. At the same time, the text in the previous step is used to generate the scene, and then the results are fused to obtain the final new scene liveness detection image. The generated image is more controllable and has better quality.

[0126] In a feasible implementation, the initial scene living body sample generation model may include at least a material locking module, a scene generation module, a fusion module and a feature extraction module;

[0127] S4004: Acquire at least one reference material image, and determine the material scene description text and material body adjustment text corresponding to the reference material image;

[0128] The reference material images are classified into live attack types during the acquisition phase, i.e., each reference material image is determined to belong to either the live material type or the attack material type. Simultaneously, each reference material image is configured with material scene description text and material subject adjustment text for the reference material image before model training.

[0129] The material body adjustment text is the material body adjustment information for the reference material image, such as adjusting the subject backlight. The material body adjustment text can be manually customized. Introducing the material body adjustment text can quickly improve the model processing capability and model robustness.

[0130] The material scene description text is scene description information based on the characteristics and commonalities of the reference material image, and can be configured by an expert end through the introduction of an expert service. The material scene description text can also be generated by using the above-mentioned scene property description generation model, by inputting the reference material image into the scene property description generation model, and the output of the model is used as the material scene description text;

[0131] S4006: Input the reference material image, the material main body adjustment text and the material scene description text into the initial scene living sample generation model to perform at least one round of model training to obtain a trained scene living sample generation model.

[0132] The initial scene living body sample generation model may include at least a material locking module, a scene generation module, a fusion module and a feature extraction module;

[0133] During the model training process, the module structure and model parameters of the feature extraction module can be controlled to remain unchanged, and the model parameters of the remaining material locking module, scene generation module, and fusion module can be adjusted.

[0134] In a feasible implementation, the initial scene living sample generation model performs at least one round of model training in the following form:

[0135] D2: Inputting the reference material image and the material scene description text into the initial scene living body sample generation model, performing image body adjustment on the reference material image based on the material body adjustment text by the material locking module to obtain a material adjustment image, performing scene generation processing based on the material scene description text by the scene generation module to obtain a reference scene image, and performing image fusion processing on the material adjustment image and the reference scene image by the fusion module to obtain a sample object detection image;

[0136] Each round of input to the material locking module is a reference material image of a living material type or an attack material type and a material body adjustment text. The material locking model determines the image body after performing image analysis on the reference material image based on the adjustment semantics of the material body adjustment text. The image body is adjusted in the reference material image based on the adjustment semantics, thereby obtaining the module output: a material adjustment image that satisfies the material body adjustment text.

[0137] Each round of input of the scene generation module is the material scene description text. The scene generation module generates AIGC images based on the material scene description text, and generates reference scene images that meet the material scene description text by AI. Optionally, the scene generation module can use the AIGC model in related technologies.

[0138] Each round of input of the fusion module is the above-mentioned material adjustment image and reference scene image, and the fusion module performs image fusion on the material adjustment image and the reference scene image to generate a new liveness detection image as a sample object detection image;

[0139] D4: Determining, by the feature extraction module, the material adjustment image features of the material adjustment image, the reference material image features of the reference material image, the material adjustment text features of the material body adjustment text, the material scene description text features of the material scene description text, the reference scene image features of the reference scene image, and the sample object image features of the sample object detection image;

[0140] Schematically, a feature extraction module is configured for the initial scene live sample generation model. The feature extraction module can remain unchanged during the model training stage. The feature extraction module is used to perform feature extraction to further generate at least one model training supervision signal, and supervise the model training process based on the model training stage signal.

[0141] D6: Adjust the model parameters of the initial scene live sample generation model based on the material adjustment image features, the reference material image features, the material adjustment text features, the material scene description text features, the reference scene image features and the sample object image features until the initial scene live sample generation model is trained to obtain a scene live sample generation model.

[0142] Schematically, based on the material adjustment image features, the reference material image features, the material adjustment text features, the material scene description text features, the reference scene image features, and the sample object image features, the material subject locking loss, the scene constraint loss, and the image fusion loss can be calculated respectively, thereby obtaining a model comprehensive loss, and the model parameters of the initial scene living sample generation model are adjusted by the model comprehensive loss;

[0143] E2: determining a material main body locking loss based on the material adjustment image feature, the reference material image feature, the material adjustment image feature, and the material adjustment text feature;

[0144] By setting the material main body locking loss as a consistency supervision signal in the material locking process, during the model training process, the material adjustment image features and the reference material image features are supervised to be as consistent as possible, and the reference material image features and the material adjustment text features are as consistent as possible. The material main body locking loss can be obtained by using relevant loss calculation functions (such as Euclidean distance loss calculation function, hinge loss calculation function, etc.);

[0145] For example, the material main body locking loss can be calculated using the following loss calculation formula:

[0146]

[0147] Among them, the LOSS lock The material main body locking loss, the f img1-g Adjust the image characteristics for the material, the f img-ori is the reference material image feature, the f text adjusting text features for the material;

[0148] E4: Determine a scene constraint loss based on the reference scene image feature and the material-adjusted text feature;

[0149] By setting the scene constraint loss as a consistency supervision signal in the scene image generation process, the reference scene image features and the material adjustment text features are supervised to be as consistent as possible during the model training process. In practical applications, the reference scene image features and the material adjustment text features are usually two high-dimensional feature vectors. The two high-dimensional feature vectors are mapped to the vector space of the same dimension and the scene constraint loss can be accurately calculated using a vector loss calculation function. The scene constraint loss can be obtained using related loss calculation functions (such as the Euclidean distance loss calculation function, the hinge loss calculation function, etc.);

[0150] For example, the scene constraint loss can be calculated using the following loss formula:

[0151]

[0152] Among them, the LOSS scene is the scene constraint loss, the f img2-g is the reference scene image feature, the f text adjusting text features for the material;

[0153] E6: Determine an image fusion loss based on the sample object image features, the material adjustment text features, and the material scene description text features;

[0154] By setting the image fusion loss as a consistency supervision signal in the image fusion process, the image features of the sample objects, the material adjustment text features, and the material scene description text features after supervision are kept as consistent as possible during the model training process. In practical applications, the above three features are usually three high-dimensional feature vectors. The three high-dimensional feature vectors are mapped to the vector space of the same dimension and the image fusion loss can be accurately calculated by the loss calculation function in vector form. The image fusion loss can be obtained by using related loss calculation functions (such as Euclidean distance loss calculation function, hinge loss calculation function, etc.);

[0155] For example, the image fusion loss can be calculated using the following loss formula:

[0156]

[0157] Among them, the LOSS fusion is the image fusion loss, the f img3-g is the sample object image feature, the f text1 Adjust the text features for the material, the f text2 Describing text features for the material scene;

[0158] E8: Adjust model parameters of the initial scene living sample generation model based on the material main body locking loss, the scene constraint loss, and the image fusion loss.

[0159] It can be understood that based on the material body locking loss, the scene constraint loss and the image fusion loss, the model comprehensive loss for the initial scene live sample generation model can be obtained. Based on the model comprehensive loss, the model back propagation method is used to adjust the model parameters of the initial scene live sample generation model until the model training end conditions are met, and a trained scene live sample generation model can be obtained.

[0160] The model training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset number threshold, etc. The specific model training termination conditions may be determined based on actual conditions and are not specifically limited here.

[0161] In one or more embodiments of this specification, for highly standardized images such as liveness detection, the existing AIGC solution cannot meet the needs. The training scene liveness sample generation model can meet the needs of automatically generating images based on descriptive text. The new scene liveness sample generation method based on material locking is introduced. By inputting reference materials (various live faces, attack materials, etc.) to confirm the subject, and fine-tuning the subject with text, the text of the previous step is used to generate the scene, and then the results are fused to obtain the final new scene liveness detection image. The image generated in this way is more controllable and of better quality. After completing the above-mentioned model training, batches of new scene samples can be generated, and finally the liveness detection model of the new scene is trained based on the generated samples.

[0162] Optionally, the initial target liveness detection model is trained using the second sample object detection image to obtain a target liveness detection model;

[0163] Furthermore, considering that the quality of generated samples still varies to a certain extent, this section proposes a new scenario liveness detection model training method based on quality adaptive selection. When creating the initial target liveness detection model, the model structure is configured to introduce a new scenario liveness detection model with quality adaptive selection, which can further improve the detection quality and detection effect of liveness detection.

[0164] Optionally, the initial target liveness detection model may include a quality adaptation module, a soft gate module, and a liveness detection module based on the reference liveness detection model;

[0165] In this specification, the use of the second sample object detection image to perform model training on the initial target liveness detection model to obtain the target liveness detection model may be:

[0166] See Figure 7 , Figure 7 This is a diagram of the training process of a target liveness detection model;

[0167] S5002: Inputting the second sample object detection image into the initial target liveness detection model, performing sample quality evaluation processing on the second sample object detection image by the quality adaptation module to obtain a sample quality score, determining a sample training weight for the second sample object detection image based on the sample quality score by using a soft gate module, and performing liveness detection on the second sample object detection image by the liveness detection module to obtain a liveness detection result;

[0168] The quality adaptation module may be, for example, a network module using an inception quality model, and the quality adaptation module is used to evaluate the image quality of the second sample object detection image input in each round and calculate the sample quality score;

[0169] Optionally, the quality adaptation module can be created based on a pre-trained Inception quality model. During the model training phase, the module architecture and model parameters of the quality adaptation module can be controlled to remain unchanged, and it is only used to calculate the image quality score for each second sample object detection image.

[0170] The soft gate module, also known as the soft gate module, performs sample weight evaluation on the second sample object detection image input in each round through the soft gate module to obtain sample training weights, and uses the sample training weights as a weight distribution supervision signal to instruct the liveness detection module to perform targeted model training;

[0171] Optionally, the soft gate module can be created based on a pre-trained soft gate network. During the model training phase, the module architecture and model parameters of the quality-controlled soft gate module remain unchanged, and are only used to calculate the sample training weight for each second sample object detection image. During the model backpropagation training process, the backpropagation weight of each sample for the model's liveness detection module is the weight output by the soft gate module. Based on the soft gate module, the amount of information output by other neurons can be controlled.

[0172] The liveness detection module can be created based on a reference liveness detection model that has been trained in a reference scene. The input of the liveness detection module is the second sample object detection image in each round. The liveness detection module performs liveness detection based on the second sample object detection image to obtain a liveness detection result.

[0173] S5004: Adjust model parameters of the liveness detection module of the initial target liveness detection model based on the sample training weight and the liveness detection result to obtain a target liveness detection model.

[0174] Schematically, the model comprehensive loss of the initial target liveness detection model can be determined by the weighted sparse loss and the liveness detection loss;

[0175] Specifically, a weight sparse loss may be determined based on the sample training weight of each second sample object detection image;

[0176] The weighted sparse loss is used to control the sum of the weights of these second sample object detection images in the training process of each batch of sample images to be as small as possible, that is, the weighted sparse loss is used to control only a part of high-quality samples to participate in the training of the living body detection module.

[0177] Specifically, a liveness detection label of the second sample object detection image may be obtained, and a liveness detection loss may be determined based on the liveness detection result and the liveness detection label;

[0178] Specifically, model parameters of the liveness detection module of the initial target liveness detection model may be adjusted based on the weighted sparse loss and the liveness detection loss.

[0179] It can be understood that based on the weight sparse loss and the liveness detection loss, the model comprehensive loss for the initial target liveness detection model can be obtained. Based on the model comprehensive loss, the model back propagation method is used to adjust the model parameters of the liveness detection module of the initial target liveness detection model until the model training end conditions are met, and a trained target liveness detection model can be obtained.

[0180] The model training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset number threshold, etc. The specific model training termination conditions may be determined based on actual conditions and are not specifically limited here.

[0181] In one or more embodiments of the present specification, the target liveness detection model trained in the above manner can overcome the quality differences of the generated sample images, and realize a training method of the liveness detection model in a new scenario based on quality adaptive selection, which can ensure the detection quality of the liveness detection model.

[0182] The following will be combined Figure 8 , the living body detection device provided in this specification is introduced in detail. It should be noted that, Figure 8 The living body detection device shown is used to execute the Figures 1 to 7 For the convenience of explanation, only the parts related to this specification are shown. For the specific technical details not disclosed, please refer to this specification. Figures 1 to 7 The embodiment shown.

[0183] See Figure 8 , which shows a schematic diagram of the structure of the liveness detection device of this specification. The liveness detection device 1 can be implemented as all or part of the user terminal through software, hardware, or a combination of both. According to some embodiments, the liveness detection device 1 includes a data processing module 11, an image generation module 12, and a liveness detection module 13, which are specifically used to:

[0184] A data processing module 11 is configured to obtain first sample object detection images in a target scene, and determine a target scene property description text based on each of the first sample object detection images;

[0185] an image generation module 12, configured to perform artificial intelligence image generation based on the target scene property description text and the first sample object detection image to obtain a plurality of second sample object detection images, wherein the number of samples of the second sample object detection images is greater than the number of samples of the first sample object detection images;

[0186] The liveness detection module 13 is used to create an initial target liveness detection model for the target scene based on the reference liveness detection model, and use the second sample object detection image to train the initial target liveness detection model to obtain a target liveness detection model. The reference liveness detection model is the liveness detection model under the reference scene.

[0187] Optionally, the data processing module is used to:

[0188] Inputting each of the first sample object images into a scene property description generation model, and outputting at least one target scene common description information and at least one target scene characteristic description information;

[0189] A target scene property description text is determined based on the target scene common description information and the target scene characteristic description information.

[0190] Optionally, the data processing module is configured to: perform description text sampling on the common description information of the target scene to obtain a common description text of the target scene;

[0191] Sampling the target scene characteristic description information to obtain a target scene characteristic description text;

[0192] A target scene property description text is obtained based on the target scene common description text and the target scene characteristic description text.

[0193] Optionally, the data processing module is configured to: determine a text frequency corresponding to each common description text based on the common description information of the target scene, and determine at least one reference common description text from each of the common description texts based on the text frequency;

[0194] Description text sampling is performed on each of the reference common description texts to obtain a common description text of the target scene.

[0195] Optionally, the data processing module is used to: create an initial scene property description generation model;

[0196] Acquire a sample object detection image corresponding to at least one sample scene, wherein the sample object detection image carries a text label describing a property of the scene;

[0197] Using the sample object detection image to perform at least one round of model training on an initial scene property description generation model to obtain a sample scene property description text;

[0198] The model parameters of the initial scene property description generation model are adjusted based on the scene property description text label and the sample scene property description text until the initial scene property description generation model completes model training to obtain a scene property description generation model.

[0199] Optionally, the initial scene property description generation model includes a basic feature encoding module, a common text description generation module and a characteristic text description generation module, and the sample scene property description text includes a sample scene common description text and a sample scene characteristic description text;

[0200] Optionally, the data processing module is used to:

[0201] Inputting the sample object detection image into the initial scene property description generation model, and performing feature extraction on the sample object detection image through the basic feature encoding module to obtain sample image features;

[0202] Performing common descriptions on the sample image features by the common text description generation module to obtain common description texts of sample scenes;

[0203] The characteristic text description generating module performs characteristic description on the sample image features to obtain a sample scene characteristic description text.

[0204] Optionally, the scene property description text tag includes a scene common description text tag and a scene characteristic description text tag.

[0205] The adjusting model parameters of the initial scene property description generation model based on the scene property description text label and the sample scene property description text includes:

[0206] Determining a common text description generation loss based on the sample scene common description text and the scene common description text label;

[0207] Determining a characteristic text description generation loss based on the sample scene characteristic description text and the scene characteristic description text label;

[0208] Model parameters of the initial scene property description generation model are adjusted based on the common text description generation loss and the characteristic text description generation loss.

[0209] Optionally, the image generation module is used to:

[0210] The target scene property description text and the first sample object detection image are input into a scene living body sample generation model, and a plurality of second sample object detection images are output.

[0211] Optionally, the image generation module is used to: create an initial scene living sample generation model,

[0212] Acquire at least one reference material image, and determine a material scene description text and a material body adjustment text corresponding to the reference material image;

[0213] The reference material image, the material main body adjustment text and the material scene description text are input into the initial scene living body sample generation model for at least one round of model training to obtain a trained scene living body sample generation model.

[0214] Optionally, the initial scene living body sample generation model includes a material locking module, a scene generation module, a fusion module and a feature extraction module, and the image generation module is used to:

[0215] Inputting the reference material image and the material scene description text into the initial scene living body sample generation model, performing image main body adjustment on the reference material image based on the material main body adjustment text by the material locking module to obtain a material adjustment image, performing scene generation processing based on the material scene description text by the scene generation module to obtain a reference scene image, and performing image fusion processing based on the material adjustment image and the reference scene image by the fusion module to obtain a sample object detection image;

[0216] Determining, by the feature extraction module, material adjustment image features of the material adjustment image, reference material image features of the reference material image, material adjustment text features of the material body adjustment text, material scene description text features of the material scene description text, reference scene image features of the reference scene image, and sample object image features of the sample object detection image;

[0217] Based on the material adjustment image features, the reference material image features, the material adjustment text features, the material scene description text features, the reference scene image features and the sample object image features, the model parameters of the initial scene live sample generation model are adjusted until the initial scene live sample generation model is trained to obtain a scene live sample generation model.

[0218] Optionally, the image generation module is configured to: determine the material body locking loss based on the material adjustment image feature and the reference material image feature, the material adjustment image feature and the material adjustment text feature;

[0219] determining a scene constraint loss based on the reference scene image features and the material-adjusted text features;

[0220] determining an image fusion loss based on the sample object image features, the material adjustment text features, and the material scene description text features;

[0221] Model parameters of the initial scene living sample generation model are adjusted based on the material main body locking loss, the scene constraint loss and the image fusion loss.

[0222] Optionally, the initial target liveness detection model includes a quality adaptation module, a soft gate module, and a liveness detection module based on the reference liveness detection model, wherein the liveness detection module is configured to:

[0223] Inputting the second sample object detection image into the initial target liveness detection model, performing sample quality evaluation processing on the second sample object detection image using the quality adaptation module to obtain a sample quality score, determining a sample training weight for the second sample object detection image using a soft gate module based on the sample quality score, and performing liveness detection on the second sample object detection image using the liveness detection module to obtain a liveness detection result;

[0224] Based on the sample training weights and the liveness detection results, model parameters of the liveness detection module of the initial target liveness detection model are adjusted to obtain a target liveness detection model.

[0225] Optionally, the living body detection module is used to:

[0226] determining a weight sparsity loss based on the sample training weight of each of the second sample object detection images;

[0227] obtaining a liveness detection label of the second sample object detection image, and determining a liveness detection loss based on the liveness detection result and the liveness detection label;

[0228] Model parameters of the liveness detection module of the initial target liveness detection model are adjusted based on the weighted sparse loss and the liveness detection loss.

[0229] Optionally, the device 1 is further used for:

[0230] The target liveness detection model is deployed in the target scene to perform liveness detection processing on the target object detection image in the target scene through the target liveness detection model.

[0231] It should be noted that the above-described embodiments of the liveness detection device, when performing the liveness detection method, illustrate the division of the functional modules described above only as an example. In actual applications, the above-described functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the liveness detection device and the liveness detection method embodiments provided in the above-described embodiments are based on the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.

[0232] The above serial numbers in this specification are for description only and do not represent the advantages or disadvantages of the embodiments.

[0233] This specification also provides a computer storage medium, which can store multiple instructions, which are suitable for being loaded and executed by a processor as described above. Figures 1 to 7 The liveness detection method of the embodiment shown in the figure can be specifically executed by referring to Figures 1 to 7 The detailed description of the illustrated embodiment will not be repeated here.

[0234] This specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 1 to 7 The liveness detection method of the embodiment shown in the figure can be specifically executed by referring to Figures 1 to 7 The detailed description of the illustrated embodiment will not be repeated here.

[0235] Please refer to Figure 9 , which shows a block diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. The electronic device described in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.

[0236] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions of the electronic device 100 and process data. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communications chip.

[0237] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system developed by Apple, including a system deeply developed based on the IOS system or other systems. The data storage area may also store data created by the electronic device during use, such as a phone book, audio and video data, chat record data, etc.

[0238] See also Figure 10As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve better operating results, the operating system allocates corresponding system resources to different third-party applications. However, the requirements for system resources in different application scenarios in the same third-party application are also different. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and the third-party application are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.

[0239] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0240] Taking the Android operating system as an example, the programs and data stored in the memory 120 are as follows: Figure 11As shown, the memory 120 may store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380. The Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to the operating system space, and the application layer 380 belongs to the user space. The Linux kernel layer 320 provides underlying drivers for various hardware components of electronic devices, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc. The system runtime library layer 340 provides major feature support for the Android system through some C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D drawing support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library (Android runtime), which mainly provides some core libraries that allow developers to write Android applications using the Java language. The application framework layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider management, package management, call management, resource management, and location management. The application layer 380 runs at least one application. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera applications, or third-party applications developed by third-party developers, such as games, instant messaging programs, and photo enhancement programs.

[0241] Taking the operating system as the IOS system as an example, the programs and data stored in the memory 120 are as follows: Figure 12As shown, the IOS system includes: a core operating system layer 420 (Core OS layer), a core service layer 440 (Core Services layer), a media layer 460 (Media layer), and a touchable layer 480 (Cocoa Touch Layer). The core operating system layer 420 includes the operating system kernel, drivers, and underlying program frameworks. These underlying program frameworks provide functions closer to the hardware for use by the program framework located in the core service layer 440. The core service layer 440 provides system services and / or program frameworks required by applications, such as the foundation framework, account framework, advertising framework, data storage framework, network connection framework, geographic location framework, motion framework, etc. The media layer 460 provides applications with audio-visual interfaces, such as graphics and image-related interfaces, audio technology-related interfaces, video technology-related interfaces, and wireless playback (AirPlay) interfaces for audio and video transmission technologies. The touchable layer 480 provides various commonly used interface-related frameworks for application development. The touchable layer 480 is responsible for user touch interaction operations on electronic devices. For example, local notification service, remote push service, advertising framework, game tool framework, message user interface (UI) framework, user interface UIKit framework, map framework, etc.

[0242] exist Figure 12 Among the frameworks shown, those relevant to most applications include, but are not limited to, the Foundation framework in the core services layer 440 and the UIKit framework in the touchable layer 480. The Foundation framework provides many basic object classes and data types, offering fundamental system services for all applications and having nothing to do with the UI. The classes provided by the UIKit framework are the foundational UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UIs, providing the application infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.

[0243] Among them, the method and principle of implementing data communication between third-party applications and the operating system in the IOS system can be referred to the Android system, and this manual will not go into details here.

[0244] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays, which are used to receive touch operations on or near the touch screen using any suitable object such as a finger or a touch pen, and to display the user interface of each application. The touch screen display is usually provided on the front panel of the electronic device. The touch screen display can be designed as a full screen, a curved screen or a special-shaped screen. The touch screen display can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in this specification.

[0245] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components, or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, and other components, which are not described in detail here.

[0246] In this specification, the execution entity of each step can be the electronic device described above. Optionally, the execution entity of each step is the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems, which is not limited in this specification.

[0247] The electronic device of this specification may also be equipped with a display device, which may be any device capable of realizing a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc. The user may use the display device on the electronic device 101 to view displayed text, images, videos and other information. The electronic device may be a smart phone, a tablet computer, a gaming device, an AR (Augmented Reality) device, a car, a data storage device, an audio player, a video player, a notebook, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, electronic clothing and the like.

[0248] exist Figure 9 In the electronic device shown, the processor 110 may be configured to call an application stored in the memory 120 and specifically perform the following operations:

[0249] Acquire first sample object detection images in a target scene, and determine target scene property description text based on each of the first sample object detection images;

[0250] performing artificial intelligence image generation based on the target scene property description text and the first sample object detection image to obtain a plurality of second sample object detection images, wherein the number of samples of the second sample object detection images is greater than the number of samples of the first sample object detection images;

[0251] An initial target liveness detection model for the target scene is created based on the reference liveness detection model, and the initial target liveness detection model is trained using the second sample object detection image to obtain a target liveness detection model. The reference liveness detection model is a liveness detection model under the reference scene.

[0252] In one embodiment, the processor 110 performs the following operations when determining the target scene property description text based on each of the first sample object detection images:

[0253] Inputting each of the first sample object images into a scene property description generation model, and outputting at least one target scene common description information and at least one target scene characteristic description information;

[0254] A target scene property description text is determined based on the target scene common description information and the target scene characteristic description information.

[0255] In one embodiment, the processor 110 performs the following steps when determining the target scene property description text based on the target scene common description information and the target scene characteristic description information:

[0256] Sampling description text of the common description information of the target scene to obtain common description text of the target scene;

[0257] Sampling the target scene characteristic description information to obtain a target scene characteristic description text;

[0258] A target scene property description text is obtained based on the target scene common description text and the target scene characteristic description text.

[0259] In one embodiment, the processor 110 performs the following steps when performing the description text sampling of the common description information of the target scene:

[0260] Determining a text frequency corresponding to each common description text based on the common description information of the target scene, and determining at least one reference common description text from each of the common description texts based on the text frequencies;

[0261] Description text sampling is performed on each of the reference common description texts to obtain a common description text of the target scene.

[0262] In one embodiment, the processor 110 further performs the following steps when executing the liveness detection method:

[0263] Create an initial scene property description generative model;

[0264] Acquire a sample object detection image corresponding to at least one sample scene, wherein the sample object detection image carries a text label describing a property of the scene;

[0265] Using the sample object detection image to perform at least one round of model training on an initial scene property description generation model to obtain a sample scene property description text;

[0266] The model parameters of the initial scene property description generation model are adjusted based on the scene property description text label and the sample scene property description text until the initial scene property description generation model completes model training to obtain a scene property description generation model.

[0267] In one embodiment, the initial scene property description generation model includes a basic feature encoding module, a common text description generation module and a characteristic text description generation module. The processor 110 performs the following steps when executing the sample scene property description text: sample scene common description text and sample scene characteristic description text:

[0268] The sample object detection image is used to perform at least one round of model training on the initial scene property description generation model to obtain a sample scene property description text, and the following steps are performed:

[0269] Inputting the sample object detection image into the initial scene property description generation model, and performing feature extraction on the sample object detection image through the basic feature encoding module to obtain sample image features;

[0270] Performing common descriptions on the sample image features by the common text description generation module to obtain common description texts of sample scenes;

[0271] The characteristic text description generating module performs characteristic description on the sample image features to obtain a sample scene characteristic description text.

[0272] In one embodiment, the scene property description text label includes a scene commonality description text label and a scene characteristic description text label. When the processor 110 adjusts the model parameters of the initial scene property description generation model based on the scene property description text label and the sample scene property description text, the processor 110 performs the following steps:

[0273] Determining a common text description generation loss based on the sample scene common description text and the scene common description text label;

[0274] Determining a characteristic text description generation loss based on the sample scene characteristic description text and the scene characteristic description text label;

[0275] Model parameters of the initial scene property description generation model are adjusted based on the common text description generation loss and the characteristic text description generation loss.

[0276] In one embodiment, the processor 110 performs the following steps after executing the artificial intelligence image generation based on the target scene property description text and the first sample object detection image to obtain a plurality of second sample object detection images:

[0277] The target scene property description text and the first sample object detection image are input into a scene living body sample generation model, and a plurality of second sample object detection images are output.

[0278] In one embodiment, the processor 110 further performs the following steps when executing the liveness detection method:

[0279] Create an initial scene live sample generation model,

[0280] Acquire at least one reference material image, and determine a material scene description text and a material body adjustment text corresponding to the reference material image;

[0281] The reference material image, the material main body adjustment text and the material scene description text are input into the initial scene living body sample generation model for at least one round of model training to obtain a trained scene living body sample generation model.

[0282] In one embodiment, the initial scene live sample generation model includes a material locking module, a scene generation module, a fusion module, and a feature extraction module. The processor 110 performs the following steps after inputting the reference material image, the material main adjustment text, and the material scene description text into the initial scene live sample generation model for at least one round of model training to obtain a trained scene live sample generation model:

[0283] Inputting the reference material image and the material scene description text into the initial scene living body sample generation model, performing image main body adjustment on the reference material image based on the material main body adjustment text by the material locking module to obtain a material adjustment image, performing scene generation processing based on the material scene description text by the scene generation module to obtain a reference scene image, and performing image fusion processing based on the material adjustment image and the reference scene image by the fusion module to obtain a sample object detection image;

[0284] Determining, by the feature extraction module, material adjustment image features of the material adjustment image, reference material image features of the reference material image, material adjustment text features of the material body adjustment text, material scene description text features of the material scene description text, reference scene image features of the reference scene image, and sample object image features of the sample object detection image;

[0285] Based on the material adjustment image features, the reference material image features, the material adjustment text features, the material scene description text features, the reference scene image features and the sample object image features, the model parameters of the initial scene live sample generation model are adjusted until the initial scene live sample generation model is trained to obtain a scene live sample generation model.

[0286] In one embodiment, the processor 110 performs the following steps when performing the step of adjusting the model parameters of the initial scene living sample generation model based on the material adjustment image features, the reference material image features, the material adjustment text features, the material scene description text features, the reference scene image features, and the sample object image features:

[0287] determining a material body locking loss based on the material adjustment image feature, the reference material image feature, the material adjustment image feature, and the material adjustment text feature;

[0288] determining a scene constraint loss based on the reference scene image features and the material-adjusted text features;

[0289] determining an image fusion loss based on the sample object image features, the material adjustment text features, and the material scene description text features;

[0290] Model parameters of the initial scene living sample generation model are adjusted based on the material main body locking loss, the scene constraint loss and the image fusion loss.

[0291] In one embodiment, the initial target liveness detection model includes a quality adaptation module, a soft gate module, and a liveness detection module based on the reference liveness detection model. When the processor 110 performs model training on the initial target liveness detection model using the second sample object detection image to obtain a target liveness detection model, the processor 110 performs the following steps:

[0292] Inputting the second sample object detection image into the initial target liveness detection model, performing sample quality evaluation processing on the second sample object detection image using the quality adaptation module to obtain a sample quality score, determining a sample training weight for the second sample object detection image using a soft gate module based on the sample quality score, and performing liveness detection on the second sample object detection image using the liveness detection module to obtain a liveness detection result;

[0293] Based on the sample training weights and the liveness detection results, model parameters of the liveness detection module of the initial target liveness detection model are adjusted to obtain a target liveness detection model.

[0294] In one embodiment, the processor 110 performs the following steps when adjusting the model parameters of the liveness detection module of the initial target liveness detection model based on the sample training weights and the liveness detection results:

[0295] determining a weight sparsity loss based on the sample training weight of each of the second sample object detection images;

[0296] obtaining a liveness detection label of the second sample object detection image, and determining a liveness detection loss based on the liveness detection result and the liveness detection label;

[0297] Model parameters of the liveness detection module of the initial target liveness detection model are adjusted based on the weighted sparse loss and the liveness detection loss.

[0298] In one embodiment, the processor 110 further performs the following steps when executing the liveness detection method:

[0299] The target liveness detection model is deployed in the target scene to perform liveness detection processing on the target object detection image in the target scene through the target liveness detection model.

[0300] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.

[0301] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the first sample object detection image and the second sample object detection image mentioned in this specification were obtained with full authorization.

[0302] The above disclosure is only a preferred embodiment of this specification, and certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.

Claims

1. A method for detecting a living body, comprising: Acquire a plurality of first sample object detection images in a target scene, and determine a target scene property description text based on each of the first sample object detection images; Inputting the first sample object detection image and the target scene property description text into a scene live sample generation model to obtain a plurality of second sample object detection images in the target scene; the scene live sample generation model includes a material locking module, a scene generation module, and a fusion module; in the scene live sample generation model, subject adjustment is performed on the first sample object detection image by the material locking module to obtain an object adjustment image; scene generation processing is performed by the scene generation module based on the target scene property description text to obtain a target scene image; and fusion processing is performed by the fusion module based on the object adjustment image and the target scene image to obtain a plurality of second sample object detection images in the target scene, wherein the number of samples of the second sample object detection images is greater than the number of samples of the first sample object detection images; An initial target liveness detection model for the target scene is created based on the reference liveness detection model, and the initial target liveness detection model is trained using the second sample object detection image to obtain a target liveness detection model. The reference liveness detection model is a liveness detection model under the reference scene.

2. The method according to claim 1, wherein determining the target scene property description text based on each of the first sample object detection images comprises: Inputting each of the first sample object detection images into a scene property description generation model, and outputting at least one target scene common description information and at least one target scene characteristic description information; A target scene property description text is determined based on the target scene common description information and the target scene characteristic description information.

3. The method according to claim 2, wherein determining the target scene property description text based on the target scene common description information and the target scene characteristic description information comprises: Sampling description text of the common description information of the target scene to obtain common description text of the target scene; Sampling the target scene characteristic description information to obtain a target scene characteristic description text; A target scene property description text is obtained based on the target scene common description text and the target scene characteristic description text.

4. The method according to claim 3, wherein sampling the description text of the common description information of the target scene comprises: Determining a text frequency corresponding to each common description text based on the common description information of the target scene, and determining at least one reference common description text from each of the common description texts based on the text frequencies; Description text sampling is performed on each of the reference common description texts to obtain a common description text of the target scene.

5. The method according to claim 2, further comprising: Create an initial scene property description generative model; Acquire a sample object detection image corresponding to at least one sample scene, wherein the sample object detection image carries a text label describing a property of the scene; Using the sample object detection image to perform at least one round of model training on an initial scene property description generation model to obtain a sample scene property description text; The model parameters of the initial scene property description generation model are adjusted based on the scene property description text label and the sample scene property description text until the initial scene property description generation model completes model training to obtain a scene property description generation model.

6. The method according to claim 5, wherein the initial scene property description generation model comprises a basic feature encoding module, a common text description generation module, and a characteristic text description generation module, and the sample scene property description text comprises a sample scene common description text and a sample scene characteristic description text. The method of performing at least one round of model training on an initial scene property description generation model using the sample object detection image to obtain a sample scene property description text includes: Inputting the sample object detection image into the initial scene property description generation model, and performing feature extraction on the sample object detection image through the basic feature encoding module to obtain sample image features; Performing common descriptions on the sample image features by the common text description generation module to obtain common description texts of sample scenes; The characteristic text description generating module performs characteristic description on the sample image features to obtain a sample scene characteristic description text.

7. The method according to claim 6, wherein the scene property description text label comprises a scene common description text label and a scene characteristic description text label. The adjusting model parameters of the initial scene property description generation model based on the scene property description text label and the sample scene property description text includes: Determining a common text description generation loss based on the sample scene common description text and the scene common description text label; Determining a characteristic text description generation loss based on the sample scene characteristic description text and the scene characteristic description text label; Model parameters of the initial scene property description generation model are adjusted based on the common text description generation loss and the characteristic text description generation loss.

8. The method according to claim 1, further comprising: Create an initial scene live sample generation model, Acquire at least one reference material image, and determine a material scene description text and a material body adjustment text corresponding to the reference material image; The reference material image, the material main body adjustment text and the material scene description text are input into the initial scene living body sample generation model for at least one round of model training to obtain a trained scene living body sample generation model.

9. The method according to claim 8, wherein the initial scene living body sample generation model comprises a material locking module, a scene generation module, a fusion module and a feature extraction module. The step of inputting the reference material image, the material main body adjustment text, and the material scene description text into the initial scene living body sample generation model for at least one round of model training to obtain a trained scene living body sample generation model includes: Inputting the reference material image and the material scene description text into the initial scene living body sample generation model, performing image main body adjustment on the reference material image based on the material main body adjustment text by the material locking module to obtain a material adjustment image, performing scene generation processing based on the material scene description text by the scene generation module to obtain a reference scene image, and performing image fusion processing based on the material adjustment image and the reference scene image by the fusion module to obtain a sample object detection image; Determining, by the feature extraction module, material adjustment image features of the material adjustment image, reference material image features of the reference material image, material adjustment text features of the material body adjustment text, material scene description text features of the material scene description text, reference scene image features of the reference scene image, and sample object image features of the sample object detection image; Based on the material adjustment image features, the reference material image features, the material adjustment text features, the material scene description text features, the reference scene image features and the sample object image features, the model parameters of the initial scene live sample generation model are adjusted until the initial scene live sample generation model is trained to obtain a scene live sample generation model.

10. The method according to claim 9, wherein the adjusting model parameters of the initial scene living body sample generation model based on the material adjustment image features, the reference material image features, the material adjustment text features, the material scene description text features, the reference scene image features, and the sample object image features comprises: determining a material body locking loss based on the material adjustment image feature, the reference material image feature, the material adjustment image feature, and the material adjustment text feature; determining a scene constraint loss based on the reference scene image features and the material-adjusted text features; determining an image fusion loss based on the sample object image features, the material adjustment text features, and the material scene description text features; Model parameters of the initial scene living sample generation model are adjusted based on the material main body locking loss, the scene constraint loss and the image fusion loss.

11. The method according to claim 1, wherein the initial target liveness detection model comprises a quality adaptation module, a soft gate module, and a liveness detection module based on the reference liveness detection model. The using the second sample object detection image to perform model training on the initial target liveness detection model to obtain the target liveness detection model includes: Inputting the second sample object detection image into the initial target liveness detection model, performing sample quality evaluation processing on the second sample object detection image using the quality adaptation module to obtain a sample quality score, determining a sample training weight for the second sample object detection image using a soft gate module based on the sample quality score, and performing liveness detection on the second sample object detection image using the liveness detection module to obtain a liveness detection result; Based on the sample training weights and the liveness detection results, model parameters of the liveness detection module of the initial target liveness detection model are adjusted to obtain a target liveness detection model.

12. The method according to claim 11, wherein adjusting the model parameters of the liveness detection module of the initial target liveness detection model based on the sample training weight and the liveness detection result comprises: determining a weight sparsity loss based on the sample training weight of each of the second sample object detection images; obtaining a liveness detection label of the second sample object detection image, and determining a liveness detection loss based on the liveness detection result and the liveness detection label; Model parameters of the liveness detection module of the initial target liveness detection model are adjusted based on the weighted sparse loss and the liveness detection loss.

13. The method according to claim 1, further comprising: The target liveness detection model is deployed in the target scene to perform liveness detection processing on the target object detection image in the target scene through the target liveness detection model.

14. A living body detection device, comprising: a data processing module, configured to obtain a plurality of first sample object detection images in a target scene, and determine a target scene property description text based on each of the first sample object detection images; an image generation module, configured to input the first sample object detection image and the target scene property description text into a scene live sample generation model to obtain a plurality of second sample object detection images in the target scene; the scene live sample generation model comprising a material locking module, a scene generation module, and a fusion module; in the scene live sample generation model, subject adjustment is performed on the first sample object detection image by the material locking module to obtain an object adjustment image; scene generation processing is performed by the scene generation module based on the target scene property description text to obtain a target scene image; and fusion processing is performed by the fusion module based on the object adjustment image and the target scene image to obtain a plurality of second sample object detection images in the target scene, wherein the number of samples of the second sample object detection images is greater than the number of samples of the first sample object detection images; A liveness detection module is used to create an initial target liveness detection model for the target scene based on a reference liveness detection model, and to train the initial target liveness detection model using the second sample object detection image to obtain a target liveness detection model. The reference liveness detection model is a liveness detection model under a reference scene.

15. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 13.

16. A computer program product, wherein the computer program product stores at least one instruction, wherein the at least one instruction is loaded by a processor and executes the method steps according to any one of claims 1 to 13.

17. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Image generation method based on multiple auxiliary information

    CN113052784A

  • Living body detection model training method and living body detection method and system

    CN115497176A