Liveness detection model training method and apparatus, device, computer-readable storage medium, and computer program product

By obtaining the description text of the image sample, the style enhancement and parameter update are carried out, and the live detection model is trained, which solves the problem of insufficient generalization ability of the model and achieves higher detection accuracy and generalization ability.

WO2025145837A1PCT designated stage expired Publication Date: 2025-07-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2024/136371
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-02
Filing Date
2024-12-03
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

When the existing live detection model faces new live attack data, the model generalization ability is insufficient, resulting in poor detection results.

Method used

By obtaining the first description text and the second description text of the image sample, style enhancement is performed, style enhancement parameters are updated, image features are style enhanced using the target style enhancement parameters, and live detection models are trained based on the enhanced image features to improve the generalization ability and detection accuracy of the model.

Benefits of technology

The training effect and generalization ability of the live detection model are improved, the detection accuracy of different styles of images is enhanced, and the accuracy of live detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136371_10072025_PF_FP_ABST
    Figure CN2024136371_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a liveness detection model training method and apparatus, a device, a computer-readable storage medium, and a computer program product, which can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving. The method comprises: performing style enhancement on image features of an image sample on the basis of a first description text of the image sample and a second description text of the image sample under a target style to obtain first enhanced image features, and performing style enhancement on the image features on the basis of style enhancement parameters to obtain second enhanced image features; updating the style enhancement parameters on the basis of the difference between the second enhanced image features and the first enhanced image features to obtain target style enhancement parameters; performing style enhancement on the image features on the basis of the target style enhancement parameters to obtain target enhanced image features of the image sample, and training a liveness detection model on the basis of the target enhanced image features.
Need to check novelty before this filing date? Find Prior Art

Description

Training method, device, equipment, computer-readable storage medium and computer program product for living body detection model

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on and claims the priority of Chinese patent application with application number 2024100146685 and application date of January 2, 2024. The entire content of the Chinese patent application is hereby incorporated into this application by reference. Technical Field

[0003] The present application relates to the field of artificial intelligence technology, and in particular to a training method, apparatus, device, computer-readable storage medium, and computer program product for a liveness detection model. Background Art

[0004] In related technologies, liveness detection is implemented using liveness detection models, which are trained by constructing training samples based on existing liveness attack data. However, with the continuous application of liveness detection, new liveness attack data is generated rapidly. This leads to problems such as insufficient model generalization and poor liveness detection results in liveness detection models trained solely on existing liveness attack data. Summary of the Invention

[0005] The embodiments of the present application provide a training method for a liveness detection model, a liveness detection method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the training effect of the liveness detection model, thereby improving the model generalization ability and liveness detection accuracy of the trained liveness detection model.

[0006] The technical solution of the embodiment of the present application is implemented as follows:

[0007] The present embodiment provides a method for training a liveness detection model, which is applied to an electronic device and includes:

[0008] Acquire an image sample for training a liveness detection model, and acquire a first description text of the image sample and a second description text of the image sample in a target style;

[0009] Based on the first description text and the second description text, performing style enhancement on the image feature of the image sample to obtain a first enhanced image feature, and based on a style enhancement parameter, performing style enhancement on the image feature to obtain a second enhanced image feature;

[0010] updating the style enhancement parameter based on a difference between the second enhanced image feature and the first enhanced image feature to obtain a target style enhancement parameter;

[0011] When training the liveness detection model, the image features are style enhanced based on the target style enhancement parameters to obtain the target enhanced image features of the image sample, and the liveness detection model is trained based on the target enhanced image features to obtain a trained liveness detection model.

[0012] The present application also provides a liveness detection method, which is applied to an electronic device and includes:

[0013] Acquire an image to be detected including a biological object and a plurality of category texts, wherein the plurality of category texts include: a first text representing a living body category, and a second text representing each of the at least one non-living body category;

[0014] Calling a liveness detection model, performing text encoding on each of the category texts to obtain category text features of each of the category texts, and performing feature extraction on the image to be detected to obtain features of the image to be detected;

[0015] Calling the liveness detection model, performing liveness detection based on the features of the image to be detected and the features of each category text, and obtaining a liveness detection result of the image to be detected;

[0016] The liveness detection result includes a predicted probability that the biological object belongs to each category, and the categories include the liveness category and the at least one non-liveness category;

[0017] The liveness detection model is trained based on the liveness detection model training method provided in the embodiment of the present application.

[0018] The present application also provides a training device for a liveness detection model, including:

[0019] a first acquisition module configured to acquire an image sample for training a liveness detection model, and acquire a first description text of the image sample and a second description text of the image sample in a target style;

[0020] a style enhancement module configured to perform style enhancement on the image features of the image sample based on the first description text and the second description text to obtain a first enhanced image feature, and to perform style enhancement on the image features based on a style enhancement parameter to obtain a second enhanced image feature;

[0021] an updating module configured to update the style enhancement parameter based on a difference between the second enhanced image feature and the first enhanced image feature to obtain a target style enhancement parameter;

[0022] The training module is configured to, when training the liveness detection model, perform style enhancement on the image features based on the target style enhancement parameters to obtain the target enhanced image features of the image sample, and train the liveness detection model based on the target enhanced image features to obtain a trained liveness detection model.

[0023] The present application also provides a living body detection device, including:

[0024] A second acquisition module is configured to acquire an image to be detected including a biological object and a plurality of category texts, wherein the plurality of category texts include: a first text representing a living body category, and a second text representing each of the at least one non-living body category;

[0025] a feature extraction module configured to call a liveness detection model, perform text encoding on each of the category texts to obtain category text features of each of the category texts, and perform feature extraction on the image to be detected to obtain features of the image to be detected;

[0026] a liveness detection module configured to call the liveness detection model, perform liveness detection based on the features of the image to be detected and the features of each category text, and obtain a liveness detection result of the image to be detected;

[0027] The liveness detection result includes a predicted probability that the biological object belongs to each category, and the categories include the liveness category and the at least one non-liveness category;

[0028] The liveness detection model is trained based on the liveness detection model training method provided in the embodiment of the present application.

[0029] An embodiment of the present application further provides an electronic device, including:

[0030] a memory configured to store computer-executable instructions;

[0031] The processor is configured to implement the method provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.

[0032] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the method provided in the embodiment of the present application is implemented.

[0033] An embodiment of the present application further provides a computer program product, comprising computer-executable instructions or a computer program, which, when executed by a processor, implements the method provided in the embodiment of the present application.

[0034] The embodiments of the present application have the following beneficial effects:

[0035] Applying the above-mentioned embodiments of the present application, first, an image sample for training a liveness detection model is obtained, and a first description text of the image sample and a second description text of the image sample in a target style are obtained; then, based on the first description text and the second description text, the image features of the image sample are style enhanced to obtain a first enhanced image feature, and based on the style enhancement parameters, the image features are style enhanced to obtain a second enhanced image feature; then, based on the difference between the second enhanced image feature and the first enhanced image feature, the style enhancement parameters are updated to obtain a target style enhancement parameter; in this way, when training the liveness detection model, the image features are style enhanced based on the target style enhancement parameters to obtain the target enhanced image feature of the image sample, and the liveness detection model is trained based on the target enhanced image feature to obtain a trained liveness detection model. Here, through text guidance, a target style enhancement parameter is learned, and the target style enhancement parameter is used to perform style enhancement on the image features of the image sample to obtain the target enhanced image feature, thereby realizing the style diversification expansion of the image features of the existing image samples and enriching the image features of the image samples. When the liveness detection model is trained through the target enhanced image features of the image samples, the training effect of the liveness detection model can be improved, thereby improving the model generalization ability and liveness detection accuracy of the trained liveness detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] FIG1 is a schematic diagram of the architecture of a training system for a liveness detection model provided in an embodiment of the present application;

[0037] FIG2 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0038] FIG3 is a flow chart of a method for training a liveness detection model according to an embodiment of the present application;

[0039] FIG4 is a schematic diagram of a process for performing style enhancement on image features according to an embodiment of the present application;

[0040] FIG5 is a schematic diagram of the structure of a machine learning model provided in an embodiment of the present application;

[0041] FIG6 is a flow chart of a method for training a liveness detection model according to an embodiment of the present application;

[0042] FIG7 is a flow chart of a method for training a liveness detection model according to an embodiment of the present application;

[0043] FIG8 is a schematic diagram of training a liveness detection model provided in an embodiment of the present application;

[0044] FIG9 is a schematic flow chart of a liveness detection method according to an embodiment of the present application;

[0045] FIG10 is a schematic diagram of the structure of a liveness detection model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0047] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0048] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0049] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0050] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0051] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0052] 1) Client: An application running in a terminal that provides various services, such as a client that supports liveness detection (such as a client that supports payment, a client that supports identity authentication, etc.).

[0053] 2) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0054] 3) Computer vision (CV) is the science of making machines "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying and measuring objects, and then further processing them to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MoE, and MAE, can be quickly and widely applied to specific downstream tasks through fine-tuning. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, text recognition (Optical Character Recognition, OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, three-dimensional (3D) technology, virtual reality, augmented reality, simultaneous positioning and map construction, etc. It also includes common biometric recognition technologies such as face recognition, fingerprint recognition, and liveness detection technology.

[0055] 4) Liveness detection is a method for determining a subject's true physiological characteristics in scenarios such as identity verification. For example, in facial recognition applications, liveness detection can verify that the user is truly alive by using facial landmark location and tracking technologies, using combinations of actions such as blinking, opening the mouth, shaking the head, and nodding. It effectively protects against common attacks such as photos, videos, face swaps, masks, occlusions, 3D animations, and screen capture, safeguarding users' interests.

[0056] 5) Model generalization refers to the ability of a machine learning model to perform well on new data beyond the training data. A model with good generalization is able to maintain high accuracy and stability when faced with unseen data. Generalization is a key factor in measuring whether the model can learn the underlying principles from the data, rather than simply fitting the training data.

[0057] Based on the above description of the nouns and terms involved in the embodiments of this application, the embodiments of this application are described in detail below. The embodiments of this application provide a training method for a liveness detection model, a liveness detection method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product to improve the training effect of the liveness detection model, thereby improving the model generalization ability and liveness detection accuracy of the trained liveness detection model.

[0058] It should be noted that the collection and processing of relevant data in this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in practice, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0059] The following describes the training system for the liveness detection model provided in an embodiment of the present application. Referring to FIG1 , FIG1 is a schematic diagram of the architecture of the training system for the liveness detection model provided in an embodiment of the present application. To support an exemplary application, the training system 100 for the liveness detection model includes: a server 200, a network 300, and a terminal 400. The terminal 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, using wireless or wired links to achieve data transmission.

[0060] Here, the terminal 400 sends an update request for the style enhancement parameters to the server 200 in response to the update instruction for the style enhancement parameters; the server 200 receives the update request for the style enhancement parameters sent by the terminal 400; in response to the update request, the image sample for training the liveness detection model is obtained, and the first description text of the image sample and the second description text of the image sample in the target style are obtained; based on the first description text and the second description text, the image features of the image sample are style enhanced to obtain the first enhanced image features, and based on the style enhancement parameters, the image features are style enhanced to obtain the second enhanced image features; based on the difference between the second enhanced image features and the first enhanced image features, the style enhancement parameters are updated to obtain the target style enhancement parameters.

[0061] In some exemplary scenarios, when the liveness detection model needs to be trained, the user can trigger a training instruction for the liveness detection model at the terminal 400; the terminal 400 responds to the training instruction for the liveness detection model and sends a training request for the liveness detection model to the server 200; the server 200 receives the training request for the liveness detection model sent by the terminal 400; in response to the training request, the image features are style enhanced based on the target style enhancement parameters to obtain the target enhanced image features of the image sample, and the liveness detection model is trained based on the target enhanced image features to obtain a trained liveness detection model.

[0062] In some exemplary scenarios, after the server 200 obtains the trained liveness detection model, it can actively send the trained liveness detection model to the terminal 400 for use by the terminal 400 when performing liveness detection; of course, the terminal 400 can also actively obtain the liveness detection model from the server 200 when performing liveness detection. In this case, the server 200 sends the liveness detection model to the terminal 400 when the terminal 400 actively obtains it.

[0063] In some exemplary scenarios, the terminal 400 may be provided with a client that supports liveness detection. A user may trigger a liveness detection instruction for an image to be detected (including a biological object) at the terminal 400. In response to the liveness detection instruction, the terminal 400 calls a trained liveness detection model to perform liveness detection on the image to be detected, and obtains a liveness detection result. The liveness detection result includes a predicted probability that the biological object in the image to be detected belongs to each category, and the category includes a liveness category and at least one non-liveness category. In this way, whether the biological object is a real live object can be determined based on the predicted probability that the biological object belongs to each category in the liveness detection result. For example, the category corresponding to the maximum predicted probability is used as the target category to which the biological object belongs. If the target category is the liveness category, then the biological object is a real live object. If the target category is the non-liveness category, then the biological object is not a real live object.

[0064] In some embodiments, the training method of the liveness detection model provided in the embodiments of the present application is implemented by an electronic device, for example, it can be implemented by a terminal alone, it can also be implemented by a server alone, or it can be implemented by a terminal and a server in collaboration. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, computer vision, biometric payment, biometric unlocking (such as access control system unlocking, terminal device unlocking, etc.), biometric identity verification, etc.

[0065] In some embodiments, the electronic device for implementing the training method of the liveness detection model provided in the embodiments of the present application can be various types of terminals or servers. Among them, the server (such as server 200) can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal (such as terminal 400) can be a laptop, a tablet computer, a desktop computer, a smart phone, an intelligent voice interaction device (such as a smart speaker), a smart home appliance (such as a smart TV), a smart watch, a car terminal, a wearable device, a virtual reality (Virtual Reality, VR) device, an aircraft, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected by wired or wireless communication, and the embodiments of the present application are not limited to this.

[0066] In some embodiments, the terminal or server can implement the training method of the living body detection model provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0067] The following describes an electronic device for implementing a training method for a liveness detection model provided in an embodiment of the present application. Referring to FIG. 2 , FIG. 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 500 provided in an embodiment of the present application can be a terminal or a server. As shown in FIG. 2 , the electronic device 500 includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It will be understood that the bus system 540 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 540 in FIG.

[0068] In some embodiments, the training device for the liveness detection model provided in the embodiments of the present application can be implemented in software. Figure 2 shows a training device 555 for the liveness detection model stored in the memory 550, which can be software in the form of programs and plug-ins, including the following software modules: a first acquisition module 5551, a style enhancement module 5552, an update module 5553 and a training module 5554. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.

[0069] The following describes the training method of the liveness detection model provided by the embodiment of the present application. As mentioned above, the training method of the liveness detection model provided by the embodiment of the present application is implemented by an electronic device, for example, it can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. Therefore, the execution body of each step will not be repeated below. Referring to Figure 3, Figure 3 is a flow chart of the training method of the liveness detection model provided by the embodiment of the present application. The training method of the liveness detection model provided by the embodiment of the present application includes:

[0070] Step 101: Obtain an image sample for training a liveness detection model, and obtain a first description text of the image sample and a second description text of the image sample in a target style.

[0071] In step 101, an image sample for training a liveness detection model is obtained. The image sample belongs to an image sample set for training a liveness detection model. The image sample set includes multiple image samples, each of which is an image of a biological object sample. The image sample can be obtained by using a photographing device to photograph the biological object sample and the background in which the biological object sample is located in a liveness detection scene, thereby obtaining the image sample. Furthermore, during the photographing process of the image photographing device, the biological object sample can perform corresponding actions in accordance with the requirements of liveness detection. For example, when the biological object sample is a hand, it can perform corresponding gestures (such as an OK gesture, a single-finger gesture, etc.) in accordance with the requirements of liveness detection. For another example, when the biological object sample is a face, it can perform actions such as blinking, opening the mouth, shaking the head, and nodding in accordance with the requirements of liveness detection. In practical applications, the target image captured by the image capture device can be directly used as an image sample, or the target image captured by the image capture device can be cropped to obtain an image sample; for example, the target image can be detected for biological objects, and a first area where the biological object sample is located in the target image can be determined, and the first area can be cropped from the target image as an image sample; of course, the image sample in the liveness detection scene can include the background where the biological object sample is located. Therefore, after determining the first area, the first area can be used as the center and expanded by a preset multiple to obtain a second area, and the second area can be cropped from the target image as an image sample. In this way, the image sample can include part of the background of the target image.

[0072] At the same time, a first description text of the image sample and a second description text of the image sample in the target style are also obtained. That is, the first description text is used to describe the image sample, and the first description text describes the image sample in the original style, which is the real style of the image sample; the second description text describes the image sample in the target style, and the target style is the style to be expanded for the image sample. By expanding the style of the image sample, an image sample with the target style is obtained. It should be noted that the target style includes but is not limited to styles in terms of lighting (such as bright ambient light, dim ambient light, etc.), background (such as single background, complex background, etc.), posture, expression, etc. of biological objects. For example, if the first description text of the image sample can be "target image captured under bright ambient light" and the target style is yellow ambient light, then the second description text can be "target image captured under yellow ambient light".

[0073] It should be noted that the style (such as the target style, original style, etc.) of the image (such as the above-mentioned image sample) used for liveness detection in the embodiments of the present application refers to the unique features and characteristics of the image in terms of visual effects. For example, the following are some styles of images used for liveness detection, including: 1) Dynamicity: Liveness detection images usually contain dynamic elements, such as facial movements, expression changes, etc., to simulate the dynamic characteristics of real faces. 2) Multi-angle: In order to better simulate the real situation, the liveness detection image may contain biological objects (such as faces) taken from different angles to improve the recognition ability of the liveness detection model under different viewing angles. 3) Natural light and shadow: The liveness detection image may include photos taken under natural light, as well as situations where shadows are generated under different lighting conditions (such as backlight, side light). 4) Multiple backgrounds: In order to improve the versatility of the liveness detection model, the liveness detection image may contain biological objects (such as faces) taken in different background environments, including indoors, outdoors, complex backgrounds, etc. 5) Facial features: Liveness detection images may highlight key parts of biological objects. Taking facial features as an example, parts such as eyes, nose, and mouth can be highlighted so that the liveness detection model can focus on these key parts for recognition. 6) Texture and details: In order to distinguish between real skin texture and the texture of non-living materials, liveness detection images may contain high-resolution, detailed details of biological objects (such as human faces). 7) Expressions and movements: Liveness detection images may contain biological objects with different expressions and movements. Taking human faces as an example, movements such as blinking, opening the mouth, and shaking the head can be included to improve the liveness detection model's ability to detect dynamic changes. 8) Reflective and translucent properties: In order to simulate the reflective and translucent properties of real biological objects (such as human faces) under light, liveness detection images may contain biological objects (such as human faces) under different lighting conditions.

[0074] Step 102: Based on the first description text and the second description text, perform style enhancement on the image features of the image sample to obtain a first enhanced image feature, and based on the style enhancement parameter, perform style enhancement on the image features to obtain a second enhanced image feature.

[0075] In step 102, the image features of the image sample can be first extracted. For example, a pre-trained first image encoder can be used to perform image encoding on the image sample to obtain the image features of the image sample. For example, the first image encoder can be constructed based on a convolutional neural network, a recurrent neural network, a Transformer network, etc., and trained using image data. For example, the first image encoder can include multiple cascaded image encoding layers, through which the image sample is downsampled to obtain image features. This can improve the extraction effect of the image features and make the image features more capable of expressing the features of the image sample. Then, based on the first description text and the second description text, the image features of the image sample are style enhanced to obtain first enhanced image features. In this way, the image features can be style enhanced by being guided by the second description text describing the image sample under the target style, generating first enhanced image features corresponding to the image sample and fused with the target style.

[0076] In some embodiments, referring to FIG4 , based on the first description text and the second description text, the image features of the image sample can be style-enhanced through the following steps to obtain a first enhanced image feature: Step 301, text encoding the first description text to obtain a first text feature, and text encoding the second description text to obtain a second text feature; Step 302, image encoding the image features of the image sample to obtain an encoded image feature; Step 303, fusing the second text feature, the first text feature and the encoded image feature to obtain a first enhanced image feature.

[0077] For step 301, the first description text and the second description text are respectively text-encoded to obtain the first text feature of the first description text and the second text feature of the second description text. For example, text encoding can be achieved by a pre-trained text encoder. For example, the text encoder can be constructed based on a convolutional neural network, a recurrent neural network, a Transformer network, etc., and trained through text data. For example, the text encoder can include multiple cascaded text encoding layers, through which the first description text is downsampled to obtain the first text feature; the processing of the second description text is the same, which can improve the extraction effect of the text feature (including the first text feature and the second text feature), so that the text feature has a stronger feature expression ability for the description text (including the first description text and the second description text). In step 302, the image features of the image sample are encoded to obtain encoded image features. For example, image encoding can be achieved using a pre-trained second image encoder. The second image encoder can be constructed based on a convolutional neural network, a recurrent neural network, a Transformer network, etc., and trained using image data. For example, the second image encoder can include multiple cascaded image encoding layers, through which the image features are downsampled to obtain encoded image features. This can further improve feature extraction and make the encoded image features more expressive of the features of the image sample. In step 303, the second text features, the first text features, and the encoded image features are fused to obtain first enhanced image features. Thus, through steps 301-303, by encoding the first description text, the second description text, and the image features, and fusing the second text features, the first text features, and the encoded image features, first enhanced image features that incorporate the target style are generated, thereby improving the style enhancement effect.

[0078] In some embodiments, step 303 can be implemented by performing the following steps: determining the text feature difference between the second text feature and the first text feature; and fusing the text feature difference with the encoded image feature to obtain the first enhanced image feature. Here, the text feature difference can be obtained by subtracting the first text feature from the second text feature; when fusing the text feature difference with the encoded image feature, the text feature difference and the encoded image feature can be added or multiplied to obtain the first enhanced image feature. In this way, by fusing the text feature difference between the second text feature and the first text feature into the encoded image feature, the target style of the image feature can be enhanced.

[0079] In step 102, the image features are also style-enhanced based on the style enhancement parameters to obtain second enhanced image features. In an embodiment of the present application, a style enhancement parameter is provided, which is used to style-enhance the image features of the image sample. The style enhancement parameter has an initial value, which can be pre-set based on experience. Based on the first description text and the second description text, the style enhancement parameter can be learned to obtain a better style enhancement parameter, so that the second enhanced image feature obtained by style-enhancing the image feature using the learned style enhancement parameter is closer to the first enhanced image feature obtained by style-enhancing the image feature based on the first description text and the second description text.

[0080] In some embodiments, based on the style enhancement parameters, the image features can be style-enhanced by the following steps to obtain second enhanced image features: fusing the style enhancement parameters and the image features to obtain a first fusion result; and performing image encoding on the first fusion result to obtain second enhanced image features. Here, fusing the style enhancement parameters and the image features can be adding the style enhancement parameters and the image features, or multiplying the style enhancement parameters and the image features. When performing image encoding on the first fusion result, the first fusion result can be image-encoded by a pre-trained second image encoder to obtain the second enhanced image features. The second image encoder can be constructed based on a convolutional neural network, a recurrent neural network, a Transformer network, etc., and trained using image data; for example, the second image encoder can include multiple cascaded image encoding layers, and the first fusion result is downsampled by the multiple cascaded image encoding layers to obtain the second enhanced image features. This can make the second enhanced image features closer to the target style, thereby improving the style enhancement effect.

[0081] Step 103: Based on the difference between the second enhanced image feature and the first enhanced image feature, the style enhancement parameters are updated to obtain target style enhancement parameters.

[0082] In this embodiment of the present application, the style enhancement parameters need to be learned to obtain more optimal style enhancement parameters, so that the second enhanced image features obtained by style enhancement of the image features using the style enhancement parameters are closer to the first enhanced image features obtained by style enhancement of the image features based on the first and second description texts. Therefore, in step 103, the style enhancement parameters can be updated based on the difference between the second enhanced image features and the first enhanced image features to obtain the target style enhancement parameters.

[0083] In some embodiments, the above-mentioned style enhancement parameters belong to a machine learning model; based on this, based on the difference between the second enhanced image features and the first enhanced image features, the style enhancement parameters can be updated through the following steps to obtain target style enhancement parameters: obtaining the loss function of the machine learning model; based on the difference between the second enhanced image features and the first enhanced image features, determining the value of the loss function; based on the value of the loss function, updating the style enhancement parameters of the machine learning model to obtain the target style enhancement parameters.

[0084] Here, the machine learning model can be pre-built, and the style enhancement parameter is a model parameter in the machine learning model. Thus, the second enhanced image feature can be understood as the predicted output when training the machine learning model, and the first enhanced image feature can be understood as the corresponding label. Therefore, a loss function of the machine learning model can be obtained, and the value of the loss function can be determined based on the difference between the second enhanced image feature and the first enhanced image feature. The machine learning model can then be trained based on the value of the loss function to update the model parameters of the machine learning model. In this way, the style enhancement parameter is updated during the training process of the machine learning model. When the training target of the machine learning model is achieved (e.g., completing a preset number of training rounds, or the loss function value meeting the training end condition, etc.), the style enhancement parameter obtained in the last round is used as the target style enhancement parameter. In practical applications, when training the machine learning model, each model parameter that needs to be updated in the machine learning model can be updated, including the above-mentioned style enhancement parameter. Alternatively, when building the machine learning model, it is possible to ensure that only one model parameter in the machine learning model needs to be learned, that is, only the style enhancement parameter that needs to be learned is included in the machine learning model, while the other model parameters have been pre-learned and frozen. This can improve the learning efficiency of the style enhancement parameter.

[0085] In some embodiments, the above-mentioned loss function includes a first loss function and a second loss function; thus, based on the difference between the second enhanced image feature and the first enhanced image feature, the value of the loss function can be determined by the following steps: based on the difference between the second enhanced image feature and the first enhanced image feature, the value of the first loss function is determined; based on the difference between the second enhanced image feature and the image feature, the value of the second loss function is determined; the value of the first loss function and the value of the second loss function are added to obtain the value of the loss function.

[0086] Here, the difference between the second enhanced image feature and the first enhanced image feature can be first determined. Then, based on the difference between the second enhanced image feature and the first enhanced image feature, the value of the first loss function can be determined. This allows the first loss function to constrain the second enhanced image feature to approach the first enhanced image feature. Next, the difference between the second enhanced image feature and the image feature can be determined, and based on the difference between the second enhanced image feature and the image feature, the value of the second loss function can be determined. This allows the second loss function to constrain the distance between the second enhanced image feature and the unstyle-enhanced image feature to not exceed a distance threshold. Thus, through the combined constraints of the first and second loss functions, it is possible to ensure that the second enhanced image feature based on the style enhancement parameter approaches the first enhanced image feature guided by the text, and also ensure that the distance between the second enhanced image feature based on the style enhancement parameter and the unenhanced image feature does not exceed the distance threshold. This improves the learning accuracy of the style enhancement parameter, enabling the learned target style enhancement parameter to more accurately enhance the image features in the target style, thereby improving the style enhancement effect of the target style. It should be noted that when determining the value of the second loss function, it can also be determined based on the difference between the second enhanced image features and the above-mentioned encoded image features, because both the image features and the encoded image features are image features of the image sample that have not been style enhanced.

[0087] As an example, see FIG5, which is a schematic diagram of the structure of the machine learning model provided in the embodiment of the present application. Here, the machine learning model includes: a text encoder (Text encoder), a first image encoder (Visual encoder1, denoted as V a ), the second image encoder (Visual encoder2, denoted as V b Specifically, the first description text (i.e., a photo taken in bright enviroment) is encoded by a text encoder to obtain the first text feature (i.e., T source ), and the second description text (i.e., a photo taken in yellow ambiant light) is encoded by the text encoder to obtain the second text feature (i.e., T style ); Subtract the first text feature from the second text feature to obtain the text feature difference (i.e., ΔT = T style -T source ); The image sample I is encoded by the first image encoder to obtain the image feature (ie F = V a (I)), and the image features are encoded by the second image encoder to obtain the encoded image features (i.e., V b(F)); the text feature difference ΔT and the encoded image feature V b (F) is added to obtain the first enhanced image feature (i.e., F gt =V b (F) + ΔT); add the style enhancement parameter A0 and the image feature F to obtain a first addition result, and perform image encoding on the first addition result to obtain a second enhanced image feature (i.e., F pred =V b (F+A0)).

[0088] Based on this, the value of the loss function can be determined by the loss function shown in the following formula (1), and then based on the value of the loss function, the style enhancement parameter A0 is updated to obtain the target style enhancement parameter A: L = 1-cos(F pred ,F gt )+L1(F pred ,V b (F)); Formula (1)

[0089] Among them, L is the value of the loss function, 1-cos(F pred ,F gt ) is the value of the first loss function, L1(F pred ,V b (F)) is the value of the second loss function. It should be noted that ΔI=V in FIG5 b (F+A)-V b (F), the value L of the loss function is also equivalent to constraining ΔI to approach ΔT; since the learning process of updating the style enhancement parameter A0 to the target style enhancement parameter A is often a multi-round iterative process, and a style enhancement parameter is obtained in each round, the style enhancement parameter shown in Figure 5 refers to the style enhancement parameter and the style enhancement parameter output in each round during the learning process. The style enhancement parameter output in the last round is the target style enhancement parameter A.

[0090] Step 104: When training the liveness detection model, the image features are style enhanced based on the target style enhancement parameters to obtain the target enhanced image features of the image sample, and the liveness detection model is trained based on the target enhanced image features to obtain a trained liveness detection model.

[0091] Thus, after the above steps 101 to 103, the learned target style enhancement parameters are obtained. In step 104, when it is necessary to train the liveness detection model, the image features of the image sample can be style-enhanced based on the target style enhancement parameters to obtain the target enhanced image features, and then the liveness detection model can be trained based on the target enhanced image features to obtain the trained liveness detection model. In this way, it is possible to add image features of different styles to the training samples for training the liveness detection model, thereby improving the training effect of the liveness detection model through the image features of the diverse styles of the training samples, improving the generalization ability of the trained liveness detection model, and thus improving the liveness detection accuracy in the liveness detection scenario.

[0092] In some embodiments, based on the target style enhancement parameters, the image features can be style-enhanced by the following steps to obtain the target enhanced image features of the image sample: fusing the target style enhancement parameters with the image features to obtain a second fusion result; performing image encoding on the second fusion result to obtain the target enhanced image features of the image sample. Here, fusing the target style enhancement parameters with the image features can be adding the target style enhancement parameters and the image features, or multiplying the target style enhancement parameters and the image features. When performing image encoding on the second fusion result, the second fusion result can be image-encoded using a pre-trained second image encoder to obtain the target enhanced image features. The second image encoder can be constructed based on a convolutional neural network, a recurrent neural network, a Transformer network, etc., and trained using image data; for example, the second image encoder can include multiple cascaded image encoding layers, and the second fusion result is downsampled by the multiple cascaded image encoding layers to obtain the target enhanced image features. This can make the second enhanced image features closer to the target style, thereby improving the style enhancement effect.

[0093] In some embodiments, before performing style enhancement on image features based on target style enhancement parameters to obtain target enhanced image features of image samples, the following steps may also be performed: obtaining a sample selection method; selecting an image sample to be style enhanced from an image sample set where the image sample is located according to the sample selection method; based on this, based on the target style enhancement parameters, the image features may be style enhanced through the following steps to obtain target enhanced image features of the image sample: when the image sample having the image feature is the image sample to be enhanced, performing style enhancement on the image feature based on the target style enhancement parameters to obtain target enhanced image features of the image sample.

[0094] Here, the image samples to be enhanced (i.e., the image samples that need to be style enhanced) can be all the image samples in the image sample set, or they can be part of the image samples in the image sample set. When the image samples to be enhanced are part of the image samples in the image sample set, part of the image samples can be selected from the image sample set as the image samples to be enhanced according to a sample selection method. For example, the sample selection method may include: presetting the ratio of the number of image samples to be enhanced to the total number of image samples in the image sample set, and then randomly selecting image samples from the image sample set as the image samples to be enhanced according to the ratio; or presetting selection conditions for the image samples to be enhanced, and then selecting image samples that meet the selection conditions as the image samples to be enhanced.

[0095] Based on this, when style enhancement is performed on the image features of an image sample, it is possible to first determine whether the image sample is an image sample to be enhanced; if the image sample is an image sample to be enhanced, the image features are style enhanced based on the target style enhancement parameters to obtain the target enhanced image features of the image sample, and then the liveness detection model is trained based on the target enhanced image features; if the image sample does not belong to the set of image samples to be enhanced, the liveness detection model is trained directly based on the image features. Of course, if the image features are style enhanced and the corresponding target enhanced image features are used to train the liveness detection model, the image features that have not been style enhanced can also be used to train the liveness detection model.

[0096] In this way, image features of different styles can be added to all or part of the image samples used to train the liveness detection model, increasing the diversity of the image features of the training samples. This, in turn, improves the training effect of the liveness detection model through the diverse image features of the training samples, enhances the generalization ability of the trained liveness detection model, and thus improves the accuracy of liveness detection in liveness detection scenarios. At the same time, by selecting image samples to be enhanced according to the sample selection method, the data processing load for style enhancement can be reduced, improving the efficiency of style enhancement, and thus improving the model training efficiency of the liveness detection model.

[0097] When the image sample with the image feature is the image sample to be enhanced, after obtaining the target enhanced image feature, the liveness detection model is trained based on the target enhanced image feature. When the image sample with the image feature is not the image sample to be enhanced, the liveness detection model is trained directly based on the image feature. In this way, combining the target enhanced image feature with the image feature to train the liveness detection model can not only improve the prediction accuracy of the trained liveness detection model, but also improve the model training efficiency of the liveness detection model.

[0098] In some embodiments, based on the target enhanced image features, a liveness detection model can be trained by the following steps to obtain a trained liveness detection model: calling the liveness detection model, performing liveness detection based on the target enhanced image features, and obtaining a liveness detection result for the image sample; training the liveness detection model based on the difference between the liveness detection result and the sample label of the image sample to obtain a trained liveness detection model. Here, in some examples, the liveness detection result can indicate whether the biological object sample in the image sample is a true live object. Furthermore, based on the difference between the liveness detection result and the sample label of the image sample, the value of the liveness detection model's loss function can be determined; when the value of the loss function exceeds a loss threshold, an error signal of the liveness detection model is determined based on the loss function, and the error signal is back-propagated through the liveness detection model. Thus, during the back-propagation of the error information, the model parameters of the liveness detection model are updated to train the liveness detection model and obtain a trained liveness detection model.

[0099] It should be noted that when training a liveness detection model, the liveness detection model is iteratively trained, that is, the liveness detection model is trained in multiple rounds, thereby updating the model parameters of the liveness detection model output in the previous round of training to the model parameters of the liveness detection model output in the current round of training, thereby obtaining the liveness detection model obtained in the current round of training. In addition, during each round of training, after obtaining the liveness detection model obtained in the current round of training, it is also determined whether the liveness detection model has achieved the training target. The training target can be that the verification index (such as the error and accuracy of liveness detection) of the liveness detection model on the validation set reaches the index threshold (which can be pre-set), or that the number of training rounds reaches the round number threshold (which can be pre-set), etc. When it is determined that the liveness detection model has not achieved the training target, the liveness detection model is trained in the next round; when it is determined that the liveness detection model has achieved the training target, the training is stopped, and the liveness detection model obtained in the last round of training is output as the final liveness detection model. In this way, the training accuracy of the liveness detection model can be improved, the training efficiency of the liveness detection model can be guaranteed, and the training effect of the liveness detection model can be improved.

[0100] In some embodiments, before training the liveness detection model based on the target enhanced image features to obtain the trained liveness detection model, the following steps may be performed: training the liveness detection model based on the image features to obtain an intermediate liveness detection model; based on this, based on the target enhanced image features, the liveness detection model may be trained by performing the following steps to obtain the trained liveness detection model: training the intermediate liveness detection model based on the target enhanced image features to obtain the trained liveness detection model.

[0101] Here, the liveness detection model can be trained not only through target-enhanced image features but also through original image features. The training process of the intermediate liveness detection model through target-enhanced image features can refer to the training process of the liveness detection model through target-enhanced image features. In this way, 1) the diversity of the image features of the training samples is increased, thereby improving the training effect of the liveness detection model through the image features of the diverse styles of the training samples; 2) the liveness detection ability of the liveness detection model for the original style image samples is guaranteed; in summary, the generalization ability of the trained liveness detection model is improved, thereby improving the liveness detection accuracy in the liveness detection scenario.

[0102] In some embodiments, before calling a liveness detection model and performing liveness detection based on target enhanced image features to obtain a liveness detection result for an image sample, the following steps may be performed: obtaining multiple category texts, the multiple category texts including: a first text representing a liveness category, and a second text representing each non-liveness category in at least one non-liveness category. Based on this, referring to FIG6 , calling a liveness detection model may be performed based on target enhanced image features to obtain a liveness detection result for an image sample through the following steps: Step 401: calling a liveness detection model to perform text encoding on each category text to obtain category text features for each category text; Step 402: calling a liveness detection model to perform liveness detection based on target enhanced image features and each category text features to obtain a predicted probability that a biological object sample in the image sample belongs to each category; wherein the categories include a liveness category and at least one non-liveness category, and the liveness detection result for the image sample includes a predicted probability that the biological object sample belongs to each category.

[0103] Here, when performing liveness detection based on target enhanced image features, it is also necessary to obtain multiple category texts. It should be noted that the multiple category texts include: a first text representing a liveness category, and a second text representing each non-liveness category; wherein the first text is used to describe biological object images of the liveness category (i.e., including images of real live objects, obtained by image acquisition of real live objects); and the second text is used to describe biological object images of the non-liveness category (i.e., not including images of real live objects, such as obtained by image acquisition of biological objects in videos or pictures played on the screen, or pre-printed biological objects, etc.). The number of non-liveness categories is at least one, such as the category of video biological objects, the category of picture biological objects, etc.

[0104] Continuing, after obtaining multiple categorical texts, in step 401, a liveness detection model is invoked to perform text encoding on each categorical text to obtain categorical text features for each categorical text. For example, text encoding can be achieved using a pre-trained text encoder. For example, the text encoder can be constructed based on a convolutional neural network, a recurrent neural network, a Transformer network, or the like, and trained using text data. For example, the text encoder can include multiple cascaded text encoding layers, through which the categorical text is downsampled to obtain categorical text features. This can improve the extraction of categorical text features and make the text features more expressive of the categorical text. In step 402, a liveness detection model is invoked to perform liveness detection based on the target enhanced image features and the categorical text features, obtaining a predicted probability that the biological object sample in the image sample belongs to each categorical category. It should be noted that the categories include a liveness category and at least one non-liveness category. In this embodiment, the liveness detection result of the image sample includes the predicted probability that the biological object sample belongs to each categorical category.

[0105] In this way, through steps 401-402, when calling the liveness detection model for liveness detection, the category text features representing the liveness category and the non-liveness category can be combined for liveness detection. The category text features can express the characteristics of the liveness category and the characteristics of the non-liveness category, thereby improving the liveness detection accuracy of the liveness detection model when performing liveness detection with reference to the category text features.

[0106] In some embodiments, referring to FIG7 , liveness detection can be performed based on target enhanced image features and text features of each category through the following steps to obtain the predicted probability that the biological object sample in the image sample belongs to each category: Step 501, feature dimensionality reduction processing is performed on the target enhanced image features to obtain intermediate enhanced image features, and feature dimensionality increase is performed on the intermediate enhanced image features to obtain target image features; Step 502, the target enhanced image features and the target image features are added to obtain added features; Step 503, liveness detection is performed based on the added features and text features of each category to obtain the predicted probability that the biological object sample in the image sample belongs to each category.

[0107] Here, in step 501, the target enhanced image features are first subjected to feature dimensionality reduction processing (for example, it can be implemented by a pre-built encoder encoder) to obtain intermediate enhanced image features, and then the intermediate enhanced image features are subjected to feature dimensionality increase processing (for example, it can be implemented by a pre-built decoder decoder) to obtain target image features. Feature dimensionality reduction processing refers to reducing the feature dimension of the feature, so that the feature dimension of the intermediate enhanced image feature is lower than the feature dimension of the target enhanced image feature; feature dimensionality increase processing refers to increasing the feature dimension of the feature, so that the feature dimension of the target image feature is higher than the feature dimension of the intermediate enhanced image feature. Then, in step 502, the target enhanced image feature and the target image feature are added to obtain the added feature. Finally, in step 503, liveness detection is performed based on the added feature and the text features of each category to obtain the predicted probability that the biological object sample in the image sample belongs to each category. Specifically, based on the added feature and the text features of each category, liveness detection can be performed, and the predicted probability that the biological object sample belongs to each category can be obtained by the following formula (2):

[0108] Among them, p i ′ is the predicted probability that the biological object sample belongs to the i-th category, T i is the category text feature of category i, E v is the additive feature, τ is the preset temperature coefficient, N is the number of categories, T j is the category text feature of the j-th category.

[0109] Based on this, after obtaining the liveness detection result (i.e., the predicted probability that the biological object sample belongs to each category), the loss function value L(θ) of the liveness detection model can be determined by the following formula (3):

[0110] Among them, y i is the true label of the biological object in the image sample being the i-th category (1 represents that the biological object sample in the image sample is the i-th category, and 0 represents that the biological object sample in the image sample is not the i-th category).

[0111] Therefore, based on the value L(θ) of the loss function, the liveness detection model is trained to obtain a trained liveness detection model.

[0112] In some embodiments, the following steps can be used to perform feature dimensionality reduction processing on the target enhanced image features to obtain intermediate enhanced image features, and then perform feature dimensionality increase processing on the intermediate enhanced image features to obtain target image features: call the feature dimension processing layer of the liveness detection model, perform feature dimensionality reduction processing on the target enhanced image features to obtain intermediate enhanced image features, and then perform feature dimensionality increase processing on the intermediate enhanced image features to obtain target image features. Based on this, based on the difference between the liveness detection results and the sample labels of the image samples, the liveness detection model can be trained to obtain a trained liveness detection model by the following steps: based on the difference between the liveness detection results and the sample labels of the image samples, the model parameters of the feature dimension processing layer are updated to train the liveness detection model to obtain a trained liveness detection model.

[0113] Here, when constructing a liveness detection model, the model parameters to be learned can be set in the feature dimension processing layer of the liveness detection model, and other model parameters can be pre-learned and frozen. The model parameters to be learned in the feature dimension processing layer include a first model parameter for feature dimensionality reduction processing and a second model parameter for feature dimensionality increase processing. In this way, rapid training of the liveness detection model can be achieved and the model training efficiency can be improved. It should be noted that the feature dimension processing layer is used to perform dimensionality reduction processing and dimensionality increase processing on the target enhanced image features. Specifically, the feature dimension processing layer includes a dimensionality reduction processing layer and a dimensionality increase processing layer; wherein, the dimensionality reduction processing layer is used to perform dimensionality reduction processing on the target enhanced image features to obtain intermediate enhanced image features, so that the secondary features in the target enhanced image features can be filtered out and the main features can be retained, thereby reducing the computing resources occupied by processing the secondary features and obtaining features that more accurately express the image. The dimensionality increase processing layer is used to increase the dimensionality of the target enhanced image features to obtain the target image features. In this way, the factors affecting the model detection results can be increased based on the main features of the target enhanced image features, so that the model can fit better, solve the problem of underfitting of the model, improve the detection accuracy of the model, and thus improve the model training effect.

[0114] As an example, see FIG8 , which is a training diagram of a liveness detection model provided in an embodiment of the present application. Here, the liveness detection model includes a text encoder (Text encoder), a first image encoder (Visual encoder1), a second image encoder (Visual encoder2), a feature dimension processing layer (including a dimensionality reduction processing layer (including the model parameters to be learned W d ) and the dimension-raising layer (including the model parameters to be learned W u)), feature splicing layer, and liveness detection layer. In this way, based on the liveness detection model shown in Figure 8, the text encoder can be called first to perform text encoding on each category text (as shown in Figure 8, there are 3 categories (i.e., the number of categories N=3) of category texts, including A photo of a normal face (corresponding to the liveness category), A photo of a printed face (corresponding to the non-liveness category, i.e., the category of biological objects in the picture), and A photo of a sereen play face (corresponding to the non-liveness category, i.e., the category of biological objects in the video)) to obtain the category text features of each category text, including category text features T1, T2, and T3. Here, T1 corresponds to 1, indicating that T1 is a category text of the liveness category, T2 corresponds to 0, indicating that T2 is a category text of the non-liveness category, and T3 corresponds to 0, indicating that T3 is a category text of the non-liveness category. Then call the first image encoder to perform image encoding on the image sample to obtain image features. Then perform one of the following operations (including operation (1) and operation (2)):

[0115] Operation (1) If it is determined that the current image sample is not the image sample to be enhanced, then the second image encoder is called to perform image encoding on the image features to obtain encoded image features; the feature dimension processing layer is called to process the encoded image features to obtain target encoded image features; the feature splicing layer is called to add the encoded image features and the target encoded image features to obtain added features, so as to add the features of the image sample that have not been processed by the feature dimension to the features processed by the feature dimension, thereby improving the expression accuracy of the features used for liveness detection for the image sample; the liveness detection layer is called to perform liveness detection based on the added features and the text features of each category to obtain the liveness detection result of the image sample; and then based on the difference between the liveness detection result and the sample label of the image sample, the model parameter W of the feature dimension processing layer is updated. d and W d , train the liveness detection model to obtain a trained liveness detection model.

[0116] Operation (2) If it is determined that the current image sample is an image sample to be enhanced, then the target style enhancement parameter is added to the image feature to obtain a second addition result, and the second addition result is image encoded to obtain the target enhanced image feature of the image sample; the target enhanced image feature is subjected to feature dimensionality reduction processing to obtain an intermediate enhanced image feature, and the intermediate enhanced image feature is subjected to feature dimensionality increase processing to obtain the target image feature; the feature splicing layer is called to add the target enhanced image feature and the target image feature to obtain the added feature, so as to add the features of the image sample that have not been processed by the feature dimension to the features processed by the feature dimension, thereby improving the expression accuracy of the features used for liveness detection for the image sample; the liveness detection layer is called to perform liveness detection based on the added features and the text features of each category to obtain the liveness detection result of the image sample; and then based on the difference between the liveness detection result and the sample label of the image sample, the model parameter W of the feature dimension processing layer is updated. d and W u , to train the liveness detection model and obtain a trained liveness detection model.

[0117] Applying the above-mentioned embodiments of the present application, first obtain an image sample for training a liveness detection model, a first description text of the image sample, and a second description text of the image sample in a target style; then, based on the first description text and the second description text, style enhancement is performed on the image features of the image sample to obtain a first enhanced image feature, and based on the style enhancement parameters, style enhancement is performed on the image features to obtain a second enhanced image feature; then, based on the difference between the second enhanced image feature and the first enhanced image feature, the style enhancement parameters are updated to obtain a target style enhancement parameter; in this way, when training the liveness detection model, style enhancement is performed on the image features based on the target style enhancement parameters to obtain the target enhanced image feature of the image sample, and the liveness detection model is trained based on the target enhanced image feature to obtain a trained liveness detection model. Here, through text guidance, a target style enhancement parameter is learned, and the target style enhancement parameter is used to perform style enhancement on the image features of the image sample to obtain the target enhanced image feature, thereby realizing the style diversification expansion of the image features of the existing image samples and enriching the image features of the image samples. When the liveness detection model is trained through the target enhanced image features of the image samples, the training effect of the liveness detection model can be improved, thereby improving the model generalization ability and liveness detection accuracy of the trained liveness detection model.

[0118] The following describes the liveness detection method provided by the embodiment of the present application. The liveness detection method provided by the embodiment of the present application is implemented by an electronic device, for example, it can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. Therefore, the execution entity of each step will not be repeated below. Referring to Figure 9, Figure 9 is a flow chart of the liveness detection method provided by the embodiment of the present application. The liveness detection method provided by the embodiment of the present application includes:

[0119] Step 201: Acquire an image to be detected including a biological object and multiple category texts.

[0120] The multiple category texts include: a first text representing a living body category, and a second text representing each non-living body category in at least one non-living body category.

[0121] In step 201, when liveness detection is required for a biological object, a user can trigger a liveness detection command on an electronic device. In response to the liveness detection command, the electronic device acquires an image to be detected that includes the biological object. Here, an image capture device is used to capture the biological object to be detected and the background in the liveness detection scene, thereby obtaining the image to be detected. Furthermore, during the capture process, the biological object can perform corresponding actions in accordance with the liveness detection requirements. For example, if the biological object is a hand, it can perform corresponding gestures in accordance with the liveness detection requirements. Alternatively, if the biological object is a face, it can perform actions such as blinking, opening the mouth, shaking the head, and nodding in accordance with the liveness detection requirements. In practical applications, the target image captured by the image capture device can be directly used as the image to be detected, or the target image captured by the image capture device can be cropped to obtain the image to be detected; for example, the target image can be used to detect biological objects, and the first area where the biological object is located in the target image can be determined, and the first area can be cropped from the target image as the image to be detected; of course, the image to be detected in the living body detection scene can include the background where the biological object is located. Therefore, after determining the first area, the first area can be used as the center and expanded by a preset multiple to obtain a second area, and the second area can be cropped from the target image as the image to be detected. In this way, the image to be detected can include part of the background.

[0122] In step 201, multiple category texts are also obtained. It should be noted that the multiple category texts include: a first text representing a living body category, and a second text representing each non-living body category; wherein the first text is used to describe biological object images of the living body category (i.e., including images of real living objects, obtained by image capture of real living objects); and the second text is used to describe biological object images of the non-living body category (i.e., excluding images of real living objects, such as biological objects in videos or pictures played on the screen, or pre-printed biological objects, etc.). The number of non-living body categories is at least one, such as categories of biological objects in videos, categories of biological objects in pictures, etc.

[0123] Step 202: calling the liveness detection model, performing text encoding on each category of text to obtain category text features of each category of text, and performing feature extraction on the image to be detected to obtain features of the image to be detected.

[0124] In step 202, the liveness detection model is first called to perform text encoding on each category of text to obtain the category text features of each category of text; then the liveness detection model is called to perform feature extraction on the image to be detected to obtain the image features to be detected. As an example, refer to Figure 10, which is a structural diagram of the liveness detection model provided by an embodiment of the present application. Here, the structure of the liveness detection model is only an example; the liveness detection model includes a text encoder (Text encoder), a first image encoder (Visual encoder1), a second image encoder (Visual encoder2), a feature dimension processing layer (including a dimensionality reduction processing layer (including model parameters W d ) and dimensionality-raising layer (including model parameters W u )), feature splicing layer, and living body detection layer.

[0125] In this way, based on the liveness detection model shown in Figure 10, the text encoder can first be called to perform text encoding on each category text (as shown in Figure 10, there are 3 categories (i.e., the number of categories N=3) of category texts, including A photo of a normal face (corresponding to the liveness category), A photo of a printed face (corresponding to the non-liveness category, i.e., the category of biological objects in the picture), and A photo of a sereen play face (corresponding to the non-liveness category, i.e., the category of biological objects in the video)) to obtain the category text features of each category text, including category text features T1, T2, and T3. Here, T1 corresponds to 1, indicating that T1 is a category text of the liveness category, T2 corresponds to 0, indicating that T2 is a category text of the non-liveness category, and T3 corresponds to 0, indicating that T3 is a category text of the non-liveness category. Then, the first image encoder is called to perform image encoding on the image to be detected to obtain the first image feature, and the second image encoder is called to perform image encoding on the first image feature to obtain the second image feature; the dimensionality reduction processing layer of the feature dimension processing layer is called to perform feature dimensionality reduction processing on the second image feature to obtain the third image feature, and the dimensionality increase processing layer of the feature dimension processing layer is called to perform feature dimensionality increase processing on the third image feature to obtain the fourth image feature; the second image feature and the fourth image feature are spliced ​​(for example, added) to obtain the image feature to be detected.

[0126] Step 203: calling the liveness detection model, performing liveness detection based on the features of the image to be detected and the features of the texts of each category, and obtaining a liveness detection result of the image to be detected.

[0127] The liveness detection result includes the predicted probability that the biological object belongs to each category, and the categories include a liveness category and at least one non-liveness category.

[0128] Among them, the liveness detection model is trained based on the liveness detection model training method provided in the embodiment of the present application.

[0129] In step 203, the liveness detection model is called to perform liveness detection based on the features of the image to be detected and the text features of each category, thereby obtaining a liveness detection result for the image to be detected. Here, the liveness detection result includes the predicted probability of the biological object belonging to each category (including the liveness category and at least one non-liveness category). In this way, based on the predicted probability of the biological object belonging to each category in the liveness detection result, it can be determined whether the biological object is a real live object; for example, the category corresponding to the maximum predicted probability is used as the target category to which the biological object belongs. If the target category is the liveness category, then the biological object is a real live object; if the target category is the non-liveness category, then the biological object is not a real live object.

[0130] Continuing with FIG10 , the liveness detection layer can be called to perform liveness detection based on the features of the image to be detected and the features of the text of each category, thereby obtaining a liveness detection result of the image to be detected. Specifically, the liveness detection layer is called to perform liveness detection based on the features of the image to be detected and the features of the text of each category, and the liveness detection result of the image to be detected can be obtained by the following formula (2):

[0131] Among them, p i is the predicted probability that the biological object belongs to the i-th category, T i is the category text feature of category i, F v is the image feature to be detected, τ is the preset temperature coefficient, N is the number of categories, T j is the category text feature of the j-th category.

[0132] Applying the above embodiments of the present application, 1) the liveness detection model is obtained by training the target enhanced image features of the image samples. The target enhanced image features are obtained by style enhancement of the image features of the image samples based on the target style enhancement parameters. The target style enhancement parameters are learned in a text-guided manner, which can diversify the styles of the image features of the image samples. Therefore, when the liveness detection model is trained by the target enhanced image features of the image samples, the training effect of the liveness detection model can be improved, thereby improving the model generalization ability and liveness detection accuracy of the trained liveness detection model. 2) The liveness detection of the image to be detected is performed in combination with the category text of each category (including the liveness category and at least one non-liveness category), which utilizes the strong semantic properties of the text modality and further improves the liveness detection accuracy of the model.

[0133] In an exemplary scenario, an embodiment of the present application can be applied to a biometric payment scenario. Specifically, for an image sample (such as a facial image sample) collected from a biometric payment scenario, a first description text of the image sample and a second description text of the image sample in a target style are obtained; based on the first description text and the second description text, the image features of the image sample are style enhanced to obtain a first enhanced image feature, and based on the style enhancement parameter, the image features are style enhanced to obtain a second enhanced image feature; based on the difference between the second enhanced image feature and the first enhanced image feature, the style enhancement parameter is updated to obtain a target style enhancement parameter; based on the target style enhancement parameter, the image features are style enhanced to obtain a target enhanced image feature of the image sample, and based on the target enhanced image feature, a liveness detection model is trained to obtain a trained liveness detection model; when performing biometric payment, an image to be detected including a biological object (such as a face) is obtained, the trained liveness detection model is called, liveness detection is performed on the image to be detected, and a liveness detection result is obtained; wherein the liveness detection result includes the predicted probability that the biological object in the image to be detected belongs to each category, and the category includes a liveness category and at least one non-liveness category. In this way, the predicted probability of the biological object belonging to each category in the liveness detection results can be used to determine whether the biological object is a genuine live object. For example, the category corresponding to the maximum predicted probability is used as the target category for the biological object. If the target category is live, the biological object is a genuine live object. If the target category is non-live, the biological object is not a genuine live object. If the biological object is a genuine live object, the payment is completed if other payment verifications are successful. This can improve the accuracy of liveness detection and thus enhance the security of biometric payment.

[0134] In an exemplary scenario, the embodiment of the present application can be applied to a biometric unlocking (such as access control system unlocking, terminal device unlocking, etc.) scenario. Specifically, for an image sample (such as a facial image sample) collected from a biometric unlocking scenario, a first description text of the image sample and a second description text of the image sample in a target style are obtained; based on the first description text and the second description text, the image feature of the image sample is style-enhanced to obtain a first enhanced image feature, and based on the style enhancement parameter, the image feature is style-enhanced to obtain a second enhanced image feature; based on the difference between the second enhanced image feature and the first enhanced image feature, the image feature is updated. The method comprises the following steps: a style enhancement parameter is used to obtain a target style enhancement parameter; based on the target style enhancement parameter, the image features are style enhanced to obtain the target enhanced image features of the image sample; and based on the target enhanced image features, a liveness detection model is trained to obtain a trained liveness detection model; when performing biometric unlocking, an image to be detected including a biological object (such as a face) is obtained, the trained liveness detection model is called, and liveness detection is performed on the image to be detected to obtain a liveness detection result; wherein the liveness detection result includes a predicted probability that the biological object in the image to be detected belongs to each category, the category including a liveness category and at least one non-liveness category. In this way, whether the biological object is a real live object can be determined based on the predicted probability that the biological object belongs to each category in the liveness detection result; for example, the category corresponding to the maximum predicted probability is used as the target category to which the biological object belongs. If the target category is the liveness category, then the biological object is a real live object; if the target category is the non-liveness category, then the biological object is not a real live object. If the biological object is a real live object, then unlocking is performed if other unlocking verifications are successful. In this way, the accuracy of liveness detection can be improved, thereby improving the security of biometric unlocking.

[0135] In an exemplary scenario, the embodiments of the present application can be applied to a biometric identity verification scenario. Specifically, for an image sample (such as a facial image sample) collected from the biometric identity verification scenario, a first description text of the image sample and a second description text of the image sample in a target style are obtained; based on the first description text and the second description text, the image features of the image sample are style enhanced to obtain a first enhanced image feature, and based on the style enhancement parameter, the image features are style enhanced to obtain a second enhanced image feature; based on the difference between the second enhanced image feature and the first enhanced image feature, the style enhancement parameter is updated to obtain To the target style enhancement parameters; based on the target style enhancement parameters, the image features are style enhanced to obtain the target enhanced image features of the image sample, and based on the target enhanced image features, the liveness detection model is trained to obtain a trained liveness detection model; when performing biometric identity verification, an image to be detected including a biological object (such as a face) is obtained, the trained liveness detection model is called, and liveness detection is performed on the image to be detected to obtain a liveness detection result; wherein the liveness detection result includes the predicted probability that the biological object in the image to be detected belongs to each category, and the category includes a liveness category and at least one non-liveness category. In this way, whether the biological object is a real live object can be determined based on the predicted probability of the biological object belonging to each category in the liveness detection result; for example, the category corresponding to the maximum predicted probability is used as the target category to which the biological object belongs. If the target category is the liveness category, then the biological object is a real live object; if the target category is the non-liveness category, then the biological object is not a real live object. If the biological object is a real live object, then if other identity verifications are also successful, the identity verification is confirmed to have passed. In this way, the accuracy of liveness detection can be improved, thereby improving the security of biometric identity verification.

[0136] The following takes face liveness detection as an example to illustrate the exemplary application of the embodiment of the present application in an actual application scenario. Face liveness detection is a key step in the face recognition process. With the implementation of face recognition systems in production and life, more and more liveness attack data are constantly generated. However, these data may differ from the attack data in the training sample set in terms of domain information such as face, lighting, background, and attack type, that is, the data distribution of the two is different. Therefore, directly migrating the liveness detection model obtained based on the training sample set to the test data may lead to a decrease in the model detection accuracy due to insufficient generalization ability. In related technologies, in order to improve the generalization performance of the model, most of the domain generalization technologies based on a single modality (i.e., image) are used, which has limitations.

[0137] Therefore, the embodiment of the present application provides a text-guided, efficient domain-generalized liveness detection method, which further improves the generalization ability of the model by introducing text modality and leveraging the text's ability to model style and content. Specifically, some enhanced image features that are different from the style of the training sample set are generated under the guidance of text modality, and the source domain image features are migrated to different specific styles, thereby enriching the image features of the model's training data, improving the style diversity of the image features of the training samples, and thus improving the generalization effect of the model. In addition, the embodiment of the present application also proposes a multimodal matching classification mechanism for liveness detection, which utilizes the strong semantic properties of the text modality, and the effect exceeds the single-modal method in the related art.

[0138] In some exemplary scenarios, liveness detection is often combined with other technologies, such as facial authentication. Liveness detection, as the first line of defense, controls a crucial aspect of authentication security. Currently, facial authentication has been applied in a variety of services, such as remote authentication, facial payment, remote authentication, and access control systems. For example, in Bank X's remote account opening process, remote facial authentication is used to verify the account holder's true identity. Liveness detection is also incorporated into this process. The specific process is as follows: First, the user captures an image of their face through a camera on the application frontend. The frontend transmits this image to the backend, which invokes a liveness detection model. The remote authentication model performs a liveness check and returns the result to the frontend. If the user is deemed live, the authentication passes; otherwise, the authentication fails. For example, facial authentication also plays a crucial role in facial payment, where liveness detection is a crucial element in ensuring payment security. High-precision liveness detection methods can reject transactions attempted by unauthorized attacks, ensuring transaction security and protecting user interests. For example, in an access control system, in order to improve authentication efficiency, the access control system directly obtains a face image at the front end, sends it to a liveness detection model packaged at the front end for direct detection, and feeds back the liveness detection result.

[0139] The following is a detailed description. First, the learning process of the target style enhancement parameter A is described. Referring to Figure 5, the model at this stage includes: the first image encoder V a , the second image encoder V b , text encoder, style enhancement parameter A. In this stage, only the target style enhancement parameter A is learned, and the parameters of the rest are completely frozen.

[0140] In this way, the first description text (i.e., a photo taken in bright environment) is encoded by the text encoder to obtain the first text feature (i.e., T source), and the second description text (i.e., a photo taken in yellow ambiant light) is encoded by the text encoder to obtain the second text feature (i.e., T style ); Subtract the first text feature from the second text feature to obtain the text feature difference (ie, ΔT = T style -T source ); The image sample I is encoded by the first image encoder to obtain the image feature (ie F = V a (I)), and the image features are encoded by the second image encoder to obtain the encoded image features (i.e., V b (F)); the text feature difference ΔT and the encoded image feature V b (F) is added to obtain the first enhanced image feature (i.e., F gt =V b (F) + ΔT); add the style enhancement parameter A0 and the image feature F to obtain a first addition result, and perform image encoding on the first addition result to obtain a second enhanced image feature (i.e., F pred =V b (F+A0)).

[0141] Based on this, the loss function value can be determined by the loss function shown below, and then based on the value of the loss function, the style enhancement parameter A0 is updated to obtain the target style enhancement parameter A: L = 1-cos(F pred ,F gt )+L1(F pred ,V b (F)); Formula (1)

[0142] Among them, L is the value of the loss function, 1-cos(F pred ,F gt ) is the value of the first loss function, which can constrain the second enhanced image features to tend to the first enhanced image features, L1(F pred ,V b (F)) is the value of the second loss function, which can constrain the distance between the second enhanced image feature and the image feature without style enhancement to not exceed the distance threshold.

[0143] Second, the training process of the living body detection model is described. In the embodiment of the present application, as shown in FIG8 , (1) the model at this stage includes: a first image encoder V a , the second image encoder V b , text encoder, target style enhancement parameter A, efficient training module Adapter (that is, the above-mentioned feature dimension processing layer, including the model parameters to be learned W d and W u). This stage only trains the model parameters W to be learned of the efficient training module Adapter. d and W u , the parameters of the remaining modules are completely frozen. (2) Image-text spatial consistency. In order to improve the robustness of the model, the embodiment of the present application fine-tunes the consistency of image-text features, performs liveness detection by semantic matching, and utilizes general semantic knowledge to facilitate model generalization. (3) Considering that fully fine-tuning the pre-trained model requires a large training cost, the embodiment of the present application designs an efficient training module Adapter to fine-tune the liveness detection task, using only a small number of parameters to achieve the effect of full fine-tuning. (4) Through the joint training of target enhanced image features and source domain data that have been enhanced in multiple styles, the model inputs more diverse image features, so the model is more robust and generalizable. (5) The model training process is as follows:

[0144] (5.1) First, the image sample I passes through the first image encoder to obtain the image feature F = V a (I) In terms of text, we set the category text that can represent the living and non-living categories (i.e., attack data), denoted as t. After the text encoder is used, we obtain the category text features, denoted as T, whose dimension is [N, C], where N is the number of categories and C is the number of channels.

[0145] (5.2) Determine whether the image feature F of the image sample I needs to be style enhanced. If it is enhanced, the image feature output by Adapter is Adapter(V b (F+A)), otherwise Adapter(V b (F)). Thus, the predicted probability of the image sample I is obtained Among them, p i ′ is the predicted probability that the image sample belongs to the i-th category, T i is the category text feature of category i, E v For the additive feature (i.e. Adapter (V b (F+A))+V b (F+A), or Adapter(V b (F))+(V b (F))), τ is the preset temperature coefficient, N is the number of categories, T j is the category text feature of the jth category. Thus, the loss function value L(θ) of the living body detection model is determined by the following formula (3):

[0146] Among them, y i is the true label of the image sample of the i-th category (1 represents that the image sample is of the i-th category, and 0 represents that the biological object sample in the image sample is not of the i-th category).

[0147] In some embodiments, the efficient training module herein can also be implemented through Side Adapter, Prompt Tuning, LoRA, Visual Prompt Tuning, and the like.

[0148] By applying the above-mentioned embodiments of the present application, the generalization ability of the liveness detection model on image features of various styles is enhanced, the training effect of the model is improved, and the liveness detection accuracy of the liveness detection model is thereby improved.

[0149] The following continues to describe an exemplary structure of the training device 555 of the liveness detection model provided by the embodiment of the present application as a software module. In some embodiments, as shown in FIG2 , the software modules in the training device 555 of the liveness detection model stored in the memory 550 may include: a first acquisition module 5551, configured to acquire an image sample for training the liveness detection model, and acquire a first description text of the image sample, and a second description text of the image sample under a target style; a style enhancement module 5552, configured to perform style enhancement on the image features of the image sample based on the first description text and the second description text, to obtain a style enhancement text of the image sample; a style enhancement module 5553, configured to enhance the style of the image features of the image sample based on the first description text and the second description text, to obtain a style enhancement text of the image sample; a style enhancement module 5554, configured to enhance the style of the image features of the image sample based on the first description text and the second description text, to obtain a style enhancement text of the image sample; a style enhancement module 5 ...; a style enhancement module 5555, configured to enhance the style of the image features of the image sample; to the first enhanced image feature, and based on the style enhancement parameter, perform style enhancement on the image feature to obtain the second enhanced image feature; the updating module 5553 is configured to update the style enhancement parameter based on the difference between the second enhanced image feature and the first enhanced image feature to obtain the target style enhancement parameter; the training module 5554 is configured to, when training the liveness detection model, perform style enhancement on the image feature based on the target style enhancement parameter to obtain the target enhanced image feature of the image sample, and train the liveness detection model based on the target enhanced image feature to obtain a trained liveness detection model.

[0150] In some embodiments, the style enhancement module 5552 is further configured to perform text encoding on the first description text to obtain a first text feature, and perform text encoding on the second description text to obtain a second text feature; perform image encoding on the image feature of the image sample to obtain an encoded image feature; and fuse the second text feature, the first text feature and the encoded image feature to obtain the first enhanced image feature.

[0151] In some embodiments, the style enhancement module 5552 is further configured to determine the text feature difference between the second text feature and the first text feature; and fuse the text feature difference with the encoded image feature to obtain the first enhanced image feature.

[0152] In some embodiments, the style enhancement module 5552 is further configured to fuse the style enhancement parameters and the image features to obtain a first fusion result; and perform image encoding on the first fusion result to obtain the second enhanced image features.

[0153] In some embodiments, the style enhancement parameters belong to a machine learning model; the update module 5553 is further configured to obtain a loss function of the machine learning model; determine a value of the loss function based on the difference between the second enhanced image feature and the first enhanced image feature; and update the style enhancement parameters of the machine learning model based on the value of the loss function to obtain the target style enhancement parameters.

[0154] In some embodiments, the loss function includes a first loss function and a second loss function; the update module 5553 is further configured to determine the value of the first loss function based on the difference between the second enhanced image feature and the first enhanced image feature; determine the value of the second loss function based on the difference between the second enhanced image feature and the image feature; and add the value of the first loss function and the value of the second loss function to obtain the value of the loss function.

[0155] In some embodiments, the training module 5554 is further configured to obtain a sample selection method before performing style enhancement on the image feature based on the target style enhancement parameter to obtain the target enhanced image feature of the image sample; according to the sample selection method, select the image sample to be style enhanced from the image sample set where the image sample is located; the training module 5554 is further configured to perform style enhancement on the image feature based on the target style enhancement parameter to obtain the target enhanced image feature of the image sample when the image sample having the image feature is the image sample to be enhanced.

[0156] In some embodiments, the training module 5554 is further configured to train the living body detection model based on the image features when the image sample having the image features is not the image sample to be enhanced.

[0157] In some embodiments, the training module 5554 is further configured to fuse the target style enhancement parameters with the image features to obtain a second fusion result; and perform image encoding on the second fusion result to obtain the target enhanced image features of the image sample.

[0158] In some embodiments, the training module 5554 is further configured to call the liveness detection model, perform liveness detection based on the target enhanced image features, and obtain the liveness detection result of the image sample; based on the difference between the liveness detection result and the sample label of the image sample, train the liveness detection model to obtain a trained liveness detection model.

[0159] In some embodiments, the training module 5554 is further configured to obtain a plurality of category texts before calling the liveness detection model, performing liveness detection based on the target enhanced image features, and obtaining the liveness detection result of the image sample. The plurality of category texts include: a first text representing a liveness category, and a second text representing each of the non-liveness categories in at least one non-liveness category; the training module 5554 is further configured to call the liveness detection model, perform text encoding on each of the category texts, and obtain category text features of each of the category texts; call the liveness detection model, perform liveness detection based on the target enhanced image features and each of the category text features, and obtain the predicted probability that the biological object sample in the image sample belongs to each category; wherein the categories include the liveness category and the at least one non-liveness category, and the liveness detection result of the image sample includes the predicted probability that the biological object sample belongs to each category.

[0160] In some embodiments, the training module 5554 is further configured to perform feature dimensionality reduction processing on the target enhanced image features to obtain intermediate enhanced image features, and perform feature dimensionality increase processing on the intermediate enhanced image features to obtain target image features; add the target enhanced image features and the target image features to obtain added features; perform liveness detection based on the added features and each of the category text features to obtain the predicted probability that the biological object sample in the image sample belongs to each category.

[0161] In some embodiments, the training module 5554 is further configured to call the feature dimension processing layer of the liveness detection model, perform feature dimensionality reduction processing on the target enhanced image features to obtain intermediate enhanced image features, and perform feature dimensionality increase processing on the intermediate enhanced image features to obtain target image features; the training module 5554 is further configured to update the model parameters of the feature dimension processing layer based on the difference between the liveness detection results and the sample labels of the image samples, so as to train the liveness detection model and obtain a trained liveness detection model.

[0162] In some embodiments, the training module 5554 is further configured to train the liveness detection model based on the image features to obtain an intermediate liveness detection model before training the liveness detection model based on the target enhanced image features to obtain a trained liveness detection model; the training module 5554 is further configured to train the intermediate liveness detection model based on the target enhanced image features to obtain a trained liveness detection model.

[0163] An embodiment of the present application also provides a liveness detection device, comprising: a second acquisition module, configured to acquire an image to be detected including a biological object and multiple category texts, wherein the multiple category texts include: a first text representing a liveness category, and a second text representing each of the non-liveness categories in at least one non-liveness category; a feature extraction module, configured to call a liveness detection model, perform text encoding on each of the category texts, obtain category text features of each of the category texts, and perform feature extraction on the image to be detected to obtain features of the image to be detected; a liveness detection module, configured to call the liveness detection model, perform liveness detection based on the features of the image to be detected and the features of each of the category texts, and obtain a liveness detection result of the image to be detected; wherein the liveness detection result includes a predicted probability that the biological object belongs to each category, and the categories include the liveness category and the at least one non-liveness category; wherein the liveness detection model is trained based on the training method of the liveness detection model provided in the embodiment of the present application.

[0164] It should be noted that the description of the device embodiments in this application is similar to the description of the method embodiments described above, and has similar beneficial effects as the method embodiments, and is not repeated here. Any technical details not fully described in the device provided in the embodiments of this application can be understood based on the description of the technical details in the method embodiments described above.

[0165] The present application also provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions or the computer program from the computer-readable storage medium and executes the computer-executable instructions or the computer program, causing the electronic device to perform the method provided in the present application.

[0166] An embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the method provided by the embodiment of the present application.

[0167] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0168] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0169] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, e.g., in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0170] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0171] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A training method for a live detection model, applied to an electronic device, the method comprising: Obtaining image samples for training a live detection model, and obtaining a first description text of the image samples and a second description text of the image samples in a target style; Based on the first description text and the second description text, performing style enhancement on the image features of the image samples to obtain first enhanced image features, and performing style enhancement on the image features based on style enhancement parameters to obtain second enhanced image features; Updating the style enhancement parameters based on the difference between the second enhanced image features and the first enhanced image features to obtain target style enhancement parameters; When training the live detection model, performing style enhancement on the image features based on the target style enhancement parameters to obtain target enhanced image features of the image samples, and training the live detection model based on the target enhanced image features to obtain a trained live detection model.

2. The method according to claim 1, wherein, The performing style enhancement on the image features of the image samples based on the first description text and the second description text to obtain first enhanced image features includes: Performing text encoding on the first description text to obtain first text features, and performing text encoding on the second description text to obtain second text features; Performing image encoding on the image features of the image samples to obtain encoded image features; Fusing the second text features, the first text features, and the encoded image features to obtain the first enhanced image features.

3. The method according to claim 2, wherein, The fusing the second text features, the first text features, and the encoded image features to obtain the first enhanced image features includes: Determining the text feature difference between the second text features and the first text features; Fusing the text feature difference and the encoded image features to obtain the first enhanced image features.

4. The method according to any one of claims 1-3, wherein, The performing style enhancement on the image features based on style enhancement parameters to obtain second enhanced image features includes: Fusing the style enhancement parameters and the image features to obtain a first fusion result; Performing image encoding on the first fusion result to obtain the second enhanced image features.

5. The method according to any one of claims 1-4, wherein, The style enhancement parameters belong to a machine learning model; the updating the style enhancement parameters based on the difference between the second enhanced image features and the first enhanced image features to obtain target style enhancement parameters includes: Obtaining the loss function of the machine learning model; Determining the value of the loss function based on the difference between the second enhanced image features and the first enhanced image features; Updating the style enhancement parameters of the machine learning model based on the value of the loss function to obtain the target style enhancement parameters.

6. The method according to claim 5, wherein, The loss function includes a first loss function and a second loss function; the determining the value of the loss function based on the difference between the second enhanced image features and the first enhanced image features includes: Determining the value of the first loss function based on the difference between the second enhanced image features and the first enhanced image features; Determine the value of the second loss function based on the difference between the second enhanced image feature and the image feature; Add the value of the first loss function and the value of the second loss function to obtain the value of the loss function.

7. The method according to any one of claims 1-6, wherein, Before enhancing the style of the image feature based on the target style enhancement parameter to obtain the target enhanced image feature of the image sample, the method further includes: Obtain a sample selection method; Select an image sample to be enhanced for style enhancement from the image sample set where the image sample is located according to the sample selection method; The enhancing the style of the image feature based on the target style enhancement parameter to obtain the target enhanced image feature of the image sample includes: When the image sample with the image feature is the image sample to be enhanced, enhance the style of the image feature based on the target style enhancement parameter to obtain the target enhanced image feature of the image sample.

8. The method according to claim 7, wherein, The method further includes: When the image sample with the image feature is not the image sample to be enhanced, train the live detection model based on the image feature.

9. The method according to any one of claims 1-8, wherein, The enhancing the style of the image feature based on the target style enhancement parameter to obtain the target enhanced image feature of the image sample includes: Fuse the target style enhancement parameter and the image feature to obtain a second fusion result; Perform image encoding on the second fusion result to obtain the target enhanced image feature of the image sample.

10. The method according to any one of claims 1-9, wherein, The training the live detection model based on the target enhanced image feature to obtain a trained live detection model includes: Call the live detection model to perform live detection based on the target enhanced image feature to obtain the live detection result of the image sample; Train the live detection model based on the difference between the live detection result and the sample label of the image sample to obtain a trained live detection model.

11. The method according to claim 10, wherein, Before calling the live detection model to perform live detection based on the target enhanced image feature to obtain the live detection result of the image sample, the method further includes: Obtain multiple category texts, where the multiple category texts include: a first text representing the live category and a second text representing each of the at least one non-live category in the non-live categories; The calling the live detection model to perform live detection based on the target enhanced image feature to obtain the live detection result of the image sample includes: Call the live detection model to perform text encoding on each of the category texts to obtain the category text features of each of the category texts; Call the live detection model to perform live detection based on the target enhanced image feature and each of the category text features to obtain the prediction probabilities of the biological object sample in the image sample belonging to each category; Wherein, the category includes the live category and the at least one non-live category, and the live detection result of the image sample includes the prediction probabilities of the biological object sample belonging to each category.

12. The method according to claim 11, wherein, Performing liveness detection based on the target enhanced image features and each of the category text features to obtain the prediction probabilities of the biological object samples in the image sample belonging to each category, including: Performing feature dimensionality reduction processing on the target enhanced image features to obtain intermediate enhanced image features, and performing feature dimensionality increase processing on the intermediate enhanced image features to obtain target image features; Adding the target enhanced image features and the target image features to obtain added features; Performing liveness detection based on the added features and each of the category text features to obtain the prediction probabilities of the biological object samples in the image sample belonging to each category.

13. The method according to claim 12, wherein The performing feature dimensionality reduction processing on the target enhanced image features to obtain intermediate enhanced image features, and performing feature dimensionality increase processing on the intermediate enhanced image features to obtain target image features, including: Invoking the feature dimension processing layer of the liveness detection model to perform feature dimensionality reduction processing on the target enhanced image features to obtain intermediate enhanced image features, and performing feature dimensionality increase processing on the intermediate enhanced image features to obtain target image features; Training the liveness detection model based on the difference between the liveness detection result and the sample label of the image sample to obtain a trained liveness detection model, including: Updating the model parameters of the feature dimension processing layer based on the difference between the liveness detection result and the sample label of the image sample to obtain a trained liveness detection model.

14. The method according to any one of claims 1-7, 9, wherein, Before training the liveness detection model based on the target enhanced image features to obtain a trained liveness detection model, the method further includes: Training the liveness detection model based on the image features to obtain an intermediate liveness detection model; Training the liveness detection model based on the target enhanced image features to obtain a trained liveness detection model, including: Training the intermediate liveness detection model based on the target enhanced image features to obtain a trained liveness detection model.

15. A liveness detection method applied to an electronic device, the method including: Obtaining a to-be-detected image including a biological object and a plurality of category texts, the plurality of category texts including: a first text representing the liveness category, and a second text representing each of the non-liveness categories in at least one non-liveness category; Invoking a liveness detection model to perform text encoding on each of the category texts to obtain category text features of each of the category texts, and performing feature extraction on the to-be-detected image to obtain to-be-detected image features; Invoking the liveness detection model to perform liveness detection based on the to-be-detected image features and each of the category text features to obtain a liveness detection result of the to-be-detected image; Wherein, the liveness detection result includes the prediction probabilities of the biological object belonging to each category, and the categories include the liveness category and the at least one non-liveness category; Wherein, the liveness detection model is trained based on the liveness detection model training method according to any one of claims 1-14.

16. A training device for a liveness detection model, the device including: A first acquisition module, configured to acquire image samples for training a live detection model, and acquire a first description text of the image samples and a second description text of the image samples in a target style; A style enhancement module, configured to enhance the image features of the image samples based on the first description text and the second description text to obtain first enhanced image features, and enhance the image features based on style enhancement parameters to obtain second enhanced image features; An update module, configured to update the style enhancement parameters based on the difference between the second enhanced image features and the first enhanced image features to obtain target style enhancement parameters; A training module, configured to, when training the live detection model, enhance the image features based on the target style enhancement parameters to obtain target enhanced image features of the image samples, and train the live detection model based on the target enhanced image features to obtain a trained live detection model.

17. An electronic device, the electronic device comprising: A memory, configured to store computer-executable instructions; A processor, configured to implement the method according to any one of claims 1 to 15 when executing the computer-executable instructions stored in the memory.

18. A computer-readable storage medium, storing computer-executable instructions or a computer program, where the computer-executable instructions or the computer program, when executed by a processor, implement the method according to any one of claims 1 to 15.

19. A computer program product, comprising computer-executable instructions or a computer program, where the computer-executable instructions or the computer program, when executed by a processor, implement the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Image recognition method and device, electronic equipment and storage medium

    CN112651451A

  • Living body detection method, training method of living body detection model and corresponding device

    CN115482591A

  • Living body detection method, device and equipment

    CN115546908A

  • Human face in-vivo detection model training method, human face in-vivo detection method and human face in-vivo detection device

    CN115761839A

  • Living body detection method and device, electronic equipment and storage medium

    CN116704620A

Cited By

  • Structured target detection method, device and equipment based on multi-modal language model

    CN120580514A