Liveness detection method and device, storage medium and equipment

By adaptively selecting the authentication action sequence and combining embarrassment perception and liveness detection models, the security and user experience issues of liveness attack detection in face recognition systems are solved, achieving efficient liveness detection and a user experience with low embarrassment.

CN116665314BActive Publication Date: 2025-12-12ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310412720.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-13
Publication Date
2025-12-12
Estimated Expiration
2043-04-13

AI Technical Summary

Technical Problem

Existing facial recognition systems have insufficient security performance in detecting liveness attacks, especially in scenarios with high security requirements. Furthermore, users may feel embarrassed when interacting with these systems, which negatively impacts the user experience.

Method used

By determining the target embarrassment score based on the target user's historical data and current environmental images, the system adaptively selects authentication action sequences and combines an embarrassment perception model and a liveness detection model for liveness detection, thereby reducing user embarrassment and improving detection effectiveness.

Benefits of technology

It effectively reduces the embarrassment of users during the liveness detection process and improves the security performance and user experience of liveness detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116665314B_ABST
    Figure CN116665314B_ABST
Patent Text Reader

Abstract

The specification discloses a living body detection method, device, storage medium and equipment, wherein the method comprises the following steps: determining a target embarrassment score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, the current environment image being an environment image in which the target user is located when performing a face recognition transaction, the user historical data being authentication mode selection data generated when the target user performs identity authentication in the past, determining an authentication action sequence corresponding to the target user in a preset authentication action set according to the target embarrassment score, acquiring an authentication action image sequence corresponding to the target user when the target user performs the authentication action sequence, and performing living body detection on the authentication action image sequence to obtain a living body detection result corresponding to the target user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of biometric identification, and in particular to a living body detection method and device, a storage medium and an equipment. BACKGROUND

[0002] With the continuous development of face recognition systems in recent years, face recognition technology is becoming mature, and its commercial application is becoming more and more widespread, such as widely used in financial transactions, access control systems, mobile terminals and other fields. However, the face is easy to be copied by photos, videos, models or masks, etc., so the false face of the legal user is an important threat to the security of the face recognition and authentication system. In order to prevent malicious people from forging and stealing the biological characteristics of others for identity authentication, "living body attack detection" has become an indispensable part of the face recognition system, which can effectively intercept non-living body type attack samples in the face recognition system. SUMMARY

[0003] The living body detection method, device, storage medium and equipment provided by the embodiments of the present specification can reduce the embarrassment of the user and improve the user experience by adaptively adjusting the living body detection authentication action of the user through embarrassment analysis of the user. The technical solution is as follows:

[0004] In a first aspect, the embodiments of the present specification provide a living body detection method, which comprises:

[0005] determining a target embarrassment score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, the current environment image being an environment image in which the target user is located when performing a face recognition transaction, and the user historical data being authentication mode selection data generated when the target user historically performs identity authentication;

[0006] determining an authentication action sequence corresponding to the target user in a preset authentication action set according to the target embarrassment score;

[0007] obtaining an authentication action image sequence corresponding to the target user when performing the authentication action sequence;

[0008] performing living body detection on the authentication action image sequence to obtain a living body detection result corresponding to the target user.

[0009] In a second aspect, the embodiments of the present specification provide a method for training an embarrassment awareness model, which comprises:

[0010] constructing a first sample training data set, the first sample training data comprising a sample environment image in which a user performs a face recognition transaction, sample historical data generated when the user historically performs identity authentication, and an embarrassment score label corresponding to the first sample training data;

[0011] inputting the first sample training data into an initial awkwardness perception model to obtain a predicted awkwardness score corresponding to the first sample training data;

[0012] performing supervised training on the awkwardness perception model based on the awkwardness perception loss function, the predicted awkwardness score, and the awkwardness score label, and iteratively updating model parameters of the awkwardness perception model until the awkwardness perception model converges, to obtain a trained awkwardness perception model.

[0013] In a third aspect, an embodiment of the present specification provides a living body detection model training method, and the method comprises:

[0014] constructing a second sample training data set, wherein the second sample training data comprises a sample authentication image sequence collected when a user performs a face recognition transaction, and a living body detection label corresponding to the second sample training data;

[0015] inputting the second sample training data into an initial living body detection model to obtain a sample living body detection result corresponding to the second sample training data;

[0016] performing supervised training on the living body detection model based on a living body detection loss function, the sample living body detection result, and the living body detection label, and iteratively updating model parameters of the living body detection model until the living body detection model converges, to obtain a trained living body detection model.

[0017] In a fourth aspect, an embodiment of the present specification provides a living body detection apparatus, comprising:

[0018] an awkwardness score determination module configured to determine a target awkwardness score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, wherein the current environment image is an environment image in which the target user is located when performing a face recognition transaction, and the user historical data is authentication mode selection data generated when the target user historically performs identity authentication;

[0019] an action sequence determination module configured to determine an authentication action sequence corresponding to the target user in a preset authentication action set according to the target awkwardness score;

[0020] an image sequence acquisition module configured to acquire an authentication action image sequence corresponding to the target user when performing the authentication action sequence;

[0021] a living body detection module configured to perform living body detection on the authentication action image sequence to obtain a living body detection result corresponding to the target user.

[0022] In a fifth aspect, an embodiment of the present specification provides an awkwardness perception model training apparatus, comprising:

[0023] a first sample set construction module, configured to construct a first sample training data set, the first sample training data including a sample environment image where a user performs a face recognition transaction, sample historical data generated when the user performs identity authentication in the past, and an embarrassment score label corresponding to the first sample training data;

[0024] an embarrassment score prediction module, configured to input the first sample training data into an initial embarrassment awareness model to obtain a predicted embarrassment score corresponding to the first sample training data;

[0025] an embarrassment model training module, configured to supervise training of the embarrassment awareness model based on an embarrassment awareness loss function, the predicted embarrassment score, and the embarrassment score label, and iteratively update model parameters of the embarrassment awareness model until the embarrassment awareness model converges, to obtain a trained embarrassment awareness model.

[0026] In a sixth aspect, an embodiment of the present specification provides a living body detection model training apparatus, including:

[0027] a second sample set construction module, configured to construct a second sample training data set, the second sample training data including a sample authentication image sequence collected when a user performs a face recognition transaction, and a living body detection label corresponding to the second sample training data;

[0028] a sample living body detection module, configured to input the second sample training data into an initial living body detection model to obtain a sample living body detection result corresponding to the second sample training data;

[0029] a living body model training module, configured to supervise training of the living body detection model based on a living body detection loss function, the sample living body detection result, and the living body detection label, and iteratively update model parameters of the living body detection model until the living body detection model converges, to obtain a trained living body detection model.

[0030] In a seventh aspect, an embodiment of the present specification provides a computer program product, the computer program product storing at least one instruction, the at least one instruction being adapted to be loaded by a processor and execute the method steps described above.

[0031] In an eighth aspect, an embodiment of the present specification provides a storage medium, the storage medium storing a computer program, the computer program being adapted to be loaded by a processor and execute the method steps described above.

[0032] In a ninth aspect, an embodiment of the present specification provides an electronic device, which can include a processor and a memory; wherein the memory stores a computer program, the computer program being adapted to be loaded by the processor and execute the method steps described above.

[0033] The technical solutions provided by some embodiments of the present specification have at least the following beneficial effects:

[0034] The method for detecting living bodies provided by the embodiments of the present specification determines a target embarrassment score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, the current environment image being an environment image in which the target user is located when performing a face recognition transaction, and the user historical data being authentication mode selection data generated when the target user performed identity authentication in the past, determines an authentication action sequence corresponding to the target user in a preset authentication action set according to the target embarrassment score, acquires an authentication action image sequence corresponding to the target user when performing the authentication action sequence, performs living body detection on the authentication action image sequence, and obtains a living body detection result corresponding to the target user. By analyzing the embarrassment of the user, the authentication action sequence is adaptively determined according to the embarrassment analysis result of the user, the embarrassment of the user in the living body detection authentication process is reduced, the living body detection effect is improved, and the user experience is improved. BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present specification, and other drawings can be obtained by those skilled in the art without creative labor.

[0036] Figure 1 A flowchart of a living body detection method provided by an embodiment of the present specification;

[0037] Figure 2 A scene diagram of a living body detection provided by an embodiment of the present specification;

[0038] Figure 3 A flowchart of a living body detection method provided by an embodiment of the present specification;

[0039] Figure 4 A flowchart of an embarrassment perception model training method provided by an embodiment of the present specification;

[0040] Figure 5 A flowchart of an embarrassment perception model training method provided by an embodiment of the present specification;

[0041] Figure 6 A structure diagram of an embarrassment perception model provided by an embodiment of the present specification;

[0042] Figure 7 A flowchart of an embarrassment perception model training method provided by an embodiment of the present specification;

[0043] Figure 8 A structural diagram of an embarrassment perception model provided for an embodiment of the present specification is shown in FIG. 1.

[0044] Figure 9 A flowchart of a living body detection model training method provided for an embodiment of the present specification is shown in FIG. 2.

[0045] Figure 10 A flowchart of a living body detection model training method provided for an embodiment of the present specification is shown in FIG. 2.

[0046] Figure 11 A structural diagram of a living body detection model provided for an embodiment of the present specification is shown in FIG. 3.

[0047] Figure 12 A flowchart of a living body detection model training method provided for an embodiment of the present specification is shown in FIG. 2.

[0048] Figure 13 A structural diagram of a living body detection model provided for an embodiment of the present specification is shown in FIG. 3.

[0049] Figure 14 A structural diagram of a living body detection device provided for an embodiment of the present specification is shown in FIG. 4.

[0050] Figure 15 A structural diagram of a living body detection device provided for an embodiment of the present specification is shown in FIG. 4.

[0051] Figure 16 A structural diagram of an embarrassment perception model training device provided for an embodiment of the present specification is shown in FIG. 5.

[0052] Figure 17 A structural diagram of a living body detection model training device provided for an embodiment of the present specification is shown in FIG. 6.

[0053] Figure 18 A structural block diagram of an electronic device provided for an embodiment of the present specification is shown in FIG. 7. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present specification will be described clearly and completely below with reference to the drawings in the embodiments of the present specification. Obviously, the described embodiments are only part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present specification.

[0055] In the description of the specification, it is understood that the terms "first", "second" and the like are only for the purpose of description and cannot be understood as indicating or implying relative importance. In the description of the specification, it is necessary to explain that, unless otherwise explicitly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units not listed, or optionally includes other steps or units inherent to the process, method, product or device. The specific meaning of the above terms in the specification can be understood by the person skilled in the art. In addition, in the description of the specification, "a plurality of" means two or more, unless otherwise specified. The association relationship of the associated objects is described, which means that there can be three relationships, for example, A and / or B can represent three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.

[0056] In related technologies, with the continuous development of face recognition systems, "liveness attack detection" has become an indispensable part of face recognition systems, which can effectively intercept non-liveness type attack samples. Common liveness detection methods can be divided into two types:

[0057] One type is a natural state silent liveness detection method, which does not require user cooperation to complete the interactive action, but only collects face images in natural state for liveness detection. Due to the limited amount of information in the natural state face image, the security performance of this method is poor, and it can only be used in low security requirement scenarios such as community access control, attendance, etc., and cannot be applied to high security requirement financial scenarios.

[0058] The second type is a liveness detection method based on user active interaction. This method requires the user to cooperate to complete the predetermined interactive action, such as "wink", "shake head" and "open mouth" and the like. Due to the introduction of additional action interaction, the information amount is increased, and the security performance of liveness detection is also improved. However, this method ignores the user experience when completing the action. For example, in public places, users are not convenient to complete the "open mouth" action, which will cause users to feel embarrassed and give up face verification, and instead use passwords, and even if the user continues to perform action authentication, the embarrassment will make the action stiff, which is easy to cause liveness detection misjudgment.

[0059] Based on this, the embodiment of the present specification proposes a living body detection method, first, based on the user historical data corresponding to the target user and the current environment image, the target user corresponding to the target embarrassment score is determined, the current environment image is the environment image when the target user performs face recognition transaction, the user historical data is the authentication mode selection data generated when the target user historically performs identity authentication, the target user corresponding to the authentication action sequence is determined in the preset authentication action set according to the target embarrassment score, the authentication action image sequence corresponding to the target user when the target user executes the authentication action sequence is obtained, the living body detection result corresponding to the target user is obtained by living body detection on the authentication action image sequence, the embarrassment of the user is analyzed, the authentication action sequence is adaptively determined according to the embarrassment analysis result of the user, the embarrassment of the user in the living body detection authentication process is reduced, the living body detection effect is improved, and the user experience is improved.

[0060] The following detailed description of the embodiments in the embodiment of the present specification. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present specification. On the contrary, they are only examples of devices and methods consistent with some aspects of the present specification as detailed in the appended claims. The flowchart shown in the drawing is only an exemplary description, and it is not necessary to perform the steps shown. For example, some steps are parallel, and there is no strict logical sequence, so the actual execution order is variable.

[0061] Please refer to Figure 1 , a flowchart of a living body detection method provided by the embodiment of the present specification. In the embodiment of the present specification, the living body detection method is applied to a living body detection model training device or an electronic device configured with a living body detection model training device. The following will be described in detail with respect to the flow shown in Figure 1 The living body detection method can specifically include the following steps:

[0062] S102, based on the user historical data corresponding to the target user and the current environment image, the target user corresponding to the target embarrassment score is determined, the current environment image is the environment image when the target user performs face recognition transaction, the user historical data is the authentication mode selection data generated when the target user historically performs identity authentication;

[0063] In the embodiment of the present specification, before the user performs face recognition authentication based on the authentication action, the authentication mode selection data corresponding to the user historically performing identity authentication and the environment image currently located by the user are obtained, the target user corresponding to the target embarrassment score is determined according to the user historical data corresponding to the target user and the current environment image located by the target user.

[0064] It should be noted that the authentication mode selection data can be understood as the authentication mode selected by the target user in the history when performing identity authentication. For example, the user can select face recognition authentication and password authentication. It can be understood that the number and frequency of times that the target user selects face recognition authentication can represent the degree of embarrassment of the target in face recognition authentication to some extent. For example, user A frequently and highly frequently uses face recognition authentication, which indicates that user A is not prone to have an embarrassing emotion, and user B only selects face recognition authentication at a low frequency when performing identity authentication, and manually switches the identity authentication mode from face recognition authentication to password authentication a large number of times, which indicates that user B is prone to have an embarrassing emotion.

[0065] The current environment image is an image of the current surrounding environment of the target user. It can be understood that the user is more prone to have an embarrassing emotion in a public place or a place with many people around, and is less prone to have an embarrassing emotion in a private place or a place with few people around.

[0066] In the embodiments of the present application, the target embarrassment score corresponding to the target user is determined based on the user historical data corresponding to the target user and the current environment image. The user historical data and the current environment image can be respectively subjected to feature extraction processing to obtain data features corresponding to the user historical data and image features corresponding to the current environment image. The data features and the image features are subjected to feature fusion processing to obtain first fusion features, and finally the target embarrassment score corresponding to the target user is predicted and generated based on the first fusion features.

[0067] In the embodiments of the present application, the target embarrassment score corresponding to the target user can be predicted based on a pre-trained embarrassment perception model. The user historical data corresponding to the target user and the current environment image are input into the embarrassment perception model to obtain the target embarrassment score corresponding to the target user. The embarrassment perception model is a neural network model trained based on deep learning, and can predict the embarrassment score corresponding to the user according to the user historical data and the current environment image.

[0068] In S104, an authentication action sequence corresponding to the target user is determined from the preset authentication action set according to the target embarrassment score.

[0069] In the embodiments of the present application, the preset authentication action set includes at least one authentication action and an action embarrassment score corresponding to each authentication action. According to the target embarrassment score, an authentication action sub-set of authentication actions with an action embarrassment score less than the target embarrassment score is determined from the preset action set. A preset number of authentication actions are randomly selected from the authentication action sub-set to generate an authentication action sequence.

[0070] It should be noted that in the traditional face swiping action mode, the user is generally required to perform two to three specified actions, such as blinking, opening the mouth, and turning the head, and the authentication action does not change. In the embodiments of the present specification, considering that the degree of embarrassment of the user when performing different authentication actions is different, the present scheme maps the authentication action and the action embarrassment score one by one through the action embarrassment score corresponding to different authentication actions, and selects the authentication action with an action embarrassment score less than the target embarrassment score from the preset authentication action set to form an authentication action sequence according to the predicted target embarrassment score of the user, so that the target user performs face authentication according to the authentication action sequence, reduces the user's embarrassment, improves the completion effect of the user's authentication action, and further improves the live detection effect.

[0071] S106, obtaining an authentication action image sequence corresponding to the target user performing the authentication action sequence;

[0072] In the embodiments of the present specification, after obtaining the authentication action sequence corresponding to the target user, the authentication action sequence is sequentially prompted on the display interface to instruct the target user to perform face authentication according to the authentication action indicated by the authentication action sequence, and the authentication action image sequence of the user performing the authentication action sequence is collected.

[0073] It can be understood that in the face authentication process, most of the face authentication process is based on the camera to collect face images and face actions. During the process of the target user performing the authentication action sequence, the authentication action image sequence of the user performing the authentication action sequence is collected. The authentication action image sequence is used for live detection.

[0074] Please refer to Figure 2 , a scene schematic diagram of live detection provided by the embodiments of the present specification. As Figure 2 shown, the electronic device is a device for performing face recognition, and an image collection device is configured thereon. When the user approaches the device and is within the image collection range of the image collection device, the image collection device will collect the face image of the approaching person. Based on the image collection device, the authentication action image sequence of the user performing the authentication action sequence can be collected, and the authentication action image sequence is subjected to live detection.

[0075] S108, performing live detection on the authentication action image sequence to obtain a live detection result corresponding to the target user.

[0076] In the embodiments of the present specification, after the authentication action image sequence corresponding to the target user performing the authentication action sequence is collected, the authentication action image sequence is subjected to live detection to obtain a live detection result corresponding to the target user.

[0077] Optionally, feature extraction processing is performed on the authentication action image sequence to obtain action completion degree features and attack clue features corresponding to the authentication action image sequence, the action completion degree features and the attack clue features are subjected to feature fusion processing to obtain second fusion features, and the second fusion features are used to predict and generate a liveness detection result corresponding to the target user.

[0078] In the embodiments of the present specification, the pre-trained liveness detection model can be used to perform liveness detection on the authentication action image sequence to obtain a liveness detection result corresponding to the target user. The authentication action image sequence is input into the liveness detection model, and the liveness detection model performs liveness detection on the authentication action image sequence to obtain a liveness detection result corresponding to the target user.

[0079] It should be noted that face spoofing and fake face are major threats to face recognition systems. Face spoofing refers to the behavior of fraudsters attacking face recognition systems by making face models, face masks, photos, videos, and the like. The liveness detection model is a model used to perform liveness detection on the authentication action image sequence collected by the face recognition system, and to distinguish whether the currently collected authentication action image sequence is of a live type or a malicious attack type. In the application process of the face recognition system, the authentication action image sequence is mainly collected by live shooting, and then the liveness detection model is used to detect whether the collected authentication action image sequence is live. If the authentication action image sequence is a face image collected by the face recognition system from a live person's face, it is determined that the authentication action image sequence is of a live type. If the authentication action image sequence is an authentication action image sequence collected by the face recognition system from a fake face made by fraudsters, it is determined that the authentication action image sequence is of an attack type.

[0080] In one embodiment, the liveness detection result can be an attack probability value of the authentication action image sequence being of an attack type. After obtaining the liveness detection result, it is determined whether the attack probability value is greater than a preset threshold. If the attack probability value is greater than the preset threshold, it is determined that the authentication action image sequence is of an attack type. If the attack probability value is not greater than the preset threshold, it is determined that the authentication action image sequence is of a live type.

[0081] In the embodiment of the present application, first, the target embarrassment score corresponding to the target user is determined based on the user historical data corresponding to the target user and the current environment image, the current environment image is the environment image in which the target user is located when performing face recognition transaction, the user historical data is the authentication mode selection data generated when the target user performs identity authentication in the past, the authentication action sequence corresponding to the target user is determined in the preset authentication action set according to the target embarrassment score, the authentication action image sequence corresponding to the target user when performing the authentication action sequence is obtained, the living body detection result corresponding to the target user is obtained by performing living body detection on the authentication action image sequence, the authentication action sequence is adaptively determined according to the embarrassment analysis result of the user through embarrassment analysis of the user, the embarrassment of the user in the living body detection authentication process is reduced, the living body detection effect is improved, and the user experience is improved.

[0082] Please refer to Figure 3 A flowchart of a living body detection method provided by the embodiment of the present application is shown. The living body detection method can specifically include the following steps:

[0083] S202, the user historical data and the current environment image are respectively subjected to feature extraction processing to obtain data features corresponding to the user historical data and image features corresponding to the current environment image;

[0084] In the embodiment of the present application, the current environment image is the environment image in which the target user is located when performing face recognition transaction, and the user historical data is the authentication mode selection data generated when the target user performs identity authentication in the past. Before the user performs face recognition authentication based on the authentication action, the authentication mode selection data corresponding to the user in the past when performing identity authentication and the environment image in which the user is currently located are obtained. The obtained user historical data and current environment image are subjected to feature extraction processing to obtain data features corresponding to the user historical data of the target user and image features corresponding to the current environment image.

[0085] The data features and the image features are data features that can represent the embarrassment degree of the target user to a certain extent.

[0086] In one embodiment, the embarrassment perception model includes a data feature extraction network and an image feature extraction network. The user historical data corresponding to the target user and the current environment image are input into the embarrassment perception model, the user historical data is subjected to feature extraction processing by the data feature extraction network to obtain data features corresponding to the user historical data, and the current environment image is subjected to feature extraction processing by the image feature extraction network to obtain image features corresponding to the current environment image.

[0087] S204, the data features and the image features are subjected to feature fusion processing to obtain first fusion features;

[0088] In the embodiments of the present specification, after feature extraction processing is respectively performed on the user historical data and the current environment image to obtain data features and image features, the data features and the image features are subjected to feature fusion processing to obtain first fused features.

[0089] Optionally, the data features and the image features are subjected to feature fusion in the channel dimension to obtain the first fused features.

[0090] In one embodiment, the embarrassment perception model comprises a first feature fusion network. After the user historical data corresponding to the target user and the current environment image are input into the embarrassment perception model, the data feature extraction network is used to perform feature extraction processing on the user historical data to obtain data features corresponding to the user historical data, the image feature extraction network is used to perform feature extraction processing on the current environment image to obtain image features corresponding to the current environment image, and then the first feature fusion network is used to perform feature fusion processing on the data features and the image features to obtain first fused features.

[0091] S206, a target embarrassment score corresponding to the target user is predicted and generated based on the first fused features;

[0092] In the embodiments of the present specification, the first fused features generated after the data features and the image features are fused are used for embarrassment score prediction to obtain a target embarrassment score corresponding to the target user.

[0093] It can be understood that the target embarrassment score is a predicted embarrassment degree of the target user in the current environment for face recognition authentication, which is calculated according to the user historical data and the current environment image. The higher the target embarrassment score is, the more likely the target user is to be embarrassed.

[0094] The target embarrassment score is used to determine an authentication action with a lower embarrassment degree for the target user from a preset set of authentication actions.

[0095] In one embodiment, the embarrassment perception model comprises a first feature fusion network. After the user historical data corresponding to the target user and the current environment image are input into the embarrassment perception model, the data feature extraction network is used to perform feature extraction processing on the user historical data to obtain data features corresponding to the user historical data, the image feature extraction network is used to perform feature extraction processing on the current environment image to obtain image features corresponding to the current environment image, and then the first feature fusion network is used to perform feature fusion processing on the data features and the image features to obtain first fused features. The first fused features are subjected to embarrassment score prediction by the first fused feature prediction network to obtain a target embarrassment score corresponding to the target user.

[0096] S208, determine a sub-set of authentication actions from the set of preset authentication actions based on the target embarrassment score, each authentication action in the sub-set of authentication actions corresponding to an action embarrassment score less than the target embarrassment score;

[0097] In the embodiments of the present disclosure, the set of preset authentication actions includes at least one authentication action and an action embarrassment score corresponding to each authentication action. After the target embarrassment score corresponding to the target user is predicted, a sub-set of authentication actions with action embarrassment scores less than the target embarrassment score is determined from the set of preset authentication actions based on the target embarrassment score.

[0098] It can be understood that the action embarrassment score corresponding to each authentication action in the set of preset authentication actions indicates the average embarrassment level of the user when performing the action. When the action embarrassment score of an authentication action is less than the target embarrassment score of the target user, it indicates that the target user can receive the authentication action and will not be too embarrassed to not want to perform or affect the execution effect of the action when performing the authentication action. When the action embarrassment score of an authentication action is greater than the target embarrassment score of the target user, it indicates that the target user cannot receive the authentication action and will be too embarrassed to not want to perform when performing the authentication action, thereby affecting the execution effect of the authentication action.

[0099] S210, randomly select a preset number of authentication actions from the sub-set of authentication actions to generate an authentication action sequence;

[0100] In the embodiments of the present disclosure, the action embarrassment score corresponding to each authentication action in the sub-set of authentication actions is less than the target embarrassment score. At this time, a preset number of authentication actions are randomly selected from the sub-set of authentication actions to generate an authentication action sequence.

[0101] For example, three authentication actions are randomly selected from a sub-set of authentication actions including ten authentication actions to generate an authentication action sequence.

[0102] S212, obtain an authentication action image sequence corresponding to the target user performing the authentication action sequence;

[0103] In the embodiments of the present disclosure, after the authentication action sequence corresponding to the target user is obtained, the authentication action sequence is sequentially prompted on the display interface to instruct the target user to perform face recognition authentication according to the authentication action indicated by the authentication action sequence, and the authentication action image sequence of the user performing the authentication action sequence is collected.

[0104] S214, perform feature extraction processing on the authentication action image sequence to obtain action completion degree features and attack clue features corresponding to the authentication action image sequence;

[0105] In the embodiments of the present specification, after the sequence of authentication action images corresponding to the sequence of authentication actions performed by the target user is collected, feature extraction processing is performed on the sequence of authentication action images to obtain action completion degree features and attack clue features corresponding to the sequence of authentication action images.

[0106] The action completion degree features are used to represent the degree of completion of the sequence of authentication actions by the target user. The attack clue features are used to detect whether there are attack features such as face spoofing, fake face, and synthetic video face in the sequence of authentication action images.

[0107] In one embodiment, the pre-trained liveness detection model can be used to perform liveness detection on the sequence of authentication action images. The liveness detection model includes a first feature extraction network and a second feature extraction network. After the sequence of authentication action images is input into the liveness detection model, the first feature extraction network is used to perform feature extraction processing on the sequence of authentication action images to obtain action completion degree features corresponding to the sequence of authentication action images. The second feature extraction network is used to perform feature extraction processing on the sequence of authentication action images to obtain attack clue features corresponding to the sequence of authentication action images.

[0108] S216, the action completion degree features and the attack clue features are subjected to feature fusion processing to obtain second fusion features;

[0109] In the embodiments of the present specification, after the sequence of authentication action images is subjected to feature extraction processing to obtain action completion degree features and attack clue features corresponding to the sequence of authentication action images, the obtained action completion degree features and attack clue features are subjected to feature fusion to obtain second fusion features.

[0110] In one embodiment, the liveness detection model includes a second feature fusion network. After the sequence of authentication action images is input into the liveness detection model, the first feature extraction network is used to perform feature extraction processing on the sequence of authentication action images to obtain action completion degree features corresponding to the sequence of authentication action images. The second feature extraction network is used to perform feature extraction processing on the sequence of authentication action images to obtain attack clue features corresponding to the sequence of authentication action images. Then, the second feature fusion network is used to perform feature fusion on the action completion degree features and the attack clue features to obtain second fusion features.

[0111] S218, the second fusion features are used to predict and generate a liveness detection result corresponding to the target user.

[0112] In an embodiment, the living body detection model comprises a second fusion feature prediction network. After the authentication action image sequence is input into the living body detection model, the first feature extraction network is used to perform feature extraction processing on the authentication action image sequence to obtain action completion degree features corresponding to the authentication action image sequence. The second feature extraction network is used to perform feature extraction processing on the authentication action image sequence to obtain attack clue features corresponding to the authentication action image sequence. The second fusion feature prediction network is used to perform living body detection on the second fusion features obtained by performing feature fusion on the action completion degree features and the attack clue features, to obtain a living body detection result corresponding to the target user.

[0113] In an embodiment, the living body detection result can be an attack probability value of the authentication action image sequence being of an attack type. After the living body detection result is obtained, it is determined whether the attack probability value is greater than a preset threshold. If the attack probability value is greater than the preset threshold, it is determined that the authentication action image sequence is of the attack type. If the attack probability value is not greater than the preset threshold, it is determined that the authentication action image sequence is of a living body type.

[0114] In the embodiments of the present disclosure, first, feature extraction processing is performed on user historical data and a current environment image to obtain data features corresponding to the user historical data and image features corresponding to the current environment image. Feature fusion processing is performed on the data features and the image features to obtain first fusion features. Then, a target embarrassment score corresponding to a target user is predicted based on the first fusion features. The target embarrassment score corresponding to the target user is accurately predicted by performing feature analysis on the user historical data and the current environment image. Then, an authentication action sub-set is determined in a preset authentication action set based on the target embarrassment score. A preset number of authentication actions are randomly selected from the authentication action sub-set to generate an authentication action sequence. The authentication action sequence is adaptively determined according to an embarrassment analysis result of the user, so as to reduce the embarrassment of the user in the living body detection authentication process, improve the living body detection effect, and improve the user experience. Next, an authentication action image sequence corresponding to the target user performing the authentication action sequence is obtained. Feature extraction processing is performed on the authentication action image sequence to obtain action completion degree features and attack clue features corresponding to the authentication action image sequence. Feature fusion processing is performed on the action completion degree features and the attack clue features to obtain second fusion features. Finally, a living body detection result corresponding to the target user is predicted based on the second fusion features. The living body is predicted from two aspects of the action completion degree and the attack clue, so as to ensure the prediction effect of the living body detection.

[0115] See Figure 4A flowchart of an embarrassment perception model training method provided in an embodiment of the present specification is shown. In the embodiment of the present specification, the embarrassment perception model training method is applied to an embarrassment perception model training device or an electronic device configured with the embarrassment perception model training device. The embarrassment perception model training method will be described in detail below with reference to the flowchart shown in the embodiment, and the embarrassment perception model training method can specifically include the following steps: Figure 4

[0116] S302, a first sample training data set is constructed, the first sample training data including a sample environment image where a user performs a face recognition transaction, sample historical data generated when the user performs identity authentication in the past, and an embarrassment score label corresponding to the first sample training data;

[0117] It should be noted that the embarrassment perception model is a deep learning model for predicting an embarrassment score corresponding to a target user.

[0118] In the embodiment of the present specification, the first sample training data set contains a plurality of first sample training data for learning and training the embarrassment perception model, the first sample training data including a sample environment image where a user performs a face recognition transaction and sample historical data generated when the user performs identity authentication in the past, and the first sample training data further including an embarrassment score label corresponding to the first sample training data.

[0119] The sample historical data is an authentication mode selected by the user when performing identity authentication in the past, for example, the user can select face recognition authentication and password authentication. It can be understood that the number and frequency of times that the user selects face recognition authentication can represent the embarrassment degree of the target user performing face recognition authentication to some extent. For example, user A frequently and highly frequently uses face recognition authentication, which indicates that user A is not prone to feeling embarrassed, and user B only selects face recognition authentication at a low frequency when performing identity authentication and manually switches the identity authentication mode from face recognition authentication to password authentication a large number of times, which indicates that user B is prone to feeling embarrassed.

[0120] Optionally, the embarrassment score label can be an embarrassment score corresponding to the first sample training data annotated according to experience.

[0121] S304, inputting the first sample training data into an initial embarrassment perception model to obtain a predicted embarrassment score corresponding to the first sample training data;

[0122] In the embodiment of the present specification, the embarrassment perception model is a deep learning model for predicting an embarrassment score corresponding to a user. The first sample training data corresponding to a certain user is input into the initial embarrassment perception model, and the embarrassment perception model predicts a predicted embarrassment score corresponding to the first sample training data according to the first sample training data.

[0123] ​In an embodiment, the embarrassment perception model comprises a data feature extraction network, an image feature extraction network, a first feature fusion network, and a first fused feature prediction network. The first sample training data is input into the embarrassment perception model. The data feature extraction network is used to perform feature extraction processing on the sample historical data to obtain sample data features corresponding to the sample historical data. The image feature extraction network is used to perform feature extraction processing on the sample environment image to obtain sample image features corresponding to the sample environment image. The first feature fusion network is used to perform feature fusion processing on the sample data features and the sample image features to obtain first sample fused features. The first fused feature prediction network is used to perform embarrassment prediction on the first sample fused features to obtain a predicted embarrassment score corresponding to the first sample training data.

[0124] In S306, the embarrassment perception model is supervised trained and the model parameters of the embarrassment perception model are iteratively updated based on the embarrassment perception loss function, the predicted embarrassment score, and the embarrassment score label until the embarrassment perception model converges, and a trained embarrassment perception model is obtained.

[0125] In the embodiments of the present specification, the embarrassment perception loss value corresponding to the predicted embarrassment score and the embarrassment score label is calculated based on the embarrassment loss function, the model parameters of the embarrassment perception model are updated based on the embarrassment perception loss value, it is judged whether the embarrassment perception model after parameter updating meets the preset convergence condition, if yes, the training is stopped, and a trained embarrassment perception model is obtained, and if not, step S304 is executed.

[0126] In the embodiments of the present specification, a first sample training data set is constructed. The first sample training data comprises a sample environment image where a user performs a face recognition transaction, sample historical data generated when the user performs identity authentication in the past, and an embarrassment score label corresponding to the first sample training data. Then, the first sample training data is input into the initial embarrassment perception model to obtain a predicted embarrassment score corresponding to the first sample training data. Finally, the embarrassment perception model is supervised trained and the model parameters of the embarrassment perception model are iteratively updated based on the embarrassment perception loss function, the predicted embarrassment score, and the embarrassment score label until the embarrassment perception model converges, and a trained embarrassment perception model is obtained. By using the embodiments, the embarrassment perception model can be deeply learned and trained based on the first sample training data comprising the sample historical data and the sample environment image, and an embarrassment perception model capable of accurately predicting the embarrassment score of a user can be obtained.

[0127] Please refer to Figure 5 A flowchart of an embarrassment perception model training method provided by the embodiments of the present specification is shown. The embarrassment perception model training method can specifically comprise the following steps:

[0128] S402, construct a first sample training data set, the first sample training data including a sample environment image where the user performs a face recognition transaction, sample historical data generated when the user performs identity authentication in the past, and an embarrassment score label corresponding to the first sample training data;

[0129] In the embodiment of the present specification, step S402 refers to the detailed description of step S302 in another embodiment of the present specification, which will not be repeated here.

[0130] It should be noted that in the embodiment of the present specification, the embarrassment perception model includes a feature extraction network, an image feature extraction network, a first feature fusion network, and a first fused feature prediction network. Please refer to Figure 6 The structure diagram of an embarrassment perception model provided in the embodiment of the present specification.

[0131] S404, performing feature extraction processing on the sample historical data using the data feature extraction network to obtain sample data features corresponding to the sample historical data;

[0132] S406, performing feature extraction processing on the sample environment image using the image feature extraction network to obtain sample image features corresponding to the sample environment image;

[0133] S408, performing feature fusion processing on the sample data features and the sample image features based on the first feature fusion network to obtain first sample fusion features;

[0134] S410, performing embarrassment prediction on the first sample fusion features based on the first fused feature prediction network to obtain a predicted embarrassment score corresponding to the first sample training data;

[0135] Steps S404-S410 are specifically as shown in the following table: Figure 6 After the first sample training data is input into the embarrassment perception model, the sample historical data in the first sample training data is input into the data feature extraction network, the sample historical data is feature extracted by the data feature extraction network, and sample data features corresponding to the sample historical data are obtained. The sample environment image in the first sample training data is input into the image feature extraction network, the sample environment image is feature extracted by the image feature extraction network, and sample image features corresponding to the sample environment image are obtained. Then, the sample data features and the sample image features are feature fused based on the first feature fusion network to obtain first sample fusion features. Finally, the first sample fusion features are input into the first fused feature prediction network, the first sample fusion features are embarrassment predicted by the first fused feature prediction network, and a predicted embarrassment score corresponding to the first sample training data is obtained.

[0136] S412, calculating an embarrassment perception loss value corresponding to the predicted embarrassment score and the embarrassment score label based on the embarrassment perception loss function;

[0137] wherein the predicted embarrassment score is an embarrassment score predicted by the embarrassment perception model based on the first sample training data, and the embarrassment score label is an embarrassment score corresponding to the first sample training data set based on experience, which can be used as a true value. The embarrassment perception loss value corresponding to the predicted embarrassment score and the embarrassment score label is calculated by the embarrassment perception loss function, so as to guide the learning direction of the model according to the embarrassment perception loss value.

[0138] S414, updating the model parameters of the embarrassment perception model based on the embarrassment perception loss value;

[0139] S416, judging whether the embarrassment perception model after parameter updating meets a preset convergence condition, if yes, stopping training to obtain the trained embarrassment perception model, and if not, executing step S404.

[0140] Optionally, the preset convergence condition can be that the embarrassment perception loss value is less than a preset loss threshold. If the embarrassment perception loss value is less than the preset loss threshold, it is determined that the model converges, if the embarrassment perception loss value is greater than or equal to the preset loss threshold, it is determined that the model does not converge, and then step S404 is executed for iterative training.

[0141] In the embodiment of the present specification, a first sample training data set is constructed, the first sample training data includes a sample environment image where the user performs a face recognition transaction, sample historical data generated when the user performs identity authentication in the past, and an embarrassment score label corresponding to the first sample training data, then the sample historical data is processed by feature extraction using a data feature extraction network to obtain sample data features corresponding to the sample historical data, the sample environment image is processed by feature extraction using an image feature extraction network to obtain sample image features corresponding to the sample environment image, the sample data features and the sample image features are processed by feature fusion based on a first feature fusion network to obtain first sample fusion features, the first sample fusion features are predicted based on a first fusion feature prediction network to obtain a predicted embarrassment score corresponding to the first sample training data, and finally, an embarrassment perception loss value corresponding to the predicted embarrassment score and the embarrassment score label is calculated based on an embarrassment perception loss function, the model parameters of the embarrassment perception model are updated based on the embarrassment perception loss value, it is judged whether the embarrassment perception model after parameter updating meets a preset convergence condition, if yes, the training is stopped, and the embarrassment perception model after training is obtained, and if not, step S404 is performed; by using the embodiment, the embarrassment perception model is trained based on the first sample training data containing the sample historical data and the sample environment image, and the embarrassment perception model that can accurately predict the embarrassment score of the user can be obtained, and the embarrassment perception model uses the data feature extraction network and the image feature extraction network to perform feature extraction processing respectively, predicts the embarrassment score based on two kinds of features, and ensures the prediction accuracy of the model for the embarrassment score.

[0142] See Figure 7 The embarrassment perception model training method provided in the embodiment of the present specification can include the following steps:

[0143] S502, a first sample training data set is constructed, the first sample training data includes a sample environment image where the user performs a face recognition transaction, sample historical data generated when the user performs identity authentication in the past, and an embarrassment score label corresponding to the first sample training data;

[0144] In the embodiment of the present specification, step S502 is described in detail in another embodiment of the present specification, and will not be repeated here.

[0145] It should be noted that in the embodiment of the present specification, the embarrassment perception model includes a feature extraction network, an image feature extraction network, a first feature fusion network, a first fusion feature prediction network, a data feature prediction network, and an image feature prediction network. See Figure 8 The structure of the embarrassment perception model provided in the embodiment of the present specification is shown in the schematic diagram.

[0146] S504, the sample historical data is subjected to feature extraction processing by using a data feature extraction network to obtain sample data features corresponding to the sample historical data;

[0147] S506, the sample data features are subjected to embarrassment prediction based on a data feature prediction network to obtain data embarrassment scores corresponding to the sample historical data;

[0148] S508, the sample environment images are subjected to feature extraction processing by using an image feature extraction network to obtain sample image features corresponding to the sample environment images;

[0149] S510, the sample image features are subjected to embarrassment prediction based on an image feature prediction network to obtain image embarrassment scores corresponding to the sample environment images;

[0150] S512, the sample data features and the sample image features are subjected to feature fusion processing based on a first feature fusion network to obtain first sample fusion features;

[0151] S514, the first sample fusion features are subjected to embarrassment prediction based on a first fusion feature prediction network to obtain predicted embarrassment scores corresponding to the first sample training data;

[0152] Steps S504-S514 are specifically as shown in Figure 8 After the first sample training data is input into the embarrassment awareness model, the sample historical data in the first sample training data is input into the data feature extraction network, the sample historical data is subjected to feature extraction by the data feature extraction network to obtain sample data features corresponding to the sample historical data, and the sample data features are subjected to embarrassment prediction based on the data feature prediction network to obtain data embarrassment scores corresponding to the sample historical data; the sample environment images in the first sample training data are input into the image feature extraction network, the sample environment images are subjected to feature extraction by the image feature extraction network to obtain sample image features corresponding to the sample environment images, and the sample image features are subjected to embarrassment prediction based on the image feature prediction network to obtain image embarrassment scores corresponding to the sample environment images; then the sample data features and the sample image features are subjected to feature fusion processing based on the first feature fusion network to obtain first sample fusion features, and finally the first sample fusion features are input into the first fusion feature prediction network, the first sample fusion features are subjected to embarrassment prediction by the first fusion feature prediction network to obtain predicted embarrassment scores corresponding to the first sample training data.

[0153] S516, a data embarrassment loss value corresponding to a data embarrassment score and a data embarrassment label is calculated based on a data loss function;

[0154] It should be noted that the embarrassment perception loss function includes a data loss function, an image loss function, and a fusion perception loss function, and the embarrassment score label includes a data embarrassment label, an image embarrassment label, and a fusion embarrassment label.

[0155] The data embarrassment label is an embarrassment score set based on experience for sample historical data, the image embarrassment label is an embarrassment score set based on experience for sample environment images, and the fusion embarrassment label is an embarrassment score set based on experience for the first sample training data.

[0156] S518, calculating an image embarrassment loss value corresponding to the image embarrassment score and the image embarrassment label based on the image loss function;

[0157] S520, calculating a fusion embarrassment loss value corresponding to the predicted embarrassment score and the fusion embarrassment label based on the fusion perception loss function;

[0158] S522, updating the model parameters of the embarrassment perception model based on the data embarrassment loss value, the image embarrassment loss value, and the fusion embarrassment loss value;

[0159] S524, determining whether the embarrassment perception model after parameter updating meets a preset convergence condition, if yes, stopping training to obtain the embarrassment perception model after training, and if not, performing step S504.

[0160] In the embodiments of the present specification, the embarrassment perception model is trained based on the first sample training data containing sample historical data and sample environment images, and an embarrassment perception model that can accurately predict the embarrassment score of a user can be obtained. The embarrassment perception model uses a data feature extraction network and an image feature extraction network to perform feature extraction processing respectively, predicts the embarrassment score based on two kinds of features, and predicts the embarrassment score based on the fusion feature of the two kinds of features. Finally, three loss values are calculated based on the data loss function, the image loss function, and the fusion perception loss function, and the model parameters are updated and trained based on the three loss values, which further ensures the prediction accuracy of the model for the embarrassment score.

[0161] Please refer to Figure 9 A flowchart of a living body detection model training method is provided in the embodiments of the present specification. In the embodiments of the present specification, the living body detection model training method is applied to a living body detection model training device or an electronic device configured with a living body detection model training device. In the following, the flowchart shown in FIG. 8 will be described in detail. The living body detection model training method can specifically include the following steps: Figure 9

[0162] ​S602, a second sample training data set is constructed, the second sample training data including a sample authentication image sequence collected when the user performs a face recognition transaction, and a living body detection label corresponding to the second sample training data;

[0163] In the embodiments of the present application, the second sample training data set includes a plurality of second sample training data for learning and training the living body detection model, the second sample training data including a sample authentication image sequence collected when the user performs a face recognition transaction, and a living body detection label corresponding to the second sample training data.

[0164] It can be understood that the sample authentication image sequence is an image sequence collected when the user performs an authentication action sequence indicated in the face recognition transaction.

[0165] The living body detection label can be an attack type or a living body type. When the living body detection label is an attack type, it indicates that the sample authentication image sequence is authentication information of an attack type, and at this time, the authentication is dangerous. When the living body detection label is a living body type, it indicates that the sample authentication image sequence is authentication information of a living body type, and at this time, the authentication is safe.

[0166] S604, the second sample training data is input into the initial living body detection model to obtain a sample living body detection result corresponding to the second sample training data;

[0167] In the embodiments of the present application, the living body detection model is a deep learning model for predicting a living body detection result corresponding to a sample authentication image sequence. The second sample training data corresponding to a certain user is input into the living body detection model, and the living body detection model predicts a sample living body detection result corresponding to the second sample training data according to the sample authentication image sequence in the second sample training data.

[0168] In one embodiment, the living body detection model includes a first feature extraction network, a second feature extraction network, a second feature fusion network, and a second fusion feature prediction network. The second sample training data is input into the living body detection model, the first feature extraction network is used to perform feature extraction processing on the sample authentication image sequence to obtain an action completion degree feature corresponding to the sample authentication image sequence, the second feature extraction network is used to perform feature extraction processing on the sample authentication image sequence to obtain an attack clue feature corresponding to the sample authentication image sequence, the second feature fusion network is used to perform feature fusion processing on the action completion degree feature and the attack clue feature to obtain a second sample fusion feature, and the first fusion feature prediction network is used to perform living body detection on the second sample fusion feature to obtain a sample living body detection result corresponding to the second sample training data.

[0169] S606, based on the living body detection loss function, the sample living body detection result and the living body detection label, the living body detection model is supervised and trained and the model parameters of the living body detection model are iteratively updated until the living body detection model converges, and a trained living body detection model is obtained.

[0170] In the embodiment of the present specification, the living body detection loss value corresponding to the sample living body detection result and the living body detection label is calculated based on the living body detection loss function, the model parameters of the living body detection model are updated based on the living body detection loss value, whether the living body detection model after parameter update meets the preset convergence condition is judged, if yes, the training is stopped, and a trained living body detection model is obtained, if not, step S604 is executed.

[0171] In the embodiment of the present specification, by constructing a second sample training data set, the second sample training data includes a sample authentication image sequence collected when a user performs face recognition transaction, and a living body detection label corresponding to the second sample training data, then the second sample training data is input into the initial living body detection model, the sample living body detection result corresponding to the second sample training data is obtained, and finally based on the living body detection loss function, the sample living body detection result and the living body detection label, the living body detection model is supervised and trained and the model parameters of the living body detection model are iteratively updated until the living body detection model converges, and a trained living body detection model is obtained; by using the embodiment, the living body detection model which can detect the living body of the user can be obtained by deep learning training of the constructed second sample training data.

[0172] Please refer to Figure 10 A flowchart of a living body detection model training method provided by the embodiment of the present specification is shown. The living body detection model training method can specifically include the following steps:

[0173] S702, a second sample training data set is constructed, the second sample training data includes a sample authentication image sequence collected when a user performs face recognition transaction, and a living body detection label corresponding to the second sample training data;

[0174] In the embodiment of the present specification, step S702 please refer to the detailed description of step S602 in another embodiment of the present specification, which will not be repeated here.

[0175] It should be noted that in the embodiment of the present specification, the living body detection model includes a first feature extraction network, a second feature extraction network, a second feature fusion network and a second fusion feature prediction network. Please refer to Figure 11 A structure diagram of a living body detection model provided by the embodiment of the present specification is shown.

[0176] S704, performing feature extraction processing on the sample authentication image sequence by using a first feature extraction network to obtain a sample authentication image sequence

[0177] S706, performing feature extraction processing on the sample authentication image sequence by using a second feature extraction network to obtain attack clue features corresponding to the sample authentication image sequence;

[0178] S708, performing feature fusion processing on the action completion degree features and the attack clue features based on a second feature fusion network to obtain second sample fusion features;

[0179] S710, performing living body detection on the second sample fusion features based on a first fusion feature prediction network to obtain a sample living body detection result corresponding to the second sample training data;

[0180] Steps S704-S710 are specifically as shown in Figure 11 After the second sample training data is input into the living body detection model, the sample authentication image sequence in the second sample training data is input into the first feature extraction network, the sample authentication image sequence is subjected to feature extraction by the first feature extraction network to obtain action completion degree features corresponding to the sample authentication image sequence; the sample authentication image sequence in the second sample training data is input into the second feature extraction network, the sample authentication image sequence is subjected to feature extraction by the second feature extraction network to obtain attack clue features corresponding to the sample authentication image sequence; then, the action completion degree features and the attack clue features are subjected to feature fusion processing based on the second feature fusion network to obtain second sample fusion features, and finally, the second sample fusion features are input into the second fusion feature prediction network, the second sample fusion features are subjected to living body detection by the second fusion feature prediction network to obtain a sample living body detection result corresponding to the second sample training data.

[0181] S712, calculating a living body detection loss value corresponding to the sample living body detection result and a living body detection label based on a living body detection loss function;

[0182] The sample living body detection result is a detection result obtained by the living body detection model performing living body detection on the second sample training data, and the living body detection label is a true value label corresponding to the second sample training data. The living body detection loss value corresponding to the sample living body detection result and the living body detection label is calculated through the living body detection loss function, so as to guide the learning direction of the living body detection model according to the living body detection loss value.

[0183] S714, updating model parameters of the living body detection model based on the living body detection loss value;

[0184] S716, determining whether the parameter-updated live body detection model meets a preset convergence condition, if yes, stopping the training, and obtaining the trained live body detection model, if not, performing step S704.

[0185] In the embodiment of the present application, the deep learning training is performed on the embarrassment perception model based on the constructed second sample training data, and the live body detection model capable of detecting the live body of the user can be obtained. The model uses the first feature extraction network and the second feature extraction network to respectively perform the action completion degree feature and the attack clue feature extraction processing, and performs the live body detection based on the two features, thereby ensuring the prediction accuracy of the model for the live body detection.

[0186] Please refer to Figure 12 A flowchart of a live body detection model training method provided in an embodiment of the present application is shown. The live body detection model training method can specifically include the following steps:

[0187] S802, constructing a second sample training data set, the second sample training data including a sample authentication image sequence collected when the user performs a face recognition transaction, and a live body detection label corresponding to the second sample training data;

[0188] In the embodiment of the present application, step S802 please refer to the detailed description of step S602 in another embodiment of the present application, which will not be repeated here.

[0189] It should be noted that in the embodiment of the present application, the live body detection model includes a first feature extraction network, a second feature extraction network, a second feature fusion network, a second fusion feature prediction network, a first feature prediction network, and a second feature prediction network. Please refer to Figure 13 A structure diagram of a live body detection model provided in an embodiment of the present application is shown.

[0190] S804, performing feature extraction processing on the sample authentication image sequence by using the first feature extraction network, to obtain the action completion degree feature corresponding to the sample authentication image sequence;

[0191] S806, performing live body detection on the action completion degree feature based on the first feature prediction network, to obtain the action completion degree prediction result corresponding to the sample authentication image sequence;

[0192] S808, performing feature extraction processing on the sample authentication image sequence by using the second feature extraction network, to obtain the attack clue feature corresponding to the sample authentication image sequence;

[0193] S810, performing live body detection on the attack clue feature based on the second feature prediction network, to obtain the attack clue detection result corresponding to the sample authentication image sequence;

[0194] S812, performing feature fusion processing on the action completion degree feature and the attack clue feature based on the second feature fusion network to obtain a second sample fusion feature;

[0195] S814, performing living body detection on the second sample fusion feature based on the first fusion feature prediction network to obtain a sample living body detection result corresponding to the second sample training data;

[0196] Steps S804-S814 are specifically as shown in the following table: Figure 13 After the second sample training data is input into the living body detection model, the sample authentication image sequence in the second sample training data is input into the first feature extraction network, the first feature extraction network performs feature extraction on the sample authentication image sequence to obtain an action completion degree feature corresponding to the sample authentication image sequence, and the first feature prediction network performs living body detection on the action completion degree feature to obtain an action completion degree prediction result corresponding to the sample authentication image sequence; the sample authentication image sequence in the second sample training data is input into the second feature extraction network, the second feature extraction network performs feature extraction on the sample authentication image sequence to obtain an attack clue feature corresponding to the sample authentication image sequence, and the second feature prediction network performs living body detection on the attack clue feature to obtain an attack clue detection result corresponding to the sample authentication image sequence; then the second feature fusion network performs feature fusion processing on the action completion degree feature and the attack clue feature to obtain a second sample fusion feature, and finally the second sample fusion feature is input into the second fusion feature prediction network, and the second fusion feature prediction network performs living body detection on the second sample fusion feature to obtain a sample living body detection result corresponding to the second sample training data.

[0197] It should be noted that in the embodiments of the present application, the living body detection loss function includes an action loss function, an attack clue loss function, and a fusion detection loss function, and the living body detection label includes an action completion degree label, an attack clue label, and a fusion living body detection label.

[0198] The action completion degree label is a label set based on the completion degree of the authentication action in the sample authentication image sequence in the second sample training data, and is the action completion degree true value corresponding to the second sample training data. The attack clue label is a label set based on the attack situation corresponding to the sample authentication image sequence in the second sample training data, such as whether there is face spoofing, video spoofing, or other attack situations in the sample authentication image sequence in the second sample training data. The fusion living body detection label is a fusion true value label set based on the action completion degree and the attack clue, and is used to represent whether the second sample training data is of an attack type or a living body type.

[0199] S816, calculating an action completion degree loss value corresponding to the action completion degree prediction result and the action completion degree label based on the action loss function;

[0200] S818, calculating an attack clue loss value corresponding to the attack clue detection result and the attack clue label based on the attack clue loss function;

[0201] S820, calculating a fusion detection loss value corresponding to the sample live body detection result and the fusion live body detection label based on the fusion detection loss function;

[0202] S822, updating the model parameters of the live body detection model based on the action completion degree loss value, the attack clue loss value and the fusion detection loss value;

[0203] S824, judging whether the live body detection model after the parameter update meets a preset convergence condition, if yes, stopping the training to obtain the trained live body detection model, and if not, executing step S804.

[0204] In the embodiment of the present application, the deep learning training is performed on the embarrassment perception model based on the constructed second sample training data, and the live body detection model capable of detecting the live body of the user can be obtained. The model uses the first feature extraction network and the second feature extraction network to respectively perform the action completion degree feature extraction processing and the attack clue feature extraction processing, respectively performs the action completion degree detection, the attack clue detection based on the two kinds of features, and performs the live body detection based on the fusion features of the two kinds of features. Finally, the three kinds of loss values are calculated according to the action loss function, the attack clue loss function and the fusion detection loss function, and the model parameters are updated and trained based on the three kinds of loss values, so as to further ensure the prediction accuracy of the live body detection model for the live body detection.

[0205] Please refer to Figure 14 , a structure schematic diagram of a live body detection device provided by an embodiment of the present application. As shown in the figure, Figure 14 According to some embodiments, the live body detection device 1 comprises an embarrassment score determination module 11, an action sequence determination module 12, an image sequence acquisition module 13 and a live body detection module 14, and specifically comprises:

[0206] The embarrassment score determination module 11 is configured to determine a target embarrassment score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, wherein the current environment image is an environment image in which the target user is located when performing a face recognition transaction, and the user historical data is authentication mode selection data generated when the target user performs identity authentication in the past;

[0207] The action sequence determination module 12 is configured to determine an authentication action sequence corresponding to the target user in a preset authentication action set according to the target embarrassment score;

[0208] An image sequence acquisition module 13 is configured to acquire an authentication action image sequence corresponding to the target user performing the sequence of authentication actions;

[0209] A living body detection module 14 is configured to perform living body detection on the authentication action image sequence to obtain a living body detection result corresponding to the target user.

[0210] Optionally, the embarrassment score determination module 11 is specifically configured to:

[0211] perform feature extraction processing on the user historical data and the current environment image respectively to obtain data features corresponding to the user historical data and image features corresponding to the current environment image;

[0212] perform feature fusion processing on the data features and the image features to obtain first fusion features;

[0213] predictively generate a target embarrassment score corresponding to the target user based on the first fusion features.

[0214] Optionally, the preset authentication action set includes at least one authentication action and an action embarrassment score corresponding to each authentication action, and the action sequence determination module 12 is specifically configured to:

[0215] determine an authentication action sub-set in the preset authentication action set based on the target embarrassment score, and an action embarrassment score corresponding to each authentication action in the authentication action sub-set is less than the target embarrassment score;

[0216] randomly select a preset number of authentication actions in the authentication action sub-set to generate an authentication action sequence.

[0217] Optionally, the living body detection module 14 is specifically configured to:

[0218] perform feature extraction processing on the authentication action image sequence to obtain an action completion degree feature and an attack clue feature corresponding to the authentication action image sequence;

[0219] perform feature fusion processing on the action completion degree feature and the attack clue feature to obtain second fusion features;

[0220] predictively generate a living body detection result corresponding to the target user based on the second fusion features.

[0221] Optionally, please refer to Figure 15 A structural schematic diagram of a living body detection device provided by an embodiment of the present application. As shown in Figure 15 The living body detection device further includes a detection result determination module 15 configured to:

[0222] Determine whether the attack probability value is greater than a preset threshold;

[0223] If the attack probability value is greater than a preset threshold, then the authentication action image sequence is determined to be an attack type;

[0224] If the attack probability value is greater than a preset threshold, then the authentication action image sequence is determined to be of the liveness type.

[0225] In the embodiments of this specification, the target embarrassment score for the target user is first determined based on the user's historical data and the current environment image. The current environment image is the environment in which the target user is located when performing facial recognition transactions. The user historical data is the authentication method selection data generated when the target user has performed identity authentication in the past. Based on the target embarrassment score, the authentication action sequence corresponding to the target user is determined from a preset authentication action set. The authentication action image sequence corresponding to the target user's execution of the authentication action sequence is obtained. Liveness detection is performed on the authentication action image sequence to obtain the liveness detection result corresponding to the target user. By performing embarrassment analysis on the user, the authentication action sequence is adaptively determined based on the embarrassment analysis result, thereby reducing the user's embarrassment during the liveness detection authentication process, improving the liveness detection effect, and enhancing the user experience.

[0226] It should be noted that the liveness detection device provided in the above embodiments is only illustrated by the division of the above functional modules when performing the liveness detection method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the liveness detection device and the liveness detection method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0227] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0228] Please see Figure 16 This is a schematic diagram of the structure of an embarrassment perception model training device provided in an embodiment of this specification. Figure 16 As shown, the embarrassment perception model training device 2 can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, the embarrassment perception model training device 2 includes a first sample set construction module 21, an embarrassment score prediction module 22, and an embarrassment model training module 23, specifically including:

[0229] The first sample set construction module 21 is configured to construct a first sample training data set, wherein the first sample training data comprises a sample environment image where a user performs a face recognition transaction, sample historical data generated when the user performs identity authentication in the past, and an embarrassment score label corresponding to the first sample training data.

[0230] The embarrassment score prediction module 22 is configured to input the first sample training data into an initial embarrassment awareness model to obtain a predicted embarrassment score corresponding to the first sample training data.

[0231] The embarrassment model training module 23 is configured to supervise training of the embarrassment awareness model based on an embarrassment awareness loss function, the predicted embarrassment score and the embarrassment score label, and iteratively update model parameters of the embarrassment awareness model until the embarrassment awareness model converges, thereby obtaining a trained embarrassment awareness model.

[0232] Optionally, the embarrassment awareness model comprises a data feature extraction network, an image feature extraction network, a first feature fusion network and a first fused feature prediction network, and the embarrassment score prediction module 22 is specifically configured to:

[0233] The data feature extraction network is adopted to perform feature extraction processing on the sample historical data to obtain sample data features corresponding to the sample historical data.

[0234] The image feature extraction network is adopted to perform feature extraction processing on the sample environment image to obtain sample image features corresponding to the sample environment image.

[0235] The first feature fusion network is adopted to perform feature fusion processing on the sample data features and the sample image features to obtain first sample fused features.

[0236] The first fused feature prediction network is adopted to perform embarrassment prediction on the first sample fused features to obtain the predicted embarrassment score corresponding to the first sample training data.

[0237] Optionally, the embarrassment model training module 23 is specifically configured to:

[0238] The embarrassment awareness loss function is adopted to calculate an embarrassment awareness loss value corresponding to the predicted embarrassment score and the embarrassment score label.

[0239] The model parameters of the embarrassment awareness model are updated based on the embarrassment awareness loss value.

[0240] determine whether the embarrassment perception model after the parameter update meets a preset convergence condition, if yes, stop the training, and obtain the embarrassment perception model after the training, and if not, execute the step of inputting the first sample training data into the initial embarrassment perception model to obtain the predicted embarrassment score corresponding to the first sample training data.

[0241] Optionally, the embarrassment perception model further comprises a data feature prediction network and an image feature prediction network, and the embarrassment score prediction module 22 is further configured to:

[0242] perform embarrassment prediction on the sample data features based on the data feature prediction network to obtain a data embarrassment score corresponding to the sample historical data;

[0243] perform embarrassment prediction on the sample image features based on the image feature prediction network to obtain an image embarrassment score corresponding to the sample environmental image.

[0244] Optionally, the embarrassment perception loss function comprises a data loss function, an image loss function and a fusion perception loss function, the embarrassment score label comprises a data embarrassment label, an image embarrassment label and a fusion embarrassment label, and the embarrassment model training module 23 is specifically configured to:

[0245] calculate a data embarrassment loss value corresponding to the data embarrassment score and the data embarrassment label based on the data loss function;

[0246] calculate an image embarrassment loss value corresponding to the image embarrassment score and the image embarrassment label based on the image loss function;

[0247] calculate a fusion embarrassment loss value corresponding to the predicted embarrassment score and the fusion embarrassment label based on the fusion perception loss function;

[0248] update the model parameters of the embarrassment perception model based on the data embarrassment loss value, the image embarrassment loss value and the fusion embarrassment loss value;

[0249] determine whether the embarrassment perception model after the parameter update meets a preset convergence condition, if yes, stop the training, and obtain the embarrassment perception model after the training, and if not, execute the step of inputting the first sample training data into the initial embarrassment perception model to obtain the predicted embarrassment score corresponding to the first sample training data.

[0250] In this embodiment, a first sample training dataset is constructed. The first sample training data includes sample environment images of the location where the user performs facial recognition, historical sample data generated during user identity authentication, and embarrassment score labels corresponding to the first sample training data. Then, the first sample training data is input into an initial embarrassment perception model to obtain the predicted embarrassment score corresponding to the first sample training data. Finally, the embarrassment perception model is trained under supervision based on the embarrassment perception loss function, the predicted embarrassment score, and the embarrassment score labels, and the model parameters of the embarrassment perception model are iteratively updated until the embarrassment perception model converges, resulting in a trained embarrassment perception model. By using this embodiment, deep learning training of the embarrassment perception model based on the first sample training data containing sample historical data and sample environment images can yield an embarrassment perception model that can accurately predict the embarrassment score of the user.

[0251] It should be noted that the awkwardness perception model training device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the awkwardness perception model training method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the awkwardness perception model training device and the awkwardness perception model training method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0252] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0253] Please see Figure 17 This is a schematic diagram of the structure of a liveness detection model training device provided in an embodiment of this specification. Figure 17 As shown, the liveness detection model training device 3 can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, the liveness detection model training device 3 includes a second sample set construction module 31, a sample liveness detection module 32, and a liveness model training module 33, specifically including:

[0254] The second sample set construction module 31 is used to construct a second sample training dataset. The second sample training dataset includes sample authentication image sequences collected by the user performing face recognition transactions, and liveness detection labels corresponding to the second sample training dataset.

[0255] The sample liveness detection module 32 is used to input the second sample training data into the initial liveness detection model to obtain the sample liveness detection result corresponding to the second sample training data;

[0256] The living body model training module 33 is configured to perform supervised training on the living body detection model based on the living body detection loss function, the sample living body detection result and the living body detection label, and iteratively update model parameters of the living body detection model until the living body detection model converges, so as to obtain a trained living body detection model.

[0257] Optionally, the living body detection model comprises a first feature extraction network, a second feature extraction network, a second feature fusion network and a second fusion feature prediction network, and the sample living body detection module 32 is specifically configured to:

[0258] perform feature extraction processing on the sample authentication image sequence by using the first feature extraction network, so as to obtain action completion degree features corresponding to the sample authentication image sequence;

[0259] perform feature extraction processing on the sample authentication image sequence by using the second feature extraction network, so as to obtain attack clue features corresponding to the sample authentication image sequence;

[0260] perform feature fusion processing on the action completion degree features and the attack clue features based on the second feature fusion network, so as to obtain second sample fusion features;

[0261] perform living body detection on the second sample fusion features based on the first fusion feature prediction network, so as to obtain a sample living body detection result corresponding to the second sample training data.

[0262] Optionally, the living body model training module 33 is specifically configured to:

[0263] calculate a living body detection loss value corresponding to the sample living body detection result and the living body detection label based on the living body detection loss function;

[0264] update the model parameters of the living body detection model based on the living body detection loss value;

[0265] determine whether the living body detection model after parameter updating meets a preset convergence condition, if yes, stop training, so as to obtain a trained living body detection model, and if not, perform the step of inputting the second sample training data into the initial living body detection model, so as to obtain a sample living body detection result corresponding to the second sample training data.

[0266] Optionally, the living body detection model further comprises a first feature prediction network and a second feature prediction network, and the sample living body detection module 32 is further configured to:

[0267] perform living body detection on the action completion degree features based on the first feature prediction network, so as to obtain an action completion degree prediction result corresponding to the sample authentication image sequence;

[0268] perform liveness detection on the attack clue feature based on the second feature prediction network to obtain an attack clue detection result corresponding to the sample authentication image sequence.

[0269] Optionally, the liveness detection loss function includes an action loss function, an attack clue loss function, and a fusion detection loss function, the liveness detection label includes an action completion degree label, an attack clue label, and a fusion liveness detection label, and the liveness model training module 33 is specifically configured to:

[0270] calculate an action completion degree loss value corresponding to the action completion degree prediction result and the action completion degree label based on the action loss function;

[0271] calculate an attack clue loss value corresponding to the attack clue detection result and the attack clue label based on the attack clue loss function;

[0272] calculate a fusion detection loss value corresponding to the sample liveness detection result and the fusion liveness detection label based on the fusion detection loss function;

[0273] update the model parameters of the liveness detection model based on the action completion degree loss value, the attack clue loss value, and the fusion detection loss value;

[0274] determine whether the liveness detection model after the parameter update meets a preset convergence condition, if yes, stop training to obtain the trained liveness detection model, and if not, perform the step of inputting the second sample training data into the initial liveness detection model to obtain the sample liveness detection result corresponding to the second sample training data.

[0275] In the embodiments of the present disclosure, by constructing a second sample training data set, the second sample training data includes a sample authentication image sequence collected when a user performs a face recognition transaction and a liveness detection label corresponding to the second sample training data, then the second sample training data is input into an initial liveness detection model to obtain a sample liveness detection result corresponding to the second sample training data, and finally the liveness detection model is supervised and trained based on a liveness detection loss function, the sample liveness detection result, and the liveness detection label, and the model parameters of the liveness detection model are iteratively updated until the liveness detection model converges, thereby obtaining a trained liveness detection model. By using the present embodiment, the liveness detection model capable of performing liveness detection on a user can be obtained by performing deep learning training on the constructed second sample training data.

[0276] It should be noted that the liveness detection model training device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the liveness detection model training method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the liveness detection model training device and the liveness detection model training method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0277] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0278] This specification also provides an embodiment of a computer storage medium that can store multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-13 The liveness detection method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-13 The specific details of the illustrated embodiments will not be elaborated here.

[0279] This specification also provides a computer program product that stores at least one instruction, said at least one instruction being loaded and executed by the processor as described above. Figures 1-13 The liveness detection method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-13 The specific details of the illustrated embodiments will not be elaborated here.

[0280] Please refer to Figure 18 This is a structural block diagram of an electronic device provided in an embodiment of this specification. The electronic device in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 can be connected via the bus 150.

[0281] The processor 110 can include one or more processing cores. The processor 110 connects various parts within the terminal through various interfaces and lines, performs various functions of the terminal 100 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Alternatively, the processor 110 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 110 can integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes an operating system, a user interface, and an application program; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 110, but can be implemented by a separate communication chip.

[0282] The memory 120 can include a random access memory (RAM) and can also include a read-only memory (ROM). Alternatively, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, codes, code sets or instruction sets.

[0283] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In the embodiment of the present specification, the input device 130 can be a temperature sensor for obtaining the operating temperature of the terminal. The output device 140 can be a speaker for outputting an audio signal.

[0284] In addition, those skilled in the art can understand that the structure of the terminal shown in the above-mentioned drawings does not constitute a limitation on the terminal, and the terminal can include more or fewer components than the drawings, or combine certain components, or different component arrangements. For example, the terminal also includes radio frequency circuitry, an input unit, a sensor, audio circuitry, a wireless fidelity (WIFI) module, a power supply, a Bluetooth module, and the like, which are not described here.

[0285] In the embodiments of the present specification, the execution subject of each step can be the terminal introduced above. Alternatively, the execution subject of each step is an operating system of the terminal. The operating system can be an Android system, an IOS system, or other operating systems, and the embodiments of the present specification do not limit this.

[0286] In Figure 18 In the electronic device of the present specification, the processor 110 can be configured to call a living body detection program stored in the memory 120 and execute to implement the living body detection method as described in the various method embodiments of the present specification.

[0287] In the embodiments of the present specification, first, a target embarrassment score corresponding to a target user is determined based on user historical data corresponding to the target user and a current environment image, the current environment image is an environment image in which the target user is located when performing a face recognition transaction, the user historical data is authentication mode selection data generated when the target user performs identity authentication in the past, a target authentication action sequence corresponding to the target user is determined in a preset authentication action set according to the target embarrassment score, an authentication action image sequence corresponding to the target user when performing the authentication action sequence is obtained, living body detection is performed on the authentication action image sequence to obtain a living body detection result corresponding to the target user, the embarrassment of the user is analyzed, and the authentication action sequence is adaptively determined according to the embarrassment analysis result of the user, so as to reduce the embarrassment of the user in the living body detection authentication process, improve the living body detection effect, and improve the user experience.

[0288] Those skilled in the art can clearly understand that the technical solutions of the present specification can be implemented by means of software and / or hardware. The "unit" and "module" in the present specification refer to software and / or hardware that can independently complete or cooperate with other components to complete a specific function, and the hardware can be, for example, a Field-Programmable Gate Array (FPGA), an Integrated Circuit (IC), and the like.

[0289] It should be noted that, for the foregoing method embodiments, the sequences of the described actions are not necessarily required to achieve the objects of the present application, and certain steps can be performed in other sequences or even concurrently. Additionally, the described embodiments are merely provided as examples, and not all of the actions described are necessarily required to achieve the objects of the present application.

[0290] In the above embodiments, the description of each embodiment focuses on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0291] In the several embodiments provided in the present specification, it should be understood that the disclosed apparatus can be implemented in other manners. For example, the described apparatus embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, or the among different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electric, mechanical or other forms.

[0292] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. In actual implementation, some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.

[0293] In addition, each functional unit in the embodiments of the present specification can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.

[0294] Those skilled in the art can understand that all or part of the steps in the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable memory, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0295] The above-described embodiments are merely illustrative for the present specification and cannot limit the scope of the present specification. That is, any equivalent changes and modifications made in accordance with the teachings of the present specification are still within the scope of the present specification. Other embodiments of the present specification will be easily suggested to those skilled in the art upon consideration of the present specification and practice of the disclosure herein. The present specification is intended to encompass any variations, uses, or adaptive changes of the present specification following the general principles of the present specification and including common knowledge or conventional technical means not described in the present specification. The present specification and embodiments are merely considered as exemplary, and the scope and spirit of the present specification are defined by the claims.

Claims

1. A living body detection method, comprising: determining a target embarrassment score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, the current environment image being an environment image in which the target user is located when performing a face recognition transaction, and the user historical data being authentication mode selection data generated when the target user historically performs identity authentication; determining an authentication action sequence corresponding to the target user in a preset authentication action set according to the target embarrassment score, the preset authentication action set including at least one authentication action and an action embarrassment score corresponding to each authentication action, and an action embarrassment score of an authentication action in the authentication action sequence being less than the target embarrassment score; obtaining an authentication action image sequence corresponding to the target user when the target user performs the authentication action sequence; performing living body detection on the authentication action image sequence to obtain a living body detection result corresponding to the target user.

2. The method of claim 1, wherein the target embarrassment score corresponding to the target user is determined based on the user historical data corresponding to the target user and the current environment image, comprising: performing feature extraction processing on the user historical data and the current environment image respectively to obtain data features corresponding to the user historical data and image features corresponding to the current environment image; performing feature fusion processing on the data features and the image features to obtain first fusion features; generating the target embarrassment score corresponding to the target user based on the first fusion features.

3. The method of claim 1, wherein the authentication action sequence corresponding to the target user is determined in the preset authentication action set according to the target embarrassment score, comprising: determining an authentication action sub-set in the preset authentication action set based on the target embarrassment score, each authentication action in the authentication action sub-set corresponding to an action embarrassment score less than the target embarrassment score; randomly selecting a preset number of authentication actions in the authentication action sub-set to generate an authentication action sequence.

4. The method of claim 1, wherein the living body detection result corresponding to the target user is obtained by performing living body detection on the authentication action image sequence, comprising: performing feature extraction processing on the authentication action image sequence to obtain action completion degree features and attack clue features corresponding to the authentication action image sequence; performing feature fusion processing on the action completion degree features and the attack clue features to obtain second fusion features; generating the living body detection result corresponding to the target user based on the second fusion features.

5. The method of claim 1, wherein the living body detection result is an attack probability value of the authentication action image sequence being an attack, and after the living body detection result corresponding to the target user is obtained by performing living body detection on the authentication action image sequence, the method further comprises: determining whether the attack probability value is greater than a preset threshold value; if the attack probability value is greater than the preset threshold value, determining that the authentication action image sequence is of an attack type; if the attack probability value is not greater than the preset threshold value, determining that the authentication action image sequence is of a living body type.

6. An embarrassment perception model training method, the method comprising: constructing a first sample training data set, the first sample training data comprising a sample environment image where a user performs a face brushing transaction, sample historical data generated when the user performs identity authentication in the past, and an embarrassment score label corresponding to the first sample training data; inputting the first sample training data into an initial embarrassment perception model to obtain a predicted embarrassment score corresponding to the first sample training data; supervising training of the embarrassment perception model based on an embarrassment perception loss function, the predicted embarrassment score, and the embarrassment score label, and iteratively updating model parameters of the embarrassment perception model until the embarrassment perception model converges, to obtain a trained embarrassment perception model; the embarrassment perception model is configured to determine a target embarrassment score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, and the target embarrassment score is configured to determine an authentication action sequence corresponding to the target user from a preset authentication action set, the preset authentication action set comprising at least one authentication action and an action embarrassment score corresponding to each authentication action, and the action embarrassment score of an authentication action in the authentication action sequence being less than the target embarrassment score.

7. The method of claim 6, wherein the embarrassment perception model comprises a data feature extraction network, an image feature extraction network, a first feature fusion network, and a first fused feature prediction network, and the inputting the first sample training data into the initial embarrassment perception model to obtain the predicted embarrassment score corresponding to the first sample training data comprises: performing feature extraction processing on the sample historical data using the data feature extraction network to obtain sample data features corresponding to the sample historical data; performing feature extraction processing on the sample environment image using the image feature extraction network to obtain sample image features corresponding to the sample environment image; performing feature fusion processing on the sample data features and the sample image features based on the first feature fusion network to obtain first sample fused features; performing embarrassment prediction on the first sample fused features based on the first fused feature prediction network to obtain the predicted embarrassment score corresponding to the first sample training data.

8. The method of claim 6, wherein the supervising training of the embarrassment perception model based on the embarrassment perception loss function, the predicted embarrassment score, and the embarrassment score label, and the iteratively updating model parameters of the embarrassment perception model until the embarrassment perception model converges to obtain the trained embarrassment perception model comprises: calculating an embarrassment perception loss value corresponding to the predicted embarrassment score and the embarrassment score label based on the embarrassment perception loss function; updating the model parameters of the embarrassment perception model based on the embarrassment perception loss value. determine whether the embarrassment perception model after updating of the parameters meets a preset convergence condition, if yes, stop training, and obtain the embarrassment perception model after training, and if not, perform the step of inputting the first sample training data into the initial embarrassment perception model to obtain the predicted embarrassment score corresponding to the first sample training data. 9.The method of claim 7, wherein the embarrassment perception model further comprises a data feature prediction network and an image feature prediction network, and the method further comprises: performing embarrassment prediction on the sample data features based on the data feature prediction network to obtain a data embarrassment score corresponding to the sample historical data; and performing embarrassment prediction on the sample image features based on the image feature prediction network to obtain an image embarrassment score corresponding to the sample environment image. 10.The method of claim 9, wherein the embarrassment perception loss function comprises a data loss function, an image loss function, and a fusion perception loss function, the embarrassment score label comprises a data embarrassment label, an image embarrassment label, and a fusion embarrassment label, and the embarrassment perception model is supervised and trained based on the embarrassment perception loss function, the predicted embarrassment score, and the embarrassment score label, and the model parameters of the embarrassment perception model are iteratively updated until the embarrassment perception model converges, and the embarrassment perception model after training is obtained, comprising: calculating a data embarrassment loss value corresponding to the data embarrassment score and the data embarrassment label based on the data loss function; calculating an image embarrassment loss value corresponding to the image embarrassment score and the image embarrassment label based on the image loss function; calculating a fusion embarrassment loss value corresponding to the predicted embarrassment score and the fusion embarrassment label based on the fusion perception loss function; updating the model parameters of the embarrassment perception model based on the data embarrassment loss value, the image embarrassment loss value, and the fusion embarrassment loss value; and determining whether the embarrassment perception model after updating of the parameters meets a preset convergence condition, if yes, stopping training, and obtaining the embarrassment perception model after training, and if not, performing the step of inputting the first sample training data into the initial embarrassment perception model to obtain the predicted embarrassment score corresponding to the first sample training data. 11.A living body detection model training method, comprising: constructing a second sample training data set, wherein the second sample training data comprises a sample authentication action image sequence collected when a user performs a face recognition transaction and a living body detection label corresponding to the second sample training data, an action embarrassment score of a sample authentication action in the sample authentication action image sequence is less than a target embarrassment score of the user when performing the face recognition transaction, and the target embarrassment score is determined based on user historical data of the user and a current environment image when performing the face recognition transaction; and inputting the second sample training data into an initial living body detection model to obtain a sample living body detection result corresponding to the second sample training data. ​ ​ ​ ​ ​ ​ ​ ​ ​ The living body detection model is supervised trained and the model parameters of the living body detection model are iteratively updated based on a living body detection loss function, the sample living body detection result and the living body detection label until the living body detection model converges, and a trained living body detection model is obtained. The initial living body detection model comprises a first feature extraction network, a second feature extraction network, a second feature fusion network and a second fusion feature prediction network, and the second sample training data is input into the initial living body detection model to obtain the sample living body detection result corresponding to the second sample training data, which comprises: The first feature extraction network is used to perform feature extraction processing on the sample authentication action image sequence to obtain the action completion degree feature corresponding to the sample authentication action image sequence; The second feature extraction network is used to perform feature extraction processing on the sample authentication action image sequence to obtain the attack clue feature corresponding to the sample authentication action image sequence; the attack clue feature is a feature representing whether there is an attack in the sample authentication action image sequence; The second feature fusion network is used to perform feature fusion processing on the action completion degree feature and the attack clue feature to obtain the second sample fusion feature; The second fusion feature prediction network is used to perform living body detection on the second sample fusion feature to obtain the sample living body detection result corresponding to the second sample training data.

12. The method of claim 11, wherein the living body detection model is supervised trained and the model parameters of the living body detection model are iteratively updated based on a living body detection loss function, the sample living body detection result and the living body detection label until the living body detection model converges, and a trained living body detection model is obtained, which comprises: The living body detection loss value corresponding to the sample living body detection result and the living body detection label is calculated based on the living body detection loss function; The model parameters of the living body detection model are updated based on the living body detection loss value; It is judged whether the living body detection model after parameter updating meets a preset convergence condition, if yes, the training is stopped to obtain the trained living body detection model, and if not, the step of inputting the second sample training data into the initial living body detection model to obtain the sample living body detection result corresponding to the second sample training data is performed.

13. The method of claim 11, wherein the living body detection model further comprises a first feature prediction network and a second feature prediction network, and the method further comprises: The first feature prediction network is used to perform living body detection on the action completion degree feature to obtain the action completion degree prediction result corresponding to the sample authentication image sequence; The second feature prediction network is used to perform living body detection on the attack clue feature to obtain the attack clue detection result corresponding to the sample authentication image sequence.

14. The method of claim 13, wherein the liveness detection loss function comprises an action loss function, an attack clue loss function, and a fusion detection loss function, the liveness detection label comprises an action completion degree label, an attack clue label, and a fusion liveness detection label, and the liveness detection model is supervised and trained based on the liveness detection loss function, the sample liveness detection result, and the liveness detection label, and the model parameters of the liveness detection model are iteratively updated until the liveness detection model converges, to obtain a trained liveness detection model, comprising: calculating an action completion degree loss value corresponding to the action completion degree prediction result and the action completion degree label based on the action loss function; calculating an attack clue loss value corresponding to the attack clue detection result and the attack clue label based on the attack clue loss function; calculating a fusion detection loss value corresponding to the sample liveness detection result and the fusion liveness detection label based on the fusion detection loss function; updating the model parameters of the liveness detection model based on the action completion degree loss value, the attack clue loss value, and the fusion detection loss value; determining whether the liveness detection model after parameter updating meets a preset convergence condition, and if so, stopping training to obtain the trained liveness detection model, and if not, performing the step of inputting the second sample training data into an initial liveness detection model to obtain a sample liveness detection result corresponding to the second sample training data.

15. A liveness detection device, comprising: an embarrassment score determination module configured to determine a target embarrassment score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, the current environment image being an environment image in which the target user is located when performing a face recognition transaction, and the user historical data being authentication mode selection data generated when the target user historically performs identity authentication; an action sequence determination module configured to determine an authentication action sequence corresponding to the target user in a preset authentication action set according to the target embarrassment score, the preset authentication action set comprising at least one authentication action and an action embarrassment score corresponding to each authentication action, and an action embarrassment score of an authentication action in the authentication action sequence being less than the target embarrassment score; an image sequence acquisition module configured to acquire an authentication action image sequence corresponding to the target user performing the authentication action sequence; a liveness detection module configured to perform liveness detection on the authentication action image sequence to obtain a liveness detection result corresponding to the target user.

16. A liveness detection model training device, comprising: a first sample set construction module configured to construct a first sample training data set, the first sample training data comprising a sample environment image in which a user performs a face recognition transaction, sample historical data generated when the user historically performs identity authentication, and an embarrassment score label corresponding to the first sample training data; the sample historical data being authentication mode selection data generated when the user historically performs identity authentication. an embarrassment score prediction module, configured to input the first sample training data into an initial embarrassment awareness model to obtain a predicted embarrassment score corresponding to the first sample training data; an embarrassment model training module, configured to supervise training of the embarrassment awareness model based on an embarrassment awareness loss function, the predicted embarrassment score, and the embarrassment score label, and iteratively update model parameters of the embarrassment awareness model until the embarrassment awareness model converges, to obtain a trained embarrassment awareness model; the embarrassment awareness model is configured to determine a target embarrassment score corresponding to a target user based on user historical data corresponding to the target user and a current environment image, and the target embarrassment score is configured to determine an authentication action sequence corresponding to the target user from a preset authentication action set, the preset authentication action set including at least one authentication action and an action embarrassment score corresponding to each authentication action, and an action embarrassment score of an authentication action in the authentication action sequence being less than the target embarrassment score.

17. A living body detection model training apparatus, comprising: a second sample set construction module, configured to construct a second sample training data set, the second sample training data including a sample authentication action image sequence collected when a user performs a face recognition transaction, and a living body detection label corresponding to the second sample training data; an action embarrassment score of a sample authentication action in the sample authentication action image sequence being less than a target embarrassment score of the user when performing the face recognition transaction; and the target embarrassment score being determined based on user historical data of the user and a current environment image when performing the face recognition transaction; a sample living body detection module, configured to input the second sample training data into an initial living body detection model to obtain a sample living body detection result corresponding to the second sample training data; the initial living body detection model including a first feature extraction network, a second feature extraction network, a second feature fusion network, and a second fusion feature prediction network; the inputting of the second sample training data into the initial living body detection model to obtain the sample living body detection result corresponding to the second sample training data including: performing feature extraction processing on the sample authentication action image sequence by using the first feature extraction network to obtain an action completion degree feature corresponding to the sample authentication action image sequence; performing feature extraction processing on the sample authentication action image sequence by using the second feature extraction network to obtain an attack clue feature corresponding to the sample authentication action image sequence; the attack clue feature representing whether there is an attack in the sample authentication action image sequence; performing feature fusion processing on the action completion degree feature and the attack clue feature by using the second feature fusion network to obtain a second sample fusion feature; and performing living body detection on the second sample fusion feature by using the second fusion feature prediction network to obtain the sample living body detection result corresponding to the second sample training data. The live body model training module is configured to supervise training of the live body detection model based on the live body detection loss function, the sample live body detection result and the live body detection label, and iteratively update model parameters of the live body detection model until the live body detection model converges, thereby obtaining a trained live body detection model.

18. A storage medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-14.

19. An electronic device, comprising: The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-14. The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-14.

20. A computer program product having stored thereon at least one instruction, the computer program product comprising: The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-14.

Citation Information

Patent Citations

  • Identity authentication method and system

    CN109756458A

  • Identity authentication method and device based on privacy protection, and equipment

    CN113987447A

  • Living body detection method and device, equipment and storage medium

    CN115862159A