Identity authentication model training method, apparatus and device, and identity authentication method, apparatus and device
By generating and style-transferring user facial images using generative artificial intelligence technology, an identity authentication model is trained to identify the risks of deepfake technology, thus solving the problem of deepfake technology abuse in identity authentication and achieving security assurance for identity authentication.
Patent Information
- Application Number
- PCT/CN2024/129140
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-14
- Filing Date
- 2024-10-31
- Publication Date
- 2025-12-18
AI Technical Summary
Existing technologies cannot accurately identify cases where users use deepfake technology to commit identity fraud during the identity authentication process, leading to the abuse of deepfake technology and affecting the security of identity authentication.
Generative artificial intelligence technology is used to generate user facial images and perform style transfer processing to generate ID card facial images. Labeled data is then used to train an identity authentication model to identify the risks of deepfake technology.
Accurately identify users using deepfake technology for identity fraud during the identity authentication process, reduce the risk of deepfake technology being abused, and ensure the security of identity authentication.
Smart Images

Figure CN2024129140_18122025_PF_FP_ABST
Abstract
Description
An identity authentication model training method, an identity authentication method, an apparatus, and a device
[0001] The present application claims priority to the Chinese patent application No. 202410775563.1, filed on June 14, 2024, and entitled "An identity authentication model training method, an identity authentication method, an apparatus, and a device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the technical field of identity verification, and in particular to an identity authentication model training method, an identity authentication method, an apparatus, and a device. BACKGROUND
[0003] Deepfake can refer to an artificial intelligence technology that learns from user image samples or user video samples by means of neural network technology, and then splices and synthesizes false content such as user voice, facial expression, and body movement. In deepfake technology, the most common way is artificial intelligence (AI) face swapping technology, which makes it possible to tamper with or generate highly realistic and difficult-to-distinguish user audio and video content. Observers often cannot distinguish the authenticity by naked eye. In recent years, as the threshold of deepfake technology is lowered, criminals begin to use deepfake technology to steal or fake personal identities for identity authentication in the process of conducting business, thereby bringing new challenges to the normal operation of the business.
[0004] Therefore, how to accurately identify the case that a user uses deepfake technology for identity fraud in the identity authentication process, so as to reduce the risk of abuse of deepfake technology and ensure the security of identity authentication, has become a technical problem to be solved.
[0005] SUMMARY
[0006] The identity authentication model training method, the identity authentication method, the apparatus, and the device provided by the embodiments of the present application can accurately identify the case that a user uses deepfake technology for identity fraud in the identity authentication process, so as to reduce the risk of abuse of deepfake technology and ensure the security of identity authentication.
[0007] To solve the above technical problems, the embodiments of the present application are implemented as follows:
[0008] The identity authentication model training method provided by the embodiments of the present application comprises the following steps.
[0009] obtain a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using a generative artificial intelligence technology; and the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample;
[0010] determine first label data possessed by the first training sample; the first label data is used to identify that the first sample user has a risk of using a deep fake technology to perform identity authentication;
[0011] train an identity authentication model by using the first training sample with the first label data, to obtain a target identity authentication model.
[0012] The identity authentication method provided by the embodiments of the present specification comprises:
[0013] obtain a user face image of a target user collected in an identity authentication process;
[0014] obtain a certificate face image in a certificate image of the target user;
[0015] perform processing on the user face image and the certificate face image by using a target identity authentication model, to obtain identity authentication result information representing whether the target user has a risk of using a deep fake technology to perform identity authentication; wherein the target identity authentication model is a model trained by using the identity authentication model training method in the embodiments of the present specification.
[0016] The identity authentication model training device provided by the embodiments of the present specification comprises:
[0017] a first obtaining module, configured to obtain a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using a generative artificial intelligence technology; and the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample;
[0018] a first determining module, configured to determine first label data possessed by the first training sample; the first label data is used to identify that the first sample user has a risk of using a deep fake technology to perform identity authentication;
[0019] a first training module, configured to train an identity authentication model by using the first training sample with the first label data, to obtain a target identity authentication model.
[0020] An identity authentication device provided by an embodiment of the present specification comprises:
[0021] A first obtaining module is configured to obtain a user face image of a target user collected in an identity authentication process.
[0022] A second obtaining module is configured to obtain a certificate face image in a certificate image of the target user.
[0023] An identity authentication module is configured to process the user face image and the certificate face image by using a target identity authentication model to obtain identity authentication result information indicating whether the target user has a risk of using deep fake technology for identity authentication.
[0024] An identity authentication model training device provided by an embodiment of the present specification comprises:
[0025] At least one processor; and
[0026] A memory in communication connection with the at least one processor; wherein
[0027] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:
[0028] Obtain a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using generative artificial intelligence technology; and the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample.
[0029] Determine first label data possessed by the first training sample; the first label data is used to identify that the first sample user has a risk of using deep fake technology for identity authentication.
[0030] Train an identity authentication model by using the first training sample with the first label data to obtain a target identity authentication model.
[0031] An identity authentication device provided by an embodiment of the present specification comprises:
[0032] At least one processor; and
[0033] A memory in communication connection with the at least one processor; wherein
[0034] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:
[0035] obtain a user face image of a target user collected in an identity authentication process;
[0036] obtain a certificate face image in a certificate image of the target user;
[0037] processing the user face image and the certificate face image by using a target identity authentication model to obtain identity authentication result information representing whether the target user has a risk of using deep fake technology for identity authentication; wherein the target identity authentication model is a model trained by using an identity authentication model training method in the embodiments of the present specification.
[0038] At least one embodiment provided in the present specification can achieve the following beneficial effects:
[0039] obtain a first user face image of a first sample user generated by using generative artificial intelligence technology, obtain a first certificate face image sample by performing style transfer processing on the first user face image sample, so as to generate a first training sample with a first label data representing that the first sample user has a risk of using deep fake technology for identity authentication according to the forged first user face image sample and the first certificate face image sample; and train the identity authentication model by using the first training sample with the first label data, so as to obtain the target identity authentication model required for identity authentication of a user with identity verification requirement. In this scheme, the first training sample used for training the target identity authentication model contains the forged face image generated by using the generative artificial intelligence technology and the image style transfer technology, so that the first training sample has consistent feature information with the face image used by the criminal when using the deep fake technology for identity fraud, thereby the target identity authentication model can accurately identify the case of using deep fake technology for identity fraud by the user in the identity authentication process, so as to reduce the risk of abuse of deep fake technology and ensure the security of identity authentication. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0041] FIG. 1 is a schematic diagram of an application scenario of an identity authentication scheme provided in an embodiment of the present specification;
[0042] FIG. 2 is a schematic diagram of a flow of a method for training an identity authentication model provided in an embodiment of the present specification;
[0043] FIG. 3 is a schematic diagram of a structure of a target identity authentication model provided in an embodiment of the present specification;
[0044] FIG. 4 is a schematic diagram of a flow of a method for identity authentication provided in an embodiment of the present specification;
[0045] FIG. 5 is a schematic diagram of a swim lane flow corresponding to the identity authentication scheme in FIG. 2 and FIG. 4 provided in an embodiment of the present specification;
[0046] FIG. 6 is a schematic diagram of a structure of an identity authentication model training apparatus corresponding to FIG. 2 provided in an embodiment of the present specification;
[0047] FIG. 7 is a schematic diagram of a structure of an identity authentication apparatus corresponding to FIG. 4 provided in an embodiment of the present specification;
[0048] FIG. 8 is a schematic diagram of a structure of an identity authentication model training device corresponding to FIG. 2 provided in an embodiment of the present specification;
[0049] FIG. 9 is a schematic diagram of a structure of an identity authentication device corresponding to FIG. 4 provided in an embodiment of the present specification. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical scheme and advantages of one or more embodiments of the present specification clearer, the technical scheme of one or more embodiments of the present specification will be described clearly and completely below in conjunction with specific embodiments of the present specification and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of one or more embodiments of the present specification.
[0051] The technical scheme provided by each embodiment of the present specification will be described in detail below in conjunction with the drawings.
[0052] With the lowering of the threshold of deepfake technology in the prior art, criminals begin to use deepfake technology to forge user face images, and use the forged user face images to make forged certificates containing face images, so that the consistency of the user face images used by the criminals in the identity authentication process and the face images contained in the certificates is high in terms of face posture, expression, makeup, etc. Because the existing identity authentication model often uses the forged user face images and the forged certificate images with high consistency of face attribute information provided by the criminals, the criminals can pass the identity authentication, so that the identity fraud of the user in the identity authentication process using deepfake technology cannot be accurately distinguished, the risk of abuse of deepfake technology exists, and the security of identity authentication in the business operation process cannot be guaranteed.
[0053] In order to solve the defects in the prior art, the present scheme provides the following embodiments:
[0054] FIG. 1 is a schematic diagram of an application scenario of an identity authentication scheme provided in the present specification.
[0055] As shown in FIG. 1, a service provider can use a device 101 to obtain a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample. The first user face image sample can be a user face image generated by using generative artificial intelligence technology. The first certificate face image sample can be a certificate face image obtained by performing style transfer processing on the first user face image sample. A first label data for identifying that the first sample user has a risk of using deepfake technology for identity authentication is set for the first training sample. The identity authentication model is trained by using the first training sample with the first label data, and a target identity authentication model can be obtained.
[0056] In the business operation process, when a target user needs to perform identity authentication, a terminal device 102 can communicate with the device 101 of the service provider, so that the device 101 of the service provider can obtain a user face image collected in real time for the target user in the identity authentication process from the terminal device 102. The device 101 of the service provider can also obtain a certificate face image in a certificate image of the target user. The target identity authentication model is used to process the user face image and the certificate face image, so as to obtain identity authentication result information indicating whether the target user has a risk of using deepfake technology for identity authentication.
[0057] Next, an identity authentication model training method and an identity authentication method provided in the present specification will be described in detail in combination with the accompanying drawings:
[0058] FIG. 2 is a flowchart of a method for training an identity authentication model according to an embodiment of the present specification. From a program perspective, the execution subject of the flowchart can be a device for training a model, or an application program loaded on the device. As shown in FIG. 2, the flowchart can include the following steps:
[0059] Step 202: obtaining a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using generative artificial intelligence technology; and the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample.
[0060] In the embodiments of the present specification, when performing identity authentication on a user, it is usually necessary to collect a user face image in real time to obtain a user face image, and compare the consistency between the user face image and the certificate face image carried in the certificate. If the consistency is high, it can be considered that the user to whom the certificate belongs is the person performing the identity authentication, so that the user can pass the identity authentication. Based on this, when a criminal uses deep fake technology to commit identity fraud, he usually forges a user face image and directly processes the forged user face image as a certificate face image in a forged certificate. Since the user face image and the certificate face image contain highly consistent facial poses, expressions, hairstyles, etc., it is easy for the criminal to successfully pass the identity authentication.
[0061] It can be seen that, in order to accurately identify the case where a user uses deep fake technology to commit identity fraud during identity authentication, a first user face image can be generated by using generative artificial intelligence technology. The first user face image is a forged user face image. In addition, a first certificate face image is obtained by performing style transfer processing on the first user face image. The first certificate face image is a forged certificate face image. At this time, the facial attribute information (for example, facial pose, expression, hairstyle) between the first user face image sample and the first certificate face image sample has a high degree of consistency, which is consistent with the features and characteristics of the image pair of the user face image and the certificate face image used by the criminal when committing identity fraud. Therefore, a first training sample can be generated according to the first user face image sample and the first certificate face image sample to train an identity authentication model.
[0062] The generative artificial intelligence (Artificial Intelligence Generated Content) technology can refer to a new artificial intelligence technology for generating new original content by learning a large-scale data set. It can generate text, pictures, sound, video, code and other content based on algorithms, models and rules. At present, there are various algorithms and models that can generate highly realistic face images or videos. Based on this, generative artificial intelligence technology can be used to generate the first user face image. It can reduce the difficulty of obtaining training samples and avoid the risk of infringing the rights and interests of users when using real user data directly.
[0063] The style transfer can refer to the fusion of the content of one image and the style of another image to generate an image with a novel appearance and a specific style. The style can refer to the texture, color and visual pattern of different spatial scales in the image. Since some certificates contain grayscale images of human faces, and are affected by the printing effect of the certificate, the certificate face images contained in the certificate often have their own specific styles. On this basis, an algorithm or model trained in advance to generate a certificate face style can be used to process the first user face image sample to obtain the first certificate face image sample.
[0064] Step 204: determining first label data possessed by the first training sample; the first label data is used to identify that the first sample user has a risk of using deep fake technology for identity authentication.
[0065] In the embodiments of the present specification, the first training sample has consistency with the features and characteristics of the image pair of the user face image and the certificate face image used by the criminal for identity fraud, so the first training sample can be set with the first label data for identifying that the first sample user to which the first training sample belongs has a risk of using deep fake technology for identity authentication.
[0066] Step 206: training an identity authentication model by using the first training sample with the first label data to obtain a target identity authentication model.
[0067] In the embodiments of the present specification, training the identity authentication model by using the first training sample with the first label data can enable the target identity authentication model learned to learn the features and knowledge of the image pair of the user face image and the certificate face image used by the criminal for identity fraud, so as to accurately identify the case of using deep fake technology for identity authentication by the user.
[0068] The method in FIG. 2 can accurately identify the case that the user uses the deep fake technology to commit identity fraud in the identity authentication process by using the target identity authentication model, so as to reduce the risk of abuse of the deep fake technology and ensure the security of identity authentication, because the first training sample used for training the target identity authentication model contains the fake face images generated by using the generative artificial intelligence technology and the image style transfer technology, and the first training sample has consistent feature information with the face images used by the criminal when committing identity fraud by using the deep fake technology.
[0069] Based on the method in FIG. 2, the present embodiment of the specification also provides some specific implementation schemes of the method, which are described below.
[0070] In the present embodiment of the specification, step 202: obtaining the first training sample containing the first user face image sample of the sample user and the first certificate face image sample, can specifically include:
[0071] Generating the user face image of the first sample user according to the preset prompt word by using the generative artificial intelligence model to obtain the first user face image sample.
[0072] Performing style transfer processing on the first user face image sample by using the pre-trained target image style transfer model to obtain the first certificate face image sample.
[0073] In the present embodiment of the specification, the existing generative artificial intelligence model can be used to generate the face image as the first user face image sample. The generative artificial intelligence model can include but is not limited to Stable Diffusion AI drawing generation tool, Midjourney online image generation platform, StyleGAN image generation model, etc., and no specific limitation is made thereto. In actual application, because the face features of users in different geographical environments often differ, the first user face image sample with specified features and characteristics can be generated by setting the preset prompt word according to actual needs. For example, the preset prompt word can be used to indicate the clothing, hairstyle, facial feature, lighting condition, background environment, etc. of the user, and no specific limitation is made thereto.
[0074] In the present embodiment of the specification, in order to enable the target image style transfer model to generate images with the face style on the certificate, the target image style transfer model usually needs to be pre-trained. Based on this, before performing style transfer processing on the first user face image sample by using the pre-trained image style transfer model, it can also include:
[0075] obtain a second training sample containing a second user face image sample of a second sample user and a second certificate face image sample; wherein the second user face image sample is a real face image obtained by image collection on the face of the second sample user, and the second certificate face image sample is a certificate face image contained in a certificate sample of the second sample user that passes the credibility verification, and the certificate sample is of the same type as the first certificate face image sample.
[0076] train the style transfer model by using the second training sample to obtain a predicted certificate face image output by the style transfer model.
[0077] optimize the parameters of the style transfer model to minimize the difference between the predicted certificate face image and the second certificate face image sample, to obtain a target image style transfer model.
[0078] In the embodiments of the present specification, since the target image style transfer model needs to generate a certificate face image with the style of the face on the certificate according to the input user face image, when training the target image style transfer model, a user face image sample and a certificate face image sample that are of good credibility and belong to the same user can be obtained as the second training sample.
[0079] Specifically, the second user face image sample contained in the second training sample can be obtained by image collection on the live face of a real user (i.e., the second sample user), and the second certificate face image sample contained in the second training sample can be a face image contained in a real and reliable certificate of the second sample user. It can be understood that the service provider has obtained the use authorization of the second sample user before training the model by using the second training sample, so as to avoid damaging the rights and interests of the user.
[0080] In actual application, the second user face image sample contained in the second training sample needs to be input into the style transfer model, at this time, the style transfer model will generate a predicted certificate face image. Since the model training target is to make the visual effect of the predicted certificate face image close to the second certificate face image sample in the second training sample, the parameters of the style transfer model can be optimized to minimize the difference between the predicted certificate face image and the second certificate face image sample, so as to obtain a target image style transfer model that has the ability to generate a certificate face style face image.
[0081] In the embodiments of the present specification, since the multi-modal model can combine and analyze rich visual information and information of other modalities (such as text), capture the relationship between multiple data, and help improve the accuracy and efficiency of the model, the multi-modal model can be widely applied.
[0082] Therefore, if the target image style transfer model is a multi-modal model, the second training sample can further include style description information for the second certificate face image sample.
[0083] Correspondingly, the parameter optimization of the style transfer model aims to minimize the difference between the predicted certificate face image and the second certificate face image sample, and can specifically include:
[0084] The parameter optimization of the style transfer model aims to minimize the difference between the predicted certificate face image and the second certificate face image sample and the style description information.
[0085] In the embodiments of the present specification, the style description information for the second certificate face image sample can include, but is not limited to, at least one of the name, type information of the certificate to which the second certificate face image sample belongs, the human-defined face style name, type information of the second certificate face image sample, visual feature information of the face region, face attribute information, and the like.
[0086] Since it is necessary to ensure that the certificate face image generated by the target image style transfer model is consistent with the style description information for the second certificate face image sample, the parameter optimization of the style transfer model aims to minimize the difference between the predicted certificate face image and the second certificate face image sample and the style description information, so as to ensure that the certificate face image generated by the trained target image style transfer model has high authenticity and meets the actual needs.
[0087] In actual applications, the style transfer model can include, but is not limited to, at least one of a styleGAN model, a CAP-VSTNet model, an InstantStyle model, an AdaIN model, and an InstantID model. Since the InstantStyle model and the InstantID model are multi-modal models, in addition to the second user face image sample and the second certificate face image sample, the second training sample for the InstantStyle model and the InstantID model can also include style description information for the second certificate face image sample. In addition, since the InstantID model has the ability to automatically extract text input information from an input image, if only one type of certificate face image needs to be generated by using the InstantID model, the second training sample can also not include the style description information for the second certificate face image sample, which is not limited in detail.
[0088] In the embodiments of the present specification, since there is usually a time interval between the certificate generation time of the user and the identity authentication time of the user, the face attribute information contained in the user face image and the certificate face image of the same user with high credibility usually has differences, but has a relatively high similarity in face features. The face attribute information contained in the user face image and the certificate face image of different users usually has great differences, and the similarity in face features is also relatively low. In the above two cases, it should be considered that the user does not have the risk of identity authentication by using deep fake technology.
[0089] Based on this, the method in FIG. 2 can further include:
[0090] obtaining a third training sample containing a third sample user's third user face image sample and a third certificate face image sample; wherein the third certificate face image sample and the third user face image sample are face images of different users; or the third certificate face image sample and the third user face image sample are face images belonging to the same user and having different face attribute information.
[0091] determining second label data possessed by the third training sample; the second label data is used to identify that the third sample user does not have the risk of identity authentication by using deep fake technology.
[0092] Correspondingly, step 206: training the identity authentication model to obtain a target identity authentication model, can include:
[0093] The identity authentication model is trained by using the first training sample with the first label data and the third training sample with the second label data, to obtain a target identity authentication model.
[0094] In the embodiments of the present specification, the face attribute information can include but is not limited to face expression, face posture, identity, makeup, and the like. In actual application, the face attribute information of the user face image and the certificate face image collected for the same user when the user performs identity authentication by using the personal certificate is often different. In this case, since the user does not perform identity fraud by using the deep fake technology, the third certificate face image sample and the third user face image sample belonging to the same user and having different face attribute information can be used as the third training sample, and the second label data indicating that the third sample user to which the third training sample belongs has no risk of performing identity authentication by using the deep fake technology is set for the third training sample.
[0095] However, in the scene of performing identity authentication by using the user face image and the certificate face image of different users, although the user has the behavior of identity fraud, it does not belong to the case of performing identity fraud by using the deep fake technology. Therefore, the third certificate face image sample and the third user face image sample belonging to different users can be used as the third training sample, and the second label data indicating that the third sample user to which the third training sample belongs has no risk of performing identity authentication by using the deep fake technology is set for the third training sample.
[0096] By training the identity authentication model by using the first training sample with the first label data and the third training sample with the second label data, the performance of the target identity authentication model obtained by training can be further improved, and details are not described herein.
[0097] In actual application, since the user face image collected in real time for the user often contains a large area of background region, and thus contains more interference information, the certificate face image in the certificate contains a small area of background region, and the background region is often a solid color region, and thus contains less interference information. Based on this, in order to improve the performance of the model, the first user face image sample in the first training sample and the third user face image sample in the third training sample can be images in the face region recognized from the collected user image by using the face recognition model, and the first certificate face image sample in the first training sample and the third certificate face image sample in the third training sample can be images in the face region recognized from the certificate image by using the face recognition model, so as to reduce the interference brought by the background region.
[0098] In an embodiment of the present specification, the identity authentication model can include a feature extraction model for feature extraction on a face image sample, and a risk identification model for identity authentication risk identification according to face image feature data extracted by the feature extraction model. The output layer of the feature extraction model is connected to the input layer of the risk identification model, thereby constituting a (target) identity authentication model.
[0099] Based on this, step 206: training the identity authentication model to obtain a target identity authentication model, which can specifically include:
[0100] The face recognition model is trained using the target training sample to obtain a trained face recognition model; wherein the target training sample includes at least one of the first training sample and the third training sample.
[0101] The model from the input layer to the specified network layer in the trained face recognition model is determined as the feature extraction model.
[0102] The output layer of the parameter-locked feature extraction model is connected to the input layer of the preset relationship model to obtain an initial identity authentication model.
[0103] The target training sample is input into the initial identity authentication model to obtain the predicted label data output by the initial identity authentication model.
[0104] The preset relationship model in the initial identity authentication model is parameter-optimized to minimize the difference between the predicted label data and the preset label data possessed by the target training sample as the target, to obtain the risk identification model; wherein the preset label data is the first label data or the second label data.
[0105] In an embodiment of the present specification, the working principle of the identity authentication model can be simply considered as: processing the features extracted from the user face image and the certificate face image to obtain a classification result reflecting whether the user has the risk of using deep fake technology for identity authentication. Based on this, in order to guarantee the performance of the identity authentication model, a feature extraction model can be trained separately for extracting face image features, and a classification model trained in combination with the extracted feature extraction model is used as a risk identification model, thereby building a target identity authentication model using the feature extraction model and the risk identification model.
[0106] Specifically, the face recognition model can be trained by using the face images in the first training samples and the third training samples. At this time, the network layer at the rear position in the trained face recognition model often has the ability to extract face features with high accuracy. Based on this, the model from the input layer to the specified network layer in the trained face recognition model can be determined as the feature extraction model in the identity authentication model. The specified network layer can be preferentially selected as the fully connected layer close to the output layer of the face recognition model, and can also be other network layers, which are not limited in particular.
[0107] In the embodiments of the present specification, the preset relation model can be used to learn the similarity between the input data, and through parameter optimization, the similarity of data with the same label is larger, and the similarity of data with different labels is smaller.
[0108] Since the feature extraction model has completed parameter optimization and can extract face image feature data with high accuracy from the first training samples and the third training samples, the output layer of the parameter-locked feature extraction model can be connected to the input layer of the preset relation model to obtain an initial identity authentication model. In the model training process, after the parameter-locked feature extraction model extracts respective face image feature data from two face images in the target training samples, the face image feature data can be input to the preset relation model, so that the preset relation model can generate prediction label data reflecting whether the user has the risk of using deep fake technology for identity authentication according to the face image feature data. And according to the difference between the prediction label data and the preset label data possessed by the target training sample, the parameters of the preset relation model are optimized, and the parameter optimization of the feature extraction model is no longer performed. The parameter-optimized preset relation model obtained based on this can be used as a risk identification model in the identity authentication model.
[0109] In actual application, the preset relation model can include various network layers such as fully connected layers, convolutional layers, pooling layers, activation layers, and loss layers, or can include a neural network block (Block) with a more complex structure encapsulated as a composite unit by a group of consecutive layers, such as a residual block structure (Residual Block) and an attention module (Attention Module) based on an attention mechanism, which are not limited in particular.
[0110] In the embodiments of the present specification, in order to further improve the accuracy of the face image feature data extracted by the feature extraction model, a feature extraction model specially used for extracting the feature data of the user face image and a feature extraction model specially used for extracting the feature data of the certificate face image can be trained.
[0111] According to the feature extraction model, the first feature extraction model can be used for feature extraction of the user face image, and the second feature extraction model can be used for feature extraction of the certificate face image.
[0112] According to the feature extraction model, the first feature extraction model can be used for feature extraction of the user face image, and the second feature extraction model can be used for feature extraction of the certificate face image.
[0113] According to the feature extraction model, the first feature extraction model can be used for feature extraction of the user face image, and the second feature extraction model can be used for feature extraction of the certificate face image.
[0114] According to the feature extraction model, the first feature extraction model can be used for feature extraction of the user face image, and the second feature extraction model can be used for feature extraction of the certificate face image.
[0115] According to the feature extraction model, the first feature extraction model can be used for feature extraction of the user face image, and the second feature extraction model can be used for feature extraction of the certificate face image.
[0116] According to the feature extraction model, the first feature extraction model can be used for feature extraction of the user face image, and the second feature extraction model can be used for feature extraction of the certificate face image.
[0117] According to the feature extraction model, the first feature extraction model can be used for feature extraction of the user face image, and the second feature extraction model can be used for feature extraction of the certificate face image.
[0118] According to the feature extraction model, the first feature extraction model can be used for feature extraction of the user face image, and the second feature extraction model can be used for feature extraction of the certificate face image.
[0119] In actual applications, the face recognition model can be a pre-trained face recognition model; or the face recognition model can include a pre-trained face recognition model after parameter locking and a light-weight multilayer perceptron, wherein an output layer of the pre-trained face recognition model is connected to an input layer of the light-weight multilayer perceptron (MLP, Multilayer Perceptron). Thus, the pre-trained face recognition model itself has good ability to extract face feature data, so as to ensure the reliability and effectiveness of the trained feature extraction model.
[0120] For ease of understanding, FIG. 3 is a structural schematic diagram of a target identity authentication model provided by an embodiment of the present specification. As shown in FIG. 3, the target identity authentication model can include a first feature extraction model 301 and a second feature extraction model 302 composed of a pre-trained face recognition model and a light-weight multilayer perceptron, and a risk identification model 303. The output layer of the pre-trained face recognition model in the first feature extraction model 301 and the second feature extraction model 302 can be connected to the input layer of the light-weight multilayer perceptron included therein, and the output layer of the light-weight multilayer perceptron in the first feature extraction model 301 and the second feature extraction model 302 can be connected to the risk identification model 303, which will not be described herein.
[0121] Based on the same idea as the scheme shown in FIG. 2, the present specification also provides an identity authentication method. FIG. 4 is a flowchart of an identity authentication method provided by an embodiment of the present specification. The execution subject of the flowchart can be a device for identity authentication, or an application program carried by the device for identity authentication. As shown in FIG. 4, the flowchart can include:
[0122] Step 402: obtaining a user face image of a target user collected in an identity authentication process.
[0123] In the embodiment of the present specification, when identity authentication is performed on a target user, a user face image of the target user is usually collected in real time, so as to compare the user face image with a certificate face image in a certificate of the target user, thereby generating an identity authentication result for the target user.
[0124] In actual application, the user face image of the target user can be collected by using the terminal device of the target user or the device provided by the service provider. If the target identity authentication model for identity authentication is mounted on the device of the target user or the device provided by the service provider, the execution subject of the scheme in FIG. 4 can be the device of the target user or the device provided by the service provider, and step 402 can be collecting the user face image of the target user in real time by using the device of the target user or the device provided by the service provider. If the target identity authentication model is deployed on the server of the service provider, the execution subject of the scheme in FIG. 4 can be the server of the service provider, and step 402 can be that the server of the service provider acquires the user face image of the target user collected by the device of the target user or the device provided by the service provider.
[0125] Step 404: acquiring a certificate face image in the certificate image of the target user.
[0126] In the embodiments of the present specification, the device or the server of the service provider can store the certificate image of the target user acquired in advance, or the certificate image of the target user can be collected and reported in real time by the target user, and then step 404 can be that a face recognition model is used to recognize a face region from the certificate image of the target user, and a certificate face image is obtained by extracting an image in the face region.
[0127] Step 406: processing the user face image and the certificate face image by using the target identity authentication model to obtain identity authentication result information representing whether the target user has a risk of identity authentication by using deep fake technology.
[0128] In the embodiments of the present specification, if the identity authentication result information generated by the target identity authentication model represents that the target user has a risk of identity authentication by using deep fake technology, the target user will usually fail in identity authentication and cannot pass the identity authentication. If the identity authentication result information generated by the target identity authentication model represents that the target user does not have a risk of identity authentication by using deep fake technology, the target user can pass the identity authentication, or the target user can be further authenticated by using other means to ensure the security of the identity authentication, which is not limited specifically.
[0129] The method in FIG. 4 can accurately identify the case that the target user uses the deep fake technology to commit identity fraud in the identity authentication process by using the target identity authentication model, so as to reduce the risk of abuse of the deep fake technology and ensure the security of identity authentication, because the first training sample used when training the target identity authentication model contains the fake face image generated by using the generative artificial intelligence technology and the image style transfer technology, and the feature information of the first training sample is consistent with the face image used by the criminal when committing identity fraud by using the deep fake technology.
[0130] Based on the method in FIG. 4, the present specification embodiment also provides some specific implementation schemes of the method, which are described below.
[0131] During the identity authentication process, a video of the target user can be collected, and the user face image of the target user can be extracted from the video. However, the user face poses in different video frames in the video can be different, but the face pose in the certificate face image of the target user is fixed. Therefore, the face image in the video frame that is consistent with the face pose in the certificate face image can be selected as the user face image of the target user, so as to reduce the interference factors and improve the accuracy of the identity authentication result.
[0132] Therefore, in step 406, the user face image of the target user collected in the identity authentication process is obtained, which can specifically include:
[0133] If multiple candidate user face images of the target user are collected in the identity authentication process, the first pose information of the face in each candidate user face image is determined by using the face attribute recognition algorithm.
[0134] The second pose information of the face in the certificate face image is determined.
[0135] The candidate user face image corresponding to the first pose information that is most consistent with the second pose information is determined as the user face image of the target user.
[0136] In actual application, if the user image of the target user collected in the identity authentication process contains a large proportion of background area, in order to reduce the interference of the background information, the face recognition model can also be used to perform face recognition processing on the collected user image, and the image in the recognized face region is taken as the user face image of the target user, which is not described herein.
[0137] FIG. 5 is a swim lane flowchart corresponding to the identity authentication scheme in FIGS. 2 and 4 provided by the present specification embodiment. As shown in FIG. 5, the identity authentication process can involve service providers and execution subjects such as target users.
[0138] In the model training stage, the service provider can obtain a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using generative artificial intelligence technology; the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample, and the first training sample has first label data; wherein the first label data can be used to identify that the first sample user has a risk of using deep fake technology for identity authentication.
[0139] The service provider can also obtain a third training sample containing a third user face image sample of a third sample user and a third certificate face image sample; wherein the third certificate face image sample and the third user face image sample are face images of different users; or the third certificate face image sample and the third user face image sample are face images belonging to the same user and having different face attribute information. And determine the second label data that the third training sample has; wherein the second label data can be used to identify that the third sample user does not have a risk of using deep fake technology for identity authentication.
[0140] Using the target training sample, the face recognition model is trained to obtain a trained face recognition model; wherein the target training sample includes at least one of the first training sample and the third training sample. The model from the input layer to the specified network layer in the trained face recognition model is determined as a feature extraction model. By connecting the output layer of the parameter-locked feature extraction model with the input layer of the preset relationship model, an initial identity authentication model is obtained, the target training sample is input into the initial identity authentication model, and the predicted label data output by the initial identity authentication model is obtained; then minimize the difference between the predicted label data and the preset label data possessed by the target training sample as the target, and optimize the parameters of the preset relationship model to obtain a target identity authentication model; wherein the preset label data is the first label data or the second label data.
[0141] In the identity authentication stage, the target user needs to cooperate to perform the identity authentication operation, so that the service provider can obtain the user face image of the target user collected in the identity authentication process, and obtain the certificate face image in the certificate image of the target user. By using the target identity authentication model to process the user face image and the certificate face image, an identity authentication result information indicating whether the target user has a risk of using deep fake technology for identity authentication is obtained.
[0142] Based on the same idea, the present specification also provides a device corresponding to the above method. FIG. 6 is a structural schematic diagram of an identity authentication model training device according to an embodiment of the present specification. As shown in FIG. 6, the device can include:
[0143] The first obtaining module 602 is configured to obtain first training samples containing first user face image samples of a first sample user and first certificate face image samples; wherein the first user face image samples are user face images generated by using generative artificial intelligence technology; and the first certificate face image samples are certificate face images obtained by performing style transfer processing on the first user face image samples.
[0144] The first determining module 604 is configured to determine first label data possessed by the first training samples; and the first label data is used to identify that the first sample user has a risk of identity authentication by using deep fake technology.
[0145] The first training module 606 is configured to train an identity authentication model by using the first training samples with the first label data, to obtain a target identity authentication model.
[0146] Based on the device of FIG. 6, the present specification also provides some specific embodiments of the device, which are described below.
[0147] Optionally, the first obtaining module can include:
[0148] The generating unit is configured to generate user face images of the first sample user according to preset prompt words by using a generative artificial intelligence model, to obtain the first user face image samples.
[0149] The style transfer processing unit is configured to perform style transfer processing on the first user face image samples by using a target image style transfer model pre-trained, to obtain the first certificate face image samples.
[0150] Optionally, the device in FIG. 6 can also include:
[0151] The second obtaining module is configured to obtain second training samples containing second user face image samples of a second sample user and second certificate face image samples; wherein the second user face image samples are real face images obtained by image collection on a face of the second sample user, the second certificate face image samples are certificate face images contained in a certificate sample of the second sample user that passes trustworthiness verification, and the certificate sample and the first certificate face image samples belong to the same type of certificate.
[0152] a second training module, configured to train the style transfer model by using the second training sample, to obtain a predicted certificate face image output by the style transfer model.
[0153] a parameter optimization module, configured to perform parameter optimization on the style transfer model, to obtain a target image style transfer model, by taking minimizing a difference between the predicted certificate face image and the second certificate face image sample as an objective.
[0154] Optionally, the second training sample can further include style description information for the second certificate face image sample.
[0155] Correspondingly, the parameter optimization module can be specifically configured to:
[0156] perform parameter optimization on the style transfer model by taking minimizing a difference between the predicted certificate face image and the second certificate face image sample and the style description information as an objective.
[0157] Optionally, the style transfer model can include at least one of a styleGAN model, a CAP-VSTNet model, an InstantStyle model, an AdaIN model, and an InstantID model.
[0158] Optionally, the apparatus in FIG. 6 can further include:
[0159] a third obtaining module, configured to obtain a third training sample including a third user face image sample of a third sample user and a third certificate face image sample; wherein the third certificate face image sample and the third user face image sample are face images of different users, or the third certificate face image sample and the third user face image sample are face images belonging to the same user and having different face attribute information.
[0160] a second determining module, configured to determine second label data possessed by the third training sample; the second label data is used to identify that the third sample user does not have a risk of identity authentication by using deep fake technology.
[0161] The first training module can be specifically configured to:
[0162] train an identity authentication model by using the first training sample with the first label data and the third training sample with the second label data, to obtain a target identity authentication model.
[0163] Optionally, the identity authentication model can comprise a feature extraction model for feature extraction on the face image sample, and a risk identification model for identity authentication risk identification according to face image feature data extracted by the feature extraction model.
[0164] Correspondingly, the first training module can comprise:
[0165] A first training unit is configured to train the face recognition model by using the target training sample to obtain a trained face recognition model, wherein the target training sample comprises at least one of the first training sample and the third training sample.
[0166] A determination unit is configured to determine a model from an input layer to a specified network layer in the trained face recognition model as the feature extraction model.
[0167] A model connection unit is configured to connect an output layer of the parameter-locked feature extraction model and an input layer of the preset relationship model to obtain an initial identity authentication model.
[0168] A second training unit is configured to input the target training sample into the initial identity authentication model to obtain predicted label data output by the initial identity authentication model.
[0169] A parameter optimization unit is configured to minimize a difference between the predicted label data and preset label data possessed by the target training sample as an objective to perform parameter optimization on the preset relationship model in the initial identity authentication model to obtain the risk identification model, wherein the preset label data is the first label data or the second label data.
[0170] Optionally, the feature extraction model can comprise a first feature extraction model for feature extraction on a user face image and a second feature extraction model for feature extraction on a certificate face image; and the risk identification model is specifically configured to perform identity authentication risk identification according to face image feature data extracted by the first feature extraction model and the second feature extraction model.
[0171] Correspondingly, the first training unit can be specifically configured to:
[0172] The first face recognition model is trained by using at least one of the first user face image sample and the third user face image sample to obtain a first trained face recognition model.
[0173] The second face recognition model is trained by using at least one of the first certificate face image sample and the third certificate face image sample to obtain a second trained face recognition model.
[0174] The determining unit can be specifically configured to:
[0175] The model constituted by the input layer to the target network layer in the first trained face recognition model is determined as the first feature extraction model.
[0176] The model constituted by the input layer to the preset network layer in the second trained face recognition model is determined as the second feature extraction model.
[0177] Optionally, the face recognition model can be a pre-trained face recognition model; or,
[0178] The face recognition model can include a lightweight multilayer perceptron and a pre-trained face recognition model after parameter locking, wherein an output layer of the pre-trained face recognition model is connected to an input layer of the lightweight multilayer perceptron.
[0179] Based on the same idea, the present specification embodiment also provides a device corresponding to the above method. FIG. 7 is a structural schematic diagram of an identity authentication device provided by the present specification embodiment corresponding to FIG. 4. As shown in FIG. 7, the device can include:
[0180] The first acquisition module 702 is configured to acquire a user face image of a target user collected in an identity authentication process.
[0181] The second acquisition module 704 is configured to acquire a certificate face image in a certificate image of the target user.
[0182] The identity authentication module 706 is configured to process the user face image and the certificate face image by using a target identity authentication model to obtain identity authentication result information representing whether the target user has a risk of identity authentication by using a deep fake technology; wherein the target identity authentication model is a model trained by using the identity authentication model training device in the present specification embodiment.
[0183] Based on the device of FIG. 7, the present specification embodiment also provides some specific implementation schemes of the device, which are described below.
[0184] Optionally, the first acquisition module can include:
[0185] The first determination unit is configured to, if multiple candidate user face images of a target user are collected in an identity authentication process, determine first pose information of a face in each of the candidate user face images by using a face attribute recognition algorithm.
[0186] The second determination unit is configured to determine second pose information of a face in the certificate face image.
[0187] The third determination unit is configured to determine the candidate user face image corresponding to the first attitude information with the highest consistency with the second attitude information as the user face image of the target user.
[0188] Based on the same idea, the present specification also provides a device corresponding to the above method.
[0189] FIG. 8 is a structural schematic diagram of an identity authentication model training device corresponding to FIG. 2 provided by an embodiment of the present specification. As shown in FIG. 8, the device 800 can include:
[0190] at least one processor 810; and
[0191] a memory 830 in communication connection with the at least one processor; wherein
[0192] The memory 830 stores instructions 820 executable by the at least one processor 810, and the instructions are executed by the at least one processor 810 to enable the at least one processor 810 to:
[0193] obtain a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using generative artificial intelligence technology; and the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample.
[0194] determine first label data possessed by the first training sample; the first label data is used to identify that the first sample user has a risk of identity authentication by using deep fake technology.
[0195] train an identity authentication model by using the first training sample with the first label data, to obtain a target identity authentication model.
[0196] Based on the same idea, the present specification also provides a device corresponding to the above method.
[0197] FIG. 9 is a structural schematic diagram of an identity authentication device corresponding to FIG. 4 provided by an embodiment of the present specification. As shown in FIG. 9, the device 900 can include:
[0198] at least one processor 910; and
[0199] a memory 930 in communication connection with the at least one processor; wherein
[0200] The memory 930 stores instructions 920 executable by the at least one processor 910, and the instructions are executed by the at least one processor 910 to enable the at least one processor 910 to:
[0201] Obtain a user face image of a target user collected in an identity authentication process.
[0202] Obtain a certificate face image in a certificate image of the target user.
[0203] Process the user face image and the certificate face image by using a target identity authentication model to obtain identity authentication result information representing whether the target user has a risk of using deep fake technology to perform identity authentication, wherein the target identity authentication model is a model trained by using an identity authentication model training method in the embodiments of the present specification.
[0204] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments mainly describes the difference from other embodiments. In particular, the device shown in FIGS. 8 and 9 is basically similar to the method embodiments, and therefore the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0205] In the 1990s, it was quite obvious to distinguish whether an improvement in a technology was in hardware (e.g., improvement in circuit structures of diodes, transistors, switches, etc.) or in software (improvement in method flow). However, as technology has evolved, many improvements in method flow today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structures by programming the improved method flow into hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented by hardware entity modules. For example, a programmable logic device (PLD) such as a field programmable gate array (FPGA) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it, rather than by asking a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented by "logic compiler" software, which is similar to software compilers used in program development, and the original code to be compiled is written in a specific programming language, called a hardware description language (HDL), of which there are many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., the most commonly used being VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. It should be clear to those skilled in the art that, by simply logically programming a method flow in one of the above hardware description languages and programming it into an integrated circuit, a hardware circuit implementing the logical method flow can be easily obtained.
[0206] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to being implemented in pure computer readable program code, the controller can also be implemented to perform the same functions in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers, etc. by logically programming the method steps. Therefore, such a controller can be considered as a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can even be considered as both a software module implementing a method and a structure within a hardware component.
[0207] The systems, apparatuses, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0208] For the sake of description, the above apparatuses are described in functional division and are described respectively. Of course, the functions of the units can be implemented in the same or multiple software and / or hardware in the implementation of the present application.
[0209] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0210] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0211] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks.
[0212] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0213] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0214] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, non-volatile memory, such as read-only memory (ROM), EPROM, and / or flash memory, etc. The memory is an example of computer readable media.
[0215] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0216] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusions, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0217] Those skilled in the art will appreciate that embodiments of the present application can be provided as a method, system or computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer usable program code.
[0218] The present application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communication network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.
[0219] The above merely provides an example of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the scope of claims of the present application.
Claims
1. A method for training an identity authentication model, comprising: obtaining a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using a generative artificial intelligence technology; and the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample; determining first label data possessed by the first training sample; the first label data is used for identifying that the first sample user has a risk of identity authentication by using a deep fake technology; training an identity authentication model by using the first training sample with the first label data to obtain a target identity authentication model.
2. The method of claim 1, wherein the first training sample containing the first user face image sample of the sample user and the first certificate face image sample specifically comprises: generating a user face image of the first sample user according to a preset prompt word by using a generative artificial intelligence model to obtain the first user face image sample; performing style transfer processing on the first user face image sample by using a pre-trained target image style transfer model to obtain the first certificate face image sample.
3. The method of claim 2, wherein before performing the style transfer processing on the first user face image sample by using the pre-trained image style transfer model, the method further comprises: obtaining a second training sample containing a second user face image sample of a second sample user and a second certificate face image sample; wherein the second user face image sample is a real face image obtained by image acquisition on a face of the second sample user, the second certificate face image sample is a certificate face image contained in a certificate sample of the second sample user that passes a credibility verification, and the certificate sample and the first certificate face image sample belong to the same type of certificate; training the style transfer model by using the second training sample to obtain a predicted certificate face image output by the style transfer model; performing parameter optimization on the style transfer model to obtain a target image style transfer model, with the goal of minimizing the difference between the predicted certificate face image and the second certificate face image sample.
4. The method of claim 3, wherein the second training sample further contains style description information for the second certificate face image sample; the parameter optimization on the style transfer model is specifically performed with the goal of minimizing the difference between the predicted certificate face image and the second certificate face image sample and the style description information. at least one of a styleGAN model, a CAP-VSTNet model, an InstantStyle model, an AdaIN model, and an InstantID model.
5. The method of claim 3 or 4, the style transfer model comprising: 6. The method of claim 1, further comprising: obtaining a third training sample containing a third user face image sample of a third sample user and a third certificate face image sample; wherein the third certificate face image sample and the third user face image sample are face images of different users, or the third certificate face image sample and the third user face image sample are face images belonging to the same user and having different face attribute information; determining second label data possessed by the third training sample; the second label data is used to identify that the third sample user does not have a risk of identity authentication by using a deep fake technology; the training of the identity authentication model to obtain a target identity authentication model, specifically comprising: training an identity authentication model by using the first training sample having the first label data and the third training sample having the second label data to obtain a target identity authentication model.
7. The method of claim 6, wherein the identity authentication model comprises a feature extraction model for feature extraction on a face image sample, and a risk identification model for identity authentication risk identification according to face image feature data extracted by the feature extraction model; the training of the identity authentication model to obtain a target identity authentication model, specifically comprising: training a face recognition model by using a target training sample to obtain a trained face recognition model; wherein the target training sample includes at least one of the first training sample and the third training sample; determining a model from an input layer to a specified network layer in the trained face recognition model as the feature extraction model; connecting an output layer of the parameter-locked feature extraction model and an input layer of a preset relationship model to obtain an initial identity authentication model; inputting the target training sample into the initial identity authentication model to obtain predicted label data output by the initial identity authentication model; performing parameter optimization on the preset relationship model in the initial identity authentication model to obtain the risk identification model, with the objective of minimizing the difference between the predicted label data and preset label data possessed by the target training sample; wherein the preset label data is the first label data or the second label data.
8. The method of claim 7, the feature extraction model comprising: a first feature extraction model for feature extraction on a user face image, and a second feature extraction model for feature extraction on a certificate face image; the risk identification model is specifically used for identity authentication risk identification according to face image feature data extracted by the first feature extraction model and the second feature extraction model; the training of the face recognition model by using the target training sample to obtain the trained face recognition model, specifically comprising: training a first face recognition model by using at least one of the first user face image sample and the third user face image sample to obtain a first trained face recognition model; The second face recognition model is trained by using at least one of the first certificate face image sample and the third certificate face image sample, to obtain a second trained face recognition model; The model from the input layer to the specified network layer in the trained face recognition model is determined as the feature extraction model, and specifically includes: The model from the input layer to the target network layer in the first trained face recognition model is determined as the first feature extraction model; The model from the input layer to the preset network layer in the second trained face recognition model is determined as the second feature extraction model.
9. The method of claim 7, wherein the face recognition model is a pre-trained face recognition model; or, The face recognition model comprises: a lightweight multi-layer perceptron and a pre-trained face recognition model after parameter locking, wherein an output layer of the pre-trained face recognition model is connected to an input layer of the lightweight multi-layer perceptron.
10. An identity authentication method, comprising: obtaining a user face image of a target user collected in an identity authentication process; obtaining a certificate face image in a certificate image of the target user; processing the user face image and the certificate face image by using a target identity authentication model to obtain identity authentication result information representing whether the target user has a risk of using deep fake technology for identity authentication; wherein the target identity authentication model is a model trained by the identity authentication model training method of any one of claims 1-9.
11. The method of claim 10, wherein the obtaining a user face image of a target user collected in an identity authentication process specifically includes: if multiple candidate user face images of a target user are collected in an identity authentication process, determining first pose information of the face in each of the candidate user face images by using a face attribute recognition algorithm; determining second pose information of the face in the certificate face image; determining the candidate user face image corresponding to the first pose information that is most consistent with the second pose information as the user face image of the target user.
12. An identity authentication model training apparatus, comprising: a first obtaining module configured to obtain a first training sample containing a first sample user's first user face image sample and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using generative artificial intelligence technology; and the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample; a first determining module configured to determine first label data possessed by the first training sample; wherein the first label data is used to identify that the first sample user has a risk of using deep fake technology for identity authentication; a first training module configured to train an identity authentication model by using the first training sample with the first label data, to obtain a target identity authentication model.
13. The apparatus of claim 12, wherein the first obtaining module comprises: The generating unit is configured to generate a user face image of the first sample user by using a generative artificial intelligence model according to a preset prompt word, to obtain the first user face image sample; The style transfer processing unit is configured to perform style transfer processing on the first user face image sample by using a pre-trained target image style transfer model, to obtain the first certificate face image sample.
14. The apparatus of claim 13, further comprising: The second obtaining module is configured to obtain a second training sample containing a second user face image sample of a second sample user and a second certificate face image sample; wherein the second user face image sample is a real face image obtained by image collection on a face of the second sample user, and the second certificate face image sample is a certificate face image contained in a certificate sample of the second sample user that passes a credibility verification, and the certificate sample is of the same type as the first certificate face image sample; The second training module is configured to train a style transfer model by using the second training sample, to obtain a predicted certificate face image output by the style transfer model; The parameter optimization module is configured to perform parameter optimization on the style transfer model to minimize a difference between the predicted certificate face image and the second certificate face image sample, to obtain a target image style transfer model.
15. The apparatus of claim 14, wherein the second training sample further contains style description information for the second certificate face image sample; The parameter optimization module is specifically configured to: perform parameter optimization on the style transfer model to minimize a difference between the predicted certificate face image and the second certificate face image sample and the style description information.
16. The apparatus of claim 14 or 15, the style transfer model comprising: at least one of a styleGAN model, a CAP-VSTNet model, an InstantStyle model, an AdaIN model, and an InstantID model.
17. The apparatus of claim 12, further comprising: The third obtaining module is configured to obtain a third training sample containing a third user face image sample of a third sample user and a third certificate face image sample; wherein the third certificate face image sample and the third user face image sample are face images of different users, or the third certificate face image sample and the third user face image sample are face images of the same user but with different face attribute information; The second determining module is configured to determine second label data of the third training sample; the second label data is used to identify that the third sample user has no risk of identity authentication by using deep fake technology; The first training module is specifically configured to: train an identity authentication model by using the first training sample with the first label data and the third training sample with the second label data, to obtain a target identity authentication model.
18. The apparatus of claim 17, wherein the identity authentication model comprises a feature extraction model for feature extraction on face image samples, and a risk identification model for identity authentication risk identification according to face image feature data extracted by the feature extraction model; the first training module comprises: a first training unit configured to train the face recognition model by using target training samples to obtain a trained face recognition model, wherein the target training samples comprise at least one of the first training samples and the third training samples; a determination unit configured to determine a model from an input layer to a designated network layer in the trained face recognition model as the feature extraction model; a model connection unit configured to connect an output layer of the parameter-locked feature extraction model and an input layer of a preset relationship model to obtain an initial identity authentication model; a second training unit configured to input the target training samples into the initial identity authentication model to obtain predicted label data output by the initial identity authentication model; a parameter optimization unit configured to minimize a difference between the predicted label data and preset label data possessed by the target training samples as an objective to perform parameter optimization on the preset relationship model in the initial identity authentication model to obtain the risk identification model, wherein the preset label data is the first label data or the second label data.
19. The apparatus of claim 18, the feature extraction model comprising: a first feature extraction model for feature extraction on user face images, and a second feature extraction model for feature extraction on certificate face images; the risk identification model is specifically configured to perform identity authentication risk identification according to face image feature data extracted by the first feature extraction model and the second feature extraction model; the first training unit is specifically configured to: train a first face recognition model by using at least one of the first user face image samples and the third user face image samples to obtain a first trained face recognition model; train a second face recognition model by using at least one of the first certificate face image samples and the third certificate face image samples to obtain a second trained face recognition model; the determination unit is specifically configured to: determine a model from an input layer to a target network layer in the first trained face recognition model as the first feature extraction model; determine a model from an input layer to a preset network layer in the second trained face recognition model as the second feature extraction model.
20. The apparatus of claim 18, wherein the face recognition model is a pre-trained face recognition model; or The face recognition model comprises: a lightweight multilayer perceptron and a parameter-locked pre-trained face recognition model, wherein an output layer of the pre-trained face recognition model is connected with an input layer of the lightweight multilayer perceptron.
21. An identity authentication apparatus, comprising: a first acquisition module configured to acquire a user face image of a target user collected in an identity authentication process; a second acquisition module configured to acquire a certificate face image in a certificate image of the target user; An identity authentication module is configured to process the user face image and the certificate face image by using a target identity authentication model to obtain identity authentication result information indicating whether the target user has a risk of identity authentication by using deep fake technology.
22. The apparatus of claim 21, wherein the first obtaining module comprises: a first determining unit configured to, if multiple candidate user face images of a target user are collected in an identity authentication process, determine first pose information of a face in each of the candidate user face images by using a face attribute recognition algorithm; a second determining unit configured to determine second pose information of a face in the certificate face image; a third determining unit configured to determine, as the user face image of the target user, the candidate user face image corresponding to the first pose information that is most consistent with the second pose information.
23. An identity authentication model training device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: obtain a first training sample containing a first user face image sample of a first sample user and a first certificate face image sample; wherein the first user face image sample is a user face image generated by using generative artificial intelligence technology; and the first certificate face image sample is a certificate face image obtained by performing style transfer processing on the first user face image sample; determine first label data possessed by the first training sample; the first label data is used to identify that the first sample user has a risk of identity authentication by using deep fake technology; train an identity authentication model by using the first training sample with the first label data to obtain a target identity authentication model.
24. An identity authentication device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: obtain a user face image of a target user collected in an identity authentication process; obtain a certificate face image in a certificate image of the target user; process the user face image and the certificate face image by using a target identity authentication model to obtain identity authentication result information indicating whether the target user has a risk of identity authentication by using deep fake technology; wherein the target identity authentication model is a model trained by using the identity authentication model training method of any one of claims 1-9.
Citation Information
Patent Citations
Method and apparatus for identity verification
CN104935438A
Face recognition system and method
CN108470169A
Anti-counterfeit detection method and device, electronic device, and storage medium
CN109359502A
Bank account opening monitoring method, server and system
CN114519635A
Vision-based epidemic prevention control method and device, storage medium and electronic equipment
CN114897655A
Cited By
Method for determining risk in identity authentication process and computing equipment
CN121706074A