A method for training an identity authentication model, an identity authentication method, apparatus, and device
Through the generation and style transfer processing of fake face image samples, the identity authentication model is trained, which solves the problem of the existing technology being difficult to identify deep forgery technology identity fraud, and realizes the security and accuracy of identity authentication.
Patent Information
- Application Number
- CN202410775563.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-06-14
AI Technical Summary
The prior art is difficult to accurately identify whether users use deep forgery technology to commit identity fraud during the identity authentication process, resulting in the abuse of deep forgery technology and the security of identity authentication cannot be guaranteed.
By obtaining user face images generated using generative artificial intelligence technology and document face images obtained through style migration processing, a training sample with identification risks is generated to train an identity authentication model and identify whether the user uses deep forgery technology to commit identity fraud.
It realizes accurate identification of users' use of deep forgery technology to perform identity fraud during the identity authentication process, reduces the risk of deep forgery technology being abused, and ensures the security of identity authentication.
Smart Images

Figure CN118711231B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of identity verification technology, and in particular to an identity authentication model training method, identity authentication method, device and equipment. Background Art
[0002] Deepfake technology refers to artificial intelligence (AI) technology that uses neural network technology to learn from user image or video samples, thereby splicing the user's voice, facial expressions, and body movements into fake content. The most common deepfake method is artificial intelligence (AI) face-swapping technology, which makes it possible to tamper with or generate highly realistic and difficult-to-distinguish user audio and video content, often indistinguishable to the naked eye. In recent years, as the threshold for deepfake technology has been lowered, criminals have begun using deepfake technology to steal others' identities or forge personal identities for authentication purposes, posing new challenges to the normal operation of businesses.
[0003] Based on this, how to accurately identify users who use deep fake technology to commit identity fraud during the identity authentication process, so as to reduce the risk of deep fake technology being abused and ensure the security of identity authentication, has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The embodiments of this specification provide an identity authentication model training method, identity authentication method, device, and equipment that can accurately identify situations where users use deep fake technology to commit identity fraud during the identity authentication process, thereby reducing the risk of deep fake technology being abused and ensuring the security of identity authentication.
[0005] To solve the above technical problems, the embodiments of this specification are implemented as follows:
[0006] The embodiment of this specification provides an identity authentication model training method, including:
[0007] Obtain a first training sample comprising a first user facial image sample of a first sample user and a first ID facial image sample; wherein the first user facial image sample is a user facial image generated using generative artificial intelligence technology; and the first ID facial image sample is an ID facial image obtained by performing style transfer processing on the first user facial image sample;
[0008] Determining first label data for the first training sample; the first label data is used to identify that the first sample user has a risk of identity authentication using deep fake technology;
[0009] The identity authentication model is trained using the first training sample having the first label data to obtain a target identity authentication model.
[0010] An identity authentication method provided in an embodiment of this specification includes:
[0011] Obtaining the target user's facial image collected during the identity authentication process;
[0012] Obtaining an ID face image from the ID image of the target user;
[0013] The user face image and the ID face image are processed using a target identity authentication model to obtain identity authentication result information indicating whether the target user has the risk of identity authentication using deep fake technology; wherein, the target identity authentication model is a model trained using the identity authentication model training method in the embodiment of this specification.
[0014] An embodiment of this specification provides an identity authentication model training device, comprising:
[0015] A first acquisition module is configured to acquire a first training sample comprising a first user facial image sample of a first sample user and a first ID facial image sample; wherein the first user facial image sample is a user facial image generated using generative artificial intelligence technology; and the first ID facial image sample is an ID facial image obtained by performing style transfer processing on the first user facial image sample;
[0016] A first determination module is configured to determine first label data of the first training sample; the first label data is used to identify that the first sample user has a risk of identity authentication using deep fake technology;
[0017] The first training module is used to train the identity authentication model using the first training sample having the first label data to obtain a target identity authentication model.
[0018] An embodiment of this specification provides an identity authentication device, including:
[0019] The first acquisition module is used to obtain the user face image of the target user collected during the identity authentication process;
[0020] A second acquisition module is used to acquire an ID face image from the ID image of the target user;
[0021] An identity authentication module is used to process the user face image and the ID face image using a target identity authentication model to obtain identity authentication result information indicating whether the target user has the risk of identity authentication using deep fake technology; wherein, the target identity authentication model is a model trained using the identity authentication model training device in the embodiment of this specification.
[0022] An embodiment of this specification provides an identity authentication model training device, comprising:
[0023] at least one processor; and,
[0024] a memory communicatively connected to the at least one processor; wherein,
[0025] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:
[0026] Obtain a first training sample comprising a first user facial image sample of a first sample user and a first ID facial image sample; wherein the first user facial image sample is a user facial image generated using generative artificial intelligence technology; and the first ID facial image sample is an ID facial image obtained by performing style transfer processing on the first user facial image sample;
[0027] Determining first label data for the first training sample; the first label data is used to identify that the first sample user has a risk of identity authentication using deep fake technology;
[0028] The identity authentication model is trained using the first training sample having the first label data to obtain a target identity authentication model.
[0029] An identity authentication device provided in an embodiment of this specification includes:
[0030] at least one processor; and,
[0031] a memory communicatively connected to the at least one processor; wherein,
[0032] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:
[0033] Obtaining the target user's facial image collected during the identity authentication process;
[0034] Obtaining an ID face image from the ID image of the target user;
[0035] The user face image and the ID face image are processed using a target identity authentication model to obtain identity authentication result information indicating whether the target user has the risk of identity authentication using deep fake technology; wherein, the target identity authentication model is a model trained using the identity authentication model training method in the embodiment of this specification.
[0036] At least one embodiment provided in this specification can achieve the following beneficial effects:
[0037] A first user facial image of a first sample user generated using generative artificial intelligence technology is obtained in advance. A first ID facial image sample is obtained by performing style transfer processing on the first user facial image sample. Based on the forged first user facial image sample and the first ID facial image sample, a first training sample with first labeled data identifying the first sample user as a risk of identity authentication using deepfake technology is generated. The first training sample with the first labeled data is used to train an identity authentication model to obtain a target identity authentication model required for authenticating users with identity verification requirements. In this solution, because the first training sample used to train the target identity authentication model contains forged facial images generated using generative artificial intelligence and image style transfer technology, the first training sample has relatively consistent feature information with facial images used by criminals to commit identity fraud using deepfake technology. This allows the target identity authentication model to accurately identify users using deepfake technology to commit identity fraud during the identity authentication process, thereby reducing the risk of deepfake technology being abused and ensuring identity authentication security. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0039] Figure 1 This is a schematic diagram of an application scenario of an identity authentication solution provided in an embodiment of this specification;
[0040] Figure 2 A flowchart of an identity authentication model training method provided in an embodiment of this specification;
[0041] Figure 3 A schematic diagram of the structure of a target identity authentication model provided in the embodiments of this specification;
[0042] Figure 4 A flowchart of an identity authentication method provided in an embodiment of this specification;
[0043] Figure 5 The embodiments of this specification provide corresponding Figure 2 and Figure 4 A swimlane flow diagram of the identity authentication scheme in [1].
[0044] Figure 6 The embodiments of this specification provide corresponding Figure 2 A structural diagram of an identity authentication model training device;
[0045] Figure 7 The embodiments of this specification provide corresponding Figure 4 A structural diagram of an identity authentication device;
[0046] Figure 8 The embodiments of this specification provide corresponding Figure 2 A structural diagram of an identity authentication model training device;
[0047] Figure 9 The embodiments of this specification provide corresponding Figure 4 A structural diagram of an identity authentication device. DETAILED DESCRIPTION
[0048] To make the purpose, technical solutions, and advantages of one or more embodiments of this specification more clear, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of one or more embodiments of this specification.
[0049] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0050] In existing technologies, as the threshold for deepfake technology has been lowered, criminals have begun to use deepfake technology to forge user facial images, and use forged user facial images to create forged ID documents containing facial images. This makes the user facial images used by criminals in the identity authentication process and the ID facial images contained in the ID documents highly consistent in terms of facial posture, expression, makeup, etc. Since existing identity authentication models often forge user facial images and forged ID images with high consistency based on the facial attribute information provided by criminals, criminals can pass identity authentication by mistakenly judging that they are using deepfake technology to commit identity fraud during the identity authentication process. This makes it impossible to accurately identify users using deepfake technology to commit identity fraud during the identity authentication process, which easily leads to the risk of abuse of deepfake technology and the inability to guarantee the security of identity authentication during business operations.
[0051] In order to solve the defects in the prior art, this solution provides the following embodiments:
[0052] Figure 1 This is a schematic diagram of an application scenario of an identity authentication solution provided in an embodiment of this specification.
[0053] like Figure 1 As shown, the service provider can use device 101 to obtain a first training sample including a first user facial image sample and a first ID facial image sample of a first sample user; wherein the first user facial image sample can be a user facial image generated using generative artificial intelligence technology; and the first ID facial image sample can be an ID facial image obtained by performing style transfer processing on the first user facial image sample. First label data is set for the first training sample to identify the first sample user as a risk of identity authentication using deepfake technology; and the identity authentication model is trained using the first training sample with the first label data to obtain a target identity authentication model.
[0054] During the service operation process, when the target user needs to perform identity authentication, the terminal device 102 can communicate with the service provider's device 101, so that the service provider's device 101 can obtain the user's facial image collected in real time during the identity authentication process from the terminal device 102. The service provider's device 101 can also obtain the ID facial image of the target user from the ID image; by using the target identity authentication model, the user facial image and the ID facial image are processed to obtain identity authentication result information indicating whether the target user is at risk of identity authentication using deep fake technology.
[0055] Next, an identity authentication model training method and an identity authentication method provided in the embodiments of the specification will be specifically described with reference to the accompanying drawings:
[0056] Figure 2 This is a flow chart of an identity authentication model training method provided in an embodiment of this specification. From a program perspective, the execution subject of this process can be a device used for model training, or an application installed on the device. Figure 2 As shown, the process may include the following steps:
[0057] Step 202: Obtain a first training sample including a first user face image sample of a first sample user and a first ID face image sample; wherein, the first user face image sample is a user face image generated using generative artificial intelligence technology; the first ID face image sample is an ID face image obtained by performing style transfer processing on the first user face image sample.
[0058] In the embodiments of this specification, when authenticating a user, it is usually necessary to collect a facial image of the user in real time to obtain the user's facial image, and then compare the consistency between the user's facial image and the ID facial image carried in the ID. If the consistency is high, it can be considered that the user to whom the ID belongs is performing the identity authentication, and the user can pass the identity authentication. Based on this, when criminals use deep fake technology to commit identity fraud, they usually forge a user's facial image and directly process the forged user's facial image as the ID facial image in the forged ID. Since the forged user's facial image is highly consistent with the facial posture, expression, hairstyle, etc. contained in the ID facial image, it is easy for criminals to successfully pass the identity authentication.
[0059] It can be seen that if it is necessary to accurately identify the situation where a user uses deep fake technology to commit identity fraud during the identity authentication process, generative artificial intelligence technology can be used to generate a first user face image, and the first user face image is a forged user face image. In addition, a first ID face image is obtained by performing style transfer processing on the first user face image, and the first ID face image is a forged ID face image. At this time, the facial attribute information (for example, facial posture, expression, hairstyle) between the first user face image sample and the first ID face image sample has a high degree of consistency, which is consistent with the features and characteristics of the image pair of user face images and ID face images used by criminals to commit identity fraud. Therefore, a first training sample can be generated based on the first user face image sample and the first ID face image sample to train the identity authentication model.
[0060] Among them, generative AI (Generated Content) technology can refer to a new type of AI technology that generates new original content by learning from large-scale data sets. It can generate text, images, sounds, videos, code, and other content based on algorithms, models, and rules. Currently, there are many algorithms and models that can generate highly realistic facial images or videos. Based on this, generative AI technology can be used to generate the first user's facial image. This can not only reduce the difficulty of obtaining training samples, but also avoid the risk of infringing on user data rights and interests brought about by directly using real user data.
[0061] Style transfer refers to the fusion of the content of one image with the style of another, thereby generating an image with a novel appearance and a specific style. Style can refer to the texture, color, and visual patterns at different spatial scales within an image. Because some ID cards contain grayscale images of faces, and due to the effects of ID printing, ID facial images often possess their own specific style. Based on this, a pre-trained algorithm or model capable of generating ID facial styles can be used to process the first user's facial image sample to obtain a first ID facial image sample.
[0062] Step 204: Determine first label data of the first training sample; the first label data is used to identify that the first sample user has a risk of using deep fake technology for identity authentication.
[0063] In the embodiment of this specification, the first training sample has consistent features and characteristics with the image pairs of user face images and ID face images used by criminals to commit identity fraud. Therefore, the first training sample can be provided with first label data for identifying that the first sample user to which the first training sample belongs has the risk of identity authentication using deep fake technology.
[0064] Step 206: Train the identity authentication model using the first training sample having the first label data to obtain a target identity authentication model.
[0065] In the embodiment of this specification, the identity authentication model is trained using the first training sample with the first label data, so that the trained target identity authentication model can learn the characteristics and knowledge of the image pairs of user face images and ID face images used by criminals to commit identity fraud, thereby being able to accurately identify situations where users use deep fake technology for identity authentication.
[0066] Figure 2In the method, since the first training sample used in training the target identity authentication model includes forged facial images generated by generative artificial intelligence technology and image style transfer technology, the first training sample is relatively consistent with the feature information of the facial images used by criminals to commit identity fraud using deep fake technology. Therefore, the target identity authentication model can be used to accurately identify the situation where users use deep fake technology to commit identity fraud during the identity authentication process, thereby reducing the risk of deep fake technology being abused and ensuring the security of identity authentication.
[0067] based on Figure 2 The method in this specification also provides some specific implementation plans of the method, which are described below.
[0068] In the embodiment of this specification, step 202: obtaining a first training sample including a first user face image sample and a first ID face image sample of a sample user may specifically include:
[0069] A generative artificial intelligence model is used to generate a user face image of the first sample user according to a preset prompt word to obtain the first user face image sample.
[0070] Using the pre-trained target image style transfer model, style transfer processing is performed on the first user face image sample to obtain the first ID face image sample.
[0071] In the embodiments of this specification, an existing generative artificial intelligence model can be used to generate a facial image as a first user facial image sample. The generative artificial intelligence model may include but is not limited to the StableDiffusion AI painting generation tool, the Midjourney online image generation platform, the StyleGAN image generation model, etc., without specific limitation. In actual applications, since the facial features of users in different geographical environments often differ, it is possible to set preset prompt words according to actual needs to generate a first user facial image sample with specified features and characteristics. For example, the preset prompt words can be used to indicate the user's clothing, hairstyle, facial features, lighting conditions, background environment and other information, without specific limitation.
[0072] In the embodiments of this specification, in order to enable the target image style transfer model to generate an image with the style of the face on the ID card, it is usually necessary to pre-train the target image style transfer model. Based on this, before performing style transfer processing on the first user face image sample using the pre-trained image style transfer model, the following steps may also be included:
[0073] Obtain a second training sample including a second user facial image sample of a second sample user and a second ID facial image sample; wherein, the second user facial image sample is a real facial image obtained by image acquisition of the face of the second sample user, and the second ID facial image sample is an ID facial image contained in an ID sample of the second sample user that has passed credibility verification, and the ID sample and the first ID facial image sample belong to the same type of ID.
[0074] The style transfer model is trained using the second training sample to obtain a predicted ID face image output by the style transfer model.
[0075] With the goal of minimizing the difference between the predicted ID face image and the second ID face image sample, the style transfer model is parameter optimized to obtain a target image style transfer model.
[0076] In the embodiments of the present specification, since the target image style transfer model needs to generate an ID face image with the style of the face on the ID card based on the input user face image, when training the target image style transfer model, user face image samples and ID face image samples that have good credibility and authenticity and belong to the same user can be obtained as second training samples.
[0077] Specifically, the second user facial image samples included in the second training sample can be obtained by capturing live facial images of a real user (i.e., the second sample user), and the second ID facial image samples included in the second training sample can be facial images included in the second sample user's authentic and trustworthy ID. It is understood that before using the second training sample for model training, the service provider typically obtains authorization from the second sample user to use it, thereby avoiding infringement of user rights.
[0078] In practical applications, the second user face image sample contained in the second training sample needs to be input into the style transfer model. At this time, the style transfer model will generate a predicted ID face image. Since the model training goal is to make the visual effect of the predicted ID face image close to the second ID face image sample in the second training sample, the style transfer model parameters can be optimized with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample to obtain a target image style transfer model that has the ability to generate face images with ID face style.
[0079] In the embodiments of this specification, since the multimodal model can combine and analyze rich visual information and information from other modalities (such as text), it can capture the relationship between multiple data, which helps to improve the accuracy and efficiency of the model and thus has been promoted and applied.
[0080] Based on this, if the target image style transfer model is a multimodal model, the second training sample may also include style description information for the second ID face image sample.
[0081] Correspondingly, optimizing the parameters of the style transfer model with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample may specifically include:
[0082] Parameters of the style transfer model are optimized with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample and the style description information.
[0083] In an embodiment of the present specification, the style description information for the second ID face image sample may include but is not limited to: the name and type information of the ID to which the second ID face image sample belongs, the name and type information of the face style to which the second ID face image sample belongs, the visual feature information and face attribute information of the face area, etc.
[0084] Since it is necessary to ensure that the ID face image generated by the target image style transfer model is consistent with the style description information for the second ID face image sample, the style transfer model can be optimized with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample and the style description information, thereby ensuring that the ID face image generated by the trained target image style transfer model has a high degree of authenticity and meets actual needs.
[0085] In practical applications, the style transfer model may include, but is not limited to, at least one of the following: a styleGAN model, a CAP-VSTNet model, an InstantStyle model, an AdaIN model, and an InstantID model. Since the InstantStyle model and the InstantID model are multimodal models, the second training samples for the InstantStyle model and the InstantID model may include, in addition to the second user face image sample and the second ID face image sample, style description information for the second ID face image sample. Furthermore, since the InstantID model has the ability to automatically extract text input information based on the input image, if only the InstantID model is needed to generate ID face images for a single type of ID, then the second training samples may not need to include style description information for the second ID face image sample. There is no specific limitation on this.
[0086] In the embodiments of this specification, since there is often a certain time interval between when a user's ID is generated and when the user authenticates their identity, a highly credible user face image and the ID face image of the same user often contain different facial attribute information, but have a relatively high degree of similarity in facial features. Conversely, the facial attribute information contained in the user face image and the ID face image of different users often differs significantly, and the similarity in facial features is also relatively low. In both cases, the user should be considered to be at no risk of identity authentication using deepfake technology.
[0087] Based on this, Figure 2 The method may further include:
[0088] Obtain a third training sample including a third user facial image sample and a third ID facial image sample of a third sample user; wherein the third ID facial image sample and the third user facial image sample are facial images of different users; or, the third ID facial image sample and the third user facial image sample are facial images belonging to the same user and having different facial attribute information.
[0089] Determine second label data of the third training sample; the second label data is used to identify that the third sample user does not have the risk of identity authentication using deep fake technology.
[0090] Correspondingly, step 206: training the identity authentication model to obtain a target identity authentication model may include:
[0091] The identity authentication model is trained using the first training sample with the first label data and the third training sample with the second label data to obtain a target identity authentication model.
[0092] In the embodiments of this specification, facial attribute information may include, but is not limited to, facial expression, facial posture, identity, makeup, and other information. In actual applications, when the same user uses a personal ID for identity authentication, the facial attribute information collected for the user's facial image and the ID facial image is often different. In this case, because the user does not use deepfake technology to commit identity fraud, a third ID facial image sample and a third user facial image sample belonging to the same user and having different facial attribute information can be used as third training samples, and a second label data is assigned to the third training sample to indicate that the third sample user to which the third training sample belongs does not pose a risk of identity authentication using deepfake technology.
[0093] In the scenario where identity authentication is performed using user facial images and ID facial images of different users, although the user has committed identity fraud, it does not constitute identity fraud using deep fake technology. Therefore, a third ID facial image sample and a third user facial image sample belonging to different users can be used as a third training sample, and a second label data is set for the third training sample to indicate that the third sample user to whom the third training sample belongs does not have the risk of identity authentication using deep fake technology.
[0094] By training the identity authentication model using the first training sample with the first label data and the third training sample with the second label data, the performance of the target identity authentication model obtained by training can be further improved, which will not be elaborated on.
[0095] In actual applications, the user face images collected in real time often contain a large background area, thus containing more interference information, while the ID face image in the ID card not only has a smaller background area, but also its background area is often a solid color area, thus containing less interference information. Based on this, in order to improve the performance of the model, the first user face image sample in the first training sample and the third user face image sample in the third training sample can be images within the face area identified from the collected user image using the face recognition model, and the first ID card face image sample in the first training sample and the third ID card face image sample in the third training sample can be images within the face area identified from the ID card image using the face recognition model, thereby reducing the interference caused by the background area.
[0096] In an embodiment of this specification, the identity authentication model may include a feature extraction model for extracting features from facial image samples, and a risk identification model for identifying identity authentication risks based on facial image feature data extracted by the feature extraction model. The output layer of the feature extraction model is connected to the input layer of the risk identification model to form a (target) identity authentication model.
[0097] Based on this, step 206: training the identity authentication model to obtain a target identity authentication model may specifically include:
[0098] The face recognition model is trained using target training samples to obtain a trained face recognition model; wherein the target training samples include at least one of the first training samples and the third training samples.
[0099] The model formed from the input layer to the specified network layer in the trained face recognition model is determined as the feature extraction model.
[0100] The output layer of the feature extraction model after parameter locking is connected to the input layer of the preset relationship model to obtain an initial identity authentication model.
[0101] The target training sample is input into the initial identity authentication model to obtain predicted label data output by the initial identity authentication model.
[0102] With the goal of minimizing the difference between the predicted label data and the preset label data of the target training sample, the preset relationship model in the initial identity authentication model is optimized to obtain the risk identification model; wherein the preset label data is the first label data or the second label data.
[0103] In the embodiments of this specification, the working principle of the identity authentication model can be simply considered as: based on the features extracted from the user's facial image and the ID facial image, it is processed to obtain a classification result reflecting the risk of the user's identity authentication using deepfake technology. Based on this, to ensure the performance of the identity authentication model, a feature extraction model can be trained separately to extract facial image features. This trained feature extraction model can be combined with the classification model to train a risk identification model, thereby constructing the target identity authentication model using the feature extraction model and the risk identification model.
[0104] Specifically, the face recognition model can be trained using the facial images in the first and third training samples. In this case, the network layers located later in the trained face recognition model often have the ability to extract facial features with greater accuracy. Based on this, the model from the input layer to the designated network layer in the trained face recognition model can be determined as the feature extraction model in the identity authentication model. The designated network layer can preferably be a fully connected layer close to the output layer of the face recognition model, but may also be other network layers, without specific limitation.
[0105] In the embodiments of this specification, a preset relation model (Relation Model) can be used to learn the similarity between input data, and through parameter optimization, the similarity of data with the same label is made greater, while the similarity of data with different labels is made smaller.
[0106] Since the feature extraction model has completed parameter optimization and is able to extract accurate facial image feature data from the first and third training samples, the output layer of the parameter-locked feature extraction model can be connected to the input layer of the preset relationship model to obtain an initial identity authentication model. During the model training process, after the parameter-locked feature extraction model extracts the corresponding facial image feature data from the two facial images in the target training samples, both can be input into the preset relationship model. This allows the preset relationship model to generate predicted label data based on the facial image feature data, reflecting the risk of a user being authenticated using deepfake technology. Based on the difference between the predicted label data and the preset label data of the target training samples, the parameters of the preset relationship model are optimized, rather than the parameters of the feature extraction model. The resulting parameter-optimized preset relationship model can then be used as the risk identification model in the identity authentication model.
[0107] In practical applications, the preset relationship model may include various network layers such as fully connected layers, convolutional layers, pooling layers, activation layers, loss layers, etc., or may also include a more complex neural network block (Block) encapsulated by a group of continuous layers as a composite unit, such as a residual block structure (Residual Block), an attention module (Attention Module) based on an attention mechanism, etc., without specific limitation.
[0108] In the embodiments of this specification, in order to further improve the accuracy of facial image feature data extracted by the feature extraction model, a feature extraction model specifically for extracting feature data of user facial images and a feature extraction model specifically for extracting feature data of ID facial images can be trained.
[0109] Based on this, the feature extraction model may include: a first feature extraction model for extracting features from user facial images, and a second feature extraction model for extracting features from ID facial images; the risk identification model is specifically used to identify identity authentication risks based on the facial image feature data extracted by the first feature extraction model and the second feature extraction model.
[0110] Correspondingly, the method of training the face recognition model using the target training sample to obtain the trained face recognition model may specifically include:
[0111] A first face recognition model is trained using at least one of the first user face image sample and the third user face image sample to obtain a first trained face recognition model.
[0112] A second face recognition model is trained using at least one of the first ID face image sample and the third ID face image sample to obtain a second trained face recognition model.
[0113] Furthermore, determining the model formed from the input layer to the specified network layer in the trained face recognition model as the feature extraction model may specifically include:
[0114] The model formed from the input layer to the target network layer in the first trained face recognition model is determined as the first feature extraction model.
[0115] The model formed from the input layer to the preset network layer in the second trained face recognition model is determined as the second feature extraction model.
[0116] Subsequently, when training the risk identification model, the output layer of the first feature extraction model and the output layer of the second feature extraction model, after the parameters are locked, can be connected to the input layer of the preset relationship model to obtain an initial identity authentication model. The predicted label data output by the initial identity authentication model is obtained by inputting the first user facial image sample into the first feature extraction model after the parameters are locked, and the first ID facial image sample into the second feature extraction model after the parameters are locked; alternatively, the predicted label data output by the initial identity authentication model is obtained by inputting the third user facial image sample into the first feature extraction model after the parameters are locked, and the third ID facial image sample into the second feature extraction model after the parameters are locked. This will not be described in detail.
[0117] In practical applications, the face recognition model can be a pre-trained face recognition model; alternatively, the face recognition model can include a lightweight multi-layer perceptron and a pre-trained face recognition model with locked parameters, wherein the output layer of the pre-trained face recognition model is connected to the input layer of the lightweight multi-layer perceptron (MLP). This leverages the pre-trained face recognition model's inherently superior ability to extract facial feature data, ensuring the reliability and effectiveness of the trained feature extraction model.
[0118] For ease of understanding, Figure 3 This is a schematic diagram of the structure of a target identity authentication model provided in the embodiment of this specification. Figure 3 As shown, the target identity authentication model may include a first feature extraction model 301 and a second feature extraction model 302 composed of a pre-trained face recognition model and a lightweight multi-layer perceptron, as well as a risk identification model 303. The output layers of the pre-trained face recognition models in the first feature extraction model 301 and the second feature extraction model 302 may be connected to the input layers of the lightweight multi-layer perceptrons they contain, and the output layers of the lightweight multi-layer perceptrons in the first feature extraction model 301 and the second feature extraction model 302 may be connected to the risk identification model 303, which will not be described in detail.
[0119] Based on Figure 2 With the same idea as the solution shown in , the embodiment of this specification also provides an identity authentication method. Figure 4 This is a flow chart of an identity authentication method provided in an embodiment of this specification. The execution subject of this process can be a device that performs identity authentication, or an application installed on the device that performs identity authentication. Figure 4 As shown, the process may include:
[0120] Step 402: Acquire the target user's facial image collected during the identity authentication process.
[0121] In the embodiments of this specification, when authenticating the target user, it is usually necessary to collect the target user's facial image in real time to compare the user's facial image with the ID facial image in the target user's ID card, thereby generating an identity authentication result for the target user.
[0122] In actual applications, the target user's terminal device or the device provided by the service provider can be used to collect the target user's face image. If the target user's terminal device or the device provided by the service provider is equipped with the target identity authentication model for identity authentication, then Figure 4The execution subject of the solution can be the terminal device of the target user or the device provided by the service provider. Step 402 can be to use the terminal device of the target user or the device provided by the service provider to collect the user's face image in real time. If the target identity authentication model is deployed on the service provider's server, then Figure 4 The execution subject of the solution may be the service end of the service provider. Specifically, step 402 may be that the service end of the service provider obtains the user face image of the target user collected from the terminal device of the target user or the device provided by the service provider.
[0123] Step 404: Obtain the ID face image in the ID image of the target user.
[0124] In the embodiments of this specification, the service provider's device or service end may store a pre-acquired ID image of the target user, or the target user may collect and report a personal ID image in real time. Step 404 may specifically utilize a face recognition model to identify a face area from the target user's ID image, extract the image within the face area, and obtain the ID face image.
[0125] Step 406: Use the target identity authentication model to process the user face image and the ID face image to obtain identity authentication result information indicating whether the target user has the risk of using deep fake technology for identity authentication; wherein, the target identity authentication model is a model trained using the identity authentication model training method provided in the embodiments of this specification.
[0126] In the embodiments of this specification, if the target authentication model generates authentication results indicating that the target user is at risk of being authenticated using deepfake technology, the target user's authentication will typically fail, and the authentication will be unsuccessful. However, if the target authentication model generates authentication results indicating that the target user is not at risk of being authenticated using deepfake technology, the target user's authentication may succeed, thereby passing the authentication. Alternatively, other means may be used to further authenticate the target user to ensure the security of the authentication, without specific limitations.
[0127] Figure 4In the method, since the first training sample used in training the target identity authentication model includes forged facial images generated by generative artificial intelligence technology and image style transfer technology, the first training sample is relatively consistent with the feature information of the facial images used by criminals to commit identity fraud using deep fake technology. Therefore, the target identity authentication model can be used to accurately identify the situation where the target user uses deep fake technology to commit identity fraud during the identity authentication process, thereby reducing the risk of deep fake technology being abused and ensuring the security of identity authentication.
[0128] based on Figure 4 The method in this specification also provides some specific implementation plans of the method, which are described below.
[0129] During the identity authentication process, a video of the target user is often captured and a facial image of the target user is extracted from it. However, the user's facial pose may vary across different video frames, while the facial pose of the target user's ID image is fixed. Therefore, the facial image in the video frame with the most consistent facial pose with the ID image can be preferentially selected as the target user's facial image, thereby reducing interference factors and improving the accuracy of the identity authentication result.
[0130] Based on this, step 406: obtaining the target user's facial image collected during the identity authentication process may specifically include:
[0131] If multiple candidate user face images of the target user are collected during the identity authentication process, the first posture information of the face in each of the candidate user face images is determined using a face attribute recognition algorithm.
[0132] Determine second posture information of the face in the ID facial image.
[0133] The candidate user face image corresponding to the first posture information having the highest degree of consistency with the second posture information is determined as the user face image of the target user.
[0134] In actual applications, if the user image of the target user collected during the identity authentication process contains a relatively large background area, in order to reduce the interference caused by background information, a face recognition model can also be used to perform face recognition processing on the collected user image, and the image within the recognized face area can be used as the user face image of the target user. This will not be elaborated on.
[0135] Figure 5 The embodiments of this specification provide corresponding Figure 2 and Figure 4 A swimlane flow diagram of the identity authentication scheme in . Figure 5 As shown, the identity authentication process may involve execution entities such as service providers and target users.
[0136] During the model training phase, the service provider may obtain a first training sample comprising a first user face image sample of a first sample user and a first ID face image sample; wherein, the first user face image sample is a user face image generated using generative artificial intelligence technology; the first ID face image sample is an ID face image obtained by performing style transfer processing on the first user face image sample, and determine the first label data of the first training sample; wherein, the first label data can be used to identify that the first sample user has the risk of identity authentication using deep fake technology.
[0137] The service provider may also obtain a third training sample comprising a third user facial image sample and a third ID facial image sample of a third sample user; wherein the third ID facial image sample and the third user facial image sample are facial images of different users; or, alternatively, the third ID facial image sample and the third user facial image sample are facial images of the same user but having different facial attribute information. The service provider may also determine second label data for the third training sample; wherein the second label data may be used to indicate that the third sample user does not pose a risk of identity authentication using deepfake technology.
[0138] The face recognition model is trained using target training samples to obtain a trained face recognition model; wherein the target training samples include at least one of the first training samples and the third training samples. The model formed from the input layer to the specified network layer in the trained face recognition model is determined as a feature extraction model. An initial identity authentication model is obtained by connecting the output layer of the feature extraction model after parameter locking with the input layer of the preset relationship model, and the target training samples are input into the initial identity authentication model to obtain predicted label data output by the initial identity authentication model; then, with the goal of minimizing the difference between the predicted label data and the preset label data of the target training samples, the preset relationship model is optimized to obtain a target identity authentication model; wherein the preset label data is the first label data or the second label data.
[0139] During the identity authentication stage, the target user needs to cooperate in performing the identity authentication operation so that the service provider can obtain the user facial image of the target user collected during the identity authentication process, as well as the ID facial image in the ID image of the target user. By using the target identity authentication model to process the user facial image and the ID facial image, the identity authentication result information indicating whether the target user has the risk of using deep fake technology for identity authentication is obtained.
[0140] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method. Figure 6 The embodiments of this specification provide corresponding Figure 2 A structural diagram of an identity authentication model training device. Figure 6 As shown, the device may include:
[0141] The first acquisition module 602 is used to obtain a first training sample including a first user face image sample of a first sample user and a first ID face image sample; wherein, the first user face image sample is a user face image generated using generative artificial intelligence technology; the first ID face image sample is an ID face image obtained by performing style transfer processing on the first user face image sample.
[0142] The first determination module 604 is used to determine first label data of the first training sample; the first label data is used to identify that the first sample user has a risk of using deep fake technology for identity authentication.
[0143] The first training module 606 is configured to train the identity authentication model using the first training sample having the first label data to obtain a target identity authentication model.
[0144] based on Figure 6 The present specification also provides some specific implementation plans of the device, which are described below.
[0145] Optionally, the first acquisition module may include:
[0146] A generation unit is used to generate a user face image of the first sample user according to a preset prompt word using a generative artificial intelligence model to obtain the first user face image sample.
[0147] The style transfer processing unit is used to use the pre-trained target image style transfer model to perform style transfer processing on the first user face image sample to obtain the first ID face image sample.
[0148] Optional, Figure 6 The device described in may further include:
[0149] The second acquisition module is used to obtain a second training sample including a second user facial image sample of a second sample user and a second ID facial image sample; wherein, the second user facial image sample is a real facial image obtained by image capture of the face of the second sample user, and the second ID facial image sample is an ID facial image contained in the ID sample of the second sample user that has passed the credibility verification, and the ID sample and the first ID facial image sample belong to the same type of ID.
[0150] The second training module is used to train the style transfer model using the second training sample to obtain the predicted ID face image output by the style transfer model.
[0151] A parameter optimization module is used to optimize the parameters of the style transfer model with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample to obtain a target image style transfer model.
[0152] Optionally, the second training sample may further include style description information for the second ID facial image sample.
[0153] Correspondingly, the parameter optimization module can be specifically used to:
[0154] Parameters of the style transfer model are optimized with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample and the style description information.
[0155] Optionally, the style transfer model may include: at least one of a styleGAN model, a CAP-VSTNet model, an InstantStyle model, an AdaIN model, and an InstantID model.
[0156] Optional, Figure 6 The device described in may further include:
[0157] The third acquisition module is used to obtain a third training sample including a third user facial image sample of a third sample user and a third ID facial image sample; wherein the third ID facial image sample and the third user facial image sample are facial images of different users; or, the third ID facial image sample and the third user facial image sample are facial images belonging to the same user and having different facial attribute information.
[0158] The second determination module is used to determine the second label data of the third training sample; the second label data is used to identify that the third sample user does not have the risk of using deep fake technology for identity authentication.
[0159] The first training module can be specifically used to:
[0160] The identity authentication model is trained using the first training sample with the first label data and the third training sample with the second label data to obtain a target identity authentication model.
[0161] Optionally, the identity authentication model may include a feature extraction model for extracting features from facial image samples, and a risk identification model for identifying identity authentication risks based on facial image feature data extracted by the feature extraction model.
[0162] Correspondingly, the first training module may include:
[0163] The first training unit is used to train the face recognition model using the target training sample to obtain a trained face recognition model; wherein the target training sample includes at least one of the first training sample and the third training sample.
[0164] A determination unit is used to determine the model formed from the input layer to the specified network layer in the trained face recognition model as the feature extraction model.
[0165] The model connection unit is used to connect the output layer of the feature extraction model after parameter locking with the input layer of the preset relationship model to obtain an initial identity authentication model.
[0166] The second training unit is used to input the target training sample into the initial identity authentication model to obtain the predicted label data output by the initial identity authentication model.
[0167] A parameter optimization unit is used to optimize the parameters of the preset relationship model in the initial identity authentication model with the goal of minimizing the difference between the predicted label data and the preset label data of the target training sample to obtain the risk identification model; wherein the preset label data is the first label data or the second label data.
[0168] Optionally, the feature extraction model may include: a first feature extraction model for extracting features from a user's facial image, and a second feature extraction model for extracting features from a document's facial image; the risk identification model is specifically used to identify identity authentication risks based on the facial image feature data extracted by the first feature extraction model and the second feature extraction model.
[0169] Correspondingly, the first training unit can be specifically used to:
[0170] A first face recognition model is trained using at least one of the first user face image sample and the third user face image sample to obtain a first trained face recognition model.
[0171] A second face recognition model is trained using at least one of the first ID face image sample and the third ID face image sample to obtain a second trained face recognition model.
[0172] The determining unit may be specifically configured to:
[0173] The model formed from the input layer to the target network layer in the first trained face recognition model is determined as the first feature extraction model.
[0174] The model formed from the input layer to the preset network layer in the second trained face recognition model is determined as the second feature extraction model.
[0175] Optionally, the face recognition model may be a pre-trained face recognition model; or,
[0176] The face recognition model may include: a lightweight multi-layer perceptron and a pre-trained face recognition model with locked parameters, wherein the output layer of the pre-trained face recognition model is connected to the input layer of the lightweight multi-layer perceptron.
[0177] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method. Figure 7 The embodiments of this specification provide corresponding Figure 4 A structural diagram of an identity authentication device. Figure 7 As shown, the device may include:
[0178] The first acquisition module 702 is used to acquire the user face image of the target user collected during the identity authentication process.
[0179] The second acquisition module 704 is configured to acquire the ID face image from the ID image of the target user.
[0180] The identity authentication module 706 is used to process the user face image and the ID face image using a target identity authentication model to obtain identity authentication result information indicating whether the target user has the risk of identity authentication using deep fake technology; wherein, the target identity authentication model is a model trained using the identity authentication model training device in the embodiment of this specification.
[0181] based on Figure 7 The present specification also provides some specific implementation plans of the device, which are described below.
[0182] Optionally, the first acquisition module may include:
[0183] The first determining unit is configured to determine first posture information of the face in each of the candidate user face images by using a face attribute recognition algorithm if multiple candidate user face images of the target user are collected during the identity authentication process.
[0184] The second determining unit is used to determine second posture information of the face in the ID face image.
[0185] The third determining unit is configured to determine the candidate user face image corresponding to the first posture information having the highest degree of consistency with the second posture information as the user face image of the target user.
[0186] Based on the same idea, the embodiments of this specification also provide devices corresponding to the above methods.
[0187] Figure 8 The embodiments of this specification provide corresponding Figure 2 A structural diagram of an identity authentication model training device. Figure 8 As shown, the device 800 may include:
[0188] at least one processor 810; and,
[0189] A memory 830 in communication with the at least one processor; wherein,
[0190] The memory 830 stores instructions 820 that can be executed by the at least one processor 810. The instructions are executed by the at least one processor 810 to enable the at least one processor 810 to:
[0191] Obtain a first training sample including a first user face image sample of a first sample user and a first ID face image sample; wherein, the first user face image sample is a user face image generated using generative artificial intelligence technology; the first ID face image sample is an ID face image obtained by performing style transfer processing on the first user face image sample.
[0192] Determine first label data of the first training sample; the first label data is used to identify that the first sample user has a risk of identity authentication using deep fake technology.
[0193] The identity authentication model is trained using the first training sample having the first label data to obtain a target identity authentication model.
[0194] Based on the same idea, the embodiments of this specification also provide devices corresponding to the above methods.
[0195] Figure 9 The embodiments of this specification provide corresponding Figure 4 A structural diagram of an identity authentication device. Figure 9 As shown, the device 900 may include:
[0196] at least one processor 910; and,
[0197] A memory 930 in communication with the at least one processor; wherein,
[0198] The memory 930 stores instructions 920 that can be executed by the at least one processor 910. The instructions are executed by the at least one processor 910 to enable the at least one processor 910 to:
[0199] Obtain the target user's face image collected during the identity authentication process.
[0200] Obtain the ID face image from the ID image of the target user.
[0201] The user face image and the ID face image are processed using a target identity authentication model to obtain identity authentication result information indicating whether the target user has the risk of identity authentication using deep fake technology; wherein, the target identity authentication model is a model trained using the identity authentication model training method in the embodiment of this specification.
[0202] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. Figure 8 and Figure 9 As for the device shown, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0203] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0204] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0205] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0206] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0207] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0208] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0209] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0210] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0211] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0212] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0213] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0214] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0215] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0216] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0217] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for training an identity authentication model, comprising: Acquire a first training sample including a first user face image sample of a first sample user and a first ID face image sample; wherein the first user face image sample is a user face image generated using generative artificial intelligence technology; and the first ID face image sample is an ID face image obtained by performing style transfer processing on the first user face image sample; Determining first label data of the first training sample; the first label data is used to identify that the first sample user has a risk of using deep fake technology for identity authentication; Acquire a third training sample including a third user face image sample and a third ID face image sample of a third sample user; wherein the third ID face image sample and the third user face image sample are face images of different users; or, the third ID face image sample and the third user face image sample are face images belonging to the same user and having different face attribute information; Determining second label data of the third training sample; the second label data is used to identify that the third sample user does not have the risk of identity authentication using deep fake technology; The identity authentication model is trained using the first training sample with the first label data and the third training sample with the second label data to obtain a target identity authentication model.
2. The method according to claim 1, wherein obtaining a first training sample including a first user face image sample of a sample user and a first ID face image sample specifically comprises: Generate a user face image of the first sample user according to a preset prompt word using a generative artificial intelligence model to obtain a sample of the first user face image; The first user face image sample is subjected to style transfer processing by using the pre-trained target image style transfer model to obtain the first document face image sample.
3. The method according to claim 2, before using the pre-trained image style transfer model to perform style transfer processing on the first user face image sample, further comprising: Acquire a second training sample including a second user face image sample and a second ID face image sample of a second sample user; wherein the second user face image sample is a real face image obtained by performing image acquisition on the face of the second sample user, and the second ID face image sample is an ID face image included in an ID sample of the second sample user that has passed credibility verification, and the ID sample and the first ID face image sample belong to the same type of ID; Using the second training sample to train the style transfer model to obtain a predicted ID face image output by the style transfer model; With the goal of minimizing the difference between the predicted ID face image and the second ID face image sample, the style transfer model is optimized to obtain a target image style transfer model.
4. The method according to claim 3, wherein the second training sample further comprises style description information for the second ID facial image sample; The step of optimizing the parameters of the style transfer model with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample specifically includes: The style transfer model is parameter optimized with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample and the style description information.
5. The method according to claim 3 or 4, wherein the style transfer model comprises: At least one of the styleGAN model, CAP-VSTNet model, InstantStyle model, AdaIN model and InstantID model.
6. The method according to claim 1, wherein the identity authentication model comprises a feature extraction model for extracting features from a facial image sample, and a risk identification model for identifying identity authentication risks based on facial image feature data extracted by the feature extraction model; The training of the identity authentication model to obtain the target identity authentication model specifically includes: Using the target training sample, training the face recognition model to obtain a trained face recognition model; wherein the target training sample includes at least one of the first training sample and the third training sample; Determine the model formed from the input layer to the specified network layer in the trained face recognition model as the feature extraction model; Connecting the output layer of the feature extraction model after parameter locking to the input layer of the preset relationship model to obtain an initial identity authentication model; Inputting the target training sample into the initial identity authentication model to obtain predicted label data output by the initial identity authentication model; With the goal of minimizing the difference between the predicted label data and the preset label data of the target training sample, the preset relationship model in the initial identity authentication model is optimized to obtain the risk identification model; wherein the preset label data is the first label data or the second label data.
7. The method of claim 6, wherein the feature extraction model comprises: A first feature extraction model for extracting features from a user's facial image, and a second feature extraction model for extracting features from a document's facial image; The risk identification model is specifically used to perform identity authentication risk identification based on the facial image feature data extracted by the first feature extraction model and the second feature extraction model; The method of training the face recognition model using the target training sample to obtain the trained face recognition model specifically includes: Using at least one of the first user face image sample and the third user face image sample to train a first face recognition model to obtain a first trained face recognition model; Using at least one of the first ID facial image sample and the third ID facial image sample to train a second facial recognition model to obtain a second trained facial recognition model; The step of determining the model formed from the input layer to the specified network layer in the trained face recognition model as the feature extraction model specifically includes: Determine the model formed from the input layer to the target network layer in the first trained face recognition model as the first feature extraction model; The model formed from the input layer to the preset network layer in the second trained face recognition model is determined as the second feature extraction model.
8. The method according to claim 6, wherein the face recognition model is a pre-trained face recognition model; or The face recognition model includes: A lightweight multi-layer perceptron and a pre-trained face recognition model with locked parameters, wherein the output layer of the pre-trained face recognition model is connected to the input layer of the lightweight multi-layer perceptron.
9. An identity authentication method, comprising: Obtaining a user face image of the target user collected during the identity authentication process; Obtaining an ID face image from an ID image of the target user; The user face image and the ID face image are processed using a target identity authentication model to obtain identity authentication result information indicating whether the target user has the risk of using deep fake technology for identity authentication; wherein the target identity authentication model is a model trained using the identity authentication model training method described in any one of claims 1-8.
10. The method according to claim 9, wherein obtaining the face image of the target user collected during the identity authentication process specifically comprises: If multiple candidate user face images of the target user are collected during the identity authentication process, the first posture information of the face in each of the candidate user face images is determined by using a face attribute recognition algorithm; Determining second posture information of the face in the ID facial image; The candidate user face image corresponding to the first posture information having the highest degree of consistency with the second posture information is determined as the user face image of the target user.
11. An identity authentication model training device, comprising: A first acquisition module is used to acquire a first training sample including a first user face image sample of a first sample user and a first ID face image sample; wherein the first user face image sample is a user face image generated using generative artificial intelligence technology; and the first ID face image sample is an ID face image obtained by performing style transfer processing on the first user face image sample; A first determination module is used to determine first label data of the first training sample; the first label data is used to identify that the first sample user has a risk of using deep fake technology for identity authentication; A third acquisition module is configured to acquire a third training sample including a third user face image sample and a third ID face image sample of a third sample user; wherein the third ID face image sample and the third user face image sample are face images of different users; or the third ID face image sample and the third user face image sample are face images belonging to the same user and having different face attribute information; A second determination module determines second label data of the third training sample; the second label data is used to identify that the third sample user does not have the risk of identity authentication using deep fake technology; The first training module is used to train the identity authentication model using the first training sample with the first label data and the third training sample with the second label data to obtain a target identity authentication model.
12. The device according to claim 11, wherein the first acquisition module comprises: A generating unit, configured to generate a user face image of the first sample user according to a preset prompt word by using a generative artificial intelligence model, and obtain a sample of the first user face image; The style transfer processing unit is used to use the pre-trained target image style transfer model to perform style transfer processing on the first user face image sample to obtain the first document face image sample.
13. The apparatus of claim 12, further comprising: A second acquisition module is used to acquire a second training sample including a second user face image sample and a second ID face image sample of a second sample user; wherein the second user face image sample is a real face image obtained by performing image acquisition on the face of the second sample user, and the second ID face image sample is an ID face image included in an ID sample of the second sample user that has passed the credibility verification, and the ID sample and the first ID face image sample belong to the same type of ID; A second training module, used to train the style transfer model using the second training sample to obtain a predicted ID face image output by the style transfer model; A parameter optimization module is used to optimize the parameters of the style transfer model with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample to obtain a target image style transfer model.
14. The device according to claim 13, wherein the second training sample further comprises style description information for the second ID facial image sample; The parameter optimization module is specifically used for: The style transfer model is parameter optimized with the goal of minimizing the difference between the predicted ID face image and the second ID face image sample and the style description information.
15. The apparatus according to claim 13 or 14, wherein the style transfer model comprises: At least one of the styleGAN model, CAP-VSTNet model, InstantStyle model, AdaIN model and InstantID model.
16. The apparatus according to claim 11, wherein the identity authentication model comprises a feature extraction model for extracting features from a facial image sample, and a risk identification model for identifying identity authentication risks based on facial image feature data extracted by the feature extraction model; The first training module includes: A first training unit, configured to train a face recognition model using a target training sample to obtain a trained face recognition model; wherein the target training sample includes at least one of the first training sample and the third training sample; A determination unit, used to determine the model formed from the input layer to the specified network layer in the trained face recognition model as the feature extraction model; A model connection unit, used to connect the output layer of the feature extraction model after parameter locking with the input layer of the preset relationship model to obtain an initial identity authentication model; A second training unit, used for inputting the target training sample into the initial identity authentication model to obtain predicted label data output by the initial identity authentication model; A parameter optimization unit is used to optimize the parameters of the preset relationship model in the initial identity authentication model with the goal of minimizing the difference between the predicted label data and the preset label data of the target training sample to obtain the risk identification model; wherein the preset label data is the first label data or the second label data.
17. The apparatus of claim 16, wherein the feature extraction model comprises: A first feature extraction model for extracting features from a user's facial image, and a second feature extraction model for extracting features from a document's facial image; The risk identification model is specifically used to perform identity authentication risk identification based on the facial image feature data extracted by the first feature extraction model and the second feature extraction model; The first training unit is specifically used for: Using at least one of the first user face image sample and the third user face image sample to train a first face recognition model to obtain a first trained face recognition model; Using at least one of the first ID facial image sample and the third ID facial image sample to train a second facial recognition model to obtain a second trained facial recognition model; The determining unit is specifically configured to: Determine the model formed from the input layer to the target network layer in the first trained face recognition model as the first feature extraction model; The model formed from the input layer to the preset network layer in the second trained face recognition model is determined as the second feature extraction model.
18. The apparatus according to claim 16, wherein the face recognition model is a pre-trained face recognition model; or The face recognition model includes: A lightweight multi-layer perceptron and a pre-trained face recognition model with locked parameters, wherein the output layer of the pre-trained face recognition model is connected to the input layer of the lightweight multi-layer perceptron.
19. An identity authentication device, comprising: The first acquisition module is used to acquire the user face image of the target user collected during the identity authentication process; A second acquisition module is used to acquire an ID face image in the ID image of the target user; An identity authentication module is used to process the user face image and the ID face image using a target identity authentication model to obtain identity authentication result information indicating whether the target user has the risk of using deep fake technology for identity authentication; wherein the target identity authentication model is a model trained using the identity authentication model training device described in any one of claims 1-8.
20. The device according to claim 19, wherein the first acquisition module comprises: A first determining unit, configured to determine first posture information of a face in each of the candidate user face images by using a face attribute recognition algorithm if multiple candidate user face images of the target user are collected during the identity authentication process; A second determining unit, used to determine second posture information of the face in the ID face image; The third determining unit is used to determine the candidate user face image corresponding to the first posture information that has the highest degree of consistency with the second posture information as the user face image of the target user.
21. An identity authentication model training device, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire a first training sample including a first user face image sample of a first sample user and a first ID face image sample; wherein the first user face image sample is a user face image generated using generative artificial intelligence technology; and the first ID face image sample is an ID face image obtained by performing style transfer processing on the first user face image sample; Determining first label data of the first training sample; the first label data is used to identify that the first sample user has a risk of using deep fake technology for identity authentication; Acquire a third training sample including a third user face image sample and a third ID face image sample of a third sample user; wherein the third ID face image sample and the third user face image sample are face images of different users; or, the third ID face image sample and the third user face image sample are face images belonging to the same user and having different face attribute information; Determining second label data of the third training sample; the second label data is used to identify that the third sample user does not have the risk of identity authentication using deep fake technology; The identity authentication model is trained using the first training sample with the first label data and the third training sample with the second label data to obtain a target identity authentication model.
22. An identity authentication device, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Obtaining a user face image of the target user collected during the identity authentication process; Obtaining an ID face image from an ID image of the target user; The user face image and the ID face image are processed using a target identity authentication model to obtain identity authentication result information indicating whether the target user has the risk of using deep fake technology for identity authentication; wherein the target identity authentication model is a model trained using the identity authentication model training method described in any one of claims 1-8.
Citation Information
Patent Citations
Anti-counterfeit detection method and device, electronic device, and storage medium
CN109359502A
Model training method, identity verification method and device
CN117373091A