Image generation model training method and device, electronic device, and storage medium
By training an image generation model using an adversarial generative network, perturbation carriers such as eyeshadow of the same user are fused into the unedited image, solving the problems of insufficient concealment and high training difficulty in existing technologies, and achieving faster and more accurate generation of perturbation images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-13
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies require complete knowledge of the identity recognition model to be attacked when generating perturbation images, which limits their application. The generated perturbation images lack concealment, and the model is difficult to train, resulting in low accuracy.
An image generation model is trained using an adversarial generative network. By fusing perturbation carriers such as eyeshadow of the same user into the unedited image, the training difficulty of the model is reduced, and the concealment and accuracy are improved. The generator and discriminator in the adversarial generative network learn from each other through game and generate accurate perturbation images.
It improves the training speed of image generation models and the accuracy of generating perturbed images, reduces the difficulty of model training, and enhances the concealment and naturalness of generated images.
Smart Images

Figure CN116152599B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a training method of an image generation model, an image recognition method and device, an electronic device and a computer readable storage medium. BACKGROUND
[0002] In the field of image processing, a user image can be recognized by using an identity recognition model to verify the identity of the user, for example, the user image can be input into a pre-trained identity recognition model to identify the identity of the user based on the identity recognition model.
[0003] At present, in order to avoid the user image being misused to bring security risks to the user privacy, it is urgent to study how to add a disturbance carrier in the user image to generate a disturbance image that can make the identity recognition model identify the user identity incorrectly. SUMMARY
[0004] The present disclosure provides a training method of an image generation model and device, an electronic device and a computer readable storage medium.
[0005] In a first aspect, the present disclosure provides a training method of an image generation model, which comprises:
[0006] obtaining a first sample image and a second sample image, wherein the first sample image contains a target region in an original plain image of a first user, and the second sample image contains a region in an original carrier image of a second user, which corresponds to the position of the target region and contains a disturbance carrier;
[0007] fusing the disturbance carrier in the second sample image into the first sample image to obtain an initial carrier migration image;
[0008] training an initial image discrimination model according to the first sample image, the initial carrier migration image and the initial image generation model to obtain a target image discrimination model, wherein the initial image discrimination model is used to discriminate whether a received image is generated by the initial image generation model;
[0009] training the initial image generation model according to the first sample image, the original plain image, the original carrier image and the target image discrimination model to obtain an image generation model, wherein the image generation model is used to generate a disturbance image corresponding to a third user, and the disturbance image is used to make a first preset identity recognition model identify the third user as the second user.
[0010] In a second aspect, the present disclosure provides an image recognition method, which comprises:
[0011] Obtain a third user's unedited image;
[0012] The unedited image is input into an image generation model to obtain a target perturbation image, wherein the image generation model is obtained according to the training method of the image generation model described in the first aspect;
[0013] The target perturbation image is input into the fourth preset identity recognition model for identity recognition processing to obtain the target recognition result, wherein the target recognition result indicates that the third user's identity is the second user, and the fourth preset identity recognition model is any model used for identity recognition.
[0014] Thirdly, this disclosure provides a training apparatus for an image generation model, the training apparatus comprising:
[0015] The acquisition unit is used to acquire a first sample image and a second sample image, wherein the first sample image contains a target region in the original bare-faced image of the first user, and the second sample image contains a region in the original carrier image of the second user that corresponds to the position of the target region and contains a disturbed carrier.
[0016] A fusion unit is used to fuse the perturbation carrier in the second sample image into the first sample image to obtain an initial carrier migration image;
[0017] The first training unit is used to train the initial image discrimination model based on the first sample image, the initial carrier migration image, and the initial image generation model to obtain the target image discrimination model, wherein the initial image discrimination model is used to determine whether the received image is generated by the initial image generation model;
[0018] The second training unit is used to train the initial image generation model based on the first sample image, the original bare-faced image, the original carrier image, and the target image discrimination model to obtain an image generation model. The image generation model is used to generate a perturbation image corresponding to the third user, and the perturbation image is used to enable the first preset identity recognition model to identify the third user as the second user.
[0019] Fourthly, this disclosure provides an image recognition device, which includes:
[0020] An image acquisition unit is used to acquire a bare-faced image of a third user.
[0021] An image generation unit is used to input the unedited image into an image generation model to obtain a target perturbation image, wherein the image generation model is obtained according to the training method of the image generation model described in the first aspect;
[0022] The identification unit is used to input the target perturbation image into a fourth preset identity recognition model for identity recognition processing to obtain a target recognition result, wherein the target recognition result indicates that the third user's identity is the second user, and the fourth preset identity recognition model is any model used for identity recognition.
[0023] Fifthly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the training method of the image generation model of the first aspect or the image recognition method of the second aspect described above.
[0024] In a sixth aspect, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the training method of the image generation model of the first aspect or the image recognition method of the second aspect.
[0025] The embodiments provided in this disclosure, after obtaining a first sample image of the target region in an original bare-faced image containing a first user, and a second sample image of the region in an original carrier image containing a second user that corresponds to the target region and contains a perturbed carrier; by first fusing the perturbed carrier in the second sample image into the first sample image, an initial carrier migration image corresponding to the first user is obtained; then, based on the initial carrier migration image, the first sample image, and the initial image generation model, the initial image discrimination model is trained to obtain a target image discrimination model for determining whether a received image is generated by the initial image generation model; then, by using the first sample image, the original bare-faced image, the original carrier image, and the target image discrimination model to train the initial image generation model, an image generation model for generating a perturbed image corresponding to a third user that enables the first preset identity recognition model to identify the second user as the second user can be obtained.
[0026] In the training method provided in this embodiment, since the image generation model is trained based on the first sample image and the initial carrier migration image corresponding to the same identity, namely the first user, during the training process, especially during the training of the target image discrimination model, only the processing of the perturbed carrier in the image can be focused on, while the processing of image elements other than the perturbed carrier in the sample image, such as eyes and eyebrows, can be neglected. This reduces the training difficulty of the target image discrimination model, improves the convergence speed and discrimination accuracy of the target image discrimination model, and further improves the training speed of the image generation model and the accuracy of the generated perturbed image by training the initial image generation model based on the target image discrimination model.
[0027] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0028] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0029] Figure 1 A flowchart illustrating the training method for the image generation model provided in this embodiment of the disclosure;
[0030] Figure 2 A schematic diagram of the framework for training a target image discrimination model provided in an embodiment of this disclosure;
[0031] Figure 3 A flowchart for training an image generation model provided in an embodiment of this disclosure;
[0032] Figure 4 A schematic diagram of a framework for training an image generation model provided in an embodiment of this disclosure;
[0033] Figure 5a This is a schematic flowchart for calculating the fourth loss value provided in an embodiment of the present disclosure;
[0034] Figure 5b This is a schematic diagram of a framework for calculating a fourth loss value provided in an embodiment of the present disclosure;
[0035] Figure 6 A flowchart of the image recognition method provided in the embodiments of this disclosure;
[0036] Figure 7 A schematic diagram illustrating an application scenario provided by an embodiment of this disclosure;
[0037] Figure 8 A block diagram of a training apparatus for an image generation model provided in an embodiment of this disclosure;
[0038] Figure 9 A block diagram of an image recognition device provided in an embodiment of this disclosure;
[0039] Figure 10 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0040] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0041] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0042] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0043] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0044] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0045] Perturbation images, also known as adversarial examples, are images containing perturbation vectors that can cause a model to misidentify a user. In related technologies, when generating perturbation images, a user's original image without perturbation vectors is often used as the original sample, while images of other users containing perturbation vectors are used as reference samples. Based on these original and reference samples, an image generation model is trained, and adversarial examples of the input images are generated based on this model. Furthermore, during the model training process, white-box attack algorithms are often used to train the target identity recognition model; that is, all information about the target identity recognition model, such as its structure and parameters, needs to be known beforehand to be utilized during image generation model training.
[0046] The methods for generating such perturbation images in related technologies are based on white-box attack algorithms, which require knowledge of all information about the identity recognition model to be attacked. Without this information, the application is limited, resulting in poor universality. Furthermore, the perturbation images generated by these methods often involve directly adding external perturbation carriers such as glasses or stickers to the user's image, leading to insufficient concealment. Additionally, these methods often train the model directly on sample images of different users, i.e., different identities. Since sample images of different identities differ not only in the perturbation carriers but also in other elements, this significantly increases the difficulty of model training and the accuracy of the generated results.
[0047] To address at least one technical problem in the related art, this disclosure provides a method for training an image generation model, please refer to... Figure 1 This is a flowchart of a training method for an image generation model provided in this embodiment. This method can be applied to electronic devices, which may be terminal devices or servers; no special limitation is made here.
[0048] like Figure 1 As shown, the training method for the image generation model provided in this embodiment may include the following steps S101-S104, which will be described in detail below.
[0049] Step S101: Obtain a first sample image and a second sample image, wherein the first sample image contains the target region in the original bare-faced image of the first user, and the second sample image contains the region in the original carrier image of the second user that corresponds to the target region and contains the disturbed carrier.
[0050] In related technologies, the perturbation carrier used to generate perturbation images is often an object detached from the human body, such as a hat or glasses. This results in insufficient concealment of the generated perturbation image. To improve the concealment of the perturbation image, in this embodiment, the perturbation carrier refers to an object that acts directly on the human skin without requiring additional clothing. The perturbation carrier can be, for example, at least one of eyeshadow, lipstick, or facial makeup. In the following description, unless otherwise specified, eyeshadow is used as an example of a perturbation carrier.
[0051] In this embodiment of the disclosure, when the perturbation carrier is eyeshadow, the target area can be the eye area in the original bare face image of the first user; the first sample image can be an image obtained by extracting the content of the eye area from the original bare face image, that is, the first sample image can be the bare face eye image of the first user; corresponding to the first sample image, the second sample image can be an eyeshadow eye image obtained by extracting the content of the eye area from the original carrier image of the second user and containing eyeshadow.
[0052] It should be noted that the original bare face image can be the bare face image of the first user in the training set, and the original carrier image can be the face image of the second user in the training set that includes eyeshadow.
[0053] Step S102: The perturbation carrier in the second sample image is fused into the first sample image to obtain the initial carrier migration image.
[0054] In related technologies, during the training of a network model for generating perturbed images, the original image of the first user and the carrier image containing the perturbed carrier of the second user are usually used as model inputs. That is, images of two different identities are used as model inputs. This requires the model to not only focus on the processing of the perturbed carrier, but also to pay attention to other elements besides the perturbed carrier, such as the processing of different elements such as skin, eyes, and mouth areas for different users, in order to avoid the generated perturbed image having too much deviation from the original image of the first user in terms of semantic features. This often increases the training difficulty of the model, causes the model to converge slowly, and also makes the generated perturbed image less accurate and natural.
[0055] To address this technical problem, in this embodiment, on one hand, an object that can be directly applied to human skin, such as eyeshadow, is used as a perturbation carrier to enhance the concealment of the generated perturbation image. On the other hand, during the training process of the image generation model, instead of using sample images of two different identities as input to the model, the perturbation carrier in the second sample image of the second user is first fused into the first sample image to obtain an initial carrier transfer image corresponding to the first user, which contains the initial perturbation carrier. Then, the first sample image and the initial carrier transfer image corresponding to the same identity (i.e., both corresponding to the first user) are used as input to the model for model training. This ensures that the images input to the model maintain as much consistency as possible in semantic features during model training, allowing the model to focus solely on processing the perturbation carrier, thereby reducing the difficulty of model training, improving model convergence speed, and enhancing the accuracy and naturalness of the generated perturbation image.
[0056] In some embodiments, fusing the perturbation carrier in the second sample image into the first sample image can be achieved by obtaining the location information of the perturbation carrier in the second sample image and, based on the location information, fusing the perturbation carrier in the second sample image into the corresponding location in the first sample image.
[0057] For example, when the perturbation carrier is eyeshadow, based on the color characteristics of the eyeshadow and the prior knowledge that eyeshadow is usually located around the eyes, the location information of the area containing eyeshadow in the second sample image can be obtained by using the key point in the second sample image, i.e., the position of the eye, as the reference position. Based on this location information, the eyeshadow in the second sample image can be fused into the first sample image. Of course, this is only an example. In actual implementation, the perturbation carrier in the second sample image can also be fused into the first sample image in other ways to obtain the initial carrier migration image. No special limitation is made here.
[0058] Step S103: Based on the first sample image, the initial carrier migration image, and the initial image generation model, train the initial image discrimination model to obtain the target image discrimination model. The initial image discrimination model is used to determine whether the received image was generated by the initial image generation model.
[0059] In this embodiment, an image generation model for generating perturbed images can be obtained by training a Generative Adversarial Network (GAN). GAN is a deep learning model and one of the most promising unsupervised learning methods on complex distributions in recent years. GAN produces reasonably good outputs through the game-like learning between at least two modules in its framework: a generator (Generative Model) and a discriminator (Discriminative Model). In GAN, the discriminator's role is generally to distinguish between real data and generated data, where generated data is obtained by inputting real data into the generator G.
[0060] In this embodiment of the disclosure, the initial image generation model is used as the generator in the Generative Adversarial Network (GAN), and the initial image discrimination model is used as the discriminator in the GAN. During the training process of the image generation model, the initial image discrimination model is first trained based on the first sample image and the initial carrier migration image, which both correspond to the same identity, i.e., the first user. This reduces the training difficulty of the image discrimination model in the GAN, improves its convergence speed and discrimination accuracy. After obtaining the target image generation model, the initial image generation model can be trained based on the target image generation model. In this way, through the mutual competition between the image discrimination model and the image generation model, a more accurate image generation model can be obtained in the end.
[0061] That is, after step S103, step S104 is executed, and the initial image generation model is trained according to the first sample image, the original bare face image, the original carrier image and the target image discrimination model to obtain the image generation model. The image generation model is used to generate a perturbation image corresponding to the third user, and the perturbation image is used to enable the first preset identity recognition model to identify the third user as the second user.
[0062] In this embodiment of the disclosure, the first preset identity recognition model can be any model used to identify the user's identity. The first preset identity recognition model can be a face recognition model, a face detection model, or other models, and no special limitation is made here.
[0063] For example, the first preset identity recognition model can be at least one of face recognition models such as irse50, facenet, and mobileface.
[0064] In addition, in this embodiment of the disclosure, the first user and the third user are both users with different identities from the second user. The first user and the third user can be the same user, or they can be any user different from the second user.
[0065] As can be seen, in this embodiment of the disclosure, since the image generation model is trained based on the first sample image and the initial carrier migration image corresponding to the same identity, namely the first user, during the training process, especially during the training of the target image discrimination model, only the processing of the perturbed carrier in the image can be focused on, while the processing of image elements other than the perturbed carrier in the sample image, such as the eyes and eyebrows, can be neglected. This reduces the training difficulty of the target image discrimination model, improves the convergence speed and discrimination accuracy of the target image discrimination model, and further improves the training speed of the image generation model and the accuracy of the generated perturbed image by training the initial image generation model based on the target image discrimination model.
[0066] In some embodiments, the step S102 described above, which involves fusing the perturbed carrier in the second sample image into the first sample image to obtain an initial carrier migration image, may include: obtaining second coordinate information of the region containing the perturbed carrier in the second sample image based on first coordinate information, wherein the first coordinate information is used to represent the position of the target key point in the second sample image, and the target key point is used to carry the perturbed carrier; and fusing the perturbed carrier into the first sample image based on the second coordinate information, using coordinate point reflection transformation and Poisson fusion processing to obtain an initial carrier migration image.
[0067] The target key point can be an object in a face image used to carry the perturbation carrier. For example, if the perturbation carrier is eyeshadow, the target key point can be the eyes, such as both the left and right eyes at the same time, or the center point of both the left and right eyes; as another example, if the perturbation carrier is lipstick, the target key point can be the center point of the mouth area.
[0068] The first coordinate information can be obtained by performing keypoint detection processing on the second sample image. For example, the first coordinate information can be obtained by inputting the second sample image into a pre-trained object detection model used to detect the location information of keypoints in face images.
[0069] The second coordinate information can be obtained based on the first coordinate information. For example, when the perturbation carrier is eyeshadow, considering that eyeshadow is usually located at key target points, i.e., the area around the eyes, the second coordinate information can be calculated by setting a preset distance threshold and using the preset distance threshold and the first coordinate information. For example, if the center points of the left and right eyes in the first coordinate information are (0,0) and the preset distance threshold is 10, the eyeshadow area can be obtained as a region composed of coordinate points (-10,10), (-10,-10), (10,110), and (10,1-10). Of course, this is only an example; in actual processing, the second coordinate information can also be obtained through other methods, which are not specifically limited here.
[0070] That is, in this embodiment of the present disclosure, in order to accurately fuse the perturbed carrier in the second sample image into the first sample image, when generating the initial carrier migration image, the second coordinate information of the region containing the perturbed carrier in the second sample image can be obtained. Then, based on the second coordinate information, the perturbed carrier can be accurately fused into the first sample image by performing affine transformation and Poisson blending, thereby obtaining the initial carrier migration image corresponding to the first user. This further reduces the difficulty of model training and improves the model convergence speed during subsequent model training.
[0071] In some embodiments, the step S103 described above, which involves training an initial image discrimination model based on a first sample image, an initial carrier migration image, and an initial image generation model to obtain a target image discrimination model, includes: inputting the first sample image into the initial image generation model for image generation processing to obtain a first perturbed carrier sample image corresponding to the first sample image and containing the perturbed carrier; and training the initial image discrimination model based on the first perturbed carrier sample image and the initial carrier migration image to obtain the target image discrimination model.
[0072] In this embodiment, training the initial image discrimination model based on the first perturbation sample image and the initial carrier migration image to obtain the target image discrimination model may include: setting the label of the first perturbation sample image as a first preset label, and setting the label of the initial carrier migration image as a second preset label, wherein the first preset label is used to indicate that the image is a generated image, and the second preset label is used to indicate that the image is not a generated image; training the initial image discrimination model based on the first perturbation sample image after setting the label and the initial carrier migration image after setting the label, and adjusting the parameters of the initial image discrimination model by using a first loss value during the training process to obtain the target image discrimination model; wherein the first loss value is used to represent the error value between the first discrimination result of the input image and the label of the input image; the first discrimination result is the result obtained by the initial image discrimination model after discriminating the input image, used to indicate whether the input image is a generated image, and the input image includes the first perturbation sample image after setting the label and the initial carrier migration image after setting the label.
[0073] It should be noted that, in this embodiment of the disclosure, in order to improve the convergence speed of the model, during the training of the target image discrimination model, the parameters of the initial image generation model can be fixed first, and the initial image discrimination model can be trained first to obtain a converged target discriminator; then, the initial image generation model can be trained again based on the target image discrimination model with converged parameters to obtain the final image generation model with converged parameters.
[0074] For easier understanding, please refer to Figure 2 This is a schematic diagram of the framework for training a target image discrimination model provided in the embodiments of this disclosure.
[0075] like Figure 2 As shown in this embodiment, the perturbation carrier in the second sample image can be fused into the first sample image to obtain an initial carrier migration image, that is, the initial carrier migration image contains the perturbation carrier of the second user, such as an eyeshadow image; then, the first sample image is input into the initial image generation model for image generation processing to obtain the first perturbation sample image.
[0076] Please continue reading. Figure 2 After obtaining the first perturbation sample image and the initial carrier migration image, the two types of images can be labeled separately. For example, the label of the first perturbation sample image can be set to 1, i.e., the first preset label, to identify it as a generated image; and the label of the initial carrier migration image can be set to 0, i.e., the second preset label, to identify it as not a generated image.
[0077] Please continue reading. Figure 2After labeling the first perturbation sample image and the initial carrier migration image, the first perturbation sample image and the initial carrier migration image can be input into the initial image discrimination model for discrimination processing. The error value between the corresponding discrimination result and its label is calculated as the first loss value. The parameters of the initial image discrimination model are then tuned according to the first loss value until the first loss value converges, and finally a target image discrimination model with a loss value that meets the preset conditions, such as being lower than the preset threshold, is obtained.
[0078] For example, the first perturbation sample image after labeling can be input into the initial image discrimination model to obtain discrimination result 11, and the error value between discrimination result 11 and its label can be calculated as loss value 1. The model parameters can be tuned based on loss value 1. Then, the initial carrier migration image after labeling can be input into the parameter-tuned initial image discrimination model to obtain discrimination result 12, and the error value between discrimination result 12 and its label can be calculated as loss value 2. The model parameters can be tuned based on loss value 2, and finally the target image discrimination model can be trained.
[0079] It should be noted that during the training of the target image discrimination model, after calculating the first loss value, the parameters of the initial image discrimination model can be updated by gradient descent to finally train the target image discrimination model.
[0080] As can be seen, the method provided in this disclosure, by first training the initial image discrimination model in the adversarial generative network with generated images and real images based on the same identity, can make the images input to the initial image discrimination model consistent with those input to the initial image generation model in terms of semantic features (e.g., in the case of eyeshadow as the perturbation carrier, the speech feature can be eyes, eyebrows, etc.). This can reduce the training difficulty of the image discrimination model, improve its training speed and accuracy, and thus improve the training speed and accuracy of the final image generation model.
[0081] Please refer to Figure 3 and Figure 4 These are, respectively, a flowchart for training an image generation model provided in the embodiments of this disclosure, and a schematic diagram of a framework for training an image generation model. The following, in conjunction with... Figure 3 and Figure 4 This paper explains how to train an image generation model to generate perturbed images based on the target image discrimination model after training the target image discrimination model.
[0082] like Figure 3 and Figure 4As shown, in some embodiments, the step S104 described above, which involves training the initial image generation model based on the first sample image, the original unedited image, the original carrier image, and the target image discrimination model, to obtain the image generation model, includes:
[0083] Step S301: Input the first sample image into the initial image generation model for image generation processing to obtain a second perturbation sample image containing the perturbation carrier, and set the label of the second perturbation sample image to a first preset label, wherein the first preset label is used to indicate that the second perturbation sample image is a generated image.
[0084] That is, such as Figure 4 As shown, after training the target image discrimination model, during the training of the initial image generation model, the first sample image of the first user can be input into the initial image generation model to generate a second perturbation sample image corresponding to the first sample image and containing the perturbation carrier, and the label of the second perturbation sample image is set as the first preset label indicating that it is an image of the generation type.
[0085] For example, when the perturbation carrier is eyeshadow, a sample image of the first user's bare eyes can be input into the initial image generation model to generate an eyeshadow sample image that contains eyeshadow, corresponding to the bare eye sample image. The label of the eyeshadow sample image is set to a first preset label. For example, the label of the eyeshadow sample image is set to 1 to identify that the eyeshadow sample image is a generated type image.
[0086] Step S302: The second perturbation sample image after setting the label is input into the target image discrimination model for discrimination processing to obtain the second discrimination result. The error value between the second discrimination result and the first preset label is calculated using the first loss function as the second loss value. The second discrimination result is used to indicate whether the second perturbation sample image is a generated image.
[0087] The second discrimination result may include the predicted type and prediction probability value of the second perturbed sample image. The predicted type indicates whether the second perturbed sample image is of the generated type, and the prediction probability value indicates the accuracy of that predicted type. For example, when the first preset label is 1, the second discrimination result may be (1, 80%), indicating that the probability that the predicted type of the second perturbed sample image is of the generated type is 80%.
[0088] In this embodiment of the disclosure, the first loss function can be a function used to calculate the binary cross-entropy (BCE) loss, that is, a function used to calculate the binary cross-entropy between the second discrimination result and the first preset label.
[0089] In the field of machine learning, binary cross-entropy is often used to evaluate the quality of a binary classification model's prediction results. In simple terms, for a sample with a label of 1, if the probability value of the prediction result of the binary classification model approaches 1, then the loss value calculated based on the binary cross-entropy should approach 0. Conversely, if the probability value of the prediction result approaches 0, then the loss value calculated based on the binary cross-entropy will be very large.
[0090] Based on the aforementioned characteristics of binary cross-entropy, in this embodiment of the disclosure, a function for calculating binary cross-entropy loss can be used as a first loss function to calculate the error value between the second discrimination result and the first preset label as a second loss value, and the parameters of the initial image generation model can be optimized based on the second loss value.
[0091] Step S303: Obtain the binary mask corresponding to the second perturbation sample image, and fuse the binary mask with the first sample image to obtain the third perturbation sample image.
[0092] That is, such as Figure 4 As shown, considering that the generated image obtained based on the initial image generation model may affect image elements other than the perturbation carrier, after generating the second perturbation sample image, in order to improve the accuracy of the result, a binary mask corresponding to the second perturbation sample image can be obtained, so as to fuse the generated perturbation carrier into the first sample image through binary mask fusion processing, thereby keeping the elements in the area other than the perturbation carrier in the second sample image as unaffected as possible.
[0093] Among them, with Represents the third perturbation sample image, with O s Representing the first sample image, with O G Let M represent the second perturbed sample image, and M represent the binary mask corresponding to the perturbed carrier. Then, it can be expressed by the formula: The third perturbation sample image is obtained. The binary mask can be obtained by a binary mask extraction algorithm, and the detailed processing method will not be described here.
[0094] Step S304: Use the second loss function to calculate the fusion loss between the third perturbed sample image and the first sample image as the third loss value.
[0095] In this embodiment of the disclosure, the second loss function can be a function used to calculate the fusion loss between the third perturbed sample image and the first sample image.
[0096] Since the fusion loss between images can usually be reflected in differences in image style, image perception, and image structure, the second loss function can be: a function that determines the third loss value representing the fusion loss by weighting and summing the loss values such as style loss, perceptual similarity, and structural similarity between the third perturbed sample image and the first sample image.
[0097] In some embodiments, the third loss value can be obtained through the following steps: inputting the third perturbed sample image and the first sample image into a preset deep feature extraction model for feature extraction processing to obtain a first feature set and a second feature set, and calculating a fifth loss value based on the style loss between the feature data in the first feature set and the corresponding feature data in the second feature set, wherein the first feature set includes feature data obtained by each network layer of the preset deep feature extraction model after performing feature extraction processing on the third perturbed sample image, and the second feature set includes feature data obtained by each network layer of the preset deep feature extraction model after performing feature extraction processing on the first sample image; calculating the perceptual similarity between the third perturbed sample image and the first sample image as a sixth loss value; calculating the structural similarity between the third perturbed sample image and the first sample image as a seventh loss value; and obtaining the third loss value based on the fifth loss value, the sixth loss value, and the seventh loss value.
[0098] A series of studies have shown that deep neural networks have the ability to extract deep features from images. Therefore, in this embodiment of the disclosure, the third perturbation sample image can be used to extract deep features. and the first sample image O s The input is fed into a deep neural network model, such as a pre-trained VGG16 model, to extract deep features between the two sample images. By comparing the intermediate layer outputs between the two images, the style loss between them can be calculated. style As the fifth loss value; where loss style It can be calculated using the following formula:
[0099]
[0100] In this formula, p represents the layer number in the deep neural network model, which is used to extract feature data from the image, and N... P VGG represents the number of channels in the intermediate layer p. p () represents the feature data obtained by the p-th layer of the deep neural network model after feature extraction of the image, gram() represents converting the feature data into the corresponding Gram matrix, β p is a hyperparameter, where j and k represent the row and column positions of elements in the Gram matrix, respectively.
[0101] In this embodiment of the disclosure, the sixth loss value is... lpips That is, the third perturbation sample image and the first sample image O s The perceptual similarity between the two images can be calculated by inputting them into a pre-trained image recognition model, fintune, using the following formula:
[0102]
[0103] In this formula, p represents the layer number of the network layer in the pre-trained image recognition model, which is used to extract feature data from the image, and H... p W p Let f represent the dimensions of the feature data output by the p-th network layer, respectively. p () represents the feature data output by the p-th layer of the image recognition model when the image is input into the model, f p () hw f represents the feature value of the element at position h-th row and w-th column in the feature data output by the p-th network layer. p The input to () could be, for example, a third perturbation sample image. Or the first sample image O s .
[0104] Sixth loss value ssim It can be achieved by measuring the third perturbation sample image. and the first sample image O s The structural similarity (SSIM) between images is obtained by considering factors such as image brightness, image contrast, and image structure. The detailed calculation process will not be elaborated here.
[0105] As can be seen, in this embodiment of the disclosure, when calculating the image fusion loss, by calculating style loss, perceptuality, structural similarity, etc., between images as loss values, edge artifacts during the image fusion process can be further reduced, and blurring within the fused image can be alleviated, thereby improving the visual quality of the generated image. It should be noted that this third loss value can also include other loss values, such as a loss value based on the Laplacian operator. grad In addition, TV (Total Variation loss) is used to make the generated image smoother, without any special restrictions here.
[0106] Step S305: The third perturbation sample image is fused into the original plain image to obtain the fourth perturbation sample image.
[0107] like Figure 4 As shown, the third perturbation sample image was obtained. Subsequently, since the third perturbation sample image may only contain the target area of the first user, such as the eye area, it is also necessary to fuse the third perturbation sample image into the first user's original bare-face image I. s For example, from the original face image, the fourth perturbation sample image is obtained.
[0108] Step S306: Based on the fourth perturbation sample image and the original carrier image, perform meta-learning processing on the second preset identity recognition model, and calculate the fourth loss value based on the third loss function during the meta-learning process. The fourth loss value represents the learning loss during the meta-learning process, and the second preset identity recognition model is any model used for identity recognition.
[0109] The second preset identity recognition model can be the same model as the first preset identity recognition model, or it can be any one or more other identity recognition models that are different from the first preset identity recognition model. That is, in this embodiment of the disclosure, the features of any one or more identity recognition models can be learned by treating them as black boxes through a preset meta-learning algorithm, and the parameters of the initial image generation model can be optimized based on the learned learning loss, so that the trained image generation model can generate perturbation images that can cause the identity recognition model to be identified incorrectly without knowing the model parameters, model structure and other features of the identity recognition model to be attacked, thereby improving the universality of the obtained image generation model.
[0110] like Figure 4 As shown, the preset meta-learning processing in this embodiment of the present disclosure may be to use a preset meta-learning algorithm to learn the recognition processing of the second preset identity recognition model, and to optimize its parameters based on the learning loss in the learning process, i.e., the fourth loss value, so that the generated perturbation image can meet the recognition conditions of the second preset identity recognition model. That is, the initial image model is made to "learn to learn" based on the preset meta-learning algorithm.
[0111] Step S307: Based on the second loss value, the third loss value, and the fourth loss value, obtain the target loss value, and adjust the parameters of the initial image generation model according to the target loss value to obtain the image generation model.
[0112] In this embodiment of the disclosure, the target loss value can be obtained by weighted summation of the second loss value, the third loss value, and the fourth loss value. The weight corresponding to each loss value can be set as needed, and no special limitation is made here.
[0113] As can be seen from the above description, in the embodiments of this disclosure, during the training of the initial image generation model, the accuracy of the second discrimination result can be improved by using the target image discrimination model pre-trained based on the initial image generation model to discriminate the second perturbation sample image generated by the initial image generation model during the training process. Then, based on the second discrimination result, a second loss value is obtained, and the parameters of the initial image generation model are optimized in reverse according to the second loss value, which can reduce the difficulty of model training. At the same time, since the fusion loss between the generated third perturbation sample image and the original first sample image is also considered during the training process, and the parameters of the initial image generation model are optimized using the third loss value representing the fusion loss, the edge artifacts in the generated perturbation image can be further reduced and the concealment of the perturbation image can be improved. Furthermore, by introducing a second preset identity model and performing preset meta-learning processing on the second preset identity model during the training process, and optimizing the parameters of the initial image generation model based on the learned learning loss as a fourth loss value, the perturbation image generated by the trained image generation model can be made more universal, that is, the transferability of the perturbation image can be improved.
[0114] In the field of machine learning, meta-learning aims to enable models to "learn how to learn," that is, to enable the model B to learn the characteristics of model A without needing to know the model parameters and structure of model A. The basic unit of meta-learning is a task, which can generally be divided into meta-training processing and meta-testing processing. The task of meta-training processing can be understood as training the model so that the model to be trained can learn the features of the model to be learned or attacked. Meta-testing processing can be understood as verifying the knowledge learned by the model to be trained in meta-training processing to measure the learning effect.
[0115] Please refer to Figure 5a and Figure 5b These are, respectively, a flowchart and a framework diagram for calculating the fourth loss value provided in the embodiments of this disclosure. Figure 5a As shown, in some embodiments, step S406 above, which involves performing preset meta-learning processing on the second preset identity recognition model based on the fourth perturbation sample image and the original carrier image, and calculating the fourth loss value based on the third loss function during the preset meta-learning processing, includes:
[0116] Step S501: Perform input transformation (DI, Diverse Input) processing on the fourth perturbation sample image to obtain the fifth perturbation sample image.
[0117] like Figure 5bAs shown in this embodiment, during the training of the initial image generation model, the fourth perturbation sample image can be transformed before performing preset meta-learning processing on any second preset identity recognition model, so as to make the sample distribution input into the second preset identity recognition model as wide as possible, avoid model training overfitting, and improve the transferability of the final training result.
[0118] Step S502: Input the fifth perturbation sample image into the second preset identity recognition model for recognition processing to obtain the first recognition result; and input the original carrier image into the second preset identity recognition model for recognition processing to obtain the second recognition result.
[0119] Step S503: Calculate the eighth loss value between the first recognition result and the second recognition result based on the third loss function, adjust the parameters of the initial image generation model according to the eighth loss value, and generate the sixth perturbation sample image corresponding to the fifth perturbation sample image based on the initial image generation model with adjusted parameters.
[0120] The recognition result of the second preset identity recognition model can be a feature vector obtained after recognizing and processing the input image. Of course, the recognition result can also include an identity prediction result, which is used to represent the identity of the user corresponding to the input image. No special limitation is made here. The input images of the second preset identity recognition model can be the fifth perturbation sample image and the original carrier image, respectively.
[0121] The third loss function can be a loss function used to calculate the cosine similarity between the first and second recognition results. Of course, this loss function can also be other functions used to measure the similarity between the first and second recognition results, without any special limitation here.
[0122] That is, by using the second recognition result as a reference, the cosine similarity is used to measure the similarity between it and the first recognition result as the eighth loss value, so as to determine the difference between the image features of the fifth perturbation sample image of the generated type and the image features of the real image during the identity recognition process; then, the parameters of the initial image generation model can be adjusted once according to the eighth loss value so that the initial image generation model can generate a sixth perturbation sample image whose feature vector is more similar to the real image.
[0123] It should be noted that during this step, after adjusting the parameters of the initial image generation model based on the eighth loss value, a sixth perturbation sample image corresponding to the current fifth perturbation sample image is only generated based on the adjusted initial image generation model and the first sample image. However, the adjusted parameters are not actually applied to the next round of training. Instead, the parameters of the initial image generation model are actually adjusted based on the obtained target loss value after the current round of training ends. Furthermore, in the process of inputting the first sample image into the adjusted initial image generation model to generate the sixth perturbation sample image, the process involves obtaining a binary mask of the generated image, fusing the binary mask into the first sample image, fusing the first sample image with the fused binary mask into the original plain image, and performing an input transformation on the fused original plain image before generating the sixth perturbation sample image. The detailed process is similar to the generation process of the fifth perturbation sample image and will not be repeated here.
[0124] like Figure 5b As shown in the present embodiment, the above steps S502-S503 are meta-training processes in the meta-learning process, that is, to enable the initial image generation model to learn the characteristics of any identity recognition model when performing recognition processing, and to try to adjust its own parameters to adapt to the recognition processing of the identity recognition model.
[0125] Step S504: Perform input transformation processing on the sixth perturbation sample image to obtain the seventh perturbation sample image.
[0126] like Figure 5b As shown, after obtaining the sixth perturbation sample image, in order to further improve the transferability of the training results, the input transformation process is still performed first.
[0127] Step S505: Input the seventh perturbation sample image into the third preset identity recognition model for recognition processing to obtain the third recognition result; and input the original carrier image into the third preset identity recognition model for recognition processing to obtain the fourth recognition result, wherein the third preset identity recognition model is any model used for identity recognition.
[0128] The third preset identity recognition model can be the same model as the first and second preset identity recognition models mentioned above, or it can be any one or more other identity recognition models, without any special restrictions here.
[0129] Step S506: Calculate the ninth loss value between the third recognition result and the fourth recognition result based on the third loss function; and step S507: Obtain the fourth loss value based on the eighth loss value and the ninth loss value.
[0130] Corresponding to the recognition result output by the second preset identity recognition model, in this embodiment of the disclosure, the recognition result obtained based on the third preset identity recognition model can be at least a feature vector obtained after recognizing and processing the input image.
[0131] In this embodiment of the disclosure, steps S504-S506 are metatest processing in the meta-learning process. That is, in the meta-training process, the parameters of the initial image generation model are adjusted and the sixth perturbation sample image is regenerated using the eighth loss value. In order to verify whether the parameters learned in the meta-training process are appropriate, metatest processing can be performed. That is, the seventh perturbation sample image obtained by input transformation of the sixth perturbation sample image and the original carrier image are respectively input into the third preset identity recognition model for recognition processing. The cosine similarity between the third recognition result and the fourth recognition result calculated based on the third loss function is used as the ninth loss value to measure whether the adjustment of the model parameters in step S503 is appropriate. Afterwards, the eighth loss value obtained in the meta-training process and the ninth loss value obtained in the metatest process can be weighted and summed to obtain the fourth loss value used to represent the learning loss in the first-order learning process. The parameters of the initial image generation model are then optimized based on the fourth loss value.
[0132] As can be seen, in the embodiments of this disclosure, by performing preset meta-learning processing on the second preset identity recognition model during the training of the image generation model, the accuracy and transferability of the perturbation image generated by the image generation model can be further improved.
[0133] Please refer to Figure 6 This is a flowchart of the image recognition method provided in the embodiments of this disclosure. That is, after training a motion image generation model for generating perturbed images based on any one or more of the above embodiments, any fourth preset identity recognition model can be attacked based on steps S601-S603 as shown in 6, so that the fourth preset identity recognition model identifies the third user as the second user.
[0134] Step S601: Obtain the unedited image of the third user.
[0135] Step S602: Input the plain image into the image generation model to obtain the target perturbation image. The image generation model is obtained according to the image generation model training method.
[0136] Step S603: Input the target perturbation image into the fourth preset identity recognition model to obtain the target recognition result, wherein the target recognition result indicates that the third user's identity is the second user, and the fourth preset identity recognition model is any model used for identity recognition.
[0137] In the image recognition method provided in this disclosure, since the target perturbation image input to the fourth preset identity recognition model, such as the Face++ model, is generated based on the image generation model obtained in any of the above embodiments, and since the perturbation carrier added to the third user's bare face image by the image generation model acts directly on the human skin rather than existing separately from the human body, and the fusion loss between images is comprehensively considered during the image generation process, it can improve the concealment and accuracy of the target perturbation image, thereby improving the success rate of the fourth preset identity recognition model in recognizing the third user as the second user; in addition, compared with the image generation model trained in related technologies that can only perform untargeted attacks on the identity recognition model to be attacked, the target perturbation image obtained based on this disclosure embodiment can enable the fourth preset identity model to recognize the user's identity as a specified identity, for example, it can enable the fourth preset identity model to accurately recognize the third user as the first user, thereby making the attack on the identity recognition model more targeted.
[0138] Please refer to Figure 7 This is a schematic diagram illustrating an application scenario provided by an embodiment of this disclosure, wherein, in Figure 7 In this example, we will use the disturbance carrier as eyeshadow to illustrate how the fourth preset identity model can identify user 2 as user 3.
[0139] like Figure 7 As shown, since the eyeshadow image of user 2 can be obtained in advance during actual implementation, the initial image generation model can be pre-trained based on the eyeshadow image of user 2 using the training method in this embodiment to obtain the image generation model for generating the perturbation image.
[0140] After obtaining the image generation model, as follows Figure 7 As shown, a bare-faced image of User 3 can be input into the image generation model to obtain a perturbed image of User 3; then, after inputting the perturbed image of User 3 into the fourth preset identity recognition model, the fourth preset identity model will output as follows. Figure 7 The result shown indicates that user 3 is identified as user 2, thus causing the fourth preset identity recognition model to incorrectly identify user 3 as user 2, thereby achieving the technical effect of protecting user 3's identity information.
[0141] In practical applications, with user authorization, an image generation model based on the training method described in this disclosure can be deployed on the user's terminal device, such as a mobile phone. When a user installs and registers an application on the mobile phone, in order to prevent the application from engaging in "price discrimination" by illegally obtaining the user's facial image, the user's eye shadow image obtained based on the image generation model can be sent to the application for recognition, so that the application will mistakenly identify the user as another user, thereby preventing the application from engaging in "price discrimination" against the user.
[0142] It should be noted that, in practical implementation, this image recognition method can also be applied to identity verification defense scenarios. For example, for financial consumer applications, they typically perform identity verification processing on users with their authorization to obtain user identity information and query user credit scores, and then provide financial services, such as lending services, to users based on the obtained credit scores. Therefore, in such applications, the accuracy of user identity verification results is particularly important for their subsequent business operations. Thus, a perturbed image can be generated by the image generation model trained based on the embodiments of this disclosure, input into the identity verification model used by such applications for recognition processing, and the parameters of the identity verification model used can be optimized based on the recognition results to avoid misidentification.
[0143] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0144] In addition, this disclosure also provides a training device for an image generation model, an image recognition device, an electronic device, and a computer-readable storage medium. All of the above can be used to implement any training method for an image generation model or any image recognition method provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding section of the method and will not be repeated here.
[0145] Figure 8 This is a block diagram of a training apparatus for an image generation model provided in an embodiment of the present disclosure.
[0146] Reference Figure 8 This disclosure provides a training device for an image generation model. The training device 800 for the image generation model includes: an acquisition unit 801, a fusion unit 802, a first training unit 803, and a second training unit 804.
[0147] The acquisition unit 801 is used to acquire a first sample image and a second sample image, wherein the first sample image contains the target region in the original bare-faced image of the first user, and the second sample image contains the region in the original carrier image of the second user that corresponds to the position of the target region and contains the disturbed carrier.
[0148] The fusion unit 802 is used to fuse the perturbation carrier in the second sample image into the first sample image to obtain an initial carrier migration image;
[0149] The first training unit 803 is used to train the initial image discrimination model based on the first sample image, the initial carrier transfer image and the initial image generation model to obtain the target image discrimination model. The initial image discrimination model is used to determine whether the received image is generated by the initial image generation model.
[0150] The second training unit 804 is used to train the initial image generation model based on the first sample image, the original bare face image, the original carrier image and the target image discrimination model to obtain the image generation model. The image generation model is used to generate a perturbation image corresponding to the third user. The perturbation image is used to enable the first preset identity recognition model to identify the third user as the second user.
[0151] In some embodiments, when the initial migration unit 802 migrates the perturbation carrier in the second sample image to the first sample image to obtain an initial carrier migration image, it can be used to: obtain the coordinate information of the pixels containing the perturbation carrier in the second sample image; and, based on the coordinate information, migrate the pixels containing the perturbation carrier to the corresponding pixels in the first sample image based on coordinate point reflection transformation and Poisson fusion processing to obtain the initial carrier migration image.
[0152] In some embodiments, when the first training unit 803 trains the initial image discrimination model based on the first sample image, the initial carrier migration image, and the initial image generation model to obtain the target image discrimination model, it can be used to: input the first sample image into the initial image generation model for image generation processing to obtain a first perturbation sample image corresponding to the first sample image and containing the perturbation carrier; and train the initial image discrimination model based on the first perturbation sample image and the initial carrier migration image to obtain the target image discrimination model.
[0153] In some embodiments, when the first training unit 802 trains the initial image discrimination model based on the first perturbation sample image and the initial carrier migration image to obtain the target image discrimination model, it can be used to: set the label of the first perturbation sample image as a first preset label, and set the label of the initial carrier migration image as a second preset label, wherein the first preset label is used to indicate that the image is a generated image, and the second preset label is used to indicate that the image is not a generated image; train the initial image discrimination model based on the first perturbation sample image after setting the label and the initial carrier migration image after setting the label, and adjust the parameters of the initial image discrimination model by using a first loss value during the training process to obtain the target image discrimination model; wherein the first loss value is used to represent the error value between the first discrimination result of the input image and the label of the input image; the first discrimination result is the result obtained by the initial image discrimination model after discriminating the input image, which is used to indicate whether the input image is a generated image, and the input image includes the first perturbation sample image after setting the label and the initial carrier migration image after setting the label.
[0154] In some embodiments, when the second training unit 804 trains the initial image generation model based on the first sample image, the original unedited image, the original carrier image, and the target image discrimination model to obtain an image generation model, it can be used to: input the first sample image into the initial image generation model for image generation processing to obtain a second perturbed sample image containing a perturbed carrier, and set the label of the second perturbed sample image as a first preset label, wherein the first preset label is used to indicate that the second perturbed sample image is a generated image; input the second perturbed sample image after setting the label into the target image discrimination model for discrimination processing to obtain a second discrimination result, and use a first loss function to calculate the error value between the second discrimination result and the first preset label as a second loss value, wherein the second discrimination result is used to indicate whether the second perturbed sample image is a generated image; obtain A binary mask corresponding to the second perturbation sample image is used, and the binary mask is fused with the first sample image to obtain the third perturbation sample image. The fusion loss between the third perturbation sample image and the first sample image is calculated using the second loss function as the third loss value. The third perturbation sample image is fused into the original plain image to obtain the fourth perturbation sample image. Based on the fourth perturbation sample image and the original carrier image, the second preset identity recognition model is subjected to preset meta-learning processing, and a fourth loss value is calculated based on the third loss function during the preset meta-learning process. The fourth loss value represents the learning loss during the preset meta-learning process, and the second preset identity recognition model is any model used for identity recognition. Based on the second loss value, the third loss value, and the fourth loss value, the target loss value is obtained, and the parameters of the initial image generation model are adjusted according to the target loss value to obtain the image generation model.
[0155] In some embodiments, the second training unit 804, when calculating the third loss value, may be used to: input the third perturbed sample image and the first sample image into a preset deep feature extraction model for feature extraction processing to obtain a first feature set and a second feature set, and calculate a fifth loss value based on the style loss between the feature data in the first feature set and the corresponding feature data in the second feature set, wherein the first feature set includes feature data obtained by each network layer of the preset deep feature extraction model after performing feature extraction processing on the third perturbed sample image, and the second feature set includes feature data obtained by each network layer of the preset deep feature extraction model after performing feature extraction processing on the first sample image; calculate the perceptual similarity between the third perturbed sample image and the first sample image as a sixth loss value; calculate the structural similarity between the third perturbed sample image and the first sample image as a seventh loss value; and obtain the third loss value based on the fifth loss value, the sixth loss value, and the seventh loss value.
[0156] In some embodiments, when the second training unit 804 performs preset meta-learning processing on the second preset identity recognition model based on the fourth perturbation sample image and the original carrier image, and calculates the fourth loss value based on the third loss function during the preset meta-learning processing, it can be used to: perform input transformation processing on the fourth perturbation sample image to obtain a fifth perturbation sample image; input the fifth perturbation sample image into the second preset identity recognition model for recognition processing to obtain a first recognition result; and input the original carrier image into the second preset identity recognition model for recognition processing to obtain a second recognition result; calculate an eighth loss value between the first recognition result and the second recognition result based on the third loss function, and adjust the system according to the eighth loss value. The parameters of the initial image generation model are determined, and a sixth perturbation sample image corresponding to the fifth perturbation sample image is generated based on the adjusted parameters of the initial image generation model. The sixth perturbation sample image is subjected to input transformation processing to obtain a seventh perturbation sample image. The seventh perturbation sample image is input into a third preset identity recognition model for recognition processing to obtain a third recognition result. The original carrier image is also input into the third preset identity recognition model for recognition processing to obtain a fourth recognition result. The third preset identity recognition model is an arbitrary model used for identity recognition. A ninth loss value is calculated based on a third loss function between the third and fourth recognition results. The fourth loss value is obtained based on the eighth and ninth loss values.
[0157] As can be seen, the training apparatus for the image generation model provided in this embodiment can train the initial image discrimination model in the adversarial generative network by first training the generated image and the real image based on the same identity in the first training unit. This can make the image input to the initial image discrimination model consistent with the input to the initial image generation model in terms of semantic features (e.g., in the case of eyeshadow as the perturbation carrier, the speech feature can be eyes, eyebrows, etc.). This can reduce the training difficulty of the image discrimination model and improve its training speed and accuracy. Then, the initial image generation model can be trained in the second training unit based on the target image discrimination model, which can further improve the training speed and accuracy of the image generation model.
[0158] Figure 9 This is a block diagram of an image recognition device provided in an embodiment of the present disclosure.
[0159] Reference Figure 9 This disclosure provides an image recognition device 900, which includes an image acquisition unit 901, an image generation unit 902, and a recognition unit 903.
[0160] The image acquisition unit 901 is used to acquire a bare-faced image of a third user.
[0161] The image generation unit 902 is used to input a plain image into an image generation model to obtain a target perturbation image, wherein the image generation model is obtained according to the training method of the image generation model.
[0162] The recognition unit 903 is used to input the target perturbation image into the fourth preset identity recognition model for identity recognition processing to obtain the target recognition result, wherein the target recognition result indicates that the third user's identity is the second user, and the fourth preset identity recognition model is any model used for identity recognition.
[0163] In the image recognition device provided in this embodiment, since the target perturbation image input to the fourth preset identity recognition model, such as the Face++ model, is generated by the image generation unit based on the training method of the image generation model in the above embodiment, and since the perturbation carrier added to the third user's bare face image by the image generation model acts directly on the human skin rather than existing outside the human body, and the fusion loss between images is comprehensively considered in the process of generating the image, it can improve the concealment and accuracy of the target perturbation image, thereby improving the success rate of the fourth preset identity recognition model in recognizing the third user as the second user.
[0164] Figure 10 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0165] ReferenceFigure 10 This disclosure provides an electronic device 1000, comprising: at least one processor 1001; at least one memory 1002; and one or more I / O interfaces 1003 connected between the processor 1001 and the memory 1002; wherein the memory 1002 stores one or more computer programs executable by the at least one processor 1001, the one or more computer programs being executed by the at least one processor 1001 to enable the at least one processor 1001 to execute the above-described image generation model training method or image recognition method.
[0166] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the above-described image generation model training method or image recognition method. The computer-readable storage medium may be volatile or non-volatile.
[0167] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the above-described image generation model training method or model adversarial method.
[0168] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0169] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0170] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0171] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0172] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0173] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0174] These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other training device for a programmable image generation model, thereby producing a machine such that, when executed by the processor of the computer or other training device for a programmable image generation model, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a training device for a programmable image generation model, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0175] Computer-readable program instructions may also be loaded onto a computer, a training device for other programmable image generation models, or other devices, causing a series of operational steps to be executed on the computer, the training device for other programmable image generation models, or other devices to produce a computer-implemented process, thereby enabling the instructions executed on the computer, the training device for other programmable image generation models, or other devices to perform the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0177] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A training method for an image generation model, characterized in that, include: Acquire a first sample image and a second sample image, wherein the first sample image contains the target region in the original bare-faced image of the first user, and the second sample image contains the region in the original carrier image of the second user that corresponds to the position of the target region and contains the disturbed carrier. The perturbation carrier in the second sample image is fused into the first sample image to obtain an initial carrier migration image; The first sample image is input into the initial image generation model for image generation processing to obtain a first perturbation sample image that contains the perturbation carrier, corresponding to the first sample image. Based on the first perturbation sample image and the initial carrier migration image, the initial image discrimination model is trained to obtain the target image discrimination model, wherein the initial image discrimination model is used to determine whether the received image is generated by the initial image generation model; The initial image generation model is trained based on the first sample image, the original unedited image, the original carrier image, and the target image discrimination model to obtain an image generation model. The image generation model is used to generate a perturbation image corresponding to the third user, and the perturbation image is used to enable the first preset identity recognition model to identify the third user as the second user.
2. The method according to claim 1, characterized in that, The step of fusing the perturbed carrier in the second sample image to the first sample image to obtain an initial carrier migration image includes: Based on the first coordinate information, the second coordinate information of the region containing the disturbance carrier in the second sample image is obtained, wherein the first coordinate information is used to represent the position of the target key point in the second sample image, and the target key point is used to carry the disturbance carrier; Based on the second coordinate information, and using coordinate point reflection transformation and Poisson fusion processing, the disturbed carrier is fused into the first sample image to obtain the initial carrier migration image.
3. The method according to claim 1, characterized in that, The step of training the initial image discrimination model based on the first perturbed sample image and the initial carrier migration image to obtain the target image discrimination model includes: The label of the first perturbation sample image is set to a first preset label, and the label of the initial carrier migration image is set to a second preset label, wherein the first preset label is used to indicate that the image is a generated image, and the second preset label is used to indicate that the image is not a generated image; Based on the first perturbation sample image after labeling and the initial carrier migration image after labeling, the initial image discrimination model is trained respectively, and the parameters of the initial image discrimination model are adjusted by using a first loss value during the training process to obtain the target image discrimination model; Wherein, the first loss value is used to represent the error value between the first discrimination result of the input image and the label of the input image; the first discrimination result is the result obtained by the initial image discrimination model after discriminating the input image, which is used to indicate whether the input image is a generated image, and the input image includes the first perturbation sample image after setting the label and the initial carrier migration image after setting the label.
4. The method according to claim 1, characterized in that, The step of training the initial image generation model based on the first sample image, the original bare-faced image, the original carrier image, and the target image discrimination model to obtain the image generation model includes: The first sample image is input into the initial image generation model for image generation processing to obtain a second perturbation sample image containing the perturbation carrier, and the label of the second perturbation sample image is set to a first preset label, wherein the first preset label is used to indicate that the second perturbation sample image is a generated image; The second perturbation sample image with the label is input into the target image discrimination model for discrimination processing to obtain a second discrimination result. The error value between the second discrimination result and the first preset label is calculated using a first loss function as the second loss value. The second discrimination result is used to indicate whether the second perturbation sample image is a generated image. Obtain the binary mask corresponding to the second perturbation sample image, and fuse the binary mask with the first sample image to obtain the third perturbation sample image; The fusion loss between the third perturbation sample image and the first sample image is calculated using the second loss function as the third loss value; The third perturbation sample image is fused into the original plain image to obtain the fourth perturbation sample image; Based on the fourth perturbation sample image and the original carrier image, a preset meta-learning process is performed on the second preset identity recognition model, and a fourth loss value is calculated based on the third loss function during the preset meta-learning process. The fourth loss value represents the learning loss during the preset meta-learning process, and the second preset identity recognition model is any model used for identity recognition. Based on the second loss value, the third loss value, and the fourth loss value, a target loss value is obtained, and the parameters of the initial image generation model are adjusted according to the target loss value to obtain the image generation model.
5. The method according to claim 4, characterized in that, The method calculates the third loss value through the following steps: The third perturbation sample image and the first sample image are input into a preset deep feature extraction model for feature extraction processing to obtain a first feature set and a second feature set. A fifth loss value is calculated based on the style loss between the feature data in the first feature set and the corresponding feature data in the second feature set. The first feature set includes feature data obtained by each network layer of the preset deep feature extraction model after performing feature extraction processing on the third perturbation sample image. The second feature set includes feature data obtained by each network layer of the preset deep feature extraction model after performing feature extraction processing on the first sample image. The perceptual similarity between the third perturbed sample image and the first sample image is calculated as the sixth loss value; The structural similarity between the third perturbation sample image and the first sample image is calculated as the seventh loss value; The third loss value is obtained based on the fifth loss value, the sixth loss value, and the seventh loss value.
6. The method according to claim 4, characterized in that, The step of performing preset meta-learning processing on the second preset identity recognition model based on the fourth perturbation sample image and the original carrier image, and calculating the fourth loss value based on the third loss function during the preset meta-learning processing, includes: The fourth perturbation sample image is subjected to input transformation processing to obtain the fifth perturbation sample image; The fifth perturbation sample image is input into the second preset identity recognition model for recognition processing to obtain a first recognition result; and the original carrier image is input into the second preset identity recognition model for recognition processing to obtain a second recognition result. The eighth loss value between the first recognition result and the second recognition result is calculated based on the third loss function. The parameters of the initial image generation model are adjusted according to the eighth loss value. A sixth perturbation sample image corresponding to the fifth perturbation sample image is generated based on the initial image generation model with adjusted parameters. The input transformation process is applied to the sixth perturbation sample image to obtain the seventh perturbation sample image; The seventh perturbation sample image is input into the third preset identity recognition model for recognition processing to obtain the third recognition result; and the original carrier image is input into the third preset identity recognition model for recognition processing to obtain the fourth recognition result, wherein the third preset identity recognition model is any model used for identity recognition. A ninth loss value is calculated based on the third loss function between the third identification result and the fourth identification result; The fourth loss value is obtained based on the eighth loss value and the ninth loss value.
7. An image recognition method, characterized in that, include: Obtain a third user's unedited image; The unedited image is input into an image generation model to obtain a target perturbation image, wherein the image generation model is obtained by the training method according to any one of claims 1-6; The target perturbation image is input into the fourth preset identity recognition model for identity recognition processing to obtain the target recognition result, wherein the target recognition result indicates that the third user's identity is the second user, and the fourth preset identity recognition model is any model used for identity recognition.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1-7.
10. A computer program product, characterized in that, Includes computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Makeup migration method and device and makeup migration network training method and device
CN114283052A