Image processing method and device and storage medium
The image processing model adds adversarial perturbation to the makeup style, which solves the contradiction between privacy protection and nature in face recognition, generates high-quality privacy-protected face images, and improves user experience and recognition accuracy.
Patent Information
- Application Number
- CN202410077380.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the field of face recognition, it is difficult to maintain the naturalness of images and the accuracy of recognition while protecting user privacy. Traditional methods often lead to makeup artifacts or high training costs.
Using an image processing model, through encoding networks, generating adversarial networks and perturbation networks, text description information is used to guide perturbation images to add adversarial perturbations to make-up styles, generating high-quality privacy-protected face images.
It realizes effective protection of face privacy without affecting the user's visual experience and can deceive unknown automated face recognition systems.
Smart Images

Figure CN120339042A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular, to an image processing method, apparatus, and storage medium. Background Art
[0002] With the development of deep learning technology, deep learning has been increasingly widely used in the field of face recognition, which has triggered related issues of face information security. When face recognition technology is misused, there may be a security risk of information leakage of user face information. Therefore, it is necessary to process face images in images or videos to protect personal privacy. Face recognition privacy protection technology is a complex issue that requires comprehensive consideration of multiple factors such as privacy protection and face recognition accuracy. Current research mainly focuses on aspects such as encryption-based, privacy protection-based, differential privacy-based, and multi-party secure computing-based. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present disclosure provides an image processing method, apparatus, and storage medium.
[0004] According to a first aspect of an embodiment of the present disclosure, an image processing method is provided, including:
[0005] Obtaining a perturbation image and a face image to be processed, where the perturbation image is used to add a perturbation to the face image;
[0006] Processing the face image according to the perturbation image through a pre-trained image processing model to obtain a target face image;
[0007] Wherein, the image processing model is a model pre-trained according to multiple face sample images, text description information respectively corresponding to each of the face sample images, and original face images respectively corresponding to each of the face sample images. The face sample images are face images that have undergone a preset modification process, and the text description information is used to describe the modification features corresponding to the preset modification process, and the text description information is used to guide the perturbation image to add a perturbation in the modification features.
[0008] Optionally, the image processing model includes an encoding network, a generative adversarial network, and a perturbation network;
[0009] The processing the face image according to the perturbation image through a pre-trained image processing model to obtain a target face image includes:
[0010] Inputting the face image into the encoding network to obtain a first latent feature, where the first latent feature characterizes the facial features of the face image;
[0011] Determine a second latent feature through the generative adversarial network and the perturbation network according to the first latent feature and the perturbed image, where the second latent feature characterizes the perturbed image feature corresponding to the face image;
[0012] Determine the target face image through the generative adversarial network according to the first latent feature and the second latent feature.
[0013] Optionally, the determining the second latent feature through the generative adversarial network and the perturbation network according to the first latent feature and the perturbed image includes:
[0014] Input the first latent feature into the generative adversarial network to obtain the simulated face feature corresponding to the face image;
[0015] Input the simulated face feature and the perturbed image into the perturbation network to obtain the second latent feature.
[0016] Optionally, the determining the target face image through the generative adversarial network according to the first latent feature and the second latent feature includes:
[0017] Perform feature merging on the first latent feature and the second latent feature to obtain target image features;
[0018] After inputting the target image features into the generative adversarial network, output the target face image through the generative adversarial network.
[0019] Optionally, the image processing model is pre-trained in the following manner:
[0020] Obtain a preset initial model, training samples, and random perturbed images for model training. The preset initial model includes an encoding network, a preset generative model to be trained, and a preset perturbation model to be trained. The training samples include multiple face sample images, text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image;
[0021] Perform model training on the preset generative model according to the multiple face sample images, the text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image to obtain a generative adversarial network;
[0022] Form a new image processing model to be trained by combining the encoding network, the generative adversarial network, and the preset perturbation model to be trained;
[0023] After training the preset perturbation model with the new image processing model to be trained according to the training samples and the random perturbation images, the image processing model is obtained.
[0024] Optionally, the training of the preset generation model according to the multiple face sample images, the text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image to obtain the generative adversarial network includes:
[0025] For each face sample image, after inputting the face sample image into the encoding network, the encoding network outputs low-dimensional image features. The low-dimensional image features are the image features after mapping the face sample image to the latent space, and the latent space represents the low-dimensional vector space on the input side of the generative adversarial network;
[0026] After inputting the low-dimensional image features into the preset generation model, the model output corresponding to the face sample image is obtained. The model output includes an output image and output text information;
[0027] After training the preset generation model according to the output image, the output text information, the original face image corresponding to the face sample image, and the text description information corresponding to the face sample image, the generative adversarial network is obtained.
[0028] Optionally, the training of the preset generation model according to the output image, the output text information, the original face image corresponding to the face sample image, and the text description information corresponding to the face sample image to obtain the generative adversarial network includes:
[0029] Determine the image loss of the model output according to the output image and the original face image corresponding to the face sample image;
[0030] Determine the text loss of the model output according to the output text information and the text description information corresponding to the face sample image;
[0031] Determine the first loss function according to the image loss and the text loss;
[0032] After training the preset generation model according to the first loss function, the generative adversarial network is obtained.
[0033] Optionally, the training of the preset perturbation model with the new image processing model to be trained according to the training samples and the random perturbation images to obtain the image processing model includes:
[0034] For each of the face sample images, after inputting the face sample image into the image processing model to be trained, obtain the low-dimensional image features output by the encoding network and the model output of the generative adversarial network, where the model output includes an output image and output text information;
[0035] After inputting the output image, the output text information, and the random perturbation image into the preset perturbation model, obtain perturbation image features guided by the output text information;
[0036] After training the preset perturbation model according to the face sample image, the random perturbation image, the low-dimensional image features, and the perturbation image features, obtain a trained perturbation network;
[0037] Combine the encoding network, the generative adversarial network, and the perturbation network to obtain the image processing model.
[0038] Optionally, the step of obtaining a trained perturbation network after training the preset perturbation model according to the face sample image, the random perturbation image, the low-dimensional image features, and the perturbation image features includes:
[0039] After merging the low-dimensional image features and the perturbation image features, input them into the generative adversarial network to obtain a model prediction image corresponding to the face sample image;
[0040] Determine a face adversarial loss according to the face sample image, the random perturbation image, and the model prediction image;
[0041] Determine a face identity preservation loss according to the low-dimensional image features, the perturbation image features, and the model prediction image;
[0042] Determine a second loss function according to the face identity preservation loss and the face adversarial loss;
[0043] After training the preset perturbation model according to the second loss function, obtain a trained perturbation network.
[0044] Optionally, the step of determining a face adversarial loss according to the face sample image, the random perturbation image, and the model prediction image includes:
[0045] Determine a first distance metric between the model prediction image and the face sample image, and a second distance metric between the model prediction image and the random perturbation image, where the smaller the distance metric, the higher the probability that the two face images corresponding to the model prediction belong to the same identity;
[0046] Use the difference between the second distance metric and the first distance metric as the face adversarial loss.
[0047] Optionally, determining the face identity preservation loss according to the low-dimensional image feature, the perturbed image feature, and the model prediction image includes:
[0048] Perform element-wise multiplication on the low-dimensional image feature and the model prediction image to obtain a first product;
[0049] Perform element-wise multiplication on the perturbed image feature and the model prediction image to obtain a second product. The larger the product, the higher the probability that the face identity corresponding to the image feature is the same as the face identity of the face sample image;
[0050] Determine the face identity preservation loss according to the difference between the second product and the first product.
[0051] Optionally, the face image includes an original image and a real-time captured image corresponding to a target face recognition task; processing the face image according to the perturbed image through a pre-trained image processing model to obtain a target face image includes:
[0052] Perform privacy protection processing on the original image according to the perturbed image through the image processing model to obtain a first face image; perform privacy protection processing on the real-time captured image according to the perturbed image through the image processing model to obtain a second face image;
[0053] The method further includes:
[0054] Determine the image similarity between the first face image and the second face image, and perform face recognition on the real-time captured image according to the image similarity.
[0055] According to a second aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, including:
[0056] An acquisition module, configured to acquire a perturbed image and a face image to be processed, where the perturbed image is used to add perturbation to the face image;
[0057] An image processing module, configured to process the face image according to the perturbed image through a pre-trained image processing model to obtain a target face image;
[0058] Among them, the image processing model is a model pre-trained according to multiple face sample images, text description information respectively corresponding to each of the face sample images, and original face images respectively corresponding to each of the face sample images. The face sample images are face images that have undergone a preset modification process. The text description information is used to describe the modification features corresponding to the preset modification process, and the text description information is used to guide the perturbation image to add perturbations to the modification features.
[0059] According to a third aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, including:
[0060] A processor;
[0061] A memory for storing instructions executable by the processor;
[0062] Among them, the processor is configured to: execute the steps of the method described in the first aspect of the present disclosure.
[0063] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the image processing method provided in the first aspect of the present disclosure are implemented.
[0064] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: The face image is processed by the image processing model pre-trained according to the perturbation image to obtain a target face image. Since during the training process of this image processing model, the text description information describing the modification features (such as makeup processing) corresponding to each face sample image is used to guide the perturbation image to add adversarial perturbations to the modification features, therefore, this image processing model can effectively hide the adversarial perturbations in the makeup style of the face image, thereby avoiding the artifact problem caused by directly adding noise to the input image, providing high-quality images, preserving the identity perceived by humans, enhancing the user's visual experience, and at the same time achieving the effect of face privacy protection.
[0065] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0067] Figure 1 is a flowchart of an image processing method shown according to an exemplary embodiment.
[0068] Figure 2 is according toFigure 1 Flowchart of an image processing method shown in the illustrated embodiment.
[0069] Figure 3 Flowchart of a training method for an image processing model shown in accordance with an exemplary embodiment.
[0070] Figure 4 Schematic diagram of a model training process shown in accordance with an exemplary embodiment.
[0071] Figure 5 is based on Figure 1 Flowchart of an image processing method shown in the illustrated embodiment.
[0072] Figure 6 Block diagram of an image processing apparatus shown in accordance with an exemplary embodiment.
[0073] Figure 7 is based on Figure 6 Block diagram of another image processing apparatus shown in the illustrated embodiment.
[0074] Figure 8 is based on Figure 6 Block diagram of another image processing apparatus shown in the illustrated embodiment.
[0075] Figure 9 Block diagram of an apparatus for image processing shown in accordance with an exemplary embodiment. Detailed implementation manners
[0076] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0077] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining authorization from the owner of the corresponding apparatus.
[0078] The present disclosure is mainly applied to the face recognition scenario, and performs privacy protection processing on the face images collected in real time to prevent malicious systems from reusing the face images. The face recognition scenario may include, for example: face recognition payment, face recognition identity authentication, face beautification, face recognition access control, and face recognition monitoring and other application scenarios.
[0079] In the above face recognition scenario, if face information is leaked, it may pose a threat to the privacy of users. It may also violate personal rights. For example, in some recognition scenarios for monitoring, tracking, etc., personal rights may be violated. At the same time, there are also data security issues. Face recognition technology requires a large amount of data for training. If this data is leaked, it may pose a threat to the privacy of users.
[0080] The following several privacy protection methods are mainly provided in the related technologies:
[0081] 1. Encryption-based face recognition technology. This technology encrypts the face image so that only authorized users can decrypt and view the face image. This method can protect privacy, but requires high-strength encryption algorithms and key management.
[0082] 2. Privacy protection-based face recognition technology. This technology processes the face image, such as blurring, occlusion, replacement, etc., to protect personal privacy. By blurring the face image and adding occlusions, such as black squares or mosaics, to the face, the face cannot be recognized. This method can protect privacy, but it will affect the usability of the image.
[0083] 3. Differential privacy-based face recognition technology. This technology adds noise to the face image to protect personal privacy. This method can protect privacy, but it will affect the accuracy of face recognition.
[0084] 4. Multi-party secure computation-based face recognition technology. This technology divides the face image into multiple parts and processes them by different computing nodes respectively to protect personal privacy. This method can protect privacy, but requires efficient multi-party secure computation algorithms and computing resources.
[0085] In addition, an ideal facial privacy protection algorithm needs to strike an appropriate balance between naturality and privacy protection. "Naturality" refers to retaining the identity perceived by humans. "Privacy protection" means that the protected image can deceive a malicious face recognition black box system. Generally speaking, the protected image needs to be similar to the given facial image, and is artifact-free and noise-free for people, and can deceive unknown automated face recognition systems.
[0086] Currently, although deep learning-based face recognition systems are very successful and have numerous practical application scenarios, they have caused serious privacy problems because they can track users without authorization or permission. Traditional deep learning techniques such as adversarial generation methods hide user identities by superimposing noise on the original face images to constrain adversarial perturbations. Since adversarial examples are usually optimized in the image space and the adversarial nature is too strong to hinder the performance of the makeup transfer module, it results in unnatural faces with makeup artifacts. Therefore, existing privacy-enhancing methods cannot generate "natural" images that can protect facial privacy without affecting the user experience.
[0087] In addition, most traditional methods aim to imitate the target identity. For each new target identity, traditional methods need to perform end-to-end re-training from scratch using a large makeup dataset. Therefore, the training cost of the model is high, and its practicality also needs to be improved.
[0088] To solve the above problems, the present disclosure provides an image processing method, apparatus, and storage medium. The following will describe the specific embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0089] Figure 1 is a flowchart of an image processing method shown according to an exemplary embodiment. This method can be applied to a server or a terminal. As Figure 1 shown, the method includes the following steps.
[0090] In step S11, a perturbed image and a face image to be processed are obtained, where the perturbed image is used to add perturbations to the face image.
[0091] Among them, adding perturbations to the face image in the present disclosure means making some changes to the face image to achieve the purpose of enhancing privacy protection. The perturbed image refers to the image that the face image needs to imitate in order to achieve privacy protection for the face image. In other words, some image features of the perturbed image can be added as perturbations to the face image to achieve privacy protection for the face image.
[0092] Generally, the perturbed image can be synthesized by an artificial intelligence model, and the synthesized perturbed image can be randomly paired with the face image to be processed for use.
[0093] In addition, the face image to be processed can be a face image collected in real time in the current face recognition task. To avoid the malicious system from reusing the face image collected in real time, the present disclosure can process the face image according to the perturbed image through a pre-trained image processing model to obtain a target face image, and perturbations are added to the target face image to achieve facial privacy protection for the face image.
[0094] In step S12, the face image is processed by an image processing model pre-trained according to the perturbed image to obtain a target face image. Wherein, the image processing model is a model pre-trained according to multiple face sample images, the text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image. The face sample image is a face image that has undergone a preset modification process. The text description information is used to describe the modification features corresponding to the preset modification process, and the text description information is used to guide the perturbed image to add perturbations in the modification features.
[0095] In this step, the face image is subjected to privacy protection processing through the image processing model according to the perturbed image. Here, the privacy protection processing may include image scrambling processing and / or image confusion processing. The target face image after privacy protection processing is only applicable to the current face recognition task. After an unauthorized face recognition system obtains the target face image, it cannot decrypt it and thus cannot be reused, thereby realizing the privacy protection of the face image.
[0096] Among them, the image processing model may include an encoding network, a generative adversarial network, and a perturbation network. The encoding network may include a pre-trained encoder, for example. The generative adversarial network may include a GAN (Generative Adversarial Network). The perturbation network may be a deep learning network built based on actual requirements. In this way, the encoding network can map the input face image to the latent space of the GAN, thereby obtaining a vector representation of the facial features characterizing the face image. The generative adversarial network can generate high-quality new face image data based on the facial features mapped in the latent space. The perturbation network can output perturbation image features based on the new face data and the perturbed image. Then, the perturbation image features and the facial features mapped in the latent space can be input into the generative adversarial network so that the generative adversarial network can generate a target face image with added perturbations, thereby achieving the purpose of privacy protection of the face image.
[0097] In the present disclosure, the image processing model is a model pre-trained according to multiple face sample images, the text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image. The face sample image is a face image that has undergone a preset modification process. Here, the preset modification process includes applying makeup to the face image, such as adding effects like thick eyebrows, red lips, and big eyes, and describing the modification features (such as thick eyebrows, red lips, big eyes, etc.) corresponding to the preset modification process through the text description information.
[0098] It should be noted that during the process of pre-training the image processing model, the text description information corresponding to each face sample image can be used to guide the perturbed image to add perturbations to the modification features corresponding to each image. In this way, the adversarial perturbation can be effectively hidden in the makeup style, thereby avoiding the artifact problem caused by directly adding noise to the original image, providing high-quality images, retaining the identity perceived by humans, enhancing the user's visual experience, and at the same time achieving the effect of face privacy protection.
[0099] Using the above method, the face image is processed by the image processing model pre-trained according to the perturbed image to obtain a target face image. Since during the training process of this image processing model, the text description information describing the modification features corresponding to each face sample image is used to guide the perturbed image to add the adversarial perturbation to the modification features, this image processing model can effectively hide the adversarial perturbation in the makeup style of the face image, thereby avoiding the artifact problem caused by directly adding noise to the input image, providing high-quality images, retaining the identity perceived by humans, enhancing the user's visual experience, and at the same time achieving the effect of face privacy protection.
[0100] As described above, the image processing model includes an encoding network, a generative adversarial network, and a perturbation network. Among them, the input of the encoding network is the face image to be processed, and the output of the encoding network is connected to the input of the generative adversarial network. The first output of the generative adversarial network (the first output is generated by the generative adversarial network based on the facial features of the face image mapped to the latent space, and is a new face image data simulating the face image, that is, the "simulated face features" described below) can be used as the input of the perturbation network, and the output of the perturbation network can be used as the input of the generative adversarial network. In this way, the generative adversarial network can obtain the target face image after privacy protection processing based on the output of the encoding network and the output of the perturbation network.
[0101] Figure 2 is according to Figure 1 shown in the flowchart of an image processing method according to the illustrated embodiment. As Figure 2 shown, step S12 includes the following sub-steps:
[0102] In step S121, the face image to be processed is input into the encoding network to obtain a first latent feature, and the first latent feature characterizes the facial features of the face image to be processed.
[0103] By performing this step, the face image can be mapped into the latent space of the generative adversarial network through the encoding network. The latent space in the generative adversarial network refers to the random noise space in the generator model, which is used to generate random vectors as the input of the model. This latent space is the input space of the generator. The generator can take this vector as the input and learn to map from this latent space to the data space to generate simulated image data (i.e., simulated face features) close to the input face image.
[0104] Among them, the first latent feature output by the encoding network is the vector representation that characterizes the facial features of the face image after mapping the face image into the latent space of the generative adversarial network.
[0105] In step S122, according to the first latent feature and the perturbation image, the second latent feature is determined through the generative adversarial network and the perturbation network. The second latent feature characterizes the perturbation image features corresponding to the face image.
[0106] The perturbation image features here refer to the perturbation features that need to be added to the face image for privacy protection of the face image.
[0107] In this step, the first latent feature can be input into the generative adversarial network to obtain the simulated face features corresponding to the face image; the simulated face features and the perturbation image are input into the perturbation network to obtain the second latent feature.
[0108] In step S123, according to the first latent feature and the second latent feature, the target face image is determined through the generative adversarial network.
[0109] In this step, the first latent feature and the second latent feature can be merged to obtain the target image features; after the target image features are input into the generative adversarial network, the target face image is output through the generative adversarial network.
[0110] It should be noted that during the training process of the image processing model, through the text description information describing the modification features corresponding to each face sample image, the perturbation image is guided to add adversarial perturbations to the modification features. Therefore, this image processing model can effectively hide the adversarial perturbations in the makeup style of the face image, thus avoiding the artifact problem caused by directly adding noise to the input image. Therefore, in this disclosure, by performing steps S121 - S123, the adversarial perturbations have been added to the makeup style of the input face image in the output target face image, so that the target face image after privacy protection can retain the identity perceived by humans and will not affect the user's visual experience.
[0111] That is to say, the method for protecting the privacy of image processing of face images provided by the present disclosure can make the protected face image (i.e., the target face image) highly similar to the original face image visually, and be artifact-free and noise-free for people. At the same time, it can also deceive unknown automated face recognition systems.
[0112] Figure 3 It is a flowchart of a method for training an image processing model shown according to an exemplary embodiment, as Figure 3 shown, and the method includes the following steps:
[0113] In step S301, a preset initial model, training samples, and random perturbation images for model training are obtained. The preset initial model includes an encoding network, a preset generation model to be trained, and a preset perturbation model to be trained. The training samples include multiple face sample images, text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image.
[0114] Among them, the original face image refers to a face image that has not undergone preset modification processing (such as adding makeup effects or beautification processing, etc.). The face sample image refers to a face image obtained by performing preset modification processing on the original face image. In the present disclosure, the modification features corresponding to each face sample image can be described by the text description information.
[0115] In a possible implementation manner, for each face sample image, the face sample image can be input into a pre-trained vision-language model, and then the text description information corresponding to the face sample image can be output through the pre-trained vision-language model. Among them, the pre-trained vision-language model can include, for example, CLIP (Contrastive Language-Image Pretraining). Among them, the CLIP model can understand images and texts simultaneously, establish connections between images and texts, and can perform cross-modal search and representation learning.
[0116] In addition, the encoding network included in the preset initial model belongs to a pre-trained encoder. In the present disclosure, there is no need to perform model training on the encoding network. When the image processing model is pre-trained in the present disclosure, only the preset generation model and the preset perturbation model need to be respectively trained. Among them, the preset generation model can include a GAN model. In an actual model training scenario, a GAN model trained on a high-resolution face image dataset can be used as the preset generation model.
[0117] The preset perturbation model may include a deep learning model built according to actual needs. For example, the structure of the preset perturbation model may include two two-dimensional convolutional layers + one single-layer LSTM (Long Short-Term Memory) + four two-dimensional convolutional layers. Among them, a prely activation function may be interspersed between the two-dimensional convolutional layers, or a residual structure may exist. Here, the model structure of the preset perturbation model is only an example, and the present disclosure does not limit this.
[0118] In step S302, based on multiple face sample images, the text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image, the preset generation model is trained to obtain a generative adversarial network.
[0119] In this step, for each of the face sample images, after inputting the face sample image into the encoding network, a low-dimensional image feature can be output through the encoding network. The low-dimensional image feature is the image feature after mapping the face sample image to the latent space, and the latent space represents the low-dimensional vector space on the input side of the generative adversarial network. After inputting the low-dimensional image feature into the preset generation model, a model output corresponding to the face sample image is obtained. The model output includes an output image and output text information. Among them, the output text information is the text description information of the output image. Then, after training the preset generation model according to the output image, the output text information, the original face image corresponding to the face sample image, and the text description information corresponding to the face sample image, the generative adversarial network is obtained.
[0120] In one implementation, the image loss of the model output can be determined according to the output image and the original face image corresponding to the face sample image; the text loss of the model output can be determined according to the output text information and the text description information corresponding to the face sample image; a first loss function can be determined according to the image loss and the text loss; after training the preset generation model according to the first loss function, the generative adversarial network is obtained.
[0121] Exemplarily, Figure 4 is a schematic diagram of a model training process shown according to an exemplary embodiment. As Figure 4 shown, IΦ in the figure represents the encoding network, Gθ represents the preset generation model to be trained, Gθ′ represents the trained generative adversarial network, and f represents the preset perturbation model to be trained. The model training process of the present disclosure may include two stages, namely Figure 4 the stage one and stage two in Figure 4In the model training process of the first stage, by executing steps S303 - S304, the following can be achieved Figure 4 The model training process of the second stage.
[0122] First, the training process of the first stage will be described. The first stage is the latent space feature initialization stage. Its essence is based on the feature mapping principle of GAN, mapping the face sample image x to the latent space W, that is, finding a latent space wi ∈ W such that the output image xi = Gθ(wi) ≈ x. To achieve this operation, as Figure 4 shown, for each of the face sample images x, the face sample image x can be input into the encoding network IΦ, and then the low - dimensional image feature wi is output through the encoding network IΦ. The low - dimensional image feature wi is the image feature after mapping the face sample image x to the latent space W, that is, wi ∈ W, and wi = IΦ(x). Then, the low - dimensional image feature wi can be input into the preset generation model Gθ to be trained, and xi = G θ (wi)1 ≈ x is obtained, and based on the preset generation model Gθ, the output text information G θ (wi)2 is obtained. In this way, the image loss L1 output by the model can be determined according to the output image G θ (wi)1 and the original face image m corresponding to the face sample image x; the text loss L2 output by the model can be determined according to the output text information G θ (wi)2 and the text description information n corresponding to the face sample image x; the first loss function L is determined according to the image loss L1 and the text loss L2, and its formula can be expressed as:
[0123] L = α·L1(m, G θ (wi)1) + L2(n, G θ (wi)2) (1)
[0124] where α represents the adjustment coefficient of the image loss L1.
[0125] In this way, based on formula (1), the preset generation model Gθ can be trained, and the fine - tuned model parameter θ′ can be expressed as:
[0126] θ′ = argmin θ [α·L1(m, G θ (wi)1) + L2(n, G θ (wi)2)] (2).
[0127] In this way, the trained generative adversarial network Gθ′ can be trained based on the fine - tuned model parameter θ′. The above examples are only for illustration, and the present disclosure is not limited thereto.
[0128] It should be noted that in the process of mapping the face sample image x to the latent space W based on the encoding network IΦ, the identity of the face sample image x needs to be retained, which can be expressed as N(x) = N(xi).
[0129] In step S303, the encoding network, the generative adversarial network, and the preset perturbation model to be trained are combined to form a new image processing model to be trained.
[0130] In step S304, after training the preset perturbation model through the new image processing model to be trained according to the training samples and the random perturbation image, the image processing model is obtained.
[0131] Among them, the random perturbation image is a face image randomly generated by a preset artificial intelligence model. In the present disclosure, the random perturbation image is used to add perturbations to the face sample image during the model training process.
[0132] In this step, for each face sample image, after inputting the face sample image into the image processing model to be trained, the low-dimensional image features output by the encoding network and the model output of the generative adversarial network can be obtained. The model output includes an output image and output text information. After inputting the output image, the output text information, and the random perturbation image into the preset perturbation model, the perturbation image features guided by the output text information are obtained. After training the preset perturbation model according to the face sample image, the random perturbation image, the low-dimensional image features, and the perturbation image features, the trained perturbation network is obtained. The encoding network, the generative adversarial network, and the perturbation network are combined to obtain the image processing model.
[0133] Among them, in the process of obtaining the trained perturbation network after training the preset perturbation model according to the face sample image, the random perturbation image, the low-dimensional image features, and the perturbation image features, the low-dimensional image features and the perturbation image features can be merged and then input into the generative adversarial network to obtain the model prediction image corresponding to the face sample image. The face adversarial loss is determined according to the face sample image, the random perturbation image, and the model prediction image. The face identity retention loss is determined according to the low-dimensional image features, the perturbation image features, and the model prediction image. The second loss function is determined according to the face identity retention loss and the face adversarial loss. After training the preset perturbation model according to the second loss function, the trained perturbation network is obtained.
[0134] In one implementation, the face adversarial loss can be determined according to the face sample image, the random perturbation image, and the model prediction image in the following manner:
[0135] Determine a first distance metric between the model-predicted image and the face sample image, and a second distance metric between the model-predicted image and the randomly perturbed image, where the smaller the distance metric, the higher the probability that the two face images corresponding to the model prediction belong to the same identity; use the difference between the second distance metric and the first distance metric as the face adversarial loss.
[0136] In one implementation, the face identity preservation loss can be determined based on the low-dimensional image feature, the perturbed image feature, and the model-predicted image in the following manner:
[0137] Element-wise multiply the low-dimensional image feature and the model-predicted image to obtain a first product; element-wise multiply the perturbed image feature and the model-predicted image to obtain a second product. The larger the product, the higher the probability that the face identity corresponding to the image feature is the same as the face identity of the face sample image; determine the face identity preservation loss based on the difference between the second product and the first product.
[0138] Exemplarily, continue with Figure 4 as an example to illustrate the model training process in stage two. In the present disclosure, based on stage two, adversarial optimization based on text guidance (i.e., text description information) can be achieved. Based on stage one, the latent space feature wi and the trained generative adversarial network Gθ′ (or referred to as the "generator") can be obtained. The main objective of the present disclosure is to generate the image compression feature wi of the latent space through the encoding network, and at the same time add the perturbed image xt to generate a protected face, so as to achieve privacy protection of the face image and retain the facial features of the original person's image. In addition, by modeling the text content, facial features imitating the makeup style of the text prompt are established.
[0139] First, in order to effectively extract the makeup style information in the face sample image x from the output text information and apply it to the face sample image x in an adversarial manner, the output text information output in stage one can be encoded by a pre-trained vision-language model, and the feature space of the encoded text is sent into the preset perturbation model f to be trained. In addition, the input of the preset perturbation model f also needs to include the output image output by the generative adversarial network Gθ′ (this output image is used to simulate face feature information) as the basic feature information of the text and the perturbation.
[0140] To ensure that the quality of the output face image is not damaged, it is solved by forcing the adversarial latent features to remain close to the initialized latent feature wi. The implementation process is as follows: During the model training process, the perturbed image feature w output by the preset interference network f is combined with the latent feature wi initialized in the first stage as the merged latent feature and fed into the generative adversarial network Gθ′. During the model training process, the output of f for the current stage will be fed into wi for the next stage of model training in an iterative manner; the purpose of doing this is to utilize the feature compression property in the latent space of the generative adversarial network Gθ′ to retain the facial features of the original face image, and perform text guidance based on the output text information fed into f in the previous training round, so as to hide the adversarial perturbation into the makeup effect in a text way for the feature information that cannot be emphasized in the image.
[0141] Among them, for the face adversarial loss during the model training process, the model prediction image G θ′ (w) has a face feature representation close to the randomly perturbed image xt and far from the face sample image x itself, that is, F(xp,x)>F(xp,xt), where F(x1,x2) = -|cos(f(x1)-f(x2))| is the cosine distance, representing the similarity of the feature representations of two images x1 and x2. Therefore, in the present disclosure, this face adversarial loss can be expressed as:
[0142] L ad =F(G θ' (w),x t )-F(G θ' (w),x) (3)
[0143] Among them, L ad represents the face adversarial loss, G θ′ (w) represents the model prediction image, x t represents the randomly perturbed image, and x represents the face sample image.
[0144] For the face identity preservation loss, in order to make the face image processed by the model artifact-free and noise-free (i.e., an image that is more natural in visual experience), the features in its latent space can be controlled using the GAN model. Since the synthesized latent features can affect image generation based on text guidance, during model training, only the latent feature representations related to the deeper layers in the GAN model can be perturbed, so as to limit the adversarial faces to the identity preservation features. In the present disclosure, the following face identity preservation loss L gan can be used to further limit the perturbed image feature w to be close to the initialized latent feature wi:
[0145] L gan =||(w⊙p i )-(wi ⊙p i )||2 (4)
[0146] Among them, ⊙ represents element-wise multiplication, pi is the model-predicted image, w represents the perturbed image feature, and wi represents the initialized latent feature.
[0147] In this way, according to the face identity retention loss L gan and the face adversarial loss L ad The determined second loss function can be expressed as:
[0148] L all = a·L ad + b·L gan (5)
[0149] Among them, L all represents the second loss function, and a and b are hyperparameters.
[0150] In this way, based on the second loss function L shown in formula (5) all After training the preset perturbation model f, a perturbation network can be obtained. After connecting the encoding network, the generative adversarial network, and the perturbation network, a trained image processing model can be obtained.
[0151] The above examples are only for illustration, and the present disclosure is not limited thereto.
[0152] In an embodiment of the present disclosure, the face image to be processed includes the original image corresponding to the target face recognition task and the real-time captured image. For example, the target face recognition task may include tasks such as face recognition for payment and face recognition for access control. The original image generally refers to the face image entered into the system by the user when first using the target face recognition task, and this face image can be stored in the database so that the real-time captured face image (i.e., the real-time captured image) can be compared with the original image stored in the database to complete the current face recognition task.
[0153] Figure 5 is a flowchart of an image processing method shown in the embodiment according to Figure 1 As shown in Figure 5 In step S12, the original image can be subjected to privacy protection processing through the image processing model according to the perturbed image to obtain a first face image; the real-time captured image can be subjected to privacy protection processing through the image processing model according to the perturbed image to obtain a second face image.
[0154] As Figure 5 shown, the method further includes the following steps:
[0155] In step S13, the image similarity between the first face image and the second face image is determined, and face recognition is performed on the real-time captured image according to the image similarity.
[0156] That is to say, based on the image processing method provided by the present disclosure, in an actual face recognition task, both the original image and the real-time captured image can be subjected to privacy protection processing according to the image processing method provided by the present disclosure. Then, after image comparison based on the two sets of processed image data (the image data of the first face image and the image data of the second face image), the current target face recognition task is completed. At the same time, the real-time captured image has undergone privacy protection processing, which can prevent unauthorized malicious systems from reusing the real-time captured image. In addition, based on the image processing method provided by the present disclosure, while protecting the privacy of the face image, the processed face image has no artifacts and noise visually, thus also providing a better visual experience for users.
[0157] Figure 6 is a block diagram of an image processing device shown according to an exemplary embodiment, as Figure 6 shown, the device includes:
[0158] An acquisition module 601, configured to acquire a perturbation image and a face image to be processed, where the perturbation image is used to add perturbations to the face image;
[0159] An image processing module 602, configured to process the face image according to the perturbation image through a pre-trained image processing model to obtain a target face image;
[0160] Wherein, the image processing model is a model pre-trained according to multiple face sample images, text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image. The face sample image is a face image subjected to a preset modification process, and the text description information is used to describe the modification features corresponding to the preset modification process, and the text description information is used to guide the perturbation image to add perturbations in the modification features.
[0161] Optionally, the image processing model includes an encoding network, a generative adversarial network, and a perturbation network;
[0162] The image processing module 602 is configured to input the face image into the encoding network to obtain a first latent feature, where the first latent feature characterizes the facial features of the face image; determine a second latent feature through the generative adversarial network and the perturbation network according to the first latent feature and the perturbation image, where the second latent feature characterizes the perturbation image features corresponding to the face image; and determine the target face image through the generative adversarial network according to the first latent feature and the second latent feature.
[0163] Optionally, the image processing module 602 is configured to input the first latent feature into the generative adversarial network to obtain the simulated face features corresponding to the face image; and input the simulated face features and the perturbation image into the perturbation network to obtain the second latent feature.
[0164] Optionally, the image processing module 602 is configured to perform feature merging on the first latent feature and the second latent feature to obtain target image features; and after inputting the target image features into the generative adversarial network, output the target face image through the generative adversarial network.
[0165] Optionally, Figure 7 is Figure 6 a block diagram of another image processing device shown in the embodiment shown in Figure 7 As shown, the device further includes:
[0166] A model training module 603, configured to pre-train the image processing model in the following manner:
[0167] Obtain a preset initial model, training samples, and random perturbation images for model training. The preset initial model includes an encoding network, a preset generative model to be trained, and a preset perturbation model to be trained. The training samples include multiple face sample images, text description information respectively corresponding to each face sample image, and original face images respectively corresponding to each face sample image;
[0168] Perform model training on the preset generative model according to the multiple face sample images, the text description information respectively corresponding to each face sample image, and the original face images respectively corresponding to each face sample image to obtain a generative adversarial network;
[0169] Form a new image processing model to be trained by combining the encoding network, the generative adversarial network, and the preset perturbation model to be trained;
[0170] After performing model training on the preset perturbation model through the new image processing model to be trained according to the training samples and the random perturbation images, obtain the image processing model.
[0171] Optionally, the model training module 603 is configured to, for each of the face sample images, after inputting the face sample image into the encoding network, output low-dimensional image features through the encoding network, where the low-dimensional image features are image features after mapping the face sample image to a latent space, and the latent space represents a low-dimensional vector space on the input side of the generative adversarial network; after inputting the low-dimensional image features into the preset generative model, obtain a model output corresponding to the face sample image, where the model output includes an output image and output text information; and after training the preset generative model according to the output image, the output text information, the original face image corresponding to the face sample image, and the text description information corresponding to the face sample image, obtain the generative adversarial network.
[0172] Optionally, the model training module 603 is configured to determine an image loss of the model output according to the output image and the original face image corresponding to the face sample image; determine a text loss of the model output according to the output text information and the text description information corresponding to the face sample image; determine a first loss function according to the image loss and the text loss; and after training the preset generative model according to the first loss function, obtain the generative adversarial network.
[0173] Optionally, the model training module 603 is configured to, for each of the face sample images, after inputting the face sample image into the image processing model to be trained, obtain the low-dimensional image features output by the encoding network and the model output of the generative adversarial network, where the model output includes an output image and output text information;
[0174] After inputting the output image, the output text information, and the random perturbation image into the preset perturbation model, obtain perturbation image features guided by the output text information;
[0175] After training the preset perturbation model according to the face sample image, the random perturbation image, the low-dimensional image features, and the perturbation image features, obtain a trained perturbation network;
[0176] Combine the encoding network, the generative adversarial network, and the perturbation network to obtain the image processing model.
[0177] Optionally, the model training module 603 is configured to merge the low-dimensional image feature and the perturbed image feature, input the merged feature into the generative adversarial network to obtain a model prediction image corresponding to the face sample image; determine a face adversarial loss based on the face sample image, the random perturbation image, and the model prediction image; determine a face identity preservation loss based on the low-dimensional image feature, the perturbed image feature, and the model prediction image; determine a second loss function based on the face identity preservation loss and the face adversarial loss; and train the preset perturbation model according to the second loss function to obtain a trained perturbation network.
[0178] Optionally, the model training module 603 is configured to determine a first distance metric between the model prediction image and the face sample image, and a second distance metric between the model prediction image and the random perturbation image, where the smaller the distance metric, the higher the probability that the two face images corresponding to the model prediction belong to the same identity; and use the difference between the second distance metric and the first distance metric as the face adversarial loss.
[0179] Optionally, the model training module 603 is configured to perform element-wise multiplication on the low-dimensional image feature and the model prediction image to obtain a first product; perform element-wise multiplication on the perturbed image feature and the model prediction image to obtain a second product, where the larger the product, the higher the probability that the face identity corresponding to the image feature is the same as the face identity of the face sample image; and determine the face identity preservation loss according to the difference between the second product and the first product.
[0180] Optionally, the face image includes an original image and a real-time acquisition image corresponding to a target face recognition task; the image processing module 602 is configured to perform privacy protection processing on the original image through the image processing model according to the perturbed image to obtain a first face image; and perform privacy protection processing on the real-time acquisition image through the image processing model according to the perturbed image to obtain a second face image. Figure 8 is a block diagram of another image processing device shown in the Figure 6 illustrated embodiment, as Figure 8 shown, the device further includes:
[0181] A face recognition module 604, configured to determine an image similarity between the first face image and the second face image, and perform face recognition on the real-time acquisition image according to the image similarity.
[0182] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0183] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the image processing method provided by the present disclosure are implemented.
[0184] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-mentioned image processing method when executed by the programmable device.
[0185] Figure 9 is a block diagram of a device for image processing shown according to an exemplary embodiment. For example, device 900 may be provided as a server. Referring to Figure 9 , device 900 includes a processing component 922, which further includes one or more processors, and memory resources represented by a memory 932 for storing instructions executable by the processing component 922, such as application programs. The application programs stored in the memory 932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 922 is configured to execute instructions to perform the above-mentioned image processing method.
[0186] Device 900 may further include a power component 926 configured to perform power management of device 900, a wired or wireless network interface 950 configured to connect device 900 to a network, and an input / output interface 958. Device 900 may operate based on an operating system stored in the memory 932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0187] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0188] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An image processing method, characterized in that, Including: Obtaining a perturbed image and a face image to be processed, where the perturbed image is used to add perturbations to the face image; Processing the face image according to the perturbed image through a pre-trained image processing model to obtain a target face image; Wherein, the image processing model is a model pre-trained according to multiple face sample images, text description information respectively corresponding to each of the face sample images, and original face images respectively corresponding to each of the face sample images. The face sample images are face images that have undergone a preset modification process, and the text description information is used to describe the modification features corresponding to the preset modification process, and the text description information is used to guide the perturbed image to add perturbations in the modification features.
2. The method according to claim 1, wherein The image processing model includes an encoding network, a generative adversarial network, and a perturbation network; The processing the face image according to the perturbed image through a pre-trained image processing model to obtain a target face image includes: Inputting the face image into the encoding network to obtain a first latent feature, where the first latent feature represents the facial features of the face image; Determining a second latent feature according to the first latent feature and the perturbed image through the generative adversarial network and the perturbation network, where the second latent feature represents the perturbed image features corresponding to the face image; Determining the target face image according to the first latent feature and the second latent feature through the generative adversarial network.
3. The method according to claim 2, characterized in that The determining the second latent feature according to the first latent feature and the perturbed image through the generative adversarial network and the perturbation network includes: Inputting the first latent feature into the generative adversarial network to obtain the simulated face features corresponding to the face image; Inputting the simulated face features and the perturbed image into the perturbation network to obtain the second latent feature.
4. The method according to claim 2, wherein The determining the target face image according to the first latent feature and the second latent feature through the generative adversarial network includes: Performing feature merging on the first latent feature and the second latent feature to obtain target image features; After inputting the target image features into the generative adversarial network, outputting the target face image through the generative adversarial network.
5. The method according to any one of claims 1-4, characterized in that The image processing model is pre-trained in the following manner: Obtaining a preset initial model, training samples, and a random perturbed image for model training. The preset initial model includes an encoding network, a preset generative model to be trained, and a preset perturbation model to be trained. The training samples include multiple face sample images, text description information respectively corresponding to each of the face sample images, and original face images respectively corresponding to each of the face sample images; Performing model training on the preset generative model according to the multiple face sample images, text description information respectively corresponding to each of the face sample images, and original face images respectively corresponding to each of the face sample images to obtain a generative adversarial network; Forming a new image processing model to be trained by combining the encoding network, the generative adversarial network, and the preset perturbation model to be trained; After training the preset perturbation model through the new image processing model to be trained based on the training samples and the random perturbation images, the image processing model is obtained.
6. The method according to claim 5, wherein The training of the preset generation model based on the multiple face sample images, the text description information respectively corresponding to each face sample image, and the original face image respectively corresponding to each face sample image to obtain the generative adversarial network includes: For each face sample image, after inputting the face sample image into the encoding network, the encoding network outputs a low-dimensional image feature, where the low-dimensional image feature is the image feature after mapping the face sample image to the latent space, and the latent space represents the low-dimensional vector space on the input side of the generative adversarial network; After inputting the low-dimensional image feature into the preset generation model, the model output corresponding to the face sample image is obtained, and the model output includes an output image and output text information; After training the preset generation model based on the output image, the output text information, the original face image corresponding to the face sample image, and the text description information corresponding to the face sample image, the generative adversarial network is obtained.
7. The method according to claim 6, characterized in that, The training of the preset generation model based on the output image, the output text information, the original face image corresponding to the face sample image, and the text description information corresponding to the face sample image to obtain the generative adversarial network includes: Determining the image loss of the model output based on the output image and the original face image corresponding to the face sample image; Determining the text loss of the model output based on the output text information and the text description information corresponding to the face sample image; Determining a first loss function based on the image loss and the text loss; After training the preset generation model based on the first loss function, the generative adversarial network is obtained.
8. The method according to claim 5, wherein The training of the preset perturbation model through the new image processing model to be trained based on the training samples and the random perturbation images to obtain the image processing model includes: For each face sample image, after inputting the face sample image into the image processing model to be trained, the low-dimensional image feature output by the encoding network and the model output of the generative adversarial network are obtained, and the model output includes an output image and output text information; After inputting the output image, the output text information, and the random perturbation image into the preset perturbation model, a perturbation image feature guided by the output text information is obtained; After training the preset perturbation model based on the face sample image, the random perturbation image, the low-dimensional image feature, and the perturbation image feature, a trained perturbation network is obtained; Combining the encoding network, the generative adversarial network, and the perturbation network to obtain the image processing model.
9. The method according to claim 8, wherein After training the preset perturbation model according to the face sample image, the random perturbation image, the low-dimensional image feature, and the perturbation image feature, the trained perturbation network obtained includes: After merging the low-dimensional image feature and the perturbation image feature, input them into the generative adversarial network to obtain a model prediction image corresponding to the face sample image; Determine the face adversarial loss according to the face sample image, the random perturbation image, and the model prediction image; Determine the face identity preservation loss according to the low-dimensional image feature, the perturbation image feature, and the model prediction image; Determine the second loss function according to the face identity preservation loss and the face adversarial loss; After training the preset perturbation model according to the second loss function, obtain the trained perturbation network.
10. The method according to claim 9, wherein The determination of the face adversarial loss according to the face sample image, the random perturbation image, and the model prediction image includes: Determine a first distance metric between the model prediction image and the face sample image, and a second distance metric between the model prediction image and the random perturbation image, where the smaller the distance metric, the higher the probability that the two face images corresponding to the model prediction belong to the same identity; Use the difference between the second distance metric and the first distance metric as the face adversarial loss.
11. The method according to claim 9, characterized in that, The determination of the face identity preservation loss according to the low-dimensional image feature, the perturbation image feature, and the model prediction image includes: Perform element-wise multiplication on the low-dimensional image feature and the model prediction image to obtain a first product; Perform element-wise multiplication on the perturbation image feature and the model prediction image to obtain a second product. The larger the product, the higher the probability that the face identity corresponding to the image feature is the same as the face identity of the face sample image; Determine the face identity preservation loss according to the difference between the second product and the first product.
12. The method according to claim 1, characterized in that, The face image includes the original image and the real-time acquisition image corresponding to the target face recognition task; the processing of the face image by the image processing model pre-trained according to the perturbation image to obtain the target face image includes: Perform privacy protection processing on the original image through the image processing model according to the perturbation image to obtain a first face image; perform privacy protection processing on the real-time acquisition image through the image processing model according to the perturbation image to obtain a second face image; The method further includes: Determine the image similarity between the first face image and the second face image, and perform face recognition on the real-time acquisition image according to the image similarity.
13. An image processing apparatus, characterized in that, Includes: An acquisition module configured to acquire a perturbation image and a face image to be processed, where the perturbation image is used to add perturbations to the face image; An image processing module configured to process the face image through an image processing model pre-trained according to the perturbation image to obtain a target face image; Among them, the image processing model is a model pre-trained according to multiple face sample images, text description information respectively corresponding to each of the face sample images, and original face images respectively corresponding to each of the face sample images. The face sample images are face images that have undergone a preset modification process. The text description information is used to describe the modification features corresponding to the preset modification process, and the text description information is used to guide the addition of perturbations to the perturbation images in the modification features.
14. An image processing apparatus, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Among them, the processor is configured to: execute the steps of the method according to any one of claims 1-12.
15. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the method according to any one of claims 1-12 are implemented.