Face image processing method, related apparatus, and storage medium

CN117831089BActive Publication Date: 2026-09-11BEIJING REALAI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211191524.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2026-09-11
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

[0004]但是,现有生成对抗眼镜、对抗扰动贴纸和对抗帽子的方法往往是简单地使得包括对抗扰动的对抗图像与目标人脸图像向最相似的方向优化,优化目标单一

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117831089B_ABST
    Figure CN117831089B_ABST
Patent Text Reader

Abstract

The application relates to the field of computer vision, and provides a face image processing method, related devices and a storage medium. The method comprises the following steps: determining an initial face image; acquiring a candidate adversarial image; determining a target face image meeting a first preset condition from a preset face image set according to the initial face image and the candidate adversarial image; and generating a face adversarial sample based on the target face image and the initial face image. The face adversarial sample obtained in the embodiment of the application is optimized based on multiple different target faces, so that the attack success rate of the face adversarial sample is high, and the attack stability in the physical world is relatively strong.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and more specifically to face image processing methods, related devices, and storage media. Background Technology

[0002] The misuse of facial recognition systems based on artificial intelligence technology is becoming increasingly serious. For example, real estate companies use facial recognition systems to collect customers' purchasing intentions and offer higher prices to customers with strong intentions; shopping malls use facial recognition systems to obtain consumers' purchasing habits and promote products to people with regular purchasing habits. Such misuse has caused unprecedented concerns about personal privacy.

[0003] Currently, there are various solutions to the problem of facial recognition abuse. For example, wearing adversarial glasses or hats, or affixing adversarial stickers to the face. These adversarial entities can mislead or confuse facial recognition systems. For instance, user A wearing an adversarial entity could cause the facial recognition system to mistakenly identify them as user B. Thus, the facial recognition system cannot accurately identify the user, which can protect facial privacy.

[0004] However, existing methods for generating adversarial glasses, adversarial perturbation stickers, and adversarial hats often simply optimize the adversarial image, including the adversarial perturbation, to be as similar as possible to the target face image, resulting in a singular optimization objective. This means that users wearing the same adversarial perturbation entity are easily identified as the same target user. For example, if these users are entering and exiting a shopping mall's facial recognition system, the system will frequently detect the target user's entry and exit, making it easy for the system's defenses to detect. Ultimately, this leads to a low success rate for adversarial perturbation attacks and low stability of attacks in the physical world. Summary of the Invention

[0005] This application provides a face image processing method, related apparatus, and storage medium, which can optimize candidate adversarial perturbations and different target faces to be more similar during the iteration process, thereby obtaining target adversarial perturbations with high attack success rate and strong attack stability in the physical world.

[0006] In a first aspect, embodiments of this application provide a face image processing method, including: Determine the initial face image; Obtain candidate adversarial images, which are updated based on historical candidate adversarial images; Based on the initial face image and the candidate adversarial image, a target face image that meets the first preset condition is determined from a preset face image set; Adversarial face samples are generated based on the target face image and the initial face image.

[0007] Secondly, embodiments of this application provide an image processing apparatus, comprising: Input / output module, used to determine the initial face image; The processing module is used to acquire candidate adversarial images, which are updated based on historical candidate adversarial images; The processing module is further configured to determine, based on the initial face image and the candidate adversarial image, a target face image that meets a first preset condition from a preset face image set; and Adversarial face samples are generated based on the target face image and the initial face image.

[0008] Thirdly, embodiments of this application provide a processing apparatus, the processing apparatus comprising: At least one processor, memory, and input / output unit; The memory is used to store computer programs, and the processor is used to invoke the computer programs stored in the memory to execute the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect.

[0010] Compared to existing technologies, the face image processing method, related apparatus, and storage medium according to embodiments of this application first determine target face images that meet a first preset condition from a preset face image set based on an initial face image and candidate adversarial images. Then, face adversarial samples are generated based on the target face image and the initial face image. That is, the face adversarial samples generated in this embodiment are not optimized directly in the direction of increasing similarity with a single attacker or decreasing similarity with a single protected subject, as is done in existing technologies. Instead, they are optimized multiple times based on different face images in the preset face image set, in the direction of being most similar to multiple target faces. Therefore, the final face adversarial sample is not merely similar or dissimilar to a single image, but similar to multiple target images that are similar or dissimilar to the single image. In other words, the face adversarial sample can obtain more adversarial features for implementing adversarial attacks based on multiple target images, rather than just a single adversarial feature optimized from the single image. Therefore, the adversarial examples generated in this application can resemble multiple target images, possessing more adversarial features capable of enabling adversarial attacks. When facing different face recognition systems, different systems can acquire different effective adversarial features, resulting in the same attack outcome. In other words, the adversarial examples generated in this application exhibit strong attack robustness and strong transferability, producing stable attack effects against multiple different face recognition models. Thus, the adversarial attacks using these examples have high success rates and high stability in the physical world, not only providing good protection for facial privacy but also allowing for transferable testing of adversarial attacks against a wider range of face recognition models. Attached Figure Description

[0011] The objectives, features, and advantages of the embodiments of this application will become readily understood by referring to the accompanying drawings and the detailed description of the embodiments. Wherein: Figure 1 This is a schematic diagram of existing anti-glasses technology; Figure 2 This is a schematic diagram of existing technologies for combating stickers; Figure 3 A schematic diagram of the process for generating adversarial hats in existing technologies; Figure 4 This is a schematic diagram of a face image processing system provided in an embodiment of this application; Figure 5 A step diagram illustrating a face image processing method provided in an embodiment of this application; Figure 6 A flowchart illustrating a face image processing method provided in an embodiment of this application; Figure 7This is a schematic diagram of wearing a mask made using an anti-disturbance face image processing method provided in an embodiment of this application; Figure 8 This application provides a method for processing facial images to obtain a scene image of a mask made against perturbation being worn on the face of a target user and then recognized by a facial recognition system. Figure 9 A schematic diagram of an image processing apparatus provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of a processing device according to an embodiment of this application; Figure 11 This application provides a partial structural diagram of a mobile phone related to a terminal device. Figure 12 This is a schematic diagram of the structure of a server provided in an embodiment of this application.

[0012] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0013] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects (e.g., first similarity and second similarity represent different similarities, and so on), and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The division of modules in the embodiments of this application is merely a logical division; in actual applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not performed. Additionally, the shown or discussed mutual coupling or direct coupling or communication connection may be through some interface, and the indirect coupling or communication connection between modules may be electrical or other similar forms, none of which are limited in the embodiments of this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separate, may or may not be physical modules, or may be distributed among multiple circuit modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of this application.

[0014] This application proposes a face image processing method, related apparatus, and storage medium, applicable to image processing systems. These systems may include an image processing device and an image recognition device, which can be integrated or deployed separately. The image processing device is at least used to determine a target face image that meets a first preset condition from a preset set of face images based on an initial face image and candidate adversarial images, and to generate adversarial face samples based on the target face image and the initial face image. The image processing device can be an application that determines the target face image and generates adversarial face samples, or a server that has this application installed. The image recognition device can be a recognition program that acquires facial features from a face image to obtain facial features, such as a face recognition model. The image recognition device can also be a terminal device with a face recognition model deployed on it.

[0015] The solutions provided in this application involve technologies such as Artificial Intelligence (AI), Natural Language Processing (NLP), and Machine Learning (ML), which are specifically illustrated through the following embodiments: Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0016] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0017] Computer vision (CV) is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in tasks such as target recognition, tracking, and measurement, and further performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0018] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning (DL) typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.

[0019] The inventors discovered that existing technologies for generating adversarial perturbations often optimize the adversarial image, including the perturbation, towards the direction most similar to a single target face image, resulting in a singular optimization target. Therefore, adversarial perturbations generated by existing technologies may be interpreted differently by different face recognition systems, leading to varying attack results. In other words, adversarial perturbations generated by existing technologies lack robustness and have weak transfer attack performance, potentially failing to produce stable attack effects against multiple different face recognition models.

[0020] like Figure 1As shown, adversarial glasses utilize a generative model to generate adversarial perturbations in the shape of glasses, which are then printed onto the glasses. This allows an initial face wearing these glasses with adversarial perturbations to cause errors in the face recognition system's results. During training, the generative model is optimized using a corresponding objective function, aiming to make the adversarial image obtained after the initial person wears the glasses, with the generated adversarial perturbations printed on them, most similar to a single target face. This method simply optimizes the adversarial image, including the adversarial perturbations, towards the most similar direction to the single target face image. This single optimization objective results in a low attack success rate and low attack stability in the physical world. Furthermore, the small size of the adversarial glasses makes it impossible to conceal key facial information, further reducing the performance of adversarial attacks.

[0021] like Figure 2 As shown, the inventors' research also revealed that the adversarial sticker obtains the adversarial image obtained by pasting stickers at various positions on the initial face using a location search method. This image is then compared to the similarity of other real faces besides the initial face. The inventors determine the position and angle at which the sticker with the highest similarity is pasted on the initial face. Then, appropriate stickers are selected and pasted onto the initial face at the position and angle with the highest similarity. This causes the face recognition system to incorrectly identify the protected face as someone else. For example, in... Figure 2 In scenario a, the initial face is James Caan. After a sticker is pasted onto his face, the facial recognition system incorrectly identifies him as John Goold. Figure 2 In scenario b, the initial face is Josie Bissett. After a sticker is applied to her face, the facial recognition system incorrectly identifies her as Patricia Arquette. Figure 2 In scenario c, the initial face is Harrison Ford. After a sticker is applied to his face, the face recognition system incorrectly identifies him as Tom Daschle. The adversarial sticker itself is a natural pattern; the optimization process only changes its position and angle to cause the face recognition system to misidentify it. The pattern of the adversarial sticker itself does not contribute to the attack process, therefore the adversarial sticker lacks algorithmic adversarial capabilities and has a low success rate. Furthermore, the adversarial sticker is small in size and cannot conceal key information about the initial face when applied.

[0022] like Figure 3As shown, the inventors' research also revealed that the adversarial hat involves printing adversarial perturbations onto a hat, enabling the initial face to malfunction when wearing the hat. First, a rectangular adversarial perturbation is initialized, and its 2D projection image (GT embedding) is pasted onto the initial face image wearing the hat, resulting in an adversarial image. Facial information features (Arc Face) are extracted from the adversarial image, and the cosine distance (Cosine Similarity loss) between the facial information features and the initial face is calculated. Simultaneously, TV loss is added to minimize the objective function, thereby maximizing the cosine distance between the facial information features of the adversarial image and the initial face. It is evident that the adversarial hat algorithm primarily utilizes the facial information features of the adversarial sample, with relatively little involvement of the adversarial perturbation. Therefore, its adversarial effectiveness is weak, and its attack success rate is low. Furthermore, although the adversarial hat has a large area, the main body of the hat is not within the face area, making it easily removed by the cropping operation in the face recognition system, thus failing to achieve the desired adversarial attack effect.

[0023] Compared with existing technologies, the embodiments of this application first obtain target face images that meet the first preset conditions from a preset face image set based on an initial face image and adversarial images. Then, adversarial face samples are generated based on the target face images. Therefore, the adversarial face samples generated in this application embodiment do not directly aim to increase the similarity with the attacker or decrease the similarity with the protected person. Instead, they are optimized based on multiple face images using a preset face image set. The optimization target is not optimized towards a single target face as in existing technologies, but towards multiple target faces. Therefore, the final adversarial face samples have a high attack success rate and high attack stability in the physical world. The adversarial face samples generated in this application may be obtained by different face recognition systems with different effective perturbation features, thus having the same attack result. That is, the adversarial face samples generated in this application have strong attack robustness and strong transferability, and can produce stable attack effects against multiple different face recognition models. Furthermore, in some embodiments of this application, adversarial examples can be materialized as preset objects occupying a large area of ​​the face, and then used to test adversarial attacks on the face recognition model or to help users protect their facial privacy. Compared to adversarial glasses and stickers, the preset objects in the embodiments of this application have a larger area and can carry more adversarial features, thus achieving better adversarial attack testing and facial privacy protection effects. Compared to adversarial hats, since the preset objects generated in the embodiments of this application are used to cover the face rather than the head, they are not easily removed by the cropping operation in the face recognition system, thus achieving a better adversarial attack effect. In the embodiments of this application, adversarial images or adversarial examples can be generated through an image processing system including an image processing device and an image recognition device.

[0024] In some implementations, the image processing device and the image recognition device are deployed separately, such as... Figure 4 As shown, the face image processing method provided in this application embodiment can be based on Figure 4 An image processing system is shown. This image processing system may include a server 01 and a terminal device 02.

[0025] The server 01 can be an image processing device in which image processing programs, such as face image processing programs, can be deployed.

[0026] The terminal device 02 can be an image recognition device, which may be equipped with a recognition model, such as an image recognition model trained using machine learning methods. This image recognition model can be a face recognition model, etc.

[0027] Server 01 can receive a preset set of face images and initial face images from an external source. Based on the initial face images, it obtains candidate adversarial images and sends each face image to terminal device 02. Terminal device 02 can process each face image using a face recognition model to obtain the face features of each face image and then feed them back to server 01. Server 01 can receive the initial features of the initial face image, the adversarial features of the candidate adversarial images, and the face features of each face image in the preset set of face images. Based on the adversarial features and the initial features, it obtains a target face image from each face image in the set of face images and determines the first similarity between the candidate adversarial image and the target face image. If the first similarity is less than a third preset value, the candidate adversarial perturbation is updated. A new target face image is obtained based on the updated candidate adversarial perturbation and the initial face image, until the first similarity between the candidate adversarial image and the target face image is not less than the third preset value. The candidate adversarial image with the first similarity not less than the third preset value is used as a face adversarial sample, and the adversarial perturbation corresponding to the candidate adversarial image can be used as the target adversarial perturbation.

[0028] It should be noted that the server involved in the embodiments of this application can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0029] The terminal devices involved in the embodiments of this application can be devices that provide voice and / or data connectivity to users, handheld devices with wireless connectivity, or other processing devices connected to a wireless modem. Examples include mobile phones (or "cellular" phones) and computers with mobile terminals, such as portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile devices that exchange voice and / or data with a wireless access network. Examples include Personal Communication Service (PCS) phones, cordless phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, Personal Digital Assistants (PDAs), and other devices.

[0030] The technical solution of this application will now be described in detail with reference to several embodiments.

[0031] Combination Figure 5 , Figure 6 This application describes a face image processing method according to embodiments of the present application, which can be applied to... Figure 4 The image processing system shown is executed by a server to update candidate adversarial perturbations to obtain a target adversarial perturbation. This target adversarial perturbation can be used to generate an object that occludes the user's face, which, when worn by the user, protects the user's facial privacy. The method includes the following steps: Step S100: Determine the initial face image.

[0032] In this embodiment, the initial face image can be a face image that meets preset privacy protection conditions, or it can be a face image for any purpose (e.g., for conducting adversarial attack tests on face recognition models). The preset privacy protection conditions can be scenarios where the user needs to protect their privacy. For example, if a user needs to protect their facial privacy in a shopping mall or real estate sales center to prevent illegal collection by the merchant, then the user's privacy needs meet the preset privacy protection conditions. The preset privacy protection conditions can also be the identity of the source user of the initial face image. For example, if the user's identity is legal and not involved in illegal or criminal activities, then it can be considered to meet the preset privacy protection conditions.

[0033] For example, real estate companies use facial recognition systems to collect customers' purchasing intentions and offer higher prices to customers with strong intentions. However, when customer facial privacy is protected, the real estate company's facial recognition system cannot verify the customer's identity based on facial information, and therefore cannot arbitrarily adjust prices based on the customer's purchasing intentions.

[0034] For example, shopping malls use facial recognition systems to learn about consumers' purchasing habits and then target those with consistent purchasing patterns with product promotions. However, when consumers' facial privacy is protected, the mall's facial recognition system cannot identify consumers based on their facial information, thus preventing malicious sales tactics.

[0035] It should be noted that adversarial attacks can include targeted and untargeted attacks, and their objectives are different. Therefore, the basis for generating the targeted adversarial images used to achieve these two attacks is also different. In a targeted attack, the initial face image can be the face image of the victim; that is, the face of the protected individual, after being recognized by the face recognition system, will be incorrectly identified as the victim (the user who created the initial face image).

[0036] In a targetless attack, the initial face image can be the face image of the protected person. That is, after the face recognition system recognizes the face of the protected person (the initial face image), it will be incorrectly identified as someone other than the protected person.

[0037] Step S200: Obtain candidate adversarial images.

[0038] Generally speaking, adversarial images can be obtained by directly superimposing adversarial perturbations onto the protected face image. Therefore, the process of iteratively updating candidate adversarial images to obtain the target adversarial image can be regarded as the process of iteratively updating candidate adversarial perturbations to obtain the target adversarial perturbation, without making any modifications to the protected face image.

[0039] Therefore, in this embodiment, an initial adversarial perturbation can be directly initialized, and then iteratively updated. During the iterative update process, multiple historical candidate adversarial perturbations may be obtained as the basis for updating candidate adversarial perturbations at the next time step, until the target adversarial perturbation is obtained. That is, the target adversarial perturbation is updated from the candidate adversarial perturbations obtained in the previous time step. For example, assuming that the target adversarial perturbation C is obtained by updating the initial adversarial perturbation c1 three times, then the first update based on the initial adversarial perturbation c1 is used to obtain the candidate adversarial perturbation c2, then the second update based on the candidate adversarial perturbation c2 is used to obtain the candidate adversarial perturbation c3, and finally the target adversarial perturbation C is obtained by updating based on the candidate adversarial perturbation c3.

[0040] In this embodiment, the initial adversarial perturbation can be generated by a preset pattern generation model based on preset latent vectors. The pattern generation model includes an encoder and a decoder. The encoder encodes the preset latent vectors, and the decoder decodes the encoded latent vectors to generate the initial adversarial perturbation. It should be noted that the size of the initial adversarial perturbation needs to be smaller than the initial face image.

[0041] In this embodiment of the application, candidate adversarial images can be obtained through the following methods ① and ②: Method ①: The candidate adversarial perturbation and the protected face image are weighted and calculated to obtain the candidate adversarial image.

[0042] In this embodiment of the application, a first weight vector and a second weight vector can be preset, so that the protected face image and the candidate adversarial perturbation can be calculated with the preset first weight vector and the second weight vector to obtain the candidate adversarial image.

[0043] For example, the first weight vector can contain multiple weight vector elements, and these elements correspond one-to-one with the pixels in the candidate adversarial perturbation. Similarly, the second weight vector can contain multiple weight vector elements, and these elements correspond one-to-one with the pixels in the protected face image. It should also be noted that the candidate adversarial perturbation is smaller than the protected face image; therefore, the first weight vector contains fewer weight vector elements than the second weight vector. Furthermore, the first and second weight vectors at the corresponding positions in the candidate adversarial perturbation and the protected face image are complementary. For instance, if the candidate adversarial perturbation is a pattern of the nose region on the face, then the sum of the weight vector elements in the first weight vector at the nose region and the corresponding weight vector elements in the second weight vector at the nose region is 1, while all weight vector elements in the second weight vector except for those at the nose region are 1.

[0044] Therefore, by adding the product of each weight vector element in the first weight vector with each pixel of the candidate adversarial perturbation, and the product of each weight vector element in the second weight vector with each pixel of the protected face image, the candidate adversarial image is obtained.

[0045] Method ②: Replace the preset region of the protected face image with the candidate adversarial perturbation to obtain the candidate adversarial image.

[0046] In this embodiment, a preset region in the protected face image can be cropped. For example, if the candidate adversarial perturbation is the nose region, then the nose region of the protected face image is the preset region. The nose region in the protected face image can be cropped out, and the candidate adversarial perturbation can be added to the cropped nose region of the protected face image to obtain the candidate adversarial image.

[0047] Both methods ① and ② described above can be used to obtain candidate adversarial images based on the face image to be protected and candidate adversarial perturbations.

[0048] Step S300: Based on the initial face image and the candidate adversarial image, determine the target face image that meets the first preset condition from the preset face image set.

[0049] In this embodiment, the preset face image set can be an open-source face image set or a set of face images obtained by temporarily capturing multiple faces. The preset face image set contains face data of multiple different faces, such as face images or photos. Since the target face image is obtained from the preset face image set, rather than a single target (original) face image as in the prior art, the adversarial sample generated based on the target face image is not similar to or dissimilar to a single target (original) face image, but rather similar to multiple target face images. That is, the external appearance of the adversarial sample is similar to multiple target face images, thus the visual representation of the face is more generic and will not be easily detected by the naked eye. Furthermore, it should be noted that in a targetless attack scenario, the preset face image set may or may not include the initial face image, while in a targeted attack scenario, the preset face image set cannot include the initial face image (the attacked face image) to prevent optimization solely towards the single target face of the attacked face.

[0050] In this embodiment, the target face image can be used as an optimization direction for candidate adversarial images. That is, in each iteration of updating candidate adversarial images, a target face image is determined, and the candidate adversarial images are optimized in a direction similar to the target face image. The target face image is a face image that meets a first preset condition among all face images in a preset face image set, that is, a face image that can help achieve the purpose of adversarial attack.

[0051] In this application embodiment, the acquisition of the target face image is divided into two cases: targeted attack and non-targeted attack. The acquisition of the target face image under the two cases of targeted attack and non-targeted attack will be described below.

[0052] When the aforementioned adversarial sample of the face is used for untargeted attacks: The target face image is determined from a preset set of face images by following the steps ad.

[0053] a. Obtain the initial features of the initial face image.

[0054] In this embodiment, the adversarial example of a face is used for a targetless attack. In this case, the initial face image is a protected face image, and the initial features can be obtained by extracting image features from the protected face image. The extraction of image features from the protected face image can be achieved through a preset face recognition model. For example, the protected face image is input into the face recognition model, and the face recognition model can perform feature extraction on the input protected face image to obtain the initial features (face features of the protected face image).

[0055] b: Obtain the adversarial features of the adversarial image.

[0056] After obtaining the adversarial image through methods ① and ② above, the method for obtaining the adversarial features can be the same as the method for obtaining the initial features. For example, the candidate adversarial image can be input into a face recognition model, and the face recognition model can extract features from the input candidate adversarial image to obtain the adversarial features.

[0057] c: Obtain the facial features of each face image in the preset face image set.

[0058] In this embodiment, facial features corresponding to each face image in the preset face image set can be obtained based on the preset face image set. The facial features of each face image in the preset face image set can be obtained using the same extraction method as adversarial features and initial features. For example, the face images in the preset face image set can be input into a face recognition model, which can extract features from each input face image to obtain the facial features corresponding to each face image. It should be noted that the facial features corresponding to each face image in the preset face image set can be saved locally or in the cloud after extraction, and can be used directly the next time without needing to perform facial feature extraction every time.

[0059] d: Select the target face image based on the first preset condition and the similarity between each of the face features and the initial feature and the adversarial feature.

[0060] In this embodiment, a third similarity between each facial feature and an adversarial feature can be used to represent the first similarity between each facial image and a candidate adversarial image, and a fourth similarity between each facial feature and an initial feature (the protected facial feature) can be used to represent the second similarity between each facial image and the initial facial image (the protected facial image). The third similarity between each facial feature and an adversarial feature, and the fourth similarity between each facial feature and an initial feature, can be obtained based on the distances (e.g., Euclidean distance, Chebyshev distance, or cosine similarity) between each facial feature and both the adversarial feature and the initial feature.

[0061] In this embodiment, a third similarity between each facial feature and adversarial feature can be used to represent the first similarity between each facial image and the candidate adversarial image. The higher the third similarity, the higher the first similarity, and vice versa. A fourth similarity between each facial feature and the initial feature (protected facial feature) can be used to represent the second similarity between each facial image and the initial facial image (protected facial image). The higher the fourth similarity, the higher the second similarity, and vice versa. The difference between the third and fourth similarities represents the difference between the first and second similarities. When the difference between the third and fourth similarities is the largest, the difference between the first and second similarities is also the largest, and at this point, the difference between the first and second similarities can be considered to be greater than a first preset value. When the difference between the third and fourth similarities is the largest, the difference between the first and second similarities is also the largest, meaning that the first similarity is the largest while the second similarity is the smallest. At this point, the target facial image has the highest similarity with the candidate adversarial image and the lowest similarity with the initial facial image (protected facial image).

[0062] Therefore, the target face image can be determined from the preset face image set based on the third and fourth similarity scores.

[0063] In this embodiment of the application, the target face image can be obtained from the preset face image set based on the following formula (1) according to the third similarity and the fourth similarity.

[0064] (1) Where T is the set of facial features composed of facial features corresponding to each facial image in the preset set of facial images, x is any facial feature in the set of facial features, O is the initial feature (the protected facial feature), and adv is the adversarial feature. The first distance from face feature x to the initial feature (the protected face feature) represents the fourth similarity score. The second distance from face feature x to adversarial feature represents the third similarity. P is the face feature with the largest difference between the second distance and the first distance among all face features in the face feature set T, i.e., the target face feature. The face image corresponding to the target face feature is the target face image.

[0065] Among these, the difference between the fourth and third similarities corresponding to the target facial features is the largest, meaning the target facial features have the highest similarity to the adversarial features, the smallest second distance, and the lowest similarity to the initial features (protected facial features), with the largest first distance. The facial image corresponding to the target facial features is the target facial image. Therefore, the target facial image has the lowest similarity to the initial facial image (protected facial image) and the highest similarity to the candidate adversarial image. Thus, when the candidate adversarial perturbation is fused with the initial facial image (protected facial image) to obtain the candidate adversarial image, it is more likely to be incorrectly identified as the target facial image by the facial recognition system compared to the initial facial image (protected facial image), thereby protecting the privacy of the initial facial image (protected facial image).

[0066] When the aforementioned adversarial sample of the face is used for a targeted attack: The target face image is determined from a preset set of face images by following the steps eh.

[0067] e. Obtain the initial features of the initial face image.

[0068] In this embodiment, the adversarial sample for faces is used for targeted attacks. In this case, the initial face image is the face image of the target, and the initial features can be obtained by extracting image features from the face image of the target. The method for extracting image features from the face image of the target can be the same feature extraction method as that used for the image features in the face image of the protected face, and will not be described in detail here.

[0069] f: Obtain the adversarial features of the adversarial image.

[0070] In the embodiments of this application, the detailed steps of step f are the same as those of step b, and will not be repeated here.

[0071] g: Obtain the facial features of each face image in the preset face image set.

[0072] In the embodiments of this application, the detailed steps of step g are the same as those of step c, and will not be repeated here.

[0073] h: Select the target face image based on the first preset condition and the similarity between each of the face features and the initial feature and the adversarial feature.

[0074] In this embodiment, a third similarity between each facial feature and an adversarial feature can be used to represent the first similarity between each facial image and the candidate adversarial image. The higher the third similarity, the higher the first similarity, and vice versa. A fourth similarity between each facial feature and the initial feature (the attacked facial feature) can be used to represent the second similarity between each facial image and the initial facial image (the attacked facial image). The higher the fourth similarity, the higher the second similarity, and vice versa. The calculation methods for each similarity are referred to in step d, and will not be elaborated here.

[0075] The sum of the third and fourth similarities represents the sum of the first and second similarities. When the sum of the third and fourth similarities is maximized, the sum of the first and second similarities is also maximized. At this point, it can be considered that the sum of the first and second similarities is greater than the second preset value. When the sum of the third and fourth similarities is maximized, the sum of the first and second similarities is also maximized. That is, the first similarity is maximized while the second similarity is also maximized. At this point, the target face image has the highest similarity with the candidate adversarial image, and also the highest similarity with the initial face image (the attacked face image). In other words, the candidate adversarial image has a high degree of similarity to the initial face image (the attacked face image).

[0076] In this embodiment of the application, the target face image can be obtained from the preset face image set based on the following formula (2) according to the third similarity and the fourth similarity.

[0077] (2) Where T is a set of facial features composed of facial features corresponding to each facial image in a preset set of facial images, x is any facial feature in the set of facial features, O is the initial feature (the facial feature of the attacked party), and adv is the adversarial feature. The first distance from face feature x to the initial feature (the attacked face feature) represents the fourth similarity score. The second distance from face feature x to adversarial feature represents the third similarity. P is the face feature in face feature set T that has the largest sum of the second distance and the first distance between each face feature. This is the target face feature, and the face image corresponding to the target face feature is the target face image.

[0078] Among these, the sum of the fourth and third similarities corresponding to the target facial features is the largest, meaning the target facial features have the highest similarity to the adversarial features, the smallest second distance, and simultaneously the highest similarity to the initial features (the attacked facial features), with the smallest first distance. The facial image corresponding to the target facial features is the target facial image. Therefore, the target facial image has the highest similarity to the initial facial image (the attacked facial image) and also the highest similarity to the candidate adversarial image. Thus, when the candidate adversarial image, obtained by fusing the candidate adversarial perturbation with the protected facial image, is recognized by the facial recognition system, it is more likely to be mistakenly identified as the attacked face compared to the protected facial image, thereby protecting the privacy of the protected facial image.

[0079] The above steps ad and eh describe the methods for obtaining target face images under untargeted and targeted attacks, respectively. After obtaining the target face image, step S400 can be performed: generating adversarial face samples based on the target face image and the initial face image.

[0080] In this embodiment, regardless of whether the attack is untargeted or targeted, steps ad or ef can ensure that the acquired target face image has a high similarity to the candidate adversarial image and a low similarity to the protected face image (in the case of untargeted attack), or a high similarity to the attacked face (in the case of targeted attack). However, it cannot guarantee that the face recognition system will incorrectly identify the candidate adversarial image as the target face image or the attacked face image. Therefore, a third preset value can be set. After obtaining the target face image, the first similarity between the target face image and the candidate adversarial image is compared to see if it reaches the third preset value. This determines whether the candidate adversarial image can cause the face recognition system to incorrectly identify it as the target face image or the attacked face image. If the third preset value is not reached, the candidate adversarial perturbation needs to be updated.

[0081] When a facial recognition system identifies an image, it does so by analyzing the image's features to determine which face it belongs to. Therefore, a first similarity score can be calculated between the adversarial features of a candidate adversarial image and the target face features of the target face image. The smaller the feature loss between the target face image and the candidate adversarial image, the easier it is for the attacked facial recognition system to incorrectly identify the target face image (in the absence of a target attack) or the attacked face image (in the presence of a target attack) when using the candidate adversarial image to attack the facial recognition system.

[0082] In this embodiment, the formula for calculating the feature loss based on the adversarial features of the candidate adversarial image and the target face features of the target face image is as follows: Loss= (3) Where adv is the adversarial feature of the candidate adversarial image, best is the target face feature corresponding to the target face image, N is the total number of vector elements in the feature vector, and Loss is the feature loss between the adversarial feature and the target face feature, representing the first similarity between the candidate adversarial image and the target face image. The smaller the feature loss, the higher the first similarity between the candidate adversarial image and the target face image, and the higher the attack success rate. Conversely, the attack success rate is lower.

[0083] According to the above formula (3), the first similarity between the candidate adversarial image obtained by the candidate adversarial perturbation and the target face image obtained by the candidate adversarial image can be calculated.

[0084] In this embodiment of the application, a preset value for feature loss can be set in advance to determine whether the feature loss between the candidate adversarial image and the target face image reaches the preset value for feature loss. When the feature loss between the candidate adversarial image and the target face image reaches the preset value for feature loss, it can be considered that the first similarity has reached the third preset value; or, when the feature loss between the candidate adversarial image and the target face image reaches the minimum (at which point the first similarity reaches the maximum), it can also be considered that the first similarity has reached the third preset value.

[0085] In this embodiment of the application, when the first similarity between the candidate adversarial image obtained from the candidate adversarial perturbation and the target face image obtained from the candidate adversarial image does not reach a third preset value, the candidate adversarial perturbation can be updated by the following step ik: i: Obtain the gradient change information of the first similarity relative to the latent vector.

[0086] j: Update the hidden vector based on the gradient change information.

[0087] k: Update the candidate adversarial perturbation based on the updated hidden vector.

[0088] In this embodiment, the direction in which the first similarity between the candidate adversarial image and the target face image is maximized is the direction in which the feature loss between the candidate adversarial image and the target face image is minimized. The candidate adversarial perturbation is generated by a preset pattern generation model based on latent vectors. Therefore, optimizers (such as Momentum optimizer, AdaGrad optimizer, RMSProp optimizer, and Adam optimizer) can be used to optimize the latent vectors of the pattern generation model in the direction that reduces the feature loss.

[0089] Considering that directly adjusting adversarial perturbations through parameter optimization (e.g., superposition) is a linear modification of the perturbation, the resulting adversarial image may only be effective against a limited number of face recognition models or the image recognition model used during generation, with poor transfer attack performance against other face recognition models, this embodiment updates the adversarial perturbation through indirect optimization to generate a more transfer-attack-capable adversarial image. Specifically, this includes updating the adversarial perturbation by updating the latent vectors of the preset pattern generation model.

[0090] In this embodiment, the latent vector can be adjusted by gradient iterative optimization, specifically by calculating the gradient of the first similarity expectation relative to the latent vector; calculating optimization parameters based on a preset step size and the direction of the gradient; then adjusting the latent vector based on the optimization parameters; and finally generating an updated adversarial perturbation based on the latent vector.

[0091] In this embodiment, the adversarial perturbation is updated through indirect optimization. Direct optimization of the adversarial perturbation is transformed into optimizing the input—the latent vector—when generating the adversarial perturbation. Changes in the latent vector lead to changes in the generated adversarial perturbation, which in turn leads to changes in the adversarial image. That is, the generative model controls and coordinates the generation process of the adversarial image, ensuring that the adversarial perturbation in the adversarial image is not directly and linearly superimposed on the initial face image, but rather generated at the semantic level. This makes the adversarial perturbation more natural and seamlessly integrated into the initial face image, less likely to be detected by the face recognition model, and possesses stronger transfer attack performance.

[0092] In this embodiment of the application, the encoder of the pattern generation model is used to encode the updated latent vector, and the decoder is used to decode it to obtain the updated candidate adversarial perturbation, that is, the candidate adversarial perturbation obtained in the new round of iteration.

[0093] After obtaining the new candidate adversarial perturbation, a new candidate adversarial image can be obtained based on the new candidate adversarial perturbation and the protected face image using the method described in step S200. Then, the target face image under the new candidate adversarial image is obtained from the preset face image set. It should be noted that the target face image at this time may be the same face image as the target face image in the previous round, or it may be a different face image. Then, the feature loss between the new candidate adversarial image and the new target face image is calculated. Based on the feature loss, it is determined whether the first similarity reaches the second preset value. If it does not reach the third preset value, steps S200-S400 are repeated until the first similarity between the candidate adversarial image corresponding to the last updated candidate adversarial perturbation and the target face image reaches the third preset value or the maximum value. Then, the candidate adversarial perturbation obtained in the last update can be used as the target adversarial perturbation, and the adversarial image obtained from the target adversarial perturbation is used as the target adversarial image. The target adversarial image at this time is the face adversarial sample.

[0094] In another embodiment of this application, the first similarity between the candidate adversarial face image and the target face image can be determined by setting the number of updates. For example, if the number of updates for the candidate adversarial perturbation is set to 100, then after the initial candidate adversarial perturbation is updated to the 100th time, the adversarial perturbation obtained from the 100th update can be considered as the target adversarial perturbation.

[0095] Compared to existing technologies, the face image processing method in this application first determines target face images that meet a first preset condition from a preset face image set based on an initial face image and candidate adversarial images. Then, it generates adversarial face samples based on the target face images and the initial face images. That is, the adversarial face samples generated in this application embodiment are not optimized directly towards increasing similarity to a single attacker or decreasing similarity to a single protected party, as is done in existing technologies. Instead, they are optimized multiple times based on different face images within the preset face image set, aiming for the most similarity to multiple target faces. Therefore, the final adversarial face sample is not merely similar or dissimilar to a single image, but similar to multiple target images that are similar or dissimilar to the single image. In other words, the adversarial face sample can acquire more adversarial features for implementing adversarial attacks based on multiple target images, rather than just a single adversarial feature optimized from a single image. Therefore, the adversarial examples generated in this application can resemble multiple target images, possessing more adversarial features capable of enabling adversarial attacks. When facing different face recognition systems, different systems can acquire different effective adversarial features, resulting in the same attack outcome. In other words, the adversarial examples generated in this application exhibit strong attack robustness and transferability, producing stable attack effects against multiple different face recognition models. Thus, the success rate and physical stability of adversarial attacks using these face examples are both high, providing excellent protection for facial privacy and allowing for transferable testing of adversarial attacks against a wider range of face recognition models.

[0096] Through the above steps S100-S400, adversarial samples that enable the face recognition system to identify erroneous faces can be obtained. At the same time, the target adversarial perturbation corresponding to the face adversarial sample can also be obtained. After materializing the target adversarial perturbation, a preset object can be obtained to cover the face of the protected person. The user can wear the preset object to protect the privacy of their face.

[0097] In order to protect the user's facial privacy in the physical world, in this embodiment of the application, after determining the target adversarial perturbation, the method further includes: materializing the target adversarial perturbation into a preset object, the preset object being used to cover the target area of ​​the protected face, the area of ​​the target area being greater than a preset ratio to the area of ​​the protected face.

[0098] like Figure 7 As shown in this embodiment, the preset object is an anti-mask. Figure 7 a is the target counter-perturbation obtained using the image processing methods in the above embodiments; Figure 7 b is the general Figure 7The target adversarial perturbation m1 in a is converted into a UV image in the shape of a mask; Figure 7 c is the general Figure 7 The diagram in section b illustrates an adversarial mask obtained by printing a UV image of the mask shape onto a mask and wearing it on the target user's face (the protected face). Compared to existing technologies, the target adversarial perturbation in this application undergoes multiple updates and optimizations. Each optimization update, moving towards a similarity to the target face image, may involve the same or different target faces. In other words, the target adversarial perturbation is optimized towards greater similarity with multiple different target faces. Therefore, the final target adversarial perturbation exhibits strong attack robustness and transferability, producing stable attack effects against multiple different recognition models. Thus, converting the target adversarial perturbation into a UV image of the mask shape and printing it on a mask provides good privacy protection against different face recognition systems when worn by the target user. Furthermore, the mask has a large area, occupying a higher percentage of the target user's face than adversarial glasses, stickers, or hats, effectively obscuring most of the target user's key facial information and providing physical protection. It should be noted that in this embodiment, the materialized preset object is a mask, but in other embodiments it can be other materialized objects, such as a face towel or a face shield.

[0099] like Figure 8 As shown, after materializing the target adversarial perturbation into a preset object, the user can wear this preset object to protect their facial privacy. Assuming the preset object is a face mask, when a user wearing this mask is recognized by the facial recognition system, the system will determine the user's identity information based on the following steps: A face privacy-preserving image is acquired, which includes at least an image of a preset object. In some embodiments, the face privacy-preserving image may be an image captured after the preset object covers the target user's face (the protected face), or it may be an image of the preset object itself. Figure 8 As shown in the image, the target user is wearing a mask. The camera of the facial recognition system acquires the facial image of the target user wearing the mask, which is the facial privacy protection image, and then transmits it to the facial recognition model to identify the target user's identity.

[0100] Extract the image features of the face privacy protection image.

[0101] The identity of the target user is determined based on the similarity between the image features of the face privacy protection image and the features in the preset face feature database.

[0102] The target user's identity is identified by the label of the target face in the preset face image set, the similarity between the image features of the target face and the image features of the face privacy protection image is greater than the third preset value, and the label of the target face is different from the real identity of the target user.

[0103] In this embodiment, after materializing the target adversarial perturbation into a preset object (e.g., a mask), the target user (e.g., the source user whose face is being protected) can cover the target area of ​​their face (e.g., the lower half of their face) with the preset object. When the target user wearing the preset object has their face captured by the face recognition system, the image captured by the face recognition system is a face privacy-protected image with the preset object covering the target user's face. Since this face privacy-protected image includes the preset object, the face recognition system is essentially capturing an adversarial image (face privacy-protected image). Subsequently, the face recognition system can obtain its image features based on the adversarial image, and based on the image features of the adversarial image, calculate the similarity between it and each feature in the preset face feature library, and select the identity label of the feature with the highest similarity as the recognition result of the adversarial image. This recognition result is the identity recognition result obtained by the face recognition system when facing the target user wearing the preset object. If the similarity between the features of the target face in the facial feature database and the adversarial image (facial privacy protection image) is greater than the third preset value, the facial recognition system will determine that the image features of the adversarial image have the highest similarity to the features of the target face in the facial feature database. However, the identity of the target face is not the real identity of the target user. In other words, when the facial recognition system encounters a target user wearing a preset object, it will mistakenly identify the user as the target face, thereby protecting the facial privacy of the target user.

[0104] The above describes a face image processing method according to an embodiment of this application. The following describes a face image processing device (e.g., a server) that performs the above face image processing method.

[0105] See Figure 9 ,like Figure 9 The diagram shows the structure of an image processing device 60, which can be applied in a server to determine an initial face image, acquire candidate adversarial images, determine a target face image that meets a first preset condition from a preset face image set based on the initial face image and the candidate adversarial images, and generate adversarial face samples based on the target face image and the initial face image. The image processing device 60 in this embodiment can achieve the above-described... Figure 5The steps of the face image processing method executed in the corresponding embodiment are described. The functions implemented by the image processing device 60 can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. The image processing device 60 may include a processing module 620 and an input / output module 610. The functional implementation of the processing module 620 and the input / output module 610 can be found in [reference needed]. Figure 5 The operations performed in the corresponding embodiments will not be described in detail here.

[0106] In this embodiment of the application, the image processing apparatus 60 includes: Input / output module 610 is used to determine the initial face image; Processing module 620 is used to acquire candidate adversarial images, which are updated based on historical candidate adversarial images; The processing module is further configured to determine, based on the initial face image and the candidate adversarial image, a target face image that meets a first preset condition from a preset face image set; and Adversarial face samples are generated based on the target face image and the initial face image.

[0107] In this embodiment of the application, the processing module 620 is configured to determine the target face image based on the following first preset condition: When the aforementioned adversarial sample of faces is used for untargeted attacks... The first preset condition includes: the difference between the first similarity and the second similarity is greater than a first preset value; When the aforementioned adversarial sample of the face is used for a targeted attack... The first preset condition includes: the sum of the first similarity and the second similarity is greater than the second preset value; Wherein, the first similarity is the similarity between the target face image and the candidate adversarial image, and the second similarity is the similarity between the target face image and the initial face image.

[0108] In this embodiment of the application, the processing module 620 is further configured to: Candidate adversarial perturbations are obtained, which are updated based on historical candidate adversarial perturbations; The candidate adversarial image is obtained based on the candidate adversarial perturbation and the initial face image; Determine the target face image that meets the first preset condition from the preset face image set; If the first similarity is less than the third preset value, the candidate adversarial perturbation is updated until the first similarity is not less than the third preset value, and the candidate adversarial image when the first similarity is not less than the third preset value is used as the face adversarial sample.

[0109] In this embodiment of the application, the processing module 620 is further configured to: Candidate adversarial perturbations are obtained, which are updated based on historical candidate adversarial perturbations; The candidate adversarial image is obtained by weighting the candidate adversarial perturbation and the initial face image; or The preset region of the initial face image is replaced with the candidate adversarial perturbation to obtain the candidate adversarial image; Wherein, when the face adversarial sample is used for a targeted attack, the initial face image is the face image of the attacked party; When the adversarial sample is used for a targetless attack, the initial face image is a protected face image.

[0110] In this embodiment of the application, the processing module 620 is further configured to: Obtain the initial features of the initial face image; Obtain the adversarial features of the candidate adversarial images; Obtain the facial features of each face image in the preset face image set; The target face image is selected based on the first preset condition and the similarity between each of the face features and the initial feature and the adversarial feature.

[0111] In this embodiment of the application, the processing module 620 is further configured to: The third similarity between each of the facial features and the adversarial features is obtained, and the fourth similarity between each of the facial features and the initial features is obtained. Based on each of the third similarity and each of the fourth similarity, the difference between the third similarity and the fourth similarity associated with each face feature is obtained, as well as the sum of the third similarity and the fourth similarity associated with each face feature. When the face adversarial sample is used for a targeted attack, the face image corresponding to the face feature with the largest sum of the third and fourth similarities is selected as the target face image; When the face adversarial sample is used for a targetless attack, the face image corresponding to the face feature with the largest difference between the third and fourth similarities is selected as the target face image.

[0112] In this embodiment of the application, if the first similarity is less than a third preset value, the processing module 620 is further configured to: Obtain the gradient change information of the first similarity relative to the latent vector of the pattern generation model that generates the candidate adversarial perturbation; The hidden vector is updated based on the gradient change information; The candidate adversarial perturbation is updated based on the updated hidden vector.

[0113] In this embodiment of the application, the processing module 620 is further configured to: The adversarial perturbation corresponding to the adversarial sample of the face is materialized into a preset object, wherein the preset object is used to cover the target region of the protected face, and the area of ​​the target region is greater than a preset ratio to the area of ​​the protected face; the preset object includes one of the following: Masks, face towels, face shields.

[0114] In this embodiment of the application, the processing module 620 is further configured to: Acquire a face privacy-protected image, wherein the face privacy-protected image is an image captured after the preset object covers the target user's face, or is an image of the preset object; Extract the image features of the face privacy-protected image; The identity of the target user is determined based on the similarity between the image features of the face privacy protection image and the features in the preset face feature database. The target user's identity is identified by the label of the target face in the preset face image set, the similarity between the image features of the target face and the image features of the face privacy protection image is greater than the third preset value, and the label of the target face is different from the real identity of the target user.

[0115] According to the image processing apparatus of this application embodiment, after materializing the target adversarial perturbation into a preset object (e.g., a mask), the target user (e.g., the source user whose face is being protected) can cover the target area of ​​their face (e.g., the lower half of their face) with the preset object. When the target user wears the preset object and their face is captured by the face recognition system, the image captured by the face recognition system is a face privacy protection image after the preset object covers the target user's face. Since the face privacy protection image includes the preset object, the face recognition system is essentially capturing an adversarial image (face privacy protection image). Subsequently, the face recognition system can obtain its image features based on the adversarial image, and based on the image features of the adversarial image, calculate the similarity between it and each feature in the preset face feature library, and select the identity label of the feature with the highest similarity as the recognition result of the adversarial image. This recognition result is the identity recognition result obtained by the face recognition system when facing the target user wearing the preset object. If the similarity between the features of the target face in the facial feature database and the adversarial image (facial privacy protection image) is greater than the third preset value, the facial recognition system will determine that the image features of the adversarial image have the highest similarity to the features of the target face in the facial feature database. However, the identity of the target face is not the real identity of the target user. In other words, when the facial recognition system encounters a target user wearing a preset object, it will mistakenly identify the user as the target face, thereby protecting the facial privacy of the target user.

[0116] The specific implementation methods of the various embodiments of the above image processing device are described in detail in the various embodiments of the face image processing method, and will not be repeated here.

[0117] The image processing apparatus in this embodiment first determines target face images that meet a first preset condition from a preset face image set based on an initial face image and candidate adversarial images. Then, it generates adversarial face samples based on the target face image and the initial face image. That is, the adversarial face samples generated in this embodiment are not optimized directly towards increasing similarity to a single attacker or decreasing similarity to a single protected party, as is done in the prior art. Instead, they are optimized multiple times based on different face images within the preset face image set, aiming for the most similarity to multiple target faces. Therefore, the final adversarial face sample is not merely similar or dissimilar to a single image, but similar to multiple target images that are similar or dissimilar to the single image. In other words, the adversarial face sample can acquire more adversarial features for implementing adversarial attacks based on multiple target images, rather than just a single adversarial feature optimized from a single image. Therefore, the adversarial examples generated in this application can resemble multiple target images, possessing more adversarial features capable of enabling adversarial attacks. When facing different face recognition systems, different systems can acquire different effective adversarial features, resulting in the same attack outcome. In other words, the adversarial examples generated in this application exhibit strong attack robustness and transferability, producing stable attack effects against multiple different face recognition models. Thus, the success rate and physical stability of adversarial attacks using these face examples are both high, providing excellent protection for facial privacy and allowing for transferable testing of adversarial attacks against a wider range of face recognition models.

[0118] After introducing the methods and apparatus in the embodiments of this application, the computer-readable storage medium in the embodiments of this application will be described next. In the embodiments of this application, the computer-readable storage medium is an optical disc, on which a computer program (i.e., a program product or instructions) is stored. When the computer program is run by a computer, it will implement the steps described in the above method embodiments, such as: determining an initial face image; obtaining candidate adversarial images; determining a target face image that meets a first preset condition from a preset face image set based on the initial face image and the candidate adversarial images; and generating adversarial face samples based on the target face image and the initial face image. The specific implementation of each step will not be repeated here.

[0119] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.

[0120] The image processing apparatus in the embodiments of this application has been described above from the perspective of modular functional entities. The server and terminal device executing the image processing method in the embodiments of this application are described below from the perspective of hardware processing.

[0121] It should be noted that, in the embodiments of the image processing apparatus of this application... Figure 9 The physical device corresponding to the input / output module 610 shown can be an input / output unit, transceiver, radio frequency circuit, communication module, and input / output (I / O) interface, etc., and the physical device corresponding to the processing module 620 can be a processor. Figure 9 The image processing apparatus shown can have, for example: Figure 10 The structure shown, when Figure 9 The image processing device shown has, for example Figure 10 When the structure shown is used, Figure 10 The processor and transceiver in the device can perform the same or similar functions as the processing module 620 and input / output module 610 provided in the aforementioned device embodiments. Figure 10 The memory stores the computer programs that the processor needs to call when executing the above image acquisition method.

[0122] This application also provides a terminal device, such as... Figure 11 As shown, for ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of this application. The terminal device can be any terminal device including mobile phones, tablets, personal digital assistants (PDAs), point-of-sale (POS) terminals, in-vehicle computers, etc. Taking a mobile phone as an example: Figure 11 This diagram illustrates a partial structural representation of a mobile phone related to the terminal device provided in this embodiment. (Reference) Figure 11The mobile phone includes components such as a radio frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090. Those skilled in the art will understand that... Figure 11 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0123] The following is combined with Figure 11 A detailed introduction to each component of a mobile phone: The RF circuit 1010 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 1080; additionally, it transmits uplink data to the base station. Typically, the RF circuit 1010 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, the RF circuit 1010 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).

[0124] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0125] The input unit 1030 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 1031), and drive the corresponding connection devices according to a pre-set program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 1080, and can also receive and execute commands sent by the processor 1080. In addition, the touch panel 1031 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1031, the input unit 1030 may also include other input devices 1032. Specifically, other input devices 1032 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0126] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1040 may include a display panel 1041, which may optionally be configured as a liquid crystal display (LCD), organic light-emitting diode (OLED), or similar display. Further, a touch panel 1031 may cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it transmits the information to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides corresponding visual output on the display panel 1041 based on the type of touch event. Although in Figure 11 In this embodiment, the touch panel 1031 and the display panel 1041 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.

[0127] The mobile phone may also include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 1041 according to the ambient light level, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0128] The audio circuit 1060, speaker 1061, and microphone 1062 provide an audio interface between the user and the mobile phone. The audio circuit 1060 converts the received audio data into electrical signals and transmits them to the speaker 1061, where the speaker 1061 converts them into sound signals for output. On the other hand, the microphone 1062 converts the collected sound signals into electrical signals, which are then received by the audio circuit 1060, converted into audio data, and then processed by the processor 1080 before being transmitted via the RF circuit 1010 to, for example, another mobile phone, or the audio data can be output to the memory 1020 for further processing.

[0129] WiFi is a short-range wireless transmission technology. Through the WiFi module 1070, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 11 The WiFi module 1070 is shown, but it is understood that it is not an essential component of a mobile phone and can be omitted as needed without changing the essence of the invention.

[0130] The processor 1080 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 1020 and calls data stored in the memory 1020 to perform various functions and process data, thereby providing overall monitoring of the phone. Optionally, the processor 1080 may include one or more processing units; optionally, the processor 1080 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 1080.

[0131] The mobile phone also includes a power supply 1090 (such as a battery) that supplies power to various components. Optionally, the power supply can be logically connected to the processor 1080 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0132] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.

[0133] In this embodiment of the application, the processor 1080 included in the mobile phone further includes a process flow for controlling the execution of the above-mentioned steps performed by the image recognition device to acquire features based on the input facial image. The processor 1080 included in the mobile phone also includes a process flow for controlling the execution of the above-mentioned steps performed by the image processing device, such as: Determine the initial face image; Obtain candidate adversarial images; Based on the initial face image and the candidate adversarial image, a target face image that meets the first preset condition is determined from a preset face image set; Adversarial face samples are generated based on the target face image and the initial face image.

[0134] This application also provides a server; please refer to [link / reference]. Figure 12 , Figure 12This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1100 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1122 (e.g., one or more processors) and memory 1132, and one or more storage media 1130 (e.g., one or more mass storage devices) for storing application programs 1142 or data 1144. The memory 1132 and storage media 1130 can be temporary or persistent storage. The program stored in the storage media 1130 may include one or more modules (…). Figure 12 (Not shown in the image), each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1122 may be configured to communicate with the storage medium 1130 and execute the series of instruction operations in the storage medium 1130 on the server 1100.

[0135] Server 1100 may also include one or more power supplies 1120, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.

[0136] The steps performed by the server in the above embodiments can be based on this Figure 12 The structure of server 1100 is shown. For example, in the above embodiment, it consists of... Figure 9 The steps performed by the image processing device shown can be based on this Figure 12 The server structure is shown. For example, the central processing unit 1122 performs the following operations by calling instructions from memory 1132: The initial face image is determined and candidate adversarial images are obtained through the input / output interface 1158. The central processing unit 1122 determines the target face image from a preset set of face images; Based on the initial face image and the candidate adversarial image, a target face image that meets the first preset condition is determined from a preset face image set; Adversarial face samples are generated based on the target face image and the initial face image.

[0137] The target adversarial perturbation corresponding to the adversarial sample can also be output through the input / output interface 1158 so that it can be materialized and superimposed on the real initial face to protect the privacy of the face.

[0138] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0139] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0140] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.

[0141] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0142] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0143] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0144] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).

[0145] The technical solutions provided in the embodiments of this application have been described in detail above. Specific examples have been used in the embodiments of this application to illustrate the principles and implementation methods of the embodiments of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.

Claims

1. A face image processing method, the method comprising: Determine the initial face image; Candidate adversarial images are obtained, which are updated based on historical candidate adversarial images, which are iteratively updated based on initial adversarial perturbations; Based on the initial face image and the candidate adversarial image, a target face image that meets the first preset condition is determined from a preset face image set; the preset face set contains face data of multiple different faces; Generate adversarial face samples based on the target face image and the initial face image; The first preset conditions include: When the face adversarial sample is used for a targetless attack, the difference between the first similarity and the second similarity is greater than a first preset value; Alternatively, when the face adversarial sample is used for a targeted attack, the sum of the first similarity and the second similarity is greater than a second preset value; Wherein, the first similarity is the similarity between the target face image and the candidate adversarial image, and the second similarity is the similarity between the target face image and the initial face image.

2. The face image processing method as described in claim 1, wherein obtaining candidate adversarial images involves determining a target face image that meets a first preset condition from a preset face image set based on the initial face image and the candidate adversarial images; Generate adversarial face examples based on the target face image and the initial face image, including: Candidate adversarial perturbations are obtained, which are updated based on historical candidate adversarial perturbations; The candidate adversarial image is obtained based on the candidate adversarial perturbation and the initial face image; Determine the target face image that meets the first preset condition from the preset face image set; If the first similarity is less than the third preset value, the candidate adversarial perturbation is updated until the first similarity is not less than the third preset value, and the candidate adversarial image when the first similarity is not less than the third preset value is used as the face adversarial sample.

3. The face image processing method as described in claim 1, wherein obtaining candidate adversarial images includes: Candidate adversarial perturbations are obtained, which are updated based on historical candidate adversarial perturbations; The candidate adversarial perturbation and the initial face image are weighted and calculated to obtain the candidate adversarial image; or The preset region of the initial face image is replaced with the candidate adversarial perturbation to obtain the candidate adversarial image; Wherein, when the face adversarial sample is used for a targeted attack, the initial face image is the face image of the attacked party; When the adversarial sample is used for a targetless attack, the initial face image is a protected face image.

4. The face image processing method as described in claim 1, wherein, The step of determining a target face image that meets a first preset condition from a preset face image set based on the initial face image and the candidate adversarial image includes: Obtain the initial features of the initial face image; Obtain the adversarial features of the candidate adversarial images; Obtain the facial features of each face image in the preset face image set; The target face image is selected based on the first preset condition and the similarity between each of the face features and the initial feature and the adversarial feature.

5. The face image processing method as described in claim 4, wherein selecting the target face image based on a first preset condition and the similarity between each of the face features and the initial feature and the adversarial feature respectively includes: The third similarity between each of the facial features and the adversarial features is obtained, and the fourth similarity between each of the facial features and the initial features is obtained. Based on each of the third similarity and each of the fourth similarity, the difference between the third similarity and the fourth similarity associated with each face feature is obtained, as well as the sum of the third similarity and the fourth similarity associated with each face feature. When the face adversarial sample is used for a targeted attack, the face image corresponding to the face feature with the largest sum of the third and fourth similarities is selected as the target face image; When the face adversarial sample is used for a targetless attack, the face image corresponding to the face feature with the largest difference between the third and fourth similarities is selected as the target face image.

6. The face image processing method as described in claim 2, wherein updating the candidate adversarial perturbation includes: Obtain the gradient change information of the first similarity relative to the latent vector of the pattern generation model that generates the candidate adversarial perturbation; The hidden vector is updated based on the gradient change information; The candidate adversarial perturbation is updated based on the updated hidden vector.

7. The face image processing method as described in claim 2, further comprising, after determining the adversarial sample of the face: The adversarial perturbation corresponding to the adversarial sample of the face is materialized into a preset object, wherein the preset object is used to cover the target region of the target user's face, and the area of ​​the target region is greater than a preset ratio to the area of ​​the target user's face; the preset object includes one of the following: Masks, face towels, face shields.

8. The face image processing method as described in claim 7, wherein after materializing the adversarial perturbation corresponding to the face adversarial sample into a preset object, the method further includes: Acquire a face privacy-protected image, wherein the face privacy-protected image is an image captured after the preset object covers the target user's face, or is an image of the preset object; Extract the image features of the face privacy-protected image; The identity of the target user is determined based on the similarity between the image features of the face privacy protection image and the features in the preset face feature database. The target user's identity is identified by the label of the target face in the preset face image set. The similarity between the image features of the target face and the image features of the face privacy protection image is greater than the third preset value. The label of the target face is different from the real identity of the target user.

9. An image processing apparatus, comprising: Input / output module, used to determine the initial face image; The processing module is used to acquire candidate adversarial images, which are updated based on historical candidate adversarial images, and the historical candidate adversarial images are iteratively updated based on initial adversarial perturbations; The processing module is further configured to determine a target face image that meets the first preset condition from a preset face image set based on the initial face image and the candidate adversarial image, wherein the preset face image set contains face data of multiple different faces; as well as Generate adversarial face samples based on the target face image and the initial face image; The first preset conditions include: When the face adversarial sample is used for a targetless attack, the difference between the first similarity and the second similarity is greater than a first preset value; Alternatively, when the face adversarial sample is used for a targeted attack, the sum of the first similarity and the second similarity is greater than a second preset value; Wherein, the first similarity is the similarity between the target face image and the candidate adversarial image, and the second similarity is the similarity between the target face image and the initial face image.

10. A processing apparatus, the processing apparatus comprising: At least one processor, memory, and input / output unit; The memory is used to store computer programs, and the processor is used to invoke the computer programs stored in the memory to execute the method as described in any one of claims 1-8.

11. A computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • White box adversarial sample generation method for scene character recognition model

    CN111476228A

  • Adversarial sample dynamic generation method and device, electronic equipment and storage medium

    CN114419704A

  • Adversarial sample generation method, related device and storage medium

    CN115081643A