A method, system and storage medium for active protection of facial images

By generating adversarial perturbations through spatial-frequency dual-domain constraint optimization, facial geometry and identity features are disrupted, solving the problems of single protection dimension and unstable effect in existing technologies, and achieving efficient privacy protection.

CN121639437BActive Publication Date: 2026-05-26ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2026-02-05
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for protecting facial images offer limited protection dimensions, making it difficult to handle diverse editing commands. Furthermore, they do not completely destroy biometric features, leading to the leakage of identity information and inconsistent protection effectiveness.

Method used

By generating adversarial perturbations and performing joint spatial-frequency dual-domain constraint optimization, the geometric position differences of facial key points are maximized. Low-frequency perturbations are restricted by frequency domain masks, while high-frequency perturbations are allowed to distribute freely. Combined with spatial domain pixel amplitude constraints, the optimal perturbation is generated to disrupt the facial geometry and identity features.

Benefits of technology

While maintaining the visual quality of the image, it effectively disrupts facial geometry and identity features, prevents generative intelligent editing attacks, and achieves efficient and stable privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121639437B_ABST
    Figure CN121639437B_ABST
Patent Text Reader

Abstract

This invention discloses an active face image protection method, system, and storage medium, relating to the field of network information security technology. The method includes: acquiring the original face image to be protected; generating an adversarial perturbation intended to be superimposed onto the original face image; iteratively optimizing the adversarial perturbation based on preset optimization objectives and constraints; wherein the optimization objective is configured to calculate and maximize the geometric position difference of facial key points between the original face image and the image after superimposing the adversarial perturbation, thereby disrupting the facial geometric structure features in the protected image; the constraints include joint spatial and frequency domain constraints, using a frequency domain mask to apply differentiated constraints to the coefficient values ​​of the adversarial perturbation in the frequency domain, and limiting the maximum amplitude of each pixel in the spatial domain; finally, superimposing the adversarial perturbation that meets the conditions onto the original face image to obtain the face protection image. This invention can effectively defend against malicious semantic editing of generative models while maintaining the visual quality of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network information security technology, and in particular to a method, system and storage medium for active protection of facial images. Background Technology

[0002] With the rapid development of generative artificial intelligence technology, image editing tools based on diffusion models can now modify facial images with high fidelity and semantic control based on natural language instructions. While this technology brings convenience, it also raises serious privacy and security risks. Criminals can use publicly available facial photos to maliciously edit them via text commands, generating highly realistic forged images that infringe on individual rights. To address this, various facial image protection methods based on adversarial perturbations have emerged, aiming to interfere with the editing process of the model by adding minute perturbations.

[0003] However, existing protection technologies suffer from the following drawbacks. First, their protection dimensions are relatively singular. Most existing methods only add perturbations in a single dimension, such as pixel space or feature space. Limiting the perturbation amplitude only in pixel space can easily lead to the perturbation being filtered out, while adding perturbations only in feature space can easily fail during the image reconstruction stage, making it difficult to cope with diverse editing commands. Second, the destruction of biometric features is incomplete. Existing technologies mainly focus on destroying facial texture details to affect the recognition model's judgment, but they neglect the fact that the preservation of facial geometry can still easily reveal personal identity information. The recognition model can still determine the original identity based on the key point positions and geometric proportions retained after editing. Third, the stability of the protection effect is poor. Existing methods have uneven defense capabilities against different types of editing commands, making it difficult to provide continuous and reliable protection in various malicious editing scenarios.

[0004] Therefore, how to construct a robust protection mechanism that can both ensure the visual quality of images and completely destroy facial geometry and identity features when encountering generative intelligent editing attacks has become an urgent technical problem to be solved. Summary of the Invention

[0005] The main objective of this invention is to provide a method, system, and storage medium for active protection of facial images, aiming to construct a robust protection mechanism that can both ensure the visual quality of images and completely destroy the geometric and identity features of faces.

[0006] To achieve the above objectives, this invention proposes an active face image protection method, comprising the following steps:

[0007] Obtain the original face image to be protected;

[0008] Generate adversarial perturbations designed to be superimposed onto the original face image;

[0009] Based on preset optimization objectives and constraints, the anti-disturbance is iteratively optimized;

[0010] The optimization objective is configured to: calculate the geometric position difference of facial key points between the original face image and the image after superimposing the adversarial perturbation, and maximize the geometric position difference to destroy the facial geometric structure features in the original face image to be protected; the constraints include a joint spatial-frequency domain constraint, which is configured to: transform the adversarial perturbation to the frequency domain, and apply differential constraints to the coefficient values ​​of the adversarial perturbation in the frequency domain using a frequency domain mask, wherein the frequency domain mask is configured to suppress the perturbation energy in the low-frequency region and not impose amplitude limits on the perturbation energy in the high-frequency region; and limit the maximum amplitude of each pixel in the adversarial perturbation in the spatial domain; and superimpose the adversarial perturbation that satisfies the optimization objective and constraints onto the original face image to obtain the face protection image.

[0011] Preferably, in the step of applying differential constraints to the coefficient values ​​of the anti-perturbation in the frequency domain using a frequency domain mask, the frequency domain mask... The specific construction process includes: determining the adjusted height of the original face image for frequency domain transformation. and adjusted width The adjusted height and adjusted width Even number; set low frequency ratio The mid-frequency ratio is 0.15. The value is 0.4; based on the adjusted height. and adjusted width Calculate the low-frequency height boundary

[0012]

[0013] Low frequency bandwidth boundary

[0014]

[0015] Mid-frequency height boundary

[0016]

[0017] and the mid-frequency bandwidth boundary

[0018]

[0019] For the frequency domain mask Each position in Assign values ​​according to the following rules:

[0020] when and At that time, it belongs to the low-frequency region, and a value is assigned. ;when or When it belongs to the high-frequency region, a value is assigned. ;

[0021] When position When the frequency is between the low-frequency region and the high-frequency region, it belongs to the mid-frequency region. The calculation formula based on the normalized distance dynamic attenuation constraint strength is as follows:

[0022]

[0023]

[0024]

[0025]

[0026] in, This indicates the floor function.

[0027] Preferably, the step of maximizing the geometric position difference specifically includes: extracting the coordinates of 68 key points from the original face image using a facial landmark detection model. ; Calculate the key point coordinates of the superimposed image after the perturbation is applied. Constructing the geometric difference loss function Its formula is:

[0028]

[0029] in, This indicates the resistance to the disturbance. Indicates the first The coordinates of each key point are obtained by minimizing the geometric difference loss function. This is to maximize the difference in the geometric position of key points.

[0030] Preferably, in the step of iteratively optimizing the adversarial disturbance based on preset optimization objectives and constraints, a multi-objective joint optimization loss function is employed. The multi-objective joint optimization loss function The formula is:

[0031]

[0032] in, To minimize the loss based on identity similarity, To maximize the loss for latent space offset, and These are the weight parameters; the goal of the iterative optimization is to find the optimal perturbation. , so that: And it satisfies: ;in, Budget for airspace disturbances.

[0033] Preferably, the identity similarity minimizes the loss. The calculation formula is:

[0034]

[0035] in, Let be the feature vector of the original face image in the deep face recognition model. The feature vector of the image after the adversarial perturbation is superimposed in the deep face recognition model.

[0036] Preferably, the latent space offset maximizes the loss. The calculation formula is:

[0037]

[0038] in, The latent encoding of the original face image in a variational autoencoder (VAE) This is the latent encoding of the image after the adversarial perturbation is superimposed in the variational autoencoder (VAE).

[0039] Preferably, the step of transforming the anti-perturbation method to the frequency domain employs a two-dimensional discrete cosine transform (DCT), and the calculation formula is as follows: ;in, For the current confrontation and disturbance, and The DCT transform matrix contains elements of the DCT transform matrix. Defined as: when hour,

[0040]

[0041] when hour,

[0042]

[0043] in, The height is adjusted to correspond to the original face image. or width , For frequency index, For spatial indexing.

[0044] Preferably, in the step of iteratively optimizing the adversarial perturbation, an intermittent projection strategy is adopted: in every five iterations, the joint spatial-frequency dual-domain constraint is executed once; in the remaining four iterations, only the constraint of spatial domain limiting the maximum amplitude of each pixel in the adversarial perturbation is executed.

[0045] Preferably, after the step of acquiring the original face image to be protected, a preprocessing step is further included: converting the original face image into a tensor and performing normalization processing; checking the height of the original face image. and width Is it an odd number; if or If the height is odd, the original face image is adjusted to an even-sized value using bilinear interpolation. and adjusted width To meet the requirements of discrete cosine transform; after obtaining the face protection image, bilinear interpolation is used to restore the face protection image to its original height. and width .

[0046] This application also discloses an active face image protection system, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the active face image protection method as described in any of the preceding claims.

[0047] This application also discloses a computer-readable storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the active face image protection method as described in any of the preceding claims.

[0048] The above technical solution has the following advantages:

[0049] This invention proposes an active face image protection method. It acquires the original face image and generates adversarial perturbations, then iteratively optimizes these perturbations using a strategy that combines spatial and frequency domain constraints with maximizing the geometric differences of key points. In the frequency domain, this method uses a differentiated mask to strictly limit low-frequency perturbations to maintain image visual quality, while allowing high-frequency perturbations to distribute freely to enhance their aggressiveness. Combined with spatial domain pixel amplitude constraints, it ensures the protected image is imperceptible to the human eye. Simultaneously, the method sets the optimization objective to maximize the geometric differences of facial key points. Without affecting the normal use of the image, it implicitly distorts facial geometric features, causing images maliciously edited by generative models to exhibit facial structural collapse or unrecognizable identities. This effectively solves the problems of existing technologies where only texture features are destroyed, leading to identity information leakage, and the lack of robustness due to a single protection dimension. This achieves efficient and stable protection of facial privacy. Attached Figure Description

[0050] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings, wherein:

[0051] Figure 1 This is a schematic diagram of frequency domain mask distribution and frequency band division provided in an embodiment of the present invention.

[0052] Figure 2 This is a schematic diagram comparing the face image protection effects provided in the embodiments of the present invention. In the figure, (a) group is wearing glasses; (b) group is growing a goatee; (c) group is wearing a shirt; (d) group is style makeup; and (e) group is wearing a headband.

[0053] Figure 3 This is a flowchart illustrating the active face image protection method provided in an embodiment of the present invention.

[0054] Figure 4 This is a structural block diagram of the active face image protection system provided in an embodiment of the present invention.

[0055] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0056] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention. The present invention provides an active face image protection method based on spatial-frequency dual-domain constraints and maximizing the differences of facial key points. This method introduces specific adversarial perturbations into face images, thereby maintaining the visual quality of the image while destroying the geometric structural features and identity features of the face, thus preventing generative artificial intelligence models from maliciously semantically editing face images.

[0057] Example 1

[0058] This embodiment provides a method for proactive protection of facial images, primarily applied to scenarios such as social media privacy protection and public security data storage. The core logic flow of this method is as follows: Figure 3 As shown, the specific steps are as follows.

[0059] Step S1, Face Image Acquisition and Preprocessing. First, acquire the original face image to be protected, denoted as... Since subsequent frequency domain constraints involve Discrete Cosine Transform (DCT), which typically requires the object being processed to have an even number of dimensions, the input image needs to be sized and adjusted. (The original face image...) Convert to tensor format and normalize to obtain the normalized image tensor. ,in The number of RGB channels is 3, where 3 represents the batch size. For height, The width is [value]. Next, check the image height. and width The parity of . If or If the height is odd, then bilinear interpolation is used to adjust the image to an even size, resulting in the adjusted height. and adjusted width This step ensures that the image meets the basic dimensional requirements for subsequent block-based or global DCT transformations, avoiding boundary artifacts or calculation errors during the transformation process. In this embodiment, if the original size is an even number, then... and .

[0060] Step S2: Extract the original facial geometric prior information. To effectively measure and disrupt the geometric structure of the face during subsequent optimization, it is necessary to extract the geometric features of the original image beforehand. This embodiment uses the face landmark detector model from the Dlib library to process the original image after step S1. The detection process is performed. The Dlib model can accurately locate 68 key feature points of the face, covering core anatomical structures such as the jawline, eyebrows, eyes, nose, and mouth. The extracted two-dimensional coordinates of the 68 key points form a set. To facilitate backpropagation and gradient calculation within the deep learning framework, this coordinate set is converted into a floating-point tensor. The keypoint tensor As a baseline for the original facial geometry, it will be the target of subsequent adversarial attack optimization. That is, the optimization process aims to make the perturbed key points of the image as far away from the baseline position as possible.

[0061] Step S3: Construct a three-band frequency domain constraint mask. To address the technical challenge of balancing visual concealment and adversarial effectiveness with a single spatial constraint, this embodiment constructs a refined three-band frequency domain mask. Its frequency domain mask distribution and frequency band division are as follows: Figure 1 As shown. The human visual system (HVS) is more sensitive to low-frequency information such as contours and color blocks, but less sensitive to high-frequency texture details. Based on this characteristic, this embodiment divides the frequency domain into three regions: low-frequency, mid-frequency, and high-frequency, and applies differentiated constraint strengths. Mask The construction process is as follows:

[0062] First, determine the frequency domain division ratio parameters and set the low-frequency ratio. The mid-frequency ratio is 0.15. The value is 0.4. These two parameters define the coverage radius of different frequency components in the spectrum. (Based on the adjusted image size) and Calculate the specific boundary coordinates of each frequency band in the frequency domain matrix.

[0063] The formula for calculating the height boundary of the low-frequency region is:

[0064]

[0065] The formula for calculating the width boundary of the low-frequency region is:

[0066]

[0067] Similarly, the formula for calculating the height boundary in the mid-frequency region is:

[0068]

[0069] The formula for calculating the width boundary of the mid-frequency region is:

[0070]

[0071] in The symbol represents the floor function, ensuring that the coordinate values ​​are integer indices. Then, a dimension is initialized. mask matrix and for each of its positions Assign values ​​according to the following rules, where For row index, For column indexes:

[0072] For the low-frequency region, i.e. when and At that time, this region corresponds to the main energy and basic contour of the image. To maintain visual naturalness to the greatest extent, the mask values ​​are set to a strongly constrained state, i.e., assigned values. Here, 1 represents imposing the most stringent restrictions, aimed at suppressing the amplitude of disturbances in this frequency band as much as possible.

[0073] For the high-frequency region, i.e. when or At that time, this region corresponds to subtle noise or sharp edges in the image. The human eye is not sensitive to changes in this region, thus allowing for a relatively large degree of freedom in perturbation. The mask values ​​are set to a weakly constrained state, i.e., assigned values. Here, 0 represents no frequency domain amplitude limit or a very small limit, allowing countermeasures to be concentrated in this frequency band.

[0074] For the mid-frequency region between low and high frequencies, to avoid ringing effects caused by frequency cutoff and to achieve a smooth transition, this embodiment designs a dynamic attenuation function based on normalized Euclidean distance. First, the normalized distance of the current point relative to the low-frequency boundary is calculated:

[0075]

[0076]

[0077] Calculate the comprehensive distance factor :

[0078]

[0079] Finally, the mask values ​​for the mid-frequency region were determined: Through the above calculations, the constraint strength in the mid-frequency region decreases linearly from 1 at the low-frequency boundary to 0 at the high-frequency boundary, achieving a smooth transition from strong to weak constraints. This design both disrupts the diffusion model's understanding of facial semantic features and maintains the visual plausibility of the image.

[0080] Step S4: Generation and iterative optimization of the adversarial perturbation. This step is the core of the invention, aiming to generate an optimal perturbation that can both deceive the AI ​​editing model and maintain visual concealment from the human eye. Initialize a perturbation tensor with all zeros or tiny random values. Its size and Consistent. Then, an iterative optimization process begins; in this embodiment, the number of iterations is set to 200. In each iteration, the following operations are performed:

[0081] First, a multi-objective joint loss function is calculated to guide the update direction of the perturbation. The first loss term, the keypoint geometric difference loss, is then constructed. The current disturbance Overlay onto the original image Obtain the current attack image The keypoint coordinates of the attacked image were extracted using the aforementioned keypoint detection model. To completely destroy the structural features of the face, this embodiment requires maximizing the original key points. Key points after the attack The difference between them. Mean squared error (MSE) is used as the metric, and a negative value is taken to fit the minimization framework of gradient descent, as shown in the following formula:

[0082]

[0083] The physical meaning of this loss function is that it forces the perturbed image to appear to the AI ​​model as having undergone drastic distortion and displacement of the facial features, thereby destroying the information of the facial geometric structure.

[0084] Construct a second loss term, identity similarity loss. The pre-trained deep face recognition model, specifically CVLFACE, is used to extract features. Features of the original images are calculated separately. and attack image features To destroy identity information, the cosine similarity between the two needs to be minimized: This ensures that the protected image is as far away from the original identity as possible in the feature space.

[0085] Construct a third loss term, latent space offset loss. The latent space code is extracted using a variational autoencoder (VAE). The original latent code is then computed. Post-attack latent encoding The Euclidean distance between them is calculated, and negative values ​​are taken to maximize it.

[0086]

[0087] This term aims to disrupt the diffusion model's ability to preserve source image information in the latent space. The total loss function is obtained by weighted summing of the three loss terms:

[0088]

[0089] in and To balance the hyperparameters of the various weights, based on this total loss function... This invention solves for the optimal perturbation by the following optimization objective. :

[0090]

[0091] in, Budget for airspace disturbances.

[0092] Based on the total loss function, calculate the perturbation. The gradient is calculated, and the perturbation is updated using an optimization algorithm, specifically PGD or Adam. After completing the backpropagation of the gradient and the initial update of the perturbation value, the perturbation must be immediately subject to joint spatial-frequency domain constraints to prevent excessive perturbation from causing image quality degradation. This embodiment employs an intermittent projection strategy to balance computational efficiency and constraint effectiveness. Specifically, a complete joint spatial-frequency domain constraint is executed every 5 iterations, while only spatial constraints are executed in the remaining 4 iterations.

[0093] Spatial constraints employ the infinity norm to ensure that every pixel in the image is protected. The disturbance amplitude does not exceed the preset threshold. The specific operation involves setting the disturbance value... Cut to Within the range.

[0094] Frequency domain constraints involve domain transformations. This is achieved using a two-dimensional discrete cosine transform (DCT) matrix. and The current airspace disturbance Transform to the frequency domain:

[0095]

[0096] The elements of the DCT transformation matrix are defined as follows (where...) For frequency index, For spatial indexing, (For the corresponding dimension size)

[0097]

[0098] Obtain frequency domain coefficients Then, based on the mask generated in step S3 The perturbation is subjected to amplitude constraints in different regions. For the low-frequency region with a mask value of 1, this operation essentially detects the amplitude; if it exceeds the allowable range, it is attenuated. Alternatively, in another implementation, a penalty term is used to pull the low-frequency perturbation towards zero. For the high-frequency region with a mask value of 0, the perturbation is allowed to remain freely. After masking, the corrected frequency domain coefficients are restored back to the spatial domain using the inverse discrete cosine transform (IDCT) to obtain the perturbation that satisfies the frequency domain constraints. .

[0099] Through the above 200 iterations of optimization, a result was finally obtained that satisfies both It also satisfies the optimal perturbation that meets the frequency domain distribution characteristics.

[0100] Step S5: Generate and output the face protection image. The optimized perturbation... Overlaying back the original face image To obtain the protected image If the image was resized in step S1, the reverse operation is required. This is done using bilinear interpolation. From size Restored to original size To obtain the final face protection image :

[0101]

[0102] The protected image is visually almost identical to the original image. However, when subjected to malicious editing based on a diffusion model, its key geometric features are implicitly distorted, and its identity features are destroyed. The resulting edited image will exhibit facial structure distortion and unrecognizable identity, thus achieving the purpose of proactive protection. The facial image protection effect provided by this invention is compared to... Figure 2 As shown in the image, the protective performance is demonstrated in various editing scenarios, such as wearing glasses and having a goatee.

[0103] Example 2

[0104] This embodiment provides a facial image active protection system. The system is built upon the logic described in Embodiment 1 and can run on a computer device as a software component, hardware circuit, or a combination of both. The system aims to provide users with automated privacy protection services, enabling them to complete anti-fraud processing with a single click before uploading photos to the public network.

[0105] The active face image protection system mainly includes an image acquisition module 201, a geometric prior extraction module 202, a dual-domain constraint construction module 203, an adversarial optimization module 204, and a protective image generation module 205. Its system structure diagram is shown below. Figure 4 As shown.

[0106] The image acquisition module 201 receives the original face image to be protected. This module supports multiple formats, can process image data in common formats, and integrates a preprocessing unit. This preprocessing unit is responsible for converting the input image data into the tensor format required for the computation graph, performing normalization operations, and checking the parity of the dimensions. When the height or width of the input image is detected to be odd, the image acquisition module 201 automatically calls a bilinear interpolation algorithm to fine-tune the image, ensuring that its length and width are both even, thereby eliminating boundary obstacles for subsequent frequency domain transformation.

[0107] The geometric prior extraction module 202 is connected to the image acquisition module 201 and is used to extract high-precision facial structure information from the preprocessed image. This module incorporates a pre-trained facial landmark detection model, specifically using a detector from the Dlib library. When image data is input into this module, it outputs a tensor containing the coordinates of 68 landmarks. These landmarks precisely cover the facial contours, eyebrows, eyes, bridge of the nose, and lip areas, forming the unique geometric topology of the face and providing accurate targeted data for subsequent differential attacks.

[0108] The dual-domain constraint construction module 203 is responsible for generating constraints to limit the perturbation characteristics. This module includes a frequency domain mask generation unit, which dynamically calculates the boundaries of low-frequency, mid-frequency, and high-frequency regions based on preset low-frequency ratios of 0.15 and mid-frequency ratios of 0.4, combined with image size information. This module generates a three-band mask with smooth transition characteristics by calculating normalized Euclidean distance. This mask presents strong constraint values ​​in the low-frequency region, linearly attenuated values ​​in the mid-frequency region, and weak constraint values ​​in the high-frequency region. This module passes this mask to the adversarial optimization module 204 to ensure that the generated perturbation, while impairing machine recognition capabilities, preserves the visual quality perceived by the human eye to the greatest extent possible.

[0109] The adversarial optimization module 204 is the core computational unit of this system. This module receives the original image from the image acquisition module 201, keypoint data from the geometric prior extraction module 202, and a frequency domain mask from the dual-domain constraint construction module 203. Internally, this module runs an iterative optimization algorithm and calculates a multi-objective joint loss function in each iteration. It integrates a deep face recognition model to calculate identity similarity loss, a variational autoencoder to calculate latent space offset loss, and calculates keypoint geometric difference loss in real time. After updating the perturbation through gradient backpropagation, this module periodically transforms the perturbation to the frequency domain and applies mask constraints according to an interval projection strategy, while simultaneously implementing infinite norm truncation in the spatial domain, thereby generating the optimal adversarial perturbation that satisfies all security and quality requirements.

[0110] The protective image generation module 205 is used to superimpose the calculated optimal perturbation onto the original image and perform necessary post-processing. If the image size was adjusted during the input stage, this module is responsible for restoring the image to its original resolution, and the final output is a visually natural but algorithmically compromised face protection image.

[0111] Example 3

[0112] This embodiment provides an electronic device and a computer-readable storage medium for implementing the above-described active face image protection method at the hardware level.

[0113] This electronic device can specifically be a smartphone, personal computer, server, or embedded security terminal. For example... Figure 5 As shown, the electronic device includes a processor 301, a memory 302, and a communication bus 303. The communication bus 303 is used to enable communication between the processor 301 and the memory 302.

[0114] The processor 301 is the control center of the electronic device, connecting all parts of the device through various interfaces and lines. The processor 301 performs various functions and processes data by running or executing software programs or modules stored in the memory 302 and by calling data stored in the memory 302, thereby performing proactive protection processing on facial images. In this embodiment, the processor 301 is configured to perform operations including but not limited to: acquiring the original face image to be protected; generating adversarial perturbations intended to be superimposed on the original face image; iteratively optimizing the adversarial perturbations based on preset optimization objectives and constraints; calculating the geometric position differences of facial key points between the original face image and the image after superimposing the adversarial perturbations, and maximizing the geometric differences; transforming the adversarial perturbations to the frequency domain, and applying differential constraints to the coefficient values ​​of the adversarial perturbations in the frequency domain using a frequency domain mask, wherein the frequency domain mask is configured to suppress perturbation energy in low-frequency regions and allow higher perturbation energy in high-frequency regions; and limiting the maximum amplitude of each pixel in the adversarial perturbations in the spatial domain; and finally superimposing the adversarial perturbations that satisfy the optimization objectives and constraints onto the original face image to obtain a face protection image.

[0115] The memory 302 can be used to store software programs and modules. The memory 302 mainly includes a program storage area and a data storage area. The program storage area can store the operating system and at least one application program required for a given function, such as program instructions for performing the aforementioned facial landmark detection, discrete cosine transform, and gradient descent optimization. The data storage area can store data created based on the use of the electronic device, such as raw facial image data, generated frequency domain mask data, and pre-trained deep learning model parameters, specifically Dlib model parameters, CVLFACE model parameters, and VAE model parameters.

[0116] This embodiment also provides a computer-readable storage medium, specifically a non-volatile storage medium. The storage medium stores a computer program, which, when executed by the processor 301 of the electronic device, causes the electronic device to perform the active face image protection method described in Embodiment 1.

[0117] It should be noted that those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0118] Furthermore, this technical solution has good scalability. Although the above embodiments specifically mention the use of the Dlib library for 68-point keypoint detection, other keypoint detection models or systems containing more or fewer keypoints can be used in other alternative embodiments, as long as they can characterize the geometric features of the face. Similarly, the frequency domain three-band division ratio mentioned in the above embodiments, namely 0.15 for low frequency and 0.4 for mid frequency, is an optimal value based on experimental data. In image processing scenarios with different resolutions or compression rates, these ratio parameters can be adaptively adjusted according to actual needs without departing from the technical concept of this invention. Through this flexible configuration, this invention can adapt to the privacy protection needs of various application scenarios such as social media, public security, and remote authentication, solving the technical problems of existing technologies such as single protection dimension, incomplete destruction of biometric features, and poor stability of protection effect.

Claims

1. A method for active protection of facial images, characterized in that, Includes the following steps: Obtain the original face image to be protected; Generate adversarial perturbations designed to be superimposed onto the original face image; Based on preset optimization objectives and constraints, the anti-disturbance is iteratively optimized; The optimization objective is configured to: calculate the geometric position difference of facial key points between the original face image and the image after superimposing the adversarial perturbation, and maximize the geometric position difference to destroy the facial geometric structure features in the original face image to be protected; the constraints include a joint spatial-frequency domain constraint, which is configured to: transform the adversarial perturbation to the frequency domain, and apply differential constraints to the coefficient values ​​of the adversarial perturbation in the frequency domain using a frequency domain mask, wherein the frequency domain mask is configured to suppress the perturbation energy in the low-frequency region and not impose amplitude limits on the perturbation energy in the high-frequency region; and limit the maximum amplitude of each pixel in the adversarial perturbation in the spatial domain; superimpose the adversarial perturbation that satisfies the optimization objective and constraints onto the original face image to obtain a face protection image; in the step of applying differential constraints to the coefficient values ​​of the adversarial perturbation in the frequency domain using a frequency domain mask, the frequency domain mask... The specific construction process includes: determining the adjusted height of the original face image for frequency domain transformation. and adjusted width The adjusted height and adjusted width Even number; set low frequency ratio The mid-frequency ratio is 0.

15. The value is 0.4; based on the adjusted height. and adjusted width Calculate the low-frequency height boundary Low frequency bandwidth boundary Mid-frequency height boundary and the mid-frequency bandwidth boundary For the frequency domain mask Each position in Assign values ​​according to the following rules: when and At that time, it belongs to the low-frequency region, and a value is assigned. ;when or When it belongs to the high-frequency region, a value is assigned. ; When position When the frequency is between the low-frequency region and the high-frequency region, it belongs to the mid-frequency region. Based on the normalized distance dynamic attenuation constraint strength, the calculation formula is as follows: in, This indicates the floor function; In the step of iteratively optimizing the adversarial disturbance based on preset optimization objectives and constraints, a multi-objective joint optimization loss function is adopted. The multi-objective joint optimization loss function The formula is: in, To minimize the loss based on identity similarity, To maximize the loss for latent space offset, and These are the weight parameters; the goal of the iterative optimization is to find the optimal perturbation. , so that: And it satisfies: ;in, Budget for airspace disturbances; This represents the geometric difference loss function.

2. The active face image protection method according to claim 1, characterized in that, The step of maximizing the geometric position difference specifically includes: extracting the coordinates of 68 key points from the original face image using a facial landmark detection model. ; Calculate the key point coordinates of the superimposed image after the perturbation is applied. Constructing the geometric difference loss function Its formula is: in, This indicates the resistance to the disturbance. Indicates the first The coordinates of each key point are obtained by minimizing the geometric difference loss function. This is to maximize the difference in the geometric position of key points.

3. The active face image protection method according to claim 1, characterized in that, The identity similarity minimization loss The calculation formula is: in, Let be the feature vector of the original face image in the deep face recognition model. The feature vector of the image after the adversarial perturbation is superimposed in the deep face recognition model.

4. The active face image protection method according to claim 1, characterized in that, The latent space offset maximizes the loss The calculation formula is: in, The latent encoding of the original face image in a variational autoencoder (VAE) This is the latent encoding of the image after the adversarial perturbation is superimposed in the variational autoencoder (VAE).

5. The active face image protection method according to claim 1, characterized in that, The step of transforming the anti-perturbation method to the frequency domain uses a two-dimensional discrete cosine transform (DCT), and the calculation formula is as follows: ;in, For the current confrontation and disturbance, and The DCT transform matrix contains elements of the DCT transform matrix. Defined as: when hour, when hour, in, The height is adjusted to correspond to the original face image. or width , For frequency index, For spatial indexing.

6. The active face image protection method according to claim 1, characterized in that, In the step of iteratively optimizing the adversarial perturbation, an intermittent projection strategy is adopted: in every five iterations, the joint spatial-frequency dual-domain constraint is executed once; in the remaining four iterations, only the constraint of spatial domain limiting the maximum amplitude of each pixel in the adversarial perturbation is executed.

7. The active face image protection method according to claim 1, characterized in that, Following the step of acquiring the original face image to be protected, a preprocessing step is also included: converting the original face image into a tensor and performing normalization processing; checking the height of the original face image. and width Is it an odd number; if or If the height is odd, the original face image is adjusted to an even-sized value using bilinear interpolation. and adjusted width To meet the requirements of discrete cosine transform; after obtaining the face protection image, bilinear interpolation is used to restore the face protection image to its original height. and width .

8. A facial image active protection system, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the active face image protection method as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the active face image protection method as described in any one of claims 1 to 7.