A face multi-attribute editing method based on low-rank adaptation and progressive scheduling

By employing a low-rank adaptive and progressive scheduling method for multi-attribute face editing, an orthogonal attribute classification system and a diffusion model with injected low-rank matrices are constructed. Combined with the progressive scheduling framework PIP-Diff and multi-level regression correction, the problems of identity drift and background inconsistency in continuous face attribute editing are solved, achieving high-fidelity age and facial fat/thin attribute editing.

CN122155933APending Publication Date: 2026-06-05NORTHWEST UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHWEST UNIV
Filing Date
2026-03-06
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing deep learning-based continuous facial attribute editing technologies struggle to achieve high-quality age and facial weight attribute editing while maintaining facial semantic consistency, especially in large-scale editing where issues such as identity drift and background inconsistency can easily arise.

Method used

A low-rank adaptive and progressive scheduling face multi-attribute editing method is adopted. By constructing a 7×7 orthogonal attribute classification system and a semantically aligned dataset, a low-rank adaptive matrix is ​​injected into the U-Net network of the diffusion model. Combined with the progressive scheduling framework PIP-Diff and a multi-level regression correction mechanism, progressive path planning and adaptive parameter control are realized to ensure the stability and high fidelity of the editing process.

Benefits of technology

It generates high-quality edited results with a wide age and facial weight attributes under low training cost, with high identity similarity and image fidelity. It solves the problem of unnatural and inconsistent edited results in existing technologies and achieves high-fidelity multi-attribute decoupled editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122155933A_ABST
    Figure CN122155933A_ABST
Patent Text Reader

Abstract

The application discloses a face multi-attribute editing method based on low-rank self-adaption and progressive scheduling, comprising the following steps: 1, based on the FFHQ dataset, constructing a face regularization dataset of a 7*7 orthogonal attribute classification system and generating a semantic-aligned attribute text description, and preprocessing the image; 2, injecting a low-rank self-adaption matrix into a U-Net network attention layer of a pre-trained diffusion model to obtain a LoRA-FA attribute editing model; 3, jointly training the LoRA-FA by using a denoising loss based on a minimum signal-to-noise ratio min-SNR weight, an identity consistency loss, a perceptual similarity loss and a prior keeping loss; 4, in the inference process of the face multi-attribute editing, adopting a progressive scheduling execution framework PIP-Diff to perform anchor point feature extraction, geometric mask generation and path planning on the images of the Celeba-HQ dataset; and 5, introducing an adaptive parameter control and a multi-level regression correction mechanism to realize stable conversion of attributes in a feedback correction cycle. The application can generate high-quality results of large-span age and face fat-thin attribute editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of computer vision and artificial intelligence technology, specifically a method for editing multiple facial attributes based on low-rank adaptive and progressive scheduling. Background Technology

[0002] In research related to face recognition and generation, facial images occupy an important position in the field of computer vision due to their highly structured nature and rich emotional expression. Facial attributes can be broadly divided into discrete attributes and continuous value attributes. Continuous value attribute editing is a research task in computer vision and artificial intelligence, aiming to adjust specific continuous attributes of facial images while maintaining the consistency of other attributes and overall semantic information. Unlike traditional discrete attribute editing (such as gender, presence or absence of glasses, or whether a smile is present), continuous value attribute editing deals with attributes that change smoothly within a certain range (such as age, facial weight, skin color, and hair volume), and has the following characteristics: 1) Editing a specified target attribute within a specific range, and attribute editing usually requires a smooth transition between the initial state and the target state; 2) Preserving the facial identity information and facial details of the source image, and ensuring that the image background is free of artifacts or blurring, so that the edited image has fidelity and realism.

[0003] The application of deep learning in image editing has promoted the development of continuous value attribute editing algorithms for facial images, enabling continuous attribute editing technology to show broad application prospects in many practical scenarios. For example, in social media platforms, it can be used to provide personalized beautification and image customization services; in film and television production, it can support character age changes or makeup simulation; in archaeology and forensic medicine, it can be used to restore the appearance of ancient people or simulate the appearance of suspects over time; and in public safety and missing persons tracking, continuous appearance prediction models can also help improve recognition efficiency and accuracy.

[0004] Currently, in deep learning-based image editing algorithms, facial attribute editing techniques based on generative models, particularly GANs and their various variants, have achieved good results. The most widely used generative models are Generative Adversarial Networks (GANs) and Denoising Diffusion Probabilistic Models (DDPMs). GAN technology has a long history in image editing research; however, GAN models are prone to convergence during training and may encounter problems such as vanishing gradients, exploding gradients, and overfitting. Denoising Diffusion Models, on the other hand, are very stable during training and have been increasingly applied to image editing tasks. With the help of additional conditions (e.g., classifier-guided image generation, text-based image generation), using Denoising Diffusion Models to edit images significantly improves the training efficiency of image editing models, bringing new possibilities to the field. However, due to the inherent diversity of Denoising Diffusion Models, maintaining facial semantic consistency remains a challenge. Modifying one specific attribute can have unexpected effects on other attributes, leading to unintended or unnatural editing results. Furthermore, existing Denoising Diffusion Models are more suited to handling discrete attributes, and their modeling and fine-tuning mechanisms for continuous attributes are still immature, lacking a unified and universal solution.

[0005] Age and facial weight are two key continuous-value attributes of the face that have a significant impact on appearance and are easily quantifiable. Their changes typically exhibit strong biological and visual correlations. Compared to other continuous-value attributes, the range of variation for age and weight is well-defined and easy to control. Therefore, studying the editing of these two attributes helps maintain consistency between facial structure and overall image during face image editing. Furthermore, age and weight attributes have broad practical needs in various application scenarios, such as personalized beautification on social media platforms and age transformation of film and television characters. It is evident that researching continuous-value facial attribute editing techniques based on these two attributes not only helps promote the development of image generation models in the direction of refined control but also has important technical support value for multiple practical application scenarios, possessing significant research significance and application prospects.

[0006] However, how to achieve high-quality and highly controllable editing of facial age and weight attributes remains a research problem with both theoretical depth and application potential. Summary of the Invention

[0007] The purpose of this invention is to provide a face multi-attribute editing method based on low-rank adaptive and progressive scheduling, which can generate high-quality results for editing facial fat and thin attributes across a wide range with low training costs.

[0008] This invention is achieved through the following technical solution: A face multi-attribute editing method based on low-rank adaptive and progressive scheduling includes the following steps: Step 1: Based on the FFHQ dataset, construct the FFHQ-Regulation-Dataset, a face regularization dataset with a 7×7 orthogonal attribute classification system, generate semantically aligned attribute text descriptions, and then preprocess the images. Step 2: Inject a low-rank adaptive matrix into the attention layer of the pre-trained diffusion model's U-Net network to obtain the LoRA-FA attribute editing model. The process is as follows: Step 2.1: Freeze the weights of the pre-trained diffusion model Stable Diffusion v1.5. ; Step 2.2: Query matrix of the cross-attention layer and self-attention layer of the pre-trained diffusion model's U-Net network. Key matrix Value matrix and output layer linear mapping Inject a low-rank adaptive matrix into the pre-trained diffusion model and update the weights to... , is represented as: In the formula; LoRA scaling factor; Both B and B are low-rank matrices, and , r is the rank parameter. For rank; and They represent and the size of B; Represents the set of real numbers; Step 3: Using the preprocessed regularized dataset FFHQ-Regulation-Dataset as the training set, the LoRA-FA attribute editing model is jointly trained using denoising loss, identity consistency loss, perceptual similarity loss and prior preservation loss based on minimum signal-to-noise ratio min-SNR weights. Step 4: In the inference process of multi-attribute face editing, the progressive scheduling execution framework PIP-Diff is used to perform anchor point feature extraction, geometric mask generation, and path planning operations on images in the Celeba-HQ dataset. The process is as follows: Step 4.1: Extract anchor features: Extract images from the Celeba-HQ dataset using the pre-trained InceptionResnetV1. A 512-dimensional deep identity feature vector is generated, which is used as the identity anchor point; simultaneously, images from the Celeba-HQ dataset are extracted using DINOv2. The intermediate layer feature map is used as the facial detail anchor point, and the identity anchor point and facial detail anchor point are used as global anchor points, with a set anchor return frequency. If the current iteration step number satisfy The system forcibly increases the weight of identity anchor points in the prompts and performs linear interpolation fusion on the latent variables of the starting anchor point of the current iteration step to obtain the updated latent variables. This will serve as the starting point for the next noise reduction step. Step 4.2: Generate a geometric mask: Use MediaPipe, a deep learning-based facial landmark detector, to extract a set of 468 3D facial key points from the current image I. Extract the corresponding index subset based on the facial contour boundaries. Calculate its center point and with the center point Based on the expansion factor Radializing the pixel coordinates in image I outwards yields the coordinates of the outwardly radiating pixels, thus constructing an expanded convex polygon mask matrix for the face. ; Simultaneously, key point sets of the left eye, right eye, eyebrows, and mouth were extracted to construct an eye protection mask. and mouth protective mask The primary mask matrix of the face image , is represented as: In the formula: and They are respectively and Mask weights; Will via the kernel size The final mask matrix is ​​obtained by applying a Gaussian smoothing filter. , is represented as: In the formula: This is a Gaussian blur operator used to feather binary masks, eliminating stitching marks caused by hard edges and enabling a smooth transition between the editing area and the background. Step 4.3, Path Planning: PIP-Diff predefines an ordered topological attribute chain of age and facial plumpness, and plans the face attribute editing task as a progressive path on the chain; Step 5: Introduce adaptive parameter control and multi-level regression correction mechanism to achieve stable attribute transformation in the feedback correction loop. The process is as follows: Step 5.1: Determine the current step size: Obtain the original image. The latent variable z is calculated based on the asymptotic path generated in step 4.3, according to the starting and ending values ​​of the target attribute. The editing task is decomposed into N steps. It is determined whether the current step has reached the ending value. If so, step 5.2 is executed; otherwise, step 5.7 is skipped. Step 5.2, Adaptive Parameter and Prompt Word Construction: Determine the denoising intensity based on the current step size progress to serve as the adaptive parameter, and generate the corresponding text prompt words; Step 5.3, Periodic Anchoring Mechanism: Introduce a periodic anchoring mechanism and set the anchoring frequency. Every experience If the step triggers the anchoring mechanism, the weight of the current latent space's identity anchor point will be forcibly increased; otherwise, the weight of the current latent space's identity anchor point will remain unchanged. Step 5.4, Latent Space Feature Fusion and LoRA-FA Attribute Editing: Inject attribute correction parameters into the attention layer of the U-Net in the LoRA-FA attribute editing model, and then... M Under the constraints, the diffusion model local repair algorithm is executed to generate candidate edited images for the current step. , is represented as: In the formula: The previous frame image, Edit the current denoised output of the model for LoRA-FA attributes; Step 5.5, Multi-level Regression Correction: Candidate images are corrected through multi-level regression. The process of performing multi-level verification is as follows: Step 5.5.1: Construct three independent perceptual metrics, including an identity extractor, a detail extractor, and an expression extractor, to calculate similarity including identity features. Global detail similarity Similarity to facial expressions Perception indicators, including those included. Step 5.5.2: Assess facial expression similarity sequentially. Global detail similarity Similarity to identity features Perform a judgment and verification, if , and If all three perception indicators are below the set threshold, it means that all three perception indicators have passed the verification, and the process proceeds directly to step 5.7; if , and If any item fails the validation, proceed to step 5.6; Step 5.6, Adaptive Three-Image Hybrid Correction Strategy: If any of the perception indicators in Step 5.5 fails the verification, the regression ratio coefficient is dynamically allocated according to the level of the verification failure, and adaptive three-image hybrid correction is performed to obtain the corrected image. , is represented as: In the formula: This is a smoothing coefficient used to prevent large single-step changes from causing visual abrupt changes or flickering. This is the anchorage coefficient; Step 5.7, Status Update and Indicator Recording: Record candidate images Or the image corrected in step 5.6 As a reference benchmark image for the next step, the perception index is updated in real time, and the LoRA scaling factor or denoising intensity of subsequent steps is dynamically adjusted according to the feedback signal until all evolutionary steps are completed, and the final multi-attribute edited face image is output.

[0009] Furthermore, the process of step 1 is as follows: Step 1.1: Construct a 7×7 orthogonal attribute classification system: Based on the principle of independence of facial semantic features, construct a 7×7 orthogonal attribute matrix, which includes two independent dimensions: age and facial plumpness. The age dimension includes infant, child, adolescent, adult, middle-aged, elderly, and very old; the facial plumpness dimension includes extremely thin, thin, slightly thin, normal, slightly plump, plump, and obese. Through the cross-combination of these two dimensions, age and facial plumpness, define 49 core attribute nodes. Step 1.2, Original Image Filtering and Pre-labeling: Select stable face samples from the public dataset FFHQ, match 15 to 20 face samples for each attribute node, and construct a face image regularization dataset FFHQ-Regulation-Dataset; Step 1.3: Generate semantically aligned attribute text descriptions: For each image in the regularization dataset FFHQ-Regulation-Dataset, the image is renamed according to its position in the 7×7 orthogonal attribute matrix, following the rule of "age category_facial fatness category_number". During the training preparation phase, the program directly reads the image file name and splits the file name with underscores, automatically extracting the two key labels of age and facial fatness attributes from the file name. After obtaining the labels, the training script generates the corresponding text prompts in real time, which are the attribute text descriptions. The attribute text descriptions are converted into feature vectors through the CLIP text encoder, serving as conditional guidance signals during the training phase, so that the image content and semantic descriptions are accurately aligned in the latent space. Step 1.4, Image Preprocessing: Perform face alignment on all images in the regularization dataset FFHQ-Regulation-Dataset, uniformly crop and scale the images to a resolution of 512×512, then perform color space normalization on the images, and use the Laplacian operator to enhance the detailed features of the facial texture region.

[0010] Furthermore, step 3 is as follows: Step 3.1: Calculate the reciprocal of the signal-to-noise ratio (SNR). And set the truncation hyperparameter To limit the penalty weight , and They are represented as follows: In the formula: This is the cumulative multiplication factor of the noise scheduler; In the formula: Minimize the function; The truncation hyperparameter represents the minimum signal-to-noise ratio (min-SNR). The reciprocal of the signal-to-noise ratio (SNR) is expressed as: Then calculate the denoising loss based on the minimum signal-to-noise ratio (min-SNR) weight. , is represented as: In the formula: It is a moment Noisy latent variables; Represents real noise; It is a U-Net network with a pre-trained diffusion model loaded with LoRA weights; These are images from the regularization dataset FFHQ-Regulation-Dataset. This is the attribute text description corresponding to the image; Represents the CLIP model; Step 3.2: Decode the predicted latent variables of the LoRA-FA attribute editing model using VAE. Then, by utilizing the feature consistency constraint of the ArcFace face recognition network, identity feature vectors are extracted, and then cosine distance is used. Measured to calculate identity consistency loss , is represented as: In the formula: For the frozen facial recognition network; Step 3.3: Utilize LPIPS network constraints to evaluate the image. and Differences in the deep perceptual feature space are then used to calculate the perceptual similarity loss. , is represented as: In the formula: It is a frozen LPIPS network; Step 3.4: Divide each batch into instance samples and class samples. Specifically: use the samples in the regularization set FFHQ-Regulation-Dataset as instance samples, the FFHQ dataset as the prior set, and the general face images within it as class samples. Calculate the denoising loss for each instance sample, then sum them using weighted averages to obtain the prior preservation loss. , is represented as: In the formula: It is the denoising loss corresponding to the instance samples from the regularization set FFHQ-Regulation-Dataset; These are prior preservation weights, used to constrain the model from deviating from the face category; It is the denoising loss corresponding to the general face images from the prior set. and Both use min-SNR weighted denoising loss; Step 3.5: Combine the denoising loss, identity consistency loss, perceptual similarity loss, and prior preservation loss based on the minimum signal-to-noise ratio (min-SNR) weights to obtain the total loss. , is represented as: In the formula: and Loss of identity consistency and perceptual similarity loss The weight parameters.

[0011] Furthermore, the updated latent variables in step 4.1 Represented as: In the formula: It is a latent variable representing the starting anchor point of the current iteration step. It is the fusion coefficient. It is the number of iterations. yes The gradient.

[0012] Further, the pixel coordinates of the outward radiation in step 4.2 are represented as follows: In the formula: Representing an image The pixel coordinates in the image; This represents the coordinates of the pixels radiating outwards.

[0013] Furthermore, the specific process of step 4.3 is as follows: Step 4.3.1: Define two one-dimensional ordered topological attribute chains. and , respectively represented as: Step 4.3.2, based on and The path planning module constructs the optimal editing path through index calculation, assuming the source attribute index is... The target attribute index is The algorithm logic of the path planning module is as follows: 1) If Then, the interval is intercepted along the positive direction. All state nodes are treated as an ordered topological attribute chain; 2) If Then, the interval is truncated in the reverse direction. The reversed nodes are treated as an ordered topological attribute chain; 3) If If the path is empty, no editing is required. PIP-Diff employs a phased, sequential execution strategy, prioritizing the gradual transformation of the age dimension. Once the age characteristics have stabilized, the gradual transformation of the facial fatness dimension is then performed based on the current results.

[0014] Furthermore, the specific process of step 5.2 is as follows: edit the span according to the attributes. And introduce the number of iterations It exhibits an exponential decay mechanism, and the attribute editing span... With the noise reduction intensity of step They are represented as follows: In the formula: As the reference strength; This represents the total number of steps. For normalized attribute editing span, indexed by source attribute With target attribute index The distance determines; It is the attenuation factor; Indicates the first noise reduction intensity of step That is, adaptive parameters.

[0015] Furthermore, the specific process of step 5.3 is as follows: introduce a periodic anchoring mechanism and set the anchoring frequency. The default setting is 3, meaning that the anchoring mechanism is triggered every 3 steps, forcibly changing the weight of the current latent space's identity anchor point. An improvement of 0.15, with a maximum improvement of 0.5, ensures that the spreading manifold remains within the neighborhood of the original identity during the shift towards the target semantics, thus preventing identity loss caused by long-path editing.

[0016] Furthermore, the specific process of step 5.5.1 is as follows: Step 5.5.1.1: Use a ResNet-50 network pre-trained on ImageNet with fully connected layers removed as the backbone network, and input the candidate edit image. and initial image After being cropped and normalized by facial bounding boxes, it is mapped to a high-dimensional feature vector. Define the currently generated candidate edit graph. With the initial image The cosine similarity is used as its identity feature similarity to quantify the closeness of two face images in deep identity semantics. Represented as: Step 5.5.1.2: Use Vision Transformer as a detail feature extractor to capture candidate editable images. and initial image High-frequency detail features, including skin texture and hair direction. Used to calculate candidate edit graphs With the initial image The cosine similarity is used as its global detail similarity to detect whether an image has undergone wax-like appearance or texture loss. Represented as: Step 5.5.1.3: Extract candidate editable images using ResNet-50. and initial image The intermediate layer features are then processed by adaptive average pooling to obtain expression features that are highly sensitive to local facial muscle tension. It is used to calculate facial expression similarity to capture subtle changes in the shape of the corners of the mouth and eyes. , is represented as: .

[0017] Furthermore, the specific process of step 5.5.2 is as follows: first, determine... If the value is lower than the threshold of 0.7, then determine... If the validation fails, proceed to step 5.6. If so, then proceed to the next step. If the value is lower than the threshold of 0.65, then determine... If the validation fails, proceed to step 5.6. If so, continue with the following checks. If the value is lower than the threshold of 0.55, then determine... If the verification fails, proceed to step 5.6; if all three perception indicators pass the verification, proceed directly to step 5.7.

[0018] The present invention has the following beneficial technical effects: The LoRA-FA attribute editing model constructed in this invention requires only a very small number of parameter fine-tuning to establish accurate attribute-semantic mapping. It has a small number of training parameters and a short training time. In the inference stage of multi-attribute face editing, it introduces a PIP-Diff progressive scheduling and multi-level verification and correction mechanism. This allows LoRA-FA to force the injection of original identity features through a periodic anchoring mechanism during the multi-step incremental evolution of the latent space, under the constraint of the facial mask. Combined with an adaptive three-image hybrid correction strategy, it dynamically adjusts the fusion ratio of the current candidate image, the previous frame image, and the anchor point image, thus enabling it to handle large attribute jumps. This invention better locks onto the core skeleton and microscopic identity features of the face, improves the high fidelity of the generated image against complex backgrounds, and solves the problem of identity drift and inconsistency with the background during large-scale attribute editing. It achieves high-fidelity decoupled editing of multiple facial attributes. Compared with the edited images generated by existing LATS, DiffusionCLIP, DeltaEdit, and Face2Diffusion, the facial images generated by this invention have higher identity similarity and SSIM scores, and the lowest LPIPS scores. It is evident that this invention achieves low-cost and high-fidelity controllable editing of age and facial fat attributes. Attached Figure Description

[0019] Figure 1 The age and facial fatness attribute matrix diagram of this invention; Figure 2 : A schematic diagram of the LoRA-FA model of the present invention; Figure 3 : Flowchart of the PIP-Diff progressive framework of the present invention; Figure 4 : A schematic diagram of the microstructure of the low-rank adaptive module LoRA of the present invention; Figure 5 : A schematic diagram of the geometric prior-based facial mask generation and its use for latent space feature fusion according to the present invention; Figure 6 The age attribute editing diagram generated by this invention; Figure 7 The facial slimness / thinness attribute editing diagram generated by this invention. Detailed Implementation

[0020] The present invention will be further described in detail below with reference to specific embodiments. These descriptions are for explanation purposes only and are not intended to limit the scope of the invention.

[0021] A face multi-attribute editing method based on low-rank adaptive and progressive scheduling includes the following steps: Step 1: Based on the FFHQ dataset, construct the FFHQ-Regulation-Dataset, a 7×7 orthogonal attribute classification system for face regularization, and generate semantically aligned attribute text descriptions. Then, preprocess the images as follows: Step 1.1: Construct a 7×7 orthogonal attribute classification system: Based on the principle of independence of facial semantic features, construct a system such as... Figure 1 The 7×7 orthogonal attribute matrix shown contains two independent dimensions: age and facial fatness. The age dimension includes infancy, child, adolescence, young adult, middle-aged, senior adult, and elderly. The facial fatness dimension includes very slim, slim, slightly slim, normal, slightly overweight, overweight, and obese. By combining these two dimensions, 49 core attribute nodes are defined, providing a discretized label benchmark for subsequent decoupled training. Step 1.2, Original Image Selection and Pre-labeling: Select stable face samples from the public dataset FFHQ. Requirements: The image is a single-person scene, the face is clearly visible, there is little occlusion, and the subject occupies a reasonable proportion in the picture. Then, match 15 to 20 face samples for each attribute node to construct a regularized dataset FFHQ-Regulation-Dataset containing 800 high-quality face images. Step 1.3: Generate semantically aligned attribute text descriptions: For each image in the regularized dataset FFHQ-Regulation-Dataset, the image is renamed according to its position in the 7×7 orthogonal attribute matrix, following the rule of "age category_facial slimness category_number", for example, child_very slim_24721. During the training preparation phase, the program directly reads the image file name and splits the file name with underscores, automatically extracting the two key labels of age and facial slimness attributes from the file name. After obtaining the labels, the training script generates the corresponding text prompts in real time, which are the attribute text descriptions. The attribute text descriptions are converted into feature vectors through the CLIP text encoder, serving as conditional guidance signals during the training phase, so that the image content and semantic descriptions are accurately aligned in the latent space. Step 1.4, Image Preprocessing: Perform face alignment on all images in the regularization dataset FFHQ-Regulation-Dataset, uniformly crop and scale the images to a resolution of 512×512, then perform color space normalization on the images, and use the Laplacian operator to enhance the detail features of the facial texture region to improve the model's sensitivity to capturing wrinkles and subcutaneous fat distribution features. Step 2: Inject the following into the attention layer (FA) of the U-Net network in the pre-trained diffusion model (Stable Diffusion v1.5): Figure 4 The low-rank adaptive matrix (LoRA) shown yields the following results: Figure 2 The LoRA-FA attribute editing model described above has the following process: Step 2.1: Freeze the weights of the pre-trained diffusion model. ; Step 2.2: Diffusion of the query matrix to the cross-attention and self-attention layers of the pre-trained U-Net network. Key matrix Value matrix and output layer linear mapping Inject a low-rank adaptive matrix into the pre-trained diffusion model and update the weights to... , is represented as: In the formula; LoRA scaling factor; Both B and B are low-rank matrices, and , r is the rank parameter. For rank; and They represent and the size of B; Represents the set of real numbers; In the fine-tuning configuration, this embodiment sets the rank parameter r to 16 and the scaling factor... Set it to 32, and set a non-zero Dropout probability to prevent overfitting during subsequent training; Step 3: Using the preprocessed regularized dataset FFHQ-Regulation-Dataset as the training set, the LoRA-FA attribute editing model is jointly trained using denoising loss, identity consistency loss, perceptual similarity loss, and prior preservation loss based on minimum signal-to-noise ratio (min-SNR) weights. The process is as follows: Step 3.1: Calculate the reciprocal of the signal-to-noise ratio (SNR). And set the truncation hyperparameter To limit the penalty weight , and They are represented as follows: In the formula: This is the cumulative multiplication factor of the noise scheduler; In the formula: Minimize the function; The truncation hyperparameter represents the minimum signal-to-noise ratio (min-SNR). The reciprocal of the signal-to-noise ratio (SNR) is expressed as: Then calculate the denoising loss based on the minimum signal-to-noise ratio (min-SNR) weight. , is represented as: In the formula: It is a moment Noisy latent variables; Represents real noise; It is a U-Net network with a pre-trained diffusion model loaded with LoRA weights; These are images from the regularization dataset FFHQ-Regulation-Dataset. This is the attribute text description corresponding to the image; Represents the CLIP model; Step 3.2: Decode the predicted latent variables of the LoRA-FA attribute editing model using VAE. Then, by utilizing the feature consistency constraint of the ArcFace face recognition network, identity feature vectors are extracted, and then cosine distance is used. Measured to calculate identity consistency loss , is represented as: In the formula: For the frozen facial recognition network; Step 3.3: Utilize LPIPS network constraints to evaluate the image. and Differences in the deep perceptual feature space are then used to calculate the perceptual similarity loss. , is represented as: In the formula: It is a frozen LPIPS network; Step 3.4: Divide each batch into instance samples and class samples. Specifically: use the samples in the regularization set FFHQ-Regulation-Dataset as instance samples, the FFHQ dataset as the prior set, and the general face images within it as class samples. Calculate the denoising loss for each instance sample, then sum them using weighted averages to obtain the prior preservation loss. , is represented as: In the formula: It is the denoising loss corresponding to the instance samples from the regularization set FFHQ-Regulation-Dataset; These are prior preservation weights, used to constrain the model from deviating from the face category; It is the denoising loss corresponding to the general face images from the prior set. and Both employ a weighted denoising loss based on the minimum signal-to-noise ratio (min-SNR) weight. Step 3.5: Combine the denoising loss, identity consistency loss, perceptual similarity loss, and prior preservation loss based on the minimum signal-to-noise ratio (min-SNR) weights to obtain the total loss. , is represented as: In the formula: and Loss of identity consistency and perceptual similarity loss The weight parameters are used to represent the importance of different loss functions to the model; Step 4: In the reasoning process of multi-attribute editing of faces, the following methods are used: Figure 3 The progressive scheduling execution framework PIP-Diff, as shown, performs anchor point feature extraction, geometric mask generation, and path planning operations. The process is as follows: Step 4.1: Extract anchor features: Extract images from the Celeba-HQ dataset using the pre-trained InceptionResnetV1. A 512-dimensional deep identity feature vector is generated, which is used as the identity anchor point; simultaneously, images from the Celeba-HQ dataset are extracted using DINOv2. The intermediate layer feature map is used as the facial detail anchor point, and the identity anchor point and facial detail anchor point are used as global anchor points, with a set anchor return frequency. When the current iteration step satisfy At this time, the weight of the identity anchor point in the prompt is forcibly increased, and the latent variables of the starting anchor point of the current iteration step are linearly interpolated and fused to calibrate the accumulated latent space drift, thus obtaining the updated latent variables. , represented as: In the formula: It is a latent variable representing the starting anchor point of the current iteration step. It is the fusion coefficient. It is the number of iterations. yes The gradient; After calibration, the updated version will be available. As the starting point for the next noise reduction; Step 4.2, Generate the geometry mask: (e.g.) Figure 5 As shown, the MediaPipe deep learning-based facial landmark detector extracts a set of 468 3D facial landmarks from the current image I. Extract the corresponding index subset based on the facial contour boundaries. Calculate its center point and with the center point Based on the expansion factor Radializing the pixel coordinates in image I outwards, as follows: In the formula: Representing an image The pixel coordinates in the image; Represents the coordinates of the pixels radiating outwards; By image Radial outwards from the pixel coordinates in the matrix to construct an expanded convex polygon mask matrix for the face. It can physically isolate the face region, so that the denoising operation of the diffusion model is only effective within the face region, thereby completely freezing the pixels in the background region. To prevent the diffusion model from altering the eye's gaze and mouth structure during redrawing, a local protection mechanism is introduced. This involves extracting local keypoint sets from the left eye, right eye, eyebrows, and mouth to construct an eye protection mask. and mouth protective mask Simultaneously, extremely low mask weights are assigned to the local protected areas to obtain the primary mask matrix of the face image. , is represented as: In the formula: and They are respectively and Mask weights; Will via the kernel size The final mask matrix is ​​obtained by applying a Gaussian smoothing filter. , is represented as: In the formula: This is a Gaussian blur operator used to feather binary masks, eliminating stitching marks caused by hard edges and enabling a smooth transition between the editing area and the background. Step 4.3, Path Planning: In order to achieve a smooth transition of the face image from the source state to the target state, PIP-Diff predefines an ordered topological attribute chain of age and facial fatness, and plans the face attribute editing task as a progressive path on the chain, so that during the reasoning process, only adjacent state changes occur at each step, with small semantic jumps and stable structure. Step 4.3.1: Define two one-dimensional ordered topological attribute chains. and , respectively represented as: Step 4.3.2: Based on these two ordered topological attribute chains and The path planning module constructs the optimal editing path through index calculation, assuming the source attribute index is... The target attribute index is The algorithm logic of the path planning module is as follows: 1) If Then, the interval is intercepted along the positive direction. All state nodes are treated as an ordered topological attribute chain; 2) If Then, the interval is truncated in the reverse direction. The reversed nodes are treated as an ordered topological attribute chain; 3) If If the path is empty, no editing is required. To avoid entanglement and conflict between age and facial fatness in the multidimensional potential subspace, PIP-Diff adopts a phased serial execution strategy, that is, it first executes the gradual transformation of the age dimension, and after the age feature stabilizes, it executes the gradual transformation of the facial fatness dimension based on the current result. Step 5, as follows Figure 3 As shown, an adaptive parameter control and multi-level regression correction mechanism are introduced to achieve stable attribute transformation in the feedback correction loop. The process is as follows: Step 5.1: Determine the current step size: Obtain the latent variable z of the original image, calculate the evolution trajectory based on the initial and final values ​​of the target attribute according to the progressive path generated in step 4.3, decompose the editing task into N steps, determine whether the current step size has reached the final value, if so, proceed to step 5.2; otherwise, jump to step 5.7. Step 5.2, Adaptive Parameter and Prompt Word Construction: The denoising intensity is determined based on the current step size progress and used as the adaptive parameter to generate the corresponding text prompt words. The specific process is as follows: Edit the span according to the attributes. And introduce the number of iterations It exhibits an exponential decay mechanism, and the attribute editing span... With the noise reduction intensity of step They are represented as follows: In the formula: As the reference strength; This represents the total number of steps. For normalized attribute editing span, indexed by source attribute With target attribute index The distance determines; It is the attenuation factor; Indicates the first noise reduction intensity of step , i.e., adaptive parameters; When the attribute editing span When the value is large, a higher denoising intensity is applied in the initial stage to break the local minimum. In this embodiment, When the value is greater than 5, a higher noise reduction intensity is assigned; as the number of steps increases... As the intensity increases, the intensity gradually decreases through the attenuation factor, allowing the editing process to smoothly transition from structural reshaping to fine-tuning of details, preventing the destruction of the generated features in the later stages; Step 5.3, Periodic Anchoring Mechanism: Introduce a periodic anchoring mechanism and set the anchoring frequency. The default setting is 3, meaning that the anchoring mechanism is triggered every 3 steps, forcibly changing the weight of the current latent space's identity anchor point. The weight is increased by 0.15, with a maximum increase of 0.5. This mechanism is similar to the periodic feedback correction in control theory, which ensures that the diffusion manifold remains within the neighborhood of the original identity as it shifts toward the target semantics, thus effectively avoiding identity loss caused by long path editing; otherwise, the weight of the identity anchor point in the current latent space remains unchanged. Step 5.4, as follows Figure 5 As shown, latent space feature fusion and LoRA-FA attribute editing: Attribute correction parameters are injected into the attention layer of the U-Net in the LoRA-FA attribute editing model. Under the constraint of the mask matrix M, the diffusion model local inpainting algorithm is executed to generate the candidate edited image for the current step. , is represented as: In the formula: The previous frame image, Edit the current denoised output of the model for LoRA-FA attributes; By constraining the mask matrix M, the editing scope is strictly focused on the effective topological region of the face, so that the core pixels of the eyes and mouth are preserved, significantly improving the stability of the hairstyle outline and background edge, and effectively blocking the spread of errors to the background region. Step 5.5, Multi-level Regression Correction: Candidate images are corrected through multi-level regression. The process of performing multi-level verification is as follows: Step 5.5.1: Construct three independent perceptual metrics, including an identity extractor, a detail extractor, and an expression extractor, to calculate similarity including identity features. Global detail similarity Similarity to facial expressions Perception indicators, including those included. Step 5.5.1.1: Identity Feature Similarity Reflecting the uniqueness of each individual, a ResNet-50 pre-trained on ImageNet with fully connected layers removed is used as the backbone network, with the input candidate edit image... and initial image After being cropped and normalized by facial bounding boxes, it is mapped to a high-dimensional feature vector. Define the currently generated candidate edit graph. With the initial image The cosine similarity is used as its identity feature similarity to quantify the closeness of two face images in deep identity semantics. Represented as: Step 5.5.1.2, Global Detail Similarity Reflecting the consistency of skin texture and lighting, the calculation incorporates a Vision Transformer as a detail feature extractor. The global self-attention mechanism of the Vision Transformer can effectively capture candidate editable images. and initial image High-frequency detail features, including skin texture and hair direction. Used to calculate the currently generated candidate edit graph With the initial image The cosine similarity is used as its global detail similarity to detect whether an image has undergone wax-like appearance or texture loss. Represented as: Step 5.5.1.3, Facial Similarity Reflecting the tension state of facial muscles, and because deep features tend to capture global identity information, ResNet-50 is used to extract candidate edit images. and initial image The intermediate layer features are used to obtain an intermediate layer feature map, which is more sensitive to shape and edges. Then, adaptive average pooling is used to obtain expression features that are highly sensitive to local facial muscle tension. It is used to calculate facial expression similarity to capture subtle changes in the shape of the corners of the mouth and eyes. , is represented as: Step 5.5.2: Assess identity feature similarity. Global detail similarity Similarity to facial expressions The following checks are performed, with the following priority: First, facial expression similarity is checked. If the value is lower than the threshold of 0.7, then determine... If the verification fails, the facial expression similarity verification is considered successful, and then a detail similarity verification is performed. Is it below the threshold of 0.65? If not, proceed with the determination. If the verification fails, then the detail similarity verification is considered to have passed, and the next step is to determine... If the value is below the threshold of 0.55, then the identity feature similarity is determined. If the verification fails, then the detail similarity verification is considered to have passed. If all three perception indicators pass the verification, proceed to step 5.7, provided the identity feature similarity is high. Global detail similarity Or facial expression similarity If any item fails the validation, proceed to step 5.6; Step 5.6, Adaptive Three-Image Hybrid Correction Strategy: If any similarity check in Step 5.5 fails, the regression ratio coefficient is dynamically allocated according to the level of failure, and adaptive three-image hybrid correction is performed to obtain the corrected image. , is represented as: In the formula: This is a smoothing coefficient used to prevent large single-step changes from causing visual abrupt changes or flickering. These are the anchoring coefficients, which force image features to regress to the neighborhood of the original identity to combat the forgetting problem generated by long sequences. These two parameters are dynamically adjusted according to the priority matrix. Step 5.7, Status Update and Indicator Recording: Record candidate images Or the image corrected in step 5.6 As a reference benchmark image for the next step, the perception index is updated in real time, and the LoRA scaling factor or denoising intensity of subsequent steps is dynamically adjusted according to the feedback signal until all evolutionary steps are completed, and the final multi-attribute edited face image is output.

[0022] The server used in this embodiment is configured with a Windows system, an NVIDIA GeForce RTX4090 graphics card, and a 13th Gen Intel(R) Core(TM) i7-13700KF CPU. It runs in an environment of PyTorch 2.3.1, CUDA 12.1, and Python 3.11.7. The shortest training time for existing models LATS, DiffusionCLIP, DeltaEdit, and Face2Diffusion is about 40 hours. The total training time of the LoRA-FA attribute editing model proposed in this embodiment is about 4 hours, which is at least 90% shorter than the existing models.

[0023] The method proposed in this embodiment achieves the following effect on editing the age attribute of facial images: Figure 6 As shown, this image illustrates the gradual change of a person from infancy to old age, clearly showing age characteristics and realistic skin texture. Even during large-scale editing, there was no image degradation, and the effect was quite good. The method proposed in this embodiment achieves the following effect on editing the facial fat and thin attributes of face images: Figure 7 As shown, the image depicts the gradual transformation of a person from extremely thin to obese. The model accurately captures the physiological characteristics of changes in facial weight, and maintains a high degree of consistency in the person's core facial features, hairstyle, clothing, background, and identity as the facial contours change, resulting in a good effect.

[0024] The quantitative analysis of the quality evaluation results of images generated by LATS, DiffusionCLIP, DeltaEdit, Face2Diffusion, and this embodiment is shown in Table 1 below. It can be seen that: First, the method in this embodiment performs best in both attribute similarity and LPIPS, indicating that it can accurately complete the semantic transfer of facial age and facial fatness attributes, and the generated images are the most natural in human visual perception, effectively avoiding common artifacts and blurring phenomena. Second, although DeltaEdit has the highest SSIM score, its LPIPS is slightly lower than that of the method in this embodiment, indicating that its editing effect still has room for improvement in terms of naturalness. Third, LATS performs poorly in LPIPS and has the lowest attribute similarity, reflecting that it is prone to producing unnatural texture distortion when processing complex facial geometric deformations, making it difficult to balance the preservation of identity information with the accurate expression of target attributes. In short, compared with LATS, DiffusionCLIP, DeltaEdit, and Face2Diffusion, the method in this embodiment can better balance the contradiction between image reconstruction and semantic editing, achieving more accurate facial attribute editing while ensuring high perceptual quality.

[0025] Table 1. Evaluation results of face attribute editing quality using different methods

Claims

1. A method for editing multiple facial attributes based on low-rank adaptive and progressive scheduling, characterized in that, Includes the following steps: Step 1: Based on the FFHQ dataset, construct the FFHQ-Regulation-Dataset, a face regularization dataset with a 7×7 orthogonal attribute classification system, generate semantically aligned attribute text descriptions, and then preprocess the images. Step 2: Inject a low-rank adaptive matrix into the attention layer of the pre-trained diffusion model's U-Net network to obtain the LoRA-FA attribute editing model. The process is as follows: Step 2.1: Freeze the weights of the pre-trained diffusion model Stable Diffusion v1.

5. ; Step 2.2: Query matrix of the cross-attention layer and self-attention layer of the pre-trained diffusion model's U-Net network. Key matrix Value matrix and output layer linear mapping Inject a low-rank adaptive matrix into the pre-trained diffusion model and update the weights to... , is represented as: In the formula; This is the LoRA scaling factor, with a value of 32. Both B and are low-rank matrices, and , r is the rank parameter. For rank; and They represent and the size of B; Represents the set of real numbers; Step 3: Using the preprocessed regularized dataset FFHQ-Regulation-Dataset as the training set, the LoRA-FA attribute editing model is jointly trained using denoising loss, identity consistency loss, perceptual similarity loss and prior preservation loss based on minimum signal-to-noise ratio min-SNR weights. Step 4: In the inference process of multi-attribute face editing, the progressive scheduling execution framework PIP-Diff is used to perform anchor point feature extraction, geometric mask generation, and path planning operations on images in the Celeba-HQ dataset. The process is as follows: Step 4.1: Extract anchor features: Extract images from the Celeba-HQ dataset using the pre-trained InceptionResnetV1. A 512-dimensional deep identity feature vector is generated, which is used as the identity anchor point; simultaneously, images from the Celeba-HQ dataset are extracted using DINOv2. The intermediate layer feature map is used as the facial detail anchor point, and the identity anchor point and facial detail anchor point are used as global anchor points, with a set anchor return frequency. If the current iteration step number satisfy The system forcibly increases the weight of identity anchor points in the prompts and performs linear interpolation fusion on the latent variables of the starting anchor point of the current iteration step to obtain the updated latent variables. This will serve as the starting point for the next noise reduction step. Step 4.2: Generate a geometric mask: Use MediaPipe, a deep learning-based facial landmark detector, to extract a set of 468 3D facial key points from the current image I. Extract the corresponding index subset based on the facial contour boundaries. Calculate its center point and with the center point Based on the expansion factor Radializing the pixel coordinates in image I outwards yields the coordinates of the outwardly radiating pixels, thus constructing an expanded convex polygon mask matrix for the face. ; Simultaneously, key point sets of the left eye, right eye, eyebrows, and mouth were extracted to construct an eye protection mask. and mouth protective mask The primary mask matrix of the face image , is represented as: In the formula: and They are respectively and Mask weights; Will via the kernel size The final mask matrix is ​​obtained by applying a Gaussian smoothing filter. , is represented as: In the formula: This is a Gaussian blur operator used to feather binary masks, eliminating stitching marks caused by hard edges and enabling a smooth transition between the editing area and the background. Step 4.3, Path Planning: PIP-Diff predefines an ordered topological attribute chain of age and facial plumpness, and plans the face attribute editing task as a progressive path on the chain; Step 5: Introduce adaptive parameter control and multi-level regression correction mechanism to achieve stable attribute transformation in the feedback correction loop. The process is as follows: Step 5.1: Determine the current step size: Obtain the original image. The latent variable z is calculated based on the asymptotic path generated in step 4.3, according to the starting and ending values ​​of the target attribute. The editing task is decomposed into N steps. It is determined whether the current step has reached the ending value. If so, step 5.2 is executed; otherwise, step 5.7 is skipped. Step 5.2, Adaptive Parameter and Prompt Word Construction: Determine the denoising intensity based on the current step size progress to serve as the adaptive parameter, and generate the corresponding text prompt words; Step 5.3, Periodic Anchoring Mechanism: Introduce a periodic anchoring mechanism and set the anchoring frequency. Every experience If the step triggers the anchoring mechanism, the weight of the current latent space's identity anchor point will be forcibly increased; otherwise, the weight of the current latent space's identity anchor point will remain unchanged. Step 5.4, Latent Space Feature Fusion and LoRA-FA Attribute Editing: Inject attribute correction parameters into the attention layer of the U-Net in the LoRA-FA attribute editing model, and then... M Under the constraints, the diffusion model local repair algorithm is executed to generate candidate edited images for the current step. , is represented as: In the formula: The previous frame image, Edit the current denoised output of the model for LoRA-FA attributes; Step 5.5, Multi-level Regression Correction: Candidate images are corrected through multi-level regression. The process of performing multi-level verification is as follows: Step 5.5.1: Construct three independent perceptual metrics, including an identity extractor, a detail extractor, and an expression extractor, to calculate similarity including identity features. Global detail similarity Similarity to facial expressions Perception indicators, including those included. Step 5.5.2: Assess facial expression similarity sequentially. Global detail similarity Similarity to identity features Perform a judgment and verification, if , and If all three perception indicators are below the set threshold, it means that all three perception indicators have passed the verification, and the process proceeds directly to step 5.7; if , and If any item fails the validation, proceed to step 5.6; Step 5.6, Adaptive Three-Image Hybrid Correction Strategy: If any of the perception indicators in Step 5.5 fails the verification, the regression ratio coefficient is dynamically allocated according to the level of the verification failure, and adaptive three-image hybrid correction is performed to obtain the corrected image. , is represented as: In the formula: This is a smoothing coefficient used to prevent large single-step changes from causing visual abrupt changes or flickering. This is the anchorage coefficient; Step 5.7, Status Update and Indicator Recording: Record candidate images Or the image corrected in step 5.6 As a reference benchmark image for the next step, the perception index is updated in real time, and the LoRA scaling factor or denoising intensity of subsequent steps is dynamically adjusted according to the feedback signal until all evolutionary steps are completed, and the final multi-attribute edited face image is output.

2. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The process of step 1 is as follows: Step 1.1: Construct a 7×7 orthogonal attribute classification system: Based on the principle of independence of facial semantic features, construct a 7×7 orthogonal attribute matrix, which includes two independent dimensions: age and facial plumpness. The age dimension includes infant, child, adolescent, adult, middle-aged, elderly, and very old; the facial plumpness dimension includes extremely thin, thin, slightly thin, normal, slightly plump, plump, and obese. Through the cross-combination of these two dimensions, age and facial plumpness, define 49 core attribute nodes. Step 1.2, Original Image Filtering and Pre-labeling: Select stable face samples from the public dataset FFHQ, match 15 to 20 face samples for each attribute node, and construct a face image regularization dataset FFHQ-Regulation-Dataset; Step 1.3: Generate semantically aligned attribute text descriptions: For each image in the regularized dataset FFHQ-Regulation-Dataset, the image is renamed according to its position in the 7×7 orthogonal attribute matrix, following the rule of "age category_facial weight category_number". During the training preparation phase, the program directly reads the image file name and splits the file name with underscores, automatically extracting the two key labels of age and facial weight attributes from the file name. After obtaining the labels, the training script generates the corresponding text prompts in real time, which are the attribute text descriptions. The attribute text descriptions are converted into feature vectors through the CLIP text encoder, serving as conditional guidance signals during the training phase, so that the image content and semantic descriptions are accurately aligned in the latent space. Step 1.4, Image Preprocessing: Perform face alignment on all images in the regularization dataset FFHQ-Regulation-Dataset, uniformly crop and scale the images to a resolution of 512×512, then perform color space normalization on the images, and use the Laplacian operator to enhance the detailed features of the facial texture region.

3. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The process of step 3 is as follows: Step 3.1: Calculate the reciprocal of the signal-to-noise ratio (SNR). And set the truncation hyperparameter To limit the penalty weight , and They are represented as follows: In the formula: This is the cumulative multiplication factor of the noise scheduler; In the formula: Minimize the function; The truncation hyperparameter represents the minimum signal-to-noise ratio (min-SNR). The reciprocal of the signal-to-noise ratio (SNR) is expressed as: Then calculate the denoising loss based on the minimum signal-to-noise ratio (min-SNR) weight. , is represented as: In the formula: It is a moment Noisy latent variables; Represents real noise; It is a U-Net network with a pre-trained diffusion model loaded with LoRA weights; These are images from the regularization dataset FFHQ-Regulation-Dataset. This is the attribute text description corresponding to the image; Represents the CLIP model; Step 3.2: Decode the predicted latent variables of the LoRA-FA attribute editing model using VAE. Then, by utilizing the feature consistency constraint of the ArcFace face recognition network, identity feature vectors are extracted, and then cosine distance is used. Measured to calculate identity consistency loss , is represented as: In the formula: For the frozen facial recognition network; Step 3.3: Utilize LPIPS network constraints to evaluate the image. and Differences in the deep perceptual feature space are then used to calculate the perceptual similarity loss. , is represented as: In the formula: It is a frozen LPIPS network; Step 3.4: Divide each batch into instance samples and class samples. Specifically: use the samples in the regularization set FFHQ-Regulation-Dataset as instance samples, the FFHQ dataset as the prior set, and the generic face images within it as class samples. Calculate the denoising loss for each instance sample, then sum them using weighted averages to obtain the prior preservation loss. , is represented as: In the formula: It is the denoising loss corresponding to the instance samples from the regularization set FFHQ-Regulation-Dataset; These are prior preservation weights, used to constrain the model from deviating from the face category; It is the denoising loss corresponding to the general face images from the prior set. and All use weighted denoising loss based on the minimum signal-to-noise ratio (min-SNR) weight; Step 3.5: Combine the denoising loss, identity consistency loss, perceptual similarity loss, and prior preservation loss based on the minimum signal-to-noise ratio (min-SNR) weights to obtain the total loss. , is represented as: In the formula: and Loss of identity consistency and perceptual similarity loss The weight parameters.

4. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The updated latent variables in step 4.1 Represented as: In the formula: It is a latent variable representing the starting anchor point of the current iteration step. It is the fusion coefficient. It is the number of iterations. yes The gradient.

5. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The pixel coordinates of the external radiation in step 4.2 are represented as follows: In the formula: Representing an image The pixel coordinates in the image; This represents the coordinates of the pixels radiating outwards.

6. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The specific process of step 4.3 is as follows: Step 4.3.1: Define two one-dimensional ordered topological attribute chains. and , respectively represented as: Step 4.3.2, based on and The path planning module constructs the optimal editing path through index calculation, assuming the source attribute index is... The target attribute index is The algorithm logic of the path planning module is as follows: 1) If Then, the interval is intercepted along the positive direction. All state nodes are treated as an ordered topological attribute chain; 2) If Then, the interval is truncated in the reverse direction. The reversed nodes are treated as an ordered topological attribute chain; 3) If If the path is empty, no editing is required. PIP-Diff employs a phased, sequential execution strategy, prioritizing the gradual transformation of the age dimension. Once the age characteristics have stabilized, the gradual transformation of the facial fatness dimension is then performed based on the current results.

7. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The specific process of step 5.2 is as follows: Edit the span according to the attributes. And introduce the number of iterations It exhibits an exponential decay mechanism, and the attribute editing span... With the noise reduction intensity of step They are represented as follows: In the formula: As the reference strength; This represents the total number of steps. For normalized attribute editing span, indexed by source attribute index of target attribute The distance determines; It is the attenuation factor; Indicates the first noise reduction intensity of step That is, adaptive parameters.

8. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The specific process of step 5.3 is as follows: introduce a periodic anchoring mechanism and set the anchoring frequency. The default setting is 3, meaning that the anchoring mechanism is triggered every 3 steps, forcibly changing the weight of the current latent space's identity anchor point. An improvement of 0.15, with a maximum improvement of 0.5, ensures that the spreading manifold remains within the neighborhood of the original identity during the shift towards the target semantics, thus preventing identity loss caused by long-path editing.

9. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The specific process of step 5.5.1 is as follows: Step 5.5.1.1: Use a ResNet-50 network pre-trained on ImageNet with fully connected layers removed as the backbone network, and input the candidate edit image. and initial image After being cropped and normalized by facial bounding boxes, it is mapped to a high-dimensional feature vector. Define the currently generated candidate edit graph. With the initial image The cosine similarity is used as its identity feature similarity to quantify the closeness of two face images in deep identity semantics. Represented as: Step 5.5.1.2: Use Vision Transformer as a detail feature extractor to capture candidate editable images. and initial image High-frequency detail features, including skin texture and hair direction. Used to calculate candidate edit graphs With the initial image The cosine similarity is used as its global detail similarity to detect whether an image has undergone wax-like appearance or texture loss. Represented as: Step 5.5.1.3: Extract candidate editable images using ResNet-50. and initial image The intermediate layer features are then processed by adaptive average pooling to obtain expression features that are highly sensitive to local facial muscle tension. It is used to calculate facial expression similarity to capture subtle changes in the shape of the corners of the mouth and eyes. , is represented as: 。 10. The face multi-attribute editing method based on low-rank adaptive and progressive scheduling according to claim 1, characterized in that, The specific process of step 5.5.2 is as follows: First, determine... If the value is lower than the threshold of 0.7, then determine... If the validation fails, proceed to step 5.

6. If so, continue with the following checks. If the value is lower than the threshold of 0.65, then determine... If the validation fails, proceed to step 5.

6. If so, continue with the following checks. If the value is lower than the threshold of 0.55, then determine... If the verification fails, proceed to step 5.6; if all three perception indicators pass the verification, proceed directly to step 5.7.