Real-time face blind restoration method and device based on identity constraint and frequency domain enhancement

The method addresses identity drift and computational inefficiency in face restoration by using identity and frequency domain constraints, achieving real-time high-frequency detail recovery in low-quality images.

CN120318122AActive Publication Date: 2025-07-15湖南马栏山视频先进技术研究院有限公司

Patent Information

Application Number
CN202510387547.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The prior art has problems such as identity feature drift, texture and edge blur, and inefficient computing in facial blind repair, which is difficult to meet the real-time needs, especially in low-quality inputs, which are seriously lost.

Method used

Using an method based on identity constraints and frequency domain enhancement, the identity encoding vector is generated through fusion of CLIP and ArcFace, combined with ControlNet, spatial structural features are extracted, and high-frequency detail components are obtained using a learning wavelet transform, and mapped to the latent space through the VQGAN encoder, iterative denoising and wavelet residual correction are used to dynamically allocate high-frequency energy weights, and finally generate multi-scale fusion images.

Benefits of technology

It realizes efficient recovery of high-frequency details while maintaining identity consistency, improves the visual quality of the image, and meets the needs of real-time processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318122A_ABST
    Figure CN120318122A_ABST
Patent Text Reader

Abstract

The invention provides a real-time face blind restoration method and device based on identity constraint and frequency domain enhancement, and relates to the technical field of image processing, and the method comprises the steps: S1, generating an identity code through the fusion of CLIP and ArcFace, extracting spatial structure features in combination with ControlNet, and obtaining a high-frequency component through the decomposition of learnable wavelets; s2, an input image is mapped to a potential space, and identity, space and high-frequency features are fused to construct a condition vector; s3, performing iterative denoising based on the consistency model, and superposing lightweight wavelet residual correction to enhance high-frequency details; s4, according to the high-frequency energy dynamic distribution weight, reconstructing a multi-scale fusion image through inverse wavelet transform; according to the method, the problems of identity distortion, detail loss and low calculation efficiency of a low-quality face image in restoration are solved, and semantic consistency, detail fidelity and real-time performance of a restoration result are improved through identity-space-frequency domain ternary constraint, lightweight residual correction and degradation perception enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and in particular to a real-time face blind restoration method and device based on identity constraint and frequency domain enhancement. Background Art

[0002] Traditional face blind restoration methods often face three challenges: identity feature drift (the restoration result is inconsistent with the real identity), high-frequency detail blurring (insufficient restoration of texture and edges), and low computational efficiency (difficult to meet real-time requirements). Existing technologies such as GAN-based generation methods can improve visual quality, but are vulnerable to degradation types (such as easy to generate artifacts when blurred and noise are mixed), and rely on prior degradation information (non-blind restoration), resulting in poor generalization in actual scenarios; diffusion model-based methods can generate high-quality details, but have many iteration steps (usually 50+ steps) and high computational costs, making it difficult to process in real time. In addition, most methods lack explicit constraints on identity features, resulting in a deviation of the restoration result from the target identity, especially in the case of low-quality inputs (such as surveillance videos, compressed images), where identity information is severely lost.

[0003] In the field of security surveillance, blurred faces need to quickly restore recognizable identity features; in the digitalization of cultural heritage, old photos need to be restored while preserving the true appearance of the people; in mobile applications, lightweight models are required to enhance selfie images in real time. Therefore, there is an urgent need for a restoration solution that does not require prior degradation, has strong identity constraints, controllable detail enhancement, and high computational efficiency. Summary of the Invention

[0004] In view of the above technical problems in the related art, the present invention proposes a real-time face blind restoration method and device based on identity constraint and frequency domain enhancement.

[0005] In a first aspect, the present invention provides a real-time face blind restoration method based on identity constraint and frequency domain enhancement, including the following steps:

[0006] S1. Multi-modal feature extraction: The low-quality input image I LQ and the reference image set are fused through CLIP and ArcFace to generate an identity encoding vector c id , then the spatial structure feature f s is extracted by combining with ControlNet, and the diagonal high-frequency detail component F hh is obtained by using a learnable wavelet transform; N is the number of reference images;

[0007] S2. Hybrid conditional latent space initialization: The low-quality input image I LQ is mapped to the latent space through the VQGAN encoder E to obtain an initial latent variable z2, and the identity encoding vector c id, spatial structure feature f s and diagonal direction high-frequency detail component F hh to construct a hybrid conditional vector c;

[0008] S3. Multi-step generation with frequency domain enhancement: Through the consistency model, perform latent space iterative denoising on the initial latent variable z2 and the hybrid conditional vector c, and superimpose wavelet residuals to correct and enhance high-frequency details to generate a multi-scale restored image

[0009] S4. Adaptive fusion in the wavelet domain: According to the multi-scale restored image and the diagonal high-frequency detail component calculate the dynamic allocation weight of high-frequency energy, and perform inverse wavelet transform reconstruction to generate a multi-scale fused image I fuse .

[0010] Specifically, step S1 specifically includes the following steps:

[0011] S11. Obtain a low-quality input image and a reference image set where H is the height of the input image, W is the width of the input image, 3 represents the RGB three channels; N is an integer greater than or equal to 1; is the i-th reference image, H ref is the height of the reference image, W ref is the width of the reference image;

[0012] S12. Input the reference image into the pre-trained CLIP model to extract the global semantic feature f CLIP ; the reference image is the i-th image selected from the reference image set ;

[0013] S13. Input the reference image into the pre-trained ArcFace model to extract the identity feature vector f ArcFace ;

[0014] S14. Concatenate the global semantic feature f CLIP and the identity feature vector f ArcFace through channel concatenation operation to generate a joint feature f joint , and output the identity encoding vector c id after being processed by a multi-layer perceptron;

[0015] S15. Input the low-quality input image I LQ into the spatial encoder based on the ControlNet architecture, and combine the identity encoding vector c id to generate the spatial structure feature f s ;

[0016] S16. Decompose the low-quality input image I through the learnable wavelet transform module to obtain the diagonal high-frequency detail component F LQ hh .

[0017] Specifically, step S2 specifically includes the following steps:

[0018] S21. Map the low-quality input image I through the VQGAN encoder E LQ to the latent space to obtain the initial latent variable where E() is the mapping function of the encoder E; H is the image height, W is the image width, and 1024 is the number of latent variable channels;

[0019] S22. Concatenate the identity encoding vector c id , the flattened spatial structure feature Flatten(f s ) and the flattened diagonal high-frequency detail component Flatten(F hh ) in the channel dimension to generate the joint conditional vector c joint , and then compress it into the dimension of the joint conditional vector c joint through the fully connected layer to obtain the mixed conditional vector c = Linear(c joint ), and finally output the initial latent variable z2 and the mixed conditional vector c; perform concatenation in the channel dimension through the following formula:

[0020]

[0021] where Flatten() is the unfolding function; Linear() is the linear mapping function corresponding to the fully connected layer.

[0022] Specifically, step S3 specifically includes:

[0023] Use the initial latent variable and the mixed conditional vector as inputs, perform multiple steps of iteration through the consistency model, and combine the wavelet residual correction term for denoising during the iteration process to generate the preliminary latent variable z t-1 ; and generate the multi-scale restored image through the multi-scale decoder D t-1 from the latent variable z k during the process of generating the preliminary latent variable z T-k

[0024] The multi-scale decoder D k ​​Designed based on the SIMO architecture, it includes k-level upsampling blocks, each level containing a transposed convolution and a LeakyReLU activation function; where T is the total number of iteration steps, and T takes the value of 4; k = 1, 2, 3; They respectively correspond to the decoding results of the highest resolution z3, the intermediate resolution z2, and the lowest resolution z1.

[0025] Specifically, step S4 specifically includes:

[0026] Using the multi-scale restored image set and the high-frequency components of each scale extracted from the diagonal high-frequency detail component as inputs, first calculate the L1 norm of the high-frequency components of each scale as the high-frequency energy Then, based on the high-frequency energy dynamically allocate the fusion weight λ through the Softmax function k ; perform a learnable discrete wavelet transform on the restored image of each scale and decompose it into a low-frequency subband and a high-frequency subband Concatenate the horizontal and vertical components of the diagonal high-frequency detail component with the high-frequency subband to form a reconstructed high-frequency subband Then generate a reconstructed image through the inverse wavelet transform: Finally, weighted-fuse the reconstructed images according to the fusion weight λ k to generate a multi-scale fused image

[0027] Specifically, the method further includes:

[0028] S5, Dynamic post-processing and output: Adjust the sharpening intensity of the multi-scale fused image I fuse based on Gaussian blur to generate a low-frequency base image I blur .

[0029] Specifically, step S5 specifically includes:

[0030] When performing Gaussian blur processing on the multi-scale fused image , first generate a 5×5 Gaussian kernel whose weight values are calculated by a two-dimensional Gaussian function, and the two-dimensional Gaussian function is as follows:

[0031]

[0032] where G(x, y) represents the weight value of the relative coordinates x, y inside the Gaussian kernel; σ is the standard deviation, taking the value of 1.0, and the origin of coordinates is located at the center of the Gaussian kernel; the weight values of each point inside the Gaussian kernel are normalized, and the sum is 1, and the specific values are:

[0033]

[0034] For the multi-scale fusion image I fuse perform a two-dimensional convolution operation with zero-padding independently on each RGB channel, where the padding width of the two-dimensional convolution operation is 2, the stride is 1, and pixel values beyond the boundary are processed as 0. The formula is expressed as follows:

[0035]

[0036] where i and j respectively represent the rows and columns of the multi-scale fusion image I fuse to index the positions of pixels, i ∈ [1, H], j ∈ [1, W]; c is the channel index, and c = 1, 2, 3 correspond to the red, green, and blue color channels respectively; dx and dy are the offsets for the two-dimensional convolution operation; dx represents the horizontal offset, with a value range of -2 to 2; dy represents the vertical offset, with a value range of -2 to 2.

[0037] Specifically, step S12 specifically includes:

[0038] Input the reference image into the pre-trained CLIP model. First, perform normalization processing on the image to adjust the resolution to A × A pixels, and generate the first input tensor through mean-variance normalization Subsequently, process it through the C-layer Transformer blocks of the image encoder, divide the image into a sequence of B × B image patches, linearly project each patch into a C-dimensional vector, and aggregate global semantic information through the self-attention mechanism; finally, extract the feature vector corresponding to the CLS token of the last layer of the Transformer block, and map it to the E-dimensional space through the fully connected layer to generate the global semantic feature f CLIP .

[0039] Specifically, step S13 specifically includes:

[0040] Input the reference image into the pre-trained ArcFace model. First, perform normalization processing to adjust the image resolution to F × F pixels, and generate the second input tensor through mean-variance normalization Subsequently, the second input tensor passes through the forward propagation of the backbone network of the pre-trained ArcFace model, successively passing through the initial convolutional layer, batch normalization BN layer, and ReLU activation, and then passing through 4 residual blocks. Finally, extract the feature vector f after the last layer of global average pooling base ; for the feature vector Input the fully connected layer corresponding to the ArcFace loss function, and apply the additive angular margin to calculate the normalized identity feature vector of the output. The backbone network is ResNet-100.

[0041] In a second aspect, the present invention provides a real-time face blind restoration device based on identity constraint and frequency domain enhancement. Based on the real-time face blind restoration method based on identity constraint and frequency domain enhancement described in the above first aspect, it includes the following units:

[0042] The multi-modal feature extraction unit is used to generate the identity encoding vector c by fusing the low-quality input image I LQ and the reference image set through the fusion of CLIP and ArcFace, and then extract the spatial structure feature f by combining with ControlNet id , and obtain the high-frequency detail component F in the diagonal direction by using the learnable wavelet transform s ; N is the number of reference images; hh

[0043] The latent space initialization unit is used to map the low-quality input image I LQ to the latent space through the VQGAN encoder E to obtain the initial latent variable z2, and fuse the identity encoding vector c id , the spatial structure feature f s and the high-frequency detail component F in the diagonal direction hh to construct the mixed conditional vector c;

[0044] The frequency domain enhancement unit is used to perform latent space iterative denoising on the initial latent variable z2 and the mixed conditional vector c through the consistency model and superimpose the wavelet residual to correct and enhance the high-frequency details to generate the multi-scale restored image

[0045] The wavelet domain fusion unit is used to calculate the high-frequency energy dynamic allocation weight according to the multi-scale restored image and the diagonal high-frequency detail component and reconstruct through the inverse wavelet transform to generate the multi-scale fusion image I fuse ;

[0046] The sharpening processing unit is used to adjust the sharpening intensity of the multi-scale fusion image I fuse based on Gaussian blur to generate the low-frequency base image I blur .

[0047] The present invention provides a real-time face blind restoration method based on identity constraint and frequency domain enhancement, including: S1. Multi-modal feature extraction: The low-quality input image I LQ and the reference image set Generate the identity encoding vector c by fusing CLIP and ArcFace id , then extract the spatial structure feature f by combining with ControlNet s , and obtain the high-frequency detail component F in the diagonal direction by using the learnable wavelet transform hh ; S2. Hybrid conditional latent space initialization: Map the low-quality input image I LQ to the latent space through the VQGAN encoder E to obtain the initial latent variable z2, and fuse the identity encoding vector c id , the spatial structure feature f s and the high-frequency detail component F in the diagonal direction hh to construct the hybrid conditional vector c; S3. Multi-step generation with frequency domain enhancement: Perform latent space iterative denoising on the initial latent variable z2 and the hybrid conditional vector c through the consistency model and superimpose the wavelet residual correction to enhance the high-frequency details to generate the multi-scale restored image S4. Adaptive fusion in the wavelet domain: Calculate the high-frequency energy dynamic allocation weight according to the multi-scale restored image and the diagonal high-frequency detail component , and reconstruct through the inverse wavelet transform to generate the multi-scale fused image I fuse ; The present invention innovatively combines multi-modal identity features (CLIP semantics + ArcFace discriminative features) with the frequency domain enhancement mechanism, that is, through the identity-space-frequency triple constraint, and then accelerates the generation through the lightweight consistency model (LCM), and introduces dynamic residual correction to achieve precise control of high-frequency details. Finally, it breaks through the trade-off dilemma between identity consistency, detail fidelity and real-time performance of traditional methods, and provides a highly available solution for actual scenarios.

[0048] In addition, the present invention also adjusts the sharpening intensity of the multi-scale fused image through the Gaussian blur degree, so that the high-frequency components (such as noise, sharp edges) in the image are suppressed, and the generated low-frequency base image only retains the global illumination and smooth area information, making the image smoother, which helps to remove unnecessary noise while maintaining the image details, and improves the overall visual perception and processing quality of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 Schematic diagram of the real-time face blind restoration method based on identity constraint and frequency domain enhancement provided by the embodiment of the present invention;

[0051] Figure 2 Schematic diagram of a real-time face blind restoration device based on identity constraint and frequency domain enhancement provided by an embodiment of the present invention;

[0052] Figure 3 Schematic diagram of a real-time face blind restoration device based on identity constraint and frequency domain enhancement provided by an embodiment of the present invention. Detailed implementation manners

[0053] The present invention can be explained in detail through the following embodiments. The purpose of providing the present invention is to protect all technical improvements within the scope of the present invention. In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0054] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0055] Embodiment 1

[0056] Refer to Figure 1 , this embodiment provides a real-time face blind restoration method based on identity constraint and frequency domain enhancement, including the following steps:

[0057] S1. Multi-modal feature extraction: The low-quality input image I LQ and the reference image set are fused by CLIP and ArcFace to generate an identity encoding vector c id , then the spatial structure feature f is extracted by combining with ControlNet s , and the diagonal direction high-frequency detail component F is obtained by using a learnable wavelet transform hh ; N is the number of reference images;

[0058] S11. Obtain the low-quality input image and the reference image set where H is the height of the input image, W is the width of the input image, 3 represents the RGB three channels; N is the number of reference images; N is an integer greater than or equal to 1;

[0059] is the i-th reference image, H ref is the height of the reference image, W ref is the width of the reference image;

[0060] It is understandable that in this embodiment, it involves The format of indicates that the dimension of variable X is Y; for example Indicates I LQ The dimension of is H×W×3.

[0061] The reference image set is N images selected from different angles or lighting conditions of the same target person; from the reference image set Select the i-th image as the reference image The selection rule can be random selection or selection according to requirements;

[0062] S12. Input the reference image into the pre-trained CLIP model to extract the global semantic feature f CLIP ; The reference image is the i-th image selected from the reference image set ;

[0063] Step S12 specifically includes:

[0064] Input the reference image into the pre-trained CLIP model. First, normalize the image to adjust the resolution to A×A pixels, and generate the first input tensor through mean-variance normalization Subsequently, the first input tensor is processed by the C-layer Transformer blocks of the visual encoder of the pre-trained CLIP model. The image is segmented into a sequence of B×B image patches. Each patch is linearly projected into a C-dimensional vector, and the global semantic information is aggregated through the self-attention mechanism; finally, the feature vector corresponding to the CLS token of the last layer of the Transformer block is extracted and mapped to the E-dimensional space through the fully connected layer to generate the global semantic feature f CLIP ;

[0065] Furthermore, the feature distribution of the global semantic feature f CLIP is also constrained by L2 normalization;

[0066] The dimension of the feature vector is D; the L2 normalization formula is: The mean of the mean-variance normalization is [0.4815, 0.4578, 0.4082], and the variance is [0.2686, 0.2613, 0.2758]; the global semantic feature f CLIP has a dimension of 512; where A = 224, B = 16, C = 12, D = 768, E = 512.

[0067] In this embodiment, the pre-trained CLIP model consists of a dual-branch structure of a visual encoder and a text encoder, and the two achieve cross-modal semantic alignment through contrastive learning: The visual encoder adopts the Vision Transformer (ViT-B / 16) architecture. The input image is first divided into non-overlapping blocks of 16×16 pixels (a total of 196 blocks). Each block is linearly projected into a 768-dimensional vector and added with learnable position encoding, and then processed through 12 Transformer modules. Finally, a 768-dimensional global feature vector corresponding to the [CLS] token is extracted; The text encoder is based on a 12-layer standard Transformer model. The input text is tokenized into a sequence with a maximum length of 77 (including the start token [SOS] and the end token [EOS]) by BPE (Byte-Pair Encoding), passed through a 512-dimensional word embedding layer, and then encoded layer by layer through the self-attention mechanism. Finally, a 512-dimensional text feature at the position of the [EOS] token is extracted; The visual and text features are respectively mapped to a shared 512-dimensional embedding space through independent linear projection layers, so that the cosine similarity of cross-modal features can be directly used for contrastive learning; This structure captures local details and global semantics of the image through block encoding and hierarchical attention mechanism, and combines the context modeling ability of the text sequence to achieve efficient alignment of images and texts in a unified semantic space. In the training stage, the input images of the model are uniformly scaled to a resolution of 224×224, and the texts are input after being normalized. After the dual-branch independent forward calculation, the feature space distribution is optimized through the symmetric cross-entropy loss, and finally a cross-modal semantic representation ability with strong generalization is formed.

[0068] CLIP model: CLIP is a deep learning model proposed by OpenAI. It is trained by pairing images with relevant text descriptions. The core idea of the CLIP model is to learn a visual and language representation such that the representations of the same concept in the visual and text domains are close in the feature space. In this way, the CLIP model can be used for a variety of visual recognition tasks without fine-tuning for specific tasks.

[0069] Based on the ViT-B / 16 architecture: This means that the visual encoder part of the CLIP model adopts the Base version (ViT-B) of Vision Transformer, and the image is divided into blocks of 16x16 pixels ( / 16). ViT-B / 16 is a powerful visual feature extractor that can handle complex relationships in images.

[0070] Mean-Variance Normalization, also often referred to as Z-score Standardization, is a data preprocessing method used to adjust the numerical range of a dataset so that the mean of each feature is 0 and the standard deviation is 1. This method can make features of different magnitudes have the same scale, thereby eliminating the influence of dimensions and improving the performance of certain machine learning algorithms.

[0071] The self-attention mechanism, also known as the internal attention mechanism, is an attention mechanism that correlates different positions of a single sequence to calculate the representation of the same sequence. This mechanism allows the model to dynamically adjust the degree of attention to each element when processing sequence data, thereby capturing complex dependencies within the sequence.

[0072] The core of the self-attention mechanism is that it does not rely on external information but instead performs information interaction and integration among the internal elements of the sequence. This means that for each element in the sequence, the self-attention mechanism calculates the correlation between that element and all other elements in the sequence, generating a weighted representation where the weights reflect the relationships between the elements.

[0073] L2 normalization refers to scaling the feature vector to a unit L2 norm, that is, the L2 norm (Euclidean norm) of the feature vector is equal to 1.

[0074] S13. Input the reference image into the pre-trained ArcFace model to extract the identity feature vector f ArcFace ;

[0075] Step S13 specifically includes:

[0076] Input the reference image into the pre-trained ArcFace model. First, perform standardization processing to adjust the image resolution to F×F pixels and generate the second input tensor through mean-variance normalization Subsequently, the second input tensor passes through the forward propagation of the backbone network of the pre-trained ArcFace model, successively going through the initial convolutional layer, batch normalization BN layer, and ReLU activation, and then passing through 4 residual blocks, and finally extracting the feature vector f after the last global average pooling base ; Input the feature vector into the fully connected layer corresponding to the ArcFace loss function, and apply additive angular margin calculation to output the normalized identity feature vector The backbone network is ResNet-100;

[0077] The mean-variance normalization formulas described in steps S12 and S13 are as follows:

[0078]

[0079] where x and y are the image pixel coordinates, c is the color channel index, and μ c is the color channel mean, σ c is the color channel variance, and ò is the numerical stability constant;

[0080] Among them, the mean-variance normalization parameters are processed independently for each channel (Channel-wise), that is, the mean and variance are defined separately for the three color channels of the RGB image.

[0081] In this embodiment, the color channel mean of the mean-variance normalization is [0.5, 0.5, 0.5], and the color channel variance is [0.5, 0.5, 0.5]; the initial convolutional layer includes a 7×7 convolutional kernel, the stride is 2, and the output channels are 64; each of the 4 residual blocks contains 3 layers of Bottleneck structures, and the number of channels is gradually increased to 512; the weight matrix of the fully connected layer The L2 norm constraint of the feature vector is 1; the feature vector f base and the identity feature vector have a dimension of 512; F = 12.

[0082] The pre-trained ArcFace model is a deep learning model specifically for face recognition. It uses ResNet-100 as the feature extractor and is optimized by the ArcFace loss function to generate highly discriminative facial feature representations. Such a model can be directly used for face recognition tasks or as the basis for other models for transfer learning.

[0083] The pre-trained ArcFace model is optimized by maximizing the inter-class difference and minimizing the intra-class difference during the training phase, and can extract discriminative identity features without fine-tuning for the subsequent joint construction of identity encoding.

[0084] ArcFace is a loss function for face recognition tasks. It is based on deep learning feature embedding and classification techniques. ArcFace increases the inter-class difference by introducing an additive angular margin, making the distance between different classes in the feature space larger, thereby improving the recognition accuracy.

[0085] The ArcFace loss function is an improvement based on the traditional softmax loss function, aiming to make the features of the same category more compact and the features of different categories more dispersed during the training process.

[0086] ResNet-100 refers to a Residual Network structure with 100 convolutional layers. ResNet is a deep neural network architecture that addresses the vanishing gradient problem in the training of deep networks by introducing residual learning.

[0087] ResNet-100 is used as the backbone network to extract high-level features of images, which are then input into the ArcFace loss function for classification.

[0088] Global Average Pooling (GAP) is a pooling technique used in deep learning networks, especially in Convolutional Neural Networks (CNNs). Its main function is to average all pixel values on each feature map, resulting in a single value that can be regarded as the global representation of the feature map.

[0089] Additive Angular Margin (AAM) is a loss function strategy used in face recognition tasks in deep learning, especially when using deep feature embeddings and cosine similarity metrics. It aims to improve the discriminative ability of the model by adding angular margins in the feature space, making the embeddings of the same category closer and the embeddings of different categories farther apart.

[0090] The L2 norm constraint, also known as Weight Decay or Tikhonov regularization, is a commonly used regularization method in the training of machine learning and deep learning models. It penalizes large weight values by adding a term proportional to the L2 norm of the model parameters (i.e., the Euclidean norm of the weight vector) to the loss function, thereby reducing the overfitting phenomenon of the model.

[0091] S14. Generate the joint feature f by concatenating the global semantic feature f CLIP with the identity feature vector f ArcFace through a channel concatenation operation, and output the identity encoding vector c after being processed by a multi-layer perceptron joint ; id ;

[0092] The multi-layer perceptron includes two fully connected layers and a GELU activation function; the channel concatenation operation is operation; in this embodiment, the identity encoding vector c idhas a dimension of 1024;

[0093] In another possible implementation, the reference image set performs operations S12 and S13 on each reference image separately, and independently extracts global semantic features and identity feature vectors Subsequently, a joint feature f is generated by element-wise averaging of the features of all reference images joint , and then mapped to the final identity encoding vector c through a multi-layer perceptron id ; The formula for element-wise averaging of the feature fusion is as follows:

[0094]

[0095] where, is the global semantic feature of the i-th reference image; is the identity feature vector of the i-th reference image.

[0096] S12 and S13 process each reference image independently (i.e., for the i-th reference image, S12 extracts S13 extracts ), and the two strictly correspond to the same input image to avoid cross-image feature misalignment.

[0097] S15. Input the low-quality input image I LQ into the spatial encoder based on the ControlNet architecture, and combine the identity encoding vector c id to generate the spatial structure feature f s ;

[0098] The spatial encoder based on the ControlNet architecture includes an initial convolutional layer and 4 downsampling blocks; the initial convolutional layer includes a 3×3 convolutional kernel, the stride of the convolutional kernel is 1, the padding is 1, and the number of channels is 64; the downsampling block includes 3×3 convolution and ReLU activation function, the stride of the convolutional kernel is 2, and the padding is 1; the spatial structure feature

[0099] Step S15 specifically includes:

[0100] Input the low-quality input image into the spatial encoder based on the ControlNet architecture, and first convert the input into an initial feature map through the initial convolutional layer Subsequently, the initial feature map combines the identity encoding vector c id and is sequentially processed through 4 downsampling blocks to output the spatial structure feature

[0101] Specifically, the subsequent initial feature map is combined with the identity encoding vector c id and sequentially processed through 4 downsampling blocks to output the spatial structure feature f s Specifically,

[0102] Each downsampling block consists of a 3×3 convolutional kernel and a ReLU activation function. The stride of the convolutional kernel is 2 and the padding is 1, gradually reducing the spatial resolution and increasing the number of channels: The first downsampling block compresses the feature map size from H×W×64 to The second to fourth downsampling blocks sequentially output feature map sizes of And in each level of downsampling, the identity encoding vector c id is dynamically projected to the current feature channel dimension (e.g., 256 dimensions for the second block) through a fully connected layer, and then extended to the same size as the feature map through spatial replication. Finally, it is added to the convolutional output feature element by element in the form of a residual to obtain the feature map of each level. The formula is expressed as:

[0103] F l = ReLU(Conv(F l-1 )) + Reshape(Linear(c id ))

[0104] where F l represents the feature map output by the l-th downsampling block; where ReLU() is the ReLU activation function; Conv() is the convolution function; l = 1, 2, 3, 4 represents the block number, and Reshape() represents expanding the projected vector into a spatial feature map; in this embodiment, the feature channel is the number of channels of each downsampling block, 128 for the first downsampling block, and 256 for the second, third, and fourth downsampling blocks. The feature map output by the last downsampling block is used as the spatial structure feature f s , that is, in this embodiment

[0105] The fully connected layer, also known as the linear transformation layer, its core role is to perform a linear mapping on the input vector (or feature map) through the weight matrix and the bias term. Linear() is the linear mapping function corresponding to the fully connected layer.

[0106] Through the conditional injection mechanism, it ensures the deep fusion of identity semantics and spatial structure, enabling f s to simultaneously encode the low-level geometric information of the input image (such as the positions of facial features) and the identity constraints of the reference image, providing fine-grained guidance for the subsequent generation stage.

[0107] The conditional injection mechanism refers to a technical strategy that deeply integrates external conditional information (such as identity semantic encoding) with the spatial structure features of images by dynamically adjusting the intermediate feature representations of the model. Its core is to enable the global identity semantic constraints to gradually guide the generation of local details in the image reconstruction process through parametric mapping and feature operations, thereby enhancing identity consistency while maintaining spatial coherence.

[0108] Downsampling is a common operation in convolutional neural networks (CNNs), also known as subsampling or pooling.

[0109] A 3×3 convolution represents a convolution operation with a convolution kernel of size 3 rows by 3 columns. The convolution kernel slides over the input data to generate the output feature map.

[0110] A stride of 2 means that the convolution kernel moves 2 pixels at a time on the input feature map. This means that the convolution kernel skips 1 pixel each time it moves, resulting in an output feature map that is smaller in size than the input feature map.

[0111] Padding of 1: Padding is adding extra boundaries at the edges of the input feature map, usually padding with 0. Here, padding of 1 means adding a 1-pixel-wide boundary in each dimension of the input feature map. This is done to keep the size of the feature map from decreasing too much after convolution.

[0112] ReLU activation function: ReLU (Rectified Linear Unit) is an activation function that sets all negative-value pixels to 0 while leaving positive-value pixels unchanged. This function is commonly used to increase the non-linearity of the network, enabling the network to learn and simulate more complex functions.

[0113] S16. Decompose the low-quality input image I LQ through the learnable wavelet transform module to obtain the low-frequency approximation component F ll and the high-frequency detail components F lh in the horizontal direction, F hl in the vertical direction, and F hh in the diagonal direction.

[0114] Step S16 specifically includes:

[0115] Input the low-quality input image into the learnable wavelet transform module. First, expand the image size to an even resolution through zero-padding to generate a preprocessed image where,

[0116] Subsequently, four groups of trainable wavelet kernels are respectively used for depthwise separable convolution operations: Each group of wavelet kernels consists of two convolutional kernels with dimensions of 2×2×3×256, denoted as corresponding to the decomposition directions of low-frequency approximation, horizontal high-frequency, vertical high-frequency, and diagonal high-frequency respectively, and perform downsampling convolution with a stride of 2 on the input image:

[0117]

[0118] where DW-Conv represents the depthwise separable convolution function, and the output size of each group of convolutions is After being processed by the GELU activation function, the four groups of outputs are grouped according to the frequency band type: The low-frequency approximation component F ll is directly generated by The horizontal high-frequency component The vertical high-frequency component The diagonal high-frequency detail component ψ k is the wavelet kernel parameter; k = 1, 2, 3, 4;

[0119] Furthermore, during the training process of the learnable wavelet transform module, the wavelet kernel parameter ψ k is optimized through backpropagation, and a perfect reconstruction constraint loss function is introduced: L recon =‖ILWT(LWT(I LQ )) - I LQ ‖1;

[0120] where LWT() and ILWT() are the learnable wavelet transform and its inverse transform respectively, ensuring lossless information during the decomposition process. Finally, four groups of components (cropped to half the resolution of the original input) are output, and the high-frequency component F hh is used as the core input for subsequent frequency domain enhancement.

[0121] The Learnable Wavelet Network (LWN) adopts a hierarchical adaptive frequency-domain decomposition structure. It performs multi-scale frequency-band separation on the input features through three levels of learnable low-pass and high-pass wavelet kernel groups. Four sub-band components, namely low-frequency (LL), horizontal high-frequency (LH), vertical high-frequency (HL), and diagonal high-frequency (HH), are generated at each level. Its core uses a parameterized 5×5 convolutional kernel to implement the wavelet transform. The kernel weights are constrained by an orthogonal regularization loss to satisfy the tight support orthogonality, ensuring the reversibility of the decomposition. Across levels, skip connections are used to fuse the low-level high-frequency details (HH) with the high-level low-frequency basis (LL), and a Sigmoid gating mechanism is introduced to dynamically allocate channel weights, suppressing noise frequency bands and enhancing identity-related texture features. This module optimizes the frequency-band division boundary of the wavelet kernel through end-to-end training, realizes degradation-aware frequency-domain feature decoupling, and combines a lightweight convolutional structure to improve computational efficiency, effectively supporting the reconstruction of high-frequency details in image inpainting.

[0122] Depthwise Separable Convolution is an efficient convolutional operation that is very popular in deep learning models, especially in the design of lightweight neural networks for mobile and edge devices. It decomposes the traditional convolutional operation into two more efficient steps: Depthwise Convolution and Pointwise Convolution.

[0123] Depthwise Convolution: In depthwise convolution, each input channel is convolved separately using different filters. If the input has C channels, then there will be C independent 3x3 (or other size) filters. Each filter acts only on the corresponding input channel, rather than all channels. This means that if there are C input channels, C feature maps will be produced.

[0124] Pointwise Convolution: After depthwise convolution, pointwise convolution takes the output of depthwise convolution as input and uses 1x1 convolutional kernels to combine these feature maps. If F output channels need to be produced, pointwise convolution will use F 1x1xC filters, where C is the number of output channels of depthwise convolution.

[0125] Step S1 first obtains a low-quality input image (where H is the image height, W is the image width, and 3 represents the three RGB channels) and a reference image set (N≥1 is the number of reference images), and inputs at least one reference image Ref i into a pre-trained CLIP model to extract global semantic features And extract identity features through a pre-trained ArcFace model Concatenate the two through a channel concatenation operation (denoted as ) to generate joint features Then process it through a multi-layer perceptron (MLP) containing two fully connected layers and a GELU activation function to output an identity encoding vector Meanwhile, input the low-quality image I LQ into a spatial encoder based on the ControlNet architecture (including 4 downsampling blocks, each block containing a 3×3 convolution and a ReLU activation), and combine it with the identity encoding vector c id to generate spatial structure features In addition, decompose I LQ through a learnable wavelet transform module (LWN) to obtain a low-frequency approximation component a horizontal high-frequency detail component F lh a vertical high-frequency detail component F hl and a diagonal high-frequency detail component where the LWN module is implemented by a trainable wavelet kernel and satisfies the perfect reconstruction constraint; finally, output the identity encoding vector c id the spatial structure feature f s and the high-frequency detail component F hh as constraints for subsequent processing

[0126] S2. Hybrid conditional latent space initialization: Map the low-quality input image I LQ to the latent space through the VQGAN encoder E to obtain an initial latent variable z2, and fuse the identity encoding vector c id the spatial structure feature f s and the diagonal high-frequency detail component F hh to construct a hybrid conditional vector c

[0127] S21. Map the low-quality input image I LQ to the latent space through the VQGAN encoder E to obtain an initial latent variable where E() is the encoder E mapping function; H is the image height, W is the image width, and 1024 is the number of latent variable channels

[0128] The VQGAN encoder E is a downsampling encoder with 4 convolutional layers, and the number of output channels is 4C = 1024, where C is the feature channel base number and takes the value of 256; the first convolutional layer includes a 3×3 convolutional kernel, with a stride of 2, a padding of 1, and the number of output channels is 64; the second convolutional layer includes a 3×3 convolutional kernel, with a stride of 2, a padding of 1, and the number of output channels is 256; the third convolutional layer includes a 3×3 convolutional kernel, with a stride of 2, a padding of 1, and the number of output channels is 256; the fourth convolutional layer includes a 3×3 convolutional kernel, with a stride of 1, a padding of 1, and the number of output channels is 1024;

[0129] Step S21 is specifically as follows:

[0130] Input the low-quality input image into the VQGAN encoder E. First, pass the input through the first convolutional layer to downsample the output feature map from the original resolution H×W×3, and the size is Subsequently, apply batch normalization and the ReLU activation function; the second convolutional layer uses the same convolutional operation configuration as the first convolutional layer, and the number of channels is increased to 256, and the size of the output feature map is further compressed to and undergoes the same batch normalization and ReLU function activation processing; the third layer expands the number of channels to 512 through a 3×3 convolution with a stride of 2, and the spatial resolution is reduced to Finally, the fourth convolutional layer increases the channel dimension to 1024 without changing the resolution to generate the initial latent variable During the encoding process, zero padding is used after each convolutional layer to ensure size alignment, and shallow detail features are fused through residual connections. The formula is expressed as:

[0131] F l = ReLU(Conv(BN(F l-1 )))+Conv skip (F l-1 )

[0132] where l = 1, 2, 3, 4 represents the layer serial number, BN() is the batch normalization function; Conv skip is a 1×1 convolution used to adjust the number of channels in the residual branch; F l is the feature map obtained from the l-th convolutional layer.

[0133] During the training phase, the VQGAN encoder E is jointly optimized through adversarial loss and perceptual loss to ensure that the latent variable z2 not only retains the structural information of the input image (such as facial feature contours) but also has the semantic expression ability that can be decoded by the generative model, providing a complete generation starting point for the subsequent consistency model (LCM).

[0134] Batch Normalization (BatchNorm for short) is a commonly used technique in deep learning, which is used to accelerate the training process of neural networks and improve performance. It normalizes the data of each small batch, making the activation values in the network have a more stable distribution, which helps to alleviate the so-called "internal covariate shift" problem.

[0135] Zero Padding is a commonly used technique when dealing with image data in deep learning, especially in Convolutional Neural Networks (CNNs). Its purpose is to add additional boundaries around the edges of the input image, and the pixel values in these boundaries are all set to 0.

[0136] S22. Concatenate the identity encoding vector c id , the flattened spatial structure feature Flatten(f s ), and the flattened high-frequency detail components in the diagonal direction Flatten(F hh ) in the channel dimension to generate a joint conditional vector c joint , and then compress it into the dimension of the joint conditional vector c joint through a fully connected layer Linear(·) to obtain a mixed conditional vector Finally, output the initial latent variable z2 and the mixed conditional vector c; concatenate them in the channel dimension through the following formula:

[0137]

[0138] where Flatten() is the unfolding function; the flattened spatial structure feature is the flattened spatial dimension; the flattened high-frequency detail components in the diagonal direction is the flattened spatial dimension; 256 is the channel dimension; the joint conditional vector represents a dimension of

[0139] S3. Multi-step generation with frequency domain enhancement: Use a consistency model to perform latent space iterative denoising on the initial latent variable z2 and the mixed conditional vector c, and superimpose wavelet residuals to correct and enhance high-frequency details to generate a multi-scale restored image

[0140] Specifically, step S3 specifically includes:

[0141] The initial latent variable and the mixed conditional vector Take [the input], perform multiple steps of iteration through the consistency model, and generate the preliminary latent variable z during the iteration by combining the wavelet residual correction term for denoising t-1 ; and during the generation of the preliminary latent variable z t-1 generate the multi-scale restored image from the latent variable z through the multi-scale decoder D k T-k The latent variable z T-k is the preliminary latent variable z generated during the iteration process t-1 , and its index relationship is k = T - t + 1;

[0142] The multi-scale decoder D k is designed based on the SIMO architecture and contains k-level upsampling blocks, each level including a transposed convolution and a LeakyReLU activation function; where, the T is the total number of iteration steps, and T takes the value of 4; k = 1, 2, 3; respectively corresponding to the decoding results of the highest resolution z3, the intermediate resolution z2, and the lowest resolution z1;

[0143] Taking the initial latent variable and the mixed conditional vector as the input, perform multiple steps of iteration through the consistency model, and generate the preliminary latent variable z during the iteration by combining the wavelet residual correction term for denoising t-1 Specifically:

[0144] Taking the initial latent variable and the mixed conditional vector as the input, perform multiple steps of iteration through the consistency model, and generate the preliminary latent variable z during the iteration by combining the wavelet residual correction term for denoising t-1 Specifically: Each step of iteration calls the LCM single-step generation function LCM_step(z t , t, c) to generate the preliminary latent variable z t-1 , and add the wavelet residual correction term at the time step t ∈ {2, 3, 4} for latent space denoising operation to enhance high-frequency details: z t is the current latent variable; the current latent variable in the first step of iteration is the initial latent variable z2,

[0145] Calling the LCM single-step generation function LCM_step(z t , t, c) to generate the preliminary latent variable z t-1 Specifically includes:

[0146] First project the mixed conditional vector c to the spatial dimension through a fully connected layer and add it element-wise to the current latent variable z t to obtain the condition-enhanced latent variable​​ Subsequently, the time step t is converted into a sine position encoding and mapped to time weights through a multi-layer perceptron (MLP) And perform a Hadamard product in the channel dimension to generate a time-conditioned latent variable Input the time-conditioned latent variable into the improved latent consistency denoising network to calculate the denoising residual Δz denoise , and finally generate the preliminary latent variable z through the consistency condition formula t-1 ;

[0147] The improved latent consistency denoising network consists of 4 residual blocks, each block contains a 3×3 depth convolutional layer, 1 layer normalization layer and 1 multi-head attention mechanism, and the multi-head attention mechanism contains 4 heads, each head has a dimension of 256;

[0148] The formula for calculating the denoising residual is as follows:

[0149]

[0150] where ResBlock() represents the residual block (Residual Block), and the number in the lower right corner (such as ResBlock1, ResBlock2) represents the processing level number of the residual block, which is used to distinguish the feature extraction modules in different stages;

[0151] The consistency condition formula is as follows:

[0152]

[0153] where is the cumulative noise attenuation coefficient (β s ∈(0,1) is the preset noise scheduling parameter), and the denominator term is used to stabilize the gradient update;

[0154] This process realizes fast denoising with degradation perception by explicitly fusing the mixed conditions and the multi-order attention mechanism.

[0155] And in the process of generating the preliminary latent variable z t-1 generate a multi-scale restored image through the multi-scale decoder D k from the latent variable z T-k Specifically: Specifically:

[0156] Input the latent variables at different iteration steps (where T = 4, k = 1, 2, 3) into the multi-scale decoder D k , and gradually reconstruct the multi-scale restored image based on the SIMO architecture:

[0157] For the highest resolution decoding path (k = 1), using the latent variable as the input, first, the size of the feature map is increased to through the first-level upsampling block (transpose convolution kernel size 4×4, stride 2, padding 1, output channels 512). After being activated by LeakyReLU (negative slope 0.2), it is aligned with the spatial structure feature in resolution through bilinear interpolation and concatenated in the channel dimension to generate a fused feature The second-level upsampling block (transpose convolution kernel 4×4, stride 2, output channels 256) further expands the resolution to and adds it to the low-frequency component Finally, the third-level upsampling block (transpose convolution kernel 4×4, stride 2, output channels 3) outputs the RGB high-resolution restored image For the middle resolution (k = 2) and low resolution (k = 3) paths, using z1 and z2 as the inputs respectively, the middle-resolution restored image and the low-resolution restored image are generated through 2-level and 1-level upsampling blocks respectively. Among them, the resolution of is

[0158]

[0159] and it is upsampled to H×W×3 through bicubic interpolation. During the decoding process, skip connections are introduced after each level of transposed convolution. The formula expression of the skip connection is: where l = 1, 2, 3 represents the upsampling block level, Interp() is the function for aligning resolution through bilinear interpolation,

[0160] is the corresponding level feature extracted from the encoder;

[0161] ConvTranspose() is the transposed convolution (also known as deconvolution), which is used to gradually restore the spatial resolution of the feature map in the decoder;

[0162] The upsampling function, usually implemented by transposed convolution (ConvTranspose) or interpolation (such as bilinear upsampling), is used to gradually enlarge the size of the feature map in the decoding path and fuse it with the shallow features provided by the skip connection to reconstruct a high-resolution image;

[0163] The design achieves multi-scale supervision through a lightweight cascaded structure to ensure that the generated image is strictly aligned with the input degradation conditions in terms of identity features, spatial structure, and high-frequency details.

[0164] The Latent Consistency Model (LCM) is a model used in the field of machine learning, especially when dealing with unsupervised learning tasks. The core idea of LCM is to utilize the distribution consistency of data in the latent space to improve the learning efficiency and quality. The latent space usually refers to the low-dimensional space obtained by mapping the original data through an encoder, which contains a compressed representation of the original data.

[0165] The Hadamard product, also known as the element-wise product or Schur product, is a mathematical operation used between two matrices or vectors of the same shape. In the Hadamard product, the elements at the corresponding positions of the two matrices or vectors are multiplied, and the corresponding position of the resulting matrix or vector is the result of multiplying these two elements.

[0166] The SIMO (Single Input Multi Output) architecture is a system design method in which there is a single input source, and through system processing, multiple output results can be generated.

[0167] S4. Wavelet domain adaptive fusion: According to the multi-scale restored images and the diagonal high-frequency detail components calculate the dynamic allocation weights of the high-frequency energy, and reconstruct the multi-scale fused image I through inverse wavelet transform fuse ;

[0168] Step S4 specifically includes:

[0169] Taking the set of multi-scale restored images output by step S3 and the high-frequency components of each scale extracted from the diagonal high-frequency detail components as inputs, first calculate the L1 norm of the high-frequency components of each scale as the high-frequency energy Then, based on the high-frequency energy ​Dynamically allocate the fusion weight λ through the Softmax function k ; For the restored images at each scale Perform a learnable discrete wavelet transform and decompose it into a low-frequency subband and high-frequency subbands The diagonal high-frequency detail components and the horizontal and vertical components of the high-frequency subbands are spliced into a reconstructed high-frequency subband Then, generate a reconstructed image through the inverse wavelet transform: Finally, according to the fusion weight λ k Weightedly fuse the reconstructed images to generate a multi-scale fused image

[0170] It can be known from step S1 that the diagonal high-frequency detail component F hh is generated through multi-scale decomposition. The multi-scale characteristics are reflected in the gradual extraction of the image frequency domain information by wavelet transforms at different levels. The multi-scale includes three decomposition levels k = 1, 2, 3:

[0171] First-order scale (k = 1): After the original input image passes through the first-level learnable wavelet transform, a high-frequency component with the same resolution as the input is generated Its frequency band covers the widest range and mainly captures pixel-level high-frequency details (such as skin texture, hair edges, local noise);

[0172] Second-order scale (k = 2): The low-frequency component decomposed in the first level is input into the second-level wavelet transform to generate a high-frequency component with a resolution downsampled by half Its frequency band is narrower and focuses on medium-scale structured high-frequency features (such as facial contours, lighting boundaries);

[0173] Third-order scale (k = 3): Further decompose the low-frequency component of the second level to obtain a high-frequency component with a resolution downsampled by half again Its frequency band is the narrowest and is associated with sparse high-frequency signals in the global low-frequency base (such as facial symmetry, large-scale lighting transitions).

[0174] Therefore, F hh is a set of multi-scale high-frequency components, and it is easy to extract the diagonal high-frequency detail component to obtain the high-frequency components at each scale

[0175] The formula for calculating the L1 norm of the high-frequency components at each scale is:

[0176] Among them, It represents the eigenvalue at the spatial position (i, j) in the high-frequency sub-band in the diagonal direction and the c-th channel in the k-th decomposition layer (k = 1, 2, 3) after the input image is decomposed at multiple scales by the learnable wavelet transform module. It is the diagonal high-frequency response value at a specific position and channel in the k-th decomposition.

[0177] The Softmax function is as follows: (satisfying );

[0178] The low-frequency sub-band and the high-frequency sub-band both contain three directions: horizontal, vertical, and diagonal. denotes concatenation in the channel dimension; IDWT() is the inverse wavelet transform function; i and j respectively represent the row (height direction) and column (width direction) of the high-frequency component feature map, used to index the position of pixels, where i ∈ [1, H] and j ∈ [1, W]; c is the channel index, and c = 1, 2, 3 respectively correspond to the red, green, and blue color channels.

[0179] The L1 norm is usually mentioned in mathematics, signal processing, and optimization problems. Specifically, the L1 norm of a vector is defined as the sum of the absolute values of the components of the vector.

[0180] S5, Dynamic Post-Processing and Output: Based on Gaussian blur, the sharpening intensity of the multi-scale fused image I fuse is adjusted to generate the low-frequency base image I blur .

[0181] Specifically, it includes: when performing Gaussian blur processing on the multi-scale fused image output in step S4, first generate a 5×5 Gaussian kernel whose weight values are calculated by the two-dimensional Gaussian function:

[0182]

[0183] where G(x, y) represents the weight value of the relative coordinates x and y inside the Gaussian kernel; σ is the standard deviation, with a value of 1.0, and the coordinate origin is located at the center of the Gaussian kernel; the weights of each point inside the Gaussian kernel are normalized, and the sum is 1. The specific values are:

[0184]

[0185] Perform a two-dimensional convolution operation with zero padding independently on each RGB channel of the multi-scale fused image I fuse . The padding width of the two-dimensional convolution operation is 2, the stride is 1, and the pixel values exceeding the boundary are processed as 0. The formula is expressed as follows:

[0186]

[0187] where \(i\) and \(j\) respectively represent the rows (height direction) and columns (width direction) of the multi-scale fusion image \(I\) fuse for indexing the positions of pixels, \(i\in[1, H]\), \(j\in[1, W]\); \(c\) is the channel index, and \(c = 1, 2, 3\) correspond to the red, green, and blue color channels respectively; \(dx\) and \(dy\) are the offsets for two-dimensional convolution operations; \(dx\) represents the horizontal offset, and its value range is from -2 to 2; \(dy\) represents the vertical offset, and its value range is from -2 to 2.

[0188] Through this operation, the high-frequency components (such as noise and sharp edges) in the image are suppressed, and the generated low-frequency base image only retains the global illumination and smooth area information, providing a benchmark for subsequent high-frequency residual calculation.

[0189] The design concept of this embodiment stems from actual industrial pain points: in the field of security monitoring, it is necessary to quickly restore recognizable identity features for blurred faces; in the digitalization of cultural heritage, it is necessary to repair old photos and retain the true appearance of people; in mobile applications, it is necessary to have a lightweight model to enhance selfie images in real time. Therefore, there is an urgent need for a restoration solution that does not require a degradation prior, has strong identity constraints, controllable details for enhancement, and is computationally efficient.

[0190] This embodiment provides a real-time face blind restoration method based on identity constraint and frequency domain enhancement, including: S1. Multi-modal feature extraction: The low-quality input image \(I\) LQ and the reference image set are used to generate an identity encoding vector \(c\) through the fusion of CLIP and ArcFace, id and then the spatial structure feature \(f\) is extracted by combining with ControlNet, s and the diagonal direction high-frequency detail component \(F\) is obtained by using a learnable wavelet transform; hh S2. Hybrid conditional latent space initialization: The low-quality input image \(I\) is mapped to the latent space through the VQGAN encoder \(E\) to obtain the initial latent variable \(z_2\), and the identity encoding vector \(c\), LQ the spatial structure feature \(f\), id and the diagonal direction high-frequency detail component \(F\) are fused to construct a hybrid conditional vector \(c\); S3. Multi-step generation with frequency domain enhancement: The initial latent variable \(z_2\) and the hybrid conditional vector \(c\) are iteratively denoised in the latent space through a consistency model, and the wavelet residual is superimposed to correct and enhance the high-frequency details to generate a multi-scale restored image s ; S4. Adaptive fusion in the wavelet domain: According to the multi-scale restored image hh and the diagonal high-frequency detail component S4. Wavelet domain adaptive fusion: According to the multi-scale restored image and the diagonal high-frequency detail component Calculate the dynamic distribution weight of high-frequency energy, and reconstruct the multi-scale fusion image I by inverse wavelet transform. fuse ; This embodiment innovatively combines multimodal identity features (CLIP semantics + ArcFace discriminative features) with the frequency domain enhancement mechanism, that is, through the identity-space-frequency domain ternary constraints, and then accelerates the generation through the lightweight consistency model (LCM), and introduces dynamic residual correction to achieve precise control of high-frequency details, ultimately breaking through the traditional method's trade-off dilemma between identity consistency, detail fidelity and real-time performance, and providing a highly available solution for actual scenarios.

[0191] In addition, this embodiment also adjusts the sharpening intensity of the multi-scale fusion image through Gaussian blur, so that the high-frequency components (such as noise and sharp edges) in the image are suppressed, and the generated low-frequency base image Only global illumination and smooth area information are retained to make the image smoother, which helps to remove unnecessary noise while maintaining image details, improving the overall look and feel of the image and processing quality.

[0192] Embodiment 2

[0193] refer to Figure 2 This embodiment provides a real-time face blindness repair device based on identity constraint and frequency domain enhancement, comprising the following units:

[0194] Multimodal feature extraction unit for low-quality input image I LQ and reference image sets Generate identity encoding vector c by fusing CLIP with ArcFace id , and then combined with ControlNet to extract the spatial structure feature f s , and use the learnable wavelet transform to obtain the high-frequency detail component F in the diagonal direction hh ; N is the number of reference images;

[0195] Latent space initialization unit for passing low-quality input image I through VQGAN encoder E LQ Map to the latent space to get the initial latent variable z2, and fuse the identity encoding vector c id , spatial structure characteristics f s And the high-frequency detail component F in the diagonal direction hh Construct the mixed condition vector c;

[0196] The frequency domain enhancement unit is used to perform latent space iterative denoising on the initial latent variable z2 and the mixed condition vector c through the consistency model and superimpose the wavelet residual correction to enhance the high-frequency details to generate a multi-scale restored image

[0197] Wavelet domain fusion unit for restoring images based on multiple scales With the diagonal high-frequency detail component Calculate the dynamic allocation weight of high-frequency energy, and perform inverse wavelet transform to reconstruct and generate the multi-scale fusion image I fuse ;

[0198] Sharpening processing unit, for sharpening the multi-scale fusion image I based on Gaussian blur fuse Adjust the sharpening intensity to generate the low-frequency base image I blur .

[0199] Embodiment III

[0200] Reference Figure 3 , Figure 3 is a schematic structural diagram of the real-time face blind restoration device based on identity constraint and frequency domain enhancement in this embodiment. The real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement in this embodiment includes a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, the steps in the above method embodiments are implemented. Alternatively, when the processor 21 executes the computer program, the functions of each module / unit in the above device embodiments are implemented.

[0201] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement. For example, the computer program can be divided into the respective modules in Embodiment II. For the specific functions of each module, please refer to the working process of the device described in the above embodiments, and details will not be repeated here.

[0202] The real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that the schematic diagram is only an example of the real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement, and does not constitute a limitation on the real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement. It may include more or fewer components than shown, or combine certain components, or different components. For example, the real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement may further include input / output devices, network access devices, buses, etc.

[0203] The processor 21 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 21 is the control center of the real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement, and connects various parts of the entire real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement through various interfaces and lines.

[0204] The memory 22 can be used to store the computer programs and / or modules. The processor 21 realizes various functions of the real-time face blind restoration device 20 based on identity constraint and frequency domain enhancement by running or executing the computer programs and / or modules stored in the memory 22, and by calling the data stored in the memory 22. The memory 22 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0205] Among them, if the modules / units integrated in the real-time face blind repair device 20 based on identity constraint and frequency domain enhancement are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0206] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0207] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process multiple processes and / or blocksFigure 1 a device with the functions specified in one or more boxes

[0208] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in the process Figure 1 one process or multiple processes and / or boxes Figure 1 a box or multiple boxes

[0209] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or boxes Figure 1 a box or multiple boxes

[0210] The parts not detailed in the present invention are prior art. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and all changes falling within the meaning and scope of the equivalent elements are intended to be included in the present invention.

Claims

1. A real-time face blind restoration method based on identity constraint and frequency domain enhancement, characterized in that It includes the following steps: S1. Multi-modal feature extraction: The low-quality input image I LQ and the reference image set are fused by CLIP and ArcFace to generate the identity encoding vector c id , then the spatial structure feature f is extracted by combining with ControlNet s , and the diagonal direction high-frequency detail component F is obtained by using the learnable wavelet transform hh ; N is the number of reference images; S2. Initialization of the hybrid conditional latent space: The low-quality input image I is mapped to the latent space through the VQGAN encoder E LQ to obtain the initial latent variable z2, and the identity encoding vector c id , the spatial structure feature f s and the high-frequency detail component F in the diagonal direction hh are used to construct the hybrid conditional vector c; S3. Multi-step Generation with Frequency Domain Enhancement: The initial latent variable z2 and the mixed conditional vector c are iteratively denoised in the latent space by the consistency model, and the wavelet residual correction is superimposed to enhance the high-frequency details to generate a multi-scale restored image. S4, Wavelet-domain adaptive fusion: Based on the multi-scale restored image and the diagonal high-frequency detail components Calculate the dynamic allocation weight of high-frequency energy, and reconstruct the multi-scale fused image I by inverse wavelet transform fuse .

2. The real-time face blind restoration method based on identity constraint and frequency domain enhancement according to claim 1, wherein Step S1 specifically includes the following steps: S11. Obtain a low-quality input image and a reference image set where H is the height of the input image, W is the width of the input image, 3 represents the three RGB channels; N is an integer greater than or equal to 1; is the i-th reference image, H ref is the height of the reference image, W ref is the width of the reference image; S12. Input the reference image into the pre-trained CLIP model to extract the global semantic feature f CLIP ; The reference image is the i-th image selected from the reference image set ; S13. Input the reference image into the pre-trained ArcFace model to extract the identity feature vector f ArcFace ; S14. Combine the global semantic feature f CLIP with the identity feature vector f ArcFace to generate a combined feature f joint through a channel concatenation operation, and output an identity encoding vector c id after processing by a multi-layer perceptron; S15. Input the low-quality input image I LQ into the spatial encoder based on the ControlNet architecture, and combine it with the identity encoding vector c id to generate the spatial structure feature f s ; S16. Decompose the low-quality input image I through the learnable wavelet transform module LQ to obtain the high-frequency detail component F in the diagonal direction hh .

3. The real-time face blind restoration method based on identity constraint and frequency domain enhancement according to claim 1, characterized in that Step S2 specifically includes the following steps: S21. Map the low-quality input image I to the latent space through the VQGAN encoder E to obtain an initial latent variable LQ where E() is the mapping function of the encoder E; H is the image height, W is the image width, and 1024 is the number of latent variable channels; ​ S22. Concatenate the identity encoding vector c id , the flattened spatial structure feature Flatten(f s ), and the flattened high-frequency detail component in the diagonal direction Flatten(F hh ) in the channel dimension to generate the joint conditional vector c joint , and then compress it into the joint conditional vector c joint through a fully connected layer to obtain the mixed conditional vector c = Linear(c joint ), and finally output the initial latent variable z2 and the mixed conditional vector c; concatenate in the channel dimension through the following formula: Among them, Flatten() is the unfolding function; Linear() is the linear mapping function corresponding to the fully connected layer.

4. The real-time face blind restoration method based on identity constraint and frequency domain enhancement according to claim 2, wherein Step S3 specifically includes: Take the initial latent variable and the mixed conditional vector as inputs, perform multiple steps of iteration through the consistency model, and combine the wavelet residual correction term during the iteration to denoise and generate the preliminary latent variable z t-1 ; and during the process of generating the preliminary latent variable z t-1 generate the multi-scale restored image k from the latent variable z T-k through the multi-scale decoder D The multi-scale decoder D k is designed based on the SIMO architecture and includes k-level upsampling blocks, each level containing a transposed convolution and a LeakyReLU activation function; where T is the total number of iterative steps, and the value of T is 4; k = 1, 2, 3; respectively corresponding to the decoding results of the highest resolution z3, the intermediate resolution z2, and the lowest resolution z1.

5. The real-time face blind restoration method based on identity constraint and frequency domain enhancement according to claim 1, wherein Step S4 specifically includes: Image set restored at multiple scales and high-frequency detail components from the diagonal High-frequency components at each scale extracted As input, first calculate the L1 norm of the high-frequency components at each scale as the high-frequency energy Then, based on the high-frequency energy Dynamically allocate the fusion weight λ through the Softmax function k ; For the restored images at each scale Perform a learnable discrete wavelet transform and decompose them into low-frequency subbands and high-frequency subbands The diagonal high-frequency detail components are concatenated with the horizontal and vertical components of the high-frequency subbands to form a reconstructed high-frequency subband Then generate a reconstructed image through the inverse wavelet transform: Finally, weight and fuse the reconstructed images according to the fusion weight λ k to generate a multi-scale fused image 6. The real-time face blind restoration method based on identity constraint and frequency domain enhancement according to claim 1, wherein The method further includes: S5. Dynamic post - processing and output: Based on Gaussian blur, adjust the sharpening intensity of the multi - scale fusion image I fuse to generate a low - frequency base image I blur .

7. The real-time face blind restoration method based on identity constraint and frequency domain enhancement according to claim 6, wherein Step S5 specifically includes: For the multi-scale fusion image When performing Gaussian blur processing, first generate a 5×5 Gaussian kernel The weight values are calculated by a two-dimensional Gaussian function, and the two-dimensional Gaussian function is as follows: Among them, G(x, y) represents the weight value of the relative coordinates x and y inside the Gaussian kernel; σ is the standard deviation, with a value of 1.0, and the coordinate origin is located at the center of the Gaussian kernel; the weights of each point inside the Gaussian kernel are normalized, and the sum is 1. The specific values are: Perform a two-dimensional convolution operation with zero padding independently on each RGB channel of the multi-scale fusion image I fuse The padding width of the two-dimensional convolution operation is 2, the stride is 1, and pixel values outside the boundary are processed as 0. The formula is expressed as follows: where i and j respectively represent the rows and columns of the multi-scale fusion image I fuse for indexing the positions of pixels, i ∈ [1, H], j ∈ [1, W]; c is the channel index, and c = 1, 2, 3 respectively correspond to the red, green, and blue color channels; dx and dy are the offsets for two-dimensional convolution operations; dx represents the horizontal offset with a value range from -2 to 2; dy represents the vertical offset with a value range from -2 to 2.

8. The real-time face blind restoration method based on identity constraint and frequency domain enhancement according to claim 2, wherein Step S12 specifically includes: The reference image is input into the pre-trained CLIP model. First, the image is normalized to adjust the resolution to A×A pixels, and the first input tensor is generated through mean-variance normalization Subsequently, it is processed through the C-layer Transformer blocks of the image encoder. The image is segmented into a sequence of B×B image patches. Each patch is linearly projected into a C-dimensional vector, and global semantic information is aggregated through the self-attention mechanism. Finally, the feature vector corresponding to the CLS token of the last layer of the Transformer block is extracted and mapped to the E-dimensional space through a fully connected layer to generate the global semantic feature f CLIP .

9. The real-time face blind restoration method based on identity constraint and frequency domain enhancement according to claim 1, wherein Step S13 specifically includes: The reference image is input into the pre-trained ArcFace model. First, normalization processing is performed to adjust the image resolution to F×F pixels, and a second input tensor is generated through mean-variance normalization Subsequently, the second input tensor passes through the forward propagation of the backbone network of the pre-trained ArcFace model, successively passing through the initial convolutional layer, batch normalization BN layer, and ReLU activation, and then passing through 4 residual blocks, and finally extracting the feature vector f after the last global average pooling base ; The feature vector is input into the fully connected layer corresponding to the ArcFace loss function, and the normalized identity feature vector is calculated and output by applying additive angular margin The backbone network is ResNet-100.

10. A real-time face blind restoration device based on identity constraint and frequency domain enhancement, characterized in that, It includes the following units: Multimodal feature extraction unit, used to process the low-quality input image I LQ and the reference image set to generate an identity encoding vector c through the fusion of CLIP and ArcFace id , then combine with ControlNet to extract spatial structure features f s , and use learnable wavelet transform to obtain the high-frequency detail component F in the diagonal direction hh ; N is the number of reference images; A latent space initialization unit for mapping a low-quality input image I to the latent space through a VQGAN encoder E to obtain an initial latent variable z2, and fusing an identity encoding vector c LQ , a spatial structure feature f id , and a diagonal direction high-frequency detail component F s to construct a mixed conditional vector c; hh ​ Frequency domain enhancement unit, which is used to perform latent space iterative denoising on the initial latent variable z2 and the mixed conditional vector c through a consistency model, and superimpose wavelet residuals to correct and enhance high-frequency details to generate a multi-scale restored image Wavelet domain fusion unit, for restoring images according to multiple scales and diagonal high-frequency detail components Calculate the dynamic allocation weight of high-frequency energy, and reconstruct the multi-scale fused image I through inverse wavelet transform fuse .

Citation Information

Patent Citations

  • An image recognition method based on improved ArcFace loss function

    CN109241995A

  • Face image quality evaluation method and device based on multi-attribute face comparison

    CN109544523A

  • Cross-age face recognition method based on feature fusion

    CN113221660A

  • Face image restoration method based on potential feature reconstruction and mask perception

    CN114331894A

  • Face recognition method and system under natural monitoring based on deep learning

    CN115512408A

Cited By

  • Multi-modal fusion feature generation method and device and storage medium

    CN120495830A

  • A multimodal fusion feature generation method, device and storage medium

    CN120495830B

  • Portrait facial feature intelligent conversion method and system based on adversarial generation

    CN120894467A

  • Image restoration method

    CN121120452A

  • Face identity exchange method, system and equipment

    CN121190619A