Text and face collaborative restoration method based on cross-modal alignment

By constructing a text-face multimodal dataset and a cross-modal network architecture, combined with optimizing the loss function, we achieved facial image restoration under complex lighting conditions, generating high-quality, identity-consistent frontal facial images, solving the problem of poor restoration effects in existing technologies and improving the face recognition rate.

CN120655787AActive Publication Date: 2025-09-16JILIN UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511149392.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-09-16
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively restoring blurred details and weakened identity features in facial image restoration under complex lighting conditions. Especially in severely degraded scenarios such as occlusion and side faces, the restoration results often suffer from structural distortion and semantic inconsistency.

Method used

A text-face multimodal dataset was constructed, and image super-resolution and text-image alignment were achieved through a cross-modal network architecture. A hybrid loss function was designed and optimized. The CLIP model was used to extract text semantics and align them with the latent identity features of low-quality images. The mask diffusion loss was combined to optimize the structural consistency of key facial areas to generate high-quality images.

Benefits of technology

It significantly improves the ability to retain the identity features of the restored facial image, generates frontal face images with consistent identity features, enhances the face recognition rate, and outperforms existing methods in visual effects and quantitative indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655787A_ABST
    Figure CN120655787A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of image processing, and provides a text and face collaborative restoration method based on cross-modal alignment, comprising the following steps: constructing a text-face multi-modal data set; based on a cross-modal network architecture, realizing image super-resolution and text and image alignment in the text-face multi-modal data set; training a cross-modal network, and designing and optimizing a mixed loss function; and reasoning to generate a high-quality image. According to the method, a deep fusion framework of text semantics and image features is constructed, so that the identity feature retention capability of the repaired face image is remarkably improved, and a key foundation is laid for improvement of the face recognition rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and in particular relates to a text and face collaborative restoration method based on cross-modal alignment. Background Art

[0002] In recent years, the application value of facial image restoration technology in public security, criminal investigation and other fields has become increasingly prominent. However, factors such as complex lighting conditions and performance limitations of imaging equipment often lead to problems such as blurred details and weakened identity features in low-quality facial images. Traditional restoration methods are difficult to effectively restore the usable information of the image. For example, restoration schemes based on optimization strategies rely on manually designed facial prior knowledge. Such methods are limited by the coverage of the prior model and the high complexity of iterative calculations. When dealing with severely degraded scenes such as occlusion and side faces, the restoration effect is difficult to meet actual needs.

[0003] With the development of deep learning technology, neural network-based restoration models have gradually become mainstream. From early convolutional neural networks (CNNs) and generative adversarial networks (GANs) to more recent Transformer and diffusion models, these methods have improved restoration capabilities by learning from large amounts of image data. However, most solutions rely solely on unimodal visual information and lack the guidance of external semantic cues. When dealing with complex degradations (such as shadows, occlusions, and pose changes), they struggle to accurately establish semantic associations between facial details. This can lead to structural distortion, loss of identity features, and semantic inconsistency in the restoration results. For example, when restoring faces from the side or those obscured by ornaments, traditional methods often generate structures that deviate from true features due to a lack of semantic constraints. Diffusion models, with their progressive generation capabilities, demonstrate robustness to complex degradation in face restoration. However, they are still limited to single-modal image input and struggle to leverage external semantic information, such as text, to optimize the restoration direction. Meanwhile, while cross-modal models such as CLIP have achieved semantically guided high-quality image synthesis in the field of text-to-image generation, these approaches are primarily targeted at generation tasks and face two major bottlenecks when directly applied to face restoration: first, it is difficult to preserve the identity features of low-quality input images, resulting in semantically correct restoration results but identity bias; second, the lack of datasets for text-low-quality-high-quality image triples poses data sparsity challenges for cross-modal model training. Therefore, integrating text semantics with image features to construct a cross-modal alignment mechanism that simultaneously achieves identity preservation and semantically consistent detail reconstruction during the restoration process has become a key breakthrough in the evolution of face restoration technology from single-modality to multi-modality. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a text and face collaborative restoration method based on cross-modal alignment, aiming to solve the problems raised in the above background technology.

[0005] The embodiment of the present invention is implemented as follows: a method for collaborative restoration of text and face based on cross-modal alignment, comprising the following steps:

[0006] Construct a text-face multimodal dataset;

[0007] Based on a cross-modal network architecture, we achieve image super-resolution and text-image alignment in a text-face multimodal dataset.

[0008] Train cross-modal networks and design and optimize hybrid loss functions;

[0009] Inference generates high-quality images.

[0010] An embodiment of the present invention provides a text and face collaborative restoration method based on cross-modal alignment. By constructing a deep fusion framework of text semantics and image features, it significantly improves the ability to retain identity features of restored facial images, thereby laying a key foundation for improving face recognition rates. Specifically, the CLIP model extracts identity-related semantics (such as age, gender, facial features, etc.) from the text and performs cross-modal alignment with the potential identity features in low-quality images to generate a joint embedding containing rich identity clues. At the same time, the stacked identity embedding mechanism integrates multi-source identity information and combines the mask diffusion loss to force optimization of the structural consistency of key facial areas (such as eyes, nose, and jawline), effectively avoiding the distortion or loss of identity features caused by insufficient single-modal information in traditional methods. At the same time, this cross-modal collaborative restoration mechanism can generate frontal facial images with consistent identity features even when the facial features are incomplete (such as those blocked by eyes, hats, etc., profile, heavy makeup, etc.), significantly enhancing the identity recognition rate of the restored image, and thus providing more reliable images for subsequent face recognition and other needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 A flowchart of a method for collaborative text and face restoration based on cross-modal alignment is provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0012] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0013] The specific implementation of the present invention is described in detail below with reference to specific embodiments.

[0014] like Figure 1 FIG. 1 is a flowchart of a method for collaborative text and face restoration based on cross-modal alignment provided by an embodiment of the present invention, comprising the following steps:

[0015] S1. Construct a text-face multimodal dataset, including:

[0016] S1.1. Obtain high-quality (HQ) face images from the public face datasets FFHQ and CelebA. For each person identity (ID), randomly select two images with different poses, clothing, or expressions to form a large number of image pairs.

[0017] S1.2. For each image pair, a random HQ image is degraded to generate a corresponding low-quality (LQ) image, forming an LQ-HQ image pair.

[0018] S1.3. Manually annotate each LQ-HQ image pair with a structured text description (Text), including character features such as gender, age, nationality, clothing, expression, and action. This forms "LQ-HQ-Text" triplet data to address the problem of missing multimodal datasets.

[0019] S1.4. Due to the high requirements for model generalization in the field of face restoration, the dataset constructed from FFHQ is used as the training set, and the dataset constructed from CelebA is used as the test set.

[0020] S2. Based on a cross-modal network architecture, we achieve image super-resolution and text-image alignment in a text-face multimodal dataset. Specifically, we:

[0021] S2.1. Using the image super-resolution module to achieve image super-resolution:

[0022] S2.1.1. Using the HQ encoder and LQ encoder, a U-Net-style downsampling architecture is used. A 512×512×3 RGB image is input and compressed into a 16×16×256 feature map through five stages of downsampling and residual block processing. Each stage consists of a convolutional layer with a stride of 2 and four residual blocks. The initial convolutional layer converts RGB to 64 channels, and subsequent downsampling layers gradually adjust the number of channels to 256. The residual block consists of two 3×3 convolutional layers with LeakyReLU activations, and skip connections are used to mitigate the vanishing gradient problem.

[0023] S2.1.2. Use the Transformer module to predict the codebook index sequence corresponding to the low-quality image: First, the 16×16×256 feature map output by the encoder is flattened into a 256×256 sequence, and a sinusoidal position encoding is added to preserve spatial information. The sequence is processed by 9 Transformer blocks, each of which contains a multi-head self-attention layer and a feedforward network. The self-attention layer calculates the correlation between elements in the sequence to capture global dependencies, and the feedforward network enhances the feature expression capability. Each block also contains residual connections and layer normalization to stabilize training. Finally, the features are mapped to 1024 dimensions (codebook size) through a linear layer, and a softmax classifier is used to predict the code index at each position;

[0024] S2.1.3. Use the HQ decoder (symmetric with the encoder) and an upsampling architecture to recover image details. Input 16×16×256 quantized features are processed through five stages of upsampling and residual blocks to gradually restore the image to a 512×512×3 RGB image. Each stage consists of a nearest neighbor interpolation layer (doubling the resolution), a 3×3 convolutional layer (adjusting the number of channels), and four residual blocks. The final convolutional layer maps the 64-channel features to 3-channel RGB, using the Tanh activation function to constrain the range to [-1, 1].

[0025] S2.2. Use the text-image alignment module to achieve text and image alignment:

[0026] S2.2.1. Use an image encoder (the ViT-L / 14 pre-trained model from CLIP) to extract the semantic embedding of the input ID image.

[0027] S2.2.2. Use a text encoder (using the dual text encoder architecture of SDXL (stable-diffusion-xl-base-1.0)) to extract multi-granularity semantic embeddings of textual cues, where the text encoder includes:

[0028] The main encoder (CLIP ViT-L / 14 text encoder) is used to process natural language descriptions;

[0029] Auxiliary encoder (OpenCLIP ViT-bigG / 14) for enhancing long text comprehension capabilities;

[0030] S2.2.3. Construct stacked ID embeddings: Fuse the embeddings of multiple input ID images with the classifier features in the text to generate fused embedding vectors. Concatenate these vectors along the length dimension to form stacked ID embeddings.

[0031] S2.3.4. Diffusion model adaptation and cross-attention mechanism: Based on the SDXL diffusion model (stable-diffusion-xl-base-1.0), the stacked ID embeddings replace the category word positions in the text embeddings to generate updated text embeddings. Through the cross-attention mechanism, the query vector interacts with the updated text embeddings, allowing the model to adaptively integrate ID information, ensuring that the generated image retains the ID information and conforms to the text description.

[0032] S3. Train the cross-modal network and design an optimized hybrid loss function:

[0033] The training strategy includes the following process:

[0034] S3.1. Learning discrete codebooks through HQ image self-reconstruction and training HQ encoder and HQ decoder simultaneously.

[0035] S3.2, freeze the discrete codebook and HQ decoder, train the LQ encoder and Transformer modules, predict the codebook index sequence, and model the global facial structure;

[0036] S3.3, fixed codebook and Transformer module, jointly trained with the image encoder, text encoder and diffusion model, and fine-tuned the parameters of the image encoder, text encoder and diffusion model using the total loss of joint training;

[0037] The hybrid loss function includes:

[0038] Codebook learning loss: L1 loss is used in the process of learning discrete codebook in step S3.1 , Perceptual Loss , fight against losses , code-level loss Joint optimization is performed to align the encoder output with the quantized representation. The expressions of several losses are as follows:

[0039] ;

[0040] ;

[0041] ;

[0042] ;

[0043] in, For original high-definition images; is the image obtained after self-reconstruction learning of the HQ image in step S3.1; represents the VGG19 feature extractor; is the discriminator; The compressed features obtained after the original image is input into the encoder, Each feature vector in the Transformer module will be replaced by the closest entry in the discrete codebook in step S3.1, thereby obtaining the quantized feature ; Indicates stopping the gradient operator; As a balance coefficient, in the embodiment of the present invention, it can be taken as 0.5;

[0044] Total loss of codebook learning :

[0045] ;

[0046] Total loss of the Transformer module: In the process of predicting the codebook index sequence through the Transformer module in step S3.2, the cross entropy loss is used and L2 loss Supervised training improves detail reconstruction capabilities. The expressions of the two losses are as follows:

[0047] ;

[0048] ;

[0049] in, is the value of the ith position in the probability distribution of the true codebook index; is the probability of the i-th codebook index predicted by the model; is the compressed feature after encoding the LQ image in step S2.1; 、 are the spatial dimensions of the feature map obtained after encoding the LQ image in step S2.1, which are height and width respectively;

[0050] Total loss of the Transformer module :

[0051] ;

[0052] in, As a balance coefficient, in the embodiment of the present invention, it can be taken as 0.5;

[0053] Mask diffusion loss : In step S3.3, the mask diffusion loss is introduced during the joint training with the diffusion model, and the binary mask of the identity-related region is generated by Mask2Former , calculate the weighted noise prediction error, the expression is as follows:

[0054] ;

[0055] in, is the real noise; Noise predicted for the diffusion model; is the time step Noisy image; Embed the stack ID in step S2.2.3; Represents the time step Make expectations, Uniform distribution , The total number of steps in the diffusion process. This layer expects the model to participate in training at different stages of the diffusion process (from strong noise addition to weak noise addition), preventing the model from only adapting to the noise pattern of a certain stage. Indicates high-definition images , reconstruct the image , Mask Expectation: This layer expects operations to cover data diversity, so that loss calculation does not rely on a single image / mask, ensuring the model's generalization to different inputs and different key area masks;

[0056] Total loss of joint training: used in joint training in step S3.3, the above losses are combined by weight:

[0057] ;

[0058] in, 、 、 is a dynamically adjustable weight coefficient.

[0059] S4. Inference to generate high-quality images: Input image, resize the input image to 512×512; input text processing, the text must contain the gender keyword "man" or "woman", otherwise the wrong gender image may be generated due to factors such as the person's hairstyle; extract multi-scale image features through the image super-resolution module and the CLIP image encoder, extract semantic features through the text encoder, and then generate stacked ID embedding through the text-image alignment module; perform 50 steps of iterative denoising in the diffusion model to output high-quality restoration results.

[0060] Comparative analysis:

[0061] Under a consistent experimental setup, using the same input image and corresponding text prompt, the cross-modal alignment-based text and face collaborative restoration method proposed in this embodiment of the present invention was compared with various existing text-to-image algorithms. The method proposed in this embodiment of the present invention was visually closer to the target image. In the case of incomplete facial features (occluded by eyes, hats, etc., profile, heavy makeup, etc.), the method proposed in this embodiment of the present invention was compared with various existing text-to-image algorithms. The method can generate frontal face images with consistent identity features. Quantitative tests were also conducted, and the method consistently outperformed competing methods on key evaluation metrics, including Face Sim., DINO, and LPIPS, demonstrating its strong capabilities in identity preservation and facial detail reconstruction. More than ten methods were compared in the experiment, and only the methods with the best performance in terms of both quantitative metrics and visual effects are shown here, as shown in Table 1.

[0062] Table 1

[0063] In a comparison of three core metrics, Face Sim., DINO, and LPIPS, the embodiments of the present invention demonstrate superiority. Compared to methods such as InstantID, MoMA (ECCV 2024), ConsistentID, and DEADiff (CVPR 2024), the method of the present invention achieves 44.80% on the Face Sim. metric, significantly surpassing DEADiff (CVPR 2024)'s 33.75%, demonstrating superior capture and restoration of facial features. The DINO metric reaches 62.04%, significantly outperforming other methods and demonstrating superior visual feature representation. And in the LPIPS metric, with a value of 0.7152, it is lower than other methods, demonstrating minimal perceptual difference between the generated content and the real content, resulting in a more realistic visual effect. These data fully demonstrate that the embodiments of the present invention lead the way in face restoration tasks in terms of feature similarity, representation, and visual quality.

[0064] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A text and face collaborative restoration method based on cross-modal alignment, characterized by: The following steps are involved: Construct a text-face multimodal dataset; Based on a cross-modal network architecture, we achieve image super-resolution and text-image alignment in a text-face multimodal dataset. Train cross-modal networks and design and optimize hybrid loss functions; Inference generates high-quality images.

2. The method for collaborative text and face restoration based on cross-modal alignment according to claim 1, characterized in that: The steps of constructing the text-face multimodal dataset specifically include: Obtain high-quality face images from a public face dataset, and randomly select two different images for each person ID to form a large number of image pairs; For each image pair, a random high-quality image is generated by downgrading the quality of the corresponding low-quality image to form an LQ-HQ image pair. A structured text description, including character features, is manually annotated for each pair of LQ-HQ images, thus forming LQ-HQ-Text triplet data.

3. The method for collaborative text and face restoration based on cross-modal alignment according to claim 2, characterized in that: The steps of achieving image super-resolution and text-image alignment in a text-face multimodal dataset based on a cross-modal network architecture specifically include: Use HQ encoder and LQ encoder to extract features and compress images; Use the Transformer module to predict the codebook index sequence corresponding to the low-quality image; The HQ decoder is used to reconstruct high-quality images based on the compressed quantized features.

4. The method for collaborative text and face restoration based on cross-modal alignment according to claim 3, characterized in that: The steps of achieving image super-resolution and text-image alignment in a text-face multimodal dataset based on a cross-modal network architecture further include: Utilize image encoder to extract semantic embedding of input ID image; Leveraging a text encoder to extract multi-granular semantic embeddings of textual cues; The embeddings of multiple input ID images are fused with the classifier features in the text to generate a fused embedding vector, which is then concatenated along the length dimension to form a stacked ID embedding. Based on the diffusion model, the stacked ID embedding replaces the category word position in the text embedding to generate an updated text embedding. Through the cross-attention mechanism, the query vector interacts with the updated text embedding, allowing the model to adaptively integrate ID information. The generated image retains the ID information and conforms to the text description.

5. The method for collaborative text and face restoration based on cross-modal alignment according to claim 4, characterized in that: The steps of training the cross-modal network and designing an optimized hybrid loss function specifically include: Learning discrete codebooks through high-quality image self-reconstruction and training HQ encoder and HQ decoder simultaneously; Freeze the discrete codebook and HQ decoder, train the LQ encoder and Transformer modules, predict the codebook index sequence, and model the global facial structure; The fixed codebook and Transformer module are jointly trained with the image, text encoder and diffusion model, and the total loss of joint training is used to fine-tune the image encoder, text encoder and diffusion model parameters.

6. The method for collaborative text and face restoration based on cross-modal alignment according to claim 5, characterized in that: The total loss of joint training is a weighted combination of the total loss of codebook learning, the total loss of the Transformer module, and the mask diffusion loss; Among them, the total loss of codebook learning is used for optimization in the process of learning discrete codebooks, the total loss of Transformer modules is used for optimization in the process of predicting codebook index sequences, and mask diffusion loss is introduced in the process of joint training with image, text encoder and diffusion model.

Citation Information

Patent Citations

  • Face structure correlation based low-resolution face image restoration method

    CN106709874A

  • Face image restoration method based on generation diffusion prior

    CN118333866A

  • Face image restoration method, system, device and medium

    CN118333910A

  • Text-aligned human motion generation method and system

    CN119941942A

  • Character loss compensation processing device and method for video subtitle extraction

    CN120475225A