A method for blind face restoration based on geometrically enhanced cross-attention mechanism

By combining geometric enhancement cross-attention mechanism with geometric priors and high-quality features, the problem of detail loss and shape distortion in existing blind face restoration methods under complex degradation conditions is solved, and high-precision, natural facial image restoration is achieved.

CN119963449BActive Publication Date: 2025-10-31ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510050585.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-10-31
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Existing methods for restoring blind faces suffer from problems such as loss of detail and distortion of facial shape when faced with complex and diverse degradation situations, making it difficult to achieve high-quality, structurally accurate image restoration.

Method used

A geometrically enhanced cross-attention mechanism is adopted, which combines geometric prior information and high-quality features. Through multi-scale feature fusion, a repair model is constructed. Geometric prior information guides facial structure repair, while high-quality features provide detailed information.

Benefits of technology

It achieves high-precision and natural facial image restoration in complex degradation scenarios, maintaining facial structural consistency and detail integrity, and improving restoration results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963449B_ABST
    Figure CN119963449B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of image processing and discloses a method for blind face restoration based on a geometrically enhanced cross-attention mechanism. The method includes: acquiring a training dataset; extracting high-quality features and geometric prior information from high-quality image samples, while simultaneously extracting degradation features from low-quality image samples, and upsampling the low-quality image samples to extract degradation features; constructing a geometrically enhanced cross-attention mechanism as the restoration model, which is based on a multi-head cross-attention mechanism, using degradation features as queries, high-quality features as keys and values, and geometric prior information as keys and values; outputting low-resolution fusion features based on the degradation features in the original low-quality image samples; outputting high-resolution fusion features based on the degradation features in the upsampled low-quality image samples; and generating a restored high-quality image. This invention achieves accurate restoration of face images with missing details, blurriness, and severe deformation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing, and is particularly applied to the restoration and enhancement of facial images. Specifically, this invention proposes a blind face restoration method based on a geometrically enhanced cross-attention mechanism (GeoMHCA). This method employs GeoMHCA, which introduces new facial geometric priors to enable deep interaction between the degraded facial image and high-quality prior features. This deep interaction can more fully capture the spatial structure and contextual information of the face, thereby achieving a high-precision and natural restoration effect on degraded facial images. Background Technology

[0002] With the rapid development of deep learning and computer vision technologies, blind face restoration has become an important research direction, aiming to recover high-quality images from face images with unknown degradation. However, existing blind face restoration methods often suffer from problems such as loss of detail and distortion of facial shape when faced with complex and diverse degradation situations, resulting in unsatisfactory restoration results and failing to meet the needs of practical applications.

[0003] High-quality facial images are of significant value in entertainment, surveillance, and human-computer interaction, making the improvement of facial restoration technology a pressing need for multifunctional vision systems. RestoreFormer++, as one of the mainstream facial restoration methods, restores facial details through multi-scale feature fusion mechanisms; however, it primarily focuses on enhancing facial details, lacking effective modeling of the overall shape and structure. Therefore, when processing severely degraded or structurally damaged facial images, RestoreFormer++ often produces suboptimal restored images, failing to accurately preserve the spatial structure of the face. Summary of the Invention

[0004] The purpose of this invention is to provide a blind face restoration method based on a geometrically enhanced cross-attention mechanism, aiming to recover high-quality, structurally accurate facial images from low-quality degraded facial images. This method comprehensively utilizes geometric prior information and high-quality features, leveraging the geometrically enhanced cross-attention mechanism to achieve multi-scale feature fusion, thereby effectively addressing complex degradation scenarios and achieving accurate restoration of facial images with missing details, blurriness, and severe deformation.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A method for blind face restoration based on a geometrically enhanced cross-attention mechanism, the method comprising:

[0007] Obtain a training dataset, which includes pairs of high-quality image samples and low-quality image samples;

[0008] High-quality features and geometric prior information are extracted from high-quality image samples, while degradation features are extracted from low-quality image samples. Degradation features are extracted after upsampling the low-quality image samples.

[0009] A geometrically enhanced cross-attention mechanism is constructed as a repair model. The geometrically enhanced cross-attention mechanism is based on a multi-head cross-attention mechanism, and degenerate features are used as queries, high-quality features are used as keys and values, and geometric prior information is used as keys and values.

[0010] The degradation features in the original low-quality image samples are used as queries, high-quality features are used as keys and values, and geometric prior information is used as keys and values. The low-resolution fused features are output using the inpainting model.

[0011] The degraded features in the upsampled low-quality image samples are used as queries, the high-quality features are used as keys and values, and the geometric prior information is used as keys and values. The repair model outputs high-resolution fused features.

[0012] High-quality restored images are generated based on low-resolution and high-resolution fusion features. The restoration model is then optimized and updated based on the loss calculated from the restored high-quality images and high-quality image samples.

[0013] Based on the trained inpainting model, low-resolution fusion features and high-resolution fusion features are output for the image to be optimized, and a high-quality inpainted image is generated.

[0014] Several alternative methods are provided below, but they are not intended as additional limitations on the overall solution above. They are merely further additions or optimizations. Provided there are no technical or logical contradictions, each alternative method can be combined individually with respect to the overall solution above, or multiple alternative methods can be combined with each other.

[0015] Preferably, the method for generating low-quality image samples is as follows: high-quality image samples are processed using a degradation model to obtain corresponding low-quality image samples.

[0016] Preferably, the use of high-quality features as keys and values ​​includes:

[0017] Using convolutional neural networks to extract high-quality features from high-quality image samples;

[0018] High-quality features are encoded and quantized using a vector quantization algorithm to construct a high-quality feature dictionary. The keys and values ​​in the high-quality feature dictionary are then used as the keys and values ​​for the geometrically enhanced cross-attention mechanism.

[0019] Preferably, the step of using geometric prior information as keys and values ​​includes:

[0020] The geometric prior information includes key points obtained by processing high-quality image samples using a key point detection algorithm, facial geometric features extracted from a 3D face model generated based on high-quality image samples, and facial region segmentation maps obtained by processing high-quality image samples using a facial parsing model.

[0021] The key points and facial geometric features are used as keys, and the facial region segmentation map is used as values.

[0022] Preferably, the processing procedure of the geometrically enhanced cross-attention mechanism is as follows:

[0023]

[0024] In the formula, Z i This represents the output feature of the i-th head in the multi-head cross-attention mechanism, where i = 1, 2, ..., N. h -1, N h The multi-head cross-attention mechanism has a specified number of heads, softmax represents the softMax function, and Q represents the number of heads. i This indicates that the query is input at the i-th header. and This represents the key and value from the high-quality features of the input i-th head. and C represents the key and value from the geometric prior information of the i-th input head. h This indicates the number of channels per head;

[0025] The output features Z of all heads i The combined output feature Z is obtained by concatenation. mh ;

[0026] The final output is as follows:

[0027]

[0028] In the formula, Z fusion This represents the final output feature of the geometrically enhanced cross-attention mechanism, i.e., low-resolution fused feature or high-resolution fused feature. FFN represents the feedforward network, and LN represents the layer normalization operation. Indicates a degenerative characteristic.

[0029] Preferably, the process of generating the restored high-quality image based on low-resolution fusion features and high-resolution fusion features includes:

[0030] The final fused feature is obtained by multi-scale fusion of low-resolution fused features and high-resolution fused features;

[0031] The final fused features are input into the decoder to generate a high-quality repaired image. The decoder is optimized and updated synchronously with the repair model.

[0032] Preferably, the step of calculating the loss based on the repaired high-quality image and high-quality image samples includes:

[0033]

[0034] In the formula, L total The loss function represents the overall loss function. L represents pixel-level loss. per L represents perceived loss. geo L represents the geometric consistency loss. comp L represents the component's resistance to loss. adv L represents the overall combat loss. id Indicating loss of identity, L p λ represents the feature matching loss. per λ represents the weight of the perceptual loss. geo λ represents the weight of the geometric consistency loss. comp λ represents the component's adversarial loss weight. adv The weight λ represents the global adversarial loss. id λ represents the identity loss weight. p This represents the feature matching loss weights.

[0035] This invention proposes a blind face restoration method based on a geometrically enhanced cross-attention mechanism, building upon RestoreFormer++. By introducing geometric prior information about the face (such as facial contours, key points, and structural features) and combining it with high-quality features (providing detailed information), this method can effectively restore facial structure and details in complex degradation scenarios. The geometric prior provides guidance on the overall shape and spatial layout of the face during the restoration process, thereby reducing detail loss and deformation, and achieving more realistic, complete, and high-precision face image restoration. Attached Figure Description

[0036] Figure 1 This is a flowchart of a blind face restoration method based on a geometrically enhanced cross-attention mechanism according to the present invention. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention.

[0039] This embodiment proposes an improved multi-head cross-attention mechanism that enhances facial feature recovery by integrating geometric and high-quality priors. This mechanism utilizes geometric prior information (such as facial contours, keypoints, and structural features) to deeply interact with degraded facial images and high-quality priors, thereby better capturing spatial relationships and contextual information of the face. By performing cross-attention modeling on multi-scale features, the realism and detail fidelity of the recovered image can be effectively improved.

[0040] like Figure 1 As shown, this embodiment provides a method for blind face restoration based on a geometrically enhanced cross-attention mechanism, applied to image-based blind face restoration, including the following steps:

[0041] Step 1: Obtain the training dataset, which includes pairs of high-quality image samples and low-quality image samples.

[0042] High-quality image samples are derived from the CelebA-HQ face attribute dataset, which contains faces with various facial attributes. Low-quality image samples are generated using a degradation model to produce low-quality images, including simulated blur, noise, compression artifacts, and other complex scenes. This prepares the model for training, aiming to improve the inpainting model and generate more realistic, higher-quality restored images.

[0043] Step 2: Extract high-quality features and geometric prior information from high-quality image samples. At the same time, use a convolutional neural network (CNN) to extract degradation features from low-quality image samples, and then upsample the low-quality image samples to extract degradation features.

[0044] Step 3: Construct a geometrically enhanced cross-attention mechanism as the repair model. The geometrically enhanced cross-attention mechanism is based on the multi-head cross-attention mechanism (MHCA), which uses degenerate features as queries, high-quality features as keys and values, and geometric prior information as keys and values.

[0045] This embodiment improves the multi-head cross-attention mechanism by introducing geometric prior information to construct a geometrically enhanced cross-attention mechanism. When constructing this mechanism, the geometric prior information of the face (such as facial geometric features, key points, and facial region segmentation maps) is combined with prior features (providing detailed information) from a high-quality feature dictionary to generate a high-quality restored image. Unlike traditional multi-head self-attention mechanisms, GeoMHCA uses the degradation features of low-quality images as the query, the prior features from the high-quality dictionary as the key and value, and introduces geometric prior information as additional keys and values ​​to provide guidance on geometric structure. Specifically, this includes:

[0046] Degradation features of low-quality images are used as queries.

[0047] Prior features in a high-quality feature dictionary serve as keys and values;

[0048] At the same time, geometric prior information is introduced as additional keys and values ​​to help maintain the facial geometry during the repair process.

[0049] The high-quality feature dictionary uses prior features from high-quality image samples as keys and values. The high-quality feature dictionary is composed of high-quality features extracted from these high-quality image samples. The generation process includes using a convolutional neural network (CNN) to extract high-quality features from these high-quality image samples. These extracted high-quality features contain rich facial details, such as local features and overall shapes of the eyes, mouth, and nose.

[0050] High-quality features are encoded and quantized using the Vector Quantization (VQ) algorithm, constructing a diverse high-quality feature dictionary. The keys and values ​​in this dictionary are then used as keys and values ​​for a geometrically enhanced cross-attention mechanism. The prior features in the high-quality feature dictionary provide high-quality prior information that matches the corresponding regions of low-quality image samples, thus supporting fine-grained restoration and ensuring that rich facial details better support the image restoration process.

[0051] Geometric prior information is introduced as additional keys and values. This information provides guidance on facial geometry, ensuring that the restored high-quality image conforms geometrically to facial anatomy. The specific method for obtaining this geometric prior information is as follows:

[0052] Facial landmark detection: High-quality image samples are processed using keypoint detection algorithms (such as the Harris corner detector) to obtain key points in the face, such as the corners of the eyes, the tip of the nose, and the corners of the mouth. These key points provide the basic geometric reference for the face.

[0053] 3D face model generation: Generate 3D face models from 2D images (e.g., deep learning-based methods), extract facial geometric features from the 3D face models (e.g., Local Binary Pattern (LBP), Histogram of Gradient Orientation (HOG), including contour, depth, and three-dimensional geometric structure information of facial regions.

[0054] Facial segmentation: Facial regions are segmented into different parts (such as eyes, nose, mouth, etc.) using a facial segmentation model (e.g., multi-task cascaded convolutional network) to generate facial region segmentation maps. These maps supplement geometric prior information, further enriching it and providing structural guidance during the restoration process, which helps maintain the realism of facial shapes and regions.

[0055] Key points and facial geometric features are used as keys for the geometric prior; facial region segmentation maps are used as values ​​for the geometric prior.

[0056] Step 4, Multi-scale feature fusion: In order to ensure that the restored high-quality image retains both global structure and rich detail information, this embodiment performs multi-scale fusion of degradation features, high-quality feature dictionary and geometric prior information at different resolutions to ensure that the model captures both details and structure.

[0057] Step 4.1: Use the degraded features in the original low-quality image samples as queries, the high-quality features as keys and values, and the geometric prior information as keys and values, and use the repair model to output low-resolution fused features.

[0058] Low resolution is primarily used to capture the overall geometry of the face, such as facial contours and key structural features. The restoration model retrieves facial shape information similar to the degraded feature structures from a high-quality feature dictionary using GeoMHCA and combines this with geometric prior information for restoration. The specific steps are as follows:

[0059] In the inputs of GeoMHCA, besides the degenerate feature inputs and high quality features It also incorporates geometric prior information.

[0060] Query: Input from degeneracy features

[0061] Keys and values: from a high-quality feature dictionary and geometric prior information Extract from.

[0062] In the calculation of GeoMHCA, geometric priors are taken into account when calculating attention:

[0063]

[0064] In the formula, Z i Let represent the output feature of the i-th head in the multi-head cross-attention mechanism, and represent the high-quality fused feature, i = 1, 2, ..., N. h -1, N h The multi-head cross-attention mechanism has a specified number of heads, softmax represents the softMax function, and Q represents the number of heads. i This represents the query for the i-th input header, which consists of degradation features extracted from low-quality image samples, representing the information that needs to be recovered. and This represents the key and value from the high-quality features of the input i-th head. and C represents the key and value from the geometric prior information of the i-th input head. h This represents the number of channels per head, typically set to C / C. h V represents the number of channels in the feature map.

[0065] The output features Z of all heads i The combined output feature Z is obtained by concatenation. mh =concat Z i i = 1, 2, ..., N h -1.

[0066] The final output of the multi-head cross-attention mechanism is added back to the input features and then processed through normalization and a feedforward network, as shown in the following formula:

[0067]

[0068] In the formula, Z fusion This represents the final output feature of the geometrically enhanced cross-attention mechanism, i.e., low-resolution fused feature or high-resolution fused feature. FFN represents a feedforward network implemented by two convolutional layers, and LN represents the layer normalization operation. Indicates a degenerative characteristic.

[0069] Step 4.2: Use the degraded features in the upsampled low-quality image samples as queries, the high-quality features as keys and values, and the geometric prior information as keys and values, and use the repair model to output high-resolution fused features.

[0070] At high resolution, the restoration model focuses on facial details such as the texture and shape of the eyes, mouth, and nose. The specific steps are as follows: After low-resolution feature fusion, the restoration model takes a higher-resolution image as input and upsamples the low-quality image samples to obtain a higher-resolution image. The improved low-quality image is used as the query, and high-quality prior information and geometric prior information are used as keys and values ​​for GeoMHCA calculation. The calculation process is consistent with that in low-resolution mode, and will not be repeated in this embodiment.

[0071] Step 5: Generate a high-quality restored image based on low-resolution fusion features and high-resolution fusion features, and optimize and update the restoration model by calculating the loss based on the restored high-quality image and high-quality image samples.

[0072] The final fused feature is obtained by multi-scale fusion of low-resolution and high-resolution fusion features. This final fused feature is then input into the decoder to generate a high-quality restored image. The high-low resolution fusion method significantly improves the effect of facial image restoration, especially in complex degraded scenes where it exhibits better geometric consistency and detail fidelity.

[0073] During the repair process, the overall loss function expression is:

[0074]

[0075] In the formula, L total The loss function represents the overall loss function. L represents pixel-level loss. per L represents perceived loss. geo L represents the geometric consistency loss. comp L represents the component's resistance to loss. adv L represents the overall combat loss. id Indicating loss of identity, L p λ represents the feature matching loss. per λ represents the weight of the perceptual loss. geo λ represents the weight of the geometric consistency loss. comp λ represents the component's adversarial loss weight. adv The weight λ represents the global adversarial loss. id λ represents the identity loss weight. p This represents the feature matching loss weights.

[0076] Among them, pixel-level loss To ensure consistency of base pixels, the following calculation is performed:

[0077]

[0078] In the formula, I hFor high-quality image samples, For the restored high-quality image, |||1 represents the L1 norm, which is the sum of absolute differences.

[0079] Among them, the perceived loss L per To improve detail quality and visual realism, the calculations are as follows:

[0080]

[0081] In the formula, Let denote the intermediate feature extractor of the pre-trained VGG network, and |||2 denotes the L2 norm.

[0082] Among them, the geometric consistency loss L geo Used to constrain geometry, calculated as follows:

[0083]

[0084] In the formula, E(·) represents the edge detection algorithm.

[0085] Among them, component adversarial loss L comp The calculation is as follows:

[0086]

[0087] In the formula, D r (·) represents a discriminator for a local region (such as key areas like the eyes and mouth), R r (·) indicates a local region extraction operation, and r represents the label or index of each local region, including the left eye region, right eye region, mouth region, and nose region.

[0088] Among them, the global adversarial loss L adv This is used to enhance the overall realism of the restored image, and is calculated as follows:

[0089]

[0090] In the formula, D(·) represents the global discriminator, which is used to distinguish between real images and generated images.

[0091] Among them, identity loss L id To maintain facial identity consistency, the calculation is as follows:

[0092]

[0093] In the formula, η(·) represents the identity feature extractor.

[0094] Among them, the feature matching loss Lp The calculation is as follows:

[0095]

[0096] By using the geometric consistency loss function, the restoration model can not only restore facial details but also maintain the consistency of facial structure, ensuring that the restored image is more visually realistic and natural.

[0097] Introducing the geometric consistency loss function L geo In this context, the aim is to enhance the quality and consistency of the restored image by utilizing geometric information. By measuring the similarity between the restored image and high-quality image samples, the restoration result is made to conform to the true or expected geometric features in terms of geometric structure, ensuring that the restored image is more natural and realistic in shape and structure.

[0098] The restoration model and decoder are trained using a large number of low-quality and high-quality image samples, simulating different types of degradation using a degradation model. The restoration model is progressively optimized through training and evaluated on a validation set. Finally, the parameter structure of the restoration model is adjusted to improve restoration results and the ability to handle complex degraded images.

[0099] Step 6: Based on the trained inpainting model, output low-resolution fusion features and high-resolution fusion features for the image to be optimized, and generate a high-quality inpainted image.

[0100] The trained restoration model can handle blind face restoration problems in various degradation scenarios. After outputting low-resolution and high-resolution fusion features of the image to be optimized, the low-resolution and high-resolution fusion features are fused at multiple scales to obtain the final fusion feature. The final fusion feature is then input into the decoder to generate a high-quality restored image.

[0101] Even in low-quality images with severe blurring, noise, or compression artifacts, the restoration model can effectively recover high-quality, natural images with complete facial structure. This makes the method of this invention highly suitable for image restoration tasks in surveillance videos and face recognition systems with poor image quality, and it can also be widely applied in fields such as photo restoration and identity verification.

[0102] This invention employs a geometrically enhanced cross-attention mechanism (GMHCA), which optimizes the restoration process of degraded facial images by introducing geometric prior information and a high-quality feature dictionary. Geometric prior information provides facial structure reference, while the high-quality feature dictionary provides detail restoration. GeoMHCA fuses these two types of prior information to ensure the realism and detail integrity of facial images in complex degradation scenes. Geometric prior information provides guidance on shape and structure, while the high-quality feature dictionary provides detail and realism. During the calculation of GeoMHCA, geometric prior information influences the allocation of attention weights, allowing the restoration model to better determine which regions to focus attention on, thus obtaining a superior restored image.

[0103] The present invention conducts the following experiments:

[0104] (1) Experimental equipment: This experiment was implemented using the PyTorch framework and trained on a server equipped with an NVIDIA RTX 3090 GPU (24GB VRAM). The system was equipped with an Intel Xeon E5-2620 CPU and 128GB of memory, and the operating system was Ubuntu 18.04.

[0105] (2) Experimental setup:

[0106] High-quality image sample set: Due to the memory limitations of the experimental equipment, the training set used in this experiment is CelebA-HQ, which is 30,000 high-resolution face images of 1024*1024 resolution obtained from the Face Attribute dataset (CelebFaces Attribute, CelebA). It also includes face images of various face attributes in CelebA-HQ. The image size was adjusted to 512*512 during training.

[0107] Low-quality image sample set: generated using a degradation model (to produce low-quality images).

[0108]

[0109] In the formula, I d For low-quality image samples, I h For high-quality image samples, k σ Let be the Gaussian blur kernel (σ is the standard deviation, σ∈[0.2,10]), ↓ r For downsampling (r is the sampling factor, r∈[1,8]), n δ For Gaussian noise (δ is the intensity, δ∈[0,20]), JPEG compression: quality factor q (q∈[60,100]), ↑ r Upsample back to the original resolution.

[0110] Test training set: The dataset used to test the qualitative and quantitative results of face restoration is Helen, which contains 2330 images with multiple faces.

[0111] (3) Training settings: The Adam optimizer is used, the initial learning rate is set to 0.0001, and the weight parameters in the loss function are set to λ. per =1.0, λ p =0.25, λ adv =0.8, λ comp =1.0, λ id =1.0, λ geo =0.5.

[0112] (4) This experiment uses the conventional multi-head cross-attention mechanism and the geometrically enhanced cross-attention mechanism of this invention, and uses the original high-quality image samples as the control group. The experimental results are shown in Table 1:

[0113] Table 1 Experimental Results

[0114]

[0115] Comparison of image quality metrics and parameter counts for different attention mechanisms on the Helen dataset: SSIM (Structural Similarity Index) ↑: Measures structural similarity. IDD (Identity Distance Index) ↓: Measures the similarity of the restored image to the real image in terms of identity features. LPIPS (Perceptual Image Quality Index) ↓: Measures perceptual image differences. PSNR (Peak Signal-to-Noise Ratio) ↑: Measures pixel-level restoration quality.

[0116] (5) Experimental conclusions:

[0117] The geometrically enhanced cross-attention mechanism outperforms the multi-head cross-attention mechanism across all evaluation metrics: SSIM indicates better structural similarity and multi-scale structural recovery. Lower LPIPS means the generated image's perceptual quality is closer to the real image. Slightly higher PSNR indicates better pixel-level restoration quality. Slightly lower IDD indicates the restored image is closer to the real image in preserving identity features. Experiments show that the geometrically enhanced cross-attention mechanism of this invention has stronger capabilities in restoring image structure and perceptual quality than the ordinary multi-head cross-attention mechanism, providing stronger support for image structural information reconstruction and identity feature preservation.

[0118] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0119] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A method for blind face restoration based on a geometrically enhanced cross-attention mechanism, characterized in that, The blind face restoration method based on geometrically enhanced cross-attention mechanism includes: Obtain a training dataset, which includes pairs of high-quality image samples and low-quality image samples; High-quality features and geometric prior information are extracted from high-quality image samples, while degradation features are extracted from low-quality image samples. Degradation features are extracted after upsampling the low-quality image samples. A geometrically enhanced cross-attention mechanism is constructed as a repair model. The geometrically enhanced cross-attention mechanism is based on a multi-head cross-attention mechanism, and degenerate features are used as queries, high-quality features are used as keys and values, and geometric prior information is used as keys and values. The degradation features in the original low-quality image samples are used as queries, high-quality features are used as keys and values, and geometric prior information is used as keys and values. The low-resolution fused features are output using the inpainting model. The degraded features in the upsampled low-quality image samples are used as queries, the high-quality features are used as keys and values, and the geometric prior information is used as keys and values. The repair model outputs high-resolution fused features. High-quality restored images are generated based on low-resolution and high-resolution fusion features. The restoration model is then optimized and updated based on the loss calculated from the restored high-quality images and high-quality image samples. Based on the trained inpainting model, low-resolution fusion features and high-resolution fusion features are output for the image to be optimized, and a high-quality inpainted image is generated.

2. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that, The method for generating low-quality image samples is as follows: high-quality image samples are processed using a degradation model to obtain corresponding low-quality image samples.

3. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that, The use of high-quality features as keys and values ​​includes: Using convolutional neural networks to extract high-quality features from high-quality image samples; High-quality features are encoded and quantized using a vector quantization algorithm to construct a high-quality feature dictionary. The keys and values ​​in the high-quality feature dictionary are then used as the keys and values ​​for the geometrically enhanced cross-attention mechanism.

4. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that, The use of geometric prior information as keys and values ​​includes: The geometric prior information includes key points obtained by processing high-quality image samples using a key point detection algorithm, facial geometric features extracted from a 3D face model generated based on high-quality image samples, and facial region segmentation maps obtained by processing high-quality image samples using a facial parsing model. The key points and facial geometric features are used as keys, and the facial region segmentation map is used as values.

5. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that, The processing procedure of the geometrically enhanced cross-attention mechanism is as follows: In the formula, Z i This represents the output feature of the i-th head in the multi-head cross-attention mechanism, where i = 1, 2, ..., N. h -1, N h The multi-head cross-attention mechanism has a specified number of heads, softmax represents the softMax function, and Q represents the number of heads. i This indicates that the query is input at the i-th header. and This represents the key and value from the high-quality features of the input i-th head. and C represents the key and value from the geometric prior information of the i-th input head. h This indicates the number of channels per head; The output features Z of all heads i The combined output feature Z is obtained by concatenation. mh ; The final output is as follows: In the formula, Z fusion This represents the final output feature of the geometrically enhanced cross-attention mechanism, i.e., low-resolution fused feature or high-resolution fused feature. FFN represents the feedforward network, and LN represents the layer normalization operation. Indicates a degenerative characteristic.

6. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that, The process of generating a high-quality restored image based on low-resolution fusion features and high-resolution fusion features includes: The final fused feature is obtained by multi-scale fusion of low-resolution fused features and high-resolution fused features; The final fused features are input into the decoder to generate a high-quality repaired image. The decoder is optimized and updated synchronously with the repair model.

7. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that, The loss is calculated based on the repaired high-quality image and high-quality image samples. include: In the formula, L total The loss function represents the overall loss function. L represents pixel-level loss. per L represents perceived loss. geo L represents the geometric consistency loss. comp L represents the component's resistance to loss. adv L represents the overall combat loss. id Indicating loss of identity, L p λ represents the feature matching loss. per λ represents the weight of the perceptual loss. geo λ represents the weight of the geometric consistency loss. comp λ represents the component's adversarial loss weight. adv The weight λ represents the global adversarial loss. id λ represents the identity loss weight. p This represents the feature matching loss weights.

Citation Information

Patent Citations

  • Real degraded image blind restoration method based on cross attention mechanism

    CN115829876A

  • Structure and texture feature guided double-encoder image restoration method

    CN116523985A