Blind person face repairing method based on geometric enhanced cross attention mechanism

By introducing a geometrically enhanced cross attention mechanism in the blind face repair method, combining geometric prior information and high-quality features, the problems of detail loss and shape distortion in the existing technology are solved, and a high-precision and natural facial image repair effect is achieved.

CN119963449AActive Publication Date: 2025-05-09ZHEJIANG UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510050585.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

When existing blind face repair methods face complex and diverse degradation, they usually have problems such as loss of details and distortion of facial shapes, resulting in unsatisfactory repair results and difficult to meet the actual application needs.

Method used

The blind face repair method based on the geometrically enhanced cross attention mechanism is adopted. By introducing geometric prior information of the face and high-quality features, the geometrically enhanced cross attention mechanism is used to achieve multi-scale feature fusion, thereby effectively responding to complex degraded scenarios and achieving accurate repair of face images with missing details, blurred and severe deformation.

Benefits of technology

It realizes high-precision and natural repair effects for degenerated face images, which can better capture the spatial structure and context information of the face, reduce details loss and deformation, and generate more realistic and complete high-quality face images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963449A_ABST
    Figure CN119963449A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of image processing, and discloses a blind person face restoration method based on a geometric enhanced cross attention mechanism, which comprises the following steps: acquiring a training data set; high-quality features and geometric prior information in the high-quality image samples are extracted, degradation features in the low-quality image samples are extracted at the same time, and the degradation features are extracted after up-sampling is carried out on the low-quality image samples; a geometric enhanced cross attention mechanism is constructed as a repair model, the geometric enhanced cross attention mechanism is based on a multi-head cross attention mechanism, the degradation features are used as queries, the high-quality features are used as key sum values, and meanwhile geometric prior information is used as key sum values; outputting a low-resolution fusion feature based on the degradation feature in the original low-quality image sample; outputting a high-resolution fusion feature based on the degradation feature in the up-sampled low-quality image sample; and a restored high-quality image is generated. According to the invention, accurate restoration of the face image with detail missing, blurring and serious deformation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and is particularly applied to the restoration and enhancement technology of face images. Specifically, the present invention proposes a blind face restoration method based on geometrically enhanced cross-attention mechanism, which adopts geometrically enhanced cross-attention mechanism (GeoMHCA) and introduces new face geometry prior to enable the degraded face image to interact deeply with high-quality prior features. This deep interaction can more fully capture the spatial structure and contextual information of the face, thereby achieving a high-precision and natural restoration effect on the degraded face image. Background Art

[0002] With the rapid development of deep learning and computer vision technology, blind face restoration has become one of the important research directions, and its goal is to restore high-quality images from unknown degraded face images. However, when faced with complex and diverse degradation situations, existing blind face restoration methods usually suffer from problems such as loss of details and distortion of facial shapes, resulting in unsatisfactory restoration effects and difficulty in meeting practical application needs.

[0003] High-quality facial images are of great value in entertainment, surveillance, and human-computer interaction. Therefore, improving the performance of face restoration technology has become an urgent need for multifunctional vision systems. As one of the current mainstream face restoration methods, RestoreFormer++ restores facial details through a multi-scale feature fusion mechanism, but it mainly focuses on enhancing facial details and lacks effective modeling of overall shape and structure. Therefore, when RestoreFormer++ processes severely degraded or structurally damaged facial images, the restored image effect is often poor and cannot accurately maintain the spatial structure of the face. Summary of the invention

[0004] The purpose of the present invention is to provide a blind face restoration method based on geometrically enhanced cross-attention mechanism, aiming to restore high-quality and structurally accurate facial images from low-quality degraded facial images. The method comprehensively utilizes geometric prior information and high-quality features, and uses the geometrically enhanced cross-attention mechanism to achieve multi-scale feature fusion, thereby effectively coping with complex degraded scenes and achieving accurate restoration of facial images with missing details, blur and severe deformation.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] A blind face restoration method based on geometrically enhanced cross-attention mechanism, the blind face restoration method based on geometrically enhanced cross-attention mechanism comprising:

[0007] Acquire a training data set, wherein the training data set includes pairs of high-quality image samples and low-quality image samples;

[0008] Extract high-quality features and geometric prior information from high-quality image samples, extract degradation features from low-quality image samples, and extract degradation features after upsampling low-quality image samples;

[0009] Constructing a geometry-enhanced criss-cross attention mechanism as a restoration model, wherein the geometry-enhanced criss-cross attention mechanism is based on a multi-head criss-cross attention mechanism and takes degraded features as queries, high-quality features as keys and values, and geometric prior information as keys and values;

[0010] The degradation features in the original low-quality image samples are used as queries, the high-quality features are used as keys and values, and the geometric prior information is used as keys and values, and the restoration model is used to output low-resolution fused features;

[0011] The degradation features in the upsampled low-quality image samples are used as queries, the high-quality features are used as keys and values, and the geometric prior information is used as keys and values, and the restoration model is used to output high-resolution fused features;

[0012] Generate a restored high-quality image based on low-resolution fusion features and high-resolution fusion features, and optimize and update the restoration model based on the loss calculated based on the restored high-quality image and high-quality image samples;

[0013] Based on the trained restoration model, low-resolution fusion features and high-resolution fusion features are output for the image to be optimized, and a high-quality restored image is generated.

[0014] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution, but are merely further supplements or preferences. Under the premise that there are no technical or logical contradictions, each optional method can be combined with the above-mentioned overall solution separately, and multiple optional methods can also be combined.

[0015] Preferably, the method for generating the low-quality image samples is: using a degradation model to process high-quality image samples to obtain corresponding low-quality image samples.

[0016] Preferably, the method of using high-quality features as keys and values ​​includes:

[0017] Use convolutional neural networks to extract high-quality features from high-quality image samples;

[0018] The high-quality features are encoded and quantized through a vector quantization algorithm to construct a high-quality feature dictionary. The keys and values ​​in the high-quality feature dictionary are used as the keys and values ​​of the geometric enhanced cross-attention mechanism.

[0019] Preferably, the method of using geometric prior information as a key and a value includes:

[0020] The geometric prior information includes key points obtained by processing high-quality image samples using a key point detection algorithm, facial geometric features extracted from a three-dimensional face model generated based on the high-quality image samples, and a facial region segmentation map obtained by processing high-quality image samples using a facial analysis model;

[0021] The key points and facial geometric features are used as keys, and the facial region segmentation map is used as a value.

[0022] Preferably, the processing process of the geometrically enhanced cross-attention mechanism is as follows:

[0023]

[0024] In the formula, Z i Represents the output features of the i-th head in the multi-head cross attention mechanism, i = 1, 2, ..., N h -1, N h represents the number of heads of the multi-head cross attention mechanism, softmax represents the softMax function, Q i represents the query for the input i-th head, and represents the keys and values ​​from the high-quality features input to the i-th head, and represents the key and value from the geometric prior information input to the i-th head, C h Indicates the number of channels per head;

[0025] The output features Z of all heads i The comprehensive output feature Z is obtained by connecting mh ;

[0026] The final output is as follows:

[0027]

[0028] In the formula, Z fusion It represents the final output features of the geometrically enhanced cross-attention mechanism, that is, low-resolution fusion features or high-resolution fusion features. FFN represents the feedforward network, and LN represents the layer normalization operation. Indicates a degenerate feature.

[0029] Preferably, the generating of the restored high-quality image based on the low-resolution fusion features and the high-resolution fusion features comprises:

[0030] The final fusion feature is obtained by multi-scale fusion of low-resolution fusion features and high-resolution fusion features;

[0031] The final fused features are input into a decoder to generate a high-quality restored image, and the decoder is optimized and updated synchronously with the restoration model.

[0032] Preferably, the calculating the loss according to the restored high-quality image and the high-quality image sample comprises:

[0033]

[0034] Where, L total represents the overall loss function, represents pixel-level loss, L per represents the perceptual loss, L geo represents the geometric consistency loss, L comp represents the component adversarial loss, L adv represents the global adversarial loss, L id represents identity loss, L p represents the feature matching loss, λ per represents the perceptual loss weight, λ geo represents the geometric consistency loss weight, λ comp represents the component adversarial loss weight, λ adv represents the weight of the global adversarial loss, λ id represents the identity loss weight, λ p Represents the feature matching loss weight.

[0035] Based on RestoreFormer++, this paper proposes a blind face restoration method based on geometrically enhanced cross-attention mechanism. By introducing geometric prior information of the face (such as facial contour, key points and structural features) and combining it with high-quality features (providing detail information), this method can effectively restore facial structure and details in complex degraded scenes. The geometric prior can provide guidance on the overall shape and spatial layout of the face for the restoration process, thereby reducing detail loss and deformation, and achieving more realistic, complete and high-precision face image restoration. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flow chart of a blind face restoration method based on a geometrically enhanced cross-attention mechanism of the present invention. DETAILED DESCRIPTION

[0037] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0039] This embodiment proposes an improved multi-head cross-attention mechanism that integrates the geometric prior and high-quality prior of the face to enhance the restoration of facial features. The mechanism uses geometric prior information (such as facial contours, key points, and structural features) to interact deeply with degraded facial images and high-quality priors to better capture the spatial relationship and contextual information of the face. By performing cross-attention modeling on multi-scale features, the realism and detail fidelity of the restored image can be effectively improved.

[0040] like Figure 1 As shown, this embodiment provides a blind face restoration method based on geometrically enhanced cross-attention mechanism, which is applied to image-based blind face restoration, and includes the following steps:

[0041] Step 1: Obtain a training data set, which includes pairs of high-quality image samples and low-quality image samples.

[0042] High-quality image samples come from the CelebA-HQ face attribute dataset, which contains faces with many different facial attributes. Low-quality image samples are generated by the degradation model: including simulated blur, noise, compression artifacts and other complex scenes. Prepare for training to improve the restoration model to generate more realistic and higher-quality restored images.

[0043] Step 2: extract high-quality features and geometric prior information from high-quality image samples, and use a convolutional neural network (CNN) to extract degradation features from low-quality image samples, and extract degradation features after upsampling the low-quality image samples.

[0044] Step 3: Construct a geometry-enhanced criss-cross attention mechanism as the restoration model. The geometry-enhanced criss-cross attention mechanism is based on the multi-head criss-cross attention mechanism (MHCA), which takes the degraded features as queries, the high-quality features as keys and values, and the geometric prior information as keys and values.

[0045] This embodiment improves the multi-head cross-attention mechanism, introduces geometric prior information, and constructs a geometrically enhanced cross-attention mechanism. When constructing the geometrically enhanced cross-attention mechanism, the geometric prior information of the face (such as facial geometric features, key points, and facial region segmentation maps) is combined with the prior features in the high-quality feature dictionary (providing detailed information) to generate a restored high-quality image. Unlike the traditional multi-head self-attention mechanism, GeoMHCA uses the degraded features of low-quality images as queries, the prior features in the high-quality dictionary as keys and values, and introduces geometric prior information as additional keys and values ​​to provide guidance on the geometric structure. Specifically including:

[0046] The degraded features of low-quality images are used as queries;

[0047] Prior features in the high-quality feature dictionary as keys and values;

[0048] At the same time, geometric prior information is introduced as additional keys and values ​​to assist in maintaining the facial geometry during the restoration process.

[0049] The prior features in the high-quality feature dictionary are used as keys and values, including: the high-quality feature dictionary is composed of high-quality features extracted from high-quality image samples. The high-quality feature dictionary generation process includes: using a convolutional neural network (CNN) to extract high-quality features from high-quality image samples, and the extracted high-quality features contain rich facial details, such as local features and overall shapes of eyes, mouths, and noses.

[0050] The high-quality features are encoded and quantized through the vector quantization (VQ) algorithm to construct a diverse high-quality feature dictionary, and the keys and values ​​in the high-quality feature dictionary are used as the keys and values ​​of the geometric enhanced cross-attention mechanism. The prior features in the high-quality feature dictionary can provide high-quality prior information that matches the corresponding areas of the low-quality image samples, thereby supporting fine restoration, thereby ensuring that rich facial details can better support the image restoration process.

[0051] The geometric prior information is introduced as additional key and value, including: The geometric prior information provides guidance on the geometric structure of the face, which is used to ensure that the restored high-quality image conforms to the facial anatomical features in terms of geometry. The specific method of obtaining the geometric prior information is as follows:

[0052] Facial landmark detection: High-quality image samples are processed through key point detection algorithms (such as Harris corner detector) to obtain key points in the face, such as the corners of the eyes, the tip of the nose, the corners of the mouth, etc. These key points provide the basic geometric reference of the face.

[0053] 3D face model generation: Generate a 3D face model using a 2D image (e.g., deep learning-based methods), extract facial geometric features in the 3D face model (e.g., local binary pattern LBP, histogram of oriented gradients HOG), including contour, depth, and three-dimensional geometric structure information of the facial area.

[0054] Facial parsing: Use a facial parsing model (such as a multi-task cascaded convolutional network) to segment the facial area into different parts (such as eyes, nose, mouth, etc.) and generate a facial area segmentation map. The facial area segmentation map serves as a supplement to the geometric prior information, thereby further enriching the geometric prior information. It can provide structural guidance during the restoration process and help maintain the authenticity of the facial shape and area.

[0055] The key points and facial geometric features are used as the key of the geometric prior; the facial region segmentation map is used as the value of the geometric prior.

[0056] Step 4, multi-scale feature fusion: In order to ensure that the restored high-quality image retains both the global structure and rich detail information, this embodiment performs multi-scale fusion of degraded features, high-quality feature dictionaries and geometric prior information at different resolutions to ensure that the model captures both details and structures.

[0057] Step 4.1: Use the degraded features in the original low-quality image sample as the query, the high-quality features as the key and value, and the geometric prior information as the key and value, and use the restoration model to output the low-resolution fused features.

[0058] Low resolution is mainly used to capture the overall geometric shape of the face, such as facial contours and main structural features. The repair model retrieves facial shape information similar to the degraded feature structure from the high-quality feature dictionary through GeoMHCA, and repairs it in combination with geometric prior information. The specific steps are as follows:

[0059] In the input of GeoMHCA, in addition to the degenerate feature input and high quality features Geometric prior information is also introduced

[0060] Query: From degenerate feature input

[0061] Keys and values: From the high-quality feature dictionary and geometric prior information Extracted from.

[0062] In the calculation of GeoMHCA, the geometric prior is taken into account when calculating the attention:

[0063]

[0064] In the formula, Z i Represents the output features of the i-th head in the multi-head cross attention mechanism, representing the fused high-quality features, i = 1, 2, ..., N h -1, N h represents the number of heads of the multi-head cross attention mechanism, softmax represents the softMax function, Q i represents the query input to the i-th head, which is the degraded features extracted from the low-quality image samples and represents the information that needs to be restored. and represents the keys and values ​​from the high-quality features input to the i-th head, and represents the key and value from the geometric prior information input to the i-th head, C h Indicates the number of channels per head, usually set to C / C h , V represents the number of channels of the feature map.

[0065] The output features Z of all heads i The comprehensive output feature Z is obtained by connecting mh =concat Z i ,i=1,2,…,N h -1.

[0066] The final multi-head cross-attention mechanism output is added back to the input features and processed by normalization and feed-forward network. The formula is as follows:

[0067]

[0068] In the formula, Z fusion represents the final output features of the geometrically enhanced cross-attention mechanism, i.e., low-resolution fusion features or high-resolution fusion features. FFN represents a feedforward network implemented by two convolutional layers. LN represents a layer normalization operation. Indicates a degenerate feature.

[0069] Step 4.2: Use the degraded features in the upsampled low-quality image samples as queries, the high-quality features as keys and values, and the geometric prior information as keys and values, and use the restoration model to output high-resolution fused features.

[0070] At high resolution, the restoration model focuses on the details of the face, such as the texture and shape of the eyes, mouth, and nose. The specific steps are as follows: After completing the low-resolution feature fusion, the restoration model will input a higher-resolution image, upsample the low-quality image samples, and obtain a higher-resolution image. The low-quality image with improved resolution is used as a query, and the high-quality prior information and geometric prior information are used as keys and values ​​to calculate GeoMHCA. The calculation process is consistent with that in low resolution, and this embodiment will not be repeated.

[0071] Step 5: Generate a restored high-quality image based on the low-resolution fusion features and the high-resolution fusion features, and optimize and update the restoration model according to the loss calculated based on the restored high-quality image and the high-quality image samples.

[0072] The final fusion feature is obtained by multi-scale fusion of low-resolution fusion features and high-resolution fusion features; the final fusion feature is input into the decoder to generate a high-quality restored image. The high-low resolution fusion method significantly improves the effect of facial image restoration, especially in complex degradation scenes with better geometric consistency and detail fidelity.

[0073] During the repair process, the overall loss function expression of the loss function is:

[0074]

[0075] Where, L total represents the overall loss function, represents pixel-level loss, L per represents the perceptual loss, L geo represents the geometric consistency loss, L comp represents the component adversarial loss, L adv represents the global adversarial loss, L id represents identity loss, L p represents the feature matching loss, λ per represents the perceptual loss weight, λ geo represents the geometric consistency loss weight, λ comp represents the component adversarial loss weight, λ adv represents the weight of the global adversarial loss, λ id represents the identity loss weight, λ p Represents the feature matching loss weight.

[0076] Among them, pixel level loss To ensure basic pixel consistency, the calculation is as follows:

[0077]

[0078] In the formula, I hFor high-quality image samples, is the high-quality image after restoration, ‖‖1 represents the L1 norm, the sum of absolute differences.

[0079] Among them, the perceptual loss L per , improves detail quality and visual realism, calculated as follows:

[0080]

[0081] In the formula, represents the intermediate feature extractor of the pre-trained VGG network and ‖‖2 represents the L2 norm.

[0082] Among them, the geometric consistency loss L geo , used to constrain the geometry, is calculated as follows:

[0083]

[0084] Where E(·) represents the edge detection algorithm.

[0085] Among them, the component adversarial loss L comp , calculated as follows:

[0086]

[0087] Where D r (·) represents the discriminator of local areas (such as eyes, mouth and other key areas), R r (·) represents the extraction operation of the local region, and r represents the label or index of each local region (region), including the left eye region (LeftEye), the right eye region (Right Eye), the mouth region (Mouth) and the nose region (Nose).

[0088] Among them, the global adversarial loss L adv , which is used to improve the overall realism of the restored image, is calculated as follows:

[0089]

[0090] Where D(·) represents the global discriminator, which is used to distinguish real images from generated images.

[0091] Among them, the identity loss L id , used to maintain facial identity consistency, is calculated as follows:

[0092]

[0093] Where η(·) represents the identity feature extractor.

[0094] Among them, the feature matching loss Lp , calculated as follows:

[0095]

[0096] Through the geometric consistency loss function, the restoration model can not only restore the facial details, but also maintain the consistency of the facial structure, ensuring that the restored image is more visually realistic and natural.

[0097] After introducing the geometric consistency loss function L geo In the context of image restoration, the purpose is to enhance the quality and consistency of the restored image by utilizing geometric information. By measuring the similarity between the restored image and the high-quality image sample, the restored result is geometrically consistent with the real or expected geometric features, ensuring that the restored image is more natural and realistic in shape and structure.

[0098] The inpainting model and decoder are trained with a large number of low-quality image samples and high-quality image samples, and the degradation model is used to simulate different types of degradation. The inpainting model is gradually optimized through training and evaluated on the validation set. Finally, the parameter structure of the inpainting model is adjusted to improve the inpainting effect and the ability to handle complex degraded images.

[0099] Step 6: Based on the trained restoration model, low-resolution fusion features and high-resolution fusion features are output for the image to be optimized, and a restored high-quality image is generated.

[0100] The trained restoration model can handle blind face restoration problems in various degradation scenarios. After outputting low-resolution fusion features and high-resolution fusion features for the optimized image, the low-resolution fusion features and high-resolution fusion features are fused at multiple scales to obtain the final fusion features. The final fusion features are input into the decoder to generate a high-quality restored image.

[0101] In the case of low-quality images with severe blur, noise or compression artifacts, the inpainting model can effectively restore high-quality, natural and high-quality images with complete facial structures. This makes the method of the present invention very suitable for image inpainting tasks in surveillance videos and face recognition systems with poor image quality, and can also be widely used in the fields of photo restoration and identity authentication.

[0102] The method of the present invention adopts a geometrically enhanced cross-attention mechanism, and optimizes the restoration process of degraded facial images by introducing geometric prior information and a high-quality feature dictionary. Geometric prior information provides a reference for facial structure, and a high-quality feature dictionary provides detail recovery. GeoMHCA integrates these two types of prior information to ensure the realism and detail integrity of facial images in complex degraded scenes. The geometric prior information provides guidance on shape and structure, while the high-quality feature dictionary provides details and realism. In the calculation process of GeoMHCA, the geometric prior information affects the distribution of attention weights, and the restoration model can better decide which areas to focus attention on, thereby obtaining a better restored image.

[0103] The present invention carries out the following experiment:

[0104] (1) Experimental equipment: This experiment is implemented using the PyTorch framework and trained on a server equipped with an NVIDIA RTX 3090 GPU (24GB of video memory). The system is equipped with an Intel Xeon E5-2620 CPU and 128GB of memory, and the operating system is Ubuntu 18.04.

[0105] (2) Experimental setup:

[0106] High-quality image sample set: Due to the memory limitation of the experimental equipment, the training set used in this experiment is CelebA-HQ, which is 30,000 high-definition face images with a resolution of 1024*1024 obtained from the face attribute dataset (CelebFaces Attribute, CelebA). It also contains face images CelebA-HQ with a variety of different face attributes. The images are resized to 512*512 during size training.

[0107] Low-quality image sample set: through the degradation model (generating low-quality images).

[0108]

[0109] In the formula, I d is a low-quality image sample, I h is a high-quality image sample, k σ is the Gaussian blur kernel, (σ is the standard deviation, σ∈[0.2,10]),↓ r is downsampling (r is the sampling factor, r∈[1,8]), n δ Gaussian noise (δ is intensity, δ∈[0,20]), JPEG compression: quality factor q (q∈[60,100]), ↑ r Upsample back to native resolution.

[0110] Test training set: The dataset used to test the qualitative and quantitative results of face restoration is Helen, which contains 2330 images containing multiple faces.

[0111] (3) Training settings: The Adam optimizer is used, the initial learning rate is set to 0.0001, and the weight parameters in the loss function are set to λ per =1.0,λ p =0.25,λ adv =0.8,λ comp =1.0,λ id =1.0,λ geo =0.5.

[0112] (4) This experiment uses the conventional multi-head cross attention mechanism and the geometric enhanced cross attention mechanism of the present invention to conduct experiments, and uses the original high-quality image samples as the control group. The experimental results are shown in Table 1:

[0113] Table 1 Experimental results

[0114]

[0115] Comparison of image quality indicators and parameters of different attention mechanisms in the Helen dataset: SSIM (Structural Similarity Index) ↑: Measures structural similarity. IDD (Identity Distance Index) ↓: Measures the similarity between the restored image and the real image in identity features. LPIPS (Perceptual Image Quality Index) ↓: Measures the perceived image difference. PSNR (Peak Signal-to-Noise Ratio) ↑: Measures the pixel-level restoration quality.

[0116] (5) Experimental conclusion:

[0117] The geometrically enhanced cross-attention mechanism is superior to the multi-head cross-attention mechanism in all evaluation indicators: SSIM shows that it has better structural similarity and multi-scale structure recovery effects. LPIPS is lower, which means that the perceptual quality of the generated image is closer to the real image. PSNR is slightly higher, indicating that the restoration quality at the pixel level is better. IDD is slightly lower, indicating that the repaired image is closer to the real image in terms of maintaining identity features. Experiments show that the geometrically enhanced cross-attention mechanism of the present invention has a stronger ability to restore the structure and perceptual quality of images, which is superior to the ordinary multi-head cross-attention mechanism, and provides stronger support for the reconstruction of image structural information and the maintenance of identity features.

[0118] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0119] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A blind face restoration method based on geometrically enhanced cross-attention mechanism, characterized in that: The blind face restoration method based on geometrically enhanced cross-attention mechanism includes: Acquire a training data set, wherein the training data set includes pairs of high-quality image samples and low-quality image samples; Extract high-quality features and geometric prior information from high-quality image samples, extract degradation features from low-quality image samples, and extract degradation features after upsampling low-quality image samples; Constructing a geometry-enhanced criss-cross attention mechanism as a restoration model, wherein the geometry-enhanced criss-cross attention mechanism is based on a multi-head criss-cross attention mechanism and takes degraded features as queries, high-quality features as keys and values, and geometric prior information as keys and values; The degradation features in the original low-quality image samples are used as queries, the high-quality features are used as keys and values, and the geometric prior information is used as keys and values, and the restoration model is used to output low-resolution fused features; The degradation features in the upsampled low-quality image samples are used as queries, the high-quality features are used as keys and values, and the geometric prior information is used as keys and values, and the restoration model is used to output high-resolution fused features; Generate a restored high-quality image based on low-resolution fusion features and high-resolution fusion features, and optimize and update the restoration model based on the loss calculated based on the restored high-quality image and high-quality image samples; Based on the trained restoration model, low-resolution fusion features and high-resolution fusion features are output for the image to be optimized, and a high-quality restored image is generated.

2. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that: The method for generating the low-quality image samples is: using a degradation model to process the high-quality image samples to obtain corresponding low-quality image samples.

3. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that: The high-quality features are used as keys and values, including: Use convolutional neural networks to extract high-quality features from high-quality image samples; The high-quality features are encoded and quantized through a vector quantization algorithm to construct a high-quality feature dictionary. The keys and values ​​in the high-quality feature dictionary are used as the keys and values ​​of the geometric enhanced cross-attention mechanism.

4. The blind face restoration method based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that: The method uses geometric prior information as keys and values, including: The geometric prior information includes key points obtained by processing high-quality image samples using a key point detection algorithm, facial geometric features extracted from a three-dimensional face model generated based on the high-quality image samples, and a facial region segmentation map obtained by processing high-quality image samples using a facial analysis model; The key points and facial geometric features are used as keys, and the facial region segmentation map is used as a value.

5. The method for blind face restoration based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that: The processing process of the geometrically enhanced cross-attention mechanism is as follows: In the formula, Z i Represents the output features of the i-th head in the multi-head cross attention mechanism, i = 1, 2, ..., N h -1, N h represents the number of heads of the multi-head cross attention mechanism, softmax represents the softMax function, Q i represents the query for the input i-th head, and represents the keys and values ​​from the high-quality features input to the i-th head, and represents the key and value from the geometric prior information input to the i-th head, C h Indicates the number of channels per head; The output features Z of all heads i The comprehensive output feature Z is obtained by connecting mh ; The final output is as follows: In the formula, Z fusion It represents the final output features of the geometrically enhanced cross-attention mechanism, that is, low-resolution fusion features or high-resolution fusion features. FFN represents the feedforward network, and LN represents the layer normalization operation. Indicates a degenerate feature.

6. The method for blind face restoration based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that: The method of generating a restored high-quality image based on the low-resolution fusion features and the high-resolution fusion features includes: The final fusion feature is obtained by multi-scale fusion of low-resolution fusion features and high-resolution fusion features; The final fused features are input into a decoder to generate a high-quality restored image, and the decoder is optimized and updated synchronously with the restoration model.

7. The method for blind face restoration based on geometrically enhanced cross-attention mechanism according to claim 1, characterized in that: The loss is calculated based on the restored high-quality image and the high-quality image sample, include: Where, L total represents the overall loss function, represents pixel-level loss, L per represents the perceptual loss, L geo represents the geometric consistency loss, L comp represents the component adversarial loss, L adv represents the global adversarial loss, L id represents identity loss, L p represents the feature matching loss, λ per represents the perceptual loss weight, λ geo represents the geometric consistency loss weight, λ comp represents the component adversarial loss weight, λ adv represents the weight of the global adversarial loss, λ id represents the identity loss weight, λ p Represents the feature matching loss weight.

Citation Information

Patent Citations

  • Real degraded image blind restoration method based on cross attention mechanism

    CN115829876A

  • Structure and texture feature guided double-encoder image restoration method

    CN116523985A

  • Progressive face image restoration method, system and device and storage medium

    CN117391995A

  • Face image restoration method, system, device and medium

    CN118333910A

  • Visual image enhancement generation method and system, device, and storage medium

    WO2022241995A1