Blind face restoration method based on image quality prior

By introducing image quality prior information into the blind face repair method, a codeword generation module with dual codebook architecture and quality prior conditions is constructed, which solves the problem of insufficient quality of portrait restoration results in the existing technology, and achieves efficient and high-quality blind face repair effects.

CN120013818AActive Publication Date: 2025-05-16TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510095520.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Existing blind face repair methods are difficult to achieve the highest quality portrait restoration, and due to the difficulty of obtaining high-quality training data, low-quality images are often mixed with, affecting the restoration results.

Method used

A blind face repair method based on image quality priors is adopted, and the input image quality evaluation model is used to evaluate the quality of the input image, and a general codebook and a high-quality codebook are constructed. Combined with the codeword generation module of quality priors, a high-quality feature codeword sequence is generated, and a repaired high-quality face image is generated through the decoder.

Benefits of technology

It significantly improves the effect of blind face repair, and achieves higher perceptual quality face picture restoration, and the generated repair results are clearer, richer details, and more natural and beautiful colors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013818A_ABST
    Figure CN120013818A_ABST
Patent Text Reader

Abstract

The invention discloses a blind face restoration method based on image quality prior, and solves the problem that the restoration result quality is not high enough in the portrait restoration field. By introducing image quality prior information, the restoration algorithm is endowed with an image quality perception capability, so that the image quality of a restoration result is remarkably improved in a reasoning stage. A double-codebook framework is designed, diversified and high-quality facial features are reserved respectively, and a high-quality feature code word sequence is generated in a codebook searching stage according to a quality score of an input image in combination with a code word generation module of a quality priori condition, preferably, a Transform network. Besides, the output space of the model is limited through discrete codebook prior, generation of adversarial samples is reduced, model parameters are directly optimized by maximizing a non-reference image quality assessment (NR-IQA) score, and the visual perception quality of the restored image is further improved. The method is superior to the prior art in the aspects of definition, facial details and color naturalness, can be applied to various restoration algorithms, and does not introduce extra inference overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to, in particular to, a blind face restoration method based on image quality prior. Background Art

[0002] Blind face restoration is the process of improving the quality of degraded facial images while preserving the identity of the face subject. Since the degradation of low-quality (LQ) facial images in practical applications is unknown and there are infinite possibilities for high-quality (HQ) outputs corresponding to given LQ inputs, blind face restoration becomes an ill-posed problem.

[0003] In order to alleviate the ill-posedness of blind face restoration, existing research has cleverly used various priors to generate rich HQ details and enhance the robustness of the algorithm to various degradations. Existing research can be divided into the following categories: methods based on geometric priors, methods based on reference priors, and methods based on generative priors. Among them, the method based on codebook reference priors has achieved good results and is also the basis of the technology of the present invention. The method based on codebook reference priors regards blind face restoration as a codeword prediction task in a discrete representation space, which usually involves two stages: a codebook prior learning stage and a codebook search stage. In the first stage, a context-rich codebook is trained using an HQ image through a vector quantization autoencoder. After obtaining the HQ codebook prior, in the second stage, various modules and strategies (such as Transformers) are used to find HQ codewords based on given LQ inputs. This discrete representation space significantly reduces the uncertainty of restoration, thereby enhancing the robustness of the model to various forms of degradation.

[0004] The quality of the restoration results of this type of method is determined by the HQ codebook, that is, by the HQ image in the codebook prior learning stage. However, due to the difficulty in obtaining large-scale high-quality data sets, some low-quality images are mixed in the commonly used HQ training data, which affects the quality of the restoration results. Therefore, it is difficult for this type of method to achieve the highest quality portrait restoration. If relatively low-quality images in the training data are manually selected and screened out, it will not only consume a lot of human resources, but also lead to a reduction in the amount of training data, thus affecting the robustness of the restoration model.

[0005] It should be noted that the information disclosed in the above background technology section is only used for understanding the background of the present application, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the invention

[0006] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a blind face restoration method based on image quality prior.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A blind face restoration method based on image quality prior includes the following steps:

[0009] S1, codebook learning stage: use the no-reference image quality assessment (NR-IQA) model to evaluate the quality of the input high-quality face image to obtain a quality score; according to the quality score, store the features of the high-quality face image in a general codebook and a high-quality codebook (HQ+ codebook), respectively, wherein the general codebook stores the features of all high-quality face images, and the HQ+ codebook stores the features of high-quality face images whose quality exceeds a preset threshold; the storage process includes feature extraction, codebook quantization and feature fusion to construct a general codebook containing rich and diverse facial features and an HQ+ codebook specifically capturing high-quality facial details;

[0010] S2, codebook search and image restoration stage: for the low-quality face image to be restored, use the quality score of the high-quality face image evaluated in step S1 to generate a high-quality feature codeword sequence that matches the quality of the low-quality face image through a codeword generation module; according to the predicted codeword sequence, retrieve the corresponding codeword from the general codebook and the HQ+ codebook, and perform feature fusion to obtain a comprehensive feature representation; input the comprehensive feature representation into a decoder to generate a restored high-quality face image; preferably, the codeword generation module is a Transformer network; wherein the codeword generation module is trained or configured using the general codebook and the HQ+ codebook constructed in step S1 to learn or establish a mapping relationship between the input image quality score and the high-quality features stored in the codebook, so that in the inference stage, a suitable high-quality feature codeword sequence can be generated according to the input quality score, thereby realizing blind face restoration based on image quality prior.

[0011] In some embodiments, the blind face restoration method based on image quality prior further includes the following steps:

[0012] S3, quality optimization stage, is used to improve the quality of the restored face image: the discrete codebook prior is used to limit the model output space to reduce the possibility of generating adversarial samples, where adversarial samples refer to those images that obtain high scores of the NR-IQA model but have poor actual visual quality; the restoration model parameters are directly fine-tuned by maximizing the NR-IQA model score to improve the visual perception quality of the generated image; wherein, the codebook prior is used to make the samples in the output space mainly reflect high-quality facial semantic information, thereby reducing the probability of outputting low visual quality images.

[0013] The present invention has the following beneficial effects:

[0014] The present invention proposes a blind face restoration method based on image quality prior, which significantly improves the effect of blind face restoration by introducing image quality prior information, and realizes face image restoration with higher perceptual quality. Specifically, the present invention designs a dual codebook architecture, which retains diversified and high-quality facial features respectively, and combines the codeword generation module of the quality prior condition, preferably the Transformer network, which can generate a high-quality feature codeword sequence according to the quality score of the input image in the inference stage, thereby outputting a clearer, more detailed, and more natural and beautiful restoration result. In addition, the present invention limits the model output space by discrete codebook prior, reduces the generation of adversarial samples, and directly optimizes the model parameters by maximizing the no-reference image quality assessment (NR-IQA) score, further improving the visual perception quality of the restored image. The score injection algorithm of the present invention has the characteristics of plug-and-play, can be applied to various existing restoration algorithms, and will not introduce additional reasoning overhead, and the model parameter quantity and reasoning speed are basically unaffected. In general, the present invention solves the problem of insufficient quality of restoration results in the field of portrait restoration by introducing image quality prior, and provides an efficient and high-quality blind face restoration method.

[0015] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a diagram of the overall architecture of the model of the blind face restoration method based on image quality prior according to an embodiment of the present invention.

[0017] Figure 2 This is a comparison chart of the experimental effects of the embodiments of the present invention and the prior art (the “original image” is a real-person image in which facial identity information cannot be seen, and the rest of the images are AI-generated images). DETAILED DESCRIPTION

[0018] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope and application of the present invention.

[0019] See also Figure 1 The embodiment of the present invention provides a blind face restoration method based on image quality prior, comprising the following steps:

[0020] Step S1, codebook learning stage: use the non-reference image quality assessment (NR-IQA) model to evaluate the quality of the input high-quality face image to obtain a quality score; according to the quality score, store the features of the high-quality face image in a general codebook and a high-quality codebook (HQ+ codebook), respectively, wherein the general codebook stores the features of all high-quality face images, and the HQ+ codebook stores the features of high-quality face images whose quality exceeds a preset threshold; the storage process includes feature extraction, codebook quantization and feature fusion to construct a general codebook containing rich and diverse facial features and an HQ+ codebook specifically capturing high-quality facial details;

[0021] In a preferred embodiment, in step S1, the dual codebook architecture in the codebook learning stage includes: a general codebook, used to learn and store diverse facial features from all high-quality face images; a high-quality codebook (HQ+ codebook), used to learn and store high-quality facial detail features from high-quality face images whose quality exceeds a preset threshold; wherein the general codebook and the HQ+ codebook perform quality scoring on the input image through a no-reference image quality assessment (NR-IQA) model, and store the image features in corresponding codebooks according to the scoring results.

[0022] Specifically, the feature fusion process in the codebook learning stage includes: extracting features from an input high-quality face image to obtain a feature representation; selectively quantizing the feature representation into a general codebook or a high-quality codebook according to the quality score of the input image; when the quality score of the input image exceeds a preset threshold, weightedly fusing the features of the general codebook and the high-quality codebook to generate a comprehensive feature representation; when the quality score of the input image does not exceed the preset threshold, only using the feature representation of the general codebook; and inputting the comprehensive feature representation into a decoder to generate a reconstructed facial image.

[0023] Specifically, in step S1, the training process of the codebook learning stage includes the optimization of the following loss functions: reconstruction loss, which is used to minimize the pixel-level difference between the input image and the reconstructed image; perceptual loss, which is used to minimize the difference between the input image and the reconstructed image in the feature space; adversarial loss, which is used to enhance the visual realism of the reconstructed image; codebook feature loss, which is used to make the features extracted by the encoder consistent with the feature representation in the codebook; by jointly optimizing the above loss functions, the model can simultaneously capture diverse facial features and high-quality facial details in the codebook learning stage.

[0024] Step S2, codebook search and image restoration stage: for the low-quality face image to be restored, use the quality score of the high-quality face image evaluated in step S1 to generate a high-quality feature codeword sequence that matches the quality of the low-quality face image through a codeword generation module; according to the predicted codeword sequence, retrieve the corresponding codeword from the general codebook and the HQ+ codebook, and perform feature fusion to obtain a comprehensive feature representation; input the comprehensive feature representation into the decoder to generate a restored high-quality face image; wherein the codeword generation module is trained or configured using the general codebook and the HQ+ codebook constructed in step S1 to learn or establish a mapping relationship between the input image quality score and the high-quality features stored in the codebook, so that in the inference stage, a suitable high-quality feature codeword sequence can be generated according to the input quality score, thereby realizing blind face restoration based on image quality prior.

[0025] In a preferred embodiment, the codeword generation module is a Transformer network. Specifically, the codebook search and quality prior condition Transformer network in the image restoration stage include: extracting the feature representation of the low-quality face image through the encoder; using the embedding module to map the quality score of the input image to a vector with the same dimension as the feature representation, and adding the vector to the low-quality feature representation to generate a quality condition enhanced feature representation; predicting a codeword sequence that matches the quality of the input image based on the quality condition enhanced feature representation through the Transformer network; according to the predicted codeword sequence, retrieving the corresponding codeword from the general codebook and the high-quality codebook, and performing weighted fusion to generate a comprehensive feature representation; inputting the comprehensive feature representation into the decoder to generate a restored high-quality face image.

[0026] In a further preferred embodiment, the training process of the Transformer network includes: freezing the parameters of the decoder and only training the encoder and Transformer network; using a feature loss function to minimize the difference between the low-quality feature representation and the feature representation stored in the codebook; using a cross entropy loss function to optimize the matching degree between the codeword sequence predicted by the Transformer network and the true codeword sequence; by jointly optimizing the above loss functions, the Transformer network can accurately predict the codeword sequence that matches the quality of the input image. The quality prior condition Transformer network controls the quality of the repaired image in the following way during the inference stage: inputting the highest quality score as a condition to guide the Transformer network to predict a codeword sequence that matches the highest quality image; generating a repair result similar to the highest quality image through the decoder to achieve high visual perception of the repaired image.

[0027] In a preferred embodiment, the blind face restoration method based on image quality prior of the present invention further comprises the following steps:

[0028] Step S3, quality optimization stage, is used to improve the quality of the restored face image: the discrete codebook prior is used to limit the model output space to reduce the possibility of generating adversarial samples, where adversarial samples refer to images that obtain high scores of the NR-IQA model but have poor actual visual quality; the restoration model parameters are directly fine-tuned by maximizing the NR-IQA model score to improve the visual perception quality of the generated image; wherein, the codebook prior is used to make the samples in the output space mainly reflect high-quality facial semantic information, thereby reducing the probability of outputting low-visual quality images, and achieving effective control and improvement of the output quality of the restoration model.

[0029] Specifically, in step S3, the quality optimization stage reduces the generation of adversarial samples in the following ways: using a discrete codebook prior to limit the model output space, limiting the output space to a limited set of high-quality facial semantic information, thereby reducing the number of adversarial samples; wherein the discrete codebook prior is obtained by learning on high-quality facial images so that the samples in the output space mainly reflect high-quality facial features; by limiting the size and distribution of the output space, the probability of optimizing the result to an adversarial sample is reduced, thereby improving the actual visual quality of the restored image.

[0030] Furthermore, the quality optimization stage also includes: directly maximizing the score of the no-reference image quality assessment (NR-IQA) model and fine-tuning the parameters of the restoration model; using a discrete codebook prior to ensure that the samples in the output space have high-quality facial semantic information, so that the process of maximizing the NR-IQA score can effectively improve the visual perception quality of the restored image; wherein the discrete codebook prior is combined with the NR-IQA score optimization to make the image generated by the restoration model consistent with the high-quality image in visual perception.

[0031] The present invention introduces image quality prior into the face restoration process, giving the restoration algorithm the ability to perceive image quality, thereby significantly improving the image quality of the restoration result in the inference stage without adjusting the training data set. Specifically, the present invention designs a blind face restoration algorithm framework based on image quality prior, introduces image quality prior into the codebook learning stage (stage one) and the codebook search stage (stage two) through different strategies, and achieves high-quality and high-fidelity blind face restoration (BFR). In the codebook learning stage, a dual codebook architecture is designed to retain diverse and high-quality facial features respectively; in the codebook search stage, a codeword generation module (preferably a Transformer network) with quality prior conditions is proposed using the dual codebook for codeword prediction, and the image quality score is used as a training target to guide the improvement of restoration quality. The present invention effectively solves the problem of insufficient quality of restoration results in the field of portrait restoration. The core of the present invention is to use the new prior information of image quality prior to guide the restoration algorithm. The embodiment proposes a three-stage implementation method around this idea. The present invention can generate restoration results that are clearer, richer in details, and more natural and beautiful in color than the prior art, such as Figure 2 As shown, the renderings of the present invention are significantly better than the latest representative technology (Difface) and the most influential technology in the field of face restoration (CodeFormer) in terms of subjective perception quality, for example, they perform better in clarity, facial details and natural structures (such as the mouth). This is due to the fact that the present invention introduces image quality priors when training the restoration model, so that the model can learn the features of high-quality images, thereby outputting higher quality results during reasoning. Specifically, the dual codebook structure improves the restoration quality by combining common facial features and high-quality facial features; the score injection algorithm enables the model to distinguish between image features of different qualities, and the highest quality restoration result can be output by inputting the highest score as a condition during reasoning; the quality optimization loss directly improves the quality of the restoration result by maximizing the quality score. Compared with the prior art, the present invention has the advantages of good effect and low inference overhead: the good effect is because the introduction of image quality priors guides the algorithm to output higher quality pictures; the low inference overhead is because the method proposed by the present invention requires very few parameters, which are basically negligible compared to the basic restoration model. In addition, the score injection algorithm of the present invention has the characteristics of plug-and-play and can be applied to various existing restoration algorithms. The concept of the present invention can also be extended to other fields. For example, in the field of dark light enhancement, the image quality prior can be replaced by the picture brightness prior to improve the illumination level of the picture.

[0032] The following further describes an algorithm example and experimental verification of a specific embodiment of the present invention.

[0033] In order to achieve high-quality and high-fidelity blind face restoration (BFR), the present invention proposes a blind face restoration method based on image quality prior and designs a blind face restoration algorithm framework based on image quality prior, which introduces image quality prior into the codebook learning stage (stage one) and the codebook search stage (stage two) through different strategies. The overall framework of the present invention is as follows Figure 1 As shown. Specifically, in the codebook learning stage, the present invention designs a dual codebook architecture to retain diverse and detailed facial features respectively. In the codebook search stage, using the dual codebook, a quality prior condition Transformer network is proposed as the preferred codeword prediction. At the same time, the image quality score is used as a training target to guide the improvement of restoration quality.

[0034] Dual codebook structure

[0035] The present invention considers that the perceived quality of input images is different, and it is obvious that those inputs with higher quality than average quality contain more valuable facial details. Therefore, the present invention proposes a dual codebook architecture, including a general codebook and a high-quality codebook (HQ+ codebook). Figure 1 As shown in (a), the universal codebook is learned from all high-quality (HQ) inputs, while the HQ+ codebook is specifically learned from a subset of images (HQ+ images) that exceed a certain quality threshold.

[0036] Specifically, the present invention first uses a no-reference image quality assessment (NR-IQA) model to evaluate the quality scores of all inputs. When the encoder extracts features from the input picture E to obtain E(x h ) and then use the universal codebook to h ) is quantified to obtain If the quality s of the current HQ input exceeds the threshold s thr , use HQ+ codebook to E(x h ) is quantified to obtain Z q Calculated by formula (1):

[0037]

[0038] Where α is a balancing weight. If the current picture quality does not reach the quality threshold, only the universal codebook is used. Then the decoder D receives Z q Takes as input and outputs the reconstructed facial image x rec .

[0039] The universal codebook is trained on all face images. The extensiveness of the data brings rich and diverse facial features, while the HQ+ codebook captures high-quality features specific to HQ+ images. and HQ+ features When D is fused with HQ+, it learns to reconstruct the HQ+ image. Therefore, in the reconstruction reasoning process, h Matching HQ+ features and combining them with universal features, D will generate HQ+ image x res , even if x h is a normal image.

[0040] It is worth noting that the dual codebook architecture of the present invention is different from using the NR-IQA model for data screening and training only on HQ+ images. Since very high-quality images are in the minority, their facial information is not rich and diverse enough, and relying solely on these images for training may lead to facial defects in the output (e.g., distorted teeth). The universal codebook of the present invention contains rich and diverse facial features, ensuring the robustness of the model to all types of inputs.

[0041] In the codebook learning stage, the training loss function includes the reconstruction loss L1 loss function Perceived loss Fighting Losses and codebook feature loss The calculation formulas of each loss function are as shown in formula (2).

[0042]

[0043]

[0044] Quality Priors Transformer

[0045] In the codebook search phase, the quality prior is introduced into the conditional Transformer as a condition for codebook prediction. The conditional Transformer model is used to predict two codeword sequences under the condition of image quality s, where s is the corresponding picture quality of the HQ image, obtained by the first stage NR-IQA. The conditional Transformer tries to find the relationship between the image and its corresponding quality. For example, during training, when the true label of the training data is the highest quality image, the input score condition is also the highest score, that is, the model models the relationship between the highest score condition and the highest quality image. Therefore, after training, when the highest score is input as a condition, the quality of the restored result will be as close as possible to the highest quality image.

[0046] First, through Z l =E(x l ) to obtain low quality (LQ) features, and then use the Embedding module to map the quality score s to a vector s∈R h*w*c In the l The same dimensions. s is directly related to Z by the following formula lAdd together, as shown in formula (3):

[0047]

[0048] take over As input, TransformerT predicts two codeword sequences c1 and c2. Then, c1 retrieves the corresponding codeword from the universal codebook in a nearest match manner to form a quantized feature c2 retrieves the codeword from the HQ+ codebook to form The following formula and Fusion to get Z f :

[0049]

[0050] Where α is a balancing weight, which is consistent with formula (1). Then Input to the decoder to generate the restored image x res Since D has learned from and The HQ+ image is reconstructed by the fusion of f When is input, the second stage produces images of similar quality to HQ+ images.

[0051] By using the quality prior as a condition, the quality of the restored image can be controlled during inference. Usually the highest quality is desired, so the highest quality score is input as a condition and the transformer will generate a code sequence that produces the highest quality image after decoding.

[0052] In the codeword prediction stage, the decoder D is frozen, so the modules that need to be trained are the encoder E and the Transformer, and the loss functions involved are feature loss and the cross entropy loss function

[0053]

[0054] Quality Optimization

[0055] Using IQA (image quality assessment) as a training objective for image processing systems is a promising but under-researched area, especially for no-reference image quality assessment (NR-IQA) methods. Some work has explored the use of full-reference image quality assessment (FR-IQA) as an objective. However, unlike FR-IQA methods, due to the lack of reference, the optimization direction of the NR-IQA objective is usually unstable, that is, any image in the image space that satisfies a high NR-IQA score may be the result of optimization. This leads to a significant disadvantage when using NR-IQA methods as an objective: the images generated by the restoration model have poor human-perceived quality, but can fool the NR-IQA model into giving a high IQA score. These images are called "adversarial samples".

[0056] In the process of optimizing the restoration model using the NR-IQA model, the model parameters are trained to maximize the IQA score of the generated image, which can easily lead to the restoration model generating "adversarial samples" of the NR-IQA model. This phenomenon is also observed in the reward optimization of large language models (LLMs) and is called "over-optimization". The root cause of over-optimization is that the IQA model is only a proxy for human preferences and is not completely consistent with human preferences. Therefore, the optimization process does not always improve the perceived quality from a human perspective.

[0057] The existence of adversarial samples makes the NR-IQA target unreliable. Therefore, the fundamental solution lies in reducing the number of adversarial samples in the output space of the restoration model. If it is ensured that there are no adversarial samples in the output space, the NR-IQA score can be directly trusted and maximized. To this end, there are two strategies: 1. Improve the consistency of the NR-IQA score with human preferences. If the score of the NR-IQA model is completely consistent with human subjective perception, there will be no situation of "high NR-IQA score but poor human subjective perception of quality". 2. Limit the size and distribution of the output space. By limiting the size of the output space and the proportion of adversarial samples through additional constraints, the probability of optimizing the results to adversarial samples is reduced. The first strategy requires a lot of manpower and material resources to improve the capabilities of NR-IQA, so it is preferred to use the second strategy to solve the over-optimization problem.

[0058] The present invention further proposes a discrete codebook prior, which limits the output space to a finite space compared to a continuous prior, which means that the number of adversarial samples is also significantly reduced. At the same time, the codebook prior learned on high-quality facial images ensures that most samples in this finite space reflect high-quality facial semantic information, thereby limiting the proportion of adversarial samples in the finite space.

[0059] Therefore, taking the NR-IQA measure as a target is very suitable for the discrete codebook-based network structure of the present invention. Based on this, directly maximizing the NR-IQA score to fine-tune the recovery model parameters further improves the restoration quality.

[0060] Experimental Results

[0061] The present invention can achieve a face image restoration effect with higher perceptual quality. The restoration result is clearer than the prior art, the facial details are richer, and the facial colors are more realistic, natural and beautiful. Figure 2 The experimental results of our invention (Ours) are compared with the latest representative technology (Difface) and the most influential technology in the field of face restoration (CodeFormer). Figure 2 The “original image” in the figure is a real-life image in which facial identity information cannot be seen, and the rest of the images are AI-generated images). The renderings of the present invention have better subjective perceived quality. For example, the results of the present invention are obviously clearer than those of Difface and Codeformer, with more facial details and more natural facial structures (such as the mouth). Because the present invention introduces image quality priors when training the restoration model, the restoration model is endowed with the ability to perceive image quality, that is, the model can learn the features that high-quality images should have, thereby outputting higher quality results during reasoning. Specifically, the dual codebook structure of the present invention improves the restoration quality by combining general facial features with high-quality facial features; the score injection algorithm enables the model to distinguish between image features of different qualities. When the highest score is input as a condition during reasoning, the model will output the highest quality restoration result; the quality optimization loss directly improves the quality of the restoration result by maximizing the quality score.

[0062] The important features of the present invention are:

[0063] 1. Dual codebook autoencoder: The present invention constructs a dual codebook autoencoder by constructing a dual codebook structure and different training data streams. Traditional codebook-based algorithms often have only one codebook, so they can only learn the average quality of the training data. The present invention uses a new codebook to learn the image quality features above the average line in the training data, so as to introduce these high-quality features during inference to achieve higher quality output. For the dual codebook structure, the fusion of general features and high-quality features can be achieved through pixel addition, but it can also be fused through other feature fusion methods, such as the attention mechanism.

[0064] 2. Quality score injection algorithm: The quality score injection algorithm introduced in the present invention can enable the restoration model to have the ability to perceive image quality, so as to input the highest quality score during inference to obtain the highest quality restoration output.

[0065] 3. Quality optimization based on discrete space: The present invention also proposes quality optimization based on discrete space, which can alleviate the over-optimization problem of traditional algorithms and truly improve the image quality output by the restoration algorithm.

[0066] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.

[0067] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.

[0068] An embodiment of the present invention further provides a processor, wherein the processor executes a computer program and at least executes the method described above.

[0069] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0070] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0071] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0072] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0073] Those skilled in the art can understand that: all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.

[0074] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0075] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0076] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0077] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0078] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art of the present invention, several equivalent substitutions or obvious variations can be made without departing from the concept of the present invention, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A blind face restoration method based on image quality prior, characterized in that: The following steps are involved: S1, codebook learning stage: use the no-reference image quality assessment (NR-IQA) model to evaluate the quality of the input high-quality face image to obtain a quality score; according to the quality score, store the features of the high-quality face image in a general codebook and a high-quality codebook (HQ+ codebook), respectively, wherein the general codebook stores the features of all high-quality face images, and the HQ+ codebook stores the features of high-quality face images whose quality exceeds a preset threshold; the storage process includes feature extraction, codebook quantization and feature fusion to construct a general codebook containing rich and diverse facial features and an HQ+ codebook specifically capturing high-quality facial details; S2, codebook search and image restoration stage: for the low-quality face image to be restored, use the quality score of the high-quality face image evaluated in step S1 to generate a high-quality feature codeword sequence that matches the quality of the low-quality face image through a codeword generation module; according to the predicted codeword sequence, retrieve the corresponding codeword from the general codebook and the HQ+ codebook, and perform feature fusion to obtain a comprehensive feature representation; input the comprehensive feature representation into a decoder to generate a restored high-quality face image; preferably, the codeword generation module is a Transformer network; wherein the codeword generation module is trained or configured using the general codebook and the HQ+ codebook constructed in step S1 to learn or establish a mapping relationship between the input image quality score and the high-quality features stored in the codebook, so that in the inference stage, a suitable high-quality feature codeword sequence can be generated according to the input quality score, thereby realizing blind face restoration based on image quality prior.

2. The blind face restoration method based on image quality prior according to claim 1, characterized in that: In step S1, the dual codebook architecture in the codebook learning phase includes: A universal codebook for learning and storing diverse facial features from all high-quality face images; High-quality codebook (HQ+ codebook), used to learn and store high-quality facial detail features from high-quality face images whose quality exceeds a preset threshold; The general codebook and the HQ+ codebook perform quality scoring on the input image through a non-reference image quality assessment (NR-IQA) model, and store the image features in corresponding codebooks according to the scoring results.

3. The blind face restoration method based on image quality prior as claimed in claim 2, characterized in that: In step S1, the feature fusion process in the codebook learning stage includes: Extract features from the input high-quality face image to obtain feature representation; Selectively quantize the feature representation into a general codebook or a high-quality codebook according to the quality score of the input image; When the quality score of the input image exceeds a preset threshold, the features of the general codebook and the high-quality codebook are weighted and fused to generate a comprehensive feature representation; When the quality score of the input image does not exceed the preset threshold, only the feature representation of the general codebook is used; The comprehensive feature representation is input into a decoder to generate a reconstructed facial image.

4. The blind face restoration method based on image quality prior according to any one of claims 1 to 3, characterized in that: In step S1, the training process of the codebook learning phase includes the optimization of the following loss function: The reconstruction loss is used to minimize the pixel-level difference between the input image and the reconstructed image; Perceptual loss, which is used to minimize the difference between the input image and the reconstructed image in the feature space; Adversarial loss, used to improve the visual realism of reconstructed images; Codebook feature loss, used to make the features extracted by the encoder consistent with the feature representation in the codebook; By jointly optimizing the above loss functions, the model can simultaneously capture diverse facial features and high-quality facial details in the codebook learning stage.

5. The blind face restoration method based on image quality prior according to any one of claims 1 to 4, characterized in that: In step S2, the quality prior condition Transformer network in the codebook search and image restoration stage includes: Extract feature representation of low-quality face images through an encoder; The quality score of the input image is mapped into a vector with the same dimension as the feature representation using an embedding module, and the vector is added to the low-quality feature representation to generate a quality-conditioned enhanced feature representation. Through the Transformer network, the feature representation is enhanced based on the quality condition, and the codeword sequence matching the input image quality is predicted; According to the predicted codeword sequence, the corresponding codewords are retrieved from the general codebook and the high-quality codebook, and weighted fusion is performed to generate a comprehensive feature representation; The comprehensive feature representation is input into a decoder to generate a restored high-quality face image.

6. The blind face restoration method based on image quality prior according to claim 5, characterized in that: In step S2, the training process of the Transformer network includes: Freeze the decoder parameters and only train the encoder and Transformer networks; Use a feature loss function to minimize the difference between the low-quality feature representation and the feature representation stored in the codebook; Use the cross entropy loss function to optimize the match between the codeword sequence predicted by the Transformer network and the actual codeword sequence; By jointly optimizing the above loss functions, the Transformer network can accurately predict the codeword sequence that matches the quality of the input image.

7. The blind face restoration method based on image quality prior according to any one of claims 1 to 6, characterized in that: In step S2, the quality prior condition Transformer network controls the quality of the repaired image in the following ways during the inference phase: Input the highest quality score as a condition to guide the Transformer network to predict the codeword sequence that matches the highest quality image; The decoder generates a restoration result similar to the highest quality image, achieving high visual perception quality of the restored image.

8. The blind face restoration method based on image quality prior according to any one of claims 1 to 7, characterized in that: The following steps are also included: S3, quality optimization stage, is used to improve the quality of the restored face image: the discrete codebook prior is used to limit the model output space to reduce the possibility of generating adversarial samples, where adversarial samples refer to those images that obtain high scores of the NR-IQA model but have poor actual visual quality; the restoration model parameters are directly fine-tuned by maximizing the NR-IQA model score to improve the visual perception quality of the generated image; wherein, the codebook prior is used to make the samples in the output space mainly reflect high-quality facial semantic information, reducing the probability of outputting low visual quality images.

9. The blind face restoration method based on image quality prior according to claim 8, characterized in that: In step S3, the quality optimization stage reduces the generation of adversarial samples by: The discrete codebook prior is used to restrict the model output space, which is limited to a limited set of high-quality facial semantic information, thereby reducing the number of adversarial samples. The discrete codebook prior is obtained by learning on high-quality facial images, so that the samples in the output space mainly reflect high-quality facial features; By limiting the size and distribution of the output space, the probability of optimizing the result to an adversarial sample is reduced, thereby improving the actual visual quality of the restored image.

10. The blind face restoration method based on image quality prior according to claim 9, characterized in that: The quality optimization stage also includes: Directly maximize the score of the No-Reference Image Quality Assessment (NR-IQA) model and fine-tune the restoration model parameters; The discrete codebook prior is used to make the samples in the output space have high-quality facial semantic information, so that the process of maximizing the NR-IQA score improves the visual perception quality of the restored image. The discrete codebook prior is combined with the NR-IQA score optimization to make the image generated by the restoration model consistent with the high-quality image in visual perception.

Citation Information

Patent Citations

  • Blind face restoration method and system

    CN112598604A

  • Face image restoration method based on semantic analysis generation guidance

    CN113888417A

  • Syntax for image / video compression with generic codebook-based representation

    WO2024226920A1