A face super-resolution method, system and computer device based on visual language prior
By constructing a visual-language prior fusion block and combining visual-language multivariate representations, the problem of ignoring language text information in existing methods is solved, and higher quality face image reconstruction and identity information preservation are achieved.
Patent Information
- Application Number
- CN202411590182.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-11-08
AI Technical Summary
Existing face super-resolution methods mainly focus on visual perception, ignoring language and text information, resulting in incomplete scene representation and affecting the face image reconstruction effect.
By constructing a visual-language prior fusion block, combining visual-language multi-representation, and fusing semantic segmentation maps, depth maps, text titles, and descriptive information, the quality of face images is improved by leveraging the complementary advantages of visual-language priors.
It significantly improves the quality and integrity of facial images, enhances objective evaluation metrics and subjective visual quality, preserves facial identity information, and improves the performance of downstream tasks.
Smart Images

Figure CN119540059B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of computer images, and particularly relates to a face super-resolution method based on visual language prior, a face super-resolution system and a computer device. BACKGROUND
[0002] Face super-resolution technology aims to restore a high-resolution face image from a low-resolution face image. In a low-cost camera and a non-ideal imaging environment, a captured face image is often of poor quality, which not only affects the visual experience, but also may adversely affect downstream tasks such as face recognition and attribute analysis. Face super-resolution technology improves the performance of these downstream tasks by improving the quality of face images, and therefore its importance has become increasingly prominent in the past few decades.
[0003] In recent years, deep learning technology has developed rapidly, and face super-resolution methods based on deep learning have made significant progress. These methods often design effective network structures to directly learn the mapping from a low-resolution face image to a high-resolution face image. However, due to the difference in the dimensions of high and low resolution face images, a low-resolution face image may correspond to multiple different high-resolution faces, which increases the difficulty of the face super-resolution task. Therefore, strategies that introduce additional prior knowledge to optimize the solution space have been proposed. For example, some studies use face visual priors such as face analysis maps and heat maps to capture semantic structural information of faces to improve the reconstruction effect.
[0004] Due to the limitations of the imaging environment, the captured face image is usually of low resolution, and the low-resolution face image cannot provide a good visual experience and affects the performance of downstream tasks. Existing face super-resolution methods mainly design efficient network structures to learn the mapping from a low-resolution face image to a high-resolution face image. However, due to the large difference in the spatial dimensions of high and low resolution face images, it is difficult to directly learn the mapping between them. To this end, some methods introduce face prior knowledge (such as face analysis maps and heat maps) to constrain the solution space and improve face image quality. However, these methods still have obvious defects: first, only a single visual prior is used, mainly focusing on visual perception, and ignoring non-visual language text information, resulting in incomplete scene representation and interfering with face image reconstruction performance. SUMMARY
[0005] The application is directed to the problems of the prior art, and proposes a face super-resolution method based on visual language prior, a face super-resolution system and a computer device. The application is implemented through the following technical solutions:
[0006] A face super-resolution method based on visual language prior, comprising the following steps:
[0007] Step one, input the low-resolution face image into the pre-trained visual-linguistic large model to extract visual-linguistic multi-modal representation;
[0008] Step two, construct a visual-linguistic prior assisted face super-resolution network to fuse visual-linguistic prior information;
[0009] Step three, input the low-resolution face image and the visual-linguistic multi-modal representation extracted in step one into the network in step two to obtain the super-resolution result and get the restored high-quality face image.
[0010] Further, in step one, the specific steps include:
[0011] Input the low-resolution face image into the base large model SAM and DAM to extract the face semantic segmentation map and depth map as visual prior information.
[0012] Further, in step one, the specific steps include:
[0013] Input the low-resolution image into the BLIP2 model to generate the corresponding text title; for generating description, combine the low-resolution face image and a specific prompt "provide useful visible features or observations about the face image" and input it into the ChatGPT-4 model to generate detailed text description of the face image.
[0014] Further, in step two, input the low-resolution face image and the visual-linguistic multi-modal representation into the feature extraction layer, CLIP encoder and image encoder respectively to extract the visual feature F0 and the visual-linguistic multi-modal representation feature E C ,E D ,F S ,F D ; the extracted features are sent to a series of visual-linguistic prior fusion blocks and base blocks;
[0015]
[0016] wherein and represent the i-th visual-linguistic prior fusion block and the base block, respectively, F i is the feature combined with the visual-linguistic prior information; the feature F L processed by L layers is input into a feature reconstructor implemented by a convolutional layer to generate the final super-resolution image I SR ; in the training process, introduce L1 loss as a constraint,
[0017]
[0018] wherein, I HR is the corresponding high-resolution face image.
[0019] Further, in step two, the visual language fusion module is specifically:
[0020] Given the low-resolution facial features and visual language prior E C ,E D ,F S ,F D The visual language prior fusion block sends the low-resolution facial features into four parallel attention mechanisms, namely SegA semantic segmentation attention, DepA depth attention, CapA caption attention, and DesA description attention, to achieve effective interaction with the four visual language priors; the depth map and the semantic segmentation map are processed by SegA and DepA, and through a series of cascaded convolution layers, the spatial attention for face structure perception is learned to enhance the facial features. CapA first applies global average pooling to the features of the title, then performs convolution processing and sigmoid activation to generate global caption attention. Compared with the title, the description of the image provides a more detailed and comprehensive perspective for the face image. Therefore, the image description is used to construct the query Q, while the facial features are used to form the key K and the value V, and their fusion is realized through the cross-attention mechanism, namely DesA. After integrating the four kinds of prior information, the four groups of result features are merged with the facial features, and a skip connection is introduced to generate the final output.
[0021] The present application also relates to a visual language prior-based face super-resolution method system, comprising a computer module, which applies the above-mentioned visual language prior-based face super-resolution method.
[0022] The present application also relates to a computer device comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the above-mentioned method when executing the computer program.
[0023] The present application also relates to a computer-readable storage medium, which stores a computer program, and wherein the computer program is executed by a processor to implement the steps of the above-mentioned method.
[0024] The present application also relates to a computer-readable storage medium, which stores a computer program, and wherein the computer program is executed by a processor to implement the steps of the above-mentioned method.
[0025] Advantages
[0026] The present application proposes a face super-resolution algorithm based on visual language priori, which not only utilizes multiple visual priors, but also integrates higher-level language priors, fully utilizes the complementary advantages of visual language priors, and more completely represents the face image, thereby effectively improving the face image quality. Through comparison of the face super-resolution performance of the present model and other methods, the model designed in the present application has the best performance. Compared with the existing mainstream face super-resolution methods (FSRNet, DIC, SISN, SFMNet, FaceFormer and WFEN), the face image recovered by the present application performs better in objective evaluation indicators and subjective visual quality. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the specific embodiments or related art, the drawings needed to be used in the specific embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0028] Figure 1 Face super-resolution model in the present application.
[0029] Figure 2 Visual language priori fusion block in the present application.
[0030] Figure 3 8 times face super-resolution result comparison chart in the present application.
[0031] Figure 4 16 times face super-resolution result comparison chart in the present application.
[0032] Figure 5 Identity distance comparison result chart of the recovered face image and the high-resolution face image in the present application.
[0033] Figure 6 Visual quality comparison chart of different models in the present application. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0035] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0036] The accompanying drawings are incorporated in and constitute a part of this specification and will be understood by those skilled in the art. Figures 1 to 6 Further implementation details are described in more detail.
[0037] The present application utilizes language-visual priors to improve the restoration quality of face images. In this section, the specific details and design concepts of the proposed framework will be elaborated.
[0038] For a given low-resolution face image I LR , the present application designs an effective reconstruction method to restore the corresponding high-quality face image from it. The reconstruction method of the present application includes the following steps:
[0039] Step one, send the low-resolution face image into the pre-trained visual-linguistic large model to extract the visual-linguistic multi-modal representation
[0040] The workflow of the proposed framework is shown in Figure 1 After inputting the low-resolution face image I LR , first send it into the visual-linguistic large model to extract the visual-linguistic multi-modal representation, the specific process is as follows:
[0041] Visual priors: semantic segmentation map and depth map. The face semantic segmentation map can accurately capture the high-level semantic details of the face structure, while the depth map provides accurate depth information of the face. Directly input the low-resolution face image into these pre-trained models to extract the face semantic segmentation map and depth map as visual prior information.
[0042] Language priors: text title and description. The text title provides a macroscopic perspective for the face image, summarizing its overall and abstract semantic content, while the text description further deepens, providing more detailed and comprehensive feature description of the face. Unlike the structural and depth information depicted in the semantic mask and depth map in the visual prior, the title and description provide higher-level semantic information, and the description contains comprehensive and detailed text content. The present application inputs the low-resolution image into the BLIP2 (image-text pre-trained model) model to generate the corresponding text title. At the same time, in order to generate the description, combine the low-resolution face image with a specific prompt, which requires the model to provide useful visible features or observations about the face image, and then input these into the ChatGPT-4 model to produce a detailed text description of the face image.
[0043] Step two, construct a visual-linguistic prior assisted face super-resolution network to fuse visual-linguistic prior information
[0044] Considering that the visual-linguistic multi-modal representation contains different information, it also has different effects on face super-resolution. Specifically, the text title summarizes the overall information of the face, while the description carefully depicts the specific content and features of the face. The semantic segmentation map contains the structure information of the face, and the depth map provides the depth information of the face. These prior information complements each other, and their comprehensive application can achieve a more comprehensive face image representation, thereby effectively promoting the face image reconstruction. Therefore, the framework proposed in the application integrates these prior information to enhance the performance of face super-resolution.
[0045] Specifically, the low-resolution face image and the visual-linguistic multi-modal representation are respectively extracted in the feature extraction layer, the CLIP (visual-linguistic contrast pre-training large model) encoder and the image encoder to obtain the visual features F0 and the visual-linguistic multi-modal representation features E C ,E D ,F S ,F D . Subsequently, these extracted features are sent into a series of visual-linguistic prior fusion blocks and basic blocks (a total of L groups) to fully exploit the potential of visual-linguistic prior.
[0046]
[0047] wherein, and respectively represent the function of the i-th visual-linguistic prior fusion block and the basic block, and F i is the feature combined with the visual-linguistic prior information. Finally, the feature F L processed by L layers is input into a feature reconstructor implemented by a convolutional layer to generate the final super-resolution image I SR . In order to ensure that the method can restore satisfactory results, L1 loss is introduced as a constraint in the training process,
[0048]
[0049] wherein, I HR is the corresponding high-resolution face image.
[0050] The visual-linguistic fusion module is specifically:
[0051] The existing face super-resolution technology mostly focuses on the field of visual perception, and improves the image quality by integrating visual priori, but often ignores the importance of language text features, which limits the integrity of image representation and the performance of face super-resolution. The present application constructs a comprehensive visual and language priori system to more richly depict the features of face images, thereby significantly improving the performance of face super-resolution. In order to effectively combine visual and language elements and fully exert the potential and complementarity of visual and language priori, the present application designs a visual and language priori fusion block (VLPFB), as shown in Figure 2 .
[0052] Given the low-resolution face features and the visual and language priori E C ,E D ,F S ,F D The goal of the visual and language priori fusion block (VLPFB) is to integrate these elements and use the complementary information between them to improve the effect of face super-resolution. The initial step of VLPFB is to send the features of the low-resolution face into four parallel and distinctive attention mechanisms, named SegA (semantic segmentation attention), DepA (depth attention), CapA (caption attention) and DesA (description attention), to achieve effective interaction with the four kinds of visual and language priori. Specifically, the depth map and the semantic segmentation map, as key elements reflecting the structure of the face and the spatial pixel information of the face, are processed by SegA and DepA, which learn the spatial attention of face structure perception through a series of cascaded convolution layers to enhance the face features. The caption, as information providing a global overview of the face image, is processed by CapA. CapA first applies global average pooling to the features of the caption, then performs convolution processing and sigmoid activation to generate global caption attention. Compared with the caption, the description of the image provides a more detailed and comprehensive perspective of the face image. Therefore, the description of the image is used to construct the query Q, while the face features are used to form the key K and the value V, and their fusion is achieved through the cross-attention mechanism, i.e. DesA. After integrating these four kinds of priori information, the four groups of result features are merged with the face features, and a skip connection is introduced to generate the final output. In the design of VLPFB, the present application adopts a targeted fusion strategy according to the characteristics of various priori information, realizes sensitive fusion of priori information, and effectively promotes the performance improvement of face super-resolution.
[0053] Step three, send the low-resolution face image and the visual and language multi-element representation extracted in step 1 into the network of step 2 to obtain the super-resolution result and restore the high-quality face image.
[0054] Effect verification
[0055] 1. Comparison of objective evaluation indicators
[0056] The comparison results of objective evaluation indicators of the present application and existing methods are shown in Table 1. In the ×16 face super-resolution task, the PSNR of the present application is 25.95 dB, which is 0.43 dB higher than the second best method WFEN. In the ×8 face super-resolution task, the present application also achieves the best objective evaluation indicators. FSRNet and DIC aim to estimate the visual prior of the face and integrate these priors into the reconstruction of the face image, but their performance improvement is limited. SISN captures face information by introducing an interleaved attention mechanism, and its performance is better than DIC and FSRNet. FaceFormer combines transformer and convolutional neural network, and its performance is better than SISN and DIC. WFEN and SFMNet use transformer and Fourier transform to capture global receptive field, which further improves the performance of SISN and FaceFormer. Although these methods perform well to some extent, they mainly focus on single visual prior and pixel-level image representation, and do not fully explore more comprehensive visual language prior. The present application integrates visual language prior into the face super-resolution task, which can more completely represent the face image, thereby effectively improving the quality of the face image.
[0057] Table 1 Comparison of objective evaluation indicators of existing face super-resolution methods
[0058]
[0059] 2. Comparison of subjective visual quality Figure 3 and Figure 4 The comparison results of visual quality of different methods in the ×8 and ×16 face super-resolution tasks are shown. SRCNN cannot reconstruct key face components. While FSRNet, DIC and SISN can restore clear face contours, they perform poorly in reconstructing key face details such as eyes and mouth in the ×8 face super-resolution task. In the more challenging ×16 face super-resolution task, the performance of these methods further deteriorates, and the results of the restoration are not satisfactory. For the ×16 face super-resolution task, FaceFormer, SFMNet and WFEN often cannot accurately restore the key parts of the face, and there is distortion. Although in some specific cases (such as Figure 5 the first sample shown), they can restore some facial features, but the generated images are often too smooth and lack the necessary details. In contrast to these methods, the face super-resolution method designed by the present application restores face images by utilizing visual language prior, providing a more comprehensive image representation, thereby being able to restore visually more pleasing, detailed and natural face images.
[0060] 3. Face identity distance comparison: In addition to improving the quality of low-resolution face images, face super-resolution methods should also preserve face identity information. That is, super-resolution face images should share the same identity information as their corresponding high-resolution face images. Therefore, further identity distance comparison was conducted. The face images recovered by different super-resolution methods and the corresponding high-resolution face images were input into the pre-trained face recognition model DeepFace to extract face identity features. Then, the cosine distance between these features was calculated as a measure of identity distance. The comparison results are shown in Figure 5 The identity distance of the present application is lower than that of the comparison method, which indicates that the present application can better preserve the identity information of face images, thereby helping to improve the accuracy and efficiency of face recognition tasks.
[0061] 4. Effectiveness analysis of VLPFB: When analyzing the effectiveness of the visual language prior fusion block (VLPFB), we adopted a step-by-step approach to evaluate its contribution. First, we created a base model Model 1 that does not include VLPFB and any visual language prior information. Then we introduced Model 2, which simply concatenates and convolves to fuse visual language prior information. Model 2 has improved performance compared to Model 1, which verifies the positive effect of visual language prior information on face reconstruction, although this improvement is not significant. To further verify the advantage of VLPFB, we replaced the concatenation in Model 2 with VLPFB, forming the model of the present application. As shown in Table 2, the present application achieved the best performance, with a significant improvement, fully demonstrating the effectiveness of VLPFB. Figure 6 The visual quality comparison between different models shows that the present application can retain more details when recovering face images, especially in the clarity and accuracy of key face features such as eyes, with significant improvement compared to Model 1 and Model 2. In summary, VLPFB not only effectively fuses visual language prior information, but also significantly improves quantitative evaluation indicators and visual quality of face images, indicating its importance and practicality in face super-resolution tasks.
[0062] Table 2 Effectiveness analysis of VLPFB
[0063]
[0064] The face super-resolution of the present application is dedicated to recovering a high-resolution face image from a given low-resolution face image. In dealing with this ill-posed and challenging problem, the prior art is either dedicated to developing an efficient network structure to learn the mapping from a low-resolution face to a high-resolution face, or to constrain the solution space by introducing prior knowledge. Although the existing methods have made significant progress, they mainly focus on visual perception, ignoring the language text features, resulting in incomplete scene representation, which in turn affects the face image reconstruction effect. With the rise of large models, they have shown excellent ability in content generation. Unlike the visual prior used in existing methods, language (text) knowledge describes a deeper understanding and higher-level abstraction, which can be regarded as a language prior, complementing the shortcomings of visual perception.
Claims
1. A face super-resolution method based on visual language prior, characterized in that, Comprising the following steps: Step one, input the low-resolution face image into the pre-trained visual-linguistic large model, and extract the visual-linguistic multi-modal representation; Step two, construct a visual-linguistic prior auxiliary face super-resolution network, and fuse visual-linguistic prior information; Step three, input the low-resolution face image and the visual-linguistic multi-modal representation extracted in step one into the network in step two to obtain the super-resolution result and get the restored high-quality face image; In step two, the low-resolution face image and the visual language multi-modal representation are respectively sent into the feature extraction layer, the CLIP encoder and the image encoder to extract features, to obtain visual features F0 and visual language multi-modal representation features E C ,E D ,F S ,F D ; the extracted features are sent into a series of visual language prior fusion blocks and basic blocks; wherein and respectively represent the i-th visual-linguistic prior fusion block and the base block, F i is the feature combined with visual-linguistic prior information; the feature F L processed by L layers is input into a feature reconstructor implemented by a convolutional layer to generate a final super-resolution image I SR ; an L1 loss is introduced as a constraint in the training process, wherein I HR is the corresponding high-resolution face image; In step two, the visual-linguistic prior fusion block is specifically: Given low-resolution face features and visual language priors E C ,E D ,F S ,F D The visual language prior fusion block sends the low-resolution face features into four parallel attention mechanisms, namely SegA semantic segmentation attention, DepA depth attention, CapA caption attention, and DesA description attention, to achieve effective interaction with the four visual language priors; the depth map and the semantic segmentation map are processed by SegA and DepA, and the spatial attention for face structure perception is learned through a cascade of convolutional layers to enhance the face features; CapA first applies global average pooling to the features of the title, then performs convolution and sigmoid activation to generate global title attention; DesA uses image description to construct query Q, and uses face features to form key K and value V, and realizes their fusion through cross attention mechanism; Then merge the four groups of result features with the face features to generate the final output.
2. The face super-resolution method based on visual language prior according to claim 1, characterized in that, In step one, it specifically includes the following steps: Input the low-resolution face image into the base large model SAM and DAM to extract the face semantic segmentation map and depth map as visual prior information.
3. The face super-resolution method based on visual language prior according to claim 1, characterized in that, In step one, it further includes the following steps: Input the low-resolution image into the BLIP2 model to generate the corresponding text title; To generate a description, combine the low-resolution face image and a specific prompt to provide useful visible features or observations about the face image, and input it into the ChatGPT-4 model to produce a detailed text description of the face image.
4. A system for face super-resolution based on visual language priors, characterized by, It comprises a computer module, which applies the visual-linguistic prior-based face super-resolution method according to any one of claims 1 to 3. 5.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-4 when the computer program is executed by the processor. The processor executes the computer program to realize the steps of the method according to any one of claims 1 to 3.
6. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the method according to any one of claims 1 to 3.