Face super-resolution method, system and device based on text information guidance and readable storage medium

Through the face super-resolution method based on text information guidance, the multimodal large language model and cross-attention mechanism are used to fuse text information, and the feature blurring and artifact problems caused by the face super-resolution method in the prior art relying on a single low-resolution image is solved, achieving higher quality and consistent face image reconstruction.

CN120031718APending Publication Date: 2025-05-23HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111785.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing super-resolution method of faces relies on a single low-resolution image, making it difficult to fully restore the fine texture and structural features of the face, resulting in blurred features, loss of details and artifacts in the generated results.

Method used

The face super-resolution method based on text information guidance is adopted to generate detailed text descriptions through a multimodal large language model, low-resolution face images and text descriptions are mapped to the latent feature space, text information is fused using the cross attention mechanism, and high-resolution images are generated through the residual diffusion generation module, and image quality and semantic consistency are finally optimized through text-aware loss.

Benefits of technology

It significantly improves the reconstruction quality of the super-resolution of faces, and the generated images have more vivid and realistic features and clearer details, and performs excellent in consistency with high-resolution face images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005256837900000031
    Figure BDA0005256837900000031
  • Figure BDA0005256837900000033
    Figure BDA0005256837900000033
  • Figure BDA0005256837900000043
    Figure BDA0005256837900000043
Patent Text Reader

Abstract

The invention relates to a face super-resolution method based on text information guidance. The method comprises the following steps of 1, text description generation, wherein text description is generated through a multi-modal large language model; step 2, potential space coding: mapping a low-resolution face image and text description to a potential feature space, and performing compact representation on the image by using a pre-trained encoder; step 3, text information fusion: embedding the generated text description into a visual feature processing process through a method based on a cross attention mechanism to form text-visual joint representation; 4, a residual diffusion generation module: in the potential space, realizing the generation of a low-resolution image to a high-resolution image through a Markov chain of residual connection; and step 5, text perception loss optimization: optimizing image quality and semantic consistency of a generated result by minimizing a potential space recovery error and a text consistency error. Compared with other generative models, the TFSR has the minimum parameter quantity, the highest sampling efficiency and the optimal FID score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer software and artificial intelligence technology, and specifically relates to a face super-resolution method, system, device and readable storage medium based on text information guidance. Background Art

[0002] With the rapid development of digital image processing technology and the increasing popularity of digital image acquisition equipment, face images are increasingly used in daily life. However, due to the limitations of sensor resolution and factors such as data transmission and storage costs, the resolution of many face images is often low, resulting in poor image quality and unclear details, which limits the application of face images in many fields. To address this problem, face image super-resolution has become a very popular research direction in the field of computer vision. Figure 1 As shown in Figure 1, face image super-resolution aims to reconstruct low-resolution (LR) face images into high-resolution (HR) images through algorithms, thereby improving image details and clarity and improving the recognition accuracy of downstream tasks such as face recognition and expression analysis. Figure 1 The comparison between the reconstruction result (TFSR) of the present invention and the low-resolution face image and the target high-resolution face image is shown.

[0003] Traditional face super-resolution methods usually rely on a single low-resolution image as input and directly interpolate, enlarge and reconstruct the image (e.g. Figure 2a ). In the absence of additional prior information, such methods are unable to fully restore the fine texture and structural features of the face. Specifically, due to limited input information, the model cannot accurately infer the subtle facial features contained in the original high-resolution image, which easily leads to feature blurring, loss of details, and obvious artifacts in the generated results. This not only affects the visual realism, but also reduces the overall performance of subsequent downstream tasks such as face recognition and expression analysis.

[0004] In order to overcome the limitation of relying on a single LR image, many methods have introduced rich and diverse prior information into the face super-resolution process in recent years (such as Figure 2b ). This prior information usually exists in the form of structured features, including facial key points, facial analysis graphs, and identity feature encoding. By integrating this prior information in the generation process, the model can have a more comprehensive understanding and guidance of the geometric layout, part proportions, and texture distribution of the face. It can not only significantly reduce artifacts and distortions caused by insufficient information, but also show better performance in detail restoration, making the output image closer to the real face appearance. At the same time, the supplement of these structured features also provides a more stable and reliable foundation for the subsequent analysis and recognition of face images.

[0005] At present, most face super-resolution (FSR) methods still over-rely on visual information in degraded low-resolution images, which not only limits their ability to accurately reconstruct details, but also easily produces artifacts and illusions when the image quality is poor, resulting in face reconstruction errors and identity distortion. Therefore, exploring rich prior information to regulate and optimize the reconstruction process has become a key direction for improving FSR quality.

[0006] To this end, some studies have begun to introduce structured features such as facial key points, parsed graphs, or identity information into the FSR model in order to provide clearer guidance when generating high-resolution images. This method improves the accuracy of the details of the generated images to a certain extent, and effectively reduces the artifacts and distortion problems caused by relying solely on low-resolution visual information. However, such methods often require additional data annotation and structured analysis tools, resulting in high data acquisition costs and limited applicable scenarios.

[0007] In sharp contrast, the potential of natural language cues in the field of FSR has not been fully explored. Inspired by the rapid development of multimodal large language models (MLLMs), we find a new idea: in addition to low-resolution images, can we use multimodal information from text to enrich the reconstruction prior (such as Figure 2c (as shown in the figure)? Text description can not only make up for the lack of visual information and provide a clearer description of facial features (such as age, gender, hairstyle, etc.), but also cover semantic dimensions such as expression, skin color, body posture, etc., thereby more effectively guiding the restoration process. At the application level, for example, in criminal investigations, if poor-quality surveillance images can be integrated with eyewitness descriptions, the accuracy and authenticity of FSR will be significantly improved. It can be seen that text-based FSR methods have great potential. However, in practice, the lack of high-quality text-image pairing data has become a bottleneck for the development of this field.

[0008] Current face super-resolution datasets often lack detailed text annotations, which makes it difficult for text-guided methods to fully demonstrate their advantages. Existing text generation methods, such as using probabilistic context-free grammars (PCFG) or BLIP models, can usually only provide simple and general descriptions and fail to capture key information such as facial expressions, feature details, or environmental background. This lack of description depth and granularity limits the further expansion and improvement of text-based FSR methods. Summary of the invention

[0009] In order to solve the above problems, the present invention further proposes a face super-resolution method, system, device and readable storage medium based on text information guidance.

[0010] The technical solution adopted by the present invention to solve the above problems is: a face super-resolution method based on text information guidance, comprising the following steps:

[0011] Step 1: Text description generation: Generate text description through a multimodal large language model;

[0012] Step 2: Latent space encoding: Map low-resolution face images and text descriptions to latent feature space and use pre-trained encoders to compactly represent images.

[0013] Step 3: Text information fusion: The generated text description is embedded into the visual feature processing process through a method based on the cross-attention mechanism to form a text-visual joint representation;

[0014] Step 4: Residual diffusion generation module: In the latent space, the Markov chain with residual connection is used to realize the generation of low-resolution to high-resolution images and reconstruct high-resolution face images;

[0015] Step 5: Text-aware loss optimization: Optimize the image quality and semantic consistency of the generated results by minimizing the latent space recovery error and text consistency error.

[0016] Furthermore, in step 1, a multimodal large language model including BLIP2, GRIT, Segment Anything and GPT-4o is used to generate a detailed text description for the image. The text content covers key facial features, texture details and scene information. The specific generation process can be expressed as:

[0017] Text=GPT(BLIP2(x),GRIT(x),SegmentAnything(x)).

[0018] Among them, x represents the input image and Text is the generated text description.

[0019] Furthermore, in step 2, the latent space encoding uses a pre-trained VQGAN model to embed the image from the pixel space into the latent feature space to reduce the computational complexity; and uses a pre-trained LongCLIP model to embed the text into the latent feature space. The process is expressed as:

[0020]

[0021] Where x and y are high-resolution face images and low-resolution face images, respectively. The VQGAN encoder E maps these images to the latent space; z 0 and z y The latent representations of high-resolution and low-resolution images, LongCLIP encoder Map the text description text into the CLIP feature space.

[0022] Furthermore, in step 3, the text fusion module introduces cross attention into the window attention mechanism of Swin Transformer to provide semantic guidance from reference text information in each generation process. The process is expressed as:

[0023]

[0024] in, represents the intermediate representation of SwinUNet, Represent the learnable projection matrices for query, key, and value vectors in the criss-cross attention mechanism, respectively.

[0025] Furthermore, in step 4, the residual diffusion generation module constructs a Markov chain through residual connections based on a potential diffusion model, so as to accurately generate a high-resolution image;

[0026] The forward diffusion process is represented as a Markov chain. The initial state of the Markov chain converges to the data distribution of a low-resolution face image, and the final state converges to the data distribution of a high-resolution face image. The Markov chain is connected in the middle by residuals.

[0027] Furthermore, in step 5, the text perception loss optimization includes the latent space recovery error and the CLIP space semantic consistency error, and the training objective is formulated as:

[0028]

[0029] Among them, f θ (z t ,z y ,z text ,t) represents the latent space representation of the image restored by the diffusion process, Indicates the mapping of the representation to the pixel space of the image, z text represents the text description corresponding to the high-resolution image x.

[0030] The present invention also relates to a face super-resolution system based on text information guidance, the system comprising a text generation module, a latent space encoding module, a text fusion module, a residual diffusion generation module and a text perception loss module.

[0031] Furthermore, a text generation module is used to generate detailed text descriptions associated with low-resolution face images; a latent space encoding module is used to encode low-resolution face images and text descriptions into latent feature representations; a text fusion module is used to embed text descriptions into visual feature processing; a residual diffusion generation module is used to generate high-resolution face images through a Markov chain; and a text-aware loss module is used to optimize the quality of generated images and text consistency.

[0032] The present invention also relates to a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0033] The present invention also relates to a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.

[0034] Beneficial Effects

[0035] Based on the deep learning method, this paper designs and proposes a new text-guided face super-resolution model TFSR. This model integrates text guidance to provide contextual support for the reconstruction of high-resolution face details, so that the generated images have more vivid and realistic features and clearer details, and the consistency with high-resolution face images is also excellent. By comparing the performance of this model with other advanced methods in the current field in terms of objective indicators and subjective perception, the model designed by this paper has the best performance. Compared with other generative models, TFSR has the least number of parameters, the highest sampling efficiency and the best FID score. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Schematic diagram of the comparison of face super-resolution reconstruction results.

[0037] Figure 2a Schematic diagram of traditional face super-resolution method.

[0038] Figure 2b Schematic diagram of the face super-resolution method based on structural information.

[0039] Figure 2c Schematic diagram of face super-resolution method based on text information.

[0040] Figure 3 This is a flowchart of the text-guided face super-resolution method of the present invention and a text generation framework diagram.

[0041] Figure 4 This is a comparison chart of the high-resolution face image after applying the method of the present invention and the results of other methods.

[0042] Figure 5 This is a comparison chart of the method of the present invention and other generation models in terms of FID score, sampling efficiency and parameter quantity. DETAILED DESCRIPTION

[0043] The following combination Figures 3 to 5 This embodiment will be described in detail.

[0044] like Figure 3 As shown, the face super-resolution method based on text information guidance of the present invention includes five stages: a text generation module, a latent space encoding module, a text fusion module, a residual diffusion generation module and a text perception loss module.

[0045] The text generation module is used to generate detailed text descriptions associated with low-resolution face images; the latent space encoding module is used to encode low-resolution face images and text descriptions into latent feature representations; the text fusion module is used to embed text descriptions into visual feature processing; the residual diffusion generation module is used to generate high-resolution face images through Markov chains; the text-aware loss module is used to optimize the quality of generated images and text consistency. The details are explained below.

[0046] The face super-resolution method based on text information guidance of the present invention comprises the following steps:

[0047] In step 1, text description generation is performed, and rich text descriptions are generated through a multimodal large language model. The specific text generation process is as follows: Figure 3 As shown in the "prompt generation" module in . Inspired by the Image2Paragraph method, the capabilities of multimodal large language models (MLLMs) and large language models (LLMs) are used to generate comprehensive text descriptions with the help of models such as BLIP2, GRIT, Segment Anything and GPT-4o. Specifically, BLIP2 generates a preliminary image caption that outlines the image content; GRIT adds detailed annotations for specific visual attributes; Segment Anything provides segmentation information and outlines detailed structural features; and GPT-4o integrates these diverse inputs into coherent paragraphs to summarize the high-level semantics and detailed visual features of the image. In the process of face restoration, combining low-resolution face images with rich text descriptions can provide deeper contextual guidance for generating more accurate and realistic face super-resolution images. Given x represents the input image and Text is the generated text description, the specific generation process can be expressed as:

[0048] Text=GPT(BLIP2(x),GRIT(x),SegmentAnything(x))

[0049] In step 2, latent space encoding is performed to map low-resolution face images and text descriptions to latent feature space, and the pre-trained encoder is used to compactly represent the image; for image encoding, given x and y are high-resolution images and low-resolution images respectively. Encoder E maps these images to latent space, and decoder D reconstructs them. 0 and z y denote the latent representations of high-resolution and low-resolution images, respectively, represents the representation of the restored image obtained by the diffusion model in the latent space, and It means that The restored image is mapped back to the pixel space. The process is expressed as:

[0050]

[0051] Specifically, the present invention adopts a pre-trained VQGAN model to perform the embedding process from the pixel space of the face image to the latent space. Of course, other pre-trained models can also be used for image embedding as needed, such as PixelCNN, VAE, etc.

[0052] For text encoding, the present invention uses a pre-trained LongCLIP encoder to map the text description into the CLIP latent feature space. Given a text description text, the LongCLIP encoder The generated text potential code is represented as z text , the text encoding process is expressed as:

[0053]

[0054] In step three, text information fusion is performed, and the generated text description is embedded into the visual feature processing process through a method based on a cross-attention mechanism to form a text-visual joint representation; the text fusion module of the present invention introduces cross-attention in the window attention mechanism of Swin Transformer, thereby providing semantic guidance from reference text information in each step of the generation process. In this way, the model not only relies on visual features, but also can use information from natural language descriptions, so that the restored image is consistent with the description and is more in line with people's expectations. The process can be expressed as:

[0055]

[0056] in represents the intermediate representation of SwinUNet, The learnable projection matrices represent the query, key, and value vectors in the cross-attention mechanism, respectively. This design enables the model to effectively integrate visual and textual information, leveraging their complementarity to obtain a more comprehensive face representation.

[0057] In step 4, the residual diffusion generation module realizes image generation. In the latent space, the generation of low-resolution to high-resolution images is realized through the residual connected Markov chain, and the high-resolution face image is reconstructed; the residual diffusion generation module is based on the latent diffusion model, and the Markov chain is constructed through residual connection to accurately generate high-resolution images, ensuring sampling efficiency and generation quality.

[0058] Let e ​​represent the residual between the HR image and the LR image in the latent space, that is, e = z 0 -z y The forward diffusion process is represented as a Markov chain. The initial state of the Markov chain converges to a data distribution approximately equal to x, and its final state converges to a data distribution approximately equal to y. The Markov chain is connected in the middle by the residual e. Specifically, the data distribution transformation from t-1 to t can be expressed as:

[0059]

[0060] where α t =β t -β t-1 ,t>1,α 1 =β 1 , transfer sequence As the time step t increases, it increases, satisfying β 1 →0 and β T →1. From z 0 to z t The conversion can be expressed as:

[0061]

[0062] In the backward sampling process, starting from the approximate distribution of the low-resolution image, we gradually sample from time step T forward until we obtain the z 0 From z y to z 0 In the process of transfer, z text Bootstrap, its posterior distribution p(z 0 |z y ,z text ) can be expressed as:

[0063]

[0064] Among them, θ represents the learnable parameters, which are implemented by SwinUNet architecture in this work and denoted as fθ (z t ,z y ,z text ,t). The parameter θ is optimized during the training process, while the mean parameter μ θ is parameterized as follows:

[0065]

[0066] In step five, text-aware loss optimization optimizes the image quality and semantic consistency of the generated result by minimizing the latent space recovery error and the text consistency error. The text-aware loss optimization includes the latent space recovery error and the CLIP space semantic consistency error, which are used to constrain the visual quality and semantic consistency of the generated image. During the training process, the goal of the present invention is to minimize two main parts: image reconstruction error and semantic consistency error. The training objective is formulated as:

[0067]

[0068] The first component is the latent space restoration distance, which minimizes the difference between the super-resolution facial image and the high-resolution image in the latent space, thereby ensuring high-quality image reconstruction. The second component is the CLIP space text information distance, which reduces the difference between the super-resolution image and its corresponding text description in the CLIP space, ensuring the semantic consistency of the generated image and its description text. In this way, the model not only focuses on the visual quality of the image, but also considers the relevance of the generated image to the natural language description, thereby improving the semantic consistency and visual effect of the image. Specifically, f θ (z t ,z y ,z text ,t) represents the latent space representation of the image restored by the diffusion process, Indicates the mapping of the representation to the pixel space of the image. text represents the natural language prompt corresponding to the high-resolution image x. When calculating the CLIP space loss, both the restored image and the corresponding text description are mapped to the CLIP space and their similarity is evaluated as shown in the formula:

[0069]

[0070] This formulation ensures that the reconstructed image is closely aligned with the semantic information provided by the natural language cues, thereby enhancing the overall quality and realism of the super-resolved image.

[0071] Effect verification

[0072] In order to verify the effectiveness of the algorithm proposed in the present invention, it is compared with existing super-resolution methods in multiple face super-resolution datasets, and the objective indicators and subjective effects are compared, and the efficiency of the present invention is tested.

[0073] Datasets and metrics

[0074] In this study, the CelebA-HQ dataset was used as the main resource for training and testing models. The CelebA-HQ dataset contains 30,000 high-resolution (HR) face images. During the training process, all images were resized to 256×256 pixels. In addition, descriptive text was generated for each image using advanced pre-trained large language models (LLM) and multimodal large language models (MLLM), totaling 30,000 text descriptions. For the CelebA-HQ test set, the images were downsampled 4 times by bicubic to generate low-resolution images for testing.

[0075] Three real-world datasets were selected for evaluation. The natural language prompts of these datasets were generated using the same procedure. The following is a brief description of these three real-world test datasets:

[0076] CelebChild-Test: This dataset contains 180 celebrity child images collected from the Internet. These images are generally of low quality, and many of them are old black-and-white photos.

[0077] LFW-Test: This dataset includes 1,711 face images of different identities, with slightly degraded features common in Internet images.

[0078] WebPhoto-Test: This dataset consists of 407 low-quality real-scene images from the Internet, including some old photos with serious details and colors degraded.

[0079] The quality of face images covers multiple aspects, such as clarity, contrast, and noise level. In order to conduct a comprehensive and rigorous quantitative evaluation of different methods, a series of widely recognized reference and non-reference evaluation metrics are adopted. For the reference evaluation metrics, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) are used to evaluate the fidelity of the image, while Learning Perceptual Patch Similarity (LPIPS) is used as a perceptual quality indicator. Frichette Embedding Distance (FID) measures the distribution similarity between the restored image and the high-quality reference image. Multi-scale Image Quality Assessment (MUSIQ) and CLIP-based Image Quality Assessment (CLIPIQA) are used as non-reference image quality assessment metrics, providing objective evaluation without relying on reference images.

[0080] Results Analysis

[0081] Comparative experiment

[0082] We evaluate the state-of-the-art super-resolution methods for faces and natural images, including DR2, DDNM, PGDiff, BFRfusion, StableSR, and PASD. The public implementations of these methods are used for testing to ensure the repeatability of the experiments. The detailed results are shown in Tables 1 and 2.

[0083] Table 1 shows the comparison results on the CelebA-HD Test dataset, while Table 2 compares the performance on three real-world datasets (CelebChild-Test, LFW-Test, and WebPhoto-Test). On all datasets, the TFSR method consistently outperforms other diffusion-based methods in all metrics. Specifically, TFSR achieves the highest PSNR, SSIM, FID, and LPIPS scores on the CelebA-HDTest dataset, and also achieves the best FID and LPIPS scores on the real-world datasets (except the FID metric for WebPhoto). This shows that TFSR excels in both pixel-level accuracy and perceptual quality, and has significant advantages in dealing with real-world degradation problems. It is worth noting that TFSR performs particularly well in the FID and LPIPS metrics, indicating that it can generate faces that are extremely realistic and perceptually consistent.

[0084] Compared with other methods, TFSR effectively integrates detailed natural language cues to enhance the super-resolution process, thereby reconstructing high-quality face images. Its balanced and excellent performance on multiple indicators highlights the combined advantages of this method in terms of accurate reconstruction and high perceptual quality. Overall, these results verify TFSR's position as a robust and leading solution in the field of face super-resolution.

[0085] Table 1 Quantitative comparison of the 4x face super-resolution reconstruction method based on CelebA-HD Test and existing methods

[0086]

[0087] Table 2 Quantitative comparison with different advanced methods on real-world datasets

[0088]

[0089] exist Figure 4The performance comparison of the TFSR method with other state-of-the-art methods is shown in Figure 1. The TFSR method combines low-resolution face images with text guidance and leverages powerful generative priors to generate more realistic features and delicate details. Among all methods, TFSR is closest to the high-resolution reference image in terms of facial details and overall appearance. As shown in the figure, TFSR performs well in recovering complex facial features, including skin texture, eye structure, and hair details. In contrast, other methods such as DDNM and PASD produce blurry and overly smooth results, while PGDiff introduces obvious artifacts in difficult areas such as hair and eyes. In the eye area (row 2), most models are able to generate relatively realistic results, but only TFSR can consistently generate clear details and accurate colors, such as blue eyeshadow and brown irises. In complex scenes, such as image inpainting to recover fine facial accessories such as flower crowns (row 4), TFSR is able to effectively reconstruct accessory details while preserving natural facial features. Other methods perform poorly in this scenario: DR2 and PGDiff produce blurry and color-inconsistent results, while BFRfusion and StableSR introduce noticeable distortions and fail to preserve the original appearance of the flower crown.

[0090] To further evaluate the sampling efficiency and performance of the TFSR model, a comparative analysis was conducted. The results are as follows: Figure 5 The comparison includes the time required to generate a single image, the FID score, and the number of parameters of different generative models. The TFSR model is significantly better than other generative models in all indicators, with the least number of parameters, the highest sampling efficiency, and the best FID score.

[0091] The method of the present invention has the following advantages:

[0092] First, the TFSR model significantly improves the reconstruction quality of face super-resolution by integrating text guidance information. Unlike traditional methods that rely only on visual information from low-resolution images, TFSR uses rich text descriptions to provide deeper contextual support for the generation process, thereby more accurately restoring facial details and features. This innovation effectively makes up for the lack of text information application in current face super-resolution methods and improves the authenticity and delicacy of the generated effect.

[0093] Secondly, TFSR introduces a special text fusion module, which effectively combines text and visual information, fully utilizing the complementarity of the two to form a more comprehensive face representation. The design of this module promotes the optimization of the generation process and further enhances the reconstruction accuracy and naturalness of the image.

[0094] In addition, the model adopts a residual-based diffusion structure, which effectively reduces the computational complexity in the hidden layer space and improves the sampling efficiency. This structure enables TFSR to perform well when processing high-resolution images, significantly reduces resource consumption, and adapts to more complex application scenarios.

[0095] Finally, the experimental results show that TFSR outperforms existing advanced methods on multiple datasets, demonstrating its comprehensive advantages in pixel-level accuracy and perceptual quality of face reconstruction, further verifying its position as a leading solution in the field of face super-resolution.

[0096] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiment when executing the computer program.

[0097] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.

[0098] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0099] Each module in the device or system of the present invention can be implemented in whole or in part by software, hardware, or a combination thereof. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to the above modules.

[0100] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0101] The above contents of the present invention are only preferred embodiments of the present invention and are not intended to limit the implementation scheme of the present invention. A person skilled in the art can easily make corresponding changes or modifications based on the main concept and spirit of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope required by the claims.

Claims

1. A face super-resolution method based on text information guidance, characterized in that: The following steps are involved: Step 1: Text description generation: Generate text description through a multimodal large language model; Step 2: Latent space encoding: Map low-resolution face images and text descriptions to latent feature space and use pre-trained encoders to compactly represent images. Step 3: Text information fusion: The generated text description is embedded into the visual feature processing process through a method based on the cross-attention mechanism to form a text-visual joint representation; Step 4: Residual diffusion generation module: In the latent space, the Markov chain with residual connection is used to realize the generation of low-resolution to high-resolution images and reconstruct high-resolution face images; Step 5: Text-aware loss optimization: Optimize the image quality and semantic consistency of the generated results by minimizing the latent space recovery error and text consistency error.

2. The face super-resolution method based on text information guidance according to claim 1, characterized in that: In step 1, a multimodal large language model including BLIP2, GRIT, Segment Anything, and GPT-4o is used to generate a detailed text description of the image. The text content covers key facial features, texture details, and scene information. The specific generation process can be expressed as: Text=GPT(BLIP2(x),GRIT(x),SegmentAnything(x)). Among them, x represents the input image and Text is the generated text description.

3. The face super-resolution method based on text information guidance according to claim 1, characterized in that: In step 2, the latent space encoding uses a pre-trained VQGAN model to embed the image from the pixel space into the latent feature space to reduce the computational complexity; the pre-trained LongCLIP model is used to embed the text into the latent feature space. The process is expressed as: Among them, x and y are high-resolution face images and low-resolution face images respectively, and the VQGAN encoder E maps these images to the latent space; z0 and z y The latent representations of high-resolution and low-resolution images, LongCLIP encoder Map the text description text into the CLIP feature space.

4. The face super-resolution method based on text information guidance according to claim 1, characterized in that: In step 3, the text fusion module introduces cross attention into the window attention mechanism of Swin Transformer to provide semantic guidance from reference text information in each generation process. The process is expressed as: in, represents the intermediate representation of SwinUNet, Represent the learnable projection matrices for query, key, and value vectors in the criss-cross attention mechanism, respectively.

5. The face super-resolution method based on text information guidance according to claim 1, characterized in that: In step 4, the residual diffusion generation module constructs a Markov chain through residual connections based on the potential diffusion model to accurately generate high-resolution images; The forward diffusion process is represented as a Markov chain. The initial state of the Markov chain converges to the data distribution of a low-resolution face image, and the final state converges to the data distribution of a high-resolution face image. The Markov chain is connected in the middle by residuals.

6. The face super-resolution method based on text information guidance according to claim 1, characterized in that: In step 5, the text perception loss optimization includes the latent space recovery error and the CLIP space semantic consistency error, and the training objective is formulated as: Among them, f θ (z t ,z y ,z text ,t) represents the latent space representation of the image restored by the diffusion process, Indicates the representation mapped to the pixel space of the image; z text represents the text description corresponding to the high-resolution image x.

7. A face super-resolution system based on text information guidance, the system comprising a text generation module, a latent space encoding module, a text fusion module, a residual diffusion generation module and a text perception loss module.

8. The face super-resolution system based on text information guidance according to claim 7, characterized in that: A text generation module is used to generate detailed text descriptions associated with low-resolution face images; a latent space encoding module is used to encode low-resolution face images and text descriptions into latent feature representations; a text fusion module is used to embed text descriptions into visual feature processing; The residual diffusion generation module is used to generate high-resolution face images through Markov chains; the text-aware loss module is used to optimize the quality of generated images and text consistency.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.