Progressive face super-resolution method and system based on artistic prior knowledge
Through a progressive face super-resolution method based on artistic prior knowledge, the structure-detail-fusion modeling process is gradually introduced, which solves the problems of structural confusion and false details in existing technologies, improves the image reconstruction quality and stability, and is suitable for security monitoring and face recognition.
Patent Information
- Application Number
- CN202511053144.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing facial image super-resolution methods have limitations in structural modeling, detail completion, and prior adaptability. Especially in high-magnification super-resolution tasks, it is difficult to ensure the structural consistency, detail authenticity, and overall naturalness of the image.
A progressive face super-resolution method based on artistic prior knowledge is adopted. Multi-scale features are extracted through the image coding module. Combined with the structure learning module, detail completion module and mask fusion module, the structure-detail-fusion collaborative modeling process is gradually introduced to simulate the human thinking process of drawing faces, guide the model to restore layer by layer from structure to details, and use dynamic dictionary units and mask fusion mechanisms to improve image quality.
It significantly improves the reconstruction quality of images, enhances the fidelity of key facial features and the consistency and naturalness of local features, avoids image distortion caused by unfounded detail generation, and is suitable for a variety of low-quality face image restoration scenarios.
Smart Images

Figure CN120563326B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and deep learning technology, and in particular to a progressive face super-resolution method and system based on artistic prior knowledge. Background Art
[0002] Face super-resolution (FSR), a key visual restoration task, aims to reconstruct facial images with high accuracy and fidelity from low-quality input. Early approaches primarily used feature modeling to upscale the image overall and attempt to recover missing details. However, these methods have limited effectiveness when processing extremely low-resolution images, particularly in restoring structure and detail in key facial regions.
[0003] With the development of deep learning, super-resolution methods based on convolutional neural networks have become mainstream, capable of learning the complex mapping relationship between images and high-resolution objects. However, due to their limited receptive field, they often struggle to capture long-range dependencies and global semantics in facial images, resulting in insufficient structural modeling capabilities. To address this, subsequent research has introduced attention mechanisms and architectures that incorporate long-range modeling capabilities, thereby improving image detail and regional consistency to a certain extent.
[0004] On the other hand, as research deepens, the introduction of auxiliary information such as structural priors, geometric information, and facial parsing maps has become an important means of improving super-resolution performance. These methods effectively enhance the consistency of facial contours and organs by guiding the model to focus on the structure of local facial regions. However, these priors are mostly based on low-frequency structural information, which has limited adaptability and is susceptible to interference, making them difficult to meet the reconstruction requirements of large-scale magnification and complex scenes.
[0005] To address the difficulty of recovering high-frequency details, prior modeling approaches based on dictionary learning have emerged in recent years. These methods construct feature dictionaries to complement and enhance image details, improving facial texture restoration and identity consistency. However, existing dictionary implementations often rely on static structures or global clustering, lacking the ability to dynamically adapt to image-specific details, resulting in unstable performance across multiple individuals or scenarios.
[0006] In summary, existing methods have certain limitations in structural modeling, detail completion, and prior adaptability. Especially in high-resolution super-resolution tasks for real applications, it is difficult to simultaneously ensure the structural consistency, detail authenticity, and overall naturalness of the image. Summary of the Invention
[0007] The present invention aims to provide a progressive face super-resolution method and system based on artistic prior knowledge, which can effectively enhance the fidelity of key facial features, ensure the consistency and naturalness of local features while improving the overall image clarity, and significantly improve the unreality and instability caused by relying on speculative generation in existing technologies, thereby improving the reconstruction quality and application reliability of facial images.
[0008] To achieve the above objectives, the present invention provides the following basic solutions.
[0009] Option 1
[0010] A progressive face super-resolution system based on state-of-the-art prior knowledge, including:
[0011] Image encoding module, used to extract multi-scale features from the input low-resolution face image and output latent features;
[0012] The structure learning module includes: a structure parsing submodule for receiving the latent features and generating a facial parsing map, wherein the facial parsing map includes spatial distribution information of key facial parts; a structure completion submodule for generating modulation parameters based on the features of the facial parsing map, and performing channel modulation on the latent features through channel division to output structure enhancement features;
[0013] The detail completion module includes: a dynamic dictionary unit for storing detail feature elements obtained through high-resolution image training; a detail fusion unit for calculating the similarity between the structure enhancement feature and the elements in the dynamic dictionary unit, performing weighted fusion on the dictionary elements based on the similarity weight, and outputting the detail enhancement feature;
[0014] The mask fusion module includes: a mask generation unit for generating a region-aware weight mask based on the facial parsing map; a feature screening unit for multiplying the weight mask with the structure enhancement feature and the detail enhancement feature point by point to achieve spatial region alignment; and a cross-region interaction unit for performing semantic spatial fusion on the aligned features to generate fused features.
[0015] The super-resolution decoder is used to receive the fused features and reconstruct a high-resolution face image.
[0016] Option 2
[0017] The progressive face super-resolution method based on state-of-the-art prior knowledge includes the following steps:
[0018] Step 1: Perform multi-scale feature extraction on the input low-resolution face image to generate latent features;
[0019] Step 2: generating a facial parsing map based on the latent features, generating modulation parameters according to the features of the facial parsing map, performing channel division and channel modulation on the latent features, and outputting structure enhancement features;
[0020] Step 3: Calculate the similarity between the structure enhancement feature and the elements in the pre-built detail dictionary, perform weighted fusion on the dictionary elements according to the similarity weight, and output the detail enhancement feature; wherein the detail dictionary is dynamically updated when trained with high-resolution images and is frozen during the inference phase and super-resolution use;
[0021] Step 4: Generate a region-aware weight mask based on the facial analysis graph, use the weight mask to spatially filter the structure enhancement features and detail enhancement features respectively, and perform region alignment semantic fusion on the filtered features to generate fused features;
[0022] Step 5: Input the fused features into a super-resolution decoder to reconstruct a high-resolution face image.
[0023] The working principle and advantages of the present invention are:
[0024] The present invention is based on a progressive face super-resolution method and system based on artistic prior knowledge. Starting from the collaborative modeling of the three stages of "structure-detail-fusion", it gradually introduces the easy-to-difficult composition process in artistic creation thinking. Through structure-guided learning to complete the basic contour, detail dictionary to enhance local texture, and mask fusion mechanism to coordinate coarse and fine information, it effectively solves the shortcomings of traditional methods in terms of structural confusion and false details, and significantly improves the authenticity, stability and expressiveness of image restoration. The key points are:
[0025] This solution can fit the real generation logic and avoid unfounded model "fantasy": This solution introduces a "from simple to complex" prior learning strategy inspired by artistic creation, simulates the human thinking process of gradually drawing a face, and guides the model to restore layer by layer from structure to details. That is, first establish accurate facial topological relationship cognition (such as the relative position of eyes, nose, and mouth) based on the structural learning module, and then combine it with the detail completion module to supplement high-frequency details based on structural constraints, cutting off the parameter path of unfounded detail generation from the data source, avoiding the image distortion problem caused by "imagining" details out of thin air in traditional methods.
[0026] This solution can enhance the ability to recover structural information and improve the accuracy of facial restoration: by introducing a structural parsing submodule and a structural completion submodule, the facial parsing map generated by the former can provide pixel-level semantic supervision, while the latter can focus on key contour features; it effectively enhances the system's ability to model the geometric layout of the face and the relative relationships of key parts, significantly improving the fidelity of the reconstructed image at the structural level.
[0027] This solution's detail completion is authentic and credible, significantly enhancing the image texture expressiveness: This solution constructs a high-quality feature dictionary (corresponding to a dynamic dictionary unit) and performs detail-level feature matching, enabling the system to accurately complete texture details in high-magnification scenes, avoiding image blur or artifacts, and restoring more recognizable facial images.
[0028] This scheme designs a coarse and fine fusion mechanism to enhance the overall naturalness and layering of the image: This scheme adaptively fuses coarse structure and detail feature information by designing a mask fusion module, so that the generated image has stronger visual realism and a natural sense of detail transition while ensuring structural consistency.
[0029] This solution is suitable for a variety of low-quality facial image restoration scenarios and has broad application prospects: this solution can be applied to various fields such as security monitoring and face recognition. It is particularly suitable for scenarios where the original image resolution is extremely low or there is serious lack of details. It has good scalability and practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Schematic diagram of the system module composition of an embodiment of the progressive face super-resolution method and system based on artistic prior knowledge of the present invention. DETAILED DESCRIPTION
[0031] The following is a further detailed description through specific implementation methods:
[0032] The embodiment is basically as shown in the attached Figure 1 Shown: A progressive face super-resolution system based on artistic prior knowledge, including an image encoding module, a structure learning module, a detail completion module, a mask fusion module, and a super-resolution decoder.
[0033] The image encoding module is used to extract multi-scale features from the input low-resolution face image and output potential features through the backbone encoder.
[0034] Specifically, the image encoding module performs multi-scale feature extraction based on the Transformer structure to retain information at each scale level and prepare for subsequent structure and detail modeling.
[0035] Furthermore, before feature extraction, the input low-resolution face image is pre-processed, wherein the pre-processing includes performing a convolution operation on the input low-resolution image to generate an embedded feature map suitable for subsequent processing.
[0036] Preferably, the image encoding module also embeds a structural supervision signal during the encoding process to enhance the characteristic response to facial contours and the relationship between the eyes, mouth, and nose, thereby improving the main encoder's ability to model key structural areas. This setting helps improve the system's perception and restoration accuracy of facial structural information, ensuring the geometric consistency of the generated image.
[0037] Specifically, the embedding of the structural supervision signal is only triggered during the training phase and removed during the inference phase.
[0038] In this embodiment, the backbone encoder of the image coding module includes an input layer, a preprocessing layer, a multi-scale coding block and an output layer. Convolutional layer, used to convert the input image into an embedded feature map; the multi-scale encoding block includes a resolution of The first Transform coding block, resolution is 16 The second Transform coding block with a resolution of The third Transform encoding block with a resolution of The fourth Transform encoding block.
[0039] When embedding the structural supervision signal, the structural supervision branch is inserted at the output end of the third Transform coding block of the image coding module to achieve the embedding of the structural supervision signal. The structural supervision branch is provided with a structural parsing header and a supervision signal layer. The structural parsing header includes the following: Convolutional layer, transposed convolutional layer and 3 Convolutional layer.
[0040] The structure supervision branch takes the features output by the third Transform coding block as input, and generates structure supervision analysis features after multi-layer convolution processing by the structure analysis head. The supervision signal layer receives the structure supervision analysis features. and the true label , the true label refers to the standard analytical features corresponding to the high-resolution image, which marks the accurate areas of each part of the face and is used to supervise the model learning. Then calculate the loss of the two , through the loss The gradient backpropagation optimizes the backbone encoder parameters, including the parameters of its multi-scale encoding blocks.
[0041] Specifically, ; Refers to the cross entropy loss.
[0042] The structure learning module includes a structure analysis submodule and a structure completion submodule.
[0043] In this embodiment, each module in the system constitutes an art prior learning framework network, and the backbone encoder of the image encoding module is located in the backbone super-resolution network of the art prior learning framework network. The backbone super-resolution network includes two branches: one is the main task branch of the image super-resolution reconstruction task, and the other is the auxiliary task branch of the facial analysis map generation.
[0044] The facial parsing map generation branch (corresponding to the structural parsing submodule, which receives the latent features and generates a facial parsing map containing the spatial distribution of key facial features) specifically uses a Transformer-based decoder to receive the latent features and generate a facial parsing map. This facial parsing map labels and delineates different regions within the facial image, accurately displaying the position, shape, and spatial distribution of various facial components (such as the eyes, nose, mouth, eyebrows, and outlines). This parsing map not only reflects the module's understanding of facial structure but also serves as a modulated signal for feature processing in image super-resolution reconstruction.
[0045] The specific steps include:
[0046] S1, receives the latent features from the image encoding module ;in, is a set of real numbers, H, W, and C are the height, width, and number of feature channels of the feature map of the potential feature, respectively.
[0047] S2, feature reconstruction based on the Transformer decoder, includes: taking the latent feature F as input, using a three-layer Transformer decoder, performing multi-head self-attention operations on each layer, performing attention calculations and generating an attention weight matrix; with the help of the attention weight matrix, the association between related regions in the features can be identified (such as the geometric constraint relationship between the left eye and the right eyebrow of the face), which helps to strengthen the expression of the feature's regional association.
[0048] S3, based on the structure-aware features output by the Transformer decoder, uses multi-layer convolution operations as feature post-processing to restore the spatial structure layer by layer and enhance boundary continuity, ultimately generating a facial parsing map.
[0049] Specifically, in this embodiment, through multi-layer convolution (such as 3 layers Convolution) operation, based on step-by-step upsampling to enlarge the size of the feature map and simultaneously refine the feature details, gradually restore the target resolution and improve the boundary and category distinction, to achieve layer-by-layer restoration of the spatial structure. For example: the structural perception feature output based on the Transformer decoder is assumed to be , after multi-layer convolution and upsampling operations, the output can be The new feature of K corresponds to the number of facial parts based on the facial structure distribution.
[0050] Conditional Random Field (CRF) or differentiable boundary loss optimization is then used to enhance boundary continuity. For example, the original convolution output may have a "jagged boundary", but after CRF post-processing, the mouth boundary can be made smoother and more continuous.
[0051] Furthermore, the facial analysis map encodes each part in the color domain, and can intuitively present the spatial contours and shape features of facial areas such as the eyes, mouth, and nose.
[0052] Specifically, in this embodiment, based on the distribution of facial structure, it can be divided into eyes (including the entire area of the left eye and the right eye), nose (including the bridge of the nose, the tip of the nose, etc.), mouth (including the upper and lower lips, the lip line area), eyebrows (including the left eyebrow and the right eyebrow), forehead (the area from the hairline to the brow bone), cheeks (the skin area on both sides of the face) and other parts, and can be fine-tuned according to the scene (for example, whether to split the "left eye" and "right eye" into independent areas).
[0053] For example, for The new feature uses softmax to calculate the probability of each pixel belonging to each facial part, takes the part category with the highest probability as the part category to which each pixel belongs, and assigns different colors to each part (such as left eye = blue, mouth = red), thereby obtaining a facial analysis map.
[0054] The structure completion submodule is configured to generate modulation parameters based on the facial parsing image features, perform channel modulation on the latent features through channel division, and output structure-enhanced features. The structure completion submodule can modulate the structural features extracted from the facial parsing image into the super-resolved features of the main branch to form structure-enhanced features.
[0055] Compared with the method of directly fusing structural graphs, this module realizes information fusion through parameterized mapping, which has stronger adaptability and generalization ability.
[0056] Specifically, the structure completion submodule implements feature enhancement through the following operations:
[0057] The latent features are divided into two subsets according to the number of channels, where half of the channels remain unchanged and the other half of the channels are modulated; for example, for the latent features , along the channel dimension C (i.e., the number of channels), it is divided into two parts, including: Subset 1: corresponding to the first C / 2 channels, used to retain the original information, for easy distinction, it is defined as ; Subset 2: corresponds to the first C / 2 channels, used for structural modulation, for easy distinction, it is defined as .
[0058] The features of the previous layer of the facial parsing image decoder are used as modulation features, and two modulation parameters are generated through different convolution layers. The modulation parameters include a scaling factor and an offset factor. In this embodiment, the modulation features are respectively obtained through two independent Convolutional layer, the output is two parameter tensors, corresponding to the scaling factor and the offset factor b .
[0059] Based on the modulation parameters generated by the facial analysis graph, an affine transformation is applied to the channel subset that needs to be modulated (i.e., subset 2), and the modulated features are obtained. . , Represents element-wise multiplication.
[0060] Finally, the modulated features With unmodulated characteristics Splicing is performed on the channel dimension to form the modulated latent features , and make a residual connection with the original latent feature F. Specifically, first calculate The residual with the original latent feature F: Then the original potential feature F and the residual Addition (skip connection): ; Then complete the residual connection. This is the structural enhancement feature.
[0061] The detail completion module includes a dynamic dictionary unit and a detail fusion unit.
[0062] The dynamic dictionary unit is used to store detail feature elements obtained through high-resolution image training. Specifically, each detail feature element represents a typical detail pattern, such as skin texture, which represents the subtle texture features of facial skin, such as small pores, wrinkles, and skin smoothness; eye details, which include the subtle structural features around the eyes, such as eyelash distribution, eyelid folds, and details of the canthus.
[0063] The dynamic dictionary unit is constructed in the following way:
[0064] Initialize the dictionary elements during the training phase. In this embodiment, the dictionary size is set to , the feature dimension is set to .dictionary Initialized using a normal distribution with mean , standard deviation .
[0065] The feature matching loss function is used to drive the iterative update of dictionary elements, making them adaptable to various types of texture details.
[0066] In the inference phase, the dictionary elements are fixed and only cosine similarity retrieval and weighted fusion are performed.
[0067] By iteratively updating elements during the training phase and freezing them during the usage phase (i.e., the inference phase), this solution ensures that the dynamic dictionary unit continuously accumulates knowledge while being able to adapt to details across images and scenes, significantly improving the accuracy and richness of the reconstructed image.
[0068] The detail fusion unit calculates the cosine similarity between the structure-enhancing features and the elements in the dynamic dictionary unit, performs weighted fusion of the dictionary elements based on the similarity weights, and outputs detail-enhancing features, thereby supplementing and enhancing image detail information. This approach avoids the "one-size-fits-all" drawbacks of traditional static dictionaries, ensuring that each image receives the most targeted detail guidance.
[0069] Specifically, for structural enhancement features Each spatial position vector , calculate its difference with dictionary D Each element in Cosine similarity of , ; N represents the size of the dictionary elements.
[0070] Similarity score Perform softmax normalization and obtain weights : .
[0071] According to weight Linearly combine the dictionary elements to obtain detail enhancement features ; Repeat the above operation for all spatial locations to generate a detail enhancement feature map .
[0072] The mask fusion module includes a mask generation unit, a feature screening unit and a cross-region interaction unit.
[0073] The mask generation unit is used to generate a region-aware weight mask based on the facial analysis map. The mask, as region guidance information, can adaptively control the degree of fusion of coarse and fine-grained information according to the structural differences of different regions of the face, thereby avoiding interference from irrelevant features.
[0074] Specifically, the partition weight mask generated by the mask generating unit satisfies:
[0075] Spatial adaptability: Using facial parsing maps to guide the generation of weighted masks, different facial regions (such as eyes, nose, and mouth) are weighted differently and dynamically, ensuring that the mask can adaptively adjust weights based on the texture refinement requirements of different regions, promoting the precise fusion of coarse and fine-grained features in space.
[0076] Specifically, the generation of weight mask guided by facial analysis map includes the following operations: , convert P into one-hot encoding Then it is processed by a small convolutional network (a two-layer CNN network is used in this embodiment) , and output weight mask , It is a 2-channel feature map, including: structural feature weights and detail feature weights ; and perform Softmax activation on the two channels of M, so that their sum is 1, that is, the constraint makes , thereby realizing spatially adaptive fusion ratio control.
[0077] Integration consistency - The mask generation unit assigns adaptive weights to spatial regions by parsing features, then concatenates the obtained coarse-grained and fine-grained features (i.e., structure-enhanced features and detail-enhanced features). An attention-driven fusion mechanism is then introduced to adaptively emphasize the complementary information of different features in the concatenated features, enabling the two types of features to be effectively and harmoniously fused, thereby significantly improving the overall expressiveness and accuracy of the fused features.
[0078] The specific operations include:
[0079] First, perform spatial filtering: enhance features with structures from the structure learning module , detail enhancement features from the detail completion module , and the weight mask from the mask generation unit For input.
[0080] Spatial filtering is achieved by point-by-point multiplication: ; This step completes the process of assigning adaptive weights to spatial regions by analyzing features.
[0081] Then perform feature splicing: .
[0082] Then calculate through the self-attention mechanism The spatial importance of is used to generate the attention map A. Specifically, the spatial importance weight is first calculated through the convolution layer: ; ; Then activate and normalize it: ; .
[0083] Output fusion features .
[0084] The feature screening unit is used to perform point-by-point multiplication of the weight mask with the structure enhancement feature and detail enhancement feature to achieve spatial region alignment and generate a filtered regional feature representation. This operation achieves spatial screening and regional alignment of information sources at the data level, ensuring that features from different sources have a good correspondence in spatial position.
[0085] The cross-region interaction unit is used to perform semantic spatial fusion on the aligned features to generate fused features. Specifically, the coarse and fine features (i.e., structure-enhanced features and detail-enhanced features) obtained through different masks are first concatenated and then attention-generated. Next, the attention scores are summed and added to the original features to generate the fused features for subsequent decoding operations.
[0086] Through semantic space fusion, the system's ability to model details in key areas (such as the eyes and mouth) is enhanced. The fused result (i.e., the fused feature) simultaneously preserves the structural integrity of coarse-grained features and the texture accuracy of fine-grained features, providing a high-quality feature foundation for final image reconstruction. This results in superior structural consistency, rich detail, and natural transitions in the final output image, significantly enhancing the layering and realism of the super-resolution result, resulting in more natural and clear super-resolution facial images.
[0087] The super-resolution decoder is used to receive the fused features and reconstruct a high-resolution face image.
[0088] Furthermore, this system adopts a multi-task joint optimization strategy during the training phase, including: using a combination of perceptual loss, pixel loss and structural consistency loss for optimization. The total loss formula is as follows:
[0089] ;
[0090] in, and is the reconstruction loss and perceptual loss of the super-resolution branch (i.e., the main task branch of the image super-resolution reconstruction task); The loss for the parsing graph branch (i.e., the auxiliary task branch for facial parsing graph generation); for the loss of the reconstruction mission; Learn the loss for the dictionary.
[0091] 、 、 、 Corresponding to various types of losses ( 、 、 、 ) is used to adjust the proportion of different losses in the total loss and can be set and optimized according to training requirements, task focus, etc. The value range is 0.01~0.1, which is used to balance high-frequency details and pixel accuracy; The value range of is 0.1~0.5, which is used to control the structural supervision strength; The value range of is 0.5~1.0, which is used to control the pixel-level fidelity; The value range of is 0.001~0.01, which is used to prevent the dictionary elements from being overly correlated. 、 、 、 The value of is determined by adjusting the existing validation set (such as the CelebA dataset), and the specific value is: 、 、 、 .
[0092] The super-resolution branch is trained with reconstruction loss and perceptual loss to jointly ensure low-level accuracy and high-level visual quality:
[0093] ;
[0094] ;
[0095] in, is the super-resolution result, is the corresponding training label, It is a feature extractor based on the pre-trained VGG network.
[0096] In order to better facilitate the optimization of the two tasks and obtain accurate structural guidance, in this embodiment, loss constraints are imposed on the generated parsing graph:
[0097] ;
[0098] in, Represents the facial parsing graph output by the parsing graph branch, Represents the corresponding facial analysis image label.
[0099] In order to enhance detail modeling, the dictionary learning branch (corresponding to the detail completion module) is jointly optimized through two objectives: one is to generate images With real images The other is a dictionary constraint loss that regularizes the self-similar matrix by minimizing the off-diagonal correlation of the matrix to reduce redundancy and encourage diverse representations.
[0100] ;
[0101] ;
[0102] Where D represents the dictionary element, and Represents the i-th and j-th elements in the dictionary; N represents the size of the dictionary elements, and B represents the batch size. is the similarity measurement function.
[0103] To generate an image, specifically the reconstructed image output by the detail completion module (dictionary learning branch).
[0104] The multi-task joint optimization strategy also includes: the structure analysis submodule uses structural supervision signals to assist in training; the detail fusion unit introduces feature matching loss in the detail completion process to enhance the system's modeling ability for high-frequency areas.
[0105] Through the above multi-task joint optimization strategy, the synchronous optimization of the overall structure and local texture can be ensured, further improving the reconstruction accuracy.
[0106] This embodiment further provides a progressive face super-resolution method based on artistic prior knowledge, the implementation details of which are consistent with the operating method of a progressive face super-resolution system based on artistic prior knowledge, including the following steps:
[0107] Step 1: Perform multi-scale feature extraction on the input low-resolution face image to generate latent features.
[0108] Step 2: Generate a facial parsing map based on the latent features, generate modulation parameters according to the facial parsing map, perform channel division and channel-by-channel modulation on the latent features, and output structure enhancement features.
[0109] Specifically, the step 2 includes: dividing the potential features into two subsets according to the number of channels, wherein half of the channels remain unchanged and the other half of the channels are modulated;
[0110] The features of the previous layer of the facial parsing map decoder are used as modulation features. These features are passed through different convolutional layers to generate two modulation parameters, including a scaling factor and an offset factor. Based on the modulation parameters generated by the facial parsing map, an affine transformation is applied to the subset of channels that need to be modulated.
[0111] The modulated features are then concatenated with the unmodulated features in the channel dimension to form the modulated latent features, which are then connected with the residual of the original latent features to obtain the structural enhancement features.
[0112] Step 3: Calculate the similarity between the structure enhancement feature and the elements in the pre-built detail dictionary, perform weighted fusion on the dictionary elements according to the similarity weight, and output the detail enhancement feature; wherein the detail dictionary is dynamically updated when trained with high-resolution images and is frozen during the inference phase and super-resolution use.
[0113] The pre-built detail dictionary corresponds to a dynamic dictionary unit, and the elements in the detail dictionary are detail feature elements obtained through high-resolution image training. The detail dictionary is constructed in the following way:
[0114] Initialize dictionary elements with normal distribution during training;
[0115] The feature matching loss function drives the iterative update of dictionary elements to adapt to various types of texture details;
[0116] In the inference phase, the dictionary elements are fixed and only similarity retrieval and weighted fusion operations are performed.
[0117] Step 4: Generate a region-aware weight mask based on the facial analysis graph, use the weight mask to spatially filter the structure enhancement features and detail enhancement features respectively, and perform region alignment and semantic fusion on the filtered features to generate fused features.
[0118] Step 5: Input the fused features into a super-resolution decoder to reconstruct a high-resolution face image.
[0119] This embodiment provides a progressive face super-resolution method and system based on artistic prior knowledge, which can effectively enhance the fidelity of key facial features, while improving the overall image clarity, ensuring the consistency and naturalness of local features, and significantly improving the unreality and instability caused by relying on speculative generation in existing technologies, thereby improving the reconstruction quality and application reliability of facial images.
[0120] Furthermore, this solution was fully validated on two public face image datasets, CelebA and Helen. Specifically, on the CelebA and Helen datasets, we applied existing mainstream face super-resolution methods (DIC, SPARNet, KDFSR, CTCNet, MOHA, SFMNet, and DPI) along with our solution to super-resolution image reconstruction. The experimental results are shown in Table 1 below.
[0121] Table 1 Experimental results
[0122]
[0123] As shown in Table 1, compared with existing mainstream face super-resolution methods, our proposed scheme achieves significant improvements on commonly used evaluation metrics (including Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM), particularly superior in restoring details in key facial regions such as the eyes and lips. Furthermore, our proposed scheme maintains its superiority across both datasets and at different magnifications, demonstrating its adaptability to diverse facial data scenarios and robustness. Furthermore, subjective evaluation results demonstrate that the images generated by our proposed scheme are more realistic and recognizable, exhibiting good visual consistency and aesthetic perceptibility. This proposed scheme balances structural integrity and detail fidelity, and offers advantages such as progressive restoration and scalability, providing a superior solution for face super-resolution tasks.
[0124] The above is only an embodiment of the present invention. Common knowledge such as the specific structure and characteristics of the scheme is not described in detail here. Ordinary technicians in the relevant field are aware of all common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the relevant field can improve and implement this scheme in combination with their own abilities under the guidance of this application. Some typical well-known structures or well-known methods should not become obstacles for ordinary technicians in the relevant field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent.
Claims
1. A progressive face super-resolution system based on artistic prior knowledge, characterized by: include: Image encoding module, used to extract multi-scale features from the input low-resolution face image and output latent features; The structure learning module includes: a structure parsing submodule for receiving the latent features and generating a facial parsing map, wherein the facial parsing map includes spatial distribution information of key facial parts; a structure completion submodule for generating modulation parameters based on the features of the facial parsing map, and performing channel modulation on the latent features through channel division to output structure enhancement features; The structure completion submodule achieves feature enhancement through the following operations: The latent features are divided into two subsets according to the number of channels, wherein half of the channels remain unchanged and the other half of the channels are modulated; Applying an affine transformation to a subset of channels to be modulated based on modulation parameters, and obtaining modulated features; the modulation parameters include a scaling factor and an offset factor; The modulated features are concatenated with the unmodulated features in the channel dimension to form the modulated latent features, which are then residually connected with the original latent features to obtain the structure-enhanced features. The detail completion module includes: a dynamic dictionary unit for storing detail feature elements obtained through high-resolution image training; a detail fusion unit for calculating the similarity between the structure enhancement feature and the elements in the dynamic dictionary unit, performing weighted fusion on the dictionary elements based on the similarity weight, and outputting the detail enhancement feature; The mask fusion module includes: a mask generation unit for generating a region-aware weight mask based on the facial parsing map; a feature screening unit for multiplying the weight mask with the structure enhancement feature and the detail enhancement feature point by point to achieve spatial region alignment; and a cross-region interaction unit for performing semantic spatial fusion on the aligned features to generate fused features. The super-resolution decoder is used to receive the fused features and reconstruct a high-resolution face image.
2. The progressive face super-resolution system based on artistic prior knowledge according to claim 1, characterized in that The image encoding module performs multi-scale feature extraction based on the Transformer structure.
3. The progressive face super-resolution system based on artistic prior knowledge according to claim 1, characterized in that The dynamic dictionary unit is constructed in the following way: Initialize dictionary elements during training; The feature matching loss function is used to drive the iterative update of dictionary elements, making them adaptable to various types of texture details. In the inference phase, the dictionary elements are fixed and only similarity retrieval and weighted fusion are performed.
4. The progressive face super-resolution system based on artistic prior knowledge according to claim 1, characterized in that The partition weight mask generated by the mask generation unit satisfies: Spatial Adaptability - Using facial parsing maps to guide the generation of weight masks to perform differentiated and dynamic weighting of different facial regions.
5. The progressive face super-resolution system based on artistic prior knowledge according to claim 4, characterized in that The partition weight mask generated by the mask generating unit also satisfies: Integration consistency - The mask generation unit assigns adaptive weights to spatial regions by parsing features, then performs feature splicing on the obtained structure enhancement features and detail enhancement features, and then introduces an attention-based fusion mechanism to adaptively emphasize the complementary information of different features in the spliced features, so that the two types of features are harmoniously fused.
6. The progressive face super-resolution system based on artistic prior knowledge according to claim 1, characterized in that The system adopts a multi-task joint optimization strategy during the training phase, including: using a combination of perceptual loss, pixel loss and structural consistency loss for optimization.
7. The progressive face super-resolution system based on artistic prior knowledge according to claim 6, characterized in that The multi-task joint optimization strategy also includes: the structure analysis submodule adopts the structure supervision signal to assist in training; and the detail fusion unit introduces feature matching loss in the detail completion process.
8. A progressive face super-resolution method based on artistic prior knowledge, characterized by: The following steps are involved: Step 1: Perform multi-scale feature extraction on the input low-resolution face image to generate latent features; Step 2: generating a facial parsing map based on the latent features, generating modulation parameters according to the features of the facial parsing map, performing channel division and channel modulation on the latent features, and outputting structure enhancement features; The step of performing channel division and channel modulation on the latent features and outputting structure enhancement features comprises the following operations: dividing the latent features into two subsets according to the number of channels, wherein half of the channels remain unchanged and the other half of the channels are modulated; Applying an affine transformation to the subset of channels that need to be modulated based on modulation parameters, wherein the modulation parameters include a scaling factor and an offset factor; The modulated features are then concatenated with the unmodulated features in the channel dimension to form the modulated latent features, which are then connected with the residuals of the original latent features to obtain the structural enhancement features. Step 3: Calculate the similarity between the structure enhancement feature and the elements in the pre-built detail dictionary, perform weighted fusion on the dictionary elements according to the similarity weight, and output the detail enhancement feature; wherein the detail dictionary is dynamically updated when trained with high-resolution images and is frozen during the inference phase and super-resolution use; Step 4: Generate a region-aware weight mask based on the facial analysis graph, use the weight mask to spatially filter the structure enhancement features and detail enhancement features respectively, and perform region alignment semantic fusion on the filtered features to generate fused features; The spatial screening of the structure enhancement feature and the detail enhancement feature using the weight mask includes the following operations: First, spatial screening is performed: the structure enhancement features from the structure learning module, the detail enhancement features from the detail completion module, and the weight mask from the mask generation unit are used as input; the weight mask includes the weights of the structure features and the weights of the detail features; Spatial screening is achieved through point-by-point multiplication, including point-by-point multiplication of structure enhancement features and structure feature weights, and point-by-point multiplication of detail enhancement features and detail feature weights; Step 5: Input the fused features into a super-resolution decoder to reconstruct a high-resolution face image.
Citation Information
Patent Citations
Progressive face super-resolution calculation method based on multi-stage enhancement framework
CN117689547A
Super-resolution image reconstruction method, system and device based on dynamic frequency domain adaptive coding and contrast constraint optimization, and medium
CN119863364A