Self-adaptive learning method and system for low-illumination face super-resolution

Through the alternating illumination-diffusion adaptive mechanism and Fourier enhancement module of the DiffLLFace framework, the compound degradation problem of low spatial resolution face images under low illumination is solved, high-quality super-resolution reconstruction is achieved, and the visual fidelity and detail recovery ability of the image are improved.

CN120689204AActive Publication Date: 2025-09-23SHANDONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510692127.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-23
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing methods find it difficult to effectively solve the compound degradation problem of low spatial resolution face images under low light conditions, resulting in poor reconstruction quality, especially in terms of brightness restoration and spatial resolution enhancement.

Method used

Using the DiffLLFace framework, through the alternating illumination-diffusion adaptive mechanism and the non-parametric Fourier enhancement module, combined with the generation capability of the diffusion model and the illumination perception trajectory modeling, super-resolution reconstruction of low-light and low-spatial resolution facial images is achieved.

Benefits of technology

It significantly improves the image reconstruction quality under low-light conditions, maintains texture details and color consistency, and demonstrates excellent generalization ability and visual fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689204A_ABST
    Figure CN120689204A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and discloses a low-illumination face super-resolution-oriented adaptive learning method and system, and the method comprises the steps: extracting the illumination features of an image, and obtaining illumination priori based on the illumination features; operating the potential features of the stable diffusion model by using illumination prior to obtain first adaptive features; fourier transform is carried out on the low-illumination low-resolution image, an illumination enhancement image is generated, feature extraction is carried out, a second adaptive feature is obtained by using the illumination enhancement image feature and the first adaptive feature, the illumination feature is corrected by using the second adaptive feature, and a corrected illumination feature is obtained; and splicing the corrected illumination features and the illumination enhanced image features, and operating the spliced features to generate a high-definition image with normal illumination. According to the method, the advantages of a diffusion model are fused, a unified architecture is designed to realize super-resolution reconstruction and illumination enhancement of a low-illumination low-spatial-resolution face image, and double challenges of illumination enhancement and resolution improvement are synchronously solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an adaptive learning method and system for low-light face super-resolution. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Face super-resolution (FSR), a classic low-level computer vision task, aims to improve the perceptual clarity and fidelity of low-spatial-resolution (LR) facial images, reconstructing photorealistic high-resolution (HR) images. In low-light or nighttime environments, facial images are not only limited by inherent spatial resolution loss, but also further degraded by compound degradation issues such as reduced visibility, contrast attenuation, and color distortion. This dual degradation severely compromises the structural integrity and texture authenticity of the image, significantly limiting its application in practical vision systems such as surveillance, security, and biometrics. This compound degradation poses a significant challenge to reconstruction algorithms, requiring simultaneous brightness enhancement and spatial resolution enhancement while ensuring high-fidelity reconstruction.

[0004] Current mainstream approaches to addressing the combined low-light and low-spatial-resolution degradation problem exhibit a polarized research landscape. On the one hand, low-light image enhancement (LLIE) aims to restore normal-light images from low-light inputs by modeling the mapping relationship between color, brightness, and contrast. While the LLIE framework achieves significant results in brightness restoration, it often overlooks the inherent spatial degradation characteristics of LR images, resulting in poor detail fidelity in the enhanced results. On the other hand, FSR technology has undergone significant development, gradually expanding from addressing ideal downsampling problems to modeling complex, unknown degradations such as noise, blur, and compression artifacts. However, existing FSR methods perform well under good lighting conditions but struggle to effectively address the unique challenges of low-light scenarios. This suggests that treating LLIE and FSR as independent tasks has significant limitations and cannot simultaneously address the dual degradation problem.

[0005] To address the above dilemma, an intuitive solution is to process the two tasks serially in a cascade manner (i.e., LLIE→FSR or FSR→LLIE). However, the essential complexity difference between single degradation and compound degradation seriously restricts the effectiveness of such methods. The first-stage processing (whether LLIE or FSR) cannot effectively decouple the intertwined degradation modes, and places excessively high demands on the quality of the results. This processing paradigm inevitably leads to the risk of error propagation, such as artifact amplification or incorrect structural priors, which in turn mislead the execution of subsequent tasks and ultimately lead to degraded reconstruction quality. The above shortcomings highlight the urgent need to establish a unified framework to jointly model low-resolution and low-light scenes. Although a small number of exploratory works have recently attempted to study the FSR problem under low-light conditions, they still have obvious shortcomings in color consistency and texture fidelity.

[0006] Existing methods primarily target degraded scenes under normal lighting conditions and lack effective modeling for complex degradation issues unique to low-light environments, such as brightness attenuation and noise interference. Researchers have proposed the IC-FSRDENet framework, which decouples the task into structure recovery and detail enhancement subtasks. However, its step-by-step processing framework can easily lead to suboptimal solutions. Summary of the Invention

[0007] To address the above problems, the present invention proposes an adaptive learning method and system for low-light face super-resolution. By integrating the advantages of the diffusion model, a unified architecture is designed to achieve super-resolution reconstruction and illumination enhancement of low-light and low-spatial resolution facial images, thereby simultaneously addressing the dual challenges of illumination enhancement and resolution improvement.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides an adaptive learning method for low-light face super-resolution, comprising the following steps: Extract the illumination features and image features of low-light and low-resolution images, use the image features as guidance information, generate illumination coefficients based on the illumination features, and obtain illumination priors based on the illumination coefficients; The potential features in the stable diffusion model are input into the UNet encoder for downsampling to obtain the intermediate layer features. The intermediate layer features are adjusted pixel by pixel using the illumination prior to obtain the first adaptive features. Performing Fourier transform on the low-light and low-resolution image to generate a light-enhanced image, extracting features of the light-enhanced image, and modulating the first adaptive feature using the features of the light-enhanced image to obtain a second adaptive feature; The second adaptive feature is used to correct the illumination feature to obtain the corrected illumination feature, the corrected illumination feature and the illumination-enhanced image feature are spliced ​​together, the spliced ​​feature is input into the control network, the feature generated by the control network is input into the Unet decoder in the stable diffusion for upsampling, the upsampled latent representation is decoded, and a high-definition face image with normal illumination and realistic texture is generated.

[0009] As an optional implementation, image features of low-light and low-resolution images are extracted, specifically: The bicubic interpolation method is used to interpolate and enlarge the low-light low-resolution image, and the features of the interpolated and enlarged image are extracted as the image features of the low-light low-resolution image.

[0010] As an optional implementation, the illumination coefficient is generated based on the illumination characteristics, specifically: Based on the illumination features, the bilateral grid algorithm is used to construct the illumination coefficient grid. The image features of low-light and low-resolution images are used as guiding information. The illumination coefficient grid is projected into the three-dimensional grid space, and then Gaussian blur is used to smooth the three-dimensional grid. Finally, the illumination coefficient is obtained through three-dimensional slicing operation.

[0011] As an optional implementation, Fourier transform is performed on the low-light, low-resolution image to generate a light-enhanced image, specifically: A Fourier transform is performed on the low-light and low-resolution image to generate amplitude and phase components. A scaling factor is introduced to amplify the amplitude, and the amplified amplitude and phase are combined to generate a light-enhanced image through inverse fast Fourier transform.

[0012] As an optional implementation, the second adaptive feature is used to correct the illumination feature, specifically: The illumination features and the second adaptive features are processed respectively through three linear layers to generate the corresponding query vector, key vector and value vector, calculate the cross attention weight, generate the attention map, and input the attention map into the feedforward network to generate the corrected illumination features.

[0013] As an optional implementation, it also includes defining a loss function to optimize the trainable parameters in the adaptive learning method, where the loss function jointly optimizes the latent diffusion model denoising objective and the latent space reconstruction objective.

[0014] In a second aspect, the present invention provides an adaptive learning system for low-light face super-resolution, comprising: an illumination prior calculation module configured to: extract illumination features and image features of a low-light, low-resolution image, use the image features as guidance information, generate illumination coefficients based on the illumination features, and obtain illumination priors based on the illumination coefficients; The first adaptive feature calculation module is configured to: input the latent features in the stable diffusion model into the UNet encoder for downsampling to obtain intermediate layer features, and perform pixel-by-pixel curve adjustment on the intermediate layer features using the illumination prior to obtain the first adaptive features; The second adaptive feature calculation module is configured to: perform Fourier transform on the low-light and low-resolution image to generate a light-enhanced image, extract features of the light-enhanced image, and modulate the first adaptive feature using the features of the light-enhanced image to obtain a second adaptive feature; The image output module is configured to: use the second adaptive feature to correct the illumination feature to obtain the corrected illumination feature, splice the corrected illumination feature and the illumination-enhanced image feature, input the spliced ​​feature into the control network, input the feature generated by the control network into the Unet decoder in the stable diffusion for upsampling, decode the upsampled latent representation, and generate a high-definition face image with normal illumination and realistic texture.

[0015] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.

[0017] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which implements the method described in the first aspect when executed by a processor.

[0018] Compared with the prior art, the present invention has the following beneficial effects: This paper proposes an adaptive learning method and system for low-light face super-resolution, and pioneers the DiffLLFace collaborative optimization framework. By fusing the generation capability of the diffusion model with the illumination perception trajectory modeling, robust face super-resolution reconstruction is achieved in low-light environments. An alternating illumination-diffusion adaptive mechanism is proposed. By dynamically coordinating the interaction process between the latent representation and illumination information, bidirectional collaborative enhancement of feature expression is achieved along the diffusion trajectory, effectively enhancing the controllability of the generation process. A non-parametric Fourier enhancement module is designed. By fusing the frequency domain structure representation and the alternating adaptive mechanism, the dual constraints of texture detail enhancement and color consistency maintenance are constructed, significantly improving visual fidelity. Verified in multiple scenarios, DiffLLFace demonstrates excellent generalization capabilities in complex natural scenes, and maintains excellent reconstruction performance under dynamic low-light conditions.

[0019] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0021] Figure 1 This is the overall architecture diagram of the DiffLLFace network provided in Example 1 of the present invention; Figure 2 This figure shows the comparison results of image processing using the method of the present invention and other cutting-edge methods on the CelebAMask-HQ dataset. DETAILED DESCRIPTION

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0024] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0026] Explanation of terms: Low-light face super-resolution is the task of restoring high-resolution and clear details from blurry, low-resolution face images collected in low-light environments.

[0027] Alternating illumination-diffusion adaptation is the core optimization mechanism, which achieves image enhancement by alternating illumination correction and diffusion model.

[0028] Adaptive learning methods refer to models that can dynamically adjust according to the characteristics of input data or task requirements to improve performance and generalization capabilities.

[0029] Example 1 like Figure 1 As shown, this embodiment provides an adaptive learning method for low-light face super-resolution, including the following steps: Extract the illumination features and image features of low-light and low-resolution images, use the image features as guidance information, generate illumination coefficients based on the illumination features, and obtain illumination priors based on the illumination coefficients; The potential features in the stable diffusion model are input into the UNet encoder for downsampling to obtain the intermediate layer features. The intermediate layer features are adjusted pixel by pixel using the illumination prior to obtain the first adaptive features. Performing Fourier transform on the low-light and low-resolution image to generate a light-enhanced image, extracting features of the light-enhanced image, and modulating the first adaptive feature using the features of the light-enhanced image to obtain a second adaptive feature; The second adaptive feature is used to correct the illumination feature to obtain the corrected illumination feature, the corrected illumination feature and the illumination-enhanced image feature are spliced ​​together, the spliced ​​feature is input into the control network, the feature generated by the control network is input into the Unet decoder in the stable diffusion for upsampling, the upsampled latent representation is decoded, and a high-definition face image with normal illumination and realistic texture is generated.

[0030] This paper proposes a novel diffusion model (DM)-based framework, DiffLLFace, that achieves super-resolution reconstruction and illumination enhancement for low-light, low-spatial-resolution (LLR) face images within a unified architecture. This framework fully exploits the synergistic potential between illumination modeling and controllable generation, leveraging a ControlNet-based framework to achieve a balance between perceptual realism and reconstruction fidelity. Specifically, we propose an alternating illumination-diffusion adaptation mechanism: we first estimate the illumination prior of the LLR image and exploit its degradation information to adjust the latent features. We then design an illumination-aware adaptation module to enhance model robustness. However, due to the lack of structural context, directly applying this prior as a control signal is difficult to achieve ideal results. Therefore, we explicitly leverage the strong representational knowledge embedded in the latent features to correct the illumination estimate and generate a refined control signal to guide the reconstruction process. By alternating feature coordination along the illumination-aware diffusion trajectory, DiffLLFace achieves dynamic co-optimization of illumination and diffusion features. To further enhance the model's ability to capture subtle textures, we design a Fourier enhancement module that requires no trainable parameters, effectively mitigating the degradation of detail caused by low brightness while suppressing noise. By incorporating Fourier transform-based representation into the alternating adaptive process, we ultimately achieve color-consistent and perceptually pleasing reconstruction.

[0031] The DiffLLFace proposed in this paper adopts a unified two-branch processing framework. Its workflow can be summarized as follows: the input low-light, low-resolution image first uses the Fourier enhancement module in the upper branch to extract appearance features, while the alternating illumination-diffusion adaptation module in the lower branch performs fine-tuning. The middle branch uses a control network to constrain the sampling trajectory of the stable diffusion model, and finally outputs a high-fidelity reconstruction result through the SD decoder.

[0032] The specific technical solutions for implementing the present invention are as follows: DiffLLFace Overview: For low-light and low-resolution input images, we propose to achieve joint optimization of low-light enhancement and super-resolution in a unified framework, and ultimately generate high-definition face images with normal lighting and realistic textures. Figure 1 As shown in Figure 2, the DiffLLFace framework, based on a control network architecture, controls the sampling trajectory of a stable diffusion model by learning illumination condition priors. Compared to existing methods that rely solely on unidirectional conditional control based on pre-trained models, this work offers two major innovations: First, through an alternating illumination-diffusion adaptation mechanism, bidirectional feature coordination is achieved between the latent diffusion representation and the illumination prior extracted from the input image, enabling precise control of the denoising process. Second, by combining Fourier-enhanced features with the corrected illumination prior, the reconstruction fidelity is improved by enhancing the representational capabilities of ControlNet.

[0033] During the training phase, only the trainable parameters of the alternating adaptive module and the ControlNet branch are updated to keep the basic knowledge of the pre-trained model intact. Our DiffLLFace adopts the joint optimization of the latent diffusion model denoising objective and the latent space reconstruction objective, which is mathematically expressed as follows:

[0034] in, Represents the total loss function, which is given by and It consists of two parts. is the noise predicted during the denoising process and real noise The difference between is the noise prediction model learned by the denoising U-Net, is the original potential representation, is the cumulative variance at time step t. Reconstruction loss It is the potential representation after denoising and the original latent representation Measure the difference between .

[0035] Alternating Lighting-Diffusion Adaptation Adjustment: Illumination information extracted from low-light, low-resolution images is crucial for modeling degradation patterns. This information can directly serve as a condition to guide the controllable generation process. However, due to insufficient structural information, the problem of color consistency remains unresolved. To this end, we propose an alternating illumination-diffusion adaptation mechanism that synergistically leverages the strengths of illumination conditions and latent features to achieve more precise control. This mechanism consists of three key steps: illumination estimation, illumination-aware adaptation, and illumination correction.

[0036] 1) Lighting Estimation First, we use a shallow network to extract illumination features , based on the bilateral grid algorithm, we constructed the illumination coefficient grid At the same time, high-resolution features are extracted by performing convolution operations on the image enlarged by bicubic interpolation and used as guidance information. ,Will Projected to pixel coordinates The three-dimensional grid space is composed of image density. Gaussian blur is then used to smooth the three-dimensional grid, and finally the illumination coefficient is obtained through three-dimensional slicing operation. :

[0037] in, Represents a slicing operation using trilinear interpolation.

[0038] 2) Light perception adaptation At the tth iteration, based on the estimated lighting prior , we have the latent features output by the previous module ( It is the feature obtained by downsampling the potential feature through the Unet encoder in the stable diffusion process at time step t) and adjusting the pixel-by-pixel curve:

[0039] in, It has been normalized to the interval [0,1] and is a trainable curve parameter obtained by performing convolutional layer and Sigmoid function operations on it. (First adaptive feature) represents the adjusted feature. It should be noted that for clarity, the operations related to the Fourier prior are omitted here and will be explained in the "Combination with Fourier-based Enhancement" section below.

[0040] 3) Lighting correction Due to the lack of image structure information, it is difficult to obtain ideal results by directly using the estimated illumination parameters as control signals. To this end, we propose to use the second adaptive feature (combined with Fourier-based enhancement) to correct for illumination characteristics , thereby enhancing the control accuracy. In the specific implementation, three linear layers are used to process the illumination features respectively. and the second adaptive feature , generate the corresponding query vector , key vector Sum value vector (The key-value vectors all come from the second adaptive feature ), and the cross attention weight is calculated by the following formula:

[0041] in, Represents the attention map, and then generates the corrected lighting features through the feedforward network As the final control signal. Experiments show that after model training, our DiffLLFace can dynamically couple illumination priors and latent features during image denoising, optimizing the information flow of the control network through a bidirectional enhancement mechanism.

[0042] Combined with Fourier-based enhancement: Fourier domain techniques have achieved significant success in both LLIE and SR tasks. On the one hand, by operating in the Fourier domain, the model can capture the global receptive field of the image, overcoming the local limitations of spatial domain methods. On the other hand, brightness information is primarily contained in the amplitude component, while structural information exists in the phase component. This demonstrates that low-light and noise issues can be addressed separately in the Fourier domain.

[0043] Therefore, this work also uses the Fourier-based enhancement method to deal with the composite degradation problem of brightness, structure and details. In the Fourier branch, we perform an LLR image on the input Perform a Fast Fourier Transform (FFT) to generate the amplitude component and phase components :

[0044] The scaling factor is then directly introduced to amplify the amplitude and the result is compared with the phase Combined, light-enhanced image is generated by inverse fast Fourier transform :

[0045] in, This strategy can significantly improve the visibility of the input image and is expected to provide structural appearance information for the model.

[0046] Afterwards, we integrate the Fourier enhancement information into the alternating adaptive module. First, multi-head self-attention (MHSA) and feed-forward network (FFN) are used to capture global dependencies and extract appearance features. Then AdaIN is used to adapt the illumination first feature in the latent space. To modulate:

[0047] in, and represents the affine transformation parameters learned by two independent 1×1 convolutional layers, which are used to adjust . Additionally, in the ControlNet branch, we simply Concatenate with the corrected illumination features along the channel dimension to update the control signal.

[0048] This paper proposes a unified framework, DiffLLFace, which synergistically leverages the generative power of the diffusion model and the illumination perception trajectory to achieve robust super-resolution reconstruction of faces under low-light conditions. Its core innovation lies in the alternating illumination-diffusion adaptation mechanism, which effectively strengthens the control flow by coordinating the bidirectional enhancement of latent representation and illumination information along the diffusion path. In addition, the non-parametric Fourier enhancement module we introduced significantly improves the visibility of the input image, thereby providing the model with reliable structural appearance information. This module works in conjunction with the alternating adaptation mechanism to effectively maintain texture details and color consistency. Experiments show that DiffLLFace not only surpasses the existing best methods on facial images, but also exhibits excellent generalization capabilities for complex natural scenes.

[0049] Figure 2The generated results of various comparison methods are visualized. We selected five cutting-edge face super-resolution methods (SISN, SCTANet, SFMNet, PGDiff, and DR2) and five advanced low-light enhancement methods (FECNet, LLFormer, LEDNet, FourierDiff, and Quadprior). The joint processing method uses IC-FSRDENet, the same task as this study, for low-light face super-resolution. It can be seen that the high-resolution images reconstructed by the cascade method generally suffer from severe color and texture distortion. Even when we use diffusion models with strong generative capabilities in both stages (such as DR2→FourierDiff and FourierDiff→PGDiff), the results are still noticeably blurry. While IC-FSRDENet produces super-resolution images with clearer textures, it still suffers from identity bias and color inconsistency. In contrast, our DiffLLFace demonstrates more natural and accurate detail recovery, and its superior visual quality is more perceptually appealing.

[0050] Figure 2 The first column shows the low-resolution input, columns 2-5 show the results of the "first super-resolution then enhancement" (FSR→LLIE) method, columns 6-9 show the results of the "first enhancement then super-resolution" (LLIE→FSR) method, columns 10-11 show the results of the combined processing method, and columns 12-13 show the real HD images. This method outperforms the current state-of-the-art methods in terms of facial texture fidelity, image detail restoration, and color accuracy.

[0051] Example 2 This embodiment provides an adaptive learning system for low-light face super-resolution, including: an illumination prior calculation module configured to: extract illumination features and image features of a low-light, low-resolution image, use the image features as guidance information, generate illumination coefficients based on the illumination features, and obtain illumination priors based on the illumination coefficients; The first adaptive feature calculation module is configured to: input the latent features in the stable diffusion model into the UNet encoder for downsampling to obtain intermediate layer features, and perform pixel-by-pixel curve adjustment on the intermediate layer features using the illumination prior to obtain the first adaptive features; The second adaptive feature calculation module is configured to: perform Fourier transform on the low-light and low-resolution image to generate a light-enhanced image, extract features of the light-enhanced image, and modulate the first adaptive feature using the features of the light-enhanced image to obtain a second adaptive feature; The image output module is configured to: use the second adaptive feature to correct the illumination feature to obtain the corrected illumination feature, splice the corrected illumination feature and the illumination-enhanced image feature, input the spliced ​​feature into the control network, input the feature generated by the control network into the Unet decoder in the stable diffusion for upsampling, decode the upsampled latent representation, and generate a high-definition face image with normal illumination and realistic texture.

[0052] It should be noted that the above modules correspond to the steps described in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0053] In further embodiments, there is also provided: An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed by the processor, wherein when the computer instructions are executed by the processor, the method described in Example 1 is performed. For the sake of brevity, no further details are given here.

[0054] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0055] The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0056] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method described in Example 1 is performed.

[0057] The method in Example 1 can be directly implemented as a hardware processor, or can be implemented using a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, it will not be described in detail here.

[0058] A computer program product includes a computer program, which implements the method described in embodiment 1 when executed by a processor.

[0059] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions contained in program modules, which are executed in a device on a real or virtual processor of a target to perform the process / method described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided between program modules as needed. The machine-executable instructions for the program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.

[0060] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on a computer, partially on a computer, as an independent software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.

[0061] In the context of the present invention, computer program code or related data can be carried by any appropriate carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, and the like.

[0062] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0063] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.

Claims

1. An adaptive learning method for low-light face super-resolution, characterized by: The following steps are involved: Extract the illumination features and image features of low-light and low-resolution images, use the image features as guidance information, generate illumination coefficients based on the illumination features, and obtain illumination priors based on the illumination coefficients; The potential features in the stable diffusion model are input into the UNet encoder for downsampling to obtain the intermediate layer features. The intermediate layer features are adjusted pixel by pixel using the illumination prior to obtain the first adaptive features. Performing Fourier transform on the low-light and low-resolution image to generate a light-enhanced image, extracting features of the light-enhanced image, and modulating the first adaptive feature using the features of the light-enhanced image to obtain a second adaptive feature; The second adaptive feature is used to correct the illumination feature to obtain the corrected illumination feature, the corrected illumination feature and the illumination-enhanced image feature are spliced ​​together, the spliced ​​feature is input into the control network, the feature generated by the control network is input into the Unet decoder in the stable diffusion for upsampling, the upsampled latent representation is decoded, and a high-definition face image with normal illumination and realistic texture is generated.

2. The adaptive learning method for low-light face super-resolution according to claim 1, characterized in that: Extract image features of low-light and low-resolution images, specifically: The bicubic interpolation method is used to interpolate and enlarge the low-light low-resolution image, and the features of the interpolated and enlarged image are extracted as the image features of the low-light low-resolution image.

3. The adaptive learning method for low-light face super-resolution according to claim 1, characterized in that Generate illumination coefficients based on illumination characteristics, specifically: Based on the illumination features, the bilateral grid algorithm is used to construct the illumination coefficient grid. The image features of low-light and low-resolution images are used as guiding information. The illumination coefficient grid is projected into the three-dimensional grid space, and then Gaussian blur is used to smooth the three-dimensional grid. Finally, the illumination coefficient is obtained through three-dimensional slicing operation.

4. The adaptive learning method for low-light face super-resolution according to claim 1, wherein: Perform Fourier transform on the low-light, low-resolution image to generate a light-enhanced image, specifically: A Fourier transform is performed on the low-light and low-resolution image to generate amplitude and phase components. A scaling factor is introduced to amplify the amplitude, and the amplified amplitude and phase are combined to generate a light-enhanced image through inverse fast Fourier transform.

5. The adaptive learning method for low-light face super-resolution according to claim 1, wherein: The second adaptive feature is used to correct the illumination feature, specifically: The illumination features and the second adaptive features are processed respectively through three linear layers to generate the corresponding query vector, key vector and value vector, calculate the cross attention weight, generate the attention map, and input the attention map into the feedforward network to generate the corrected illumination features.

6. The adaptive learning method for low-light face super-resolution according to claim 1, wherein: It also includes defining a loss function to optimize the trainable parameters in the adaptive learning method. The loss function uses the latent diffusion model denoising objective and the latent space reconstruction objective to jointly optimize.

7. An adaptive learning system for low-light face super-resolution, characterized by: include: an illumination prior calculation module configured to: extract illumination features and image features of a low-light, low-resolution image, use the image features as guidance information, generate illumination coefficients based on the illumination features, and obtain illumination priors based on the illumination coefficients; The first adaptive feature calculation module is configured to: input the latent features in the stable diffusion model into the UNet encoder for downsampling to obtain intermediate layer features, and perform pixel-by-pixel curve adjustment on the intermediate layer features using the illumination prior to obtain the first adaptive features; The second adaptive feature calculation module is configured to: perform Fourier transform on the low-light and low-resolution image to generate a light-enhanced image, extract features of the light-enhanced image, and modulate the first adaptive feature using the features of the light-enhanced image to obtain a second adaptive feature; The image output module is configured to: use the second adaptive feature to correct the illumination feature to obtain the corrected illumination feature, splice the corrected illumination feature and the illumination-enhanced image feature, input the spliced ​​feature into the control network, input the feature generated by the control network into the Unet decoder in the stable diffusion for upsampling, decode the upsampled latent representation, and generate a high-definition face image with normal illumination and realistic texture.

8. An electronic device, characterized in that: The method comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.

9. A computer-readable storage medium, characterized in that Used to store computer instructions, which, when executed by a processor, complete the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer program, which is used to implement the method according to any one of claims 1 to 6 when the computer program is executed by a processor.

Citation Information

Patent Citations

  • Face image super-resolution reconstruction method under low illumination condition

    CN117830096A

  • Low-illumination image enhancement method based on curve wavelet attention and Fourier

    CN118822908A

  • Deferred neural lighting in augmented image generation

    US20240386656A1