The present invention discloses a face super-resolution method and
system based on a dual guided
diffusion model. The method comprises the following steps: collecting low-resolution non-frontal face images and high-resolution frontal face images to construct training data pairs; preliminarily restoring the low-resolution non-frontal face images to obtain rough frontal face images; mapping the face images in pixel space to implicit space, enabling a
diffusion model to be calculated in the implicit space, pre-training the unconditional
diffusion model, using the training results as initialization parameters of the diffusion model, and freezing the
encoder of a denoising network; extracting facial prior features from the rough frontal face images, and capturing the spatial and semantic correlations between the facial prior features and denoising features through a
hybrid cross-attention mechanism; extracting
facial identity coding information and embedding it into a denoising network; initializing a
Gaussian noise map, iteratively denoising using the trained diffusion model, mapping the denoising results in the implicit space to the pixel space, and finally reconstructing a high-resolution frontal face image.