Digital human image detail reconstruction method and system based on super-resolution algorithm

CN122175787APending Publication Date: 2026-06-09GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GIANT MOBILE TECH CO LTD
Filing Date
2026-03-09
Publication Date
2026-06-09

Smart Images

  • Figure CN122175787A_ABST
    Figure CN122175787A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image reconstruction, in particular to a digital human image detail reconstruction method and system based on a super-resolution algorithm. The method comprises the following steps: S1: inputting a low-quality face image into a U-Net deep learning model, removing degradation in the image through a 7-time down-sampling and 7-time up-sampling process of the U-Net deep learning model; S2: combining image features of the input image with a pre-trained face generation model of StyleGAN2 to generate an image containing high-quality details; and S3: merging the image containing high-quality details and spatial information features of the input image to obtain a digital human image after detail reconstruction. Thus, a face repair scheme with natural effect, high restoration degree and efficient use is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image reconstruction technology, and in particular to a method and system for reconstructing details of digital human images based on a super-resolution algorithm. Background Technology

[0002] The existing technology has the following drawbacks:

[0003] (1) Relying on low-quality input to extract key information results in poor reliability:

[0004] Some technologies require extracting key information such as the position and outline of facial features from blurry or damaged input images to assist in restoration. However, low-quality images themselves cannot provide accurate and complete information, which leads to problems such as missing details, deformed facial features, and distorted outlines in the restored images, and even produces obvious restoration traces, making it impossible to restore a natural facial state.

[0005] (2) Lack of rich detailed references, resulting in a limited range of repair effects:

[0006] Existing technologies rely on clear reference images or pre-built facial feature libraries to supplement the repair details. However, in practical use, it is difficult to find clear reference images that match the identity and pose of the face in the image to be repaired. At the same time, the content of the library is limited, and the coverage of facial features, texture styles, hairstyles, etc. is not comprehensive enough. This results in the repaired facial details (such as eyebrows, eyelashes, skin texture, etc.) appearing rigid and monotonous, lacking the diversity and uniqueness of real faces, and even resulting in a situation where "a thousand people look the same".

[0007] (3) The repair process is cumbersome, time-consuming, and inefficient:

[0008] Some technologies require repeated parameter adjustments and iterative optimizations to achieve repair results. This is not only complex to operate, but also consumes a lot of time and computing resources, making it impossible to achieve rapid repair and difficult to adapt to scenarios that require immediate processing (such as real-time optimization of images captured by mobile phones and rapid clarification of surveillance footage).

[0009] (4) Lacks color optimization function, limiting its adaptability to various scenarios:

[0010] Many existing technologies can only restore facial details and cannot simultaneously handle color-related issues. For black and white old photos, faded and yellowed images, additional tools are needed for coloring and brightening, which is cumbersome. Furthermore, some technologies are prone to color deviation and unnatural colors when processing colors, making it difficult to restore realistic and vivid color effects.

[0011] Therefore, it is necessary to provide a method and system for digital human image detail reconstruction based on super-resolution algorithms, thereby providing a face restoration solution that is natural in effect, highly accurate in restoration, and efficient in use. Summary of the Invention

[0012] The purpose of this invention is to provide a method and system for reconstructing details of digital human images based on super-resolution algorithms, thereby providing a face restoration solution that is natural in effect, highly accurate in reproduction, and efficient in use.

[0013] To address the problems existing in the prior art, this invention provides a method for detail reconstruction of digital human images based on a super-resolution algorithm, comprising the following steps:

[0014] S1: Input the low-quality facial image into the U-Net deep learning model, and remove the degradation in the image through the process of 7 downsampling and 7 upsampling by the U-Net deep learning model;

[0015] S2: Combine the input image features with the StyleGAN2 pre-trained face generation model to generate images containing high-quality details;

[0016] S3: Merge the spatial information features of the image containing high-quality details and the input image to obtain a digital human image after detail reconstruction.

[0017] Optionally, in the digital human image detail reconstruction method based on super-resolution algorithm, after removing degradation in the image through S1, the results are latent features and multi-resolution spatial features.

[0018] Optionally, in the digital human image detail reconstruction method based on the super-resolution algorithm, the pre-trained face generation model of StyleGAN2 can generate a wide variety of facial details, including the geometric structure, texture and color of the face.

[0019] Optionally, in the digital human image detail reconstruction method based on super-resolution algorithm, S1 uses a Latent code mapping mechanism to map degraded image features into the space of the pre-trained face generation model of StyleGAN2.

[0020] Optionally, in the digital human image detail reconstruction method based on super-resolution algorithm, the image is gradually repaired at different resolutions during the repair process.

[0021] Optionally, in the digital human image detail reconstruction method based on super-resolution algorithm, the difference between the repaired image and the real image is compared by reconstructing loss to ensure that the repaired image is close to the original image in terms of pixels and perception.

[0022] Optionally, in the digital human image detail reconstruction method based on super-resolution algorithm, adversarial loss is used to determine the authenticity of the restored image.

[0023] This invention also provides a digital human image detail reconstruction system based on a super-resolution algorithm. The digital human image detail reconstruction system constructed using the method described above includes:

[0024] The degradation removal module is configured to input low-quality facial images into the U-Net deep learning model and remove degradation in the image through the U-Net deep learning model's process of 7 downsampling and 7 upsampling.

[0025] The generative face prior module is configured to combine the input image features with the StyleGAN2 pre-trained face generation model to generate images containing high-quality details;

[0026] The fusion module is configured to merge the spatial information features of the image containing high-quality details and the input image to obtain a digital human image with reconstructed details.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] (1) This invention proposes a novel repair approach, the core of which is to utilize the rich facial information contained in the pre-trained face generation model. This information covers the facial features, skin texture, hair style and natural color matching of different faces, eliminating the need to extract unreliable information from blurry images.

[0029] (2) Through innovative network structure design, this invention cleverly combines the rich facial information with the only effective information remaining in the blurred image, which not only preserves the natural and realistic facial details provided by the generation model, but also does not lose the unique features of the face in the original image.

[0030] (3) This invention can also complete facial detail restoration and color optimization in one process, such as colorizing old black and white photos and restoring vibrant colors to faded photos, without the need for multiple operations, which greatly improves the restoration efficiency. Through specialized optimization design, it ensures that the restored image not only looks clear and natural, conforming to people's perception of real faces, but also accurately restores the identity features of the face in the original image, without "face swapping" distortion, and at the same time reduces unnecessary traces that may be generated during the restoration process.

[0031] (4) The present invention can also adapt to various complex blurry situations. Whether it is blurry due to a single cause or low-quality images with multiple problems superimposed, it can be effectively processed. It is applicable to various practical scenarios such as old photo restoration, mobile phone image optimization, and monitoring screen clarification, providing people with more practical and higher quality blind face restoration technology support. Attached Figure Description

[0032] Figure 1 This is a flowchart illustrating the digital human image detail reconstruction method provided in an embodiment of the present invention. Detailed Implementation

[0033] The specific embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. The advantages and features of the present invention will become clearer from the following description. It should be noted that the drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the present invention.

[0034] In the following, if the methods described herein include a series of steps, the order of these steps presented herein is not necessarily the only order in which these steps can be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0035] To address the problems existing in the prior art, this invention provides a method for detail reconstruction of digital human images based on a super-resolution algorithm, such as... Figure 1 As shown, it includes the following steps:

[0036] S1: Input the low-quality facial image into the U-Net deep learning model. The U-Net deep learning model performs 7 downsampling and 7 upsampling processes to remove degradations (such as blurring and noise) from the image. After removing degradations, the image yields Latent features (which are used to map to a pre-trained generative model to help generate more realistic details) and multi-resolution spatial features (which are used for spatial adjustment in the subsequent restoration process).

[0037] Furthermore, U-net is a deep learning model for image semantic segmentation that combines a convolutional neural network (CNN) and an encoder-decoder structure, enabling it to effectively handle image segmentation tasks. During the training of the U-net model, the cross-entropy loss function is widely used to help the network learn the correct pixel classification.

[0038] S2: Combine the input image features with the StyleGAN2 pre-trained face generation model to generate images containing high-quality details;

[0039] StyleGAN2's pre-trained facial generation model can generate a wide variety of facial details, including facial geometry, texture, and color.

[0040] S3: Merge the spatial information features of the image containing high-quality details and the input image to obtain a digital human image after detail reconstruction, ensuring that the restored image is both realistic and accurate.

[0041] The technology used in this invention is as follows:

[0042] S1 uses a Latent code mapping mechanism to map degraded image features into the space of the pre-trained face generation model of StyleGAN2, preserving the semantic information of the face and thus avoiding the loss of details.

[0043] During the restoration process, the image is gradually restored at different resolutions, from small to large, to ensure that the overall structure and details in the image are well restored.

[0044] The loss function is designed as follows:

[0045] By comparing the differences between the restored image and the real image through reconstruction loss, the restored image is ensured to be close to the original image in terms of pixels and perception.

[0046] Adversarial loss is used to determine the realism of the restored image. Specifically, a discriminator (whose task is to determine whether the restored image looks like a real facial photograph. If the discriminator thinks the restored image does not look realistic enough, it will give feedback to tell the model where it did not do well) ensures that the restored image looks like a real facial image.

[0047] This invention also provides a digital human image detail reconstruction system based on a super-resolution algorithm. The digital human image detail reconstruction system constructed using the method described above includes:

[0048] The degradation removal module is configured to input low-quality facial images into the U-Net deep learning model. Through the U-Net deep learning model's 7 downsampling and 7 upsampling processes, degradations (such as blurring and noise) in the image are removed. After removing degradations from the image, the resulting features are Latent features (which are used to map to a pre-trained generative model to help generate more realistic details) and multi-resolution spatial features (which are used for spatial adjustment in the subsequent restoration process).

[0049] Furthermore, U-net is a deep learning model for image semantic segmentation that combines a convolutional neural network (CNN) and an encoder-decoder structure, enabling it to effectively handle image segmentation tasks. During the training of the U-net model, the cross-entropy loss function is widely used to help the network learn the correct pixel classification.

[0050] The generative face prior module is configured to combine the input image features with the StyleGAN2 pre-trained face generation model to generate images containing high-quality details;

[0051] StyleGAN2's pre-trained facial generation model can generate a wide variety of facial details, including facial geometry, texture, and color.

[0052] The fusion module is configured to merge the spatial information features of the image containing high-quality details and the input image to obtain a digital human image after detail reconstruction, ensuring that the restored image is both realistic and accurate.

[0053] This invention presents a digital human image detail reconstruction method based on super-resolution algorithms, which can significantly improve the detail and quality of low-quality facial images, especially excelling in removing degradation issues such as blur, noise, and compression distortion. By employing deep learning techniques, particularly Generative Adversarial Networks (GANs) and multi-level restoration mechanisms, this invention can effectively restore various details in facial images, including the fine features of key areas such as skin texture, hair, eyes, and mouth, making the restored image appear more natural, realistic, and layered. The restoration effect is particularly significant for low-resolution, blurry, compressed, or heavily noisy images, restoring the realism and delicacy of the image while maintaining the consistency of facial features. To ensure high precision during the restoration process, this invention designs a multi-resolution restoration strategy, ensuring the gradual restoration of facial image details from coarse to fine, avoiding the detail loss or distortion problems common in traditional methods when enlarging images. Simultaneously, the generative model combines features from the input image, enabling the restored image to perfectly reproduce the natural features of the face while preserving the original facial structure and texture, without obvious restoration traces. More importantly, the identity feature preservation mechanism during the restoration process ensures that the restored face remains identical to the person in the original image, preventing identity errors or changes in facial features due to restoration, thus ensuring high reliability and credibility of the image. The method of this invention also possesses strong adaptability, capable of handling various types of image degradation, ensuring ideal restoration results regardless of the environment. Furthermore, for various complex degradation types that may exist in images, the training model of this invention is specifically optimized for adaptability to mixed degradation scenarios, capable of handling different types of image damage, such as Gaussian blur, low resolution, JPEG compression, and noise. By designing a multi-objective joint loss function, this invention not only ensures pixel accuracy during image reconstruction but also maintains perceptual consistency, making the restored image not only more visually natural but also achieving optimal results in terms of detail realism and identity consistency. Overall, the facial image restoration technology of this invention has significant advantages in improving image quality, detail restoration, realism, and identity consistency, and can be widely applied in fields such as face recognition, digital portrait reconstruction, and virtual avatar generation, bringing more efficient, accurate, and realistic image restoration solutions to related industries.

[0054] The application scenarios of this invention are as follows:

[0055] (1) Virtual Reality (VR) and Augmented Reality (AR): In virtual reality (VR) and augmented reality (AR) technologies, the image quality of digital humans is a key factor in enhancing user immersion and interactive experience. Using the super-resolution algorithm in this invention, the details of low-resolution virtual characters can be significantly improved, making them more vivid and realistic, especially in highly interactive VR / AR environments. For example, in scenarios such as virtual meetings, online education, and virtual tourism, the quality of interaction between users and digital humans directly affects the experience. Through the technology of this invention, users can not only see clearer and more realistic digital human images, but also feel more natural facial expressions and body movements, which is crucial for enhancing interactive experience and improving the sense of immersion in virtual scenes.

[0056] (2) Digital Humans (Virtual Idols) and AI Anchors: Digital humans and virtual idols have become an important part of modern entertainment, live streaming, and social media. Maintaining the detail and clarity of digital human images is particularly crucial in the production of AI anchors and virtual idols. Utilizing the super-resolution algorithm of this invention, the detail of virtual character images can be further enhanced on the basis of low-resolution or early rendering, making facial expressions, skin texture, and eye luster of virtual idols clearer, increasing intimacy with the audience. This technology can significantly improve the realism of virtual anchors in live streaming, providing viewers with a higher quality audiovisual experience, and even providing optimal effects on devices with different resolutions, ensuring smooth display of content on various platforms.

[0057] (3) Video Surveillance and Security Monitoring: In the security field, video surveillance systems are a key technology for ensuring public safety and protecting private property. In some high-risk areas or important locations, video surveillance cameras often need to shoot from a distance, and the image quality is limited by the camera resolution and ambient light. The super-resolution algorithm in this invention can improve the clarity of low-resolution surveillance images, especially images taken at night or in adverse environments, thereby helping security personnel to better identify facial features, clothing colors, or other details in the images. This technology can be widely applied in public safety, traffic monitoring, airport security checks, and other fields to improve recognition accuracy and help police and security personnel react quickly.

[0058] (4) Education and Training: The application of digital avatars in education and training is increasing. Whether in online classrooms, virtual laboratories, or corporate training courses, high-quality digital lecturers or virtual mentors are needed. The super-resolution algorithm of this invention can enhance the expressiveness of virtual avatars in the teaching process, making them more realistic and clear. For example, in a virtual laboratory, a digital lab instructor can demonstrate operating steps and answer students' questions with high-definition images, helping students better understand complex experimental processes. In corporate training, high-definition digital mentors can provide a more intuitive teaching experience, enhance employee learning enthusiasm, and improve training efficiency.

[0059] Compared with the prior art, the present invention has the following advantages:

[0060] (1) This invention proposes a novel repair approach, the core of which is to utilize the rich facial information contained in the pre-trained face generation model. This information covers the facial features, skin texture, hair style and natural color matching of different faces, eliminating the need to extract unreliable information from blurry images.

[0061] (2) Through innovative network structure design, this invention cleverly combines the rich facial information with the only effective information remaining in the blurred image, which not only preserves the natural and realistic facial details provided by the generation model, but also does not lose the unique features of the face in the original image.

[0062] (3) This invention can also complete facial detail restoration and color optimization in one process, such as colorizing old black and white photos and restoring vibrant colors to faded photos, without the need for multiple operations, which greatly improves the restoration efficiency. Through specialized optimization design, it ensures that the restored image not only looks clear and natural, conforming to people's perception of real faces, but also accurately restores the identity features of the face in the original image, without "face swapping" distortion, and at the same time reduces unnecessary traces that may be generated during the restoration process.

[0063] (4) The present invention can also adapt to various complex blurry situations. Whether it is blurry due to a single cause or low-quality images with multiple problems superimposed, it can be effectively processed. It is applicable to various practical scenarios such as old photo restoration, mobile phone image optimization, and monitoring screen clarification, providing people with more practical and higher quality blind face restoration technology support.

[0064] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the technical solutions and content disclosed in the present invention without departing from the scope of the present invention shall be deemed to have remained within the protection scope of the present invention.

Claims

1. A method for detail reconstruction of digital human images based on super-resolution algorithm, characterized in that, Includes the following steps: S1: Input the low-quality facial image into the U-Net deep learning model, and remove the degradation in the image through the process of 7 downsampling and 7 upsampling by the U-Net deep learning model; S2: Combine the input image features with the StyleGAN2 pre-trained face generation model to generate images containing high-quality details; S3: Merge the spatial information features of the image containing high-quality details and the input image to obtain a digital human image after detail reconstruction.

2. The method for reconstructing digital human image details based on super-resolution algorithm as described in claim 1, characterized in that, After removing degradation in the image using S1, we obtain Latent features and multi-resolution spatial features.

3. The method for reconstructing digital human image details based on super-resolution algorithm as described in claim 1, characterized in that, StyleGAN2's pre-trained facial generation model can generate a wide variety of facial details, including facial geometry, texture, and color.

4. The method for detail reconstruction of digital human images based on super-resolution algorithm as described in claim 1, characterized in that, In S1, a Latent code mapping mechanism is used to map degraded image features into the space of the pre-trained face generation model of StyleGAN2.

5. The method for reconstructing digital human image details based on super-resolution algorithm as described in claim 1, characterized in that, During the restoration process, the image is gradually restored at different resolutions.

6. The method for detail reconstruction of digital human images based on super-resolution algorithm as described in claim 1, characterized in that, By comparing the differences between the restored image and the real image through reconstruction loss, the restored image is ensured to be close to the original image in terms of pixels and perception.

7. The method for reconstructing digital human image details based on super-resolution algorithm as described in claim 1, characterized in that, The authenticity of the restored image is determined by adversarial loss.

8. A digital human image detail reconstruction system based on super-resolution algorithm, characterized in that, A digital human image detail reconstruction system is constructed using the method described in any one of claims 1-7, comprising: The degradation removal module is configured to input low-quality facial images into the U-Net deep learning model and remove degradation in the image through the U-Net deep learning model's process of 7 downsampling and 7 upsampling. The generative face prior module is configured to combine the input image features with the StyleGAN2 pre-trained face generation model to generate images containing high-quality details; The fusion module is configured to merge the spatial information features of the image containing high-quality details and the input image to obtain a digital human image with reconstructed details.