Portrait super-resolution processing method and system combining local repair and global optimization

By combining local repair and global optimization portrait super-resolution processing methods, using multi-step dice denoising and local redrawing technology, the problem of difficult balance between global structure and local details in image processing is solved, and efficient and accurate image repair effects are achieved, especially in key areas such as the face and hands of the characters, which improves the visual quality and reality of the image.

CN120374380APending Publication Date: 2025-07-25SHANGHAI YIQIUCHUN CULTURE MEDIA CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510411242.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When processing character images, existing image super-resolution methods are difficult to effectively balance the global structure and local details, resulting in the loss or excessive modification of the details in key areas, affecting the clarity and naturalness of the image.

Method used

Combining the portrait super-resolution processing method with local repair and global optimization, multi-step dice denoising and local redrawing are performed to accurately repair key areas through super-resolution model, semantic guidance generation model, diffusion model, block diffusion model, image understanding model and image segmentation model.

Benefits of technology

It significantly improves the detail accuracy and quality of the image, improves visual effects, avoids striped artifacts and waste of computing resources, especially in key areas such as the face and hands of the figure, to meet the needs of high-quality image optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374380A_ABST
    Figure CN120374380A_ABST
Patent Text Reader

Abstract

The invention discloses a local repair and global optimization combined portrait super-resolution processing method and system, which are used in the technical field of image enhancement, and the method comprises the steps: carrying out the preprocessing of an input image, and obtaining the processed input image and the semantic guidance information of the input image; based on the semantic guidance information of the input image, redrawing the processed input image by using a diffusion model to obtain an initial redrawn image; performing local redrawing on the initial redrawing image according to a block diffusion model to obtain a local redrawing image, and integrating the local redrawing image to obtain a secondary redrawing image; and according to the secondary redrawing image, carrying out target region identification and segmentation by using an image understanding model and an image segmentation model, and repairing the target region by using a weak redrawing technology to obtain a super-resolution image. The detail precision and quality of the image are remarkably improved, the structural integrity of the image can be kept, and the visual effect and reality sense of the image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image enhancement. Specifically, it particularly relates to a method and system for portrait super-resolution processing that combines local repair and global optimization. Background Art

[0002] With the development of computer vision and artificial intelligence technologies, image super-resolution and optimization processing have become important research directions in the field of image enhancement. Super-resolution technology can improve image quality and enhance image details by recovering higher-resolution images from low-resolution images. However, traditional image super-resolution methods often have certain limitations. Especially when dealing with complex scenes and images rich in details, many challenges still remain.

[0003] Existing image optimization technologies usually rely on global image redrawing based on diffusion models. The core idea of such methods is to perform a one-time optimization process on the entire image, mostly enhancing image details through multiple steps of iteration. However, when dealing with large-sized images, diffusion models are prone to generating artifacts in local areas, especially in parts with complex details, such as key areas like the face and hands of a person. These artifacts not only affect the clarity of the image but may also cause distortion of the overall structure of the image.

[0004] In addition, when dealing with human images, existing technologies usually adopt a global optimization method to uniformly process the entire image. Since global optimization does not perform independent cropping and optimization on each part of the person, especially when the proportion of the person in the picture is less than 25%, details such as the face and hands often suffer from detail loss or breakdown during the optimization process. Traditional methods lack pertinence in repairing key parts and are difficult to effectively avoid the loss of human details. Especially during the repair process, it is easy to over-modify non-target areas, thus affecting the overall naturalness of the image.

[0005] Existing technologies have not yet effectively combined the advantages of local repair and global optimization. There are problems in the processing process such as being unable to accurately distinguish key areas from background areas, resulting in difficulties in balancing local details and global structure. Therefore, existing image optimization methods have significant deficiencies in terms of detail accuracy, repair effect, and processing efficiency. Especially when performing delicate repair on human images, they cannot meet the requirements of high-quality image optimization.

[0006] Regarding the problems in the related technologies, no effective solutions have been proposed yet. Summary of the Invention

[0007] In order to overcome the above problems, the present invention aims to propose a method and system for portrait super-resolution processing that combines local repair and global optimization, aiming to solve the problem of difficultly effectively balancing the optimization of global structure and local details when processing images.

[0008] To this end, the specific technical solution adopted by the present invention is as follows:

[0009] According to one aspect of the present invention, there is provided a portrait super-resolution processing method combining local repair and global optimization, the method comprising the following steps:

[0010] S1. Using a super-resolution model and a semantic guidance generation model, preprocess the input image to obtain a processed input image and semantic guidance information of the input image;

[0011] S2. Based on the semantic guidance information of the input image, use a diffusion model to redraw the processed input image to obtain an initial redrawn image;

[0012] S3. According to the block diffusion model, perform local redrawing on the initial redrawn image to obtain a locally redrawn image, and obtain a second redrawn image by integrating the locally redrawn images;

[0013] S4. According to the second redrawn image, use an image understanding model and an image segmentation model to perform target area recognition and segmentation, and use weak redrawing technology to repair the target area to obtain a super-resolution image.

[0014] Optionally, using a super-resolution model and a semantic guidance generation model to preprocess the input image to obtain a processed input image and semantic guidance information of the input image includes the following steps:

[0015] S11. Based on the super-resolution model, use super-resolution technology to magnify the input image to obtain a processed input image;

[0016] S12. For the processed input image, extract hierarchical features in the image through a convolutional neural network in the semantic guidance generation model;

[0017] S13. Based on the hierarchical features, through integrated analysis of the processed input image, identify visual information and convert it into structured text prompt words to obtain semantic guidance information of the input image.

[0018] Optionally, based on the semantic guidance information of the input image, using a diffusion model to redraw the processed input image to obtain an initial redrawn image includes the following steps:

[0019] S21. Input the processed input image into the diffusion model, and use the diffusion model to add noise to the processed input image to obtain a noisy image;

[0020] S22. According to the semantic guidance information of the input image, use the redrawing intensity controller in the diffusion model to perform multiple denoising redraws on the noisy image to obtain a preliminary redrawn image.

[0021] Optionally, according to the block diffusion model, perform local redrawing on the initial redrawn image to obtain a locally redrawn image. The steps for obtaining the second redrawn image by integrating the locally redrawn image are as follows:

[0022] S31. Enlarge the initial redrawn image and cut the enlarged initial redrawn image into a number of uniform sub - blocks;

[0023] S32. Use the local redrawing processing module in the block diffusion model to perform local detail enhancement on the sub - blocks to obtain a locally redrawn image;

[0024] S33. Integrate the locally redrawn images to obtain a second redrawn image.

[0025] Optionally, use the local redrawing processing module in the block diffusion model to perform local detail enhancement on the sub - blocks to obtain a locally redrawn image;

[0026] S321. Input each sub - block into the local redrawing processing module in the block diffusion model and perform noise processing on each sub - block;

[0027] S322. According to the sub - blocks with noise processing, simulate the image degradation process and perform denoising processing on the sub - blocks with noise processing through multiple iterations;

[0028] S323. During the denoising process, use the block diffusion model to perform local detail enhancement on the current small block to obtain a locally redrawn image.

[0029] Optionally, according to the second redrawn image, use an image understanding model and an image segmentation model to perform target area recognition and segmentation, and use weak redrawing technology to repair the target area to obtain a super - resolution image, including:

[0030] S41. Use the image understanding model to locate the human body area in the second redrawn image;

[0031] S42. Use the image segmentation model to segment the target area in the human body area to obtain the target area;

[0032] S43. Extract the contour and perform weak redrawing on the obtained target area to obtain a repaired target area;

[0033] S44. Paste the repaired target area back into the second redrawn image to obtain a super - resolution image.

[0034] Optionally, using the image understanding model to locate the human body area in the second redrawn image includes:

[0035] S411. Based on the self-supervised learning strategy, extract the global and local feature information in the secondary redrawn image through the convolutional layer and Transformer structure in the image understanding model;

[0036] S412. According to the global and local feature information in the secondary redrawn image, use the attention mechanism to capture the structure and details in the secondary redrawn image;

[0037] S413. According to the structure and details in the secondary redrawn image, generate the position and contour of the human body area.

[0038] Optionally, use an image segmentation model to segment the target area in the human body area to obtain the target area;

[0039] S421. Use the encoder-decoder structure in the image segmentation model to analyze each pixel in the human body area, and generate a segmentation mask in combination with the prompt information and context relevance;

[0040] S422. According to the segmentation mask, identify the target area and irrelevant area in the human body area;

[0041] S423. Segment the target area from the secondary redrawn image to obtain the segmented target area.

[0042] Optionally, perform contour extraction and weak redrawing processing on the obtained target area to obtain the repaired target area;

[0043] S431. Extract the contour of the target area according to the segmentation mask of the target area, and calculate the bounding box of the target area according to the contour of the target area;

[0044] S432. According to the bounding box of the target area, crop out the target area through a cropping operation to form a sub-image;

[0045] S433. Apply weak redrawing processing to the sub-image, combine the local features and semantic prompts of the target area, and optimize each pixel of the sub-image through multiple iterations to obtain the repaired target area.

[0046] According to another aspect of the present invention, a portrait super-resolution processing system combining local repair and global optimization is also provided. The system includes: an input preprocessing module, an image diffusion module, an image block diffusion module, and a secondary optimization module;

[0047] The input preprocessing module is used to preprocess the input image by using a super-resolution model and a semantic guidance generation model to obtain the processed input image and the semantic guidance information of the input image;

[0048] An image diffusion module, which is used to redraw the processed input image based on the semantic guidance information of the input image by using a diffusion model to obtain an initial redrawn image;

[0049] An image block diffusion module, which is used to locally redraw the initial redrawn image according to the block diffusion model to obtain a locally redrawn image, and integrate the locally redrawn images to obtain a second redrawn image;

[0050] A secondary optimization module, which is used to identify and segment the target area based on the second redrawn image by using an image understanding model and an image segmentation model, and use weak redrawing technology to repair the target area to obtain a super-resolution image.

[0051] Compared with the prior art, the present application has the following beneficial effects:

[0052] 1. The present invention combines a super-resolution processing method of local repair and global optimization. Through multi-step denoising and redrawing intensity control, the detail accuracy and quality of the image are significantly improved. Compared with traditional super-resolution methods, the present invention can enhance the visual effect and realism of the image while maintaining the integrity of the image structure. Especially in key areas such as the human face and hands, it demonstrates excellent detail performance.

[0053] 2. The present invention divides the image into multiple small blocks and independently performs local redrawing processing on each block, greatly improving the processing efficiency. Image blocking not only avoids stripe artifacts that may occur during the image processing process but also reduces the consumption of computing resources through local processing, making the entire image processing process more efficient.

[0054] 3. In each step of image repair, the present invention uses weak redrawing technology and precise cropping and repair methods, which not only improve the repair effect but also avoid the delay and resource waste of large-scale computing. The optimization of the system makes the processing process efficient and accurate, saving time and computing costs while improving the quality of the processing effect.

[0055] 4. The present invention has a wide range of applications and strong practicability. It is particularly suitable for portrait image processing scenarios that require high-quality detail repair. The present invention can accurately repair details such as the face and hands, avoiding excessive modification of non-target areas, effectively enhancing the realism and detail presentation of the image in practical applications, making the image more in line with the actual needs of users, and having broad commercial value and application potential.

[0056] 5. The present invention combines advanced image processing technologies, such as super-resolution models, diffusion models, and deep learning models, and has strong scalability. When processing other types of images, relevant models and parameters can be flexibly adjusted to adapt to different application scenarios. The present invention can easily adapt to images of different sizes and complexities, and can ensure accurate and efficient processing effects for both high-definition portraits and other images with high detail requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] With the following description of the embodiments, the above-mentioned characteristics, features, and advantages of the present invention, as well as the implementation methods and means, become more clearly understandable. The embodiments are described in detail in conjunction with the accompanying drawings. Shown herein in schematic diagrams:

[0058] Figure 1 is a flowchart of a portrait super-resolution processing method combining local repair and global optimization according to an embodiment of the present invention;

[0059] Figure 2 is a schematic block diagram of a portrait super-resolution processing system combining local repair and global optimization according to an embodiment of the present invention.

[0060] In the figures:

[0061] 1. Input preprocessing module; 2. Image diffusion module; 3. Image block diffusion module; 4. Secondary optimization module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0063] According to an embodiment of the present invention, a portrait super-resolution processing method and system combining local repair and global optimization are provided.

[0064] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. As Figure 1 shown, according to an embodiment of the present invention, a portrait super-resolution processing method combining local repair and global optimization is provided, and the method includes the following steps:

[0065] S1. Using a super-resolution model and a semantic guidance generation model, preprocess the input image to obtain the processed input image and the semantic guidance information of the input image.

[0066] Preferably, using a super-resolution model and a semantic guidance generation model to preprocess the input image, and obtaining the processed input image and the semantic guidance information of the input image includes the following steps:

[0067] S11. Based on the super-resolution model, use super-resolution technology to magnify the input image to obtain the processed input image;

[0068] S12. For the processed input image, extract the hierarchical features in the image through the convolutional neural network in the semantic guidance generation model;

[0069] S13. Based on the hierarchical features, through the integrated analysis of the processed input image, identify the visual information and convert it into a structured text prompt to obtain the semantic guidance information of the input image.

[0070] It should be explained that the input low-resolution image is first magnified by the super-resolution model. Usually, the goal is to increase the image resolution by a fixed multiple, such as a 4-fold increase. The principle of using the super-resolution model for image magnification is mainly based on the super-resolution technology in deep learning. These models learn to recover the details and textures of the high-resolution image from the low-resolution image through training.

[0071] Specifically, these models usually adopt the convolutional neural network (CNN) architecture. After training with a large number of low-resolution and high-resolution image pairs, they can capture the subtle features and structural information of the image. During the magnification process, the model will predict and generate the missing high-frequency details, thereby improving the clarity and detail richness of the image.

[0072] This process is carried out through the existing super-resolution model, which can effectively increase the number of pixels of the image and restore the details. Especially in high-resolution images, there are usually more details and clarity, which is convenient for subsequent image generation and processing.

[0073] The semantic guidance generation model uses the Gemini-1.5-pro API. Gemini-1.5-pro is a powerful image content understanding tool that can identify information such as the main objects, actions, and scene elements in the image, and generate a structured semantic description based on these elements.

[0074] Specifically, Gemini-1.5-pro utilizes advanced deep learning techniques to comprehensively analyze images. First, it extracts low-level features such as edges, textures, and colors in the image through a convolutional neural network, and then passes these features to a high-level semantic model (such as a Transformer-based architecture) to integrate and analyze the overall content of the image, identifying key information such as main objects, scenes, and actions. This process ultimately converts visual information into structured text prompts, providing accurate semantic guidance for subsequent image generation or processing.

[0075] For example, if the input image is a person holding an object on the street, the generated semantic prompt might be: "A figure is standing with an object in hand against the background of a street." Through this process, the generated semantic prompts will cover multiple aspects of the image content, including but not limited to the actions of the subject, the setting of the scene, and the relative positions of the objects. These multi-dimensional descriptions will provide important semantic guidance for subsequent image processing and diffusion models.

[0076] S2. Based on the semantic guidance information of the input image, use the diffusion model to redraw the processed input image to obtain an initial redrawn image.

[0077] Preferably, based on the semantic guidance information of the input image, using the diffusion model to redraw the processed input image to obtain an initial redrawn image includes the following steps:

[0078] S21. Input the processed input image into the diffusion model, and use the diffusion model to add noise to the processed input image to obtain a noisy image;

[0079] S22. According to the semantic guidance information of the input image, use the redrawing intensity controller in the diffusion model to perform multiple denoising redraws on the noisy image to obtain a preliminary redrawn image.

[0080] It should be explained that the core idea of the diffusion model is to simulate the degradation process of the image by introducing noise and gradually remove the noise in the reverse process to restore or enhance the original image. Multi-step denoising is an important step in the diffusion model. The model can obtain higher-quality images through multiple iterations of denoising. Each step of denoising will refine the image and gradually improve its details.

[0081] First, the model adds noise to the input image, gradually destroying the details of the input image. Then, the model learns how to recover the input image from the noise, i.e., the denoising process. During the denoising process, the model generates the content of the missing area according to the semantic guidance information of the input image (such as prompts or image parts). This semantic guidance information of the input image ensures that the filled content is consistent with the style and semantics of the original image. The denoising process is usually divided into multiple steps, and each step makes subtle adjustments to the input image. Through multiple iterations, the model gradually recovers the details of the missing area until a natural and coherent image is generated. The model can adjust the intensity of redrawing to control the similarity between the generated content and the original image. An appropriate redrawing intensity helps to retain the stability of the main structure while enhancing the naturalness and integrity of the texture details. Through the above steps, the diffusion model can generate natural and coherent content in the missing area of the image, achieving low redrawing processing.

[0082] By adjusting the redrawing intensity parameter (usually between 0.3 and 0.7), the degree of detail enhancement can be controlled. A lower redrawing intensity (e.g., 0.3) will generate less detail and retain the structure of the image; a higher redrawing intensity (e.g., 0.7) will produce more detail and enhance the texture and texture of the image. By controlling this parameter, the detail generation of the image can be flexibly adjusted so that the generated details are neither distorted nor lose the stability of the main structure.

[0083] S3. According to the block diffusion model, perform local redrawing on the initial redrawn image to obtain a locally redrawn image, and obtain a secondary redrawn image by integrating the locally redrawn images.

[0084] Preferably, according to the block diffusion model, performing local redrawing on the initial redrawn image to obtain a locally redrawn image, and obtaining a secondary redrawn image by integrating the locally redrawn images includes the following steps:

[0085] S31. Enlarge the initial redrawn image and cut the enlarged initial redrawn image into a number of uniform sub-blocks;

[0086] S32. Use the local redrawing processing module in the block diffusion model to perform local detail enhancement processing on the sub-blocks to obtain a locally redrawn image;

[0087] S33. Integrate the locally redrawn images to obtain a secondary redrawn image.

[0088] Preferably, use the local redrawing processing module in the block diffusion model to perform local detail enhancement processing on the sub-blocks to obtain a locally redrawn image;

[0089] S321. Input each sub-block into the local redrawing processing module in the block diffusion model and perform noise processing on each sub-block;

[0090] S322. Simulate the image degradation process based on the sub - image blocks processed with added noise, and perform denoising processing on the sub - image blocks processed with added noise through multiple iterations;

[0091] S323. During the denoising process, use the block - diffusion model to enhance the local details of the current small block to obtain a locally redrawn image.

[0092] It should be noted that through the magnification module of the TTP_Toolset plug - in, the initial redrawn image can be divided into multiple small blocks instead of processing the entire large image. The TTP_Toolset plug - in can divide the initial redrawn image into uniform sub - regions (such as sizes of 2×2, 2×3, 3×3, etc.), and these small blocks can be independently processed for subsequent image processing.

[0093] The initial redrawn image is divided into multiple smaller sub - image blocks. The purpose of this approach is to avoid generating complex artifacts (such as stripe artifacts) when processing large images. The size of the image blocks (such as 2×2, 2×3, 3×3, etc.) can be adjusted according to needs, and each sub - image block can be independently processed to reduce the computational complexity during processing and avoid visual distortion.

[0094] In the diffusion model, after the initial redrawn image is divided into multiple sub - image blocks, each sub - image block will independently enter the local redrawing processing module. This module first applies slight noise to each sub - image block to simulate the image degradation process, and then gradually eliminates the noise through multiple iterative denoising processes while enhancing the local details. In each iteration, according to the semantic guidance information of the input image, the model will finely adjust the pixel - level information of the current sub - image block, optimize the texture details, and obtain a locally redrawn image. Independently processing each sub - image block enables the model to precisely control the intensity of detail enhancement, effectively avoiding stripe artifacts and other distortion phenomena that may occur in the overall image processing. Finally, all locally redrawn images will be seamlessly integrated back into the original initial redrawn image to obtain a second - redrawn image, ensuring the naturalness and consistency of the overall visual effect of the second - redrawn image.

[0095] S4. Based on the second - redrawn image, use the image understanding model and the image segmentation model to identify and segment the target area, and use the weak redrawing technology to repair the target area to obtain a super - resolution image.

[0096] Preferably, based on the second - redrawn image, using the image understanding model and the image segmentation model to identify and segment the target area, and using the weak redrawing technology to repair the target area to obtain a super - resolution image includes:

[0097] S41. Use the image understanding model to locate the human body area in the second - redrawn image;

[0098] S42. Use an image segmentation model to segment the target area in the human body area to obtain the target area;

[0099] S43. Perform contour extraction and weak redrawing processing on the obtained target area to obtain the repaired target area;

[0100] S44. Paste the repaired target area back into the second redrawn image to obtain a super-resolution image.

[0101] Preferably, using an image understanding model to locate the human body area in the second redrawn image includes:

[0102] S411. Based on a self-supervised learning strategy, extract global and local feature information in the second redrawn image through the convolutional layer and Transformer structure in the image understanding model;

[0103] S412. According to the global and local feature information in the second redrawn image, use the attention mechanism to capture the structure and details in the second redrawn image;

[0104] S413. According to the structure and details in the second redrawn image, generate the position and contour of the human body area.

[0105] Preferably, use an image segmentation model to segment the target area in the human body area to obtain the target area;

[0106] S421. Use the encoder-decoder structure in the image segmentation model to analyze each pixel in the human body area, and generate a segmentation mask by combining the prompt information and context relevance;

[0107] S422. According to the segmentation mask, identify the target area and irrelevant area in the human body area;

[0108] S423. Segment the target area from the second redrawn image to obtain the segmented target area.

[0109] Preferably, perform contour extraction and weak redrawing processing on the obtained target area to obtain the repaired target area;

[0110] S431. Extract the contour of the target area according to the segmentation mask of the target area, and calculate the bounding box of the target area according to the contour of the target area;

[0111] S432. According to the bounding box of the target area, cut out the target area through a cropping operation to form a sub-image;

[0112] S433. Apply weak redrawing processing to the sub-image, combine the local features and semantic prompts of the target area, and optimize each pixel of the sub-image through multiple iterations to obtain the repaired target area.

[0113] It should be noted that the G-DINO (Global-Local DINO) model is an image understanding model based on self-supervised learning. It can roughly locate the human body area by analyzing the global and local information in the image. In this way, the system can identify the approximate position and contour of the person in the secondarily redrawn image (for example, detect the area of the human body), but will not perform precise segmentation or redrawing of the details.

[0114] The G-DINO model achieves rough localization of the human body area by performing deep feature extraction on the secondarily redrawn image. The model adopts a self-supervised learning strategy. First, it uses convolutional layers and Transformer structures to extract global and local feature information in the secondarily redrawn image, and captures the structure and details in the image through the self-attention mechanism, and then generates candidate regions containing the approximate position and contour of the human body. This process combines multi-scale feature fusion, which not only ensures the global view but also takes into account local details, enabling the model to quickly identify the main human body area in various complex scenarios.

[0115] Based on the rough localization result of the G-DINO model, the SAM model (image segmentation model) performs more refined pixel-level segmentation on the human body area, especially high-precision segmentation on key parts such as the face and hands. The SAM model adopts an encoder-decoder structure based on Transformer. By deeply analyzing each pixel in the input area and combining prompt information and context relevance, it generates an accurate segmentation mask to ensure that the subtle structures such as the face and hands in the human body area of the secondarily redrawn image can be accurately extracted, and effectively isolate the background and other irrelevant areas.

[0116] The pixel-level segmentation ability of the SAM model can ensure that the facial and hand features of the person in the secondarily redrawn image are accurately extracted without being affected by other irrelevant areas.

[0117] After the pixel-level segmentation of the face and hands in the secondarily redrawn image is completed, the system first uses the binary mask generated by the segmentation to extract the contours of the target area, and calculates the precise bounding boxes based on these contours. Through the cropping operation, the system separates the facial and hand areas in the original image to form independent human body area images for subsequent processing. This step not only focuses on the key details but also reduces the interference of background information, thus providing a clearer target area for fine repair.

[0118] After cropping out the human body region images, the system applies weak redrawing processing to these human body region images. Weak redrawing uses low-intensity image redrawing technology, usually controlling the redrawing intensity at 0.5 or below, to ensure that only subtle pixel-level adjustments and texture improvements are made to the target area, without causing excessive interference to the original structure. Through this low-intensity processing method, the model can finely repair the details of the face and hands, enhance the local clarity and naturalness of the human body region images, and effectively avoid stripe artifacts or other distortion phenomena caused by excessive redrawing.

[0119] During the weak redrawing process, the model combines the local features of the target area and the semantic guidance information of the previously acquired input image, optimizes the details of each pixel through multiple iterations, gradually improves the texture and light and shadow effects within the human body region images, and precisely enhances key biometric features (such as the face and hands), thereby achieving high-quality super-resolution image processing.

[0120] After the repair process is completed, the system will accurately paste the repaired face and hand regions back into the original image according to the previous cropping frame information. Specifically, during the cropping process, the system will record the position information of the cropping frame, including the X and Y coordinates and the width and height of the region. The repaired target area will reverse-apply these cropping frame parameters, seamlessly embed the repaired details into the original image, and perform fusion processing on the embedding edges of the weakly redrawn human body region images and the secondarily redrawn images. Image processing techniques (such as Gaussian blur, Poisson fusion, etc.) can be used to smoothly transition the regions, ensuring the seamless connection of the target area with the surrounding images. Through this process, the repaired details are naturally integrated into the original image, avoiding excessive modification of other irrelevant regions, while enhancing the clarity and naturalness of the repaired region, ensuring the overall consistency and authenticity of the image.

[0121] According to another embodiment of the present invention, as Figure 2 shown, there is also provided a portrait super-resolution processing system that combines local repair and global optimization, and this system includes: an input preprocessing module 1, an image diffusion module 2, an image block diffusion module 3, and a secondary optimization module 4;

[0122] The input preprocessing module 1 is used to preprocess the input image by using a super-resolution model and a semantic guidance generation model to obtain the processed input image and the semantic guidance information of the input image;

[0123] The image diffusion module 2 is used to redraw the processed input image based on the semantic guidance information of the input image by using a diffusion model to obtain an initial redrawn image;

[0124] The image block diffusion module 3 is used to perform local redrawing on the initial redrawn image according to the block diffusion model to obtain a locally redrawn image, and obtain a secondarily redrawn image by integrating the locally redrawn images;

[0125] The secondary optimization module 4 is used to perform target area recognition and segmentation on the basis of the secondarily redrawn image by using an image understanding model and an image segmentation model, and to repair the target area by using a weak redrawing technique to obtain a super-resolution image.

[0126] It should be noted that, as shown in Table 1, the portrait super-resolution processing method and system that combine local repair and global optimization are provided with the optimization effect of the human image, and the technical performance comparison results between the present invention and the prior art are shown.

[0127] Table 1 Technical performance comparison table

[0128]

[0129]

[0130] In summary, by means of the above technical solutions of the present invention, the super-resolution processing method of the present invention that combines local repair and global optimization significantly improves the detail accuracy and quality of the image through multi-step denoising and redrawing intensity control. Compared with the traditional super-resolution method, the present invention can improve the visual effect and realism of the image while maintaining the integrity of the image structure, especially in key areas such as the human face and hands, showing excellent detail performance ability.

[0131] The present invention divides the image into multiple small blocks and independently performs local redrawing processing on each block, greatly improving the processing efficiency. Image block division not only avoids the stripe artifacts that may occur during the image processing process, but also reduces the consumption of computing resources through local processing, making the entire image processing process more efficient.

[0132] In each step of image repair, the present invention uses a weak redrawing technique and an accurate cropping and repair method, which not only improves the repair effect, but also avoids the delay and resource waste of large-scale computing. The optimization of the system makes the processing process efficient and accurate, saving time and computing costs, and at the same time improving the quality of the processing effect.

[0133] The present invention has a wide range of applications and strong practicability, and is particularly suitable for portrait image processing scenarios that require high-quality detail repair. The present invention can accurately repair details such as the face and hands, avoiding excessive modification of non-target areas, and can effectively improve the realism and detail presentation of the image in practical applications, making the image more in line with the actual needs of users, and having broad commercial value and application potential.

[0134] The present invention combines advanced image processing technologies, such as super-resolution models, diffusion models, and deep learning models, and has strong scalability. When processing other types of images, relevant models and parameters can be flexibly adjusted to adapt to different application scenarios. The present invention can easily adapt to images of different sizes and complexities. Whether it is high-definition portraits or other images with high detail requirements, the processing effect can be guaranteed to be accurate and efficient.

[0135] Although the present invention has been disclosed above in preferred embodiments, the embodiments are only for illustrative purposes and are not intended to limit the present invention. Those skilled in the art can make several modifications and refinements without departing from the spirit and scope of the present invention. The scope of protection claimed by the present invention shall be subject to what is described in the claims.

Claims

1. A portrait super-resolution processing method combining local repair and global optimization, characterized in that, The method includes the following steps: S1. Using a super-resolution model and a semantic-guided generation model, preprocess the input image to obtain the processed input image and the semantic guidance information of the input image; S2. Based on the semantic guidance information of the input image, use a diffusion model to redraw the processed input image to obtain an initial redrawn image; S3. According to the block diffusion model, perform local redrawing on the initial redrawn image to obtain a locally redrawn image, and integrate the locally redrawn images to obtain a second redrawn image; S4. According to the second redrawn image, use an image understanding model and an image segmentation model to identify and segment the target area, and use weak redrawing technology to repair the target area to obtain a super-resolution image.

2. The method for portrait super-resolution processing combining local repair and global optimization according to claim 1, wherein The step of using a super-resolution model and a semantic-guided generation model to preprocess the input image to obtain the processed input image and the semantic guidance information of the input image includes the following steps: S11. Based on the super-resolution model, use super-resolution technology to magnify the input image to obtain the processed input image; S12. For the processed input image, extract the hierarchical features in the image through the convolutional neural network in the semantic-guided generation model; S13. Based on the hierarchical features, through integrated analysis of the processed input image, identify visual information and convert it into structured text prompts to obtain the semantic guidance information of the input image.

3. The method for portrait super-resolution processing combining local repair and global optimization according to claim 1, wherein The step of using a diffusion model to redraw the processed input image based on the semantic guidance information of the input image to obtain an initial redrawn image includes the following steps: S21. Input the processed input image into the diffusion model, and use the diffusion model to add noise to the processed input image to obtain a noisy image; S22. According to the semantic guidance information of the input image, use the redrawing intensity controller in the diffusion model to perform multiple denoising redraws on the noisy image to obtain a preliminary redrawn image.

4. The method for portrait super-resolution processing combining local repair and global optimization according to claim 1, wherein The step of performing local redrawing on the initial redrawn image according to the block diffusion model to obtain a locally redrawn image and integrating the locally redrawn images to obtain a second redrawn image includes the following steps: S31. Magnify the initial redrawn image and cut the magnified initial redrawn image into a number of uniform sub-blocks; S32. Use the local redrawing processing module in the block diffusion model to perform local detail enhancement processing on the sub-blocks to obtain locally redrawn images; S33. Integrate the locally redrawn images to obtain a second redrawn image.

5. The method for portrait super-resolution processing combining local repair and global optimization according to claim 4, characterized in that, The step of using the local redrawing processing module in the block diffusion model to perform local detail enhancement processing on the sub-blocks to obtain locally redrawn images; S321. Input each sub-block into the local redrawing processing module in the block diffusion model and perform noise processing on each sub-block; S322. According to the sub-blocks with noise processing, simulate the image degradation process and perform denoising processing on the sub-blocks with noise processing through multiple iterations; S323. During the denoising process, use the block diffusion model to perform local detail enhancement on the current small block to obtain a locally redrawn image.

6. The method for portrait super-resolution processing combining local repair and global optimization according to claim 1, characterized in that The above-mentioned obtaining a super-resolution image by using the image understanding model and the image segmentation model to identify and segment the target area according to the second redrawn image, and using the weak redrawing technology to repair the target area includes: S41. Use the image understanding model to locate the human body area in the second redrawn image; S42. Use the image segmentation model to segment the target area in the human body area to obtain the target area; S43. Extract the contour and perform weak redrawing processing on the obtained target area to obtain the repaired target area; S44. Paste the repaired target area back into the second redrawn image to obtain the super-resolution image.

7. The method for portrait super-resolution processing combining local repair and global optimization according to claim 6, characterized in that, The above-mentioned using the image understanding model to locate the human body area in the second redrawn image includes: S411. Based on the self-supervised learning strategy, extract the global and local feature information in the second redrawn image through the convolutional layer and the Transformer structure in the image understanding model; S412. According to the global and local feature information in the second redrawn image, use the attention mechanism to capture the structure and details in the second redrawn image; S413. According to the structure and details in the second redrawn image, generate the position and contour of the human body area.

8. The method for portrait super-resolution processing combining local repair and global optimization according to claim 7, characterized in that The above-mentioned using the image segmentation model to segment the target area in the human body area to obtain the target area; S421. Use the encoder-decoder structure in the image segmentation model to analyze each pixel in the human body area, and combine the prompt information and context relevance to generate a segmentation mask; S422. According to the segmentation mask, identify the target area and the irrelevant area in the human body area; S423. Segment the target area from the second redrawn image to obtain the segmented target area.

9. The method for portrait super-resolution processing combining local repair and global optimization according to claim 8, characterized in that, The above-mentioned extracting the contour and performing weak redrawing processing on the obtained target area to obtain the repaired target area; S431. Extract the contour of the target area according to the segmentation mask of the target area, and calculate the bounding box of the target area according to the contour of the target area; S432. According to the bounding box of the target area, cut out the target area through a cropping operation to form a sub-image; S433. Apply weak redrawing processing to the sub-image, combine the local features and semantic prompts of the target area, and optimize each pixel of the sub-image through multiple iterations to obtain the repaired target area.

10. A portrait super-resolution processing system that combines local repair and global optimization, which is used to implement the portrait super-resolution processing method that combines local repair and global optimization described in any one of claims 1-9, and is characterized in that, The system includes: an input preprocessing module, an image diffusion module, an image block diffusion module, and a secondary optimization module; The input preprocessing module is used to preprocess the input image by using the super-resolution model and the semantic guidance generation model to obtain the processed input image and the semantic guidance information of the input image; The image diffusion module is used to redraw the processed input image based on the semantic guidance information of the input image by using the diffusion model to obtain the initial redrawn image; The image block diffusion module is used to perform local redrawing on the initial redrawn image according to the block diffusion model to obtain a locally redrawn image, and obtain the second redrawn image by integrating the locally redrawn images; The secondary optimization module is used to perform target area recognition and segmentation on the basis of the secondarily redrawn image by using an image understanding model and an image segmentation model, and to repair the target area by using a weak redrawing technique to obtain a super-resolution image.

Citation Information

Cited By

  • Image processing method and related device

    CN121073837A

  • An image processing method and related apparatus

    CN121073837B

  • AI local picture accurate modification method and system based on guide amplification

    CN121904228A

  • An AI local picture precise modification method and system based on guided amplification

    CN121904228B