Image restoration method and device

By acquiring and extracting the target image data of the face area, and optimizing and redrawing the text prompt words with the target grayscale map and feature prompt words, the problem of facial distortion in the face image generation of the text generation image is solved, and high-quality image repair is achieved.

CN120013824APending Publication Date: 2025-05-16CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411875589.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When mainstream text-generating image models generate images involving faces, facial areas are prone to distortion, resulting in a decrease in image quality, limiting the practical application of the technology.

Method used

By obtaining the initial image data to be repaired, the target image data of the face area is extracted, the target grayscale map and face feature prompt words are generated, and the text prompt words are spliced ​​for redrawing, and the image data is optimized to repair facial distortion.

Benefits of technology

It realizes accurate repair of facial distortion areas, effectively retains the original properties of the face, improves image quality, and promotes the development of image generation technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013824A_ABST
    Figure CN120013824A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image restoration method and device, and the method comprises the steps: obtaining to-be-restored initial image data, extracting target image data corresponding to a face region in the initial image data, generating a target gray-scale map of the target image data and a face feature prompt word, obtaining an initial text cue word used for generating the initial image data, splicing the face feature cue word and the initial text cue word into a target text cue word, and redrawing the target image data according to the target text cue word, the target grey-scale map and the target image data, according to the method and the device, the initial image data is restored, and the face distortion image is restored by optimizing the cue word, the grey-scale map and the original target image data, so that accurate restoration of the face distortion area can be completed, the original attribute of the face is effectively reserved, the image quality is improved, and the development of the image generation technology is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and image generation, and in particular to a method and device for image restoration. Background Art

[0002] Image generation technology is an important field in computer vision research, with broad application prospects and great potential in many industries such as art creation, product design, advertising marketing, and media entertainment. In recent years, with the progress of multimodal research and the emergence of diffusion models, text-generated image technology (text-generated image) has achieved remarkable results. This technology generates corresponding images by analyzing the descriptive text prompts entered by the user, which greatly expands the application scope of image generation and gives rise to a variety of AI products such as virtual fitting, image style transfer, and poster-assisted design.

[0003] However, when the generated images involve human faces, mainstream text-based image models are often prone to distortion in the facial area, resulting in a decrease in image quality, which seriously restricts the practical application of this technology. Summary of the invention

[0004] In view of the above problems, a method and apparatus for image restoration is proposed to overcome the above problems or at least partially solve the above problems, including:

[0005] A method for image restoration, the method comprising:

[0006] Acquire initial image data to be restored, and extract target image data corresponding to the face area in the initial image data;

[0007] Generate a target grayscale image and facial feature prompt words of the target image data;

[0008] Acquire initial text prompt words used to generate the initial image data, and splice the facial feature prompt words with the initial text prompt words into target text prompt words;

[0009] The target image data is redrawn according to the target text prompt word, the target grayscale image and the target image data to obtain restored initial image data.

[0010] Optionally, the extracting target image data corresponding to the face area in the initial image data includes:

[0011] Obtaining a plurality of face detection frames in the initial image data through a preset face detection model;

[0012] Calculating a plurality of clarity values ​​corresponding to the image data in the plurality of face detection frames;

[0013] Determining a plurality of candidate face detection frames from the plurality of face detection frames according to the plurality of clarity values;

[0014] Performing an expansion operation on the multiple candidate face detection frames to obtain corresponding multiple object detection frames;

[0015] The corresponding image data within the multiple target detection frames are cropped to obtain corresponding multiple target image data.

[0016] Optionally, determining a plurality of candidate face detection frames from the plurality of face detection frames according to the plurality of clarity values ​​comprises:

[0017] Acquire a plurality of target clarity values ​​greater than a preset clarity threshold value among the plurality of clarity values;

[0018] The multiple face detection frames corresponding to the multiple target clarity values ​​are determined as multiple candidate face detection frames.

[0019] Optionally, generating a target grayscale image of the target image data includes:

[0020] Obtaining an initial grayscale image of the target image data through a preset semantic segmentation model;

[0021] An edge feathering operation is performed on the initial grayscale image to obtain a target grayscale image of the target image data.

[0022] Optionally, redrawing the target image data according to the target text prompt word, the target grayscale image and the target image data includes:

[0023] According to the target text prompt word, a control condition vector is obtained;

[0024] According to the target image data, obtaining initial image features;

[0025] The target image data is redrawn according to the control condition vector, the target grayscale image and the initial image feature.

[0026] Optionally, redrawing the target image data according to the control condition vector, the target grayscale image and the initial image feature comprises:

[0027] Adding Gaussian noise to the initial image feature to obtain an initial noise image feature;

[0028] Based on a preset number of iterations, performing denoising and weighting operations on the initial noise image features in an iterative manner;

[0029] In response to completion of the preset number of iterations, a redrawn image feature is obtained;

[0030] According to the redrawing image feature, redrawing image data of the target image data is obtained.

[0031] Optionally, the performing denoising and weighting operations on the initial noise image features includes:

[0032] De-noising the initial noise image feature according to the control condition vector to obtain an initial denoised image feature;

[0033] According to the target grayscale image, the initial denoising image features and the initial image features are weightedly combined to obtain initial redrawing image features.

[0034] A device for image restoration, comprising:

[0035] A target image acquisition module, used to acquire the initial image data to be restored, and extract the target image data corresponding to the face area in the initial image data;

[0036] A repair data acquisition module is used to generate a target grayscale image and facial feature prompt words of the target image data;

[0037] A text prompt word acquisition module, used to acquire initial text prompt words used to generate the initial image data, and to splice the facial feature prompt words with the initial text prompt words into target text prompt words;

[0038] The redrawing and repairing module is used to redraw the target image data according to the target text prompt word, the target grayscale image and the target image data to obtain the repaired initial image data.

[0039] An electronic device comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the image restoration method as claimed in any one of claims 1 to 7.

[0040] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the image restoration method according to any one of claims 1 to 7 is implemented.

[0041] The embodiments of the present invention have the following advantages:

[0042] In an embodiment of the present invention, initial image data to be repaired is obtained, and target image data corresponding to a face area in the initial image data is extracted, a target grayscale image and a face feature prompt word of the target image data are generated, an initial text prompt word used to generate the initial image data is obtained, and the face feature prompt word and the initial text prompt word are spliced ​​into a target text prompt word, and the target image data is redrawn according to the target text prompt word, the target grayscale image and the target image data to obtain the repaired initial image data, thereby realizing the repair of a facially distorted image by optimizing the prompt word, the grayscale image and the original target image data, thereby completing the accurate repair of the facially distorted area, effectively retaining the original attributes of the face, improving the image quality, and promoting the development of image generation technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0044] Figure 1 is a flowchart of steps of an image restoration method provided by some embodiments of the present invention;

[0045] Figure 2 is a flowchart of steps of a facial image restoration solution provided by some embodiments of the present invention;

[0046] Figure 3 is a schematic diagram of a process of repairing a facial image by redrawing provided by some embodiments of the present invention;

[0047] Figure 4 It is a structural block diagram of an image restoration device provided by some embodiments of the present invention. DETAILED DESCRIPTION

[0048] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] At present, common image restoration algorithms based on generative adversarial networks (GANs) often focus on image quality issues such as image denoising and old photo restoration, but have limited effect in dealing with severe distortions caused by the Wenshengtu model (such as artifacts such as distorted facial features). Another type of image restoration algorithm based on the diffusion model can produce obvious changes in the distorted area, but there is a problem of disharmony between the restored face and the surrounding background area. Therefore, it is urgent to develop a method that can effectively repair facial distortion of portraits while ensuring that the restoration results are harmonious and reasonable.

[0050] In one embodiment of the invention, the image quality is improved by optimizing the distorted facial area. First, a segmentation mask is introduced to prevent the background area that is not a human face from being erroneously redrawn. Then, on this basis, the detection frame optimization strategy, the segmentation mask optimization strategy, and the text prompt optimization strategy are used to make the repair result more harmoniously integrated with the background area, so that the distorted facial area in the portrait can be effectively repaired while ensuring the harmony of the repair result, thereby effectively improving the image quality.

[0051] Reference Figure 1 , shows a flowchart of a method for image restoration provided by some embodiments of the present invention, which may specifically include the following steps:

[0052] Step 101, obtaining initial image data to be restored, and extracting target image data corresponding to the face area in the initial image data.

[0053] like Figure 2 As shown, when generating an image based on a text prompt word, the text prompt word can be first input into a text-generated image model (which can be referred to as a text-generated image model) to obtain a corresponding image. Furthermore, if multiple images are generated by multiple sets of text prompt words, the initial image data to be restored can be obtained from the multiple images.

[0054] In addition, the image restoration method of the present invention can also be used as an image optimization method, that is, after generating an image according to the text prompt word, the generated image can be used as the initial image data to be restored, and then optimized by the image restoration method of the present invention.

[0055] After the initial image data to be restored is acquired, the target image data corresponding to the face area in the initial image data to be restored may be extracted.

[0056] In some embodiments of the present invention, Figure 2 As shown in the corresponding steps of the detection module, the extraction of target image data corresponding to the face area in the initial image data includes:

[0057] Sub-step 11, obtaining multiple face detection frames in the initial image data through a preset face detection model.

[0058] A plurality of face detection frames in the initial image data may be obtained through a preset face detection model. Specifically, the initial image data to be repaired is input into the face detection model to obtain a plurality of face detection frames. The face detection model may be a YOLO face detection model (You Only Look Once).

[0059] Sub-step 12, calculating multiple clarity values ​​of the corresponding image data in the multiple face detection frames.

[0060] After obtaining a plurality of face detection frames, a plurality of clarity values ​​corresponding to the image data in the plurality of face detection frames may be calculated.

[0061] Sub-step 13: determining a plurality of candidate face detection frames from the plurality of face detection frames according to the plurality of clarity values.

[0062] After obtaining a plurality of clarity values, a plurality of face detection frames may be screened according to the clarity values, thereby determining a plurality of candidate face detection frames from the plurality of face detection frames.

[0063] In some embodiments of the present invention, determining a plurality of candidate face detection frames from the plurality of face detection frames according to the plurality of clarity values ​​comprises:

[0064] Sub-step 21 : acquiring a plurality of target clarity values ​​greater than a preset clarity threshold value among the plurality of clarity values.

[0065] After obtaining a plurality of clarity values, a plurality of target clarity values ​​greater than a preset clarity threshold value may be obtained from the plurality of clarity values, that is, only clarity values ​​greater than the preset clarity threshold value are retained.

[0066] Sub-step 22: determining a plurality of face detection frames corresponding to the plurality of target definition values ​​as a plurality of candidate face detection frames.

[0067] A plurality of face detection frames corresponding to a plurality of target clarity values ​​may be determined as a plurality of candidate face detection frames.

[0068] In a specific operation, after calculating multiple clarity values ​​of corresponding image data in multiple face detection frames with the help of the Laplace operator, multiple face detection frames are screened according to the clarity values, and only face detection frames whose clarity values ​​are greater than a preset clarity threshold are retained. The algorithm may be as follows:

[0069]

[0070] In formula (2), B and B' are the face detection frame sets before and after screening, I is the initial image data to be repaired, |B i | represents the number of pixels in the i-th face detection frame, Δ represents the Laplace operator, and T is the preset clarity threshold.

[0071] Through the above screening operation, multiple candidate face detection frames can be determined from multiple face detection frames, and the blurred corresponding image data in the multiple face detection frames are filtered out, which can effectively avoid the inharmonious problem of the repaired image when a clear face is redrawn on a blurred background.

[0072] Sub-step 14: performing an expansion operation on the multiple candidate face detection frames to obtain corresponding multiple object detection frames.

[0073] After determining a plurality of candidate face detection frames, an expansion operation may be performed on the plurality of candidate face detection frames to obtain a plurality of corresponding target detection frames.

[0074] Specifically, multiple candidate face detection frames can be expanded, that is, expanded outward, keeping the center coordinates unchanged, and expanding the length and width to r times (r>1) of the original, so as to obtain multiple corresponding target detection frames. The purpose of the expansion operation is to provide sufficient context information for redrawing the distorted area. At the same time, considering that the mainstream diffusion model used for subsequent redrawing is often trained based on a large amount of 1:1 image data, in order to better exert the ability of the diffusion model, multiple candidate face detection frames are further expanded outward into squares, while keeping the center coordinates unchanged before and after the expansion.

[0075] In addition, if the expanded detection frame exceeds the image boundary, the pixels in the excess part are filled with black.

[0076] In a specific implementation, the calculation method for performing the expansion operation on multiple candidate face detection frames may be as follows:

[0077]

[0078] In formula (2) to formula (5), (x1, y1) and (x'1, y'1) are the coordinates of the upper left corner of the candidate face detection frame before and after the expansion operation, and (x2, y2) and (x'2, y'2) are the coordinates of the lower right corner of the candidate face detection frame before and after the expansion operation.

[0079] Sub-step 15, cropping the corresponding image data within the multiple target detection frames to obtain corresponding multiple target image data.

[0080] After performing an expansion operation on the multiple candidate face detection frames to obtain the corresponding multiple target detection frames, the corresponding image data in the multiple target detection frames may be cropped to obtain the corresponding multiple target image data.

[0081] Specifically, the initial image data to be repaired is cropped according to the multiple target detection frames to obtain the corresponding image data within the multiple target detection frames, that is, the multiple target image data.

[0082] Step 102: Generate a target grayscale image and facial feature prompt words of the target image data.

[0083] After acquiring the target image data, a corresponding target grayscale image and facial feature prompt words can be generated according to the target image data. The facial feature prompt words can include prompt words of character attributes such as gender, age, and expression.

[0084] In some embodiments of the present invention, Figure 2 As shown in the steps corresponding to the segmentation module in the figure, generating a target grayscale image of the target image data includes:

[0085] Sub-step 31, obtaining an initial grayscale image of the target image data through a preset semantic segmentation model.

[0086] The target image data can be input into a preset semantic segmentation model to obtain an initial grayscale image corresponding to the target image data.

[0087] Specifically, the target image data can be input into a preset semantic segmentation model to obtain a segmentation mask of the target image data, and the segmentation mask can be used to indicate the image area that needs to be redrawn. It should be noted that the segmentation mask can be a special grayscale image used to represent the segmentation results of different areas in the image, that is, the segmentation mask can be an initial grayscale image corresponding to the target image data.

[0088] In addition, in the present invention, a fine segmentation mask can be constructed through a preset semantic segmentation model, thereby effectively avoiding the problem of inharmonious repaired images caused by redrawing non-facial areas of the background.

[0089] Sub-step 32, performing an edge feathering operation on the initial grayscale image to obtain a target grayscale image of the target image data.

[0090] After obtaining the initial grayscale image corresponding to the target image data, an edge feathering operation may be performed on the initial grayscale image to obtain a target grayscale image corresponding to the target image data.

[0091] Specifically, the segmentation mask corresponding to the target image data may be subjected to an edge feathering operation so as to harmoniously blend the redrawn content and the background content. The segmentation mask after the feathering operation is the target grayscale image corresponding to the target image data. The edge feathering operation may be performed as follows:

[0092] M'=M*G (6)

[0093] In formula (6), M and M' are the segmentation masks before and after feathering, G is the Gaussian kernel function, and * represents the convolution operation.

[0094] As an example, the generation of facial feature prompt words for target image data can be performed through a prompt word inversion tool. Specifically, the corresponding image data in multiple target detection frames can be input into the prompt word inversion tool to obtain prompt words (i.e., facial feature prompt words) including character attributes such as gender, age, and expression. The prompt word inversion tool can use a model based on CLIP (Contrastive Language–Image Pretraining, a deep learning model) or a multimodal large language model.

[0095] Step 103, obtaining initial text prompt words used to generate the initial image data, and splicing the facial feature prompt words with the initial text prompt words into target text prompt words.

[0096] like Figure 2 The text prompt word optimization in the redrawing module shown above can first obtain the initial text prompt word used to generate the initial image data, that is, the initial image data is the image data to be optimized or repaired generated by the initial text prompt word. Then, the facial feature prompt word obtained by the prompt word reverse deduction tool and the initial text prompt word are spliced ​​into the target text prompt word, that is, the initial text prompt word is optimized.

[0097] The optimized initial text prompt word, namely the target text prompt word, is used for face area redrawing, which helps to maintain the character attributes of the redrawn content, thereby promoting the harmonious integration of the redrawn content and the non-face area content.

[0098] Step 104 , redrawing the target image data according to the target text prompt word, the target grayscale image and the target image data to obtain restored initial image data.

[0099] like Figure 2 In the distorted area redrawing step of the redrawing module, after obtaining the target text prompt word, the target grayscale image and the target image data, that is, after obtaining the optimized initial text prompt word, the segmentation mask after the feathering operation and the target image data, the target image data can be redrawn according to the optimized initial text prompt word, the segmentation mask after the edge feathering operation and the target image data to obtain the redrawn image data, and then the repaired initial image data can be obtained.

[0100] As an example, the corresponding redrawn image may be pasted back to the original image data to be repaired according to the position determined by the target detection frame to replace the target image data, thereby obtaining the final repaired original image data.

[0101] In some embodiments of the present invention, Figure 3 As shown, the target image data is redrawn according to the target text prompt word, the target grayscale image and the target image data, including:

[0102] Sub-step 41, obtaining a control condition vector according to the target text prompt word.

[0103] Before officially starting to redraw the target image data, the control condition vector can be obtained according to the target text prompt word. Specifically, the target text prompt word can be input into a text encoder to obtain the control condition vector.

[0104] Sub-step 42, obtaining initial image features according to the target image data.

[0105] The initial image features may be obtained according to the target image data. Specifically, the target image data may be input into an image encoder to obtain the initial image features.

[0106] Sub-step 43, redrawing the target image data according to the control condition vector, the target grayscale image and the initial image feature.

[0107] After obtaining the initial image features and the control condition vector, the target image data is formally redrawn, that is, the target image data is redrawn according to the control condition vector, the target grayscale image and the initial image features.

[0108] Specifically, the control condition vector, the segmentation mask after the edge feathering operation, and the initial image features may be input into the diffusion model for processing.

[0109] In some embodiments of the present invention, Figure 3 As shown, redrawing the target image data according to the control condition vector, the target grayscale image and the initial image feature includes:

[0110] Sub-step 51, adding Gaussian noise to the initial image feature to obtain an initial noise image feature.

[0111] During the redrawing process, Gaussian noise may be first added to the initial image features to obtain initial noise image features.

[0112] Sub-step 52, based on a preset number of iterations, in an iterative manner, performs denoising and weighting operations on the initial noise image features.

[0113] Based on a preset number of iterations, the initial noise image features can be subjected to denoising and weighting operations in an iterative manner.

[0114] In a specific implementation, a number of iterations, that is, a preset number of time steps T, can be preset, and then the diffusion model can perform an iterative denoising operation on the initial noisy image features, that is, the initial image features after adding Gaussian noise, based on the preset time steps.

[0115] For example, starting from the first time step, the initial noisy image features are denoised and weighted, and the obtained image features are used as input for the denoising and weighting operations in the second time step. Then, the corresponding image features obtained in the second time step are used as input for the third time step, and it is iterated in this way.

[0116] Sub-step 53, in response to completion of the preset number of iterations, obtaining a redrawn image feature.

[0117] During the above iterative denoising operation, the response and iterative time step number reaches T, indicating that the iterative denoising operation has been completed, and the image features output in the last time step are the redrawn image features.

[0118] Sub-step 54, obtaining redrawn image data of the target image data according to the redrawn image feature.

[0119] After obtaining the redrawn image feature, the redrawn image of the target image data can be obtained according to the redrawn image feature. Specifically, the redrawn image feature can be input into an image decoder for decoding, and then the redrawn image data corresponding to the target image data can be obtained.

[0120] In some embodiments of the present invention, the performing denoising and weighting operations on the initial noisy image features includes:

[0121] Sub-step 61, denoising the initial noisy image features according to the control condition vector to obtain initial denoised image features.

[0122] First of all, it should be noted that this step describes the denoising operation performed at the first time step of the iterative denoising operation, that is, the denoising operation is performed on the initial noise image features. The same denoising operation can also be performed at other time steps, but the object of the denoising operation is the noise image features corresponding to the corresponding time step.

[0123] Specifically, the diffusion model can obtain the corresponding denoised image features according to the control condition vector and the noise image features corresponding to a certain time step. That is, the diffusion model can generate the corresponding denoised image features according to the control condition vector and the noise image features corresponding to the corresponding time step.

[0124] For example, the denoising operation may be to denoise the initial noise image features according to the control condition vector to obtain the initial denoised image features. That is, the diffusion model may generate the corresponding denoised image features according to the control condition vector and the initial noise image features corresponding to the first time step.

[0125] Sub-step 62, performing weighted combination on the initial denoising image features and the initial image features according to the target grayscale image, to obtain initial redrawing image features.

[0126] It should also be noted that this step describes the weighted operation performed at the first time step of the iterative denoising operation, that is, a weighted operation is performed on the initial denoised image features. The same weighted operation can also be performed at other time steps, but the object of the weighted operation is the denoised image features corresponding to the corresponding time step.

[0127] Specifically, the diffusion model can perform a weighted combination of the denoised image features and the initial image features corresponding to the corresponding time step according to the segmentation mask after the edge feathering operation, and obtain the corresponding intermediate redrawn image features.

[0128] For example, the diffusion model can perform a weighted combination of the initial denoised image features and the initial image features corresponding to the first time step according to the segmentation mask after the edge feathering operation to obtain the initial redrawn image features.

[0129] The weighted combination formula of the weighted operation can be as follows:

[0130]

[0131] In formula (7), ⊙ represents the element-by-element multiplication operation, c text is the control condition vector, x known is the initial image feature, m is the segmentation mask after edge feathering operation, F is the diffusion model, x t-1 is the intermediate redrawn image feature corresponding to the corresponding time step.

[0132] In practical applications, this method can be used as the post-processing part of the text-generated image model, and thus integrated into various applications and products derived from the text-generated image model and whose generated content may contain human faces, such as creative design, poster generation, virtual fitting, etc.

[0133] Take a creative design website as an example: the front end obtains the text prompt words input by the user through the user interface, and then calls the text-image model of the back end to generate an image, and then obtains the repaired and optimized image through the image repair method of the present invention, and finally returns the repaired and optimized image to the front end through the interface and provides it to the user.

[0134] In an embodiment of the present invention, initial image data to be repaired is obtained, and target image data corresponding to a face area in the initial image data is extracted, a target grayscale image and a face feature prompt word of the target image data are generated, an initial text prompt word used to generate the initial image data is obtained, and the face feature prompt word and the initial text prompt word are spliced ​​into a target text prompt word, and the target image data is redrawn according to the target text prompt word, the target grayscale image and the target image data to obtain the repaired initial image data, thereby realizing the repair of a facially distorted image by optimizing the prompt word, the grayscale image and the original target image data, thereby completing the accurate repair of the facially distorted area, effectively retaining the original attributes of the face, improving the image quality, and promoting the development of image generation technology.

[0135] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0136] Reference Figure 4 , shows a schematic diagram of the structure of an image restoration device provided by some embodiments of the present invention, which may specifically include the following modules:

[0137] The target image acquisition module 401 is used to acquire the initial image data to be restored, and extract the target image data corresponding to the face area in the initial image data;

[0138] A repair data acquisition module 402 is used to generate a target grayscale image and facial feature prompt words of the target image data;

[0139] The text prompt word acquisition module 403 is used to acquire the initial text prompt words used to generate the initial image data, and to splice the facial feature prompt words with the initial text prompt words into target text prompt words;

[0140] The redrawing and repairing module 404 is used to redraw the target image data according to the target text prompt word, the target grayscale image and the target image data to obtain the repaired initial image data.

[0141] In one embodiment of the present invention, the target image acquisition module 401 includes:

[0142] A face detection frame acquisition submodule, used to obtain multiple face detection frames in the initial image data through a preset face detection model;

[0143] A clarity calculation submodule, used to calculate a plurality of clarity values ​​corresponding to the image data in the plurality of face detection frames;

[0144] A candidate face detection frame acquisition submodule, used for determining a plurality of candidate face detection frames from the plurality of face detection frames according to the plurality of clarity values;

[0145] The target detection frame acquisition submodule is used to perform an expansion operation on the multiple candidate face detection frames to obtain corresponding multiple target detection frames;

[0146] The target image acquisition submodule is used to crop the corresponding image data within multiple target detection frames to obtain corresponding multiple target image data.

[0147] In one embodiment of the present invention, the candidate face detection frame acquisition submodule includes:

[0148] A target clarity acquisition unit, configured to acquire a plurality of target clarity values ​​greater than a preset clarity threshold value among the plurality of clarity values;

[0149] The candidate face detection frame acquisition unit is used to determine the multiple face detection frames corresponding to the multiple target clarity values ​​as multiple candidate face detection frames.

[0150] In one embodiment of the present invention, the repair data acquisition module 402 includes:

[0151] An initial grayscale image acquisition submodule is used to obtain an initial grayscale image of the target image data through a preset semantic segmentation model;

[0152] The target grayscale image acquisition submodule is used to perform an edge feathering operation on the initial grayscale image to obtain a target grayscale image of the target image data.

[0153] In one embodiment of the present invention, the redraw repair module 404 includes:

[0154] A vector acquisition submodule, used to obtain a control condition vector according to the target text prompt word;

[0155] A feature acquisition submodule, used to obtain initial image features according to the target image data;

[0156] The redrawing submodule is used to redraw the target image data according to the control condition vector, the target grayscale image and the initial image feature.

[0157] In one embodiment of the present invention, the redrawing submodule includes:

[0158] A noise image feature acquisition unit, used for adding Gaussian noise to the initial image feature to obtain an initial noise image feature;

[0159] An iterative redrawing unit, configured to perform a denoising operation and a weighting operation on the initial noise image features in an iterative manner based on a preset number of iterations;

[0160] A redrawing image feature acquisition unit, configured to obtain a redrawing image feature in response to completion of the preset number of iterations;

[0161] The image decoding unit is used to obtain the redrawing image data of the target image data according to the redrawing image feature.

[0162] In one embodiment of the present invention, the iterative redrawing unit includes:

[0163] A denoising operation subunit, configured to denoise the initial noise image feature according to the control condition vector to obtain an initial denoised image feature;

[0164] The weighted operation subunit is used to perform a weighted combination of the initial denoising image features and the initial image features according to the target grayscale image to obtain initial redrawing image features.

[0165] In an embodiment of the present invention, initial image data to be repaired is obtained, and target image data corresponding to a face area in the initial image data is extracted, a target grayscale image and a face feature prompt word of the target image data are generated, an initial text prompt word used to generate the initial image data is obtained, and the face feature prompt word and the initial text prompt word are spliced ​​into a target text prompt word, and the target image data is redrawn according to the target text prompt word, the target grayscale image and the target image data to obtain the repaired initial image data, thereby realizing the repair of a facially distorted image by optimizing the prompt word, the grayscale image and the original target image data, thereby completing the accurate repair of the facially distorted area, effectively retaining the original attributes of the face, improving the image quality, and promoting the development of image generation technology.

[0166] Some embodiments of the present invention further provide an electronic device, comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the above method is implemented when the computer program is executed by the processor.

[0167] Some embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored, and the computer program implements the above method when executed by a processor.

[0168] Some embodiments of the present invention further provide a computer program product, including a computer program, which implements the above method when executed by a processor.

[0169] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0170] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0171] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0172] Those skilled in the art will appreciate that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0173] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0174] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0176] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0177] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the above elements.

[0178] The above is a detailed introduction to the provided method and device for image restoration. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A method for image restoration, characterized in that: The method comprises: Acquire initial image data to be restored, and extract target image data corresponding to the face area in the initial image data; Generate a target grayscale image and facial feature prompt words of the target image data; Acquire initial text prompt words used to generate the initial image data, and splice the facial feature prompt words with the initial text prompt words into target text prompt words; The target image data is redrawn according to the target text prompt word, the target grayscale image and the target image data to obtain restored initial image data.

2. The method according to claim 1, characterized in that The extracting target image data corresponding to the face area in the initial image data includes: Obtaining a plurality of face detection frames in the initial image data through a preset face detection model; Calculating a plurality of clarity values ​​corresponding to the image data in the plurality of face detection frames; Determining a plurality of candidate face detection frames from the plurality of face detection frames according to the plurality of clarity values; Performing an expansion operation on the multiple candidate face detection frames to obtain corresponding multiple object detection frames; The corresponding image data within the multiple target detection frames are cropped to obtain corresponding multiple target image data.

3. The method according to claim 2, characterized in that The step of determining a plurality of candidate face detection frames from the plurality of face detection frames according to the plurality of clarity values ​​comprises: Acquire a plurality of target clarity values ​​greater than a preset clarity threshold value among the plurality of clarity values; The multiple face detection frames corresponding to the multiple target clarity values ​​are determined as multiple candidate face detection frames.

4. The method according to claim 1, characterized in that The step of generating a target grayscale image of the target image data comprises: Obtaining an initial grayscale image of the target image data through a preset semantic segmentation model; An edge feathering operation is performed on the initial grayscale image to obtain a target grayscale image of the target image data.

5. The method according to claim 1, characterized in that The step of redrawing the target image data according to the target text prompt word, the target grayscale image and the target image data comprises: According to the target text prompt word, a control condition vector is obtained; According to the target image data, obtaining initial image features; The target image data is redrawn according to the control condition vector, the target grayscale image and the initial image feature.

6. The method according to claim 5, characterized in that The step of redrawing the target image data according to the control condition vector, the target grayscale image and the initial image feature comprises: Adding Gaussian noise to the initial image feature to obtain an initial noise image feature; Based on a preset number of iterations, performing denoising and weighting operations on the initial noise image features in an iterative manner; In response to completion of the preset number of iterations, a redrawn image feature is obtained; According to the redrawing image feature, redrawing image data of the target image data is obtained.

7. The method according to claim 6, characterized in that The performing denoising and weighting operations on the initial noise image features includes: De-noising the initial noise image feature according to the control condition vector to obtain an initial denoised image feature; According to the target grayscale image, the initial denoising image features and the initial image features are weightedly combined to obtain initial redrawing image features.

8. An image restoration device, characterized in that: The device comprises: A target image acquisition module, used to acquire the initial image data to be restored, and extract the target image data corresponding to the face area in the initial image data; A repair data acquisition module is used to generate a target grayscale image and facial feature prompt words of the target image data; A text prompt word acquisition module, used to acquire initial text prompt words used to generate the initial image data, and to splice the facial feature prompt words with the initial text prompt words into target text prompt words; The redrawing and repairing module is used to redraw the target image data according to the target text prompt word, the target grayscale image and the target image data to obtain the repaired initial image data.

9. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the method for image restoration according to any one of claims 1 to 7 is implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the image restoration method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Image generation method and device, computer equipment, storage medium and program product

    CN121010661A

  • Image generation method and device, computer device, storage medium and program product

    CN121010661B