An Image Inpainting Method, Device, Equipment and Medium Based on Latent Mask
By performing potential mask processing and grouping pre-training on the images, and selecting the optimal model for fine-tuning in combination with the damage ratio of the images to be repaired, the problem that traditional image repair methods fail to effectively utilize the visible information in the image is achieved, achieving higher quality and more efficient image repair effects.
Patent Information
- Application Number
- CN202411845482.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-12-16
AI Technical Summary
Traditional image repair methods fail to effectively utilize the information of visible parts of the image to be repaired, resulting in poor image quality after repair, especially in edge processing and tone texture.
An image repair method based on latent masks is adopted, and an image repair model with different mask ratios is obtained by masking and pre-training the images in a large data set. Then select the optimal repair model according to the damage ratio of the image to be repaired and fine-tune it to improve the repair effect.
By utilizing the visible information of the image to be repaired, the quality and accuracy of image repair are significantly improved, the adaptability and generalization capabilities of the model are enhanced, and the repair cost and time are reduced.
Smart Images

Figure CN119295355B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image restoration, and particularly relates to an image restoration method, device, equipment and medium based on a latent mask. Background Art
[0002] In the field of image processing, image restoration is a crucial technology aimed at restoring damaged or missing parts of an image to make it look complete and natural. Traditional image restoration methods generally follow a fixed process: First, a large existing dataset is used to pre-train a model to enable it to have basic image understanding and generation capabilities; subsequently, for a specific application scenario or dataset, the model is fine-tuned to further improve its performance in that field.
[0003] However, these traditional methods have a significant limitation in the specific execution steps of image restoration. That is, after the model is pre-trained and fine-tuned, when performing actual image restoration operations, the weights of the model will remain unchanged. This strategy ignores the important information contained in the visible parts of the image to be restored. This information is of extremely high value to the model because it directly reflects the unique features and context information of the image, and is the key basis for the model to learn and restore the missing parts.
[0004] Since traditional methods fail to make full use of the information in these visible parts, in some cases, the quality of the restored image is not satisfactory. Specifically, the degree of fusion between the restored area and the surrounding environment is not high enough, the edge processing is not natural enough, or the overall tone and texture are significantly different from the original image. These problems not only affect the visual effect of the image, but also limit the application of image restoration technology in more practical scenarios. Summary of the Invention
[0005] Aiming at the problem that traditional image restoration methods cannot effectively utilize the information in the visible parts of the image to be restored, resulting in poor quality of the restored image in some cases, the present invention provides an image restoration method, device, equipment and medium based on a latent mask.
[0006] In a first aspect, the technical solution of the present application provides an image restoration method based on a latent mask, including:
[0007] Mask the images in the large dataset and group them according to different mask ratios, and then pre-train the encoder-decoder architecture model to obtain image restoration models with different mask ratios;
[0008] Evaluate the image restoration models with different mask ratios to obtain the optimal image restoration model for each group of mask ratio images;
[0009] Evaluate the damaged ratio of the image to be repaired, and select the optimal image repair model corresponding to the masked ratio image based on the damaged ratio as the basic model for the image to be repaired;
[0010] After masking the visible part of the image to be repaired, fine-tune the basic model;
[0011] Input the image to be repaired into the fine-tuned model for image repair.
[0012] As a further limitation of the technical solution of the present invention, the steps of pre-training the encoder-decoder architecture model to obtain image repair models with different masked ratios after masking the images in the large dataset and grouping them according to different masked ratios include:
[0013] Encode the images in the large dataset using the encoder based on the attention mechanism;
[0014] Select some images in the large dataset for masking and group them according to different masked ratios;
[0015] Use the decoder to restore the masked part of the images with each masked ratio;
[0016] Obtain the difference between the image restored by the decoder and the original masked part of the image, which is represented by the mean squared error loss function; and optimize the loss function by the gradient descent method;
[0017] When it is determined that the difference between the image restored by the decoder and the original masked part of the image is less than the set threshold, the training is completed to obtain image repair models with different masked ratios; wherein, the image repair model contains basic repair knowledge.
[0018] By encoding the image through an encoder based on the attention mechanism, the model can capture key information and features in the image more accurately. This helps to more precisely restore the masked part during the subsequent decoding process, thereby improving the accuracy of image inpainting. Group training is performed on images with different masking ratios, enabling the model to adapt to various degrees of image damage. This training method enhances the generalization ability of the model, enabling it to give relatively satisfactory inpainting results when faced with images of different damage levels. The mean squared error loss function is used to represent the difference between the image restored by the decoder and the original masked part of the image, and the loss function is optimized using the gradient descent method. This training strategy can efficiently guide the update of model parameters, enabling the model to continuously approach the optimal solution during training, thereby optimizing the training process. The image inpainting model obtained through pre-training contains basic inpainting knowledge, which provides strong support for subsequent model fine-tuning. In practical applications, according to the damage situation of the image to be inpainted, the corresponding image inpainting model can be selected for fine-tuning, thereby further improving the pertinence and accuracy of inpainting.
[0019] As a further limitation of the technical solution of the present invention, the steps of encoding the images in the large dataset by the encoder based on the attention mechanism include:
[0020] The picture is cut into multiple small blocks and each small block is straightened into a one-dimensional vector;
[0021] The one-dimensional vectors are combined to obtain a two-dimensional vector;
[0022] The two-dimensional vector is subjected to three matrix operations to obtain three basic matrices Q, K, and V for the attention mechanism operations; enabling each original one-dimensional vector to contain the information of the entire picture.
[0023] As a further limitation of the technical solution of the present invention, the steps of using the decoder to restore the masked part of the images with each masking ratio include:
[0024] The decoder reversely generates a two-dimensional vector from the one-dimensional vector, and restores the two-dimensional vector back to multiple small picture blocks and then splices them into a picture.
[0025] As a further limitation of the technical solution of the present invention, the steps of evaluating the image inpainting models with different masking ratios to obtain the optimal image inpainting model for each group of masked ratio images include:
[0026] Perform masking processing on some images with a set masking ratio, input the decoder of the image inpainting model for images with different masking ratios to restore the masked images and calculate the difference loss value;
[0027] Select the image inpainting model with the smallest difference value as the optimal image inpainting model for the images with this masking ratio.
[0028] As a further limitation of the technical solution of the present invention, after masking the visible part of the image to be repaired, the steps of fine-tuning the basic model include:
[0029] Take the visible part of the image to be repaired as the original image, mask the original image, and let the decoder of the basic model predict the masked part and calculate the difference value. Use the mean square error loss function for evaluation, and use the gradient descent algorithm to optimize the loss function until the model contains the feature information of the picture to obtain the fine-tuned model. The fine-tuned model contains the feature information of the repaired picture.
[0030] As a further limitation of the technical solution of the present invention, the steps of inputting the image to be repaired into the fine-tuned model for image repair include;
[0031] Receive the image to be repaired, and based on the basic repair knowledge and the feature information of the repaired picture, the decoder restores the part of the image that needs to be repaired and outputs the repaired image.
[0032] In a second aspect, the technical solution of the present invention also provides an image repair device based on latent masking, including a pre-training module, an optimal matching module, a basic model confirmation module, a fine-tuning module, and an execution repair module;
[0033] The pre-training module is used to mask the images in the large data set and group them according to different masking ratios, and then pre-train the encoder-decoder architecture model to obtain image repair models with different masking ratios;
[0034] The optimal matching module is used to evaluate the image repair models with different masking ratios to obtain the optimal image repair model for each group of masking ratio images;
[0035] The basic model confirmation module is used to evaluate the damaged ratio of the image to be repaired, and based on the damaged ratio, select the optimal image repair model of the corresponding masking ratio image as the basic model of the image to be repaired;
[0036] The fine-tuning module is used to mask the visible part of the image to be repaired and then fine-tune the basic model;
[0037] The execution repair module is used to input the image to be repaired into the fine-tuned model for image repair.
[0038] As a further limitation of the technical solution of the present invention, the pre-training module includes an encoding unit, a masking processing unit, a decoding unit, a loss optimization unit, and an image repair model output unit;
[0039] The encoding unit is used to encode the images in the large data set based on the encoder of the attention mechanism;
[0040] A mask processing unit, which is used to select some images in the large dataset for mask processing and group them according to different mask ratios;
[0041] A decoding unit, which is used to use a decoder to restore the masked part of the images with each group of mask ratios;
[0042] A loss optimization unit, which is used to obtain the difference between the image restored by the decoder and the image of the originally masked part, expressed by the mean squared error loss function; and optimize the loss function by the gradient descent method;
[0043] An image inpainting model output unit, which is used to determine that when the difference between the image restored by the decoder and the image of the originally masked part is less than a set threshold, the training is completed, and an image inpainting model with different mask ratios is obtained; wherein, the image inpainting model contains basic inpainting knowledge.
[0044] The present invention ensures that the model can learn the information of the picture to be inpainted before inpainting the image through the latent mask technology, rather than only based on the previous dataset. The application of the contrast loss function: solves the problem of unbalanced training data and improves the robustness of the model.
[0045] As a further limitation of the technical solution of the present invention, an encoding unit is specifically used to cut the picture into multiple small blocks and straighten each small block into a one-dimensional vector; combine the one-dimensional vectors to obtain a two-dimensional vector; perform three matrix operations on the two-dimensional vector to obtain three basic matrices for the attention mechanism operations of Q, K, and V; make each original one-dimensional vector contain the information of the entire picture.
[0046] As a further limitation of the technical solution of the present invention, a decoding unit is specifically used to reversely generate a two-dimensional vector from the one-dimensional vector, and restore the two-dimensional vector back to multiple small picture blocks and then splice them into a picture.
[0047] As a further limitation of the technical solution of the present invention, an optimal matching module is specifically used to perform mask processing on some images with a set mask ratio, input the decoder of the image inpainting model with different mask ratio images to restore the masked image and calculate the difference loss value; select the image inpainting model with the smallest difference value as the optimal image inpainting model for the image with this mask ratio.
[0048] As a further limitation of the technical solution of the present invention, a fine-tuning module is specifically used to use the visible part of the image to be inpainted as the original image, mask the original image, and let the decoder of the basic model predict the masked part and calculate the difference value, evaluate it using the mean squared error loss function, and optimize the loss function using the gradient descent algorithm until the model contains the feature information of the picture to obtain a fine-tuned model, and the fine-tuned model contains the feature information of the inpainted picture.
[0049] As a further limitation of the technical solution of the present invention, the repair module is specifically configured to receive an image to be repaired, and based on basic repair knowledge and the repair picture feature information decoder, restore the part of the image that needs to be repaired and output the repaired image.
[0050] In a third aspect, the technical solution of the present invention further provides an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to execute the potential mask-based image repair method as described in the first aspect.
[0051] In a fourth aspect, the technical solution of the present invention further provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the potential mask-based image repair method as described in the first aspect.
[0052] It can be seen from the above technical solutions that the present invention has the following advantages: By processing the images in the large dataset with different mask ratios and separately training image repair models with different mask ratios, it can be ensured that the model has stronger adaptability to images with different degrees of damage. When repairing an image to be repaired, the optimal repair model can be selected according to its damage ratio, thereby significantly improving the repair effect. By pre-training on a large-scale dataset and performing grouped training for different mask ratios, this method can obtain multiple image repair models with different generalization abilities. These models can jointly handle image repair tasks in various complex scenarios and improve the generalization performance of the model.
[0053] This method can select the optimal repair model for fine-tuning according to the damage ratio of the image to be repaired. This adaptive repair mechanism enables this method to flexibly handle images with different degrees of damage and improve the flexibility and accuracy of repair. Since this method adopts a combination of pre-training and fine-tuning, there is no need to train the model from scratch when repairing new images, which greatly reduces the repair cost and time. At the same time, since the model has been fully pre-trained, it can converge quickly in the fine-tuning stage, further improving the repair efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 It is a schematic flow chart of the method according to an embodiment of the present invention.
[0056] Figure 2 It is a schematic block diagram of the device according to an embodiment of the present invention. Detailed implementation manners
[0057] In traditional image inpainting methods, after model pre-training and fine-tuning, the model weights are no longer changed during the image inpainting step. This approach cannot effectively utilize the information in the visible part of the image to be inpainted, and this information is the most helpful for the model to learn the features of the image, resulting in poor-quality inpainted images in some cases. For this reason, this application proposes a method in which the model first learns the information in the visible part of the image to be inpainted and then performs image inpainting. To enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0058] To facilitate a clear description of the technical solutions in the embodiments of this application, the following briefly introduces some terms and technologies involved in the embodiments of this application:
[0059] Encoder: In an encoder-decoder architecture, the encoder is responsible for encoding the input data (such as an image) into a fixed-length continuous representation, usually for feature extraction.
[0060] Masked: A technique used in an encoder-decoder architecture to train the model by covering part of the data, enabling it to predict or recover the masked data.
[0061] Latent Masked: First mask and decode the visible part of the image to be inpainted, and then repair the missing part of the image after training is completed.
[0062] Decoder: Corresponding to the encoder, the decoder is responsible for decoding the output of the encoder (i.e., the latent representation) back into the form of the original data, or generating new data, such as the missing part in image inpainting.
[0063] Image Inpainting: An image processing technique that focuses on filling in the missing or damaged parts of an image to restore the integrity and beauty of the image.
[0064] Self-attention mechanism: A mechanism that allows the model to dynamically focus on other elements in the data sequence when processing the data sequence, which is particularly useful when processing images because it can capture long-range dependencies between different parts of the image.
[0065] The technical solution of this application provides an image inpainting method based on a latent mask. First, by adding a masking technique, the model of the encoder-decoder architecture is trained to learn how to predict the missing image. Then, based on the visible part of the pre-inpainted image, mask-based fine-tuning training is used to learn the information of the image. Finally, the decoder predicts the image of the damaged part. A model with better inpainting effect is obtained, avoiding the problem of the decline in transfer ability caused by only training on the pre-training dataset in the traditional image inpainting method, and effectively improving the model's learning of the prior knowledge of the pre-inpainted image; as Figure 1 shown, the method includes:
[0066] Step 1: After masking the images in the large dataset and grouping them according to different masking ratios, pre-train the encoder-decoder architecture model to obtain image inpainting models with different masking ratios;
[0067] This step specifically includes:
[0068] Step 11: Encode the images in the large dataset by the encoder based on the attention mechanism;
[0069] Step 12: Select some images in the large dataset for masking and group them according to different masking ratios;
[0070] Step 13: Use the decoder to restore the masked part of the images with each masking ratio;
[0071] Step 14: Obtain the difference between the image restored by the decoder and the original masked part of the image, which is represented by the mean squared error loss function; and optimize the loss function by the gradient descent method;
[0072] Step 15: When the difference between the image restored by the decoder and the original masked part of the image is less than the set threshold, the training is completed to obtain image inpainting models with different masking ratios; among them, the image inpainting model contains basic inpainting knowledge.
[0073] Step 2: Evaluate the image inpainting models with different masking ratios to obtain the optimal image inpainting model for each group of images with a masking ratio;
[0074] Specifically, this step includes the following:
[0075] Perform masking processing on part of the image by setting the masking ratio, and input the image restoration model decoder of images with different masking ratios to restore the masked image and calculate the difference loss value;
[0076] Select the image restoration model with the smallest difference value as the optimal image restoration model for the image with this masking ratio.
[0077] Step 3: Evaluate the damaged ratio of the image to be restored, and select the optimal image restoration model of the corresponding masking ratio image as the basic model for this image to be restored;
[0078] Step 4: After masking the visible part of the image to be restored, fine-tune the basic model;
[0079] Take the visible part of the image to be restored as the original image, mask this original image, and let the decoder of the basic model predict the masked part and calculate the difference value. Use the mean square error loss function for evaluation, and use the gradient descent algorithm to optimize the loss function until the model contains the feature information of the picture to obtain the fine-tuned model. The fine-tuned model contains the feature information for restoring the picture.
[0080] Step 5: Input the image to be restored into the fine-tuned model for image restoration. Specifically, receive the image to be restored, and based on the basic restoration knowledge and the feature information of the restored picture, the decoder restores the part of the image that needs to be restored and outputs the restored image.
[0081] By encoding the image through an encoder based on the attention mechanism, the model can more accurately capture the key information and features in the image. This helps to more precisely restore the masked part in the subsequent decoding process, thereby improving the accuracy of image restoration. Group training is performed on images with different masking ratios, enabling the model to adapt to various degrees of image damage. This training method enhances the generalization ability of the model, enabling it to give relatively satisfactory restoration results when facing images with different degrees of damage. Use the squared error loss function to represent the difference between the image restored by the decoder and the original masked part of the image, and optimize the loss function through the gradient descent method. This training strategy can efficiently guide the update of model parameters, enabling the model to continuously approach the optimal solution during training, thereby optimizing the training process. The image restoration model obtained through pre-training contains basic restoration knowledge, which provides strong support for subsequent model fine-tuning. In practical applications, according to the damaged situation of the image to be restored, the corresponding image restoration model can be selected for fine-tuning, thereby further improving the pertinence and accuracy of restoration.
[0082] In some embodiments, the steps of encoding the images in the large dataset by the encoder based on the attention mechanism include:
[0083] Slice the image into multiple small pieces and straighten each small piece into a one-dimensional vector;
[0084] Combine the one-dimensional vectors to obtain a two-dimensional vector;
[0085] Perform three matrix operations on the two-dimensional vector to obtain three basic matrices for the attention mechanism operations of Q, K, and V; enabling each original one-dimensional vector to contain the information of the entire image.
[0086] Correspondingly, the steps of using the decoder to restore the masked part of the image for each group of mask ratios include:
[0087] The decoder reversely generates a two-dimensional vector through the one-dimensional vector, and restores the two-dimensional vector back to multiple small image blocks and then stitches them into an image.
[0088] It should be noted that the present invention constructs an image restoration scheme divided into two key stages: training and repair. This method is applicable to the encoder-decoder architecture model based on the attention mechanism, which is one of the current mainstream architecture models and has certain applicability.
[0089] Before introducing the training, first introduce the structure and function of the model. The model applicable to this method is the encoder-decoder architecture model. And the overall model realizes information extraction based on the attention mechanism.
[0090] First, introduce the encoder of the model. The role of the encoder is to extract information and concentrate the information. An image is usually considered as a three-dimensional matrix, while the encoder based on the attention mechanism can only process two-dimensional matrices. So the first approach is to perform matrix transformation. The common practice is to slice the image into multiple small pieces, then straighten each small piece into a one-dimensional vector, and then combine these one-dimensional vectors into a two-dimensional vector. This two-dimensional vector can be regarded as the most primitive data information, and it first needs to go through 3 matrix operations to obtain three basic matrices of Q, K, and V for the attention mechanism operations. The Q, K, and V matrices have the same matrix specifications as the original matrix. Among them, Q represents the information that should be focused on, K represents the information provided, and V represents the information that can be obtained after focusing. The formula of the attention mechanism is:
[0091]
[0092] Calculate the attention scores using the Q and K matrices. This score indicates how much attention should be given to each element in the sequence when generating the output of the current element. This is achieved by calculating the dot product of the query and key vectors and then applying the softmax function, which converts these scores into a probability distribution. dKis the dimension of the key vector. The output at each position is a weighted sum of the value vectors, where the weights are determined by the attention scores calculated in the previous step. This means that each output vector is a weighted combination of all elements in the sequence, and the weights reflect the relevance of each element to the current element. In this way, the output vector of each patch is not just its own representation, but a weighted representation of the entire sequence (in this case, the entire image). So after the operation of the attention mechanism, each original one-dimensional vector is considered to contain information about the entire picture.
[0093] A single encoder layer often does not extract information very thoroughly. Therefore, the encoder layer as a whole usually consists of multiple operations of the attention mechanism to ensure the accuracy of information extraction. The two-dimensional vector information extracted by multiple encoder layers is compressed into a one-dimensional vector and contains information about the entire picture. Then the decoder of the model is introduced. The role of the decoder is exactly the opposite of that of the encoder. Its main task is to generate a two-dimensional vector from the one-dimensional vector in reverse, and then these two-dimensional vectors are restored to multiple small image patches and stitched together to form an image. The principle of being able to repair the mask and the image at the same time lies in that the scale can be larger than the original input when becoming a two-dimensional vector, so as to generate some information to fill in the places masked or in need of repair in the image.
[0094] In the training stage, the model first conducts training on a large public dataset. The training method is to mask a part of the image and then use the decoder to try to restore the masked part of the image. The difference between the image restored by the decoder and the originally masked part of the image during this process is called the training loss. In order to more accurately quantify this loss value, a more suitable loss function is used for calculation. Currently, the commonly used loss function for image repair is the mean squared error loss function.
[0095] The mean squared error loss function is a commonly used performance function in the training of artificial neural networks. It is used to measure the difference between the predicted output and the actual target value. The formula for the mean squared error loss function is shown as follows.
[0096]
[0097] In the formula, SELoss represents the mean squared error. p represents the total number of training samples, that is, the number of samples or instances in the dataset used to train the neural network. k represents the total number of output units, which corresponds to the number of output neurons in the last layer of the neural network, that is, the number of predicted target variables. represents the k-th target value of the p-th training pattern. It represents the k-th expected output of the p-th training pattern. represents the k-th output obtained from the p-th pattern.
[0098] By applying the gradient descent method, the image generated by the decoder becomes more and more like the original appearance of the masked part. The gradient descent algorithm selected is the stochastic gradient descent algorithm, and the two cores of the stochastic gradient descent algorithm are the learning rate and momentum. Generally, when the model is pre-trained, a relatively large learning rate is adopted, while during the fine-tuning period, a smaller learning rate is required.
[0099] During the entire training period, in order to perfectly adapt to different scenarios and various image inpainting tasks with different repair difficulties, multiple masking ratios must be carefully adjusted and deeply trained. The common increasing masking ratio is 10%. Taking the masking ratio of 10% as a group, multiple groups of control experiments are carried out respectively to select the best masking ratio suitable for different scenarios. The preset masking table is shown in Table 1.
[0100] Table 1: Masking Table
[0101]
[0102] In terms of the selection of the masking style, it is recommended to use the random masking method. When using random masking, after shuffling the mask, the front part of the mask can be fixed. In this way, the ideal effect with a time complexity close to a constant can be successfully achieved, thus effectively avoiding the masking step from occupying too much time resources. At the same time, during the use of random masking, position information needs to be added to the attention model. Adding position information can ensure that the shuffling operation will not have an adverse impact on the accuracy of the model.
[0103] After training is completed on a large-scale dataset, an image inpainting model is obtained. At this time, multiple image inpainting models with different masking ratios are obtained, so in-depth evaluation needs to be carried out on image inpainting tasks with different ratios. The evaluation process is also carried out on a public dataset. At this stage, although a part of the image is still masked, and then the decoder is allowed to restore this part of the image and calculate the difference loss value, at this time, the model will no longer perform the gradient descent operation. The difference value only exists as an important evaluation criterion at this stage, and its role is to accurately judge how much masking ratio is more suitable for inpainting images damaged to a specific degree. Finally, the most suitable pre-training masking ratio for a certain scenario can be judged. For example, when inpainting an image with a 10% damage ratio, the models with masking ratios of 10%, 20%, and 30% during training can be compared to select the model with the best accuracy. Thanks to the fact that the operation matrices Q, K, and V of the attention mechanism all come from the initial matrix and have the same specifications, different masking ratios for training and evaluation will not cause problems in the training and evaluation of the model. When performing subsequent inpainting, the ratio of the image to be inpainted can be evaluated first, and then the model with the best accuracy can be selected.
[0104] During the repair phase of the model, the images used in the repair task should be pictures that the model has never encountered during the training phase. In this way, the real scenario of repairing images can be simulated highly realistically.
[0105] When performing image repair, it is first necessary to evaluate the damaged proportion of the picture. Then, select the model with the best accuracy for this proportion in the previous evaluation as the basic model. Then let the model start to learn the basic information of the image to be repaired, and this step is achieved through fine-tuning. Fine-tuning is a technique in the field of machine learning. Its basic idea is to further train on the basis of a pre-trained model (such as BERT, GPT, etc.) to adapt to specific downstream tasks (such as text classification, sentiment analysis, etc.). In this experiment, fine-tuning is to learn the basic knowledge of the image to be repaired.
[0106] During fine-tuning, first take the visible part of the image to be repaired as the original image, then mask this image, and let the decoder predict the masked part. At this time, the prediction difference value is still evaluated using the mean squared error loss function, and the same stochastic gradient descent is used to reduce this loss value. After fine-tuning for a period of time, the model basically knows the information contained in this image. At this time, perform the image repair task. Then, use the repaired picture as the training set for fine-tuning. During this process, the decoder needs to restore the masked part and evaluate the loss value of the model and perform operations related to gradient descent. (In this way, the model can fully learn the feature information of this picture so that it can repair this picture more excellently.
[0107] After training is completed, no masking operation will be performed on the picture anymore, and the decoder needs to restore the part of the picture that needs to be repaired. Based on the basic repair knowledge learned from pre-training and the repair picture feature information learned during fine-tuning, the repair effect can be better than the traditional repair effect. When repairing an image, it is necessary to save the pre-trained weights as a basic weight, and then load, re-train, and re-evaluate each time the user calls. This way, there is no need to save multiple weights, saving space and ensuring that the model can repair based on the feature information of the repaired picture.
[0108] As Figure 2 shown, an embodiment of the present invention further provides an image repair device based on latent masking, including a pre-training module, an optimal matching module, a basic model confirmation module, a fine-tuning module, and an execution repair module;
[0109] The pre-training module is used to perform masking processing on the images in the large dataset and group them according to different masking ratios, and then pre-train the encoder-decoder architecture model to obtain image repair models with different masking ratios;
[0110] The optimal matching module is used to evaluate image restoration models with different mask ratios to obtain the optimal image restoration model for each group of images with mask ratios.
[0111] The basic model confirmation module is used to evaluate the damage ratio of the image to be restored, and select the optimal image restoration model of the corresponding mask ratio image as the basic model of the image to be restored based on the damage ratio.
[0112] The fine-tuning module is used to perform mask processing on the visible part of the image to be restored, and then fine-tune the basic model.
[0113] The execution repair module is used to input the image to be restored into the fine-tuned model for image restoration.
[0114] In some embodiments, the pre-training module includes an encoding unit, a mask processing unit, a decoding unit, a loss optimization unit, and an image restoration model output unit.
[0115] The encoding unit is used to encode the images in the large dataset based on the encoder of the attention mechanism.
[0116] The mask processing unit is used to select some images in the large dataset for mask processing and group them according to different mask ratios.
[0117] The decoding unit is used to use the decoder to restore the masked part of the images with each mask ratio.
[0118] The loss optimization unit is used to obtain the difference between the image restored by the decoder and the original masked part of the image, which is represented by the mean squared error loss function; and optimize the loss function by the gradient descent method.
[0119] The image restoration model output unit is used to determine that when the difference between the image restored by the decoder and the original masked part of the image is less than the set threshold, the training is completed, and image restoration models with different mask ratios are obtained; among them, the image restoration model contains basic restoration knowledge.
[0120] In some embodiments, the encoding unit is specifically used to cut the picture into multiple small blocks and straighten each small block into a one-dimensional vector; combine the one-dimensional vectors to obtain a two-dimensional vector; perform three matrix operations on the two-dimensional vector to obtain three basic matrices for the attention mechanism operations of Q, K, and V; so that each original one-dimensional vector contains the information of the entire picture. The decoding unit is specifically used to generate a two-dimensional vector in reverse through the one-dimensional vector, and restore the two-dimensional vector back to multiple small picture blocks and then splice them into a picture.
[0121] In some embodiments, the optimal matching module is specifically configured to perform masking processing on a partial image with a set masking ratio, input the decoder of the image restoration model with images of different masking ratios to restore the masked image and calculate the difference loss value; and select the image restoration model with the smallest difference value as the optimal image restoration model for the image with this masking ratio.
[0122] In some embodiments, the fine-tuning module is specifically configured to use the visible part of the image to be restored as the original image, mask the original image, and let the decoder of the basic model predict the masked part and calculate the difference value, evaluate it using the mean square error loss function, and use the gradient descent algorithm to optimize the loss function until the model contains the feature information of the picture to obtain the fine-tuned model, and the fine-tuned model contains the feature information of the restored picture.
[0123] In some embodiments, the execution restoration module is specifically configured to receive the image to be restored, and based on the basic restoration knowledge and the feature information decoder of the restored picture, restore the part of the image that needs to be restored and output the restored image.
[0124] An embodiment of the present invention further provides an electronic device, which includes: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The communication bus can be used for information transmission between the electronic device and the sensor. The processor can call the logical instructions in the memory to execute the following method: Step 1: Perform masking processing on the images in the large dataset and group them according to different masking ratios, and then pre-train the encoder-decoder architecture model to obtain image restoration models with different masking ratios; Step 2: Evaluate the image restoration models with different masking ratios to obtain the optimal image restoration model for each group of images with a masking ratio; Step 3: Evaluate the damage ratio of the image to be restored, and based on the damage ratio, select the optimal image restoration model of the corresponding masking ratio image as the basic model for the image to be restored; Step 4: After masking the visible part of the image to be restored, fine-tune the basic model; Step 5: Input the image to be restored into the fine-tuned model for image restoration.
[0125] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0126] An embodiment of the present invention provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and these computer instructions cause the computer to execute the methods provided in the above method embodiments. For example, it includes: Step 1: Perform masking processing on the images in the large dataset and group them according to different masking ratios, and then pre-train the encoder-decoder architecture model to obtain image inpainting models with different masking ratios; Step 2: Evaluate the image inpainting models with different masking ratios to obtain the optimal image inpainting model for each group of masking ratio images; Step 3: Evaluate the damaged ratio of the image to be inpainted, and based on the damaged ratio, select the optimal image inpainting model corresponding to the masking ratio image as the basic model for the image to be inpainted; Step 4: Perform masking processing on the visible part of the image to be inpainted, and then fine-tune the basic model; Step 5: Input the image to be inpainted into the fine-tuned model for image inpainting.
[0127] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0128] An embodiment of the image inpainting device based on latent masking provided by an embodiment of the present invention. This device and the image inpainting method based on latent masking in the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the image inpainting device based on latent masking, reference can be made to the embodiment of the image inpainting method based on latent masking.
[0129] Although the present invention has been described in detail by referring to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, and all should be covered within the protection scope of the present invention.
Claims
1. A latent mask-based image restoration method, characterized in that: include: After masking the images in the large data set and grouping them according to different mask ratios, the encoder-decoder architecture model is pre-trained to obtain image restoration models with different mask ratios. The image inpainting models with different mask ratios are evaluated to obtain the optimal image inpainting model for each group of mask ratio images; specifically, the method comprises: performing mask processing with a set mask ratio on part of the images, inputting the decoder of the image inpainting model with different mask ratio images to restore the masked images and calculate the difference loss value; and selecting the image inpainting model with the smallest difference value as the optimal image inpainting model for the mask ratio image; Evaluate the damage ratio of the image to be repaired, and select the optimal image repair model corresponding to the mask ratio image as the basic model of the image to be repaired based on the damage ratio; After masking the visible part of the image to be repaired, fine-tuning the basic model; specifically, the method includes: taking the visible part of the image to be repaired as the original image, masking the original image, and letting the decoder of the basic model predict the masked part and calculate the difference value, evaluating it using the mean square error loss function, and optimizing the loss function using the gradient descent algorithm until the model contains the feature information of the picture to obtain a fine-tuned model, and the fine-tuned model contains the feature information of the repaired picture; The image to be repaired is input into the fine-tuned model for image repair.
2. The image restoration method based on latent mask according to claim 1, characterized in that: After masking the images in the large data set and grouping them according to different mask ratios, the encoder-decoder architecture model is pre-trained to obtain the image restoration model with different mask ratios. The steps include: The encoder based on the attention mechanism encodes the images in the large data set; Select some images from the large data set for mask processing and group them according to different mask ratios; Use the decoder to restore the masked portion of the image for each set of mask ratios; Obtain the difference between the image restored by the decoder and the image of the original masked part, and express it using the square error loss function; and optimize the loss function through the gradient descent method; When the difference between the image restored by the decoder and the image of the original masked part is less than a set threshold, the training is completed and an image restoration model with different mask ratios is obtained; wherein the image restoration model includes basic restoration knowledge.
3. The image restoration method based on latent mask according to claim 2, characterized in that: The steps of encoding images in a large data set based on the attention mechanism encoder include: Divide the image into multiple small blocks and straighten each small block into a one-dimensional vector; Combining the one-dimensional vectors to obtain a two-dimensional vector; Perform three matrix operations on the two-dimensional vector to obtain the basic matrices Q, K, and V for the three attention mechanism operations; so that each original one-dimensional vector contains the information of the entire image.
4. The image restoration method based on latent mask according to claim 3, characterized in that: The steps of restoring the masked portion of the image of each set of mask ratios using the decoder include: The decoder generates a two-dimensional vector by reversely converting the one-dimensional vector, and restores the two-dimensional vector back to multiple small image blocks and then splices them into an image.
5. The image restoration method based on latent mask according to claim 4, characterized in that: The steps of inputting the image to be repaired into the fine-tuned model for image repair include: The image to be repaired is received, and based on basic repair knowledge and repair picture feature information decoder, the part of the image that needs to be repaired is restored and the repaired image is output.
6. An image restoration device based on latent mask, characterized in that: It includes pre-training module, optimal matching module, basic model confirmation module, fine-tuning module and execution repair module; The pre-training module is used to mask the images in the large data set and group them according to different mask ratios, and then pre-train the encoder-decoder architecture model to obtain image restoration models with different mask ratios; The optimal matching module is used to evaluate the image restoration models with different mask ratios to obtain the optimal image restoration model for each group of mask ratio images; specifically, it is used to perform mask processing with a set mask ratio on part of the image, input the decoder of the image restoration model with different mask ratio images to restore the masked image and calculate the difference loss value; select the image restoration model with the smallest difference value as the optimal image restoration model for the mask ratio image; the basic model confirmation module is used to evaluate the damaged ratio of the image to be restored, and select the optimal image restoration model of the corresponding mask ratio image as the basic model of the image to be restored based on the damaged ratio; A fine-tuning module is used to fine-tune the basic model after masking the visible part of the image to be repaired; specifically, it is used to take the visible part of the image to be repaired as the original image, mask the original image, and let the decoder of the basic model predict the masked part and calculate the difference value, use the mean square error loss function for evaluation, and use the gradient descent algorithm to optimize the loss function until the model contains the feature information of the picture to obtain a fine-tuned model, and the fine-tuned model contains the feature information of the repaired picture; The execution restoration module is used to input the image to be restored into the fine-tuned model for image restoration.
7. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the latent mask-based image restoration method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the latent mask-based image restoration method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Portrait picture restoration method, device and equipment
CN111127366A
Restoration model adjustment method and device, equipment and storage medium
CN116977195A