Image restoration model training method and device
By degrading the image and text enhancement processing, training sample pairs are generated, and image repair models are trained using the generative adversarial network to solve the defect problem in text image repair, and the generalization ability and repair effect of the model are improved.
Patent Information
- Application Number
- CN202510369971.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
When processing images with text, existing image repair solutions are prone to text defects, such as strange distortions, white edges, etc., and the repair effect is not good.
By performing degradation processing on images and text enhancement processing, training sample pairs are generated, and image repair models are trained using generative adversarial networks to simulate various text addition methods to improve the generalization ability and robustness of the model.
It significantly improves the ability of the image repair model to identify and repair text, prevent text defects, achieve better repair effects, and avoid text distortion and video jitter.
Smart Images

Figure CN120298262A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of Internet technologies, and in particular, to a method and apparatus for training an image restoration model. Background Art
[0002] Image restoration refers to restoring the blurred, unclear or damaged parts of an image so that the overall restored image is as close as possible to the original image.
[0003] In some cases, users may need to use the editing function to add text to an image, which may result in the situation where the image is blocked by the text. However, the existing image restoration solutions have poor restoration effects on images with added text. For example, there are various defects in the text after image restoration, such as strange distortion changes, white edges, etc. How to improve the restoration effect of images with text has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of the above problems, the present application is proposed to provide an image restoration model training method, apparatus, computing device, computer storage medium and computer program product that overcome the above problems or at least partially solve the above problems.
[0005] According to one aspect of the embodiments of the present application, there is provided an image restoration model training method, including:
[0006] Obtain a first image, and perform degradation processing on the first image to obtain a second image associated with the first image;
[0007] Perform text enhancement processing on the first image and the second image associated with the first image respectively to obtain a third image and a fourth image associated with the third image, where the third image is an image after adding text content to the first image, and the fourth image is an image after adding text content with the same text attributes at the same position of the second image associated with the first image. Each pair of the third image and the fourth image associated with the third image is a training sample pair for training the image restoration model;
[0008] Train a generative adversarial network according to the training sample pair to obtain an image restoration model.
[0009] Further, performing text enhancement processing on the first image and the second image associated with the first image respectively to obtain a third image and a fourth image associated with the third image further includes:
[0010] Randomly determine the text content length according to a preset text length range, and randomly select the text content with the text content length from a preset text content library;
[0011] Add the text content to the first image and the second image respectively to obtain a third image and a fourth image associated with the third image.
[0012] Furthermore, the text attributes include: font color, font format, text border width, text border color;
[0013] Before adding the text content to the first image and the second image respectively, the method further includes:
[0014] Randomly select a font color from a preset font color library and apply the selected font color to the text content; and / or
[0015] Randomly select a font format from a preset font format library and apply the selected font format to the text content; and / or
[0016] Randomly select a text border color from a preset text border color library, randomly determine the text border width according to a preset text border width range, and add a border to the text content according to the text border color and the text border width.
[0017] Furthermore, the text attributes further include: text rotation angle;
[0018] Before adding the text content to the first image and the second image respectively, the method further includes:
[0019] Randomly select and determine the text content placement coordinates from a preset text content coordinate library, and randomly select a text rotation angle from a preset text rotation angle range;
[0020] Adding the text content to the first image and the second image respectively further includes:
[0021] Add the text content to the first image and the second image respectively according to the text content placement coordinates and the text rotation angle.
[0022] Furthermore, the generative adversarial network includes: a generator and a discriminator;
[0023] Training the generative adversarial network according to training samples to obtain an image inpainting model further includes:
[0024] Obtain the image label corresponding to the third image in the training sample pair;
[0025] Iteratively adversarially train the generator and the discriminator in the generative adversarial network according to the third image, the image label, and the fourth image until the training stops when the iteration stop condition is reached, and obtain the image inpainting model.
[0026] Further, the generator and discriminator in the generative adversarial network are iteratively adversarially trained based on the third image, the image label, and the fourth image, and the training is stopped until the iteration stop condition is reached, and obtaining the image restoration model further includes:
[0027] S1. Use the generator to perform image restoration processing on the fourth image to obtain the restored fourth image;
[0028] S2. Use the discriminator to perform authenticity discrimination on the restored fourth image to obtain an authenticity prediction result;
[0029] S3. Calculate a restoration loss function according to the authenticity prediction result and the image label corresponding to the third image;
[0030] S4. Adjust the parameters of the generator and the discriminator according to the restoration loss function;
[0031] Iteratively execute S1 - S4 until the iteration stop condition is reached to generate an image restoration model.
[0032] Further, the method further includes: obtaining an image to be restored, wherein text content is added to the image to be restored;
[0033] Input the image to be restored into the image restoration model for image restoration to obtain a restored image.
[0034] Further, the first image is a face image.
[0035] According to another aspect of the embodiments of the present application, there is provided an image restoration model training device, including:
[0036] An acquisition module, adapted to acquire a first image;
[0037] A degradation processing module, adapted to perform degradation processing on the first image to obtain a second image associated with the first image;
[0038] A text enhancement processing module, adapted to perform text enhancement processing on the first image and the second image associated with the first image respectively to obtain a third image and a fourth image associated with the third image, wherein the third image is the image after adding text content to the first image, and the fourth image is the image after adding text content with the same text attributes at the same position of the second image associated with the first image, and each group of the third image and the fourth image associated with the third image is a training sample pair for training the image restoration model;
[0039] A training module, adapted to train the generative adversarial network according to the training sample pairs to obtain an image restoration model.
[0040] According to another aspect of the embodiments of the present application, a computing device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0041] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above image restoration model training method.
[0042] According to still another aspect of the embodiments of the present application, a computer storage medium is provided, in which at least one executable instruction is stored, and the executable instruction causes the processor to perform the operations corresponding to the above image restoration model training method.
[0043] According to yet another aspect of the embodiments of the present application, a computer program product is provided, including at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above image restoration model training method.
[0044] The image restoration model training method and device provided according to the embodiments of the present application have diversity and randomness in the way of adding text, simulating various situations in which text may appear, which helps to recognize various complex situations during subsequent model training, realizes data augmentation, enables the model to have the ability to recognize text during the image restoration process, thereby significantly improving the generalization ability and robustness of the image restoration model, preventing the appearance of defective restoration results regarding text during the image restoration process, solving the problem of text defects that occur after the existing technology model repairs pictures, and can also help the model better recognize image features and achieve better restoration effects. For example, the image restoration model of the prior art has defects in text processing during image restoration, specifically manifested as changes in the internal color of the text, the appearance of white edges, and the blurring of the text lines. And due to the occlusion of the subtitles, the image restoration effect is also not good. However, the image restoration model of the present application has no impact on the text, and the image restoration effect is also better. The present application can avoid strange distortion changes of the text while restoring the image features, and at the same time, the effect is robust and there will be no situation of jitter in the restored video results.
[0045] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the embodiments of the present application more obvious and understandable, the following specifically describes the specific implementation manners of the embodiments of the present application. Description of the Drawings
[0046] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to limit the embodiments of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0047] Figure 1 A schematic flowchart of a method for training an image restoration model according to an embodiment of the present application is shown;
[0048] Figure 2 A schematic flowchart of a method for training an image restoration model according to another embodiment of the present application is shown;
[0049] Figure 3 A schematic diagram for placing coordinates of text content;
[0050] Figure 4 A schematic flowchart of a method for training a face restoration model according to still another embodiment of the present application is shown;
[0051] Figure 5 A block diagram of the structure of an image restoration model training device according to an embodiment of the present application is shown;
[0052] Figure 6 A schematic diagram of the structure of a computing device according to an embodiment of the present application is shown. Detailed Embodiments
[0053] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0054] First, the noun terms related to one or more embodiments of the present application are explained.
[0055] Image Restoration refers to restoring a blurred or damaged image through deep learning methods, which can make the blurred parts in the image clearer.
[0056] Face Restoration refers to restoring a blurred or damaged image through deep learning methods, which can make the blurred facial features and skin in the face clearer and improve the expression and details.
[0057] The degradation technique refers to using some technical methods (such as blurring, downsampling, adding noise, etc.) to simulate the process of image quality degradation in the real world, generating low-quality images from high-quality images.
[0058] Figure 1 FIG. shows a schematic flow chart of an image inpainting model training method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0059] Step S101, obtain a first image, perform degradation processing on the first image to obtain a second image associated with the first image.
[0060] Specifically, the first image is a high-quality image, and a high-quality image refers to an image that performs excellently in terms of resolution, clarity, color accuracy, contrast, noise control, etc. Therefore, the main image quality indicators include: resolution, clarity, color depth, color space, contrast, noise level, etc. The source of the first image has multiple ways, for example, an image obtained by shooting with an image acquisition device, or an image resource extracted from a specific image database, etc. The first image can be an image of any scene. For example, it can be a face image, a landscape image, an animal image, etc. Among them, the face in the face image is a false face and does not involve any real face.
[0061] Degradation processing is an image processing operation whose purpose is to artificially reduce the quality of an image, making the image produce an effect similar to the image quality degradation that may occur in actual application scenarios. Common degradation processing methods include but are not limited to adding noise, reducing resolution, blurring, and adjusting the color balance of the image. Therefore, degradation processing methods such as adding noise, reducing resolution, blurring, and adjusting the color balance of the image can be used to process the image data of the first image.
[0062] Among them, adding noise is to randomly introduce some interfering pixel points in the image to simulate the image noise phenomenon caused by factors such as sensor errors and signal transmission interference in reality; reducing resolution means reducing the number of pixels in the image, making the image become blurred. Among them, reducing resolution can be achieved by downsampling compression; blurring is to make the edges and details of the image unclear through convolution operations, etc., simulating situations such as camera shake and out-of-focus; adjusting the color balance of the image will cause the color of the image to deviate from the original true color.
[0063] After performing the above degradation processing on the first image, a second image associated with the first image can be obtained. The second image is a low-quality image after the degradation processing of the first image. It is closely associated with the first image and remains consistent in content, but shows a significant decline in image quality and has the characteristics after specific degradation processing. For example, if Gaussian noise is added to the first image during the degradation processing, obvious Gaussian-distributed noise points will appear in the second image; if the resolution is reduced, the second image will be more blurred than the first image, and more details will be lost.
[0064] By simulating a rich variety of low-quality image samples, it provides a rich set of negative samples for subsequent model training, enabling the subsequently trained image restoration model to better adapt to different types and degrees of image degradation problems, enhancing the generalization ability of the model, and also enabling the trained image restoration model to more effectively restore the original appearance of the image when dealing with real-world low-quality images.
[0065] Step S102: Perform text enhancement processing on the first image and the second image associated with the first image respectively to obtain a third image and a fourth image associated with the third image. Among them, the third image is the image after adding text content to the first image, and the fourth image is the image after adding text content with the same text attributes at the same position as the second image associated with the first image. Each pair of the third image and the fourth image associated with the third image is a training sample pair for training the image restoration model.
[0066] Text enhancement processing refers to the operation process of adding text content to an image. In this embodiment, the same text enhancement processing is performed on the first image and the second image associated with the first image. The text enhancement processing of the first image and the second image associated with the first image can be carried out simultaneously or successively. For example, text information can be added to the corresponding position of the first image first to obtain the third image. Here, the position can be any position in the image frame of the first image.
[0067] After adding text to the first image to obtain the third image, for the second image, text content with the same text attributes needs to be added at the same position where text is added to the first image to obtain the fourth image. Text attributes include the font, font size, color, style (such as bold, italic, underlined, etc.), text rotation angle, and text arrangement method of the text.
[0068] For example, if a text "xxx" in red, with a font size of 16 pixels, in Song typeface, and a transparency of 0.8 is added to the upper left corner of the first image, then at the upper left corner of the corresponding second image, the text content "xxx" also needs to be added with exactly the same text attributes. In this way, each pair of the third image and the fourth image that appear in pairs constitutes a pair of training samples for training the image restoration model.
[0069] By ensuring that the text content with the same text attributes is added at the same position in the first image and the second image, the third image and the fourth image are made to have consistency and relevance in terms of text addition. This relevance enables the model to learn, during subsequent model training, the feature differences of the same text content under different image qualities between the high-quality image (the third image) and the corresponding low-quality image (the fourth image), as well as how to restore from the low-quality image to the high-quality image state.
[0070] Step S103: Train a generative adversarial network according to the training sample pairs to obtain an image restoration model.
[0071] When training the image restoration model, these pairs of the third image and the fourth image associated with the third image constitute the training sample pairs.
[0072] The purpose of training is to use the high-quality third image with added text as a reference standard to train the model on how to restore the fourth image with the same text content but in a low-quality state to make it as close as possible to the visual effect and quality of the third image. During the training process, the model will learn the changes in image features in the fourth image due to degradation processing, as well as the impact of these changes and the added text content on the restoration. Through learning a large number of training sample pairs, the model gradually masters the ability to restore clear images from low-quality images, so that in practical applications, it can effectively restore images with reduced quality for various reasons and containing text content, avoiding problems and defects that occur during restoration and the poor overall restoration effect of the image, and improving the usability and readability of the image.
[0073] According to the image restoration model training method provided by the embodiments of the present application, by performing text enhancement on the image, the model is enabled to recognize text during the image restoration process, thereby significantly improving the generalization ability and robustness of the image restoration model, preventing text-related defective restoration results during the image restoration process, solving the problem of text defects in the existing technology after the model restores pictures, and also helping the model better recognize image features to achieve a better restoration effect, and making the image restoration technology more in line with the requirements of actual application scenarios.
[0074] Figure 2 shows a schematic flowchart of an image restoration model training method according to another embodiment of the present application, asFigure 2 As shown in the figure, the method includes the following steps:
[0075] Step S201: Obtain a first image, perform degradation processing on the first image to obtain a second image associated with the first image.
[0076] Specifically, there are various sources of the first image. For example, an image obtained by shooting with an image acquisition device (such as a photo taken by a digital camera, a video frame image, a video frame image recorded by a surveillance camera), or an image resource extracted from a specific image database, etc. These images have relatively high quality, and are characterized by rich detail information, accurate color restoration, and a resolution that can meet the needs of normal visual observation and information extraction.
[0077] After obtaining the first image, it is necessary to perform degradation processing on it to obtain a low-quality second image associated with the first image. Degradation processing is an operation that simulates the situation of image quality degradation in reality, aiming to generate a low-quality version of the image for subsequent training of the image restoration model. In this embodiment, methods such as applying a blur kernel, downsampling compression, and adding noise can be used to perform degradation processing on the high-quality first image.
[0078] Downsampling compression is a method of reducing the resolution by reducing the number of image pixels. Taking average pooling downsampling as an example, the image is divided into small blocks of the same size. Within each small block, the average value of all pixel values is calculated, and then this average value is used to represent the pixel value of the entire small block. This causes each small block to be compressed into a single pixel point, and the number of pixels in the horizontal and vertical directions of the image is significantly reduced. For example, an image with a resolution of 2000×2000 pixels, after 2×2 average pooling downsampling, the resolution becomes 1000×1000 pixels. In this process, every 4 adjacent original pixels are combined into 1 new pixel, and the new pixel value is the average value of these 4 original pixel values. For the image after downsampling compression, the detail information is significantly lost, the edges are no longer sharp, and the fine textures become blurred.
[0079] Applying a blur kernel is to use a two-dimensional matrix (i.e., the blur kernel) to perform weighted summation on the image pixels through convolution operations, changing the pixel values to achieve the image blurring effect. Different shapes and numerical distributions of the blur kernel will produce different degrees and styles of blurring. For example, a Gaussian blur kernel, whose numerical distribution conforms to the Gaussian function, has a large central value and becomes smaller as it gets farther from the central value. When using a Gaussian blur kernel to perform convolution operations on an image, each pixel in the image is multiplied by the value at the corresponding position of the Gaussian blur kernel, and the product results of the surrounding pixels are accumulated to finally obtain a new pixel value. This makes the differences between pixels in the image become smooth, and the original clear edges and details become blurred. For example, for a clear face image, after applying a Gaussian blur kernel, the facial contours and facial features of the face will become blurred.
[0080] Adding noise to an image introduces random interference at the pixel level. Common types of noise include Gaussian noise, salt-and-pepper noise, etc. Taking Gaussian noise as an example, it is a type of noise with a normal distribution characteristic. When adding Gaussian noise to an image, a random noise value is added to each pixel point according to the probability density function of the Gaussian distribution. The magnitude and sign of this noise value are both random, and the mean and variance of its distribution determine the intensity of the noise. When the noise intensity is high, the pixel values in the image are greatly affected, and the originally clear pixel values become chaotic after being interfered by the noise. For example, after adding strong Gaussian noise to a clear building image, the edges and surface textures of the building are covered by the noise and become blurred, and the overall resolution of the image also appears to decrease visually because the detailed information is difficult to distinguish after being interfered by the noise. Salt-and-pepper noise randomly sets some pixel values in the image to the maximum value (white) or the minimum value (black), just like sprinkling salt and pepper on the image. This destroys the continuity and details of the image, causing a large number of isolated black and white noise points to appear in the originally clear image, also reducing the resolution and visual quality of the image. By adding this noise, the resolution of the high-quality first image is reduced both visually and in terms of its actual information-carrying capacity, further simulating the quality degradation that an image may suffer in an actual scenario.
[0081] Through the above degradation processing operation, the originally high-quality first image is transformed into a low-quality second image. The second image is significantly lower than the first image in terms of resolution, clarity, detail presentation, etc., truly simulating the possible image quality degradation in an actual scenario, providing crucial low-quality sample data for the training of subsequent image restoration models so that the models can learn how to restore low-quality images to a high-quality state.
[0082] After obtaining the first image and the second image associated with the first image, it is necessary to perform text enhancement processing on the first image and the second image associated with the first image. For example, the methods in steps S202 - S207 can be used to achieve text enhancement processing. It should be noted that only some of the methods in the steps can also be used to achieve text enhancement processing.
[0083] Step S202, randomly determine the text content length according to the preset text length range, and randomly select the text content with the determined text content length from the preset text content library.
[0084] Specifically, when performing text enhancement processing on an image, it is necessary to obtain appropriate text content from the preset text content library and add it to the image.
[0085] The preset text length range can be set according to the actual application scenario and requirements. In the actual application scenario, the text is added to the image by using the editing function of the image software for a simple description of the image. Therefore, the text content will not be too long. To ensure that all possible text addition requirements can be fully met and to take into account the possibility of slightly longer explanatory text in special cases, the maximum value of the preset text length range can be set relatively loosely. For example, the preset text length range can be set from 1 to 20 characters. After determining the text length range, the length of the text content is randomly determined by a random number generation algorithm. For example, by using the random number generation function in a programming language, a random integer is generated within the preset text length range, and this random integer is the selected length of the text content.
[0086] To add text content, it is necessary to pre - construct a text content library. The text content library collects various text information, which can include texts in different languages, various characters, etc. For example, letters, Chinese characters, punctuation marks, special symbols, and numbers, etc. The text content library collects text content of different lengths.
[0087] After determining the length of the text content, appropriate - length text content can be randomly selected from the preset text content library for subsequent addition to the first image and the second image, so as to construct a more diverse and targeted training sample pair for training the image restoration model. When selecting, text content that matches the determined text content length can be preferentially selected. If not available, multiple text contents can be selected, and the total length of the multiple text contents is equal to the randomly determined text content length.
[0088] In this embodiment, a combination rule for text content can also be preset. After randomly selecting appropriate - length text content from the preset text content library, a combination rule can be randomly selected, and the selected text content is combined according to the selected combination rule. For example, when the selected text content includes Chinese characters, numbers, and letters, they can be combined in the following order: Chinese characters, letters, numbers. Here is just an example for illustration and has no restrictive effect.
[0089] Step S203: Randomly select a font color from the preset font color library and apply the selected font color to the text content.
[0090] Specifically, a font color library is pre - constructed. The font color library contains a rich variety of font colors, which cover common basic colors such as red, blue, green, black, white, etc., to various soft intermediate tones and bright colors.
[0091] In order to increase the richness of the training sample pairs, the font color of the selected text content may be set, that is, a font color may be randomly selected from a preset font color library and the selected font color may be applied to the text content.
[0092] For example, an index value can be set for each font color in the preset font color library. When randomly selecting a font color from the preset font color library, the specific color index can be determined with the help of a random number generation mechanism. Taking Python as an example, if the preset font color library is stored as a list, a random index value can be generated through a code such as random.randint(0,len(font_color_list)-1), which corresponds to a specific font color in the list. Once the font color is selected, it will be applied to the previously selected text content, so that the text content presents a randomly determined color.
[0093] Step S204: randomly select a font format from a preset font format library, and apply the selected font format to the text content.
[0094] The preset font format library is also a carefully prepared resource library, which contains many different styles of font formats. These font formats include common regular fonts such as Songti, Heiti, Kaiti, as well as artistic handwriting, cartoon fonts, and retro fonts with specific styles, technological fonts, etc. Each font format has unique stroke shape, thickness and style characteristics.
[0095] When randomly selecting a font format from the preset font format library, similar to the way of selecting font color, a random number is used to generate an index value corresponding to a font format in the font format library. Assuming that the font format library is stored in the form of a dictionary, where the key is the font name and the value is the corresponding font format data, the corresponding font format in the dictionary is obtained through the randomly generated index value. Subsequently, the selected font format is applied to the text content, so that the text shows diverse characteristics in font style, laying the foundation for the subsequent construction of richer training sample pairs.
[0096] Step S205, randomly selecting a text border color from a preset text border color library, randomly determining a text border width according to a preset text border width range, and performing border processing on the text content according to the text border color and the text border width.
[0097] The preset text border color library stores a variety of color options for text borders, ranging from simple black and white borders to colored borders such as golden yellow, sky blue, and lavender. When randomly selecting a text border color from the preset text border color library, the random number generation mechanism is still used to make a selection in the border color library. At the same time, the preset text border width range also needs to be set according to actual needs, for example, set to be between 1 and 10 pixels. The text border width is randomly determined within this range through a random number generation algorithm. For example, in Python, random.randint(1, 10) can be used to generate an integer representing the border width. After determining the text border color and width, border addition processing is performed on the previously selected text content.
[0098] The border addition process can utilize image processing algorithms to modify the pixel values at the edge pixel positions of the text according to the determined border width and color, thereby forming a border with a specific color and width around the text, further enriching the presentation form of the text in the image.
[0099] In an alternative embodiment, the method further includes: randomly generating a random number within a preset numerical range (for example, the lower limit is 0 and the upper limit is 1) using a random number generation algorithm; determining whether the random number is greater than a preset threshold. If the random number is greater than the preset threshold, step S205 is executed; if the random number is less than or equal to the preset threshold, no border is added to the selected text content, that is, step S205 is skipped and step S206 is executed.
[0100] Step S206: Randomly select the coordinates for placing the text content from the preset text content coordinate library and randomly select the text rotation angle from within the preset text rotation angle range.
[0101] The preset text content coordinate library stores a series of selectable coordinate information, which is used to determine the placement position of the text content in the image. For example, at 1 / 5 of the height and 1 / 5 of the width of the image, at 1 / 2 of the height and 1 / 2 of the width, at 2 / 3 of the height and 1 / 2 of the width, as Figure 3 shown, which illustrates some coordinate methods.
[0102] When randomly selecting and determining the placement coordinates of the text content from the preset text content coordinate library, a set of coordinate values can be randomly extracted from the text content coordinate library, and this set of coordinate values represents the starting position of the text content in the image. At the same time, the preset text rotation angle range is generally set to a reasonable interval, such as between 0 degrees and 360 degrees. The text rotation angle is randomly selected within this range through a random number generation function. For example, an angle value of 0 degrees, 10 degrees, 15 degrees, or 30 degrees is generated. This rotation angle determines the angle at which the text will be tilted when placed on the image, further increasing the variation possibilities of the text in the image.
[0103] In an alternative implementation manner of the present application, the position for adding the text content can also be determined by the following method: randomly determine the starting coordinate point of the text content within the picture pixel coordinate range of the first image. For example, when determining the starting coordinate in the x-axis direction, a random integer is generated within the range of 0 to width (image width) through a random function algorithm, and this integer is the value of the starting coordinate point of the text content on the x-axis. Similarly, for the starting coordinate in the y-axis direction, a random integer is generated within the range of 0 to height (image height) using the random function algorithm as the value of the starting coordinate point of the text content on the y-axis.
[0104] In actual operation, some factors also need to be considered to ensure the rationality of text placement. For example, to avoid part of the text content exceeding the image boundary, the generated starting coordinates need to be appropriately adjusted according to the actual size of the text content. If the width of the text content is text_width pixels and the height is text_height pixels, then when generating the starting coordinate x_start on the x-axis, it is necessary to ensure that x_start + text_width <= width; when generating the starting coordinate y_start on the y-axis, it is necessary to ensure that y_start + text_height <= height. If the generated coordinates do not meet this condition, they need to be regenerated until the requirements are met.
[0105] Step S207, according to the text content placement coordinates and the text rotation angle, add the text content to the first image and the second image respectively to obtain a third image and a fourth image associated with the third image.
[0106] Among them, the third image is the image after adding the text content to the first image, and the fourth image is the image after adding the text content with the same text attributes at the same position of the second image associated with the first image. Each pair of the third image and the fourth image associated with the third image is a training sample pair for training an image restoration model;
[0107] According to the text content placement coordinates and text rotation angle determined in step S206, add the text content that has undergone a series of previous settings (including selecting text content, determining font color, font format, adding borders, etc.) to the first image and the second image respectively. During the addition process, accurately place the text at the corresponding position of the image according to the determined coordinate values, and at the same time rotate the text according to the selected rotation angle. For example, if the placement coordinates are (200, 300) and the rotation angle is 45 degrees, then the text will start from this coordinate and be rotated 45 degrees on the image. Through such operations, the third image is obtained after adding text to the first image, and the fourth image associated with the third image is obtained after adding text with the same settings to the second image. This pair of images (the third image and the fourth image) are consistent in terms of text content, its style, position, angle, etc., and only differ in the quality of the image itself (the second image is obtained by degrading the first image), thus constructing a more targeted and diverse training sample pair for training the image restoration model, which helps the model more comprehensively learn the ability to restore from a low-quality image to a high-quality image under different text settings.
[0108] Through the above steps, the way of adding text can be made diverse and random, simulating various situations where text may appear, which helps the subsequent model training to recognize various complex situations, thus solving the problem of text flaws in the existing technology after the model restores pictures, and can also help the model better identify image features to achieve a better restoration effect.
[0109] After determining the training sample pair, a generative adversarial network can be trained according to the training sample pair to obtain an image restoration model. Among them, the generative adversarial network includes: a generator and a discriminator. Specifically, it can be realized through step S208 - step S209:
[0110] Step S208, obtain the image label corresponding to the third image in the training sample pair.
[0111] The image label is usually an identifier indicating that the image is a real, high-quality image. Therefore, in this step, the image label is a description or annotation of the third image, which represents the real attributes of the third image. For example, use "1" to indicate that the third image is a real high-quality image, and "0" can be used to indicate a generated fake image. The role of the image label is to provide a judgment standard for the discriminator to help the discriminator distinguish between real images and restored images generated by the generator.
[0112] Step S209, perform iterative adversarial training on the generator and discriminator in the generative adversarial network according to the third image, the image label, and the fourth image until the training stops when the iterative stop condition is reached, and obtain the image restoration model.
[0113] The generative adversarial network consists of two main parts: a generator and a discriminator. The goal of the generator is to restore a low-quality fourth image to a state close to that of a high-quality third image, while the task of the discriminator is to distinguish whether the input image is a real third image or a restored image generated by the generator. During the training process, the generator and the discriminator compete with and learn from each other, and continuously adjust their own parameters to improve performance.
[0114] The iteration stop condition determines when the model training ends. For example, the iteration stop condition can be any of the following:
[0115] Reaching the maximum number of iterations: Set a maximum number of iterations. When the training reaches this number, stop the training. For example, set the maximum number of iterations to 100 times. When the training reaches the 100th iteration, stop the training.
[0116] Convergence of the loss function: Monitor the loss function values of the generator and the discriminator. When the values of the loss function no longer change significantly within a certain number of iterations, that is, when reaching the convergence state, stop the training. For example, the change rate of the loss function value can be calculated. When the change rate is less than a certain threshold, it is considered that the loss function has converged.
[0117] Stable performance of the validation set: Use a validation set to evaluate the performance of the model (including the accuracy, robustness, stability, and practicality of the model). When the performance on the validation set (such as the quality index of image restoration) no longer improves significantly within a certain number of iterations, stop the training.
[0118] In an alternative implementation, the generator and the discriminator in the generative adversarial network are iteratively trained against each other based on the third image, the image label, and the fourth image until the iteration stop condition is reached and the training is stopped, and an image restoration model is obtained. Further steps include:
[0119] S1. Use the generator to perform image restoration processing on the fourth image to obtain the restored fourth image;
[0120] S2. Use the discriminator to perform authenticity discrimination on the restored fourth image to obtain an authenticity prediction result;
[0121] S3. Calculate the restoration loss function according to the authenticity prediction result and the image label corresponding to the third image;
[0122] S4. Adjust the parameters of the generator and the discriminator according to the restoration loss function;
[0123] Iteratively execute S1 - S4 until the iteration stop condition is reached to generate an image restoration model.
[0124] Specifically, the generator is a neural network model based on deep learning, and its architecture is designed to learn how to restore low-quality images to high-quality states. When faced with the fourth image (which is a low-quality image that has been degraded and has text added), the generator processes the fourth image pixel by pixel according to the image features and restoration patterns it has learned internally. The generator contains multiple convolutional layers, deconvolutional layers, and other types of neural network layers. The convolutional layers are responsible for extracting low-level features in the fourth image, such as edges and textures; the deconvolutional layers are used to gradually upsample these features to restore the resolution and details of the image. During the processing, the generator continuously adjusts the connection weights between its internal neurons to optimize the restoration effect. Finally, the generator outputs the restored fourth image, which should visually be closer to the high-quality third image, and the clarity, detail richness, and text presentation quality of the image should be significantly improved.
[0125] The discriminator is also a carefully constructed neural network model, and its core task is to authenticate the restored fourth image output by the generator. The input to the discriminator includes not only the restored fourth image but also the third image (a high-quality image with the same text added) as a real sample. The internal network structure of the discriminator is designed to effectively extract various features of the image, including the overall structure of the image, texture details, color distribution, and text features, etc. It converts the input image into feature vectors through a series of operations such as convolution and pooling, and then uses fully connected layers to classify and judge these feature vectors. The discriminator judges the similarity degree of the features between the restored fourth image and the real third image and outputs a authenticity prediction result. This prediction result is usually presented in the form of a probability value, with a value range between 0 and 1. For example, if the discriminator believes that the restored fourth image is very similar to the real third image, the probability value it outputs may be close to 1, indicating that the image is very likely to be real; on the contrary, if the discriminator identifies obvious unrealistic features in the restored fourth image, such as unreasonable restoration of blurred areas or distortion in the text part, the probability value it outputs may be close to 0, indicating that the image is very likely to be a fake image generated by the generator.
[0126] The restoration loss function is used to measure the difference degree between the restored fourth image output by the generator and the real third image. According to the authenticity prediction result output by the discriminator and the image label corresponding to the third image (in this case, the label corresponding to the third image is real, usually represented by 1), the restoration loss function can be calculated. The value of this loss function will be used as a feedback signal to guide the parameter adjustment of the generator and the discriminator. The loss function is used to evaluate the differences between the generated restored fourth image and the third image in terms of pixels, features, styles, etc.
[0127] According to the calculated repair loss function, it is necessary to adjust the parameters of the generator and the discriminator. This process is achieved with the help of optimization algorithms. Common optimization algorithms include Stochastic Gradient Descent (SGD), Adagrad, Adadelta, Adam, etc. By continuously adjusting the parameters of the discriminator, it can better identify the unrealistic features in the images generated by the generator, thus prompting the generator to further optimize the repair effect.
[0128] The above steps S1 to S4 will be continuously iterated. In each iteration, the generator and the discriminator will operate and adjust their parameters based on new training sample pairs (i.e., different pairs of the third image and the fourth image). As the number of iterations increases, the generator gradually learns a more accurate image repair pattern and can better repair the low-quality fourth image to a state close to the real third image; the discriminator also continuously improves its discrimination ability and can more keenly distinguish the subtle differences between the repaired images generated by the generator and the real images. When the iteration stop condition is reached, the parameters and model structure learned by the generator at this time constitute the finally expected image repair model. This image repair model has the ability to restore low-quality images to high-quality images under different text settings and can be applied to actual scenarios to effectively repair various degraded and text-containing damaged images.
[0129] In an alternative embodiment of the present application, the method further includes: obtaining an image to be repaired, where the image to be repaired refers to a low-quality image with a repair requirement and has text added to it; inputting the image to be repaired into the image repair model for image repair to obtain a repaired image, and the quality of the repaired image becomes higher.
[0130] The image in the embodiments of the present application can be a face image. That is to say, the first image is a face image, and the image repair method provided in the embodiments of the present application can be used for face image repair.
[0131] In addition, the image repair model can be used for videos. Specifically, obtain a video to be repaired, perform decoding processing on the video to be repaired to obtain video frame images to be repaired, input the video frame images to be repaired into the image repair model for image repair to obtain repaired video frame images, and perform encoding processing on the repaired video frame images to obtain a repaired video. This can avoid strange distortion changes of text while achieving the repair of image features, and at the same time, the effect is robust and there will be no jitter in the repaired video results.
[0132] According to the image restoration model training method provided by the embodiments of the present application, the text addition method has diversity and randomness, simulating various situations where text may appear, which helps the subsequent model training to recognize various complex situations, realizes data augmentation, enables the model to have the ability to recognize text during the image restoration process, thereby significantly improving the generalization ability and robustness of the image restoration model, preventing the appearance of text-related defective restoration results during the image restoration process, solving the problem of text defects in the existing technology after the model restores pictures, and can also help the model better recognize facial features to achieve a better restoration effect. For example, the image restoration model of the prior art has defects in text processing during image restoration, specifically manifested as changes in the internal color of the text, the appearance of white edges, and the blurring of text lines, and due to the occlusion of subtitles, the image restoration effect is also not good. However, the image restoration model of the present application does not affect the text, and the restoration effect of the image is also better. The present application can avoid strange distortion changes in the text while restoring the image features, and at the same time, the effect is robust and there will be no situation of jitter in the restored video results.
[0133] Face restoration aims to restore high-quality face images from low-quality input face images and can be used in scenarios such as old photo restoration, old video restoration, and the restoration of blurred surveillance. Face restoration is a special restoration task, and the dataset for its model training is centered around a large number of face samples. By using a large number of clear and blurred face datasets for model training, the model is enabled to have the ability to restore blurred faces. With the popularity of short videos on major platforms, the application of face restoration scenarios in short videos has gradually spread. Due to the rich shooting devices for short videos and the uneven quality of the finished video, the application of face restoration is also very necessary.
[0134] One of the characteristics of short videos is that users may use the editing function to add text descriptions before uploading the video to achieve the video effect, which may lead to the situation where the face is blocked by text sometimes. Since face restoration is a special restoration, the restoration effect of its model is obvious, but for non-face scenarios, unexpected defect situations will occur. In addition, the principle of video restoration is to restore each frame of the video and then encode it into a video. The unstable text restoration effect on the face will cause the jitter of the restored video effect, greatly affecting the visual experience. The above are the disadvantages and drawbacks of the background technology.
[0135] Figure 4 The flowchart of the face restoration model training method according to another embodiment of the present application is shown, as Figure 4 shown, the method includes the following steps:
[0136] Step S401, obtain a first face image, perform degradation processing on the first face image to obtain a second face image associated with the first face image.
[0137] Commonly used face datasets include FFHQ, CelebA, AFLW dataset, etc. This application mainly uses the FFHQ dataset with relatively high resolution, as well as some face data of yellow race generated by stylegan2. The face images here are fake faces and do not involve any real faces. The first face image is the face image in the above dataset. The first face image covers different ages, genders, races, expressions, etc. When degrading the first face image, conditions such as lighting conditions and different degrees of damage need to be covered.
[0138] Step S402: Randomly determine the text content length according to the preset text length range, and randomly select the text content with the text content length from the preset text content library.
[0139] Step S403: Randomly select the font color from the preset font color library, and apply the selected font color to the text content.
[0140] Step S404: Randomly select the font format from the preset font format library, and apply the selected font format to the text content.
[0141] Step S405: Randomly select the text border color from the preset text border color library, randomly determine the text border width according to the preset text border width range, and add a border to the text content according to the text border color and text border width.
[0142] Step S406: Randomly select and determine the text content placement coordinates from the preset text content coordinate library, and randomly select the text rotation angle from the preset text rotation angle range.
[0143] Step S407: According to the text content placement coordinates and text rotation angle, add the text content to the first face image and the second face image respectively to obtain the third face image and the fourth face image associated with the third face image.
[0144] By training on these datasets, the model can learn the prior knowledge and repair skills of faces.
[0145] Step S408: Obtain the image label corresponding to the third face image in the training sample pair.
[0146] Step S409: Iteratively and adversarially train the generator and discriminator in the generative adversarial network according to the third face image, the image label, and the fourth face image until the training stops when the iteration stop condition is reached, and obtain the face image repair model.
[0147] In addition, the face image restoration model can be used for videos. Specifically, a video to be restored is obtained, decoded to obtain video frame images to be restored, the video frame images to be restored are input into the face image restoration model for face restoration to obtain restored video frame images, and the restored video frame images are encoded to obtain a restored video.
[0148] For the specific implementation manners of the above steps, reference can be made to Figure 2 the embodiments shown, which will not be elaborated here.
[0149] The embodiments of the present application have diversity and randomness in the way of adding text, simulating various situations where text may appear, which helps to recognize various complex situations during subsequent model training, realizes data augmentation, enables the model to have the ability to recognize text during face restoration, thereby significantly improving the generalization ability and robustness of the face model, preventing the appearance of defective restoration results regarding text during face restoration, solving the problem of text defects in the existing technology after the model restores pictures, and can also help the model better recognize face features and achieve better restoration effects. For example, the image restoration model of the existing technology has defects in text processing during image restoration, specifically manifested as changes in the internal color of the text, the appearance of white edges, and the blurring of the text lines. Moreover, due to the occlusion of the subtitles, the restoration effect of the face is also not good. However, the face image restoration model of the present application has no impact on the text, and the restoration effect of the face is also better. The present application can avoid strange distortion changes of the text while the face features are restored, and at the same time, the effect is robust and there will be no situation of jitter in the restored video result.
[0150] Figure 5 shows a structural block diagram of an image restoration model training device according to an embodiment of the present application, as Figure 5 shown. The device includes:
[0151] An acquisition module 501, adapted to acquire a first image;
[0152] A degradation processing module 502, adapted to perform degradation processing on the first image to obtain a second image associated with the first image;
[0153] A text enhancement processing module 503, adapted to perform text enhancement processing on the first image and the second image associated with the first image respectively to obtain a third image and a fourth image associated with the third image, where the third image is the image after adding text content to the first image, and the fourth image is the image after adding text content with the same text attributes at the same position of the second image associated with the first image. Each group of the third image and the fourth image associated with the third image is a training sample pair for training an image restoration model;
[0154] The training module 504 is adapted to train a generative adversarial network according to training sample pairs to obtain an image inpainting model.
[0155] Optionally, the text enhancement processing module is further adapted to: randomly determine the text content length according to a preset text length range, and randomly select text content with the text content length from a preset text content library;
[0156] Add the text content to the first image and the second image respectively to obtain a third image and a fourth image associated with the third image.
[0157] Optionally, the text attributes include: font color, font format, text border width, text border color;
[0158] The text enhancement processing module is further adapted to: randomly select a font color from a preset font color library and apply the selected font color to the text content; and / or
[0159] Randomly select a font format from a preset font format library and apply the selected font format to the text content; and / or
[0160] Randomly select a text border color from a preset text border color library, randomly determine the text border width according to a preset text border width range, and add a border to the text content according to the text border color and the text border width.
[0161] Optionally, the text attributes further include: text rotation angle;
[0162] The text enhancement processing module is further adapted to: randomly select and determine the text content placement coordinates from a preset text content coordinate library, and randomly select a text rotation angle from a preset text rotation angle range;
[0163] Adding the text content to the first image and the second image respectively further includes:
[0164] Add the text content to the first image and the second image respectively according to the text content placement coordinates and the text rotation angle.
[0165] Optionally, the generative adversarial network includes: a generator and a discriminator;
[0166] The training module is further adapted to: obtain an image label corresponding to the third image in the training sample pair;
[0167] Iteratively adversarially train the generator and the discriminator in the generative adversarial network according to the third image, the image label and the fourth image until the iteration stop condition is reached, and then stop the training to obtain an image inpainting model.
[0168] Optionally, the training module is further adapted to: S1, perform image inpainting processing on the fourth image using a generator to obtain the inpainted fourth image;
[0169] S2, perform authenticity discrimination on the inpainted fourth image using a discriminator to obtain an authenticity prediction result;
[0170] S3, calculate a repair loss function according to the authenticity prediction result and the image label corresponding to the third image;
[0171] S4, adjust the parameters of the generator and the discriminator according to the repair loss function;
[0172] Iteratively execute S1 - S4 until the iteration stop condition is reached to generate an image inpainting model.
[0173] Optionally, the apparatus further includes: a repair module, adapted to obtain an image to be repaired, wherein text content is added to the image to be repaired;
[0174] Input the image to be repaired into the image inpainting model for image inpainting to obtain the inpainted image.
[0175] Optionally, the first image is a face image.
[0176] The descriptions of the above modules refer to the corresponding descriptions in the method embodiments and will not be elaborated herein.
[0177] According to the image inpainting model training apparatus provided by the embodiments of the present application, in terms of the text addition method, it has diversity and randomness, simulating various situations where text may appear, which helps to recognize various complex situations during subsequent model training, realizes data augmentation, enables the model to have the ability to recognize text during the image inpainting process, thus significantly improving the generalization ability and robustness of the image inpainting model, preventing the appearance of defective repair results regarding text during the image inpainting process, solving the problem of text defects in the existing technology after the model repairs the picture, and can also help the model better recognize face features to achieve a better repair effect. For example, the image inpainting model of the existing technology has defects in text processing during image inpainting, specifically manifested as changes in the internal color of the text, the appearance of white edges, and the blurring of the text lines, and due to the occlusion of the subtitles, the image inpainting effect is also not good. However, the image inpainting model of the present application has no impact on the text, and the image inpainting effect is also better. The present application can avoid strange distortion changes of the text while repairing the image features, and at the same time, the effect is robust and there will be no situation of jitter in the repaired video result.
[0178] An embodiment of the present application provides a non-volatile computer storage medium. The computer storage medium stores at least one executable instruction or computer program, and the executable instruction or computer program can cause a processor to execute the operations corresponding to the image restoration model training method in any of the above method embodiments.
[0179] An embodiment of the present application provides a computer program product. The computer program product includes at least one executable instruction or computer program, and the executable instruction or computer program can cause a processor to execute the operations corresponding to the image restoration model training method in any of the above method embodiments.
[0180] Figure 6 The structure diagram of an embodiment of the computing device of the present application is shown. The specific implementation of the computing device is not limited in the specific embodiments of the present application.
[0181] As Figure 6 shown, the computing device may include: a processor 602, a communication interface 604, a memory 606, and a communication bus 608.
[0182] Among them: The processor 602, the communication interface 604, and the memory 606 communicate with each other through the communication bus 608. The communication interface 604 is used to communicate with network elements of other devices such as clients or other servers. The processor 602 is used to execute the program 610, and specifically can execute the relevant steps in the above method embodiment for training the image restoration model of the computing device.
[0183] Specifically, the program 610 may include program code, and the program code includes computer operation instructions.
[0184] The processor 602 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0185] The memory 606 is used to store the program 610. The memory 606 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0186] The program 610 can specifically be used to cause the processor 602 to execute the image restoration model training method in any of the above method embodiments. For the specific implementation of each step in the program 610, reference can be made to the corresponding steps and units in the above image restoration model training embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.
[0187] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings provided herein. The structure required to construct such systems will be apparent from the above description. In addition, the embodiments of the present application are not directed to any specific programming language. It should be understood that the content of the embodiments of the present application described herein can be implemented using various programming languages, and the description of a specific language above is for the purpose of disclosing the best mode of the embodiments of the present application.
[0188] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0189] Similarly, it should be understood that, in order to streamline the present disclosure and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the embodiments of the present application, the various features of the embodiments of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed embodiments of the present application require more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0190] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and set in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise clearly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0191] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means being within the scope of the embodiments of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0192] Each component embodiment of the embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present application. The embodiments of the present application can also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the embodiments of the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0193] It should be noted that the above embodiments illustrate the embodiments of the present application rather than limit them, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The embodiments of the present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.
Claims
1. A method for training an image inpainting model, comprising: Obtaining a first image, degrading the first image to obtain a second image associated with the first image; Performing text enhancement processing on the first image and the second image associated with the first image respectively to obtain a third image and a fourth image associated with the third image, wherein the third image is an image after adding text content to the first image, and the fourth image is an image after adding text content with the same text attributes at the same position of the second image associated with the first image, and each pair of the third image and the fourth image associated with the third image is a training sample pair for training the image inpainting model; Training a generative adversarial network according to the training sample pair to obtain an image inpainting model.
2. The method according to claim 1, wherein, The performing text enhancement processing on the first image and the second image associated with the first image respectively to obtain a third image and a fourth image associated with the third image further includes: Randomly determining the text content length according to a preset text length range, and randomly selecting text content of the text content length from a preset text content library; Adding the text content to the first image and the second image respectively to obtain a third image and a fourth image associated with the third image.
3. The method according to claim 2, wherein, The text attributes include: font color, font format, text border width, text border color; Before adding the text content to the first image and the second image respectively, the method further includes: Randomly selecting a font color from a preset font color library and applying the selected font color to the text content; and / or Randomly selecting a font format from a preset font format library and applying the selected font format to the text content; and / or Randomly selecting a text border color from a preset text border color library, randomly determining the text border width according to a preset text border width range, and adding a border to the text content according to the text border color and the text border width.
4. The method according to claim 2 or 3, wherein The text attributes further include: text rotation angle; Before adding the text content to the first image and the second image respectively, the method further includes: Randomly selecting and determining the text content placement coordinates from a preset text content coordinate library, and randomly selecting a text rotation angle from a preset text rotation angle range; The adding the text content to the first image and the second image respectively further includes: Adding the text content to the first image and the second image respectively according to the text content placement coordinates and the text rotation angle.
5. The method according to any one of claims 1 to 4, wherein The generative adversarial network includes: a generator and a discriminator; The training the generative adversarial network according to the training sample pair to obtain an image inpainting model further includes: Obtaining an image label corresponding to the third image in the training sample pair; Performing iterative adversarial training on the generator and the discriminator in the generative adversarial network according to the third image, the image label and the fourth image until the iterative stop condition is reached and the training is stopped to obtain an image inpainting model.
6. The method according to claim 5, wherein, Iteratively and adversarially training the generator and discriminator in the generative adversarial network according to the third image, the image label, and the fourth image until the training is stopped when the iterative stop condition is reached, obtaining the image restoration model further includes: S1. Using the generator to perform image restoration processing on the fourth image to obtain the restored fourth image; S2. Using the discriminator to perform authenticity discrimination on the restored fourth image to obtain an authenticity prediction result; S3. Calculating a restoration loss function according to the authenticity prediction result and the image label corresponding to the third image; S4. Adjusting the parameters of the generator and the discriminator according to the restoration loss function; Iteratively execute S1 - S4 until the iterative stop condition is reached to generate an image restoration model.
7. The method according to any one of claims 1-6, wherein, The method further includes: Obtaining an image to be restored, wherein text content is added to the image to be restored; Inputting the image to be restored into the image restoration model for image restoration to obtain a restored image.
8. The method according to any one of claims 1-7, wherein, The first image is a face image.
9. An image restoration model training device, including: An acquisition module, adapted to acquire a first image; A degradation processing module, adapted to perform degradation processing on the first image to obtain a second image associated with the first image; A text enhancement processing module, adapted to perform text enhancement processing on the first image and the second image associated with the first image respectively to obtain a third image and a fourth image associated with the third image, wherein the third image is the image after adding text content to the first image, the fourth image is the image after adding text content with the same text attribute at the same position of the second image associated with the first image, and each pair of the third image and the fourth image associated with the third image is a training sample pair for training the image restoration model; A training module, adapted to train a generative adversarial network according to the training sample pairs to obtain an image restoration model.
10. A computing device, comprising: A processor, a memory, a communication interface, and a communication bus, the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the image restoration model training method according to any one of claims 1 - 8.
11. A computer storage medium, in which at least one executable instruction is stored, and the executable instruction causes the processor to execute the operations corresponding to the image restoration model training method according to any one of claims 1 - 8.
12. A computer program product, including at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the image restoration model training method according to any one of claims 1 - 8.