Intelligent Colorization Method, System, Electronic Device and Storage Medium for Image Files
By constructing a color mathematical model and combining the self-attention mechanism and residual network, the image archives are intelligently colored, which solves the problem of poor visual effects of image archives in the prior art, and achieves more efficient image processing effects and more realistic color image generation.
Patent Information
- Application Number
- CN202210937534.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-08-05
AI Technical Summary
The prior art has problems such as monochromatic, distortion, low pixels, and low resolution in the digital processing of image files, resulting in poor visual effects, low efficiency of manual adjustment and interpolation technology and may lead to high-frequency information loss and noise amplification, which will lead to blurred reconstruction images.
An intelligent coloring method of image archives is adopted. By constructing a color mathematical model, using self-attention mechanism and residual network, the image archives are colored and processed to improve the processing effect of image archives. The method includes the steps of image acquisition, color mathematical model construction and image coloring, and generates more realistic and detailed color images through feature extraction, feature mapping and three-way processing.
Through intelligent coloring methods, the visual effect of image files is significantly improved, information utilization is improved, image detail characteristics are enhanced, and better coloring effect is achieved, avoiding the low efficiency of manual adjustment and interpolation technology and the loss of high-frequency information.
Smart Images

Figure CN117593219B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and relates to an image restoration system, and in particular to an intelligent colorization method, system, electronic device and storage medium for image archives. Background Art
[0002] As original documents recording the true face of history, archives play an important role in the economic and social development. During the process of storage and utilization, archives are inevitably affected by environmental factors such as temperature, humidity, and light. Paper archives and photo archives gradually fade, bleed, distort, turn yellow, become damaged, mildew, and negatives become invalid. To rescue and protect the information of these precious historical archives, archivists use devices such as scanners or digital cameras to digitize the archives, converting them into digital images stored on carriers such as magnetic tapes, disks, and optical discs, which are convenient for retrieval and utilization, thereby protecting the original archives. However, due to reasons such as low clarity of early scanning and shooting devices, these digital images have defects such as being monochromatic, distorted, having low pixels, and low resolution, and the visual effects of the digital images are poor. To solve such problems, artificial adjustment and interpolation techniques are usually used to improve the resolution of image archives.
[0003] Artificial adjustment uses tools such as Photoshop to adjust the brightness, saturation, histogram, mean, clarity, color difference, hue, pixels, etc. of electronic image archives to improve the clarity and visual effects of electronic archive images. However, due to the artificial adjustment of image parameters and diverse image information, it requires a large amount of labor costs and time cycles, and the efficiency is too low. Interpolation techniques are used to improve the clarity of images, including nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, resampling and other schemes. By calculating the unknown pixel points in the new image from the known pixel points of the original image, some details of the image will be lost after restoration, such as the loss of high-frequency information. And during the interpolation process, noise may be amplified, resulting in side effects such as blurred reconstructed images. Interpolation techniques not only have high computational complexity, but also different interpolation methods will cause great differences in the results of image reconstruction.
[0004] In view of this, there is an urgent need to design a new image processing method today to overcome at least some of the above-mentioned defects existing in the existing image processing methods. Summary of the Invention
[0005] The present invention provides an intelligent colorization method, system, electronic device and storage medium for image archives, which can perform colorization processing on image archives and improve the processing effect of image archives.
[0006] To solve the above technical problems, according to one aspect of the present invention, the following technical solution is adopted:
[0007] An intelligent colorization method for image archives, the intelligent colorization method for image archives includes:
[0008] Image acquisition step; Obtain image file information;
[0009] Color mathematical model construction step; Construct a color mathematical model;
[0010] Image coloring step; Input the image file information obtained in the image acquisition step into the color mathematical model to color the image file.
[0011] As an implementation manner of the present invention, in the color mathematical model construction step, by directly calculating the relationship between any two pixel points in the image, the global geometric features of the noise image are obtained, and the dependence relationship between the global features is learned; Model the long-distance dependence relationship between target image regions;
[0012] Extract features from the image, perform feature mapping on the extracted features; Perform three-way processing on the feature mapping;
[0013] The first-way processing is convolved through the first convolutional unit to obtain the first feature space f(x);
[0014] f(x) = W f X
[0015] The second-way processing is convolved through the second convolutional unit to obtain the second feature space g(x);
[0016] g(x) = W g X
[0017] where x represents the input feature map, f and g are two convolutions respectively, and W f is the first weight matrix, and W g is the second weight matrix;
[0018] Normalize each feature vector along the feature channel direction, as shown in formula (2), to obtain the normalized result norm(x);
[0019]
[0020] Then, transpose f(x) and perform matrix multiplication with g(x), and then perform softmax on each row of it. According to formula (3), the weight β of all feature points to a certain feature point can be obtained j,i , forming an attention matrix;
[0021]
[0022] where β j,i is an element in the attention matrix, and s ijis each element in the matrix obtained by performing matrix multiplication between the transpose of the normalized f(x) and the normalized g(x); the sum of each row in the attention matrix is 1, and the i numbers in the j-th row respectively represent the contributions of all other points to it at the i-th point. In this way, the contribution values of all other pixel points to this pixel point can be obtained through each row, and the contribution value of a certain pixel point to all other pixel points can be obtained through each column;
[0023] The third path of processing convolves through the third convolutional unit to obtain the third feature space h(x), and then multiplies it by the transpose of the attention matrix according to formula (4) to obtain the final result;
[0024]
[0025] where h(x) is the convolution result of the third convolutional unit, x i is an element in the input feature map x, and W h is the third weight matrix, and W v is the fourth weight matrix.
[0026] As an implementation manner of the present invention, in the step of constructing the color mathematical model, the self-attention mechanism and the residual network are combined, so that the detailed features of the output color image are richer, the information utilization rate is higher, and a better coloring effect is achieved;
[0027] The color mathematical model includes a first network model and a second network model. The first network model is provided with a first generation network formed by a first generator and a first discrimination network formed by a first discriminator. The second network model is provided with a second generation network formed by a second generator and a second discrimination network formed by a second discriminator;
[0028] The first generation network represents a generator from the X domain to the Y domain, denoted as: G: X -> Y. The second generation network represents a generator from the Y domain to the X domain, denoted as F: Y -> X. The first discrimination network Dx is used to identify whether the input image is X; the second discrimination network Dy is used to identify whether the input image is Y;
[0029] The first discrimination network Dx and the second discrimination network Dy both include four convolutional layers and two self-attention mechanisms, and train the images generated by the corresponding generation network and the existing colored images, and judge the quality of the generated images according to the similarity between the images generated by the generation network and the original images.
[0030] As an implementation manner of the present invention, the loss function of the color mathematical model includes: an adversarial loss for controlling the style of the generated image to approximate the target image, and a cyclic consistency loss for retaining the contour information of the input to better retain the content structure of the input and capture the features of the target domain; that is:
[0031] Loss = Loss GAN + Loss cycle (5)
[0032] Loss GAN Ensure that the generator network and the discriminator network evolve mutually, thereby ensuring that the generator network can generate more realistic images, Loss cycle Ensure that the output image of the generator network is only different in color from the input image, but the content is the same; Loss GAN The specific calculation formula is as in (6), Loss cycle The specific calculation formula is as in (7);
[0033]
[0034]
[0035] When training the generator, D X and D Y parameters are fixed, only the parameters of G and F are adjustable; adjust the parameters of G so that D Y the score D Y (G(x)) of the image G(x) generated by G is as high as possible; adjust the parameters of F so that D X the score D X (F(y)) of the image F(y) generated by F is as high as possible; the ultimate goal is to make F(G(x)) = x and G(F(y)) = y, ensuring that the color of the image generated by the generator becomes more and more realistic;
[0036] When training the discriminator, the parameters of G and F are fixed, and D X and D T parameters are adjustable; when training the discriminator D X , maximize the value of D X (x); at the same time, minimize the value of D X (F(y)), thereby improving the discrimination ability of the discriminator.
[0037] According to another aspect of the present invention, the following technical solution is adopted: An intelligent colorization system for image files, the intelligent colorization system for image files includes:
[0038] An image acquisition module for acquiring image file information;
[0039] A color mathematical model construction module for constructing a color mathematical model;
[0040] An image colorization module for inputting the image file information acquired by the image acquisition module into the color mathematical model to colorize the image file.
[0041] As an implementation manner of the present invention, the color mathematical model construction module obtains the global geometric features of the noise image by directly calculating the relationship between any two pixel points in the image, learns the dependency relationship between the global features, and models the long-distance dependency relationship between the target image regions.
[0042] As an implementation manner of the present invention, the color mathematical model construction module extracts features from the image, performs feature mapping on the extracted features, and performs three-way processing on the feature mapping;
[0043] The first-way processing is convolved through the first convolutional unit to obtain the first feature space f(x);
[0044] f(x) = W f X
[0045] The second-way processing is convolved through the second convolutional unit to obtain the second feature space g(x);
[0046] g(x) = W g X
[0047] where x represents the input feature map, f and g are two convolutions respectively, and W f is the first weight matrix, and W g is the second weight matrix;
[0048] Normalize each feature vector along the feature channel direction, as shown in formula (2), to obtain the normalized result norm(x);
[0049]
[0050] Then, transpose f(x) and perform matrix multiplication with g(x), and then perform softmax on each row. According to formula (3), the weight β of all feature points with respect to a certain feature point can be obtained j,i , forming an attention matrix;
[0051]
[0052] where β j,i is an element in the attention matrix, and s ij is each element in the matrix obtained by performing matrix multiplication on the transpose of the normalized f(x) and the normalized g(x); the sum of each row in the attention matrix is 1, and the i numbers in the j-th row respectively represent the contributions of all other points to it at the i-th point. In this way, the contribution values of all other pixel points to this pixel point can be obtained through each row, and the contribution value of a certain pixel point to all other pixel points can be obtained through each column;
[0053] The third - path processing obtains the third feature space h(x) through convolution, and then multiplies it by the transpose of the attention matrix according to formula (4) to obtain the final result;
[0054]
[0055] where h(x) is the convolution result of the third convolution unit, and x i is an element in the input feature map x, and W h is the third weight matrix, and W v is the fourth weight matrix.
[0056] As an implementation manner of the present invention, the color mathematical model construction module combines the self - attention mechanism and the residual network, making the detailed features of the output color image richer, the information utilization rate higher, and thus achieving a better coloring effect;
[0057] The color mathematical model includes a first network model and a second network model. The first network model is provided with a first generation network formed by a first generator and a first discriminant network formed by a first discriminator. The second network model is provided with a second generation network formed by a second generator and a second discriminant network formed by a second discriminator;
[0058] The first generation network represents a generator from the X domain to the Y domain, denoted as: G: X -> Y. The second generation network represents a generator from the Y domain to the X domain, denoted as F: Y -> X. The first discriminant network Dx is used to identify whether the input image is X; the second discriminant network Dy is used to identify whether the input image is Y;
[0059] The first discriminant network Dx and the second discriminant network Dy both include four convolutional layers and two self - attention mechanisms, training the images generated by the corresponding generation network and the existing colored images, and judging the quality of the generated images according to the similarity between the images generated by the generation network and the original images;
[0060] As an implementation manner of the present invention, the loss function of the color mathematical model consists of two parts, including an adversarial loss that controls the style of the generated image to approximate the target image and a cycle - consistency loss that preserves the contour information of the input to better preserve the content structure of the input and capture the features of the target domain; that is:
[0061] Loss = Loss GAN + Loss cycle (5)
[0062] Loss GAN ensures that the generation network and the discriminant network evolve mutually, and thus ensures that the generation network can generate more realistic pictures. Loss cycleEnsure that the output image of the generation network is the same in content as the input image, except for the color; Loss GAN The specific calculation formula is as shown in (6), Loss cycle The specific calculation formula is as shown in (7);
[0063]
[0064]
[0065]
[0066] When training the generator, the parameters of D X and D Y are fixed, and only the parameters of G and F are adjustable; adjust the parameters of G to make D Y give a higher score D Y (G(x)) for the image G(x) generated by G. Adjust the parameters of F to make D X give a higher score D X (F(y)) for the image F(y) generated by F; the ultimate goal is to make F(G(x)) = x and G(F(y)) = y, ensuring that the color of the image generated by the generator becomes more and more realistic;
[0067] When training the discriminator, the parameters of G and F are fixed, and the parameters of D X and D Y are adjustable; when training the discriminator D X , maximize the value of D X (x), and at the same time, minimize the value of D X (F(y)), thereby improving the discrimination ability of the discriminator. The process of updating the [formula] is similar, which will not be elaborated here.
[0068] According to another aspect of the present invention, the following technical solution is adopted: an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps of the above method are implemented.
[0069] According to another aspect of the present invention, the following technical solution is adopted: a storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the steps of the above method are implemented.
[0070] The beneficial effects of the present invention are as follows: The intelligent colorization method, system, electronic device, and storage medium for image files proposed by the present invention can perform colorization processing on image files and improve the processing effect of image files. Brief Description of the Drawings
[0071] Figure 1It is a flowchart of an image file repair method in an embodiment of the present invention.
[0072] Figure 2 It is a flowchart of an image file repair method in another embodiment of the present invention.
[0073] Figure 3 It is a schematic diagram of the principle of the normalized self-attention mechanism in an embodiment of the present invention.
[0074] Figure 4 It is a schematic diagram of a generation network and a discriminator network in an embodiment of the present invention.
[0075] Figure 5 It is a schematic diagram of the composition of the GAN network structure in an embodiment of the present invention.
[0076] Figure 6 It is a schematic diagram of the composition of a discriminator in an embodiment of the present invention.
[0077] Figure 7 It is a schematic diagram of the principle of the image super-resolution reconstruction steps in an embodiment of the present invention.
[0078] Figure 8 It is a schematic diagram of a typical residual structure.
[0079] Figure 9 It is a schematic diagram of a residual structure in an embodiment of the present invention.
[0080] Figure 10 It is a schematic diagram of the structure of a feature extraction module in an embodiment of the present invention.
[0081] Figure 11 It is a schematic diagram of the structure of a super-resolution reconstruction module in an embodiment of the present invention.
[0082] Figure 12 It is a schematic diagram of sub-pixel upsampling in an embodiment of the present invention.
[0083] Figure 13 It is a schematic diagram of the composition of an image file repair system in an embodiment of the present invention.
[0084] Figure 14 It is a schematic diagram of the composition of an electronic device in an embodiment of the present invention. Detailed implementation manners
[0085] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0086] To further understand the present invention, the preferred implementation manners of the present invention will be described below in conjunction with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.
[0087] The description of this part only targets several typical embodiments, and the present invention is not limited to the scope described in the embodiments. The mutual replacement of the same or similar prior art means and some technical features in the embodiments is also within the scope of the description and protection of the present invention.
[0088] The expressions of the steps in the various embodiments in the specification are only for convenience of description, and the implementation manner of the present application is not limited by the order of step implementation.
[0089] "Connection" in the specification includes both direct connection and indirect connection.
[0090] The present invention discloses an intelligent colorization method for image files. Figure 1 It is a flowchart of the image file repair method in an embodiment of the present invention; please refer to Figure 1 The intelligent colorization method for image files includes:
[0091]
Step S1
[0092]
Step S3
[0093]
Step S5
[0094] Although the colorful landscape appears black and white in the photo, due to the different landscape colors, the grayscales presented in the black and white photo are also different. Based on the different grayscales of the black and white photos, the algorithm can roughly distinguish the colors of objects. Image and video colorization is to perform colorization operations on images and videos again on the basis of image and video super-resolution restoration, improving the overall visual effect of images and videos. In one embodiment, the photo and video colorization adopted takes the self-attention mechanism generative adversarial network EnhanceSAGAN algorithm as the core.
[0095] Deep convolutional networks can improve the details of high-resolution images generated by GAN network models, but due to the limitations of the local receptive field of the underlying convolutional network, problems will arise if long-range dependency areas are to be generated. For example, in the process of dyeing black and white images, color coordination is very important, but because general convolution kernels are difficult to cover a large area, when performing convolution on large areas of the image, it is impossible to take into account the mutual influence, so the resulting rendering will lack color coordination and integrity. Therefore, the focus of rendering black and white images is to find a method that can utilize global information to cover all areas. Traditional methods often make the number of parameters and calculations too large. The existing methods use deeper convolutional networks, or directly use fully connected layers to obtain global information. This approach is both complicated and time-consuming.
[0096] Adversarial learning in some GAN models can easily learn the texture features of image information: such as target size and target shape, but it is not easy to learn specific structural and geometric features inside the image information, such as the detailed features, textures and patterns inside the target. Adversarial networks can enhance discriminative local details, and the self-attention mechanism shows a better balance between the ability to simulate image long-range correlation dependencies, computational efficiency and statistical efficiency than other generative adversarial networks, which is crucial for the coloring of movie frames. The essence of the mechanism is to take the weighted sum of the features at all positions as the new features of the position, where the weights are calculated only at a very small computational cost.
[0097] The introduction of the self-attention generative adversarial network introduces the self-attention mechanism into the image generation of the GAN network. This method solves this problem simply and efficiently. Therefore, the present invention adopts an improved Self-Attention GAN (SAGAN) model, namely: EnhanceSAGAN. The algorithm improves the self-attention mechanism, proposes a normalized self-attention mechanism, and embeds the normalized self-attention mechanism into the generative adversarial network.
[0098] Since the calculation of the correlation between features is easily affected by noise, for example, if a value in the feature vector is affected by noise and suddenly becomes larger or smaller, then when calculating the correlation between the vector and other vectors, the result will be inaccurate. To solve this problem, EnhanceSAGAN adopts a normalized self-attention mechanism, that is, all feature vectors are normalized before calculating the correlation, so as to eliminate the influence of noise as much as possible.
[0099] Figure 3 is a schematic diagram of the principle of the normalized self-attention mechanism in one embodiment of the present invention; please refer to Figure 3, in one embodiment, step S3 includes: extracting features from the image, performing feature mapping on the extracted features; performing three-way processing on the feature mapping:
[0100] The first-way processing is convolved through a first convolutional unit (which can be a 1×1 convolution) to obtain a first feature space f(x);
[0101] f(x) = W f X
[0102] The second-way processing is convolved through a second convolutional unit (which can be a 1×1 convolution) to obtain a second feature space g(x);
[0103] g(x) = W g X
[0104] where x represents the input feature map, f and g are two convolutions respectively, and W f is the first weight matrix, and W g is the second weight matrix.
[0105] Normalize each feature vector along the feature channel direction, as shown in formula (2), to obtain the normalized result norm(x);
[0106]
[0107] Then, transpose f(x) and perform matrix multiplication with g(x), and then perform softmax on each row. According to formula (3), the weights β of all feature points with respect to a certain feature point can be obtained j,i , forming an attention matrix;
[0108]
[0109] where β j,i is an element in the attention matrix, and s ij is each element in the matrix obtained by matrix multiplication of the transpose of the normalized f(x) and the normalized g(x); the sum of each row in the attention matrix is 1, and the i numbers in the j-th row respectively represent the contributions of all other points to the i-th point. In this way, the contributions of all other pixel points to this pixel point can be obtained through each row, and the contributions of a certain pixel point to all other pixel points can be obtained through each column.
[0110] The third-way processing is convolved through a third convolutional unit (which can be a 1×1 convolution) to obtain a third feature space h(x), and then multiplied by the transpose of the attention matrix according to formula (4) to obtain the final result (the result is the feature vector input to the next module);
[0111]
[0112] Among them, h(x) is the convolution result of the third convolution unit, and x i is an element in the input feature map x, and W h is the third weight matrix, and W v is the fourth weight matrix.
[0113] In the step of constructing the color mathematical model, the EnhanceSAGAN network combines the self-attention mechanism and the residual network, making the detailed features of the output color image richer and the information utilization rate higher, thereby achieving a better coloring effect. Please refer to Figure 4 , the color mathematical model includes a first network model and a second network model. The first network model is provided with a first generation network formed by a first generator and a first discriminant network formed by a first discriminator. The second network model is provided with a second generation network formed by a second generator and a second discriminant network formed by a second discriminator. The first generation network represents the generator from the X domain to the Y domain, denoted as: G: X -> Y. The second generation network represents the generator from the Y domain to the X domain, denoted as F: Y -> X; The first discriminant network Dx is used to identify whether the input image is X; The second discriminant network Dy is used to identify whether the input image is Y. The first discriminant network Dx and the second discriminant network Dy both include four convolutional layers and two self-attention mechanisms, train the images generated by the corresponding generation network and the existing colored images, and judge the quality of the generated images according to the similarity between the images generated by the generation network and the original images.
[0114] In one embodiment, the generation network in EnhanceSAGAN is composed of two single identical GAN networks, and the structure of a single GAN network is as Figure 5 shown. The GAN network is mainly composed of three operations: convolution, transposed convolution, and residual block. The BN layer and the activation function are performed after the convolution or transposed convolution operation. Among them, c is the convolutional layer, r is the residual layer, d is the transposed convolutional layer, and G_SA represents the added improved self-attention mechanism. The settings of each convolutional layer are as Figure 5 shown by the numbers in. Taking c1 as an example, the size of the c1 convolutional kernel is 7×7, the number of layers is 64, and the stride is 1.
[0115] The discriminator D x and D y in EnhanceSAGAN can be as Figure 6 shown, composed of four convolutional layers and two self-attention mechanisms, train the images generated by the generator and the existing colored images, and judge the quality of the generated images according to the similarity between the images generated by the generator and the original images; among them, D_SA represents the added improved self-attention mechanism.
[0116] In one embodiment of the present invention, the loss function of the color mathematical model (EnhanceSAGAN loss function) includes: an adversarial loss that controls the style of the generated image to approximate the target image, and a cycle consistency loss that preserves the contour information of the input to better preserve the content structure of the input and capture the features of the target domain; that is:
[0117] Loss = Loss GAN +Loss cycle (5)
[0118] Loss GAN ensures that the generator network and the discriminator network evolve mutually, thereby ensuring that the generator network can generate more realistic pictures. Loss cycle ensures that the output picture of the generator network is the same in content as the input picture, except for the color; Loss GAN The specific calculation formula is as shown in (6), and Loss cycle The specific calculation formula is as shown in (7);
[0119]
[0120]
[0121] When training the generator, the parameters of D X and D Y are fixed, and only the parameters of G and F are adjustable; by adjusting the parameters of G, the score D Y of the picture G(x) generated by G by D Y (G(x)) is as high as possible; by adjusting the parameters of F, the score D X of the picture F(y) generated by F by D X (F(y)) is as high as possible; the ultimate goal is to make F(G(x)) = x and G(F(y)) = y, ensuring that the color of the pictures generated by the generator becomes more and more realistic.
[0122] When training the discriminator, the parameters of G and F are fixed, and the parameters of D X and D Y are adjustable; when training the discriminator D X , maximize the value of D X (x); at the same time, minimize the value of D X (F(y)), thereby enhancing the discrimination ability of the discriminator; the process of updating the formula is similar and will not be elaborated here.
[0123] In an embodiment of the present invention, in step S3, EnhanceSAGAN can directly calculate the relationship between any two pixel points in the image, and can obtain the global geometric features of the noise image in one step, which enables it to better learn the dependence relationship between global features. By using an improved self-attention mechanism to model the long-range dependence relationship between target image regions, this model can indeed better approximate the original image distribution compared to other newly proposed generative adversarial network models.
[0124] In an embodiment of the present invention, the intelligent colorization method for image files may further include: a color pre-discrimination step of pre-discriminating the colors of each region according to the grayscale of each region in the image file information obtained by the image acquisition module. In the image colorization step, the image file is colored by combining the color data pre-discriminated in the color pre-discrimination step and the color data obtained by inputting the image file information obtained in the image acquisition step into the color mathematical model.
[0125] Figure 2 It is a flowchart of an image file repair method in another embodiment of the present invention; please refer to Figure 2 In an embodiment of the present invention, the intelligent colorization method for image files further includes step S2: an image super-resolution reconstruction step of performing image super-resolution reconstruction on the denoised image to obtain the corresponding image file. Step S2 can be between step S1 and step S3. After obtaining the image, image super-resolution reconstruction is performed, and then the image after image super-resolution reconstruction is colored.
[0126] In an embodiment, the image super-resolution reconstruction step includes:
[0127] An image denoising step of denoising the image obtained in the image acquisition step;
[0128] A feature extraction step of extracting image features;
[0129] A feature fusion step of fusing the extracted image features;
[0130] An image super-resolution reconstruction step of performing image super-resolution reconstruction according to the features fused in the feature fusion step to obtain a repaired image super-resolution colored image.
[0131] Figure 7 It is a schematic diagram of the principle of the image super-resolution reconstruction step in an embodiment of the present invention; please refer to Figure 7, in an embodiment of the present invention, the image super-resolution reconstruction step adopts surrounding the computer vision task target, and by fusing multi-object features and super-resolution reconstruction, so as to achieve a real and credible reconstruction of the task target; the real and accurate reconstruction of the input images and videos in the computer vision task can effectively improve the performance of various computer vision algorithms in the actual application scenario.
[0132] When facing different computer vision tasks, multiple (1 to N) target features can be set and extracted. During network training, the loss functions (LOSS) calculated and set for multiple target features are fused with the image / video LOSS generated by the super-resolution reconstruction module, and then backpropagated to each training module (training) to correct the relevant parameters of the deep neural network; after the network training is completed, only the "feature extraction module" and the "super-resolution reconstruction module" are required to achieve real and accurate super-resolution reconstruction of the input images and videos.
[0133] In the feature extraction step, the set feature layer corresponding to the image / video to be processed is extracted. The feature extraction module takes the low-resolution image as the input and outputs a set of feature images through the convolutional neural network. In order to extract more features and obtain better performance, the feature extraction network of the MFFSR algorithm adopts a residual structure, making the entire network deeper. Increasing the width and depth of the network can improve the network performance to a certain extent. Generally, a deeper network has better effects than a shallower network. A typical residual structure is as Figure 8 shown.
[0134] As the neural network becomes deeper and deeper, the main problems encountered by deep learning for network depth are gradient disappearance and gradient explosion. The traditional corresponding solutions are data initialization (normalized initialization) and (batch normalization) regularization. However, although this solves the gradient problem and the depth is increased, it brings another problem, that is, the problem of network performance degradation. As the depth increases, the error rate rises. The residual is used to design and solve the degradation problem, which also solves the gradient problem and further improves the network performance. The residual structure mainly includes two parts. One part directly outputs the original data, and the other part selectively passes through the convolutional layer, BN layer, and activation function layer, and outputs residual data with the same size as the original data. Finally, the two parts of data are concatenated and other operations are performed to obtain the final output of the residual structure. Conv in the figure refers to the convolutional operation, addition refers to the identity addition operation, BN refers to the batch normalization operation, and ReLU is the activation function. The residual block can be defined as:
[0135] x (l+1) =x l +F(x l , wl ) (3-1)
[0136] Among them, x l represents the output of the l-th local residual learning, and x (l+1) represents the output of the (l + 1)-th local residual learning. F(x l , w l ) represents the residual mapping of l + 1, and w l is its mapping weight.
[0137] When the neural network does not introduce an activation function, the neural network can only learn a linear model, and the expression ability of the linear model is insufficient. Introducing an activation function is to add non-linear factors. The most commonly used activation function in the residual structure is ReLU, and its expression is as follows:
[0138] f(x) = max(0, x) (3-2)
[0139] Among them, x represents the data result output by the previous layer network, and f(x) represents the output of the activation function layer.
[0140] Batch Normalization (BN) is a very important technology in deep learning. It can correct the distribution of intermediate data, alleviate the problem of gradient dispersion, reduce the dependence of network performance on initialization parameters, and accelerate network training. In this way, the training of deep networks can be made easier and the network convergence can be accelerated. The BN layer is widely used in many classification tasks based on neural networks, and there have also been many attempts in super-resolution networks. However, in image super-resolution and image generation, the performance of BN is not very good. After adding the BN layer to the network, the training speed becomes slow and unstable, and even the final result diverges.
[0141] The poor effect of the BN layer in the super-resolution task is mainly because image super-resolution requires that the output image of the network needs to be consistent with the input in terms of color, contrast, and brightness, and only the resolution and some details are changed. And Batch Norm is similar to a kind of contrast stretching for images. After any image passes through Batch Norm, the color distribution of it will be normalized. That is to say, it destroys the original contrast information of the image. Therefore, the addition of Batch Norm instead affects the quality of the network output. In one embodiment, the residual block used in the present invention removes some unnecessary modules in the above residual block structure, such as the BN layer, as Figure 9 shown. Since the memory consumption of the BN layer is the same as that of the convolutional layer, after removing the BN layer, the memory usage required for calculation will also be reduced. Therefore, these unnecessary operations can be removed with limited computing resources, which can be used to construct a deeper network model to obtain higher accuracy and better performance.
[0142] Figure 10 Shown is the structure of the feature extraction module adopted by the present invention, including an input layer, a residual layer, and an output layer. Both the input layer and the output layer are convolutional layers. The residual layer is composed of connected residual blocks, and the output of each residual block serves as the input to the next residual block. In order to deepen the network and better mine the image super-resolution features, in addition to the local residuals between the residual blocks, this study also adds the feature image extracted by the input layer to the output of the residual layer through a skip connection to form a global residual.
[0143] In the image super-resolution reconstruction step, image upsampling plays an important role in super-resolution reconstruction. There are traditional upsampling methods based on interpolation and upsampling methods based on deep learning. Nowadays, it has gradually become a trend to use CNNs to learn end-to-end upsampling methods.
[0144] Traditional interpolation-based upsampling mainly includes nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, Sinc, and Lanczos resampling, etc. These methods are relatively simple and easy to interpret.
[0145] Nearest neighbor interpolation mainly selects the pixel value of the closest pixel at each position. This method is fast in operation, but prone to blocky effects and poor image quality. Bilinear interpolation first performs linear interpolation on one axis of the image and then on the other axis. Its receptive field is 2*2, and the operation speed is relatively fast. Compared with nearest neighbor interpolation, it has better performance. Bicubic interpolation performs cubic interpolation on each of the two axes. Its receptive field is 4×4. Compared with bilinear interpolation, the result is smoother, but the speed is much slower. Bicubic interpolation is also the mainstream method for constructing SR datasets (i.e., degrading HR images into LR images). Interpolation-based upsampling methods only utilize the information of the image itself and do not bring additional information, and will produce many side effects when applied to SR, such as high computational complexity, amplified noise, blurring, etc.
[0146] Deep learning-based upsampling methods mainly use the method of sub-pixel convolution for image upsampling, avoiding the disadvantages of interpolation-based upsampling methods and being able to learn in an end-to-end form. The upsampling method based on the sub-pixel convolutional layer mainly utilizes transposed convolution, also known as permutation convolution, which is a special forward convolution. First, the size of the data image is enlarged by padding with 0 according to a certain ratio, then the convolutional kernel is rotated, and finally, forward convolution is performed to achieve the purpose of increasing the image size.
[0147] The super-resolution reconstruction module of the present invention employs upsampling based on a sub-pixel convolutional layer. The original image is input, and after feature extraction, upsampling is performed to obtain the final super-resolution image. To improve the module efficiency, the module is based on a model with an upsampling factor of 2 and is migrated to operations with upsampling factors of 3 and 4. Similarly, the network models with scaling factors of 3 and 4 are also trained based on their 2-fold models. Based on the pre-trained network, the training efficiency of the model with a larger reconstruction factor can be improved, and the model generation time can be reduced. The specific model structure is as Figure 11 shown.
[0148] During the sub-pixel convolutional upsampling process, taking the processing of a 4×4 pixel image with an upsampling factor of 2 as an example, in order to make the output feature map the same size as the original picture, zero-padding is performed on the original image, and the size of the convolutional kernel is 3×3. The number of output feature maps is determined by the magnification factor. For example, if magnified by L times, the number of output feature maps is L×L. As Figure 12 the output feature map is 4, and then the pixel points in the feature map are arranged periodically to achieve the operation of magnifying the original image by 2 times; the process is as Figure 12 shown.
[0149] The example uses a 3×3 kernel to achieve 2-fold upsampling. The specific operation process is as follows: First, double the input, and set the newly inserted value to 0; then use a 3×3 convolutional kernel, set the stride to 1, and the padding to 1. In this way, the input 3×3 image becomes 5×5 in size after the first step of inserting 0 values, and after calculation by the convolutional layer, an image of 6×6 size will be generated. Make full use of the stride, padding, and convolution to achieve N-fold image upsampling.
[0150] Calculate the image video loss function based on the super-resolution image obtained according to the image super-resolution reconstruction steps. In the method based on deep learning, the loss function plays a crucial role in the iterative optimization of network parameters. Different loss functions can guide the network to learn in different target directions. A good loss function can accelerate network convergence and improve network performance. On the contrary, an inappropriate loss function will make the network training slow or oscillate.
[0151] In the field of image super-resolution, the loss function is used to measure the reconstruction error and guide model optimization. Therefore, multiple loss functions (such as pixel loss, content loss, texture loss, adversarial loss, etc.) are applied to the field of image video super-resolution, aiming to accurately balance the reconstruction error and then generate high-quality reconstructed images.
[0152] Pixel loss mainly measures the pixel-level differences between two images, mainly including the L1 loss function or the L2 loss function. Most researchers will use the pixel loss function as one of the objective optimization functions. Among them, the L1 loss function, namely the least absolute error (LAE), calculates the error of the pixel values at the corresponding positions of the reconstructed image and the real image, making the reconstructed image as close as possible to the real high-resolution image. The specific formula is as follows:
[0153]
[0154] Among them, L 1 represents the average absolute error of each iteration, N represents the learning samples of the minimum batch, represents the real high-resolution image, represents the high-resolution image reconstructed and output by the super-resolution model.
[0155] The L2 loss function, namely the mean square error loss (MSE), calculates the sum of the squares of the absolute differences between the pixels of the reconstructed image and the real image. The specific formula of the L2 loss function is as follows:
[0156]
[0157] Among them, each parameter has the same meaning as the parameter in the L 1 formula, except for the difference in finally calculating the average absolute error and the mean square error. For example, since the goal of the L2 loss function is to optimize the average value of pixel points, it usually produces fuzzy predictions, and the reconstructed high-resolution image tends to be smooth. Moreover, the L1 loss function has a faster convergence speed than the L2 loss function, and at the same time, the generated image has rich detail information. This model uses the L1 loss function as the objective optimization function.
[0158] The corresponding feature information loss function is calculated through the feature information loss function acquisition module. The feature information loss function mainly includes content loss and feature loss. Among them, the content loss is used to evaluate the visual effect of the image. The semantic information of the image is extracted by using the image classification network of the pre-trained model, and then the semantic difference between the images is measured. Denote the pre-trained network as The feature extracted at the l-th layer is The content loss is defined as the Euclidean distance between the high-level features of the two images. The specific formula is as follows:
[0159]
[0160] Among them, h 1 、w 1 and c 1 respectively represent the height, width and number of channels of the features of the first layer. Compared with the pixel loss, the content loss makes the reconstructed image It is visually close to the true value image I. The image reconstructed based on this loss function has a better visual effect and is thus widely used in the SR field. The most commonly used pre-trained classification models are VGG and ResNet.
[0161] The texture loss is also called the style reconstruction loss. The image reconstructed based on this has a similar style to the target image (such as color, texture, contrast). The image texture can be represented by the correlation between different channel features, so the Gram matrix is defined. Among them represents the inner product of the feature vectors i and j (after vectorization) of the l-th layer:
[0162]
[0163] where vec(·) represents vectorization. represents the i-th channel of the feature of image I at the l-th layer. The texture loss is defined as follows:
[0164]
[0165] The texture of the image generated based on the texture loss is more realistic and has a better visual effect. However, the size of the cropped image needs to be determined according to experience. If the size is too small, there will be ghosting in the texture part, and if it is too large, there will be ghosting in the whole picture.
[0166] The adversarial loss comes from the GAN series of networks. GAN contains a generator and a discriminator. The generator is used to generate text or images, and the discriminator is used to judge whether the generated content is generated or real. In super-resolution, the generator is regarded as the super-resolution model, and the discriminator judges whether the super-resolved image is generated or real (a binary classifier). The common adversarial loss function is defined as follows:
[0167]
[0168]
[0169] Compared with the pixel-level loss, the PSNR of the images reconstructed by the adversarial loss and the content loss is lower, but the visual effect is very good. This is because the discriminator can extract some difficult-to-learn latent information from the image, and through the adversarial network, the generator can learn accordingly. Therefore, the generated HR images are more realistic. However, the training of GAN is more difficult and is extremely prone to instability.
[0170] The loss function adopted in this method is a feature information loss function improved on the basis of the perceptual loss function. The feature information loss function is defined according to the relevant characteristics of the feature information. The feature information loss function consists of two parts: the content loss and the adversarial loss, and can be specifically expressed as the weighted sum of these two loss functions:
[0171]
[0172] Among them, represents the content loss function, represents the adversarial loss function. The content loss function is selected as the pixel-level loss based on MSE (see Equation 3-3) and the high-level semantic loss based on the VGG network (see Equation 3-4). The adversarial loss function supplements the generation component to the perceptual loss, and the adversarial loss function is defined as:
[0173]
[0174] Among them, D represents the discriminator network, G represents the generator network, and I LR represents the input low-resolution image, represents the super-resolution image reconstructed by the generator network, represents the probability of being judged as a real high-resolution image.
[0175] The loss function of each concatenated feature will extract feature information at different depths, and this information will promote the reconstruction performance of the super-resolution network to varying degrees. Concatenating the outputs of each loss function in parallel can improve the utilization rate of features at different levels and effectively enhance the network's feature information selection ability. Therefore, a loss function fusion mechanism is set up to fuse the image and video loss functions obtained by the image and video loss function acquisition module and the feature information loss functions obtained by each feature information loss function acquisition module to obtain the fused loss function.
[0176] The loss function fusion mechanism mainly concatenates many loss functions and sets specific weights to ensure that the pictures generated by the super-resolution model can preserve the original high- and low-frequency details of the image. Loss function fusion connects the output of the upper layer and the output of each loss function to the loss function fusion layer for global loss function fusion. The global loss function fusion layer can strengthen the information flow between the shallow and deep layers of features, improve the utilization rate of model features, and use the fused loss function for backpropagation to train the image and video super-resolution and super-sharpness reconstruction network, providing more high-frequency and low-detail information for the final image reconstruction.
[0177] The present invention also discloses an intelligent image file coloring system, Figure 13 which is a schematic diagram of the composition of the image file repair system in an embodiment of the present invention; please refer to Figure 13 , the intelligent image file coloring system includes: an image acquisition module 1, a color mathematical model construction module 3, and an image coloring module 5.
[0178] The image acquisition module 1 is used to acquire image file information; the color mathematical model construction module 3 is used to construct a color mathematical model; the image coloring module 5 is used to input the image file information acquired by the image acquisition module into the color mathematical model to color the image file.
[0179] In an embodiment of the present invention, the color mathematical model construction module 3 obtains the global geometric features of the noise image by directly calculating the relationship between any two pixel points in the image, learns the dependency relationship between the global features; and models the long-distance dependency relationship between the target image regions.
[0180] The color mathematical model construction module 3 extracts features from the image, performs feature mapping on the extracted features; and performs three-way processing on the feature mapping;
[0181] The first-way processing is convolved by the first convolution unit to obtain the first feature space f(x);
[0182] f(x) = W f X
[0183] The second-way processing is convolved by the second convolution unit to obtain the second feature space g(x);
[0184] g(x) = W g X
[0185] where x represents the input feature map, f and g are two convolutions respectively, and W f is the first weight matrix, and W g is the second weight matrix.
[0186] Normalize each feature vector along the feature channel direction, as shown in formula (2), to obtain the normalized result norm(x);
[0187]
[0188] Then, transpose f(x) and perform matrix multiplication with g(x), and then perform softmax on each row. According to formula (3), the weight β of all feature points to a certain feature point can be obtained j,i , forming an attention matrix;
[0189]
[0190] The sum of each row in the attention matrix is 1. The i numbers in the j-th row respectively represent the contributions of all other points to it at the i-th point. In this way, the contribution values of all other pixel points to this pixel point can be obtained through each row, and the contribution value of a certain pixel point to all other pixel points can be obtained through each column.
[0191] The third - path processing obtains the third feature space h(x) through convolution, and then multiplies it by the transpose of the attention matrix according to formula (4) to obtain the final result (the result is the feature vector input to the next module);
[0192]
[0193] where h(x) is the convolution result of the third convolution unit, x i is an element in the input feature map x, W h is the third weight matrix, W v is the fourth weight matrix.
[0194] In an embodiment of the present invention, the color mathematical model construction module 3 combines the self - attention mechanism and the residual network, making the detailed features of the output color image richer, the information utilization rate higher, and thus achieving a better coloring effect. The color mathematical model includes a first network model and a second network model. The first network model is provided with a first generation network formed by a first generator and a first discrimination network formed by a first discriminator. The second network model is provided with a second generation network formed by a second generator and a second discrimination network formed by a second discriminator. The first generation network G: X -> Y, and the second generation network is F: Y -> X; the first discrimination network Dx is used to identify whether the input image is X; the second discrimination network Dy is used to identify whether the input image is Y. The first discrimination network Dx and the second discrimination network Dy both include four convolutional layers and two self - attention mechanisms, training the images generated by the corresponding generation network and the existing colored images, and judging the quality of the generated images according to the similarity between the images generated by the generation network and the original images.
[0195] In an embodiment of the present invention, the loss function of the color mathematical model consists of two parts, including an adversarial loss that controls the style of the generated image to approximate the target image and a cycle - consistency loss that preserves the contour information of the input to better preserve the content structure of the input and capture the features of the target domain; that is:
[0196] Loss = Loss GAN +Loss cycle (5)
[0197] Loss aAN ensures that the generation network and the discrimination network evolve mutually, thereby ensuring that the generation network can generate more realistic pictures. Loss cycle ensures that the output picture of the generation network has the same content as the input picture, only different in color; Loss GAN The specific calculation formula is as in (6), Loss cycle The specific calculation formula is as in (7);
[0198]
[0199]
[0200] When training the generator, the parameters of D X and D Y are fixed, and only the parameters of G and F are adjustable; adjust the parameters of G to make the score D Y of the image G(x) generated by G as high as possible. Adjust the parameters of F to make the score D Y (G(x)) as high as possible. Adjust the parameters of F to make the score D X of the image F(y) generated by F as high as possible; the ultimate goal is to make F(G(x)) = x and G(F(y)) = y, ensuring that the colors of the images generated by the generator become more and more realistic. When training the discriminator, the parameters of G and F are fixed, and the parameters of D X and D X are adjustable; when training the discriminator D Y , maximize the value of D X (x), and at the same time, minimize the value of D X (F(y)) to improve the discrimination ability of the discriminator. The process of updating the [formula] is similar and will not be elaborated here. X
[0201] The present invention also discloses an electronic device. Figure 14 is a schematic diagram of the composition of the electronic device in an embodiment of the present invention; please refer to Figure 14 . At the hardware level, the electronic device includes a memory, a processor, and at least one network interface; the processor can be a microprocessor, and the memory can include a memory, such as a random access memory (RAM), and can also include a non-volatile memory, etc. Of course, the electronic device can also be provided with other hardware as needed.
[0202] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect Standard) bus, or an EISA (Extended Industry Standard Architecture) bus, etc.; the bus can include an address bus, a data bus, a control bus, etc. The memory is used to store programs (which can include an operating system program and application programs); the program can include program code, and the program code can include computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0203] In one embodiment, the processor can read the corresponding program from the non-volatile memory into the memory and then run it; the processor can execute the program stored in the memory and is specifically used to perform the following operations (as Figure 1 shown):
[0204]
Step S1
[0205]
Step S3
[0206]
Step S5
[0207] Of course, between step S1 and step S3, step S2, an image super-resolution reconstruction step, can also be set to perform image super-resolution reconstruction on the denoised image to obtain the corresponding image file. Or, step S2 can be set after step S5.
[0208] The present invention further discloses a storage medium on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the following steps of the image edge super-resolution enhancement method of the present invention are implemented (as Figure 1 shown):
[0209]
Step S1
[0210]
Step S3
[0211]
Step S5
[0212] Of course, between step S1 and step S3, step S2, an image super-resolution reconstruction step, can also be set to perform image super-resolution reconstruction on the denoised image to obtain the corresponding image file. Or, step S2 can be set after step S5.
[0213] In summary, the intelligent image file coloring method, system, electronic device and storage medium proposed by the present invention can perform coloring processing on image files and improve the processing effect of image files.
[0214] It should be noted that the present application can be implemented in software and / or a combination of software and hardware; for example, it can be implemented using an application specific integrated circuit (ASIC), a general purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium; for example, a RAM memory, a magnetic or optical drive, or a floppy disk and the like. In addition, some steps or functions of the present application can be implemented using hardware; for example, as a circuit that cooperates with the processor to execute each step or function.
[0215] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0216] The description and application of the present invention here are illustrative and do not intend to limit the scope of the present invention to the above embodiments. The effects or advantages involved in the embodiments may not be reflected in the embodiments due to various factors. The description of the effects or advantages is not used to limit the embodiments. The deformations and changes of the embodiments disclosed here are possible, and the substitutions and equivalent components of the embodiments are well known to those of ordinary skill in the art. Those skilled in the art should clearly understand that the present invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials, and parts without departing from the spirit or essential characteristics of the present invention. Other deformations and changes can be made to the embodiments disclosed here without departing from the scope and spirit of the present invention.
Claims
1. An intelligent colorization method for image files, characterized in that, the intelligent colorization method for image files includes: Image acquisition step: acquiring image file information; Color mathematical model construction step: constructing a color mathematical model; Image colorization step: inputting the image file information acquired in the image acquisition step into the color mathematical model to colorize the image file; In the color mathematical model construction step, by combining the self-attention mechanism and the residual network, the detailed features of the output color image are made more abundant, the information utilization rate is higher, and thus a better coloring effect is achieved; The color mathematical model includes a first network model and a second network model. The first network model is provided with a first generation network formed by a first generator and a first discrimination network formed by a first discriminator. The second network model is provided with a second generation network formed by a second generator and a second discrimination network formed by a second discriminator; The first generation network represents a generator from the X domain to the Y domain, denoted as: G: X -> Y. The second generation network represents a generator from the Y domain to the X domain, denoted as F: Y -> X; The first discrimination network Dx is used to identify whether the input image is X; The second discrimination network Dy is used to identify whether the input image is Y; Both the first discrimination network Dx and the second discrimination network Dy include four convolutional layers and two self-attention mechanisms, and are used to train the images generated by the corresponding generation network and the existing colored images, and judge the quality of the generated images according to the similarity between the images generated by the generation network and the original images.
2. The intelligent colorization method for image files according to claim 1, characterized in that: In the color mathematical model construction step, by directly calculating the relationship between any two pixel points in the image, the global geometric features of the noise image are obtained, and the dependence relationship between the global features is learned; Modeling the long-distance dependence relationship between target image regions; Performing feature extraction on the image and performing feature mapping on the extracted features; Performing three-way processing on the feature mapping; The first-way processing is convolved through the first convolutional unit to obtain the first feature space f(x); f(x) = W f X The second-way processing is convolved through the second convolutional unit to obtain the second feature space g(x); g(x) = W g X Among them, x represents the input feature map, f and g are two convolutions respectively, and W f is the first weight matrix, and W g is the second weight matrix; Normalize each feature vector along the feature channel direction, as shown in formula (2), to obtain the normalized result norm(x); Then, after transposing f(x) and performing matrix multiplication with g(x), and then applying softmax to each row of the result, all the weights β of the feature points with respect to a certain feature point can be obtained according to formula (3). j,i , which constitutes the attention matrix; Among them, β j,i is an element in the attention matrix, and s ij is each element in the matrix obtained by performing matrix multiplication on the transpose of the normalized f(x) and the normalized g(x); the sum of each row in the attention matrix is 1, and the i numbers in the j-th row respectively represent the contributions of all other points to it at the i-th point. In this way, the contribution values of all other pixel points to this pixel point can be obtained through each row, and the contribution value of a certain pixel point to all other pixel points can be obtained through each column; The third-way processing is convolved through the third convolutional unit to obtain the third feature space h(x), and then multiplied by the transpose of the attention matrix according to formula (4) to obtain the final result; Among them, h(x) is the convolution result of the third convolution unit, and x i is an element in the input feature map x, and W h is the third weight matrix, and W v is the fourth weight matrix.
3. The intelligent colorization method for image files according to claim 1, characterized in that: The loss function of the color mathematical model includes: an adversarial loss for controlling the style of the generated image to approximate the target image, and a cycle consistency loss for retaining the contour information of the input so as to better retain the content structure of the input and capture the features of the target domain; that is: Loss=Loss GAN +Loss cycle (5) Loss GAN Ensure that the generation network and the discriminator network evolve with each other, thereby ensuring that the generation network can produce more realistic images, Loss cycle Ensure that the output image of the generation network is only different in color from the input image, but the content is the same; Loss GAN The specific calculation formula is as shown in (6), Loss cycle The specific calculation formula is as shown in (7); When training the generator, D X and D Y parameters are fixed, and only the parameters of G and F are adjustable; adjust the parameters of G to make D Y give a higher score D Y (G(x)) to the image G(x) generated by G; adjust the parameters of F to make D X give a higher score D X (F(y)) to the image F(y) generated by F; the ultimate goal is to make G(F(x)) = x and G(F(y)) = y, ensuring that the colors of the images generated by the generator become more and more realistic; When training the discriminator, the parameters of G and F are fixed, and the parameters of D are adjustable. X and D Y When training the discriminator D X , maximize the value of D X (x); at the same time, minimize the value of D X (F(y)) to improve the discrimination ability of the discriminator.
4. An intelligent colorization system for image files, characterized in that, the intelligent colorization system for image files includes: An image acquisition module for acquiring image file information; A color mathematical model construction module for constructing a color mathematical model; An image coloring module for inputting the image file information obtained by the image acquisition module into the color mathematical model to color the image file; The color mathematical model construction module combines the self-attention mechanism and the residual network, making the detailed features of the output color image richer, the information utilization rate higher, and thus achieving a better coloring effect; The color mathematical model includes a first network model and a second network model. The first network model is provided with a first generation network formed by a first generator and a first discrimination network formed by a first discriminator. The second network model is provided with a second generation network formed by a second generator and a second discrimination network formed by a second discriminator; The first generation network represents a generator from the X domain to the Y domain, denoted as: G:X -> Y. The second generation network represents a generator from the Y domain to the X domain, denoted as F:Y -> X. The first discrimination network Dx is used to identify whether the input image is X. The second discrimination network Dy is used to identify whether the input image is Y; Both the first discrimination network Dx and the second discrimination network Dy include four convolutional layers and two self-attention mechanisms, training the images generated by the corresponding generation network and the existing colored images, and judging the quality of the generated images according to the similarity between the images generated by the generation network and the original images.
5. The intelligent image file coloring system according to claim 4, wherein: The color mathematical model construction module directly calculates the relationship between any two pixel points in the image, obtains the global geometric features of the noise image, and learns the dependence relationship between the global features; modeling the long-distance dependence relationship between the target image regions; The color mathematical model construction module extracts features from the image and performs feature mapping on the extracted features; Performing three-way processing on the feature mapping; The first-way processing is convolved through the first convolutional unit to obtain the first feature space f(x); f(x) = W f x The second-way processing is convolved through the second convolutional unit to obtain the second feature space g(x); g(x) = W g x Among them, x represents the input feature map, and f and g are two convolutions respectively, and W f is the first weight matrix, and W g is the second weight matrix; Normalize each feature vector along the feature channel direction, as shown in formula (2), to obtain the normalized result norm(x); Then, after transposing f(x) and performing matrix multiplication with g(x), Softmax is applied to each row of the result. The output values are converted into a probability distribution ranging from [0,1] and summing to 1 through the Softmax function. The specific calculation method is as shown in formula (3); all the weights β of the feature points for a certain feature point can be obtained according to formula (3). j,i , which constitutes the attention matrix. where β j,i is an element in the attention matrix, and s ij is each element in the matrix obtained by performing matrix multiplication of the transpose of the normalized f(x) and the normalized g(x); the sum of each row in the attention matrix is 1, and the i numbers in the j-th row respectively represent the contributions of all other points to it at the i-th point. In this way, the contribution values of all other pixel points to this pixel point can be obtained through each row, and the contribution value of a certain pixel point to all other pixel points can be obtained through each column; The third-way processing is convolved through the third convolutional unit to obtain the third feature space h(x), and then multiplied by the transpose of the attention matrix according to formula (4) to obtain the final result; Among them, h(x) is the convolution result of the third convolution unit, and x i is an element in the input feature map x, and W h is the third weight matrix, and W v is the fourth weight matrix.
6. The intelligent image file coloring system according to claim 4, wherein: The loss function of the color mathematical model includes: an adversarial loss for controlling the style of the generated image to approximate the target image, and a cycle consistency loss for retaining the contour information of the input so as to better retain the content structure of the input and capture the features of the target domain; that is: Loss=Loss GAN +Loss cycle (5) Loss GAN Ensure that the generation network and the discriminator network evolve with each other, thereby ensuring that the generation network can generate more realistic images, Loss cycle Ensure that the output image of the generation network is only different in color from the input image, but the content is the same; Loss GAN The specific calculation formula is as shown in (6), Loss cycle The specific calculation formula is as shown in (7); When training the generator, D X and D Y parameters are fixed, and only the parameters of G and F are adjustable; adjust the parameters of G to make D Y give a higher score D Y (G(x)) to the images G(x) generated by G; adjust the parameters of F to make D X give a higher score D X (F(y)) to the images F(y) generated by F; the ultimate goal is to make G(F(y)) = x and F(F(y)) = y, ensuring that the colors of the images generated by the generator become more and more realistic; When training the discriminator, the G and F parameters are fixed, and the D X and D Y parameters are adjustable; when training the discriminator D X , maximize the value of D X (x); at the same time, minimize the value of D X (F(y)) to improve the discrimination ability of the discriminator.
7. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.
8. A storage medium, on which computer program instructions are stored, wherein, When the computer program instructions are executed by the processor, the steps of the method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Image coloring and model training method and device, electronic equipment and storage medium
CN113362409A